Preparation of terpenoid compounds

WO2026068770A3PCT designated stage Publication Date: 2026-08-06FIRMENICH SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
FIRMENICH SA
Filing Date
2025-09-26
Publication Date
2026-08-06

AI Technical Summary

Technical Problem

The perfumery industry faces a need for sustainable processes to produce terpenoid compounds like ambroxide, which are typically prepared through chemical methods that pose environmental and waste management challenges.

Method used

A multi-step enzymatic process using polypeptides with specific enzyme activities, such as alcohol dehydrogenase, enal-cleaving, Baeyer-Villiger monooxygenase, esterase, and terpene cyclase, to convert precursor compounds into ambroxide and other terpenoid compounds in vivo or through bioconversion methods.

Benefits of technology

This process provides a sustainable and efficient method for producing terpenoid compounds with high yield and selectivity, suitable for use in perfumery and flavor applications.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present invention provides a process for the preparation of a compound of formula (I). Included in this invention is an in vivo process for the preparation of compound of formula (I) in recombinant cells. The present invention also provides recombinant cells which may be used in said process. The present invention also provides polypeptides which can be used in the present invention. The present invention also provides further processes for the preparation of further compounds.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PREPARATION OF TERPENOID COMPOUNDS

[0002] Technical field

[0003] The present invention relates to polypeptides and processes for the enzymatic preparation of ambroxide and other terpenoid compounds.

[0004] Background

[0005] In the perfumery industry, there is a constant need to provide methods for the preparation of compounds for industrial use in fragrances. Key amongst such compounds are ingredients relating to the amber class of olfactory compounds which are naturally found in ambergris and which can function as fixatives to allow scent to endure much longer.

[0006] A key compound in ambergris is ambroxide, a terpenoid compound. Terpenoids are found in most organisms (microorganisms, animals and plants). These compounds are made up of five carbon units called isoprene units and are classified by the number of these units present in their structure. Thus monoterpenes, sesquiterpenes and diterpenes are terpenes containing 10, 15 and 20 carbon atoms respectively. Sesquiterpenes, for example, are widely found in the plant kingdom. Many sesquiterpene molecules are known for their flavor and fragrance properties and their cosmetic, medicinal and antimicrobial effects. Numerous sesquiterpene hydrocarbons and sesquiterpenoids have been identified. Commercially relevant compounds include Cetalox® ((3aRS,9aRS,9bRS)-3a,6,6,9a-tetramethyl- 1 ,2,3a,4,6,7,8,9,9a,9b-decahydronaphtho[2,1-b]furan; origin: Firmenich SA, Geneva, Switzerland) or Ambrox® ((3aR,5aS,9aS,9bR)-3a,6,6,9a-tetramethyldodecahydronaphtho[2,1-b]furan; origin: Firmenich SA, Geneva, Switzerland), these compounds replicating ambroxide.

[0007] Chemical routes to the preparation of these compounds are known in the art. However, given the environmental and waste problems associated with chemical production of such compounds, there is a need to develop more sustainable processes for the production of ambroxide and other terpenoid compounds.

[0008] This problem is addressed by the present invention which provides polypeptides and processes for producing such compounds by in vivo and / or bioconversion methods. Said methods may use a multi-step enzymatic process. Summary

[0009] A first aspect of the invention provides a process for the preparation of a compound of formula (I) in the form of any one of its stereoisomers or a mixture thereof, comprising:

[0010] (i) contacting a compound of formula (II) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having alcohol dehydrogenase (ADH) enzyme activity to produce a compound of formula (III);

[0011] (ii) contacting the compound of formula (III) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having enal-cleaving enzyme activity to produce a compound of formula (IV);

[0012] (iii) contacting the compound of formula (IV) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having Baeyer-Villiger monooxygenase (BVMO) enzyme activity to produce a compound of formula (V);

[0013] (iv) contacting the compound of formula (V) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having esterase enzyme activity to produce a compound of formula (VI); and

[0014] (v) contacting the compound of formula (VI) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having terpene cyclase enzyme activity to produce a compound of formula (I), wherein the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323.

[0015] A second aspect of the invention provides a process for the preparation of a compound of formula (I): in the form of any one of its stereoisomers or a mixture thereof, comprising contacting a compound of formula (VI) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having terpene cyclase enzyme activity to produce a compound of formula (I), wherein the polypeptide having terpene cyclase enzyme activity has at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 322 and 323. and optionally, further comprising one or more step(s) selected from:

[0016] (i) contacting the compound of formula (II) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having alcohol dehydrogenase (ADH) enzyme activity to produce a compound of formula (III);

[0017] (ii) contacting the compound of formula (III) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having enal- cleaving enzyme activity to produce a compound of formula (IV);

[0018] (iii) contacting the compound of formula (IV) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having Baeyer-Villiger monooxygenase (BVMO) enzyme activity to produce a compound of formula (V); and

[0019] (iv) contacting the compound of formula (V) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having esterase enzyme activity to produce a compound of formula (VI).

[0020] A third aspect of the invention provide a process for the preparation of a compound of formula (I): in the form of any one of its stereoisomers or a mixture thereof, comprising: contacting a compound of formula (V) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having esterase enzyme activity to produce a compound of formula (VI); and contacting the compound of formula (VI) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having terpene cyclase enzyme activity to produce a compound of formula (I), wherein the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323.

[0021] In a fourth aspect of the invention, the process of the third aspect of the invention further comprises: contacting a compound of formula (IV) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having BVMO enzyme activity to produce a compound of formula (V).

[0022] In yet a fifth aspect of the invention, the process of the fourth aspect of the invention further comprises: contacting a compound of formula (III) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having enal- cleaving enzyme activity to produce a compound of formula (IV).

[0023] In said aspects of the invention, the polypeptide having terpene cyclase enzyme activity is a squalene cyclase (SHC) enzyme.

[0024] An embodiment of said aspects of the invention is wherein more than 97% of the compound of formula (I) is in the form of formula (la): (formula la); and / or formula (lb): (formula lb).

[0025] A further embodiment of said aspects of the invention is wherein:

[0026] (i) the compound of formula (I) is in the form of formula (la): (formula la)

[0027] (ii) the compound of formula (II) is in the form of formula (Ila):

[0028] (formula Ila);

[0029] (iii) the compound of formula (III) is in the form of formula (Illa): (formula Illa);

[0030] (iv) the compound of formula (IV) is in the form of formula (IVa): (formula IVa);

[0031] (v) the compound of formula (V) is in the form of formula (Va): (formula Va); (vi) the compound of formula (VI) is in the form of formula (Via):

[0032] (formula Via).

[0033] A further embodiment of said aspects of the invention is wherein the process further comprises one or more step(s) selected from:

[0034] 1) preparing geranylgeranyl-diphosphate (GGPP) from isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP) using one or more polypepides having prenyltransferase enzyme activity; and,

[0035] 2) preparing a compound of formula (II) from GGPP using one or more polypeptides having phosphatase enzyme activity.

[0036] A further embodiment of said aspects of the invention is wherein the process is an in vivo or a bioconversion process.

[0037] A further embodiment of said aspects of the invention is wherein the process is performed using a recombinant cell comprising, capable of functionally expressing or expressing the polypeptide having terpene cyclase enzyme activity to produce compound of formula (I). In particular, the process comprises gowing the recombinant cell comprising, capable of functionally expressing or expressing the polypeptide having terpene cyclase enzyme activity to produce compound of formula (I). In yet a further embodiment, said polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; in particular, a heterologous squalene cyclase enzyme.

[0038] A further embodiment of said aspects of the invention is wherein the process is performed in a recombinant cell comprising, capable of functionally expressing or expressing: (i) the polypeptide having ADH enzyme activity, (ii) the polypeptide having enal-cleaving enzyme activity, (iii) the polypeptide having BVMO enzyme activity, (iv) the polypeptide having esterase enzyme activity, and (v) the polypeptide having terpene cyclase enzyme activity. In particular, the process comprises gowing the recombinant cell comprising, capable of functionally expressing or expressing: (i) the polypeptide having ADH enzyme activity, (ii) the polypeptide having enal-cleaving enzyme activity, (iii) the polypeptide having BVMO enzyme activity, (iv) the polypeptide having esterase enzyme activity, and (v) the polypeptide having terpene cyclase enzyme activity. In yet a further embodiment, one or more of said polypeptides may be a heterologous polypeptide. In particular, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In yet a further embodiment, the polypeptides having ADH enzyme activity, enal-cleaving enzyme activity, BVMO enzyme activity, esterase enzyme activity and terpene cyclase enzyme activity are heterologous polypeptides.

[0039] A further aspect of the invention provides a recombinant cell comprising, capable of producing or producing a compound of formula (I), wherein the recombinant cell comprises, is capable of functionally expressing or expresses: (i) a polypeptide having ADH enzyme activity, (ii) a polypeptide having enal-cleaving enzyme activity, (iii) a polypeptide having BVMO enzyme activity, (iv) a polypeptide having esterase enzyme activity, and (v) a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323. In yet a further embodiment, one or more of said polypeptides may be a heterologous polypeptide. In particular, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In yet a further embodiment, the polypeptides having ADH enzyme activity, enal-cleaving enzyme activity, BVMO enzyme activity, esterase enzyme activity and terpene cyclase enzyme activity are heterologous polypeptides.

[0040] A further aspect of the invention provides a recombinant cell comprising, capable of producing or producing a compound of formula (I), wherein the recombinant cell comprises, is capable of functionally expressing or expresses a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323. In yet a further embodiment, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme.

[0041] A further aspect of the invention provides a cell culture fermentation medium comprising the recombinant cell of the invention as described herein above. The cell culture fermentation medium may further comprise a compound of formula (I) and optionally, one or more compound(s) selected from compounds of formula (II), formula (III), formula (IV), formula (V) and formula (VI).

[0042] A further aspect of the invention provides a reaction mixture comprising a compound of formula (I) and one or more compound(s) selected from compounds of formula (II), formula (III), formula (IV) and formula (V), wherein the reaction mixture further comprises a polypeptide having terpene cyclase enzyme activity having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323.

[0043] A further aspect of the invention provides a reaction mixture comprising a compound of formula (I) and a polypeptide having terpene cyclase enzyme activity having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323.

[0044] In yet futher embodiments of the above aspects of the invention, the polypeptide having terpene cyclase enzyme activity has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. A further aspect of the invention provides a compound of formula (I) obtained or obtainable by a process of the invention, from a recombinant cell of the invention, from a cell culture fermentation medium of the invention or from a reaction mixture of the invention as described herein above.

[0045] A further aspect of the invention provides the use of said compound of formula (I) as a perfumery, flavor or aroma ingredient, or as a precursor for making said ingredient.

[0046] A further aspect of the invention provides a mutant terpene cyclase enzyme having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323, wherein said polypeptide has amino acid alanine at position 437 and amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Said mutant terpene cyclase enzyme is a mutant squalene cyclase enzyme.

[0047] Description of the drawings

[0048] Figure 1. Biosynthetic pathway of (2E)-geranyl-diphosphate (GPP), (2E,6E)-farnesyl-diphosphate (FPP) and (2E,6E,10E)-geranylgeranyl-diphosphate (GGPP) from isopentenyl-diphosphate (IPP) and dimethylallyl-diphosphate (DMAPP).

[0049] Figure 2. Biosynthetic pathways of (2E,6E,10E)-geranylgeraniol from (2E,6E,10E)-geranylgeranyl- diphosphate (GGPP). Pi, inorganic phosphate; PPi, inorganic pyrophosphate.

[0050] Figure 3. New biochemical pathway to (3E,7E)-homofarnesol.

[0051] Figure 4. GC-MS analysis of terpenoids and derivatives produced using E. coli cells engineered to produce (3E,7E)-homofarnesol and expressing the proteins PsAerADH (SEQ ID NO: 11), SCH24-BVMO1 (SEQ ID NO: 23), SCH24-EST1 (SEQ ID NO: 27), CcrGGPPS2-del57 (SEQ ID NO: 1), PgpB (SEQ ID NO: 3) encoded by the plasmid pHFOL-5.

[0052] Figure 5. GC-MS analysis of terpenoids and derivatives produced using E. coli cells expressing the proteins PsAerADH (SEQ ID NO: 11), SCH24-EST1 (SEQ ID NO: 27), CcrGGPPS2-del57 (SEQ ID NO: 1), PgpB (SEQ ID NO: 3) encoded by the plasmid pF-Facetone-7 (A) and using the same cells expressing in addition a BVMO enzyme (AflaBVMOI , SEQ ID NO: 26) (B).

[0053] Figure 6. GC-MS chromatogram of YST403_HFOL strain engineered to produce (3E,7E)-homofarnesol. The final product (3E,7E)-homofarnesol, as well as the pathway intermediates (5E,9E)-farnesylacetone and (2E,6E,10E)-geranylgeraniol are shown.

[0054] Figure 7. New biochemical pathway of compound of formula la using squalene cyclases. Figure 8: GC-MS chromatograms in single ion monitoring mode (221 Da) of S. cerevisiae strain YST403 engineered to express the (3E,7E)-homofarnesol biosynthetic pathway enzymes and wild-type squalene cyclase BmeSCH (SEQ ID NO: 315) (A), SCHJC_C003990 (SEQ ID NO: 321) (B) or BsuTC (SEQ ID NO: 317) (C); GC-MS chromatogram of the control strain YST403_HFOL expressing only the (3E,7E)- homofarnesol biosynthetic pathway enzymes (D). Peaks that correspond to (3E,7E)-homofarnesol and compound of formula (la) are indicated.

[0055] Figure 9: GC-MS chromatograms in single ion monitoring mode (221 Da) of S. cerevisiae strain YST403 engineered to express the (3E,7E)-homofarnesol biosynthetic pathway enzymes and mutant squalene cyclase BmeSCH_V1 (SEQ ID NO: 316) (A), TARA_SAMEA2623601_METAG_s18_g109_V1 (SEQ ID NO: 323) (B), CttSHC_V1 (SEQ ID NO: 320) (C) or BsuTC_V1 (SEQ ID NO: 318) (D); GC-MS chromatogram of the control strain YST403_HFOL expressing only the (3E,7E)-homofarnesol biosynthetic pathway enzymes (E). Peaks that correspond to (3E,7E)-homofarnesol and compound of formula (la) are indicated.

[0056] Figure 10. GC-MS chromatograms in single ion monitoring mode (221 Da) of S. cerevisiae stain YST403 engineered to express the (3E,7E)-homofarnesol biosynthetic pathway enzymes and wild-type squalene cyclase BmeSCH (SEQ ID NO: 315) or mutant squalene cyclase BmeSCH_V1 (SEQ ID NO: 316). Peaks that correspond to (3E,7E)-homofarnesol and compound of formula (la) are indicated.

[0057] Abbreviations used

[0058] ADH alcohol dehydrogenase

[0059] BVMO Baeyer-Villiger Monooxygenase bp base pair kb kilo base

[0060] DNA deoxyribonucleic acid cDNA complementary DNA

[0061] DMAPP dimethylallyl diphosphate

[0062] FMO Flavin Monooxygenase

[0063] FPP farnesyl diphosphate

[0064] GPP geranyldiphosphate

[0065] GGPP geranylgeranyl diphosphate

[0066] GGPS geranylgeranyl diphosphate synthase

[0067] GC gas chromatograph

[0068] IPP isopentenyl diphosphate iMS mass spectrometer I mass spectrometry

[0069] MVA mevalonic acid

[0070] PP diphosphate, pyrophosphate

[0071] PCR polymerase chain reaction RNA ribonucleic acid

[0072] SHC squalene cyclase mRNA messenger ribonucleic acid miRNA micro RNA siRNA small interfering RNA rRNA ribosomal RNA tRNA transfer RNA

[0073] TPP terpenyl diphosphate

[0074] Definitions

[0075] General terms

[0076] For the descriptions herein and the appended claims, the use of “or” means “and / or” unless stated otherwise. Similarly, “comprise”, “comprises”, “comprising”, “include”, “includes”, and “including” are interchangeable and not intended to be limiting.

[0077] It is to be further understood that where descriptions of various embodiments use the term "comprising," those skilled in the art would understand that in some specific instances, an embodiment can be alternatively described using language "consisting essentially of or "consisting of.

[0078] The terms "purified", "substantially purified", and "isolated" as used herein refer to the state of being free of other, dissimilar compounds with which a compound of the invention is normally associated in its natural state, so that the "purified", "substantially purified", and "isolated" subject comprises at least 0.5%, 1 %, 5%, 10%, or 20%, or at least 50% or 75% of the mass, by weight, of a given sample. In one embodiment, these terms refer to the compound of the invention comprising at least 95, 96, 97, 98, 99 or 100%, of the mass, by weight, of a given sample. As used herein, the terms "purified," "substantially purified," and "isolated" when referring to a nucleic acid or protein, or nucleic acids or proteins, also refers to a state of purification or concentration different than that which occurs naturally, for example in a prokaryotic or eukaryotic environment, like, for example in a bacterial or fungal cell, or in the mammalian organism, especially human body. Any degree of purification or concentration greater than that which occurs naturally, including (1) the purification from other associated structures or compounds or (2) the association with structures or compounds to which it is not normally associated in said prokaryotic or eukaryotic environment, are within the meaning of "isolated”. The nucleic acid or protein or classes of nucleic acids or proteins, described herein, may be isolated, or otherwise associated with structures or compounds to which they are not normally associated in nature, according to a variety of methods and processes known to those of skill in the art.

[0079] The term “about” indicates a potential variation of ± 25% of the stated value, in particular ± 15%, ± 10 %, more particularly ± 5%, ± 2% or ± 1 %. The term "substantially" describes a range of values from about 80 to 100%, such as, for example, 85- 99.9%, in particular 90 to 99.9%, more particularly 95 to 99.9%, or 98 to 99.9% and especially 99 to 99.9%.

[0080] “Predominantly” refers to a proportion in the range of above 50%, as for example in the range of 51 to 100%, particularly in the range of 75 to 99,9%, more particularly 85 to 98,5%, like 95 to 99%.

[0081] A “main product” in the context of the present invention designates a single compound or a group of at least 2 compounds, like 2, 3, 4, 5 or more, particularly 2 or 3 compounds, which single compound or group of compounds is “predominantly” prepared by a reaction as described herein, and is contained in said reaction in a predominant proportion based on the total amount of the constituents of the product formed by said reaction. Said proportion may be a molar proportion, a weight proportion or, preferably based on chromatographic analytics, an area proportion calculated from the corresponding chromatogram of the reaction products.

[0082] A “side product” in the context of the present invention designates a single compound or a group of at least 2 compounds, like 2, 3, 4, 5 or more, particularly 2 or 3 compounds, which single compound or group of compounds is not “predominantly” prepared by a reaction as described herein.

[0083] Because of the reversibility of enzymatic reactions, the present invention relates, unless otherwise stated, to the enzymatic or biocatalytic reactions described herein in both directions of reaction.

[0084] The term "stereoisomers" includes conformational isomers and in particular configuration isomers.

[0085] Included in general are, according to the invention, all “stereoisomeric forms” of the compounds described herein, such as “constitutional isomers” and “stereoisomers”.

[0086] “Stereoisomeric forms” encompass in particular, “stereoisomers” and mixtures thereof, e.g. configuration isomers (optical isomers), such as enantiomers, or geometric isomers (diastereomers), such as E- and Z- isomers, and combinations thereof. If one or more asymmetric centers are present in one molecule, the invention encompasses all combinations of different conformations of these asymmetry centers, e.g. enantiomeric pairs.

[0087] “Stereoselectivity” describes the ability to produce a particular stereoisomer of a compound in a stereoisomerically pure form or to specifically convert a particular stereoisomer in an enzyme catalyzed method as described herein out of a plurality of stereoisomers. More specifically, this means that a product of the invention is enriched with respect to a specific stereoisomer, or an educt may be depleted with respect to a particular stereoisomer. This may be quantified via the purity %ee-parameter calculated according to the formula:

[0088] %ee = [XA-XB] / [ XA+XB]*100, wherein XA and XB represent the molar ratio (Molen bruch) of the stereoisomers A and B.

[0089] The terms “selectively converting” or “increasing the selectivity” in general means that a particular stereoisomeric form, as for example the E-form, of an unsaturated hydrocarbon, is converted in a higher proportion or amount (compared on a molar basis) than the corresponding other stereoisomeric form, as for example Z-form, either during the entire course of said reaction (i.e. between initiation and termination of the reaction), at a certain point of time of said reaction, orduring an “interval” of said reaction. In particular, said selectivity may be observed during an “interval” corresponding 1 to 99%, 2 to 95%, 3 to 90%, 5 to 85%, 10 to 80%, 15 to 75%, 20 to 70%, 25 to 65%, 30 to 60%, or 40 to 50% conversion of the initial amount of the substrate. Said higher proportion or amount may, for example, be expressed in terms of: a higher maximum yield of an isomer observed during the entire course of the reaction or said interval thereof; a higher relative amount of an isomer at a defined % degree of conversion value of the substrate; and / or an identical relative amount of an isomer at a higher % degree of conversion value; each of which preferably being observed relative to a reference method, said reference method being performed under otherwise identical conditions with known chemical or biochemical means.

[0090] “Yield" and I or the "conversion rate" of a reaction according to the invention is determined over a defined period of, for example, 4, 6, 8, 10, 12, 16, 20, 24, 36 or 48 hours, in which the reaction takes place. In particular, the reaction is carried out under precisely defined conditions, for example at “standard conditions” as herein defined.

[0091] If the present disclosure refers to features, parameters and ranges thereof of different degree of preference (including general, not explicitly preferred features, parameters and ranges thereof) then, unless otherwise stated, any combination of two or more of such features, parameters and ranges thereof, irrespective of their respective degree of preference, is encompassed by the disclosure of the present description.

[0092] Biochemical and biological terms

[0093] The term "domain" refers to a set of amino acids or a partial sequence of amino acids residues conserved at specific positions along an alignment of sequences of evolutionarily related proteins. While amino acids at other positions can vary between protein homologues, amino acids that are highly conserved at specific positions of such domain indicate amino acids that are likely essential in the structure, stability or function of a protein. Identified by their high degree of conservation in aligned sequences of a family of protein homologues, they can be used as identifiers to determine if any polypeptide in question belongs to a previously identified polypeptide family.

[0094] The term "motif " or consensus sequence" or "signature" refers to a short-conserved region in the sequence of evolutionarily related proteins. Motifs are frequently highly conserved parts of domains, but may also include only part of the domain. Signatures are predictive models which describe protein families, domains or sites.

[0095] The sequences of motifs can be described using the standard IUPAC one-letter codes for the amino acids. Ambiguities are indicated by listing the acceptable amino acids for a given position between brackets. For example, [LWI] stands for L (Leucine), W (Tryptophan) or I (Isoleucine). X represent positions where independently of each other any natural amino acid residue is present.

[0096] A “protein family” is defined as a group of proteins that share a common evolutionary origin reflected by their related functions, similarities in sequence, or similar primary, secondary or tertiary structure. Proteins within protein families are usually homologous and have similar structure of conserved functional domains and motifs.

[0097] Specialist databases exist for the identification of protein domains, for example, SMART (http: / / smart.embl- heidelberg.de / smart / set_mode. cgi?GENOMIC=1) (Schultz et al. (1998) Proc. Natl. Acad. Sci. USA 95, 5857-5864; Letunic et al. (2020) Nucleic Acids Res 49, D458-D460), InterPro (Paysan-Lafosse et al, Nucleic Acids Research, Nov 2022; Mulder et al., (2003) Nucl. Acids. Res. 31 , 315-318), or Pfam (Bateman et al., Nucleic Acids Research 30(1): 276-280 (2002)).

[0098] Useful tools to search or predict protein domains or protein family signatures in protein sequence are for example the NCBI conserved domain search tool (https: / / www.ncbi.nlm.nih.gov / Structure / cdd / wrpsb.cgi) or the InterProScan tool (http: / / www.ebi.ac.uk / interpro / search / sequence / ). Domains or motifs may also be identified using routine techniques, such as by sequence alignment.

[0099] The term "Pfam" refers to a large collection of protein domains and protein families maintained by the Pfam Consortium and available at several sponsored world wide web sites, such as the InterPro consortium web site https: / / www.ebi.ac.uk / interpro / (European Molecular Biology Laboratory-European Bioinformatics Institute (EMBL_EBI). The latest release of Pfam is Pfam 35.0 (November 2021), based on the UniProt Reference Proteomes (El-Gebali S. et al, 2019, Nucleic Acids Res. 47, Database issue D427-D432). Pfam domains and families are identified using multiple sequence alignments and hidden Markov models (HMMs). Pfam-A family or domain assignments, are high quality assignments generated by a curated seed alignment using representative members of a protein family and profile hidden Markov models based on the seed alignment (Unless otherwise specified, matches of a queried protein to a Pfam domain or family are Pfam-A matches). All identified sequences belonging to the family are then used to automatically generate a full alignment for the family (Sonnhammer (1998) Nucleic Acids Research 26, 320-322; Bateman (2000) Nucleic Acids Research 26, 263-266; Bateman (2004) Nucleic Acids Research 32, Database Issue, D138-D141 ; Finn (2006) Nucleic Acids Research Database Issue 34, D247-251 ; Finn (2010) Nucleic Acids Research Database Issue 38, D211-222). By accessing the Pfam database, for example, using any of the above- reference websites, protein sequences can be queried against the HMMs using HMMER homology search software (e.g., HMMER2, HMMER3, or a higher version, hmmer.janelia.org / ). Significant matches that identify a queried protein as being in a pfam family (or as having a particular Pfam domain) are those in which the bit score is greater than or equal to the gathering threshold for the Pfam domain. Expectation values (e-values) can also be used as a criterion for inclusion of a queried protein in a Pfam or for determining whether a queried protein has a particular Pfam domain, where low e-values, much less than 1 .0, for example less than 0.1 , or less.

[0100] InterPro is another database of protein families providing a classification of protein sequences into families and identifies functionally important domains and conserved sites (Blum et al, Nucleic Acids Res. 2021 49(D1):D344-D354). The protein signatures are provided by multiple databases such as Pfam or SMART (Simple Modular Architecture Research Tool). InterProScan is a software that allows protein and nucleic acid sequences to be searched against InterPro's signatures.

[0101] The “E-value” (expectation value) is the number of hits that would be expected to have a score equal to or better than this value, by chance alone. This means that a good E-value which gives a confident prediction is much less than 1. E-values around 1 is what is expected by chance. Thus, the lower the E-value, the more specific the search for domains will be. Only positive numbers are allowed.

[0102] A “precursor” compound or molecule of a target compound or molecule as described herein is converted to said target compound, preferably through the enzymatic action of a suitable polypeptide performing at least one structural or functional change on said precursor molecule. For example, a “diphosphate precursor” (as for example a “terpenyl diphosphate precursor”) is converted to said target compound (as for example a terpene alcohol) via enzymatic removal of the diphosphate moiety, for example by removal of mono- or diphosphate groups by a phosphatase enzyme. For example, a “non-cyclic precursor” (like a “non-cyclic terpenyl precursor”) may be converted to the cyclic target molecule (like a cyclic terpene compound) through the action of a cyclase or synthase enzyme, irrespective of the particular enzymatic mechanism of such enzyme, in one or more steps.

[0103] The enzyme nomenclature or enzyme classification (EC) established by the International Union of Biochemistry and Molecular Biology (IUBMB) is a system of naming and categorizing enzymes based on their catalytic activity and biochemical properties. The enzyme nomenclature is widely used in biochemistry to classify and categorize based on their function. The E.C. classification assigns each enzyme a number reflecting the reaction or the type of reaction catalyzed by this enzyme.

[0104] The enzyme classification can be explored using the ‘ExplorEnz’ database (https: / / www. enzymedatabase. org / ) or International Union of Biochemistry and Molecular Biology (IUBMB) web site (https: / / iubmb.qmul.ac.uk). Information can be found about the classification and nomenclature of enzymes, their functions and properties. The database can be searched to find information for a specific enzyme family or enzyme. The terms “biological function,” “function”, “biological activity” or “activity” of a terpenyl synthase refer to the ability of a terpenyl diphosphate synthase as described herein to catalyze the formation of at least one terpenyl diphosphate from the corresponding precursor terpene.

[0105] The terms “biological function,” “function”, “biological activity” or “activity” of a terpenyl diphosphate phosphatase refer to the ability of the terpenyl diphosphate phosphatase as described herein to catalyze the removal of a diphosphate group from said terpenyl compound to form the corresponding terpene alcohol.

[0106] As used herein, the term “host cell”, “recombinant cell” or “transformed cell” refers to a cell (or organism) altered to harbor at least one nucleic acid molecule, for instance, a recombinant gene encoding a desired protein or nucleic acid sequence which upon transcription yields at least one functional polypeptide of the present invention. The host cell is particularly a bacterial cell, a fungal cell or a plant cell or plants. The host cell may contain a recombinant gene or several genes, as for example organized as an operon, which has been integrated into the nuclear organelle genomes of the host cell. Alternatively, the host may contain the recombinant gene extra-chromosomally. Methods of introducing recombinant nucleic acid sequences into such host cells are well known in the art and constitute routine laboratory methodologies which do not need to be further described herein.

[0107] The term "heterologous" when used with respect to a polynucleotide (such as DNA or RNA), polypeptide or protein refers to a polynucleotide, polypeptide or protein that does not occur naturally as part of the recombinant cell, genome or DNA or RNA in which it is present, or that is found in a different number of copies, or under the control of a different control sequence, or in a cell or location or locations in the genome or DNA or RNA that differ from that in which it is found in nature. Heterologous polynucleotides, polypeptides or proteins are not endogenous to the cell into which they are introduced but have been obtained from another cell or synthetically or recombinantly produced.

[0108] The term "homologous" when used to indicate the relation between a given (recombinant) polynucleotide or polypeptide and a given host organism or host cell such as the recombinant cell as disclosed herein, is understood to mean that in nature the polynucleotide or polypeptide molecule is produced by a recombinant cell, host cell or organism of the same species, such as of the same variety or strain.

[0109] The term “organism” refers to any non-human multicellular or unicellular organism such as a plant, or a microorganism. Particularly, a micro-organism is a bacterium, a yeast, an algae or a fungus.

[0110] The term “plant” is used interchangeably to include plant cells including plant protoplasts, plant tissues, plant cell tissue cultures giving rise to regenerated plants, or parts of plants, or plant organs such as roots, stems, leaves, flowers, pollen, ovules, embryos, fruits and the like. Any plant can be used to carry out the methods of an embodiment herein. Detailed description

[0111] As described above, many sesquiterpene molecules are known for their flavor and fragrance properties and their cosmetic, medicinal and antimicrobial effects. Numerous sesquiterpene hydrocarbons and sesquiterpenoids have been identified. Commercially relevant compounds include Cetalox® ((3aRS,9aRS,9bRS)-3a,6,6,9a-tetramethyl-1 ,2,3a,4,6,7,8,9,9a,9b-decahydronaphtho[2,1-b]furan; origin: Firmenich SA, Geneva, Switzerland) or Ambrox® ((3aR,5aS,9aS,9bR)-3a,6,6,9a- tetramethyldodecahydronaphtho[2,1-b]furan; origin: Firmenich SA, Geneva, Switzerland), these compounds replicating ambroxide.

[0112] The present inventors sought to identify improved processes for the preparation of a compound of formula (I), also known as 3a, 6, 6, 9a tetramethyldodecahydronaphtho[2,1-b]furan.

[0113] To prepare an improved process for the preparation of a compound of formula (I), they developed a deep understanding of the biochemical route to the production of these compounds by a multi-enzymatic reaction from precursor compounds. This multi-enzymatic reaction is the first time the preparation of this compound has been performed by such a step-wise reaction and constitutes a significant scientific and commercial advance in the preparation of sesquiterpene compound of formula (I). In particular, the combination of enzymes and their order in the process has not been described before in the prior art.

[0114] Included in this invention is an in vivo process for the preparation of compound of formula (I) in recombinant cells. This is the first time a wholly in vivo process for this production of this compound has been demonstrated by creating a biosynthetic pathway to the compound of formula (I) in recombinant cells.

[0115] Accordingly, this invention provides a solution to the the problem of the preparation of such compound.

[0116] First aspect for the preparation of a compound of formula (I)

[0117] A first aspect of the invention provides a process for the preparation of a compound of formula (I) in the form of any one of its stereoisomers or a mixture thereof, comprising:

[0118] (i) contacting a compound of formula (II) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having alcohol dehydrogenase (ADH) enzyme activity to produce a compound of formula (III);

[0119] (ii) contacting a compound of formula (III) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having enal-cleaving enzyme activity to produce a compound of formula (IV);

[0120] (iii) contacting the compound of formula (IV) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having Baeyer-Villiger monooxygenase (BVMO) enzyme activity to produce a compound of formula (V);

[0121] (iv) contacting the compound of formula (V) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having esterase enzyme activity to produce a compound of formula (VI); and,

[0122] (v) contacting the compound of formula (VI) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having terpene cyclase enzyme activity to produce a compound of formula (I).

[0123] For the sake of clarity, by the expression “any one of its stereoisomers”, or the similar, it is meant the normal meaning understood by a person skilled in the art, i.e. that the invention compound can be a pure stereoisomer such as an enantiomer or a diastereomer (e.g. in relation to the configuration E or Z of any of the double bonds or in relation to the configuration R or S of any of the chiral carbon centers).

[0124] According to any of the aspects or embodiments of the invention, said compound can be in the form of any of its steroisomers or of a mixture thereof, e.g. the invention relates to compositions of matter comprising one or more forms of the compound of formula (I), having the same chemical structure but differing by the configuration of the chiral centers.

[0125] In particular, compound (I) can be in the form of a mixture comprising stereoisomer la (formula la) and wherein said stereoisomer la represents at least 50 %, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% or more of the total mixture. (formula la)

[0126] Alternatively, compound (I) can be in the form of a mixture comprising stereoisomer lb (formula lb) and wherein said stereoisomer lb represents at least 50 %, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% or more of the total mixture. (formula lb)

[0127] In one embodiment, more than 97% of the compound of formula (I) is in the form of formula (la) and / or (lb).

[0128] According to any of the aspects or embodiments of the invention, compounds of formula (II) to (VI) can be in the form of its E or Z isomer or of a mixture thereof. In particular, any one of compounds of formula (II) to (VI) can be in the form of a mixture consisting of stereoisomer E and Z and wherein said stereoisomer Ila, Illa, IVa, Va or Via represent at least 50 % of the total mixture, or even at least 75% (i.e a mixture E / Z comprised between 75 / 25 and 100 / 0).

[0129] Step (i) of the process of the invention

[0130] Step (i) of the process of the invention relates to contacting a compound of formula (II) with a polypeptide having ADH enzyme activity. (formula II)

[0131] The compound of formula (II) is also known as geranylgeraniol, i.e. 3,7,11 ,15-tetramethylhexadeca- 2,6,10,14-tetraen-1-ol, CAS No 7614-21-3.

[0132] The compound of formula (II) may be present in any one of its stereoisomers or a mixture thereof. Specifically, the compound may have the following structures and isoforms.

[0133] (formula Ila)

[0134] (2E,6E,10E)-Geranylgeraniol; (2E,6E,10E)-3,7,11 ,15-tetramethylhexadeca-2,6,10,14-tetraen-1-ol; CAS

[0135] No 24034-73-9.

[0136] (2Z,6E,10E)-Geranylgeraniol; (2Z,6E,10E)-3,7,11 ,15-tetramethylhexadeca-2,6,10,14-tetraen-1-ol; CAS

[0137] No 57784-25-5.

[0138] (2Z,6Z,10Z)-Geranylgeraniol; (2Z,6Z,10Z)-3,7,11 ,15-tetramethylhexadeca-2,6,10,14-tetraen-1-ol; CAS No 1945-42-2.

[0139] Step (i) relates to the use of a polypeptide having ADH enzyme activity.

[0140] An “alcohol dehydrogenase” (ADH) in the context of the present invention refers to a polypeptide having the ability to oxidize an alcohol to the corresponding aldehyde in the presence of NAD+or NADP+as cofactor. Such enzymes are members of the E.C. families 1.1.1.1 (NAD+dependent) or 1.1.1.2 (NADP+dependent). More particularly, an ADH of the invention has the ability to oxidize linear terpenoid alcohols to the respective carbonyl compounds in particular to the corresponding aldehydes, like geranylgeraniol to geranylgeranial. ADHs, as used herein, may either be endogenously present in the respective biocatalytic process or may be exogenous.

[0141] “Alcohol dehydrogenase enzyme activity” is determined under “standard conditions” as described herein below: It can be determined using recombinant alcohol dehydrogenase (ADH) polypeptide expressing host cells, disrupted ADH polypeptide expressing cells, fractions of these or enriched or purified ADH polypeptide, in a culture medium or reaction medium, preferably buffered, having a pH in the range of 6 to 11 , preferably 7 to 9, at a temperature in the range of about 20 to 45 °C, like about 25 to 40 °C, preferably 25 to 32 °C and in the presence of a reference substrate, here in particular geranylgeraniol, either added at an initial concentration in the range of 1 to 100 pM, preferably 5 to 50 pM, in particular 30 to 40 pM, or endogenously produced by the host cell. For in-vitro assays a cofactor selected from NADH and NADPH has to be added in a suitable easily to be determined concentration. The conversion reaction to form the respective aldehyde compounds, like geranylgeranial is conducted from 10 min to 5 h, preferably about 1 to 2 h. The oxidation product may then be determined in conventional matter, for example after extraction with an organic solvent, like ethyl acetate.

[0142] A further method to evaluate the oxidation of geranylgeraniol to geranylgeranial by ADHs is described in Example 3.

[0143] A preferred embodiment of the invention is wherein the polypeptide having said ADH enzyme activity comprises at least one or more sequence motifs selected from:

[0144] . CHTD (SEQ ID NO: 228) as for example in SEQ ID NO: 11 , 12, 13, 14, 17, 18, 19 or 20;

[0145] . GHEGxG (SEQ ID NO: 229) as for example in SEQ ID NO: 11 , 12, 13, 14, 17, 18, 19 or 20;

[0146] . LxCGxxTGxGA (SEQ ID NO: 230) as for example in SEQ ID NO: 11 , 12, 13, 14, 17, 18, 19 or 20;

[0147] . Gx[VI]GL (SEQ ID NO: 231) as for example in SEQ ID NO: 11 , 12, 13, 14, 15, 17, 18 , 19 or 20;

[0148] . LxxxG[LVI][PA] (SEQ ID NO: 232) as for example in SEQ ID NO: 11 , 12, 15, 17, 18 , 19 or 20;

[0149] . GxVxAl (SEQ ID NO: 233) as for example in SEQ ID NO: 16 or 21 ; and

[0150] . YxATKxA (SEQ ID NO: 234) as for example in SEQ ID NO: 16 or 21 ; wherein in the above motifs, residues x represent independently of each other any natural amino acid residue in a polypeptide having ADH activity. Ambiguities are indicated by listing the acceptable amino acids for a given position between brakets. For example, [VI] stands for V (valine), or I (isoleucine).

[0151] Preferably, the polypeptide having said ADH enzyme activity comprises: CHTD (SEQ ID NO: 228), GHEGxG (SEQ ID NO: 229), LxCGxxTGxGA (SEQ ID NO: 230) and Gx[VI]GL (SEQ ID NO: 231) motifs, as for example in SEQ ID NO: 11 , 12, 13, 14, 17, 18, 19 or 20;

[0152] Preferably, the polypeptide having said ADH activity comprises: CHTD (SEQ ID NO 228), GHEGxG (SEQ ID NO: 229), LxCGxxTGxGA (SEQ ID NO: 230), Gx[VI]GL (SEQ ID NO: 231) and LxxxG[LVI][PA] (SEQ ID NO: 232) motifs, as for example in SEQ ID NO: 11 , 12, 17, 18 ,19 or 20.

[0153] A preferred embodiment of the invention is wherein the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 11 to 21 . Preferably, the polypepeptide has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to the amino acid sequence provided in SEQ ID NO: 11 or 21 . Preferably, the polypeptide has the amino acid sequence provided in SEQ ID NO: 11 or 21. Step (ii) of the process of the invention

[0154] Step (ii) of the process of the invention relates to contacting a compound of formula (III) with a polypeptide having enal-cleaving enzyme activity. (formula III)

[0155] The compound of formula (III) is also known as geranylgeranial, i.e. 3,7,11 ,15-tetramethylhexadeca- 2,6,10,14-tetraenal; CAS No 32480-11-8.

[0156] The compound of formula (III) may be present in any one of its stereoisomers or a mixture thereof. Specifically, the compound may have the following structures and isoforms:

[0157] (2Z,6E,10Z)-geranylgeranial; (2Z,6E,10Z)-3,7,11 ,15-Tetramethyl-2,6,10,14-hexadecatetraenal.

[0158]

[0159] (2Z,6E,10Z)-geranylgeranial; (2Z,6E,10Z)-3,7,11 ,15-Tetramethyl-2,6,10,14-hexadecatetraenal.

[0160] (2Z,6Z,10Z)-geranylgeranial; (2Z,6Z,10Z)-3,7,11 ,15-Tetramethyl-2,6,10,14-hexadecatetraenal.

[0161] Step (ii) relates to the use of a polypeptide having enal-cleaving enzyme activity.

[0162] An “enal-cleaving enzyme” or “enal-cleaving protein” or “enal-cleaving polypeptide” in the context of the present invention designates an “a,p-unsaturated aldehyde carbon-carbon double bond-cleaving enzyme”, which also may be called a “a,p-unsaturated aldehyde C=C bond-cleaving enzyme” or “a,p-unsaturated aldehyde C=C-cleaving enzyme” or a “enal C=C-cleaving enzyme”. The enal-cleaving protein of the invention, based on protein domain organization, may also be described as a member of the “DUF4334 protein family” and / or as a member of the “GXWXG protein family” (SEQ ID NO: 263). Examples of such enzymes can be found in literature; for example, in WO2021005097.

[0163] More particularly, an enal cleaving enzyme of the invention has the ability to cleave terpenoid compounds containing an a,p-unsaturated aldehyde group, in particular geranylgeranial to farnesylacetone. “Enal-cleaving enzyme activity” is determined under “standard conditions” as described herein below. It can be determined using recombinant enal-cleaving polypeptide expressing host cells, disrupted enal-cleaving polypeptide expressing cells, fractions of these or enriched or purified enal-cleaving polypeptide, in a culture medium or reaction medium, preferably buffered, having a pH in the range of 6 to 1 1 , preferably 7 to 9, at a temperature in the range of about 20 to 45 °C, like about 25 to 40 °C, preferably 25 to 32 °C and in the presence of a reference substrate, here in particular geranylgeranial, either added at an initial concentration in the range of 1 to 100 pM, preferably 5 to 50 pM, in particular 30 to 40 pM, or endogenously produced by the host cell. The conversion reaction to form the respective cleavage product, like farnesylacetone is conducted from 10 min to 5 h, preferably about 1 to 2 h. The cleavage product may then be determined in conventional matter, for example after extraction with an organic solvent, like ethyl acetate.

[0164] The polypeptide having said enal-cleaving enzyme activity may be selected from the group of polypeptides containing: a) at least one DUF4334 protein family domain having the Pfam ID number PF14232 (in particular within the C-terminal region of their amino acid sequence); b) at least one GXWXG (SEQ ID NO: 263) protein family domain having the Pfam ID number PF14231 (in particular within the N-terminal region of their amino acid sequence); and / or c) a domain retaining at least 90% sequence identity to PF14232 or PF14231 .

[0165] In particular, a polypeptide of the invention having enal-cleaving enzyme activity is identified as a member of the DUF4334 protein family comprising said domain PF14232 if it matches with said domain with an e- value of less than 1x1 O’5, or less than 1x1 O’10, or less than 1x1 O’15, or less than 1x1 O’20, or less than 1x10’25, or less than 1x1 O’30, or less than or equal to 1x1 O’35, in particular in a range of 1x1 O’20to 1x1 O’32and more particular in a range of 1x1 O’25to 1x1 O’31.

[0166] In particular, a polypeptide having enal-cleaving enzyme activity is identified as a member of GXWXG (SEW ID NO: 263) protein family comprising said domain PF14231 if it matches with an e-value of less than 1x1 O’5, or less than 1x1 O’10, or less than 1x1 O’15, or less than 1x1 O’20, or less than 1x1 O’25, or less than 1x1 O’30, or less than or equal to 1x1 O’35, in particular in a range of 1x1 O’20to 1x1 O’30.

[0167] As the query sequence the sequence of a polypeptide having enal-cleaving enzyme activity is applied.

[0168] For example, the following website may be applied for the search and calculating such e-value: http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or http: / / www.ebi.ac.uk / Tools / pfa / pfamscan / .

[0169] Furthermore, the polypeptide having said enal-cleaving enzyme activity may be selected from the group of polypeptides that comprise at least one or more sequence motifs / domains selected from:

[0170] . G-[Y or “-“]-x-W-x-G-x-x-[F,L or l]-x-[T,S or R]-G-[H or D] (also expressed as GxxWxGxxxxxGx) set forth in SEQ ID NO: 235, or any partial motif thereof comprising up to 10 or up to 5 consecutive amino acid residues, as for example corresponding to residues in positions 1-8 or 9-13 of SEQ ID NO: 235 Here, X2 can be Y or can be deleted; X3 can be any naturally occurring amino acid; X5 can be any naturally occurring amino acid; X7 can be any naturally occurring amino acid; X8 can be any naturally occurring amino acid; X9 can be F, L, or I; X10 can be any naturally occurring amino acid; X1 1 can be R, S, or T; X13 can be H or D;

[0171] . W-[Y, A or V]-G-K-x-[F or Y]-x-[S or D] (also expressed as WxGKxxxx) set forth in SEQ ID NO: 236, or any partial motif thereof comprising up to 4 consecutive amino acid residues, as for example corresponding to residues in positions 1-4 or 5-8 of SEQ ID NO: 236. Here, X2 can be A, V, or Y; X5 can be any naturally occurring amino acid; X6 can be F or Y; X7 can be any naturally occurring amino acid; X8 can be D or S;

[0172] . [G or S]-x-[A or G]-x-[L or V]-x-x-x-x-[F, Y or L]-R-G-x-V (also expressed as xxxxxxxxxxRGxV) set forth in SEQ ID NO:237, or any partial motif thereof comprising up to 10 or up to 5 consecutive amino acid residues, as for example corresponding to residues in positions 1-8 or 9-14 of SEQ ID NO:237. Here X1 can be G or S; X2 can be any naturally occurring amino acid; X3 can be A or G; X4 can be any naturally occurring amino acid; X5 can be L or V; X6 can be any naturally occurring amino acid; X7 can be any naturally occurring amino acid; X8 can be any naturally occurring amino acid; X9 can be any naturally occurring amino acid; X10 can be F, L, or Y; X13 can be any naturally occurring amino acid; and

[0173] . [M or L]-[V or l]-Y-D-x-x-P-[l or V]-x-D-[H or S]-[F or L] (also expressed as xxYDxxPxxDxx) set forth in SEQ ID NO:238, or any partial motif thereof comprising up to 10 or up to 5 consecutive amino acid residues, as for example corresponding to residues in positions 1 -6 or 7-12 of SEQ ID NO:238. Here X1 can be L or M; X2 can be I or V; X5 can be any naturally occurring amino acid; X6 can be any naturally occurring amino acid; X8 can be I or V; X9 can be any naturally occurring amino acid; X11 can be H or S; X12 can be F or L; wherein the numbering of X (e.g. X2) corresponds to its position in the relevant sequence. For example, X2 corresponds to X at position 2 in the relevant sequence; and in the above motifs, residues x represent independently of each other any natural amino acid residue, and wherein optionally in each of the above motifs, 1 , 2, 3, 4 or 5 amino acid residues different from the x residues may be modified, for example by amino acid substitution, in particular by conservative substitutions, provided that the enzyme retains, at least to analytically detectable extent, enal-cleaving enzyme activity. The function of the square brackets has been described above.

[0174] A preferred embodiment of the invention is wherein the polypeptide having enal-cleaving activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to the sequence provided in SEQ ID NO: 22.

[0175] Step (Hi) of the process of the invention

[0176] Step (iii) of the process of the invention relates to contacting a compound of formula (IV) with a polypeptide having BVMO enzyme activity. (formula IV)

[0177] The compound of formula (IV) is also known as farnesylacetone, i.e. 6,10,14-trimethylpentadeca-5,9,13- trien-2-one; CAS No 762-29-8.

[0178] The compound of formula (IV) may be present in any one of its stereoisomers or a mixture thereof. Specifically, the compound may have the following structures and isoforms:

[0179] (5Z,9E)-farnesylacetone; (5Z,9E)-6,10,14-trimethylpentadeca-5,9,13-trien-2-one; CAS No 1117-51-7.

[0180] (5Z,9Z)-farnesylacetone; (5Z,9Z)-6,10,14-trimethylpentadeca-5,9,13-trien-2-one; Cas No 3796-69-8.

[0181] Step (iii) relates to the use of a polypeptide having BVMO enzyme activity.

[0182] “Baeyer-Villiger monooxygenases” (BVMOs) are flavoenzymes and belong to the class of refers to a polypeptide having oxidoreductase activity (EC 1.14.13.X). They catalyze the oxidation of linear, cyclic (aromatic or non-aromatic) aldehydes or ketones to the corresponding esters or lactones, highly similar to the chemical Baeyer-Villiger oxidation. During the enzymatic oxidation one atom of molecular oxygen is incorporated into a carbon-carbon bond of a non-activated carbonyl compound. The BVMOs require NADPH or NADH as cofactor or accept both. They also require molecular oxygen as co-substrate. More particularly, a BVMO of the invention has the ability to oxidize terpene- derived aldehydes or ketones, like for example linear terpenoid carbonyl compounds, in particular farnesylacetone to the respective carbonyl ester.

[0183] “BVMO enzyme activity” is determined under “standard conditions” as described herein below: It can be determined using recombinant BVMO expressing host cells, disrupted BVMO expressing cells, fractions of these or enriched or purified BVMO enzyme, in a culture medium or reaction medium, preferably buffered, having a pH in the range of 6 to 1 1 , preferably 7 to 9, at a temperature in the range of about 20 to 45 °C, like about 25 to 40 °C, preferably 25 to 32 °C and in the presence of a reference substrate, here in particular farnesylacetone, either added at an initial concentration in the range of 1 to 100 pM, preferably 5 to 50 pM, in particular 30 to 40 pM, or endogenously produced by the host cell and in the presence of molecular oxygen. For in-vitro assays a cofactor selected from NADH and NADPH has to be added in a suitable easily to be determined concentration range of the conversion reaction to form the respective enzyme product, like homofarnesyl acetate in the case of farnesylacetone is conducted from 10 min to 5 h, preferably about 1 to 2 h. The BVMO product may then be determined in conventional matter, for example after extraction with an organic solvent, like ethyl acetate.

[0184] A further method to screen for BVMOs and evaluate the conversion of farnesylacetone to homofarnesyl acetate is described in Example 2.

[0185] The polypeptide having BVMO enzyme activity may be selected from:

[0186] (1) the group of polypeptides containing a flavin-containing monooxygenase (FMO) protein family domain having the Pfam ID number PF00743 within their amino acid sequence; or a domain retaining at least 90%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to PF00743;

[0187] In particular, a polypeptide having BVMO activity is identified as member of the FMO protein family comprising said domain PF00743 if it matches with said domain with an e-value of less than 1x10-5or less than 1x1010, or less than or equal to 1x10-15, or less than or equal to 1x10-18, in particular in a range of 1x1010to 1x10-18and more particular, in a range of 1x10-14to 1x10-17. As the query sequence, the sequence of a polypeptide having BVMO activity is applied.

[0188] For example, the following website may be applied for the search and calculating such e-value: http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or http: / / www.ebi.ac.uk / Tools / pfa / pfamscan / . and / or

[0189] (2) the group of polypeptides that comprise at least one or more of the sequence motifs / domains selected from:

[0190] . GxGxxG (SEQ ID NO: 239), as for example in any of SEQ ID NOs: 23 to 26. Here, X4 can be any naturally occurring amino acid, particularly A or I. The numbering of X corresponds to its position in the sequence.

[0191] . [GS]GxWxxxxYPGxxxD (SEQ ID NO: 240), as for example in any of SEQ ID NOs: 23 to 26;

[0192] . Gxxx[FY]xGxxx[HS]xxxW (SEQ ID NO: 241), as for example in any of SEQ ID NOs: 23 to 26; and . [KQ]x[VI]xx[IV]GxG (SEQ ID NO: 242), as for example in any of SEQ ID NOs: 23 to 26. wherein in the above motifs, residues x represent independently of each other any natural amino acid residue, and wherein optionally in each of the above motifs, 1 , 2, 3, 4 or 5 of the conserved amino acid residues (i.e. different from the x residues) may be modified, for example by amino acid substitution, in particular by conservative substitutions, provided that the enzymes retains, at least to analytically detectable extent, BVMO enzyme activity. The function of the square brackets has been described above.

[0193] A preferred embodiment of the invention is wherein the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 23 to 26. A preferred embodiment of the invention is wherein the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 25 or 26. Preferably, the polypeptide has the amino acid sequence provided in SEQ ID NO: 25 or 26.

[0194] Alternatively, the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 216 to 227.

[0195] Step (iv) of the process of the invention

[0196] Step (iv) of the process of the invention relates to contacting a compound of formula (V) with a polypeptide having esterase enzyme activity. (formula V)

[0197] The compound of formula (V) is also known as homofarnesyl acetate, i.e. 4,8,12-trimethyltrideca-3, 7,11- trien-1-yl acetate; CAS No 109813-25-4.

[0198] The compound of formula (V) may be present in any one of its stereoisomers or a mixture thereof.

[0199] Specifically, the compound may have the following structures and isoforms: (formula Va)

[0200] (3E,7E)- homofarnesyl acetate; (3E,7E)-4,8,12-trimethyltrideca-3,7,11-trien-1-yl acetate; CAS No 944346- 19-4.

[0201] (3Z,7Z)- homofarnesyl acetate; (3Z,7Z)-4,8,12-trimethyltrideca-3,7,11-trien-1-yl acetate.

[0202] Step (iv) relates to the use of a polypeptide having esterase enzyme activity.

[0203] An “esterase” refers to a polypeptide having hydrolase activity that splits esters into an acid and an alcohol in a chemical reaction with water (hydrolysis). Esterases in the context of the present invention are selected from the class of carboxylic ester hydrolases (EC 3.1.1.-), which splits off acyl groups, like acetyl or formyl groups, from the respective ester substrate. More particularly, an esterase of the invention has the ability to cleave terpenyl ester compounds, like homofarnesyl acetate, to form the corresponding alcohol, in particular homofarnesol.

[0204] “Esterase enzyme activity” is determined under “standard conditions” as described herein below: It can be determined using recombinant esterase polypeptide expressing host cells, disrupted esterase polypeptide expressing cells, fractions of these or enriched or purified esterase polypeptide, in a culture medium or reaction medium, preferably buffered, having a pH in the range of 6 to 11 , preferably 7 to 9, at a temperature in the range of about 20 to 45 °C, like about 25 to 40°C, preferably 25 to 32 °C and in the presence of a reference substrate, here in particular homofarnesyl acetate, either added at an initial concentration in the range of 1 to 100 pM preferably 5 to 50 pM, in particular 30 to 40 pM, or endogenously produced by the host cell. The conversion reaction to form the respective alcohol, in particular homofarnesol is conducted from 10 min to 5 h, preferably about 1 to 2 h. The detection and quantification of esterase product may then be determined in conventional matter, for example after extraction with an organic solvent, like ethyl acetate. A further method to evaluate the conversion of homofarnesyl acetate to homofarnesol by an esterase is described in Example 4.

[0205] A preferred embodiment of the invention is wherein the polypeptide having said esterase activity comprises at least one or more sequence motifs selected from:

[0206] . AxWxVxxRLAPE (SEQ ID NO: 243), as for example in SEQ ID NO: 27 or 28;

[0207] . GASAGGGLxA (SEQ ID NO: 244), as for example in SEQ ID NO: 27 or 28;

[0208] . VxQLLxYPMLDDR (SEQ ID NO: 245), as for example in SEQ ID NO: 27 or 28; and,

[0209] . ARxxDLSGLPxT (SEQ ID NO: 246), as for example in SEQ ID NO: 27 or 28; wherein in the above motifs, residues x represent independently of each other any natural amino acid residue, and wherein optionally in each of the above motifs, 1 , 2, 3 or 4 amino acid residues different from the x residues may be modified, for example by amino acid substitution, in particular by conservative substitutions, provided that the enzyme retains, at least to analytically detectable extent, esterase enzyme activity.

[0210] Preferably, the polypeptide having said esterase activity comprises: AxWxVxxRLAPE (SEQ ID NO: 243), GASAGGGLxA (SEQ ID NO: 244), VxQLLxYPMLDDR (SEQ ID NO: 245), and ARxxDLSGLPxT (SEQ ID NO: 246), as for example in SEQ ID NO: 27 or 28.

[0211] A preferred embodiment of the invention is wherein the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 27 or 28. A preferred embodiment of the invention is wherein the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO 28. Preferably, the polypeptide has the amino acid sequence provided in SEQ ID NO: 28.

[0212] Step (v) of the process of the invention

[0213] Step (v) of the process of the invention relates to contacting a compound of formula (VI) with a polypeptide having terpene cyclase enzyme activity.

[0214] The compound of formula (VI) is also known as homofarnesol, 4 ,8,12-trimethyltrideca-3,7,11-trien-1-ol CAS No 35826-67-6.

[0215] The compound of formula (VI) may be present in any one of its stereoisomers or a mixture thereof. Specifically, the compound may have the following structures and isoforms:

[0216] (3Z,7E)-homofarnesol; (3Z,7E)-4,8, 12-trimethyltrideca-3, 7, 11 -trien-1 -ol; CAS No 138152-06-4.

[0217] (3Z,7Z)-homofarnesol; (3Z,7z)-4,8,12-trimethyltrideca-3,7,11 -trien-1 -ol; CAS No 138152-08-6.

[0218] Terpene cyclases are divided into two categories depending on the way the initial carbocation is generated. In class I (or type I) terpene cyclase, the diphosphate group of the linear terpenoid precursor is abstracted to form an allylic carbocation on the terpene moiety. In class II, the initial carbocation is formed by protonation of a double bond or epoxy group in the terpene carbon chain. Thus, class I cyclase necessarily use substrates with a diphosphate group, while class II cyclase (since they do not need a diphosphate group for the generation of the initial carbocation) can use terpenoids as substrates.

[0219] For all terpene cyclases, the generated reactive carbocation species triggers the subsequent cascade reaction including carbocation reactions with double bonds, alkyl-shifts, hydride shifts or carbon-carbon bound formation. The reaction can be terminated by deprotonation of a carbon atom adjacent to the carbocation or by quenching of the carbocation with a hydroxyl group or molecule of water.

[0220] The type II activity in terpene cyclases is associated with aspartate-rich conserved motifs.

[0221] Typical examples of class II terpene cyclases are the class II diterpene cyclases catalyzing the protonation- initiated cyclization of geranylgeranyl-diphosphate into for example, labdadienyl-diphosphate intermediates or other cyclic diphosphate intermediates (Peters, R. J. (2010). Nat. Prod. Rep. 27, 1521-1530; Zerbe, P. et al (2015). Plant J. 83, 783-793). Squalene cyclases (SHCs) constitute a classical example of class II terpene cyclases where the substrate does not contain a diphosphate functional group. The squalene cyclase enzyme family comprise squalene cyclases and 2,3-oxidosqualene cyclases and enzymes catalyzing mechanistically related cyclization reactions. Squalene cyclases catalyze a protonation-initiated cyclization cascade of a linear terpene to a cyclic compound. Thus, squalene cyclases are class II terpene cyclases. The squalene family includes for example squalene-hopene cyclases catalyzing the cyclization of squalene to hopene (EC 5.4.99.17) and squalene-hopanol cyclases catalyzing the cyclization of squalene to hopan-22-ol (EC 4.2.1.129). Tetraprenyl-p-curcumene-sporulenol cyclases catalyze similar class II cyclization of linear terpene substrate (EC 4.2.1 .137). It was shown that tetraprenyl-p-curcumene-sporulenol cyclase can also catalyze the cyclization of squalene (Sato, T., et al. (2011). Journal of the American Chemical Society 133(44): 17540-17543), thus tetraprenyl-p-curcumene-sporulenol cyclases are also members of the squalene cyclase family.

[0222] Squalene cyclase polypeptides have typically a length between 600 and 800 amino acids and are membrane-associated proteins. They bind to the surface of cellular membranes but do not contain a transmembrane region. Squalene cyclases are classified in the IPR018333 family of the InterPro protein sequence classification database (https: / / www.ebi.ac.uk / interpro / entry / lnterPro / IPR018333 / ) (InterPro release 93.0, 2nd March 2023). The structure of squalene cyclases is organized in two domains comprising several alpha-helices, recognized as the p-domain and y-domain or the py-domain architecture (Christianson DW, Chem. Rev, 2017, 117, 11570-11648). The two domains have characteristic sequence signatures as described in the Pfam database under the Pfam Squalene-hopene cyclase N-terminal domain (PF13249) and Squalene-hopene cyclase C-terminal domain and (PF13243) (Pfam 35.0 released, 19 November 2021). The presence of the IPR018333, PF13249 or PF13243 protein sequences signatures can be predicted using the NCBI conserved domain search tool (https: / / www.ncbi.nlm.nih.gov / Structure / cdd / wrpsb.cgi) or the InterProScan tool (http: / / www.ebi.ac.uk / interpro / search / sequence / ).

[0223] The squalene cyclase polypeptide contains characteristic conserved amino acid motifs located along the sequence and associated with the protein architecture and enzymatic reaction. In particular, the squalene cyclase contains at least one or more amino acid motifs selected from:

[0224] ■ [SP][TP][VIL]WDTx[LWI] (SEQ ID NO: 247),

[0225] . PGG[WF][GYA]F (SEQ ID NO: 248),

[0226] . PDxDD[TAS][TIAS] (SEQ ID NO: 249),

[0227] . [MIL]QxxxG[GA][WF]x[AS][FY] (SEQ ID NO: 250),

[0228] . Qxxx[GH]xWxG[RK]WGxx[YF]xYG (SEQ ID NO: 251), . Qxx[DN]G[GS][WF][GS]ExxxS (SEQ ID NO: 252), and . [STA]xx[SFN][QC]T[AGT]W[AS][LIV]xx[LQ] (SEQ ID NO: 253).

[0229] The motif sequences are described using the standard IUPAC one-letter codes for the amino acids. Ambiguities are indicated by listing the acceptable amino acids for a given position between brackets. For example, [SP] or [S or P] stands for S (serine), or P (proline). The “x” represents positions where independently of each other any natural amino acid residue is present. The function of the square brackets has been described above.

[0230] Meroterpenoids are hybrid secondary metabolites derived from mixed biosynthetic pathways and are partially derived from a terpenoid co-substrate (Cornforth, J.W. Terpenoid biosynthesis. Chem. Br. 1968, 4, 102-106). The non-terpenoid part can originate for example from polyketides, alkaloids, phenols, or amino acids biosynthetic pathway. Large chemical diversity is found among meroterpenoids, in particular in bacteria and in fungi.

[0231] The meroterpenoids biosynthetic pathways follow several modular biosynthetic steps. In the first step, the building blocks are generated from the corresponding biosynthetic pathway (e.g. terpenoids, polyketides). The terpenoid and non-terpenoid parts are assembled by prenyltransferases. The precursors of the terpenoid parts are generally linear terpenoid-diphosphates such as geranyl-diphosphate, farnesyl- diphosphate or geranylgeranyl-diphosphate.

[0232] In the following step, the linear polyene terpenoid part of the hybrid precursor is cyclized to form a monocyclic or polycyclic structure. This cyclization is catalyzed by a specific class of non-canonical class II terpene cyclases named meroterpenoid cyclase first discovered in fungi (T. Itoh et al, 2010, 2, 858-864). The first discovered representative meroterpenoid cyclase is Pyr4 from Aspergillus fumigatus Af293 (Itoh, T„ et al. (2010). Nature Chemistry 2(10): 858-864).

[0233] In many meroterpenoids, the linear terpenoid precursor is first activated by a stereoselective epoxidation by a monooxygenase of one of the double-bonds. The meroterpenoid cyclases catalyze then the protonation of the epoxide moiety generating a reactive carbocation species and triggering a subsequent cascade reaction similar to other terpene cyclases. Some meroterpenoid cyclases can convert the isoprenic precursors to cyclized products without the involvement of a prior epoxidation step. These meroterpenoid cyclases are able to directly protonate the terminal double bond generating a reactive carbocation and catalyzing a cyclization. For example, MacJ from the fungi Penicillium terrestry w as the first identified fungi meroterpenoid cyclase using a type II double-bond protonation initiations reaction (Tang, M.-C., et al. (2017). Organic Letters 19(19): 5376-5379). Another example of meroterpenoid cyclase which initiates polyene cyclization by direct double bond protonation is DmtA1 from bacteria (Streptomyces youssoufiensis OUC68199) (Yao et al, Nat. Commun., 2018, 9, 4091).

[0234] Like other class II terpene cyclases, the carbocation generated by meroterpenoid cyclases triggers a cascade reaction generally starting by the attack of a double bond and generating monocyclic or polycyclic structure with a tertiary carbocation. The reaction is terminated either by deprotonation to form a double bond or by reacting with a water molecule to generate a tertiary alcohol. Typical cyclic structures found in meroterpenoids compounds contain drimane or labdane scafolds. The largest group of meroterpenoid cyclases are compact membrane-integrated proteins containing several (generally seven) transmembrane helices. This protein architecture based on transmembrane helices can easily be predicted using for example the TMHMM 2.0 server available at https: / / dtu.biolib.com / DeepTMHMM (Krogh, A., et al. (2001) J Mol Biol 305(3): 567-580.). In addition to the protein architecture, meroterpenoid cyclases differ from other class II cyclases such as the squalene cyclases by their smaller polypeptide size. The bacterial and fungal meroterpenoid cyclase polypetides have a length ranging from 150 to 550 residues. The transmembrane helices are located over a portion of the polypetide covering 180 to 300 amino acid and carry the catalytic domains.

[0235] Recently meroterpenoid cyclases having a protein architecture different from the membrane-integrated meroterpenoid cyclases were described. For example, MstE from the bacteria Scytonema sp. PCC 1002 is a soluble cyclase having a structure similar to canonical cyclases such as diterpene synthases and squalene cyclases, but nevertheless different, since it is a monodomain protein with only an a-domain (Moosmann, P., et al. (2020). Nat Chem 12(10): 968-972). Soluble bacterial meroterpenoid cyclase polypetides have length ranging from 150 to 550.

[0236] Several meroterpenoid cyclases catalyze reactions of cyclisation of the terpenoid part of the meroterpenoid hybrid precursor to labdane cyclic structures. However, the cyclization of linear terpenoids by meroterpenoid cyclases to labdane compounds have so far not been shown.

[0237] Meroterpenoid cyclase polypeptides contain characteristic conserved amino acid motifs located along the sequence and associated with the protein architecture or enzymatic reaction as follows:

[0238] Membrane-integrated meroterpenoid cyclase of bacterial origin containing at least one or more amino acid motifs selected from:

[0239] . [W]xxx[D]xx[ILVMN] (SEQ ID NO: 254);

[0240] . PxxAxxxNxxWE (SEQ ID NO: 255);

[0241] . MxxxFxxMLxxR (SEQ ID NO: 256); and

[0242] . RxxxxGQS (SEQ ID NO: 257).

[0243] Membrane-integrated meroterpenoid cyclase of fungal origin containing a least one or more amino acid motifs selected from:

[0244] . [WY]Exx[YFW] (SEQ ID NO: 258); and

[0245] . [DNE]xSYxxP (SEQ ID NO: 259).

[0246] Soluble meroterpenoid cyclases of bacterial origin containing a least one or more amino acid motifs selected from:

[0247] . GxWxxxW[WG]xxxxY (SEQ ID NO: 260);

[0248] . WxxxHxxV[TSA] (SEQ ID NO: 261); and

[0249] . GxWxD[FY] (SEQ ID NO: 262). The motif sequences are described using the standard IUPAC one-letter codes for the amino acids. Residues x represent independently of each other any natural amino acid residue, and wherein optionally in each of the above motifs, 1 , 2, 3 or 4 amino acid residues different from the x residues may be modified, for example by amino acid substitution, in particular by conservative substitutions, provided that the enzyme retains, at least to analytically detectable extent, its enzyme activity. The function of the square brackets has been described above.

[0250] Meroterpenoid cyclase polypeptides can be searched in sequences databases using for example the BLAST search tools (Tatiana et al, FEMS Microbiol Lett., 1999, 174:247-250, 1999) using as query sequences MacJ (SEQ ID NO: 71), DmTA1 (SEQ ID NO: 77) or MstE (SEQ ID NO: 76). The selection can further be refined by selecting sequences with appropriate length or containing the characteristic amino acid motifs as described above. The selection can also be refined bases using a prediction of the protein architecture, in particular by predicting the presence of the transmembrane helices as described above.

[0251] Particular examples of suitable standard conditions for each of the above-described enzyme activities may be taken from the Examples section below.

[0252] As discussed above, the terpene cyclase may be a squalene cyclase (SHC) or a meroterpenoid cyclase (MeroTPS).

[0253] For the avoidance of doubt, SHCs and meroterpenoid cyclases are distinct classes of enzymes which can be distinguished by physical characteristics. Furthermore, meroterpenoid cyclases can be classified as (i) bacterial membrane-integrated meroterpenoid cyclases; (ii) fungal membrane-integrated meroterpenoid cyclases; (iii) bacterial soluble meroterpenoid cyclases.

[0254] Table A below outlines the differences between SHCs and the different types of meroterpenoid cyclases.

[0255] Table A: enzyme characteristics Hence the skilled person can, from the information provided herein, readily identify whether an enzyme is a SHC enzyme or a class of meroterpenoid cyclase enzyme.

[0256] A preferred embodiment of the invention is wherein the terpene cyclase is a squalene cyclase (SHC) enzyme.

[0257] In PCT / EP2024 / 066253, SHC enzymes were shown to catalyze step (v) of the process. These are provided herein in SEQ ID NOs: 29 to 49, 78 to 89 and 265 to 279. The nucleic acid sequence encoding said SHC enzymes are provided herein in SEQ ID NOs: 123 to 153, 193 to 215 and 290 to 304.

[0258] In the context of the present invention, as provided in the accompanying Examples, the inventors have suprisingly found additional SHC enzymes which can be used in step (v) of the process of the invention.

[0259] Accordingly, a preferred embodiment of the invention is wherein the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323. In a further preferred embodiment, the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, the polypeptide having terpene cyclase enzyme activity is a polypeptide comprising the amino acid sequence of any of SEQ ID NOs: 315 to 323. Said polypeptide having terpene cyclase enzyme activity is a squalene cyclase enzyme.

[0260] A preferred embodiment of the invention is wherein the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 316, 318, 320 and 323. In a further preferred embodiment, the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 316, 318, 320 and 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, the polypeptide having terpene cyclase enzyme activity is a polypeptide comprising the amino acid sequence of any of SEQ ID NOs: 316, 318, 320 and 323. Said polypeptide having terpene cyclase enzyme activity is a squalene cyclase enzyme.

[0261] A preferred embodiment of the invention is wherein the nucleic acid sequence encoding the polypeptide having terpene cyclase enzyme activity (i.e. the SHC enzyme) according to the invention is a nucleic acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 326 to 334. A preferred embodiment of the invention is wherein the nucleic acid sequence encoding the polypeptide having terpene cyclase enzyme activity (i.e. the SHC enzyme) according to the invention is a nucleic acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 327, 329, 331 and 334.

[0262] Embodiments of the first aspect of the invention

[0263] The first aspect of the process of the invention comprises using (i) a polypeptide having ADH enzyme activity, (ii) a polypeptide having enal-cleaving enzyme activity; (iii) a polypeptide having BVMO enzyme activity; (iv) a polypeptide having esterase enzyme activity; and (v) a polypeptide having terpene cyclase enzyme activity.

[0264] An embodiment of the first aspect of the invention is a process forthe preparation of a compound of formula (I) in the form of any one of its stereoisomers or a mixture thereof, comprising:

[0265] (i) contacting a compound of formula (II) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having alcohol dehydrogenase (ADH) enzyme activity to produce a compound of formula (III);

[0266] (ii) contacting the compound of formula (III) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having enal- cleaving enzyme activity to produce a compound of formula (IV);

[0267] (iii) contacting the compound of formula (IV) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having Baeyer- Villiger monooxygenase (BVMO) enzyme activity to produce a compound of formula (V);

[0268] (iv) contacting the compound of formula (V) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having esterase enzyme activity to produce a compound of formula (VI); and

[0269] (v) contacting the compound of formula (VI) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having terpene cyclase enzyme activity to produce a compound of formula (I), wherein the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323.

[0270] Alternatively, in step (v) of the above embodiment of the invention:

[0271] . the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, said polypeptide comprises the amino acid sequence of any of SEQ ID NOs: 315 to 323; or

[0272] . the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 316, 318, 320 and 323. In a further embodiment, said polypeptide has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 316, 318, 320 and 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, said polypeptide comprises the amino acid sequence of any of SEQ ID NOs: 316, 318, 320, and 323. A further embodiment of this first aspect of the invention is wherein:

[0273] (i) the polypeptide having ADH enzyme activity comprises at least one or more motifs selected from CHTD (SEQ ID NO: 228), GHEGxG (SEQ ID NO: 229), LxCGxxTGxGA (SEQ ID NO: 230), Gx[VI]GL (SEQ ID NO: 231), LxxxG[LVI][PA] (SEQ ID NO: 232), GxVxAl (SEQ ID NO: 233) and YxATKxA (SEQ ID NO: 234); wherein, residues x represent independently of each other any natural amino acid residue;

[0274] (ii) the polypeptide having enal-cleaving enzyme activity is selected from the group of polypeptides containing:

[0275] (a) at least one DUF4334 protein family domain having the Pfam ID number PF14232,

[0276] (b) at least one GXWXG (SEQ ID NO: 263) protein family domain having the Pfam ID number PF142311, and / or

[0277] (c) a domain retaining at least 90% sequence identity to PF14232 or PF14231 ;

[0278] (iii) the polypeptide having BVMO enzyme activity is selected from the group of polypeptides comprising:

[0279] (a) a flavin-containing monooxygenase (FMO) protein family domain having the Pfam ID number PF00743 within their amino acid sequence or a domain retaining at least 90%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to PF00743, and / or

[0280] (b) at least one or more of motifs selected from:

[0281] . GxGxxG (SEQ ID NO: 239),

[0282] . [GS]GxWxxxxYPGxxxD (SEQ ID NO: 240),

[0283] . Gxxx[FY]xGxxx[HS]xxxW (SEQ ID NO: 241), and

[0284] . [KQ]x[VI]xx[IV]GxG (SEQ ID NO: 242); wherein, residues x represent independently of each other any natural amino acid residue; and / or,

[0285] (iv) the polypeptide having esterase enzyme activity comprises at least one or more motifs selected from AxVVxVxxRLAPE (SEQ ID NO: 243), GASAGGGLxA (SEQ ID NO: 244), VxQLLxYPMLDDR (SEQ ID NO: 245) and ARxxDLSGLPxT (SEQ ID NO: 246); wherein, residues x represent independently of each other any natural amino acid residue.

[0286] A further embodiment of this first aspect of the invention is wherein:

[0287] (i) the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 11 to 21 ;

[0288] (ii) the polypeptide having enal-cleaving enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22;

[0289] (iii) the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 23 to 26 and 216 to 227; preferably, to any of SEQ ID NOs: 23 to 26; and / or,

[0290] (iv) the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 27 and 28. A further embodiment of this first aspect of the invention is wherein:

[0291] (i) the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 11 or 21 ; preferably the ADH enzyme has the sequence of SEQ ID NO: 11 or 21 ;

[0292] (ii) the polypeptide having enal-cleaving enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22; preferably the enal-cleaving enzyme has the sequence of SEQ ID NO: 22;

[0293] (iii) the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 25 or 26; preferably the BVMO enzyme has the sequence of SEQ ID NO: 25 or 26; and / or,

[0294] (iv) the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO 28; preferably the esterase enzyme has the sequence of SEQ ID NO: 28.

[0295] Second aspect for the preparation of a compound of formula (I)

[0296] A second aspect of the invention provides a process fo the preparation of a compound of formula (I) in the form of any one of its stereoisomers or a mixture thereof, comprising contacting a compound of formula (VI) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having terpene cyclase enzyme activity to produce a compound of formula (I); and optionally, further comprising one or more step(s) selected from:

[0297] (i) contacting the compound of formula (II) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having alcohol dehydrogenase (ADH) enzyme activity to produce a compound of formula (III); (ii) contacting the compound of formula (III) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having enal- cleaving enzyme activity to produce a compound of formula (IV);

[0298] (iii) contacting the compound of formula (IV) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having Baeyer- Villiger monooxygenase (BVMO) enzyme activity to produce a compound of formula (V); and

[0299] (iv) contacting the compound of formula (V) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having esterase enzyme activity to produce a compound of formula (VI).

[0300] As can be appreciated in this second aspect of the invention, the process of preparing a compound of formula (I) from compound of formula (VI) comprises step (v) of the process for preparing a compound of formula (I) in the first aspect of the invention above. Accordingly, the embodiments set out in relation to step (v) of the process of preparing a compound of formula (I) in the first aspect of the invention can be used in the present aspect of the invention and are incorporated herein to the present aspect of the invention.

[0301] As can be also appreciated in this second aspect of the invention, optional steps (i) to (iv) correspond to the respective steps (i) to (iv) of the process for preparing a compound of formula (I) in the first aspect of the invention above. Accordingly, the embodiments set out in relation to steps (i) to (iv) of the process of preparing a compound of formula (I) in the first aspect of the invention can be used in the present aspect of the invention and are incorporated herein to the present aspect of the invention.

[0302] Embodiments of the second aspect of the invention

[0303] The second aspect of the process of the invention comprises using a polypeptide having terpene cyclase enzyme activity and optionally, one or more polypeptides selected from (i) a polypeptide having ADH enzyme activity, (ii) a polypeptide having enal-cleaving enzyme activity; (iii) a polypeptide having BVMO enzyme activity; and (iv) a polypeptide having esterase enzyme activity.

[0304] An embodiment of this second aspect of the invention is wherein the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323. In a further preferred embodiment, the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, the polypeptide having terpene cyclase enzyme activity is a polypeptide comprising the amino acid sequence of any of SEQ ID NOs: 322 and 323.

[0305] Alternatively, a further embodiment of this second aspect of the invention is wherein:

[0306] . the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to the sequence provided in SEQ ID NO: 323. In a further embodiment, said polypeptide has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to the sequence provided in SEQ ID NO: 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, said polypeptide comprises the amino acid sequence of SEQ ID NO: 323.

[0307] A further embodiment of this second aspect of the invention is wherein:

[0308] (i) the polypeptide having ADH enzyme activity comprises at least one or more motifs selected from CHTD (SEQ ID NO: 228), GHEGxG (SEQ ID NO: 229), LxCGxxTGxGA (SEQ ID NO: 230), Gx[VI]GL (SEQ ID NO: 231), LxxxG[LVI][PA] (SEQ ID NO: 232), GxVxAl (SEQ ID NO: 233) and YxATKxA (SEQ ID NO: 234); wherein, residues x represent independently of each other any natural amino acid residue;

[0309] (ii) the polypeptide having enal-cleaving enzyme activity is selected from the group of polypeptides containing:

[0310] (a) at least one DUF4334 protein family domain having the Pfam ID number PF14232,

[0311] (b) at least one GXWXG (SEQ ID NO: 263) protein family domain having the Pfam ID number PF142311, and / or

[0312] (c) a domain retaining at least 90% sequence identity to PF14232 or PF14231 ;

[0313] (iii) the polypeptide having BVMO enzyme activity is selected from the group of polypeptides comprising:

[0314] (a) a flavin-containing monooxygenase (FMO) protein family domain having the Pfam ID number PF00743 within their amino acid sequence or a domain retaining at least 90%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to PF00743, and / or

[0315] (b) at least one or more of motifs selected from:

[0316] . GxGxxG (SEQ ID NO: 239), . [GS]GxWxxxxYPGxxxD (SEQ ID NO: 240),

[0317] . Gxxx[FY]xGxxx[HS]xxxW (SEQ ID NO: 241), and

[0318] . [KQ]x[VI]xx[IV]GxG (SEQ ID NO: 242); wherein, residues x represent independently of each other any natural amino acid residue; and / or,

[0319] (iv) the polypeptide having esterase enzyme activity comprises at least one or more motifs selected from AxVVxVxxRLAPE (SEQ ID NO: 243), GASAGGGLxA (SEQ ID NO: 244), VxQLLxYPMLDDR (SEQ ID NO: 245) and ARxxDLSGLPxT (SEQ ID NO: 246); wherein, residues x represent independently of each other any natural amino acid residue.

[0320] A further embodiment of this second aspect of the invention is wherein:

[0321] (i) the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 11 to 21 ;

[0322] (ii) the polypeptide having enal-cleaving enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22;

[0323] (iii) the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 23 to 26 and 216 to 227; preferably, to any of SEQ ID NOs: 23 to 26; and / or,

[0324] (iv) the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 27 and 28.

[0325] A further embodiment of this second aspect of the invention is wherein:

[0326] (i) the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 11 or 21 ; preferably the ADH enzyme has the sequence of SEQ ID NO: 11 or 21 ;

[0327] (ii) the polypeptide having enal-cleaving enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22; preferably the enal-cleaving enzyme has the sequence of SEQ ID NO: 22;

[0328] (iii) the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 25 or 26; preferably the BVMO enzyme has the sequence of SEQ ID NO: 25 or 26; and / or,

[0329] (iv) the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO 28; preferably the esterase enzyme has the sequence of SEQ ID NO: 28. Third and further aspects for the preparation of a compound of formula (I)

[0330] A third aspect of the invention provides a process for the preparation of a compound of formula (I): in the form of any one of its stereoisomers or a mixture thereof, comprising: contacting a compound of formula (V) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having esterase enzyme activity to produce a compound of formula (VI); and contacting the compound of formula (VI) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having terpene cyclase enzyme activity to produce a compound of formula (I).

[0331] In a fourth aspect of the invention, the process of the third aspect of the invention further comprises: contacting a compound of formula (IV) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having BVMO enzyme activity to produce a compound of formula (V).

[0332] In a fifth aspect of the invention, the process of the fourth aspect of the invention further comprises: contacting a compound of formula (III) in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having enal- cleaving enzyme activity to produce a compound of formula (IV).

[0333] As can be appreciated in said third and further aspects of the invention, the steps comprising the enal- cleaving enzyme, the BVMO enzyme, the esterase enzyme and the terpene cyclase enzyme correspond to the respective steps (ii) to (v) of the process for preparing a compound of formula (I) in the first aspect of the invention. Accordingly, the embodiments set out in relation to steps (ii) to (v) of the process for preparing a compound of formula (I) in the first aspect of the invention can be used in the present aspects of the invention and are incorporated herein to the present aspects of the invention.

[0334] Embodiments of the third and further aspects of the invention

[0335] In one embodiment, in said third and further aspects of the invention, the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323.

[0336] Alternatively,

[0337] . the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, said polypeptide comprises the amino acid sequence of any of SEQ ID NOs: 315 to 323; or

[0338] . the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 316, 318, 320 and 323. In a further embodiment, said polypeptide has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 316, 318, 320 and 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, said polypeptide comprises the amino acid sequence of any of SEQ ID NOs: 316, 318, 320, and 323; or

[0339] . the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 322 and 323; or

[0340] . the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, said polypeptide comprises the amino acid sequence of any of SEQ ID NOs: 322 and 323; or . the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to the sequence provided in SEQ ID NO: 323. In a further embodiment, said polypeptide has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to the sequence provided in SEQ ID NO: 323 and has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, said polypeptide comprises the amino acid sequence of SEQ ID NO: 323.

[0341] A further embodiment of these aspects of the invention is wherein:

[0342] (ii) the polypeptide having ADH enzyme activity comprises at least one or more motifs selected from CHTD (SEQ ID NO: 228), GHEGxG (SEQ ID NO: 229), LxCGxxTGxGA (SEQ ID NO: 230), Gx[VI]GL (SEQ ID NO: 231), LxxxG[LVI][PA] (SEQ ID NO: 232), GxVxAl (SEQ ID NO: 233) and YxATKxA (SEQ ID NO: 234); wherein, residues x represent independently of each other any natural amino acid residue;

[0343] (ii) the polypeptide having enal-cleaving enzyme activity is selected from the group of polypeptides containing:

[0344] (a) at least one DUF4334 protein family domain having the Pfam ID number PF14232,

[0345] (b) at least one GXWXG (SEQ ID NO: 263) protein family domain having the Pfam ID number PF142311, and / or

[0346] (c) a domain retaining at least 90% sequence identity to PF14232 or PF14231 ;

[0347] (iii) the polypeptide having BVMO enzyme activity is selected from the group of polypeptides comprising:

[0348] (a) a flavin-containing monooxygenase (FMO) protein family domain having the Pfam ID number PF00743 within their amino acid sequence or a domain retaining at least 90%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to PF00743, and / or

[0349] (b) at least one or more of motifs selected from:

[0350] . GxGxxG (SEQ ID NO: 239),

[0351] . [GS]GxWxxxxYPGxxxD (SEQ ID NO: 240),

[0352] . Gxxx[FY]xGxxx[HS]xxxW (SEQ ID NO: 241), and

[0353] . [KQ]x[VI]xx[IV]GxG (SEQ ID NO: 242); wherein, residues x represent independently of each other any natural amino acid residue; and / or,

[0354] (iv) the polypeptide having esterase enzyme activity comprises at least one or more motifs selected from AxVVxVxxRLAPE (SEQ ID NO: 243), GASAGGGLxA (SEQ ID NO: 244), VxQLLxYPMLDDR (SEQ ID NO: 245) and ARxxDLSGLPxT (SEQ ID NO: 246); wherein, residues x represent independently of each other any natural amino acid residue.

[0355] A further embodiment of these aspects of the invention is wherein:

[0356] (i) the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 11 to 21 ; (ii) the polypeptide having enal-cleaving enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22;

[0357] (iii) the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 23 to 26 and 216 to 227; preferably, to any of SEQ ID NOs: 23 to 26; and / or,

[0358] (iv) the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 27 and 28.

[0359] A further embodiment of these aspects of the invention is wherein:

[0360] (i) the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 11 or 21 ; preferably the ADH enzyme has the sequence of SEQ ID NO: 11 or 21 ;

[0361] (ii) the polypeptide having enal-cleaving enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22; preferably the enal-cleaving enzyme has the sequence of SEQ ID NO: 22;

[0362] (iii) the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 25 or 26; preferably the BVMO enzyme has the sequence of SEQ ID NO: 25 or 26; and / or,

[0363] (iv) the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO 28; preferably the esterase enzyme has the sequence of SEQ ID NO: 28.

[0364] Forms of compounds produced by a process according to the aspects of the invention for the preparation of a compound of formula (I)

[0365] The first, second, third and further aspects for the preparation of a compound of formula (I) relate to a process for the preparation of a compound of formula (I).

[0366] The compound of formula (I) formula (I)

[0367] The compound of formula (I) is also known as 3a, 6, 6, 9a tetramethyldodecahydronaphtho[2,1-b]furan; CAS No 3738-00-9.

[0368] The compound of formula (I) may be present in any one of its stereoisomers or a mixture thereof. Specifically, the compound may have the following structures and isoforms: (formula la)

[0369] (3aR,5aS,9aS,9bR)-3a,6,6,9a-tetramethyldodecahydronaphtho[2,1-b]furan; CAS No 6790-58-5. (formula lb)

[0370] (3aS,5aR,9aR,9bS)-3a,6,6,9a-tetramethyldodecahydronaphtho[2,1-b]furan; CAS No 234431-64-2. (formula Ic)

[0371] (3aR,5aS,9aS,9bS)-3a,6,6,9a-tetramethyldodecahydronaphtho[2,1-b]furan. (formula Id)

[0372] (3aS,5aS,9aS,9bS)-3a,6,6,9a-tetramethyldodecahydronaphtho[2,1-b]furan.

[0373] A compound of formula (I) can have different enantiomeric forms. Some isomers have a preferred olfactory profile to alternative isomeric forms of the compound. In particular, the olfactory preferred form of compound (I) is the compound (la) and / or (lb).

[0374] To improve the yield or ratio of compound (la) and / or (lb) to other isomeric forms of compound (I), achieving high selectivity in the generation of compound (VI) plays an important role. Specifically, the ratio of the E / Z isomers at the 3,4-double bond of compound (VI) has significant importance. The preferred form of compound (VI) is that of formula (Via).

[0375] Despite extensive efforts, using chemical methods to obtain a compound of formula (VI) with a E / Z ratio at the 3,4-double bond higher higher than 90:10 remains a challenge and is so far not accessible at large scale (Eichhorn, E. and F. Schroeder (2023). J Agric Food Chem).

[0376] The present invention achieves this goal by the use of highly selective enzymes. In the present invention, high selectivity of production of a compound of formula (VI) in the form of formula (Via) has been achieved, as demonstrated in the accompanying Examples. Chemical methods (as described in Eichhorn, E. and F. Schroeder (2023). J Agric Food Chem) do not allow to achieve such high selectivity. Compounds of the pathway with high a E / Z ratio can be obtained using enzymatic pathways. In particular geranylgeranyl- diphopshate and geranylgeraniol with high E / Z ratio of the double bonds can be achieve using enzymes and in particular, geranylgeranyl-diphopshate synthase. The configuration of the double bonds is retained in all intermediates of the pathway.

[0377] This subsequently leads to a process in which olfactory preferred forms of a compound of formula (I) are prepared. Hence, the process of the invention to prepare a compound of formula (I) lacking substantial amount of undesirable side products is of significant commercial importance. As can be shown in the accompanying Examples, more than 97% of the compound of formula (I) is in the form of formula (la) and / or (lb).

[0378] Accordingly, an embodiment of the aspects of the invention for the preparation of a compound of formula (I) is wherein more than 97% of the compound of formula (I) is in the form of formula (la) and / or (lb).

[0379] Also included in the scope of the present invention is a compound of formula (I) obtained or obtainable by an in vivo process. For the reasons outlined herein, this is an important advance in the preparation of this compound, from technical and commercial aspects. The compound of formula (I) may be prepared via any in vivo process, preferably using recombinant cells expressing enzymes which can be used in the pathway to synthesise this molecule. An embodiment of the present invention is wherein more than 97% of the compound of formula (I) is in the form of formula (la) and / or (lb).

[0380] A further embodiment of the aspects of the invention for the preparation of a compound of formula (I) is wherein:

[0381] (i) the compound of formula (I) is in the form of formula (la): (formula la)

[0382] (ii) the compound of formula (II) is in the form of formula (Ila):

[0383] (formula Ila);

[0384] (iii) the compound of formula (III) is in the form of formula (Illa): (formula Illa); (iv) the compound of formula (IV) is in the form of formula (IVa): (formula IVa);

[0385] (v) the compound of formula (V) is in the form of formula (Va): (formula Va);

[0386] (vi) the compound of formula (VI) is in the form of formula (Va):

[0387] (formula Via).

[0388] Preparation of a compound of formula (II)

[0389] Some aspects of the invention for the preparation of a compound of formula (I) provides a process of preparing a compound of formula (I) from the sequential biocatalytic conversion of a compound of formula (II).

[0390] The compound of formula (II) can be used as the starting substrate in the process of the invention, for example in the form of a purified compound preparation.

[0391] However, a preferred embodiment of the invention is wherein the process of the invention further comprises providing a compound of formula (II) by the sequential biocatalytic conversion of precursor compounds to the compound of formula (II).

[0392] Compounds of formulas (I) to (VI) are terpenoids.

[0393] Terpenoids is a large family of structurally diverse natural compounds. All terpenoids derive biosynthetically from two five-carbon units, isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP). IPP and DMAPP can be produced from different biosynthetic pathways such as the 2-C-methyl-D-erythritol-4- phosphate (MEP) pathway, the mevalonate (MVA) pathway or alternative MVA pathways (Dellas, N., et al. (2013) eLife 2: e00672). Alternatively, IPP and DMAPP can also be formed by successive enzymatic phosphorylation or by enzymatic pyrophosphorylation of their corresponding alcohols, isoprenol and prenol (Ma, X, et al. (2022). J Agric Food Chem 70(11): 3512-3520).

[0394] These terpene building blocks are condensed successively to form linear terpenoid precursors with various length and multiple of five carbon such as geranyl-diphosphate (GPP), farnesyl-diphosphate (FPP) or geranylgeranyl-diphosphate (GGPP) containing 10, 15 and 15 carbons, respectively. The condensation of the IPP and DMAPP is performed by a class of enzyme named prenyltransferases. Prenyltransferase enzymes catalyze the initial condensation reaction between IPP and DMAPP to give GPP, and the subsequent addition of IPP molecules to give FPP and then GGPP (Ogura, K., and Koyama, T. (1998). Chem. Rev. 98, 1263-1276). The successive condensation of DMAPP and IPP to GGPP can be performed by: i. the successive action of 3 prenyltransferases, a GPP synthase, a FPP synthase catalyzing the addition of one IPP to GPP and a GGPP synthase catalyzing the addition of one IPP to FPP; ii. the combination of 2 prenyltransferases, for example a FPP synthase catalyzing the condensation of one DMAPP and two IPP and a GGPP synthase catalyzing the addition of one IPP to FPP; and / or

[0395] Hi. the action of 1 prenyltransferase, for example a GGPP synthase capable of catalyzing the the successive condensation of tree IPP and one DMAPP.

[0396] Some terpenoids have simple linear structures, for example geranylgeraniol is a linear diterpene (with 20 carbons) containing a terminal hydroxyl group. Geranylgeraniol can be made from GGPP using either: i. an enzyme having pyrophosphatase activity such as a phosphatase; and / or ii. two enzymes having phosphatase activity and successively cleaving the two phosphate groups of GGPP.

[0397] Alternatively, geranylgeraniol can be made from GGPP using an enzyme from the class I terpene cyclase family (as described below) able to cleave the pyrophosphate group of GGPP but lacking the ability to catalyze the successive cyclization.

[0398] These enzymatic pathways thus use enzymes having phosphatase activities and allow the cleavage of a diphosphate group of GGPP and release of geranylgeraniol.

[0399] Alternatively, geranylgeraniol can be made from another linear diterpene. For example, from geranyllinallol or from a corresponding polyene for example from beta-springene. Enzymes such as dehydrataseisomerases can be used for such reaction. Dehydratase-isomerases can catalyze the reversibly isomerization reactions between linear terpene compounds having a terminal alcohol group or terminal double bound (Nestl, B. M., et al. (2017). Nature Chemical Biology 13(3): 275-281) (see figure 2).

[0400] Alternatively, geranylgeraniol can also be synthesized using chemical methods. For example, geranylgeraniol can be obtained from farnesene or farnesol by chain extension (Organic Syntheses, Vol. 84, p. 43-57 (2007).

[0401] The pathways leading to IPP and DMAPP and to geranyl-diphosphate (GPP), farnesyl-diphosphate (FPP) or geranylgeranyl-diphosphate (GGPP) is involved in the synthesis of terpenoids, a diverse class of molecules that play essential roles in primary metabolism and various cellular processes. Terpenoids are involved in numerous biological functions, including the synthesis of sterols, such as cholesterol in animals and phytosterols in plants, as well as the production of hormones, vitamins (such as vitamin E and K), and signaling molecules (such as ubiquinone and dolichol). Additionally, terpenoids are crucial for the formation of membrane lipids and post-translational modifications of proteins.

[0402] Therefore, the pathway leading to GPP, FPP and GGPP as described above is an important component of primary metabolism in all organisms as it provides the necessary precursors for the synthesis of essential terpenoid compounds involved in various physiological processes essential for the growth of the cells.

[0403] The majority of terpenoids compounds contain cyclic carbon scaffolds. The diversity of monocyclic and polycyclic carbon skeletons is due to the enzymatic conversion of the linear terpenoid precursors by Terpene Cyclases (TC), also referred to as terpene synthases. The cyclization reaction starts from a carbocation which reacts with electron rich double bonds leading to new carbon-bound formation. The outcome of the reaction is defined by the substrate folding and the nature of amino acid side chains in the enzyme active site.

[0404] Hence in a preferred embodiment of the invention, the process according to the aspects of the invention for the preparation of a compound of formula (I) further comprises one or more step(s) selected from:

[0405] 1) preparing (producing) geranylgeranyl-diphosphate (GGPP) from IPP and DMAPP using one or more polypeptides having prenyltransferase enzyme activity; and,

[0406] 2) preparing (producing) a compound of formula (II) from GGPP using one or more polypeptides having phosphatase enzyme activity.

[0407] Step 1) of this embodiment of the invention requires the use of one or more prenyltransferase enzyme(s).

[0408] The term ‘prenyltransferases’ represents a group of enzymes having the ability to condense successively five-carbon units such as isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP) to form linear terpenyl-diphosphate compounds such as geranyl-diphosphate (GPP), farnesyl-diphosphate (FPP) or geranylgeranyl-diphosphate (GGPP) containing 10, 15 and 15 carbons, respectively. Some prenyl transferases can add 5-carbon units to linear terpenyl-diphosphate compounds thereby extending the carbon chain length. An example of prenyl transferase are geranyl-diphosphate synthases (GGPP synthases) having the ability of producing GGPP from IPP and DMAPP or by adding 5 carbons to FPP. The term ‘Prenyltransferases’ also refers to a group of enzymes having the ability to transfer an isoprenoid subunit from a terpenyl-diphosphate compound, generally from a linear terpenyl-diphosphate compounds, to the non-terpenoid scaffold during the biosynthesis of meroterpenoids.

[0409] Step 1) of this embodiment of the invention may be performed by: i. the successive action of 3 prenyltransferases: a GPP synthase, a FPP synthase catalyzing the addition of one IPP to GPP and a GGPP synthase catalyzing the addition of one IPP to FPP; ii. the combination of 2 prenyltransferases: for example, a FPP synthase catalyzing the condensation of one DMAPP and two IPP and a GGPP synthase catalyzing the addition of one IPP to FPP; and / or iii. the action of 1 prenyltransferase for example a GGPP synthase capable of catalyzing the the successive condensation of tree IPP and one DMAPP.

[0410] Examples of prenyltransferase enzymes that can be used in this step of the process of the invention are well known in the art. For example, a GGPP synthase from Blakeslea trispora can be used for the biosynthesis of GGPP (sun et al, Biotechnol. Lett. 34 (11), 2077-2082 (2012))

[0411] Preferably, the prenyltransferase is a GGPP synthase. Preferably, the GGPP synthase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 1 or 2. Preferably, the GGPP synthase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 2. Preferably, the GGPP synthase has the amino acid sequence of SEQ ID NO: 2.

[0412] Step 2) of this embodiment of the invention requires the use of one or more enzymes having phosphatase activity.

[0413] The term “phosphatase” represents a group of enzymes that are known to remove phosphate or diphosphate groups from a precursor containing a phosphate or diphosphate group. A particular subgroup of phosphatases has the ability of removing phosphate or diphosphate group from a terpenyl precursor releasing inorganic phosphate and the corresponding terpenyl alcohol. For example, some phosphatases are known to remove the diphosphate group of GGPP to form geranylgeraniol. Phosphatases acting on terpenyl diphosphate are found in several enzyme classes.

[0414] The term “protein tyrosine phosphatase” represents a group of enzymes that are generally known to remove phosphate groups from phosphorylated tyrosine residues on proteins. A particular subgroup of said family as described in WQ202001 1883A1 are enzymes useful to remove diphosphate groups from phosphorylated terpene molecules (terpenyl diphosphate). In particular, phosphatases from the protein tyrosine phosphatase family having the Pfam ID number PF13350 can dephosphorylate GGPP to geranylgeraniol. Polypeptides can be scanned for matches against the Pfam protein family signature databases.

[0415] Phosphatases and in particular GGPP phosphatases can also be obtained from other protein families; for example, from the Phosphatidic Acid Phosphatases of type 2 (PAP2) protein family (IPR00326), Nudix- Hydrolase protein family (IPR015797) and Haloacid Dehalogenase-like (HAD-like) Hydrolases protein family (IPR041492). A method to screen for phosphatases and evaluate the conversion of geranylgeranyl diphosphate to geranylgeraniol is described in Example 5.

[0416] Examples of phosphatases enzymes that can be used in this step of the process of the invention are well known in the art. Preferably, the phosphatase is a GGPP phosphatase. Preferably, the GGPP phosphatase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any one of SEQ ID NOs: 3 to 10. Preferably, the GGPP phosphatase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 3. Preferably, the GGPP phosphatase has the amino acid sequence of SEQ ID NO: 3.

[0417] As described above, this embodiment of the invention provides a compound of formula (II) by the sequential biocatalytic conversion of precursor compounds to the compound of formula (II).

[0418] In this embodiment of the invention, the precursor compounds to the compound of formula (II) are provided from IPP and DMAPP.

[0419] A further embodiment of the invention is wherein the process according to the aspects of the invention for the preparation of a compound of formula (I) further comprises the preparation of IPP and DMAPP.

[0420] One means for the preparation of IPP and DMAPP is via the “mevalonate pathway”. The “mevalonate pathway” also known as the “isoprenoid pathway” or “HMG-CoA reductase pathway” is an essential metabolic pathway present in eukaryotes, archaea, and some bacteria. The mevalonate pathway begins with acetyl-CoA and produces two five-carbon building blocks called isopentenyl pyrophosphate (IPP) and dimethyl allyl pyrophosphate (DMAPP). Combining the mevalonate pathway with enzyme activity to generate the terpene precursors GPP, FPP or GGPP allows the recombinant cellular production of terpenes. The pathway is well known in the art. The list of enzymes required for the conversion of acetyl- CoA to IPP and DMAPP is provided below:

[0421] . Acetyl-CoA acetyltransferase (ACAT);

[0422] . 3-hydroxy-3-methylglutaryl-CoA synthase (HMG-CoA synthase);

[0423] . 3-hydroxy-3-methylglutaryl-CoA reductase (HMG-CoA reductase);

[0424] . Mevalonate kinase;

[0425] . Phosphomevalonate kinase;

[0426] . Mevalonate diphosphate decarboxylase;

[0427] . Isopentenyl diphosphate isomerase.

[0428] An alternative means for the preparation of IPP, and DMAPP is via the methylerythritol phosphate (MEP). The pathway is well known in the art. The list of enzymes required for the conversion of glyceraldehyde 3- phosphate (GAP) and pyruvate to IPP and DMAPP is provided below:

[0429] . 1-Deoxy-D-xylulose 5-phosphate synthase (DXS);

[0430] . 1-Deoxy-D-xylulose 5-phosphate reductoisomerase (DXR);

[0431] . 2-C-methyl-D-erythritol 4-phosphate cytidylyltransferase (MCT, IspD);

[0432] . 4-diphosphocytidyl-2-C-methyl-D-erythritol kinase (CMK, IspE);

[0433] . 2-C-methyl-D-erythritol 2,4-cyclodiphosphate synthase (MDS, IspF);

[0434] . 4-hydroxy-3-methylbut-2-en-1-yl diphosphate synthase (HDS, IDS); . 4-hydroxy-3-methylbut-2-en-1-yl diphosphate reductase (HDR).

[0435] Further alternative pathways for the preparation of IPP and DMAPP are known, see for example: Rinaldi, M. A., et al. (2022). Natural Product Reports 39(1): 90-118. https: / / doi.org / 10.1039 / D1 NP00025J (see part 3 of this article).

[0436] The inventors have therefore provided a complete biocatalytic route for the preparation of a compound of formula (I) from acetyl-CoA or glyceraldehyde 3-phosphate (GAP) and pyruvate. This multistep biocatalytic process has for the first time been described herein and constitutes a significant advance in the preparation of such compound; in particular, enabling in an in vivo process forthe preparation of a compound of formula (I).

[0437] Recombinant cells of the invention

[0438] The present invention also provides a recombinant cell wherein more than 97% of the compound of formula (I) is in the form of formula (la) and / or (lb). In particular, the present invention provides a recombinant cell comprising, capable of producing or producing a compound of formula (I) and optionally, one or more compound(s) of formula (II), formula (III), formula (IV), formula (V) and / or formula (VI).

[0439] Methods for preparing a recombinant cell comprising, capable of producing or producing a compound of formula (I) and optionally, one or more compound(s) of formula (II), formula (III), formula (IV), formula (V) and / or formula (VI) are provided herein.

[0440] In one embodiment, the invention provides a recombinant cell comprising, capable of producing or producing a compound of formula (I), wherein the recombinant cell comprises, is capable of functionally expressing or expresses a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323. Optionally, said recombinant cell may further comprise (i) a polypeptide having ADH enzyme activity, (ii) a polypeptide having enal-cleaving enzyme activity, (iii) a polypeptide having BVMO enzyme activity, and / or (iv) a polypeptide having esterase enzyme activity. In a further embodiment, one or more of said polypeptides may be a heterologous polypeptide. In particular, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In yet a further embodiment, the polypeptides having ADH enzyme activity, enal-cleaving enzyme activity, BVMO enzyme activity, esterase enzyme activity and terpene cyclase enzyme activity are heterologous polypeptides.

[0441] In another embodiment, the present invention provides a recombinant cell comprising, capable of producing or producing a compound of formula (I), wherein the recombinant cell comprises, is capable of functionally expressing or expresses: . a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323; and,

[0442] . one or more polypeptide(s) selected from (i) a polypeptide having ADH enzyme activity, (ii) a polypeptide having enal-cleaving enzyme activity, (iii) a polypeptide having BVMO enzyme activity, and (iv) a polypeptide having esterase enzyme activity. In a further embodiment, one or more of said polypeptides may be a heterologous polypeptide. In particular, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In yet a further embodiment, the polypeptides having ADH enzyme activity, enal-cleaving enzyme activity, BVMO enzyme activity, esterase enzyme activity and terpene cyclase enzyme activity are heterologous polypeptides.

[0443] In another embodiment, the present invention provides a recombinant cell comprising, capable of producing or producing a compound of formula (I), wherein the recombinant cell comprises, is capable of functionally expressing or expresses:

[0444] . a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323; and, . a polypeptide having esterase enzyme activity.

[0445] Optionally, said recombinant cell may further comprise (i) a polypeptide having ADH enzyme activity, (ii) a polypeptide having enal-cleaving enzyme activity, and (iii) a polypeptide having BVMO enzyme activity. In a further embodiment, one or more of said polypeptides may be a heterologous polypeptide. In particular, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In yet a further embodiment, the polypeptides having ADH enzyme activity, enal-cleaving enzyme activity, BVMO enzyme activity, esterase enzyme activity and terpene cyclase enzyme activity are heterologous polypeptides.

[0446] In another embodiment, the present invention provides a recombinant cell comprising, capable of producing or producing a compound of formula (I), wherein the recombinant cell comprises, is capable of functionally expressing or expresses:

[0447] . a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323;

[0448] . a polypeptide having esterase enzyme activity; and,

[0449] . a polypeptide having BVMO enzyme activity.

[0450] Optionally, said recombinant cell may further comprise (i) a polypeptide having ADH enzyme activity and (ii) a polypeptide having enal-cleaving enzyme activity. In a further embodiment, one or more of said polypeptides may be a heterologous polypeptide. In particular, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In yet a further embodiment, the polypeptides having ADH enzyme activity, enal-cleaving enzyme activity, BVMO enzyme activity, esterase enzyme activity and terpene cyclase enzyme activity are heterologous polypeptides.

[0451] In another embodiment, the present invention provides a recombinant cell comprising, capable of producing or producing a compound of formula (I), wherein the recombinant cell comprises, is capable of functionally expressing or expresses:

[0452] . a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323;

[0453] . a polypeptide having esterase enzyme activity;

[0454] . a polypeptide having BVMO enzyme activity; and, . a polypeptide having enal-cleaving enzyme activity.

[0455] Optionally, said recombinant cell may further comprise (i) a polypeptide having ADH enzyme activity. In a further embodiment, one or more of said polypeptides may be a heterologous polypeptide. In particular, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In yet a further embodiment, the polypeptides having ADH enzyme activity, enal-cleaving enzyme activity, BVMO enzyme activity, esterase enzyme activity and terpene cyclase enzyme activity are heterologous polypeptides.

[0456] In another embodiment, the present invention provides a recombinant cell comprising, capable of producing or producing a compound of formula (I), wherein the recombinant cell comprises, is capable of functionally expressing or expresses:

[0457] . a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323;

[0458] . a polypeptide having esterase enzyme activity;

[0459] . a polypeptide having BVMO enzyme activity;

[0460] . a polypeptide having enal-cleaving enzyme activity; and,

[0461] . a polypeptide having ADH enzyme activity.

[0462] In a further embodiment, one or more of said polypeptides may be a heterologous polypeptide. In particular, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In yet a further embodiment, the polypeptides having ADH enzyme activity, enal-cleaving enzyme activity, BVMO enzyme activity, esterase enzyme activity and terpene cyclase enzyme activity are heterologous polypeptides.

[0463] A futher embodiment of the recombinant cell of the invention is wherein:

[0464] (i) the polypeptide having ADH enzyme activity comprises at least one or more motifs selected from CHTD (SEQ ID NO: 228), GHEGxG (SEQ ID NO: 229), LxCGxxTGxGA (SEQ ID NO: 230), Gx[VI]GL (SEQ ID NO: 231), LxxxG[LVI][PA] (SEQ ID NO: 232), GxVxAl (SEQ ID NO: 233) and YxATKxA (SEQ ID NO: 234); wherein, residues x represent independently of each other any natural amino acid residue; (ii) the polypeptide having enal-cleaving enzyme activity is selected from the group of polypeptides containing:

[0465] (a) at least one DUF4334 protein family domain having the Pfam ID number PF14232,

[0466] (b) at least one GXWXG (SEQ ID NO: 263) protein family domain having the Pfam ID number PF142311, and / or

[0467] (c) a domain retaining at least 90% sequence identity to PF14232 or PF14231 ;

[0468] (iii) the polypeptide having BVMO enzyme activity is selected from the group of polypeptides comprising:

[0469] (a) a flavin-containing monooxygenase (FMO) protein family domain having the Pfam ID number PF00743 within their amino acid sequence or a domain retaining at least 90%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to PF00743, and / or

[0470] (b) at least one or more of motifs selected from:

[0471] . GxGxxG (SEQ ID NO: 239),

[0472] . [GS]GxWxxxxYPGxxxD (SEQ ID NO: 240),

[0473] . Gxxx[FY]xGxxx[HS]xxxW (SEQ ID NO: 241), and

[0474] . [KQ]x[VI]xx[IV]GxG (SEQ ID NO: 242); wherein, residues x represent independently of each other any natural amino acid residue; and / or,

[0475] (iv) the polypeptide having esterase enzyme activity comprises at least one or more motifs selected from AxVVxVxxRLAPE (SEQ ID NO: 243), GASAGGGLxA (SEQ ID NO: 244), VxQLLxYPMLDDR (SEQ ID NO: 245) and ARxxDLSGLPxT (SEQ ID NO: 246); wherein, residues x represent independently of each other any natural amino acid residue.

[0476] A futher embodiment of the recombinant cell of the invention is wherein:

[0477] (i) the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 11 to 21 ;

[0478] (ii) the polypeptide having enal-cleaving enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22;

[0479] (iii) the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 23 to 26 and 216 to 227; preferably to SEQ ID NOs: 23 to 26; and / or

[0480] (iv) the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 27 and 28.

[0481] A futher embodiment of the recombinant cell of the invention is wherein:

[0482] (i) the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 11 or 21 ; preferably, the ADH enzyme has the sequence of SEQ ID NO: 11 or 21 ; (ii) the polypeptide having enal-cleaving enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22; preferably, the enal-cleaving enzyme has the sequence of SEQ ID NO: 22;

[0483] (iii) the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 25 or 26; preferably, the BVMO enzyme has the sequence of SEQ ID NO: 25 or 26; and / or,

[0484] (iv) the polypeptide having esterase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO 28; preferably, the esterase enzyme has the sequence of SEQ ID NO: 28.

[0485] A futher embodiment of the recombinant cell of the invention is wherein the polypeptide having terpene cyclase enzyme activity as described herein above has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82.

[0486] A further embodiment of the recombinant cell of the invention is wherein:

[0487] (i) the compound of formula (I) is in the form of formula (la): (formula la)

[0488] (ii) the compound of formula (II) is in the form of formula (Ila):

[0489] (formula Ila);

[0490] (iii) the compound of formula (III) is in the form of formula (Illa): (formula Illa);

[0491] (iv) the compound of formula (IV) is in the form of formula (IVa):

[0492] (v) the compound of formula (V) is in the form of formula (Va): (formula Va); (vi) the compound of formula (VI) is in the form of formula (Via):

[0493] (formula Via).

[0494] The recombinant cell may be any such cell suitable for the production of a compound of formula (I).

[0495] A list of suitable cells for the production of a compound of formula (I) is provided above in relation to the process of the invention and are also cells for this aspect of the invention.

[0496] Preferably, the cell is a bacterium or a fungal cell, in particular a yeast. Preferably, the cell is a unicellular organism, a cultured cell derived from a multi-cellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism. Preferably, the cell is a bacterial cell of the genus Escherichia, preferably E. coli, or a yeast cell of the genus Saccharomyces, preferably S. cerevisiae, of the genus Yarrowia, preferably Y. lipolytica, or of the genus Pichia, preferably P. pastoris.

[0497] Methods of introducing recombinant nucleic acid sequences into such host cells are well known in the art and constitute routine laboratory methodologies which do not need to be further described herein.

[0498] A further embodiment of this aspect of the invention is wherein the recombinant cell further comprises, is further capable of functionally expressing or further expresses one or more prenyltransferase enzyme(s). Preferably, the recombinant cell further comprises, is further capable of functionally expressing or further expresses one or more enzymes having phosphatase activity.

[0499] As mentioned above in relation to the process of the invention, the recombinant cell of the invention may further comprise, may further be capable of functionally expressing or may further express enzymes to provide a compound of formula (II). Such a process requires the presence of one or more prenyltransferase enzyme(s) and one or more enzymes having phosphatase activity.

[0500] Preferably, the prenyltransferase is a GGPP synthase. Preferably, the GGPP synthase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 1 or 2. Said prenyltransferase may be a heterologous prenyltransferase.

[0501] Preferably, the phosphatase is a GGPP phosphatase. Preferably, the GGPP phosphatase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 3 to 10. Said phosphatase may be a heterologous phosphatase.

[0502] Further embodiments of the invention are wherein the recombinant cell of the invention comprises enzymes for the IPP and DMAPP. As mentioned above in relation to the process of the invention, the recombinant cell of the invention may further comprise, may further be capable of functionally expressing or may further express enzymes to provide a compound of formula (II) via the “mevalonate pathway”, methylerythritol phosphate (MEP) pathway or alternative pathways to the preparation of IPP and DMAPP.

[0503] In one embodiment of the invention, the recombinant cell comprises, is capable of functionally expressing or expresses enzymes of the mevalonate pathway:

[0504] . Acetyl-CoA acetyltransferase (ACAT),

[0505] . 3-hydroxy-3-methylglutaryl-CoA synthase (HMG-CoA synthase),

[0506] . 3-hydroxy-3-methylglutaryl-CoA reductase (HMG-CoA reductase),

[0507] . Mevalonate kinase,

[0508] . Phosphomevalonate kinase,

[0509] . Mevalonate diphosphate decarboxylase,

[0510] . Isopentenyl diphosphate isomerase,

[0511] . Dimethylallyl diphosphate synthase.

[0512] In a further embodiment, one or more of said enzymes may be a heterologous enzyme.

[0513] In one embodiment of the invention, the recombinant cell comprises, is capable of functionally expressing or expresses enzymes of the MEP pathway:

[0514] . 1-Deoxy-D-xylulose 5-phosphate synthase (DXS),

[0515] . 1-Deoxy-D-xylulose 5-phosphate reductoisomerase (DXR),

[0516] . 2-C-methyl-D-erythritol 4-phosphate cytidylyltransferase (MCT, IspD),

[0517] . 4-diphosphocytidyl-2-C-methyl-D-erythritol kinase (CMK, IspE),

[0518] . 2-C-methyl-D-erythritol 2,4-cyclodiphosphate synthase (MDS, IspF),

[0519] . 4-hydroxy-3-methylbut-2-en-1-yl diphosphate synthase (HDS, IDS),

[0520] . 4-hydroxy-3-methylbut-2-en-1-yl diphosphate reductase (HDR).

[0521] In a further embodiment, one or more of said enzymes may be a heterologous enzyme.

[0522] Said recombinant cells of the invention may be used in an in vivo process or a bioconversion process for the preparation of a compound of formula (I).

[0523] Accordingly, further embodiments of the process for preparing a compound of formula (I) according to the first, second, third and further aspects of the invention is a process further comprising growing a recombinant cell of the invention as described in any of the above embodiments under growth conditions suitable for the production of the compound of formula (I).

[0524] For example, in one embodiment, the process for the preparation of a compound of formula (I) may comprise growing the recombinant cell comprising, capable of functionally expressing or expressing the polypeptide having terpene cyclase enzyme activity to produce the compound of formula (I). For example, in the first, third and further aspects of the process of the invention for preparing a compound of formula (I), the polypeptide having terpene cyclase enzyme activity may have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323. For example, in the second aspect of the process of the invention for preparing a compound of formula (I), the polypeptide having terpene cyclase enzyme activity may have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 322 and 323. A further embodiment is wherein the polypeptide having terpene cyclase enzyme activity has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. In a further embodiment, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme.

[0525] In another embodiment, the process for the preparation of a compound of formula (I) may comprise growing the recombinant cell comprising, capable of functionally expressing or expressing the polypeptide having esterase enzyme activity and the polypeptide having terpene cyclase enzyme activity to produce the compound of formula (I). In said embodiment, the polypeptide having terpene cyclase enzyme activity may have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323. A further embodiment is wherein the polypeptide having terpene cyclase enzyme activity has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Embodiments for the polypeptide having esterase enzyme activity are described herein above for the recombinant cells and are incorporated herein to the present aspects of the process for the preparation of a compound of formula (I). In a further embodiment, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In a further embodiment, the polypeptide having esterase enzyme activity is a heterologous polypeptide.

[0526] In another embodiment, the process for the preparation of a compound of formula (I) may be an in vivo process comprising growing the recombinant cell comprising, capable of producing or producing a compound of formula (I), wherein the recombinant cell comprises, is capable of functionally expressing or is functionally expressing:

[0527] . a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323;

[0528] . a polypeptide having esterase enzyme activity;

[0529] . a polypeptide having BVMO enzyme activity;

[0530] . a polypeptide having enal-cleaving enzyme activity;

[0531] . a polypeptide having ADH enzyme activity; and optionally, one or more of:

[0532] . a polypeptide having prenyltransferase enzyme activity; and,

[0533] . a polypeptide having having phosphatase enzyme activity. In said embodiment, the recombinant cell may further comprise, may be capable of functionally expressing or may be functionally expressing one or more polypeptides of a pathway for the preparation of IPP and DMAPP; for example, one or more polypeptides of the “mevalonate pathway” or one or more polypeptides of the MEP pathway as detailed above. Embodiments for each of the polypeptides (i.e. enzymes) are described herein above for the recombinant cells and are incorporated herein to the present aspects of the process for the preparation of a compound of formula (I). A further embodiment is wherein the polypeptide having terpene cyclase enzyme activity has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. In a further embodiment, one or more of said polypeptides may be a heterologous polypeptide. In particular, the polypeptide having terpene cyclase enzyme activity is a heterologous terpene cyclase enzyme; more in particular, a heterologous squalene cyclase enzyme. In yet a further embodiment, the polypeptides having ADH enzyme activity, enal-cleaving enzyme activity, BVMO enzyme activity, esterase enzyme activity and terpene cyclase enzyme activity are heterologous polypeptides.

[0534] Reaction conditions for the preparation of a compound of formula (I)

[0535] The first, second, third and further aspects for the preparation of a compound of formula (I) may be an in vivo process or a bioconversion process.

[0536] The term in vivo process (or whole-cell production, or in-vivo production, or in-vivo biosynthesis) refers to a process of using a metabolically active cell where the primary metabolism is active to produce the precursors for the processes of the invention (preferably a microbial cell) to convert a carbon source to a new compound, such as the conversion of a carbon source to a terpene or terpene-derived compound.

[0537] Preferred sources of carbon are sugars, such as mono-, di- or polysaccharides. Very good sources of carbon are for example glucose, fructose, mannose, galactose, ribose, sorbose, ribulose, lactose, maltose, sucrose, raffinose, starch or cellulose. Sugars can also be added to the media via complex compounds, such as molasses, or other by-products from sugar refining. It may also be advantageous to add mixtures of various sources of carbon. Other possible sources of carbon are oils and fats such as soybean oil, sunflower oil, peanut oil and coconut oil, fatty acids such as palmitic acid, stearic acid or linoleic acid, alcohols such as glycerol, methanol or ethanol and organic acids such as acetic acid or lactic acid.

[0538] The cells thus contain all enzymes of one or more biosynthetic pathways. At least some of the enzymes involved in the process are part of the cell's primary metabolism. For example, the cells may contain the enzymes of a pathway to convert a carbon source (e.g., glucose, glycerol, isoprenol, prenol, CO2) to terpenoid precursors (e.g., IPP, DMAPP, FPP, GGPP) and a pathway converting the terpene precursor to a terpene or terpene derived molecule - such as a compound of formula (I) - as described herein above in the aspects of the process of the invention for the preparation of a compound of formula (i). The enzymes may be present naturally in the cell or the cells can be transformed to produce the enzymes. Alternatively, the processes of the present invention may be performed under bioconversion, also known as biotransformation conditions. Bioconversion processes refer to processes of conversion of compounds to different products using a biological process or agent such enzymes or whole cells (preferably a microbial cell). Bioconversion does not include the use of a cell's primary metabolism (as defined above) to produce the precursors forthe processes of the invention. A bioconversion process can comprise multistep reactions each performed by a different enzyme. The compounds used in bioconversion process can be extracted from a natural source or produced using a separate chemical or a biochemical process.

[0539] The at least one polypeptide / enzyme which is present during the bioconversion method of the invention or an individual step of the multistep method as defined herein above, can be present in living cells naturally or recombinantly producing the enzyme or enzymes, in harvested cells, dead cells, in permeabilized cells, in crude cell extracts, in purified extracts, or in essentially pure or completely pure form, i.e. under bioconversion conditions. Such extracts may comprise membrane fraction or a liquid fraction prepared from the recombinant host cell that expresses at least one polypeptide / enzyme. The cells may be immobilized on a suitable substrate as is known in the art. At least one polypeptide / enzyme may be present in solution or as an enzyme immobilized on a carrier. One or several enzymes may simultaneously be present in soluble and / or immobilized forms.

[0540] It can be understood by the skilled person that there may be advantages for the use of an in vivo process.

[0541] In particular, a bioconversion process involves multiple steps, typically:

[0542] Preparation or isolation of the starting compound to be transformed. The compound can be prepared using a chemical or biochemical process or by extraction from a natural source.

[0543] Production of the enzymes or (living) cells used for the bioconversion.

[0544] Biotransformation reaction by contacting the compound with the enzymes or (living) cells. Product Recovery and Refinement.

[0545] In comparison, an in vivo process requires a limited number of steps, generally limited to:

[0546] The cultivation of the microorganism under conditions suitable for the production of the desired compound.

[0547] Harvesting the cells or growing medium and purification of the desired compound.

[0548] In a bioconversion process such as the bioconversion of a compound of formula (VI), the addition of a detergent is often required to facilitate the solubilization of the compound or to maximize the contact with the biocatalyst. In an in vivo process, the reactants and enzymes are produced in the cells and the addition of a detergent is not needed.

[0549] Therefore, the in vivo process is usually more efficient and cost-effective than a bioconversion process. Laboratory methods that can be used in in vivo and bioconversion processes of the invention are well known in the art. There follows a discussion on some of the methods that can be used.

[0550] The bioconversion processes according to the invention can be performed in common reactors, which are known to those skilled in the art, and in different ranges of scale, e.g. from a laboratory scale (few milliliters to dozens of liters of reaction volume) to an industrial scale (several liters to thousands of cubic meters of reaction volume). If the polypeptide is used in a form encapsulated by non-living, optionally permeabilized cells, in the form of a more or less purified cell extract or in purified form, a chemical reactor can be used. The chemical reactor usually allows controlling the amount of at least one enzyme, the amount of at least one substrate, the pH, the temperature and the circulation of the reaction medium.

[0551] Where the process of the invention is in v / vo then it is preferred that the reaction is performed in a fermenter, where parameters necessary for suitable living conditions for the living cells (e.g. culture medium with nutrients, temperature, aeration, presence or absence of oxygen or other gases, antibiotics, and the like) can be controlled.

[0552] The term "fermentative production" or "fermentation" refers to the ability of a microorganism (assisted by enzyme activity contained in or generated by said microorganism) to produce a chemical compound in cell culture utilizing at least one carbon source added to the incubation.

[0553] The term "fermentation broth" or "fermentation medium" is understood to mean a liquid, particularly aqueous or aqueous / organic solution which is based on a fermentative process and has not been worked up or has been worked up, for example, as described herein.

[0554] Those skilled in the art are familiar with chemical reactors or bioreactors, e.g. with procedures for up-scaling chemical or biotechnological methods from laboratory scale to industrial scale, or for optimizing process parameters, which are also extensively described in the literature (for biotechnological methods see e.g. Crueger und Crueger, Biotechnologie - Lehrbuch der angewandten Mikrobiologie, 2. Ed., R. Oldenbourg Verlag, Munchen, Wien, 1984).

[0555] The culture medium that is to be used must satisfy the requirements of the particular strains in an appropriate manner. Descriptions of culture media for various microorganisms are given in the handbook "Manual of Methods for General Bacteriology" of the American Society for Bacteriology (Washington D. C., USA, 1981).

[0556] These media that can be used according to the invention may comprise one or more sources of carbon, sources of nitrogen, inorganic salts, vitamins and / or trace elements.

[0557] Sources of nitrogen are usually organic or inorganic nitrogen compounds or materials containing these compounds. Examples of sources of nitrogen include ammonia gas or ammonium salts, such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate or ammonium nitrate, nitrates, urea, amino acids or complex sources of nitrogen, such as corn-steep liquor, soybean flour, soy-bean protein, yeast extract, meat extract and others. The sources of nitrogen can be used separately or as a mixture.

[0558] Inorganic salt compounds that may be present in the media comprise the chloride, phosphate or sulfate salts of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper and iron.

[0559] Inorganic sulfur-containing compounds, for example sulfates, sulfites, di-thionites, tetrathionates, thiosulfates, sulfides, but also organic sulfur compounds, such as mercaptans and thiols, can be used as sources of sulfur.

[0560] Phosphoric acid, potassium dihydrogenphosphate or dipotassium hydrogenphosphate or the corresponding sodium-containing salts can be used as sources of phosphorus.

[0561] Chelating agents can be added to the medium, in order to keep the metal ions in solution. Especially suitable chelating agents comprise dihydroxyphenols, such as catechol or protocatechuate, or organic acids, such as citric acid.

[0562] The fermentation media used according to the invention may also contain other growth factors, such as vitamins or growth promoters, which include for example biotin, riboflavin, thiamine, folic acid, nicotinic acid, pantothenate and pyridoxine. Growth factors and salts often come from complex components of the media, such as yeast extract, molasses, corn-steep liquor and the like. In addition, suitable precursors can be added to the culture medium. The precise composition of the compounds in the medium is strongly dependent on the particular experiment and must be decided individually for each specific case. Information on media optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (1997) Growing media can also be obtained from commercial suppliers, such as Standard 1 (Merck) or BHI (Brain heart infusion, DIFCO) etc.

[0563] All components of the medium are sterilized, either by heating (20 min at 1 .5 bar and 121 °C) or by sterile filtration. The components can be sterilized either together or if necessary, separately. All the components of the medium can be present at the start of growing, or optionally can be added continuously or by batch feed.

[0564] The temperature of the culture is normally between 15°C and 45°C, preferably 25°C to 40°C and can be kept constant or can be varied during the experiment. The pH value of the medium should be in the range from 5 to 8.5, preferably around 7.0. The pH value for growing can be controlled during growing by adding basic compounds such as sodium hydroxide, potassium hydroxide, ammonia or ammonia water or acid compounds such as phosphoric acid or sulfuric acid. Antifoaming agents, e.g. fatty acid polyglycol esters, can be used for controlling foaming. To maintain the stability of plasmids, suitable substances with selective action, e.g. antibiotics, can be added to the medium. Oxygen or oxygen-containing gas mixtures, e.g. the ambient air, are fed into the culture in order to maintain aerobic conditions. The temperature of the culture is normally from 20°C to 45°C. Culture is continued until a maximum of the desired product has formed. This is normally achieved within 1 hour to 160 hours.

[0565] Where the process of the invention is a bioconversion, cells containing the at least one enzyme can be permeabilized by physical or mechanical means, such as ultrasound or radiofrequency pulses, French presses, or chemical means, such as hypotonic media, lytic enzymes and detergents present in the medium, or combination of such methods. Examples for detergents are SDS, digitonin, n-dodecylmaltoside, octylglycoside, Triton® X-100, Tween ® 20, deoxycholate, CHAPS (3-[(3-

[0566] Cholamidopropyl)dimethylammonio]-1-propansulfonate), Nonidet ® P40

[0567] (Ethylphenolpoly(ethyleneglycolether), and the like. As stated above, where the process of the invention is an in vivo process, then a detergent is not required for the reasons stated herein.

[0568] The conversion reaction can be carried out batch wise, semi-batch wise or continuously. Reactants (and optionally nutrients) can be supplied at the start of reaction or can be supplied subsequently, either semi- continuously or continuously.

[0569] The bioconversion reaction of the invention, depending on the particular reaction type, may be performed in an aqueous, aqueous-organic or non-aqueous reaction medium.

[0570] An aqueous or aqueous-organic medium may contain a suitable buffer in order to adjust the pH to a value in the range of 5 to 1 1 , like 6 to 10.

[0571] In an aqueous-organic medium an organic solvent miscible, partly miscible or immiscible with water may be applied. Non-limiting examples of suitable organic solvents are listed below. Further examples are mono- or polyhydric, aromatic or aliphatic alcohols, in particular polyhydric aliphatic alcohols like glycerol.

[0572] The non-aqueous medium may contain is substantially free of water, i.e. will contain less than about 1 wt.- % or 0.5 wt.-% of water.

[0573] Bioconversion methods may also be performed in an organic non-aqueous medium. As suitable organic solvents there may be mentioned aliphatic hydrocarbons having for example 5 to 8 carbon atoms, like pentane, cyclopentane, hexane, cyclohexane, heptane, octane or cyclooctane; aromatic carbohydrates, like benzene, toluene, xylenes, chlorobenzene or dichlorobenzene, aliphatic acyclic and ethers, like diethylether, methyl-tert.-butylether, ethyl-tert.-butylether, dipropylether, diisopropylether, dibutylether; or mixtures thereof. The concentration of the reactants / substrates may be adapted to the optimum bioconversion reaction conditions, which may depend on the specific enzyme applied. For example, the initial substrate concentration may be in the 0,1 to 0,5 M, as for example 10 to 100 mM.

[0574] The bioconversion reaction temperature may be adapted to the optimum reaction conditions, which may depend on the specific enzyme applied. For example, the reaction may be performed at a temperature in a range of from 0 to 70°C, as for example 20 to 50 or 25 to 40°C. Examples for reaction temperatures are about 30°C, about 35°C, about 37°C, about 40°C, about 45°C, about 50°C, about 55°C and about 60°C.

[0575] The bioconversion may proceed until equilibrium between the substrate and then product(s) is achieved, but may be stopped earlier. Usual process times are in the range from 1 minute to 25 hours, in particular 10 min to 6 hours, as for example in the range from 1 hour to 4 hours, in particular 1.5 hours to 3.5 hours. These parameters are non-limiting examples of suitable process conditions.

[0576] Advantageously, microorganisms such as bacteria, fungi or yeasts are used as host organisms. Advantageously, gram-positive or gram-negative bacteria are used, preferably bacteria of the families Enterobacteriaceae, Pseudomonadaceae, Rhizobiaceae, Streptomycetaceae, Streptococcaceae or Nocardiaceae, especially preferably bacteria of the genera Escherichia, Pseudomonas, Streptomyces, Lactococcus, Nocardia, Burkholderia, Salmonella, Agrobacterium, Clostridium or Rhodococcus. The genus and species Escherichia coll is quite especially preferred. Furthermore, other advantageous bacteria are to be found in the group of alpha-Proteobacteria, beta-Proteobacteria or gamma-Proteobacteria. Advantageously, also yeasts of families like Saccharomyces or Pichia are suitable hosts.

[0577] Preferably, the cell is a bacterium or a fungal cell, in particular a yeast cell. Preferably, the cell is a unicellular organism, a cultured cell derived from a multi-cellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism. Preferably, the cell is a bacterial cell of the genus Escherichia, preferably E. coll, or a yeast cell of the genus Saccharomyces, preferably S. cerevisiae, of the genus Yarrowia, preferably Y. lipolytica, or of the genus Pichia, preferably P. pastoris.

[0578] Alternatively, entire plants or plant cells may serve as natural or recombinant host. As non-limiting examples, the following plants or cells derived therefrom may be mentioned: the genera Nicotiana, in particular Nicotiana benthamiana and Nicotiana tabacum (tobacco); as well as Arabidopsis, in particular Arabidopsis thaliana.

[0579] Cell culture fermentation media of the invention

[0580] The present invention also provides a cell culture fermentation medium comprising compound of formula (I), wherein more than 97% of the compound of formula (I) is in the form of formula (la) and / or (lb). A further aspect of the invention provides a cell culture fermentation medium comprising the recombinant cell as described herein above. The cell culture fermentation medium may further comprise a compound of formula (I) and one or more compound(s) selected from compounds of formula (II), formula (III), formula

[0581] (IV), formula (V) and formula (VI). Embodiments for the recombinant cell as described herein above are incorporated herein to the present aspect of invention providing a cell culture fermentation medium.

[0582] As discussed herein, the present inventors were able to innovate a biosynthetic pathway to the compound of formula (I) in recombinant cells. Such cells are then grown under conditions suitable for the production of said compound using cell culture fermentation media as appropriate for specific cell types.

[0583] The cell culture fermentation media can be a nutrient rich broth for the growth and maintenance of the cells during the production phase. Yeast culture conditions for maintaining and propagating various strains can require specific formulations of complex media for use in cloning and protein expression, and can be appreciated by those of skill in the art. Commercially available culture media can be used from ThermoFisher for example. The media can be YPD broth or can have a yeast nitrogen base. Yeast can be grown in YPD or synthetic media at 30°C.

[0584] Lysogeny broth (LB) is typically used for bacterial cells. The bacterial cells can have antibiotic resistance to prevent the growth of other cells in the culture media and contamination. The cells can have an antibiotic gene cassettes for resistance to antibiotics such as chloramphenicol, penicillin, kanamycin and ampicillin, for example.

[0585] Reaction mixture comprising compounds of the invention

[0586] The process of the invention may be a bioconversion process.

[0587] The present invention also provides a reaction mixture comprising compound of formula (I), wherein more than 97% of the compound of formula (I) is in the form of formula (la) and / or (lb). The reaction mixture may further comprise one or more compounds of formula (II), formula (III), formula (IV), formula (V), and / or formula (VI).

[0588] Additional components of the reaction mixture may include detergents, co-factors, cells, cell-debris, cell culture media, and other such components well known to the person skilled in the art.

[0589] In one embodiment, the invention provides a reaction mixture comprising a compound of formula (I) and one or more compound(s) selected from compounds of formula (II), formula (III), formula (IV) and formula

[0590] (V), wherein the reaction mixture further comprises a polypeptide having terpene cyclase enzyme activity having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323. A further embodiment is wherein the polypeptide having terpene cyclase enzyme activity has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82.

[0591] In another embodiment, the invention provides a reaction mixture comprising a compound of formula (I) and a polypeptide having terpene cyclase enzyme activity having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323. A further embodiment is wherein the polypeptide having terpene cyclase enzyme activity has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82.

[0592] Product isolation

[0593] The methodology of the present invention can further include a step of recovering an end product or an intermediate product, optionally in stereoisomerically or enantiomerically substantially pure form. The term “recovering” includes extracting, harvesting, isolating or purifying the compound from culture or reaction media. Recovering the compound can be performed according to any conventional isolation or purification methodology known in the art including, but not limited to, treatment with a conventional resin (e.g., anion or cation exchange resin, non-ionic adsorption resin, etc.), treatment with a conventional adsorbent (e.g., activated charcoal, silicic acid, silica gel, cellulose, alumina, etc.), alteration of pH, solvent extraction (e.g., with a conventional solvent such as an alcohol, ethyl acetate, hexane and the like), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, lyophilization and the like.

[0594] Identity and purity of the isolated product may be determined by known techniques, like High Performance Liquid Chromatography (HPLC), gas chromatography (GC), Spektroskopy (like IR, UV, NMR), Colouring methods, TLC, NIRS, enzymatic or microbial assays. (see for example: Patek et al. (1994) Appl. Environ. Microbiol. 60:133-140; Malakhova et al. (1996) Biotekhnologiya 11 27-32; und Schmidt et al. (1998) Bioprocess Engineer. 19:67-70. Ullmann's Encyclopedia of Industrial Chemistry (1996) Bd. A27, VCH: Weinheim, S. 89-90, S. 521-540, S. 540-547, S. 559-566, 575-581 und S. 581-587; Michal, G (1999) Biochemical Pathways: An Atlas of Biochemistry and Molecular Biology, John Wiley and Sons; Fallon, A. et al. (1987) Applications of HPLC in Biochemistry in: Laboratory Techniques in Biochemistry and Molecular Biology, Bd. 17.).

[0595] The compounds produced in any of the processes described herein can be converted to derivatives such as, but not limited to hydrocarbons, esters, amides, glycosides, ethers, epoxides, aldehydes, ketones, alcohols, diols, acetals or ketals. The terpene compound derivatives can be obtained by a chemical method such as, but not limited to oxidation, reduction, alkylation, acylation and / or rearrangement. Alternatively, the terpene compound derivatives can be obtained using a biochemical method by contacting the terpene compound with an enzyme such as, but not limited to an oxido reductase, a monooxygenase, a dioxygenase, a transferase or a terpene cyclase. The biochemical conversion can be performed in vitro using isolated enzymes, enzymes from lysed cells or bioconversion using whole cells. Phosphatase enzymes for use in the process of the invention

[0596] As described herein above, one embodiment of the process of the first aspect of the invention further comprises one or more biocatalytic steps prior to step (i) to prepare a compound of formula (II), said biocatalytic step(s) comprising:

[0597] (a) preparing geranylgeranyl-diphosphate (GGPP) from IPP and DMAPP using one or more prenyltransferase enzyme(s); and,

[0598] (b) preparing a compound of formula (II) from GGPP using one or more enzymes having phosphatase activity.

[0599] When preparing the process of the invention, the present inventors identified phosphatase enzymes which can be used in step (b) of said process. This is the first time these enzymes were shown to be capable of catalysing the step of preparing a compound of formula (II) from GGPP.

[0600] Hence, included in this aspect of the invention is the use of a phosphatase for the preparation of a compound of formula (II).

[0601] A further aspect of the invention provides the use of a phosphatase enzyme comprising 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in any of SEQ ID NOs: 3 to 10 for the preparation of a compound of formula (II) from GGPP.

[0602] Preferably, the phosphatase enzyme comprises 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in any of SEQ ID NOs: 3, 4, 5, 7 and 8. More preferably, the phosphatase enzyme comprises the sequences provided in any of SEQ ID NOs: 3, 4, 5, 7 and 8.

[0603] ADH enzymes for use in the process of the invention

[0604] As described herein, the process of the first aspect of the invention comprises as step (i) the preparation of a compound of formula (III) by contacting a compound of formula (II) with an ADH enzyme.

[0605] When preparing the process of the invention, the present inventors identified ADH enzymes which can be used in step (i) of said process. This is the first time these enzymes were shown to be capable of catalysing the step of preparing a compound of formula (III) from a compound of formula (II).

[0606] Hence, a further aspect of the invention provides the use of an ADH enzyme comprising 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in any of SEQ ID NOs: 11 to 21 for the preparation of compound of formula (III) from a compound of formula (II). Preferably, the ADH enzyme comprises the sequence provided in any of SEQ ID NOs: 11 to 21.

[0607] Yet, another further aspect of the invention provides the use of an ADH enzyme comprising 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 11 or 21 for the preparation of compound of formula (III) from a compound of formula (II). Preferably, the ADH enzyme comprises the sequence provided in SEQ ID NOs: 11 or 21 .

[0608] Enal-cleavinq enzymes for use in the process of the invention

[0609] As described herein, the process of the first aspect of the invention comprises as step (ii) the preparation of a compound of formula (IV) by contacting a compound of formula (III) with an enal-cleaving enzyme.

[0610] When preparing the process of the invention, the present inventors identified an enal-cleaving enzyme which can be used in step (ii) of said process. This is the first time this enzyme was shown to be capable of catalysing the step of preparing a compound of formula (IV) from a compound of formula (III).

[0611] Hence, a further aspect of the invention provides the use of a polypeptide having enal-cleaving enzyme activity to produce a compound of formula (IV). A further aspect of the invention provides the use of a polypeptide having enal-cleaving enzyme activity for the preparation of a compound of formula (IV) from a compound of formula (III). A further aspect of the invention is the use of a polypeptide having enal-cleaving enzyme activity to produce a compound of formula (IV), (V), (VI), (I) and / or a derivative thereof.

[0612] Hence, a further aspect of the invention provides the use of a polypeptide having enal-cleaving enzyme activity comprising 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22 to produce a compound of formula (IV). A further aspect of the invention provides the use of a polypeptide having enal-cleaving enzyme activity comprising 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22 for the preparation of a compound of formula (IV) from a compound of formula (III). A further aspect of the invention provides the use of a polypeptide having enal-cleaving enzyme activity comprising 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 22 to produce a compound of formula (IV), (V), (VI), (I) and / or a derivative thereof.

[0613] Preferably, the polypeptide having enal-cleaving enzyme activity comprises the sequence provided in SEQ ID NO: 22.

[0614] BVMO enzymes for use in the process of the invention

[0615] As described herein, the process of the first aspect of the invention comprises as step (iii) the preparation of a compound of formula (V) by contacting a compound of formula (IV) with a BVMO enzyme. When preparing the process of the invention, the present inventors identified BMVO enzymes which can be used in step (iii) of said process. This is the first time these enzymes were shown to be capable of catalysing the step of preparing a compound of formula (V) from a compound of formula (IV).

[0616] Hence, a further aspect of the invention provides the use of a BVMO enzyme comprising 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 23 to 26 and 216 to 227 for the preparation of compound of formula (V) from a compound of formula (IV).

[0617] Preferably, the BVMO enzyme comprises 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 23 to 26. More preferably, the BVMO enzyme comprises the sequence provided in any of SEQ ID NOs: 23 to 26.

[0618] Preferably, the BVMO enzyme comprises 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 25 or 26. More preferably, the BVMO enzyme comprises the sequence provided in SEQ ID NO: 25 or 26.

[0619] Esterase enzymes for use in the process of the invention

[0620] As described herein, the process of the first aspect of the invention comprises as step (iv) the preparation of a compound of formula (VI) by contacting a compound of formula (V) with an esterase enzyme.

[0621] When preparing the process of the invention, the present inventors identified esterase enzymes which can be used in step (iv) of said process. This is the first time these enzymes were shown to be capable of catalysing the step of preparing a compound of formula (VI) from a compound of formula (V).

[0622] Hence a further aspect of the invention provides the use of an esterase enzyme comprising 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to SEQ ID NO: 27 or 28 for the preparation of compound of formula (VI) from a compound of formula (V).

[0623] Preferably, the esterase enzyme comprises the sequence provided in SEQ ID NO: 27 or 28.

[0624] Terpene cyclase enzymes according to the invention

[0625] When preparing the process of the invention, the present inventors identified new polypeptide sequences having terpene cyclase enzyme activity which are squalene cyclase (SHC) enzymes and can be used in said process. These polypeptides are therefore also part of the present invention. When investigating said enzymes, the present inventors identified enzymes having amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to sequence provided in SEQ ID NO: 82 as being of particular utility.

[0626] Accordingly therefore, a further aspect of the present invention is a mutant terpene cyclase enzyme having an amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to sequence provided in SEQ ID NO: 82. Said mutant terpene cyclase enzyme is a SHC enzyme.

[0627] Terpene cyclase enzymes, in particular SHC enzymes, have been described above; for example, in relation to step (v) of the process of the first aspect of the process of the invention. Using the information provided in that section of the present specification, the skilled person can identify any SHC enzyme which can be modified using standard laboratory techniques to arrive at the mutant SHC enzyme of this aspect of the invention.

[0628] A preferred embodiment of this aspect of the invention is wherein the mutant terpene cyclase enzyme is a polypeptide having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323, wherein said polypeptide has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. In a further preferred embodiment of this aspect of the invention is wherein the mutant terpene cyclase enzyme is a polypeptide having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 316, 318, 320 and 323, wherein said polypeptide has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, the mutant terpene cyclase enzyme is a polypeptide comprising the amino acid sequence of any of SEQ ID NOs: 316, 318, 320 and 323. Also included in this aspect of the invention is polypeptide fragments, variants and functional equivalents thereof. In said embodiment, said mutant terpene cyclase enzyme is a mutant SHC enzyme

[0629] Another preferred embodiment of this aspect of the invention is wherein the mutant terpene cyclase enzyme is a polypeptide having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323, wherein said polypeptide has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82. Alternatively, the the mutant terpene cyclase enzyme is a polypeptide comprising the amino acid sequence of any of SEQ ID NOs: 323. Also included in this aspect of the invention is polypeptide fragments, variants and functional equivalents thereof. Said mutant terpene cyclase enzyme is a mutant SHC enzyme.

[0630] The polypeptides of this aspect of the invention are SHC enzymes which may be used in step (v) of the process of the first aspect of the invention or any futher aspects of the invention comprising prepraring a compound of formula (I) from a compound of formula (VI). A further aspect of the invention provides a nucleic acid sequence encoding a mutant SHC enzyme having an amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to sequence provided in SEQ ID NO: 82. Methods of preparing such nucleic acid sequence are known in the art.

[0631] A preferred embodiment of this aspect of the invention is wherein the nucleic acid sequence encoding the mutant terpene cyclase enzyme is a nucleic acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 326 to 334, wherein said nucleic acid sequence encodes a polypeptide having amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to sequence provided in SEQ ID NO: 82. Preferably, the nucleic acid sequence encoding the mutant SHC enzyme is a nucleic acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 327, 329, 331 and 334, wherein said nucleic acid sequence encodes a polypeptide having amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to sequence provided in SEQ ID NO: 82. Preferably, the nucleic acid sequence is a nucleic sequence provided in any of SEQ ID NOs: 327, 329, 331 and 334. Also included in this aspect of the invention is expression vectors, cassettes and other such related technology comprising nucleic acid sequences of the invention.

[0632] A preferred embodiment of this aspect of the invention is wherein the nucleic acid sequence encoding the mutant SHC enzyme is a nucleic acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 333 to 334, wherein said nucleic acid sequence encodes a polypeptide having amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to sequence provided in SEQ ID NO: 82. Preferably, the nucleic acid sequence is a nucleic sequence provided in SEQ ID NO: 334. Also included in this aspect of the invention is expression vectors, cassettes and other such related technology comprising nucleic acid sequences of the invention.

[0633] Furthermore, as described herein, in the process of the invention comprising the preparation of a compound of formula (I) by contacting a compound of formula (VI) with a terpene cyclase enzyme, the present inventors identified terpene cyclase enzymes, in particular SHC enzymes, which can be used in said process.

[0634] Hence, a further aspect of the invention provides the use of a terpene cyclase enzyme, in particular a SHC enzyme, comprising 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 315 to 323, preferably to any of SEQ IS NOs: 322 and 323, for the preparation of compound of formula (I) from a compound of formula (VI). Even more preferably, said enzyme further comprises amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to sequence provided in SEQ ID NO: 82. Polypeptides and nucleic acids of the invention or used in the process of the invention

[0635] The generic terms “polypeptide” or “peptide”, which may be used interchangeably, refer to a natural or synthetic linear chain or sequence of consecutive, peptidically linked amino acid residues, comprising about 10 to up to more than 1 .000 residues. Short chain polypeptides with up to 30 residues are also designated as “oligopeptides”.

[0636] The term “protein” refers to a macromolecular structure consisting of one or more polypeptides. The amino acid sequence of its polypeptide(s) represents the “primary structure” of the protein. The amino acid sequence also predetermines the “secondary structure” of the protein by the formation of special structural elements, such as alpha-helical and beta-sheet structures formed within a polypeptide chain. The arrangement of a plurality of such secondary structural elements defines the “tertiary structure” or spatial arrangement of the protein. If a protein comprises more than one polypeptide chains said chains are spatially arranged forming the “quaternary structure” of the protein. A correct spatial arrangement or “folding” of the protein is prerequisite of protein function. Denaturation or unfolding destroys protein function. If such destruction is reversible, protein function may be restored by refolding.

[0637] A typical protein function referred to herein is an “enzyme function”, i.e. the protein acts as biocatalyst on a substrate, for example a chemical compound, and catalyzes the conversion of said substrate to a product. An enzyme may show a high or low degree of substrate and / or product specificity.

[0638] A “polypeptide” referred to herein as having a particular “activity” thus implicitly refers to a correctly folded protein showing the indicated activity, as for example a specific enzyme activity.

[0639] Thus, unless otherwise indicated the term “polypeptide” also encompasses the terms “protein” and “enzyme”.

[0640] Similarly, the term “polypeptide fragment” encompasses the terms “protein fragment" and “enzyme fragment”.

[0641] The term “isolated polypeptide” refers to an amino acid sequence that is removed from its natural environment by any method or combination of methods known in the art and includes recombinant, biochemical and synthetic methods.

[0642] “Target peptide” refers to an amino acid sequence which targets a protein, or polypeptide to intracellular organelles, i.e., mitochondria, or plastids, or to the extracellular space (secretion signal peptide). A nucleic acid sequence encoding a target peptide may be fused to the nucleic acid sequence encoding the amino terminal end, e.g., N-terminal end, of the protein or polypeptide, or may be used to replace a native targeting polypeptide. The present invention also relates to "functional equivalents" (also designated as “analogs” or “functional mutations”) of the polypeptides specifically described herein.

[0643] For example, "functional equivalents" refer to polypeptides which, in a test used for determining enzymatic activity display at least a 1 to 10 %, or at least 20 %, or at least 50 %, or at least 75 %, or at least 90 % higher or lower activity, as that of the polypeptides specifically described herein.

[0644] "Functional equivalents”, according to the invention, also cover particular mutants, which, in at least one sequence position of an amino acid sequences stated herein, have an amino acid that is different from that concretely stated one, but nevertheless possess one of the aforementioned biological activities, as for example enzyme activity. "Functional equivalents" thus comprise mutants obtainable by one or more, like 1 to 20, in particular 1 to 15 or 5 to 10 amino acid additions, substitutions, in particular conservative substitutions, deletions and / or inversions, where the stated changes can occur in any sequence position, provided they lead to a mutant with the profile of properties according to the invention. Functional equivalence is in particular also provided if the activity patterns coincide qualitatively between the mutant and the unchanged polypeptide, i.e. if, for example, interaction with the same agonist or antagonist or substrate, however at a different rate, (i.e. expressed by a EC50 or IC50 value or any other parameter suitable in the present technical field) is observed. Examples of suitable (conservative) amino acid substitutions are shown in the following table: "Functional equivalents" in the above sense are also "precursors" of the polypeptides described herein, as well as "functional derivatives" and "salts" of the polypeptides.

[0645] "Precursors" are in that case natural or synthetic precursors of the polypeptides with or without the desired biological activity.

[0646] The expression "salts" means salts of carboxyl groups as well as salts of acid addition of amino groups of the protein molecules. Salts of carboxyl groups can be produced in a known way and comprise inorganic salts, for example sodium, calcium, ammonium, iron and zinc salts, and salts with organic bases, for example amines, such as triethanolamine, arginine, lysine, piperidine and the like. Salts of acid addition, for example salts with inorganic acids, such as hydrochloric acid or sulfuric acid and salts with organic acids, such as acetic acid and oxalic acid, are also covered by the invention.

[0647] "Functional derivatives" of polypeptides according to the invention can also be produced on functional amino acid side groups or at their N-terminal or C-terminal end using known techniques. Such derivatives comprise for example aliphatic esters of carboxylic acid groups, amides of carboxylic acid groups, obtainable by reaction with ammonia or with a primary or secondary amine; N-acyl derivatives of free amino groups, produced by reaction with acyl groups; or O-acyl derivatives of free hydroxyl groups, produced by reaction with acyl groups.

[0648] ’’Functional equivalents” naturally also comprise polypeptides that can be obtained from other organisms, as well as naturally occurring variants. For example, areas of homologous sequence regions can be established by sequence comparison, and equivalent polypeptides can be determined on the basis of the concrete parameters of the invention.

[0649] "Functional equivalents" also comprise “fragments”, like individual domains or sequence motifs, of the polypeptides according to the invention, or N- and or C-terminally truncated forms, which may or may not display the desired biological function. Preferably such “fragments” retain the desired biological function at least qualitatively.

[0650] "Functional equivalents" are, moreover, fusion proteins, which have one of the polypeptide sequences stated herein or functional equivalents derived there from and at least one further, functionally different, heterologous sequence in functional N-terminal or C-terminal association (i.e. without substantial mutual functional impairment of the fusion protein parts). Non-limiting examples of these heterologous sequences are e.g. signal peptides, histidine anchors or enzymes.

[0651] “Functional equivalents” which are also comprised in accordance with the invention are homologs to the specifically disclosed polypeptides. These have at least 60%, preferably at least 75%, in particular at least 80 or 85%, such as, for example, 90, 91 , 92, 93, 94, 95, 96, 97, 98 or 99%, homology (or identity) to one of the specifically disclosed amino acid sequences, calculated by the algorithm of Pearson and Lipman, Proc. Natl. Acad, Sci. (USA) 85(8), 1988, 2444-2448. A homology or identity, expressed as a percentage, of a homologous polypeptide according to the invention means in particular an identity, expressed as a percentage, of the amino acid residues based on the total length of one of the amino acid sequences described specifically herein.

[0652] The identity data, expressed as a percentage, may also be determined with the aid of BLAST alignments, algorithm blastp (protein-protein BLAST), or by applying the Clustal settings specified herein below.

[0653] In the case of a possible protein glycosylation, "functional equivalents" according to the invention comprise polypeptides as described herein in deglycosylated or glycosylated form as well as modified forms that can be obtained by altering the glycosylation pattern.

[0654] Functional equivalents or homologues of the polypeptides according to the invention can be produced by mutagenesis, e.g. by point mutation, lengthening or shortening of the protein or as described in more detail below.

[0655] Functional equivalents or homologs of the polypeptides according to the invention can be identified by screening combinatorial databases of mutants, for example shortening mutants. For example, a variegated database of protein variants can be produced by combinatorial mutagenesis at the nucleic acid level, e.g. by enzymatic ligation of a mixture of synthetic oligonucleotides. There are a great many methods that can be used for the production of databases of potential homologues from a degenerated oligonucleotide sequence. Chemical synthesis of a degenerated gene sequence can be carried out in an automatic DNA synthesizer, and the synthetic gene can then be ligated in a suitable expression vector. The use of a degenerated genome makes it possible to supply all sequences in a mixture, which code for the desired set of potential protein sequences. Methods of synthesis of degenerated oligonucleotides are known to a person skilled in the art.

[0656] In the prior art, several techniques are known for the screening of gene products of combinatorial databases, which were produced by point mutations or shortening, and for the screening of cDNA libraries for gene products with a selected property. These techniques can be adapted for the rapid screening of the gene banks that were produced by combinatorial mutagenesis of homologues according to the invention. The techniques most frequently used for the screening of large gene banks, which are based on a high- throughput analysis, comprise cloning of the gene bank in expression vectors that can be replicated, transformation of the suitable cells with the resultant vector database and expression of the combinatorial genes in conditions in which detection of the desired activity facilitates isolation of the vector that codes for the gene whose product was detected. Recursive Ensemble Mutagenesis (REM), a technique that increases the frequency of functional mutants in the databases, can be used in combination with the screening tests, in order to identify homologues. An embodiment provided herein provides orthologs and paralogs of polypeptides disclosed herein as well as methods for identifying and isolating such orthologs and paralogs. A definition of the terms “ortholog” and “paralog” is given below and applies to amino acid and nucleic acid sequences.

[0657] The polypeptides of the invention include all active forms, including active subsequences, e.g., catalytic domains or active sites, of an enzyme of the invention. In one aspect, the invention provides catalytic domains or active sites as set forth below. In one aspect, the invention provides a peptide or polypeptide comprising or consisting of an active site domain as predicted through use of a database such as Pfam (http: / / pfam.wustl.edu / hmmsearch.shtml) (which is a large collection of multiple sequence alignments and hidden Markov models covering many common protein families, The Pfam protein families database, A. Bateman, E. Birney, L. Cerruti, R. Durbin, L. Etwiller, S. R. Eddy, S. Griffiths-Jones, K. L. Howe, M. Marshall, and E. L. L. Sonnhammer, Nucleic Acids Research, 30(1):276-280, 2002) or equivalent, as for example InterPro and SMART databases (http: / / www.ebi.ac.uk / interpro / scan.html, http: / / smart.embl- heidelberg.de / ).

[0658] The invention also encompasses “polypeptide variant” having the desired activity, wherein the variant polypeptide is selected from an amino acid sequence having at least 40%, 45%, 50%. 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or more sequence identity to a specific, in particular natural, amino acid sequence as referred to by a specific SEQ ID NO and contains at least one substitution modification relative to said SEQ ID NO.

[0659] Coding nucleic acid sequences applicable according to the invention

[0660] In this context the following definitions apply:

[0661] The terms “nucleic acid sequence,” “nucleic acid,” “nucleic acid molecule” and “polynucleotide” are used interchangeably meaning a sequence of nucleotides. A nucleic acid sequence may be a single-stranded or double-stranded deoxyribonucleotide, or ribonucleotide of any length, and include coding and non-coding sequences of a gene, exons, introns, sense and anti-sense complimentary sequences, genomic DNA, cDNA, miRNA, siRNA, mRNA, rRNA, tRNA, recombinant nucleic acid sequences, isolated and purified naturally occurring DNA and / or RNA sequences, synthetic DNA and RNA sequences, fragments, primers and nucleic acid probes. The skilled artisan is aware that the nucleic acid sequences of RNA are identical to the DNA sequences with the difference of thymine (T) being replaced by uracil (U). The term “nucleotide sequence” should also be understood as comprising a polynucleotide molecule or an oligonucleotide molecule in the form of a separate fragment or as a component of a larger nucleic acid.

[0662] An “isolated nucleic acid” or “isolated nucleic acid sequence” relates to a nucleic acid or nucleic acid sequence that is in an environment different from that in which the nucleic acid or nucleic acid sequence naturally occurs and can include those that are substantially free from contaminating endogenous material. The term “naturally-occurring” as used herein as applied to a nucleic acid refers to a nucleic acid that is found in a cell of an organism in nature and which has not been intentionally modified by a human in the laboratory.

[0663] A “fragment” of a polynucleotide or nucleic acid sequence refers to contiguous nucleotides that is particularly at least 15 bp, at least 30 bp, at least 40 bp, at least 50 bp and / or at least 60 bp in length of the polynucleotide of an embodiment herein. Particularly the fragment of a polynucleotide comprises at least 25, more particularly at least 50, more particularly at least 75, more particularly at least 100, more particularly at least 150, more particularly at least 200, more particularly at least 300, more particularly at least 400, more particularly at least 500, more particularly at least 600, more particularly at least 700, more particularly at least 800, more particularly at least 900, more particularly at least 1000 contiguous nucleotides of the polynucleotide of an embodiment herein. Without being limited, the fragment of the polynucleotides herein may be used as a PCR primer, and / or as a probe, or for anti-sense gene silencing or RNAi.

[0664] “Recombinant nucleic acid sequences” are nucleic acid sequences that result from the use of laboratory methods (for example, molecular cloning) to bring together genetic material from more than on source, creating or modifying a nucleic acid sequence that does not occur naturally and would not be otherwise found in biological organisms.

[0665] “Recombinant DNA technology” refers to molecular biology procedures to prepare a recombinant nucleic acid sequence as described, for instance, in Laboratory Manuals edited by Weigel and Glazebrook, 2002, Cold Spring Harbor Lab Press; and Sambrook et al., 1989, Cold Spring Harbor, NY, Cold Spring Harbor Laboratory Press.

[0666] The term “gene” means a DNA sequence comprising a region, which is transcribed into a RNA molecule, e.g., an mRNA in a cell, operably linked to suitable regulatory regions, e.g., a promoter. A gene may thus comprise several operably linked sequences, such as a promoter, a 5’ leader sequence comprising, e.g., sequences involved in translation initiation, a coding region of cDNA or genomic DNA, introns, exons, and / or a 3’non-translated sequence comprising, e.g., transcription termination sites.

[0667] “Polycistronic” refers to nucleic acid molecules, in particular mRNAs, that can encode more than one polypeptide separately within the same nucleic acid molecule.

[0668] A “chimeric gene” refers to any gene which is not normally found in nature in a species, in particular, a gene in which one or more parts of the nucleic acid sequence are present that are not associated with each other in nature. For example, the promoter is not associated in nature with part or all of the transcribed region or with another regulatory region. The term “chimeric gene” is understood to include expression constructs in which a promoter or transcription regulatory sequence is operably linked to one or more coding sequences or to an antisense, i.e., reverse complement of the sense strand, or inverted repeat sequence (sense and antisense, whereby the RNA transcript forms double stranded RNA upon transcription). The term "chimeric gene" also includes genes obtained through the combination of portions of one or more coding sequences to produce a new gene.

[0669] A “3’ UTR” or “3’ non-translated sequence” (also referred to as “3’ untranslated region,” or “3’end”) refers to the nucleic acid sequence found downstream of the coding sequence of a gene, which comprises, for example, a transcription termination site and (in most, but not all eukaryotic mRNAs) a polyadenylation signal such as AAUAAA or variants thereof. After termination of transcription, the mRNA transcript may be cleaved downstream of the polyadenylation signal and a poly(A) tail may be added, which is involved in the transport of the mRNA to the site of translation, e.g., cytoplasm.

[0670] The term “primer” refers to a short nucleic acid sequence that is hybridized to a template nucleic acid sequence and is used for polymerization of a nucleic acid sequence complementary to the template.

[0671] The term “selectable marker” refers to any gene which upon expression may be used to select a cell or cells that include the selectable marker. Examples of selectable markers are described below. The skilled artisan will know that different antibiotic, fungicide, auxotrophic or herbicide selectable markers are applicable to different target species.

[0672] The invention also relates to nucleic acid sequences that code for polypeptides as defined herein.

[0673] In particular, the invention also relates to nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, e.g. cDNA, genomic DNA and mRNA), coding for one of the above polypeptides and their functional equivalents, which can be obtained for example using artificial nucleotide analogs.

[0674] The invention relates both to isolated nucleic acid molecules, which code for polypeptides according to the invention or biologically active segments thereof, and to nucleic acid fragments, which can be used for example as hybridization probes or primers for identifying or amplifying coding nucleic acids according to the invention.

[0675] The present invention also relates to nucleic acids with a certain degree of “identity” to the sequences specifically disclosed herein. "Identity" between two nucleic acids means identity of the nucleotides, in each case over the entire length of the nucleic acid.

[0676] The “identity” between two nucleotide sequences (the same applies to peptide or amino acid sequences) is a function of the number of nucleotide residues (or amino acid residues) or that are identical in the two sequences when an alignment of these two sequences has been generated. Identical residues are defined as residues that are the same in the two sequences in a given position of the alignment. The percentage of sequence identity, as used herein, is calculated from the optimal alignment by taking the number of residues identical between two sequences dividing it by the total number of residues in the shortest sequence and multiplying by 100. The optimal alignment is the alignment in which the percentage of identity is the highest possible. Gaps may be introduced into one or both sequences in one or more positions of the alignment to obtain the optimal alignment. These gaps are then taken into account as non-identical residues for the calculation of the percentage of sequence identity. Alignment for the purpose of determining the percentage of amino acid or nucleic acid sequence identity can be achieved in various ways using computer programs and for instance publicly available computer programs available on the world wide web.

[0677] Particularly, the BLAST program (Tatiana et al, FEMS Microbiol Let., 1999, 174:247-250, 1999) set to the default parameters, available from the National Center for Biotechnology Information (NCBI) website at ncbi.nlm.nih.gov / BLAST / bl2seq / wblast2.cgi, can be used to obtain an optimal alignment of protein or nucleic acid sequences and to calculate the percentage of sequence identity.

[0678] Alternatively, the identity may be determined according to Chenna, et al. (2003), the web page: http: / / www.ebi.ac.Uk / Tools / clustalw / index.html# and the following settings:

[0679] DNA Gap Open Penalty 15.0

[0680] DNA Gap Extension Penalty 6.66

[0681] DNA Matrix Identity

[0682] Protein Gap Open Penalty 10.0

[0683] Protein Gap Extension Penalty 0.2

[0684] Protein matrix Gonnet

[0685] Protein / DNA ENDGAP -1

[0686] Protein / DNA GAPDIST 4

[0687] All the nucleic acid sequences mentioned herein (single-stranded and double-stranded DNA and RNA sequences, for example cDNA and mRNA) can be produced in a known way by chemical synthesis from the nucleotide building blocks, e.g. by fragment condensation of individual overlapping, complementary nucleic acid building blocks of the double helix. Chemical synthesis of oligonucleotides can, for example, be performed in a known way, by the phosphoamidite method (Voet, Voet, 2ndedition, Wiley Press, New York, pages 896-897). The accumulation of synthetic oligonucleotides and filling of gaps by means of the Klenow fragment of DNA polymerase and ligation reactions as well as general cloning techniques are described in Sambrook et al. (1989), see below.

[0688] The nucleic acid molecules according to the invention can in addition contain non-translated sequences from the 3’ and / or 5’ end of the coding genetic region.

[0689] The invention further relates to the nucleic acid molecules that are complementary to the concretely described nucleotide sequences or a segment thereof. The nucleotide sequences according to the invention make possible the production of probes and primers that can be used for the identification and / or cloning of homologous sequences in other cellular types and organisms. Such probes or primers generally comprise a nucleotide sequence region which hybridizes under “stringent” conditions (as defined herein elsewhere) on at least about 12, preferably at least about 25, for example about 40, 50 or 75 successive nucleotides of a sense strand of a nucleic acid sequence according to the invention or of a corresponding antisense strand.

[0690] “Homologous” sequences include orthologous or paralogous sequences. Methods of identifying orthologs or paralogs including phylogenetic methods, sequence similarity and hybridization methods are known in the art and are described herein.

[0691] “Paralogs” result from gene duplication that gives rise to two or more genes with similar sequences and similar functions. Paralogs typically cluster together and are formed by duplications of genes within related plant species. Paralogs are found in groups of similar genes using pair-wise Blast analysis or during phylogenetic analysis of gene families using programs such as CLUSTAL. In paralogs, consensus sequences can be identified characteristic to sequences within related genes and having similar functions of the genes.

[0692] “Orthologs”, or orthologous sequences, are sequences similar to each other because they are found in species that descended from a common ancestor. For instance, plant species that have common ancestors are known to contain many enzymes that have similar sequences and functions. The skilled artisan can identify orthologous sequences and predict the functions of the orthologs, for example, by constructing a polygenic tree for a gene family of one species using CLUSTAL or BLAST programs. A method for identifying or confirming similar functions among homologous sequences is by comparing of the transcript profiles in host cells or organisms, such as plants or microorganisms, overexpressing or lacking (in knockouts / knockdowns) related polypeptides. The skilled person will understand that genes having similar transcript profiles, with greater than 50% regulated transcripts in common, or with greater than 70% regulated transcripts in common, or greater than 90% regulated transcripts in common will have similar functions. Homologs, paralogs, orthologs and any other variants of the sequences herein are expected to function in a similar manner by making the host cells, organism such as plants or microorganisms producing terpene cyclase proteins.

[0693] A nucleic acid molecule according to the invention can be recovered by means of standard techniques of molecular biology and the sequence information supplied according to the invention. For example, cDNA can be isolated from a suitable cDNA library, using one of the concretely disclosed complete sequences or a segment thereof as hybridization probe and standard hybridization techniques (as described for example in Sambrook, (1989)).

[0694] In addition, a nucleic acid molecule comprising one of the disclosed sequences or a segment thereof, can be isolated by the polymerase chain reaction, using the oligonucleotide primers that were constructed on the basis of this sequence. The nucleic acid amplified in this way can be cloned in a suitable vector and can be characterized by DNA sequencing. The oligonucleotides according to the invention can also be produced by standard methods of synthesis, e.g. using an automatic DNA synthesizer.

[0695] To test a function of variant DNA sequences according to an embodiment herein, the sequence of interest is operably linked to a selectable or screenable marker gene and expression of said reporter gene is tested in transient expression assays, for example, with microorganisms or with protoplasts or in stably transformed plants.

[0696] The invention also relates to derivatives of the concretely disclosed or derivable nucleic acid sequences.

[0697] Thus, further nucleic acid sequences according to the invention can be derived from the sequences specifically disclosed herein and can differ from it by one or more, like 1 to 20, in particular 1 to 15 or 5 to 10 additions, substitutions, insertions or deletions of one or several (like for example 1 to 10) nucleotides, and furthermore code for polypeptides with the desired profile of properties.

[0698] The invention also encompasses nucleic acid sequences that comprise so-called silent mutations or have been altered, in comparison with a concretely stated sequence, according to the codon usage of a special original or host organism.

[0699] According to a particular embodiment of the invention variant nucleic acids may be prepared in order to adapt its nucleotide sequence to a specific expression system. For example, bacterial expression systems are known to more efficiently express polypeptides if amino acids are encoded by particular codons. Due to the degeneracy of the genetic code, more than one codon may encode the same amino acid sequence, multiple nucleic acid sequences can code for the same protein or polypeptide, all these DNA sequences being encompassed by an embodiment herein. Where appropriate, the nucleic acid sequences encoding the polypeptides described herein may be optimized for increased expression in the host cell. For example, nucleic acids of an embodiment herein may be synthesized using codons particular to a host for improved expression.

[0700] The invention also encompasses naturally occurring variants, e.g. splicing variants or allelic variants, of the sequences described therein.

[0701] Allelic variants may have at least 60 % homology at the level of the derived amino acid, preferably at least 80 % homology, quite especially preferably at least 90 % homology over the entire sequence range (regarding homology at the amino acid level, reference should be made to the details given above for the polypeptides). Advantageously, the homologies can be higher over partial regions of the sequences. The invention also relates to sequences that can be obtained by conservative nucleotide substitutions (i.e. as a result thereof the amino acid in question is replaced by an amino acid of the same charge, size, polarity and / or solubility).

[0702] The invention also relates to the molecules derived from the concretely disclosed nucleic acids by sequence polymorphisms. Such genetic polymorphisms may exist in cells from different populations or within a population due to natural allelic variation. Allelic variants may also include functional equivalents. These natural variations usually produce a variance of 1 to 5 % in the nucleotide sequence of a gene. Said polymorphisms may lead to changes in the amino acid sequence of the polypeptides disclosed herein. Allelic variants may also include functional equivalents.

[0703] Furthermore, derivatives are also to be understood to be homologs of the nucleic acid sequences according to the invention, for example animal, plant, fungal or bacterial homologs, shortened sequences, singlestranded DNA or RNA of the coding and noncoding DNA sequence. For example, homologs have, at the DNA level, a homology of at least 40 %, preferably of at least 60 %, especially preferably of at least 70 %, quite especially preferably of at least 80 % over the entire DNA region given in a sequence specifically disclosed herein.

[0704] Moreover, derivatives are to be understood to be, for example, fusions with promoters. The promoters that are added to the stated nucleotide sequences can be modified by at least one nucleotide exchange, at least one insertion, inversion and / or deletion, though without impairing the functionality or efficacy of the promoters. Moreover, the efficacy of the promoters can be increased by altering their sequence or can be exchanged completely with more effective promoters even of organisms of a different genus.

[0705] Generation of functional polypeptide mutants

[0706] Moreover, a person skilled in the art is familiar with methods for generating functional mutants, that is to say nucleotide sequences which code for a polypeptide with at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or more sequence identity to anyone of amino acid related to SEQ ID NOs as disclosed herein and / or encoded by a nucleic acid molecule comprising a nucleotide sequence having at least 50% sequence identity to anyone of the nucleotide related to SEQ ID NOs as disclosed herein.

[0707] Depending on the technique used, a person skilled in the art can introduce entirely random or else more directed mutations into genes or else noncoding nucleic acid regions (which are for example important for regulating expression) and subsequently generate genetic libraries. The methods of molecular biology required for this purpose are known to the skilled worker and for example described in Sambrook and Russell, Molecular Cloning.3rd Edition, Cold Spring Harbor Laboratory Press 2001 . Methods for modifying genes and thus for modifying the polypeptide encoded by them have been known to the skilled worker for a long time, such as, for example:

[0708] - site-specific mutagenesis, where individual or several nucleotides of a gene are replaced in a directed fashion (Trower MK (Ed.) 1996; In vitro mutagenesis protocols. Humana Press, New Jersey),

[0709] - saturation mutagenesis, in which a codon for any amino acid can be exchanged or added at any point of a gene (Kegler-Ebo DM, Docktor CM, DiMaio D (1994) Nucleic Acids Res 22:1593; Barettino D, Feigenbutz M, Valcarel R, Stunnenberg HG (1994) Nucleic Acids Res 22:541 ; Barik S (1995) Mol Biotechnol 3:1),

[0710] - error-prone polymerase chain reaction, where nucleotide sequences are mutated by error-prone DNA polymerases (Eckert KA, Kunkel TA (1990) Nucleic Acids Res 18:3739);

[0711] - the SeSaM method (sequence saturation method), in which preferred exchanges are prevented by the polymerase. Schenk et al., Biospektrum, Vol. 3, 2006, 277-279

[0712] - the passaging of genes in mutator strains, in which, for example owing to defective DNA repair mechanisms, there is an increased mutation rate of nucleotide sequences (Greener A, Callahan M, Jerpseth B (1996) An efficient random mutagenesis technique using an E.coli mutator strain. In: Trower MK (Ed.) In vitro mutagenesis protocols. Humana Press, New Jersey), or

[0713] - DNA shuffling, in which a pool of closely related genes is formed and digested and the fragments are used as templates for a polymerase chain reaction in which, by repeated strand separation and reassociation, full-length mosaic genes are ultimately generated (Stemmer WPC (1994) Nature 370:389; Stemmer WPC (1994) Proc Natl Acad Sci USA 91 :10747).

[0714] Using so-called directed evolution (described, inter alia, in Reetz MT and Jaeger K-E (1999), Topics Curr Chem 200:31 ; Zhao H, Moore JC, Volkov AA, Arnold FH (1999), Methods for optimizing industrial polypeptides by directed evolution, In: Demain AL, Davies JE (Ed.) Manual of industrial microbiology and biotechnology. American Society for Microbiology), a skilled worker can produce functional mutants in a directed manner and on a large scale. To this end, in a first step, gene libraries of the respective polypeptides are first produced, for example using the methods given above. The gene libraries are expressed in a suitable way, for example by bacteria or by phage display systems.

[0715] The relevant genes of host organisms which express functional mutants with properties that largely correspond to the desired properties can be submitted to another mutation cycle. The steps of the mutation and selection or screening can be repeated iteratively until the present functional mutants have the desired properties to a sufficient extent. Using this iterative procedure, a limited number of mutations, for example 1 , 2, 3, 4 or 5 mutations, can be performed in stages and assessed and selected for their influence on the activity in question. The selected mutant can then be submitted to a further mutation step in the same way. In this way, the number of individual mutants to be investigated can be reduced significantly.

[0716] The results according to the invention also provide important information relating to structure and sequence of the relevant polypeptides, which is required for generating, in a targeted fashion, further polypeptides with desired modified properties. In particular, it is possible to define so-called “hot spots”, i.e. sequence segments that are potentially suitable for modifying a property by introducing targeted mutations.

[0717] Information can also be deduced regarding amino acid sequence positions, in the region of which mutations can be affected that should probably have little effect on the activity, and can be designated as potential “silent mutations”.

[0718] Constructs for expressing polypeptides of the invention and / or used in the process of the invention

[0719] In this context the following definitions apply:

[0720] “Expression of a gene” encompasses “heterologous expression” and “over-expression” and involves transcription of the gene and translation of the mRNA into a protein. Overexpression refers to the production of the gene product as measured by levels of mRNA, polypeptide and / or enzyme activity in transgenic cells or organisms that exceeds levels of production in non-transformed cells or organisms of a similar genetic background.

[0721] “Expression vector” as used herein means a nucleic acid molecule engineered using molecular biology methods and recombinant DNA technology for delivery of foreign or exogenous DNA into a host cell. The expression vector typically includes sequences required for proper transcription of the nucleotide sequence. The coding region usually codes for a protein of interest but may also code for an RNA, e.g., an antisense RNA, siRNA and the like.

[0722] An “expression vector” as used herein includes any linear or circular recombinant vector including but not limited to viral vectors, bacteriophages and plasmids. The skilled person is capable of selecting a suitable vector according to the expression system. In one embodiment, the expression vector includes the nucleic acid of an embodiment herein operably linked to at least one “regulatory sequence”, which controls transcription, translation, initiation and termination, such as a transcriptional promoter, operator or enhancer, or an mRNA ribosomal binding site and, optionally, including at least one selection marker. Nucleotide sequences are “operably linked” when the regulatory sequence functionally relates to the nucleic acid of an embodiment herein.

[0723] An “expression system” as used herein encompasses any combination of nucleic acid molecules required for the expression of one, or the co-expression of two or more polypeptides either in vivo of a given expression host, or in vitro. The respective coding sequences may either be located on a single nucleic acid molecule or vector, as for example a vector containing multiple cloning sites, or on a polycistronic nucleic acid, or may be distributed over two or more physically distinct vectors. As a particular example there may be mentioned an operon comprising a promotor sequence, one or more operator sequences and one or more structural genes each encoding an enzyme as described herein. As used herein, the terms "amplifying" and "amplification" refer to the use of any suitable amplification methodology for generating or detecting recombinant of naturally expressed nucleic acid, as described in detail, below. For example, the invention provides methods and reagents (e.g., specific degenerate oligonucleotide primer pairs, oligo dT primer) for amplifying (e.g., by polymerase chain reaction, PCR) naturally expressed (e.g., genomic DNA or mRNA) or recombinant (e.g., cDNA) nucleic acids of the invention in vivo, ex vivo or in vitro.

[0724] “Regulatory sequence” refers to a nucleic acid sequence that determines expression level of the nucleic acid sequences of an embodiment herein and is capable of regulating the rate of transcription of the nucleic acid sequence operably linked to the regulatory sequence. Regulatory sequences comprise promoters, enhancers, transcription factors, promoter elements and the like.

[0725] A “promoter”, a “nucleic acid with promoter activity” or a “promoter sequence” is understood as meaning, in accordance with the invention, a nucleic acid which, when functionally linked to a nucleic acid to be transcribed, regulates the transcription of said nucleic acid. “Promoter” in particular refers to a nucleic acid sequence that controls the expression of a coding sequence by providing a binding site for RNA polymerase and other factors required for proper transcription including without limitation transcription factor binding sites, repressor and activator protein binding sites. The meaning of the term promoter also includes the term “promoter regulatory sequence”. Promoter regulatory sequences may include upstream and downstream elements that may influences transcription, RNA processing or stability of the associated coding nucleic acid sequence. Promoters include naturally-derived and synthetic sequences. The coding nucleic acid sequences is usually located downstream of the promoter with respect to the direction of the transcription starting at the transcription initiation site.

[0726] In this context, a “functional” or “operative” linkage is understood as meaning for example the sequential arrangement of one of the nucleic acids with a regulatory sequence. For example, the sequence with promoter activity and of a nucleic acid sequence to be transcribed and optionally further regulatory elements, for example nucleic acid sequences which ensure the transcription of nucleic acids, and for example a terminator, are linked in such a way that each of the regulatory elements can perform its function upon transcription of the nucleic acid sequence. This does not necessarily require a direct linkage in the chemical sense. Genetic control sequences, for example enhancer sequences, can even exert their function on the target sequence from more remote positions or even from other DNA molecules. Preferred arrangements are those in which the nucleic acid sequence to be transcribed is positioned behind (i.e. at the 3’-end of) the promoter sequence so that the two sequences are joined together covalently. The distance between the promoter sequence and the nucleic acid sequence to be expressed recombinantly can be smaller than 200 base pairs, or smaller than 100 base pairs or smaller than 50 base pairs.

[0727] In addition to promoters and terminator, the following may be mentioned as examples of other regulatory elements: targeting sequences, enhancers, polyadenylation signals, selectable markers, amplification signals, replication origins and the like. Suitable regulatory sequences are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).

[0728] The term “constitutive promoter” refers to an unregulated promoter that allows for continual transcription of the nucleic acid sequence it is operably linked to.

[0729] As used herein, the term “operably linked” refers to a linkage of polynucleotide elements in a functional relationship. A nucleic acid is “operably linked” when it is placed into a functional relationship with another nucleic acid sequence. For instance, a promoter, or rather a transcription regulatory sequence, is operably linked to a coding sequence if it affects the transcription of the coding sequence. Operably linked means that the DNA sequences being linked are typically contiguous. The nucleotide sequence associated with the promoter sequence may be of homologous or heterologous origin with respect to the plant to be transformed. The sequence also may be entirely or partially synthetic. Regardless of the origin, the nucleic acid sequence associated with the promoter sequence will be expressed or silenced in accordance with promoter properties to which it is linked after binding to the polypeptide of an embodiment herein. The associated nucleic acid may code for a protein that is desired to be expressed or suppressed throughout the organism at all times or, alternatively, at a specific time or in specific tissues, cells, or cell compartment. Such nucleotide sequences particularly encode proteins conferring desirable phenotypic traits to the host cells or organism altered or transformed therewith. More particularly, the associated nucleotide sequence leads to the production of the product or products of interest as herein defined in the cell or organism. Particularly, the nucleotide sequence encodes a polypeptide having an enzyme activity as herein defined.

[0730] The nucleotide sequence as described herein above may be part of an “expression cassette”. The terms “expression cassette” and “expression construct” are used synonymously. The (preferably recombinant) expression construct contains a nucleotide sequence which encodes a polypeptide according to the invention and which is under genetic control of regulatory nucleic acid sequences.

[0731] In a process applied according to the invention, the expression cassette may be part of an “expression vector”, in particular of a recombinant expression vector.

[0732] An “expression unit” is understood as meaning, in accordance with the invention, a nucleic acid with expression activity which comprises a promoter as defined herein and, after functional linkage with a nucleic acid to be expressed or a gene, regulates the expression, i.e. the transcription and the translation of said nucleic acid or said gene. It is therefore in this connection also referred to as a “regulatory nucleic acid sequence”. In addition to the promoter, other regulatory elements, for example enhancers, can also be present.

[0733] An “expression cassette” or “expression construct” is understood as meaning, in accordance with the invention, an expression unit which is functionally linked to the nucleic acid to be expressed or the gene to be expressed. In contrast to an expression unit, an expression cassette therefore comprises not only nucleic acid sequences which regulate transcription and translation, but also the nucleic acid sequences that are to be expressed as protein as a result of transcription and translation.

[0734] The terms “expression” or “overexpression” describe, in the context of the invention, the production or increase in intracellular activity of one or more polypeptides in a microorganism, which are encoded by the corresponding DNA. To this end, it is possible for example to introduce a gene into an organism, replace an existing gene with another gene, increase the copy number of the gene(s), use a strong promoter or use a gene which encodes for a corresponding polypeptide with a high activity; optionally, these measures can be combined.

[0735] Preferably such constructs according to the invention comprise a promoter 5’-upstream of the respective coding sequence and a terminator sequence 3’-downstream and optionally other usual regulatory elements, in each case in operative linkage with the coding sequence.

[0736] Nucleic acid constructs according to the invention comprise in particular a sequence coding for a polypeptide for example derived from the amino acid related SEQ ID NOs as described therein or the reverse complement thereof, or derivatives and homologs thereof and which have been linked operatively or functionally with one or more regulatory signals, advantageously for controlling, for example increasing, gene expression.

[0737] In addition to these regulatory sequences, the natural regulation of these sequences may still be present before the actual structural genes and optionally may have been genetically modified so that the natural regulation has been switched off and expression of the genes has been enhanced. The nucleic acid construct may, however, also be of simpler construction, i.e. no additional regulatory signals have been inserted before the coding sequence and the natural promoter, with its regulation, has not been removed. Instead, the natural regulatory sequence is mutated such that regulation no longer takes place and the gene expression is increased.

[0738] A preferred nucleic acid construct advantageously also comprises one or more of the already mentioned “enhancer” sequences in functional linkage with the promoter, which sequences make possible an enhanced expression of the nucleic acid sequence. Additional advantageous sequences may also be inserted at the 3’-end of the DNA sequences, such as further regulatory elements or terminators. One or more copies of the nucleic acids according to the invention may be present in a construct. In the construct, other markers, such as genes which complement auxotrophisms or antibiotic resistances, may also optionally be present so as to select for the construct.

[0739] Examples of suitable regulatory sequences are present in promoters such as cos, tac, trp, tet, trp-tet, Ipp, lac, Ipp-lac, laclq, T7, T5, T3, gal, trc, ara, rhaP (rhaPBAD)SP6, lambda-PR or in the lambda-Pi. promoter, and these are advantageously employed in Gram-negative bacteria. Further advantageous regulatory sequences are present for example in the Gram-positive promoters amy and SPO2, in the yeast or fungal promoters ADC1 , MFalpha, AC, P-60, CYC1 , GAPDH, TEF, rp28, ADH. Artificial promoters may also be used for regulation.

[0740] For expression in a host organism, the nucleic acid construct is inserted advantageously into a vector such as, for example, a plasmid or a phage, which makes possible optimal expression of the genes in the host. Vectors are also understood as meaning, in addition to plasmids and phages, all the other vectors which are known to the skilled worker, that is to say for example viruses such as SV40, CMV, baculovirus and adenovirus, transposons, IS elements, phasmids, cosmids and linear or circular DNA or artificial chromosomes. These vectors are capable of replicating autonomously in the host organism or else chromosomally. These vectors are a further development of the invention. Binary or cpo-integration vectors are also applicable.

[0741] Suitable plasmids are, for example, in E. coli pLG338, pACYC184, pBR322, pUC18, pUC19, pKC30, pRep4, pHS1 , pKK223-3, pDHE19.2, pHS2, pPLc236, pMBL24, pLG200, pUR290, plN-lll113-B1 , Agt11 or pBdCI, in Streptomyces plJ1 O1 , plJ364, plJ702 or plJ361 , in Bacillus pUB110, pC194 or pBD214, in Corynebacterium pSA77 or pAJ667, in fungi pALS1 , plL2 or pBB116, in yeasts 2alphaM, pAG-1 , YEp6, YEp13 or pEMBLYe23 or in plants pLGV23, pGHIac+, pBIN19, pAK2004 or pDH51 . The abovementioned plasmids are a small selection of the plasmids which are possible. Further plasmids are well known to the skilled worker and can be found for example in the book Cloning Vectors (Eds. Pouwels P. H. et al. Elsevier, Amsterdam-New York-Oxford, 1985, ISBN 0 444 904018).

[0742] In a further development of the vector, the vector which comprises the nucleic acid construct according to the invention or the nucleic acid according to the invention can advantageously also be introduced into the microorganisms in the form of a linear DNA and integrated into the host organism’s genome via heterologous or homologous recombination. This linear DNA can consist of a linearized vector such as a plasmid or only of the nucleic acid construct or the nucleic acid according to the invention.

[0743] For optimal expression of heterologous genes in organisms, it is advantageous to modify the nucleic acid sequences to match the specific “codon usage” used in the organism. The “codon usage” can be determined readily by computer evaluations of other, known genes of the organism in question.

[0744] An expression cassette according to the invention is generated by fusing a suitable promoter to a suitable coding nucleotide sequence and a terminator or polyadenylation signal. Customary recombination and cloning techniques are used for this purpose, as are described, for example, in T. Maniatis, E.F. Fritsch and J. Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1989) and in T.J. Silhavy, M.L. Berman and L.W. Enquist, Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1984) and in Ausubel, F.M. et al., Current Protocols in Molecular Biology, Greene Publishing Assoc, and Wiley Interscience (1987). For expression in a suitable host organism, the recombinant nucleic acid construct or gene construct is advantageously inserted into a host-specific vector which makes possible optimal expression of the genes in the host. Vectors are well known to the skilled worker and can be found for example in “cloning vectors” (Pouwels P. H. et al., Ed., Elsevier, Amsterdam-New York-Oxford, 1985).

[0745] An alternative embodiment of an embodiment herein provides a method to “alter gene expression” in a host cell. For instance, the polynucleotide of an embodiment herein may be enhanced or overexpressed or induced in certain contexts (e.g. upon exposure to certain temperatures or culture conditions) in a host cell or host organism.

[0746] Alteration of expression of a polynucleotide provided herein may also result in ectopic expression which is a different expression pattern in an altered and in a control or wild-type organism. Alteration of expression occurs from interactions of polypeptide of an embodiment herein with exogenous or endogenous modulators, or as a result of chemical modification of the polypeptide. The term also refers to an altered expression pattern of the polynucleotide of an embodiment herein which is altered below the detection level or completely suppressed activity.

[0747] In one embodiment, provided herein is also an isolated, recombinant or synthetic polynucleotide encoding a polypeptide or variant polypeptide provided herein.

[0748] In one embodiment, several polypeptide encoding nucleic acid sequences are co-expressed in a single host, particularly under control of different promoters. In another embodiment, several polypeptide encoding nucleic acid sequences can be present on a single transformation vector or be co-transformed at the same time using separate vectors and selecting transformants comprising both chimeric genes. Similarly, one or polypeptide encoding genes may be expressed in a single plant, cell, microorganism or organism together with other chimeric genes.

[0749] Recombinant production of polypeptides according to the invention

[0750] The invention further relates to methods for recombinant production of polypeptides according to the invention or functional, biologically active fragments thereof, wherein a polypeptide-producing microorganism is cultured, optionally the expression of the polypeptides is induced by applying at least one inducer inducing gene expression and the expressed polypeptides are isolated from the culture. The polypeptides can also be produced in this way on an industrial scale, if desired.

[0751] The microorganisms produced according to the invention can be cultured continuously or discontinuously in the batch method or in the fed-batch method or repeated fed-batch method. A summary of known cultivation methods can be found in the textbook by Chmiel (Bioprozesstechnik 1. Einfuhrung in die Bioverfahrenstechnik [Bioprocess technology 1. Introduction to bioprocess technology] (Gustav Fischer Verlag, Stuttgart, 1991)) or in the textbook by Storhas (Bioreaktoren und periphere Einrichtungen [Bioreactors and peripheral equipment] (Vieweg Verlag, Braunschweig / Wiesbaden, 1994)). Standard laboratory methods can be used for this purpose and are known in the art and also further described herein.

[0752] If the polypeptides are not secreted in the culture medium, the cells can also be lysed and the product can be obtained from the lysate by known methods for isolation of proteins. The cells can optionally be disrupted with high-frequency ultrasound, high pressure, for example in a French press, by osmolysis, by the action of detergents, lytic enzymes or organic solvents, by means of homogenizers or by a combination of several of the aforementioned methods.

[0753] The polypeptides can be purified by known chromatographic techniques, such as molecular sieve chromatography (gel filtration), such as Q-sepharose chromatography, ion exchange chromatography and hydrophobic chromatography, and with other usual techniques such as ultrafiltration, crystallization, saltingout, dialysis and native gel electrophoresis. Suitable methods are described for example in Cooper, T. G., Biochemische Arbeitsmethoden [Biochemical processes], Verlag Walter de Gruyter, Berlin, New York or in Scopes, R., Protein Purification, Springer Verlag, New York, Heidelberg, Berlin.

[0754] For isolating the recombinant protein, it can be advantageous to use vector systems or oligonucleotides, which lengthen the cDNA by defined nucleotide sequences and therefore code for altered polypeptides or fusion proteins, which for example serve for easier purification. Suitable modifications of this type are for example so-called "tags" functioning as anchors, for example the modification known as hexa-histidine anchor or epitopes that can be recognized as antigens of antibodies (described for example in Harlow, E. and Lane, D., 1988, Antibodies: A Laboratory Manual. Cold Spring Harbor (N.Y.) Press). These anchors can serve for attaching the proteins to a solid carrier, for example a polymer matrix, which can for example be used as packing in a chromatography column, or can be used on a microtiter plate or on some other carrier.

[0755] At the same time these anchors can also be used for recognition of the proteins. For recognition of the proteins, it is moreover also possible to use usual markers, such as fluorescent dyes, enzyme markers, which form a detectable reaction product after reaction with a substrate, or radioactive markers, alone or in combination with the anchors for derivatization of the proteins.

[0756] Uses of a compound of formula (I) and other compounds of the invention

[0757] A further aspect of the invention comprises the use of a compound of formula (I) or another (intermediate) compound obtained or obtainable from the embodiments of the invention as a perfumery, flavor or aroma ingredient and / or as a precursor for making said ingredient.

[0758] As mentioned above, the invention comprises the use of a compound of formula (I) as a perfuming ingredient. In other words, it concerns a method or a process to confer, enhance, improve or modify the odor properties of a perfuming composition or of a perfumed article or of a surface, which method comprises adding to said composition or article an effective amount of at least a compound of formula (I), e.g. to impart its typical note. Understood that the final hedonic effect may depend on the precise dosage and on the organoleptic properties of the invention’s compound, but anyway the addition of the invention’s compound will impart to the final product its typical touch in the form of a note, touch or aspect depending on the dosage.

[0759] By “use of a compound of formula (I)” it has to be understood here also the use of any composition containing a compound (I) and which can be advantageously employed in the perfumery industry.

[0760] Said compositions, which in fact can be advantageously employed as perfuming ingredients, are also an object of the present invention.

[0761] Therefore, another object of the present invention is a perfuming composition comprising: i) as a perfuming ingredient, at least one invention’s compound as defined above; ii) at least one ingredient selected from the group consisting of a perfumery carrier and a perfumery base; and iii) optionally at least one perfumery adjuvant.

[0762] By “perfumery carrier” it is meant here a material which is practically neutral from a perfumery point of view, i.e. that does not significantly alter the organoleptic properties of perfuming ingredients. Said carrier may be a liquid or a solid.

[0763] As liquid carrier one may cite, as non-limiting examples, an emulsifying system, i.e. a solvent and a surfactant system, or a solvent commonly used in perfumery. A detailed description of the nature and type of solvents commonly used in perfumery cannot be exhaustive. However, one can cite as non-limiting examples, solvents such as butylene or propylene glycol, glycerol, dipropyleneglycol and its monoether, 1 ,2,3-propanetriyl triacetate, dimethyl glutarate, dimethyl adipate 1 ,3-diacetyloxypropan-2-yl acetate, diethyl phthalate, isopropyl myristate, benzyl benzoate, benzyl alcohol, 2-(2-ethoxyethoxy)-1 -ethanol, triethyl citrate or mixtures thereof, which are the most commonly used. For the compositions which comprise both a perfumery carrier and a perfumery base, other suitable perfumery carriers than those previously specified, can be also ethanol, water / ethanol mixtures, limonene or other terpenes, isoparaffins such as those known under the trademark Isopar™ (origin: Exxon Chemical) or glycol ethers and glycol ether esters such as those known under the trademark Dowanol™ (origin: Dow Chemical Company), or hydrogenated castors oils such as those known under the trademark Cremophor® RH 40 (origin: BASF).

[0764] Solid carrier is meant to designate a material to which the perfuming composition or some element of the perfuming composition can be chemically or physically bound. In general, such solid carriers are employed either to stabilize the composition, or to control the rate of evaporation of the compositions or of some ingredients. Solid carriers are of current use in the art and a person skilled in the art knows how to reach the desired effect. However, by way of non-limiting examples of solid carriers, one may cite absorbing gums or polymers or inorganic materials, such as porous polymers, cyclodextrins, wood-based materials, organic or inorganic gels, clays, gypsum talc or zeolites.

[0765] As other non-limiting examples of solid carriers, one may cite encapsulating materials. Examples of such materials may comprise wall-forming and plasticizing materials, such as mono, di- or trisaccharides, natural or modified starches, hydrocolloids, cellulose derivatives, polyvinyl acetates, polyvinylalcohols, proteins or pectins, or yet the materials cited in reference texts such as H. Scherz, Hydrokolloide: Stabilisatoren, Dickungs- und Geliermittel in Lebensmitteln, Band 2 der Schriftenreihe Lebensmittelchemie, Lebensmittelqualitat, Behr's Verlag GmbH & Co., Hamburg, 1996. The encapsulation is a well-known process to a person skilled in the art, and may be performed, for instance, by using techniques such as spray-drying, agglomeration or yet extrusion; or consists of a coating encapsulation, including coacervation and complex coacervation techniques.

[0766] As non-limiting examples of solid carriers, one may cite in particular the core-shell capsules with resins of aminoplast, polyamide, polyester, polyurea or polyurethane type or a mixture thereof (all of said resins are well known to a person skilled in the art) using techniques like phase separation process induced by polymerization, interfacial polymerization, coacervation or altogether (all of said techniques have been described in the prior art), optionally in the presence of a polymeric stabilizer or of a cationic copolymer.

[0767] Resins may be produced by the polycondensation of an aldehyde (e.g. formaldehyde, 2,2- dimethoxyethanal, glyoxal, glyoxylic acid or glycolaldehyde and mixtures thereof) with an amine such as urea, benzoguanamine, glycoluryl, melamine, methylol melamine, methylated methylol melamine, guanazole and the like, as well as mixtures thereof. Alternatively, one may use preformed resins alkylolated polyamines such as those commercially available under the trademark Urac® (origin: Cytec Technology Corp.), Cymel® (origin: Cytec Technology Corp.), Urecoll® or Luracoll® (origin: BASF).

[0768] Other resins are the ones produced by the polycondensation of an a polyol, like glycerol, and a polyisocyanate, like a trimer of hexamethylene diisocyanate, a trimer of isophorone diisocyanate or xylylene diisocyanate or a Biuret of hexamethylene diisocyanate or a trimer of xylylene diisocyanate with trimethylolpropane (known with the tradename of Takenate®, origin: Mitsui Chemicals), among which a trimer of xylylene diisocyanate with trimethylolpropane and a Biuret of hexamethylene diisocyanate are preferred.

[0769] Some of the seminal literature related to the encapsulation of perfumes by polycondensation of amino resins, namely melamine-based resins with aldehydes includes articles such as those published by K. Dietrich et al. Acta Polymerica, 1989, vol. 40, pages 243, 325 and 683, as well as 1990, vol. 41 , page 91 . Such articles already describe the various parameters affecting the preparation of such core-shell microcapsules following prior art methods that are also further detailed and exemplified in the patent literature. US 4’396'670, to the Wiggins Teape Group Limited is a pertinent early example of the latter. Since then, many other authors have enriched the literature in this field and it would be impossible to cover all published developments here, but the general knowledge in encapsulation technology is very significant. More recent publications of pertinence, which disclose suitable uses of such microcapsules, are represented for example by the article of K. Bruyninckx and M. Dusselier, ACS Sustainable Chemistry & Engineering, 2019, vol. 7, pages 8041-8054.

[0770] By “perfumery base” what is meant here is a composition comprising at least one perfuming co-ingredient.

[0771] Said perfuming co-ingredient is not of formula (I). Moreover, by “perfuming co-ingredient” it is meant here a compound, which is used in a perfuming preparation or a composition to impart a hedonic effect. In other words such a co-ingredient, to be considered as being a perfuming one, must be recognized by a person skilled in the art as being able to impart or modify in a positive or pleasant way the odor of a composition, and not just as having an odor. The perfuming ingredient may impart an additional benefit beyond that of modifying or imparting an odor, such as long-lasting, blooming, malodour counteraction, antimicrobial effect, antiviral effect, microbial stability, or pest control.

[0772] The nature and type of the perfuming co-ingredients present in the base do not warrant a more detailed description here, which in any case would not be exhaustive, the skilled person being able to select them on the basis of his general knowledge and according to the intended use or application and the desired organoleptic effect. In general terms, these perfuming co-ingredients belong to chemical classes as varied as alcohols, lactones, aldehydes, ketones, esters, ethers, acetates, nitriles, terpenoids, nitrogenous or sulphurous heterocyclic compounds and essential oils, and said perfuming co-ingredients can be of natural or synthetic origin.

[0773] In particular one may cite perfuming co-ingredients which are commonly used in perfume formulations, such as:

[0774] Aldehydic ingredients: decanal, dodecanal, 2-methyl-undecanal, 10-undecenal, octanal, nonanal and / or nonenal;

[0775] Aromatic-herbal ingredients: eucalyptus oil, camphor, eucalyptol, 5- methyltricyclo[6.2.1 ,0~2,7~]undecan-4-one, 1-methoxy-3-hexanethiol, 2-ethyl-4,4-dimethyl-1 ,3-oxathiane, 2,2,7 / 8,9 / 10-Tetramethylspiro[5.5]undec-8-en-1-one, menthol and / or alpha-pinene;

[0776] Balsamic ingredients: coumarin, ethylvanillin and / or vanillin;

[0777] Citrus ingredients: dihydromyrcenol, citral, orange oil, linalyl acetate, citronellyl nitrile, orange terpenes, limonene, 1-p-menthen-8-yl acetate and / or 1 ,4(8)-p-menthadiene;

[0778] Floral ingredients:methyl dihydrojasmonate, linalool, citronellol, phenylethanol, 3-(4-tert- butylphenyl)-2-methylpropanal, hexylcinnamic aldehyde, benzyl acetate, benzyl salicylate, tetrahydro-2- isobutyl-4-methyl-4(2H)-pyranol, beta ionone, methyl 2-(methylamino)benzoate, (E)-3-methyl-4-(2,6,6- trimethyl-2-cyclohexen-1-yl)-3-buten-2-one, (1 E)-1-(2,6,6-trimethyl-2-cyclohexen-1-yl)-1-penten-3-one, 1- (2,6,6-trimethyl-1 ,3-cyclohexadien-1 -yl)-2-buten-1 -one, (2E)-1 -(2,6,6-trimethyl-2-cyclohexen-1 -yl)-2- buten-1 -one, (2E)-1 -[2,6,6-trimethyl-3-cyclohexen-1 -y l]-2- buten- 1 -one, (2E)-1 -(2,6,6-trimethyl-1 - cyclohexen-1 -yl)-2-buten-1 -one, 2,5-dimethyl-2-indanmethanol, 2,6,6-trimethyl-3-cyclohexene-1 - carboxylate, 3-(4,4-dimethyl-1-cyclohexen-1-yl)propanal, 3-(3,3 / 1 ,1-dimethyl-5-indanyl)propanal, hexyl salicylate, 3,7-dimethyl-1 ,6-nonadien-3-ol, 3-(4-isopropylphenyl)-2-methylpropanal, verdyl acetate, geraniol, p-menth-1-en-8-ol, 4-(1 ,1-dimethylethyl)-1-cyclohexyle acetate, 1 ,1-dimethyl-2-phenylethyl acetate, 4-cyclohexyl-2-methyl-2-butanol, amyl salicylate , high cis methyl dihydrojasmonate, 3-methyl-5- phenyl-1 -pentanol, verdyl proprionate, geranyl acetate, tetrahydro linalool, cis-7-p-menthanol, propyl (S)- 2-(1 ,1-dimethylpropoxy)propanoate, 2-methoxynaphthalene, 2, 2, 2-trichloro-1 -phenylethyl acetate, 4 / 3-(4- hydroxy-4-methylpentyl)-3-cyclohexene-1-carbaldehyde, amylcinnamic aldehyde, 8-decen-5-olide, 4- phenyl-2-butanone, isononyle acetate, 4-(1 , 1 -dimethylethyl)-1 -cyclohexyl acetate, verdyl isobutyrate and / or mixture of methylionones isomers;

[0779] Fruity ingredients: gamma-undecalactone, 2,2,5-trimethyl-5-pentylcyclopentanone, 2-methyl-4- propyl-1 ,3-oxathiane, 4-decanolide, ethyl 2-methyl-pentanoate, hexyl acetate, ethyl 2-methylbutanoate, gamma-nonalactone, allyl heptanoate, 2-phenoxyethyl isobutyrate, ethyl 2-methyl-1 ,3-dioxolane-2-acetate, diethyl 1 ,4-cyclohexanedicarboxylate, 3-methyl-2-hexen-1-yl acetate, 1-[3,3-dimethylcyclohexyl]ethyl [3- ethyl-2-oxiranyl]acetate and / or diethyl 1 ,4-cyclohexane dicarboxylate;

[0780] Green ingredients: 2-methyl-3-hexanone (E)-oxime, 2,4-dimethyl-3-cyclohexene-1-carbaldehyde, 2-tert-butyl-1 -cyclohexyl acetate, styrallyl acetate, allyl (2-methylbutoxy)acetate, 4-methyl-3-decen-5-ol, diphenyl ether, (Z)-3-hexen-1-ol and / or 1 -(5, 5-dimethyl-1-cyclohexen-1-yl)-4-penten-1 -one;

[0781] Musk ingredients: 1 ,4-dioxa-5,17-cycloheptadecanedione, (Z)-4-cyclopentadecen-1-one, 3- methylcyclopentadecanone, 1-oxa-12-cyclohexadecen-2-one, 1-oxa-13-cyclohexadecen-2-one, (9Z)-9- cycloheptadecen-1-one, 2-{(1 S)-1-[(1 R)-3,3-dimethylcyclohexyl]ethoxy}-2-oxoethyl propionate, 3-methyl- 5-cyclopentadecen-1-one, 4,6,6,7,8,8-hexamethyl-1 ,3,4,6,7,8-hexahydrocyclopenta[g]isochromene, (1 S,1 'R)-2-[1-(3',3'-dimethyl-1 '-cyclohexyl)ethoxy]-2-methylpropyl propanoate, oxacyclohexadecan-2-one and / or (1 S,1 'R)-[1-(3',3'-dimethyl-1 '-cyclohexyl)ethoxycarbonyl]methyl propanoate;

[0782] Woody ingredients: 1-[(1 RS,6SR)-2,2,6-trimethylcyclohexyl]-3-hexanol, 3,3-dimethyl-5-[(1 R)- 2,2,3-trimethyl-3-cyclopenten-1-yl]-4-penten-2-ol, 3,4'-dimethylspiro[oxirane-2,9'- tricyclo[6.2.1 ,02,7]undec[4]ene, (l-ethoxyethoxy)cyclododecane, 2,2,9, 11-tetramethylspiro[5.5]undec-8- en-1-yl acetate, 1 -(octahydro-2, 3, 8, 8-tetramethyl-2-naphtalenyl)-1 -ethanone, patchouli oil, terpenes fractions of patchouli oil, Clearwood®, (1 'R,E)-2-ethyl-4-(2',2',3'-trimethyl-3'-cyclopenten-1 '-yl)-2-buten-1- ol, 2-ethyl-4-(2,2,3-trimethyl-3-cyclopenten-1-yl)-2-buten-1-ol, methyl cedryl ketone, 5-(2,2,3-trimethyl-3- cyclopentenyl)-3-methylpentan-2-ol, 1-(2,3,8,8-tetramethyl-1 ,2,3,4,6,7,8,8a-octahydronaphthalen-2- yl)ethan-1-one and / or isobornyl acetate;

[0783] Other ingredients (e.g. amber, powdery spicy or watery): dodecahydro-3a,6,6,9a-tetramethyl- naphtho[2,1-b]furan and any of its stereoisomers, heliotropin, anisic aldehyde, eugenol, cinnamic aldehyde, clove oil, 3-(1 ,3-benzodioxol-5-yl)-2-methylpropanal, 7-methyl-2H-1 ,5-benzodioxepin-3(4H)-one, 2,5,5- trimethyl-1 ,2,3,4,4a,5,6,7-octahydro-2-naphthalenol, 1-phenylvinyl acetate, 6-methyl-7-oxa-1-thia-4- azaspiro[4.4]nonane and / or 3-(3-isopropyl-1-phenyl)butanal.

[0784] A perfumery base according to the invention may not be limited to the above mentioned perfuming coingredients, and many other of these co-ingredients are in any case listed in reference texts such as the book by S. Arctander, Perfume and Flavor Chemicals, 1969, Montclair, New Jersey, USA, or its more recent versions, or in other works of a similar nature, as well as in the abundant patent literature in the field of perfumery. It is also understood that said co-ingredients may also be compounds known to release in a controlled manner various types of perfuming compounds also known as properfume or profragrance. Nonlimiting examples of suitable properfume may include 4-(dodecylthio)-4-(2,6,6-trimethyl-2-cyclohexen-1-yl)- 2-butanone, 4-(dodecylthio)-4-(2,6,6-trimethyl-1-cyclohexen-1-yl)-2-butanone, trans-3-(dodecylthio)-1- (2,6,6-trimethyl-3-cyclohexen-1-yl)-1-butanone, 2-(dodecylthio)octan-4-one, 2-phenylethyl oxo(phenyl)acetate, 3,7-dimethylocta-2,6-dien-1-yl oxo(phenyl)acetate, (Z)-hex-3-en-1-yl oxo(phenyl)acetate, 3,7-dimethyl-2,6-octadien-1 -yl hexadecanoate, bis(3,7-dimethylocta-2,6-dien-1 -yl) succinate (2-((2-methylundec-1 -en-1 -yl)oxy)ethyl)benzene, 1 -methoxy-4-(3-methyl-4-phenethoxybut-3-en- 1 -yl)benzene, (3-methyl-4-phenethoxybut-3-en-1 -yl)benzene, 1 -(((Z)-hex-3-en-1 -yl)oxy)-2-methylundec-1 - ene, (2-((2-methylundec-1-en-1-yl)oxy)ethoxy)benzene, 2-methyl-1-(octan-3-yloxy)undec-1-ene, 1- methoxy-4-(1 -phenethoxyprop-1 -en-2-yl)benzene, 1 -methyl-4-(1 -phenethoxyprop-1 -en-2-yl)benzene, 2- (1 -phenethoxyprop-1 -en-2-yl)naphthalene, (2-phenethoxyvinyl)benzene, 2-(1 -((3,7-dimethyloct-6-en-1 - yl)oxy)prop-1-en-2-yl)naphthalene, (2-((2-pentylcyclopentylidene)methoxy)ethyl)benzene, 4-allyl-2- methoxy-1 -((2-methoxy-2-phenylvinyl)oxy)benzene, (2-((2- pentylcyclopentylidene)methoxy)ethyl)benzene, (2-((2-heptylcyclopentylidene)methoxy)ethyl)benzene, 1- isopropyl-4-methyl-2-((2-pentylcyclopentylidene)methoxy)benzene, 2-methoxy-1-((2- pentylcyclopentylidene)methoxy)-4-propylbenzene, 3-methoxy-4-((2-methoxy-2- phenylvinyl)oxy)benzaldehyde, 4-((2-(hexyloxy)-2-phenylvinyl)oxy)-3-methoxybenzaldehyde or a mixture thereof.

[0785] By “perfumery adjuvant”, it is meant here an ingredient capable of imparting additional added benefit such as a color, a particular light resistance, chemical stability, etc. A detailed description of the nature and type of adjuvant commonly used in perfuming composition cannot be exhaustive, but it has to be mentioned that said ingredients are well known to a person skilled in the art. One may cite as specific non-limiting examples the following: viscosity agents (e.g. surfactants, thickeners, gelling and / or rheology modifiers), stabilizing agents (e.g. preservatives, antioxidant, heat / light and or buffers or chelating agents, such as BHT), coloring agents (e.g. dyes and / or pigments), preservatives (e.g. antibacterial or antimicrobial or antifungal or anti irritant agents), abrasives, skin cooling agents, fixatives, insect repellants, ointments, vitamins and mixtures thereof.

[0786] It is understood that a person skilled in the art is perfectly able to design optimal formulations forthe desired effect by admixing the above-mentioned components of a perfuming composition, simply by applying the standard knowledge of the art as well as by trial and error methodologies.

[0787] An invention’s composition consisting of at least one compound of formula (I) and at least one perfumery carrier consists of a particular embodiment of the invention as well as a perfuming composition comprising at least one compound of formula (I), at least one perfumery carrier, at least one perfumery base, and optionally at least one perfumery adjuvant. According to a particular embodiment, the compositions mentioned above, comprise more than one compound of formula (I) and enable the perfumer to prepare accords or perfumes possessing the odor tonality of various compounds of the invention, creating thus new building block for creation purposes.

[0788] For the sake of clarity, it is also understood that any mixture resulting directly from a chemical synthesis, e.g. a reaction medium without an adequate purification, in which the compound of the invention would be involved as a starting, intermediate or end-product could not be considered as a perfuming composition according to the invention as far as said mixture does not provide the inventive compound in a suitable form for perfumery. Thus, unpurified reaction mixtures are generally excluded from the present invention unless otherwise specified.

[0789] The invention’s compound can also be advantageously used in all the fields of modern perfumery, i.e. fine or functional perfumery, to positively impart or modify the odor of a consumer product into which said compound (I) is added. Consequently, another object of the present invention consists of a perfumed consumer product comprising, as a perfuming ingredient, at least one compound of formula (I), as defined above.

[0790] The invention’s compound can be added as such or as part of an invention’s perfuming composition.

[0791] For the sake of clarity, “perfumed consumer product” is meant to designate a consumer product which delivers at least a pleasant perfuming effect to the surface or space to which it is applied (e.g. skin, hair, textile, or home surface). In other words, a perfumed consumer product according to the invention is a perfumed consumer product which comprises a functional formulation, as well as optionally additional benefit agents, corresponding to the desired consumer product, and an olfactive effective amount of at least one invention’s compound. For the sake of clarity, said perfumed consumer product is a non-edible product.

[0792] The nature and type of the constituents of the perfumed consumer product do not warrant a more detailed description here, which in any case would not be exhaustive, the skilled person being able to select them on the basis of his general knowledge and according to the nature and the desired effect of said product.

[0793] Non-limiting examples of suitable perfumed consumer products include a perfume, such as a fine perfume, a splash or eau de parfum, a cologne or a shave or after-shave lotion; a fabric care product, such as a liquid or solid detergent, a fabric softener, a liquid or solid scent booster, a fabric refresher, an ironing water, a paper, a bleach, a carpet cleaner, a curtain-care product; a body-care product, such as a hair care product (e.g. a shampoo, a coloring preparation or a hairspray, a color-care product, a hair shaping product, a dental care product), a disinfectant, an intimate care product; a cosmetic preparation (e.g. a skin cream or lotion, a vanishing cream or a deodorant or antiperspirant (e.g. a spray or roll on), a hair remover, a tanning or sun or after sun product, a nail product, a skin cleansing, a makeup); or a skin-care product (e.g. a soap, a shower or bath mousse, oil or gel, or a hygiene product or a foot / hand care products); an air care product, such as an air freshener or a “ready to use” powdered air freshener which can be used in the home space (rooms, refrigerators, cupboards, shoes or car) and / or in a public space (halls, hotels, malls, etc..); or a home care product, such as a mold remover, a furnisher care product, a wipe, a dish detergent or a hard-surface (e.g. a floor, bath, sanitary or a window-cleaning) detergent; a leather care product; a car care product, such as a polish, a wax or a plastic cleaner.

[0794] Some of the above-mentioned perfumed consumer products may represent an aggressive medium for the invention’s compounds, so that it may be necessary to protect the latter from premature decomposition, for example by encapsulation or by chemically binding it to another chemical which is suitable to release the invention’s ingredient upon a suitable external stimulus, such as an enzyme, light, heat or a change of pH.

[0795] The proportions in which the compounds according to the invention can be incorporated into the various aforementioned products or compositions vary within a wide range of values. These values are dependent on the nature of the article to be perfumed and on the desired organoleptic effect as well as on the nature of the co-ingredients in a given base when the compounds according to the invention are mixed with perfuming co-ingredients, solvents or additives commonly used in the art.

[0796] For example, in the case of perfuming compositions, typical concentrations are in the order of 0.001 % to 10 % by weight, or even more, of the compounds of the invention based on the weight of the composition into which they are incorporated. In the case of perfumed consumer product, typical concentrations are in the order of 0.01 % to 1 % by weight, or even more, of the compounds of the invention based on the weight of the consumer product into which they are incorporated.

[0797] Additionally, the intermediates compounds produced in any of the embodiments described herein can be converted to derivatives such as, but not limited to hydrocarbons, alcohols, diols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides and / or an esters. These derivatives can be obtained by a chemical method such as, but not limited to oxidation, reduction, alkylation, acylation and / or rearrangement. Alternatively, the derivatives can be obtained using a biochemical method by contacting the terpene compound with an enzyme such as, but not limited to an oxidoreductase, a monooxygenase, a dioxygenase, a transferase. The biochemical conversion can be performed in-vitro using isolated enzymes, enzymes from lysed cells or in vivo using whole cells. The conversion can be a cyclization reaction realized by chemical or biochemical method. The derivatives can be used as perfumery, flavor or aroma ingredients.

[0798] The numerous possible variations that will become immediately evident to a person skilled in the art after having considered the disclosure provided herein also fall within the scope of the invention.

[0799] The invention will now be described in further details by way of the following Examples. Said Examples are illustrative only and are not intended to limit the scope of the embodiments as described herein. EXAMPLES

[0800] Materials and Methods

[0801] Unless otherwise stated, all chemical and biochemical materials and microorganisms or cells employed herein are commercially available products.

[0802] Unless otherwise specified, recombinant proteins are cloned and expressed by standard methods, such as, for example, as described by Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular cloning: A Laboratory Manual, 2ndEdition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.

[0803] General methods for Escherichia coli, genetic modification, cultivation and compound analysis

[0804] Recombinant E. coli for production of terpenoid precursors

[0805] An E. coli strain was engineered to produce farnesyl-pyrophosphate (FPP) by chromosomal integration of recombinant genes encoding mevalonate pathway enzymes.

[0806] An upper pathway operon (operon 1 from acetyl-CoA to mevalonate) was designed consisting of the atoB gene from E. coli encoding an acetoacetyl-CoA thiolase, and the mvaA and mvaS genes from Staphylococcus aureus encoding a HMG-CoA synthase and a HMG-CoA reductase, respectively.

[0807] As a lower mevalonate pathway operon (operon 2 from mevalonate to IPP / DMAPP), a natural operon from the gram-negative bacteria Streptococcus pneumoniae was selected, encoding a mevalonate kinase (mvaK1), a phosphomevalonate kinase (mvaK2), a phosphomevalonate decarboxylase (mvaD), and an isopentenyl diphosphate isomerase (fni).

[0808] A codon optimized Saccharomyces cerevisiae FPP synthase encoding gene (ERG20) was introduced at the 3’-end of the upper pathway operon to convert isopentenyl-diphosphate (IPP) and dimethylallyldiphosphate (DMAPP) into FPP.

[0809] The above-described operons were synthesized by DNA 2.0 and integrated into the araA gene of the Escherichia coli strain BL21 (DE3). The heterologous pathway was introduced in two separate recombination steps using the CRISPR / Cas9 genome engineering system. The first operon (lower pathway; operon 2) to be integrated carries a spectinomycin (Spec) marker which was used to screen for Spec resistant candidate integrants. The second operon was designed to displace the Spec marker of the previously integrated operon and was accordingly screened for Spec candidate integrants following the second recombination event. Guide RNA expression vectors targeting the araA gene were designed and synthetized by DNA 2.0. PCR was used to verify operon integration by designing PCR primers to amplify across the araA gene integration target and across recombination junctions of integrants. One clone yielding correct PCR results was then fully sequenced and archived as strain DP1205. Medium composition for E. coli cultivation

[0810] The mineral AM medium used in the shake flask and lab-scale fermentation experiments consists of : KH2PO44.2 g / L; K2HPO4 • 3H2O 15.7 g / L; (NH4)2SO42.0 g / L; Citric acid 1 .7 g / L; EDTA 8.4 mg / L; glycerol 30 g / L; yeast extract 5 g / L were dissolved in diH2O; dodecane at 10% (v / v) and sterilized at 121 °C for 30 min. Concentrated stocks of MgSO4 • 7H2O 1 M, 5 mL / L; vitamin (thiamine • HCI 4.5 g / L), 1 mL / L; and batch trace metal solution; 10 mL / L were added aseptically to the medium and the pH was adjusted at 7 with NaOH 5M. Batch trace metal solution in (per L 1 M HCL): CoCI2 • 6H2O 0.25 g / L; MnCI2 • 4H2O 1 .5 g / L; CuCI2 • 2H20 0.15 g / L; H3BO3 0.3 g / L; Na2MoC4 • 2H2O 0.25 g / L; Zn(CHCOO)2 -2H2O 1.3 g / L; Fe(lll)citrate 10 g / L. For the fermentations carried out in feed-batch mode with the AM medium, 20 L glycerol feed solution containing 700 g / L glycerol, 12 g / L MgSO4 • 7H2O; 13 mg / L EDTA and 10 mL / L feed trace solution was prepared. The feed trace metal solution was prepared by dissolving 0.4 g CoCI2 • 6H2O, 2.35 g MnCI2 •4H2O, 0.25 g CuCI2 • 2H20, 0.5 g H3BO3; 0.4 g Na2MoC4 • 2H2O; 1 .6 g Zn(CHCOO)2 «2H2O and 10 g Fe(lll)citrate • H2O in 1 L HCI 1 M.

[0811] Cultivation of engineered bacteria cells under conditions enabling production of terpene compounds

[0812] The DP1205 E. coli cells engineered to produce increased levels of the terpenoid precursor farnesyl diphosphate (FPP) were transformed with one or two expression plasmids carrying genes encoding for enzymes from a homofarnesol biosynthetic pathway enzymes and / or for terpene cyclases. The transformed cells were cultured with the appropriate antibiotics (kanamycin (50 pg / mL) and / or carbenicillin (50 pg / mL) and / or chloramphenicol (34 pg / mL)) and / or streptomycin (50 pg / mL)) on LB-agarose plates. Single colonies were used to inoculate 5 mL liquid LB medium supplemented with the same antibiotics, 4 g / L glucose and 10% (v / v) n-dodecane. The next day, 2 mL of AM medium supplemented with the same antibiotics and 10% (v / v) n-dodecane were inoculated with 0.2 mL of the overnight culture. The cultures were incubated at 37°C until an optical density of 3 was reached. The expression of the recombinant proteins was then induced by addition of 0.1 mM IPTG and the cultures were incubated for 72 h at 25°C.

[0813] The cultures were then extracted with one volume of methyl tert-butyl ether (MTBE) and the composition of the organic phase was analyzed by GC-MS as described below. For quantification, an internal standard (a-longipinene (Sigma-Aldrich, Missouri, USA)) was added to the extract prior to GC-MS analysis and concentrations of the components were estimated based on comparison of the peak areas.

[0814] GC-MS analysis methods

[0815] Samples were analyzed using an Agilent 6890N GC system coupled with a 5975B series Mass Selective Detector (MSD) and equipped with a split / splitless injector (Agilent Technologies, CA) and a CombiPAL autosampler (PAL LSI 85 autosampler, Agilent Technologies, CA) injection system. The GC inlet temperature was set to 240°C and 1 .0 pL of sample was injected in split mode with a ratio of 25:1 (23.304 PSI) and analyzed on a DB-5ms capillary column (30 m x 0.25 mm inner diameter x 0.25 pm film thickness; Agilent J&W) using helium as a carrier gas at a constant flow of 1 .2 mL / min. The initial temperature of the oven was set at 80°C (hold 1 min) and was programmed to 300°C (10°C / min) and then to 300°C (30°C / min; hold 1 min).

[0816] General methods for Saccharomyces cerevisiae, genetic modification, cultivation and compound analysis.

[0817] A Saccharomyces cerevisiae strain producing increased levels of the terpenoid precursor farnesyl diphosphate (FPP) (as described in WO2018 / 114839) was used as base strain for the expression of homofarnesol biosynthetic pathway genes and terpene cyclases. In short, the strain contains all the endogenous mevalonate pathway genes integrated in its genome under the control of the native GAL1 or GAL10 promoters. Further increase of the Farnesyl diphosphate precursor pool in this strain was achieved through downregulation of the squalene synthase gene (ERG9) via replacement of its native promoter.

[0818] All genes (synthesized by ATUM, California, USA or Twist Bioscience, California, USA), together with relevant regulatory elements (eg. promoters and terminators) were introduced in the base strain through genomic integration or through plasmids constructed in vivo using the yeast homologous recombination machinery (Kuijpers et al., Microb Cell Fact., 2013, 12:47). All yeast transformations were performed with the lithium acetate method (Gietz and Woods, Methods Enzymol., 2002, 350:87-96).

[0819] Successfully transformed yeast colonies were grown for three days at 30 °C on media containing 6.7 g / L of Yeast Nitrogen Base without amino acids (BD Difco, New Jersey, USA), the appropriate antibiotic or nutrients according to the marker gene used, 20 g / L glucose and 20 g / L agar.

[0820] For metabolite production and analysis, single colonies of the modified yeast strains were inoculated in 2 mL of culture medium (Westfall et al., Proc Natl Acad Sci USA, 2012, 109:E111-118) with addition of 2% galactose and 10% (v / v) n-dodecane (Sigma-Aldrich, Missouri, USA). The cultures were incubated for 3 days at 30°C and shaking at 200 rpm. After the incubation period, the cultures were extracted with two volumes of MTBE (supplemented with a-longipinene standard for quantification as described above) and the composition of the organic phase was analyzed by GC-MS using an Agilent 7890A GC system coupled with a 5975C series Mass Selective Detector (MSD) and equipped with a split / splitless injector and a GC Injector 80 injection system (Agilent Technologies, CA). The GC inlet temperature was set to 260°C and 1 .0 pl of sample was injected in splitless mode and analyzed on a HP-5 GC column (30 m x 0.25 mm x 0.25 pm; Agilent J&W) using helium as a carrier gas at a constant flow of 1 .2 mL / min. The initial temperature of the oven was set at 100°C and was programmed to 300°C (10°C / min). Example 1. In vivo production of (3E,7E)-homofarnesol and biosynthetic intermediates in engineered bacterial cells expressing a GGPP synthase, a phosphatase, an alcohol dehydrogenase, a BVMO, an enal-cleaving enzyme and an esterase.

[0821] The reaction scheme in Figure 3 depicts a biochemical pathway that can be used to produce (3E,7E)- homofarnesol in vivo. The universal isoprenoid precursors isopentenyl-diphosphate (IPP) and dimethylallyldiphosphate (DMAPP) are condensed to form (2E,6E,10E)-geranylgeranyl-diphosphate (GGPP). This reaction can be catalyzed by a GGPP synthase or using a combination of a farnesyl-diphosphate synthase (FPP synthase) and a GGPP synthase. GGPP can be converted to (2E,6E,10E)-geranylgeraniol by a terpene synthase or by a phosphatase as for example described in W02020011883A1. (2E,6E,10E)- geranylgeraniol undergoes then several enzymatic degradations steps. In the proposed pathway, (2E,6E,10E)-geranylgeraniol is first oxidized by an alcohol dehydrogenase (ADH) and cleaved to yield (5E,9E)-farnesylacetone. The following cleavage reaction can be catalyzed by an enal-cleaving enzyme (ENase) such as a protein containing a GXWXG (SEQ ID NO: 263) and DUF4334 domain as described in WO2021005097. (5E,9E)-farnesylacetone is further converted by a Baeyer-Villiger monooxygenase (BVMO) to (3E,7E)-homofarnesyl acetate. In the last step, the ester is hydrolyzed by an esterase to finally form (3E,7E)-homofarnesol.

[0822] To validate the (3E,7E)-homofarnesol pathway, E. coli cells were engineered to express the necessary enzymes. A plasmid was assembled to contain two operons.

[0823] The first operon was designed to contain three cDNAs encoding for:

[0824] - PsAerADH (SEQ ID NO: 11), an alcohol dehydrogenase from Pseudomonas aeruginosa (GeneBank accession number: WP_079868259.1) having the ability to oxidize (2E,6E,10E)-geranylgeraniol to (2E,6E,10E)-geranylgeranial,

[0825] - SCH24-BVMO1 (SEQ ID NO: 23), a Baeyer-Villiger Monooxygenase (BVMO) from Filobasidium magnum and described in WO2021005097, which allow the oxidation of (5E,9E)-farnesylacetone to (3E,7E)- homofarnesyl acetate, and

[0826] - SCH24-EST1 (SEQ ID NO: 27), an esterase from Filobasidium magnum and described in WO2021005097, which hydrolyses (3E,7E)-homofarnesylacetate to (3E,7E)-homofarnesol and acetic acid.

[0827] The second operon was constructed to contain 4 cDNAs encoding for:

[0828] - SCH94-03944 (SEQ ID NO: 22), a protein containing an enal-cleaving enzyme from Rhodococcus erythropolis and described in WO2021005097, which cleaves (2E,6E,10E)-geranylgeranial into (5E,9E)- farnesylacetone and acetaldehyde,

[0829] - CcrGGPPS2-del57, a truncated version of the (2E,6E,10E)-geranylgeranyl diphosphate synthase from Cistus creticus (GeneBank: AAM21639.1) (SEQ ID NO: 1), and

[0830] -Two copies of a cDNAs encoding for PgpB (SEQ ID NO: 3), a phosphatase from Escherichia coli (GeneBank: WP_089622241 .1), which converts (2E,6E,10E)-geranylgeranyl diphosphate into (2E,6E,10E)-geranylgeraniol by cleaving of the diphosphate group. The cDNAs encoding for PsAerADH, SCH24-BVMO1 , SCH24-EST1 , SCH94-03944, CcrGGPPS2-del57 and PgpB were codon optimized for Escherichia coli (SEQ ID NOs:102, 115, 120, 113, 90 and 93) and an RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 264) was placed upstream of each of these cDNAs. The first operon containing the cDNAs for PsAerADH, SCH24-BVMO1 and SCH24-EST1 were under the control of a T5 promotor and rrnB T1 terminator. The second operon containing the cDNA for SCH94- 03944, CcrGGPPS2-del57 and the two cDNAs for PgpB was under the control of a T5 promotor and rrnB terminator. Both operons were synthesized and cloned into a vector backbone containing a pUC origin of replication, a kanamycin resistance gene and a Lacl gene resulting in the vector pHFOL-5.

[0831] The farnesyl-diphosphate (FPP) producing E. coli strain DP1205, described in WO2021005097, was transformed with the vector pHFOL-5 described above. When cultivated in the conditions enabling production of terpene compounds, the resulting cells were capable of producing (3E,7E)-homofarnesol (Figure 4). Under the conditions described in the “materials and methods” section, 69 mg / L of (3E,7E)- homofarnesol were produced in the tube assay in the culture media.

[0832] The product profile of the cells (Figure 4) also shows the accumulation of several metabolic intermediates, such as (5E,9E)-farnesylacetone and (2E,6E,10E)-geranylgeraniol. The different enzymatic steps of the pathway can be optimized to limit the accumulation of intermediates and increase the concentration of the final product.

[0833] The following examples show methods of identifying suitable enzymes for each of the enzymatic step in the pathway to increase (3E,7E)-homofarnesol production and limit the accumulation of the metabolic intermediates.

[0834] Example 2. Screening of Baeyer-Villiger monooxygenases to improve the enzymatic conversion of (5E,9E)-farnesylacetone to (3E,7E)-homofarnesylacetate and the in vivo production of (3E,7E)- homofarnesol.

[0835] Example 1 shows that for the in vivo production of (3E,7E)-homofarnesol, (5E,9E)-farnesylacetone can also be detected due to the insufficient activity of the Baeyer-Villiger monooxygenase (SCH24-BVMO1 (SEQ ID NO: 23)) in this strain.

[0836] In this example, an in vivo screening of different BVMOs was conducted to identify enzyme candidates with higher efficiency compared to SCH24-BVMO1 (SEQ ID NO: 23). For the screening, a modified version of the vector pHFOL-5 (described in example 1) was created by removing SCH24-BVMO1 . The new vector was called pF-Facetone-7. The E. coli strain DP1205 was transformed with the vector pF-Facetone-7, providing a strain that was capable of producing (5E,9E)-farnesylacetone when cultivated under conditions enabling the production of terpene compounds. Up to 410 mg / L of (5E,9E)-farnesylacetone were produced in the culture media in the tube assay (Figure 5A). When the cells are further transformed with a vector expressing an active BVMO, (5E,9E)-farnesylacetone is converted to (3E,7E)-homofarnesyl acetate, which itself is converted to (3E,7E)-homofanesol by the SCH24-EST1 esterase (Figure 5B).

[0837] In the next step, codon optimized cDNAs encoding for BVMOs were designed and cloned in the pJ423 expression plasmid (ATUM, Newark, California). The DP1205 E. coli cells were co-transformed with one of these plasmids and plasmid pF-Facetone-7. The activity of the BVMOs was determined by quantifying the production of (3E,7E)-homofarnesol for each BVMO tested and compared to SCH24-BVMO1 (SEQ ID NO: 23).

[0838] The table below (Table 1) shows the relative activity of (5E,9E)-farnesylacetone conversion of some BVMOs identified in this screening.

[0839] Table 1 : Selected BVMOs and relative activity of conversion of (5E,9E)-farnesylacetone to (3E,7E)- homofarnesyl acetate, which was further converted to (3E,7E)-homofarnesol.

[0840] AraBVMOI and AflavBVMOI were found to produce significantly higher amounts of (3E,7E)-homofarnesol than SCH24-BVMO1 (SEQ ID NO: 23). AflavBVMOI (SEQ ID NO: 26) increases the production of (3E,7E)- homofarnesol by 59 % compared to the reference BVMO.

[0841] Example 3. In vivo screening of alcohol dehydrogenases to improve the efficiency of the enzymatic conversion of (2E,6E,10E)-geranylgeraniol to (2E,6E,10E)-geranylgeranial and the in vivo production of (3E,7E)-homofarnesol.

[0842] In this Example, alcohol dehydrogenases (ADHs) were tested in vivo for their efficiency to oxidize (2E,6E,10E)-geranylgeraniol to (2E,6E,10E)-geranylgeranial. Alcohol dehydrogenases catalyze the reversible oxidation of alcohols.

[0843] To avoid the reverse alcohol-dehydrogenase reaction in this in vivo screening assay, the enzyme enal- cleaving enzyme SCH94-03944 described in Example 1 was co-expressed in the E. coli cells to enzymatically convert the (2E,6E,10E)-geranylgeranial form to (5E,9E)-farnesylacetone. The catalytic efficiency of the ADHs tested was then correlated to the amount of (2E,6E,10E)-geranylgeraniol being converted to (5E,9E)-farnesylacetone. The ADH candidates were codon optimized and cloned into the pJ423 expression plasmid (ATUM, Newark, California). The DP1205 E. coli cells were co-transformed with one of these plasmids and with the plasmid pJ401-SCH94-3944-PgpB-CcrGGPPS, which contained the necessary genes to produce (5E,9E)-farnesylacetone as described below, except the ADH encoding gene.

[0844] Thus, the plasmid pJ401-SCH94-3944-PgpB-CcrGGPPS2-del57, contained an operon harboring 3 cDNAs that encoded for:

[0845] - The enal-cleaving enzyme SCH94-03944 (SEQ ID NO: 22), a protein containing a GXWXG (SEQ ID NO: 263) and DUF4334 domain described in WO2021005097, which cleaves (2E,6E,10E)-geranylgeranial into (5E,9E)-farnesylacetone and acetaldehyde

[0846] - PgpB (SEQ ID NO: 3), a phosphatase from Escherichia coli (GeneBank: WP_089622241 .1), which converts (2E,6E,10E)-geranylgeranyl diphosphate into (2E,6E,10E)-geranylgeraniol by cleaving off diphosphate, and

[0847] - CcrGGPPS2-del57 (SEQ ID NO: 1), a truncated version of the (2E,6E,10E)-geranylgeranyl diphosphate synthase from Cistus criticus (GeneBank: AAM21639.1).

[0848] The cDNAs encoding for SCH94-03944, PgpB and CcrGGPPS2-del57 were codon optimized for the expression in Escherichia coli (SEQ ID NOs: 113, 93 and 90). An operon was designed containing successively the three cDNAs and an RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 264) placed upstream of each cDNA. The operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California).

[0849] The resulting strains (DP1205 containing the plasmid pJ401-SCH94-3944-PgpB-CcrGGPPS2-del57 and a plasmid pJ423 harboring an alcohol dehydrogenase candidate) were cultivated under conditions enabling the production of terpenes. The amounts of (2E,6E,10E)-geranylgeraniol and (5E,9E)-farnesylacetone were measured and the conversion rate was calculated for each alcohol dehydrogenase tested.

[0850]

[0851] Table 2: ADHs selected in vivo for their activity on (2E,6E,10E)-geranylgeraniol to produce (2E,6E,10E)- geranylgeranial, which was further converted to (5E,9E)-farnesylacetone.

[0852] The results are shown in Table 2. The highest conversion rates from (2E,6E,10E)-geranylgeraniol were detected in the culture medium in the tube assay for the alcohol dehydrogenases ThTerpADHI (SEQ ID NO: 12), Ppseudo-alkJ (SEQ ID NO: 15) and CymB (SEQ ID NO: 17). The highest amount of (5E.9E)- farnesylacetone was 172 mg / L and was produced with the ADH Ppseudo-alkJ (SEQ ID NO: 15).

[0853] Example 4. In vivo testing of esterases to catalyze the hydrolysis of (3E,7E)-homofarnesyl acetate to (3E,7E)-homofarnesol and acetic acid.

[0854] In this example, different esterases were tested in vivo fortheir ability to catalyze the hydrolysis of (3E,7E)- homofarnesyl acetate to (3E,7E)-homofarnesol and acetic acid.

[0855] The esterases were tested in a two-plasmid system in E. coli DP1205. The first plasmid consisted of an operon containing 3 cDNAs encoding for:

[0856] - The enal-cleaving enzyme SCH94-03944 (SEQ ID NO: 22), a protein containing a GXWXG (SEQ ID NO: 263) and DUF4334 domain described in W02021 / 005097, which cleaves (2E,6E,10E)-geranylgeranial into (5E,9E)-farnesylacetone and acetaldehyde,

[0857] - PgpB (SEQ ID NO: 3), a phosphatase from Escherichia coli (GeneBank: WP_089622241 .1), which converts (2E,6E,10E)-geranylgeranyl diphosphate into (2E,6E,10E)-Geranylgeraniol by cleaving off diphosphate, and

[0858] - CcrGGPPS2-del57 (SEQ ID NO: 1), a truncated version of the (2E,6E,10E)-geranylgeranyl diphosphate synthase from Cistus criticus (GeneBank: AAM21639.1).

[0859] The cDNAs encoding for SCH94-03944, PgpB and CcrGGPPS2-del57 were codon optimized for the expression in E. coli (SEQ ID NOs: 113, 93 and 90) and contained an upstream placed RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 264) The operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California).

[0860] The second operon consisted of 3 cDNAs encoding for:

[0861] - PsAerADH, an alcohol dehydrogenase from Pseudomonas aeruginosa (GeneBank: WP_079868259.1) (SEQ ID NO: 11) having the ability to oxidize (2E,6E,10E)-Geranylgeraniol to (2E,6E,10E)-geranylgeranial,

[0862] - SCH24-BVMO1 (SEQ ID NO: 23), a Baeyer-Villiger Monooxygenase (BVMO) described in WO2021005097, which allows the oxidation of (5E,9E)-farnesylacetone to (3E,7E)-homofarnesyl acetate, and

[0863] - an esterase gene candidate.

[0864] The cDNAs encoding for PsAerADH, SCH24-BVMO1 and an esterase gene candidate were codon optimized for the expression in Escherichia coli (SEQ ID NOs: 102 and 115) and contained an upstream RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 264). The operons were synthesized and cloned into the pJ424 expression plasmid (ATUM, Newark, California).

[0865] The DP1205 E. coli cells were co-transformed with the plasmids pJ401-SCH94-03944-PgpB-CcrGGPPS2- del57 and pJ424-PsAerADH-SCH24-BVMO1 -Esterase. In the resulting strains, the esterases SCH24- EST1 (SEQ ID NO: 27) from Filobasidium magnum and SCH23-EST1 (SEQ ID NO: 28) from Hyphozyma roseonigra were found to be able to convert (3E,7E)-homofarnesyl acetate to (3E,7E)-homofarnesol.

[0866] Example 5. In vivo screening of different phosphatases to catalyze the hydrolysis from (2E,6E,10E)- geranylgeranyl-disphosphate to (2E,6E,10E)-geranylgeraniol.

[0867] In this example, phosphatases from different protein families were tested in vivo for their ability to catalyze the hydrolysis of (2E,6E,10E)-geranylgeranyl diphosphate to (2E,6E,10E)-geranylgeraniol and diphosphate. The phosphatases were tested in DP1205 E. coli cells, which contained a geranylgeranyldiphosphate synthase in an expression vector. Codon optimized versions of the DNA fragments coding for the phosphatase genes, each cloned into a second expression vector, were then introduced into this strain by transformation. Under terpene production enabling conditions, we identified 8 phosphatases showing activity on (2E,6E,10E)-geranylgeranyl-diphosphate and production of (2E,6E,10E)-geranylgeraniol. The results are shown in Table 3. The phosphatases PgpB (SEQ ID NO: 3), PeSubTPPI (SEQ ID NO: 7) and TalVeTPP (SEQ ID NO: 8) showed the highest activity. PgpB (SEQ ID NO: 3) was found to be the most efficient enzyme for the production of (2E,6E,10E)-geranylgeraniol.

[0868] Table 3: Phosphatases selected in vivo for their activity to convert (2E,6E,10E)-geranylgeraniol diphosphate to (2E,6E,10E)-geranylgeraniol. The phosphatases were ranked and divided into 4 groups (- No activity; + Minor activity, ++ High activity; +++ very high activity) based on the productivity.

[0869] Example 6. Production of (3E,7E)-homofarnesol in engineered fungal cells.

[0870] The (3E,7E)-homofarnesol biosynthetic pathway (Figure 3) genes were introduced in a Saccharomyces cerevisiae strain producing high levels of the terpenoid precursor farnesyl diphosphate (FPP) (as described herein above) in a two-step process. First, single copies of the genes coding for the alcohol dehydrogenase SCH23-ADH1 (SEQ ID NO: 21) (WO2021005097), the esterase SCH23-EST1 (SEQ ID NO: 28) (WO2021005097) and the Baeyer-Villiger monooxygenase AflavBVMO (SEQ ID NO: 26) were integrated in the genome of the yeast strain. The resulting strain, YST403, was subsequently used for the in vivo construction of a 2-micron plasmid containing the geranylgeranyl diphosphate synthase CarG (SEQ ID NO: 2) (from Blakeslea trispora, NCBI accession JQ289995.1), the phosphatase PgpB (SEQ ID NO: 3) (from Escherichia coli, NCBI accession WP_089622241 .1) and the enal-cleaving enzyme SCH94-03944 (SEQ ID NO: 22). All genes coding for the different enzymes were codon optimized for their expression in S. cerevisiae (SEQ ID NOs: 112, 122, 119, 92, 94 and 1 14) and controlled by galactose inducible promoters. After cultivation, it was confirmed that the final strain, termed YST403_HFOL, produced (3E,7E)- homofarnesol. The product profile (Figure 6) also shows the accumulation of several metabolic intermediate such as of (5E,9E)-farnesylacetone (44 mg / L) and (2E,6E,10E)-geranylgeraniol (52 mg / L). The pathway can be optimized by increasing the enzymatic activity of each enzymatic step in order to limit the accumulation of intermediates and thereby to increase the concentration of the final product. Example 7. In vivo production of compound of formula (la) and biosynthetic intermediates in fungal cells engineered to produce (3E,7E)-homofarnesol and expressing different wild type or mutant squalene cyclases.

[0871] In this example, compound of formula (la) is being produced in vivo in fungal cells by using the artificial biochemical pathway shown in Figure 7. This pathway encompasses the production of (3E,7E)- homofarnesol from a simple carbon source and its conversion to compound of formula (la) by a squalene cyclase enzyme according to the invention.

[0872] Squalene cyclase candidate enzymes were selected from NCBI, UniProt and Ocean Microbiomics databases. They belong to the squalene cyclase family (InterPro ID IPR018333) and are annotated in the different databases as: BmeSHC (SEQ ID NO: 315), BsuTC (SEQ ID NO: 317), CttSHC (SEQ ID NO: 319), SCHJC_C003990 (SEQ ID NO: 321), and TARA_SAMEA2623601_METAG_s18_g109 (SEQ ID NO: 322). Additionally, mutants of seleted candidate enzymes were generated. For this, the amino acid positions relative to F437 and G600 in the squalene cyclase from Alicyclobacillus acidocaldarius (AAcSHC, SEQ ID NO: 82) were selected and, if appropriate, modified to alanine and methionine, respectively. This process led to the mutants BmeSHC_V1 (SEQ ID: 316), BsuTC_V1 (SEQ ID: 318), CttSHC_V1 (SEQ ID: 320), and TARA_SAMEA2623601_METAG_s18_g109_V1 (SEQ ID: 323).

[0873] The squalene cyclase candidate enzymes (SEQ ID NO: 315, 316, 317, 318, 320, 321 , 323) were tested for (3E,7E)-homofarnesol cyclization and compound of formula (la) production in vivo in fungal cells. For this, the DNA encoding the squalene cyclase candidate enzymes were ordered codon optimized for expression in S. cerevisie (SEQ ID NOs: 326, 327, 328, 329, 331 , 332, 334) and were introduced in the strain YST403_HFOL as described herein above in a plasmid containing genes encoding CarG (SEQ ID NO: 2), PgpB (SEQ ID NO: 3) and SCH94-03944 (SEQ ID NO: 22).

[0874] As shown in Figure 8 (wild-type enzymes) and Figure 9 (mutant enzymes), the tested squalene cyclase candidate enzymes were all shown to catalyze the conversion of (3E,7E)-homofarnesol to the compound of formula (la). Interestingly, the mutations at amino acid positions relative to F437 and G600 in AAcSHC improved the production of the compound of formula (la). For example, as shown in Figure 10, a 10-fold improvement was observed when the mutant variant BmeSHC_V1 was expressed in comparison to the wild-type BmeSCH.

Claims

CLAIMS1 . A process for the preparation of a compound of formula (I)in the form of any one of its stereoisomers or a mixture thereof, comprising: contacting a compound of formula (II)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having alcohol dehydrogenase (ADH) enzyme activity to produce a compound of formula (III);(ii) contacting the compound of formula (III)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having enal- cleaving enzyme activity to produce a compound of formula (IV);(iii) contacting the compound of formula (IV)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having Baeyer-Villiger monooxygenase (BVMO) enzyme activity to produce a compound of formula(V);(iv) contacting the compound of formula (V)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having esterase enzyme activity to produce a compound of formula (VI); and(v) contacting the compound of formula (VI)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having terpene cyclase enzyme activity to produce a compound of formula (I), wherein the polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323.

2. A process for the preparation of a compound of formula (I)in the form of any one of its stereoisomers or a mixture thereof, comprising contacting a compound of formula (VI)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having terpene cyclase enzyme activity to produce a compound of formula (I), wherein the polypeptide having terpene cyclase enzyme activity has at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of SEQ ID NOs: 322 and 323; and optionally, further comprising one or more steps selected from:(i) contacting the compound of formula (II)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having alcohol dehydrogenase (ADH) enzyme activity to produce a compound of formula (III);(ii) contacting the compound of formula (III)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having enal- cleaving enzyme activity to produce a compound of formula (IV);(iii) contacting the compound of formula (IV)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having Baeyer-Villiger monooxygenase (BVMO) enzyme activity to produce a compound of formula (V); and(iv) contacting the compound of formula (V)in the form of any one of its stereoisomers or a mixture thereof, with a polypeptide having esterase enzyme activity to produce a compound of formula (VI).

3. The process of any of the previous claims, wherein the polypeptide having terpene cyclase enzyme activity is a squalene cyclase enzyme.

4. The process of any of the previous claims, wherein more than 97% of the compound of formula (I) is in the form of formula (la):(formula la) and / or formula (lb):116(formula lb).

5. The process of any of the previous claims, wherein:(i) the compound of formula (I) is in the form of formula (la):(formula la)(ii) the compound of formula (II) is in the form of formula (Ila):(formula Ila);(iii) the compound of formula (III) is in the form of formula (Illa):(formula Illa);(iv) the compound of formula (IV) is in the form of formula (IVa):(formula IVa);(v) the compound of formula (V) is in the form of formula (Va):(formula Va);(vi) the compound of formula (VI) is in the form of formula (Via):(formula Via).1176. The process of any of the previous claims, wherein:(i) the polypeptide having ADH enzyme activity comprises at least one or more motifs selected from CHTD (SEQ ID NO: 228), GHEGxG (SEQ ID NO: 229), LxCGxxTGxGA (SEQ ID NO: 230), Gx[VI]GL (SEQ ID NO: 231), LxxxG[LVI][PA] (SEQ ID NO: 232), GxVxAl (SEQ ID NO: 233) and YxATKxA (SEQ ID NO: 234); wherein, residues x represent independently of each other any natural amino acid residue;(ii) the polypeptide having enal-cleaving enzyme activity is selected from the group of polypeptides containing:(a) at least one DUF4334 protein family domain having the Pfam ID number PF14232,(b) at least one GXWXG (SEQ ID NO: 263) protein family domain having the Pfam ID number PF142311, and / or(c) a domain retaining at least 90% sequence identity to PF14232 or PF14231 ;(iii) the polypeptide having BVMO enzyme activity is selected from the group of polypeptides comprising:(a) a flavin-containing monooxygenase (FMO) protein family domain having the Pfam ID number PF00743 within their amino acid sequence or a domain retaining at least 90%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to PF00743, and / or(b) at least one or more of motifs selected from:. GxGxxG (SEQ ID NO: 239),. [GS]GxWxxxxYPGxxxD (SEQ ID NO: 240),. Gxxx[FY]xGxxx[HS]xxxW (SEQ ID NO: 241), and. [KQ]x[VI]xx[IV]GxG (SEQ ID NO: 242); wherein, residues x represent independently of each other any natural amino acid residue; and / or,(iv) the polypeptide having esterase enzyme activity comprises at least one or more motifs selected from AxWxVxxRLAPE (SEQ ID NO: 243), GASAGGGLxA (SEQ ID NO: 244), VxQLLxYPMLDDR (SEQ ID NO: 245) and ARxxDLSGLPxT (SEQ ID NO: 246); wherein, residues x represent independently of each other any natural amino acid residue.

7. The process of any of the previous claims, wherein the process further comprises one or more steps selected from:1) preparing geranylgeranyl-diphosphate (GGPP) from isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP) using one or more polypeptides having prenyltransferase enzyme activity; and,2) preparing compound of formula (II) from GGPP using one or more polypeptides having phosphatase enzyme activity.

8. The process of any of the previous claims, wherein the process is an in vivo or a bioconversion process.1189. The process of any of the previous claims, wherein the process comprises growing a recombinant cell capable of functionally expressing the polypeptide having terpene cyclase enzyme activity to produce compound of formula (I).

10. A recombinant cell producing a compound of formula (I), wherein the recombinant cell comprises (i) a polypeptide having ADH enzyme activity, (ii) a polypeptide having enal-cleaving enzyme activity, (iii) a polypeptide having BVMO enzyme activity, (iv) a polypeptide having esterase enzyme activity, and (v) a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323.

11. A recombinant cell producing a compound of formula (I), wherein the recombinant cell comprises a polypeptide having terpene cyclase enzyme activity, wherein said polypeptide having terpene cyclase enzyme activity has at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323.

12. The recombinant cell of any of claims 10 and 11 , further comprising a polypeptide having prenyltransferase enzyme activity and / or a polypeptide having phosphatase enzyme activity.

13. A cell culture fermentation medium comprising the recombinant cell of any of claims 10 to 12.

14. The process of claim 9, the recombinant cell of any of claims 10 to 12, the cell culture fermentation medium of claim 13, wherein the recombinant cell is a bacterial cell, a plant cell, a fungal cell such as a yeast cell; preferably, the recombinant cell is of the genus Escherichia, Saccharomyces, Yarrowia or Pichia.

15. A reaction mixture comprising a compound of formula (I) and one or more compounds selected from compounds of formula (II), formula (III), formula (IV) and formula (V), wherein the reaction mixture further comprises a polypeptide having terpene cyclase enzyme activity having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323.

16. A reaction mixture comprising a compound of formula (I) and a polypeptide having terpene cyclase enzyme activity having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 322 and 323.

17. The process of any of claims 1 to 9 and 14, the recombinant cell of any of claims 10 to 12 and 14, the cell culture fermentation medium of any of claims 13 and 14, the reaction mixture of any of claims 15 and 16, wherein the polypeptide having terpene cyclase enzyme activity has amino acid alanine at position 437 and / or amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82.11918. A mutant squalene cyclase enzyme having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or more sequence identity to any of the sequences provided in SEQ ID NOs: 315 to 323, wherein said polypeptide has amino acid alanine at position 437 and amino acid methionine at position 600 relative to the sequence provided in SEQ ID NO: 82.

19. A compound of formula (I) obtained or obtainable by the process of any of claims 1 to 9, 14 and 17, from the recombinant cell of any of claims 10 to 12, 14 and 17, from the cell culture fermentation medium of any of claims 13, 14 and 17, from the reaction mixture of any of claims 15, 16 and 17.

20. Use of the compound of formula (I) of claim 19 as a perfumery, flavor or aroma ingredient, or as a precursor for making said ingredient.120