Polypeptides for making terpenoid compounds
By employing a multi-step enzymatic process and biotransformation method, and utilizing peptides with terpene cyclase activity, the environmental pollution problem inherent in chemical methods has been solved. This has enabled the efficient preparation of high-purity ambroxol and other terpene compounds, suitable for the fragrance industry.
Patent Information
- Application Number
- CN202480051166.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-12
- Filing Date
- 2024-06-12
- Publication Date
- 2026-06-05
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to polypeptides and methods for the enzymatic preparation of ambroxide and other terpenoid compounds. Background Technology
[0002] In the perfumery industry, there is a constant need for methods to prepare compounds for use in the fragrances industry. A key component in these compounds is an amber-like olfactory compound, naturally found in ambergris, which acts as a fixative, making the fragrance last longer.
[0003] The key compound in ambergris is ambroxol—a terpene compound. Terpene compounds are found in most living organisms (microorganisms, animals, and plants). These compounds consist of five-carbon units called isoprene units and are classified according to the number of these units present in their structure. Thus, monoterpenes, sesquiterpenes, and diterpenes are terpenes containing 10, 15, and 20 carbon atoms, respectively. Sesquiterpenes, for example, are widely distributed in the plant kingdom. Many sesquiterpene molecules are known for their flavor and aroma properties, as well as their cosmetic, medicinal, and antibacterial effects. A variety of sesquiterpene hydrocarbons and sesquiterpene classes have been identified. Commercially relevant compounds include Cetalox® ((3aRS,9aRS,9bRS)-3a,6,6,9a-tetramethyl-1,2,3a,4,6,7,8,9,9a,9b-decahydronaphtho[2,1-b]furan; source: Firmenich SA, Geneva, Switzerland) or Ambrox® ((3aR,5aS,9aS,9bR)-3a,6,6,9a-tetramethyldodecanonaphtho[2,1-b]furan; source: Firmenich SA, Geneva, Switzerland), which are similar to ambroxol.
[0004] Chemical routes for preparing these compounds are known in the art. However, given the environmental and waste problems associated with the chemical production of such compounds, there is a need to develop more sustainable processes for producing ambroxol and other terpenoids.
[0005] This invention aims to solve the above-mentioned problems and provides peptides and processes for producing such compounds through in vivo and / or biotransformation methods. The methods may employ multi-step enzymatic processes. Summary of the Invention
[0006] The first aspect of the present invention provides a method for preparing a compound of formula (I).
[0007] (I)
[0008] The compound is in the form of any of its stereoisomers or mixtures thereof, and the method comprises:
[0009] (i) Contacting the compound of formula (VI) with a polypeptide having terpene cyclase activity to produce the compound of formula (I),
[0010] (VI)
[0011] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0012] One embodiment of the present invention is wherein more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib):
[0013] (Formula Ia);
[0014] (Formula Ib).
[0015] Another embodiment of the invention is that the compound of formula (VI) is in the form of formula (VIa):
[0016] (Formula VIa).
[0017] Another aspect of the invention is that the polypeptide with terpene cyclase activity is: a polypeptide that is not a squalene cyclase (SHC), and / or a polypeptide that is a squalene cyclase. In the context of this invention, the polypeptide that is not a SHC enzyme is a meroterpenoid cyclase.
[0018] In one embodiment, the mixed-origin terpene cyclase is selected from at least one or more of the following polypeptides:
[0019] (a) A bacterial membrane-integrated mixed-origin terpene cyclase containing at least one or more amino acid motifs selected from:
[0020] . [W]xxx[D]xx[ILVMN] (SEQ ID NO: 254);
[0021] .PxxAxxxNxxWE(SEQ ID NO: 255);
[0022] . MxxxFxxMLxxR (SEQ ID NO: 256); and
[0023] .RxxxxGQS (SEQ ID NO: 257);
[0024] (b) A fungal-derived membrane-integrated mixed-origin terpene cyclase containing at least one or more amino acid motifs selected from:
[0025] [WY]Exx[YFW] (SEQ ID NO: 258); and
[0026] . [DNE]xSYxxP (SEQ ID NO: 259);
[0027] (c) A bacterial soluble mixed-origin terpene cyclase containing at least one or more amino acid motifs selected from:
[0028] .GxWxxxW[WG]xxxxY (SEQ ID NO: 260);
[0029] . WxxxHxxV[TSA] (SEQ ID NO: 261); and
[0030] .GxWxD[FY] (SEQ ID NO: 262);
[0031] Residue x can be any natural amino acid residue that is independent of each other.
[0032] In a preferred embodiment, the mixed-origin terpene cyclase is a membrane-integrated mixed-origin terpene cyclase. The enzyme preferably produces compounds of formula (I) in the form of formula (Ia).
[0033] In another preferred embodiment, the mixed-origin terpene cyclase is a soluble mixed-origin terpene cyclase. The enzyme preferably produces a compound of formula (I) in the form of formula (Ib).
[0034] In another embodiment, the squalene cyclase comprises at least one motif selected from [SP][TP][VIL]WDTx[LWI] (SEQ ID NO: 247), PGG[WF][GYA]F (SEQ ID NO: 248), PDxDD[TAS][TIAS] (SEQ ID NO: 249), [MIL]QxxxG[GA][WF]x[AS][FY] (SEQ ID NO: 250), Qxxx[GH]xWxG[RK]WGxx[YF]xYG (SEQ ID NO: 251), Qxx[DN]G[GS][WF][GS]ExxxS (SEQ ID NO: 252), and [STA]xx[SFN][QC]T[AGT]W[AS][LIV]xx[LQ] (SEQ ID NO: 253); wherein residue x independently represents any native amino acid residue.
[0035] In another embodiment of the invention, the method further includes one or more steps prior to step (i), said steps including:
[0036] (a) Contacting the compound of formula (V) with a polypeptide having esterase activity to produce the compound of formula (VI),
[0037] (V)
[0038] The compound is in the form of any of its stereoisomers or mixtures thereof;
[0039] (b) Contacting compound (IV) with a polypeptide having Bayer-Villiger monooxygenase (BVMO) enzyme activity to produce compound (V),
[0040] (IV)
[0041] The compound is in the form of any of its stereoisomers or mixtures thereof;
[0042] (c) Contacting the compound of formula (III) with a polypeptide having enal lyase activity to produce the compound of formula (IV),
[0043] (III)
[0044] The compound is in the form of any of its stereoisomers or mixtures thereof;
[0045] (d) Contacting compound (II) with a polypeptide having alcohol dehydrogenase (ADH) activity to produce compound (III),
[0046] (II)
[0047] The compound is in the form of any of its stereoisomers or mixtures thereof;
[0048] (e) Using one or more polypeptides with phosphatase activity, to produce a compound of formula (II) from geraniol geraniol diphosphate (GGPP); and / or
[0049] (f) Using one or more polypeptides with isoprenyl transferase activity, GGPP is generated from isoprenyl diphosphate (IPP) and dimethyl allyl diphosphate (DMAPP).
[0050] In another embodiment of the invention, the method is an in vivo method or a biotransformation method.
[0051] Another aspect of the invention provides a recombinant cell comprising, producing, or capable of producing a compound of formula (I), wherein more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
[0052] Another aspect of the invention provides a cell culture fermentation medium containing the recombinant cells of the invention.
[0053] Another aspect of the invention provides a reaction mixture comprising a compound of formula (I), wherein more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
[0054] Another aspect of the present invention provides a compound that is obtained or can be obtained by the method of the present invention as described above, or obtained or can be obtained from the recombinant cells of the present invention, or obtained or can be obtained from the cell culture fermentation medium of the present invention, or obtained or can be obtained from the reaction mixture of the present invention.
[0055] Another aspect of the invention provides the use of the compound as a flavoring ingredient.
[0056] Another aspect of the invention provides the use of a mixed-source terpene cyclase for producing a compound of formula (I) and / or its derivatives.
[0057] Another aspect of the invention provides a variant of a hybrid terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 56 to 70.
[0058] Another aspect of the invention provides a variant of a hybrid terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 56 to 61, 69, and 70, wherein the variant of the hybrid terpene cyclase has an amino acid substitution at the 9th amino acid position relative to the sequence provided in SEQ ID NO: 51.
[0059] Another aspect of the invention provides a variant squalene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 29, 31, 33, 34, 36 to 38, 40, 41, 43 to 46, 48, 49, 265 to 279, wherein the polypeptide has, relative to the sequence provided in SEQ ID NO: 82, alanine at amino acid position 437 and methionine at amino acid position 600. Preferably, the variant squalene cyclase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 29, 31, 33, 34, 36 to 38, 40, 41, 43, 44, 36 to 38, 40, 41, 43 to 46, 48, 49, 265 to 274, and 276 to 279, wherein the polypeptide has alanine at amino acid position 437 and methionine at amino acid position 600 relative to the sequence provided in SEQ ID NO: 82. Attached Figure Description
[0060] Figure 1 Biosynthetic pathways from isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP) to (2E)-geranyl diphosphate (GPP), (2E,6E)-farnesyl diphosphate (FPP) and (2E,6E,10E)-geranylgeranyl diphosphate (GGPP).
[0061] Figure 2 Biosynthetic pathway from (2E,6E,10E)-geranylgeranyl diphosphate (GGPP) to (2E,6E,10E)-geranylgeranyl. Pi, inorganic phosphate ester (salt); PPi, inorganic pyrophosphate ester (salt).
[0062] Figure 3 The neochemical pathway to (3E,7E)-homofarnesol.
[0063] Figure 4GC-MS analysis was performed on terpenoid compounds and their derivatives produced by Escherichia coli cells engineered to produce (3E,7E)-gafarnesol and express proteins PsAerADH (SEQ ID NO: 11), SCH24-BVMO1 (SEQ ID NO: 23), SCH24-EST1 (SEQ ID NO: 27), CcrGGPPS2-del57 (SEQ ID NO: 1), and PgpB (SEQ ID NO: 3) encoded by plasmid pHFOL-5.
[0064] Figure 5 GC-MS analysis of terpenoid compounds and their derivatives produced using E. coli cells (which express proteins PsAerADH (SEQ ID NO: 11), SCH24-EST1 (SEQ ID NO: 27), CcrGGPPS2-del57 (SEQ ID NO: 1), and PgpB (SEQ ID NO: 3) encoded by plasmid pF-Facetone-7 (A)) and the same cells additionally expressing the BVMO enzyme (AflaBVMO1, SEQ ID NO: 26 (B)).
[0065] Figure 6 GC-MS chromatogram of strain YST403_HFOL, engineered to produce (3E,7E)-gafarnesol. The figure shows the final product (3E,7E)-gafarnesol, as well as the pathway intermediates (5E,9E)-farnesylacetone and (2E,6E,10E)-geraniol.
[0066] Figure 7 A biochemical pathway for the synthesis of Ia compounds using squalene cyclase.
[0067] Figure 8 GC-MS analysis of terpenoid compounds and their derivatives produced by *Escherichia coli* DP1205 expressing PsAerADH (SEQ ID NO: 11), AflavBVMO1 (SEQ ID NO: 26), SCH24-EST1 (SEQ ID NO: 27), CcrGGPPS2-del57 (SEQ ID NO: 1), PgpB (SEQ ID NO: 3), wild-type A0A5P9HJ69 (SEQ ID NO: 42) (A), and the A0A5P9HJ69_V1 variant (SEQ ID NO: 43) (B). The figure shows the compound of formula (Ia) and the pathway intermediate (3E,7E)-gafarnesol.
[0068] Figure 9The amount of compound (Ia) produced by *Escherichia coli* DP1205 expressing PsAerADH (SEQ ID NO: 11), AflavBVMO1 (SEQ ID NO: 26), SCH24-EST1 (SEQ ID NO: 27), CcrGGPPS2-del57 (SEQ ID NO: 1), PgpB (SEQ ID NO: 3) and different wild-type or mutant squalene cyclases (titer).
[0069] Figure 10 The amount of compound (Ia) produced in Saccharomyces cerevisiae cells expressing geraniol geraniol diphosphate synthase CarG (SEQ ID NO: 2), phosphatase PgpB (SEQ ID NO: 3), alcohol dehydrogenase SCH23-ADH1 (SEQ ID NO: 21), Bayer-Villiger monooxygenase AflavBVMO1 (SEQ ID NO: 26), enal lyase SCH94-03944 (SEQ ID NO: 22) and different wild-type or mutant squalene cyclases.
[0070] Figure 11 GC-MS chromatograms of culture extracts from *Saccharomyces cerevisiae* engineered to express the (3E,7E)-gafarnesol biosynthetic pathway gene and the squalene cyclase mutant OYT72085_V1 (SEQ ID NO: 48) (A) or the wild-type squalene cyclase OYT72085.1 (SEQ ID NO: 47) (B). Peaks corresponding to (3E,7E)-gafarnesol and compound (Ia) are marked in the figure.
[0071] Figure 12 The amount of compound (Ia) produced by *Escherichia coli* DP1205 expressing PsAerADH (SEQ ID NO: 11), AflavBVMO1 (SEQ ID NO: 26), SCH24-EST1 (SEQ ID NO: 27), CcrGGPPS2-del57 (SEQ ID NO: 1), PgpB (SEQ ID NO: 3) proteins and bacterial membrane-integrated mixed-origin terpene cyclase.
[0072] Figure 13 A neochemical pathway for the synthesis of compound type (Ia) using mixed-origin terpene cyclases.
[0073] Figure 14 The amount of compound (Ia) produced by yeast strain YST403, which was engineered to produce (3E,7E)-gafarnesol and express different bacterial membrane-integrated mixed-origin terpene cyclases.
[0074] Figure 15 GC-MS chromatograms of yeast strain YST403, engineered to produce (3E,7E)-gafarnesol and express WP_234754442.1 (SEQ ID NO:51) (A), and GC-MS chromatograms of control strain YST403 expressing only the (3E,7E)-gafarnesol pathway enzyme (B). The figures show the compound of formula (Ia) and the pathway intermediate (3E,7E)-gafarnesol.
[0075] Figure 16 The single-ion monitoring mode (221 Da) GC-MS chromatogram of yeast strain YST403, engineered to express (3E,7E)-gafarnesol and A0A2P1DP74.1 (MacJ) (SEQ ID NO:71) (A), and the GC-MS chromatogram of control strain YST403 expressing only the (3E,7E)-gafarnesol pathway enzyme (B). The figures show the compound of formula (Ia) and the pathway intermediate (3E,7E)-gafarnesol.
[0076] Figure 17 A single-ion monitoring mode (221 Da) chiral GC-MS chromatogram of Escherichia coli DP1205, engineered to express (3E,7E)-gafarnesol and OKH29475.1 (SEQ ID NO: 74) (C), was obtained and compared with real standards of compound (B) of formula (Ia) and compound (A) of formula (Ib).
[0077] Figure 18 The single-ion monitoring mode (221 Da) GC-MS chromatograms of strain YST403, engineered to produce (3E,7E)-gafarnesol and express soluble mixed-source terpene cyclases OKH29475.1 (SEQ ID NO: 74) and NEQ07043.1 (SEQ ID NO: 75), and the GC-MS chromatogram of (3E,7E)-gafarnesol produced by the control strain YST403 HFOL are shown in the figures. Compound (Ia) is also shown in the figures.
[0078] Figure 19 The predicted structure of WP_234754442.1 (SEQ ID NO: 51), using ESMFold, shows a porous structure consisting of 7 helices. The N-terminus and C-terminus, as well as the inferred active site entrance, are also marked in the figure, which is located on the same side as the N-terminus.
[0079] Figure 20The amount of compound (Ia) produced by different mutants of *E. coli* DP1205 expressing PsAerADH (SEQ ID NO: 11), AflavBVMO1 (SEQ ID NO: 26), SCH24-EST1 (SEQ ID NO: 27), CcrGGPPS2-del57 (SEQ ID NO: 1), PgpB (SEQ ID NO: 3), and the bacterial membrane-integrated hybrid terpene cyclase WP_234754442.1 (SEQ ID NO: 51). The production amount is expressed relative to the wild-type enzyme WP_234754442.1 (SEQ ID NO: 51).
[0080] Figure 21 The amount of compound (Ia) produced by yeast strain YST403, engineered to produce (3E,7E)-gafarnesol and expressing the bacterial membrane-integrated hybrid terpene cyclase WP_234754442.1 (SEQ ID NO: 51) and its mutants WP_234754442.1_S9C (SEQ ID NO: 56) and WP_234754442.1_S9M (SEQ ID NO: 57). The amount produced is expressed as a percentage relative to the wild-type enzyme WP_234754442.1 (SEQ ID NO: 51).
[0081] Figure 22 The amounts of compound (Ia) (left), compound (Ic) (middle), and compound (Id) produced by biotransformation of chemically synthesized high farnesol with bacterial membrane-integrated mixed-source terpene cyclases WP_051467941.1 (SEQ ID NO: 50), WP_234754442.1 (SEQ ID NO: 51), WP_190963420.1 (SEQ ID NO: 52) and squalene-hope cyclase AacSHC_M132R_I432T_A224V (SEQ ID NO: 78) in the presence or absence of 0.06 (w / v) sodium dodecyl sulfate (SDS) were measured.
[0082] Figure 23GC-MS analysis of terpenoid compounds and derivatives produced by biotransformation of chemically synthesized (3E,7E)-gafarnesol containing the (3Z,7E)-gafarnesol impurity using *E. coli* Bl21(DE3)Star cells expressing the bacterial membrane-integrated hybrid terpene cyclase WP_234754442.1 (SEQ ID NO: 51) (A) and the mutant squalene cyclase AAcSHC_M132R_A224V_I432T (SEQ ID NO: 78) (B). The figure shows compounds of formula (Ia) and (Id) generated by cyclization of (3E,7E)-gafarnesol, and compound of formula (Ic) generated by cyclization of (3Z,7E)-gafarnesol.
[0083] Figure 24 GC-MS analysis of the biotransformation of chemically synthesized (3E,7E)-gafarnesol using *E. coli* Bl21(DE3)Star cells expressing genes encoding squalene cyclases ZmSHC_F437A_G600M (SEQ ID NO: 88), AacSHC_F437A_G600M (SEQ ID NO: 81), A0A0T6LPP7-V1 (SEQ ID NO: 265), A0A7V0I7Y5-V1 (SEQ ID NO: 266), UPI00248B5E40-V1 (SEQ ID NO: 267), and UPI002800B5BA-V1 (SEQ ID NO: 268). The control group was the absence of the aforementioned squalene cyclases.
[0084] Figure 25 The amount of compound (Ia) produced by *Escherichia coli* DP1205 expressing PsAerADH (SEQ ID NO: 11), AflavBVMO1 (SEQ ID NO: 26), SCH24-EST1 (SEQ ID NO: 27), CcrGGPPS2-del57 (SEQ ID NO: 1), PgpB (SEQ ID NO: 3), and various bacterial membrane-integrating mixed-origin terpene cyclases. The amount produced is expressed relative to enzyme WP_234754442.1 (SEQ ID NO: 51). Detailed Implementation
[0085] abbreviations used
[0086] ADH: alcohol dehydrogenase
[0087] BVMO: Bayer-Villiger monooxygenase
[0088] bp: base pairs
[0089] kb: kb base
[0090] DNA: deoxyribonucleic acid
[0091] cDNA: complementary DNA
[0092] DMAPP: Dimethyl allyl diphosphate
[0093] FMO: Flavin monooxygenase
[0094] FPP: Farnesyl diphosphate
[0095] GPP: Geraniol diphosphate
[0096] GGPP: Geraniol Geraniol diphosphate
[0097] GGPS: Geraniol Geraniol diphosphate synthase
[0098] GC: Gas Chromatography
[0099] IPP: Isoprene diphosphate
[0100] iMS: Mass Spectrometer / Mass Spectrometry
[0101] MVA: Mevaleric acid
[0102] PP: diphosphate (ester), pyrophosphate (ester)
[0103] PCR: Polymerase Chain Reaction
[0104] RNA: Ribonucleic Acid
[0105] SHC: Squalene cyclase
[0106] mRNA: messenger ribonucleic acid
[0107] miRNA: microRNA
[0108] siRNA: Small interfering RNA
[0109] rRNA: ribosomal RNA
[0110] tRNA: Transfer RNA
[0111] TPP: Terpenoid diphosphate
[0112] definition
[0113] General terminology:
[0114] In connection with the description and appended claims herein, the use of “or” means “and / or” unless otherwise stated. Similarly, the tenses of “containing,” “containing,” “comprising,” and “including” are interchangeable and not restrictive.
[0115] It should be further understood that, where the term “comprising” is used in the description of various implementation schemes, those skilled in the art will understand that, in certain specific cases, the language of “substantially consisting of” or “consisting of” can be used instead to describe the implementation schemes.
[0116] As used herein, the terms “purified,” “substantially purified,” and “isolated” refer to a state free from other different compounds (with which the compounds of the present invention are typically associated in their natural state), and thus “purified,” “substantially purified,” and “isolated” articles comprise at least 0.5%, 1%, 5%, 10%, or 20%, or at least 50% or 75% by weight of a given sample. In one embodiment, these terms mean that the compounds of the present invention comprise at least 95%, 96%, 97%, 98%, 99%, or 100% by weight of a given sample. As used herein, when referring to nucleic acids or proteins, the terms “purified,” “substantially purified,” and “isolated” for nucleic acids or proteins also refer to a purified or concentrated state that differs from that naturally occurring in, for example, prokaryotic or eukaryotic environments, such as in bacterial or fungal cells, or in mammals, particularly humans. Any degree of purification or concentration greater than that of naturally occurring purification or concentration, including (1) purification from other related structures or compounds, or (2) association with structures or compounds that are typically unrelated in the said prokaryotic or eukaryotic environment, is within the meaning of “isolated.” Nucleic acids, proteins, or classes of nucleic acids or proteins described herein may be isolated or associated with structures or compounds that are otherwise unrelated in nature, according to various methods and processes known to those skilled in the art.
[0117] The term “about” indicates a possible variation of ±25% in the value, particularly ±15%, ±10%, more particularly ±5%, ±2%, or ±1%.
[0118] The term “basically” describes a value range of approximately 80 to 100%, such as 85 to 99.9%, particularly 90 to 99.9%, even more particularly 95 to 99.9%, or 98 to 99.9%, particularly 99 to 99.9%.
[0119] "Mainly" refers to a proportion in the range of more than 50%, such as in the range of 51% to 100%, especially in the range of 75% to 99.9%; especially in the range of 85% to 98.5%, such as 95% to 99%.
[0120] In the context of this invention, "major product" means a single compound or a group of at least two compounds, such as two, three, four, five or more, particularly two or three compounds, which are "majorly" prepared by the reaction described herein and are contained in the reaction in a major proportion based on the total amount of the components of the products formed by the reaction. The proportion may be a molar proportion, a weight proportion, or preferably an area proportion calculated from the corresponding chromatograms of the reaction products based on chromatographic analysis.
[0121] In the context of this invention, "byproduct" means a single compound or a group of at least two compounds, such as two, three, four, five or more, particularly two or three compounds, which are not "mainly" prepared by the reaction described herein.
[0122] Due to the reversibility of enzymatic reactions, unless otherwise stated, this invention relates to enzymatic or biocatalytic reactions described herein in both reaction directions.
[0123] The term "stereoisomer" includes conformational isomers, and in particular configurational isomers.
[0124] According to the present invention, all "stereoisomers" of the compounds described herein are generally included, such as "structural isomers" and "stereoisomers".
[0125] "Stereoisomeric forms" specifically include "stereoisomers" and mixtures thereof, such as configurational isomers (optical isomers), such as enantiomers, or geometric isomers (diastereomers), such as E- and Z-isomers, and combinations thereof. If one or more asymmetric centers are present in a molecule, the invention includes all combinations of different conformations of these asymmetric centers, such as enantiomer pairs.
[0126] "Stereoselectivity" describes the ability to produce a specific stereoisomer of a compound in its stereoisomeric pure form, or the ability to specifically convert a specific stereoisomer from a variety of stereoisomers using the enzymatic catalytic method described herein. More specifically, this means that the product of the invention is enriched relative to a specific stereoisomer, or that the precipitate can be depleted relative to a specific stereoisomer. This can be quantified by a purity parameter %ee calculated according to the following formula:
[0127] %ee = [X A -X B ] / [X A +X B ]×100,
[0128] Where X A and X BMolenbruch represents the molar ratio of stereoisomers A and B.
[0129] The terms "selective conversion" or "increased selectivity" generally refer to the conversion of a specific stereoisomer, such as the E-form of the unsaturated hydrocarbon, at a higher proportion or amount (compared to molar amounts) than the corresponding other stereoisomers, such as the Z-form, during the entire course of the reaction (i.e., between the start and end of the reaction), at a certain point in time of the reaction, or during a "segment" of the reaction. Specifically, during the "segment," the selectivity can be observed corresponding to conversions of 1 to 99%, 2 to 95%, 3 to 90%, 5 to 85%, 10 to 80%, 15 to 75%, 20 to 70%, 25 to 65%, 30 to 60%, or 40 to 50% of the initial substrate amount. The higher proportion or amount can be expressed, for example, as follows:
[0130] -High maximum yield of isomers observed throughout the entire reaction process or during a specified period;
[0131] - Higher relative abundance of isomers at a defined percentage substrate conversion value; and / or
[0132] -At higher conversion percentage values, the same relative content of isomers;
[0133] Each of these preferred methods is observed relative to a reference method, which is performed under otherwise identical conditions using known chemical or biochemical methods.
[0134] The “yield” and / or “conversion” of the reaction according to the invention are determined within a specified time period, such as 4, 6, 8, 10, 12, 16, 20, 24, 36 or 48 hours (during which the reaction takes place). In particular, the reaction is carried out under precisely defined conditions, such as the “standard conditions” as defined herein.
[0135] If this disclosure relates to features, parameters, and ranges of different priorities (including superior, non-preferred features, parameters, and ranges), then unless otherwise stated, any combination of two or more of these features, parameters, and ranges is covered in the disclosure of this invention regardless of their respective priority.
[0136] Biochemical and biological terminology
[0137] The term "domain" refers to a group of amino acids or portions of an amino acid sequence that are conserved at a specific position along an alignment of an evolution-related protein sequence. While amino acids at other positions may vary between protein homologs, highly conserved amino acids at specific positions within such domains represent amino acids that may be essential for the protein's structure, stability, or function. Identified by their high conservation in aligned sequences of protein homolog families, they can be used as identifiers to determine whether any polypeptide under discussion belongs to a previously identified polypeptide family.
[0138] The terms "motif," "shared sequence," or "signature" refer to a short, conserved region in an evolutionarily relevant protein sequence. A motif is typically a highly conserved portion of a domain, but may also include only a part of the domain. A signature is a predictive model describing a protein family, domain, or site.
[0139] Motif sequences can be described using standard IUPAC single-letter amino acid codes. Ambiguities are indicated by listing the acceptable amino acids at a given position in square brackets. For example, [LWI] represents L (leucine), W (tryptophan), or I (isoleucine). X indicates a position where any naturally occurring amino acid residue can exist independently.
[0140] A "protein family" is defined as a group of proteins that share a common evolutionary origin, reflected in their related functions, sequence similarities, or similar primary, secondary, or tertiary structures. Proteins within a protein family are typically homologous and possess similar conserved functional domains and motifs.
[0141] Several databases exist for identifying protein domains, such as SMART (http: / / smart.embl-heidelberg.de / smart / set_mode.cgi?GENOMIC=1) (Schultz et al. (1998) Proc. Natl.Acad. Sci. USA 95, 5857-5864; Letunic et al. (2020) Nucleic Acids Res 49, D458-D460), InterPro (Paysan-Lafosse et al, Nucleic Acids Research, Nov 2022; Mulder et al., (2003) Nucl . Acids. Res. 31, 315-318) or Pfam (Bateman et al., Nucleic Acids Research 30(1): 276-280 (2002)).
[0142] Practical tools for searching or predicting characteristic sequences of protein domains or protein families in protein sequences include, for example, the NCBI Conserved Domain Search Tool (https: / / www.ncbi.nlm.nih.gov / Structure / cdd / wrpsb.cgi) or the InterProScan tool (http: / / www.ebi.ac.uk / interpro / search / sequence / ). Domains or motifs can also be identified using conventional techniques such as sequence alignment.
[0143] The term "Pfam" refers to a large collection of protein domains and families maintained by the Pfam Consortium, available on several sponsored World Wide Web sites, such as the InterPro Consortium website https: / / www.ebi.ac.uk / interpro / (European Laboratory for Molecular Biology - European Institute for Bioinformatics (EMBL_EBI)). The latest version of Pfam is Pfam35.0 (November 2021), based on UniProt Reference Proteomes (El-Gebali S. et al, 2019, Nucleic Acids Res. 47, Database issue D427-D432). Pfam domains and families are identified using multiple sequence alignments and Hidden Markov Models (HMMs). Pfam-A family or domain assignments are high-quality assignments generated using representative members of a protein family through curated seed alignments and seed alignment-based profile HMMs (unless otherwise stated, a match between a queried protein and a Pfam domain or family is a Pfam-A match). Then, a complete alignment of the family is automatically generated using all identified sequences belonging to that family (Sonnhammer (1998) Nucleic Acids Research 26, 320-322; Bateman (2000) Nucleic Acids Research 26, 263-266; Bateman (2004) Nucleic Acids Research 32, Database Issue, D138-D141; Finn (2006) Nucleic Acids Research Database Issue 34, D247-251; Finn (2010) Nucleic Acids Research Database Issue 38, D211-222). For example, protein sequences can be queried against HMM using HMMER homology search software (e.g., HMMER2, HMMER3, or later, hmmer.janelia.org / ) by accessing the Pfam database through any of the aforementioned reference websites. An important match for identifying a queried protein as belonging to the pfam family (or having a specific Pfam domain) is a match where the bit score is greater than or equal to the collection threshold of the Pfam domain. The expected value (e-value) can also be used as a criterion for including the queried protein in the Pfam or determining whether the queried protein has a specific Pfam domain; where the e-value is low, much less than 1.0, for example, less than 0.1 or smaller.
[0144] InterPro is another protein family database that classifies protein sequences into different families and identifies functionally important domains and conserved sites (Blum et al, Nucleic Acids Res. 2021 49(D1):D344-D354). Protein signature sequences are provided by multiple databases, such as Pfam or SMART (Simple Modular Structure Study Tool). InterProScan is software that allows users to retrieve protein and nucleic acid sequences based on InterPro signature sequences.
[0145] The "E-value" (expected value) refers to the number of hits for which a score equal to or higher than this value is expected by chance. This means that a good E-value that provides a reliable prediction is much less than 1. An E-value near 1 represents a chance expectation. Therefore, the lower the E-value, the more specific the search for the domain. Only positive numbers are allowed.
[0146] The "precursor" compounds or molecules of the target compounds or molecules described herein are preferably converted into the target compounds through the enzymatic action of a suitable polypeptide that alters at least one structure or function of the precursor molecule. For example, a "diphosphate precursor" (e.g., a "terpenoid diphosphate precursor") can be converted into the target compound (e.g., a terpene alcohol) by enzymatically removing the diphosphate moiety, such as by removing a monophosphate or diphosphate group, using a phosphatase. For example, an "acyclic precursor" (e.g., an acyclic terpene precursor) can be converted into a cyclic target molecule (e.g., a cyclic terpene compound) through one or more steps by the action of a cyclase or synthase, regardless of the specific enzymatic mechanism of such enzyme.
[0147] The enzyme nomenclature or enzyme classification (EC) system, established by the International Union of Biochemistry and Molecular Biology (IUBMB), is a system for naming and classifying enzymes based on their catalytic activity and biochemical characteristics. Enzyme nomenclature is widely used in the field of biochemistry to classify enzymes according to their function. The EC classification assigns a number to each enzyme, reflecting the reaction or type of reaction catalyzed by that enzyme.
[0148] Enzyme classification can be explored using the "ExplorEnz" database (https: / / www.enzyme-database.org / ) or the International Union of Biochemistry and Molecular Biology (IUBMB) website (https: / / iubmb.qmul.ac.uk). These databases contain information on enzyme classification and nomenclature, function, and properties. Users can search the databases to find information on specific enzyme families or enzymes.
[0149] The terms “biological function,” “function,” “biological activity,” or “activity” for terpene synthases refer to the ability of the terpene diphosphate synthases described herein to catalyze the formation of at least one terpene diphosphate from the corresponding precursor terpene.
[0150] The terms “biological function,” “function,” “biological activity,” or “activity” for terpenoid diphosphate phosphatases refer to the ability of the terpenoid diphosphate phosphatases described herein to catalyze the removal of diphosphate groups from said terpenoid compounds to form the corresponding terpenoid alcohols.
[0151] As used herein, the terms "host cell," "recombinant cell," or "transformed cell" refer to a cell (or organism) that has been modified to carry at least one nucleic acid molecule, such as a recombinant gene encoding a desired protein or nucleic acid sequence, which, upon transcription, produces at least one functional polypeptide of the present invention. Host cells can be, in particular, bacterial, fungal, or plant cells. Host cells may contain recombinant genes or several genes integrated into the nuclear genome of the host cell, such as genes organized as operons. Alternatively, host cells may also contain recombinant genes extrachromosomally. Methods for introducing recombinant nucleic acid sequences into such host cells are well known in the art and are routine laboratory methods that do not require further description herein.
[0152] The term "organism" refers to any non-human multicellular or single-celled organism, such as plants or microorganisms. In particular, microorganisms are bacteria, yeast, algae, or fungi.
[0153] The term "plant" is used interchangeably to include plant cells, including plant protoplasts, plant tissues, plant cell tissue cultures that produce regenerated plants, or parts of a plant, or plant organs such as roots, stems, leaves, flowers, pollen, ovules, embryos, fruits, etc. Any plant may be used to implement the methods described herein.
[0154] Detailed description
[0155] As mentioned above, many sesquiterpenes are known for their flavor and aroma properties, as well as their cosmetic, medicinal, and antibacterial effects. A variety of sesquiterpenes and sesquiterpene compounds have been identified. Commercially relevant compounds include Cetalox® ((3aRS,9aRS,9bRS)-3a,6,6,9a-tetramethyl-1,2,3a,4,6,7,8,9,9a,9b-decahydronaphtho[2,1-b]furan; source: Firmenich SA, Geneva, Switzerland) or Ambrox® ((3aR,5aS,9aS,9bR)-3a,6,6,9a-tetramethyldodecanonaphtho[2,1-b]furan; source: Firmenich SA, Geneva, Switzerland), which are similar to ambroxol.
[0156] The inventors sought to improve the method for preparing compound (I) (also known as 3a,6,6,9a-tetramethyldodecano[2,1-b]furan).
[0157] To develop an improved method for preparing formula (I) compounds, researchers have gained a deeper understanding of the biochemical pathway by which precursor compounds are generated through a multi-enzyme reaction. This multi-enzyme reaction, the first of its kind to be used in a stepwise manner, represents a significant scientific and commercial advancement in the preparation of formula (I) sesquiterpenes. In particular, the combination of enzymes and their sequence in this method have not been described in the prior art.
[0158] This invention includes an in vivo method for preparing a compound of formula (I) in recombinant cells. This is the first demonstration of a method for completely producing the compound in vivo via a biosynthetic pathway that constructs the compound of formula (I) in recombinant cells.
[0159] Therefore, the present invention provides a solution to the problem of preparing such compounds.
[0160] The first aspect of the present invention provides a method for preparing a compound of formula (I).
[0161] (I)
[0162] The compound is in the form of any of its stereoisomers or mixtures thereof, and the method comprises:
[0163] (i) Contacting compound (II) with a polypeptide having ADH enzyme activity to produce compound (III),
[0164] (II)
[0165] The compound is in the form of any of its stereoisomers or mixtures thereof;
[0166] (ii) Contacting the compound of formula (III) with a polypeptide having enal lyase activity to produce the compound of formula (IV),
[0167] (III)
[0168] The compound is in the form of any of its stereoisomers or mixtures thereof;
[0169] (iii) Contacting compound (IV) with a polypeptide having BVMO enzyme activity to produce compound (V),
[0170] (IV)
[0171] The compound is in the form of any of its stereoisomers or mixtures thereof;
[0172] (iv) Contacting the compound of formula (V) with a polypeptide having esterase activity to produce the compound of formula (VI),
[0173] (V)
[0174] The compound is in the form of any of its stereoisomers or mixtures thereof; and
[0175] (v) Contacting the compound of formula (VI) with a polypeptide having terpene cyclase activity to produce the compound of formula (I),
[0176] (VI)
[0177] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0178] For clarity, the use of phrases such as "any of its stereoisomers" or similar expressions refers to the common meaning understood by those skilled in the art, that is, the compounds of the present invention can be pure stereoisomers, such as enantiomers or diastereomers (e.g., related to the E or Z configuration of any double bond, or the R or S configuration of any chiral carbon center).
[0179] According to any form or embodiment of the invention, the compound may be in the form of any stereoisomer thereof or a mixture thereof. For example, the invention relates to a composition of substances comprising one or more forms of a compound of formula (I) having the same chemical structure but different chiral center configurations.
[0180] In particular, compound (I) may be in the form of a mixture containing stereoisomer Ia (formula Ia), wherein stereoisomer Ia accounts for at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% or more of the total mixture.
[0181] (Formula Ia)
[0182] Alternatively, compound (I) may be in the form of a mixture containing stereoisomer Ib (formula Ib), wherein stereoisomer Ib comprises at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% or more of the total mixture.
[0183] (Formula Ib)
[0184] In one embodiment, more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
[0185] According to any form or embodiment of the invention, the compounds of formulas (II) to (VI) may be in the form of their E or Z isomers, or mixtures thereof. In particular, any compound of formulas (II) to (VI) may be in the form of a mixture of stereoisomers E and Z, wherein said stereoisomers IIa, IIIa, IVa, Va or VIa constitute at least 50% to at least 75% of the total mixture (i.e., the ratio of the E / Z mixture is 75 / 25 to 100 / 0).
[0186] Step (i) of the method of the present invention
[0187] Step (i) of the method of the present invention involves contacting a compound of formula (II) with a polypeptide having ADH enzyme activity.
[0188] (Formula II)
[0189] The compound of formula (II) is also known as geraniol, namely 3,7,11,15-tetramethylhexadec-2,6,10,14-tetraen-1-ol, CAS number 7614-21-3.
[0190] Compounds of formula (II) may exist as any of their stereoisomers or mixtures thereof. Specifically, the compound may have the following structures and isomers.
[0191] (Form IIa)
[0192] (2E,6E,10E)-geraniol; (2E,6E,10E)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraen-1-ol; CAS No. 24034-73-9.
[0193] (Formula IIb)
[0194] (2Z,6E,10E)-geraniol; (2Z,6E,10E)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraen-1-ol; CAS No. 57784-25-5.
[0195] (Formula IIc)
[0196] (2E,6Z,10E)-geraniol; (2E,6Z,10E)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraen-1-ol; CAS No. 83689-05-8.
[0197] (Formula IId)
[0198] (2E,6E,10Z)-geraniol; (2E,6E,10Z)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraen-1-ol; CAS No. 68690-77-7.
[0199] (Formula IIe)
[0200] (2Z,6Z,10E)-geraniol; (2Z,6Z,10E)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraen-1-ol; CAS No. 83689-06-9.
[0201] (Formula IIf)
[0202] (2Z,6E,10Z)-geraniol; (2Z,6E,10Z)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraen-1-ol; CAS No. 83689-07-0.
[0203] (Formula IIg)
[0204] (2E,6Z,10Z)-geraniol; (2E,6Z,10Z)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraen-1-ol; CAS No. 83689-08-1.
[0205] (Formula IIh)
[0206] (2Z,6Z,10Z)-geraniol; (2Z,6Z,10Z)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraen-1-ol; CAS No. 1945-42-2.
[0207] Step (i) involves the use of a polypeptide with ADH enzyme activity.
[0208] In the context of this invention, "alcohol dehydrogenase" (ADH) refers to the enzyme that dehydrogenases produced by the NAD+ reaction of NAD+ enzymes. + or NADP +These are polypeptides that, in the presence of a cofactor, possess the ability to oxidize alcohols to the corresponding aldehydes. These enzymes belong to the EC family 1.1.1.1 (NAD... + (Dependency-dependent) or 1.1.1.2 (NADP) + (Dependency-dependent). More particularly, the ADH of the present invention can oxidize linear terpene alcohols to the corresponding carbonyl compounds, especially to the corresponding aldehydes, such as geraniol to geraniol. The ADH used herein can be endogenously present in the corresponding biocatalytic process or exogenous.
[0209] The "alcohol dehydrogenase activity" is determined under the "standard conditions" described below: The determination can be performed using host cells expressing recombinant alcohol dehydrogenase (ADH) peptides, lysed cells expressing ADH peptides, fractions or enriched or purified ADH peptides, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at pH 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (here, particularly geraniol), at an initial concentration of 1 to 100 µM, preferably 5 to 50 µM, especially 30 to 40 µM, or generated endogenously by the cell host. For in vitro assays, a cofactor selected from NADH and NADPH must be added at an easily measurable concentration. The conversion reaction to form the corresponding aldehyde compound (e.g., geraniol) proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. The oxidation product can then be determined using conventional methods, such as extraction with an organic solvent like ethyl acetate.
[0210] Another method for evaluating the oxidation of geraniol to geraniol using ADH is described in Example 3.
[0211] A preferred embodiment of the present invention is that the polypeptide having said ADH enzyme activity comprises one or more sequence motifs selected from the following:
[0212] CHTD (SEQ ID NO: 228), such as the motif in SEQ ID NO: 11, 12, 13, 14, 17, 18, 19 or 20;
[0213] . GHEGxG (SEQ ID NO: 229), such as the motif in SEQ ID NO: 11, 12, 13, 14, 17, 18, 19 or 20;
[0214] . LxCGxxTGxGA (SEQ ID NO: 230), such as the motif in SEQ ID NO: 11, 12, 13, 14, 17, 18, 19 or 20;
[0215] . Gx[VI]GL (SEQ ID NO: 231), such as the motif in SEQ ID NO: 11, 12, 13, 14, 15, 17, 18, 19 or 20;
[0216] . LxxxG[LVI][PA] (SEQ ID NO: 232), such as the motif in SEQ ID NO: 11, 12, 15, 17, 18, 19 or 20;
[0217] . GxVxAI (SEQ ID NO: 233), such as the motif in SEQ ID NO: 16 or 21; and
[0218] YxATKxA (SEQ ID NO: 234), for example, the motif in SEQ ID NO: 16 or 21;
[0219] In the above motif, residue x independently represents any native amino acid residue in a polypeptide with ADH activity. Uncertainty is indicated by listing the acceptable amino acids at a given position in square brackets. For example, [VI] represents V (valine) or I (isoleucine).
[0220] Preferably, the polypeptide having the ADH enzyme activity comprises: CHTD (SEQ ID NO: 228), GHEGxG (SEQ ID NO: 229), LxCGxxTGxGA (SEQ ID NO: 230) and Gx[VI]GL (SEQ ID NO: 231) motifs, such as those in SEQ ID NO: 11, 12, 13, 14, 17, 18, 19 or 20;
[0221] Preferably, the polypeptide having the ADH activity comprises the motifs CHTD (SEQ ID NO: 228), GHEGxG (SEQ ID NO: 229), LxCGxxTGxGA (SEQ ID NO: 230), Gx[VI]GL (SEQ ID NO: 231), and LxxxG[LVI][PA] (SEQ ID NO: 232), such as those motifs in SEQ ID NO: 11, 12, 17, 18, 19, or 20.
[0222] In a preferred embodiment of the present invention, the polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 11 to 21. Preferably, the polypeptide has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with the amino acid sequence provided in SEQ ID NO: 11 or 21. Preferably, the polypeptide has the amino acid sequence provided in SEQ ID NO: 11 or 21.
[0223] Step (ii) of the method of the present invention
[0224] Step (ii) of the method of the present invention involves contacting a compound of formula (III) with a polypeptide having enal lyase activity.
[0225] (Formula III)
[0226] The compound of formula (III) is also known as geraniol, namely 3,7,11,15-tetramethylhexadec-2,6,10,14-tetraenal; CAS number 32480-11-8.
[0227] Compounds of formula (III) may exist as any of their stereoisomers or mixtures thereof. Specifically, the compound may have the following structures and isomers:
[0228] (Formula IIIa)
[0229] (2E,6E,10E)-geraniol; (2E,6E,10E)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraenal; CAS No. 13920-12-2.
[0230] (Formula IIIb)
[0231] (2Z,6E,10E)-geraniol; (2Z,6E,10E)-3,7,11,15-tetramethylhexadec-2,6,10,14-tetraenal; CAS No. 57784-38-0.
[0232] (Formula IIIc)
[0233] (2Z,6Z,10E)-geraniol; (2Z,6Z,10E)-3,7,11,15-tetramethyl-2,6,10,14-hexadecathatetraenal.
[0234] (Formula IIId)
[0235] (2Z,6E,10Z)-geraniol; (2Z,6E,10Z)-3,7,11,15-tetramethyl-2,6,10,14-hexadecathatetraenal.
[0236] (Formula IIIe)
[0237] (2Z,6Z,10E)-geraniol; (2Z,6Z,10E)-3,7,11,15-tetramethyl-2,6,10,14-hexadecathatetraenal.
[0238] (Equation IIIf)
[0239] (2Z,6E,10Z)-geraniol; (2Z,6E,10Z)-3,7,11,15-tetramethyl-2,6,10,14-hexadecathatetraenal.
[0240] (Formula IIIg)
[0241] (2E,6Z,10Z)-geraniol; (2E,6Z,10Z)-3,7,11,15-tetramethyl-2,6,10,14-hexadecathatetraenal.
[0242] (Formula IIIh)
[0243] (2Z,6Z,10Z)-geraniol; (2Z,6Z,10Z)-3,7,11,15-tetramethyl-2,6,10,14-hexadecathatetraenal.
[0244] Step (ii) involves the use of a polypeptide with enal lyase activity.
[0245] In the context of this invention, "enal lyase," "enal lyase protein," or "enal lyase polypeptide" all refer to "α,β-unsaturated aldehyde C=C bond lyase," and may also be called "α,β-unsaturated aldehyde C=C bond lyase," "α,β-unsaturated aldehyde C=C lyase," or "enal C=C lyase." Based on the protein domain composition, the enal lyase protein of this invention may also be described as a member of the "DUF4334 protein family" and / or the "GXWXG protein family" (SEQ ID NO: 263). Examples of such enzymes can be found in the literature, for example, WO2021005097.
[0246] More specifically, the enaldehyde lyase of the present invention has the ability to cleave terpenoid compounds containing α,β-unsaturated aldehyde groups, particularly geraniol to farnesylacetone.
[0247] The "enal lyase activity" is determined under the "standard conditions" described below. The determination can be performed using host cells expressing recombinant enal lysing peptides, lysed cells expressing enal lysing peptides, fractions or enriched or purified enal lysing peptides, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at a pH of 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (here, particularly geraniol), at an initial concentration of 1 to 100 µM, preferably 5 to 50 µM, especially 30 to 40 µM, or generated endogenously by the cell host. The conversion reaction to form the corresponding lysate (e.g., farnesylacetone) proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. The lysate can then be determined using conventional methods, such as extraction with an organic solvent like ethyl acetate.
[0248] The polypeptide having the enal lyase activity may be selected from polypeptides containing the following:
[0249] a) At least one DUF4334 protein family domain with Pfam ID number PF14232 (particularly in the C-terminal region of its amino acid sequence).
[0250] b) At least one GXWXG (SEQ ID NO: 263) protein family domain with Pfam ID number PF14231 (particularly within the N-terminal region of its amino acid sequence); and / or
[0251] c) A structural domain that maintains at least 90% sequence identity with PF14232 or PF14231.
[0252] In particular, the polypeptides of the present invention possessing enal lyase activity are those with an e value less than 1 × 10⁻⁶. -5 Less than 1×10 -10Less than 1×10 -15 Less than 1×10 -20 Less than 1×10 -25 Less than 1×10 -30 or less than or equal to 1×10 -35 Especially in 1×10 -20 Up to 1×10 -32 Within the range, especially in 1×10 -25 Up to 1×10 -31 If the domain matches within the specified range, it is identified as a member of the DUF4334 protein family, which includes the PF14232 domain.
[0253] In particular, peptides with enal lyase activity, if combined with less than 1×10 -5 Less than 1×10 -10 Less than 1×10 -15 Less than 1×10 -20 Less than 1×10 -25 Less than 1×10 -30 or less than or equal to 1×10 -35 Especially in 1×10 -20 Up to 1×10 -30 If the e-value matches within the range, it is identified as a member of the GXWXG (SEQ ID NO: 263) protein family containing the PF14231 domain.
[0254] The query sequence is a polypeptide sequence with enal lyase activity.
[0255] For example, you can use the following websites to search for and calculate such e values: http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or http: / / www.ebi.ac.uk / Tools / pfa / pfamscan / .
[0256] Furthermore, the polypeptide having the enal lyase activity may be selected from a group of polypeptides containing at least one or more sequence motifs / domains selected from the following:
[0257] The G-[Y or "-"]-xWxGxx-[F, L or I]-x-[T, S or R]-G-[H or D] (also represented as GxxWxGxxxxxGx) described in SEQ ID NO: 235, or containing at most 10 or at most 5 consecutive amino acid residues, such as any partial motif corresponding to residues at positions 1-8 or 9-13 in SEQ ID NO: 235. X2 can be Y or can be omitted; X3 can be any naturally occurring amino acid; X5 can be any naturally occurring amino acid; X7 can be any naturally occurring amino acid; X8 can be any naturally occurring amino acid; X9 can be F, L or I; X10 can be any naturally occurring amino acid; X11 can be R, S or T; X13 can be H or D;
[0258] The W-[Y, A, or V]-GKx-[F or Y]-x-[S or D] (also represented as WxGKxxxx) described in SEQ ID NO: 236, or containing up to four consecutive amino acid residues, such as any partial motif corresponding to residues 1-4 or 5-8 in SEQ ID NO: 236. Wherein, X2 can be A, V, or Y; X5 can be any naturally occurring amino acid; X6 can be F or Y; X7 can be any naturally occurring amino acid; X8 can be D or S;
[0259] The [G or S]-x-[A or G]-x-[L or V]-xxxx-[F, Y or L]-RGxV (also represented as xxxxxxxxxxRGxV) as described in SEQ ID NO: 237, or containing at most 10 or at most 5 consecutive amino acid residues, for example, any partial motif corresponding to residues at positions 1-8 or 9-14 in SEQ ID NO: 237. Wherein, X1 can be G or S; X2 can be any naturally occurring amino acid; X3 can be A or G; X4 can be any naturally occurring amino acid; X5 can be L or V; X6 can be any naturally occurring amino acid; X7 can be any naturally occurring amino acid; X8 can be any naturally occurring amino acid; X9 can be any naturally occurring amino acid; X10 can be F, L or Y; X13 can be any naturally occurring amino acid; and
[0260] The [M or L]-[V or I]-YDxxP-[I or V]-xD-[H or S]-[F or L] (also represented as xxYDxxPxxDxx) described in SEQ ID NO:238, or containing at most 10 or at most 5 consecutive amino acid residues, for example, any partial motif corresponding to residues at positions 1-6 or 7-12 in SEQ ID NO:238. Wherein, X1 can be L or M; X2 can be I or V; X5 can be any naturally occurring amino acid; X6 can be any naturally occurring amino acid; X8 can be I or V; X9 can be any naturally occurring amino acid; X11 can be H or S; X12 can be F or L.
[0261] Here, the number of X (e.g., X2) corresponds to its position in the relevant sequence. For example, X2 corresponds to X at position 2 in the relevant sequence; and
[0262] In the above structural motifs, residue x independently represents any native amino acid residue; and in each of the above structural motifs, 1, 2, 3, 4, or 5 amino acid residues different from residue x may optionally be modified, for example by amino acid substitution, particularly by conserved substitution, provided that the enzyme retains enal lyase activity at least within the analytically detectable range. The function of the square brackets has been described above.
[0263] In a preferred embodiment of the present invention, the polypeptide having enal cleavage activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with the sequence provided in SEQ ID NO:22.
[0264] Step (iii) of the method of the present invention
[0265] Step (iii) of the method of the present invention involves contacting the compound of formula (IV) with a polypeptide having BVMO enzyme activity.
[0266] (Form IV)
[0267] The compound of formula (IV) is also known as farnesylacetone, namely 6,10,14-trimethylpentadecano-5,9,13-trien-2-one; CAS number 762-29-8.
[0268] Compounds of formula (IV) may exist as any of their stereoisomers or mixtures thereof. Specifically, the compound may have the following structures and isomers:
[0269] (Form IVa)
[0270] (5E,9E)-Farnesylacetone; (5E,9E)-6,10,14-Trimethylpentadecano-5,9,13-trien-2-one; CAS No. 1117-52-8.
[0271] (Formula IVb)
[0272] (5Z,9E)-Farnesylacetone; (5Z,9E)-6,10,14-Trimethylpentadecano-5,9,13-trien-2-one; CAS No. 1117-51-7.
[0273] (Formula IVc)
[0274] (5E,9Z)-Farnesylacetone; (5E,9Z)-6,10,14-Trimethylpentadecano-5,9,13-trien-2-one; CAS No. 3053-35-3.
[0275] (Formula IVd)
[0276] (5Z,9Z)-Farnesylacetone; (5Z,9Z)-6,10,14-Trimethylpentadecan-5,9,13-trien-2-one; CAS No. 3796-69-8.
[0277] Step (iii) involves the use of a polypeptide with BVMO enzyme activity.
[0278] "Bayer-Villiger monooxygenases" (BVMOs) are flavinases belonging to a class of polypeptides known as oxidoreductases (EC 1.14.13.X). They catalyze the oxidation of straight-chain, cyclic (aromatic or non-aromatic) aldehydes or ketones to the corresponding esters or lactones, highly similar to chemical Bayer-Villiger oxidation. In the enzymatic oxidation process, an atom of molecular oxygen binds to the carbon-carbon bond of the unactivated carbonyl compound. BVMOs require NADPH or NADH as cofactors, or both. They also require an oxygen molecule as a co-substrate. More specifically, a BVMO of the present invention is capable of oxidizing terpene-derived aldehydes or ketones, such as straight-chain terpene carbonyl compounds, particularly farnesylacetone, to the corresponding carbonyl esters.
[0279] BVMO enzyme activity is determined under the "standard conditions" described below: the assay can be performed using host cells expressing recombinant BVMO, lysed BVMO-expressing cells, fractions or enriched or purified BVMO enzymes thereof, at a temperature of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at a pH of 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (here, particularly farnesylacetone), at an initial concentration of 1 to 100 µM, preferably 5 to 50 µM, especially 30 to 40 µM, or produced endogenously by the cell host, in the presence of molecular oxygen. For in vitro assays, a cofactor selected from NADH and NADPH must be added at an easily measurable conversion reaction concentration to generate the corresponding enzymatic product. For example, in the conversion reaction of farnesylacetone, perfarnesyl acetate is added, and the reaction is carried out for 10 minutes to 5 hours, preferably about 1 to 2 hours. The BVMO product can then be determined using conventional methods, such as extraction with an organic solvent like ethyl acetate.
[0280] Another method for screening BVMOs and evaluating the conversion of farnesylacetone to high farnesylacetate is described in Example 2.
[0281] Peptide chains with BVMO enzyme activity can be selected from:
[0282] (1) A group of polypeptides whose amino acid sequence contains a domain of the flavin monooxygenase (FMO) protein family with Pfam ID number PF00743; or a domain having at least 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with PF00743.
[0283] In particular, peptides with BVMO activity, if their e-value is less than 1 × 10⁻⁶, are considered to have BVMO activity. -5 or less than 1×10 -10 or less than or equal to 1×10 -15 or less than or equal to 1×10 -18 Especially in 1×10 -10 Up to 1×10 -18 Within the range, especially in 1×10 -14 Up to 1×10 -17 If the domain matches within the specified range, it is identified as a member of the FMO protein family containing the PF00743 domain. A peptide sequence with BVMO activity is used as the query sequence.
[0284] For example, you can use the following websites to search for and calculate such e values: http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or http: / / www.ebi.ac.uk / Tools / pfa / pfamscan / .
[0285] and / or
[0286] (2) A polypeptide comprising at least one or more sequence motifs / domains selected from the following:
[0287] . GxGxxG (SEQ ID NO: 239), for example, in any of the sequences of SEQ ID NO: 23 to 26. Wherein, X4 can be any naturally occurring amino acid, particularly A or I. The number of X corresponds to its position in the sequence.
[0288] [GS]GxWxxxxYPGxxxD (SEQ ID NO: 240), for example, in any of the sequences SEQ ID NO: 23 to 26;
[0289] . Gxxx[FY]xGxxx[HS]xxxW (SEQ ID NO: 241), for example, in any of the sequences of SEQ ID NO: 23 to 26; and
[0290] [KQ]x[VI]xx[IV]GxG (SEQ ID NO: 242), for example, in any of the sequences of SEQ ID NO: 23 to 26.
[0291] In the above structural motifs, residue x independently represents any native amino acid residue; and in each of the above structural motifs, 1, 2, 3, 4, or 5 amino acid residues different from residue x may optionally be modified, for example by amino acid substitution, particularly by conserved substitution, provided that the enzyme retains BVMO activity at least within the analytically detectable range. The function of square brackets has been described above.
[0292] In a preferred embodiment of the present invention, the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 23 to 26. In another preferred embodiment of the present invention, the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 25 or 26. Preferably, the polypeptide has the amino acid sequence provided in SEQ ID NO: 25 or 26.
[0293] Alternatively, the polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 216 to 227.
[0294] Step (iv) of the method of the present invention
[0295] Step (iv) of the method of the present invention involves contacting the compound of formula (V) with a polypeptide having esterase activity.
[0296] (Form V)
[0297] The compound of formula (V) is also known as gofarnesyl acetate, namely 4,8,12-trimethyldecadec-3,7,11-trien-1-yl acetate; CAS number 109813-25-4.
[0298] The compound of formula (V) can exist as any of its stereoisomers or mixtures thereof. Specifically, the compound can have the following structures and isomers:
[0299] (Equation Va)
[0300] (3E,7E)-gofarnesyl acetate; (3E,7E)-4,8,12-trimethyldecadec-3,7,11-trien-1-yl acetate; CAS No. 944346-19-4.
[0301] (Formula Vb)
[0302] (3Z,7E)-gofarnesyl acetate; (3Z,7E)-4,8,12-trimethyldecadec-3,7,11-trien-1-yl acetate; CAS No. 1467099-77-9.
[0303] (Formula Vc)
[0304] (3E,7Z)-gofarnesyl acetate; (3E,7Z)-4,8,12-trimethyldeca-3,7,11-trien-1-yl acetate.
[0305] (Formula Vd)
[0306] (3Z,7Z)-Gaofarnesyl acetate; (3Z,7Z)-4,8,12-Trimethyldeca-3,7,11-trien-1-yl acetate.
[0307] Step (iv) involves the use of a polypeptide with esterase activity.
[0308] "Esterase" refers to a polypeptide with hydrolytic activity that breaks down esters into acids and alcohols in a chemical reaction (hydrolysis) with water. In the context of this invention, esterases are selected from carboxylic acid ester hydrolases (EC 3.1.1-) that cleave acyl groups, such as acetyl or formyl groups, from the corresponding ester substrate. More particularly, the esterases of this invention have the ability to cleave terpene ester compounds, such as gofarnesyl acetate, to form the corresponding alcohols, particularly gofarnesol.
[0309] Esterase activity is determined under the "standard conditions" described below: The determination can be performed using host cells expressing recombinant esterase peptides, lysed cells expressing esterase peptides, fractions or enriched or purified esterase peptides thereof, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at a pH of 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (here, particularly gofarnesol acetate), at an initial concentration of 1 to 100 µM, preferably 5 to 50 µM, especially 30 to 40 µM, or generated endogenously by the cell host. The conversion reaction to form the corresponding alcohol (particularly gofarnesol) proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. The esterase product can then be determined using conventional methods, such as extraction with an organic solvent like ethyl acetate.
[0310] Another method for evaluating the conversion of gofarnesyl acetate to gofarnesol using esterases is described in Example 4.
[0311] A preferred embodiment of the present invention is that the polypeptide having said esterase activity comprises at least one or more sequence motifs selected from the following:
[0312] . AxVVxVxxRLAPE (SEQ ID NO: 243), for example in SEQ ID NO: 27 or 28;
[0313] . GASAGGGLxA (SEQ ID NO: 244), for example in SEQ ID NO: 27 or 28;
[0314] VxQLLxYPMLDDR (SEQ ID NO: 245), for example in SEQ ID NO: 27 or 28; and,
[0315] ARxxDLSGLPxT (SEQ ID NO: 246), for example in SEQ ID NO: 27 or 28;
[0316] In the above structural motifs, residue x independently represents any natural amino acid residue; and in each of the above structural motifs, 1, 2, 3 or 4 amino acid residues different from residue x may optionally be modified, for example by amino acid substitution, particularly by conserved substitution, provided that the enzyme retains esterase activity at least within the range detectable by analysis.
[0317] Preferably, the polypeptide having the esterase activity comprises: AxVVxVxxRLAPE (SEQ ID NO: 243), GASAGGGLxA (SEQ ID NO: 244), VxQLLxYPMLDDR (SEQ ID NO: 245), and ARxxDLSGLPxT (SEQ ID NO: 246), for example, in SEQ ID NO: 27 or 28.
[0318] In a preferred embodiment of the present invention, the polypeptide having esterase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 27 or 28. Preferably, the polypeptide has the amino acid sequence provided in SEQ ID NO: 28.
[0319] Step (v) of the method of the present invention
[0320] Step (v) of the method of the present invention involves contacting a compound of formula (VI) with a polypeptide having terpene cyclase activity.
[0321] (Form VI)
[0322] Compound (VI) is also known as gafarnesol, 4,8,12-trimethyl-1,8,11-trien-1-ol, CAS No. 35826-67-6.
[0323] Compounds of formula (VI) may exist as any of their stereoisomers or mixtures thereof. Specifically, the compound may have the following structures and isomers:
[0324] (Form VIa)
[0325] (3E,7E)-Gaofarneol; (3E,7E)-4,8,12-Trimethyldecadec-3,7,11-trien-1-ol; CAS No. 459-89-2.
[0326] (Formula VIb)
[0327] (3Z,7E)-Gaofarneol; (3Z,7E)-4,8,12-Trimethyldecadec-3,7,11-trien-1-ol; CAS No. 138152-06-4.
[0328] (Formula VIc)
[0329] (3E,7Z)-Gaofarneol; (3E,7Z)-4,8,12-Trimethyldecadec-3,7,11-trien-1-ol; CAS No. 2032064-12-1.
[0330] (Form VId)
[0331] (3Z,7Z)-Gaofarneol; (3Z,7z)-4,8,12-Trimethyldecadec-3,7,11-trien-1-ol; CAS No. 138152-08-6.
[0332] Based on the different methods of initial carbocation formation, terpene cyclases can be divided into two categories. In type I (or class I) terpene cyclases, the diphosphate group of the linear terpene precursor is first stripped, forming an allyl carbocation in the terpene moiety. In type II cyclases, the initial carbocation is formed through protonation of the double bond or epoxy group in the terpene carbon chain. Therefore, type I cyclases must use substrates containing diphosphate groups, while type II cyclases (since they do not require diphosphate groups to form the initial carbocation) can use terpene compounds as substrates.
[0333] For all terpene cyclases, the generated active carbocation initiates a subsequent cascade of reactions, including the reaction of the carbocation with a double bond, alkyl migration, hydride migration, or carbon-carbon bond formation. This reaction can be terminated by deprotonation of the carbon atom adjacent to the carbocation or by quenching the carbocation with a hydroxyl group or water molecules.
[0334] Type II activity of terpene cyclases is associated with conserved motifs rich in aspartic acid.
[0335] A typical example of class II terpene cyclases is class II diterpene cyclases, which catalyze the cyclization of gerany-gerany diphosphate initiated by protonation to produce, for example, labdadienyl diphosphate intermediates or other cyclic diphosphate intermediates (Peters, RJ (2010). Nat. Prod. Rep. 27, 1521-1530; Zerbe, P. et al (2015). Plant J. 83, 783-793).
[0336] Squalene cyclases (SHCs) are classic examples of class II terpene cyclases, whose substrates do not contain diphosphate functional groups. The squalene cyclase family includes squalene cyclases, 2,3-oxidized squalene cyclases, and enzymes that catalyze cyclization reactions with related catalytic mechanisms. Squalene cyclases catalyze a protonation-initiated cyclization cascade of linear terpenes to generate cyclic compounds. Therefore, squalene cyclases belong to class II terpene cyclases. For example, the squalene family also includes squalene-hopene cyclase (EC 5.4.99.17), which catalyzes the cyclization of squalene to hopene, and squalene-hopanol cyclase (EC 4.2.1.129), which catalyzes the cyclization of squalene to hopan-22-ol. Tetrasopentenyl-β-curcumene-sporeenol cyclase catalyzes a type II cyclization reaction similar to that of linear terpene substrates (EC 4.2.1.137). Studies have shown that tetrasopentenyl-β-curcumene-sporeenol cyclase can also catalyze the cyclization of squalene (Sato, T., et al. (2011). Journal of the American Chemical Society 133(44): 17540-17543), therefore, tetrasopentenyl-β-curcumene-sporeenol cyclase is also a member of the squalene cyclase family.
[0337] Squalene cyclase polypeptides are typically 600 to 800 amino acids long and are membrane-associated proteins. They bind to the cell membrane surface but do not contain transmembrane regions. Squalene cyclases are classified as belonging to the IPR018333 family in the InterPro protein sequence classification database (https: / / www.ebi.ac.uk / interpro / entry / InterPro / IPR018333 / ) (InterPro version 93.0, March 2, 2023). The structure of squalene cyclases is organized into two domains containing multiple α-helices, referred to as the β-domain and γ-domain, or βγ-domain (Christianson DW, Chem. Rev, 2017, 117, 11570-11648). These two domains possess characteristic sequence features, as described in the Pfam database for the N-terminal domain (PF13249) and C-terminal domain (PF13243) of squalene-hope cyclase (Pfam version 35.0, November 19, 2021). The presence of the protein sequence features IPR018333, PF13249, or PF13243 can be predicted using the NCBI Conserved Domain Search Tool (https: / / www.ncbi.nlm.nih.gov / Structure / cdd / wrpsb.cgi) or the InterProScan tool (http: / / www.ebi.ac.uk / interpro / search / sequence / ).
[0338] The squalene cyclase polypeptide contains a characteristic conserved amino acid motif distributed along the sequence and associated with protein structure and enzymatic reactions. Specifically, the squalene cyclase contains at least one or more amino acid motifs selected from the following:
[0339] . [SP][TP][VIL]WDTx[LWI] (SEQ ID NO: 247),
[0340] .PGG[WF][GYA]F (SEQ ID NO: 248),
[0341] .PDxDD[TAS][TIAS] (SEQ ID NO: 249),
[0342] . [MIL]QxxxG[GA][WF]x[AS][FY] (SEQ ID NO: 250),
[0343] .Qxxx[GH]xWxG[RK]WGxx[YF]xYG (SEQ ID NO: 251),
[0344] . Qxx[DN]G[GS][WF][GS]ExxxS (SEQ ID NO: 252), and
[0345] . [STA]xx[SFN][QC]T[AGT]W[AS][LIV]xx[LQ] (SEQ ID NO: 253)
[0346] Motif sequences are described using standard IUPAC single-letter amino acid codes. Uncertainty is indicated by listing the acceptable amino acids at a given position using square brackets. For example, [SP] or [S or P] represents S (serine) or P (proline), respectively. "x" indicates a position where any naturally occurring amino acid residue can exist independently. The function of square brackets has been described previously.
[0347] Mixed-origin terpenoids are a mixture of secondary metabolites derived from multiple biosynthetic pathways, partly from terpenoid co-substrate (Cornforth, JW Terpenoid biosynthesis. Chem. Br. 1968, 4, 102-106). The non-terpenoid portion may originate from biosynthetic pathways involving polyketides, alkaloids, phenols, or amino acids. Mixed-origin terpenoids exhibit rich chemical diversity, particularly in bacteria and fungi.
[0348] The biosynthetic pathway of mixed-origin terpenoids follows several modular biosynthetic steps. The first step involves the generation of structural units (e.g., terpenoids, polyketides) via the corresponding biosynthetic pathway. The terpenoid and non-terpenoid moieties are assembled by isoprenyltransferases. The precursors for the terpenoid moieties are typically linear terpenoid diphosphates, such as geranyl diphosphate, farnesyl diphosphate, or geranylgeranyl diphosphate.
[0349] Next, the linear polyene terpenoids of the mixed precursors partially cyclize to form monocyclic or polycyclic structures. This cyclization reaction is catalyzed by a special class of atypical class II terpene cyclases, known as mixed-origin terpene cyclases, which were first discovered in fungi (T. Itoh et al, 2010, 2, 858-864). The first representative mixed-origin terpene cyclase discovered was Pyr4 from Aspergillus fumigatus Af293 (Itoh, T., et al. (2010). Nature Chemistry 2(10): 858-864).
[0350] In many mixed-origin terpenoid compounds, the linear terpene precursor is first activated by a monooxygenase through stereoselective epoxidation of one of the double bonds. Subsequently, the mixed-origin terpene cyclase catalyzes the protonation of the epoxide moiety, generating an active carbocation, and initiating a subsequent cascade reaction similar to that of other terpene cyclases. Some mixed-origin terpene cyclases can convert isoprene precursors into cyclized products without a prior epoxidation step. These mixed-origin terpene cyclases can directly protonate the terminal double bond, generating an active carbocation and catalyzing the cyclization reaction. For example, the MacJ enzyme from the fungus *Penicillium terrestry* was the first fungal mixed-origin terpene cyclase identified to initiate the reaction using type II double bond protonation (Tang, M.-C., et al. (2017). Organic Letters 19(19): 5376-5379). Another example of a mixed-origin terpene cyclase that initiates polyene cyclization via direct double bond protonation is DmtA1 from the bacterium Streptomyces youssoufiensis OUC68199 (Yao et al, Nat. Commun., 2018, 9, 4091).
[0351] Like other class II terpene cyclases, the carbocations produced by mixed-source terpene cyclases initiate a cascade reaction, which typically begins with an attack on the double bond, resulting in a monocyclic or polycyclic structure with a tertiary carbocation. The reaction terminates in two ways: either through deprotonation to form a double bond, or through a reaction with water molecules to form a tertiary alcohol. Typical cyclic structures in mixed-source terpenoids contain a drimane or labdane skeleton.
[0352] The largest class of hybrid terpene cyclases are compact, membrane-integrated proteins containing multiple (typically seven) transmembrane helices. The structure of these transmembrane helices can be readily predicted, for example, using the TMHMM 2.0 server available at https: / / dtu.biolib.com / DeepTMHMM (Krogh, A., et al. (2001) J Mol Biol305(3): 567-580). Besides their protein structure, hybrid terpene cyclases differ from other class II cyclases (such as squalene cyclases) in their smaller polypeptide chains. Bacterial and fungal hybrid terpene cyclase polypeptides range in length from 150 to 550 amino acid residues. The transmembrane helix is located on the polypeptide chain, covering 180 to 300 amino acids, and contains a catalytic domain.
[0353] Recently, some hybrid terpene cyclases with protein structures different from membrane-integrated hybrid terpene cyclases have been discovered. For example, MstE from the bacterium *Scytonema* sp. PCC 1002 is a soluble cyclase whose structure is similar to, but different from, typical cyclases such as diterpene synthases and squalene cyclases because it is a single-domain protein containing only the α-domain (Moosmann, P., et al. (2020). Nat Chem 12(10): 968-972). The length of soluble bacterial hybrid terpene cyclase polypeptides ranges from 150 to 550 amino acids.
[0354] Various mixed-origin terpene cyclases can catalyze the cyclization of terpenoid portions in mixed-origin terpene precursors to form hemisarane rings. However, no studies have yet reported that mixed-origin terpene cyclases can catalyze the cyclization of linear terpenoid compounds to form hemisarane compounds.
[0355] Mixed-origin terpene cyclase polypeptides contain characteristic conserved amino acid motifs located in the sequence and associated with protein structure or enzymatic reactions, as shown below:
[0356] Bacterial-derived membrane-integrated mixed-origin terpene cyclases containing at least one or more amino acid motifs selected from the following:
[0357] . [W]xxx[D]xx[ILVMN] (SEQ ID NO: 254);
[0358] .PxxAxxxNxxWE(SEQ ID NO: 255);
[0359] . MxxxFxxMLxxR (SEQ ID NO: 256); and
[0360] .RxxxxGQS (SEQ ID NO: 257).
[0361] A fungal-derived membrane-integrated mixed-origin terpene cyclase containing at least one or more amino acid motifs selected from the following:
[0362] [WY]Exx[YFW] (SEQ ID NO: 258); and
[0363] . [DNE]xSYxxP (SEQ ID NO: 259).
[0364] A bacterially derived soluble mixed-origin terpene cyclase containing at least one or more amino acid motifs selected from the following:
[0365] .GxWxxxW[WG]xxxxY (SEQ ID NO: 260);
[0366] . WxxxHxxV[TSA] (SEQ ID NO: 261); and
[0367] .GxWxD[FY] (SEQ ID NO: 262).
[0368] The motif sequences are described using standard IUPAC single-letter amino acid codes. Residue x independently represents any native amino acid residue, and in each of the above motifs, optionally one, two, three, or four amino acid residues different from residue x may be modified, for example by amino acid substitution, particularly by conserved substitution, provided that the enzyme retains its enzymatic activity at least within the analytically detectable range. The function of square brackets has been described above.
[0369] The selection of heterologous terpene cyclase polypeptides can be searched in sequence databases using tools such as BLAST (Tatiana et al., FEMS Microbiol Lett., 1999, 174:247-250, 1999), with MacJ (SEQ ID NO: 71), DmTA1 (SEQ ID NO: 77), or MstE (SEQ ID NO: 76) as query sequences. This selection can be further optimized by choosing sequences of appropriate length or containing the aforementioned characteristic amino acid motifs. Furthermore, the selection can be further optimized by predicting protein architecture, particularly the presence of transmembrane helices.
[0370] Specific examples of suitable standard conditions for each of the above enzyme activities can be obtained from the Examples section below.
[0371] As discussed above, terpene cyclases can be squalene cyclases (SHC) or mixed-origin terpene cyclases (MeroTPS).
[0372] To avoid ambiguity, SHCs and mixed-origin terpene cyclases are two different classes of enzymes that can be distinguished by their physical properties. Furthermore, mixed-origin terpene cyclases can be classified as: (i) bacterial membrane-integrated mixed-origin terpene cyclases; (ii) fungal membrane-integrated mixed-origin terpene cyclases; and (iii) bacterial soluble mixed-origin terpene cyclases.
[0373] Table A below lists the differences between SHC and different types of mixed-origin terpene cyclases (meroTPS).
[0374]
[0375] Table A: Characteristics of Enzymes
[0376] Therefore, those skilled in the art can easily determine whether an enzyme is an SHC enzyme or a mixed-origin terpene cyclase based on the information provided herein.
[0377] In a preferred embodiment of the present invention, the terpene cyclase is SHC.
[0378] As shown in the accompanying examples, the inventors have demonstrated that the SHC enzyme can be used in step (v) of the method of the present invention. A preferred embodiment of the invention is that the polypeptide having SHC enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NOs: 29 to 49, 265 to 274, and 276 to 279. Preferably, the polypeptide having SHC enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NOs: 29 to 49, 265 to 274, and 276 to 279. More preferably, the polypeptide having SHC enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 48, 265, 266, 267, 268, 274, 276, and 279. More preferably, the polypeptide having SHC enzyme activity has the amino acid sequence of any one of SEQ ID NO: 48, 265, 266, 267, 268, 274, 276, and 279.
[0379] Alternatively, the polypeptide having SHC enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 79 to 89.
[0380] In a preferred embodiment of the present invention, the terpene cyclase is a mixed-origin terpene cyclase.
[0381] In constructing the method of the present invention, the inventors sought to compare the isomer distribution of the compound of formula (I) synthesized by SHC and mixed-source terpene cyclase.
[0382] To their surprise, they found that when a mixed-origin terpene cyclase was used in the method of the present invention, the resulting compound (I) exhibited isomer preference, with a better olfactory character for isomers of compound (I) than that produced by the SHC enzyme.
[0383] In particular, they demonstrated that less than 1% of the compounds of formula (I) produced by the method of the present invention are isomers of formula I(c) and / or I(d). Therefore, the use of mixed-origin terpene cyclases has a surprising technical advantage compared to the use of SHC enzymes.
[0384] Preferably, the mixed-origin terpene cyclase is a membrane-integrated mixed-origin terpene cyclase. As can be seen from the accompanying examples, such mixed-origin terpene cyclases preferably produce compounds of formula (Ia) rather than those of formulas (Ic) and / or (Id). Examples of membrane-integrated mixed-origin terpene cyclases include those provided in any of SEQ ID NO: 50 to 73 and 280 to 289. SEQ ID NO: 50 to 70 and 280 to 289 are bacterial membrane-integrated mixed-origin terpene cyclases, while SEQ ID NO: 71 to 73 are fungal membrane-integrated mixed-origin terpene cyclases.
[0385] Preferably, the mixed-origin terpene cyclase is a soluble mixed-origin terpene cyclase. As can be seen from the accompanying examples, such mixed-origin terpene cyclases preferably produce compounds of formula (Ib) rather than compounds of formula (Ic) and / or formula (Id). Examples of soluble mixed-origin terpene cyclases include those provided in SEQ ID NO: 74 or 75.
[0386] This is the first time that a compound of type (I) has been prepared using a mixed-origin terpene cyclase, and the preference for this isomer form is surprising and has significant technical implications for commercial applications.
[0387] Furthermore, the inventors unexpectedly discovered that mixed-source terpene cyclases can be used in biotransformation methods without the need to add any detergent to the reaction. The absence of detergent in the biotransformation reaction simplifies the preparation of compound (I), making it more cost-effective.
[0388] As shown in the accompanying examples, the inventors have demonstrated that the mixed-origin terpene cyclase can be used in step (v) of the method of the present invention. A preferred embodiment of the invention is that the polypeptide having mixed-origin terpene cyclase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 50 to 75 and 280 to 289. Preferably, the mixed-origin terpene cyclase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 57, 71, 74, 280, 281, 282, 283, 286, 287, and 288. Preferably, the mixed-source terpene cyclase activity has the sequence provided in SEQ ID NO: 57, 71, 74, 280, 281, 282, 283, 286, 287 or 288.
[0389] Various forms of the compound of formula (I) produced by the method of the present invention.
[0390] This invention provides a method for preparing compound (I).
[0391] Compound of formula (I):
[0392] Formula (I)
[0393] The compound of formula (I) is also known as 3a,6,6,9a-tetramethyldodecano[2,1-b]furan; CAS number 3738-00-9.
[0394] The compound of formula (I) may exist as any of its stereoisomers or mixtures thereof. Specifically, the compound may have the following structures and isomers:
[0395] (Formula Ia)
[0396] (3aR,5aS,9aS,9bR)-3a,6,6,9a-tetramethyldodecano[2,1-b]furan; CAS No. 6790-58-5.
[0397] (Formula Ib)
[0398] (3aS,5aR,9aR,9bS)-3a,6,6,9a-tetramethyldodecanonaphtho[2,1-b]furan; CAS No. 234431-64-2.
[0399] (Formula Ic)
[0400] (3aR,5aS,9aS,9bS)-3a,6,6,9a-tetramethyldodecanonaphtho[2,1-b]furan.
[0401] (Formula ID)
[0402] (3aS,5aS,9aS,9bS)-3a,6,6,9a-tetramethyldodecanonaphtho[2,1-b]furan.
[0403] Compound (I) can have different enantiomers. Some isomers have superior olfactory characteristics than other isomers of the compound. In particular, the olfactorily preferred forms of compound (I) are compounds (Ia) and / or (Ib).
[0404] To improve the yield or proportion of compounds (Ia) and / or (Ib) relative to other isomers of compound (I), achieving high selectivity in the formation of compound (VI) is crucial. Specifically, the ratio of E / Z isomers at the 3,4-double bond of compound (VI) is significant. The preferred form of compound (VI) is formula (VIa).
[0405] Despite extensive efforts, obtaining compounds of formula (VI) with an E / Z ratio greater than 90:10 at the 3,4-double bond using chemical methods remains a challenge and has not yet been achieved on a large scale (Eichhorn, E. and F. Schroeder (2023). J Agric Food Chem).
[0406] This invention achieves this objective by using highly selective enzymes. As shown in the accompanying examples, this invention enables the highly selective production of compound (VI) in the form of (VIa). Chemical methods (as described in Eichhorn, E. and F. Schroeder (2023). J Agric Food Chem) cannot achieve such high selectivity. Enzymatic pathways can yield pathway compounds with high E / Z ratios. In particular, using enzymes, especially geraniol-geraniol diphosphate synthase, geraniol-geraniol and geraniol with high double bond E / Z ratios can be obtained. All intermediates in this pathway retain the double bond configuration.
[0407] This provides a method for preparing compounds of formula (I) in an olfactorily preferred form. Therefore, the method of preparing compounds of formula (I) free from a large number of undesirable byproducts has significant commercial value. As shown in the accompanying examples, over 97% of the compounds of formula (I) are in the form of formula (Ia) and / or formula (Ib).
[0408] Therefore, in one embodiment of the present invention, more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
[0409] The scope of this invention also includes compounds of formula (I) obtained by in vivo methods or obtainable by such methods.
[0410] Understandably, this is the first time a compound of formula (I) has been prepared via an in vivo method. For the reasons described herein, this represents a significant advancement in the field of compound preparation, both technically and commercially. Compounds of formula (I) can be prepared by any in vivo method, preferably using recombinant cells expressing enzymes that can be used to synthesize this molecular pathway. In one embodiment of the invention, more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
[0411] Implementation scheme of the method of the present invention.
[0412] The method of the present invention includes using (i) a polypeptide having ADH enzyme activity; (ii) a polypeptide having enal lyase activity; (iii) a polypeptide having BVMO enzyme activity; (iv) a polypeptide having esterase activity; and (v) a polypeptide having terpene cyclase activity.
[0413] Another preferred embodiment of the method of the present invention is wherein:
[0414] (i) The polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 11 to 21.
[0415] (ii) The polypeptide having enal lyase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 22.
[0416] (iii) The polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 23 to 26 and 216 to 227 (preferably with any of the sequences in SEQ ID NO: 23 to 26);
[0417] (iv) The polypeptide having esterase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any sequence in SEQ ID NO: 27 and 28; and / or,
[0418] (v) The polypeptide having terpene cyclase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 29 to 75, 79 to 89 and 265 to 289 (preferably with any of the sequences in SEQ ID NO: 29 to 75, 265 to 274 and 276 to 289).
[0419] Another preferred embodiment of the present invention is wherein:
[0420] (i) The polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 11 or 21; preferably, the ADH enzyme has the sequence of SEQ ID NO: 11 or 21.
[0421] (ii) The polypeptide having enal lyase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 22; preferably, the enal lyase has the sequence of SEQ ID NO: 22.
[0422] (iii) The polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 25 or 26; preferably, the BVMO enzyme has the sequence of SEQ ID NO: 25 or 26.
[0423] (iv) The polypeptide having esterase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 28; preferably, the esterase has the sequence of SEQ ID NO: 28; and / or,
[0424] (v) A polypeptide having terpene cyclase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 48, 57, 71, 74, 265, 266, 267, 268, 274, 276, 279, 280, 281, 282, 283, 286, 287, or 288. Preferably, the terpene cyclase has the sequence of SEQ ID NO: 48, 57, 71, 74, 265, 266, 267, 268, 274, 276, 279, 280, 281, 282, 283, 286, 287, or 288.
[0425] Alternatively, in this embodiment of the invention:
[0426] (v) The polypeptide having terpene cyclase activity is an SHC enzyme and has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 48, 265, 266, 267, 268, 274, 276, or 279. Preferably, the SHC enzyme has the sequence of SEQ ID NO: 48, 265, 266, 267, 268, 274, 276, or 279.
[0427] Alternatively, in this embodiment of the invention:
[0428] (v) The polypeptide having terpene cyclase activity is a mixed-source terpene cyclase and has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 57, 71, 74, 280, 281, 282, 283, 286, 287, or 288. Preferably, the mixed-source terpene cyclase has the sequence of SEQ ID NO: 57, 71, 74, 280, 281, 282, 283, 286, 287, or 288.
[0429] Another embodiment of the present invention is wherein:
[0430] (i) Compounds of formula (I) are in the form of formula (Ia):
[0431] (Formula Ia)
[0432] (ii) Compounds of formula (II) are in the form of formula (IIa):
[0433] (Form IIa)
[0434] (iii) Compounds of formula (III) are in the form of formula (IIIa):
[0435] (Formula IIIa)
[0436] (iv) Compounds of formula (IV) are in the form of formula (IVa):
[0437] (Form IVa)
[0438] (v) Compounds of formula (V) are in the form of formula (Va):
[0439] (Equation Va)
[0440] (vi) Compounds of formula (VI) are in the form of formula (VIa):
[0441] (Formula VIa).
[0442] Preparation of compound (II)
[0443] The present invention provides a method for preparing compound (I) by continuous biocatalytic transformation of compound (II).
[0444] Compounds of formula (II) can be used as starting substrates in the methods of the present invention, for example, in the form of purified compound formulations.
[0445] However, in a preferred embodiment of the present invention, the method of the present invention further includes providing compound (II) by sequentially biocatalytically converting the precursor compound into compound (II).
[0446] The compounds of formulas (I) to (VI) are terpenoids.
[0447] Terpenoids are a large family of structurally diverse natural compounds. All terpenoids are biosynthesized from two five-carbon units—isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP). IPP and DMAPP can be generated through different biosynthetic pathways, such as the 2-C-methyl-D-erythritol-4-phosphate (MEP) pathway, the mevalonate (MVA) pathway, or the MVA substitution pathway (Dellas, N. et al. 2013, eLife 2: e00672). Alternatively, IPP and DMAPP can also be generated through the sequential enzymatic phosphorylation or enzymatic pyrophosphorylation of their corresponding alcohols—3-methyl-3-butenol and 3-methyl-2-butenol (Ma, X, et al. (2022). J AgricFood Chem 70(11): 3512-3520).
[0448] These terpene structural units condense sequentially to form linear terpene precursors with different lengths and multiples of 5 carbon atoms, such as geraniyl diphosphate (GPP), farnesyl diphosphate (FPP), or geraniylgeraniyl diphosphate (GGPP) containing 10, 15, and 15 carbon atoms, respectively. The condensation reaction of IPP and DMAPP is carried out by an enzyme called isoprenyltransferase. Isoprenyltransferase catalyzes the initial condensation reaction between IPP and DMAPP to generate GPP, followed by further addition of the IPP molecule to generate FPP and then GGPP (Ogura, K., and Koyama, T. (1998). Chem. Rev. 98, 1263-1276). The sequential condensation reaction of DMAPP and IPP to generate GGPP can proceed as follows:
[0449] i. The sequential action of three isoprenyltransferases: GPP synthase, FPP synthase which catalyzes the addition of one IPP to a GPP, and GGPP synthase which catalyzes the addition of one IPP to a FPP.
[0450] ii. A combination of two isoprenyltransferases, such as an FPP synthase catalyzing the condensation of one DMAPP and two IPPs, and a GGPP synthase catalyzing the addition of one IPP to an FPP; and / or
[0451] iii. The role of an isoprenyltransferase, such as a GGPP synthase capable of catalyzing the sequential condensation of three IPPs and one DMAPP.
[0452] Some terpenoids have simple linear structures; for example, geraniol is a linear diterpene containing a terminal hydroxyl group (20 carbon atoms). Geraniol can be prepared from GGPP via any of the following methods:
[0453] i. Enzymes with pyrophosphatase activity, such as phosphatases; and / or
[0454] ii. Two phosphatase-active enzymes that sequentially cleave the two phosphate groups of GGPP.
[0455] Alternatively, geraniol can be prepared from GGPP using an enzyme from the class I terpene cyclase family (described below). This enzyme is capable of cleaving the pyrophosphate group of GGPP, but lacks the ability to catalyze continuous cyclization.
[0456] Therefore, these enzymatic pathways utilize enzymes with phosphatase activity to cleave the diphosphate group of GGPP and release geraniol.
[0457] Alternatively, geraniol can also be prepared from another linear diterpene. For example, it can be prepared from geraniol or the corresponding polyene (e.g., β-springene). Enzymes such as dehydratase-isomerases can be used for this type of reaction. Dehydratase-isomerases can catalyze reversible isomerization reactions between linear terpenoids with terminal alcohol groups or terminal double bonds (Nestl, BM, et al. (2017). Nature Chemical Biology 13(3): 275-281) (see...) Figure 2 ).
[0458] Alternatively, geraniol can be synthesized chemically. For example, geraniol can be obtained by chain extension of farnesene or farnesol (Organic Syntheses, Vol. 84, p. 43-57 (2007)).
[0459] The pathways that generate IPP, DMAPP, and geranyl diphosphate (GPP), farnesyl diphosphate (FPP), or geranylgeranyl diphosphate (GGPP) are involved in the synthesis of terpenoids. Terpenoids are a diverse class of molecules that play crucial roles in primary metabolism and various cellular processes. They are involved in a variety of biological functions, including the synthesis of sterols (such as cholesterol in animals and phytosterols in plants), and the production of hormones, vitamins (such as vitamin E and vitamin K), and signaling molecules (such as ubiquinone and polyterpenols). Furthermore, terpenoids are also essential for membrane lipid formation and post-translational modifications of proteins.
[0460] Therefore, as mentioned above, the pathways to GPP, FPP, and GGPP are an important part of primary metabolism in all organisms because they provide the necessary raw materials for the synthesis of essential terpenoid compounds that participate in various physiological processes that are crucial for cell growth.
[0461] Most terpenoids contain a cyclic carbon framework. The diversity of monocyclic and polycyclic carbon skeletons stems from the enzymatic conversion of linear terpene precursors by terpene cyclases (TCs, also known as terpene synthases). The cyclization reaction begins with a carbocation, which reacts with an electron-rich double bond to form a new carbon bond. The outcome depends on the folding pattern of the substrate and the nature of the amino acid side chains at the enzyme's active site.
[0462] Therefore, in a preferred embodiment of the present invention, the method of the present invention further includes one or more steps prior to step (i), said steps including:
[0463] (a) Using one or more isoprenyltransferases, geraniylgeraniyl diphosphate (GGPP) is prepared from IPP and DMAPP; and / or,
[0464] (b) Using one or more enzymes with phosphatase activity, GGPP is used to prepare (generate) compound (II).
[0465] Step (a) of this embodiment of the invention requires the use of one or more isoprenyl transferases.
[0466] The term "isoprenyltransferase" refers to a group of enzymes that can sequentially condense five-carbon units, such as isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP), to form linear terpene diphosphates, such as geranyl diphosphate (GPP), farnesyl diphosphate (FPP), or geranylgeranyl diphosphate (GGPP), containing 10, 15, and 15 carbon atoms, respectively. Some isopentenyltransferases can add five-carbon units to linear terpene diphosphates, thereby extending the carbon chain length. Geranyl diphosphate synthase (GGPP synthase) is an example of an isopentenyltransferase, capable of generating GGPP from IPP and DMAPP, or by adding five carbon atoms to FPP. The term "isoprenyltransferase" also refers to a group of enzymes that can transfer isoprenoid subunits from terpenoid diphosphates (usually linear terpenoid diphosphates) to non-terpenoid frameworks during the biosynthesis of mixed-origin terpenoids.
[0467] Step (a) of this embodiment of the invention can be performed in the following manner:
[0468] i. The sequential action of three isoprenyltransferases: GPP synthase, FPP synthase which catalyzes the addition of one IPP to a GPP, and GGPP synthase which catalyzes the addition of one IPP to a FPP.
[0469] ii. A combination of two isoprenyltransferases, such as an FPP synthase catalyzing the condensation of one DMAPP and two IPPs, and a GGPP synthase catalyzing the addition of one IPP to an FPP; and / or
[0470] iii. The role of an isoprenyltransferase, such as a GGPP synthase capable of catalyzing the sequential condensation of three IPPs and one DMAPP.
[0471] Examples of isoprenyltransferases that can be used in this step of the method of the present invention are well known in the art. For example, GGPP synthase from Blakeslea trispora can be used for the biosynthesis of GGPP (sun et al, Biotechnol. Lett. 34 (11), 2077-2082 (2012)).
[0472] Preferably, the isoprenyltransferase is a GGPP synthase. Preferably, the GGPP synthase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 1 or 2. Preferably, the GGPP synthase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 2. Preferably, the GGPP synthase has the amino acid sequence of SEQ ID NO: 2.
[0473] Step (b) of this embodiment of the invention requires the use of one or more enzymes with phosphatase activity.
[0474] The term "phosphatase" refers to a group of enzymes known to remove phosphate or bisphosphate groups from precursors containing phosphate or bisphosphate groups. A specific subgroup of phosphatases has the ability to remove phosphate or bisphosphate groups from terpene precursors, thereby releasing inorganic phosphate and the corresponding terpene alcohol. For example, some phosphatases are known to remove the bisphosphate group from GGPP to produce geraniol. Phosphatases acting on terpene diphosphates exist in multiple enzyme classes.
[0475] The term "protein tyrosine phosphatase" refers to a group of enzymes that are generally known to remove phosphate groups from phosphorylated tyrosine residues on proteins. As described in WO2020011883A1, a specific subgroup of this family consists of enzymes capable of removing diphosphate groups from phosphorylated terpene molecules (terpenoid diphosphates). Specifically, the protein tyrosine phosphatase family with Pfam ID PF13350 dephosphorylates GGPP to geraniol. Peptides can be compared to the Pfam protein family signature database to detect the presence of a match.
[0476] Phosphatases, particularly GGPP phosphatases, can also be obtained from other protein families; for example, the phosphatidylphosphatase type 2 (PAP2) protein family (IPR00326), the Nudix hydrolase protein family (IPR015797), and the haloacid dehalogenase-like (HAD-like) hydrolase protein family (IPR041492). A method for screening phosphatases and evaluating the conversion of geraniol to geraniol is described in Example 5.
[0477] Examples of phosphatases that can be used in this step of the method of the present invention are well known in the art.
[0478] Preferably, the phosphatase is a GGPP phosphatase. Preferably, the GGPP phosphatase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any one of SEQ ID NO: 3. Preferably, the GGPP phosphatase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 3. Preferably, the GGPP phosphatase has the amino acid sequence of SEQ ID NO: 3.
[0479] As described above, this embodiment of the present invention provides compound (II) by sequentially biocatalyzing a precursor compound into compound (II).
[0480] In this embodiment of the invention, the precursor compound of formula (II) is provided by IPP and DMAPP.
[0481] Another embodiment of the present invention further includes the preparation of IPP and DMAPP.
[0482] One method for preparing IPP and DMAPP is via the mevalonate pathway. The mevalonate pathway, also known as the isoprene pathway or HMG-CoA reductase pathway, is an important metabolic pathway present in eukaryotes, archaea, and certain bacteria. This pathway uses acetyl-CoA as a starting substrate to generate two five-carbon structural units called isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP). By combining the mevalonate pathway with enzyme activity, terpene precursors GPP, FPP, or GGPP are generated, enabling the recombinant cellular production of terpenes. This pathway is known in the art. The enzymes required for the conversion of acetyl-CoA to IPP and DMAPP are listed below:
[0483] Acetyl-CoA acetyltransferase (ACAT)
[0484] 3-Hydroxy-3-methylglutaryl-CoA synthase (HMG-CoA synthase)
[0485] 3-Hydroxy-3-methylglutaryl-CoA reductase (HMG-CoA reductase)
[0486] Mevalonate kinase
[0487] Mevalonate kinase
[0488] Mevalonate diphosphate decarboxylase
[0489] Isopentenyl diphosphate isomerase.
[0490] An alternative method for preparing IPP and DMAPP is via methyl erythritol phosphate (MEP). This route is well known in the art. The list of enzymes required to convert glyceraldehyde-3-phosphate (GAP) and pyruvate into IPP and DMAPP is as follows:
[0491] 1-Deoxy-D-xyulose 5-phosphate synthase (DXS)
[0492] 1-Deoxy-D-xyulose-5-phosphate reductase (DXR)
[0493] 2-C-methyl-D-erythritol 4-phosphate cytidine transferase (MCT, IspD)
[0494] 4-Cytidine diphosphate-2-C-methyl-D-erythritol kinase (CMK, IspE)
[0495] 2-C-methyl-D-erythritol 2,4-cyclic diphosphate synthase (MDS, IspF)
[0496] 4-Hydroxy-3-methylbut-2-en-1-yl diphosphate synthase (HDS, IDS)
[0497] 4-Hydroxy-3-methylbut-2-en-1-yl diphosphate reductase (HDR)
[0498] Other alternative routes for the preparation of IPP and DMAPP are also known, see, for example: Rinaldi, MA, et al. (2022). Natural Product Reports 39(1): 90-118. https: / / doi.org / 10.1039 / D1NP00025J (see Part 3 of this document).
[0499] Therefore, the inventors have provided a complete biocatalytic route for the preparation of compound (I) from acetyl-CoA or glyceraldehyde-3-phosphate (GAP) and pyruvate. This multi-step biocatalytic method is described for the first time and represents a significant advance in the field of compound preparation.
[0500] Reaction conditions of the method of the present invention
[0501] The method of the present invention can be an in vivo method or a biotransformation method.
[0502] The term in vivo method (or whole-cell production, in vivo generation, or in vivo biosynthesis) refers to a method that utilizes metabolically active cells (where primary metabolism is active to produce the precursors of the method of the present invention) (preferably microbial cells) to convert a carbon source into a new compound, such as converting a carbon source into a terpene or a terpene-derived compound.
[0503] Preferred carbon sources are sugars, such as monosaccharides, disaccharides, or polysaccharides. Good carbon sources include, for example, glucose, fructose, mannose, galactose, ribose, sorbitol, ribulose, lactose, maltose, sucrose, raffinose, starch, or cellulose. Sugars can also be added to the culture medium via complex compounds (e.g., molasses) or other byproducts of saccharification. Adding mixtures of different carbon sources is also advantageous. Other possible carbon sources are oils and fats, such as soybean oil, sunflower oil, peanut oil, and coconut oil; fatty acids such as palmitic acid, stearic acid, or linoleic acid; alcohols such as glycerol, methanol, or ethanol; and organic acids such as acetic acid or lactic acid.
[0504] Therefore, these cells contain all the enzymes involved in one or more biosynthetic pathways. At least some of the enzymes involved in this pathway are part of the cell's primary metabolism. For example, the cell may contain enzymes required for pathways that convert carbon sources (e.g., glucose, glycerol, 3-methyl-3-butenol, 3-methyl-2-butenol, CO2) into terpene precursors (e.g., IPP, DMAPP, FPP, GGPP), and enzymes required for pathways that convert terpene precursors into terpenes or terpene-derived molecules (e.g., compounds of formula (I)). These enzymes may be naturally present in the cell or can be produced by modifying the cell.
[0505] It should be noted that prior to this invention, it was not possible to prepare compound (I) using in vivo methods. Prior art at the time of the invention's conception can be found, for example, in WO2016170099 and WO2010139719. Therefore, existing methods for preparing compound (I) are not in vivo methods, and thus this invention provides significant improvements over the methods disclosed therein.
[0506] Alternatively, the method of the present invention can also be carried out under bioconversion (also known as biotransformation) conditions. Bioconversion refers to the method of converting a compound into different products using biological methods or reagents (e.g., enzymes or whole cells, preferably microbial cells). Bioconversion does not include utilizing primary cellular metabolism (as described above) to produce the precursors required for the method of the present invention. A bioconversion method may comprise a multi-step reaction, each step carried out by a different enzyme. The compounds used in a bioconversion method can be extracted from natural sources or prepared by separate chemical or biochemical methods.
[0507] The at least one polypeptide / enzyme present in the biotransformation method of the present invention or a single step of the multi-step method described above can be present in living cells that naturally or recombinantly produce the one or more enzymes, harvested cells, dead cells, permeabilized cells, crude cell extracts, purified extracts, or in substantially pure or completely pure form, i.e., present under biotransformation conditions. Such extracts may contain membrane or liquid components prepared from recombinant host cells expressing at least one polypeptide / enzyme. Cells can be immobilized on suitable substrates according to methods known in the art. The at least one polypeptide / enzyme can be present in solution or as an enzyme immobilized on a carrier. The one or more enzymes can be present simultaneously in soluble and / or immobilized forms.
[0508] Those skilled in the art will understand that using an in vivo method may have advantages.
[0509] In particular, biotransformation methods typically involve multiple steps:
[0510] - Preparation or isolation of the starting compound to be transformed. This compound can be prepared by chemical or biochemical methods, or by extraction from natural sources.
[0511] - The production of enzymes or (living) cells used for biotransformation.
[0512] - A biotransformation reaction that occurs by bringing the compound into contact with an enzyme or (living) cell.
[0513] - Product recycling and purification.
[0514] In contrast, in vivo methods require a limited number of steps, typically limited to:
[0515] - Cultivate microorganisms under conditions suitable for producing the desired compounds.
[0516] - Harvest cells or culture medium and purify the desired compounds.
[0517] In biotransformation methods, such as the biotransformation of compounds of formula (VI), detergents are typically added to promote the dissolution of the compounds or to maximize their contact with the biocatalyst. However, in in vivo methods, both reactants and enzymes are produced within the cells, thus eliminating the need for detergents.
[0518] Therefore, in vivo methods are generally more efficient and economical than biotransformation methods.
[0519] Laboratory methods that can be used in in vivo and biotransformation are well known in the art. Some of these methods will be discussed below.
[0520] The biotransformation method according to the invention can be carried out in common reactors known to those skilled in the art, and can be carried out on various scales, from laboratory scale (a few milliliters to tens of liters of reaction volume) to industrial scale (a few liters to thousands of cubic meters of reaction volume). A chemical reactor can be used if the peptide is used in the form of encapsulation with non-living, optionally permeabilized cells, as a more or less purified cell extract, or in a purified form. Chemical reactors typically allow control over the amount of at least one enzyme, at least one substrate, pH, temperature, and circulation of the reaction medium.
[0521] If the method of the present invention is carried out in vivo, the reaction is preferably carried out in a fermenter, wherein parameters necessary for the survival of living cells (e.g., nutrient-rich culture medium, temperature, aeration, aerobic or anaerobic or other gases, antibiotics, etc.) can be controlled.
[0522] The term “fermentation production” or “fermentation” refers to the ability of microorganisms (assisted by enzyme activity contained in or produced by said microorganisms) to produce compounds in cell cultures using at least one carbon source added to incubation.
[0523] The terms “fermentation broth” or “fermentation medium” should be understood to refer to a liquid, particularly an aqueous solution or aqueous / organic solution, that is based on a fermentation process and has been subjected to or has been subjected to post-processing, such as that described herein.
[0524] Those skilled in the art are familiar with chemical or bioreactors, for example, procedures for scaling up chemical or biotechnological methods from laboratory to industrial scale or optimizing process parameters, which are also extensively described in the literature (for biotechnological methods, see, for example, Crueger und Crueger, Biotechnologie - Lehrbuchder angewandten Mikrobiologie, 2. Ed., R. Oldenbourg Verlag, Munich, Wien, 1984).
[0525] The culture medium used must be appropriately suited to the requirements of each strain. Descriptions of culture media for various microorganisms are provided in the "Manual of Methods for General Bacteriology" of the American Society for Bacteriology (Washington DC, USA, 1981).
[0526] These culture media, which can be used according to the present invention, typically contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.
[0527] Nitrogen sources are typically organic or inorganic nitrogen compounds or materials containing these compounds. Examples of nitrogen sources include ammonia or ammonium salts, such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate, or ammonium nitrate, nitrates, urea, amino acids, or complex nitrogen sources, such as corn steep liquor, soybean flour, soybean protein, yeast extract, meat extract, etc. Nitrogen sources can be used alone or as a mixture.
[0528] Inorganic salt compounds that can be present in culture media include chlorides, phosphorus, or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper, and iron.
[0529] Inorganic sulfur-containing compounds, such as sulfates, sulfites, dithionites, tetrathionites, thiosulfates, and sulfides, as well as organic sulfur compounds, such as thiols and thiols, can be used as sulfur sources.
[0530] Phosphoric acid, potassium dihydrogen phosphate, or dipotassium hydrogen phosphate, or their corresponding sodium-containing salts, can be used as phosphorus sources.
[0531] Chelating agents can be added to the culture medium to retain metal ions in solution. Particularly suitable chelating agents include dihydroxyphenols such as catechol or protocatechuic acid, or organic acids such as citric acid.
[0532] The fermentation medium used according to the present invention typically also contains other growth factors, such as vitamins or growth promoters, including, for example, biotin, riboflavin, thiamine, folic acid, niacin, pantothenic acid, and pyridoxine. The growth factors and salts are often derived from components of complex culture media, such as yeast extract, molasses, corn steep liquor, etc. Furthermore, suitable precursors may be added to the medium. The exact composition of the compounds in the medium depends largely on the specific experiment and is determined individually for each case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, APractical Approach" (1997). Growth media are also available from commercial suppliers, such as Standard 1 (Merck) or BHI (Brain Heart Infusion, DIFCO), etc.
[0533] All components of the culture medium are sterilized by heating (at 1.5 bar and 121°C for 20 minutes) or by aseptic filtration. These components can be sterilized together or individually as needed. All components of the culture medium can be given at the start of the culture or added continuously or in batches.
[0534] The culture temperature is typically between 15°C and 45°C, preferably between 25°C and 40°C, and can be varied or kept constant during the experiment. The pH of the culture medium should be in the range of 5 to 8.5, preferably around 7.0. The pH during growth can be controlled by adding alkaline compounds (e.g., sodium hydroxide, potassium hydroxide, ammonia, or ammonia solution) or acidic compounds (e.g., phosphoric acid or sulfuric acid). Antifoaming agents such as fatty acid polyethylene glycol esters can be used to control foaming. To maintain plasmid stability, suitable selective substances such as antibiotics can be added to the culture medium. To maintain aerobic conditions, oxygen or an oxygen-containing gas mixture (e.g., ambient air) is supplied to the culture. The culture temperature is typically between 20°C and 45°C. Continue culturing until the maximum amount of the desired product is formed. This usually takes between 1 and 160 hours.
[0535] If the method of the present invention is a biotransformation, cells containing at least one enzyme can be permeated by physical or mechanical means (e.g., ultrasound or radiofrequency pulses, French press) or chemical means (e.g., hypotonic media present in the culture medium, lysin, and detergent) or a combination of these methods. Examples of detergents include sodium dodecyl sulfate (SDS), digitalis saponins, n-dodecyl maltodextrin, octyl glycoside, Triton® X-100, Tween® 20, deoxycholate, CHAPS (3-[(3-cholamidopropyl)dimethylammonium]-1-propanesulfonate), Nonidet® P40 (ethylphenol poly(ethylene glycol ether)), etc. As mentioned above, if the method of the present invention is an in vivo method, detergents are not required for the reasons described herein.
[0536] The conversion reaction can be carried out in batches, semi-batches, or continuously. Reactants (and optional nutrients) can be provided at the start of the reaction, or they can be provided semi-continuously or continuously thereafter.
[0537] Depending on the specific reaction type, the biotransformation reaction of the present invention can be carried out in aqueous, aqueous-organic, or non-aqueous reaction media.
[0538] Aqueous or aqueous-organic media may contain suitable buffer solutions to adjust the pH to 5 to 11, such as 6 to 10.
[0539] In aqueous-organic media, organic solvents that are miscible, partially miscible, or immiscible with water can be used. Some non-limiting examples of suitable organic solvents are listed below. Other examples include monohydric or polyhydric alcohols, aromatic or aliphatic alcohols, particularly polyaliphatic alcohols such as glycerol.
[0540] Non-aqueous media can be substantially free of water, that is, may contain less than about 1% by weight or 0.5% by weight of water.
[0541] Biotransformation methods can also be carried out in organic non-aqueous media. Suitable organic solvents include: aliphatic hydrocarbons having, for example, 5 to 8 carbon atoms, such as pentane, cyclopentane, hexane, cyclohexane, heptane, octane, or cyclooctane; aromatic hydrocarbons, such as benzene, toluene, xylene, chlorobenzene, or dichlorobenzene; aliphatic acyclic ethers, such as diethyl ether, methyl tert-butyl ether, ethyl tert-butyl ether, dipropyl ether, diisopropyl ether, dibutyl ether; or mixtures thereof.
[0542] The reactant / substrate concentration can be adjusted according to the optimal biotransformation reaction conditions, depending on the specific enzyme used. For example, the initial substrate concentration can be 0.1 to 0.5 M, such as 10 to 100 mM.
[0543] The temperature of a biotransformation reaction can be adjusted to achieve optimal reaction conditions, depending on the specific enzyme used. For example, the reaction can be carried out at temperatures ranging from 0 to 70°C (e.g., 20 to 50°C or 25 to 40°C). Examples of reaction temperatures include approximately 30°C, approximately 35°C, approximately 37°C, approximately 40°C, approximately 45°C, approximately 50°C, approximately 55°C, and approximately 60°C.
[0544] Biotransformation processes can continue until substrate and product reach equilibrium, but can also be terminated prematurely. Typical process times range from 1 minute to 25 hours, particularly from 10 minutes to 6 hours, for example from 1 hour to 4 hours, and especially from 1.5 hours to 3.5 hours. These parameters are merely non-limiting examples of suitable process conditions.
[0545] Advantageously, microorganisms such as bacteria, fungi, or yeast can be used as host organisms. It is advantageous to use Gram-positive or Gram-negative bacteria, preferably those belonging to the families Enterobacteriaceae, Pseudomonadaceae, Rhizobiaceae, Streptomycetaceae, Streptococcaceae, or Nocardiaceae, with particular preference for bacteria of the genera Escherichia, Pseudomonas, Streptomyces, Lactococcus, Nocardia, Burkholderia, Salmonella, Agrobacterium, Clostridium, or Rhodococcus. The genus and species *Escherichia coli* are highly preferred. Additionally, other advantageous bacteria have been found in the alpha-proteobacteria, beta-proteobacteria, or gamma-proteobacteria groups. Advantageously, yeasts such as *Saccharomyces* or the Pichia family are also suitable hosts.
[0546] Preferably, the cell is a bacterial cell or a fungal cell, particularly a yeast cell. Preferably, the cell is a single-celled organism, a cultured cell derived from a multicellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism. Preferably, the cell is a bacterial cell of the genus *Escherichia*, preferably *Escherichia coli*, or a yeast cell of the genus *Saccharomyces*, preferably *Saccharomyces cerevisiae*, or a yeast cell of the genus *Yarrowia*, preferably *Y. lipolytica*, or a yeast cell of the genus *Pichia*, preferably *P. pastoris*.
[0547] Alternatively, the whole plant or plant cells can be used as a natural or recombinant host. As non-limiting examples, the following plants or cells derived from them may be mentioned: the genus *Nicotiana*, particularly *Nicotiana abenthamiana* and *Nicotiana tabacum* (tobacco); and the genus *Arabidopsis*, particularly *Arabidopsis thaliana*.
[0548] Product separation
[0549] The method of the present invention may further include the step of recovering a final product or intermediate product, which may optionally be a substantially pure form of a stereoisomer or enantiomer. The term "recovery" includes the extraction, harvesting, separation, or purification of a compound from a culture medium or reaction medium. The recovery of a compound can be carried out according to any conventional separation or purification method known in the art, including but not limited to: treatment with conventional resins (e.g., anion or cation exchange resins, nonion adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silica, silica gel, cellulose, alumina, etc.), pH alteration, solvent extraction (e.g., with conventional solvents such as alcohols, ethyl acetate, hexane, etc.), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, lyophilization, etc.
[0550] The identity and purity of the separated products can be determined using known techniques, such as high-performance liquid chromatography (HPLC), gas chromatography (GC), spectroscopy (e.g., IR, UV, NMR), staining methods, TLC, NIRS, enzyme or microbial assays (see, for example: Patek et al. (1994) Appl. Environ. Microbiol. 60:133-140; Malakhova et al. (1996) Biotekhnologiya 11 27-32; und Schmidt et al. (1998) BioprocessEngineer. 19:67-70. Ullmann's Encyclopedia of Industrial Chemistry (1996) Bd.A27, VCH: Weinheim, S. 89-90, S. 521-540, S. 540-547, S. 559-566, 575-581 undS. 581-587; Michal, G (1999) Biochemical Pathways: An Atlas of Biochemistryand Molecular Biology, John Wiley and Sons; Fallon, A. et al. (1987)Applications of HPLC in Biochemistry in: Laboratory Techniques inBiochemistry and Molecular Biology, Bd. 17.).
[0551] Compounds produced by any of the methods described herein can be converted into derivatives, such as, but not limited to, hydrocarbons, esters, amides, glycosides, ethers, epoxides, aldehydes, ketones, alcohols, diols, acetals, or ketals. Terpene derivatives can be obtained by chemical methods, such as, but not limited to, oxidation, reduction, alkylation, acylation, and / or rearrangement. Alternatively, terpene derivatives can be obtained by biochemical methods by contacting the terpene compound with an enzyme, such as, but not limited to, oxidoreductases, monooxygenases, dioxygenases, transferases, or terpene cyclases. Biochemical transformation can be performed in vitro using isolated enzymes, enzymes derived from lysed cells, or whole-cell biotransformation.
[0552] The recombinant cells of the present invention
[0553] As discussed herein, the inventors have innovatively developed an in vivo method for preparing a compound of formula (I). To this end, the inventors constructed a biosynthetic pathway for compound (I) in recombinant cells. Prior to this invention, no one knew of an in vivo method for preparing this compound using recombinant cells. Therefore, prior to this invention, recombinant cells that produce compound (I) were not known in the art. This compound can be present within the recombinant cells or released into the reaction medium.
[0554] Furthermore, as described above in this invention, high selectivity in producing compounds of formula (VI) in the form of formula (VIa) has been achieved, as shown in the accompanying examples.
[0555] This provides a method for preparing compounds of formula (I) in an olfactorily preferred form. Therefore, the method of preparing compounds of formula (I) free from a large number of undesirable byproducts has significant commercial value. As shown in the accompanying examples, over 97% of the compounds of formula (I) are in the form of formula (Ia) and / or formula (Ib).
[0556] Therefore, one embodiment of the present invention is a recombinant cell in which more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
[0557] Therefore, another aspect of the invention provides a recombinant cell that comprises, produces, or is capable of producing a compound of formula (I), as well as one or more compounds of formula (II), formula (III), formula (IV), formula (V), and / or formula (VI).
[0558] This document provides a method for preparing recombinant cells comprising, producing, or capable of producing a compound of formula (I), and one or more compounds of formulas (II), (III), (IV), (V), and / or (VI). Preferably, the cells comprise: (i) a polypeptide having ADH enzyme activity; (ii) a polypeptide having enal lyase activity; (iii) a polypeptide having BVMO enzyme activity; (iv) a polypeptide having esterase activity; and (v) a polypeptide having terpene cyclase activity.
[0559] A preferred embodiment of the present invention is wherein:
[0560] (i) The polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 11 to 21.
[0561] (ii) The polypeptide having enal lyase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 22.
[0562] (iii) The polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 23 to 26 and 216 to 227 (preferably with any of the sequences in SEQ ID NO: 23 to 26);
[0563] (iv) The polypeptide having esterase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any sequence in SEQ ID NO: 27 and 28; and / or,
[0564] (v) The polypeptide having terpene cyclase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 29 to 75, 79 to 89 and 265 to 289 (preferably with any of the sequences in SEQ ID NO: 29 to 75, 265 to 274 and 276 to 289).
[0565] Another preferred embodiment of the present invention is wherein:
[0566] (i) The polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 11 or 21; preferably, the ADH enzyme has the sequence of SEQ ID NO: 11 or 21.
[0567] (ii) The polypeptide having enal lyase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 22; preferably, the enal lyase has the sequence of SEQ ID NO: 22.
[0568] (iii) The polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 25 or 26; preferably, the BVMO enzyme has the sequence of SEQ ID NO: 25 or 26.
[0569] (iv) The polypeptide having esterase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 28; preferably, the esterase has the sequence of SEQ ID NO: 28; and / or,
[0570] (v) A polypeptide having terpene cyclase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 48, 57, 71, 74, 265, 266, 267, 268, 274, 276, 279, 280, 281, 282, 283, 286, 287, or 288. Preferably, the terpene cyclase has the sequence of SEQ ID NO: 48, 57, 71, 74, 265, 266, 267, 268, 274, 276, 279, 280, 281, 282, 283, 286, 287, or 288.
[0571] Alternatively, in this embodiment of the invention:
[0572] (v) The polypeptide having terpene cyclase activity is an SHC enzyme and has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 48, 265, 266, 267, 268, 274, 276, or 279. Preferably, the SHC enzyme has the sequence of SEQ ID NO: 48, 265, 266, 267, 268, 274, 276, or 279.
[0573] Alternatively, in this embodiment of the invention:
[0574] (v) The polypeptide having terpene cyclase activity is a mixed-source terpene cyclase and has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 57, 71, 74, 280, 281, 282, 283, 286, 287, or 288. Preferably, the mixed-source terpene cyclase has the sequence of SEQ ID NO: 57, 71, 74, 280, 281, 282, 283, 286, 287, or 288.
[0575] Another embodiment of the present invention is wherein:
[0576] (i) Compounds of formula (I) are in the form of formula (Ia):
[0577] (Formula Ia)
[0578] (ii) Compounds of formula (II) are in the form of formula (IIa):
[0579] (Form IIa)
[0580] (iii) Compounds of formula (III) are in the form of formula (IIIa):
[0581] (Formula IIIa)
[0582] (iv) Compounds of formula (IV) are in the form of formula (IVa):
[0583] (Form IVa)
[0584] (v) Compounds of formula (V) are in the form of formula (Va):
[0585] (Equation Va)
[0586] (vi) Compounds of formula (VI) are in the form of formula (VIa):
[0587] (Formula VIa).
[0588] The recombinant cell can be any cell suitable for producing compound (I).
[0589] The foregoing provides a list of suitable cells for producing compounds of formula (I) in relation to the method of the present invention, which are also the cells of this morphology of the present invention.
[0590] Preferably, the cell is a bacterial cell or a fungal cell, particularly a yeast cell. Preferably, the cell is a single-celled organism, a cultured cell derived from a multicellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism. Preferably, the cell is a bacterial cell of the genus *Escherichia*, preferably *Escherichia coli*, or a yeast cell of the genus *Saccharomyces*, preferably *Saccharomyces cerevisiae*, or a yeast cell of the genus *Yarrowia*, preferably *Y. lipolytica*, or a yeast cell of the genus *Pichia*, preferably *P. pastoris*.
[0591] Methods for introducing recombinant nucleic acid sequences into such host cells are well known in the art and constitute routine laboratory methods, and need not be described further herein.
[0592] In one embodiment of this invention, the cell further comprises one or more isoprenyltransferases. Preferably, the cell further comprises one or more enzymes with phosphatase activity.
[0593] As described above regarding the method of the present invention, the cells of the present invention may also contain other enzymes to provide the compound of formula (II). The method requires the presence of one or more isoprenyltransferases and one or more enzymes with phosphatase activity.
[0594] Preferably, the isoprenyltransferase is a GGPP synthase. Preferably, the GGPP synthase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 1 or 2.
[0595] Preferably, the phosphatase is a GGPP phosphatase. Preferably, the GGPP phosphatase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any one of SEQ ID NO: 3 to 10.
[0596] Another embodiment of the invention is that the cells of the invention contain enzymes for IPP and DMAPP.
[0597] As described above regarding the method of the present invention, the cells of the present invention may also contain more enzymes to provide compounds of formula (II) via the mevalonate pathway, the methyl erythritol phosphate (MEP) pathway, or alternative pathways for the preparation of IPP and DMAPP.
[0598] In one embodiment of the invention, the cell contains enzymes of the mevalonate pathway:
[0599] Acetyl-CoA acetyltransferase (ACAT)
[0600] 3-Hydroxy-3-methylglutaryl-CoA synthase (HMG-CoA synthase)
[0601] 3-Hydroxy-3-methylglutaryl-CoA reductase (HMG-CoA reductase)
[0602] Mevalonate kinase
[0603] Mevalonate kinase
[0604] Mevalonate diphosphate decarboxylase
[0605] Isopentenyl diphosphate isomerase
[0606] Dimethylallyl diphosphate synthase
[0607] In another embodiment of the invention, the cell contains enzymes of the MEP pathway:
[0608] 1-Deoxy-D-xyulose 5-phosphate synthase (DXS)
[0609] 1-Deoxy-D-xyulose-5-phosphate reductase (DXR)
[0610] 2-C-methyl-D-erythritol 4-phosphate cytidine transferase (MCT, IspD)
[0611] 4-Cytidine diphosphate-2-C-methyl-D-erythritol kinase (CMK, IspE)
[0612] 2-C-methyl-D-erythritol 2,4-cyclic diphosphate synthase (MDS, IspF)
[0613] 4-Hydroxy-3-methylbut-2-en-1-yl diphosphate synthase (HDS, IDS)
[0614] 4-Hydroxy-3-methylbut-2-en-1-yl diphosphate reductase (HDR)
[0615] As described above, when a mixed-origin terpene cyclase is used in the method of the present invention, the resulting compound (I) exhibits isomer preference, tending to generate isomers of compound (I) with preferred olfactory characteristics, namely formula (Ia) and / or formula (Ib), rather than formula (Ic) and / or formula (Id). Therefore, using a mixed-origin terpene cyclase offers significant technical advantages compared to using SHC enzymes. In particular, they demonstrated that the content of the formula I(c) and / or formula I(d) isomers in the compound (I) produced by the method of the present invention is less than 1%. Therefore, using a mixed-origin terpene cyclase offers significant technical advantages compared to using SHC enzymes.
[0616] Therefore, a preferred embodiment of the present invention is that the recombinant cells contain a mixed-source terpene cyclase as a terpene cyclase.
[0617] Preferably, the mixed-origin terpene cyclase is a membrane-integrated mixed-origin terpene cyclase. As can be seen from the accompanying examples, such mixed-origin terpene cyclases preferably produce compounds of formula (Ia) rather than compounds of formula (Ic) and / or formula (Id). Examples of membrane-integrated mixed-origin terpene cyclases include those provided in any one of SEQ ID NO: 50 to 73 and 280 to 289.
[0618] Preferably, the mixed-origin terpene cyclase is a soluble mixed-origin terpene cyclase. As can be seen from the accompanying examples, such mixed-origin terpene cyclases preferably produce compounds of formula (Ib) rather than compounds of formula (Ic) and / or formula (Id). Examples of soluble mixed-origin terpene cyclases include those provided in SEQ ID NO: 74 or 75.
[0619] Preferably, the recombinant cells of the present invention contain more than 97% of the compound of formula (I), which is present in the form of compound (Ia) and / or formula (b), rather than in the form of I(c) and / or I(d).
[0620] The cell culture fermentation medium of the present invention
[0621] As discussed herein, the inventors have innovatively developed an in vivo method for preparing a compound of formula (I). To this end, the inventors constructed a biosynthetic pathway for the compound of formula (I) in recombinant cells, and subsequently cultured these recombinant cells in a suitable cell culture fermentation medium. Prior to this invention, no one knew of an in vivo method for preparing this compound using recombinant cells in a cell culture fermentation medium. Therefore, prior to this invention, cell culture fermentation media comprising the recombinant cells of this invention and / or the compound of formula (I) were not known in the art.
[0622] Furthermore, as described above, in this invention, a high selectivity for producing compounds of formula (VI) in the form of formula (VIa) has been achieved, as shown in the accompanying examples.
[0623] This provides a method for preparing compounds of formula (I) in an olfactorily preferred form. Therefore, the method of preparing compounds of formula (I) free from a large number of undesirable byproducts has significant commercial value. As shown in the accompanying examples, over 97% of the compounds of formula (I) are in the form of formula (Ia) and / or formula (Ib).
[0624] Therefore, one embodiment of the present invention is a cell culture fermentation medium comprising a compound of formula (I), wherein more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
[0625] Therefore, another aspect of the present invention provides a cell culture fermentation medium comprising the recombinant cells as described above. This cell culture fermentation medium may also comprise a compound of formula (I), and / or one or more compounds of formulas (II), (III), (IV), (V), and / or (VI).
[0626] As discussed herein, the inventors have innovatively developed a pathway for the biosynthesis of compound (I) in recombinant cells for the first time. These cells are then cultured under conditions suitable for the production of said compound using a cell culture fermentation medium appropriate for a specific cell type.
[0627] Cell culture fermentation media can be nutrient-rich broths used for cell growth and maintenance during the production phase. The culture conditions used to maintain and propagate various yeast strains may require complex media with specific formulations for cloning and protein expression, as those skilled in the art should understand. For example, commercially available media from Thermo Fisher Scientific can be used. These media can be YPD broth or yeast nitrogen source media. Yeast can be cultured at 30°C in YPD medium or synthetic media.
[0628] Lysozyme (LB) is commonly used to culture bacterial cells. Bacterial cells can develop antibiotic resistance to prevent the growth and contamination of other cells in the culture medium. These cells can carry antibiotic gene cassettes, thus developing resistance to antibiotics such as chloramphenicol, penicillin, kanamycin, and ampicillin.
[0629] Reaction mixture containing the compounds of the present invention
[0630] The method of the present invention can be a biotransformation method.
[0631] As discussed herein, the inventors have innovatively developed a biotransformation method for preparing compounds of formula (I). To this end, the inventors constructed a biosynthetic pathway for compounds of formula (I). Prior to this invention, it was not possible to prepare this compound via biotransformation.
[0632] Furthermore, as described above, in this invention, a high selectivity for producing compounds of formula (VI) in the form of formula (VIa) has been achieved, as shown in the accompanying examples.
[0633] This yields a biotransformation method for preparing compounds of formula (I) with an olfactorily preferred form. Therefore, the present invention, through biotransformation, provides compounds of formula (I) free from a large number of undesirable byproducts and has significant commercial value. As shown in the accompanying examples, over 97% of the compounds of formula (I) are in the form of formula (Ia) and / or formula (Ib).
[0634] Therefore, one embodiment of the present invention is a reaction mixture comprising a compound of formula (I), wherein more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib). The reaction mixture may also comprise one or more compounds of formula (II), formula (III), formula (IV), formula (V) and / or formula (VI).
[0635] Other components of the reaction mixture may include detergents, cofactors, cells, cell debris, cell culture media, and other such components well known to those skilled in the art.
[0636] Preparation of compounds of formula (VI)
[0637] As described above, in order to construct an improved method for preparing compounds of formula (I), the inventors gained a deep understanding of the biochemical pathway by which precursor compounds are generated from this compound via a multi-enzyme reaction. This multi-enzyme reaction is the first time that this stepwise reaction has been used to prepare the compound, representing a significant scientific and commercial advancement in the preparation of sesquiterpenoid compounds of formula (I).
[0638] In addition to preparing compounds of formula (I), the inventors have also devised a method for preparing compounds of formula (VI). This compound is an important commercial precursor for preparing compounds of formula (I) through subsequent biotransformation or chemical methods.
[0639] Therefore, another aspect of the present invention provides a method for preparing compound of formula (VI).
[0640] (VI)
[0641] The compound is in the form of any of its stereoisomers or mixtures thereof, including:
[0642] (i) Contacting compound (IV) with a polypeptide having BVMO enzyme activity to produce compound (V),
[0643] (IV)
[0644] The compound is in the form of any of its stereoisomers or mixtures thereof; and
[0645] (ii) Contacting the compound of formula (V) with a polypeptide having esterase activity to produce the compound of formula (VI),
[0646] (V)
[0647] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0648] One embodiment of the present invention includes the following prior steps:
[0649] (a) Contacting compound (III) with a polypeptide having enal lyase activity to produce compound (IV),
[0650] (III)
[0651] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0652] Another embodiment of the present invention includes the following prior steps:
[0653] (a) Contacting compound (II) with a polypeptide having ADH enzyme activity to produce compound (III),
[0654] (II)
[0655] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0656] As discussed herein, the inventors have innovatively developed a pathway for the biosynthesis of compound (I) in recombinant cells. In constructing this pathway, the inventors also prepared cells capable of producing compound (VI). Therefore, this aspect of the invention has not been disclosed in the prior art.
[0657] As can be understood, the method for preparing compound (VI) includes steps (i) to (iv) of the method for preparing compound (I). Therefore, all embodiments relating to the method for preparing compound (I) herein are applicable to this aspect of the invention, except for polypeptides relating to step (v) of the method for preparing compound (I).
[0658] For the avoidance of doubt, compound (VI) is also known as 4,8,12-trimethyl-1-tridec-3,7,11-trien-1-ol; CAS No. 35826-67-6.
[0659] Compounds of formula (VI) may exist as any of their stereoisomers or mixtures thereof. Specifically, the compound may have the following structures and isomers:
[0660] (Form VIa)
[0661] (3E,7E)-Gaofarneol; (3E,7E)-4,8,12-Trimethyldecadec-3,7,11-trien-1-ol; CAS No. 459-89-2.
[0662] (Formula VIb)
[0663] (3Z,7E)-Gaofarneol; (3Z,7E)-4,8,12-Trimethyldecadec-3,7,11-trien-1-ol; CAS No. 138152-06-4.
[0664] (Formula VIc)
[0665] (3E,7Z)-Gaofarneol; (3E,7Z)-4,8,12-Trimethyldecadec-3,7,11-trien-1-ol; CAS No. 2032064-12-1.
[0666] (Form VId)
[0667] (3Z,7Z)-Gaofarneol; (3Z,7Z)-4,8,12-Trimethyldecadec-3,7,11-trien-1-ol; CAS No. 138152-08-6.
[0668] Furthermore, as described above, the present invention achieves the highly selective production of compounds of formula (VIa) as shown in the accompanying examples. Therefore, a preferred embodiment of the invention is wherein the method prepares compounds of formula (VI), wherein more than 99% of the compounds of formula (VI) are in the form of formula (VIa).
[0669] Another preferred embodiment of the present invention is wherein:
[0670] (i) The polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 25 or 26; preferably, the BVMO enzyme has the sequence of SEQ ID NO: 25 or 26; and / or
[0671] (ii) The polypeptide having esterase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 28; preferably, the esterase has the sequence of SEQ ID NO: 28.
[0672] Another preferred embodiment of the present invention is wherein:
[0673] (i) The polypeptide having ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 11 or 21; preferably, the ADH enzyme has the sequence of SEQ ID NO: 11 or 21; and / or
[0674] (ii) The polypeptide having enal lyase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 22; preferably, the enal lyase has the sequence of SEQ ID NO: 22.
[0675] Preferably, the method is an in vivo method or a biotransformation method.
[0676] This embodiment of the invention also includes a recombinant cell comprising, producing, or capable of producing a compound of formula (VI), and one or more compounds of formulas (II), (III), (IV), and / or (V). The recombinant cell comprises: (i) a polypeptide having ADH enzyme activity; (ii) a polypeptide having enal lyase activity; (iii) a polypeptide having BVMO enzyme activity; and (iv) a polypeptide having esterase activity.
[0677] This embodiment of the invention also includes a method for preparing a formula (VI) compound, the method comprising culturing recombinant cells of this embodiment of the invention under growth conditions suitable for producing a formula (VI) compound as described above.
[0678] This embodiment of the invention also includes a cell culture fermentation medium containing recombinant cells of this embodiment. The cell culture fermentation medium may also contain a compound of formula (VI), and / or one or more compounds of formulas (II), (III), (IV), and / or (V).
[0679] This embodiment of the invention also includes a reaction mixture comprising a compound of formula (VI); preferably, in the form of formula (VIa). More preferably, more than 99% of the compound of formula (VI) is in the form of formula (VIa). The reaction mixture may also comprise one or more compounds of formula (II), formula (III), formula (IV) and / or formula (V).
[0680] This embodiment of the invention also includes compounds of formula (VI), which can be obtained by or can be obtained by the method of this embodiment of the invention, or can be obtained from recombinant cells, cell culture fermentation or reaction mixtures of this embodiment of the invention.
[0681] Preparation of compound (V)
[0682] In addition to preparing compound (I), the inventors have also devised a method for preparing compound (V). This compound is an important commercial precursor for preparing compound (I) through subsequent biotransformation or chemical methods.
[0683] Therefore, another aspect of the present invention provides a method for preparing compound (V).
[0684] (V)
[0685] The compound is in the form of any of its stereoisomers or mixtures thereof, including:
[0686] (i) Contacting the compound of formula (III) with a polypeptide having enal lyase activity to produce the compound of formula (IV),
[0687] (III)
[0688] The compound is in the form of any of its stereoisomers or mixtures thereof; and
[0689] (ii) Contacting compound (IV) with a polypeptide having BVMO enzyme activity to produce compound (V),
[0690] (IV)
[0691] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0692] One embodiment of the present invention includes the following prior steps:
[0693] (a) Contacting compound (II) with a polypeptide having ADH enzyme activity to produce compound (III),
[0694] (II)
[0695] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0696] As discussed herein, the inventors have innovatively developed a pathway for the biosynthesis of compound (I) in recombinant cells. In constructing this pathway, the inventors also prepared cells capable of producing compound (V). Therefore, this aspect of the invention has not been disclosed in the prior art.
[0697] As can be understood, the method for preparing compound (V) includes steps (i) to (iii) of the method for preparing compound (I). Therefore, all embodiments relating to the method for preparing compound (I) herein are applicable to this form of the invention, except for polypeptides relating to steps (iv) and (v) of the method for preparing compound (I).
[0698] For the avoidance of doubt, the compound of formula (V) is also called 4,8,12-trimethyl-1,8,12-trimethyl-1,7,11-trien-1-yl acetate; CAS number 109813-25-4.
[0699] The compound of formula (V) can exist as any of its stereoisomers or mixtures thereof. Specifically, the compound can have the following structures and isomers:
[0700] (Equation Va)
[0701] (3E,7E)-gofarnesyl acetate; (3E,7E)-4,8,12-trimethyldecadec-3,7,11-trien-1-yl acetate; CAS No. 944346-19-4.
[0702] (Formula Vb)
[0703] (3Z,7E)-gofarnesyl acetate; (3Z,7E)-4,8,12-trimethyldecadec-3,7,11-trien-1-yl acetate; CAS No. 1467099-77-9.
[0704] (Formula Vc)
[0705] (3E,7Z)-gofarnesyl acetate; (3E,7Z)-4,8,12-trimethyldeca-3,7,11-trien-1-yl acetate.
[0706] (Formula Vd)
[0707] (3Z,7Z)-Gaofarnesyl acetate; (3Z,7Z)-4,8,12-Trimethyldeca-3,7,11-trien-1-yl acetate.
[0708] Preferably, the method is an in vivo method or a biotransformation method.
[0709] This embodiment of the invention also includes a recombinant cell that comprises, produces, or is capable of producing a compound of formula (V), and one or more compounds of formulas (II), (III), and / or (IV). The recombinant cell comprises: (i) a polypeptide having ADH enzyme activity; (ii) a polypeptide having enal lyase activity; and (iii) a polypeptide having BVMO enzyme activity.
[0710] This embodiment of the invention also includes a method for preparing the compound of formula (V), the method comprising culturing the recombinant cells of this embodiment of the invention under growth conditions suitable for producing the compound of formula (V) as described above.
[0711] This embodiment of the invention also includes a cell culture fermentation medium containing recombinant cells of this embodiment. The cell culture fermentation medium may also contain a compound of formula (V), and / or one or more compounds of formulas (II), (III), and / or (IV).
[0712] This embodiment of the invention also includes a reaction mixture comprising a compound of formula (V). The reaction mixture may also comprise one or more compounds of formula (II), formula (III), and / or formula (IV).
[0713] This embodiment of the invention also includes compounds of formula (V), which can be obtained by or can be obtained by the method of this embodiment of the invention, or can be obtained from recombinant cells, cell culture fermentation or reaction mixtures of this embodiment of the invention.
[0714] Preparation of compounds of formula (IV)
[0715] In addition to preparing the formula (I) compound, the inventors have also devised a method for preparing the formula (IV) compound. This compound is an important commercial precursor for preparing the formula (I) compound through subsequent biotransformation or chemical methods.
[0716] Therefore, another aspect of the present invention provides a method for preparing a (IV) compound.
[0717] (IV)
[0718] The compound is in the form of any of its stereoisomers or mixtures thereof, including:
[0719] (i) Contacting compound (II) with a polypeptide having ADH enzyme activity to produce compound (III),
[0720] (II)
[0721] The compound is in the form of any of its stereoisomers or mixtures thereof; and
[0722] (ii) Contacting the compound of formula (III) with a polypeptide having enal lyase activity to produce the compound of formula (IV),
[0723] (III)
[0724] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0725] As discussed herein, the inventors have innovatively developed a pathway for the biosynthesis of compound (I) in recombinant cells. In constructing this pathway, the inventors also prepared cells capable of producing compound (IV). Therefore, this aspect of the invention has not been disclosed in the prior art.
[0726] As can be understood, the method for preparing compound (IV) includes steps (i) to (ii) of the method for preparing compound (I). Therefore, all embodiments relating to the method for preparing compound (I) herein are applicable to this form of the invention, except for polypeptides relating to steps (iii), (iv) and (v) of the method for preparing compound (I).
[0727] For the avoidance of doubt, the compound of formula (IV) is also called farnesylacetone, namely 6,10,14-trimethylpentadecano-5,9,13-trien-2-one; CAS number 762-29-8.
[0728] Compounds of formula (IV) may exist as any of their stereoisomers or mixtures thereof. Specifically, the compound may have the following structures and isomers:
[0729] (Form IVa)
[0730] (5E,9E)-Farnesylacetone; (5E,9E)-6,10,14-Trimethylpentadecano-5,9,13-trien-2-one; CAS No. 1117-52-8.
[0731] (Formula IVb)
[0732] (5Z,9E)-Farnesylacetone; (5Z,9E)-6,10,14-Trimethylpentadecano-5,9,13-trien-2-one; CAS No. 1117-51-7.
[0733] (Formula IVc)
[0734] (5E,9Z)-Farnesylacetone; (5E,9Z)-6,10,14-Trimethylpentadecano-5,9,13-trien-2-one; CAS No. 3053-35-3.
[0735] (Formula IVd)
[0736] (5Z,9Z)-Farnesylacetone; (5Z,9Z)-6,10,14-Trimethylpentadecan-5,9,13-trien-2-one; CAS No. 3796-69-8.
[0737] Preferably, the method is an in vivo method or a biotransformation method.
[0738] This embodiment of the invention also includes a recombinant cell that comprises, produces, or is capable of producing a compound of formula (IV), and one or more compounds of formula (II) and / or formula (III). The recombinant cell comprises: (i) a polypeptide having ADH enzyme activity; and (ii) a polypeptide having enal lyase activity.
[0739] This embodiment of the invention also includes a method for preparing a formula (IV) compound, the method comprising culturing recombinant cells of this embodiment of the invention under growth conditions suitable for producing a formula (IV) compound as described above.
[0740] This embodiment of the invention also includes a cell culture fermentation medium containing recombinant cells of this embodiment. The cell culture fermentation medium may also contain compounds of formula (IV), and / or one or more compounds of formula (II) and / or formula (III).
[0741] This embodiment of the invention also includes a reaction mixture comprising a compound of formula (IV). The reaction mixture may also comprise one or more compounds of formula (II) and / or formula (III).
[0742] This embodiment of the invention also includes compounds of formula (IV), which can be obtained by or can be obtained by the method of this embodiment of the invention, or can be obtained from recombinant cells, cell culture fermentation or reaction mixtures of this embodiment of the invention.
[0743] Another form of preparative compound (I)
[0744] Another aspect of the present invention provides a method for preparing a compound of formula (I).
[0745] (I)
[0746] The compound is in the form of any of its stereoisomers or mixtures thereof, including:
[0747] (i) Contacting the compound of formula (VI) with a polypeptide having terpene cyclase activity to produce the compound of formula (I),
[0748] (VI)
[0749] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0750] One embodiment of this invention is wherein more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
[0751] Another embodiment of this form of the invention is that the compound of formula (VI) is in the form of formula (VIa).
[0752] Another aspect of the invention is that the polypeptide having terpene cyclase activity is: a polypeptide that is not a squalene cyclase (SHC), and / or a polypeptide that is a squalene cyclase. In the context of this invention, the polypeptide that is not a SHC enzyme is a mixed-origin terpene cyclase.
[0753] Therefore, another aspect of the present invention is that the polypeptide having terpene cyclase activity is a mixed-origin terpene cyclase and / or squalene cyclase.
[0754] As can be understood from this aspect of the invention, the method for preparing compound (I) from compound (VI) includes step (v) of the method for preparing compound (I) described in the first aspect above. Therefore, all embodiments relating to step (v) of the method for preparing compound (I) in the first aspect of the invention herein are applicable to this aspect of the invention.
[0755] In one embodiment of this form of the invention, the squalene cyclase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 29 to 49 and 265 to 279 (preferably, with any of the sequences in SEQ ID NO: 29 to 49, 265 to 274, and 276 to 279). More preferably, the squalene cyclase comprises the sequence provided in any of SEQ ID NO: 29 to 49, 265 to 274, and 276 to 279.
[0756] In another embodiment, the squalene cyclase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 29, 31, 33, 34, 36 to 38, 40, 41, 43 to 46, 48, 49, 265 to 274, and 276 to 279. Preferably, the squalene cyclase comprises the sequence provided in any of SEQ ID NO: 29, 31, 33, 34, 36 to 38, 40, 41, 43 to 46, 48, 49, 265 to 274, and 276 to 279.
[0757] In constructing the method of the present invention, the inventors sought to compare the isomer distribution of the compound of formula (I) synthesized by SHC and mixed-source terpene cyclase.
[0758] To their surprise, they found that when a mixed-origin terpene cyclase was used in the method of the present invention, the resulting compound (I) exhibited isomer preference, with a better olfactory character for isomers of compound (I) than that produced by the SHC enzyme.
[0759] In particular, they demonstrated that less than 1% of the compounds of formula (I) produced by the method of the present invention are isomers of formula I(c) and / or I(d). Therefore, the use of mixed-origin terpene cyclases has a surprising technical advantage compared to the use of SHC enzymes.
[0760] Preferably, the mixed-origin terpene cyclase is a membrane-integrated mixed-origin terpene cyclase. As can be seen from the accompanying examples, such mixed-origin terpene cyclases preferably produce compounds of formula (Ia) rather than compounds of formula (Ic) and / or formula (Id). Examples of membrane-integrated mixed-origin terpene cyclases include those provided in any of SEQ ID NO: 50 to 73 and 280 to 289.
[0761] Preferably, the mixed-origin terpene cyclase is a soluble mixed-origin terpene cyclase. As can be seen from the accompanying examples, such mixed-origin terpene cyclases preferably produce compounds of formula (Ib) rather than compounds of formula (Ic) and / or formula (Id). Examples of soluble mixed-origin terpene cyclases include those provided in SEQ ID NO: 74 or 75.
[0762] This is the first time that a compound of type (I) has been prepared using a mixed-origin terpene cyclase, and the preference for this isomer form is surprising and has significant technical implications for commercial applications.
[0763] Therefore, another aspect of the present invention provides a method for preparing compound (I):
[0764] (I)
[0765] The compound is in the form of any of its stereoisomers or mixtures thereof, including:
[0766] (i) Contacting the compound of formula (VI) with a polypeptide having terpene cyclase activity to produce the compound of formula (I),
[0767] (VI)
[0768] The compound is in the form of any of its stereoisomers or mixtures thereof, wherein the polypeptide having terpene cyclase activity is not an SHC enzyme.
[0769] In a preferred embodiment of this form of the invention, the polypeptide that is not an SHC enzyme is a mixed-origin terpene cyclase.
[0770] As can be understood, mixed-origin terpene cyclases (or polypeptides having mixed-origin terpene cyclase activity) have been described in detail above in conjunction with the first aspect of the present invention, and these contents are also incorporated into this aspect of the present invention.
[0771] Preferably, the mixed-source terpene cyclase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 50 to 75 and 280 to 289. Preferably, the mixed-source terpene cyclase comprises the sequence provided in any of SEQ ID NO: 50 to 75 and 280 to 289.
[0772] Preferably, the mixed-source terpene cyclase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 57, 71, 74, 280, 281, 282, 283, 286, 287, or 288. Preferably, the mixed-source terpene cyclase comprises the sequence provided by any one of SEQ ID NO: 57, 71, 74, 280, 281, 282, 283, 286, 287, or 288.
[0773] Preferably, the hybrid terpene cyclase is a membrane-integrated hybrid terpene cyclase. As can be seen from the accompanying examples, such hybrid terpene cyclases preferably produce compounds of formula (Ia) rather than those of formulas (Ic) and / or (Id). Therefore, in a preferred embodiment, the hybrid terpene cyclase is a membrane-integrated hybrid terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 50 to 73 and 280 to 289.
[0774] Preferably, the mixed-origin terpene cyclase is a soluble mixed-origin terpene cyclase. As can be seen from the accompanying examples, such mixed-origin terpene cyclases preferably produce compounds of formula (Ib) rather than those of formula (Ic) and / or formula (Id). Therefore, in another preferred embodiment, the mixed-origin terpene cyclase is a soluble mixed-origin terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 74 and 75.
[0775] As can be understood from this aspect of the invention, in order to prepare the compound of formula (I), the method may further include one or more steps performed prior to step (i), which have been described in detail above in conjunction with the first aspect of the invention, and these descriptions are incorporated into this aspect of the invention.
[0776] Preferably, the method is an in vivo method or a biotransformation method.
[0777] This embodiment of the invention also includes a recombinant cell that comprises, produces, or is capable of producing a compound of formula (I); optionally, more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or (Ib). In one embodiment, the recombinant cell comprises or is capable of functionally expressing, or is functionally expressing, a polypeptide having SHC enzyme activity as described above and / or a polypeptide having mixed-origin terpene cyclase activity.
[0778] This embodiment of the invention also includes a method for preparing a compound of formula (I), the method comprising culturing recombinant cells of this embodiment of the invention under growth conditions suitable for producing a compound of formula (I).
[0779] This embodiment of the invention also includes a cell culture fermentation medium containing recombinant cells of this embodiment. The cell culture fermentation medium may further contain a compound of formula (I); optionally, more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib). The cell culture fermentation medium may also contain a compound of formula (VI).
[0780] This embodiment of the invention also includes a reaction mixture comprising a compound of formula (I); optionally, more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or (Ib). The reaction mixture may also comprise a compound of formula (VI).
[0781] This embodiment of the invention also includes compounds of formula (I), which are obtained or can be obtained by the method of this embodiment of the invention, or can be obtained from recombinant cells, cell culture fermentation media or reaction mixtures of this embodiment of the invention.
[0782] This embodiment of the invention also includes compound (I), wherein more than 97% of the compound is in the form of formula (Ia) and / or (Ib).
[0783] Another aspect of the invention is the use of mixed-source terpene cyclases to produce compounds of formula (I) and / or their derivatives.
[0784] Another form of preparative compound (I)
[0785] As described above, in order to construct an improved method for preparing compounds of formula (I), the inventors gained a deep understanding of the biochemical pathway by which precursor compounds are generated from this compound via a multi-enzyme reaction. This multi-enzyme reaction is the first time that this stepwise reaction has been used to prepare the compound, representing a significant scientific and commercial advancement in the preparation of sesquiterpenoid compounds of formula (I).
[0786] In addition to the method for preparing compound (I) from compound (II), the inventors have also devised a method for preparing compound (I) from compound (V). This is also a method for preparing compound (I) with significant commercial value.
[0787] Therefore, another aspect of the present invention provides a method for preparing compound (I):
[0788] (I)
[0789] The compound is in the form of any of its stereoisomers or mixtures thereof, including:
[0790] (i) Contacting the compound of formula (V) with a polypeptide having esterase activity to produce the compound of formula (VI),
[0791] (V)
[0792] The compound is in the form of any of its stereoisomers or mixtures thereof; and
[0793] (ii) Contacting the compound of formula (VI) with a polypeptide having terpene cyclase activity to produce the compound of formula (I),
[0794] (VI)
[0795] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0796] The foregoing description of esterases (or polypeptides with esterase activity) and terpene cyclases (or polypeptides with terpene cyclase activity) in conjunction with the first embodiment of the present invention is incorporated herein by reference.
[0797] In this embodiment of the invention, the method for preparing the compound of formula (I) uses the compound of formula (V) as a starting material. This method can be an in vivo method or a biotransformation method. Preferably, the method is a biotransformation method.
[0798] The preferred embodiment of this form is, wherein:
[0799] (i) The polypeptide having esterase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 27 or 28 (preferably with SEQ ID NO: 28); and / or
[0800] (ii) The polypeptide having terpene cyclase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or higher sequence identity with any of the sequences in SEQ ID NO: 29 to 75, 79 to 89 and 265 to 289 (preferably with any of the sequences in SEQ ID NO: 29 to 75, 265 to 274 and 276 to 289; more preferably with any of the sequences in SEQ ID NO: 29 to 75, 265 to 274 and 276 to 289; even more preferably with any of the sequences in SEQ ID NO: 48, 57, 71, 74, 265, 266, 267, 268, 274, 276, 279, 280, 281, 282, 283, 286, 287 or 288).
[0801] This embodiment of the invention also includes a recombinant cell comprising one or more compounds of formula (I), formula (V), and / or formula (VI). The recombinant cell may also comprise (i) a polypeptide with esterase activity and (ii) a polypeptide with terpene cyclase activity.
[0802] This embodiment of the invention also includes a method for preparing a compound of formula (I), the method comprising culturing recombinant cells of this embodiment of the invention under growth conditions suitable for producing a compound of formula (I).
[0803] This embodiment of the invention also includes a cell culture fermentation medium containing recombinant cells of this embodiment. The cell culture fermentation medium may also contain one or more compounds of formula (I), formula (V), and / or formula (VI).
[0804] This embodiment of the invention also includes a reaction mixture comprising a compound of formula (I). The reaction mixture may also comprise one or more compounds of formula (V) and / or formula (VI).
[0805] This embodiment of the invention also includes compounds of formula (I), which can be obtained by or can be obtained by the method of this embodiment of the invention, or can be obtained from recombinant cells, cell culture fermentation media or reaction mixtures of this embodiment of the invention.
[0806] Furthermore, as described above, the present invention achieves the highly selective production of compound (VIa) in the form of compound (VIa), as demonstrated in the accompanying examples. Subsequently, a biotransformation method is provided, by which compound (I) having an olfactorily preferred form is prepared. Therefore, the biotransformation method of the present invention for preparing compound (I) free from a large number of undesirable byproducts has significant commercial value.
[0807] Therefore, and as shown in the accompanying embodiments, one embodiment of this form of the invention is that more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
[0808] Another embodiment of this type of invention is wherein the method includes the following prior steps:
[0809] (a) Contacting compound (IV) with a polypeptide having BVMO enzyme activity to produce compound (V),
[0810] (IV)
[0811] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0812] Another embodiment of this type of invention includes an additional prior step:
[0813] (a) Contacting compound (III) with a polypeptide having enal lyase activity to produce compound (IV),
[0814] (III)
[0815] The compound is in the form of any of its stereoisomers or mixtures thereof.
[0816] Examples of BVMO and enal lyase described above in the first aspect of the invention can also be used in this aspect of the invention.
[0817] Phosphatase used in the method of the present invention
[0818] As described above, one embodiment of the method of the first aspect of the present invention further includes one or more biocatalytic steps prior to step (i) for preparing a compound of formula (II), said biocatalytic steps comprising:
[0819] (a) Preparation of geranylgeranyl diphosphate (GGPP) from IPP and DMAPP using one or more isoprenyltransferases; and / or,
[0820] (b) Using one or more enzymes with phosphatase activity, compound (II) is prepared from GGPP.
[0821] In constructing the method of the present invention, the inventors screened for phosphatases that could be used in step (b) of the method. This is the first time that these enzymes have been demonstrated to catalyze the step of preparing compound (II) from GGPP.
[0822] Therefore, this aspect of the invention includes the use of phosphatases for the preparation of compounds of formula (II).
[0823] Another aspect of the invention provides the use of a phosphatase for preparing a compound of formula (II) by GGPP, the phosphatase having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 3 to 10.
[0824] Preferably, the phosphatase has 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 3, 4, 5, 7, and 8. More preferably, the phosphatase comprises the sequences provided in SEQ ID NO: 3, 4, 5, 7, and 8.
[0825] ADH enzyme used in the method of the present invention
[0826] As described herein, the method of the first aspect of the present invention includes step (i): preparing compound (III) by contacting compound (II) with an ADH enzyme.
[0827] In constructing the method of the present invention, the inventors screened out ADH enzymes that could be used in step (i) of the method. This is the first time that these enzymes have been demonstrated to catalyze the step of preparing compound (III) from compound (II).
[0828] Therefore, another aspect of the invention provides the use of an ADH enzyme for the preparation of a compound of formula (III) from a compound of formula (II), the ADH enzyme having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 11 to 21. Preferably, the ADH enzyme comprises the sequences provided in SEQ ID NO: 11 to 21.
[0829] Another aspect of the invention provides the use of an ADH enzyme for the preparation of a compound of formula (III) from a compound of formula (II), the ADH enzyme having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 11 or 21. Preferably, the ADH enzyme comprises the sequence provided in SEQ ID NO: 11 or 21.
[0830] Enal lyase used in the method of the present invention
[0831] As described herein, the method of the first aspect of the present invention includes step (ii): preparing compound (IV) by contacting compound (III) with an enal lyase.
[0832] In constructing the method of the present invention, the inventors screened for an enal lyase that could be used in step (ii) of the method. This is the first demonstration that this enzyme can catalyze the step of preparing compound (IV) from compound (III).
[0833] Therefore, another aspect of the invention provides the use of a polypeptide having enal lyase activity for the preparation of compound (IV). Another aspect of the invention provides the use of a polypeptide having enal lyase activity for the preparation of compound (IV) from compound (III). Another aspect of the invention is the use of a polypeptide having enal lyase activity for the preparation of compounds (IV), (V), (VI), (I), and / or their derivatives.
[0834] Therefore, another aspect of the invention provides the use of a polypeptide having enal lyase activity for the preparation of compound (IV), the polypeptide having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 22. Another aspect of the invention provides the use of a polypeptide having enal lyase activity for the preparation of compound (IV) from compound (III), the polypeptide having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 22. Another aspect of the invention provides the use of a polypeptide having enal lyase activity for the preparation of compounds of formula (IV), (V), (VI), (I) and / or derivatives thereof, the polypeptide having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 22.
[0835] Preferably, the polypeptide having enal lyase activity comprises the sequence provided in SEQ ID NO: 22.
[0836] BVMO enzyme used in the method of the present invention
[0837] As described herein, the method of the first aspect of the present invention includes step (iii): preparing compound (V) by contacting compound (IV) with BVMO enzyme.
[0838] In constructing the method of the present invention, the inventors screened for BMVO enzymes that could be used in step (iii) of the method. This is the first demonstration that these enzymes are capable of catalyzing the preparation of compound (V) from compound (IV).
[0839] Therefore, another aspect of the invention provides the use of a BVMO enzyme for the preparation of a compound of formula (V) from a compound of formula (IV), the enzyme having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 23 to 26 and 216 to 227.
[0840] Preferably, the BVMO enzyme has 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 23 to 26. More preferably, the BVMO enzyme comprises any of the sequences provided in SEQ ID NO: 23 to 26.
[0841] Preferably, the BVMO enzyme has 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 25 or 26. More preferably, the BVMO enzyme comprises the sequence provided in SEQ ID NO: 25 or 26.
[0842] Esterase used in the method of the present invention
[0843] As described herein, the method of the first aspect of the present invention includes step (iv): preparing compound (VI) by contacting compound (V) with an esterase.
[0844] In constructing the method of the present invention, the inventors screened for esterases that could be used in step (iv) of the method. This is the first demonstration that these enzymes are capable of catalyzing the preparation of compound (VI) from compound (V).
[0845] Therefore, another aspect of the invention provides the use of an esterase for preparing a compound of formula (VI) from a compound of formula (V), the esterase having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with SEQ ID NO: 27 or 28.
[0846] Preferably, the esterase comprises the sequence provided in SEQ ID NO: 27 or 28.
[0847] SHC enzyme according to the present invention
[0848] In constructing the method of this invention, the inventors screened for novel polypeptide sequences encoding SHC enzymes that can be used in the method. Therefore, these polypeptides are also part of this invention.
[0849] In studying SHC enzymes, the inventors screened and identified enzymes with amino acids alanine at position 437 and methionine at position 600 relative to the sequence provided in SEQ ID NO: 82, which have particular uses.
[0850] Therefore, another embodiment of the invention is a mutant SHC enzyme, wherein the 437th amino acid is alanine and the 600th amino acid is methionine, relative to the sequence provided in SEQ ID NO: 82. This is the first demonstration that such a mutant combination can function in the reaction described in step (v) of the method of the first embodiment of the invention.
[0851] SHC enzymes have been described above; for example, they have been described in step (v) of the method of the first embodiment of the present invention. Using the information provided in this part of the specification, those skilled in the art can screen any SHC enzyme and modify it using standard laboratory techniques to obtain the mutant SHC enzyme described in this embodiment of the present invention.
[0852] A preferred embodiment of this form of the invention is that the mutant SHC enzyme is a polypeptide having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 29, 31, 33, 34, 36 to 38, 40, 41, 43 to 46, 48, 49, 265 to 274, and 276 to 279, wherein the polypeptide has alanine at amino acid position 437 and methionine at amino acid position 600 relative to the sequence provided in SEQ ID NO: 82. In a preferred embodiment, the mutant SHC enzyme is a polypeptide having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 29, 31, 33, 34, 36 to 38, 40, 41, 43 to 46, 48, 49, 265, 266, 267, 268, 274, 276, and 279, wherein the polypeptide has alanine at amino acid position 437 and methionine at amino acid position 600 relative to the sequence provided in SEQ ID NO: 82. Preferably, the mutant SHC enzyme is a polypeptide having the amino acid sequence provided in any one of SEQ ID NO: 29, 31, 33, 34, 36 to 38, 40, 41, 43 to 46, 48, 49, 265 to 274, and 276 to 279. Preferably, the mutant SHC enzyme is a polypeptide having the amino acid sequence provided in any one of SEQ ID NO: 29, 31, 33, 34, 36 to 38, 40, 41, 43 to 46, 48, 49, 265, 266, 267, 268, 274, 276, and 279. This embodiment of the invention also includes polypeptide fragments, variants, and functional equivalents thereof.
[0853] The polypeptide of this invention is an SHC enzyme, which can be used in step (v) of the method of the first form of the invention or any other form of the invention including the preparation of compound (I) from compound (VI).
[0854] Another aspect of the invention provides a nucleic acid sequence encoding a mutant SHC enzyme, wherein the mutant SHC enzyme, relative to the sequence provided in SEQ ID NO: 82, has alanine at amino acid position 437 and methionine at amino acid position 600. Methods for preparing such a nucleic acid sequence are known in the art.
[0855] A preferred embodiment of this form of the invention is that the nucleic acid sequence encoding the mutant SHC enzyme is a nucleic acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 123, 125, 126, 128 to 131, 134 to 137, 139 to 142, 144 to 148, 150 to 153, 290 to 299, and 301 to 304, wherein the polypeptide encoded by the nucleic acid sequence has, relative to the sequence provided in SEQ ID NO: 82, amino acid position 437 being alanine and amino acid position 600 being methionine. Preferably, the nucleic acid sequence is any one of SEQ ID NO: 123, 125, 126, 128 to 131, 134 to 137, 139 to 142, 144 to 148, 150 to 153, 290 to 299, and 301 to 304. This embodiment of the invention also includes expression vectors, expression cassettes, and other related technologies comprising the nucleic acid sequence of the invention.
[0856] Furthermore, as described herein, in the method of the present invention comprising preparing compound (I) by contacting compound (VI) with a terpene cyclase, the terpene cyclase may be an SHC enzyme. In constructing the method of the present invention, the inventors screened for SHC enzymes suitable for use in the method. This is the first demonstration that these enzymes are capable of catalyzing the step of preparing compound (I) from compound (VI).
[0857] Therefore, another aspect of the invention provides the use of an SHC enzyme for the preparation of a compound of formula (I) from a compound of formula (VI), the SHC enzyme having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 29 to 49, 79 to 89, and 265 to 279.
[0858] Preferably, the SHC enzyme has 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 29 to 49 and 265 to 279. More preferably, the SHC enzyme comprises the sequence provided in any of SEQ ID NO: 29 to 49 and 265 to 279.
[0859] Preferably, the SHC enzyme has 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 29 to 49, 265 to 274, and 276 to 279. More preferably, the SHC enzyme comprises the sequence provided in any of SEQ ID NO: 29 to 49, 265 to 274, and 276 to 279.
[0860] The mixed-origin terpene cyclase according to the present invention
[0861] In developing the method of this invention, the inventors screened for novel polypeptide sequences encoding mixed-origin terpene cyclases that can be used in the method. Therefore, these polypeptides are also part of this invention.
[0862] One embodiment of this form of the invention is a mutant hybrid terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 56 to 70. Preferably, the mutant hybrid terpene cyclase has the amino acid sequence provided in any of SEQ ID NO: 56 to 70. This form of the invention also includes polypeptide fragments, variants, and functional equivalents thereof.
[0863] Therefore, this embodiment of the invention provides a nucleic acid sequence encoding a mutant hybrid terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 165 to 185. Preferably, the nucleic acid sequence is any of the nucleic acid sequences provided in SEQ ID NO: 165 to 185. This embodiment of the invention also includes expression vectors, expression cassettes, and other related technologies comprising the nucleic acid sequence of the invention.
[0864] In studying mixed-origin terpene cyclases, the inventors screened and identified an enzyme having an amino acid substitution at the 9th amino acid position relative to the sequence provided in SEQ ID NO:51.
[0865] Therefore, another aspect of the invention is a mutant mixed-origin terpene cyclase having an amino acid substitution at the 9th amino acid position relative to the sequence provided in SEQ ID NO: 51. This is the first demonstration that a mixed-origin terpene cyclase with this mutation can function in the step of preparing compound (I) from compound (VI).
[0866] The mixed-origin terpene cyclase has been described above in step (v) of the method of the first embodiment of the present invention. Using the information provided in this part of the specification, those skilled in the art can screen for any mixed-origin terpene cyclase and modify it using standard laboratory techniques to obtain the mutant mixed-origin terpene cyclase of this embodiment of the present invention.
[0867] Preferably, the mutant mixed-origin terpene cyclase has a substitution at the 9th amino acid position relative to the sequence provided in SEQ ID NO: 51, where cysteine, methionine, or threonine is introduced.
[0868] Therefore, a preferred embodiment of this form of the invention is that the mutant hybrid terpene cyclase is a polypeptide having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 56 to 61, 69, and 70, wherein the mutant hybrid terpene cyclase has an amino acid substitution at the 9th amino acid position relative to the sequence provided in SEQ ID NO: 51. Preferably, the mutant hybrid terpene cyclase has the amino acid sequence provided in any of SEQ ID NO: 56 to 61, 69, and 70. This form of the invention also includes polypeptide fragments, variants, and functional equivalents thereof.
[0869] The polypeptide of this invention is a mixed-origin terpene cyclase, which can be used in step (v) of the method of the first form of the invention or any other form of the invention including the preparation of compound (I) from compound (VI).
[0870] Another aspect of the invention provides a nucleic acid sequence encoding a mutant hybrid terpene cyclase having an amino acid substitution at the 9th amino acid position relative to the sequence provided in SEQ ID NO: 51. Methods for preparing such a nucleic acid sequence are known in the art.
[0871] A preferred embodiment of this embodiment of the invention is that the nucleic acid sequence encoding the mutant hybrid terpene cyclase is a nucleic acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 165 to 176, 184, and 185, wherein the polypeptide encoded by the nucleic acid sequence has an amino acid substitution at the 9th amino acid position relative to the sequence provided in SEQ ID NO: 51. Preferably, the nucleic acid sequence is any of the nucleic acid sequences provided in SEQ ID NO: 165 to 176, 184, and 185. This embodiment of the invention also includes expression vectors, expression cassettes, and other such related technologies comprising the nucleic acid sequence of the invention.
[0872] Furthermore, as described herein, in the method of the present invention comprising preparing compound (I) by contacting a compound of formula (VI) with a terpene cyclase, the terpene cyclase may be a mixed-source terpene cyclase. In constructing the method of the present invention, the inventors screened for mixed-source terpene cyclases suitable for use in the method. This is the first demonstration that these mixed-source terpene cyclases are capable of catalyzing the step of preparing compound (I) from compound (VI).
[0873] Therefore, another aspect of the invention provides the use of a mixed-source terpene cyclase for the preparation of a compound of formula (I) from a compound of formula (VI), the mixed-source terpene cyclase having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 50 to 75 and 280 to 289.
[0874] Preferably, the mixed-source terpene cyclase comprises the sequence provided in any one of SEQ ID NO: 50 to 75 and 280 to 289.
[0875] The polypeptides and nucleic acids of the present invention or used in the methods of the present invention
[0876] The commonly used terms “polypeptide” or “peptide” refer to a natural or synthetic, continuous, peptide-linked linear chain or sequence of amino acid residues containing approximately 10 to more than 1,000 residues. Short-chain polypeptides with up to 30 residues are also called “oligopeptides”.
[0877] The term "protein" refers to a large molecular structure composed of one or more polypeptides. The amino acid sequence of the polypeptide represents the protein's "primary structure." The amino acid sequence also predetermines the protein's "secondary structure" by forming specific structural elements, such as the α-helices and β-sheets formed within the polypeptide chain. The arrangement of multiple such secondary structural elements defines the protein's "tertiary structure," or spatial arrangement. If a protein contains more than one polypeptide chain, these chains are arranged spatially to form the protein's "quaternary structure." Proper spatial arrangement, or "folding," is a prerequisite for protein function. Denaturation or unfolding disrupts protein function. If this disruption is reversible, protein function can be restored by refolding.
[0878] The typical protein function referred to in this article is "enzyme function," which means that a protein acts as a biocatalyst on a substrate, such as a compound, and catalyzes the conversion of said substrate into a product. Enzymes can exhibit high or low levels of substrate and / or product specificity.
[0879] Therefore, the term "peptide" as used in this article to refer to a specific "activity" implicitly refers to a correctly folded protein that exhibits the indicated activity, such as a specific enzyme activity.
[0880] Therefore, unless otherwise stated, the term "polypeptide" also covers the terms "protein" and "enzyme".
[0881] Similarly, the term "peptide fragment" encompasses the terms "protein fragment" and "enzyme fragment".
[0882] The term "isolated polypeptide" refers to an amino acid sequence extracted from its natural environment by any method known in the art or a combination of such methods, including recombinant, biochemical and synthetic methods.
[0883] A "target peptide" is an amino acid sequence that targets a protein or polypeptide to intracellular organelles (i.e., mitochondria or plastids) or to the extracellular space (secretory signal peptides). The nucleic acid sequence encoding the target peptide can be fused to the nucleic acid sequence encoding the amino terminus (e.g., the N-terminus) of a protein or polypeptide, or it can be used to replace the natural targeting peptide.
[0884] The present invention also relates to “functional equivalents” (also referred to as “analogs” or “functional mutations”) of the polypeptides specifically described herein.
[0885] For example, a “functional equivalent” refers to a polypeptide that, in a test used to determine enzyme activity, shows at least 1 to 10%, or at least 20%, or at least 50%, or at least 75%, or at least 90% higher or lower activity than the polypeptide specifically described herein.
[0886] According to the invention, "functional equivalents" also encompass specific mutants that have an amino acid at at least one sequence position in the amino acid sequence described herein that differs from the specifically stated amino acid, but still possess one of the aforementioned biological activities, such as enzyme activity. Thus, "functional equivalents" include mutants obtainable by the addition, substitution, particularly conserved substitution, deletion, and / or inversion of one or more, for example, 1 to 20, 1 to 15, or 5 to 10 amino acids, wherein said changes can occur at any sequence position, as long as they result in the mutant possessing the general profile of the properties of the invention. Functional equivalence is also particularly provided if the activity pattern qualitatively overlaps between the mutant and the unaltered polypeptide, i.e., if, for example, an interaction with the same agonist or antagonist or substrate is observed, but at different rates (i.e., by EC...). 50 or IC 50 (Values or any other parameters suitable in this technical field). The table below shows examples of suitable (conservative) amino acid substitutions:
[0887]
[0888] The “functional equivalents” in the above sense are also the “precursors” of the polypeptides described in this article, as well as the “functional derivatives” and “salts” of the polypeptides.
[0889] In this case, a "precursor" is a natural or synthetic precursor of a polypeptide, which may or may not have the desired biological activity.
[0890] The term "salt" as used in this invention refers to salts of the carboxyl group of the protein molecule and salts formed by the acid addition of the amino group. Salts of the carboxyl group can be produced in known ways and include inorganic salts such as sodium, calcium, ammonium, iron, and zinc salts, as well as salts formed with organic bases such as amines, such as triethanolamine, arginine, lysine, piperidine, etc. Salts formed by acid addition, such as those formed with inorganic acids such as hydrochloric acid or sulfuric acid, and salts formed with organic acids such as acetic acid and oxalic acid, are also covered by this invention.
[0891] The “functional derivatives” of the polypeptides according to the invention can also be generated at the side groups of functional amino acids or their N-terminus or C-terminus using known techniques. Such derivatives include, for example: aliphatic esters of carboxylic acid groups, amides of carboxylic acid groups, which can be obtained by reacting with ammonia or with primary or secondary amines; N-acyl derivatives of free amino groups, which are generated by reacting with acyl groups; or O-acyl derivatives of free hydroxyl groups, which are generated by reacting with acyl groups.
[0892] "Functional equivalents" naturally include polypeptides that can be obtained from other organisms as well as naturally occurring variants. For example, the area of homologous sequence regions can be determined by sequence comparison, and equivalent polypeptides can be determined based on the specific parameters of this invention.
[0893] "Functional equivalents" also include "fragments" of the polypeptide according to the invention, such as a single domain or sequence motif, or in the form of an N-terminal and / or C-terminal truncation, which may or may not exhibit the desired biological function. Preferably, such "fragments" at least qualitatively retain the desired biological function.
[0894] Furthermore, a “functional equivalent” is a fusion protein having one of the polypeptide sequences described herein or a functional equivalent derived therefrom, and having at least one additional functionally distinct heterologous sequence in functional N-terminal or C-terminal association (i.e., without substantial mutual functional impairment of the fusion protein portion). Non-limiting examples of such heterologous sequences are, for example, signal peptides, histidine anchors, or enzymes.
[0895] The invention also includes “functional equivalents” that are homologs of the specifically disclosed polypeptides. They have at least 60%, preferably at least 75%, particularly at least 80 or 85%, such as 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% homology (or identity) with the specifically disclosed amino acid sequence, calculated using the algorithm described in Pearson and Lipman, Proc. Natl. Acad. Sci. (USA) 85(8), 1988, 2444-2448. The homology or identity of the homologous polypeptides according to the invention, expressed as a percentage, refers in particular to identity expressed as a percentage of amino acid residues based on the total length of one of the amino acid sequences specifically described herein.
[0896] Identity data expressed as a percentage can also be determined using BLAST alignment, the blastp (protein-protein BLAST) algorithm, or by applying the Clustal settings detailed below.
[0897] In the case of possible protein glycosylation, the “functional equivalents” according to the invention include polypeptides in deglycosylated or glycosylated forms as described herein, as well as modified forms that can be obtained by changing the glycosylation pattern.
[0898] Functional equivalents or homologs of the polypeptides according to the present invention can be generated by mutagenesis, for example by point mutation, elongation or shortening of the protein, or as described in more detail below.
[0899] Functional equivalents or homologs of the polypeptides according to the invention can be identified by screening a database of mutants, such as shortened mutants. For example, a database of protein variant diversity can be generated by combinatorial mutagenesis at the nucleic acid level, for example by enzymatic ligation of a mixture of synthetic oligonucleotides. Numerous methods are available for generating a database of potential homologs from degenerate oligonucleotide sequences. The chemical synthesis of degenerate gene sequences can be performed in an automated DNA synthesizer, and the synthesized gene can then be ligated into a suitable expression vector. The use of degenerate genomes makes it possible to provide all sequences in a mixture that encode the desired set of potential protein sequences. Methods for synthesizing degenerate oligonucleotides are known to those skilled in the art.
[0900] In the prior art, several techniques are known for screening gene products from combinatorial databases generated by point mutations or shortening, and for screening cDNA libraries containing gene products with selected properties. These techniques can be applied to rapidly screen gene libraries generated by combinatorial mutagenesis of homologs according to the present invention. The most commonly used high-throughput analysis-based techniques for screening large gene libraries involve cloning the gene library in a reproducible expression vector, transforming suitable cells with the resulting vector database, and expressing the combinatorial gene under specific conditions, under which detection of the desired activity facilitates the isolation of vectors encoding the gene (whose product is detected). Recursive integration mutagenesis (REM) is a technique for increasing the frequency of functional mutants in a database and can be used in conjunction with screening tests to identify homologs.
[0901] The embodiments provided herein offer orthologs and paralogs of the disclosed peptides, as well as methods for identifying and isolating such orthologs and paralogs. The definitions of the terms "ortholog" and "paralog" are given below and apply to both amino acid and nucleic acid sequences.
[0902] The polypeptides of this invention comprise all active forms of the enzymes of this invention, including active subsequences, such as catalytic domains or active sites. In one embodiment, this invention provides the catalytic domains or active sites as described below. In one embodiment, the present invention provides a peptide or polypeptide comprising or composed of an active site domain predicted by using a database such as Pfam (http: / / pfam.wustl.edu / hmmsearch.shtml) (a large collection covering multiple sequence alignments and hidden Markov models for many common protein families, Pfam Protein Family Database, A. Bateman, E. Birney, L. Cerruti, R. Durbin, L. Etwiller, SR Eddy, S. Griffiths-Jones, KL Howe, M. Marshall, and ELL Sonnhammer, NucleicAcids Research, 30(1):276-280, 2002) or equivalent resources such as the InterPro and SMART databases (http: / / www.ebi.ac.uk / interpro / scan.html, http: / / smart.embl-heidelberg.de / ).
[0903] The present invention also covers “peptide variants” having the desired activity, wherein the variant peptide is selected from amino acid sequences having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with a particular (particularly natural) amino acid sequence (represented by a particular SEQ ID NO), and contains at least one substitution modification relative to the SEQ ID NO.
[0904] According to the applicable coding nucleic acid sequence
[0905] In the context of this article, the following definitions apply:
[0906] The terms “nucleic acid sequence,” “nucleic acid,” “nucleic acid molecule,” and “polynucleotide” are used interchangeably and refer to a sequence of nucleotides. A nucleic acid sequence can be a single-stranded or double-stranded deoxyribonucleotide or ribonucleotide of any length and includes coding and non-coding sequences of genes, exons, introns, sense and antisense complementary sequences, genomic DNA, cDNA, miRNA, siRNA, mRNA, rRNA, tRNA, recombinant nucleic acid sequences, isolated and purified naturally occurring DNA and / or RNA sequences, synthetic DNA and RNA sequences, fragments, primers, and nucleic acid probes. Those skilled in the art understand that the nucleic acid sequence of RNA is identical to that of DNA, except that thymine (T) is replaced by uracil (U). The term “nucleotide sequence” should also be understood to include polynucleotide or oligonucleotide molecules that are present as individual fragments or as components of larger nucleic acids.
[0907] "Isolated nucleic acid" or "isolated nucleic acid sequence" refers to a nucleic acid or nucleic acid sequence that exists in an environment different from that of naturally occurring nucleic acids or nucleic acid sequences, and may include those that are substantially free of endogenous contaminants.
[0908] As used in this article, the term "naturally occurring" for nucleic acids refers to a nucleic acid that is found in the cells of organisms in nature and has not been intentionally modified by humans in a laboratory.
[0909] A “fragment” of a polynucleotide or nucleic acid sequence refers to a continuous sequence of nucleotides, particularly of a length of at least 15 bp, at least 30 bp, at least 40 bp, at least 50 bp, and / or at least 60 bp, according to one embodiment of this invention. Specifically, the polynucleotide fragment comprises at least 25, more particularly at least 50, more particularly at least 75, more particularly at least 100, more particularly at least 150, more particularly at least 200, more particularly at least 300, more particularly at least 400, more particularly at least 500, more particularly at least 600, more particularly at least 700, more particularly at least 800, more particularly at least 900, and more particularly at least 1000 consecutive nucleotides of a polynucleotide sequence according to one embodiment of this invention. Without limitation, the polynucleotide fragments described herein can be used as PCR primers and / or probes, or for antisense gene silencing or RNAi.
[0910] "Recombinant nucleic acid sequences" refer to nucleic acid sequences generated by combining genetic material from more than one source using laboratory methods (such as molecular cloning), thereby creating or modifying nucleic acid sequences that are not naturally occurring and cannot be found in biological organisms in any other way.
[0911] “Recombinant DNA technology” refers to molecular biological methods used to prepare recombinant nucleic acid sequences, as described, for example, in Laboratory Manuals, edited by Weigel and Glazebrook, 2002, Cold Spring Harbor Lab Press; and Sambrook et al., 1989, Cold Spring Harbor, NY: Cold Spring Harbor Laboratory Press.
[0912] The term "gene" refers to a DNA sequence that contains a region that is transcribed into an RNA molecule, such as mRNA, in a cell and is operatively linked to a suitable regulatory region, such as a promoter. Therefore, a gene can contain multiple operatively linked sequences, such as a promoter, a 5' leader sequence (containing, for example, a sequence involved in translation initiation), a coding region of cDNA or genomic DNA, introns, exons, and / or a 3' untranslated sequence (containing, for example, a transcription termination site).
[0913] "Polycistronic" refers to nucleic acid molecules, especially mRNA, that can encode more than one polypeptide within the same nucleic acid molecule.
[0914] A "chimeric gene" is any gene that is not normally found in species in nature, particularly a gene in which one or more portions of the nucleic acid sequence are unrelated in nature. For example, a promoter that is unrelated in nature to part or all of the transcriptional region or to another regulatory region. The term "chimeric gene" should be understood to include expression constructs in which a promoter or transcriptional regulatory sequence is operatively linked to one or more coding sequences or antisense (i.e., the inverse complementary strand of the sense strand) or inverted repeat sequences (sense and antisense, whereby the RNA transcript forms a double-stranded RNA post-transcriptionally). The term "chimeric gene" also includes genes obtained by combining portions of one or more coding sequences to produce new genes.
[0915] The “3' URT” or “3' untranslated sequence” (also known as the “3' untranslated region” or “3' end”) refers to a nucleic acid sequence found downstream of the gene coding sequence that contains, for example, a transcription termination site and (in most, but not all, eukaryotic mRNAs) a polyadenylation signal, such as AAUAAA or its variants. After transcription termination, the mRNA transcript can be cleaved downstream of the polyadenylation signal and a poly(A) tail can be added, which is involved in the transport of mRNA to the translation site, such as the cytoplasm.
[0916] The term "primer" refers to a short nucleic acid sequence that is hybridized to a template nucleic acid sequence and used for the polymerization of nucleic acid sequences complementary to that template.
[0917] The term "selectable marker" refers to any gene that, after expression, can be used to select one or more cells containing that selectable marker. Examples of selectable markers are described below. Those skilled in the art will understand that different antibiotic, fungicide, auxotrophic, or herbicide selectable markers may be applicable to different target species.
[0918] This invention also relates to nucleic acid sequences encoding polypeptides as defined herein.
[0919] In particular, the present invention also relates to nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, such as cDNA, genomic DNA and mRNA) encoding one of the aforementioned polypeptides and their functional equivalents, which can be obtained, for example, by using artificial nucleotide analogs.
[0920] This invention relates to both isolated nucleic acid molecules encoding polypeptides or biologically active regions thereof according to the invention, and nucleic acid fragments that can be used, for example, as hybridization probes or primers for identifying or amplifying nucleic acids encoded according to the invention.
[0921] This invention also relates to nucleic acids that have a degree of "identity" with the sequences specifically disclosed herein. "Identity" between two nucleic acids refers to the identity of nucleotides along the entire length of the nucleic acid in each case.
[0922] The “identity” between two nucleotide sequences (and similarly, peptide or amino acid sequences) is a function of the number of nucleotide residues (or amino acid residues) that produce the alignment of the two sequences, or the number of identical residues in the two sequences. Identical residues are defined as the same residues at a given position in the alignment of the two sequences. The percentage of sequence identity used in this paper is calculated from the best alignment by dividing the number of identical residues between the two sequences by the total number of residues in the shortest sequence and multiplying by 100. The best alignment is the alignment with the highest probability of identity percentage. Vacancies can be introduced into one or more positions in the alignment of one or both sequences to obtain the best alignment. These vacancies are then considered as dissimilar residues used to calculate the percentage of sequence identity. Alignments used to determine the percentage of identity of amino acid or nucleic acid sequences can be performed in various ways using computer programs, such as those publicly available on the Internet.
[0923] Specifically, the BLAST program (Tatiana et al., FEMS Microbiol Lett., 1999, 174:247-250, 1999), which is available from the National Center for Biotechnology Information (NCBI) at http: / / www.ncbi.nlm.nih.gov / BLAST / bl2seq / wblast2.cgi with default parameters, can be used to obtain the best alignment of protein or nucleic acid sequences and calculate the percentage of sequence identity.
[0924] Alternatively, identity can be determined according to the method of Chenna et al. (2003), webpage: http: / / www.ebi.ac.uk / Tools / clustalw / index.html# and the following settings:
[0925] DNA gap openness deduction: 15.0
[0926] DNA gap extension deduction: 6.66
[0927] DNA matrix: identity
[0928] Open protein gaps: Deduction 10.0
[0929] Protein gap extension deduction: 0.2
[0930] Protein matrix: Gonnet
[0931] Protein / DNA ENDGAP: -1
[0932] Protein / DNA GAPDIST: 4
[0933] All nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, such as cDNA and mRNA) mentioned in this article can be produced from nucleotide structural units by chemical synthesis in a known manner, for example, by fragment condensation of individual overlapping complementary nucleic acid structural units of a double helix. The chemical synthesis of oligonucleotides can be carried out, for example, by the phosphoramide process (Voet, Voet, 2nd edition, Wiley Press, New York, pages 896-897) in a known manner. The accumulation of synthetic oligonucleotides, and the filling of vacancies by means of the Klenow fragment of DNA polymerase and ligation reactions, as well as general cloning techniques, are described in Sambrook et al. (1989), see below.
[0934] The nucleic acid molecules according to the present invention may also contain untranslated sequences from the 3' and / or 5' ends of the coding genetic region.
[0935] The present invention also relates to nucleic acid molecules that are complementary to the nucleotide sequences or segments thereof specifically described.
[0936] The nucleotide sequences according to the invention enable the generation of probes and primers that can be used to identify and / or clone homologous sequences in other cell types and organisms. Such probes or primers typically contain a nucleotide sequence region that hybridizes to at least about 12, preferably at least about 25, such as about 40, 50 or 75 consecutive nucleotides of the sense strand or corresponding antisense strand of the nucleic acid sequence according to the invention under “stringent” conditions (as defined elsewhere herein).
[0937] "Homologous" sequences include orthologous or paralogous sequences. Methods for identifying orthologous or paralogous sequences include phylogenetic methods, sequence similarity methods, and hybridization methods known in the art and described herein.
[0938] "Paralleloids," or paralogous sequences, arise from gene replication, resulting in two or more genes with similar sequences and functions. Paralogs typically cluster together and form through gene replication within related plant species. Paralogs are identified in groups of similar genes using pairwise BLAST analysis or procedures such as CLUSTAL during phylogenetic analysis of gene families. In paralogs, the shared sequence can be identified as a sequence characteristic of the related gene and possessing a similar function.
[0939] "Orthologs," or orthologous sequences, are sequences that are similar to each other because they are found in species descended from a common ancestor. For example, plant species known to share a common ancestor contain many enzymes with similar sequences and functions. For instance, by constructing a phylogenetic tree of a gene family for a species using procedures such as CLUSTAL or BLAST, a technician can identify orthologous sequences and predict the functions of orthologs. One method for identifying or confirming similar functions between homologous sequences is by comparing transcript profiles in host cells or organisms (such as plants or microorganisms) that overexpress or lack (in gene knockout / reduction) the relevant polypeptide. A technician can understand that genes with similar transcript profiles (having a common transcript with greater than 50% regulation, or a common transcript with greater than 70% regulation, or a common transcript with greater than 90% regulation) will have similar functions. By causing host cells, organisms such as plants or microorganisms to produce terpene synthase proteins, homologs, parahomologs, orthologs, and any other variants of the sequences described herein are expected to function in a similar manner.
[0940] Nucleic acid molecules according to the invention can be isolated using standard molecular biology techniques and the sequence information provided according to the invention. For example, cDNA can be isolated from a suitable cDNA library using one of the specifically disclosed complete sequences or fragments thereof as hybridization probes and standard hybridization techniques (e.g., described in Sambrook, (1989)).
[0941] Alternatively, nucleic acid molecules containing one or a fragment of the disclosed sequence can be isolated by polymerase chain reaction using oligonucleotide primers constructed based on that sequence. The nucleic acid amplified in this manner can be cloned into a suitable vector and characterized by DNA sequencing. The oligonucleotides according to the invention can also be prepared using standard synthetic methods, for example, using an automated DNA synthesizer.
[0942] To test the function of a variant DNA sequence according to one embodiment of the present invention, the target sequence is operatively linked to an optional or screenable marker gene, and the expression of the reporter gene is tested in a transient expression analysis using microorganisms or protoplasts or in stably transformed plants.
[0943] The present invention also relates to derivatives of specifically disclosed or derivable nucleic acid sequences.
[0944] Therefore, the additional nucleic acid sequences according to the invention may be derived from the sequences specifically disclosed herein and may be distinguished by one or more (e.g., 1 to 10) nucleotides, such as 1 to 20, particularly 1 to 15 or 5 to 10, additions, substitutions, insertions or deletions, and may also encode polypeptides having the desired properties.
[0945] The invention also includes nucleic acid sequences containing so-called silent mutations or altered sequences, depending on the codon usage of a particular original or host organism, compared to the specifically stated sequences.
[0946] According to a specific embodiment of the invention, variant nucleic acids can be prepared to adapt their nucleotide sequences to a specific expression system. For example, bacterial expression systems are known to express polypeptides more efficiently if the amino acids are encoded by specific codons. Due to the degeneracy of the genetic code, more than one codon can encode the same amino acid sequence, and multiple nucleic acid sequences can encode the same protein or polypeptide; all these DNA sequences are covered in one embodiment herein. Where appropriate, the nucleic acid sequence encoding the polypeptide described herein can be optimized to increase expression in host cells. For example, the nucleic acid of one embodiment herein can be synthesized using host-specific codons to improve expression.
[0947] The present invention also covers naturally occurring variants of the sequences described herein, such as splice variants or allelic variants.
[0948] The allele variants have at least 60% homology across the entire amino acid range at the derived amino acid level, preferably at least 80%, and very particularly preferably at least 90% homology (for details regarding homology at the amino acid level, please refer to the information given above for peptides). Advantageously, the homology may be even higher in certain regions of the sequence.
[0949] The present invention also relates to sequences that can be obtained by conserved nucleotide substitution (i.e., as a result, the amino acid in question is replaced by an amino acid having the same charge, size, polarity and / or solubility).
[0950] This invention also relates to molecules derived from specifically disclosed nucleic acids through sequence polymorphism. Such genetic polymorphism can exist in cells from different populations or from cells within a single population due to natural allelic variations. Allelic variants may also include functional equivalents. These natural variations typically produce changes of 1–5% in the nucleotide sequence of a gene. The polymorphism can lead to alterations in the amino acid sequence of the polypeptides disclosed herein. Allelic variants may also include functional equivalents.
[0951] Furthermore, derivatives should also be understood as homologs of the nucleic acid sequences according to the present invention, such as homologs of animals, plants, fungi, or bacteria, shortened sequences, single-stranded DNA or RNA encoding or non-coding DNA sequences. For example, at the DNA level, the homolog has at least 40%, preferably at least 60%, particularly preferably at least 70%, and very particularly preferably at least 80% homology in the entire DNA region given in the sequence specifically disclosed herein.
[0952] Furthermore, derivatives should be understood as, for example, fusions with promoters. Promoters added to the nucleotide sequence can be modified by at least one nucleotide exchange, at least one insertion, inversion, and / or deletion, without impairing the function or effectiveness of the promoter. Moreover, the effectiveness of promoters can be increased by altering their sequence, or by completely exchanging them with more effective promoters or even promoters from different genera of organisms.
[0953] Generation of functional polypeptide mutants
[0954] Furthermore, those skilled in the art are familiar with methods for generating functional mutants, namely, a nucleotide sequence encoding a polypeptide having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any amino acid-related SEQ ID NO disclosed herein; and / or encoded by a nucleic acid molecule containing a nucleotide sequence having at least 50% sequence identity with any nucleotide-related SEQ ID NO disclosed herein.
[0955] Depending on the techniques used, those skilled in the art can introduce completely random or more targeted mutations into gene or non-coding nucleic acid regions (e.g., those important for regulating expression) and subsequently generate a genetic library. The molecular biological methods required for this purpose are known to those skilled in the art, for example, as described in Sambrook and Russell, Molecular Cloning, 3rd Edition, Cold Spring Harbor Laboratory Press, 2001.
[0956] Methods for modifying genes and thereby modifying the polypeptides encoded by them have long been known to those skilled in the art, for example:
[0957] - Site-specific mutagenesis, in which single or multiple nucleotides of a gene are replaced in a directed manner (Trower MK (Ed.) 1996; In vitro mutagenesis protocols. Humana Press, New Jersey).
[0958] - Saturation mutagenesis, in which the codon of any amino acid can be exchanged or added at any site in the gene (Kegler-Ebo DM, Docktor CM, DiMaio D (1994) Nucleic Acids Res 22:1593; Barettino D, Feigenbutz M, Valcárel R, Stunnenberg HG (1994) Nucleic Acids Res 22:541; Barik S (1995) Mol Biotechnol 3:1).
[0959] - Error-prone polymerase chain reaction, in which the nucleotide sequence is mutated by error-prone DNA polymerase (Eckert KA, KunkelTA (1990) Nucleic Acids Res 18:3739).
[0960] -SeSaM method (sequence saturation method), in which preferred exchanges are prevented by polymerase. Schenk et al., Biospektrum, Vol. 3, 2006, 277-279.
[0961] - Gene propagation in mutant strains, where, for example, due to defects in DNA repair mechanisms, the mutation rate of nucleotide sequences increases (Greener A, Callahan M, Jerpseth B (1996) An efficient random mutagenesis technique using an E. coli mutator strain. In: Trower MK (Ed.) In vitro mutagenesis protocols. Humana Press, New Jersey), or
[0962] -DNA shuffling, in which a group of closely related genes are formed and digested, and these fragments are used as templates for polymerase chain reactions, in which the full-length mosaic gene is eventually generated through repeated strand separation and recombination (Stemmer WPC (1994) Nature 370:389; Stemmer WPC (1994) Proc Natl Acad Sci USA 91:10747).
[0963] Using so-called directed evolution (particularly described in Reetz MT and Jaeger KE (1999), TopicsCurr Chem 200:31; Zhao H, Moore JC, Volkov AA, Arnold FH (1999), Methods for optimizing industrial polypeptides by directed evolution, In: Demain AL, Davies JE (Ed.) Manual of industrial microbiology and biotechnology. American Society for Microbiology), skilled workers can mass-produce functional mutants in a directed manner. To this end, in the first step, gene libraries of the respective polypeptides are first generated, for example, using the methods given above. The gene libraries are expressed in a suitable manner, for example, through bacterial or phage display systems.
[0964] The relevant genes in the host organism expressing the functional mutant (whose function largely corresponds to the desired trait) can be submitted to another mutation cycle. The mutation and selection or screening steps can be repeated iteratively until the functional mutant of the present invention possesses a sufficient degree of the desired trait. Using this iterative process, a limited number of mutations, such as 1, 2, 3, 4, or 5 mutations, can be performed in stages, and their effects on the activity under study can be evaluated and selected. The selected mutants can then be subjected to further mutation steps in the same manner. This significantly reduces the number of individual mutants to be studied.
[0965] The results of this invention also provide important information regarding the structure and sequence of the relevant polypeptides, which is essential for the targeted generation of other polypeptides with desired modified properties. In particular, so-called "hot spots" can be defined as sequence segments potentially suitable for modification of properties by introducing targeted mutations.
[0966] Information about the location of amino acid sequences can also be derived, where mutations that may have little effect on activity can occur, and these can be designated as potential “silent mutations”.
[0967] Used to express the polypeptides of the present invention and / or to constructs for use in the methods of the present invention.
[0968] In the context of this article, the following definitions apply:
[0969] "Gene expression" encompasses both "heterologous expression" and "overexpression," and involves gene transcription and the translation of mRNA into proteins. Overexpression refers to the production of gene products, measured as mRNA, peptide, and / or enzyme activity levels, in transgenic cells or organisms exceeding the levels found in non-transformed cells or organisms with similar genetic backgrounds.
[0970] As used herein, an "expression vector" refers to a nucleic acid molecule engineered using molecular biology methods and recombinant DNA technology to deliver foreign or exogenous DNA into a host cell. Expression vectors typically include the sequences required for correct transcription of the nucleotide sequence. The coding region usually encodes the target protein, but it can also encode RNA, such as antisense RNA, siRNA, etc.
[0971] As used herein, “expression vector” includes any linear or circular recombinant vector, including but not limited to viral vectors, bacteriophages, and plasmids. Those skilled in the art can select a suitable vector based on the expression system. In one embodiment, the expression vector includes a nucleic acid of the embodiments described herein, operably linked to at least one “regulatory sequence” that controls transcription, translation, initiation, and termination, such as a transcription promoter, operon, or enhancer, or an mRNA ribosome binding site, and optionally includes at least one selection marker. When the regulatory sequence functionally relates to the nucleic acid of the embodiments described herein, the nucleotide sequence is “operably linked.”
[0972] As used herein, “expression system” encompasses any combination of nucleic acid molecules required to express one, or co-express two or more, polypeptides in vivo or in vitro in a given expression host. The respective coding sequences may reside on a single nucleic acid molecule or vector, such as a vector containing multiple cloning sites, or on a polycistronic nucleic acid, or may be distributed across two or more physically distinct vectors. As a specific example, an operon may be mentioned comprising a promoter sequence, one or more operon sequences, and one or more structural genes, each encoding the enzyme described herein.
[0973] As used herein, the terms “amplifying” and “amplification” refer to the use of any suitable amplification method to generate or detect recombinants of naturally expressed nucleic acids, as described in detail below. For example, the present invention provides methods and reagents (e.g., specific degenerate oligonucleotide primer pairs, oligo-dT primers) for amplifying (e.g., by polymerase chain reaction, PCR) naturally expressed (e.g., genomic DNA or mRNA) or recombinant nucleic acids (e.g., cDNA) of the present invention in vivo, in vitro, or in vitro.
[0974] A "regulatory sequence" refers to a nucleic acid sequence that determines the expression level of the nucleic acid sequence in the embodiment described herein and can regulate the transcription rate of a nucleic acid sequence operatively linked to that regulatory sequence. Regulatory sequences include promoters, enhancers, transcription factors, promoter elements, etc.
[0975] According to the present invention, "promoter," "nucleic acid with promoter activity," or "promoter sequence" should be understood as referring to a nucleic acid that, when functionally linked to a nucleic acid to be transcribed, regulates the transcription of said nucleic acid. "Promoter" specifically refers to a nucleic acid sequence that controls the expression of a coding sequence by providing a binding site for RNA polymerase and other factors suitable for transcription, including but not limited to transcription factor binding sites, repressor and activator protein binding sites. The term "promoter" also includes the term "promoter regulatory sequence." A promoter regulatory sequence may include upstream and downstream elements that may affect transcription, RNA processing, or the stability of the associated coding nucleic acid sequence. Promoters include naturally derived and synthetic sequences. The coding nucleic acid sequence is typically located downstream of the promoter relative to the transcription direction initiating from the transcription start site.
[0976] In this context, "functional" or "operationally" linked is understood, for example, to refer to the sequential arrangement of one of the nucleic acids having a regulatory sequence. For example, a sequence with promoter activity, and the nucleic acid sequence to be transcribed, along with optional other regulatory elements (e.g., nucleic acid sequences that ensure transcription) and, for example, a terminator, arranged such that each regulatory element can perform its function after transcription of the nucleic acid sequence. This does not necessarily require a direct chemical link. Genetic control sequences, such as enhancer sequences, can even act on the target sequence from more distant locations or even from other DNA molecules. A preferred arrangement is one where the nucleic acid sequence to be transcribed is located downstream (i.e., at the 3' end) of the promoter sequence, thereby covalently linking the two sequences together. The distance between the promoter sequence and the nucleic acid sequence to be recombined can be less than 200 base pairs, or less than 100 base pairs, or less than 50 base pairs.
[0977] In addition to promoters and terminators, other examples of regulatory elements include: target sequences, enhancers, polyadenylation signals, selectable markers, amplification signals, and origins of replication. Suitable regulatory sequences are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).
[0978] The term "constitutive promoter" refers to an unregulated promoter that allows for the continuous transcription of nucleic acid sequences that are operatively linked to it.
[0979] As used herein, the term "operably linked" refers to the linking of polynucleotide elements in a functional relationship. A nucleic acid is "operably linked" when it is in a functional relationship with another nucleic acid sequence. For example, if a promoter or transcriptional regulatory sequence can influence the transcription of a coding sequence, then the promoter or transcriptional regulatory sequence is operably linked to the coding sequence. Operable linking means that the linked DNA sequences are typically adjacent. The nucleotide sequence associated with the promoter sequence can be homologous or heterologous in origin relative to the plant to be transformed. The sequence can also be wholly or partially synthetic. Regardless of origin, the nucleic acid sequence associated with the promoter sequence will be expressed or silenced depending on the nature of the promoter linked after binding to the polypeptide of the embodiments described herein. The associated nucleic acid can encode a protein that needs to be expressed or repressed throughout the organism or in a specific tissue, cell, or cell compartment at all times or alternatively at specific times. This nucleotide sequence specifically encodes a protein that confers the desired phenotypic trait to the host cell or organism altered or transformed by it. More specifically, the associated nucleotide sequence results in the production of one or more target products as defined herein in the cell or organism. In particular, the nucleotide sequence encodes a polypeptide having enzymatic activity as defined herein.
[0980] The nucleotide sequences described above can be part of an "expression cassette". The terms "expression cassette" and "expression construct" are used synonymously. A (preferred recombinant) expression construct contains a nucleotide sequence that encodes a polypeptide according to the invention and that the polypeptide is genetically controlled by a regulatory nucleic acid sequence.
[0981] In the method applied according to the present invention, the expression cassette may be part of an "expression vector", particularly part of a recombinant expression vector.
[0982] According to the present invention, "expression unit" should be understood as a nucleic acid with expressive activity, which contains a promoter as defined herein, and regulates expression upon functional linkage with a nucleic acid or gene to be expressed, i.e., transcription and translation of said nucleic acid or gene. Therefore, it is also referred to in this respect as a "regulatory nucleic acid sequence". In addition to promoters, other regulatory elements, such as enhancers, may also be present.
[0983] According to the present invention, an "expression cassette" or "expression construct" should be understood as an expression unit functionally linked to a nucleic acid or gene to be expressed. Therefore, in contrast to an expression unit, an expression cassette contains not only nucleic acid sequences that regulate transcription and translation, but also nucleic acid sequences that are expressed as proteins due to transcription and translation.
[0984] In the context of this invention, the terms "expression" or "overexpression" describe the generation or increase of intracellular activity of one or more polypeptides encoded by corresponding DNA in a microorganism. For this purpose, for example, a gene may be introduced into the organism, an existing gene may be replaced with another gene, the copy number of a gene may be increased, a strong promoter may be used, or a gene encoding a corresponding polypeptide with high activity may be used. Optionally, these measures may be combined.
[0985] Preferably, such constructs according to the invention include a promoter upstream of the respective coding sequence 5' and a terminator sequence downstream of the respective coding sequence 3', as well as optionally other common regulatory elements, each operatively connected to the coding sequence.
[0986] The nucleic acid constructs according to the invention specifically comprise a sequence encoding a polypeptide, such as derived from the amino acid-related SEQ ID NO or its inverse complementary sequence as described herein, or derivatives and homologs thereof, and is operatively or functionally linked to one or more regulatory signals for advantageous control, for example, increasing gene expression.
[0987] In addition to these regulatory sequences, the natural regulation of these sequences may still exist before the actual structural gene, and optionally may have been genetically modified so that natural regulation is turned off and gene expression is enhanced. However, nucleic acid constructs can also have simpler constructions, i.e., no additional regulatory signals are inserted before the coding sequence, and the natural promoter with regulatory function has not been removed. Instead, the natural regulatory sequences are mutated so that regulation no longer occurs and gene expression increases.
[0988] Preferred nucleic acid constructs advantageously also include one or more previously mentioned "enhancer" sequences functionally linked to a promoter, which enable enhanced expression of the nucleic acid sequence. Other advantageous sequences, such as other regulatory elements or terminators, may also be inserted at the 3' end of the DNA sequence. One or more copies of the nucleic acid according to the invention may be present in the construct. Optionally, other markers, such as genes complementary to auxotrophic or antibiotic resistance genes, may also be present in the construct for selection.
[0989] Examples of suitable regulatory sequences exist in promoters, such as cos, tac, trp, tet, trp-tet, lpp, lac, lpp-lac, and lacI. q , T7, T5, T3, gal, trc, ara, rhaP (rhaP BAD SP6, lambda-P R Or lambda-P LIn promoters, they are advantageously used in Gram-negative bacteria. Other advantageous regulatory sequences are found, for example, in the Gram-positive promoters amy and SpO2, and in yeast or fungal promoters ADC1, MFalpha, AC, P-60, CYC1, GAPDH, TEF, rp28, and ADH. Artificial promoters can also be used for regulation.
[0990] To facilitate expression in a host organism, nucleic acid constructs are advantageously inserted into vectors, such as plasmids or phages, enabling optimal gene expression in the host. Besides plasmids and phages, vectors should be understood to include all other vectors known to those skilled in the art, such as viruses like SV40, CMV, baculoviruses, and adenoviruses, transposons, IS elements, phages, granules, and linear or circular DNA or artificial chromosomes. These vectors are capable of autonomous replication in the host organism or replication via chromosomes. These vectors represent a further development of the invention. Binary or CPO integration vectors are also suitable.
[0991] Suitable plasmids include, for example, those for *E. coli* pLG338, pACYC184, pBR322, pUC18, pUC19, pKC30, pRep4, pHS1, pKK223-3, pDHE19.2, pHS2, pPLc236, pMBL24, pLG200, pUR290, and pIN-III. 113 -B1, λgt11, or pBdCI; Streptomyces pIJ101, pIJ364, pIJ702, or pIJ361; Bacillus pUB110, pC194, or pBD214; Corynebacterium pSA77 or pAJ667; Fungi pALS1, pIL2, or pBB116; Yeast 2alphaM, pAG-1, YEp6, YEp13, or pEMBLYe23; or Plant pLGV23, pGHlac + The plasmids mentioned above are a small selection of possible plasmids, including pBIN19, pAK2004, and pDH51. Other plasmids are well known to those skilled in the art and can be found, for example, in the book Cloning Vectors (Eds. Pouwels PH et al. Elsevier, Amsterdam-New York-Oxford, 1985, ISBN 0 444 904018).
[0992] In further development of the vector, vectors containing the nucleic acid constructs of the present invention or the nucleic acids of the present invention can also be advantageously introduced into microorganisms in the form of linear DNA and integrated into the genome of the host organism via heterologous or homologous recombination. This linear DNA can consist of linearized vectors such as plasmids, or solely of the nucleic acid constructs or nucleic acids of the present invention.
[0993] To achieve optimal expression of heterologous genes in an organism, it is advantageous to modify the nucleic acid sequence to match the specific "codon usage habits" used in the organism. These "codon usage habits" can be readily determined through computer evaluation of other known genes in the organism under discussion.
[0994] The expression cassette according to the invention is generated by fusing a suitable promoter to a suitable coding nucleotide sequence and a terminator or polyadenylation signal. Conventional recombination and cloning techniques are used for this purpose, as described in, for example, T. Maniatis, EF Fritsch and J. Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1989) and TJ Silhavy, ML Berman and LW Enquist, Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1984) and Ausubel, FM et al., Current Protocols in Molecular Biology, Greene Publishing Assoc. and Wiley Interscience (1987).
[0995] To facilitate expression in a suitable host organism, recombinant nucleic acid constructs or gene constructs are advantageously inserted into host-specific vectors, enabling optimal gene expression in the host. Vectors are well-known to those skilled in the art and can be found, for example, in "cloning vectors" (Pouwels PH et al., Ed., Elsevier, Amsterdam-New York-Oxford, 1985).
[0996] Alternative embodiments of the present invention provide a method for “altering gene expression in host cells.” For example, in certain contexts (e.g., exposure to certain temperatures or culture conditions), polynucleotides of the present invention can be enhanced, overexpressed, or induced in host cells or host organisms.
[0997] The altered expression of the polynucleotides described herein can also result in ectopic expression, which is a different expression pattern in altered and control or wild-type organisms. The alteration in expression occurs due to the interaction of the peptide of one embodiment of this invention with an exogenous or endogenous regulator or due to chemical modification of the peptide. The term also refers to the altered expression pattern of the polynucleotides of the embodiments described herein, which is altered to below detectable levels or whose activity is completely inhibited.
[0998] In one embodiment, this document also provides isolated, recombinant, or synthetic polynucleotides encoding the polypeptide or variant polypeptide provided herein.
[0999] In one embodiment, multiple nucleic acid sequences encoding polypeptides are co-expressed in a single host, particularly under the control of different promoters. In another embodiment, multiple nucleic acid sequences encoding polypeptides may be present on a single transformation vector, or separate vectors may be used and transformants containing two chimeric genes may be selected for simultaneous co-transformation. Similarly, one or more polypeptide-encoding genes may be expressed together with other chimeric genes in a single plant, cell, microorganism, or organism.
[1000] Recombination generation of peptides according to the present invention
[1001] The present invention further relates to a method for recombinantly producing polypeptides or functional biologically active fragments thereof according to the invention, wherein microorganisms that produce polypeptides are cultured, and optionally, expression of the polypeptides is induced by applying at least one inducer for gene expression, and the polypeptides are isolated from the culture. If desired, the polypeptides can also be produced on an industrial scale in this manner.
[1002] The microorganisms produced according to the present invention can be cultured continuously or discontinuously using batch culture, fed-batch culture, or repeated fed-batch culture. An overview of known culture methods can be found in Chmiel's textbook (Bioprozesstechnik 1. Einführungin die Bioverfahrenstechnik [Bioprocess technology 1. Introduction to bioprocess technology] (Gustav Fischer Verlag, Stuttgart, 1991)) or Storhas's textbook (Bioreaktoren und periphere Einrichtungen [Bioreactors and peripheralequipment] (Vieweg Verlag, Braunschweig / Wiesbaden, 1994)). Standard laboratory methods can be used for this purpose; these methods are known in the art and will be further described herein.
[1003] If the polypeptide is not secreted in the culture medium, the cells can also be lysed, and the product can be obtained from the lysate using known methods for protein separation. Cells can be optionally destroyed by high-frequency ultrasound, high pressure (e.g., in a French press), by osmosis, by the action of detergents, lysing enzymes, or organic solvents, by a homogenizer, or by a combination of these methods.
[1004] Peptides can be purified using known chromatographic techniques, such as molecular sieve chromatography (gel filtration), Q-agarose chromatography, ion exchange chromatography, and hydrophobic chromatography, as well as other conventional techniques such as ultrafiltration, crystallization, salting out, dialysis, and natural gel electrophoresis. Suitable methods are described, for example, in Cooper, TG, Biochemische Arbeitsmethoden [Biochemical processes], Verlag Walter de Gruyter, Berlin, New York, or Scopes, R., Protein Purification, Springer Verlag, New York, Heidelberg, Berlin.
[1005] For the isolation of recombinant proteins, the use of a carrier system or oligonucleotide may be advantageous, which extends cDNA by a defined nucleotide sequence and thus encodes an altered polypeptide or fusion protein, for example, for easier purification. Suitable modifications of this type are, for example, so-called “tags” that act as anchors, such as modifications known as hexahistine anchors or epitopes that can be recognized as antibody antigens (e.g., described in Harlow, E. and Lane, D., 1988, Antibodies: A Laboratory Manual. Cold Spring Harbor (NY) Press). These anchors can be used to attach proteins to solid supports, such as polymer matrices, which can be used, for example, as packing material in chromatographic columns, or on microplates or other supports.
[1006] These anchors can also be used to identify proteins. To identify proteins, conventional markers such as fluorescent dyes, enzyme markers (which react with a substrate to form a detectable reaction product), or radiolabels can be used, alone or in combination with anchors, to derivatize proteins.
[1007] Uses of compounds of formula (I) and other compounds of this invention
[1008] Another aspect of the invention includes the use of the compound of formula (I) or other (intermediate) compounds obtained or obtainable from embodiments of the invention as flavoring, flavoring or aroma components, or as precursors for the manufacture of said components.
[1009] As described above, the present invention includes the use of compounds of formula (I) as fragrance ingredients. In other words, the present invention relates to a method or process for imparting, enhancing, improving, or modifying the odor properties of a fragrance composition or a fragranced article or surface, the method comprising adding an effective amount of at least one compound of formula (I) to said composition or article, for example, to impart a typical fragrance note. It should be understood that the final pleasurable effect may depend on the precise dosage and sensory properties of the compounds of the present invention, but in any case, the addition of the compounds of the present invention will, depending on the dosage, impart a typical style to the final product in the form of a note, touch, or aspect.
[1010] By “use of the compound of formula (I)”, it must also be understood herein to mean the use of any composition which includes the compound of formula (I) and can be advantageously used in the fragrance industry.
[1011] In fact, the composition described herein can be advantageously used as a flavoring ingredient, which is also an object of the present invention.
[1012] Therefore, another object of the present invention is a flavoring composition comprising:
[1013] i) at least one of the inventive compounds as defined above as a flavoring ingredient;
[1014] ii) at least one ingredient selected from the group consisting of a fragrance carrier and a fragrance base; and
[1015] iii) Optionally, at least one flavoring adjuvant.
[1016] The term "fragrance carrier" refers to a material that is neutral from the perspective of fragrance, meaning it does not significantly alter the sensory properties of the flavoring ingredients. The carrier can be liquid or solid.
[1017] As liquid carriers, emulsion systems, i.e., solvent and surfactant systems, or solvents commonly used in the fragrance industry, can be listed as non-limiting examples. A detailed description of the properties and types of solvents commonly used in the fragrance industry is impossible to exhaustively cover. However, solvents such as butanediol or propylene glycol, glycerol, dipropylene glycol and its monoethers, 1,2,3-propanetriyltriacetate, dimethyl glutarate, dimethyl adipate, 1,3-diacetoxypropyl-2-yl acetate, diethyl phthalate, isopropyl myristate, benzyl benzoate, benzyl alcohol, 2-(2-ethoxyethoxy)-1-ethanol, triethyl citrate, or mixtures thereof, are the most commonly used. Alternatively, solvents of natural origin such as glycerol or various vegetable oils such as palm oil, sunflower oil, or linseed oil can also be used. For compositions containing both a fragrance carrier and a fragrance base, in addition to those previously described in detail, other suitable fragrance carriers may be ethanol, water / ethanol mixtures, limonene or other terpenes, isoparaffins, such as those marketed under the trademark Isopar. TM Those known to the public (source: Exxon Chemical), or glycol ethers and glycol ether esters, such as those under the trademark Dowanol TM Those that are well-known (Source: Dow Chemical Company), or hydrogenated castor oil, such as those known under the trademark Cremophor® RH 40 (Source: BASF).
[1018] A solid carrier is a material to which a flavoring composition or certain components of a flavoring composition can be chemically or physically bonded. Generally, such solid carriers are used to stabilize compositions or to control the evaporation rate of compositions or certain components. Solid carriers are currently used in the art, and those skilled in the art know how to achieve the desired effects. However, as non-limiting examples of solid carriers, absorbent adhesives or polymers or inorganic materials, such as porous polymers, cyclodextrins, wood-based materials, organic or inorganic gels, clay, gypsum, talc, or zeolite, can be listed.
[1019] Other non-limiting examples of solid carriers include encapsulating materials. Examples of such materials may include wall-forming and plasticizing materials, such as glucose syrup, natural or modified starch, hydrocolloids, cellulose derivatives, polyvinyl acetate, polyvinyl alcohol, proteins, or pectins, or materials listed in references such as H. Scherz, Hydrokolloide: Stabilisatoren, Dickungs- und Geliermittel in Lebensmitteln, Band 2 der Schriftenreihe Lebensmittelchemie, Lebensmittelqualität, Behr's Verlag GmbH & Co., Hamburg, 1996. Encapsulation is a method well known to those skilled in the art and can be carried out, for example, using techniques such as spray drying, agglomeration, or extrusion; or consisting of coating encapsulation including agglomeration and composite agglomeration techniques.
[1020] As a non-limiting example of a solid carrier, a core-shell capsule may be specifically cited, which uses resins of the type of amino plastics, polyamides, polyesters, polyureas, or polyurethanes, or mixtures thereof (all of which are well known to those skilled in the art), and is carried out by a phase separation method initiated by using techniques such as polymerization, interfacial polymerization, coagulation, or these techniques together (all of which have been described in the prior art), and optionally in the presence of a polymer stabilizer or a cationic copolymer.
[1021] Resins can be produced by polycondensation of aldehydes (such as formaldehyde, 2,2-dimethoxyacetaldehyde, glyoxal, glyoxylic acid, or hydroxyacetaldehyde and mixtures thereof) with amines such as urea, benzoguanidine, glycyrrhizin, melamine, hydroxymethylmelamine, methylated hydroxymethylmelamine, guanidineazole, and mixtures thereof. Alternatively, pre-formed resins can be used to alkylate polyamines, such as those commercially available under the trademarks Urac® (Source: Cytec Technology Corp.), Cymel® (Source: Cytec Technology Corp.), Urecoll®, or Luracoll® (Source: BASF).
[1022] Other resins are produced by the polycondensation of polyols such as glycerol with polyisocyanates such as hexamethylene diisocyanate, isophorone diisocyanate or phenyl diisocyanate trimer or hexamethylene diisocyanate biuret, or phenyl diisocyanate trimer with trimethylolpropane (known under the trade name Takenate®, source: Mitsui Chemicals), with preference given to phenyl diisocyanate trimer with trimethylolpropane and hexamethylene diisocyanate biuret.
[1023] Some research literature relating to the encapsulation of fragrances via the condensation of melamine resins with aldehydes includes articles such as those by K. Dietrich et al., Acta Polymerica, 1989, vol. 40, pages 243, 325, and 683, and 1990, vol. 41, page 91. These articles have described various parameters affecting the preparation of such core-shell microcapsules according to existing methods, which are further detailed and exemplified in patent literature. US 4'396'670 of WigginsTeape Group Limited is a relevant early example of the latter. Since then, many other authors have enriched the literature in this field, and it is impossible to cover all published advances here, but general knowledge of encapsulation techniques is essential. More recent targeted publications also address suitable uses of such microcapsules, with representative articles such as K. Bruyninckx and M. Dusselier, ACS Sustainable Chemistry & Engineering, 2019, vol. 7, pages 8041-8054.
[1024] The term "fragrance base" refers to a composition containing at least one flavoring ingredient.
[1025] The stated flavoring agent does not conform to formula (I). Furthermore, by "flavoring agent," it refers to a compound used in flavoring preparations or compositions to impart a pleasant effect; that is, primarily used to impart or modify odor. In other words, for such an agent to be considered a flavoring agent, it must be recognized by those skilled in the art as capable of imparting or modifying the odor of a composition in an active or pleasant manner, and not merely possessing an odor. In addition to modifying or imparting odor, flavoring agents can provide additional benefits such as persistence, bursting, odor neutralization, antimicrobial activity, antiviral activity, microbial stability, or pest control.
[1026] The nature and type of fragrance additives present in the base material are not guaranteed to be described in greater detail here, as it is impossible to be exhaustive in any way. Those skilled in the art can select them based on their common sense and according to the intended use or application and the desired sensory effect. Generally, these fragrance additives belong to different chemical classifications, such as alcohols, lactones, aldehydes, ketones, esters, ethers, acetates, nitriles, terpenes, nitrogen- or sulfur-containing heterocyclic compounds, and essential oils, and the fragrance additives can be of natural or synthetic origin.
[1027] In particular, examples of flavoring agents commonly used in fragrance formulations can be listed, such as:
[1028] - Aldehyde components: decanal, dodecanal, 2-methylundecaldehyde, 10-undecenal, octanal, nonanal and / or nonenal;
[1029] - Aromatic herbal ingredients: eucalyptus oil, camphor, eucalyptol, 5-methyltricyclo[6.2.1.0~2,7~]undec-4-one, 1-methoxy-3-hexethiol, 2-ethyl-4,4-dimethyl-1,3-oxathiane, 2,2,7 / 8,9 / 10-tetramethylspiro[5.5]undec-8-en-1-one, menthol and / or α-pinene;
[1030] - Balsam ingredients: coumarin, ethyl vanillin and / or vanillin;
[1031] - Citrus flavoring components: dihydromyrcenol, citral, orange oil, linalyl acetate, citronellol, orange terpene, limonene, 1-p-menthene-8-yl acetate and / or 1,4(8)-p-menthadiene;
[1032] - Floral fragrance components: Methyl dihydrojasmonate, linalool, citronellol, phenethyl alcohol, 3-(4-tert-butylphenyl)-2-methylpropanal, hexylcinnamaldehyde, benzyl acetate, benzyl salicylate, tetrahydro-2-isobutyl-4-methyl-4(2H)-pyranol, β-ionone (β-violaceone), methyl 2-(methylamino)benzoate, (E)-3-methyl-4-(2,6,6-trimethyl-2-cyclohexen-1-yl)-3-buten-2-one, (1E)-1-(2,6,6-trimethyl-2-cyclohexen-1-yl)-1-penten-3- Ketones, 1-(2,6,6-trimethyl-1,3-cyclohexadien-1-yl)-2-buten-1-one, (2E)-1-(2,6,6-trimethyl-2-cyclohexen-1-yl)-2-buten-1-one, (2E)-1-[2,6,6-trimethyl-3-cyclohexen-1-yl]-2-buten-1-one, (2E)-1-(2,6,6-trimethyl-1-cyclohexen-1-yl)-2-buten-1-one, 2,5-dimethyl-2-indane methanol, 2,6,6-trimethyl-3-cyclohexene-1-carboxylate, 3-(4 4-Dimethyl-1-cyclohexen-1-ylpropanal, 3-(3,3 / 1,1-dimethyl-5-indanyl)propanal, hexyl salicylate, 3,7-dimethyl-1,6-nonadien-3-ol, 3-(4-isopropylphenyl)-2-methylpropanal, tricyclodecenyl acetate, geraniol, p-menth-1-en-8-ol, 4-(1,1-dimethylethyl)-1-cyclohexyl acetate, 1,1-dimethyl-2-phenylethyl acetate, 4-cyclohexyl-2-methyl-2-butanol, pentyl salicylate, methyl jasmonate of high-cis dihydrojasmonate, 3-methyl-5-benzene Mixtures of 1-pentanol, tricyclodecenyl propionate, geraniol acetate, tetrahydrolinalool, cis-7-p-menthol, (S)-2-(1,1-dimethylpropoxy)propionate, 2-methoxynaphthalene, 2,2,2-trichloro-1-phenylethyl acetate, 4 / 3-(4-hydroxy-4-methylpentyl)-3-cyclohexene-1-carboxaldehyde, pentylcinnamaldehyde, 8-decen-5-lactone, 4-phenyl-2-butanone, isononyl acetate, 4-(1,1-dimethylethyl)-1-cyclohexyl acetate, tricyclodecenyl isobutyrate, and / or methyl ionone isomers;
[1033] - Fruity flavor components: γ-undecyl lactone, 2,2,5-trimethyl-5-pentylcyclopentanone, 2-methyl-4-propyl-1,3-oxothiacyclohexane, 4-decyl lactone, ethyl 2-methyl-valerate, hexyl acetate, ethyl 2-methylbutyrate, γ-nonyl lactone, allyl heptaate, 2-phenoxyethyl isobutyrate, ethyl 2-methyl-1,3-dioxolane-2-ethyl acetate, diethyl 1,4-cyclohexanedicarboxylate, 3-methyl-2-hexen-1-yl acetate, 1-[3,3-dimethylcyclohexyl]ethyl [3-ethyl-2-epoxyethylene]acetate and / or diethyl 1,4-cyclohexanedicarboxylate;
[1034] - Green fragrance components: 2-methyl-3-hexanone (E)-oxime, 2,4-dimethyl-3-cyclohexen-1-carboxaldehyde, 2-tert-butyl-1-cyclohexyl acetate, styrax acetate, (2-methylbutoxy) allyl acetate, 4-methyl-3-decen-5-ol, diphenyl ether, (Z)-3-hexen-1-ol and / or 1-(5,5-dimethyl-1-cyclohexen-1-yl)-4-penten-1-one;
[1035] Musk components: 1,4-dioxa-5,17-cycloheptadecanedione, (Z)-4-cyclopentadecane-1-one, 3-methylcyclopentadecaneone, 1-oxa-12-cyclohexadecene-2-one, 1-oxa-13-cyclohexadecene-2-one, (9Z)-9-cycloheptadecene-1-one, 2-{1S)-1-[(1R)-3,3-dimethylcyclohexyl]ethoxy}-2-oxoethyl ester of propionic acid, 3-methyl-5-cycloheptadecene-2-one Pentadecen-1-one, 4,6,6,7,8,8-hexamethyl-1,3,4,6,7,8-hexahydrocyclopenta[G]-2-benzopyran, (1S,1'R)-2-[1-(3',3'-dimethyl-1'-cyclohexyl)ethoxy]-2-methylpropyl ester, oxetane-2-one and / or (1S,1'R)-[1-(3',3'-dimethyl-1'-cyclohexyl)ethoxycarbonyl]methyl ester;
[1036] - Costus root components: 1-[(1RS,6SR)-2,2,6-trimethylcyclohexyl]-3-hexanol, 3,3-dimethyl-5-[(1R)-2,2,3-trimethyl-3-cyclopenten-1-yl]-4-penten-2-ol, 3,4'-dimethylspiro[ethylene oxide-2,9'-tricyclo[6.2.1.02,7]undec[4]ene, (1-ethoxyethoxy)cyclododecane, 2,2,9,11-tetramethylspiro[5.5]undec-8-en-1-yl ester, 1-(octahydro-2,3,8,8-tetramethyl-2-naphthyl)-1-ethylone, patchouli Terpene fractions of sesame oil and patchouli oil, Clearwood®, (1'R,E)-2-ethyl-4-(2',2',3'-trimethyl-3'-cyclopenten-1'-yl)-2-buten-1-ol, 2-ethyl-4-(2,2,3-trimethyl-3-cyclopenten-1-yl)-2-buten-1-ol, methyl cypressone, 5-(2,2,3-trimethyl-3-cyclopentenyl)-3-methylpentan-2-ol, 1-(2,3,8,8-tetramethyl-1,2,3,4,6,7,8,8a-octahydronaphthyl-2-yl) ethyl-1-one and / or isobornyl acetate;
[1037] - Other components (e.g., amber, powdery, spicy, or watery): dodecahydro-3a,6,6,9a-tetramethylnaphtho[2,1-b]furan and any stereoisomers thereof, jasmine aldehyde, anisaldehyde, eugenol, cinnamaldehyde, clove oil, 3-(1,3-benzodioxacyclopenten-5-yl)-2-methylpropanal, 7-methyl-2H-1,5-benzodioxacycloheptan-3(4H)-one, 2,5,5-trimethyl-1,2,3,4,4a,5,6,7-octahydro-2-naphthol, 1-phenylvinyl acetate, 6-methyl-7-oxa-1-thia-4-azaspiro[4,4]nonane and / or 3-(3-isopropyl-1-phenyl)butanal.
[1038] The fragrance base according to the invention may not be limited to the aforementioned flavoring auxiliary ingredients, and many other such auxiliary ingredients are listed in the references, such as S. Arctander, Perfume and Flavor Chemicals, 1969, Montclair, New Jersey, USA, or later versions thereof, or other works of similar nature, as well as a large body of patent literature in the fragrance industry. It should also be understood that the auxiliary ingredients may also be compounds known to release various types of flavoring compounds in a controlled manner, also referred to as fragrance precursors or aroma compounds.Non-limiting examples of suitable flavoring precursors may include 4-(dodecylthio)-4-(2,6,6-trimethyl-2-cyclohexen-1-yl)-2-butanone, 4-(dodecylthio)-4-(2,6,6-trimethyl-1-cyclohexen-1-yl)-2-butanone, trans-3-(dodecylthio)-1-(2,6,6-trimethyl-3-cyclohexen-1-yl)-1-butanone, 2-(dodecylthio)oct-4-one, 2-phenylethyl oxo(phenyl)acetic acid, 3,7-dimethyloct-2,6-dien-1-yl oxo(phenyl)acetic acid, and oxo... (Phenylacetic)acetic acid (Z)-hex-3-en-1-yl ester, hexadecanoic acid 3,7-dimethyl-2,6-octadien-1-yl ester, bis(3,7-dimethyloct-2,6-dien-1-yl) succinate, (2-((2-methylundec-1-en-1-yl)oxy)ethyl)benzene, 1-methoxy-4-(3-methyl-4-phenethoxybut-3-en-1-yl)benzene, (3-methyl-4-phenethoxybut-3-en-1-yl)benzene, 1-(((Z)-hex-3-en-1-yl)oxy)-2-methylundec-1-ene, (2-((2-methylundec-1-yl)-ethyl ... C-1-en-1-yl)oxy)ethoxy)benzene, 2-methyl-1-(octane-3-yloxy)undec-1-ene, 1-methoxy-4-(1-phenethoxyprop-1-en-2-yl)benzene, 1-methyl-4-(1-phenethoxyprop-1-en-2-yl)benzene, 2-(1-phenethoxyprop-1-en-2-yl)naphthalene, (2-phenethoxyvinyl)benzene, 2-(1-((3,7-dimethyloct-6-en-1-yl)oxy)prop-1-en-2-yl)naphthalene, (2-((2-pentylcyclopentyl)methoxy)ethyl)benzene, 4-allyl-2-methoxy 1-((2-methoxy-2-phenylvinyl)oxy)benzene, (2-((2-pentylcyclopentyl)methoxy)ethyl)benzene, (2-((2-heptylcyclopentyl)methoxy)ethyl)benzene, 1-isopropyl-4-methyl-2-((2-pentylcyclopentyl)methoxy)benzene, 2-methoxy-1-((2-pentylcyclopentyl)methoxy)-4-propylbenzene, 3-methoxy-4-((2-methoxy-2-phenylvinyl)oxy)benzaldehyde, 4-((2-(hexyloxy)-2-phenylvinyl)oxy)-3-methoxybenzaldehyde or mixtures thereof.
[1039] By "fragrance adjuvant," we mean an ingredient that imparts additional benefits, such as color, specific lightfastness, chemical stability, etc. A detailed description of the properties and types of adjuvants commonly used in fragrance compositions cannot be exhaustive, but it must be mentioned that the ingredients are well known to those skilled in the art. Specific, non-limiting examples include: viscosity agents (e.g., surfactants, thickeners, gelling agents, and / or rheology modifiers), stabilizers (e.g., preservatives, antioxidants, heat / light and / or buffers or chelating agents, such as BHT), colorants (e.g., dyes and / or pigments), preservatives (e.g., antibacterial or antimicrobial or antifungal or antiirritant agents), abrasives, skin coolants, fixatives, insect repellents, ointments, vitamins, and mixtures thereof.
[1040] It is understood that those skilled in the art are fully capable of designing the optimal formulation for the desired effect simply by applying standard knowledge in the art and by mixing the above-mentioned components of the fragrance composition through trial and error.
[1041] In addition to flavoring compositions comprising at least one compound of formula (I), at least one flavor carrier, at least one flavor base and optionally at least one flavor adjuvant, an inventive composition consisting of at least one compound of formula (I) and at least one flavor carrier is also a specific embodiment of the present invention.
[1042] According to a particular embodiment, the above composition comprises more than one compound of formula (I) and enables perfumers to prepare blends or fragrances having the odor notes of different compounds of the present invention, thereby creating new structural units for creative purposes.
[1043] For clarity, it should also be understood that any mixture obtained directly from chemical synthesis (where the compounds of the present invention serve as starting materials, intermediates, or end products), such as reaction media that have not been adequately purified, cannot be considered a flavoring composition according to the present invention, provided that the mixture does not provide the compounds of the present invention in a suitable form for use in fragrances. Therefore, unless otherwise stated, unpurified reaction mixtures are generally excluded from the present invention.
[1044] The compounds of the present invention can also be advantageously used in all areas of the modern fragrance industry, namely the fine fragrance industry or the functional fragrance industry, to actively impart or modify the aroma of consumer products containing said compound (I). Therefore, another object of the present invention is a flavored consumer product comprising at least one compound of formula (I) as defined above as a flavoring ingredient.
[1045] The compounds of the present invention may be added as is or as part of the flavoring composition of the present invention.
[1046] For clarity, a "fragrant consumer product" refers to a consumer product that provides a pleasant fragrance effect to at least the surface or space on which it is applied (e.g., skin, hair, textiles, or household surfaces). In other words, a fragrant consumer product according to the invention is a manufactured article comprising a functional formulation, and optional additional beneficial agents corresponding to the desired consumer product, and at least one compound of the invention in an olfactoryally effective amount. For clarity, the fragrant consumer product is a non-edible product.
[1047] The nature and type of ingredients in flavored consumer products cannot be described in greater detail here, as it is impossible to be exhaustive in any way. Skilled personnel can select them based on their common sense and the characteristics of the product and the desired effect.
[1048] Non-restricted examples of suitable scented consumer products include perfumes, such as fine perfumes, sprays, or eau de perfumes; colognes, shaving lotions, or aftershaves; fabric care products, such as liquid or solid detergents, fabric softeners, liquid or solid fragrance enhancers, fabric fresheners, ironing solutions, paper products, bleach, carpet cleaners, and curtain care products; body care products, such as hair care products (e.g., shampoos, colorants, or hair sprays, color-care products, hair styling products, and dental care products), disinfectants, and feminine hygiene products; cosmetic preparations (e.g., skin creams or lotions, deodorants, or antiperspirants, such as sprays or roll-ons), hair removal agents, tanning agents, and sunscreens. Or after-sun products, nail products, skin cleansers, cosmetics; or skin care products (such as soap, shower or bath mousse, bath oil or shower gel, or hygiene products or foot / hand care products); air care products, such as air fresheners or "ready-to-use" powdered air fresheners that can be used in home spaces (rooms, refrigerators, cabinets, shoes or cars) and / or public spaces (lobby, hotels, shopping malls, etc.); or home care products, such as mold removers, furniture conditioners, wipes, dishwashing liquids or hard surface cleaners (such as floor, bathroom, sanitary ware or window cleaners); leather care products; car care products, such as polishes, waxes or plastic cleaners.
[1049] Some of the aforementioned scented consumer products may represent corrosive media for the compounds of the present invention, and therefore it may be necessary to protect them from premature decomposition, for example by encapsulation or by chemically binding them to another chemical substance suitable for releasing the components of the present invention upon exposure to suitable external stimuli such as enzymes, light, heat or pH changes.
[1050] The compounds according to the invention can be incorporated into the various products or compositions described above in proportions that vary over a wide range. When the compounds according to the invention are mixed with commonly used flavoring agents, solvents, or additives in the art, these values depend on the nature of the product to be flavored, the desired sensory effect, and the nature of the auxiliary agents in the given base.
[1051] For example, in the case of a flavored composition, the typical concentration of the compounds of the present invention is 0.001% to 10% by weight, or even more, based on the weight of the flavored composition to which they are incorporated. In the case of a flavored consumer product, the typical concentration of the compounds of the present invention is 0.01% to 1% by weight, or even more, based on the weight of the consumer product to which they are incorporated.
[1052] Furthermore, the intermediate compounds generated in any of the embodiments described herein can be converted into derivatives, such as, but not limited to, hydrocarbons, alcohols, glycols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides, and / or esters. These derivatives can be obtained by chemical methods, such as, but not limited to, oxidation, reduction, alkylation, acylation, and / or rearrangement reactions. Alternatively, these derivatives can be obtained by biochemical methods, such as contacting terpenoid compounds with enzymes (e.g., but not limited to, oxidoreductases, monooxygenases, dioxygenases, and transferases). Biochemical conversion can be carried out in vitro using isolated enzymes, enzymes derived from lysed cells, or in vivo using intact cells. This conversion can be a cyclization reaction achieved by chemical or biochemical methods. The derivative can be used as a flavoring, aroma, or fragrance component.
[1053] Those skilled in the art, upon considering the contents disclosed herein, will immediately recognize that many possible variations also fall within the scope of this invention.
[1054] The present invention will be described in more detail below through embodiments. These embodiments are merely exemplary and are not intended to limit the scope of the implementation schemes described herein.
[1055] Example
[1056] Materials and Methods
[1057] Unless otherwise stated, all chemical and biochemical materials, as well as microorganisms or cells, used in this article are commercially available products.
[1058] Unless otherwise stated, recombinant proteins are cloned and expressed using standard methods, such as those described, for example, in Sambrook, J., Fritsch, EF and Maniatis, T., Molecular cloning: A Laboratory Manual, 2nd Edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
[1059] Recombinant Escherichia coli strains were engineered to produce terpene precursors by integrating the gene encoding the mevalonate pathway enzyme into the chromosome.
[1060] Escherichia coli strains were engineered to produce farnesyl pyrophosphate (FPP) by integrating a recombinant gene encoding a mevalonate pathway enzyme into the chromosome.
[1061] An upstream pathway operon (operon 1 from acetyl-CoA to mevalonic acid) was designed, which consists of the atoB gene from Escherichia coli encoding acetyl-CoA thiolase, and the mvaA and mvaS genes from Staphylococcus aureus encoding HMG-CoA synthase and HMG-CoA reductase, respectively.
[1062] As the downstream mevalonate pathway operator (operon 2 from mevalonate to farnesyl pyrophosphate), a natural operator from the Gram-negative bacterium Streptococcus pneumoniae was selected, which encodes mevalonate kinase (mvaK1), phosphate mevalonate kinase (mvaK2), phosphate mevalonate decarboxylase (mvaD), and isopentenyl diphosphate isomerase (fni).
[1063] The codon-optimized Saccharomyces cerevisiae FPP synthase encoding gene (ERG20) was introduced into the 3' end of the upstream pathway operon to convert isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP) into FPP.
[1064] The aforementioned operon was synthesized using DNA 2.0 and integrated into the araA gene of *E. coli* strain BL21(DE3). The introduction of the heterologous pathway involved two separate recombination steps, performed using the CRISPR / Cas9 genome editing system. The first operon to be integrated (downstream pathway; operon 2) carried a spectinomycin (Spec) marker, which was used to screen for Spec-resistant candidate integrons. A second operon was designed to replace the Spec marker of the previously integrated operon, and Spec candidate integrons were screened accordingly after the second recombination event. A guide RNA expression vector targeting the araA gene was designed and synthesized using DNA 2.0. PCR primers were designed to amplify the araA gene integration target site and the integron recombination ligation site, and operon integration was verified using PCR. A clone that produced the correct PCR result was then fully sequenced and stored as strain DP1205.
[1065] Components of a culture medium used for Escherichia coli culture.
[1066] The mineral AM medium used in shake flask and laboratory-scale fermentation experiments consisted of the following components: KH₂PO₄ 4.2 g / L; K₂HPO₄·3H₂O 15.7 g / L; (NH₄)₂SO₄ 2.0 g / L; citric acid 1.7 g / L; EDTA 8.4 mg / L; glycerol 30 g / L; yeast extract 5 g / L dissolved in deionized water; dodecane concentration 10% (v / v); and sterilized at 121°C for 30 minutes. Concentrated stock solution MgSO₄·7H₂O (1 M, 5 mL / L), vitamins (thiamine hydrochloride, 4.5 g / L, 1 mL / L), and batch-prepared trace metal solution (10 mL / L) were aseptically added to the medium, and the pH was adjusted to 7 with 5 M NaOH solution. Batch-prepared trace metal solutions (per liter of 1M HCl): CoCl2·6H2O 0.25 g / L; MnCl2·4H2O 1.5 g / L; CuCl2·2H2O 0.15 g / L; H3BO3 0.3 g / L; Na2MoO4·2H2O 0.25 g / L; Zn(CHCOO)2·2H2O 1.3 g / L; ferric citrate(III) 10 g / L. For fed-batch fermentation using AM medium, a 20 L glycerol feed solution was prepared, containing 700 g / L glycerol, 12 g / L MgSO4·7H2O, 13 mg / L EDTA, and a 10 mL / L feed trace solution. A feed-weighted trace metal solution was prepared by dissolving 0.4g CoCl2·6H2O, 2.35g MnCl2·4H2O, 0.25g CuCl2·2H2O, 0.5g H3BO3, 0.4g Na2MoO4·2H2O, 1.6g Zn(CHCOO)2·2H2O, and 10g citric acid Fe(III)·H2O in 1L 1M HCl.
[1067] Engineered bacterial cells were cultured under conditions that enabled the production of terpenoid compounds.
[1068] DP1205 *E. coli* strains (as described in WO2018 / 114839), engineered to increase the level of the terpene precursor farnesyl diphosphate (FPP), were transformed with one or two expression plasmids carrying genes encoding enzymes of the high farnesol biosynthesis pathway and / or terpene cyclases. Transformed cells were cultured on LB agarose plates containing appropriate antibiotics (kanamycin (50 µg / mL), carbenicillin (50 µg / mL), chloramphenicol (34 µg / mL), and / or streptomycin (50 µg / mL)). Single colonies were picked and inoculated into 5 mL of liquid LB medium supplemented with the same antibiotics, 4 g / L glucose, and 10% (v / v) n-dodecane. The following day, 0.2 mL of the overnight culture was inoculated into 2 mL of AM medium supplemented with the same antibiotics and 10% (v / v) n-dodecane. The cultures were incubated at 37°C until the optical density reached 3. Then, 0.1 mM IPTG was added to induce recombinant protein expression, and the mixture was cultured at 25°C for another 72 hours.
[1069] The culture was then extracted with 1 volume of methyl tert-butyl ether (MTBE), and the composition of the organic phase was analyzed by GC-MS as described below. For quantitative analysis, an internal standard (α-pinene (Sigma-Aldrich, Missouri, USA)) was added to the extract prior to GC-MS analysis, and the concentration of each component was estimated based on the comparison of peak areas.
[1070] GC-MS analysis method.
[1071] Sample analysis was performed using an Agilent 6890N gas chromatograph coupled with a 5975B series mass selective detector (MSD) and equipped with a split / splitless injector (Agilent Technologies, California) and a CombiPAL autosampler (PAL LSI 85 autosampler, Agilent Technologies, California). The GC injector temperature was set to 240°C, the injection volume to 1.0 µL, and the split ratio to 25:1 (pressure 23.304 PSI). A DB-5ms capillary column (30 m × 0.25 mm inner diameter × 0.25 μm film thickness; Agilent J&W) was used, with helium as the carrier gas and a constant flow rate of 1.2 mL / min. The oven was initially set to 80°C (hold for 1 minute), then programmed to 300°C (10°C / min), and then further programmed to 300°C (30°C / min; hold for 1 minute).
[1072] General methods, gene modification, culture, and compound analysis of Saccharomyces cerevisiae.
[1073] A *Saccharomyces cerevisiae* strain capable of increasing levels of the terpene precursor farnesyl diphosphate (FPP) (as described in WO2018 / 114839) was used as the base strain to express genes for the high farnesol biosynthesis pathway and terpene cyclases. In short, this strain integrates all endogenous mevalonate pathway genes into its genome, controlled by the native GAL1 or GAL10 promoter. By replacing its native promoter, the expression of the squalene synthase gene (ERG9) was downregulated, thereby further increasing the farnesyl diphosphate precursor library in this strain.
[1074] All genes (synthesized by ATUM Corporation, California, or Twist Bioscience, California, USA) and their associated regulatory elements (e.g., promoters and terminators) were introduced into the basal strain via genome integration or by plasmids constructed in vivo using the yeast homologous recombination mechanism (Kuijpers et al., Microb Cell Fact., 2013, 12:47). All yeast transformations were performed using the lithium acetate method (Gietz and Woods, Methods Enzymol., 2002, 350:87-96).
[1075] The successfully transformed yeast colonies were cultured at 30°C for three days on a medium containing 6.7 g / L of amino acid-free yeast nitrogen source (BDDifco, New Jersey, USA), appropriate antibiotics or nutrients added according to the marker gene used, 20 g / L glucose, and 20 g / L agar.
[1076] To produce and analyze metabolites, single colonies of the modified yeast strain were inoculated into 2 mL of medium (Westfall et al., Proc Natl Acad Sci USA, 2012, 109:E111-118) supplemented with 2% galactose and 10% (v / v) n-dodecane (Sigma-Aldrich, Missouri, USA). The cultures were incubated at 30°C on a shaker at 200 rpm for 3 days. After incubation, the cultures were extracted with twice the volume of MTBE (quantified with α-pinene standard as described above), and the organic phase composition was analyzed by GC-MS (using an Agilent 7890A gas chromatograph coupled with a 5975B series mass selective detector (MSD) and equipped with a split / splitless injector and GC Injector 80 injection system) (Agilent Technologies, California). The GC injection port temperature was set to 260℃. A 1.0 µL sample was injected in splitless mode and analyzed on an HP-5 column (30 m × 0.25 mm × 0.25 µm; Agilent J&W) using helium as the carrier gas at a constant flow rate of 1.2 mL / min. The column oven temperature was initially set to 100℃ and increased to 300℃ at a rate of 10℃ / min.
[1077] Example 1. In vivo production of (3E,7E)-gafarnesol and biosynthetic intermediates in engineered bacterial cells expressing GGPP synthase, phosphatase, alcohol dehydrogenase, BVMO, enal lyase and esterase.
[1078] Figure 3The reaction scheme shown describes a biochemical pathway for the in vivo production of (3E,7E)-hofarnesol. The universal isoprene precursor isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP) condense to form (2E,6E,10E)-geraniylgeraniyl diphosphate (GGPP). This reaction can be catalyzed by GGPP synthase or by a combination of farnesyl diphosphate synthase (FPP synthase) and GGPP synthase. GGPP can be converted to (2E,6E,10E)-geraniylgeraniol by terpene synthases or phosphatases (e.g., as described in WO2020011883A1). (2E,6E,10E)-geraniylgeraniol then undergoes several enzymatic degradation steps. In the proposed pathway, (2E,6E,10E)-geraniol is first oxidized and cleaved by an alcohol dehydrogenase (ADH) to generate (5E,9E)-farnesylacetone. Subsequent cleavage can be catalyzed by an enal lyase, such as a protein containing the GXWXG (SEQ ID NO: 263) and DUF4334 domains, as described in WO2021005097. (5E,9E)-farnesylacetone is further converted to (3E,7E)-hofarnesylacetate by Bayer-Villiger monooxygenase (BVMO). Finally, this ester is hydrolyzed by an esterase to ultimately form (3E,7E)-hofarnesol.
[1079] To validate the (3E,7E)-gafarnesol pathway, *E. coli* was engineered to express the required enzyme. A plasmid containing two operons was assembled.
[1080] The first operon was designed to contain three cDNAs, which encode:
[1081] - PsAerADH (SEQ ID NO: 11), an alcohol dehydrogenase from *Pseudomonas aeruginosa* (GeneBank accession number: WP_079868259.1), possesses the ability to oxidize (2E,6E,10E)-geraniol to (2E,6E,10E)-geraniol.
[1082] - SCH24-BVMO1 (SEQ ID NO: 23), a Bayer-Villiger monooxygenase (BVMO) from *Filobasidium magnum*, described in WO2021005097, oxidizes (5E,9E)-farnesylacetone to (3E,7E)-hofarnesylacetate, and
[1083] - SCH24-EST1 (SEQ ID NO: 27), an esterase from Ustilago maydis, described in WO2021005097, hydrolyzes (3E,7E)-gofarnesylacetic acid to (3E,7E)-gofarnesol and acetic acid.
[1084] The second operon is constructed to contain 4 cDNAs, which encode:
[1085] - SCH94-03944 (SEQ ID NO: 22), a protein containing an enal lyase from Rhodococcus erythropolis, described in WO2021005097, which cleaves (2E,6E,10E)-geranylgeranyl to (5E,9E)-farnesylacetone and acetaldehyde.
[1086] - CcrGGPPS2-del57, a truncated version of (2E,6E,10E)-geranylgeranyl diphosphate synthase from *Cistus creticus* (GeneBank: AAM21639.1) (SEQ ID NO: 1), and
[1087] - Two copies of cDNA encoding PgpB (SEQ ID NO: 3), PgpB is a phosphatase from Escherichia coli (GeneBank: WP_089622241.1) that converts (2E,6E,10E)-geraniol to (2E,6E,10E)-geraniol by cleaving diphosphate groups.
[1088] The cDNAs encoding PsAerADH, SCH24-BVMO1, SCH24-EST1, SCH94-03944, CcrGGPPS2-del57, and PgpB have been codon-optimized for *E. coli* (SEQ ID NO: 102, 115, 120, 113, 90, and 93), and an RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 264) has been inserted upstream of each of these cDNAs. The first operon of the cDNAs containing PsAerADH, SCH24-BVMO1, and SCH24-EST1 is controlled by the T5 promoter and the rrnB T1 terminator. The second operon of the cDNAs containing SCH94-03944, CcrGGPPS2-del57, and the two PgpB cDNAs is controlled by the T5 promoter and the rrnB terminator. Both operons were synthesized and cloned into a vector backbone containing the pUC origin of replication, kanamycin resistance gene, and LacI gene, thus obtaining the vector pHFOL-5.
[1089] The farnesyl diphosphate (FPP) of the producing strain DP1205, described in WO2021005097, was transformed using the aforementioned vector pHFOL-5. When cultured under conditions conducive to terpene production, the resulting cells produced (3E,7E)-hofarnesol (…). Figure 4 Under the conditions described in the "Materials and Methods" section, 69 mg / L of (3E,7E)-gafarnesol was produced in the culture medium determined by the tube method.
[1090] Cell product analysis ( Figure 4 The study also revealed the accumulation of various metabolic intermediates, such as (5E,9E)-farnesylacetone and (2E,6E,10E)-geraniol. The accumulation of these intermediates can be limited and the concentration of the final product increased by optimizing the different enzymatic steps in this pathway.
[1091] The following examples demonstrate how to identify the appropriate enzymes for each enzymatic step in the pathway to increase the yield of (3E,7E)-gafarnesol and limit the accumulation of metabolic intermediates.
[1092] Example 2. Screening of Bayer-Verilog monooxygenases to enhance the enzymatic conversion of (5E,9E)-farnesylacetone to (3E,7E)-hofarnesylacetate and the in vivo production of (3E,7E)-hofarnesol.
[1093] Example 1 shows that (5E,9E)-farnesylacetone can also be detected in vivo when (3E,7E)-gafarnesol is produced, due to insufficient activity of Bayer-Villiger monooxygenase (SCH24-BVMO1 (SEQ ID NO: 23)) in this strain.
[1094] In this embodiment, different BVMOs were screened in vivo to identify enzyme candidates with higher efficiency than SCH24-BVMO1 (SEQ ID NO: 23). During the screening process, a modified version of the vector pHFOL-5 (described in Example 1) was constructed by removing SCH24-BVMO1 and named pF-Facetone-7. The pF-Facetone-7 vector was transformed into Escherichia coli DP1205 strain to obtain a strain capable of producing (5E,9E)-farnesylacetone under conditions that allow for the production of terpenoid compounds. The highest concentration of (5E,9E)-farnesylacetone (FVMO1) produced in the culture medium was 410 mg / L in vitro. Figure 5 A). When cells are further transformed by a vector expressing active BVMO, (5E,9E)-farnesylacetone is converted to (3E,7E)-gafarnesylacetate, which is then converted to (3E,7E)-gafarnesol by SCH24-EST1 esterase. Figure 5B).
[1095] In the next step, a codon-optimized cDNA encoding BVMO was designed and cloned into the pJ423 expression plasmid (ATUM, Newark, CA). DP1205 E. coli cells were co-transformed with one of these plasmids and plasmid pF-Facetone-7. BVMO activity was determined by quantifying the amount of (3E,7E)-gafarnesol produced by each tested BVMO and comparing it with SCH24-BVMO1 (SEQ ID NO: 23).
[1096] The table below (Table 1) shows the relative activities of (5E,9E)-farnesylacetone conversion of some BVMOs identified in this screening.
[1097]
[1098] Table 1: Relative activities of selected BVMO and (5E,9E)-farnesylacetone to (3E,7E)-gafarnesylacetate, which is further converted to (3E,7E)-gafarnesol.
[1099] The (3E,7E)-gafarnesol produced by AraBVMO1 and AflavBVMO1 was found to be significantly higher than that produced by SCH24-BVMO1 (SEQ ID NO: 23). Compared with the reference strain BVMO, AflavBVMO1 (SEQ ID NO: 26) increased the yield of (3E,7E)-gafarnesol by 59%.
[1100] Example 3. In vivo screening of alcohol dehydrogenases to enhance the enzymatic conversion of (2E,6E,10E)-geraniol to (2E,6E,10E)-geraniol and the in vivo production of (3E,7E)-gafarnesol.
[1101] In this embodiment, the efficiency of alcohol dehydrogenase (ADH) in oxidizing (2E,6E,10E)-geranylgeraniol to (2E,6E,10E)-geranylgeranialdehyde in vivo was tested. Alcohol dehydrogenase catalyzes the reversible oxidation reaction of alcohols.
[1102] To avoid reverse alcohol dehydrogenase reactions in this in vivo screening assay, the enal lyase SCH94-03944 described in Example 1 was co-expressed in *E. coli* cells to enzymatically convert (2E,6E,10E)-geraniylgeraniol to (5E,9E)-farnesylacetone. The catalytic efficiency of the ADH was then correlated with the amount of (2E,6E,10E)-geraniylgeraniol converted to (5E,9E)-farnesylacetone. The ADH candidate was codon-optimized and cloned into the pJ423 expression plasmid (ATUM, Newark, CA). DP1205 *E. coli* cells were co-transformed with one of these plasmids and plasmid pJ401-SCH94-3944-PgpB-CcrGGPPS, which contains the genes required for the production of (5E,9E)-farnesylacetone, as described below, except for the ADH-encoding gene.
[1103] Therefore, plasmid pJ401-SCH94-3944-PgpB-CcrGGPPS2-del57 contains one operon carrying three cDNAs, which encode:
[1104] - Enaldehyde lyase SCH94-03944 (SEQ ID NO: 22), a protein containing GXWXG (SEQ ID NO: 263) and DUF4334 domains, has been described in WO2021005097. It can cleave (2E,6E,10E)-geraniylgeranialdehyde into (5E,9E)-farnesylacetone and acetaldehyde.
[1105] - PgpB (SEQ ID NO: 3), a phosphatase from *Escherichia coli* (GeneBank: WP_089622241.1), converts (2E,6E,10E)-geranylgeranyl diphosphate to (2E,6E,10E)-geranylgeranyl by cleaving diphosphate, and...
[1106] - CcrGGPPS2-del57 (SEQ ID NO: 1), a truncated form of (2E,6E,10E)-geraniol-geraniol diphosphate synthase from Rosa rockosa (GeneBank: AAM21639.1).
[1107] The cDNAs encoding SCH94-03944, PgpB, and CcrGGPPS2-del57 were codon-optimized for expression in *E. coli* (SEQ ID NO: 113, 93, and 90). An operon was designed containing these three cDNAs sequentially and an RBS sequence (AAGGAGGTAAAAAA) placed upstream of each cDNA (SEQ ID NO: 264). The operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, CA).
[1108] The resulting strains (DP1205 containing plasmid pJ401-SCH94-3944-PgpB-CcrGGPPS2-del57 and plasmid pJ423 carrying candidate alcohol dehydrogenases) were cultured under terpene-producing conditions. The amounts of (2E,6E,10E)-geraniol and (5E,9E)-farnesylacetone were determined, and the conversion rate of each tested alcohol dehydrogenase was calculated.
[1109]
[1110] Table 2: ADHs selected in vivo that are active in the conversion of (2E,6E,10E)-geraniol to (2E,6E,10E)-geraniol, which is further converted to (5E,9E)-farnesylacetone.
[1111] The results are shown in Table 2. In the culture medium determined by the tube method, the highest conversion rates of (2E,6E,10E)-geraniol were detected for alcohol dehydrogenases ThTerpADH1 (SEQ ID NO: 12), Ppseudo-alkJ (SEQ ID NO: 15), and CymB (SEQ ID NO: 17). The highest amount of (5E,9E)-farnesylacetone, 172 mg / L, was produced using the alcohol dehydrogenase Ppseudo-alkJ (SEQ ID NO: 15).
[1112] Example 4. In vivo testing of esterase-catalyzed hydrolysis of (3E,7E)-gafarnesyl acetate to (3E,7E)-gafarnesol and acetic acid.
[1113] In this embodiment, the ability of different esterases to catalyze the hydrolysis of (3E,7E)-gafarnesyl acetate to (3E,7E)-gafarnesol and acetic acid in vivo was tested.
[1114] Esterases were tested in *E. coli* DP1205 using a dual-plasmid system. The first plasmid contained one operon with three cDNAs encoding:
[1115] - Enal lyase SCH94-03944 (SEQ ID NO: 22), a protein containing GXWXG (SEQ ID NO: 263) and DUF4334 domains, described in WO2021 / 005097, cleaves (2E,6E,10E)-geraniylgeranialdehyde into (5E,9E)-farnesylacetone and acetaldehyde.
[1116] - PgpB (SEQ ID NO: 3), a phosphatase from *Escherichia coli* (GeneBank: WP_089622241.1), converts (2E,6E,10E)-geranylgeranyl diphosphate to (2E,6E,10E)-geranylgeranyl by cleaving diphosphate, and...
[1117] - CcrGGPPS2-del57 (SEQ ID NO: 1), a truncated form of (2E,6E,10E)-geraniol-geraniol diphosphate synthase from Rosa rockosa (GeneBank: AAM21639.1).
[1118] The cDNA encoding SCH94-03944, PgpB, and CcrGGPPS2-del57 was codon-optimized for expression in *E. coli* (SEQ ID NO: 113, 93, and 90) and contained an upstream-placed RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 264). This operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, CA).
[1119] The second operon consists of three cDNAs that encode:
[1120] - PsAerADH, an alcohol dehydrogenase from *Pseudomonas aeruginosa* (GeneBank accession number: WP_079868259.1) (SEQ ID NO: 11), has the ability to oxidize (2E,6E,10E)-geranylgeraniol to (2E,6E,10E)-geranylgeranialdehyde.
[1121] - SCH24-BVMO1 (SEQ ID NO: 23), a Bayer-Villiger monooxygenase (BVMO) described in WO2021005097, oxidizes (5E,9E)-farnesylacetone to (3E,7E)-hofarnesylacetate, and
[1122] - Esterase gene candidates.
[1123] The cDNA encoding PsAerADH, SCH24-BVMO1, and esterase gene candidates was codon-optimized for expression in *E. coli* (SEQ ID NO: 102 and 115) and contained the upstream RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 264). The operon was synthesized and cloned into the pJ424 expression plasmid (ATUM, Newark, CA).
[1124] Escherichia coli DP1205 cells were co-transformed with plasmids pJ401-SCH94-03944-PgpB-CcrGGPPS2-del57 and pJ424-PsAerADH-SCH24-BVMO1-esterases. In the resulting strains, esterases SCH24-EST1 (SEQ ID NO: 27) from *Filobasidium magnum* and SCH23-EST1 (SEQ ID NO: 28) from *Hyphozyma roseonigra* were found to convert (3E,7E)-gafarnesyl acetate to (3E,7E)-gafarnesol.
[1125] Example 5. Screening different phosphatases in vivo to catalyze the hydrolysis of (2E,6E,10E)-geraniol to (2E,6E,10E)-geraniol.
[1126] In this embodiment, the ability of phosphatases from different protein families to catalyze the hydrolysis of (2E,6E,10E)-geranyl-geranyl diphosphate to (2E,6E,10E)-geranyl-geranyl and diphosphate in vivo was tested. These phosphatases were tested in *E. coli* DP1205 cells containing geranyl-geranyl diphosphate synthase in their expression vectors. DNA fragments encoding the codon-optimized phosphatases (each cloned into a second expression vector) were then introduced into the strain via transformation. Under terpene-producing conditions, eight phosphatases were identified that were active for (2E,6E,10E)-geranyl-geranyl diphosphate and capable of producing (2E,6E,10E)-geranyl-geranyl. The results are shown in Table 3. Phosphatases PgpB (SEQ ID NO: 3), PeSubTPP1 (SEQ ID NO: 7), and TalVeTPP (SEQ ID NO: 8) showed the highest activity. PgpB (SEQ ID NO: 3) was found to be the most efficient enzyme for producing (2E,6E,10E)-geraniol.
[1127]
[1128] Table 3: Phosphatases selected in vivo based on their activity in converting (2E,6E,10E)-geranylgeranyl diphosphate to (2E,6E,10E)-geranylgeranyl. These phosphatases were ranked and divided into four groups based on their production efficiency: - Inactive; + Low activity; ++ High activity; +++ Very high activity.
[1129] Example 6. Production of (3E,7E)-gafarnesol in engineered fungal cells.
[1130] The (3E,7E)-gafarnesol biosynthesis pathway ( Figure 3 The gene was introduced into a Saccharomyces cerevisiae strain that can efficiently produce the terpene precursor farnesyl diphosphate (FPP) in a two-step process. First, a single copy of the gene encoding the alcohol dehydrogenase SCH23-ADH1 (SEQ ID NO: 21) (WO2021005097), the esterase SCH23-EST1 (SEQ ID NO: 28) (WO2021005097), and the Bayer-Villiger monooxygenase AflavBVMO (SEQ ID NO: 26) was integrated into the genome of the yeast strain. The resulting strain YST403 was subsequently used to construct a 2-micron plasmid containing geranyl-geranyl diphosphate synthase CarG (SEQ ID NO: 2) (from *Bacillus trispora*, NCBI accession number JQ289995.1), phosphatase PgpB (SEQ ID NO: 3) (from *Escherichia coli*, NCBI accession number WP_089622241.1), and enal lyase SCH94-03944 (SEQ ID NO: 22). All genes encoding different enzymes were codon-optimized for expression in *Saccharomyces cerevisiae* (SEQ ID NO: 112, 122, 119, 92, 94, and 114) and regulated by a galactose-inducible promoter. Post-culture, the final strain, named YST403_HFOL, produced (3E,7E)-gafarnesol. Product analysis ( Figure 6 The study also showed that the strain accumulated multiple metabolic inte...
Claims
1. A method for preparing a compound of formula (I), (I) The compound is in the form of any of its stereoisomers or mixtures thereof, and the method comprises: (i) Contacting the compound of formula (VI) with a polypeptide having terpene cyclase activity to produce the compound of formula (I), (WE) The compound is in the form of any of its stereoisomers or mixtures thereof.
2. The method according to claim 1, wherein more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib): (Form) (Formula Ib).
3. The method according to claim 1 or 2, wherein the compound of formula (VI) is in the form of formula (VIa): (Formula VIa).
4. The method according to any one of the preceding claims, wherein the polypeptide having terpene cyclase activity is a mixed-source terpene cyclase and / or squalene cyclase.
5. The method of claim 4, wherein the mixed-origin terpene cyclase is selected from at least one or more of the following polypeptides: (a) A bacterial membrane-integrated mixed-origin terpene cyclase containing at least one or more amino acid motifs selected from: . [W]xxx[D]xx[ILVMN] (SEQ ID NO: 254); .PxxAxxxNxxWE(SEQ ID NO: 255); . MxxxFxxMLxxR (SEQ ID NO: 256); and .RxxxxGQS (SEQ ID NO: 257); (b) A fungal-derived membrane-integrated mixed-origin terpene cyclase containing at least one or more amino acid motifs selected from: [WY]Exx[YFW] (SEQ ID NO: 258); and . [DNE]xSYxxP (SEQ ID NO: 259); (c) A bacterial soluble mixed-origin terpene cyclase containing at least one or more amino acid motifs selected from: .GxWxxxW[WG]xxxxY (SEQ ID NO: 260); . WxxxHxxV[TSA] (SEQ ID NO: 261); and .GxWxD[FY] (SEQ ID NO: 262); Residue x can be any natural amino acid residue that is independent of each other.
6. The method according to any one of claims 4 and 5, wherein the mixed-origin terpene cyclase is a membrane-integrated mixed-origin terpene cyclase.
7. The method according to any one of claims 4 to 6, wherein the mixed-origin terpene cyclase is a membrane-integrated mixed-origin terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any one of the sequences in SEQ ID NO: 50 to 73 and 280 to 289.
8. The method according to any one of claims 6 and 7, wherein the enzyme preferably produces a compound of formula (I) in the form of formula (Ia).
9. The method according to any one of claims 4 and 5, wherein the mixed-origin terpene cyclase is a soluble mixed-origin terpene cyclase.
10. The method according to any one of claims 4, 5 and 9, wherein the mixed-origin terpene cyclase is a soluble mixed-origin terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any one of the sequences in SEQ ID NO: 74 and 75.
11. The method according to any one of claims 9 and 10, wherein the enzyme preferably produces a compound of formula (I) in the form of formula (Ib).
12. The method according to any one of claims 1 to 5, wherein the polypeptide having terpene cyclase activity is a mixed terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any one of SEQ ID NO: 57, 71, 74, 280, 281, 282, 283, 286, 287 and 288.
13. The method of claim 4, wherein the squalene cyclase comprises at least one motif selected from [SP][TP][VIL]WDTx[LWI] (SEQ ID NO: 247), PGG[WF][GYA]F (SEQ ID NO: 248), PDxDD[TAS][TIAS] (SEQ ID NO: 249), [MIL]QxxxG[GA][WF]x[AS][FY] (SEQ ID NO: 250), Qxxx[GH]xWxG[RK]WGxx[YF]xYG (SEQ ID NO: 251), Qxx[DN]G[GS][WF][GS]ExxxS (SEQ ID NO: 252), and [STA]xx[SFN][QC]T[AGT]W[AS][LIV]xx[LQ] (SEQ ID NO: 253); wherein residue x independently represents any natural amino acid residue.
14. The method according to any one of claims 4 and 13, wherein the squalene cyclase has 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any one of SEQ ID NO: 29 to 49, 265 to 274 and 276 to 279.
15. The method according to any one of the preceding claims, wherein the method further comprises one or more steps prior to step (i), said steps comprising: (a) Contacting the compound of formula (V) with a polypeptide having esterase activity to produce the compound of formula (VI), (V) The compound is in the form of any of its stereoisomers or mixtures thereof; (b) Contacting compound (IV) with a polypeptide having Bayer-Villiger monooxygenase (BVMO) enzyme activity to produce compound (V), (IV) The compound is in the form of any of its stereoisomers or mixtures thereof; (c) Contacting the compound of formula (III) with a polypeptide having enal lyase activity to produce the compound of formula (IV), (III) The compound is in the form of any of its stereoisomers or mixtures thereof; (d) Contacting compound (II) with a polypeptide having alcohol dehydrogenase (ADH) activity to produce compound (III), (II) The compound is in the form of any of its stereoisomers or mixtures thereof; (e) Using one or more polypeptides with phosphatase activity, producing a compound of formula (II) from geraniol geraniol diphosphate (GGPP); and / or (f) Using one or more polypeptides with isoprenyl transferase activity, GGPP is generated from isoprenyl diphosphate (IPP) and dimethyl allyl diphosphate (DMAPP).
16. The method of claim 15, wherein: (a) The polypeptide with esterase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 27 and 28. (b) The polypeptide having BVMO enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 23 to 26 and 216 to 227. (c) The polypeptide having enal lyase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO:
22. (d) The polypeptide with ADH enzyme activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 11 to 21. (e) The polypeptide with phosphatase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences in SEQ ID NO: 3 to 10; and / or (f) The polypeptide having isoprenyl transferase activity has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with SEQ ID NO: 1 or 2.
17. The method according to any one of the preceding claims, wherein the method is an in vivo method or a biotransformation method.
18. The method according to any one of the preceding claims, wherein the method is carried out in recombinant cells capable of functionally expressing the following polypeptides: a polypeptide having terpene cyclase activity as defined in any one of claims 4 to 14; and optionally, one or more polypeptides as defined in any one of claims 15 and 16.
19. The method according to claim 18, wherein the recombinant cell is a bacterial cell, plant cell, or fungal cell such as a yeast cell; preferably, the recombinant cell belongs to the genera *Escherichia coli*, *Saccharomyces cerevisiae*, *Yersinia*, or *Pichia pastoris*.
20. A recombinant cell comprising, producing or capable of producing a compound of formula (I), wherein more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
21. The recombinant cell of claim 20, wherein the cell comprises a polypeptide having terpene cyclase activity as defined in any one of claims 4 to 14.
22. A cell culture fermentation medium comprising the recombinant cells of any one of claims 20 and 21.
23. A reaction mixture comprising a compound of formula (I), wherein more than 97% of the compound of formula (I) is in the form of formula (Ia) and / or formula (Ib).
24. A compound of formula (I) obtained or obtainable by the method of claims 1 to 19, wherein the compound is derived from recombinant cells of any one of claims 20 and 21, from the cell culture fermentation medium of claim 22, or from the reaction mixture of claim 23.
25. A compound of formula (I), wherein more than 97% of the compound is in the form of formula (Ia) and / or formula (Ib).
26. Use of the compound of formula (I) according to any one of claims 24 and 25 as a flavoring ingredient.
27. Use of mixed-source terpene cyclases for the production of compounds of formula (I) and / or their derivatives.
28. A variant of a hybrid terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 56 to 70.
29. A variant hybrid terpene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 56 to 61, 69, and 70, wherein, This variant of the mixed-source terpene cyclase has an amino acid substitution at the 9th amino acid position relative to the sequence provided in SEQ ID NO:
51.
30. A variant squalene cyclase having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity with any of the sequences provided in SEQ ID NO: 29, 31, 33, 34, 36 to 38, 40, 41, 43 to 46, 48, 49, 265 to 274, and 276 to 279, wherein the polypeptide is alanine at amino acid position 437 and methionine at amino acid position 600 relative to the sequence provided in SEQ ID NO: 82.
Citation Information
Patent Citations
Phonophoretic cannabidiol composition and transdermal delivery system
CA3053353A1
Paper file
CA61014A
Process for the production of microcapsules
US4396670A
Biocatalytic production of ambroxan
WO2010139719A2
Enzymes and applications thereof
WO2016170099A1