Biocatalytic method for controlled degradation of terpene compounds
A new class of polypeptides with enal cleavage and BVMO activities enables the controlled degradation of terpenes, addressing the need for terpene-degrading enzymes and providing a biocatalytic pathway for synthesizing valuable terpene-derived compounds.
Patent Information
- Application Number
- JP2025025299
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-11-13
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-03
Smart Images

Figure 2025084810000287 
Figure 2025084810000288 
Figure 2025084810000289
Abstract
Description
Technical Field
[0001] This specification provides a biocatalytic method for producing terpene degradation products useful as starting materials for the production of perfume ingredients such as, for example, ambrox. In particular, novel terpene degradation polypeptides (enal cleavage polypeptides) and novel peptides (oxygenases) that can be applied to a novel type of fully enzymatic multi-step degradation pathway that enables controlled stepwise conversion and degradation of linear or cyclic terpene substrates, as well as mutants and variants derived therefrom, are provided. By the aforementioned novel biosynthetic strategy, it becomes possible to completely biochemically synthesize valuable terpene-derived compounds such as, for example, manooloxy or gamma ambrol. The present invention also provides a recombinant host organism having a set of genetic information essential for the functional expression of a set of enzymes required to catalyze a combination of enzymatic conversion and degradation steps.
[0002] Background of the Invention Terpenes are found in most organisms (microorganisms, animals, and plants). These compounds are composed of units of five carbon atoms, so-called isoprene units, and are classified according to the number of these units present in the isoprene structure. Thus, hemiterpenes, monoterpenes, sesquiterpenes, and diterpenes are terpenes containing 5, 10, 15, and 20 carbon atoms, respectively (i.e., 1, 2, 3, and 4 isoprene units). For example, sesquiterpenes are widely present in the plant kingdom. Many sesquiterpene molecules are known for their flavor and fragrance properties, as well as their effects in cosmetics, pharmaceuticals, and antibacterial activity. Numerous sesquiterpene hydrocarbons and sesquiterpenoids have been identified.
[0003] The production by biosynthesis of terpenoids involves enzymes called terpene synthases. These enzymes convert acyclic terpene precursors into one or more terpene products. In particular, diterpene synthases cyclize the precursor geranylgeranyl diphosphate (GGPP) to produce diterpenoids. In the cyclization reaction of GGPP, often two enzyme polypeptides, namely type I and type II diterpene synthases, need to function in combination in two consecutive enzyme reactions. Type II diterpene synthases catalyze the cyclization / rearrangement of GGPP initiated by protonation of the terminal double bond of GGPP, resulting in a cyclic diterpene diphosphate intermediate. This intermediate is then further converted by type I diterpene synthases that catalyze ionization-initiated cyclization.
[0004] Diterpene synthases are present in plants or other organisms and use substrates such as GGPP, but have different product profiles. Genes and cDNAs encoding diterpene synthases have been cloned and the properties of the corresponding recombinant enzymes have been elucidated.
[0005] Enzymes that catalyze the specific or preferential cleavage or removal of diphosphate groups from terpene diphosphate intermediates, particularly cyclic terpene diphosphate intermediates such as copalyl diphosphate (CPP) or labdadienol diphosphate (LPP) of diterpenoids, have only recently been described in a previous European patent application (EP application number 18182783.3). The number or carbon atoms of the terpene diphosphate do not change by the aforementioned enzyme.
[0006] However, terpene-derived compounds that can be regarded as degradation products of terpene precursors such as acyclic or cyclic sesquiterpenoids or diterpenoids are needed, and these compounds can be further converted chemically and / or enzymatically into end products and applied, for example, as fragrance components.
[0007] The problem to be solved by the present invention is to provide a polypeptide exhibiting enzymatic terpene-degrading activity or a polypeptide that can be converted into a derivative capable of degrading such terpenes.
[0008] Another problem to be solved by the present invention is to establish a new complete biocatalytic degradation pathway for producing defined terpene degradation products.
[0009] Summary of the Invention Surprisingly, the above problems could be solved by providing a new class of polypeptides having enal cleavage activity that enables for the first time a carbonyl-functionalized terpene compound with specifically shortened two carbon atoms and its respective biocatalytic process. For example, by this new class of enzymes, copalal, a labdane-type compound containing a diterpene carbon skeleton and having a terminal aldehyde group, can be converted into manooloxy, a respective dinor-labdane-type compound that maintains a carbon skeleton shortened by two carbon atoms, i.e., composed of 18 carbon atoms.
[0010] Surprisingly, by providing a new class of polypeptides having Baeyer-Villiger monooxygenase (BVMO) activity that enables specific oxidation of terpene compounds to their esters (Baeyer-Villiger oxidation) and their respective biocatalytic processes, the above-mentioned problems in alternative approaches could also be solved. For example, the novel class of BVMO enables the conversion of copalal, a labdane-type compound having a diterpene carbon skeleton and a terminal aldehyde group, into their respective norlabdane-type formate esters. By the aforementioned Baeyer-Villiger oxidation, the labdane-type compound can be easily converted into their respective norlabdanes by the action of a polypeptide having esterase activity. In this step, as a result, one carbon atom is shortened. If the terminal aldehyde group is replaced by a terminal keto group, although it is a shortening by the same method, this time, shortening of one or more carbonate moieties is possible. By repeating the combination of the oxygenation step by the BVMO catalyst and the cleavage step by the esterase catalyst, it becomes possible to gradually shorten the hydrocarbon chain of the terpene molecule.
[0011] By combining the degradation steps catalyzed by the above enal cleavage enzyme and BVMO enzyme, it becomes possible to construct a completely new biochemical degradation pathway that can be applied to a greater variety of carbonyl-functionalized compounds, particularly cyclic or acyclic terpenes or terpenoids.
[0012] The aforementioned biocatalytic steps may be combined with several other preceding (upstream) or consecutive (downstream) enzyme steps, enabling the provision of a biocatalytic multi-step process for the complete enzymatic synthesis of a number of valuable and complex terpene molecules from their respective precursors of terpene molecules.
[0013] The subsequent scheme shows two specific embodiments of two alternative routes of the present invention (the "enal cleavage polypeptide route" and the "BMVO route") that enable the degradation of copalol of the labdane-type aldehyde to manool oxide, and these routes are described in more detail in subsequent sections of this specification. This scheme also shows the degradation of manool oxide to γ-ambrinol by applying a further BMVO-based degradation step.
[0014]
Chemical formula
[0015] Completely analogous to the reaction sequences exemplified above, this basic biosynthetic strategy can be applied to any other isomer of copalol or any other labdane-type aldehyde and can provide structurally related isomers of manool oxide, γ-ambril acetate, and γ-ambrinol.
[0016] Also, as will be described in more detail below, it can be applied to structurally different monocyclic or acyclic carbonyl compounds.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
[0018] Abbreviations used ADH Alcohol dehydrogenase BVMO Baeyer-Villiger monooxygenase bp Base pair kb Kilobase CPP Copalyl diphosphate CPS Copalyl diphosphate synthase DNA Deoxyribonucleic acid cDNA Complementary DNA DMAPP Dimethylallyl diphosphate DTT Dithiothreitol FMO Flavin monooxygenase FPP Farnesyl diphosphate GPP Geranyl diphosphate GGPP Geranylgeranyl diphosphate GGPS Geranylgeranyl diphosphate synthase GC Gas chromatograph IPP Isopentenyl diphosphate LPP Lambda-diol diphosphate LPS Lambda-diol diphosphate synthase MS Mass spectrometer / mass spectrometry MVA Mevalonic acid PP Diphosphate, pyrophosphate PCR Polymerase chain reaction RNA Ribonucleic acid mRNA Messenger ribonucleic acid miRNA MicroRNA siRNA Small interfering RNA rRNA Ribosomal RNA tRNA Transfer RNA TPP Terpenyl diphosphate Definitions a) General terms In the description of this specification and the appended claims, the use of "or" means "and / or" unless otherwise specified. Similarly, "comprise", "comprises", "comprising", "include", "includes", and "including" are interchangeable and are not intended to be limiting.
[0019] Furthermore, when the term "comprising" is used in the description of various embodiments, those skilled in the art will understand that in some specific examples, the embodiments can alternatively be described using the phrases "consisting essentially of" or "consisting of".
[0020] As used herein, the terms "purified," "substantially purified," and "isolated" refer to a state in which the compounds of the present invention are free of other heterogeneous compounds that are normally associated in their natural state, and the "purified," "substantially purified," and "isolated" subject constitutes at least 0.5 wt%, 1 wt%, 5 wt%, 10 wt%, or 20 wt%, or at least 50 wt% or 75 wt% of the mass of a given sample. In one embodiment, these terms refer to the compounds of the present invention constituting at least 95, 96, 97, 98, 99 or 100 wt% of the mass of a given sample. As used herein, the terms "purified," "substantially purified," and "isolated" when referring to a nucleic acid or protein, or in reference to a nucleic acid or protein, also refer to a state of purification or concentration that is different from that which occurs naturally, for example, in a prokaryotic or eukaryotic environment, such as in bacterial or fungal cells, or in a mammalian organism, particularly a human body. A degree of purity or concentration higher than that which occurs naturally, including (1) purification from other related structures or compounds, or (2) association with structures or compounds that are not normally associated in the aforementioned prokaryotic or eukaryotic environment, is within the meaning of "isolated." The nucleic acids or proteins, or classes of nucleic acids or proteins, described herein may be isolated according to various methods and processes known to those skilled in the art, or alternatively, may be associated with structures or compounds that are not normally associated in nature.
[0021] The term "about" indicates potential variation values of ±25%, particularly ±15%, ±10%, more particularly ±5%, ±2% or ±1% of the stated value.
[0022] The term "substantially" represents values in the range of about 80 - 100%, such as 85 - 99.9%, particularly 90 - 99.9%, more particularly 95 - 99.9%, or 98 - 99.9%, particularly 99 - 99.9%.
[0023] "Predominantly" refers to a proportion in a range exceeding 50%, such as in the range of, for example, 51 to 100%, particularly 75 to 99.9%, more particularly 85 to 98.5%, for example in the range of 95 to 99%.
[0024] "Major product" in the context of the present invention refers to a single compound or at least two compounds, such as a group of 2, 3, 4, 5 or more, particularly 2 or 3 compounds, and this single compound or group of compounds is "predominantly" prepared by the reaction described herein and is contained in the aforementioned reaction in a predominant proportion based on the total amount of the components of the product formed by the aforementioned reaction. The aforementioned proportion may be a molar ratio, a weight ratio, or preferably an area ratio calculated from the corresponding chromatogram of the reaction product based on chromatographic analysis.
[0025] "By-product" in the context of the present invention refers to a single compound or at least two compounds, such as a group of 2, 3, 4, 5 or more, particularly 2 or 3 compounds, and this single compound or group of compounds is not "predominantly" prepared by the reaction described herein.
[0026] Since enzyme reactions are reversible, the present invention relates to reactions in both directions of the enzyme reactions or biocatalytic reactions described herein, unless otherwise specified.
[0027] "Functional variants" of the polypeptides described herein include "functional equivalents" of the said polypeptides as defined below.
[0028] "Stereoisomers" include conformational isomers, particularly configurational isomers.
[0029] Generally included are, according to the present invention, all "stereoisomeric forms" of the compounds described herein, such as "structural isomers" and "stereoisomers".
[0030] The term "stereoisomeric form" particularly encompasses "stereoisomers" and mixtures thereof, such as configurational isomers (optical isomers) like enantiomers, or geometric isomers (diastereomers) like E- and Z-isomers, and combinations thereof. If there is one or more asymmetric centers within a single molecule, the present invention encompasses all combinations of different conformations of these asymmetric centers, such as enantiomer pairs.
[0031] "Stereoselectivity" means the ability to produce a specific stereoisomer of a compound in a stereochemically pure form, or in the enzyme-catalyzed processes described herein, the ability to specifically convert a particular stereoisomer out of a plurality of stereoisomers. More specifically, this means that the product of the present invention is enriched with respect to a specific stereoisomer, or that the starting material (educt) may be depleted with respect to a specific stereoisomer. This can be quantified via the %ee purity parameter calculated according to the following formula: %ee = [X A - X B / [X A + X B * 100 wherein X A and X B represent the molar ratios of stereoisomers A and B.
[0032] The terms "selectively convert" or "enhance selectivity" generally mean that a specific stereoisomeric form of an unsaturated hydrocarbon, such as the E-form, is converted at a higher ratio or amount (compared on a molar basis) than the corresponding other stereoisomeric form, such as the Z-form, during the entire course of the aforementioned reaction (i.e., from the start to the end of the reaction), at a certain point in the aforementioned reaction, or during an "interval" of the aforementioned reaction. In particular, the aforementioned selectivity can be observed during an "interval" corresponding to a conversion of 1 - 99%, 2 - 95%, 3 - 90%, 5 - 85%, 10 - 80%, 15 - 75%, 20 - 70%, 25 - 65%, 30 - 60%, or 40 - 50% of the initial amount of the substrate. The aforementioned higher ratio or amount can be expressed, for example, from the following viewpoints: - A higher maximum yield of the isomers observed during the entire process of the reaction or during the aforementioned interval thereof; - A higher relative amount of the isomers at a defined degree of conversion (%) of the substrate; and / or - The same relative amount of the isomers at a higher degree of conversion (%). All of these are preferably observed as compared to a reference method, which is carried out under otherwise identical conditions using known chemical or biochemical means.
[0033] Generally, in accordance with the present invention, all "isomeric forms" of the compounds described herein, such as structural isomers, in particular stereoisomers, and mixtures thereof, such as optical isomers or geometric isomers, such as E- and Z-isomers, and combinations thereof, etc. are also included. When several asymmetric centers are present in the molecule, the present invention includes all combinations of different conformations of these asymmetric centers, such as enantiomer pairs, or any mixture of stereoisomeric forms.
[0034] The "yield" and / or "conversion rate" of the reaction according to the present invention is determined over a defined period, for example 4, 6, 8, 10, 12, 16, 20, 24, 36 or 48 hours, during which the reaction is carried out. In particular, the reaction is carried out under precisely defined conditions, such as the "standard conditions" defined herein.
[0035] Different yield parameters ("yield" or Y P / S , "Specific Productivity Yield"; or Space-Time Yield (STY)) are well known in the art and are determined as described in the literature.
[0036] In this specification, "yield" and "Y P / S " (each represented by the mass of the product produced / mass of the material consumed) are used as synonyms.
[0037] Specific productivity represents the amount of product produced per hour per gram of biomass and per liter of fermentation broth. The wet cell weight, denoted as WCW, represents the amount of biologically active microorganisms in a biochemical reaction. This value is given as grams of product per gram of WCW per hour (i.e., g / g WCW -1 h -1 ). Alternatively, the amount of biomass can be expressed as the dry cell weight, denoted as DCW. Furthermore, the biomass concentration can be more easily determined by measuring the optical density (OD 600 ) at 600 nm and estimating the corresponding wet cell weight or dry cell weight using experimentally determined correlation coefficients, respectively.
[0038] When the present disclosure refers to features, parameters and ranges thereof of different priorities (including general features, parameters and ranges thereof which are not explicitly prioritized), unless otherwise specified, any combination of two or more of such features, parameters and ranges thereof is included in the disclosure herein, regardless of their respective priorities.
[0039] b) Biochemical terms The term "domain" refers to a partial sequence of a series of amino acids or amino acid residues that are conserved at specific positions along the alignment of evolutionarily related protein sequences. Amino acids at other positions may vary between protein homologs, but amino acids that are highly conserved at specific positions in the domain indicate amino acids that are likely essential for the structure, stability or function of the protein. Identified from the high conservation of amino acids in the aligned sequences of a family of protein homologs, it can be used as an identifier to determine whether any polypeptide in question belongs to a pre-identified polypeptide family.
[0040] The term "motif" or "consensus sequence" or "signature" refers to a short conserved region in the sequences of evolutionarily related proteins. A motif is often highly conserved as part of a domain, but may also contain only a part of a domain.
[0041] A "protein family" is defined as a group of proteins with a common evolutionary origin, reflected by their related functions, sequence similarity, or similar primary, secondary, or tertiary structures. Proteins within a protein family usually have homology and have similar structures of conserved functional domains and motifs.
[0042] There are specialized databases for identifying domains, for example, SMART (Schultz et al. (1998) Proc. Natl. Acad. Sci. USA 95, 5857-5864; Letunic et al. (2002) Nucleic Acids Res 30, 242-244), InterPro (Mulder et al., (2003) Nucl. Acids. Res. 31, 315-318), Prosite (Bucher and Bairoch (1994), A generalized profile syntax for biomolecular sequences motifs and its function in automatic sequence interpretation. (In) ISMB-94; Proceedings 2nd International Conference on Intelligent Systems for Molecular Biology. Altman R., Brutlag D., Karp P., Lathrop R., Searls D., Eds., pp 53-61, AAAI Press, Menlo Park; Hulo et al., Nucl. Acids. Res. 32:D134-D137, (2004)), or Pfam (Bateman et al., Nucleic Acids Research 30(1): 276-280 (2002)). Domains or motifs may be identified using conventional techniques such as sequence alignment.
[0043] The term "Pfam" refers to a large collection of protein domains and protein families maintained by the Pfam Consortium and available on several sponsored World Wide Web sites such as http: / / pfam.xfam.org / (European Molecular Biology Laboratory-European Bioinformatics Institute (EMBL_EBI)). The latest release of Pfam is Pfam 32.0 (September 2018) and is based on the UniProt Reference Proteomes (El-Gebali S. et al, 2019, Nucleic Acids Res. 47, Database issue D427-D432). Pfam domains and families are identified using multiple sequence alignments and hidden Markov models (HMMs). Assignments of Pfam-A families or domains are high-quality assignments generated by a seed alignment vetted with representative members of the protein family and a profile hidden Markov model based on the seed alignment (unless otherwise specified, a match of a queried protein to a Pfam domain or family is a Pfam-A match). A full alignment of the family is then automatically generated using all identified sequences belonging to the family (Sonnhammer (1998) Nucleic Acids Research 26, 320-322; Bateman (2000) Nucleic Acids Research 26, 263-266; Bateman (2004) Nucleic Acids Research 32, Database Issue, D138-D141; Finn (2006) Nucleic Acids Research Database Issue 34, D247-251; Finn (2010) Nucleic Acids Research Database Issue 38, D211-222).For example, when accessing the Pfam database using any of the above reference websites, a protein sequence can be queried against an HMM using HMMER homology search software (e.g., HMMER2, HMMER3, or a higher version (hmmer.janelia.org / )). A significant match that identifies that the queried protein belongs to a pfam family (or has a specific Pfam domain) is one where the bit score is above the collection threshold of the Pfam domain. The expected value (e-value) can also be used as a criterion for including the queried protein in Pfam or for determining whether the queried protein has a specific Pfam domain, where a low e-value is a value much lower than 1.0, e.g., less than or below 0.1.
[0044] The "E-value" (expected value) is the expected number of hits that would have a score equal to or greater than this value by chance. That is, a good E-value for a reliable prediction is a value much lower than 1. E-values around 1 are values expected by chance. Thus, the lower the E-value, the more specific the domain search becomes. Only positive numbers are allowed (definition by Pfam).
[0045] The "precursor" molecules of the target compounds described herein are preferably converted to the aforementioned target compounds by the enzymatic action of an appropriate polypeptide that effects at least one structural change on the aforementioned precursor molecules. For example, a "diphosphate precursor" (e.g., a "terpenyl diphosphate precursor") is converted to the aforementioned target compound (e.g., a terpene alcohol) through the enzymatic removal of the diphosphate moiety, e.g., by the removal of a monophosphate or diphosphate group by a phosphatase enzyme. For example, an "acyclic precursor" (such as an acyclic terpenyl precursor) can be converted to a cyclic target molecule (such as a cyclic terpene compound) in one or more steps through the action of a cyclic or synthetic enzyme, regardless of the specific enzymatic mechanism of such an enzyme.
[0046] The term "protein tyrosine phosphatase" refers to a group of enzymes that are generally known to remove phosphate groups from phosphorylated tyrosine residues on proteins. A particular subgroup of the aforementioned family described herein are enzymes useful for dephosphorylating phosphorylated terpene molecules.
[0047] "Terpene synthase" refers to a polypeptide that converts a terpene precursor molecule into a terpene target molecule, particularly referring to a processed target terpene alcohol or terpene hydrocarbon, etc. Non-limiting examples of such terpene precursor molecules are, for example, acyclic compounds selected from farnesyl pyrophosphate (FPP), geranylgeranyl pyrophosphate (GGPP), or a mixture of isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP). When the resulting terpene contains a diphosphate moiety, the synthase is called a "terpenyl diphosphate synthase".
[0048] The term "terpenyl diphosphate synthase" or "polypeptide having terpenyl diphosphate synthase activity" or "terpenyl diphosphate synthase protein" or "having the ability to produce terpenyl diphosphate" refers to a polypeptide that can catalyze the synthesis of terpenyl diphosphate starting from acyclic terpene pyrophosphates, particularly GPP, FPP or GGPP or a combination of IPP and DMAPP, in the form of any of its stereoisomers, or a mixture thereof. Terpenyl diphosphate may be the sole product or may be part of a mixture of terpenyl phosphates. The aforementioned mixture may contain terpenyl monophosphate and / or terpene alcohol. The above definition also applies to the group of "bicyclic terpenyl diphosphate synthases" that produce bicyclic terpenyl diphosphates such as CPP or LPP. Examples of such "terpenyl diphosphate synthases" include copalyl diphosphate synthase (CPS). Copalyl diphosphate may be the sole product or may be part of a mixture of copalyl phosphates. The aforementioned mixture may contain copalyl monophosphate and / or other terpenyl diphosphates. Another example of such "terpenyl diphosphate synthase" is labdadienol diphosphate synthase (LPS). Labdadienol diphosphate may be the sole product or may be part of a mixture of labdadienol phosphates. The aforementioned mixture may contain labdiene monophosphate and / or terpenyl diphosphate.
[0049] The term "terpenyl diphosphate phosphatase" or "polypeptide having terpenyl diphosphate phosphatase activity" or "terpenyl diphosphate phosphatase protein" or "having the ability to produce terpene alcohol" relates to a polypeptide that can catalyze the removal of a diphosphate moiety or a monophosphate moiety (regardless of the specific enzyme mechanism) to form a dephosphorylated compound, particularly the corresponding alcohol compound of the aforementioned terpenyl moiety. The terpene alcohol may be present in the product as any of its stereoisomers or as a mixture thereof. The terpene alcohol may be the sole product or may be part of a mixture with other terpene compounds, such as, for example, dephosphorylated analogs of the respective (e.g., acyclic) precursors of the aforementioned terpenyl diphosphate. The above definition also applies to the group of "bicyclic terpenyl diphosphate phosphatases" that produce bicyclic terpene alcohols such as copalol or labda-diol.
[0050] Examples of such "terpenyl diphosphate phosphatase" enzymes include copalyl diphosphate phosphatase (CPP phosphatase). Copalol may be the sole product or may be part of a mixture with, for example, dephosphorylated precursors such as farnesol and / or geranylgeraniol; and / or by-products resulting from the secondary activity of the enzyme in the reaction mixture, such as mixtures with such alcohols or esters or aldehydes of other cyclic or acyclic diterpenes. Another example of such a "terpenyl diphosphate phosphatase" enzyme is labda-diol diphosphate phosphatase (LPP phosphatase). Labda-diol may be the sole product or may be part of a mixture with, for example, dephosphorylated precursors such as farnesol and / or geranylgeraniol; and / or by-products resulting from the secondary activity of the enzyme in the reaction mixture, such as mixtures with such alcohols or esters or aldehydes of other cyclic or acyclic diterpenes.
[0051] As used herein, the terms "enal cleavage enzyme", "enal cleavage protein", or "enal cleavage polypeptide" refer to an "α,β-unsaturated aldehyde carbon-carbon double bond cleavage enzyme", which may also be referred to as an "α,β-unsaturated aldehyde C=C bond cleavage enzyme", an "α,β-unsaturated aldehyde C=C cleavage enzyme", or an "enal C=C cleavage enzyme". The enal cleavage proteins of the present invention can also be represented as members of the "DUF4334 protein family" and / or as members of the "GXWXG protein family" based on the domain composition of the proteins.
[0052] More specifically, the enal cleavage enzymes of the present invention have the ability to cleave labdane-type carbonyl compounds, such as labdanal, especially copalal, into their respective dinorlabdane carbonyl compounds. "Baeyer-Villiger monooxygenases" (BVMOs) are flavin enzymes and belong to the class of polypeptides (EC 1.14.13.X) having oxidoreductase activity. These enzymes are catalysts that oxidize linear, cyclic (aromatic or non-aromatic) aldehydes or ketones to the corresponding esters or lactones, and are very similar to chemical Baeyer-Villiger oxidations. During enzymatic oxidation, one atom of molecular oxygen is incorporated into the carbon-carbon bond of the non-activated carbonyl compound. BVMOs require NADPH or NADH, or both, as cofactors. BVMOs also require molecular oxygen as a cosubstrate. More specifically, the BVMOs of the present invention have the ability to oxidize terpene-derived aldehydes or ketones, such as labdane-type carbonyl compounds, such as labdanal, especially copalal and / or manool oxide, to their respective carbonyl esters.
[0053] "Esterase" refers to a polypeptide having hydrolase activity that decomposes esters into an acid and an alcohol in a chemical reaction (hydrolysis) with water. The esterase in the context of the present invention is selected from the class of carboxylic ester hydrolases (EC 3.1.1.-) and cleaves an acyl group such as an acetyl group or a formyl group from each ester substrate. More specifically, the esterase of the present invention has the ability to cleave a labdane-type ester compound such as γ-ambrinol acetate to form each labdane-type alcohol such as γ-ambrinol.
[0054] "Alcohol dehydrogenase" (ADH) in the context of the present invention refers to a polypeptide having the ability to oxidize an alcohol to the corresponding aldehyde in the presence of NAD + or NADP + as a cofactor. Such an enzyme is a member of the E.C. family 1.1.1.1 (NAD + -dependent) or 1.1.1.2 (NADP + -dependent). More specifically, the ADH of the present invention has the ability to oxidize labdane-type alcohols to their respective labdane-type carbonyl compounds (aldehydes or ketones) such that coparol is oxidized to copal and / or labdanediol is oxidized to copal, each aldehyde of labdanediol or other labdane-type derivatives, for example, the nor- or dinor-labdane-type derivatives of coparol or labdanediol. ADH as used herein may be endogenous to each biocatalytic process or may be exogenous.
[0055] The "enal cleavage activity" is determined under the "standard conditions" described later in this specification. The enal cleavage activity can be determined using a host cell expressing a recombinant enal cleavage polypeptide, a disrupted cell expressing an enal cleavage polypeptide, these fractions, or an enriched or purified enal cleavage polypeptide, in a preferably buffered culture medium or reaction medium having a pH in the range of 6 to 11, preferably 7 to 9, at a temperature of about 20 to 45 °C, for example about 25 to 40 °C, preferably 25 to 32 °C, by adding at an initial concentration of 1 to 100 μM mg / ml, preferably 5 to 50 μM, particularly 30 to 40 μM, or in the presence of a reference substrate endogenously produced by the host cell, here particularly copalal. The conversion reaction for forming each cleavage product such as manool oxide is carried out for 10 minutes to 5 hours, preferably about 1 to 2 hours. The cleavage product can then be determined by conventional methods after extraction with an organic solvent such as ethyl acetate, for example.
[0056] The "BVMO activity" is determined under the "standard conditions" described later in this specification. The BVMO activity can be determined using a host cell expressing recombinant BVMO, a cell expressing disrupted BVMO, these fractions, or a concentrated or purified BVMO enzyme, in a preferably buffered culture medium or reaction medium having a pH in the range of 6 to 11, preferably 7 to 9, at a temperature of about 20 to 45 °C, such as about 25 to 40 °C, preferably 25 to 32 °C, with an initial concentration of 1 to 100 μM mg / ml, preferably 5 to 50 μM, particularly 30 to 40 μM, or by adding a reference substrate endogenously produced by the host cell, here in particular copalol and / or manool oxide, in the presence of molecular oxygen. In an in vitro assay, a coenzyme selected from NADH and NADPH must be added in a suitable concentration range that allows for easy determination. The conversion reaction to form the respective cleavage products, such as formyl esters 1a and / or 1b in the case of copalol and γ-ambril acetate in the case of manool oxide, is carried out for 10 minutes to 5 hours, preferably about 1 to 2 hours. The oxidation products can then be determined by conventional methods after extraction with an organic solvent such as ethyl acetate.
[0057] "Terpenyl diphosphate synthase activity" (e.g., CPS or LPS activity) is determined under the "standard conditions" described later in this specification. Terpenyl diphosphate synthase activity can be determined using host cells expressing recombinant terpenyl diphosphate synthase, disrupted cells expressing terpenyl diphosphate synthase, fractions thereof, or concentrated or purified terpenyl diphosphate synthase in a preferably buffered culture medium or reaction medium having a pH in the range of 6 to 11, preferably 7 to 9, at a temperature of about 20 to 45 °C, for example about 25 to 40 °C, preferably 25 to 32 °C, with an initial concentration of 1 to 100 μM mg / ml, preferably 5 to 50 μM, particularly 30 to 40 μM, added or in the presence of a reference substrate endogenously produced by the host cell, here particularly GGPP. The conversion reaction for forming terpenyl diphosphate is carried out for 10 minutes to 5 hours, preferably about 1 to 2 hours. In the absence of endogenous phosphatase, one or more exogenous phosphatases, such as alkaline phosphatase, are added to the reaction mixture to convert the terpenyl diphosphate formed by the synthase to the respective terpene alcohol. The terpene alcohol can then be determined by conventional methods after extraction with an organic solvent such as ethyl acetate.
[0058] "Terpenyl diphosphate PhosphataseThe "activity" (e.g., CPP or LPP phosphatase activity) is determined under the "standard conditions" described later in this specification. The terpenyl diphosphate phosphatase activity can be determined using a host cell expressing recombinant terpenyl diphosphate phosphatase, a disrupted cell expressing terpenyl diphosphate phosphatase, these fractions, or a concentrated or purified terpenyl diphosphate phosphatase enzyme, in a preferably buffered culture medium or reaction medium having a pH in the range of 6 - 11, preferably 7 - 9, and adding at an initial concentration of 1 - 100 μM mg / ml, preferably 5 - 50 μM, particularly 30 - 40 μM, or can be determined in the presence of a reference substrate endogenously produced by the host cell, here for example CPP or LPP. The conversion reaction for forming terpenyl diphosphate is carried out for 10 minutes to 5 hours, preferably about 1 - 2 hours. Then, the terpene alcohol can be extracted with an organic solvent such as ethyl acetate and then determined by conventional methods.
[0059] Specific examples of the standard conditions suitable for each of the above enzyme activities can be obtained from the later "Experimental Section".
[0060] The terms "biological function", "function", "biological activity", or "activity" of terpenyl synthase refer to the ability of the terpenyl diphosphate synthase described in this specification to catalyze the formation of at least one terpenyl diphosphate from the corresponding precursor terpene.
[0061] The terms "biological function", "function", "biological activity", or "activity" of terpenyl diphosphate phosphatase refer to the ability of the terpenyl diphosphate phosphatase described in this specification to catalyze the removal of a diphosphate group from the aforementioned terpenyl compound to form the corresponding terpene alcohol.
[0062] The "mevalonate pathway", also called the "isoprenoid pathway" or the "HMG-CoA reductase pathway", is an essential metabolic pathway present in eukaryotes, archaea, and some bacteria. The mevalonate pathway starts from acetyl-CoA and generates two C5 building blocks called isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP). The main enzymes are acetoacetyl-CoA thiolase (atoB), HMG-CoA synthase (mvaS), HMG-CoA reductase (mvaA), mevalonate kinase (MvaK1), phosphomevalonate kinase (MvaK2), mevalonic acid diphosphate decarboxylase (MvaD), and isopentenyl diphosphate isomerase (idi). By combining enzyme activities in the mevalonate pathway that generate terpene precursors such as GPP, FPP, or GGPP, especially FPP synthase (ERG20), etc., recombinant cell production of terpenes becomes possible.
[0063] As used herein, the terms "host cell" or "transformed cell" refer to a cell (or organism) that has been modified to harbor at least one nucleic acid molecule, e.g., a recombinant gene encoding a desired protein or a nucleic acid sequence that results in at least one functional polypeptide of the present invention, particularly a terpene diphosphate synthase protein or a terpene diphosphate phosphatase enzyme as defined hereinabove. Host cells are particularly bacterial cells, fungal cells, or plant cells or plants. The host cell may contain recombinant genes or multiple genes incorporated into the genome of the host cell nucleus or organelle, e.g., configured as an operon. Alternatively, the host may contain the recombinant gene extrachromosomally.
[0064] "Organism" refers to any non-human multicellular or unicellular organism such as a plant, or a microorganism. In particular, microorganisms are bacteria, yeast, algae, or fungi.
[0065] The term "plant" is used interchangeably to include plant cells including plant protoplasts, plant tissues, plant cell tissue cultures that give rise to regenerated plants, or parts of plants, or plant organs such as roots, stems, leaves, flowers, pollen, ovules, embryos, fruits, etc. Any plant can be used to carry out the methods of the embodiments herein.
[0066] A particular organism or cell "is capable of producing FPP" if it naturally produces FPP or if it does not naturally produce FPP but has been transformed to produce FPP using the nucleic acids described herein. Organisms or cells transformed to produce more FPP than naturally occurring organisms or cells are also included within "organisms or cells capable of producing FPP".
[0067] A particular organism or cell "is capable of producing GGPP" if it naturally produces GGPP or if it does not naturally produce GGPP but has been transformed to produce GGPP using the nucleic acids described herein. Organisms or cells transformed to produce more GGPP than naturally occurring organisms or cells are also included within "organisms or cells capable of producing GGPP".
[0068] A particular organism or cell "is capable of producing terpenyl diphosphate" if it naturally produces the terpenyl diphosphate as defined herein or if it does not naturally produce the aforementioned diphosphate but has been transformed to produce the aforementioned diphosphate using the nucleic acids described herein. Organisms or cells transformed to produce more terpenyl diphosphate than naturally occurring organisms or cells are also included within "organisms or cells capable of producing terpenyl diphosphate".
[0069] A particular organism or cell is said to "be able to produce terpene alcohol" if it naturally produces the terpene alcohol as defined herein, or if it does not naturally produce the aforementioned alcohol but is transformed to produce the aforementioned alcohol using the nucleic acids described herein. Organisms or cells transformed to produce more terpene alcohol than naturally occurring organisms or cells are also included in "organisms or cells capable of producing terpene alcohol". The same applies to particular organisms that "are able to produce labdane-type alcohol".
[0070] A particular organism or cell is said to "be able to produce an ester" if it naturally produces the ester as defined herein, or if it does not naturally produce the aforementioned ester but is transformed to produce the aforementioned ester using the nucleic acids described herein. Organisms or cells transformed to produce more ester than naturally occurring organisms or cells are also included in "organisms or cells capable of producing an ester".
[0071] A particular organism or cell is said to "be able to produce a target product" (e.g., an ester, an alcohol, or a carbonyl compound, more particularly a labdane-type compound) if it naturally produces the aforementioned target product, or if it does not naturally produce the aforementioned target product but is transformed to produce the aforementioned target product using the nucleic acids described herein. Organisms or cells transformed to produce more target product than naturally occurring organisms or cells are also included in "organisms or cells capable of producing a target product".
[0072] The term "fermentative production" or "fermentation" means the ability of a microorganism to produce a compound in cell culture using at least one carbon source added to the incubation (aided by the enzyme activity contained in or produced by the aforementioned microorganism).
[0073] The term "fermentation broth", as described herein for example, is understood to mean a liquid, particularly an aqueous solution or an aqueous / organic solution, which is based on a fermentation process and which is either unprocessed or processed.
[0074] The "enzymatic catalysis" method or "biocatalysis" method, as defined herein, means that the aforementioned method is carried out under the catalysis of an enzyme (including variants of the enzyme). Therefore, this method can be carried out in the presence of the aforementioned enzyme in isolated (purified, concentrated) or crude form, or in the presence of a cell system, particularly a natural or recombinant microbial cell containing the aforementioned enzyme in active form and having the ability to catalyze the conversion reaction disclosed herein.
[0075] c) Chemical terms: The term "α,β-unsaturated carbonyl" compound refers to an organic molecule containing an aldehyde group or a ketone group of the general formula R a R b C=C(R c )-C=O, wherein the C=C bond may be in any stereoisomeric configuration, and the residues R a , R b and R c may be the same or different and may have the meanings defined below for a particular α,β-unsaturated carbonyl compound.
[0076] The "lubdan" compound in the context of the present invention has the following basic structure consisting of 20 carbon atoms in its carbon skeleton. The numbering of the depicted carbon atoms is applied to further define specific positions within the aforementioned carbon skeleton.
[0077]
Chemical formula
[0078] The term "lubdan" refers to this basic C 20It encompasses any compound of the structure in any stereoisomeric form and variants of this structure containing one or more unsaturated C-C bonds, especially one or more C=C bonds, at any position within the carbocyclic ring and / or side chain. Also included are its variants containing substituents selected from the group of one or more substituents such as -OH, =O, -O-CO_R (wherein R is a linear or branched alkyl, especially lower alkyl, more particularly methyl, ethyl, n- or i-propyl, or n-, i- or t-butyl C 1 ~C 4 alkyl, and -COOH may be present at any of the indicated primary, secondary or tertiary C atoms).
[0079] Such "labdane"-derived compounds of "labdane" include compounds in which the basic C 20 carbon skeleton is modified by removing one or more carbon atoms. Examples include the following: Norlabdane (C 19 skeleton), dinorlabdane (C 18 skeleton), trinorlabdane (C 17 skeleton), tetranorlabdane (C 16 skeleton). The position of the removed carbon atoms is indicated by stating the carbon number. For example, in norlabdane where the 15th carbon is missing, it is denoted as "15-norlabdane".
[0080] Such "labdane"-derived compounds of "labdane" also include compounds in which the basic C 20 carbon skeleton is modified by inserting a heteroatom between two C atoms of the labdane skeleton. For example, by inserting an ether bridge between the 14th and 15th positions, labdane is converted to norlabdane, especially norlabdane ester.
[0081] Non-limiting examples of substituted labdane or substituted labdane-derived structures are shown below:
Chemical formula
[0082] As used herein, "diphosphate" and "pyrophosphate" are synonymous.
[0083] "Terpenes" are a diverse class of organic compounds produced by a variety of plants, especially conifers, and some insects. Terpenes are hydrocarbons. Although sometimes used interchangeably with "terpenes," "terpenoids" or "isoprenoids" are modified terpenes because they contain additional functional groups, usually oxygen-containing.
[0084] "Terpenoids" ("isoprenoids") are a broad and diverse class of naturally occurring organic chemicals derived from terpenes. Although sometimes used interchangeably with the term "terpenes," "terpenoids" contain additional functional groups, usually containing O, such as hydroxyl, carbonyl, or carboxyl groups. Most are polycyclic structures with oxygen-containing functional groups. Unless otherwise noted, in the context of this specification, the terms "terpene" and "terpenoid" may be used interchangeably.
[0085] Terpenes (and terpenoids) can be classified by the number of isoprene units in the molecule, with the prefix in the name indicating the number of terpene units required to assemble the molecule. Hemiterpenes consist of one isoprene unit. Monoterpenes consist of two isoprene units and have the molecular formula C 10 H 16 Sesquiterpenes consist of three isoprene units and have the molecular formula C 15 H 24 Diterpenes consist of four isoprene units and have the molecular formula C 20 H 32 It is.
[0086] "Terpenyls" are C 5 It refers to acyclic and cyclic chemical hydrocarbyl residues derived from the building block isoprene, and specifically includes one or more such building blocks.
[0087] The term "cyclic terpene" or "cyclic terpenyl" or "cyclic diterpene" or "cyclic diterpenyl" refers to a terpene compound or terpenyl residue containing at least one, for example, 1, 2, 3, 4 or 5 carbon cyclic condensed rings and / or non-condensed rings in its structure, preferably 2 carbon cyclic condensed rings.
[0088] The term "bicyclic terpene" or "bicyclic terpenyl" or "bicyclic diterpene" or "bicyclic bicycloditerpenyl" refers to a terpene compound or terpenyl residue containing 2 carbon cyclic rings in its structure, preferably 2 carbon cyclic condensed rings.
[0089] In the context of the present invention, the "derivative of terpenes" or "derivative of terpenoids" particularly refers to a compound obtained from a terpene or terpenoid by chemical and / or enzymatic modification. More specifically, such derivatives include "derivatives with decomposed hydrocarbon chains".
[0090] A terpene or terpenoid with a "decomposed hydrocarbon chain" has a reduced number of carbon atoms in the carbon skeleton of the precursor compared to the undecomposed precursor.
[0091] A "hydrocarbyl" residue is a chemical group essentially composed of carbon atoms and hydrogen atoms, which may be an acyclic, linear or branched saturated or unsaturated moiety, or a cyclic saturated or unsaturated moiety, aromatic or non-aromatic moiety. The hydrocarbyl residue contains 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10 or 1 to 5 carbon atoms in the case of an acyclic structure. The hydrocarbyl residue contains 4 to 30, 4 to 25, 4 to 20, 4 to 15, 4 to 10, or particularly 4, 5, 6 or 7 carbon atoms in the case of a cyclic structure.
[0092] The aforementioned hydrocarbyl residue may be unsubstituted or may have at least one, for example, 1 to 5, preferably 0, 1 or 2 substituents.
[0093] Specific examples of such hydrocarbyl residues are acyclic linear or branched alkyl or alkenyl residues as defined below; or, for example, cyclic (e.g., bicyclic) or acyclic terpene-type compounds, and monocyclic or polycyclic, particularly monocyclic or bicyclic, saturated or unsaturated non-aromatic moieties such as those found in labdane-type compounds.
[0094] "Alkyl" residues represent linear or branched saturated hydrocarbon residues. "Alkyl" residues contain from 1 to 30, 1 to 25, 1 to 20, 1 to 15, or 1 to 10, or 1 to 7, 1 to 6, 1 to 5, or 1 to 4 carbon atoms.
[0095] "Alkenyl" residues represent linear or branched, mono- or polyvalent unsaturated hydrocarbon residues. "Alkenyl" residues contain from 2 to 30, 2 to 25, 2 to 20, 2 to 15, or 2 to 10, or 2 to 7, 2 to 6, 2 to 5, or 2 to 4 carbon atoms. "Alkenyl" residues may have up to 10, for example 1, 2, 3, 4 or 5, C=C double bonds.
[0096] The term "lower alkyl" or "short-chain alkyl" represents a saturated straight-chain or branched hydrocarbon radical having 1 to 4, 1 to 5, 1 to 6, or 1 to 7, particularly 1 to 4 carbon atoms. Examples include: methyl, ethyl, n-propyl, 1-methylethyl, n-butyl, 1-methylpropyl, 2-methylpropyl, 1,1-dimethylethyl, n-pentyl, 1-methylbutyl, 2-methylbutyl, 3-methylbutyl, 2,2-dimethylpropyl, 1-ethylpropyl, n-hexyl, 1,1-dimethylpropyl, 1,2-dimethylpropyl, 1-methylpentyl, 2-methylpentyl, 3-methylpentyl, 4-methylpentyl, 1,1-dimethylbutyl, 1,2-dimethylbutyl, 1,3-dimethylbutyl, 2,2-dimethylbutyl, 2,3-dimethylbutyl, 3,3-dimethylbutyl, 1-ethylbutyl, 2-ethylbutyl, 1,1,2-trimethylpropyl, 1,2,2-trimethylpropyl, 1-ethyl-1-methylpropyl, 1-ethyl-2-methylpropyl; and n-heptyl, as well as their mono- or multi-branched analogs.
[0097] "Long-chain alkyl" represents, for example, a saturated straight-chain or branched hydrocarbyl radical having 8 to 30, for example 8 to 20 or 8 to 15 carbon atoms, such as octyl, nonyl, decyl, undecyl, dodecyl, tridecyl, tetradecyl, pentadecyl, hexadecyl, heptadecyl, octadecyl, nonadecyl, eicosyl, henicosyl, docosyl, tricosyl, tetracosyl, pentacosyl, hexacosyl, heptacosyl, octacosyl, nonacosyl, squalyl, structural isomers, particularly their mono- or multi-branched isomers.
[0098] "Long-chain alkenyl" represents a mono- or polyunsaturated analog of the above-mentioned "long-chain alkyl" group.
[0099] "Short-chain alkenyl" (or "lower alkenyl") represents a mono- or polyunsaturated, particularly monounsaturated, straight-chain or branched hydrocarbon radical having 2 to 4, 2 to 6, or 2 to 7 carbon atoms and one double bond at any position, and examples include C 2 ~C 6 -alkenyl, such as ethenyl, 1-propenyl, 2-propenyl, 1-methylethenyl, 1-butenyl, 2-butenyl, 3-butenyl, 1-methyl-1-propenyl, 2-methyl-1-propenyl, 1-methyl-2-propenyl, 2-methyl-2-propenyl, 1-pentenyl, 2-pentenyl, 3-pentenyl, 4-pentenyl, 1-methyl-1-butenyl, 2-methyl-1-butenyl, 3-methyl-1-butenyl, 1-methyl-2-butenyl, 2-methyl-2-butenyl, 3-methyl-2-butenyl, 1-methyl-3-butenyl, 2-methyl-3-butenyl, 3-methyl-3-butenyl, 1,1-dimethyl-2-propenyl, 1,2-dimethyl-1-propenyl, 1,2-dimethyl-2-propenyl, 1-ethyl-1-propenyl, 1-ethyl-2-propenyl, 1-hexenyl, 2-hexenyl, 3-hexenyl, 4-hexenyl, 5-hexenyl, 1-methyl-1-pentenyl, 2-methyl-1-pentenyl, 3-methyl-1-pentenyl, 4-methyl-1-pentenyl, 1-methyl-2-pentenyl, 2-methyl-2-pentenyl, 3-methyl-2-pentenyl, 4-methyl-2-pentenyl, 1-methyl-3-pentenyl, 2-methyl-3-pentenyl, 3-methyl-3-pentenyl, 4-methyl-3-pentenyl, 1-methyl-4-pentenyl, 2-methyl-4-pentenyl, 3-methyl-4-pentenyl, 4-methyl-4-pentenyl, 1,1-dimethyl-2-butenyl, 1,1-dimethyl-3-butenyl, 1,2-dimethyl-1-butenyl, 1,2-dimethyl-2-butenyl, 1,2-dimethyl-3-butenyl, 1,3-dimethyl-1-butenyl 1,3-dimethyl-2-butenyl, 1,3-dimethyl-3-butenyl, 2,2-dimethyl-3-butenyl, 2,3-Dimethyl-1-butenyl, 2,3-dimethyl-2-butenyl, 2,3-dimethyl-3-butenyl, 3,3-Dimethyl-1-butenyl, 3,3-dimethyl-2-butenyl, 1-ethyl-1-butenyl, 1-ethyl-2-butenyl, 1-ethyl-3-butenyl, 2-ethyl-1-butenyl, 2-ethyl-2-butenyl, 2-ethyl-3-butenyl, 1,1,2-Trimethyl-2-propenyl, 1-ethyl-1-methyl-2-propenyl, 1-ethyl-2-methyl-1-propenyl and 1-ethyl-2-methyl-2-propenyl are exemplified.
[0100] "Alkylene" represents a straight-chain or mono- or multi-branched hydrocarbon bridging group having 1 to 10 carbon atoms, for example, -CH 2 , -(CH 2 ) 2 -, -(CH 2 ) 3 -, -(CH 2 ) 4 -, -(CH 2 ) 2 -CH(CH 3 )-, -CH 2 -CH(CH 3 )-CH 2 -, (CH 2 ) 4 -, -(CH 2 ) 5 -, -(CH 2 ) 6 , -(CH 2 ) 7 -, -CH(CH 3 )-CH 2 -CH 2 -CH(CH 3 )- or -CH(CH 3 )-CH 2 -CH 2 -CH 2 -CH(CH 3 )- selected from C 1 ~C 7 -alkylene groups, especially, -CH 2 -, -(CH 2 ) 2 -, -(CH2 ) 3 -, -(CH 2 ) 4 -, -(CH 2 ) 2 -CH(CH 3 )-, -CH 2 -CH(CH 3 )-CH 2 -selected C 1 ~C 4 -alkylene groups are exemplified.
[0101] The "alkylidene" group represents a substituent of a linear or branched hydrocarbon bonded to the main body of the molecule via a double bond. The "alkylidene" group contains 1 to 6 carbon atoms. Such "C 1 ~C 6 -alkylidene" examples include methylidene (=CH 2 ), ethylidene, (=CH-CH 2 ), n-propylidene, n-butylidene, n-pentylidene, n-hexylidene, and its structural isomers such as isopropylidene.
[0102] "Alkenylidene" represents a mono-unsaturated analog of the above alkylidene having more than 2 carbon atoms and may also be called "C 3 ~C 6 -alkenylidene". Examples include n-propenylidene, n-butenylidene, n-pentenylidene, and n-hexenylidene.
[0103] The "substituent" of the above residue contains one heteroatom such as O or N. Preferably, the substituents are independently selected from -OH, C=O, or -COOH. Most preferably, the aforementioned substituent is -OH.
[0104] "Monocyclic or polycyclic hydrocarbyl residue" includes one, two or three fused (annelated) or non-fused, optionally substituted, saturated or unsaturated hydrocarbon ring groups (or "carbocyclic" groups). Each ring may independently contain 3 to 8, particularly 5 to 7, more particularly 6 ring carbon atoms. Examples of monocyclic residues are "cycloalkyl" groups which are carbocyclic radicals having 3 to 7 carbocyclic atoms such as cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, cyclooctyl, etc.; and the corresponding "cycloalkenyl" groups. "Cycloalkenyl" (or "mono- or polyunsaturated cycloalkyl") particularly represents a monocyclic, mono- or polyunsaturated carbocyclic group having 5 to 8, preferably at most 6 carbocyclic members, and examples thereof include monounsaturated cyclopentenyl, cyclohexenyl, cycloheptenyl and cyclooctenyl radicals.
[0105] Examples of polycyclic residues include groups in which one, two or three such cycloalkyl and / or cycloalkenyl are joined together, for example by annelation, to form a polycyclic cycloalkyl or cycloalkenyl ring. As a non-limiting example, a bicyclic decalinyl residue composed of two fused 6-membered carbocyclic rings can be mentioned.
[0106] The number of substituents of such monocyclic or polycyclic hydrocarbyl residues may vary from 1 to 10, particularly 1 to 5 substituents. Suitable substituents for such cyclic residues are selected from lower alkyl, lower alkenyl, alkylidene, alkenylidene, or residues containing one heteroatom such as O or N, for example -OH or -COOH. In particular, the substituents are independently selected from -OH, -COOH, methyl and methylidene.
[0107] The unsaturated cyclic groups may contain one or more, for example 1, 2 or 3 C=C bonds and may be aromatic or particularly non-aromatic.
[0108] The above-mentioned monocyclic or polycyclic saturated or unsaturated group may contain at least one, for example, 1, 2, 3 or 4 ring heteroatoms such as O, N or S.
[0109]
Table 1
[0110]
Table 2-1
[0111]
Table 2-2
[0112]
Table 2-3
[0113]
Table 2-4
[0114]
Table 2-5
[0115]
Table 2-6
[0116]
Table 2-7
[0117]
Table 2-8
[0118]
Table 2-9
[0119]
Table 2-10
[0120]
Table 2-11
[0121]
Table 2-12
[0122]
Table 2-13
[0123] Detailed Description a. Specific Embodiments of the Invention i) The present invention relates to the following specific embodiments of a biocatalytic method comprising the use of a polypeptide having BMVO activity 1. A biocatalytic method for preparing an ester compound, comprising (1) General formula I
Chemical formula
[0124] 2. In the carbonyl compound of general formula I, "a" represents a chemical double bond, and Z represents =O (refer to manool oxide) or =C(R 4 )-C(R 5 )=O (refer to copalal), or "a" represents a chemical single bond, and Z represents -C(R 5 )=O (refer to norlabdane compounds 3a, 3b), wherein, R 4 and R 5 are, independently of each other, H or lower alkyl such as C 1 ~C 4 -alkyl, especially H or methyl, according to the biocatalytic method described in Embodiment 1.
[0125] 3. The biocatalytic method according to any one of the preceding embodiments, wherein the carbonyl compound of general formula I has a labdane-type structure, especially a labdane structure, a norlabdane structure or a dinorlabdane structure.
[0126] 4. The carbonyl ester formed is of formula II
Chemical formula
[0127] 5. The carbonyl ester group E is -O-C(O)-R 1 , -C(R 1 ) 2 -O-C(O)R 5 , -C(R 1 )=C(R 4 )-O-C(R 5 )=O; and E and R 2 together with the carbon atom to which they are attached form a cyclic ester group, where the cyclic ester ring is, for example, of formula IIa and IIb
Chemical formula
[0128] 6. R 2 represents the group Cyc-A-, where A represents a straight-chain or branched C 1 ~C 4 -alkylene bridge, particularly methylene, and Cyc represents a monocyclic or polycyclic, particularly bicyclic, saturated or unsaturated hydrocarbyl residue, particularly a bicyclic annelated hydrocarbyl residue containing 5 to 7, particularly 6, ring atoms per ring, where Cyc is optionally substituted with 1 to 10, particularly 1 to 5 substituents, and particularly, the aforementioned substituents are C 1 ~C 4 -alkyl, C 1 ~C 4-alkylidene, C 3 ~C 6 -alkenylidene, C 2 ~C 4 -alkenyl, oxo(=O), hydroxy, or amino, especially C such as methyl 1 ~C 4 -alkyl, and C such as methylidene 1 ~C 4 The biocatalytic method according to any one of the preceding embodiments, which can be independently selected from -alkylidene.
[0129] 7.R 2 The method according to any one of the preceding embodiments, wherein the Cyc residue of forms an optionally substituted decalinyl residue, especially a bicyclic residue, which can be obtained by terpene cyclization.
[0130] 8.Cyc-A is of formula IIIa, IIIb or IIIc
Chemical formula
[0131] 9. The polypeptide having BVMO activity is (1) A group of polypeptides containing a flavin-containing monooxygenase (FMO) protein family domain having the Pfam ID number PF00743 within its amino acid sequence; or a domain retaining at least 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with PF00743; In particular, the polypeptide of the present invention having BVMO activity is such that it is less than 1×10 -5 or less than 1×10 -10 or less than 1×10 -15 or less, or less than 1×10 -18 or less, especially in the range of 1×10 -10 ~1×10 -18 and more particularly in the range of 1×10 -14 ~1×10 -17When the e-value within the range matches the aforementioned domain PF00743, it is identified as a member of the FMO protein family containing the aforementioned domain. As the query sequence, the sequence of a polypeptide having BVMO activity is applied. For example, the following websites can be used for such e-value searches and calculations: http: / / pfam.xfam.org / , http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or http: / / www.ebi.ac.uk / Tools / pfa / pfamscan / and / or (2) Selected from the group of polypeptides containing at least 1, 2, 3, 4, 5, 6, 7 or all of the following sequence motifs / domains: GAGxSGL shown in SEQ ID NO: 197; EKNxxxxGTWxENRYPGCACDVPxHxYXXSFE shown in SEQ ID NO: 198; Or any partial motif thereof containing up to 15, up to 10 or up to 5 consecutive amino acid residues corresponding to, for example, residues 1 to 10, 11 to 20 or 21 to 32 of SEQ ID NO: 198; LxNAxGILNxWxxPxIPG shown in SEQ ID NO: 199 Or any partial motif thereof containing up to 15, up to 10 or up to 5 consecutive amino acid residues corresponding to, for example, residues 1 to 10 or 11 to 18 of SEQ ID NO: 199; LxxKxVxxIGxGSSGIQIxPxI shown in SEQ ID NO: 200; Or any partial motif thereof containing up to 15, up to 10 or up to 5 consecutive amino acid residues corresponding to, for example, residues 1 to 10 or 11 to 18 of SEQ ID NO: 200; GCRRxTPGxxYLExL shown in SEQ ID NO: 201 Or any partial motif thereof comprising up to 15, up to 10, or up to 5 consecutive amino acid residues corresponding to, for example, the residues at positions 1 to 10 or positions 11 to 15 of SEQ ID NO: 201; CATGFDxxxxPRFxxxG shown in SEQ ID NO: 202 Or any partial motif thereof comprising up to 15, up to 10, or up to 5 consecutive amino acid residues corresponding to, for example, the residues at positions 1 to 10 or positions 11 to 17 of SEQ ID NO: 202; PNxFxxxGPNxPxxNGxV shown in SEQ ID NO: 203 Or any partial motif thereof comprising up to 15, up to 10, or up to 5 consecutive amino acid residues corresponding to, for example, the residues at positions 1 to 10 or positions 11 to 18 of SEQ ID NO: 203; AxWPGSxLHYxEAxxxPRxED shown in SEQ ID NO: 204 Or any partial motif thereof comprising up to 15, up to 10, or up to 5 consecutive amino acid residues corresponding to, for example, the residues at positions 1 to 10 or positions 11 to 21 of SEQ ID NO: 204, wherein the motif residue x represents, independently of one another, any natural amino acid residue, and optionally, in each of the above motifs 1 to 5, for example, 1, 2, 3, 4, or 5 of the conserved amino acid residues (i.e., different from the x residue) may be modified, for example, by amino acid substitution, particularly by conservative substitution, provided that the BVMO enzyme activity is retained to at least an analytically detectable level. and / or (3) A group of polypeptides consisting of (a) A polypeptide comprising the amino acid sequence of SCH23 - BVMO1 shown in SEQ ID NO: 2; (b) A polypeptide comprising the amino acid sequence of SCH24 - BVMO1 shown in SEQ ID NO: 6; (c) A polypeptide comprising the amino acid sequence of SCH25 - BVMO1 shown in SEQ ID NO: 10; (d) A polypeptide comprising the amino acid sequence of SCH46-BVMO1 shown in SEQ ID NO: 13; (e) A polypeptide comprising the amino acid sequence of AspWeBVMO shown in SEQ ID NO: 16 (manno-ol oxidase and its isomers which are preferred substrates); (f) A polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to any one of the amino acid sequences of (a) to (e) above The method according to any one of the preceding embodiments, selected from the above.
[0132] For the five specific reference BVMO polypeptides described above, the protein family domain having the Pfam ID number PF00743 may be located at the amino acid residue positions shown in the following table (see also the alignment depicted in Figure 32 and the sequence portion within the frame thereof).
[0133] [Table 3]
[0134] The numbering of amino acid residues refers to the residue number in each sequence number of each protein sequence in the attached sequence listing.
[0135] Another specific embodiment refers to a polypeptide variant of the novel polypeptide of the present invention having BVMO activity identified above by any one of the specific amino acid sequences of SEQ ID NOs: 2, 6, 10, and 13, wherein the polypeptide variant is selected from amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 2, 6, 10, 13, and 16 and comprises at least one substitution modification relative to any one of the unmodified SEQ ID NOs: 2, 6, 10, 13, and 16.
[0136] 10. The method according to any one of the preceding embodiments, carried out in vitro or in vitro.
[0137] 11. A method according to embodiment 10 carried out in vivo, comprising, before step (1), recombinantly expressing one or more polypeptides having the enzymatic activity necessary to perform the enzymatic step by a BVMO catalyst, particularly in a non-human host cell.
[0138] 12. The method according to embodiment 11, wherein a non-human host cell is transformed with a nucleic acid encoding at least one polypeptide having BVMO activity.
[0139] 13. The method according to embodiment 11 or 12, wherein the aforementioned non-human host cell is a eukaryotic cell or a prokaryotic cell, particularly a plant cell, a bacterial cell or a fungal cell, particularly a yeast cell.
[0140] 14. The method according to any one of embodiments 11 to 13, wherein the aforementioned non-human host cell is a unicellular organism, a cultured cell derived from a multicellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism.
[0141] 15. The method according to any one of embodiments 10 to 13, wherein the aforementioned non-human host cell is a bacterium of the genus Escherichia, particularly Escherichia coli, and the aforementioned yeast is a yeast of the genus Saccharomyces or Pichia, particularly Saccharomyces cerevisiae, or a plant cell.
[0142] 16. The carbonyl compound of general formula I is a) labdane aldehyde, particularly copalal (or any stereoisomerically different form thereof, for example, including the cis form or the trans form, or a form containing a mixture of the cis form and the trans form), which is converted by the aforementioned BVMO into the respective norlabdane formate, particularly (5S,9S,10S)-15-norlabda-8(20),13-dien-14-yl-formate or any stereoisomerically different form thereof; b) gynol labdane ketone, especially manool oxy or any stereoisomerically different form thereof, which is converted by the aforementioned BVMO into γ-ambryl acetate or any stereoisomerically different form thereof; or c) norlabdane aldehyde, especially the C of copal 1 degradation analog or any stereoisomerically different form thereof, especially of the formula
Chem.
Chem.
[0143] Further specific examples of the conversion of each carbonyl compound to an ester by the BVMO catalyst are summarized in the following schematic overview:
Chem.
[0144] 17. The method according to a) of embodiment 16, wherein prior to step (1), it includes the biocatalytic oxidation of labdane alcohol to labdane aldehyde, especially the oxidation of copalol to copalal, This labdane alcohol is optionally formed by the biocatalytic conversion of at least one terpene diphosphate precursor selected from IPP, DMAPP, FPP and GGPP, in particular by one step or a combination of at least two steps known in the prior art.
[0145] The above-mentioned labdane alcohol can be obtained, for example, by using a biocatalyst a) It can be produced in one step in a cyclization / dephosphorylation reaction from geranylgeranyl diphosphate (GGPP), b) It can be produced by forming labdane diphosphate, such as copalyldiphosphate (CPP), in two steps by a cyclization reaction from GGPP, and then dephosphorylating this to obtain the labdane alcohol. c) It can be produced by directly converting IPP and DMAPP into labdane diphosphate, such as CPP, by the action of a GGPP synthase / CPP synthase with two functions, and then dephosphorylating this. The GGPP used in these steps can also be provided by different biocatalytic steps: d) A GGPP synthase that directly produces GGPP from IPP and DMAPP is available; or e) GGPP can be provided from IPP and DMAPP via FPP by the action of an FPP synthase, followed by the conversion of FPP to GGPP by the action of a GGPP synthase.
[0146] 18. Catalyzing the above-mentioned biocatalytic oxidation of labdane alcohol, especially copalol to copalyl, by an exogenous or endogenous polypeptide having alcohol dehydrogenase (ADH) (EC 1.1.1.-) activity; and / or The above-mentioned biocatalytic formation of labdane alcohol i) The biocatalytic dephosphorylation of labdane diphosphate to labdane alcohol, especially the dephosphorylation of copalyldiphosphate (CPP) to copalol, catalyzed by a polypeptide having terpene diphosphate (TPP) phosphatase activity, and / or ii) Biocatalytic cyclization of a terpenyl diphosphate precursor, such as geranylgeranyl diphosphate (GGPP), to CPP, catalyzed by a polypeptide having CPP synthase activity such as SmCPS2 (SEQ ID NO: 185); or, for example, biocatalytic cyclization of IPP and DMAPP to CPP, catalyzed by a polypeptide having two functions, prenyltransferase and copalyl diphosphate synthase activity, such as PvCPS, and / or iii) Biocatalytic formation of GGPP from FPP or biocatalytic formation of GGPP from IPP and DMAPP, each catalyzed by a polypeptide having GGPP synthase activity The method according to embodiment 17, comprising at least one step selected from the above.
[0147] 19. The aforementioned biocatalytic oxidation, particularly the oxidation of copalol to copalyl, is a) A polypeptide comprising the amino acid sequence of SCH23-ADH1_wt shown in SEQ ID NO: 134 b) A polypeptide comprising the amino acid sequence of SCH24-ADH1_wt shown in SEQ ID NO: 140 c) A polypeptide comprising the amino acid sequence of SCH94-3945_wt shown in SEQ ID NO: 161 d) A polypeptide comprising the amino acid sequence of SCH80-0540_wt shown in SEQ ID NO: 164 e) A polypeptide comprising the amino acid sequence of AzTolADH1_wt shown in SEQ ID NO: 167 f) A polypeptide comprising the amino acid sequence of CdGeoA_wt shown in SEQ ID NO: 179 g) A polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any one of the amino acid sequences from a) to f) and having ADH activity Catalyzed by a polypeptide having alcohol dehydrogenase (ADH) activity selected from the above, and / or The aforementioned biocatalytic dephosphorylation, particularly the dephosphorylation of copalyldiphosphate (CPP) to copalol, is a) a polypeptide comprising the amino acid sequence of AspWE TPP shown in SEQ ID NO: 170, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; b) a polypeptide comprising the amino acid sequence of TalCeTPP shown in SEQ ID NO: 176, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and; c) a polypeptide comprising the amino acid sequence of TalVeTPP shown in SEQ ID NO: 194, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; catalyzed by a polypeptide having terpenyldiphosphate (TPP) phosphatase activity selected from Further suitable phosphatases are also disclosed in the applicant's previous EP application number 18182783.3, which is incorporated by reference, and / or The aforementioned biocatalytic cyclization, particularly the cyclization of geranylgeranyl diphosphate (GGPP) to CPP, is a polypeptide having copalyldiphosphate synthase activity comprising the amino acid sequence of SmCPS2 shown in SEQ ID NO: 185, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto catalyzed by a polypeptide selected from The aforementioned biocatalytic cyclization, particularly the cyclization of IPP and DNMAPP to CPP, is a polypeptide having prenyltransferase activity and copalyldiphosphate synthase activity comprising the amino acid sequence of PvCPS shown in SEQ ID NO: 173, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto catalyzed by a polypeptide selected from and / or the aforementioned in vivo catalytic formation of GGPP is catalyzed by a polypeptide having GGPP synthase activity, a) a polypeptide comprising the amino acid sequence of carG shown in SEQ ID NO: 182, or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; b) a polypeptide comprising the amino acid sequence of CrtE shown in SEQ ID NO: 191, or a polypeptide comprising an amino acid sequence having at least 70, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; c) a polypeptide comprising the amino acid sequence of PvCPS shown in SEQ ID NO: 173, or a polypeptide comprising an amino acid sequence having at least 70, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto selected from the method according to embodiment 18.
[0148] 20. As step (3), further comprising treating the carbonyl ester formed in step (1) or isolated in step (2) to obtain its derivative using chemical synthesis or biocatalytic synthesis, or a combination of both, and optionally isolating the derivative of step (3), wherein the aforementioned derivative is particularly selected from hydrocarbons, alcohols, diols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides and / or esters, the method according to any of the preceding embodiments.
[0149] 21. The method according to embodiment 20, wherein step (3) comprises hydrolyzing the carbonyl ester compound with an esterase activity EC 3.1.1 (carboxylic acid ester hydrolase) to obtain the corresponding de-esterified product (which may be an alcohol or an isomerized product thereof), and optionally isolating the derivative of step (3).
[0150] 22. The de-esterification product of step (3) is subjected, in a further step (4), to an enzymatic redox reaction, where in particular the redox reaction comprises oxidizing the alcohol group formed in step (3) to the corresponding keto group by the enzymatic action of an exogenous or endogenous alcohol dehydrogenase (ADH) (EC 1.1.1.-), according to the method of embodiment 21.
[0151] 23. The esterase is a) a polypeptide comprising the amino acid sequence of SCH23-esterase shown in SEQ ID NO: 20; b) a polypeptide comprising the amino acid sequence of SCH24-esterase shown in SEQ ID NO: 24; c) a polypeptide comprising the amino acid sequence of SCH25-esterase shown in SEQ ID NO: 28; d) a polypeptide comprising the amino acid sequence of SCH46-esterase shown in SEQ ID NO: 31; or e) a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any one of the amino acid sequences from a) to d) and having esterase activity selected from the group consisting of, according to the method of embodiment 21.
[0152] 24. a) Norlabdane ester, in particular norlabdane formate, is de-esterified by the aforementioned esterase to obtain a norlabdane carbonyl compound, in particular a compound of the formula
Chemical formula
Chem.
Chem.
[0153] 25. The ADH is a) a polypeptide comprising the amino acid sequence of SCH23-ADH2_wt shown in SEQ ID NO: 137 b) a polypeptide comprising the amino acid sequence of SCH24-ADH2_wt shown in SEQ ID NO: 143 c) a polypeptide comprising the amino acid sequence of RrhSecADH_wt shown in SEQ ID NO: 146 d) a polypeptide comprising the amino acid sequence of SCH80-06135_wt shown in SEQ ID NO: 155 e) a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any one of the amino acid sequences from a) to d) and having ADH activity selected from the group consisting of, according to the method described in Embodiment 22.
[0154] 26. The resulting ginolrabdan alcohol, especially the formula
Chem.
[0155] ii) The present invention relates to the following specific embodiments of a biocatalytic method comprising the use of a polypeptide having enal dissociation activity: 27. General formula IV
Chemical formula
Chemical formula
[0156] 28. The aforementioned polypeptide having enal cleavage activity as described above is a) at least one DUF4334 protein family domain having Pfam ID number PF14232 (especially within the C-terminal region of their amino acid sequences); and / or b) at least one GXWXG protein family domain having Pfam ID number PF14231 (especially within the N-terminal region of their amino acid sequences); or c) a domain retaining at least 90% sequence identity with PF14232 or PF14231 The method according to embodiment 27, selected from the group of polypeptides comprising.
[0157] In particular, the polypeptide of the present invention having enal cleavage activity is such that it is 1×10 -5 less than, or 1×10 -10 less than, or 1×10 -15 less than, or 1×10 -20 less than, or 1×10 -25 less than, or 1×10-30 less than, or 1×10 -35 or less e-value, particularly 1×10 -20 ~1×10 -32 , more particularly 1×10 -25 ~1×10 -31 When matching the aforementioned domain with an e-value in the range of, it is identified as a member of the DUF4334 protein family containing the aforementioned domain PF14232. For example, the following websites can be used for searching and calculating such e-values: http: / / pfam.xfam.org / , http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or http: / / www.ebi.ac.uk / Tools / pfa / pfamscan / Particularly, the polypeptide of the present invention having enal cleavage activity, when it is 1×10 -5 less than, 1×10 -10 less than, or 1×10 -15 less than, or 1×10 -20 less than, or 1×10 -25 less than, or 1×10 -30 less than, or 1×10 -35 or less e-value, particularly 1×10 -20 ~1×10 -30 When matching with an e-value in the range of, it is identified as a member of the GXWXG protein family containing the aforementioned domain PF14232. As the query sequence, the sequence of the polypeptide having enal cleavage activity is applied. For example, the following websites can be used for searching and calculating such e-values: http: / / pfam.xfam.org / , http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or http: / / www.ebi.ac.uk / Tools / pfa / pfamscan / and / or the aforementioned polypeptide having the aforementioned enal cleavage activity G-[Y or "-"]-x-W-x-G-x-x-[F, L or I]-x-[T, S or R]-G-[H or D] shown in SEQ ID NO: 205, or any partial motif thereof containing up to 10 or up to 5 consecutive amino acid residues corresponding to the residues at positions 1 to 8 or 9 to 13 of SEQ ID NO: 205, for example; W-[Y, A or V]-G-K-x-[F or Y]-x-[S or D] shown in SEQ ID NO: 206, or any partial motif thereof containing up to 4 consecutive amino acid residues corresponding to the residues at positions 1 to 4 or 5 to 8 of SEQ ID NO: 206, for example; [G or S]-x-[A or G]-x-[L or V]-x-x-x-x-[F, Y or L]-R-G-x-V shown in SEQ ID NO: 207, or any partial motif thereof containing up to 10 or up to 5 consecutive amino acid residues corresponding to the residues at positions 1 to 8 or 9 to 14 of SEQ ID NO: 207, for example; [M or L]-[V or I]-Y-D-x-x-P-[I or V]-x-D-[H or S]-[F or L] shown in SEQ ID NO: 208, or any partial motif thereof containing up to 10 or up to 5 consecutive amino acid residues corresponding to the residues at positions 1 to 6 or 7 to 12 of SEQ ID NO: 208, for example selected from the group of polypeptides comprising at least one sequence motif / domain selected from: wherein the motif residue x represents, independently of each other, any natural amino acid residue, and optionally, in each of motifs 1 to 5 above, for example, 1, 2, 3, 4 or 5 amino acid residues different from the x residue may be modified, for example by amino acid substitution, particularly by conservative substitution, provided that the enzyme is retained to at least an analytically detectable extent. and / or the aforementioned polypeptides having the aforementioned enal cleavage activity consist of the group of polypeptides comprising the respective amino acid sequences: a) SCH94-3944 shown in SEQ ID NO: 34 b) SCH80-05241 shown in SEQ ID NO: 38 c) Pdigit7033 shown in SEQ ID NO: 42 d) PitalDUF4334-1 shown in SEQ ID NO: 46 e) AspWeDUF4334 shown in SEQ ID NO: 49 f) RhoagDUF4334-2 shown in SEQ ID NO: 53 g) RhoagDUF4334-3 shown in SEQ ID NO: 56 h) RhoagDUF4334-4 shown in SEQ ID NO: 59 i) CnecaDUF4334 shown in SEQ ID NO: 62 j) Rins-DUF4334 shown in SEQ ID NO: 69 k) CgatDUF4334 shown in SEQ ID NO: 72 l) GclavDUF4334 shown in SEQ ID NO: 75 m) TcurvaDUF4334 shown in SEQ ID NO: 81 n) PprotDUF4334 shown in SEQ ID NO: 87, and o) a polypeptide comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with any one of the amino acid sequences from a) to n) and retaining the aforementioned enzyme activity for degrading the terpene precursor of formula (1) The method according to embodiment 27, selected from
[0158] Another specific embodiment refers to a polypeptide variant of the novel polypeptide of the present invention having enal dissociation activity identified above by any of the specific amino acid sequences of SEQ ID NOs: 34, 38, 42, 46, 49, 53, 56, 59, 62, 69, 72, 75, 81 and 87, wherein the polypeptide variant has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity with any of SEQ ID NOs: 34, 38, 42, 46, 49, 53, 56, 59, 62, 69, 72, 75, 81 and 87 and is selected from amino acid sequences containing at least one substitution modification with respect to any of SEQ ID NOs: 34, 38, 42, 46, 49, 53, 56, 59, 62, 69, 72, 75, 81 and 87.
[0159] In the above 14 specific reference enal cleavage polypeptides, the protein family domains having Pfam ID numbers PF14232 and PF14231 may be located at the amino acid residue positions shown in the following table (see also the alignment depicted in Figure 31 and the sequence portion within the frame thereof).
[0160]
Table 4
[0161] The numbering of amino acid residues refers to the residue numbers in the respective sequence numbers of each protein sequence in the attached sequence listing.
[0162] 29. The aforementioned enal cleavage polypeptide consists of the following polypeptides and variants containing the respective amino acid sequences in the following groups: a) SCH94-3944-T51A_variant shown in SEQ ID NO: 91 b) SCH94-3944-H53A_variant shown in SEQ ID NO: 93 c) SCH94-3944-L59A_variant shown in SEQ ID NO: 95 d) SCH94-3944-W64A_variant shown in SEQ ID NO: 97 e) The SCH94-3944-S71A_variant shown in SEQ ID NO: 101 f) The SCH94-3944-R106A_variant shown in SEQ ID NO: 103 g) The SCH94-3944-Y115A_variant shown in SEQ ID NO: 105 h) The SCH94-3944-D116A_variant shown in SEQ ID NO: 107 i) The SCH94-3944-M136A_variant shown in SEQ ID NO: 111 j) The SCH94-3944-K139A_variant shown in SEQ ID NO: 113 k) The SCH94-3944-R156A_variant shown in SEQ ID NO: 119, and l) A polypeptide comprising an amino acid sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with any one of the amino acid sequences from a) to l) and retaining the aforementioned enzyme activity for decomposing the terpene precursor of formula (1) and retaining the positions of the aforementioned mutated amino acid sequences The method according to embodiment 28, selected from
[0163] 30. Applying a compound of formula V, for example a terpene-type compound, wherein R 1 represents H or methyl, R 2 represents H or a) An acyclic, linear or branched, saturated or unsaturated hydrocarbyl residue having 1 to 20 carbon atoms, particularly 1 to 10 carbon atoms, 1 to 15 carbon atoms or 1 to 20 carbon atoms; or b) The group Cyc-A-, wherein A represents a straight-chain or branched C 1 ~C 4 -alkylene bridge, particularly methylene, and Cyc is a monocyclic or polycyclic, particularly bicyclic, saturated or unsaturated hydrocarbyl residue, particularly a bicyclic cyclic hydrocarbyl residue containing 5 to 7, particularly 6, ring atoms per ring, optionally substituted with 1 to 10, 1 to 5 substituents, and the substituents are C 1 ~C 4 -alkyl, C1 ~C 4 -alkylidene, C 2 ~C 4 -alkenyl, oxo, hydroxy, or amino, especially C such as methyl 1 ~C 4 -alkyl, and C such as methylidene 1 ~C 4 represents a hydrocarbyl residue independently selected from -alkylidene, each R 3 represents H R 4 represents H or methyl, R 5 represents H or methyl, The method according to any one of embodiments 27 to 29.
[0164] 31. The compound of general formula V has a labdane-type structure and / or Cyc-A is of formula IIIa, IIIb or IIIc
Chemical formula
[0165] 32. The precursor of formula (V) is selected from labdane-type compounds such as farnesal, geranylgeraniol, citral, dodecanal, 8-hydroxy-labda-13-en-15-al, and copalal, each being in the form of a mixture of its stereoisomers or in its stereoisomerically pure form, the method according to any one of embodiments 27 to 31.
[0166] 33. The decomposition product of formula (IV) is selected from geranyl acetone, farnesyl acetone, methyl heptenone, decanal; or manool oxy, or 8-hydroxy-14,15-dinorlabdan-13-one, each being in the form of a mixture of its stereoisomers or in its stereoisomerically pure form, the method according to any one of embodiments 27 to 32.
[0167] Further specific inventive examples of the conversion of carbonyl compounds by enal cleavage enzyme catalysis to their respective cleavage products are summarized in the following schematic overview: [Chemical formula] In the formula, the parameter "n" is an integer of 1 to 20, 1 to 15, 1 to 10, or 1, 2, 3, 4, or 5.
[0168] 34. The method according to any one of embodiments 27 to 33, which is carried out in vitro or in vitro.
[0169] 35. The method according to embodiment 34, which is carried out in vivo and includes recombinantly expressing one or more polypeptides having the enzymatic activity necessary to perform the chain cleavage step, particularly in a non-human host cell, prior to step (1).
[0170] 36. The method according to embodiment 35, wherein the non-human host cell is transformed with a nucleic acid encoding at least one polypeptide having enal cleavage activity.
[0171] 37. The method according to embodiment 35 or 36, wherein the aforementioned non-human host cell is a eukaryotic cell or a prokaryotic cell, particularly a plant cell, a bacterial cell, or a fungal cell, particularly a yeast cell.
[0172] 38. The method according to any one of embodiments 35 to 37, wherein the non-human host cell is a unicellular organism, a cultured cell derived from a multicellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism.
[0173] 39. The method according to any one of embodiments 35 to 38, wherein the aforementioned non-human host cell is a bacterium of the genus Escherichia, preferably Escherichia coli, and the aforementioned yeast is a yeast of the genus Saccharomyces or Pichia, preferably Saccharomyces cerevisiae, or a plant cell.
[0174] 40. As step (3), treating the compound of formula IV formed in step (1) or isolated in step (2) to obtain its derivatives using chemical synthesis and / or biocatalytic synthesis, or a combination of both, and optionally isolating the derivative of step (3), wherein the aforementioned derivatives can be selected in particular from hydrocarbons, alcohols, diols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides and / or esters, the method according to any one of embodiments 27 to 39.
[0175] 41. The method according to embodiment 40, wherein step (3) comprises treating the compound of formula IV formed in step (1) or isolated in step (2) with a polypeptide having Baeyer-Villiger monooxygenase (BVMO) activity to form the respective carbonyl esters.
[0176] 42. Further comprising hydrolyzing the carbonyl ester compound with an esterase to obtain the corresponding de-esterified product which may be an alcohol or its isomerized product, and optionally isolating the derivative of step (3), the method according to embodiment 41.
[0177] 43. The method according to embodiment 41, wherein the polypeptide having BVMO activity is as defined above in embodiment 9.
[0178] 44. The esterase is a) a polypeptide comprising the amino acid sequence of SCH23-esterase shown in SEQ ID NO: 20; b) a polypeptide comprising the amino acid sequence of SCH24-esterase shown in SEQ ID NO: 24; c) a polypeptide comprising the amino acid sequence of SCH25-esterase shown in SEQ ID NO: 28; d) a polypeptide comprising the amino acid sequence of SCH46-esterase shown in SEQ ID NO: 31; or e) a polypeptide having an amino acid sequence with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to any one of the amino acid sequences of a) to d) and having esterase activity The method according to embodiment 42, selected from the group consisting of
[0179] 45. The method according to any one of embodiments 41 to 44, wherein the carbonyl compound is dinorlabdane ketone, particularly manool oxy, and is converted by the aforementioned BVMO into the respective tetranorlabdanil acetate, particularly γ-ambryl acetate.
[0180] 46. The method according to embodiment 45, wherein tetranorlabdanil acetate, particularly γ-ambryl acetate, is de-esterified by the aforementioned esterase to form the respective tetranorlabdane, particularly γ-ambrol.
[0181] 47. The method according to any one of the preceding embodiments 27 to 46, wherein prior to step (1), it includes the biocatalytic oxidation of labdane alcohol to labdane aldehyde, particularly the oxidation of copalol to copalal, This labdane alcohol is optionally formed by the biocatalytic conversion of at least one terpene diphosphate precursor selected from IPP, DMAPP, FPP, and GGPP, particularly formed by one step or a combination of at least two steps known in the prior art.
[0182] The aforementioned labdane alcohol can be produced, for example, using a biocatalyst, a) It can be produced in one step in a cyclization reaction / dephosphorylation reaction from geranylgeranyl diphosphate (GGPP), b) It can be produced by forming labdane diphosphate, such as copalyl diphosphate (CPP), in two steps by a cyclization reaction from GGPP, and then dephosphorylating this to obtain labdane alcohol. c) From IPP and DMAPP, it can be directly converted to geranylgeranyl diphosphate, for example, CPP, by the action of geranylgeranyl diphosphate synthase / copalyl diphosphate synthase having two functions, and then generated by dephosphorylating this. The GGPP used in these steps can also be provided by different biochemical steps: d) Geranylgeranyl diphosphate synthase that directly generates GGPP from IPP and DMAPP is available; or e) GGPP can be provided from IPP and DMAPP via FPP by the action of FPP synthase, and subsequently FPP is converted to GGPP by the action of geranylgeranyl diphosphate synthase.
[0183] 48. Catalyze the aforementioned biochemical oxidation of labdanol, especially copalol to copalyl alcohol, by an exogenous or endogenous polypeptide having alcohol dehydrogenase (ADH) (EC 1.1.1.-) activity; and / or The aforementioned biochemical formation of labdanol is i) The biochemical dephosphorylation of geranylgeranyl diphosphate to labdanol, especially the dephosphorylation of copalyl diphosphate (CPP) to copalol, catalyzed by a polypeptide having terpenyl diphosphate (TPP) phosphatase activity, and / or ii) The biochemical cyclization of a terpenyl diphosphate precursor, such as geranylgeranyl diphosphate (GGPP), to CPP, catalyzed by a polypeptide having CPP synthase activity such as SmCPS2 (SEQ ID NO: 185); or, for example, the biochemical cyclization of IPP and DMAPP to CPP, catalyzed by a polypeptide having two functions with prenyltransferase and copalyl diphosphate synthase activity such as PvCPS, and / or iii) The biochemical formation of GGPP from FPP or the biochemical formation from IPP and DMAPP, each catalyzed by a polypeptide having geranylgeranyl diphosphate synthase activity The method according to embodiment 47, comprising at least one step selected from the above.
[0184] 49. The above-described biocatalytic oxidation, particularly the oxidation of copalol to copalyl alcohol, is a) a polypeptide comprising the amino acid sequence of SCH23-ADH1_wt shown in SEQ ID NO: 134 b) a polypeptide comprising the amino acid sequence of SCH24-ADH1_wt shown in SEQ ID NO: 140 c) a polypeptide comprising the amino acid sequence of SCH94-3945_wt shown in SEQ ID NO: 161 d) a polypeptide comprising the amino acid sequence of SCH80-0540_wt shown in SEQ ID NO: 164 e) a polypeptide comprising the amino acid sequence of AzTolADH1_wt shown in SEQ ID NO: 167 f) a polypeptide comprising the amino acid sequence of CdGeoA_wt shown in SEQ ID NO: 179 g) a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any one of the amino acid sequences from a) to f) and having ADH activity and is catalyzed by a polypeptide having alcohol dehydrogenase (ADH) activity selected from and / or the above-described biocatalytic dephosphorylation, particularly the dephosphorylation of copalyl diphosphate (CPP) to copalol, is d) a polypeptide comprising the amino acid sequence of AspWE TPP shown in SEQ ID NO: 170, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity therewith; e) a polypeptide comprising the amino acid sequence of TalCeTPP shown in SEQ ID NO: 176, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity therewith; and; f) a polypeptide comprising the amino acid sequence of TalVeTPP shown in SEQ ID NO: 194, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity therewith; catalyzed by a polypeptide having terpenyl diphosphate (TPP) phosphatase activity selected from Further suitable phosphatases are also disclosed in the applicant's previous EP application number 18182783.3, which is incorporated by reference, and / or The above-mentioned biocatalytic cyclization, especially the cyclization of geranylgeranyl diphosphate (GGPP) to CPP, is a polypeptide having copalyl diphosphate synthase activity, comprising the amino acid sequence of SmCPS2 shown in SEQ ID NO: 185, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto catalyzed by a polypeptide selected from The above-mentioned biocatalytic cyclization, especially the cyclization of IPP and DNMAPP to CPP, is a polypeptide having prenyltransferase activity and copalyl diphosphate synthase activity, comprising the amino acid sequence of PvCPS shown in SEQ ID NO: 173, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto catalyzed by a polypeptide selected from and / or The above-mentioned biocatalytic formation of GGPP is catalyzed by a polypeptide having GGPP synthase activity, d) a polypeptide comprising the amino acid sequence of carG shown in SEQ ID NO: 182, or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; e) a polypeptide comprising the amino acid sequence of CrtE shown in SEQ ID NO: 191, or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; f) a polypeptide comprising the amino acid sequence of PvCPS shown in SEQ ID NO: 173, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto The method according to embodiment 48, selected from
[0185] iii) The present invention relates to the following specific embodiments of enal lyase and the corresponding coding sequences 50. An isolated polypeptide having the enal cleavage activity as defined in any of embodiments 28 and 29, particularly the activity of an α,β-unsaturated aldehyde C=C bond cleavage enzyme The polypeptides of the present invention include all active forms including the active partial sequences of enzymes having enal cleavage activity, such as catalytic domains or active sites
[0186] 51. A nucleic acid sequence encoding the polypeptide of embodiment 50, particularly a nucleic acid sequence selected from SEQ ID NOs: 33, 35, 36, 37, 39, 40, 41, 43, 44, 45, 47, 48, 50, 51, 52, 54, 55, 57, 58, 60, 61, 63, 64, 68, 70, 71, 73, 74, 76, 80, 82, 86, 88, 92, 94, 96, 98, 102, 104, 106, 108, 112, and 120, and a nucleic acid sequence having a degree of sequence identity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% with any one of the foregoing sequences from SEQ ID NOs: 33, 35, 36, 37, 39, 40, 41, 43, 44, 45, 47, 48, 50, 51, 52, 54, 55, 57, 58, 60, 61, 63, 64, 68, 70, 71, 73, 74, 76, 80, 82, 86, 88, 92, 94, 96, 98, 102, 104, 106, 108, 112, and 120, an isolated nucleic acid molecule
[0187] 52. An expression cassette comprising the nucleotide sequence of at least one nucleic acid molecule of embodiment 50
[0188] An expression vector comprising the nucleotide sequence of at least one nucleic acid molecule according to embodiment 51, or at least one expression cassette according to embodiment 52.
[0189] 54. The expression vector according to embodiment 53, wherein the vector is a prokaryotic vector, a viral vector or a eukaryotic vector.
[0190] 55. The expression vector according to embodiment 53 or 54, which is a plasmid or a combination of two or more plasmids.
[0191] 56. A recombinant non-human host cell comprising at least one nucleic acid molecule defined in embodiment 51, or at least one expression cassette according to embodiment 52, or an expression vector according to any one of embodiments 53 to 55.
[0192] 57. The host cell according to embodiment 56, wherein at least one nucleic acid molecule or at least one expression cassette is stably integrated into the genome of the cell.
[0193] 58. The host cell according to embodiment 56 or 57, which is a prokaryotic cell or a eukaryotic cell, particularly a plant cell, a bacterium or a fungal cell, particularly yeast.
[0194] 59. The host cell according to any one of embodiments 56 to 58, which is a unicellular organism, a cultured cell derived from a multicellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism.
[0195] 60. The host cell according to embodiment 59, which is a bacterium of the genus Escherichia, preferably Escherichia coli, or a yeast cell of the genus Saccharomyces, preferably Saccharomyces cerevisiae, or a yeast cell of the genus Pichia, preferably Pichia pastoris.
[0196] 61. A method for producing at least one polypeptide having enal cleavage activity according to embodiment 51, the method comprising (i) In any one of the non-human host cells described in Embodiments 57 to 60, expressing the at least one polypeptide described above; (ii) Optionally, isolating the at least one polypeptide described above from the non-human host cell used in step (i); A method comprising.
[0197] 62. Before step (i): By introducing at least one nucleic acid molecule defined in Embodiment 51, or at least one expression cassette described in Embodiment 52, or at least one expression vector described in any one of Embodiments 53 to 55 into a non-human cell, preparing the non-human host cell used in step (i), and thus further obtaining a host cell capable of expressing or overexpressing at least one polypeptide having enal cleavage activity according to Embodiment 50. The method according to Embodiment 61 further comprising.
[0198] 63. A method for preparing a mutant polypeptide having enal cleavage activity, the method comprising: (i) Providing a nucleic acid molecule according to Embodiment 51; (ii) Modifying the nucleotide sequence of the nucleic acid molecule described above, particularly the nucleotide sequence encoding the polypeptide described in Embodiment 50, so that at least one mutant nucleic acid molecule is obtained; (iii) Recombinantly expressing the mutant nucleic acid molecule described above in a non-human host cell; (iv) Screening the expression product obtained in step (iii) for at least one mutant polypeptide having enal cleavage activity; (v) Optionally, repeating steps (ii) to (iv) using the mutant nucleic acid molecule until the expression product contains a mutant polypeptide having the desired enal cleavage activity; (vi) Optionally, isolating the mutant polypeptide having the desired enal cleavage activity. A method comprising.
[0199] iv) The present invention relates to the following specific embodiments regarding BVMO enzymes and corresponding coding sequences. 64. An isolated polypeptide having BVMO activity as defined in embodiment 9. The polypeptide of the present invention includes all active forms including the active partial sequence of the enzyme having BVMO activity, such as the catalytic domain or the active site.
[0200] 65. A nucleic acid sequence encoding the polypeptide according to embodiment 64, in particular a nucleic acid sequence selected from SEQ ID NOs: 1, 3, 4, 5, 7, 8, 9, 11, 12, 14, 15, 17 and 18, and any one of the aforementioned sequences of SEQ ID NOs: 1, 3, 4, 5, 7, 8, 9, 11, 12, 14, 15, 17 and 18 having a degree of sequence identity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99%. An isolated nucleic acid molecule.
[0201] 66. An expression cassette comprising the nucleotide sequence of at least one nucleic acid molecule according to embodiment 65.
[0202] 67. An expression vector comprising the nucleotide sequence of at least one nucleic acid molecule according to embodiment 65, or at least one expression cassette according to embodiment 66.
[0203] 68. The expression vector according to embodiment 67, wherein the vector is a prokaryotic vector, a viral vector or a eukaryotic vector.
[0204] 69. The expression vector according to embodiment 67 or 68, which is a plasmid or a combination of two or more plasmids.
[0205] 70. A recombinant non-human host cell comprising at least one nucleic acid molecule as defined in embodiment 65, or at least one expression cassette according to embodiment 66, or at least one expression vector according to any one of embodiments 67 to 69.
[0206] 71. The host cell according to embodiment 70, wherein at least one nucleic acid molecule or at least one expression cassette is stably integrated into the genome of the cell.
[0207] 72. The host cell according to embodiment 70 or 71, which is a prokaryotic cell or a eukaryotic cell, particularly a plant cell, a bacterium or a fungal cell, particularly yeast.
[0208] 73. The host cell according to any one of embodiments 70 to 72, which is a unicellular organism, a cultured cell derived from a multicellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism.
[0209] 74. The host cell according to embodiment 72, which is a bacterium of the genus Escherichia, preferably Escherichia coli, or a yeast cell of the genus Saccharomyces, preferably Saccharomyces cerevisiae, or a yeast cell of the genus Pichia, preferably Pichia pastoris.
[0210] 75. A method for producing at least one polypeptide having BVMO activity according to embodiment 64, comprising: (i) expressing the aforementioned at least one polypeptide in a non-human host cell according to any one of embodiments 70 to 74; and (ii) optionally, isolating the aforementioned at least one polypeptide from the non-human host cell used in step (i). The method comprising.
[0211] 76. Before step (i): By introducing into a non-human cell at least one nucleic acid molecule defined in embodiment 65, or at least one expression cassette described in embodiment 66, or at least one expression vector according to any one of embodiments 67 to 69, preparing the non-human host cell used in step (i), and thus obtaining a host cell capable of expressing or overexpressing at least one polypeptide having BVMO activity according to embodiment 64. The method according to embodiment 75, further comprising.
[0212] 77. A method for preparing a mutant polypeptide having BVMO activity, the method comprising: (i) providing a nucleic acid molecule according to embodiment 65; (ii) modifying the nucleotide sequence of the nucleic acid molecule described above, particularly the nucleotide sequence encoding the polypeptide described in embodiment 64, so as to obtain at least one mutant nucleic acid molecule; (iii) recombinantly expressing the mutant nucleic acid molecule described above in a non-human host cell; (iv) screening the expression product obtained in step (iii) for at least one mutant polypeptide having BVMO activity; (v) optionally repeating steps (ii) to (iv) using the mutant nucleic acid molecule until the expression product contains a mutant polypeptide having the desired BVMO activity; (vi) optionally isolating the mutant polypeptide having the desired BVMO activity and including the method.
[0213] v) The present invention relates to the following specific embodiments regarding a biocatalytic multi-step in vivo method for converting a rubdan compound by applying a polypeptide having enal cleavage activity and / or BVMO activity 78. An in vivo method for preparing rubdan-type terpenoids, the method comprising the following series of reaction steps: (1) optionally, by the enzymatic action of an exogenous or endogenous ADH polypeptide, particularly the ADH defined in any of embodiments 19 or 49, converting rubdan alcohol, particularly copalol, to the respective rubdan aldehyde, particularly copalal; (2) converting the aforementioned rubdan aldehyde of step (1), particularly copalal, to the respective dinor rubdan carbonyl compound, particularly manool oxy, by the action of a polypeptide having enal cleavage activity, particularly the polypeptide defined in any of embodiments 28 and 29; (3) Optionally, converting the aforementioned dinorlabdane carbonyl compound of step (2), especially manool oxy, into the respective tetranorlabdanil acetate, especially γ-ambril acetate, by the action of a polypeptide having BVMO activity, especially BVMO as defined in embodiment 9; (4) Optionally, converting the aforementioned tetranorlabdanil acetate of step (3), especially γ-ambril acetate, into the respective tetranorlabdanol, especially γ-ambrinol, by the action of a polypeptide having esterase activity, especially esterase as defined in either embodiment 23 or 44; and optionally, (5) Isolating the product of step (2), (3) or (4) A method comprising providing a recombinant host that expresses a series of polypeptides having the enzyme activities necessary to catalyze.
[0214] 79. An in vivo method for preparing labdane-type cycloterpenes, The method comprises the following series of reaction steps: (1) Optionally, converting labdanol, especially copalol, into the respective labdanal, especially copalal, by the enzymatic action of an exogenous or endogenous ADH polypeptide, especially ADH as defined in either embodiment 19 or 49; (2) Converting the aforementioned labdanal of step (1), especially copalal, into the respective norlabdane ester compound, especially [(4-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decalin-1-yl]-2-methyl-but-1-enyl] formate (compounds 1a, 1b), by the action of a polypeptide having BVMO activity, especially BVMO as defined in any of embodiment 9; (3) Optionally, converting the aforementioned labdane ester compound of step (2), particularly [(4-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decalin-1-yl]-2-methyl-but-1-enyl)] formate (compounds 1a, 1b), by the action of a polypeptide having esterase activity, particularly the esterase defined in any of embodiments 23 or 44, into the respective norlabdane aldehyde, particularly 4-[(1S,8aS)-5,5,8a-trimethyl-2-methylene-decalin-1-yl]-2-methyl-butan-al (compounds 3a, 3b); (4) Converting the aforementioned norlabdane aldehyde of step (3), particularly 4-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decalin-1-yl]-2-methyl-butan-al (compounds 3a, 3b), by the action of a polypeptide having BVMO activity, particularly the BVMO defined in any of embodiments 9, into the respective dinorlabdane ester, particularly [(3-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decalin-1-yl]-1-methyl-propyl)] formate (compounds 4a, 4b); (5) Converting the aforementioned dinorlabdane ester of step (4), particularly [(3-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decalin-1-yl]-1-methyl-propyl)] formate (compounds 4a, 4b), by the action of a polypeptide having esterase activity, particularly the esterase defined in any of embodiments 23 or 44, into the respective dinorlabdane alcohol, particularly 4-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decalin-1-yl]butan-2-ol (compounds 5a, 5b); (6) Optionally, the aforementioned dinorlabdan alcohol of step (5), particularly 4-[(1S,8aS)-5,5,8a-trimethyl-2-methylene-decalin-1-yl]butan-2-ol (compounds 5a, 5b), is converted by the action of an exogenous or endogenous polypeptide having ADH activity, particularly ADH as defined in any of embodiments 19 or 49, into the respective dinorlabdan carbonyl compound, particularly manool oxy; (7) Optionally, the aforementioned dinorlabdan carbonyl compound of step (6), particularly manool oxy, is converted by the action of a polypeptide having BVMO activity, particularly BVMO as defined in any of embodiments 9, into the respective tetranorlabdanil acetate, particularly γ-ambril acetate; (8) The aforementioned tetranorlabdanil acetate of step (7), particularly γ-ambril acetate, is converted by the action of a polypeptide having esterase activity, particularly esterase as defined in any of embodiments 23 or 44, into the respective tetranorlabdan alcohol, particularly γ-ambrinol; and optionally, (9) isolating the product of step (5), (6), (7) or (8); A method comprising providing a recombinant host that expresses a series of polypeptides having the enzyme activities necessary to catalyze.
[0215] 80. The ADH applied in steps (1) and (6) is the same or different, exogenous or endogenous, and / or The BVMO applied in steps (2), (4) and (7) is the same or different, and / or The esterases applied in steps (3), (5) and (8) are the same or different, The method according to embodiment 79.
[0216] 81. The recombinant host, prior to step (1), the following series of reaction steps: (i) Catalytically forming geranylgeranyl diphosphate (GGPP) by the action of a polypeptide having GGPP synthase activity, particularly the GGPP synthase as defined in any one of Embodiments 19 and 49; (ii) Catalytically cyclizing GGPP to the aforementioned labdane diphosphate, particularly copalyldiphosphate (CPP), by the action of a polypeptide having labdane diphosphate synthase activity, particularly a polypeptide comprising the CPP synthase activity as defined in any one of Embodiments 19 and 49; (iii) Catalytically dephosphorylating the aforementioned labdane diphosphate to the aforementioned labdane alcohol, particularly CPP to copalol, by the action of a polypeptide having labdane diphosphate phosphatase activity, particularly a polypeptide having the TPP phosphatase activity as defined in any one of Embodiments 19 and 49; The method according to any one of Embodiments 78 to 80, further comprising expressing and applying a series of polypeptides having the enzyme activities necessary to catalyze the above.
[0217] 82. The method according to any one of Embodiments 78 to 81, further comprising expressing and applying a recombinant host that additionally expresses at least one polypeptide that catalyzes an enzymatic step of the mevalonate pathway or the MEP pathway.
[0218] 83. The method according to any one of Embodiments 78 to 82, applying a recombinant host in which the coding sequences of the respective catalytically active polypeptides are carried on one or more expression vectors and / or stably integrated into the genome of the host.
[0219] 84. The method according to any one of Embodiments 1 to 49 and 78 to 83, which is performed in vivo, and the method includes, prior to step (1), introducing one or more nucleic acid molecules encoding one or more polypeptides having the enzyme activities necessary to perform each step or multiple steps of the respective biocatalytic conversion into a non-human host organism or cell, and optionally stably integrating them into their respective genomes.
[0220] 85. A non-human host organism or cell that endogenously produces FPP and / or GGPP; or a mixture of IPP and DMAPP; or a non-human host organism genetically modified to produce an increased amount of FPP and / or GGPP and / or a mixture of IPP and DMAPP, the method according to any one of embodiments 1 to 49 and 78 to 83, which is carried out by applying such a non-human host organism.
[0221] Among these host cells or organisms applicable in the present invention, there are those that do not naturally produce FPP or GGPP, or a mixture of IPP and DMAPP. Such organisms or cells that do not naturally produce acyclic terpene pyrophosphate precursors, such as FPP or GGPP, or a mixture of IPP and DMAPP, may be genetically modified to produce the aforementioned precursors. Such organisms or cells can be transformed in this way, for example, prior to modification with the nucleic acids described herein. Methods for transforming organisms to produce acyclic terpene pyrophosphate precursors, such as FPP or GGPP, or a mixture of IPP and DMAPP, are already known in the art. For example, introducing the enzyme activity of the mevalonate pathway is a suitable strategy for causing an organism to produce FPP or GGPP, or a mixture of IPP and DMAPP.
[0222] 86. A recombinant microorganism as defined in any of embodiments 78 to 85.
[0223] vi) The present invention relates to the following specific embodiments related to further converting the chemical intermediate compounds obtained by the biocatalytic methods described herein into further end products of particular interest. 87. A method for preparing an epoxy-tetranorlabdane compound, particularly ambrinol, the method comprising (1) A step of providing tetranorlabdane alcohol, especially γ-ambrinol, or tetranorlabdane acetate, especially γ-ambrylol acetate, or dinorlabdane carbonyl compound, especially manool oxy, comprising applying a biocatalytic method comprising one or more method steps defined in any of claims 1 to 49 or 78 to 83, and optionally isolating the aforementioned product; (2) A step of converting the aforementioned product of step (1) into epoxy-tetranorlabdane, especially ambrlox, by applying one or more chemical and / or biochemical conversion steps A method comprising.
[0224] 88. A method for preparing diepoxy-dinorlabdane, especially Z11, the method comprising (1) A method for forming the aforementioned dinorlabdane carbonyl compound, especially manool oxy, by applying a method comprising one or more method steps defined in any of claims 1 to 49 or 78 to 84, to provide a dinorlabdane carbonyl compound, especially manool oxy, and optionally isolating the aforementioned dinorlabdane carbonyl compound, especially manool oxy; (2) A step of converting the aforementioned dinorlabdane carbonyl compound, especially manool oxy, of step (1) into the aforementioned diepoxy-dinorlabdane, especially Z-11, by applying one or more chemical and / or biochemical conversion steps A method comprising.
[0225] b. Polypeptides applicable according to the present invention In this regard, the following definitions apply: The general terms "polypeptide" or "peptide" can be used interchangeably and refer to a linear chain or sequence of natural or synthetic, continuously peptide-bonded amino acid residues, containing from about 10 residues to more than 1,000 residues at most. Short-chain polypeptides of up to 30 residues are also referred to as "oligopeptides".
[0226] The term "protein" refers to a macromolecular structure composed of one or more polypeptides. The amino acid sequence of the (multiple) polypeptides represents the "primary structure" of the protein. The amino acid sequence also predetermines the "secondary structure" of the protein through the formation of special structural elements such as α-helical structures and β-sheet structures formed within the polypeptide chain. The arrangement of such multiple secondary structural elements defines the "tertiary structure" or spatial arrangement of the protein. When a protein contains two or more polypeptide chains, the aforementioned chains are spatially arranged to form the "quaternary structure" of the protein. The correct spatial arrangement of the protein, i.e., "folding", is a prerequisite for the protein to function. Through denaturation or unfolding, the function of the protein is destroyed. If such destruction is reversible, the function of the protein can be restored by refolding.
[0227] The typical protein function referred to in this specification is the "enzymatic function", that is, the protein acts as a biocatalyst on a substrate, for example, a compound, and catalyzes the conversion of the aforementioned substrate into a product. Enzymes may have a high or low degree of specificity for substrates and / or products.
[0228] Therefore, the "polypeptide" referred to as having a specific "activity" in this specification implicitly refers to a correctly folded protein exhibiting the specified activity, for example, a specific enzymatic activity.
[0229] Therefore, unless otherwise specified, the term "polypeptide" also encompasses the terms "protein" and "enzyme".
[0230] Similarly, the term "polypeptide fragment" encompasses the terms "protein fragment" and "enzyme fragment".
[0231] The term "isolated polypeptide" refers to an amino acid sequence removed from its natural environment by any method or combination of methods known in the art, including recombinant, biochemical, and synthetic methods.
[0232] "Target peptide" refers to an amino acid sequence that targets a protein or polypeptide to an intracellular organelle, namely a mitochondrion or a plastid, or to the extracellular space (secretory signal peptide). The nucleic acid sequence encoding the target peptide may be fused to the nucleic acid sequence encoding the amino terminus, e.g., the N-terminus, of the protein or polypeptide, or may be used to replace the native targeting polypeptide.
[0233] The present invention also relates to "functional equivalents" (also referred to as "analogs" or "functional variants") of the polypeptides specifically described herein.
[0234] For example, "functional equivalents" refer to polypeptides that exhibit an activity that is at least 1-10%, or at least 20%, or at least 50%, or at least 75%, or at least 90% higher or lower compared to the activity of the polypeptides specifically described herein in a test used to determine enzymatic terpene diphosphate synthase activity or terpene diphosphate phosphatase activity.
[0235] According to the present invention, "functional equivalents" include specific variants that have amino acids different from those specifically described at at least one sequence position of the amino acid sequences described herein, but nevertheless have one of the aforementioned biological activities, for example, enzymatic activity. Thus, "functional equivalents" are variants that can be obtained by addition, substitution, particularly conservative substitution, deletion, and / or inversion of one or more, for example, 1 to 20, particularly 1 to 15 or 5 to 10 amino acids, where the described changes may be made at any sequence position, provided that they result in variants having the profile of properties according to the present invention. In particular, when the activity patterns are qualitatively identical between the variant and the unchanged polypeptide, i.e., for example, interactions with the same agonist or antagonist or substrate, which are observed at different rates (i.e., represented by EC 50 or IC 50 values, or any other suitable parameter in the current art), functional equivalence is also provided. Examples of suitable (conservative) amino acid substitutions are shown in the following table.
[0236]
Table 5
[0237] "Functional equivalents" in the above sense are also "precursors" of the polypeptides described herein, as well as "functional derivatives" and "salts" of the polypeptides.
[0238] A "precursor" is, in that case, the natural or synthetic precursor of a polypeptide that has or does not have the desired biological activity.
[0239] The term "salt" means not only acid addition salts of the amino groups of the protein molecules according to the invention, but also salts of the carboxyl groups. Salts of carboxyl groups can be produced by known methods and include inorganic salts such as sodium, calcium, ammonium, iron and zinc salts, as well as salts with organic bases such as amines like triethanolamine, arginine, lysine, piperidine and the like. Acid addition salts, such as salts with inorganic acids like hydrochloric acid or sulfuric acid, and salts with organic acids like acetic acid and oxalic acid are also included within the scope of the present invention.
[0240] "Functional derivatives" of the polypeptides according to the invention can also be produced using known techniques on functional amino acid side groups, or at their N- or C-terminus. Such derivatives include, for example, aliphatic esters of carboxylic acid groups, amides of carboxylic acid groups obtainable by reaction with ammonia or primary or secondary amines; N-acyl derivatives of free amino groups produced by reaction with an acyl group; or O-acyl derivatives of free hydroxyl groups produced by reaction with an acyl group.
[0241] "Functional equivalents" naturally include not only naturally occurring variants, but also polypeptides obtainable from other organisms. For example, areas of homologous sequence regions can be established by sequence comparison and equivalent polypeptides can be determined based on the specific parameters of the present invention.
[0242] "Functional equivalents" also include "fragments" such as individual domains or sequence motifs of the polypeptides according to the invention, or forms with the N- and C-termini truncated, which may or may not exhibit the desired biological function. Preferably, such "fragments" retain at least qualitatively the desired biological function.
[0243] "Functional equivalent" further refers to a fusion protein in which one of the polypeptide sequences described herein or a functional equivalent derived therefrom is functionally linked at the N-terminus or C-terminus to at least one additional heterologous sequence that is functionally different (i.e., there is no substantial mutual functional impairment of the fusion protein moieties). Non-limiting examples of such heterologous sequences are, for example, signal peptides, histidine anchors or enzymes.
[0244] "Functional equivalents" similarly included according to the present invention are also homologues of the specifically disclosed polypeptides. These have at least 60%, preferably at least 75%, particularly at least 80 or 85%, for example 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% homology (or identity) with one of the specifically disclosed amino acid sequences and are calculated by the algorithm of Pearson and Lipman, Proc. Natl. Acad, Sci. (USA) 85(8), 1988, 2444-2448. Homology or identity expressed as a percentage of the homologous polypeptides according to the invention means, in particular, identity expressed as a percentage of amino acid residues based on the full length of one of the amino acid sequences specifically described herein.
[0245] Identity data expressed as a percentage can be determined by using the BLAST alignment, the algorithm blastp (protein-protein BLAST), or by applying the Clustal settings defined later herein.
[0246] Where there is a possibility of protein glycosylation, "functional equivalents" according to the present invention include not only modified forms that can be obtained by altering the glycosylation pattern, but also deglycosylated or glycosylated forms of the polypeptides described herein.
[0247] Functional equivalents or homologs of the polypeptides according to the invention can be generated by mutagenesis, for example by point mutations, by extension or shortening of the protein, or as will be described in more detail below.
[0248] Functional equivalents or homologs of the polypeptides according to the invention can be identified by screening a combinatorial database of mutants, for example mutants of truncated mutants. For example, a database rich in protein variant changes can be created by combinatorial mutagenesis at the nucleic acid level, for example by enzymatic ligation of a mixture of synthetic oligonucleotides. There are a great many ways that can be used to create a database of potential homology from degenerate oligonucleotide sequences. Chemical synthesis of degenerate gene sequences can be carried out on an automated DNA synthesizer and the synthetic genes can then be ligated into a suitable expression vector. Using a degenerate genome makes it possible to mix and supply all the sequences encoding a desired set of potential protein sequences. Methods for the synthesis of degenerate oligonucleotides are known to those skilled in the art.
[0249] In the prior art, several techniques are known for screening combinatorial databases of gene products generated by point mutations or deletions, and for screening cDNA libraries for gene products having selected properties. These techniques can be adapted to the rapid screening of gene banks produced by combinatorial mutagenesis of homologs according to the present invention. The techniques most frequently used for screening large-scale gene banks are based on high-throughput analysis and include cloning of the gene bank in a replicable expression vector, transformation of appropriate cells with the resulting vector database, and expression of combinatorial genes under conditions that facilitate isolation of the vector encoding the gene of the detected product by detection of the desired activity. REM (recursive ensemble mutagenesis), a technique for increasing the frequency of functional variants in a database, can be used in combination with screening tests to identify homology.
[0250] Embodiments provided herein provide orthologs and paralogs of the polypeptides disclosed herein and methods for identifying and isolating such orthologs and paralogs. Definitions of the terms "ortholog" and "paralog" are provided below and apply to amino acid and nucleic acid sequences.
[0251] The polypeptides of the present invention include all active forms that include the active partial sequences of the enzymes of the present invention, such as catalytic domains or active sites. In one aspect, the present invention provides the catalytic domains or active sites described below. In one aspect, the present invention uses databases such as Pfam (http: / / pfam.wustl.edu / hmmsearch.shtml) (which is a large collection of multiple sequence alignments and hidden Markov models covering many common protein families, The Pfam protein families database, A. Bateman, E. Birney, L. Cerruti, R. Durbin, L. Etwiller, S. R. Eddy, S. Griffiths-Jones, K. L. Howe, M. Marshall, and E. L. L. Sonnhammer, Nucleic Acids Research, 30(1):276-280, 2002) or equivalent databases, such as InterPro and SMART databases ( http: / / www.ebi.ac.uk / interpro / scan.html , http: / / smart.embl-heidelberg.de / ) to provide a peptide or polypeptide that contains or consists of an active site domain predicted by the use thereof.
[0252] The present invention also encompasses "polypeptide variants" having the desired activity, where the variant polypeptide is selected from amino acid sequences having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity with a specific, particularly natural amino acid sequence referred to by a specific sequence number, and includes at least one substitution modification with respect to the aforementioned sequence number.
[0253] c. Coding nucleic acid sequences applicable according to the present invention In this regard, the following definitions apply: The terms "nucleic acid sequence", "nucleic acid", "nucleic acid molecule" and "polynucleotide" are used interchangeably to mean a sequence of nucleotides. A nucleic acid sequence may be single-stranded or double-stranded deoxyribonucleotides or ribonucleotides of any length, and includes coding and non-coding sequences of genes, exons, introns, sense and antisense complementary sequences, genomic DNA, cDNA, miRNA, siRNA, mRNA, rRNA, tRNA, recombinant nucleic acid sequences, isolated and purified naturally occurring DNA and / or RNA sequences, synthetic DNA and RNA sequences, fragments, primers and nucleic acid probes. One of ordinary skill in the art will recognize that an RNA nucleic acid sequence is identical to a DNA sequence, except that thymine (T) is replaced by uracil (U). The term "nucleic acid sequence" should also be understood to include polynucleotide molecules or oligonucleotide molecules in the form of separate fragments or as components of larger nucleic acids.
[0254] "Isolated nucleic acid" or "isolated nucleic acid sequence" refers to a nucleic acid or nucleic acid sequence that is in an environment different from its natural environment and can include those substantially free of contaminating endogenous substances.
[0255] As used herein, the term "naturally occurring" when applied to a nucleic acid refers to a nucleic acid found in the cells of organisms in nature and not intentionally modified by humans in the laboratory.
[0256] A "fragment" of a polynucleotide or nucleic acid sequence refers to a continuous nucleotide sequence where the length of the polynucleotide in the embodiments herein is, in particular, at least 15 bp, at least 30 bp, at least 40 bp, at least 50 bp and / or at least 60 bp. In particular, a fragment of a polynucleotide includes at least 25, more particularly at least 50, more particularly at least 75, more particularly at least 100, more particularly at least 150, more particularly at least 200, more particularly at least 300, more particularly at least 400, more particularly at least 500, more particularly at least 600, more particularly at least 700, more particularly at least 800, more particularly at least 900, more particularly at least 1000 consecutive nucleotides of the polynucleotide in the embodiments herein. Without limitation, fragments of the polynucleotides herein can be used as PCR primers and / or as probes, or for antisense gene silencing or RNAi.
[0257] As used herein, the term "hybridization" or "hybridizes under certain conditions" is intended to describe hybridization and washing conditions under which nucleotide sequences that are significantly identical or homologous to each other remain bound to each other. Such conditions may be conditions under which sequences having at least about 70%, such as at least about 80%, such as at least about 85%, 90%, or 95% identity remain bound to each other. Definitions of low stringency, medium, and high stringency hybridization conditions are provided later herein. Appropriate hybridization conditions can also be selected by one of ordinary skill in the art with minimal experimentation, as exemplified in Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6). Additionally, stringency conditions are described in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).
[0258] A "recombinant nucleic acid sequence" is a nucleic acid sequence that is the result of bringing together genetic material from multiple sources and making or modifying a nucleic acid sequence that does not occur naturally and is not found in organisms by other means, using laboratory methods such as molecular cloning.
[0259] "Recombinant DNA technology" refers to the molecular biology procedures for preparing recombinant nucleic acid sequences, as described, for example, in Weigel and Glazebrook, eds., Laboratory Manuals, 2002, Cold Spring Harbor Lab Press; and Sambrook et al., 1989, Cold Spring Harbor, NY, Cold Spring Harbor Laboratory Press.
[0260] The term "gene" means a DNA sequence that includes regions transcribed into RNA molecules such as mRNA within a cell and is operably linked to appropriate regulatory regions, such as a promoter. Thus, a gene can include a plurality of operably linked sequences, such as a promoter, for example, a 5' leader sequence including a sequence involved in translation initiation, a coding region of cDNA or genomic DNA, introns, exons, and / or a 3' untranslated sequence including, for example, a transcription termination site.
[0261] "Polycistronic" refers to a nucleic acid molecule, particularly mRNA, that can encode two or more polypeptides separately within the same nucleic acid molecule.
[0262] "Chimeric gene" refers to any gene that is not normally found in nature, particularly a gene in which one or more portions of nucleic acid sequences that are not normally associated with each other in nature are present. For example, a promoter is not associated in nature with a part or all of the transcribed region or another regulatory region. The term "chimeric gene" is understood to include an expression construct in which a promoter or transcriptional regulatory sequence is operably linked to one or more coding sequences or antisense, that is, the reverse complementary strand of the sense strand, or an inverted repeat sequence (sense and antisense, whereby the RNA transcript forms double-stranded RNA during transcription). The term "chimeric gene" also includes a gene obtained by combining parts of one or more coding sequences to generate a new gene.
[0263] "3' UTR" or "3' untranslated sequence" (also referred to as "3' untranslated region" or "3' end") refers to a nucleic acid sequence found downstream of the coding sequence of a gene, and this nucleic acid sequence includes, for example, a transcription termination site and a polyadenylation signal such as AAUAAA or its variants (although not all, but most eukaryotic mRNAs). After transcription termination, the mRNA transcript can be cleaved downstream of the polyadenylation signal and a poly(A) tail can be added, which is involved in the translation site of the mRNA, for example, the transport to the cytoplasm.
[0264] The term "primer" refers to a short nucleic acid sequence that hybridizes to a template nucleic acid sequence and is used for polymerization of a nucleic acid sequence complementary to the template.
[0265] The term "selectable marker" refers to any gene that, when expressed, can be used to select cells containing the selectable marker. Examples of selectable markers are shown below. One of ordinary skill in the art will know that selectable markers for different antibiotics, fungicides, auxotrophs, or herbicides are applicable to different target species.
[0266] The present invention also relates to nucleic acid sequences encoding polypeptides as defined herein.
[0267] In particular, the present invention also relates to nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, such as cDNA, genomic DNA, and mRNA) encoding one of the above polypeptides and functional equivalents thereof, which can be obtained, for example, using artificial nucleotide analogs.
[0268] The present invention relates to both an isolated nucleic acid molecule encoding a polypeptide according to the present invention or a biologically active segment thereof, and a nucleic acid fragment that can be used, for example, as a hybridization probe or primer for identifying or amplifying a coding nucleic acid according to the present invention.
[0269] The present invention also relates to nucleic acids having a certain degree of "identity" to the sequences specifically disclosed herein. "Identity" between two nucleic acids means nucleotide identity over the full length of the nucleic acid in each case.
[0270] "Identity" between two nucleotide sequences (the same applies to peptide or amino acid sequences) refers to the nucleotide residues (or amino acid residues) that are identical in the two sequences when an alignment of these two sequences is generated, or a function of the number thereof. Identical residues are defined as residues that are identical in two sequences at a given position in the alignment. As used herein, the percentage of sequence identity is calculated by taking the number of identical residues between two sequences from an optimal alignment, dividing it by the number of residues in the shortest sequence, and multiplying by 100. An optimal alignment is an alignment that gives the highest percentage of identity. To obtain an optimal alignment, gaps may be introduced into one or both of the sequences at one or more positions in the alignment. These gaps are considered as non-identical residues in the calculation of the percentage of sequence identity. Alignments for determining the percentage of identity of amino acid or nucleic acid sequences can be achieved in various ways using computer programs, such as publicly available computer programs available on the World Wide Web.
[0271] In particular, an optimal alignment of a protein or nucleic acid sequence can be obtained and the percentage of sequence identity can be calculated using the BLAST program (Tatiana et al, FEMS Microbiol Lett., 1999, 174:247-250, 1999) set with default parameters available from the website of the National Center for Biotechnology Information (NCBI) of the United States (ncbi.nlm.nih.gov / BLAST / bl2seq / wblast2.cgi).
[0272] In another example, identity can be determined using the Clustal method (Higgins DG, Sharp PM. ((1989))) with the following settings using the Vector NTI Suite 7.1 program of Informax, Inc. (USA).
[0273] Multiple alignment parameters: Gap opening penalty 10 Gap extension penalty 10 Gap separation penalty range 8 Gap separation penalty off %identity for alignment delay 40 Residue specific gaps off Hydrophilic residue gap off Transition weighing 0 Pairwise alignment parameters: FAST algorithm on K-tuple size 1 Gap penalty 3 Window size 5 Number of best diagonals 5 Alternatively, identity may be determined according to the web page of Chenna, et al. (2003): http: / / www.ebi.ac.uk / Tools / clustalw / index.html# and the following settings.
[0274] DNA Gap Open Penalty 15.0 DNA Gap Extension Penalty 6.66 DNA Matrix Identity Protein Gap Open Penalty 10.0 Protein Gap Extension Penalty 0.2 Protein matrix Gonnet Protein / DNA ENDGAP -1 Protein / DNA GAPDIST 4 All nucleic acid sequences described herein (single-stranded and double-stranded DNA and RNA sequences, such as cDNA and mRNA) can be generated by chemical synthesis from nucleotide building blocks, for example, by fragment condensation of individual overlapping complementary nucleic acid building blocks of a double helix, in a known manner. Chemical synthesis of oligonucleotides can be carried out in a known manner, for example, by the phosphoramidite method (Voet, Voet, 2nd edition, Wiley Press, New York, pages 896-897). The accumulation of synthetic oligonucleotides, the filling of gaps and ligation reactions by the Klenow fragment of DNA polymerase, and general cloning techniques are described in Sambrook et al. (1989) (see below).
[0275] The nucleic acid molecules according to the invention can further comprise untranslated sequences from the 3' and / or 5' ends of the coding gene region.
[0276] The invention further relates to nucleic acid molecules complementary to the specifically described nucleotide sequences or segments thereof.
[0277] The nucleotide sequences according to the invention enable the generation of probes and primers that can be used for the identification and / or cloning of homologous sequences in other cell types and organisms. Such probes or primers generally comprise nucleotide sequence regions that hybridize under "stringent" conditions (as defined separately herein) on at least about 12, preferably at least about 25, for example about 40, 50 or 75 consecutive nucleotides of the sense strand or the corresponding antisense strand of the nucleic acid sequence according to the invention.
[0278] "Homologous" sequences include ortholog or paralog sequences. Methods for identifying orthologs or paralogs, including phylogenetic methods, sequence similarity and hybridization methods, are known in the art and are described herein.
[0279] "Paralogs" result from gene duplication that produces two or more genes with similar sequences and similar functions. Paralogs are typically formed by duplication of genes within related plant species and are clustered together. Paralogs are found in groups of similar genes using pairwise Blast analysis or during phylogenetic analysis of gene families using programs such as CLUSTAL. In paralogs, consensus sequences that are characteristic of the sequences within related genes and have similar functions of the genes can be identified.
[0280] "Orthologs", or orthologous sequences, are sequences that are similar to each other because they are found in species that are descendants of a common ancestor. For example, it is known that plant species with a common ancestor contain many enzymes with similar sequences and functions. A person skilled in the art can identify ortholog sequences and predict the functions of orthologs, for example, by constructing a phylogenetic tree of a certain gene family using CLUSTAL or BLAST programs. As a method for identifying or confirming similar functions between homologous sequences, examples include comparing transcriptional profiles in which related polypeptides are overexpressed or deleted (knocked out / knocked down) in host cells or organisms such as plants or microorganisms. A person skilled in the art will understand that genes with similar transcriptional profiles, whether more than 50% of the regulated transcripts are common, more than 70% of the regulated transcripts are common, or more than 90% of the regulated transcripts are common, have similar functions. Homologs, paralogs, orthologs, and any other variants of the sequences herein are expected to function similarly by causing a terpene synthase protein to be produced in an organism such as a host cell, plant, or microorganism.
[0281] The term "selectable marker" refers to any gene that can be used, upon expression, to select cells or a plurality of cells containing the selectable marker. Examples of selectable markers are shown below. One skilled in the art will know that selectable markers for different antibiotics, fungicides, auxotrophs or herbicides are applicable to different target species.
[0282] The nucleic acid molecules according to the invention can be retrieved by standard techniques of molecular biology and the sequence information provided by the invention. For example, one or a segment of one of the specifically disclosed complete sequences can be used as a hybridization probe and standard hybridization techniques (such as those described in Sambrook, (1989)) to isolate cDNA from an appropriate cDNA library.
[0283] Furthermore, nucleic acid molecules containing one or a segment of one of the disclosed sequences can be isolated by polymerase chain reaction using oligonucleotide primers constructed based on this sequence. The nucleic acid amplified in this way can be cloned into an appropriate vector and characterized by DNA sequencing. The oligonucleotides according to the invention can also be generated using standard synthetic methods, for example, an automated DNA synthesizer.
[0284] The nucleic acid sequences according to the invention or derivatives, homologs or parts of these sequences can be isolated from other bacteria, for example, via genomic or cDNA libraries, by, for example, conventional hybridization techniques or PCR techniques. These DNA sequences hybridize with the sequences according to the invention under standard conditions.
[0285] "Hybridize" means the ability of a polynucleotide or oligonucleotide to bind to a substantially complementary sequence under standard conditions, while non-specific binding does not occur between non-complementary partners under these conditions. For this reason, the sequences can be 90 - 100% complementary. The property that complementary sequences can specifically bind to each other is utilized, for example, in Northern or Southern blotting, or in primer binding in PCR or RT-PCR.
[0286] For hybridization, short oligonucleotides with conserved regions are preferably used. However, it is also possible to use longer fragments or complete sequences of the nucleic acids according to the present invention for hybridization. These "standard conditions" vary depending on the nucleic acid used (oligonucleotide, longer fragment or complete sequence), or depending on which type of nucleic acid - DNA or RNA - is used for hybridization. For example, the melting temperature of a DNA:DNA hybrid is approximately 10 °C lower than that of a DNA:RNA hybrid of the same length.
[0287] For example, depending on the particular nucleic acid, standard conditions mean a temperature of 42°C to 58°C in an aqueous buffer having a concentration of 0.1 to 5×SSC (1×SSC = 0.15 M NaCl, 15 mM sodium citrate, pH 7.2), or alternatively in the presence of 50% formamide, for example at 42°C in 5×SSC in the presence of 50% formamide. Advantageously, the hybridization conditions for DNA:DNA hybrids are 0.1×SSC and a temperature of about 20°C to 45°C, preferably about 30°C to 45°C. The hybridization conditions for DNA:RNA hybrids are advantageously 0.1×SSC and a temperature of about 30°C to 55°C, preferably about 45°C to 55°C. The temperatures described for these hybridizations are examples of melting temperatures calculated for nucleic acids of about 100 nucleotides in length and a G+C content of 50% in the absence of formamide. The experimental conditions for DNA hybridization are described in textbooks of relevant genetics, such as Sambrook et al., 1989, and can be calculated, for example, using formulas known to those skilled in the art depending on the length of the nucleic acid, the type of hybrid, or the G+C content. Those skilled in the art can obtain further information on hybridization from the following textbooks: Ausubel et al. (eds), (1985), Brown (ed) (1991).
[0288] "Hybridization" can be carried out, in particular, under stringent conditions. Such hybridization conditions are described, for example, in Sambrook (1989), or Current Protocols in Molecular Biology, John Wiley & Sons, N.Y. (1989), 6.3.1-6.3.6.
[0289] As used herein, the term hybridize or hybridizes under certain conditions is intended to describe hybridization and washing conditions under which nucleotide sequences that are significantly identical or homologous to each other remain bound to each other. Such conditions may be conditions under which sequences having at least about 70%, such as at least about 80%, such as at least about 85%, 90%, or 95% identity remain bound to each other. Definitions of low stringency, moderate, and high stringency hybridization conditions are provided herein.
[0290] Appropriate hybridization conditions can also be selected by one of ordinary skill in the art with a minimum of experimentation, as exemplified in Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6). Additionally, stringency conditions are described in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).
[0291] As used herein, the defined conditions of low stringency are as follows. A filter containing DNA is pretreated for 6 hours at 40°C in a solution containing 35% formamide, 5×SSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml of denatured salmon sperm DNA. Hybridization is carried out with the following modifications to the same solution: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml of salmon sperm DNA, 10% (w / v) dextran sulfate, and 5 - 20×106 cpm of 32P-labeled probe. The filter is incubated in the hybridization mixture for 18 - 20 hours at 40°C and then washed at 55°C for 1.5 hours in a solution containing 2×SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS. The wash solution is replaced with a fresh solution and incubated at 60°C for a further 1.5 hours. The filter is blotted and dried and subjected to autoradiography.
[0292] As used herein, the defined conditions of medium stringency are as follows. A filter containing DNA is pretreated for 7 hours at 50°C in a solution containing 35% formamide, 5×SSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml of denatured salmon sperm DNA. Hybridization is carried out with the following modifications to the same solution: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml of salmon sperm DNA, 10% (w / v) dextran sulfate, and 5 - 20×106 cpm of 32P-labeled probe. The filter is incubated in the hybridization mixture for 30 hours at 50°C and then washed at 55°C for 1.5 hours in a solution containing 2×SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS. The wash solution is replaced with a fresh solution and incubated at 60°C for a further 1.5 hours. The filter is blotted and dried and subjected to autoradiography.
[0293] As used herein, the defined conditions of high stringency are as follows. Prehybridization of a filter containing DNA is carried out at 65°C for 8 hours to overnight in a buffer composed of 6×SSC, 50 mM Tris-HCl (pH 7.5), 1 mM EDTA, 0.02% PVP, 0.02% Ficoll, 0.02% BSA, and 500 μg / ml of denatured salmon sperm DNA. The filter is hybridized at 65°C for 48 hours in a prehybridization mixture containing 100 μg / ml of denatured salmon sperm DNA and 5 - 20×106 cpm of 32P-labeled probe. Washing of the filter is carried out at 37°C for 1 hour in a solution containing 2×SSC, 0.01% PVP, 0.01% Ficoll, and 0.01% BSA. Subsequently, washing is carried out at 50°C for 45 minutes in 0.1×SSC.
[0294] If the above conditions are inappropriate (e.g., conditions such as those employed for heterologous hybridization), other conditions of low, medium, and high stringency well-known in the art (e.g., conditions such as those employed for heterologous hybridization) may be used.
[0295] The detection kit for a nucleic acid sequence encoding the polypeptide of the present invention may include a primer and / or a probe specific to the nucleic acid sequence encoding the polypeptide, and a related protocol for detecting the nucleic acid sequence encoding the polypeptide in a sample using the primer and / or the probe. Such a detection kit can be used to determine whether a plant, organism, microorganism, or cell is modified with a sequence encoding the polypeptide, i.e., whether it is transformed.
[0296] To test the function of a variant DNA sequence according to an embodiment of the present specification, the sequence of interest is operably linked to a selectable or screenable marker gene, and the expression of the aforementioned reporter gene is tested in a transient expression assay, for example, using microorganisms, or using protoplasts, or using stably transformed plants.
[0297] The present invention also relates to derivatives of specifically disclosed nucleic acid sequences or nucleic acid sequences that can be derivatized.
[0298] Accordingly, further nucleic acid sequences according to the present invention can be derived from the sequences specifically disclosed herein, and differ therefrom by the addition, substitution, insertion or deletion of one or more nucleotides, for example, 1 to 20, particularly 1 to 15 or 5 to 10 nucleotides, and can further encode polypeptides having the desired profile of properties.
[0299] The present invention also encompasses nucleic acid sequences containing so-called silent mutations, or nucleic acid sequences modified according to the codon usage frequency of a particular primary or host organism, compared to the specifically described sequences.
[0300] According to a particular embodiment of the present invention, variant nucleic acids may be prepared to adapt the nucleotide sequence to a particular expression system. For example, in a bacterial expression system, it is known that polypeptides are expressed more efficiently when the amino acids are encoded by particular codons. Due to the degeneracy of the genetic code, multiple codons can encode the same amino acid sequence, and multiple nucleic acid sequences can encode the same protein or polypeptide, and all of these DNA sequences are encompassed by the embodiments of the present specification. Where appropriate, the nucleic acid sequences encoding the polypeptides described herein may be optimized to enhance expression in the host cell. For example, the nucleic acids of the embodiments herein may be synthesized using codons specific to the host to improve expression.
[0301] The present invention also encompasses naturally occurring variants of the sequences described herein, such as splicing variants or allelic variants.
[0302] Allelic variants can have at least 60% homology, preferably at least 80% homology, particularly preferably at least 90% homology over the entire sequence range at the level of the derived amino acids (for homology at the amino acid level, see the details described above for polypeptides). Advantageously, the homology can be high over partial regions of the sequence.
[0303] The present invention also relates to sequences obtainable by conservative nucleotide substitutions (i.e., substitutions that result in the amino acid in question being replaced by an amino acid of the same charge, size, polarity and / or solubility).
[0304] The present invention also relates to molecules derived by sequence polymorphisms from the specifically disclosed nucleic acids. Such genetic polymorphisms can be present in cells from different populations or within a population due to natural allelic mutations. Allelic variants can also include functional equivalents. These natural mutations usually result in a 1 - 5% difference in the nucleotide sequence of the gene. The aforementioned polymorphisms can result in changes in the amino acid sequence of the polypeptides disclosed herein. Allelic variants can also include functional equivalents.
[0305] Furthermore, it should be understood that derivatives are homologs of the nucleic acid sequences according to the invention, such as homologs from animals, plants, fungi or bacteria, truncated sequences, single - stranded DNA or RNA of coding and non - coding DNA sequences. For example, homologs have at least 40%, preferably at least 60%, particularly preferably at least 70%, very preferably at least 80% homology over the entire DNA region given by the sequences specifically disclosed herein at the DNA level.
[0306] Furthermore, the derivative should be understood to be, for example, a fusion with a promoter. The promoter added to the described nucleotide sequence can be changed by at least one nucleotide exchange, at least one insertion, inversion, and / or deletion without impairing the functionality or effectiveness of the promoter. Furthermore, the effectiveness of the promoter can be enhanced by changing their sequences, or can be completely exchanged with a more effective promoter even in organisms of different genera.
[0307] d. Generation of functional polypeptide variants Furthermore, those skilled in the art are proficient in methods for generating functional variants, i.e., nucleotide sequences encoding polypeptides having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity with any of the amino acid-related sequence numbers disclosed herein and / or encoded by nucleic acid molecules containing nucleotide sequences having at least 70% sequence identity with any of the amino acid-related sequence numbers disclosed herein.
[0308] Depending on the techniques used, those skilled in the art can introduce completely random mutations or other more directed mutations into genes or other non-coding nucleic acid regions (for example, important for regulating expression), and then generate a genetic library. The molecular biology methods required for this purpose are known to those skilled in the art and are described, for example, in Sambrook and Russell, Molecular Cloning. 3rd Edition, Cold Spring Harbor Laboratory Press 2001.
[0309] Methods for modifying genes and, consequently, the polypeptides encoded by those genes have long been known to those skilled in the art and include, for example, the following. - Site-directed mutagenesis, a method of directed replacement of individual or several nucleotides of a gene (Trower MK (Ed.) 1996; In vitro mutagenesis protocols. Humana Press, New Jersey), - Saturation mutagenesis, a method by which codons of any amino acid can be exchanged or added at any position of a gene (Kegler-Ebo DM, Docktor CM, DiMaio D (1994) Nucleic Acids Res 22:1593; Barettino D, Feigenbutz M, Valca’rel R, Stunnenberg HG (1994) Nucleic Acids Res 22:541; Barik S (1995) Mol Biotechnol 3:1), - Error-prone polymerase chain reaction, a method of mutating a nucleotide sequence with an error-prone DNA polymerase (Eckert KA, Kunkel TA (1990) Nucleic Acids Res 18:3739), - SeSaM method (sequence saturation mutagenesis), a method that prevents preferential exchange by a polymerase (Schenk et al., Biospektrum, Vol. 3, 2006, 277-279), - For example, in mutant strains with an increased mutation rate of nucleotide sequences due to defects in the DNA repair mechanism, gene passage (Greener A, Callahan M, Jerpseth B (1996) An efficient random mutagenesis technique using an E.coli mutator strain. In: Trower MK (Ed.) In vitro mutagenesis protocols. Humana Press, New Jersey), or - DNA shuffling, which forms a pool of related genes, digests them, uses the fragments as templates for polymerase chain reaction, and repeats strand separation and recombination to finally generate full-length mosaic genes (Stemmer WPC (1994) Nature 370:389; Stemmer WPC (1994) Proc Natl Acad Sci USA 91:10747).
[0310] Using so-called directed evolution (especially as described in Reetz MT and Jaeger K-E (1999), Topics Curr Chem 200:31; Zhao H, Moore JC, Volkov AA, Arnold FH (1999), Methods for optimizing industrial polypeptides by directed evolution, In: Demain AL, Davies JE (Ed.) Manual of industrial microbiology and biotechnology. American Society for Microbiology), a person skilled in the art can produce functional variants in a directed and large-scale manner. For this purpose, in the first step, a gene library of each polypeptide is first created using, for example, the methods shown above. The gene library is expressed in an appropriate manner, for example, by a bacterial or phage display system.
[0311] The relevant genes of a host organism that express functional variants having characteristics that approximately correspond to the desired characteristics can be subjected to another mutation cycle. The steps of mutation and selection or screening can be repeatedly iterated until the existing functional variants have the desired characteristics to a sufficient degree. Using this iterative procedure, a limited number of mutations, for example 1, 2, 3, 4, or 5 mutations, can be performed stepwise, their effects on the activity in question can be evaluated, and they can be selected. The selected variants can then be subjected to further mutation steps in the same manner. In this way, the number of individual variants to be investigated can be significantly reduced.
[0312] The results according to the present invention also provide important information regarding the structure and sequence of the relevant polypeptide, which is necessary to further generate polypeptides having the desired altered characteristics in a targeted manner. In particular, it is possible to define so-called "hot spots", i.e., sequence segments that are potentially suitable for altering characteristics by introducing target mutations.
[0313] Information can also be obtained regarding the position of the amino acid sequence, in which region mutations can be made that would presumably have little effect on the activity and can be referred to as potential "silent mutations".
[0314] e. A construct for expressing the polypeptide of the present invention In this regard, the following definitions apply: "Gene expression" includes "heterologous expression" and "overexpression" and includes the transcription of a gene and the translation of mRNA into protein. "Overexpression" refers to the production of a gene product measured by the level of mRNA, polypeptide, and / or enzyme activity in a transfected cell or organism being higher than the production level in a non-transformed cell or organism having a similar genetic background.
[0315] As used herein, the "expression vector" means a nucleic acid molecule designed using molecular biological methods and recombinant DNA techniques for delivering foreign or exogenous DNA into a host cell. An expression vector typically contains sequences necessary for proper transcription of a nucleotide sequence. The coding region usually encodes a protein of interest, but may also encode an RNA, such as antisense RNA, siRNA, etc.
[0316] The "expression vector" as used herein includes any linear or circular recombinant vector, including but not limited to viral vectors, bacteriophages, and plasmids. Those skilled in the art can select an appropriate vector according to the expression system. In one embodiment, the expression vector contains a nucleic acid of an embodiment herein operably linked to at least one "regulatory sequence" that controls transcription, translation, initiation, and termination, such as a transcription promoter, operator or enhancer, or an mRNA ribosome binding site, and optionally contains at least one selectable marker. A nucleotide sequence is "operably linked" when the regulatory sequence is functionally related to the nucleic acid of an embodiment herein.
[0317] The "expression system" as used herein includes any combination of nucleic acid molecules necessary for the expression of one polypeptide, or the co-expression of two or more polypeptides, either in vivo or in vitro in a given expression host. Each coding sequence may be arranged on a single nucleic acid molecule or vector, such as a vector containing a multiple cloning site, on a polycistronic nucleic acid, or dispersed on two or more physically different vectors. Specific examples include an operon containing a promoter sequence, one or more operator sequences, and one or more structural genes each encoding an enzyme described herein.
[0318] As used herein, the terms "amplify" and "amplification" refer to the use of any suitable amplification method to generate or detect recombinant forms of naturally expressed nucleic acids, as detailed below. For example, the present invention provides methods and reagents (e.g., specific degenerate oligonucleotide primer pairs, oligo dT primers) for amplifying (e.g., by polymerase chain reaction, PCR) nucleic acids of the present invention that are naturally expressed (e.g., genomic DNA or mRNA) or recombinant (e.g., cDNA) in vivo, ex vivo or in vitro.
[0319] "Regulatory sequence" refers to a nucleic acid sequence that can determine the expression level of a nucleic acid sequence in an embodiment herein and regulate the transcription rate of a nucleic acid sequence operably linked to the regulatory sequence. Regulatory sequences include promoters, enhancers, transcription factors, promoter elements, and the like.
[0320] "Promoter", "nucleic acid having promoter activity" or "promoter sequence" is understood to mean, according to the present invention, a nucleic acid that regulates the transcription of a nucleic acid when functionally linked to the nucleic acid to be transcribed. In particular, a "promoter" refers to a nucleic acid sequence that controls the expression of a coding sequence by providing a binding site for RNA polymerase and other factors necessary for proper transcription, including but not limited to transcription factor binding sites, repressor and activator protein binding sites. The meaning of the term promoter also includes the term "promoter regulatory sequence". Promoter regulatory sequences can include upstream and downstream elements that may affect the transcription, RNA processing or stability of the associated coding nucleic acid sequence. Promoters include naturally occurring sequences and synthetic sequences. The coding nucleic acid sequence is usually located downstream of the promoter with respect to the direction of transcription starting from the transcription start site.
[0321] In this context, a "functional" or "operable" linkage is understood to mean, for example, a continuous arrangement of one of the nucleic acids and the regulatory sequence. For example, a sequence having promoter activity, and the sequence of the nucleic acid sequence to be transcribed, and optionally further regulatory elements, such as a nucleic acid sequence that ensures transcription of the nucleic acid, and for example a terminator, are linked such that each of the regulatory elements can perform its function during transcription of the nucleic acid sequence. This does not necessarily require a direct linkage in the chemical sense. Gene regulatory sequences, such as enhancer sequences, can even exert their function on the target sequence from a more distant position or from other DNA molecules. A preferred arrangement is one in which the nucleic acid sequence to be transcribed is located behind (i.e., at the 3'-end) the promoter sequence and the two sequences are covalently linked. The distance between the promoter sequence and the nucleic acid sequence recombinantly expressed can be less than 200 base pairs, or less than 100 base pairs, or less than 50 base pairs.
[0322] In addition to promoters and terminators, examples of other regulatory elements can include the following: target sequences, enhancers, polyadenylation signals, selectable markers, amplification signals, origins of replication, and the like. Suitable regulatory sequences are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).
[0323] The term "constitutive promoter" refers to an unregulated promoter that allows continuous transcription of the nucleic acid sequence to which it is operably linked.
[0324] As used herein, the term "operably linked" means the linkage of polynucleotide elements that are in a functional relationship. Nucleic acids are "operably linked" when placed in a functional relationship with another nucleic acid sequence. For example, a promoter, or rather a transcriptional regulatory sequence, is operably linked to a coding sequence if it affects the transcription of the coding sequence. Being operably linked means that the linked DNA sequences are typically contiguous. The nucleotide sequences associated with a promoter sequence may be of homologous or heterologous origin with respect to the plant to be transformed. This sequence may also be wholly or partially synthetic. Regardless of origin, the nucleic acid sequence associated with a promoter sequence, after binding to the polypeptide of the embodiments herein, is expressed or silenced according to the characteristics of the promoter to which it is linked. The associated nucleic acid may encode a protein that is desired to be expressed or repressed either constitutively throughout the organism, or at a particular time, or in a particular tissue, cell, or cell compartment. Such nucleotide sequences in particular encode a protein that confers a desired phenotypic trait on the host cell or organism in which it is used to change or transform it. More particularly, the associated nucleotide sequence results in the production of the product or products of interest as defined herein in a cell or organism. In particular, this nucleotide sequence encodes a polypeptide having the enzymatic activity as defined herein.
[0325] The nucleotide sequences described herein may be part of an "expression cassette". The terms "expression cassette" and "expression construct" are used synonymously. A (preferably recombinant) expression construct comprises a nucleotide sequence encoding a polypeptide according to the invention and under the genetic control of a regulatory nucleic acid sequence.
[0326] In the processes applied according to the invention, the expression cassette may be part of an "expression vector", in particular part of a recombinant expression vector.
[0327] The "expression unit" is understood to mean, according to the present invention, a nucleic acid that contains a promoter defined herein and is functionally linked to a nucleic acid or gene to be expressed, and has an expression activity that regulates the expression, i.e., transcription and translation, of the aforementioned nucleic acid or the aforementioned gene. Therefore, the expression unit is also referred to as a "regulatory nucleic acid sequence" in this context. In addition to the promoter, other regulatory elements, such as enhancers, may also be present.
[0328] The "expression cassette" or "expression construct" is understood to mean, according to the present invention, an expression unit functionally linked to a nucleic acid to be expressed or a gene to be expressed. Therefore, in contrast to the expression unit, the expression cassette includes not only the nucleic acid sequence that regulates transcription and translation, but also the nucleic acid sequence that is to be expressed as a protein as a result of transcription and translation.
[0329] The terms "expression" or "overexpression" in the context of the present invention refer to the production of one or more polypeptides encoded by the corresponding DNA in a microorganism or the enhancement of intracellular activity. For this purpose, for example, it is possible to introduce a gene into an organism, replace an existing gene with another gene, increase the copy number of a gene, use a strong promoter, or use a gene encoding a corresponding polypeptide with high activity, and optionally combine these means.
[0330] Preferably, such a construct according to the present invention contains a promoter upstream of the 5' of each coding sequence, a terminator sequence downstream of the 3', and optionally other normal regulatory elements, and is operably linked to the coding sequence in each case.
[0331] The nucleic acid construct according to the invention particularly comprises a sequence encoding a polypeptide derived from, for example, the amino acid-related sequence numbers described herein, or its reverse complement, or its derivatives and homologs, and is advantageously operably or functionally linked to one or more regulatory signals in order to control, for example, increase gene expression.
[0332] In addition to these regulatory sequences, the natural regulation of these sequences may still be present in front of the actual structural gene and, optionally, the natural regulation can be genetically modified so that it is switched off and the expression of the gene is enhanced. However, the nucleic acid construct may also have a simpler structure, i.e., when no additional regulatory signals are inserted in front of the coding sequence and the natural promoter with its regulation has not been removed. Instead, the natural regulatory sequences are mutated so that the control is lost and the gene expression level is increased.
[0333] Preferred nucleic acid constructs advantageously also comprise one or more of the aforementioned "enhancer" sequences functionally linked to a promoter, which sequences enable enhanced expression of the nucleic acid sequence. Advantageous sequences such as further regulatory elements or terminators may be inserted at the 3'-end of the DNA sequence. One or more copies of the nucleic acid according to the invention may be present in the construct. The construct may optionally also contain other markers, such as genes complementing auxotrophy or antibiotic resistance, to select the construct.
[0334] Examples of suitable regulatory sequences are promoters such as cos, tac, trp, tet, trp-tet, lpp, lac, lpp-lac, lacI q , T7, T5, T3, gal, trc, ara, rhaP(rhaP BAD )SP6, lambda-P R , or lambda-P LThey are present in promoters which are advantageously used in Gram-negative bacteria. Further advantageous regulatory sequences are present, for example, in the promoters amy and SPO2 of Gram-positive bacteria, in the promoters ADC1, MFalpha, AC, P-60, CYC1, GAPDH, TEF, rp28, ADH of yeast or fungi. It is also possible to regulate using artificial promoters.
[0335] For expression in a host organism, the nucleic acid construct is advantageously inserted into a vector such as, for example, a plasmid or a phage, enabling optimal expression of the gene in the host. A vector means, in addition to plasmids and phages, all other vectors known to the person skilled in the art, namely viruses such as SV40, CMV, baculovirus and adenovirus, transposons, IS elements, fosmids, cosmids and linear or circular DNA or artificial chromosomes. These vectors can replicate autonomously in the host organism and can otherwise also replicate on the chromosome. These vectors are further developments of the present invention. Binary vectors or cpo integration vectors are also applicable.
[0336] Suitable plasmids are, for example, pLG338, pACYC184, pBR322, pUC18, pUC19, pKC30, pRep4, pHS1, pKK223-3, pDHE19.2, pHS2, pPLc236, pMBL24, pLG200, pUR290, pIN-III 113 -B1, λgt11 or pBdCI, pIJ101, pIJ364, pIJ702 or pIJ361 of Streptomyces, pUB110, pC194 or pBD214 of Bacillus, pSA77 or pAJ667 of Corynebacterium, pALS1, pIL2 or pBB116 of fungi, 2alphaM, pAG-1, YEp6, YEp13 or pEMBLYe23 of yeast, or pLGV23, pGHlac of plants +、in pBIN19, pAK2004 or pDH51. The plasmids described above are only a small part of the possible plasmids. Further plasmids are well known to the person skilled in the art and can be found, for example, in the book Cloning Vectors (Eds. Pouwels P. H. et al. Elsevier, Amsterdam - New York - Oxford, 1985, ISBN 0 444 904018).
[0337] As a further development of the vector, the nucleic acid construct according to the invention or the vector comprising the nucleic acid according to the invention can advantageously be introduced into a microorganism in the form of linear DNA and integrated into the genome of the host organism via heterologous or homologous recombination. This linear DNA can consist of a linearized vector such as a plasmid or can consist only of the nucleic acid construct or nucleic acid according to the invention.
[0338] For optimal expression of a heterologous gene in an organism, it is advantageous to modify the nucleic acid sequence so as to match the specific "codon usage frequency" used in the organism. This "codon usage frequency" can be easily determined by computer - evaluating other known genes of the organism in question.
[0339] The expression cassette according to the present invention is generated by fusing a suitable promoter to a suitable coding nucleotide sequence and a terminator or polyadenylation signal. For this purpose, conventional recombinant and cloning techniques described, for example, in T. Maniatis, E.F. Fritsch and J. Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1989), T.J. Silhavy, M.L. Berman and L.W. Enquist, Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1984), Ausubel, F.M. et al., Current Protocols in Molecular Biology, Greene Publishing Assoc. and Wiley Interscience (1987) are used.
[0340] For expression in a suitable host organism, the recombinant nucleic acid construct or gene construct is preferably inserted into a host-specific vector that allows optimal expression of the gene in the host. Vectors are well known to those skilled in the art and can be found, for example, in "cloning vectors" (Pouwels P. H. et al., Ed., Elsevier, Amsterdam-New York-Oxford, 1985).
[0341] Alternative embodiments of the embodiments herein provide a method of "altering gene expression" in a host cell. For example, the polynucleotides of the embodiments herein may be enhanced or overexpressed or induced in a host cell or host organism under certain circumstances (e.g., upon exposure to certain temperatures or culture conditions).
[0342] When the expression of the polynucleotides provided herein changes, ectopic expression may occur, which is an expression pattern different between the changed organism and a control or wild-type organism. The change in expression results from the interaction between the polypeptides of the embodiments herein and exogenous or endogenous modulators, or as a result of chemical modification of the polypeptide. This term also refers to the situation where the expression pattern of the polynucleotides of the embodiments herein has changed below the detection level or the activity has been completely suppressed.
[0343] In one embodiment, provided herein is also an isolated, recombinant or synthetic polynucleotide encoding a polypeptide or variant polypeptide provided herein.
[0344] In one embodiment, several polypeptides encoding nucleic acid sequences are co-expressed in a single host, particularly under the control of different promoters. In another embodiment, several polypeptides encoding nucleic acid sequences can be present on a single transformation vector or, using separate vectors, co-transformed while selecting transformants containing both chimeric genes. Similarly, one or polypeptides encoding a gene can be expressed together with other chimeric genes in a single plant, cell, microorganism or organism.
[0345] f. Hosts applicable to the present invention Depending on the context, the term "host" can mean a wild-type host or a genetically modified recombinant host, or both.
[0346] In principle, all prokaryotes or eukaryotes can be regarded as hosts or recombinant host organisms for the nucleic acids or nucleic acid constructs according to the present invention.
[0347] Using the vector according to the present invention, for example, a recombinant host can be produced that is transformed with at least one vector according to the present invention and can be used to produce the polypeptide according to the present invention. Advantageously, the recombinant construct according to the present invention described above is introduced into a suitable host system for expression. Preferably, the nucleic acids described are expressed in their respective expression systems using common cloning and transfection methods known to those skilled in the art, such as coprecipitation, protoplast fusion, electroporation, retroviral transfection, etc. Suitable systems are described, for example, in Current Protocols in Molecular Biology, F. Ausubel et al., Ed., Wiley Interscience, New York 1997, or Sambrook et al. Molecular Cloning: A Laboratory Manual. 2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
[0348] Advantageously, a microorganism such as a bacterium, fungus or yeast is used as the host organism. Advantageously, Gram-positive or Gram-negative bacteria are used, preferably bacteria of the family Enterobacteriaceae, Pseudomonadaceae, Rhizobiaceae, Streptomycetaceae, Streptococcaceae or Nocardiaceae, particularly preferably bacteria of the genus Escherichia, Pseudomonas, Streptomyces, Lactococcus, Nocardia, Burkholderia, Salmonella, Agrobacterium, Clostridium or Rhodococcus. Strains of Escherichia coli are highly preferred. Furthermore, other advantageous bacteria can be found in the groups of the class Alphaproteobacteria, Betaproteobacteria or Gammaproteobacteria. Advantageously, yeasts such as the family Saccharomyces or Pichia are also suitable hosts.
[0349] Alternatively, the whole plant or plant cells can also be used as natural or recombinant hosts. Non-limiting examples include the following plants or cells derived therefrom: the genus Nicotiana, particularly Nicotiana benthamiana and Nicotiana tabacum, and the genus Arabidopsis, particularly Arabidopsis thaliana.
[0350] Depending on the host organism, the organisms used in the method according to the invention are grown or cultured in a manner known to those skilled in the art. The culture may be batch, semi-batch or continuous. The nutrients may be present at the start of the fermentation or supplied semi-continuously or continuously later. This will also be described in detail below.
[0351] g. Recombinant production of the polypeptide according to the invention The invention further relates to a method for recombinantly producing a polypeptide according to the invention or a functional, biologically active fragment thereof, wherein a microorganism producing the polypeptide is cultured and the expression of the polypeptide is induced by applying at least one inducer that optionally induces gene expression, and the expressed polypeptide is isolated from the culture. The polypeptide can also be produced on an industrial scale in this way if necessary.
[0352] The microorganisms produced according to the invention can be cultured continuously or discontinuously by a batch method or a fed-batch method or a repeated fed-batch method. An overview of known culture methods can be found in the textbook by Chmiel (Bioprozesstechnik 1. Einfuehrung in die Bioverfahrenstechnik [Bioprocess technology 1. Introduction to bioprocess technology] (Gustav Fischer Verlag, Stuttgart, 1991)) or the textbook by Storhas (Bioreaktoren und periphere Einrichtungen [Bioreactors and peripheral equipment] (Vieweg Verlag, Braunschweig / Wiesbaden, 1994)).
[0353] The culture medium to be used must appropriately meet the requirements of each strain. The American Society for Microbiology's manual "Manual of Methods for General Bacteriology" (Washington D.C., USA, 1981) describes culture media for various microorganisms.
[0354] These media that can be used according to the present invention usually contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.
[0355] Preferred carbon sources are saccharides such as monosaccharides, disaccharides or polysaccharides. Very excellent carbon sources are, for example, glucose, fructose, mannose, galactose, ribose, sorbose, ribulose, lactose, maltose, sucrose, raffinose, starch or cellulose. Saccharides can also be added to the medium via complex compounds such as molasses, or by-products of sugar refining. It is also advantageous to add a mixture of different carbon sources. Other possible carbon sources are oils and fats such as soybean oil, sunflower oil, peanut oil and coconut oil, fatty acids such as palmitic acid, stearic acid or linoleic acid, alcohols such as glycerol, methanol or ethanol, and organic acids such as acetic acid or lactic acid.
[0356] The nitrogen source is usually an organic or inorganic nitrogen compound, or a material containing these compounds. Examples of nitrogen sources include ammonia gas or ammonium salts such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate or ammonium nitrate, nitrates, urea, amino acids or complex nitrogen sources such as corn steep liquor, soybean flour, soy protein, yeast extract, meat extract and the like. The nitrogen source can be used alone or in combination.
[0357] Inorganic salt compounds that may be present in the medium include chlorides, phosphates or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper and iron.
[0358] As sulfur sources, sulfur-containing inorganic compounds such as sulfates, sulfites, dithionites, tetrathionates, thiosulfates, sulfides, and organic sulfur compounds such as mercaptans and thiols can be used.
[0359] As phosphorus sources, phosphoric acid, potassium dihydrogen phosphate or dipotassium hydrogen phosphate, or their corresponding sodium-containing salts can be used.
[0360] To retain metal ions in the solution, a chelating agent can be added to the medium. Particularly suitable chelating agents include dihydroxy phenols such as catechol or protocatechuic acid, or organic acids such as citric acid.
[0361] The fermentation medium used according to the present invention usually also contains other growth factors such as vitamins or growth promoters, including, for example, biotin, riboflavin, thiamine, folic acid, nicotinic acid, pantothenic acid and pyridoxine. Growth factors and salts often originate from components of complex media such as yeast extract, molasses, corn steep liquor, etc. Furthermore, appropriate precursors can also be added to the culture medium. The exact composition of the compounds in the medium strongly depends on each experiment and is determined individually for each specific case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (Ed. P.M. Rhodes, P.F. Stanbury, IRL Press (1997) p. 53-73, ISBN 0 19 963577 3). The growth medium can be obtained from commercial suppliers such as Standard 1 (Merck) or BHI (Brain Heart Infusion Medium, DIFCO).
[0362] All components of the medium are sterilized by heat (20 minutes at 1.5 bar and 121 °C) or by sterile filtration. Each component can be sterilized together or separately as required. All components of the medium may be present at the start of the culture or added continuously or in batches.
[0363] The culture temperature is usually from 15 °C to 45 °C, preferably from 25 °C to 40 °C, and can be varied or kept constant during the experiment. The pH of the medium is in the range of 5 to 8.5, preferably about 7.0. The pH during growth can be controlled by adding basic compounds such as sodium hydroxide, potassium hydroxide, ammonia or aqueous ammonia, or acidic compounds such as phosphoric acid or sulfuric acid. To control foaming, an antifoaming agent such as a fatty acid polyglycol ester can be used. To maintain the stability of the plasmid, an appropriate selective substance, such as an antibiotic, can be added to the medium. To maintain aerobic conditions, oxygen or an oxygen-containing gas mixture, such as ambient air, is supplied to the culture. The temperature of the culture is usually in the range of 20 °C to 45 °C. The culture is continued until the maximum amount of the desired product is formed. Usually, this goal is achieved within 10 to 160 hours.
[0364] The fermentation broth is then further processed. Depending on the requirements, the biomass can be completely or partially removed from the fermentation broth or left entirely in the fermentation broth by separation techniques such as centrifugation, filtration, decantation or combinations of these methods.
[0365] If the polypeptide is not secreted into the culture medium, the cells can also be lysed and the product obtained from the lysate by known methods for protein isolation. Optionally, the cells can be disrupted by high-frequency ultrasound, high pressure, such as high pressure in a French press, osmotic lysis, the action of detergents, lytic enzymes or organic solvents, a homogenizer, or a combination of several of the aforementioned methods.
[0366] The polypeptide can be purified by known chromatography techniques such as molecular sieve chromatography (gel filtration), such as Q-sepharose chromatography, ion exchange chromatography and hydrophobic chromatography, as well as other conventional techniques such as ultrafiltration, crystallization, salting out, dialysis and native gel electrophoresis. Suitable methods are described, for example, in Cooper, T. G., Biochemische Arbeitsmethoden [Biochemical processes], Verlag Walter de Gruyter, Berlin, New York, or Scopes, R., Protein Purification, Springer Verlag, New York, Heidelberg, Berlin.
[0367] It may be advantageous to use a vector system or oligonucleotides to isolate the recombinant protein, which oligonucleotides extend the cDNA by a defined nucleotide sequence and thus encode a modified polypeptide or fusion protein, whereby, for example, it can be purified more easily. Suitable modifications of this type are, for example, so-called "tags" that function as anchors, such as modifications known as hexahistidine anchors or epitopes that can be recognized as antigens of antibodies (for example, described in Harlow, E. and Lane, D., 1988, Antibodies: A Laboratory Manual. Cold Spring Harbor (N.Y.) Press). These anchors can serve to attach the protein to a solid support, such as a polymeric matrix, which can be used, for example, as a filling for a chromatography column or on a microtiter plate or other support.
[0368] At the same time, these anchors can also be used for protein recognition. Furthermore, for protein recognition, normal markers such as fluorescent dyes, enzyme markers that form detectable reaction products after reaction with substrates, or radioactive markers can be used alone or in combination with anchors for protein derivatization.
[0369] h. Immobilization of polypeptides The enzyme or polypeptide according to the present invention can be used in the free state or immobilized in the methods described herein. An immobilized enzyme is an enzyme immobilized on an inert carrier. Suitable carrier materials and the enzymes immobilized thereon are known from European Patent Application Publication No. 1149849, European Patent Application Publication No. 1069183, German Patent Application Publication (DE-OS) No. 100193773 and the documents cited therein. In this regard, the disclosures of these documents are incorporated herein by reference in their entirety. Suitable carrier materials include, for example, clays, clay minerals such as kaolinite, diatomaceous earth, perlite, silica, aluminum oxide, sodium carbonate, calcium carbonate, cellulose powder, anion exchanger materials, synthetic polymers such as polystyrene, acrylic resins, phenol formaldehyde resins, polyurethanes and polyolefins such as polyethylene and polypropylene. For the production of supported enzymes, the carrier material usually takes the form of finely divided particles, preferably in a porous form. The particle size of the carrier material is usually 5 mm or less, particularly 2 mm or less (particle size distribution curve). Similarly, when using dehydrogenase as a catalyst for the whole cell, the free or immobilized form can be selected. The carrier materials are, for example, calcium alginate and carrageenan. The enzyme or cell can also be directly cross-linked with glutaraldehyde (cross-linking to CLEA). Other corresponding immobilization techniques are described, for example, in J. Lalonde and A. Margolin “Immobilization of Enzymes” in K. Drauz and H. Waldmann, Enzyme Catalysis in Organic Synthesis 2002, Vol. III, 991-1032, Wiley-VCH, Weinheim. Further information on biotransformation and bioreactors for carrying out the method according to the present invention is also described, for example, in Rehm et al. (eds) Biotechnology, 2nd Edn, Vol 3, Chapter 17, VCH, Weinheim.
[0370] i. Reaction conditions of the method for producing a biocatalyst of the present invention The reaction of the present invention can be carried out under in vivo or in vitro conditions.
[0371] At least one polypeptide / enzyme present between the individual steps of the method of the present invention or the multi-step method defined herein may be present in living cells that produce the enzyme or enzymes naturally or recombinantly, in harvested cells, i.e., under in vivo conditions, or in dead cells, in permeabilized cells, in crude cell extracts, in purified extracts, or in an essentially pure or completely pure form, i.e., under in vitro conditions. At least one enzyme may be present in solution or as an enzyme immobilized on a carrier. One or more enzymes may be present simultaneously in soluble and / or immobilized forms.
[0372] The method according to the invention can be carried out in a common reactor known to those skilled in the art, in a range of different scales, for example from laboratory scale (reaction volumes from a few millilitres to several tens of litres) to industrial scale (reaction volumes from several litres to several thousand cubic metres). When the polypeptide is used in an encapsulated form by abiotic, optionally permeabilized cells, in the form of more or less purified cell extracts, or in a purified form, a chemical reactor can be used. The chemical reactor enables the control of, usually, the amount of at least one enzyme, the amount of at least one substrate, the pH, the temperature, and the circulation of the reaction medium. When at least one polypeptide / enzyme is present within living cells, the process is a fermentation. In this case, the biocatalyst production is carried out in a bioreactor (fermenter), and the parameters necessary for the appropriate survival conditions of the living cells (for example, a culture medium containing nutrients, temperature, aeration, the presence or absence of oxygen or other gases, antibiotics, etc.) can be controlled. A person skilled in the art is familiar with chemical reactors or bioreactors, and, for example, the procedures for scaling up chemical or biotechnology methods from laboratory scale to industrial scale, or the procedures for optimizing process parameters, are also well described in the literature (for biotechnology methods, see, for example, Crueger und Crueger, Biotechnologie‐Lehrbuch der angewandten Mikrobiologie, 2. Ed., R. Oldenbourg Verlag, Muenchen, Wien, 1984).
[0373] Cells containing at least one enzyme can be permeabilized by ultrasonic or high-frequency pulses, physical or mechanical means such as a French press, or chemical means such as hypotonic solutions, lytic enzymes, and detergents present in the medium, or combinations thereof. Examples of detergents are digitonin, n-dodecyl maltoside, octyl glucoside, Triton® X-100, Tween® 20, deoxycholate, CHAPS (3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate), Nonidet® P40 (ethylphenol poly(ethylene glycol ether)), and the like.
[0374] For the bioconversion reaction of the present invention, instead of living cells, a biomass of non-viable cells containing the necessary biocatalyst may be applied.
[0375] When at least one enzyme is immobilized, it is attached to an inert carrier as described above.
[0376] The conversion reaction can be carried out batchwise, semi-batchwise or continuously. The reactants (and optionally nutrients) may be supplied at the start of the reaction or subsequently supplied semi-continuously or continuously.
[0377] The reaction of the present invention can be carried out in an aqueous, aqueous-organic or non-aqueous reaction medium depending on the particular reaction type.
[0378] The aqueous or aqueous-organic medium may contain a suitable buffer to adjust the pH to a value in the range of 5 to 11, for example 6 to 10.
[0379] Organic solvents that are miscible, partially miscible or immiscible with water can be applied to the aqueous-organic medium. Non-limiting examples of suitable organic solvents are shown below. Further examples are monovalent or polyvalent, aromatic or aliphatic alcohols, especially polyvalent aliphatic alcohols such as glycerol.
[0380] The non-aqueous medium substantially does not contain water, that is, it contains less than about 1% by weight or 0.5% by weight of water.
[0381] The biocatalytic process may be carried out in an organic non-aqueous medium. Suitable organic solvents include aliphatic hydrocarbons having, for example, 5 to 8 carbon atoms such as pentane, cyclopentane, hexane, cyclohexane, heptane, octane or cyclooctane; aromatic hydrocarbons such as benzene, toluene, xylene, chlorobenzene or dichlorobenzene; aliphatic acyclic ethers such as diethyl ether, methyl-tert-butyl ether, ethyl-tert-butyl ether, dipropyl ether, diisopropyl ether, dibutyl ether; or mixtures thereof.
[0382] The concentration of the reactant / substrate can be adjusted according to the optimal reaction conditions, which depends on the specific enzyme applied. For example, the initial substrate concentration is 0.1 to 0.5 M, for example 10 to 100 mM.
[0383] The reaction temperature may be adjusted according to the optimal reaction conditions that may depend on the specific enzyme applied. For example, the reaction may be carried out at a temperature in the range of 0 to 70 °C, for example 20 to 50 °C or 25 to 40 °C. Examples of reaction temperatures are about 30 °C, about 35 °C, about 37 °C, about 40 °C, about 45 °C, about 50 °C, about 55 °C and about 60 °C.
[0384] The process may proceed until equilibrium between the substrate and the product is achieved, or it may stop earlier. The normal process time is in the range of 1 minute to 25 hours, especially in the range of 10 minutes to 6 hours, for example in the range of 1 hour to 4 hours, especially in the range of 1.5 hours to 3.5 hours. These parameters are non-limiting examples of suitable process conditions.
[0385] When the host is a transgenic plant, optimal growth conditions, such as optimal light, water and nutrient conditions, can be provided.
[0386] k. Isolation of the product The method of the present invention can further include the step of recovering the final product or intermediate product, optionally, in a stereoisomerically or enantiomerically substantially pure form. The term "recovery" includes extracting, harvesting, isolating or purifying the compound from the culture or reaction medium. The recovery of the compound can be carried out according to any conventional isolation or purification method known in the art, including but not limited to treatment with conventional resins (e.g., anion or cation exchange resins, nonionic adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silicic acid, silica gel, cellulose, alumina, etc.), pH change, solvent extraction (e.g., treatment with conventional solvents such as alcohol, ethyl acetate, hexane, etc.), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, freeze-drying, etc.
[0387] The identity and purity of the isolated product can be determined by known techniques such as high performance liquid chromatography (HPLC), gas chromatography (GC), spectroscopy (IR, UV, NMR, etc.), colorimetry, TLC, NIRS, enzymatic or microbial assays, etc. (see, for example, the following. Patek et al. (1994) Appl. Environ. Microbiol. 60:133-140; Malakhova et al. (1996) Biotekhnologiya 11 27-32; and Schmidt et al. (1998) Bioprocess Engineer. 19:67-70. Ullmann’s Encyclopedia of Industrial Chemistry (1996) Bd. A27, VCH: Weinheim, P. 89-90, P. 521-540, P. 540-547, P. 559-566, 575-581 and P. 581-587; Michal, G (1999) Biochemical Pathways: An Atlas of Biochemistry and Molecular Biology, John Wiley and Sons; Fallon, A. et al. (1987) Applications of HPLC in Biochemistry in: Laboratory Techniques in Biochemistry and Molecular Biology, Bd. 17.).
[0388] Any cyclic terpene compound produced by any method described herein can be converted into derivatives including, but not limited to, hydrocarbons, esters, amides, glycosides, ethers, epoxides, aldehydes, ketones, alcohols, diols, acetals or ketals. Derivatives of terpene compounds can be obtained by chemical methods including, but not limited to, oxidation, reduction, alkylation, acylation and / or rearrangement. Alternatively, terpene compound derivatives can be obtained using biochemical methods by contacting the terpene compound with enzymes including, but not limited to, oxidoreductases, monooxygenases, dioxygenases, transferases. Biochemical conversions can be carried out in vitro using isolated enzymes, enzymes from lysed cells, or in vivo using whole cells.
[0389] l. Fermentative production of terpene / terpenoid compounds, such as labdane-type compounds The present invention also relates to a method for the fermentative production of terpene / terpenoid compounds, such as labdane-type compounds.
[0390] The fermentation used according to the present invention can be carried out, for example, in stirred fermenters, bubble columns and loop reactors. An inclusive overview of the types of possible methods, including the type of stirrer and geometric design, is described in "Chmiel: Bioprozesstechnik: Einfuehrung in die Bioverfahrenstechnik, Band 1". In the process of the present invention, typical variants available are known to those skilled in the art or are the following variants described, for example, in "Chmiel, Hammes and Bailey: Biochemical Engineering", for example batch, fed-batch, repeated fed-batch or other continuous fermentations with and without biomass recycle. Depending on the production strain, sparging with air, oxygen, carbon dioxide, hydrogen, nitrogen or a suitable gas mixture can be carried out to achieve a good yield (YP / S).
[0391] The medium to be used must meet the requirements of a specific strain in an appropriate manner. The American Society for Microbiology's handbook "Manual of Methods for General Bacteriology" (Washington D.C., USA, 1981) describes culture media for various microorganisms.
[0392] These media that can be used according to the present invention may contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.
[0393] Preferred carbon sources are saccharides such as monosaccharides, disaccharides or polysaccharides. Very good carbon sources are, for example, glucose, fructose, mannose, galactose, ribose, sorbose, ribulose, lactose, maltose, sucrose, raffinose, starch or cellulose. The saccharides can also be added to the medium via complex compounds such as molasses, or by-products of sugar refining. It is also advantageous to add a mixture of different carbon sources. Other possible carbon sources are oils and fats such as soybean oil, sunflower oil, peanut oil and coconut oil, fatty acids such as palmitic acid, stearic acid or linoleic acid, alcohols such as glycerol, methanol or ethanol, and organic acids such as acetic acid or lactic acid.
[0394] The nitrogen source is usually an organic or inorganic nitrogen compound, or a material containing these compounds. Examples of nitrogen sources include ammonia gas or ammonium salts such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate or ammonium nitrate, nitrates, urea, amino acids or complex nitrogen sources such as corn steep liquor, soybean meal, soy protein, yeast extract, meat extract, etc. The nitrogen source can be used alone or in combination.
[0395] Inorganic salt compounds that can be present in the medium include chlorides, phosphates, or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper, and iron.
[0396] As a sulfur source, sulfur-containing inorganic compounds such as sulfates, sulfites, dithionites, tetrathionates, thiosulfates, sulfides, and organic sulfur compounds such as mercaptans and thiols can also be used.
[0397] As a phosphorus source, phosphoric acid, potassium dihydrogen phosphate, or dipotassium hydrogen phosphate, or their corresponding sodium-containing salts can be used.
[0398] To retain metal ions in solution, a chelating agent can be added to the medium. Particularly suitable chelating agents include dihydroxyphenols such as catechol or protocatechuic acid, or organic acids such as citric acid.
[0399] The fermentation medium used according to the present invention may also contain other growth factors such as vitamins or growth promoters, including, for example, biotin, riboflavin, thiamine, folic acid, nicotinic acid, pantothenic acid, and pyridoxine. Growth factors and salts are often derived from components of complex media such as yeast extract, molasses, and corn steep liquor. Furthermore, appropriate precursors can be added to the culture medium. The exact composition of the compounds in the medium strongly depends on each experiment and must be determined individually for each specific case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (1997). Growth media can be obtained from commercial suppliers such as Standard 1 (Merck) or BHI (Brain Heart Infusion Medium, DIFCO).
[0400] All components of the medium are sterilized by heat (20 minutes at 1.5 bar and 121 °C) or by sterile filtration. Each component can be sterilized together or separately, if necessary. All components of the medium may be present at the start of the culture or may optionally be added continuously or in batches.
[0401] The culture temperature is usually from 15 °C to 45 °C, preferably from 25 °C to 40 °C, and can be varied or kept constant during the experiment. The pH of the medium is from 5 to 8.5, preferably about 7.0. The pH during growth can be controlled by adding basic compounds such as sodium hydroxide, potassium hydroxide, ammonia or aqueous ammonia, or acidic compounds such as phosphoric acid or sulfuric acid. To control foaming, an antifoaming agent such as a fatty acid polyglycol ester can be used. To maintain the stability of the plasmid, an appropriate selective substance, such as an antibiotic, can be added to the medium. To maintain aerobic conditions, oxygen or an oxygen-containing gas mixture, such as ambient air, is supplied to the culture. The temperature of the culture is usually in the range of 20 °C to 45 °C. The culture is continued until the maximum amount of the desired product is formed. Usually, this goal is achieved within 1 hour to 160 hours.
[0402] The method of the present invention can further include a step of recovering the aforementioned terpene alcohol.
[0403] The term "recovery" includes extracting, harvesting, isolating or purifying the compound from the culture medium. The recovery of the compound can be carried out according to any conventional isolation or purification method known in the art, including but not limited to treatment with conventional resins (such as anion or cation exchange resins, nonionic adsorption resins, etc.), treatment with conventional adsorbents (such as activated carbon, silicic acid, silica gel, cellulose, alumina, etc.), pH change, solvent extraction (such as treatment with conventional solvents such as alcohol, ethyl acetate, hexane, etc.), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, freeze-drying, etc.
[0404] Prior to the intended isolation, the biomass of the broth can be removed. The process of removing biomass is known to those skilled in the art and includes, for example, filtration, sedimentation, and flotation. As a result, the biomass can be removed using, for example, a centrifuge, separator, decanter, filter, or flotation device. In order to maximize the recovery of valuable products, it is often desirable to wash the biomass, for example, in the form of diafiltration. The choice of method depends on the biomass content and characteristics in the fermentation broth, as well as the interaction between the biomass and the valuable product.
[0405] In one embodiment, the fermentation broth can be sterilized or pasteurized. In a further embodiment, the fermentation broth is concentrated. Depending on the requirements, this concentration can be carried out batchwise or continuously. The pressure and temperature ranges should be selected such that, firstly, no product damage occurs and, secondly, the use of equipment and energy is minimized. Particularly in the case of multi-stage evaporation, energy savings can be achieved by skillfully selecting the pressure and temperature levels.
[0406] The following examples are merely illustrative and are not intended to limit the scope of the embodiments described herein.
[0407] Numerous possible variations that will become immediately apparent to those skilled in the art after considering the disclosure provided herein are also included within the scope of the present invention.
[0408] Experimental section The present invention will then be described in more detail using the following examples.
[0409] a) Materials: Unless otherwise specified, all chemicals, biochemicals, and microorganisms or cells used herein are commercially available.
[0410] Recombinant proteins are cloned and expressed by standard methods, such as those described in Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular cloning: A Laboratory Manual, 2nd Edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989, unless otherwise specified.
[0411] b) General methods Preparation of cell-free protein fraction The expression vector was transformed into Escherichia coli KRX cells (Promega Corporation, Madison, Wisconsin, USA), and the transformed cells were selected on LB agar plates supplemented with the appropriate antibiotic. Subsequently, the cells were cultured in 25 mL of liquid LB medium supplemented with the appropriate antibiotic at 37 °C until the OD reached 1. Expression of the recombinant protein was induced using 1 mM isopropyl-1-thio-β-D-galactopyranoside and 0.1% (w / v) L-rhamnose monohydrate, and the cells were incubated at 25 °C for 24 h with moderate shaking.
[0412] Bacterial cells were harvested by centrifugation (5000 g, 12 min) and lysed by sonication on ice (Vibra cell X 130 sonicator equipped with a 6 mm diameter tip microprobe from Sonics; three 20 s pulses of 20 kHz at 80% of the maximum output) in 1.8 mL of 50 mM MOPSO buffer pH 7.4 containing 15% glycerol. The lysate was removed by centrifugation (3500 g, 8 min, 4 °C), and the resulting supernatant was stored frozen and used as an enzyme source for in vitro assays.
[0413] In vitro enzyme assay A protein fraction containing one of the recombinant proteins was assayed in a screw-cap tube (volume 11 mL) sealed with borosilicate glass and PTFE (Wheaton, Millville, NJ 08332, USA) consisting of 20 μL of cell-free extract, 160 - 320 mg / L of substrate (using a 40 g / L substrate stock solution in DMSO), 1 mM of coenzyme if necessary, and 50 mM of MOPSO (pH 7.4) in a final volume of 0.5 - 1 mL, and incubated at 24 °C for 4 hours with shaking at 230 rpm. The assay was extracted with 1 volume of methyl-tert-butyl-ether (MTBE) and analyzed by GC-MS as described below.
[0414] Whole-cell bioconversion assay Bioconversion of the compound was performed using Escherichia coli cells expressing the recombinant enzyme. The expression vector was transformed into E. coli KRX cells (Promega Corporation, Madison, WI, USA), and the transformed cells were selected on LB agar plates supplemented with the appropriate antibiotic. First, the cells were cultured overnight at 30 °C in 5 mL of LB medium supplemented with 1% glucose and the appropriate antibiotic. The next day, 20 mL of TB medium (Terrific Broth) supplemented with the appropriate antibiotic was inoculated to an initial optical density of 0.2 - 0.75. This culture was placed in an Erlenmeyer flask and cultured at 37 °C until the optical density reached 1 - 4, and expression of the recombinant protein was induced by adding 0.1 mM of isopropyl-1-thio-β-D-galactopyranoside IPTG and 0.1% L-rhamnose. Then, the culture was dispensed into 12 mL glass tubes at 0.5 - 1 mL aliquots and incubated at 20 °C with moderate shaking.
[0415] Ninety minutes after inducing the expression of the recombinant protein, the substrate was added to each tube. The substrate was added using a 40 g / L stock solution dissolved in DMSO to a final concentration of 0.25 - 1 g / L. Alternatively, an emulsion was prepared by dissolving 150 mg / mL of Tween® 80 (Sigma - Aldrich) and 300 mg / mL of the substrate in water, and added to the assay so that the final concentration of the substrate reached 12 mg / mL.
[0416] After 8 - 48 hours of incubation, the culture was extracted with 1 volume of MTBE and analyzed by GC - MS described below.
[0417] Culture of artificial bacterial cells under conditions enabling the production of terpene compounds DP1205 Escherichia coli cells were transformed with one or two expression plasmids carrying terpene biosynthesis genes and / or terpene - modifying enzymes, and the transformed cells were cultured on LB - agarose plates using appropriate antibiotics (kanamycin (50 μg / mL) and / or chloramphenicol (34 μg / mL)). Using a single colony, 5 mL of liquid LB medium supplemented with the same antibiotics, 4 g / L of glucose, and 10% (v / v) dodecane was inoculated. The next day, 0.2 mL of the overnight - cultured culture was inoculated into 2 mL of TB medium supplemented with the same antibiotics and 10% (v / v) dodecane. The culture was incubated at 37 °C until the optical density reached 3. Then, 1 mM IPTG was added to induce the expression of the recombinant protein, and the culture was incubated at 20 °C for 72 hours.
[0418] Then, the culture was extracted with 1 volume of (MTBE), and the composition of the organic phase was analyzed by GC - MS described below. For quantification, an internal standard substance (α - longipinene (Aldrich)) was added to the extract before GC - MS analysis, and the concentration of the components was estimated based on the comparison of peak areas.
[0419] GC - MS analysis method Samples of the whole-cell bioconversion assay were analyzed using an Agilent 7890A GC system (Agilent Technologies, California) coupled with a 5975C series mass selective detector (MSD) and equipped with a split / splitless injector.
[0420] The GC inlet temperature was set at 230 °C, 1.0 μL of the sample was injected in split mode (split ratio 20:1), and analyzed at a constant flow rate of 1 mL / min using helium as the carrier gas with a DB-5ms capillary column (30 m × 0.25 mm inner diameter × 0.25 μm film thickness, Agilent J&W). The initial oven temperature was set at 80 °C and programmed to 240 °C (10 °C / min; hold 1 min), then to 300 °C (20 °C / min; hold 1 min).
[0421] Samples of the in vitro assay were analyzed using an Agilent 6890N GC system equipped with a 5975 series mass selective detector (MSD), a split / splitless injector (Agilent Technologies, California), and a CombiPAL autosampler (CTC Analytics, Zwingen, Switzerland) injection system. The GC inlet temperature was set at 250 °C, 1.0 μL of the sample was injected in pulsed splitless mode (pulse pressure 1.56 bar, pulse time 0.6 min), and analyzed at a constant flow rate of 1.2 mL / min using helium as the carrier gas with a DB-1ms capillary column (30 m × 0.25 mm inner diameter × 0.25 μm film thickness, Agilent J&W). The initial oven temperature was 100 °C (hold 1 min) and programmed to 260 °C (10 - 20 °C / min), then to 300 °C (30 °C / min; hold 1 min). For compounds with low molecular weight, the analysis was performed under the same conditions except that the initial oven temperature was lowered to 80 °C.
[0422] Manipulation of recombinant strains that degrade terpene compounds A recombinant strain capable of producing or converting a compound was engineered by introducing a nucleotide sequence encoding one or more of the following proteins: - A Baeyer-Villiger monooxygenase (BVMO) selected from the following SCH23-BVMO1 from Hyphozyma roseonigra (SEQ ID NO: 2), SCH24-BVMO1 from Filobasidium magnum (SEQ ID NO: 6), SCH25-BVMO1 from Papiliotrema laurentii (SEQ ID NO: 10) and SCH46-BVMO1 from Bensingtonia ciliata (SEQ ID NO: 13); - An esterase selected from the following SCH23-EST from Hyphozyma roseonigra (SEQ ID NO: 20), SCH24-EST from Filobasidium magnum (SEQ ID NO: 24), SCH25-EST from Papiliotrema laurentii (SEQ ID NO: 28); and - An enal cleavage enzyme (lyase) selected from the following SCH94-3944 from Rhodococcus erythropolis (SEQ ID NO: 34), SCH80-05241 from Rhodococcus rhodochrous (SEQ ID NO: 38), Pdigit7033 from Penicillium digitatum (SEQ ID NO: 42), PitalDUF4334-1 from Penicillium italicum (SEQ ID NO: 46), AspWeDUF4334 from Aspergillus wentii (SEQ ID NO: 49), RhoagDUF4334-2 from Rhodococcus hoagii strain PAM2288 (SEQ ID NO: 53), RhoagDUF4334-3 from Rhodococcus hoagii strain N128 (SEQ ID NO: 56), RhoagDUF4334 - 4 Rhodococcus hoagii NBRC10125 (SEQ ID NO: 59), CnecaDUF4334 Cupriavidus necator (SEQ ID NO: 62), Rins - DUF4334 Ralstonia insidiosa (SEQ ID NO: 69), CgatDUF4334 Cryptococcus gattii EJB2 (SEQ ID NO: 72), GclavDUF4334 Grosmannia clavigera kw1407 (SEQ ID NO: 75), TcurvaDUF4334 Thermomonospora curvata (SEQ ID NO: 81), PprotDUF4334 Pseudomonas protegens (SEQ ID NO: 87).
[0423] Bacterial host cells for in vitro enzyme assays or whole - cell bioconversion assays were selected from Escherichia coli KRX cells (Promega Corporation, Madison, Wisconsin, USA) and Escherichia coli BL21 Star (TM) (DE3) cells (ThermoFisher).
[0424] To biochemically produce terpene compounds using one or more enzymes selected from the above - mentioned enzymes, host cells were engineered to overproduce farnesyl pyrophosphate (FPP) using the mevalonate enzyme pathway and further transformed to express sesquiterpene or diterpene biosynthetic enzymes.
[0425] Engineering of recombinant E. coli strains for FPP production by chromosomal integration of genes encoding mevalonate pathway enzymes E. coli strains were engineered to produce farnesyl pyrophosphate (FPP) by chromosomal integration of recombinant genes encoding mevalonate pathway enzymes. See also the structural scheme and recombination events depicted in Figure 1.
[0426] An upper pathway operon (operon 1 from acetyl-CoA to mevalonic acid) consisting of the atoB gene encoding Escherichia coli-derived acetoacetyl-CoA thiolase and the mbaA and mbaS genes encoding Staphylococcus aureus-derived HMG-CoA synthase and HMG-CoA reductase was designed.
[0427] As the lower mevalonic acid pathway operon (operon 2 from mevalonic acid to farnesyl pyrophosphate), a native operon from the gram-negative bacterium Streptococcus pneumoniae encoding mevalonate kinase (mvaK1), phosphomevalonate kinase (mvaK2), phosphomevalonate decarboxylase (mvaD), and isopentenyl diphosphate isomerase (fni) was selected.
[0428] To convert isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP) to FPP, a codon-optimized budding yeast FPP synthase coding gene (ERG20) was introduced at the 3' end of the upper pathway operon.
[0429] The above operons were synthesized by DNA2.0 and integrated into the araA gene of Escherichia coli strain BL21(DE3). The introduction of the heterologous pathway was performed in two recombination steps using the CRISPR / Cas9 genome editing system. The first operon to be integrated (lower pathway; operon 2) carried a spectinomycin (Spec) marker, which was used to screen for Spec-resistant candidate integrants. The second operon was designed to replace the Spec marker of the previously integrated operon, and after the second recombination event, Spec candidate integrants were screened (see Figure 1). A guide RNA expression vector targeting the araA gene was designed and synthesized by DNA2.0. To verify the integration of the operon using PCR, PCR primers were designed to amplify across the integration target of the araA gene and the recombination junction of the integrant. Subsequently, one clone that gave the correct PCR result was fully sequenced and stored as strain DP1205.
[0430] Engineering of Recombinant Bacterial Cells for Copalol Production An operon was constructed that contains two cDNAs encoding the following: - AspWeTPP (SEQ ID NO: 170) (GenBank accession number OJJ34585.1), a protein having terpene diphosphate phosphatase activity from Aspergillus wentii, which has the ability to dephosphorylate terpene diphosphate compounds such as coparyl PP; and - PvCPS (SEQ ID NO: 173) (GenBank accession number BBF88128.1), a protein having prenyltransferase and copalyl diphosphate synthase activity from Talaromyces verruculosus. PvCPS catalyzes the production of coparyl PP from IPP and DMAPP.
[0431] The cDNAs encoding AspWeTPP and PvCPS were codon-optimized (SEQ ID NOs: 171 and 174). An operon was designed that contains these two cDNAs and an RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 196) placed upstream of the cDNA. This operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, CA) to obtain plasmid pJ401-CPOL-4.
[0432] By transforming E. coli cells such as DP1205 E. coli with this plasmid pJ401-CPOL-4 and culturing under conditions that allow the production of terpene compounds, recombinant cells capable of producing copalol can be obtained.
[0433] Engineering of Recombinant Bacterial Cells for Copalol Production An operon was constructed that contains three cDNAs encoding the following: - Aspergillus wentii (SEQ ID NO: 170)-derived protein with terpenyl diphosphate phosphatase activity having the ability to dephosphorylate terpenyl diphosphate compounds such as copalyyl PP, AspWeTPP (GenBank accession number OJJ34585.1); - Alcohol dehydrogenase (ADH) activity from Azoarcus tolucasticus-derived protein having the ability to oxidize terpene alcohols such as copalol to their respective carbonyl compounds such as copalal, AzTolADH1 (SEQ ID NO: 167) (GenBank accession number WP_018990713.1); and - Protein having prenyltransferase and copalyyl diphosphate synthase activity from Talaromyces verruculosus having the ability to produce cyclic terpenyl diphosphate compounds such as copalyyl diphosphate from IPP and DMAPP, PvCPS (SEQ ID NO: 173) (GenBank accession number BBF88128.1).
[0434] The cDNAs encoding AspWeTPP, AzTolADH1 and PvCPS were codon-optimized (SEQ ID NOs: 171, 168 and 174). An operon was designed that continuously contained these three cDNAs and the RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 196) placed upstream of each cDNA. This operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California) to obtain plasmid pJ401-CPAL-1.
[0435] By transforming Escherichia coli cells such as DP1205 with this plasmid pJ401-CPOL-4, recombinant cells capable of producing copalol can be obtained when cultured under conditions that allow the production of terpene compounds.
[0436] Engineering of recombinant bacterial cells for the production of farnesal An operon was constructed that contained two cDNAs encoding the following: - TalCeTPP (SEQ ID NO: 176) (GenBank: GAM42000.1), a protein having terpenyl diphosphate phosphatase activity derived from Talaromyces cellulolyticus, which has the ability to dephosphorylate terpenyl diphosphate compounds such as farnesyl diphosphate; and - CdGeoA (NCBI accession number WP_043683915.1) (SEQ ID NO: 179), a protein having alcohol dehydrogenase (ADH) activity derived from Castellaniella defragrans, which has the ability to oxidize terpene alcohols such as farnesol to their respective carbonyl compounds such as farnesal.
[0437] The cDNAs encoding TalCeTPP and CdGeoA were codon-optimized (SEQ ID NOs: 177 and 180). An operon was designed that continuously contained the two cDNAs and the RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 196) placed upstream of each cDNA. This operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California) to obtain plasmid pJ401-FAL-1.
[0438] By transforming Escherichia coli cells such as DP1205 with this plasmid pJ401-FAL-1, recombinant cells capable of producing farnesal can be obtained when cultured under conditions allowing the production of terpene compounds.
[0439] Engineering of Recombinant Bacterial Cells for Production of Labda-diol An operon was constructed that contained three cDNAs encoding the following: - TalVeTPP (Genbank accession number KUL89334.1) (SEQ ID NO: 194), a protein having terpenyl diphosphate phosphatase activity derived from Talaromyces verruculosus, which has the ability to dephosphorylate terpenyl diphosphate compounds such as labda-diol PP, - A protein, SsLPS (Genbank accession number AET21247.1) (SEQ ID NO: 188), having labdadienol - phytyl diphosphate (LPP) synthase activity derived from Salvia sclarea, which has the ability to produce cyclic terpene diphosphate compounds such as labdadienol diphosphate from GGPP; and - CrtE, a geranylgeranyl diphosphate synthase derived from Pantoea agglomerans (GenBank accession number AAA24819.1) (SEQ ID NO: 191), which has the ability to produce GGPP from FPP.
[0440] The cDNAs encoding TalVeTPP, SsLPS, and CrtE were codon - optimized (SEQ ID NOs: 195, 189, and 192). An operon was designed that continuously contains these three cDNAs and the RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 196) placed upstream of each cDNA. This operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California) to obtain plasmid pJ401 - LOH - 2.
[0441] By transforming E. coli cells such as DP1205 E. coli with this plasmid pJ401 - LOH - 2, recombinant cells capable of producing labdendiol can be obtained when cultured under conditions allowing the production of terpene compounds.
[0442] Transformation, selection, and culture of yeast cells Transformation of all yeast cells was carried out using the lithium acetate protocol described in Gietz and Woods, Methods Enzymol., 2002, 350:87-96. The transformation mixtures were plated on SmUra- or SmLeu- medium plates containing 6.7 g / L of yeast nitrogen base without amino acids (BD Difco, New Jersey, USA), 1.92 g / L of dropout supplement without uracil (Sigma Aldrich, Missouri, USA) or 1.6 g / L of dropout supplement without leucine (Sigma Aldrich, Missouri, USA), 20 g / L of glucose and 20 g / L of agar. The plates were incubated at 30 °C for 3 - 4 days.
[0443] Engineering of yeast cells to increase the level of endogenous farnesyl diphosphate To increase the endogenous farnesyl diphosphate (FPP) pool in Saccharomyces cerevisiae cells, an extra copy of the yeast endogenous genes involved in the mevalonate pathway, from ERG10 encoding acetyl-CoA C-acetyltransferase to ERG20 encoding FPP synthase, was integrated into the genome of the Saccharomyces cerevisiae strain CEN.PK2-1C (Euroscarf, Frankfurt, Germany) under the control of a galactose-inducible promoter, in the same manner as described in Paddon et al., Nature, 2013, 496:528-532. Briefly, three cassettes were integrated into the LEU2, TRP1, and URA3 loci, respectively. The first cassette contained the genes ERG20 and truncated HMG1 (tHMG1 described in Donald et al., Proc Natl Acad Sci USA, 1997, 109:E111-8) under the control of the GAL10 / GAL1 bidirectional promoter, and the genes ERG19 and ERG13 also under the control of the GAL10 / GAL1 promoter. This cassette was flanked by two 100-nucleotide regions corresponding to the upstream and downstream regions of LEU2. The second cassette contained the genes IDI1 and tHMG1 under the control of the GAL10 / GAL1 promoter, and the gene ERG13 under the control of the promoter region of GAL7. This cassette was flanked by two 100-nucleotide regions corresponding to the upstream and downstream regions of TRP1. The third cassette contained the genes ERG10, ERG12, tHMG1, and ERG8, all under the control of the GAL10 / GAL1 promoter. This cassette was flanked by two 100-nucleotide regions corresponding to the upstream and downstream regions of URA3. All genes in the three cassettes contained 200 nucleotides of their own terminator regions. Also, an extra copy of GAL4 under the control of a mutant version of its own promoter was integrated upstream of the ERG9 promoter region, as described in Griggs and Johnston, Proc Natl Acad Sci USA, 1991, 88:8597-8601. Furthermore, the expression of ERG9 was modified by promoter exchange.The GAL7, GAL10, and GAL1 genes were deleted using a cassette containing the HIS3 gene with its own promoter and terminator. The resulting strain was mated with the haploid strain CEN.PK2-1D (Euroscarf, Frankfurt, Germany), named YST045, which was induced for sporulation according to Solis-Escalante et al., FEMS Yeast Res, 2015, 15:2. Spore isolation was achieved by resuspending the asci in 200 μL of 0.5 M sorbitol supplemented with 2 μL of zymolyase (1000 U / mL, -1 , Zymo research, Irvine, California) and incubating at 37 °C for 20 min. Subsequently, it was plated on a medium containing 20 g / L peptone, 10 g / L yeast extract, and 20 g / L agar, and one germinated spore was isolated and named YST075.
[0444] Engineering of recombinant yeast cells for copalol production For copalol production, the expression of GGPP synthase carG (from Blakeslea trispora, NCBI accession number JQ289995.1) (SEQ ID NO: 182), copalyl-diphosphate synthase SmCPS2 (from Salvia miltiorrhiza, NCBI accession number ABV57835.1) (SEQ ID NO: 185), and copalyl-diphosphate phosphatase TalVeTPP (from Talaromyces verruculosus, NCBI accession number KUL89334.1) (SEQ ID NO: 194) in different artificial yeast cells was achieved using a plasmid system constructed in vivo using yeast endogenous homologous recombination as previously described by Kuijpers et al., Microb Cell Fact., 2013, 12:47. This plasmid was composed of six DNA fragments used for the co-transformation of Saccharomyces cerevisiae. The fragments were as follows: a) Primers using plasmid pESC-LEU (Agilent Technologies, California, USA) as a template [Chemical formula] (SEQ ID NO: 124) and [Chemical formula] (SEQ ID NO: 125) The LUE2 yeast marker constructed by PCR using b) Primers using plasmid pESC-URA as a template [Chemical formula] (SEQ ID NO: 126) and [Chemical formula] (SEQ ID NO: 127) The Amp E. coli marker constructed by PCR using c) Primers using pESC-URA as a template [Chemical formula] (SEQ ID NO: 128) and [Chemical formula] (SEQ ID NO: 129) The yeast origin of replication obtained by PCR using d) Primers using pESC-URA as a template [Chemical formula] (SEQ ID NO: 130) and [Chemical formula] (SEQ ID NO: 131) The E. coli origin of replication obtained by PCR using e) The last 60 nucleotides of fragment "d", 200 nucleotides downstream of the stop codon of the yeast gene PGK1, the GGPP synthase coding sequence carG, the GAL10 / GAL1 bidirectional yeast promoter, the coding sequence of TalVeTPP, 200 nucleotides downstream of the stop codon of the yeast gene CYC1, and the sequence [Chemical formula] A fragment composed of (SEQ ID NO: 132), which was obtained by DNA synthesis (ATUM, Menlo Park, CA 94025); and f) A fragment composed of the last 60 nucleotides of fragment "e", 200 nucleotides downstream of the stop codon of the yeast gene CYC1, the SmCPS2 copalyl-pyrophosphate synthase coding sequence, the GAL10 / GAL1 bidirectional yeast promoter, and 60 nucleotides corresponding to the start of fragment "a", which was obtained by DNA synthesis (ATUM, Menlo Park, CA 94025).
[0445] Optionally, GGPP synthase carG and copalyl pyrophosphate synthase were replaced with PvCPS having two functions.
[0446] Engineering of recombinant yeast cells for the production of manool oxide To degrade copalol to manool oxide, genomic integration of the YST075 strain was performed using different alcohol dehydrogenases (ADH), Baeyer-Villiger monooxygenases (BVMO), and esterases (EST). Each integration cassette was formed by four fragments: 1) 658 bp corresponding to the upstream of the NDT80 gene and the sequence [Chemical formula] A fragment containing (SEQ ID NO: 121), which was obtained by PCR using the genomic DNA of the YST075 strain as a template; 2) The sequence [Chemical formula] (SEQ ID NO: 121), CYC1 terminator region, one of the genes encoding BVMO, intergenic region between GAL1 gene and GAL10 gene, one of the genes encoding esterase, terminator region of ADH1 gene, and sequence [Chemical formula] A fragment containing (SEQ ID NO: 122), which was obtained by DNA synthesis (ATUM, Menlo Park, CA 94025); 3) Sequence [Chemical formula] (SEQ ID NO: 122), PGK1 terminator region, one of the genes encoding alcohol dehydrogenase, promoter regions of GAL1 gene and GAL10 gene, one of the genes encoding alcohol dehydrogenase, CYC1 terminator region, and sequence [Chemical formula] A fragment containing (SEQ ID NO: 123). This fragment may contain one or two alcohol dehydrogenases depending on the experiments conducted. They were obtained by DNA synthesis (ATUM, Menlo Park, CA 94025); and 4) Sequence [Chemical formula] A fragment containing (SEQ ID NO: 123) and 405 bp corresponding to the NDT80 gene, which was obtained by PCR using the genomic DNA of YST075 strain as a template.
[0447] Engineering of recombinant yeast cells for the degradation of copalol to manoolol To degrade copalol to manool oxide, genomic integration of the YST075 strain was performed using alcohol dehydrogenase and different enal cleavage polypeptides, and each integration cassette was formed by three fragments: 1) A 658 bp fragment corresponding to the upstream region of the NDT80 gene and the sequence
Chem.
Chem.
Chem.
Chem.
Chem.
[0448] In each case, copalol production was achieved by expressing the biosynthetic pathway in a plasmid system as described above.
[0449] Engineering of Recombinant Yeast Cells for the Production of γ-Ambrinol Acetate To degrade copalol to γ-ambrinol acetate, genomic integration of the YST075 strain was performed using alcohol dehydrogenase, enal cleavage polypeptide, and different Baeyer-Villiger monooxygenases (BVMOs). Each integration cassette was formed by three fragments: 1) A 658 bp fragment corresponding to the upstream region of the NDT80 gene and the sequence
Chem.
Chem.
Chem.
Chem.
Chem.
[0450] In all cases, copalol production was achieved by expressing the biosynthetic pathway in a plasmid system as described above.
[0451] Engineering of Recombinant Yeast Cells for γ-Ambroal Production To degrade copalol to γ-ambroal, genomic integration of the YST075 strain was performed using alcohol dehydrogenase, enal cleavage polypeptide, Baeyer-Villiger monooxygenase (BVMO), and different esterases (ESTs). Each integration cassette was formed by four overlapping fragments: 1) A fragment containing at least 300 bp corresponding to the upstream region of the BUD9 gene and at least 60 bp of overlapping sequence for in vivo assembly. This fragment was obtained by PCR using the genomic DNA of the YST075 strain as a template; 2) A fragment containing the terminator region of the ADH1 gene, one of the genes encoding a tested esterase, and the intergenic region between the GAL1 gene and the GAL10 gene. This fragment was flanked by sequences enabling in vivo assembly. This fragment was obtained by DNA synthesis (ATUM, Menlo Park, CA 94025); 3) A fragment containing the URA3 yeast marker with its own promoter and terminator, flanked by sequences enabling homologous recombination. This fragment was obtained by PCR; and 4) A fragment containing at least 300 bp corresponding to the downstream region of the BUD9 gene and at least 60 bp of overlapping sequence enabling in vivo assembly. This fragment was obtained by PCR using the genomic DNA of the YST075 strain as a template.
[0452] In both cases, copalol production was achieved by expressing the biosynthetic pathway in a plasmid system as described above.
[0453] Culturing of artificial yeast cells and GC-MS analysis under conditions enabling production of terpene compounds Cells were cultured under the same conditions as described in Westfall et al., Proc Natl Acad Sci USA, 2012, 109:E111-118, and terpene and its derivatives were produced from artificial yeast cells using 10% dodecane or 10% isopropyl myristate (IPM) as an overlay of organic substances. Then, the culture was extracted with 2 volumes of MTBE, and the composition of the organic phase was analyzed by GC-MS using an Agilent 7890A GC system coupled with a 5975C series mass selective detector (MSD) and equipped with a split / splitless injector and a GC injector 80 injection system (Agilent Technologies, California). The GC inlet temperature was set at 260 °C, 1.0 μl of the sample was injected in splitless mode, and analysis was performed at a constant flow rate of 1.2 mL / min using helium as the carrier gas with an HP-5 GC column (30 m × 0.25 mm × 0.25 μm; Agilent J&W). The initial oven temperature was set at 100 °C and programmed to 300 °C (10 °C / min).
[0454] c) Examples Example 1 : In vivo conversion of manool oxide to γ-ambrinol acetate using BVMO Codon-optimized cDNAs encoding SCH23-BVMO1 (SEQ ID NO: 2) from Hyphozyma roseonigra, SCH24-BVMO1 (SEQ ID NO: 6) from Filobasidium magnum, and SCH46-BVMO1 (SEQ ID NO: 13) from Bensingtonia ciliata were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, CA) to obtain plasmids pJ414-SCH23-BVMO1, pJ414-SCH24-BVMO1, and pJ414-SCH46-BVMO1. KRX E. coli cells (Promega Corporation, Madison, WI, USA) were transformed with these expression plasmids. The transformed cells were grown and used in the above-described whole-cell bioconversion assay using manool oxide as a substrate. A negative control consisting of cells transformed with an empty plasmid was also included. Conversion of manool oxide to γ-ambrinol acetate was observed in the presence of SCH23-BVMO1, SCH24-BVMO1, or SCH46-BVMO1 recombinant protein (Figure 2). No conversion was observed in the negative control. This experiment demonstrates that SCH23-BVMO1, SCH24-BVMO1, or SCH46-BVMO1 can catalyze the following conversion: [Chemical formula]
[0455] These results demonstrate that SCH23-BVMO1, SCH24-BVMO1, and SCH46-BVMO1 catalyze the Baeyer-Villiger oxidation of manool oxide.
[0456] Example 2 In vivo conversion of copalol to compound 4 using BVMO Codon-optimized cDNAs encoding SCH23-BVMO1 (SEQ ID NO: 3) from Hyphozyma roseonigra, SCH24-BVMO1 (SEQ ID NO: 7) from Filobasidium magnum, and SCH46-BVMO1 (SEQ ID NO: 14) from Bensingtonia ciliata were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, California) to obtain plasmids pJ414-SCH23-BVMO1, pJ414-SCH24-BVMO1, and pJ414-SCH46-BVMO1. These expression plasmids were used to transform KRX E. coli cells (Promega Corporation, Madison, Wisconsin, USA). The cells were cultured and used in the above-described whole-cell bioconversion assay using a mixture of cis-coumaric acid and trans-coumaric acid as a substrate. A negative control consisting of cells transformed with an empty plasmid was also included. Conversion of cis-coumaric acid and trans-coumaric acid was observed in the presence of the recombinant proteins of SCH23-BVMO1, SCH24-BVMO1, or SCH46-BVMO1. GC-MS analysis (Figure 3) of the bioconversion products after 42 hours of incubation showed the formation of four major products, two stereoisomers 3a and 3b and two stereoisomers 4a and 4b.
[0457] In the time point measurement of bioconversion, it has been shown that compounds 1a and 1b are formed as intermediate products. Figure 4 compares the GC-MS analysis of the conversion of cis - copalal and trans - copalal by SCH23 - BVMO1 at different times. A similar evolution of the product profile is observed with SCH24 - BVMO1 and SCH46 - BVMO1. The sequential formation of these compounds indicates that trans - copalal and cis - copalal are converted to compounds 4a and 4b in several steps. Compounds 1a and 1b as well as compounds 4a and 4b are formate esters. Such functional groups can be formed from aldehyde compounds by Baeyer - Villiger monooxygenases. Therefore, the following reaction scheme, including enzymatic and non - enzymatic (chemical reactions), can be depicted to explain the conversion of trans - copalal by SCH23 - BVMO1, SCH24 - BVMO1 or SCH46 - BVMO1.
[0458] [Chemical formula]
[0459] In this scheme, the recombinant enzyme catalyzes two Baeyer - Villiger - type oxidations on two different aldehydes. First, the α,β - unsaturated aldehyde group of trans - copalal is oxidized in the first Baeyer - Villiger - type oxidation by the recombinant enzyme to form compound 1a. The enol formate functional group of compound 1a is unstable under the experimental conditions and is partially hydrolyzed to form compound 2a. Since the latter compound is rapidly converted to compound 3 (3a and 3b) via keto - enol tautomerism, it is not detected by GC - MS analysis. Compound 3 (3a and 3b) serves as a substrate for the same enzyme and catalyzes the second Baeyer - Villiger oxidation to form compound 4 (4a and 4b).
[0460] The following reaction scheme depicts a similar reaction in the conversion of cis - coparol by SCH23 - BVMO1, SCH24 - BVMO1 or SCH46 - BVMO1.
[0461] [Chemical formula]
[0462] These results indicate that SCH23 - BVMO1, SCH24 - BVMO1 and SCH46 - BVMO1 catalyze the Baeyer - Villiger oxidation of the lambda - aldehyde compound.
[0463] Example 3 : In vitro conversion of manoolol by BVMO and esterase For this experiment, the following recombinant proteins were used: SCH23 - BVMO1 from Hyphozyma roseonigra (SEQ ID NO: 2), SCH24 - BVMO1 from Filobasidium magnum (SEQ ID NO: 6), SCH23 - EST from Hyphozyma roseonigra (SEQ ID NO: 20) and SCH24 - EST from Filobasidium magnum (SEQ ID NO: 24). Condon - optimized cDNAs encoding SCH23 - BVMO1 (SEQ ID NO: 3) and SCH24 - BVMO1 (SEQ ID NO: 7) were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, California) to obtain plasmids pJ414 - SCH23 - BVMO1 and pJ414 - SCH24 - BVMO1. Condon - optimized cDNAs encoding SCH23 - EST (SEQ ID NO: 21) and SCH24 - EST (SEQ ID NO: 25) were synthesized and cloned into the pJ431 expression plasmid (ATUM, Newark, California) to obtain plasmids pJ414 - SCH23 - EST, pJ414 - SCH24 - EST.
[0464] KRX Escherichia coli cells (Promega Corporation, Madison, Wisconsin, USA) were transformed with each of these expression plasmids. The transformed cells were grown, and cell-free lysates were prepared as described. In vitro enzyme assays were performed using any of these protein fractions, or combinations of two of these protein fractions. The conditions for the in vitro assay were as described above with the addition of 160 mg / L manool oxy, 60 μM flavin adenine dinucleotide (FAD), and 500 μM reduced β-nicotinamide adenine dinucleotide phosphate (NADPH).
[0465] Using the crude fractions containing recombinant SCH23-BVMO1 and SCH24-BVMO1 proteins, the conversion of manool oxy to γ-ambrinol acetate was observed. When using the control lysate obtained from Escherichia coli cells transformed with the empty plasmid, no conversion was detected (Figure 5). From these experiments, the following enzyme reactions can be derived.
[0466] [Chemical formula]
[0467] In vitro enzyme assays were also performed using the protein fraction containing the recombinant esterase enzyme and the combination of the protein fraction containing the recombinant BVMO and the recombinant esterase enzyme. These assays were performed as described above using manool oxy as the substrate. GC-MS analysis of the formed products (Figures 6 and 7) showed the conversion of manool oxy to γ-ambril acetate in the presence of the BVMO enzyme (SCH23-BVMO1 or SCH24-BVMO1), and the further conversion of γ-ambril acetate to γ-ambrinol in the presence of the esterase enzyme (SCH23-EST or SCH24-EST) in the assay. No conversion of the substrate was observed when using the esterase in the absence of BVMO (Figures 6 and 7).
[0468] This experiment shows that in the presence of BVMO and esterase, manool oxidation is converted to γ-ambrinol according to the reaction scheme depicted below.
[0469]
Chemical formula
[0470] Example 4 : In vitro conversion of compounds 4a and 4b to compounds 5a and 5b using esterase Codon-optimized cDNAs encoding SCH23-EST (SEQ ID NO: 21) from Hyphozyma roseonigra, SCH24-EST (SEQ ID NO: 25) from Filobasidium magnum, and SCH46-EST (SEQ ID NO: 32) from Bensingtonia ciliata were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, CA) to obtain plasmids pJ414-SCH23-EST1, pJ414-SCH24-EST1, and pJ414-SCH46-EST1. KRX E. coli cells (Promega Corporation, Madison, WI, USA) were transformed with these expression plasmids. The transformed cells were grown, and cell-free lysates were prepared as described. Using these protein fractions, in vitro enzyme assays were performed according to the above conditions.
[0471] As shown in Figure 8, the conversion of the two stereoisomers 4a and 4b to compounds 5a and 5b was observed using crude fractions containing recombinant SCH23-EST1, SCH24-EST1, and SCH25-EST1 proteins. In contrast, no conversion was detected when using a lysate containing the recombinant BVMO enzyme (therefore, these proteins were used in the control reactions of this experimental series). Under these conditions, the enzyme activities of SCH23-EST1 and SCH25-EST1 were higher than that of SCH24-EST.
[0472] From these experiments, the following enzymatic reactions can be derived: [Chemical formula]
[0473] Example 5 : In vitro conversion of copalol using BVMO and esterase For this experiment, the following recombinant proteins were used: SCH23 - BVMO1 from Hyphozyma roseonigra (SEQ ID NO: 2), SCH24 - BVMO1 from Filobasidium magnum (SEQ ID NO: 6), SCH25 - BVMO1 from Papiliotrema laurentii (SEQ ID NO: 10), SCH23 - EST from Hyphozyma roseonigra (SEQ ID NO: 20), SCH24 - EST from Filobasidium magnum (SEQ ID NO: 24), SCH25 - EST from Papiliotrema laurentii (SEQ ID NO: 28).
[0474] Codon - optimized cDNAs encoding SCH23 - BVMO1 (SEQ ID NO: 3), SCH24 - BVMO1 (SEQ ID NO: 7) and SCH25 - BVMO1 (SEQ ID NO: 11) were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, California) to obtain plasmids pJ414 - SCH23 - BVMO1, pJ414 - SCH24 - BVMO1 and pJ414 - SCH25 - BVMO1. Codon - optimized cDNAs encoding SCH23 - EST (SEQ ID NO: 21), SCH24 - EST (SEQ ID NO: 25) and SCH25 - EST (SEQ ID NO: 29) were synthesized and cloned into the pJ431 expression plasmid (ATUM, Newark, California) to obtain plasmids pJ414 - SCH23 - EST, pJ414 - SCH25 - EST.
[0475] Escherichia coli cells (Promega Corporation, Madison, Wisconsin, USA) were transformed with these expression plasmids. The transformed cells were grown, and cell-free lysates were prepared as described. In vitro enzyme assays were performed using protein fractions containing recombinant BVMO enzyme or recombinant esterase enzyme, or by combining protein fractions containing recombinant BVMO enzyme and esterase enzyme. The assays were performed as described above with the addition of a mixture of 320 mg / L of cis - copalol and trans - copalol as substrate, 60 μM of flavin adenine dinucleotide (FAD), and 500 μM of reduced β - nicotinamide adenine dinucleotide phosphate (NADPH).
[0476] Figure 9 compares the products resulting from the conversion of copalol in the presence of SCH23 - BVMO1 alone and in combination with different esterase enzymes. In the presence of SCH23 - BVMO1, the major products are the formate compounds 1a, 1b and 4a, 4b. When the assay was performed in the additional presence of SCH23 - EST or SCH25 - EST, the major products of the conversion are compounds 5a and 5b, indicating that these two esterase enzymes can efficiently hydrolyze the formate intermediates produced by the BVMO enzyme. In the additional presence of SCH24 - EST, hydrolysis of the same intermediates (1a, 1b and 4a, 4b) was observed, but with this enzyme, hydrolysis of intermediates 1a and 1b by SCH24 - EST seemed to be more efficient than hydrolysis of intermediates 4a and 4b.
[0477] Similar conversions of cis and trans - copalol were observed when SCH24 - BVMO1 was combined with esterases SCH23 - EST or SCH24 - EST (Figure 10). In control experiments, no conversion was observed when copalol was incubated with esterase alone.
[0478] From these experiments, the following enzymatic pathway can be inferred.
[0479] [Chem.]
[0480] Example 6 : In Vivo Production of 14,15-Dinor-labdane Compounds 5a and 5b and Biosynthetic Intermediates in Artificial Bacterial Cells Expressing BVMO and Esterase In this experiment, Escherichia coli cells were transformed with plasmid pJ401-CPAL-1 (described above) and copalal was produced as described in the experimental section. When DP1205 Escherichia coli was transformed and cultured under the conditions described in the experimental section, the formation of trans-copalal and cis-copalal was confirmed (Figure 11, upper chromatogram). The detection of two double bond isomers of copalal is due to the relatively easy isomerization of (E)-α,β-unsaturated aldehyde (Konning et al, Org. Lett., 2012, 14 (20), pp 5258-5261). The additional detection of labda-8(20)-en-15-ol is due to the endogenous enoate reductase activity of Escherichia coli.
[0481] Next, bacterial cells were transformed with a second expression plasmid carrying a codon-optimized cDNA encoding SCH24-BVMO1 (SEQ ID NO: 7) from Filobasidium magnum (ATCC® 20918™) or SCH46-BVMO1 (SEQ ID NO: 14) from Bensingtonia ciliata. These plasmids were prepared by cloning the optimized cDNA into the pJ423 expression plasmid (ATUM, Newark, CA), yielding plasmids pJ423-SCH23-BVMO and pJ423-SCH46-BVMO, respectively. Cells transformed with the two plasmids were cultured, and the production levels of terpene compounds and terpene derivatives were analyzed using the conditions described in the Experimental section. Under these conditions, compounds 1a and 1b, 3a and 3b, and 4a and 4b were detected in the solvent extract of the culture broth (Figure 11). These results indicate that the combination of these enzymes can introduce into recombinant cells the biosynthesis of labdane diterpenes such as copalol and the sequential enzymatic cleavage of two carbon-carbon bonds in the side chain.
[0482] Similarly, bacterial cells were co-transformed using plasmid pJ401-CPAL-1 and a second plasmid having a gene encoding BVMO and a gene encoding an esterase: pJ423-SCH24-BVMO-SCH24-EST prepared by inserting a synthetic operon composed of a codon-optimized cDNA encoding SCH24-BVMO1 (SEQ ID NO: 7) and a codon-optimized cDNA encoding SCH24-EST (SEQ ID NO: 25) into a pJ423 expression plasmid (ATUM, Newark, CA), or pJ423-SCH46-BVMO-SCH46-EST, which is a plasmid prepared by inserting a synthetic operon composed of a codon-optimized cDNA encoding SCH46-BVMO (SEQ ID NO: 14) and a codon-optimized cDNA encoding SCH46-EST (SEQ ID NO: 32) into a pJ423 expression plasmid (ATUM, Newark, CA). These cells were cultured, and the production amounts of terpene compounds and terpene derivatives were analyzed under the conditions described in the experimental section. Under these conditions, compounds 5a and 5b were detected, and it was confirmed that the amounts of pathway intermediates (compounds 1a, 1b, 3a, 3b, 4a, and 4b) were decreased.
[0483] This series of experiments shows that the following biosynthetic pathway can be introduced into host cells transformed to express diterpene biosynthetic enzymes in combination with BVMO and an esterase.
[0484]
Chemical formula
[0485] Example 7 : In vivo conversion of compounds 5a and 5b to manool oxide using alcohol dehydrogenase.
[0486] For this experiment, the following alcohol dehydrogenases were evaluated for the oxidation of compounds 5a and 5b to manool oxide: RrhSecADH from Rhodococcus rhodochrous (SEQ ID NO: 146), SCH80-00043 (SEQ ID NO: 149) derived from Rhodococcus rhodochrous, SCH80-04254 (SEQ ID NO: 152) derived from Rhodococcus rhodochrous, SCH80-06135 (SEQ ID NO: 155) derived from Rhodococcus rhodochrous, SCH80-06582 (SEQ ID NO: 158) derived from Rhodococcus rhodochrous. (see also WO 2005 / 026338; the above ADHs are merely non-limiting examples and may be replaced by other known ADHs).
[0487] Codon-optimized cDNAs encoding each of these proteins were synthesized, cloned into the vector pJ401, and plasmids pJ401-RrhSecADH, pJ401-SCH80-00043, pJ401-SCH80-04254, pJ401-SCH80-06135, and pJ401-SCH80-06582 (ATUM, Newark, California) were obtained.
[0488] KRX E. coli cells (Promega Corporation, Madison, Wisconsin, USA) were transformed with these expression plasmids. The transformed cells were cultured and used in the above-described whole-cell bioconversion assay using a mixture of compounds 5a and 5b as a substrate. Five hours after induction of recombinant protein expression, an emulsion containing 50 mg / mL of tween80 and 25 mg / mL of the substrate in water was used to add the substrate to a final concentration of 0.55 mg / mL. A negative control consisting of cells transformed with an empty plasmid was also included. Oxidation reactions were observed only in the presence of the recombinant proteins of SCH80-06135 and RrhSecADH (Figure 12), indicating that these enzymes can catalyze the following reactions.
[0489]
Chemical formula
[0490] Example 8 : In vivo production of γ-ambrol, a tetranor-labdane compound, and biosynthetic intermediates in an artificial bacterial cell expressing BVMO, esterase, and alcohol dehydrogenase.
[0491] In this experiment, Escherichia coli cells DP1205, which were used to create a background strain producing copalol (cis and trans-isomers) as described above using plasmid pJ401-CPAL-1 (described previously), were transformed.
[0492] Subsequently, this strain was co-transformed with plasmid pJ423-SCH24-BVMO-SCH24-EST (described previously) to further express BVMO and esterase in the same cells. This recombinant organism produces 14,15-dinor-labdane compounds, similar to the observation results obtained in the previous section.
[0493] To continue the degradation of the side chain until the formation of the tetranor-λ-lactone derivative, the secondary alcohol groups of compounds 5a and 5b must be oxidized to the corresponding ketones. Therefore, plasmids containing nucleotide sequences encoding BVMO, esterase, and the appropriate alcohol dehydrogenase (identified in Example 7) were constructed. For the alcohol dehydrogenase, a codon-optimized cDNA encoding RrhSecADH from Rhodococcus sp. (accession number WP_043801412.1) (SEQ ID NO: 147) was synthesized, and a synthetic operon combining the RrhSecADH cDNA with the cDNAs encoding SCH24-BVMO and SCH24-EST was designed. This operon was cloned into the pJ423 expression plasmid to obtain the pJ423-secADH-23BVMO-EST plasmid. When Escherichia coli DP1205 co-transformed with the vector pJ401-CPAL-1 and the vector pJ423-secADH-23BVMO-EST was cultured under the above conditions, γ-ambrinol was detected by GC-MS analysis of the culture broth (Figure 13). These data indicate that when compound 5 (5a and 5b) is oxidized to manooxy in the presence of the appropriate ADH, BVMO can catalyze the following pathway steps to provide γ-ambrinol.
[0494] This series of experiments demonstrates that the following biosynthetic pathway can be introduced into recombinant host cells.
[0495]
Chemical Structure
[0496] Example 9 : In vivo production of manooxy in budding yeast cells using alcohol dehydrogenase (ADH), Baeyer-Villiger monooxygenase (BVMO), and esterase (EST) from Hyphozyma roseonigra or Cryptococcus albidus For the production of manool oxide, genes encoding GGPP synthase carG (from Blakesleatrispora, NCBI accession number JQ289995.1) (SEQ ID NO: 182), copalyl-pyrophosphate synthase SmCPS2 (from Salvia miltiorrhiza, NCBI accession number ABV57835.1) (SEQ ID NO: 185), copalyl-pyrophosphate phosphatase TalVeTPP (from Talaromycesverruculosus, NCBI accession number KUL89334.1) (SEQ ID NO: 194), and either alcohol dehydrogenase SCH23-ADH1 (SEQ ID NO: 134), Baeyer-Villiger monooxygenase SCH23-BVMO1 (SEQ ID NO: 2), esterase SCH23-EST (SEQ ID NO: 20) and alcohol dehydrogenase SCH23-ADH2 (from Hyphozyma roseonigra) (SEQ ID NO: 137) or alcohol dehydrogenase SCH24-ADH1 (SEQ ID NO: 140), Baeyer-Villiger monooxygenase SCH24-BVMO1 (SEQ ID NO: 6), esterase SCH24-EST1 (SEQ ID NO: 24) and alcohol dehydrogenase SCH24-ADH2 (from Filobasidium magnum) (SEQ ID NO: 143) were expressed in the artificial budding yeast strain YST075 as described in the section on the general method above. All genes were codon-optimized for their expression in budding yeast (SCH23-ADH1, SEQ ID NO: 135; SCH23-BVMO1, SEQ ID NO: 4; SCH23-EST, SEQ ID NO: 22; SCH23-ADH2, SEQ ID NO: 138; SCH24-ADH1, SEQ ID NO: 141; SCH24-BVMO1, SEQ ID NO: 8; SCH24-EST, SEQ ID NO: 26; SCH24-ADH2, SEQ ID NO: 144; carG, SEQ ID NO: 183; SmCPS2, SEQ ID NO: 186; and TalVeTPP, SEQ ID NO: 195).
[0497] YST120 (including SCH23-ADH1, SCH23-BVMO1, SCH23-EST and SCH23-ADH2) and YST121 (including SCH24-ADH1a, SCH24-BVMO1, SCH24-EST and SCH24-ADH2), which also possess plasmid systems for copalol biosynthesis, were obtained and cultured under the conditions described in the section of the general method above.
[0498] Under these conditions, copalol was identified in all cultures. Only the strains containing SCH23-ADH1 or SCH24-ADH1 were able to convert copalol to copalarol (Figure 14A). Furthermore, farnesal was detected in the cultures expressing alcohol dehydrogenase (Figure 14B). In all cultures, the accumulation of nerolidol and farnesol was identified (Figure 14A).
[0499] Furthermore, manool oxide was identified in the cultures containing YST120 and YST121 strains that possess plasmids with copalol biosynthesis genes (Figure 14C). Neither γ-ambrinol acetate nor γ-ambrinol was identified. However, the presence of manool oxide suggests that BVMO, EST and ADH are functionally expressed in the artificial yeast cells. The inventors speculate that the amount of manool oxide obtained may have been limited for BVMO to catalyze the conversion to γ-ambrinol acetate.
[0500] Example 10 : In vivo production of manool oxide in budding yeast cells using alcohol dehydrogenase (ADH), Baeyer-Villiger monooxygenase (BVMO) and esterase (EST) from Hyphozyma roseonigra or Cryptococcus albidus For the production of manool oxide, the genes encoding GGPP synthase carG (from Blakeslea trispora, NCBI accession number JQ289995.1), copalyl-pyrophosphate synthase SmCPS2 (from Salvia miltiorrhiza, NCBI accession number ABV57835.1), copalyl-pyrophosphate phosphatase TalVeTPP (from Talaromyces verruculosus, NCBI accession number KUL89334.1), alcohol dehydrogenase SCH23-ADH1, and either Baeyer-Villiger monooxygenase SCH23-BVMO1 and esterase SCH23-EST (from Hyphozyma roseonigra) or Baeyer-Villiger monooxygenase SCH24-BVMO1 and esterase SCH24-EST (from Cryptococcus albidus) were expressed in the artificial budding yeast strain YST075 as described in the section on general methods.
[0501] The resulting strains were named YST177 (containing carG, SmCPS2, TalVeTPP, SCH23-ADH1, SCH23-BVMO1, and SCH23-EST) and YST178 (containing carG, SmCPS2, TalVeTPP, SCH23-ADH1, SCH24-BVMO1, and SCH24-EST) and cultured as described in the section on general methods above. The cultures were analyzed by GC-MS as described above.
[0502] Coparol, copalal, nerolidol, farnesol and farnesal were identified in the culture after extraction. Artificial cells lacking alcohol dehydrogenase SCH23-ADH2 or SCH24-ADH2 were expected to accumulate intermediate 5a (or 5b) and be unable to produce manool oxide. Interestingly, manool oxide was identified (Figure 15) and molecule 5a (or 5b) was not detected. These results suggest that SCH23-ADH2 and SCH24-ADH2 may contribute to the production of manool oxide in yeast cells, but are not essential for the production of manool oxide under the conditions tested. The inventors speculate that the endogenous alcohol dehydrogenase activity of yeast is involved in this conversion.
[0503] Example 11 : Characterization of the enzyme SCH94-3944 from Rhodococcus erytheropolis with carbon-carbon bond cleavage activity In this experiment, DP1205 E. coli cells that produce a background strain producing coparol were transformed using plasmid pJ401-CPOL-4 (described above). This transformant produced coparol at a concentration of up to 500 mg / L in the culture medium as the main product in a tube assay (Figure 16).
[0504] This strain was then further transformed with a second plasmid carrying one or more E. coli codon-optimized cDNAs from R. erytheropolis. Two cDNAs were selected: - SCH94-3945 (SEQ ID NO: 161) encoding a putative alcohol dehydrogenase, - SCH94-3944 (SEQ ID NO: 34) encoding a 157-amino acid protein containing two protein family domains: the "GXWXG" protein domain (pfam14231, http: / / pfam.xfam.org / ) and the domain of unknown function "DUF4334" (pfam14232, http: / / pfam.xfam.org / ).
[0505] The expression vector was prepared using pJ423 as a background and containing either the codon-optimized cDNA encoding SCH94-3945 (pJ423-SCH94-3945), or the codon-optimized cDNA encoding SCH94-3944 (pJ423-SCH94-3944), or a bicistronic operon (pJ423-SCH94-3944-3945) composed of the optimized cDNA encoding SCH94-3945 and SCH94-3944.
[0506] When cells were transformed with vector pJ401-CPOL-4 and vector pJ423-SCH94-3944, no difference was observed compared to cells transformed with pJ401-CPOL-4 alone, indicating that the SCH94-3944 recombinant protein does not convert copalol. When cells were transformed with vector pJ401-CPOL-4 and vector pJ423-SCH94-3945, the formation of cis-copalol and trans-copalol was observed, indicating that SCH94-3945 is an alcohol dehydrogenase capable of oxidizing copalol to copalarol (Figure 16).
[0507] When cells were transformed with vector pJ401-CPOL-4 and vector pJ423-SCH94-3944-3945, in the tube assay, the formation of manool oxide was observed as the major product at concentrations up to 1 g / L in the culture. Under these assay conditions, the conversion of cis and trans-copalol was almost complete (Figure 16).
[0508] This experiment shows that the enzyme of SCH94-3944 can cleave the α-β carbon-carbon double bond of copalarol and catalyze the direct conversion to manool oxide, a 14,15-dinor-labdane compound of cis-copalol and trans-copalol, as shown in the following scheme.
[0509]
Chemical formula
[0510] Example 12 :In vivo conversion of cis- and trans-farnesal using an enal cleavage polypeptide from Rhodococcus erythropolis In this experiment, Escherichia coli cells DP1205 were transformed with plasmid pJ401-FAL-1 (described above) to generate a background strain that produces cis-farnesal and trans-farnesal as major products at a concentration of up to 500 mg / L in the culture under tube assay conditions (Figure 17).
[0511] This strain was then further transformed with plasmid pJ423-SCH94-3944 carrying the cDNA encoding SCH94-3944 from R. erytheropolis. Analysis of the compounds produced by these cells by GC-MS showed the formation of geranylacetone (Figure 17). Thus, this experiment shows that the SCH94-3944 enzyme can cleave the α-β carbon-carbon double bond of the acyclic compound farnesal and catalyze the direct conversion of cis- and trans-farnesal to geranylacetone, as shown in the following scheme.
[0512]
Chemical formula
[0513] Under the applied test conditions, no conversion with farnesol was observed.
[0514] Example 13 .In vivo conversion of citral using an enal cleavage polypeptide from Rhodococcus erythropolis Using Escherichia coli KRX (Promega) cells transformed with plasmid pJ423-SCH94-3944, biochemical conversion of the compound was performed, thereby overexpressing the recombinant protein of SCH94-3944. Using a 2:1 substrate:Tween 80 emulsion, the substrate was added to the cell culture to a final concentration of 12 g / L. Bioconversion was performed as described in the Experimental section. The negative control was performed using cells transformed with the pJ423 expression plasmid without insert. Several substrates were tested: citral (a mixture composed of geranial and neral), citronellal (2,3-dihydrocitral) and (E)-2-dodecanal. The cells were incubated for 24 h in the presence of various compounds and the products of the conversion were analyzed as described in the Experimental section.
[0515] In the presence of the SCH94-3944 recombinant protein, both geranial and neral were converted to methylheptenone (Figure 18), indicating that this enzyme can cleave the α-β carbon-carbon double bond of acyclic monoterpene aldehydes as shown in the following scheme.
[0516] [Chemical formula]
[0517] Formula [Chemical formula] For citronellal of the formula, no conversion was obtained in the presence of the SCH94-3944 recombinant protein (Figure 18), indicating that unsaturation of the α,β-carbon bond is required for catalysis.
[0518] (E)-2-dodecanal [Chemical formula] When [substance] was used, conversion to decanal was observed. However, when compared with citral, the conversion yield was significantly lower (Figure 18). This observation suggests that the absence of the 3-methyl group has an adverse effect on the enzymatic conversion by the SCH94-3944 protein.
[0519] Example 14 : In vivo conversion of copalal and farnesal using GXWXG- and DUF4334 domain-containing proteins from other organisms The SCH94-3944 protein sequence contains the GXWXG protein family domain and the DUF4334 protein family domain. Proteins with similar domain structures were searched in other organisms to test whether the enzymatic activity related to SCH94-3944 could also be related to these homologous enzymes.
[0520] In this experiment, DP1205 Escherichia coli cells were transformed using plasmid pJ401-CPAL-1 (described above) to generate a background strain that produces copalal (cis and trans isomers) as described in the previous section. In this strain, the FPP synthase is expressed from a genome-integrated operon. Since the terpene phosphatase AspWeTPP can dephosphorylate not only GGPP but also FPP, and AzeTolADH1 can oxidize farnesol, when DP1205 cells were transformed with pJ401-CPAL-1, a large amount of trans-farnesol was detected in addition to copalal (Figure 19).
[0521] Next, this strain was co-transformed with a second plasmid carrying a gene encoding a protein containing the GXWXG protein family domain and the DUF4334 protein family domain. Several proteins were selected: - SCH80-05241 (SEQ ID NO: 38) from Rhodococcus rhodochrous ((registered trademark) ATCC 12674 (trademark)), - Pdigit7033 (SEQ ID NO: 42) from Penicillium digitatum, - PitalDUF3443-1 derived from Penicillium italicum (SEQ ID NO: 46), - AspWeDUF3443 derived from Aspergillus wentii (SEQ ID NO: 49), - RhoagDUF4334-2 derived from Rhodococcus hoagii strain PAM2288 (SEQ ID NO: 53), - RhoagDUF4334-3 derived from Rhodococcus hoagii strain N128 (SEQ ID NO: 56), - RhoagDUF4334-4 derived from Rhodococcus hoagii NBRC 10125 (SEQ ID NO: 59), - CnecaDUF4334 derived from Cupriavidus necator (SEQ ID NO: 62), - Rins-DUF4334 derived from Ralstonia insidiosa (SEQ ID NO: 69), - CgatDUF4334 derived from Cryptococcus gattii EJB2 (SEQ ID NO: 72), - GclavDUF4334 derived from Grosmannia clavigera kw1407 (SEQ ID NO: 75), - TcurvaDUF4334 derived from Thermomonospora curvata (SEQ ID NO: 81), and - PprotDUF4334 derived from Pseudomonas protegens (SEQ ID NO: 87).
[0522] Codon-optimized cDNAs encoding each of these proteins were designed and cloned into the pJ423 expression plasmid (ATUM, Newark, CA). DP1205 E. coli cells were co-transformed with one of these plasmids and the plasmid pJ401-CPAL-1. Figures 20 and 21 show the conversion of cis-copalal and trans-copalal to manool oxide in the presence of each of the recombinant proteins containing the GXWXG domain and the DUF4334 domain. Under these assay conditions, the conversion of copalal was almost complete for each recombinant enzyme except for the GclavDUF4334 enzyme, in which only minor conversion was observed. Figures 22 and 23 show the conversion of cis-farnesal and trans-farnesal to geranylacetone. The conversion of farnesal was also complete for each enzyme except for GclavDUF4334, in which only about 50% of farnesal was converted.
[0523] This experiment shows that a protein containing a GXWXG protein family domain in the N-terminal region and a DUF4334 protein family domain in the C-terminal region can catalyze enal cleavage activity against copalal and farnesal as shown in the following scheme.
[0524] [Chemical formula]
[0525] Example 15 : Variant of SCH94-3944 with a single amino acid modification Alignment of the amino acid sequences of GXWXG- and DUF4334-domain-containing proteins with enal cleavage activity showed the amino acids conserved along the amino acid sequence and within the two aforementioned protein domains (Figure 24). Conserved residues of protein families are often important for enzyme activity.
[0526] To evaluate the involvement of conserved residues of an enzyme containing a GXWXG domain and a DUF4334 domain in enzyme activity, artificial mutants of the SCH94-3944 protein with the conserved residues individually substituted with alanine residues were designed. The following residues were mutated: W44, T51, H53, L59, W64, K67, S71, R106, Y115, D116, D122, M136, K139, F152, L154, and R156. The modified proteins were designated SCH94-3944-W44A, SCH94-3944-T51A, SCH94-3944-H53A, SCH94-3944-L59A, SCH94-3944-W64A, SCH94-3944-K67A, SCH94-3944-S71A, SCH94-3944-R106A, SCH94-3944-Y115A, SCH94-3944-D116A, SCH94-3944-D122A, SCH94-3944-M136A, SCH94-3944-K139A, SCH94-3944-F152A, SCH94-3944-L154A, and SCH94-3944-R156A.
[0527] Codon-optimized cDNAs encoding each of these proteins were designed and cloned into the pJ423 expression plasmid (ATUM, Newark, CA). DP1205 E. coli cells were co-transformed with one of these plasmids and the plasmid pJ401-CPAL-1. In the presence of the recombinant proteins SCH94-3944-W44A, SCH94-3944-K67A, SCH94-3944-D122A, SCH94-3944-F152A or SCH94-3944-L154A, no conversion of copalal and farnesal was observed. For the enzymes SCH94-3944-T51A, SCH94-3944-H53A, SCH94-3944-L59A, SCH94-3944-W64A, SCH94-3944-S71, SCH94-3944-R106A, SCH94-3944-Y115A, SCH94-3944-D116A, SCH94-3944-M136A, SCH94-3944-K139A and SCH94-3944-R156A, conversion of copalal and farnesal was observed, but at a lower efficiency than the wild-type SCH94-3944 protein. Figure 25 shows the activity of each single amino acid variant enzyme against wild-type SCH94-3944.
[0528] Example 16 : In vivo production of γ-ambrinol acetate by the combination of enal cleavage activity and BVMO activity in E. coli cells In this experiment, DP1205 E. coli cells were transformed with the plasmid pJ401-CPAL-1 (described above) to generate a background strain that produces copalal (cis and trans-isomers) as described above.
[0529] This strain was then co-transformed with a second plasmid carrying a codon-optimized nucleotide sequence encoding either an enzyme with enal cleavage activity or an enzyme with BVMO activity, or a second vector carrying an operon composed of a codon-optimized cDNA encoding an enal cleavage polypeptide and a codon-optimized cDNA encoding BVMO: - pJ423-AspWeBVMO, which contains an optimized DNA sequence encoding AspWeBVMO (SEQ ID NO: 17); - pJ423-SCH94-3944, which contains an optimized DNA sequence encoding SCH94-3944 (SEQ ID NO: 35); - pJ423-SCH94-3944-SCH23-BVMO, which contains optimized DNA sequences encoding SCH94-3944 and SCH23-BVMO1 (SEQ ID NOs: 35 and 3); - pJ423-SCH94-3944-SC...
Claims
1. 1. An isolated polypeptide having enal cleavage activity, comprising: (1) A group of polypeptides including: a) at least one DUF4334 protein family domain having the Pfam ID number PF14232; and / or b) at least one GXWXG protein family domain having the Pfam ID number PF14231; or c) at least one domain that retains at least 90% sequence identity with PF14232 or PF14231 and / or (2) A group of polypeptides, each polypeptide containing at least one sequence motif / domain selected from the following: G-[Y or -]-x-W-x-G-x-x-[F, L or I]-x-[T, S or R]-G-[H or D] shown in SEQ ID NO: 205; W-[Y, A or V]-G-K-x-[F or Y]-x-[S or D] shown in SEQ ID NO: 206; [G or S]-x-[A or G]-x-[L or V]-x-x-x-x-[F,Y or L]-R-G-x-V as shown in SEQ ID NO: 207; [M or L]-[V or I]-Y-D-xx-x-P-[I or V]-x-D-[H or S]-[F or L] shown in SEQ ID NO: 208; where the residues x, independently of each other, represent any naturally occurring amino acid residue. and / or (3) A group of polypeptides comprising an amino acid sequence selected from the following: a) SCH94-3944 as depicted in SEQ ID NO: 34; b) SCH80-05241 as set forth in SEQ ID NO: 38; c) Pdigit7033 as set forth in SEQ ID NO:42; d) PitalDUF4334-1 as set forth in SEQ ID NO: 46; e) AspWeDUF4334 as set forth in SEQ ID NO: 49; f) RhoagDUF4334-2 as set forth in SEQ ID NO:53; g) RhoagDUF4334-3 as set forth in SEQ ID NO:56; h) RhoagDUF4334-4 as set forth in SEQ ID NO:59; i) CnecaDUF4334 as set forth in SEQ ID NO:62; j) Rins-DUF4334 as set forth in SEQ ID NO: 69; k) CgatDUF4334 as set forth in SEQ ID NO: 72; l) GclavDUF4334 as set forth in SEQ ID NO: 75; m) TcurvaDUF4334 as set forth in SEQ ID NO: 81; n) PprotDUF4334 as set forth in SEQ ID NO: 87; and o) a polypeptide comprising an amino acid sequence having at least 40% sequence identity with any one of the amino acid sequences of a) to n) and retaining the enzyme activity of degrading a terpene precursor of the following formula (V):
2. An isolated polypeptide selected from the group consisting of
2. 2. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding the polypeptide of claim 1.
3. General formula IV 【Chemistry 1】 [In the formula, R 1 represents H or lower alkyl; R 2 represents H, a linear or branched, saturated or unsaturated, optionally substituted hydrocarbyl residue, or the residue Cyc-A-, Where: Cyc represents an optionally substituted, saturated or unsaturated, monocyclic or polycyclic hydrocarbyl residue, A represents a chemical bond or an optionally substituted linear or branched alkylene bridge; R 3 are each independently H or lower alkyl, the method comprising the steps of: (1) General formula V 【Chemistry 2】 [In the formula, R 1 , R 2 and R 3 is as defined above, R 4 represents H or lower alkyl; R 5 represents H or lower alkyl. the corresponding undecomposed precursor of contacting a polypeptide having enal cleavage activity, wherein said compound of formula V may be present in essentially stereoisomerically pure form or as a mixture of stereoisomers; (2) optionally isolating said degradation product of formula IV obtained in step (1), wherein said compound of formula IV is obtained in stereoisomerically pure form or as a mixture of stereoisomers; A method comprising:
4. Applying a terpene precursor of formula V, Where: R 1 represents H or methyl, R 2 is H or an acyclic, linear or branched, saturated or unsaturated hydrocarbyl residue having 1 to 20 carbon atoms; or The cyclic group Cyc-A-, Where: A is a linear or branched C 1 ~C 4 -alkylene bridge, Cyc is a monocyclic or polycyclic, saturated or unsaturated hydrocarbyl residue, optionally substituted with 1 to 10 substituents, the substituents being C 1 ~C 4 -Alkyl, C 1 ~C 4 -Alkylidene, C 2 ~C 4 represents a hydrocarbyl residue independently selected from alkenyl, oxo, hydroxy, or amino, Each R 3 represents H R 4 represents H or methyl, R 5 represents H or methyl, The method of claim 3.
5. said compound of general formula V has a labdan type structure and / or Cyc-A has the formula IIIa, IIIb or IIIc 【Chemistry 3】 The method according to claim 3 or 4, wherein the amino acid sequence represents one of the residues:
6. 6. The method according to any one of claims 3 to 5, wherein the polypeptide having enal cleavage activity is selected from the polypeptides defined in claim 1.
7. 7. The method according to any one of claims 3 to 6, further comprising, as step (3), treating the compound of formula IV formed in step (1) or isolated in step (2) to obtain a derivative thereof using chemical synthesis or biocatalytic synthesis or a combination of both, and optionally, as step (4), isolating the derivative of step (3), wherein the derivative may in particular be selected from hydrocarbons, alcohols, diols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides and / or esters.
8. 8. The method of claim 7, wherein step (3) further comprises treating the compound of formula IV formed in step (1) or isolated in step (2) with a polypeptide having Baeyer-Villiger monooxygenase (BVMO) activity to form the respective carbonyl ester (EC.1.13.14.-), optionally hydrolyzing the carbonyl ester compound with an esterase (EC 3.1.1) to obtain the corresponding de-esterified product, and optionally isolating the derivative of step (3).
9. 1. A biocatalytic process for preparing an ester compound, comprising the steps of: (1) General formula I 【Chemistry 4】 [In the formula, "a" represents a single or double bond; When "a" represents a double bond, "x" is 1, or when "a" represents a single bond, "x" is 2; R 1 represent, independently of each other, H or lower alkyl, R 2 represents H, a linear or branched, saturated or unsaturated, optionally substituted hydrocarbyl residue, or the group Cyc-A-, Where: Cyc represents an optionally substituted, saturated or unsaturated, monocyclic or polycyclic hydrocarbyl residue, A represents a chemical bond or an optionally substituted linear or branched alkylene bridge; R 3 are each independently H or C 1 ~C 15 or a lower alkyl group, When "a" represents a single bond, Z represents a hydrocarbyl residue containing a carbonyl group; When "a" represents a double bond, Z together with the carbon atom to which it is attached forms a carbonyl group, or an alkylidene residue having a terminal carbonyl group, or When "a" represents a double bond and Z, together with the carbon atom to which it is attached, forms a carbonyl group, R 2 and R 1 may together with the carbon atom to which they are attached form a cyclic, saturated or unsaturated, optionally substituted carbocyclic ring group; wherein said carbonyl compounds of general formula I are provided in stereoisomerically pure form or as a mixture of stereoisomers, with a polypeptide having Baeyer-Villiger monooxygenase (BVMO) (EC 1.13.14.-) activity to form the respective carbonyl ester; (2) optionally isolating the carbonyl ester formed in step (1), wherein the carbonyl ester compound is obtained in stereoisomerically pure form or as a mixture of stereoisomers. A method comprising:
10. In said carbonyl compounds of general formula I, "a" represents a chemical double bond, Z is =O or =C(R 4 )-C(R 5 )=O, or "a" represents a single chemical bond and Z is -C(R 5 )=O, Where: R 4 and R 5 10. The method of claim 9, wherein, independently of each other, represent H or lower alkyl.
11. 11. The method according to claim 9 or 10, wherein said carbonyl compound of general formula I has a labdan type structure.
12. R 2 represents the group Cyc-A-, where A is a linear or branched C 1 ~C 4 -alkylene bridge, Cyc represents a monocyclic or polycyclic, saturated or unsaturated hydrocarbyl residue, where Cyc is optionally substituted with 1 to 10 substituents, said substituents being C 1 ~C 4 -Alkyl, C 1 ~C 4 -Alkylidene, C 2 ~C 4 12. The method of any one of claims 9 to 11, wherein said alkyl radicals are independently selected from: alkenyl, oxo, hydroxy, or amino.
13. Cyc-A is represented by the formula IIIa, IIIb or IIIc 【Chemistry 5】 The method of claim 12, wherein the bicyclic residue represents one of:
14. The polypeptide having BVMO activity is (1) a group of polypeptides that contain within their amino acid sequence a flavin-containing monooxygenase (FMO) protein family domain having the Pfam ID number PF00743, or a domain that retains at least 90% sequence identity with PF00743; and / or (2) A group of polypeptides comprising at least one sequence motif / domain selected from the following: GAGxSGL as set forth in SEQ ID NO: 197; EKNxxxxGTWxENRYPGCACDVPxHxYXXSFE as set forth in SEQ ID NO:198; LxNAxGILNxWxxPxIPG as set forth in SEQ ID NO: 199; LxxKxVxxIGxGSSGIQIxPxI as set forth in SEQ ID NO: 200; GCRRxTPGxxYLExL as set forth in SEQ ID NO: 201; CATGFDxxxxPRFxxxG as set forth in SEQ ID NO: 202; PNxFxxxGPNxPxxNGxV as set forth in SEQ ID NO: 203; AxWPGSxLHYxEAxxxPRxED as set forth in SEQ ID NO: 204; where the residues x, independently of each other, represent any naturally occurring amino acid residue. and / or (3) The group of polypeptides consisting of: (a) a polypeptide comprising the amino acid sequence of SCH23-BVMO1 set forth in SEQ ID NO:2; (b) a polypeptide comprising the amino acid sequence of SCH24-BVMO1 set forth in SEQ ID NO:6; (c) a polypeptide comprising the amino acid sequence of SCH25-BVMO1 set forth in SEQ ID NO: 10; (d) a polypeptide comprising the amino acid sequence of SCH46-BVMO1 set forth in SEQ ID NO: 13; (e) a polypeptide comprising the amino acid sequence of AspWeBVMO set forth in SEQ ID NO: 16; (f) a polypeptide comprising an amino acid sequence having at least 70% identity with any one of the amino acid sequences of a) to e).
14. The method according to any one of claims 9 to 13, wherein the
15. 15. The method according to any one of claims 9 to 14, further comprising as step (3) treating the carbonyl ester formed in step (1) or isolated in step (2) to obtain a derivative thereof using chemical synthesis or biocatalytic synthesis or a combination of both, and optionally isolating the derivative of step (3), wherein the derivative may in particular be selected from hydrocarbons, alcohols, diols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides and / or esters.
16. 16. The method according to claim 15, wherein step (3) comprises hydrolysing said carbonyl ester compound with a polypeptide having esterase activity (EC 3.1.1.) to obtain a corresponding de-esterified product, optionally isolating said derivative of step (3), and optionally subjecting the de-esterified product to an enzymatic redox reaction by the enzymatic action of a polypeptide having alcohol dehydrogenase (ADH) (EC 1.1.1.-) activity in a further step (4).
17. 15. An isolated polypeptide having BVMO activity as defined in claim 14.
18. 18. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding the polypeptide of claim 17.
19. 1. An in vivo method for preparing labdane-type terpenes comprising: The method comprises the following series of reaction steps: (1) optionally converting the labdan alcohol to the respective labdan aldehyde by the enzymatic action of an ADH polypeptide; (2) converting the labdan aldehyde of step (1) into the respective dinorlabdane carbonyl compound by the action of a polypeptide having enal cleavage activity as defined in claim 1; (3) optionally converting the dinorlabdane carbonyl compounds of step (2) to their respective tetranorlabdanyl acetates by the action of a polypeptide having BVMO activity as defined in claim 17; (4) optionally converting the tetranorlabdanyl acetate of step (3) into the respective tetranorlabdanyl alcohol by the action of a polypeptide having esterase activity; and, optionally, (5) isolating the product of step (2), (3) or (4). The method comprises providing a recombinant host expressing a set of polypeptides having the enzymatic activities necessary to catalyze
20. 1. An in vivo method for preparing labdane-type cycloterpenes comprising: The method comprises the following series of reaction steps: (1) optionally converting the labdan alcohol to the respective labdan aldehyde by the enzymatic action of an ADH polypeptide; (2) converting the labdan aldehyde of step (1) into the respective labdan ester compound by the action of a polypeptide having BVMO activity as defined in claim 17; (3) converting the labdane ester compounds of step (2) into their respective norlabdane aldehydes, optionally by the action of a polypeptide having esterase activity; (4) converting the norlabdane aldehyde of step (3) into the respective norlabdane ester by the action of a polypeptide having BVMO activity as defined in claim 17; (5) converting the norlabdane esters of step (4) into their respective dinorlabdane alcohols by the action of a polypeptide having esterase activity; (6) optionally converting the dinorlabdane alcohols of step (5) to their respective dinorlabdane carbonyl compounds by the action of a polypeptide having ADH activity; (7) optionally converting the dinorlabdane carbonyl compounds of step (6) into their respective tetranorlabdanyl acetates by the action of a polypeptide having BVMO activity as defined in claim 17; (8) converting the tetranorlabdanyl acetate of step (7) into the respective tetranorlabdane alcohol by the action of a polypeptide having esterase activity; and, optionally, (9) isolating the product of step (5), (6), (7) or (8); The method comprises providing a recombinant host expressing a set of polypeptides having the enzymatic activities necessary to catalyze
21. 1. A method for preparing an epoxy-tetranorlabdane compound, the method comprising: (1) providing a tetranorlabdane alcohol or acetate, or a dinorlabdane carbonyl compound, by applying a biocatalytic process comprising one or more process steps as defined in any one of claims 3 to 16, 19 or 20, and optionally isolating said product; (2) converting said product of step (1) into epoxy-tetranorlabdane by applying one or more chemical and / or biochemical conversion steps; A method comprising:
22. 1. A method for preparing a diepoxy-dinorlabdane, the method comprising: (1) providing a dinorlabdane carbonyl compound by applying a process comprising one or more process steps as defined in any one of claims 3 to 16, 19 or 20, and optionally isolating said dinorlabdane carbonyl compound; (2) converting said dinorlabdane carbonyl compound to said diepoxy-dinorlabdane by applying one or more chemical and / or biochemical conversion steps; A method comprising:
Citation Information
Patent Citations
Culture and mixture for producing diol
JP1995132082A
Method for producing alkyl methacrylate
JP2016504029A
Novel Baeyer-Villiger monooxygenase BVMOgm1 from Metagenome
KR1020150142243A
Biotransformation of ricinoleic acid into ω-hydroxyundec-9-enoic acid by fed-batch fermentation using glucose and glycerol
KR1020180058699A
Polypeptide for the enzymatic detoxification of zearalenone, isolated polynucleotide, and associated additive, use and method
US20180298352A1