Enzyme-mediated process for the preparation of amber aldehyde or amber aldehyde homologues

CN115667538BActive Publication Date: 2026-09-04GIVAUDAN SA
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202180038965.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-15
Filing Date
2021-04-14
Publication Date
2026-09-04
Estimated Expiration
2041-04-14

AI Technical Summary

Technical Problem

然而,天然泪柏醇的供应是有限的

Benefits of technology

[0021] Certain embodiments of the present invention may provide one or more of the following advantages:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115667538B_ABST
    Figure CN115667538B_ABST
Patent Text Reader

Abstract

An enzyme-mediated process for the production of ambrox and ambrox homologues, products of the process and uses of said products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to methods for preparing amberketal and its homologues using squalene-hopaene cyclase (SHC) or enzyme variants. This invention also relates to compositions prepared by this method, various uses of said compositions, and consumer products comprising said compositions. Background Technology

[0002] Ambroxanal provides a strong and persistent amber and woody aroma, and can be used alone or in combination with other woody or ambery components in fragrance compositions. Ambroxanal has traditionally been prepared from basalt alcohol through numerous chemical transformations. However, the supply of natural basalt alcohol is limited. Therefore, it is desirable to provide a new, efficient, and cost-effective synthetic route to obtain ambroxanal and its homologues. Summary of the Invention

[0003] According to a first aspect of the present invention, a method for preparing a compound of formula (I) is provided.

[0004]

[0005] The method involves contacting the compound of formula (II) with squalene-hopaene cyclase (SHC) or an enzyme variant.

[0006]

[0007] Where R is H, methyl, or ethyl.

[0008] In some embodiments, the method includes contacting a compound of formula (I) (the E,Z-compound of formula (II)) having an E-configuration double bond between C-8 and C-9 and a Z-configuration double bond between C-4 and C-5 with squalene-hopaene cyclase (SHC) (wild-type or variant enzyme).

[0009] In some embodiments, the method includes contacting a mixture of a compound of formula (II) comprising two double bonds of formula (E, E-compound) having both double bonds in the E-configuration and a compound of formula (II) comprising two double bonds of formula (II) having a double bond between C-8 and C-9 in the E-configuration and a double bond between C-4 and C-5 in the Z-configuration (E, Z-compound) with squalene-hopaene cyclase (SHC) or an enzyme variant.

[0010] In some embodiments, the weight ratio of the E,Z-compound to the E,E-compound is in the range of about 99:1 to about 10:90. For example, the weight ratio of the E,Z-compound of formula (II) to the E,E-compound of formula (II) may be in the range of about 95:5 to about 50:50, or about 80:20 to about 50:50, or about 80:20 to about 60:40.

[0011] According to a second aspect of the invention, a composition is provided comprising a compound of formula (I) and a compound of formula (III), consisting mainly of a compound of formula (I) and a compound of formula (III), or consisting of a compound of formula (I) and a compound of formula (III).

[0012]

[0013] Where R is H, methyl, or ethyl.

[0014] According to a third aspect of the invention, a composition is provided comprising a compound of formula (I), a compound of formula (IV), and a compound of formula (III), consisting mainly of a compound of formula (I), a compound of formula (IV), and a compound of formula (III), or consisting of a compound of formula (I), a compound of formula (IV), and a compound of formula (III).

[0015]

[0016] Where R is H, methyl, or ethyl.

[0017] According to a fourth aspect of the invention, a compound or composition is provided that can be obtained by or through the methods of the first aspect of the invention. This composition may include, for example, any embodiment thereof as defined in the second aspect of the invention.

[0018] According to a fifth aspect of the invention, the composition of the second or third aspect of the invention is provided for use as a flavoring composition or for use in a flavoring composition.

[0019] According to a sixth aspect of the invention, a consumer product comprising a composition of the second or third aspect of the invention is provided.

[0020] In some embodiments of any aspect of the invention, R is methyl. When R is methyl, the compound of formula (I) may be called (+)-ambroacetal.

[0021] Certain embodiments of the present invention may provide one or more of the following advantages:

[0022] ●Biocatalytic routes for the production of (+)-amyral and (+)-amyral homologues;

[0023] ● Milder reaction conditions (e.g., lower temperatures);

[0024] ●High selectivity;

[0025] ● Use alternative raw materials (e.g., alternative raw materials used in chemical synthesis).

[0026] Details, embodiments, and preferred embodiments provided with respect to any one or more of the aspects described herein will be further described herein, and are equally applicable to all aspects of the invention. Unless otherwise indicated herein or clearly contradicted by the context, the invention covers any combination of all possible variations of the embodiments, embodiments, and preferred embodiments described herein. Attached Figure Description

[0027] Figure 1 The reaction scheme for preparing the compound of formula (II) is shown. For the compound, R is H, methyl, or ethyl.

[0028] Figure 2 The conversion rates [%] of hydroxyfarethiacetone to (+)-amyrosine acetal using wild-type and variant SHC enzymes are shown, and an overview of the performance of the tested SHC enzymes under their optimal reaction conditions is given as shown in Table 2.

[0029] Sequence Overview

[0030] SEQ ID NO:1 is the amino acid sequence of wild-type Alicyclobacillus acidocaldarius AacSHC.

[0031] SEQ ID NO:2 is the nucleotide sequence of wild-type Acidothermic cyclophosphamide Bac.

[0032] SEQ ID NO:3 is SEQ ID NO:1 having substitutions for M132R, A224V, I432T, A557T and R613S and may be referred to herein as SHC enzyme variant #65.

[0033] SEQ ID NO:4 is the nucleotide sequence of SHC variant #65.

[0034] SEQ ID NO:5 corresponds to SEQ ID NO:1 having substitutions for M132R, A224V, I432T, Y81H, A557T, and R613S, and may be referred to herein as SHC enzyme variant #66.

[0035] SEQ ID NO:6 is the nucleotide sequence of the SHC variant SHC#66.

[0036] SEQ ID NO:7 corresponds to SEQ ID NO:1 having substitutions for M132R, A224V, I432T, T90A and R613S and may be referred to herein as SHC enzyme variant #90C7.

[0037] SEQ ID NO:8 is the nucleotide sequence of SHC variant #90C7.

[0038] SEQ ID NO:9 corresponds to SEQ ID NO:1 having substitutions for M132R, A224V, I432T, Y81H, H431L and A557T and may be referred to herein as SHC enzyme variant #110B8.

[0039] SEQ ID NO:10 is the nucleotide sequence of SHC variant #110B8.

[0040] SEQ ID NO:11 corresponds to SEQ ID NO:1 having substitutions for M132R, A224V, I432T, A172T, and M277K, and may be referred to herein as SHC enzyme variant #115A7.

[0041] SEQ ID NO:12 is the nucleotide sequence of SHC variant #115A7.

[0042] SEQ ID NO:13 corresponds to SEQ ID NO:1 with the mutations M132R, A224V and I432T and may be referred to herein as SHC enzyme variant 215G2.

[0043] SEQ ID NO:14 is the nucleotide sequence of SHC variant #215G2.

[0044] SEQ ID NO:15 is the amino acid sequence of wild-type Zymomonas mobilis ZmoSHC1.

[0045] SEQ ID NO:16 is the amino acid sequence of wild-type motile fermentation monoclonal bacteria ZmoSHC2.

[0046] SEQ ID NO:17 is the amino acid sequence of wild-type Bradyrhizobium japonicum BjaSHC.

[0047] SEQ ID NO:18 is the amino acid sequence of wild-type Thermosynechococcus elongatus TelSHC.

[0048] SEQ ID NO:19 is the amino acid sequence of wild-type Acetobacter pasteurianus ApaSHC1.

[0049] SEQ ID NO:20 is the amino acid sequence of wild-type pathogenic Gluconobacter morbifer GmoSHC.

[0050] SEQ ID NO:21 is the amino acid sequence of wild-type Bacillus megaterium BmeSHC.

[0051] SEQ ID NO:22 corresponds to SEQ ID NO:1 having substitutions for M132R, A224V, I432T, A557T, and H431L, and may be referred to herein as SHC enzyme variant #49. Detailed Implementation

[0052] This invention is based, at least in part, on the surprising discovery that squalene-hopaene cyclase (SHC) and enzyme variants can be used to prepare (+)-amyral and amyral homologues from polyunsaturated alcohols of formula (II). Particularly surprising is that substrates in which the alkenyl chain as defined herein is substituted with a hydroxymethyl group undergo an enzymatic polycyclization reaction terminated by internal ketalization.

[0053] Therefore, this paper provides a method for preparing compounds of formula (I) in the first aspect.

[0054]

[0055] The method involves contacting the compound of formula (II) with squalene-hopaene cyclase (SHC) or an enzyme variant.

[0056]

[0057] Where R is H, methyl, or ethyl.

[0058] In some embodiments, both double bonds are in the E-configuration (E,E-compounds of formula (II)).

[0059] In some embodiments, the double bond between C-8 and C-9 is E-configured and the double bond between C-4 and C-5 is Z-configured (E,Z-compounds of formula (II)).

[0060] In the specific implementation scheme, R stands for methyl.

[0061] The method described herein uses an enzymatic method to convert compounds of formula (II) into compounds of formula (I) using SHC enzymes or enzyme variants (biotransformation reaction).

[0062] Compound of formula (II)

[0063] The compounds of formula (II) exist in four different stereoisomers, for example, as compounds of formula (II) having E, E- or E, Z- configurations.

[0064] In some embodiments, the method includes contacting the E,Z-compound of formula (II) with squalene-hopaene cyclase (SHC) or an enzyme variant in the absence of any other stereoisomer of formula (II).

[0065] In other embodiments, the compound of formula (II) may be, for example, a mixture of stereoisomers. In some embodiments, the mixture comprises the E, E-compound of formula (II) and one or more other stereoisomers of formula (II). In some embodiments, the mixture comprises the E, Z-compound of formula (II) and one or more other stereoisomers of formula (II).

[0066] In some embodiments, the method includes contacting a mixture comprising an E,E-compound of formula (II) and an E,Z-compound of formula (II), a mixture consisting primarily of an E,E-compound of formula (II) and an E,Z-compound of formula (II), or a mixture consisting of an E,E-compound of formula (II) and an E,Z-compound of formula (II) with an SHC enzyme or an enzyme variant. In some embodiments, the composition does not contain any other stereoisomers of formula (II).

[0067] The weight ratio of the E,Z-compound of formula (II) to all other stereoisomers of formula (II) can, for example, be equal to or greater than about 10:90. For example, the weight ratio of the E,Z-compound of formula (II) to all other stereoisomers of formula (II) can be equal to or greater than about 20:80, or equal to or greater than about 30:70, or equal to or greater than about 40:60, or equal to or greater than about 50:50, or equal to or greater than about 60:40, or equal to or greater than about 70:30, or equal to or greater than about 80:20, or equal to or greater than about 90:10, or equal to or greater than about 95:5, or equal to or greater than about 99:1.

[0068] The weight ratio of the E,Z-compound of formula (II) to all other stereoisomers of formula (II) may, for example, be equal to or less than about 99:1. For example, the weight ratio of the E,Z-compound of formula (II) to other stereoisomers of formula (II) may be equal to or less than about 95:5, or equal to or less than about 90:10, or equal to or less than about 85:15, or equal to or less than about 80:20, or equal to or less than about 60:40.

[0069] For example, the weight ratio of the E,Z-compound of formula (II) to all other stereoisomers of formula (II) may be in the range of about 10:90 to about 99:1, or about 10:90 to about 90:10, or about 20:80 to about 80:20, or about 50:50 to about 80:20, or about 60:40 to about 80:20.

[0070] The weight ratio of the E,Z-compound of formula (II) to the E,E-compound of formula (II) can, for example, be equal to or greater than about 10:90. For example, the weight ratio of the E,Z-compound of formula (II) to the E,E-compound of formula (II) can be equal to or greater than about 20:80, or equal to or greater than about 30:70, or equal to or greater than about 40:60, or equal to or greater than about 50:50, or equal to or greater than about 60:40, or equal to or greater than about 70:30, or equal to or greater than about 80:20, or equal to or greater than about 90:10, or equal to or greater than about 95:5, or equal to or greater than about 99:1.

[0071] The weight ratio of the E,Z-compound of formula (II) to the E,E-compound of formula (II) can, for example, be equal to or less than about 99:1. For example, the weight ratio of the E,Z-compound of formula (II) to the E,E-compound of formula (II) can be equal to or less than about 95:5, or equal to or less than about 90:10, or equal to or less than about 85:15, or equal to or less than about 80:20, or equal to or less than about 70:30, or equal to or less than 60:40.

[0072] For example, the weight ratio of the E,Z-compound of formula (II) to the E,E-compound of formula (II) can be in the range of about 10:90 to about 99:1, or about 10:90 to about 90:10, or about 20:80 to about 80:20, or about 50:50 to about 80:20, or about 60:40 to about 80:20.

[0073] The amount of each stereoisomer in a mixture of stereoisomers can be identified, for example, by gas chromatography or NMR spectroscopy.

[0074] When R is methyl, the compound of formula (II) can be called hydroxyfarnesy acetone, including E,E-hydroxyfarnesy acetone, E,Z-hydroxyfarnesy acetone, Z,E-hydroxyfarnesy acetone and Z,Z-hydroxyfarnesy acetone and mixtures thereof.

[0075] In some embodiments, not all compounds of formula (II) are converted into compounds of formula (I) or reaction byproducts. Therefore, in some embodiments, compositions described herein, such as those obtained by or obtainable by the methods described herein, may comprise compounds of formula (I) and compounds of formula (II).

[0076] In some embodiments, any unconverted compound of formula (II) in a mixture prepared by the methods described herein can be separated from the other reaction products, such that the composition is free of any compound of formula (II).

[0077] In alternative embodiments, all compounds of formula (II) are converted into compounds of formula (I) or byproducts of the reaction by the methods described herein.

[0078] The number of stereoisomers of the compound of formula (II) present can affect the reaction rate. SHC enzymes or enzyme variants can convert the E,Z-compound of formula (II) from a complex mixture of stereoisomers of the compound of formula (II) (this mixture may contain only two stereoisomers, such as the E,Z-compound and E,E-compound of formula (II), or it may contain three (i.e., the E,Z- and E,E- and Z,E- or Z,Z-compound of formula (II), or even all four stereoisomers) into the compound of formula (I). However, lower conversion rates can be observed, as other stereoisomers can compete with the E,Z-compound of formula (II) to approach the SHC enzyme or enzyme variant. The viewpoint is consistent with that of the body, and therefore it can act as a competitive inhibitor for converting the E,Z- compounds of formula (II) into compounds of formula (I) and / or also as an alternative substrate. Therefore, the substrate of the compound of formula (II) can comprise 2-4 isomers, preferably a mixture of isomers of two isomers. Preferably, the substrate of the compound of formula (II) comprises a mixture of isomers of the E,Z- and E,E- compounds of formula (II), mainly a mixture of isomers of the E,Z- and E,E- compounds of formula (II), or a mixture of isomers of the E,Z- and E,E- compounds of formula (II).

[0079] Compounds of formula (II) can be synthesized according to the general method described by Fujiwara et al. (Tetrahedron Letters, 1995 Vol 36(46), 8435-8438). Alternatively, compounds of formula (II) can be synthesized as follows: Figure 1 The information is briefly described below, where R is H, methyl, or ethyl.

[0080] Compound of formula (I)

[0081] The compounds of formula (I) contain a plurality of chiral carbon atoms, and one or more stereoisomers of the compounds of formula (I) may also exist, including enantiomers and diastereomers. In addition to the compounds of formula (I), products prepared by the methods described herein may include one or more stereoisomers of the compounds of formula (I). The resulting stereoisomers may depend on the stereoisomers of the compounds of formula (II) used.

[0082] For example, in addition to compounds of formula (I), compounds of formula (IV) can also be prepared.

[0083]

[0084] Where R is H, methyl, or ethyl.

[0085] Compounds of formula (IV) in which R is methyl are also called (-)-epi-8-ambrosal acetal.

[0086] In some embodiments, other stereoisomers of the compound of formula (I) are not prepared by this method or are not present in the product of this method. For example, in some embodiments, the compound of formula (IV) is not prepared by this method or is not present in the product of this method.

[0087] The methods described herein can, for example, prepare compounds of formula (I) and one or more other stereoisomers of compounds of formula (I) (e.g., compounds of formula (IV)). Therefore, compositions described herein, such as those obtained by or that can be obtained by the methods described herein, may comprise compounds of formula (I) and one or more stereoisomers of compounds of formula (I) (e.g., compounds of formula (IV)).

[0088] The weight ratio of the compound of formula (I) to all other stereoisomers of formula (I) can, for example, be equal to or greater than about 50:50. For example, the weight ratio of the compound of formula (I) to all other stereoisomers of formula (I) can be equal to or greater than about 55:45, or equal to or greater than about 60:40, or equal to or greater than about 65:35, or equal to or greater than about 70:30, or equal to or greater than about 75:25, or equal to or greater than about 80:20, or equal to or greater than about 85:15, or equal to or greater than about 90:10, or equal to or greater than about 95:5, or equal to or greater than about 99:1.

[0089] For example, the compound of formula (I) may be the only stereoisomer of formula (I) prepared or present in the composition. In other words, the total weight of all stereoisomers of the compound of formula (I) in the composition may contain 100 wt% of the compound of formula (I), or the weight ratio of the compound of formula (I) to all other stereoisomers of formula (I) may be 100:0. Alternatively, the weight ratio of the compound of formula (I) to all other stereoisomers of formula (I) may be less than about 100 wt%. For example, the weight ratio of the compound of formula (I) to all other stereoisomers of formula (I) may be equal to or less than about 99:1, or equal to or less than about 98:2, or equal to or less than about 97:3.

[0090] For example, the weight ratio of the compound of formula (I) to all other stereoisomers of formula (I) may be from about 50:50 to about 100:0, or from about 60:40 to about 99:1, or from about 70:30 to about 98:2, or from about 80:20 to about 97:3, or from about 90:10 to about 97:3.

[0091] The amount of each stereoisomer of the compound of formula (I) in a mixture of stereoisomers can be identified, for example, by gas chromatography on a chiral column or by NMR spectroscopy in the presence of a shift reagent.

[0092] The compound of formula (I) obtained by the method described herein can be, for example, in an amorphous or crystalline form.

[0093] The compounds of formula (I) prepared by the methods described herein can be separated by steam extraction / distillation or organic solvent extraction using an aqueous miscible solvent (to separate the reaction products and unreacted substrates from the biocatalyst remaining in the aqueous phase), followed by solvent evaporation, to obtain the crude reaction products as determined by gas chromatography (GC) analysis. Steam extraction / distillation and organic solvent extraction methods are known to those skilled in the art.

[0094] As an example, the obtained compound of formula (I) can be extracted from the entire reaction mixture using an organic solvent such as an aqueous miscible solvent (e.g., toluene). Alternatively, the obtained compound of formula (I) can be extracted from the solid phase of the reaction mixture (obtained by, for example, centrifugation or filtration) using an aqueous miscible solvent (e.g., ethanol) or an aqueous miscible solvent (e.g., toluene). As another example, the compound of formula (I) can exist in the solid phase as a crystal or in an amorphous form, and can also be separated from the remaining solid phase (cell material or fragments thereof) and liquid phase by filtration. As yet another example, at temperatures above the melting point of the compound of formula (I), an oil layer can form on top of the aqueous phase, which can be removed and collected. To ensure complete recovery of the compound of formula (I) after removal of the oil layer, an organic solvent can be added to the aqueous phase containing biomass to extract any residual compound of formula (I) (e.g., (+)-ambrosal) contained in, on, or around the biomass. The organic layer can be combined with the oil layer, and then the whole can be further processed to separate and purify the compound of formula (I). The compound of formula (I) can be further selectively crystallized to remove byproducts and any unreacted compounds of formula (II) from the final product. The term "selective crystallization" refers to a method step in which the compound of formula (I) is crystallized from a solvent while the byproducts remain dissolved in the crystallization solvent to a degree such that the separated crystalline material contains only compounds of formula (I), or if it contains any byproducts, they are present only in olfactorily acceptable amounts. For example, the compound of formula (I) contains no or substantially no byproducts such as compounds of formula (III) or (IIIa). The selective crystallization step can use water-miscible solvents such as ethanol. Selective crystallization of the compound of formula (I) can be affected by the presence of unreacted compounds of formula (II) and the ratio of the compound of formula (I) to other detectable byproducts. Selective crystallization of the compound of formula (I) is still possible even if only 10% of the compound of formula (II) is converted to the compound of formula (I).

[0095] The olfactory purity of the final compound product of formula (I) can be determined using a 10% ethanol extract in water or by testing the crystalline material. The olfactory purity, quality, and sensory characteristics of the final compound product of formula (I) are tested relative to a commercially available reference compound of formula (I). The compound material of formula (I) is also tested by experts in applied studies to determine whether the material meets the specifications regarding its sensory characteristics.

[0096] Examples of suitable water-miscible and non-water-miscible organic solvents for the extraction and / or selective crystalline (I) compounds include, but are not limited to, aliphatic hydrocarbons, preferably those having 5 to 8 carbon atoms, such as pentane, cyclopentane, hexane, cyclohexane, heptane, octane, or cyclooctane; halogenated aliphatic hydrocarbons, preferably those having one or two carbon atoms, such as dichloromethane, chloroform, carbon tetrachloride, dichloroethane, or tetrachloroethane; aromatic hydrocarbons, such as benzene, toluene, xylene, chlorobenzene, or dichlorobenzene; aliphatic acyclic and cyclic ethers or alcohols, preferably those having 4 to 8 carbon atoms, such as ethanol, isopropanol, diethyl ether, methyl tert-butyl ether, ethyl tert-butyl ether, dipropyl ether, diisopropyl ether, dibutyl ether, tetrahydrofuran, or esters, such as ethyl acetate or n-butyl acetate, or ketones, such as methyl isobutyl ketone, or dioxane, or mixtures thereof. Particularly preferred solvents are heptane, methyl tert-butyl ether (also known as MTBE, tert-butyl methyl ether, tert-butyl methyl ether, and tBME), diisopropyl ether, tetrahydrofuran, ethyl acetate, and / or mixtures thereof. Preferably, a water-miscible solvent such as ethanol is used to extract the compound of formula (I) from the solid phase of the reaction mixture. The use of ethanol is advantageous because it is easy to handle, non-toxic, and environmentally friendly.

[0097] As used herein, the term "isolated" refers to a biotransformation product, such as a compound of formula (I), which has been isolated or purified from its accompanying components. An entity produced in a cellular system different from its natural origin is "isolated" because it necessarily does not contain the naturally accompanying components. The degree of isolation or purity can be measured by any suitable method, such as gas chromatography (GC), HPLC, or NMR analysis.

[0098] In some embodiments, the compound of formula (I) (e.g., (+)-amyral) is isolated and purified from the obtained crude product (e.g., to a purity of at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%).

[0099] Ideally, the concentration of the compound of formula (I) in the reaction solution obtained by the method described herein can be from about 1 mg / L to about 20,000 mg / L (20 g / L) or higher, for example from about 20 g / L to about 200 g / L or 100-500 g / L (including 150 g / L, 250 g / L, 300 g / L, 350 g / L, 400 g / L or 450 g / L).

[0100] Compounds of formula (III)

[0101] The method described herein can, for example, prepare compounds of formula (III) as byproducts, wherein R is H, methyl, or ethyl:

[0102]

[0103] Compounds of formula (III) can be, for example, compounds having the relative configuration of formula (IIIa), wherein R is H, methyl, or ethyl:

[0104]

[0105] The methods described herein can, for example, prepare one or more stereoisomers of a compound of formula (III). The compositions described herein can comprise one or more stereoisomers of a compound of formula (III). The methods described herein can, for example, prepare compounds of formula (III) having the relative configuration shown in formula (IIIa). Therefore, the compositions described herein can comprise compounds of formula (III) having the relative configuration shown in formula (IIIa). In some embodiments, the only compound of formula (III) prepared by the methods described herein and thus present in the compositions described herein is a compound having the relative configuration shown in formula (IIIa).

[0106] In some embodiments, at least about 50 wt% of all compounds of formula (III) have the relative configuration shown in formula (IIIa). For example, at least about 60 wt%, or at least about 70 wt%, or at least about 80 wt%, or at least about 90 wt% of all compounds of formula (III) may have the relative configuration shown in formula (IIIa).

[0107] In some embodiments, the compound having the configuration shown in formula (IIIa) is the only stereoisomer of formula (III). For example, 100 wt% of all compounds of formula (III) have the relative configuration shown in formula (IIIa). In some embodiments, equal to or less than about 99 wt%, or equal to or less than about 95 wt%, or equal to or less than about 90 wt%, or equal to or less than about 85 wt%, or equal to or less than about 80 wt%, or equal to or less than about 75 wt% of all compounds of formula (III) have the relative configuration shown in formula (IIIa).

[0108] For example, about 50 wt% to about 100 wt%, or about 60 wt% to about 99 wt%, or about 70 wt% to about 95 wt% of all compounds of formula (III) have the relative configuration shown in formula (IIIa).

[0109] The amounts of different isomers of the compound of formula (III) in a mixture of stereoisomers can be identified, for example, by gas chromatography on a chiral column or by NMR spectroscopy in the presence of a shift reagent.

[0110] The product obtained by the method described in this article

[0111] This document also provides products made by the methods described herein. Therefore, this document also provides compositions that are obtained by or can be obtained by the methods described herein (including all embodiments thereof).

[0112] This document provides a composition comprising, being primarily composed of, or consisting of, compounds of formula (I) and (III) or compounds of formula (III). The composition may, for example, further comprise one or more stereoisomers of other formula (I), such as compounds of formula (IV). The composition may, for example, comprise one or more stereoisomers of other formula (III). For example, the composition may comprise a compound having the relative configuration of formula (IIIa). The composition may, for example, further comprise any unreacted compound of formula (II). In some embodiments, R is methyl.

[0113] In one specific embodiment, a composition is provided comprising a compound of formula (I), a compound of formula (IV), and a compound of formula (III), consisting mainly of a compound of formula (I), a compound of formula (IV), and a compound of formula (III), or consisting of a compound of formula (I), a compound of formula (IV), and a compound of formula (III).

[0114] This document also provides compositions comprising a compound of formula (I) and one or more stereoisomers of the compound of formula (I), such as a compound of formula (IV). The composition may, for example, further comprise a compound of formula (III). The composition may further comprise any unreacted compound of formula (II). In some embodiments, R is methyl.

[0115] The weight ratio of the compound of formula (I) to the compound of formula (III) in the compositions described herein can, for example, range from about 60:40 to about 99:1. For example, the weight ratio of the compound of formula (I) to the compound of formula (III) can range from about 65:35 to about 99:1, or about 70:30 to about 99:1, or about 75:25 to about 99:1, or about 80:20 to about 99:1, or about 85:15 to about 99:1, or about 90:10 to about 99:1, or about 95:5 to about 99:1. For example, the weight ratio of the compound of formula (I) to the compound of formula (III) can range from about 65:35 to about 98:2, or about 70:30 to about 97:3, or about 75:25 to about 96:4, or about 80:20 to about 95:5, or about 85:15 to about 90:10.

[0116] The weight ratio of the compound of formula (I) to the compound of formula (II) in the crude reaction product described herein can, for example, range from about 90:10 to about 100:0. For example, the weight ratio of the compound of formula (I) to the compound of formula (II) in the composition described herein can range from about 92:8 to about 100:0, or about 94:6 to about 100:0, or about 95:5 to about 100:0, or about 96:4 to about 99.5:0.5, or about 97:3 to about 99.0:1.0, or about 98:2 to about 99.0:1.0.

[0117] Flavoring composition

[0118] This article also provides information on the use of the reaction products described herein in or as part of a fragrance composition.

[0119] Therefore, this document also provides fragrance compositions comprising one or more compounds of formula (I). A "fragrance composition" refers to any composition comprising one or more compounds of formula (I) and a matrix material.

[0120] As used herein, “matrix material” includes all known fragrance ingredients selected from a wide range of natural products and currently available synthetic molecules, such as essential oils, alcohols, aldehydes and ketones, ethers and acetals, esters and lactones, macrocyclic and heterocyclic compounds, and / or mixtures with one or more ingredients or excipients (e.g., carrier materials, diluents, and other auxiliaries commonly used in the art) that are typically used in conjunction with flavor enhancers in fragrance compositions.

[0121] Fragrance ingredients known in the art are readily available from major fragrance manufacturers. Non-limiting examples of such ingredients include:

[0122] - Essential oils and extracts, such as castoreum, costus root oil, oakmoss absolute oil, geranium oil, tree moss absolute oil, basil oil, fruit oils (such as bergamot oil and mandarin orange oil), myrtle oil, palmarosa oil, patchouli oil, orange leaf oil, jasmine oil, rose oil, sandalwood oil, wormwood oil, lavender oil and / or ylang-ylang oil;

[0123] - Alcohols, such as cinnamyl alcohol ((E)-3-phenylprop-2-en-1-ol); cis-3-hexenol ((Z)-hex-3-en-1-ol); citronellol (3,7-dimethyloct-6-en-1-ol); dihydromyrcenol (2,6-dimethyloct-7-en-2-ol); Ebanol TM ((E)-3-methyl-5-(2,2,3-trimethylcyclopentan-3-en-1-yl)pentan-4-en-2-ol); eugenol (4-allyl-2-methoxyphenol); ethyl linalool ((E)-3,7-dimethylnon-1,6-dien-3-ol); farnesol ((2E,6Z)-3,7,11-trimethyldodec-2,6,10-trien-1-ol); geraniol ((E)-3,7-dimethyloct-2,6-dien-1-ol); Super Muguet TM (E)-6-ethyl-3-methyloct-6-en-1-ol; linalool (3,7-dimethyloct-1,6-dien-3-ol); menthol (2-isopropyl-5-methylcyclohexanol); nerol (3,7-dimethyl-2,6-octadien-1-ol); phenylethyl alcohol (2-phenylethanol); Rhodinol TM (3,7-Dimethyloct-6-en-1-ol); Sandalore TM (3-Methyl-5-(2,2,3-trimethylcyclopentan-3-en-1-yl)pentan-2-ol); terpineol (2-(4-methylcyclohexane-3-en-1-yl)propan-2-ol); or Timberol TM (1-(2,2,6-trimethylcyclohexyl)hex-3-ol); 2,4,7-trimethyloct-2,6-dien-1-ol and / or [1-methyl-2(5-methylhex-4-en-2-yl)cyclopropyl]-methanol;

[0124] - Aldehydes and ketones, such as anisaldehyde (4-methoxybenzaldehyde); α-pentylcinnamaldehyde (2-benzylidene heptanal); Georgywood TM (1-(1,2,8,8-tetramethyl-1,2,3,4,5,6,7,8-octahydronaphth-2-yl)acetone); hydroxycitronellol (7-hydroxy-3,7-dimethyloctanal); Iso E (1-(2,3,8,8-tetramethyl-1,2,3,4,5,6,7,8-octahydronaphth-2-yl)acetone); ((E)-3-methyl-4-(2,6,6-trimethylcyclohex-2-en-1-yl)but-3-en-2-one); 3-(4-isobutyl-2-methylphenyl)propanal; maltol; methyl cypressone; methyl ionone; verbenone; and / or vanillin;

[0125] -Ethers and acetals, for example (3a,6,6,9a-tetramethyl-2,4,5,5a,7,8,9,9b-octahydro-1H-benzo[e][1]benzofuran); geranyl methyl ether ((2E)-1-methoxy-3,7-dimethyloctyl-2,6-diene); rose ether (4-methyl-2-(2-methylprop-1-en-1-yl)tetrahydro-2H-pyran); and / or (2',2',3,7,7-pentamethylspiro[bicyclo[4.1.0]heptane-2,5'-[1,3]dioxane]);

[0126] - Macrocyclic compounds, such as asterolone ((Z)-oxetane-10-en-2-one); ethylidene tridecanoate (1,4-dioxetane-5,17-dione); and / or (16-oxacyclohexadecane-1-one); and

[0127] - Heterocyclic compounds, such as isobutylquinoline (2-isobutylquinoline).

[0128] As used in this article, “carrier material” refers to a material that is actually neutral from the perspective of flavor enhancers, that is, a material that does not significantly alter the sensory properties of flavor enhancers.

[0129] The term "diluent" refers to any diluent that is typically used in conjunction with flavor enhancers, such as diethyl phthalate (DEP), dipropylene glycol (DPG), isopropyl myristate (IPM), triethyl citrate (TEC), and alcohols (e.g., ethanol).

[0130] The term "auxiliary agent" refers to an ingredient that can be used in a fragrance composition for reasons not particularly related to the olfactory properties of the composition. For example, an auxiliary agent may be an ingredient used as an adjuvant in processing a fragrance ingredient or a composition containing said ingredient, or it may improve the handling or storage of a fragrance ingredient or a composition containing said fragrance ingredient, such as an antioxidant auxiliary agent, which may be selected from, for example... TT (BASF) Q (BASF), tocopherol (including its isomers, CAS 59-02-9; 364-49-8; 18920-62-2; 121854-78-2), 2,6-bis(1,1-dimethylethyl)-4-methylphenol (BHT, CAS 128-37-0) and related phenols, hydroquinone (CAS 121-31-9).

[0131] It can also be an ingredient that provides additional benefits (such as imparting color or texture). It can also be an ingredient that imparts lightfastness or chemical stability to one or more ingredients contained in a fragrance composition.

[0132] A detailed description of the properties and types of auxiliaries commonly used in fragrance compositions containing auxiliaries is not exhaustive, but it must be mentioned that the ingredients are well known to those skilled in the art.

[0133] The various applications of compounds of formula (I) include, but are not limited to, fine fragrances or consumer products such as fabric care, toiletries, beauty care and cleaning products, detergent products and soap products, including virtually all products in which the currently available (+)-amyraldehyde ingredient is used commercially.

[0134] This document also provides consumer products that comprise compositions or fragrance compositions as described herein, including any embodiments thereof. Consumer products can be, for example, cosmetics (e.g., perfumes or eau de parfums), cleaning products, detergent products, or soap products.

[0135] intermediates and raw materials

[0136] This document also provides intermediates and raw materials for use in the methods described herein.

[0137] This document also provides compounds comprising formula (II), compounds primarily composed of formula (II), or mixtures comprising formula (II). For example, a mixture may comprise a compound of formula (II) in which both double bonds are in the E-configuration (E,E-compound) and a compound of formula (II) in which the double bond between C-8 and C-9 is in the E-configuration and the double bond between C-4 and C-5 is in the Z-configuration (E,Z-compound), or is primarily composed of them or thereof. In one embodiment, the mixture may comprise three stereoisomers of formula (II) (i.e., E,Z-, E,E-, and Z,E- or Z,Z-compounds of formula (II)) or even all four stereoisomers of formula (II), or is primarily composed of them or thereof.

[0138] The weight ratio of E,Z-compound to E,E-compound in formula (II) can, for example, be equal to or greater than about 10:90. For example, the weight ratio of E,Z-compound to E,E-compound in formula (II) can be equal to or greater than about 20:80, or equal to or greater than about 30:70, or equal to or greater than about 40:60, or equal to or greater than about 50:50, or equal to or greater than about 60:40, or equal to or greater than about 70:30, or equal to or greater than about 80:20, or equal to or greater than about 90:10, or equal to or greater than about 95:5, or equal to or greater than 99:1.

[0139] The weight ratio of E,Z-compound to E,E-compound in formula (II) can, for example, be equal to or less than about 99:1. For example, the weight ratio of E,Z-compound to E,E-compound in formula (II) can be equal to or less than about 95:5, or equal to or less than about 90:10, or equal to or less than about 85:15, or equal to or less than about 80:20, or equal to or less than about 70:30, or equal to or less than 60:40.

[0140] For example, the weight ratio of the E,Z-compound of formula (II) to the E,E-compound of formula (II) can be in the range of about 10:90 to about 99:1, or about 10:90 to about 90:10, or about 20:80 to about 80:20, or about 50:50 to about 80:20, or about 60:40 to about 80:20.

[0141] SHC enzymes or enzyme variants

[0142] The method described herein uses an SHC enzyme or a variant enzyme to enzymatically convert a compound of formula (II) into a compound of formula (I).

[0143] As used herein, the term "SHC enzyme" refers to wild-type squalene-hopaene cyclases that are naturally found in, for example, thermophilic bacteria such as Bacillus pyrolyticus.

[0144] As used herein, the term "variant" should be understood as a polypeptide that differs from the polypeptide from which it is derived by one or more variations in its amino acid sequence. The polypeptide from which a variant is derived is also called the parent or reference polypeptide. Typically, variants are artificially generated, preferably by genetic technology. Typically, the polypeptide from which a variant is derived is a wild-type enzyme or a wild-type enzyme. However, variants used in this disclosure may also be derived from homologs, orthologs, or paralogs of the parent polypeptide, or from artificially constructed variants. Variations in the amino acid sequence can be amino acid exchanges (substitutions), insertions, deletions, N-terminal truncation, or C-terminal truncation, or any combination of these variations, which may occur at one or more sites.

[0145] As used herein, the term "SHC enzyme variant" refers to an enzyme derived from a wild-type SHC enzyme but with one or more amino acid alterations compared to the wild-type SHC enzyme and therefore not naturally occurring. These one or more amino acid alterations may, for example, alter (e.g., increase) the enzyme activity of a substrate (e.g., a compound of formula (II)). Alternatively, a variant SHC enzyme may be derived from an existing SHC enzyme variant.

[0146] Assays for determining and quantifying the activity of SHC enzymes and / or SHC enzyme variants are described herein and are known in the art. As an example, the activity of SHC enzymes and / or SHC enzyme variants can be determined by incubating a purified SHC enzyme or enzyme variant, or an extract from a host cell or intact recombinant host organism that has produced an SHC enzyme or enzyme variant, with a suitable substrate under appropriate conditions, followed by analysis of the substrate and reaction products (e.g., by gas chromatography (GC) or HPLC). Further details regarding the assays of SHC enzyme and / or SHC enzyme variant activity and the analysis of reaction products are provided in the examples. These assays may include the production of SHC enzyme variants in recombinant host cells (e.g., *Escherichia coli*).

[0147] As used herein, the term "activity" refers to the ability of an enzyme to react with a substrate to provide a target product. This activity can be determined in a so-called activity assay by an increase in the target product, a decrease in the substrate (or starting material), or a combination of these parameters as a function of time. The SHC enzymes of this disclosure are characterized by their ability to convert compounds of formula (II) (e.g., hydroxyfarethionylacetone) into compounds of formula (I) (e.g., (+)-amyrosal).

[0148] As used herein, “bioactivity” means any activity that a polypeptide may exhibit, including but not limited to: enzymatic activity; binding activity to another compound (e.g., binding to another polypeptide, particularly to a receptor, or to a nucleic acid); inhibitory activity (e.g., enzyme inhibitory activity); activating activity (e.g., enzyme activating activity); or toxicity. It is not required that the variant exhibit the same degree of such activity as the parent or wild-type polypeptide. In other embodiments, the SHC enzyme variants used herein show better substrate conversion yields than reference SHC enzymes (e.g., wild-type SHC enzymes or known SHC enzyme variants). In further embodiments, the SHC enzyme variants used herein may show altered (e.g., increased) productivity relative to reference SHC enzymes (e.g., 215G2 AacSHC relative to wild-type AacSHC). The term “productivity” refers to the amount (in grams) of recoverable product per liter of reaction per hour of bioconversion time (i.e., time after substrate addition).

[0149] As used herein, the term "amino acid alteration" refers to the insertion, deletion, or substitution of one or more amino acids (which may be conserved or non-conserved) between two amino acids in an amino acid sequence relative to a reference amino acid sequence. Substitution involves replacing the amino acids in the reference sequence with the same number of amino acids in the variant sequence. The reference amino acid sequence may be, for example, a wild-type (WT) amino acid sequence (e.g., SEQ ID NO:1, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, or SEQ ID NO:21), or may be, for example, an SHC enzyme variant sequence (e.g., AacSHC variant 215G2-SEQ ID NO:13).

[0150] Amino acid alterations can be readily identified by comparing the amino acid sequence of an SHC enzyme variant with that of a reference amino acid sequence (e.g., wild-type or variant SHC amino acid sequence).

[0151] Suitable sources of SHC enzymes include, for example, *Bacillus cyclophosphamide* (Aac), *Zomonas molybdate* (Zmo), *Bja* (Bja), *Glucosobacterium virulence* (Gmo), *Burkholderia ambifaria*, *Bacillus anthracis*, *Bme* (Bme), *Methylococcus capsulatus*, *Frankia alni*, *Acetobacter pastoris* (Apa), *Synechococcus slenderus* (Tel), *Streptomyces coelicolor* (Sco), *Rhodopseudomonas palustris* (Rpa), *Teredinibacter urnerae* (Ttu), and *Pelobacter* (Pelobacter). *Carbinolicus carbinolicus* (Pca) and *Tetrahymenapyriformis* (see, for example, WO2010 / 139719, US2012 / 01345477, WO2012 / 066059, WO2016 / 170099; WO2018 / 157021 and JP2009060799, the contents of which are incorporated herein by reference). Suitable enzymes are also described, for example, in Neumann & Simon 1986, Biol Chem Hoppe-Seyler 367, 723-729; Seckler & Pollala 1986, Biochem Biophys Act 356-363; Ochs et al. 1990, J Bacteriol 174, 298-302; Seitz et al. 2012, J Molecular Catalysis B: Enzymatic 84, 72-77; and Seitz's 2012 PhD paper (http: / / elib.uni-stuttgart.de / handle / 11682 / 1400), the contents of which are incorporated herein by reference. These SHC enzymes and variants can be used in accordance with the methods described herein.

[0152] Specifically, the SHC enzyme (e.g., SHC enzyme variants from which can be derived) can be *Aac* SHC enzyme, *Zmo* SHC enzyme, *Bja* SHC enzyme, *Bme* SHC enzyme, or *Gmo* SHC enzyme. In one embodiment, the SHC enzyme (e.g., SHC enzyme variants from which can be derived) can be *Aac* SHC enzyme. In another embodiment, the SHC enzyme (e.g., SHC enzyme variants from which can be derived) can be *Bme* SHC enzyme.

[0153] For ease of reference, the names "AacSHC" can refer to Aac-SCH367 (Aac) SHC enzyme, "ZmoSHC" can refer to Zmo-SCH367 (Zmo) SHC enzyme, "BmeSHC" can refer to Bme-SCH367 (Bme) SHC enzyme, "BjaSHC" can refer to Bja-SCH367 (Bja) SHC enzyme, and "GmoSHC" can refer to Gmo-SCH367 (Gmo) SHC enzyme.

[0154] The sequences of AacSHC, ZmoSHC, and BjaSHC enzymes are published in BASF WO2010 / 139719, US2012 / 01345477A1, Seitz et al. (as cited above), and Seitz's 2012 PhD paper (as cited above). Two different sequences of ZmoSHC, designated ZmoSHC1 and ZmoSHC2, are disclosed. The sequence of GmoSHC enzyme is published in WO2018 / 157021. Table 1 discloses the sources and accession numbers of wild-type SHC enzymes.

[0155] Table 1. Sources and accession numbers of wild-type SHC enzymes.

[0156]

[0157] The amino acid sequences of wild-type AacSHC, wild-type ZmoSHC1, wild-type ZmoSHC2, wild-type BjaSHC, wild-type GmoSHC, wild-type TelSHC, wild-type ApaSHC1, and wild-type BmeSHC are also disclosed in this paper (SEQ ID NO:1, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, and SEQ ID NO:21, respectively).

[0158] The SHC enzyme or enzyme variant described herein and used in the methods described herein may be, for example, based on one of the wild-type amino acid sequences (SEQ ID NO 1, 13, 15, 16, 17, 18, 19, 20, 21) or a variant, homology, mutant, derivative, or fragment thereof. The SHC enzyme or enzyme variant may, for example, have at least 30%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70% of the wild-type amino acid sequence disclosed herein. 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the amino acid sequence is identical.

[0159] The SHC enzyme or enzyme variant described herein and used in the methods described herein may, for example, have a selectivity equal to or greater than about 75%. For example, the SHC enzyme or enzyme variant may, for example, have a selectivity equal to or greater than about 80%, or equal to or greater than about 85%, or equal to or greater than about 90%, or equal to or greater than about 95%. For example, the SHC enzyme or enzyme variant may, for example, have a selectivity of up to 100%, for example less than 100%, for example equal to or less than about 99.5%, or equal to or less than about 99.0%, or equal to or less than about 98.0%, or equal to or less than about 97.0%.

[0160] The "percentage (%) identity" of a gene's nucleotide sequence is defined as the percentage of nucleotides in a candidate DNA sequence that are identical to nucleotides in the DNA sequence after sequence alignment and the introduction of gaps (if necessary) to achieve maximum percentage sequence identity, and without considering any conserved substitutions as part of sequence identity. Alignments for determining percentage nucleotide sequence identity can be performed in various ways within the scope of the art, such as using publicly available computer software. Those skilled in the art can determine appropriate parameters for measuring alignments, including any algorithms required to achieve maximum alignment across the full length of the sequences being compared. The terms "peptide" and "protein" are used interchangeably herein and refer to any amino acid-linked peptide chain, regardless of length or post-translational modifications.

[0161] As used herein, the term "derivative" includes, but is not limited to, variants. The terms "derivative" and "variant" are used interchangeably herein.

[0162] The specific SHC enzymes and enzyme variants that can be used in the methods described herein are further described below.

[0163] Wild-type SHC enzymes

[0164] The methods described herein can, for example, use SHC enzymes that have 100% sequence identity with wild-type SHC enzymes. Wild-type SHC enzymes do not necessarily need to be obtained directly from their natural organisms and can be synthesized in the laboratory, for example, using recombinant DNA technology.

[0165] Wild-type SHC enzymes can be derived from, for example, *Bacillus cyclophosphamide* (Aac), *Zmo*, *Bja*, *Gmo*, *Burkholderia bifidum*, *Bacillus anthracis*, *Bme*, *Methylcoccus capsulatum*, *Frankella altissima*, *Acetobacter pastoris* (Apa), *Synechococcus slenderus* (Tel), *Streptomyces coli* (Sco), *Rhodopseudomonas palustris* (Rpa), *Teredinibacter turnerae* (Ttu), *Methylobacterium methylprednis* (Pca), and *Tetrahymena piriformis* (see, for example, WO2010 / 139719, US2012 / 01345477, WO2012 / 066059, the contents of which are incorporated herein by reference).

[0166] Specifically, wild-type SHC enzymes can be Aac (acid-thermal cyclophosphamide) SHC enzymes, Zmo (molymphoma) SHC enzymes, Bme (megabramycosis) SHC enzymes, Bja (Japanese stomatologic rhizobium) SHC enzymes, or Gmo (gastropathogenic glucosamine) SHC enzymes. Specifically, wild-type SHCs can be Aac (acid-thermal cyclophosphamide) SHC enzymes, Bme (megabramycosis) SHC enzymes, or Zmo (molymphoma) SHC enzymes (e.g., ZmoSHC1).

[0167] For ease of reference, the names "AacSHC" can refer to the SHC enzyme of Bacillus cyclophosphamide, "ZmoSHC" can refer to the SHC enzymes ZmoSHC1 and ZmoSHC2 of Fermentomonas motilityis, "BjaSHC" can refer to the SHC enzyme of Chlorobacterium tumefaciens, "BmeSHC" can refer to the SHC enzyme of Bacillus megaterium, "GmoSHC" can refer to the SHC enzyme of Staphylococcus aureus, "TelSHC" can refer to the SHC enzyme of Synechococcus slenderus, and "ApaSHC1" can refer to the SHC enzyme of Acetobacter pasteurellium.

[0168] The amino acid sequence of the wild-type SHC enzyme can be, for example, AacSHC (SEQ ID NO:1), ZmoSHC1 (SEQ ID NO:15), ZmoSHC2 (SEQ ID NO:16), BjaSHC (SEQ ID NO:17), GmoSHC (SEQ ID NO:20), BmeSHC (SEQ ID NO:21), TelSHC (SEQ ID NO:18), or ApaSHC1 (SEQ ID NO:19). For example, the wild-type SHC enzyme can be BmeSHC (SEQ ID NO:21), ZmoSHC1 (SEQ ID NO:15), or AacSHC (SEQ ID NO:1).

[0169] SHC variant enzyme

[0170] The methods described herein can, for example, use SHC enzyme variants (i.e., SHC enzymes with less than 100% sequence identity to wild-type SHC enzymes).

[0171] The methods described herein may, for example, use SHC enzyme variants as described in WO2016 / 170099 or WO2018 / 157021, the contents of which are incorporated herein by reference. For example, the SHC enzyme variant used in the methods described herein may be SHC enzyme variant 215G2 as described in WO2016 / 170099.

[0172] SHC enzyme variants may, for example, have an amino acid sequence that is at least about 70.0% identical to the amino acid sequence of the wild-type SHC enzyme. For example, an SHC enzyme variant may have an amino acid sequence that is at least about 75.0%, or at least about 80.0%, or at least about 85.0%, or at least about 90.0%, or at least about 95.0%, or at least about 95.5%, or at least about 96.5%, or at least about 97.0%, or at least about 97.5%, or at least about 98.0%, or at least about 98.5%, or at least about 99.0%.

[0173] The amino acid sequence identity between the SHC enzyme variant and the wild-type SHC enzyme is less than 100%, for example, equal to or less than about 99.5%, or equal to or less than about 99.0%.

[0174] For example, SHC enzyme variants may have about 70.0% to about 99.5%, or about 80.0% to about 99.0%, or about 85.0% to about 98.5%, or about 90.0% to about 98.0% of the amino acid sequence of wild-type SHC enzymes.

[0175] Wild-type SHC enzymes can be derived from, for example, *Bacillus cyclophosphamide* (Aac), *Zmo*, *Bja*, *Gmo*, *Burkholderia bifidum*, *Bacillus anthracis*, *Bme*, *Methylcoccus capsulatum*, *Frankella altissima*, *Acetobacter pastoris* (Apa), *Synechococcus slenderus* (Tel), *Streptomyces coli* (Sco), *Rhodopseudomonas palustris* (Rpa), *Ttu*, *Pca*, and *Tetrahymena piriformis* (see, for example, WO2010 / 139719, US2012 / 01345477, WO2012 / 066059, the contents of which are incorporated herein by reference).

[0176] The amino acid sequence of the wild-type SHC enzyme can be, for example, AacSHC (SEQ ID NO:1), ZmoSHC1 (SEQ ID NO:15), ZmoSHC2 (SEQ ID NO:16), BjaSHC (SEQ ID NO:17), GmoSHC (SEQ ID NO:20), BmeSHC (SEQ ID NO:21), TelSHC (SEQ ID NO:18), or ApaSHC1 (SEQ ID NO:19). For example, the wild-type SHC enzyme can be AacSHC (SEQ ID NO:1).

[0177] Therefore, in some embodiments, the SHC enzyme or SHC enzyme variant may have an amino acid sequence that has at least about 70.0% identity with SEQ ID NO:1, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:20, SEQ ID NO:18, SEQ ID NO:19 or SEQ ID NO:21. For example, an SHC enzyme or an SHC enzyme variant has an amino acid sequence that is at least about 75.0%, or at least about 80.0%, or at least about 85.0%, or at least about 90.0%, or at least about 95.0%, or at least about 95.5%, or at least about 96.5%, or at least about 97.0%, or at least about 97.5%, or at least about 98.0%, or at least about 98.5%, or at least about 99.0% identical to SEQ ID NO:1, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:20, SEQ ID NO:18, SEQ ID NO:19, or SEQ ID NO:21.

[0178] For example, an SHC enzyme variant may have an amino acid sequence having less than 100% identity with SEQ ID NO:1, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:20, SEQ ID NO:18, SEQ ID NO:19, or SEQ ID NO:21, for example, equal to or less than about 99.5%, or equal to or less than about 99.0%.

[0179] For example, the SHC enzyme variant may have about 70.0% to about 99.5%, or about 80.0% to about 99.0%, or about 85.0% to about 98.5%, or about 90.0% to about 98.0% of the same as SEQ ID NO:1, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:20, SEQ ID NO:18, SEQ ID NO:19, or SEQ ID NO:21.

[0180] The "percentage (%) identity" of a polypeptide or nucleotide sequence is defined, respectively, as the percentage of amino acids or nucleotides in a candidate sequence that are identical to those in a reference sequence after sequence alignment and the introduction of vacancies (if necessary) to achieve maximum percentage sequence identity, and without considering any conserved substitutions as part of sequence identity. Alignments for determining percentage sequence identity can be performed in various ways within the scope of the art, such as using publicly available computer software. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms required to achieve maximum alignment across the full length of the sequences being compared. The terms "polypeptide" and "protein" are used interchangeably herein and refer to any amino acid-linked peptide chain, regardless of length or post-translational modifications.

[0181] The similarity of nucleotide and amino acid sequences, i.e., the percentage of sequence identity, can be determined by sequence alignment. Such alignment can be performed using several algorithms known in the art, preferably the mathematical algorithm of Karlin and Altschul (Karlin & Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877), using hmmalign (HMMER package, http: / / hmmer.wustl.edu / ), or using, for example, the CLUSTAL algorithm (Thompson, JD, Higgins, DG & Gibson, TJ (1994) Nucleic Acids Res. 22, 4673-80) available at https: / / www.ebi.ac.uk / Tools / msa / clustalo / , or the GAP program (mathematical algorithm of the University of Iowa) or the mathematical algorithm of Myers and Miller (1989-Cabios 4:11-17). The preferred parameters used are the default parameters, as set at https: / / www.ebi.ac.uk / Tools / msa / clustalo / .

[0182] Percentage sequence identity can be calculated using algorithms such as BLAST, BLAT, or BlastZ (or BlastX). Similar algorithms were introduced in the BLASTN and BLASTP programs by Altschul et al. (1990) J. Mol. Biol. 215, 403-410. BLAST polynucleotide searches can be performed using the BLASTN program with a score of 100 and a word length of 12 to obtain polynucleotide sequences homologous to nucleic acids encoding the relevant protein. BLAST protein searches can be performed using the BLASTP program with a score of 50 and a word length of 3 to obtain amino acid sequences homologous to peptides.

[0183] To obtain gap alignments for comparative purposes, gap BLAST can be used as described by Altschul et al. (1997) Nucleic Acids Res. 25, 3389-3402. When using BLAST and gap BLAST procedures, the default parameters of each procedure are used. Sequence matching analysis can be supplemented by established homology mapping techniques such as Shuffle-LAGAN (Brudno M., Bioinformatics 2003b, 19Suppl 1:154-162) or Markov random fields. When percentages of sequence identity are referred to in this application, unless otherwise specified, these percentages are calculated relative to the full length of longer sequences.

[0184] In the specific implementation, CLUSTAL O (version 1.2.4) is used to determine the % identity between two sequences.

[0185] Compared to the parental SHC enzyme, the SHC enzyme variant may, for example, have increased enzymatic activity in converting compounds of formula (II) to compounds of formula (I). Increased enzyme activity can refer to any aspect of the enzymatic conversion of compounds of formula (II) to compounds of formula (I), including, for example, an increase in the overall conversion rate of compounds of formula (II) to compounds of formula (I), an increase in the conversion rate of compounds of formula (II) (e.g., in the first 4 hours, or the first 6 hours, or the first 12 hours of the reaction), an increase in the yield of compounds of formula (I), and / or a decrease in the yield of byproducts. Typically, increased enzyme activity can be defined by increased productivity. This can be defined based on the amount of compounds of formula (I) produced per hour, per gram of biocatalyst, and per liter of reaction.

[0186] Compared to the parental SHC enzyme, SHC enzyme variants can, for example, provide an increased conversion rate of compounds of formula (II). Therefore, the method described herein can have an increased conversion level of compounds of formula (II) compared to the method using the parental SHC enzyme. Compared to the parental SHC enzyme, SHC enzyme variants can, for example, provide an increased conversion rate of compounds of formula (II). Therefore, the method described herein can have an increased conversion rate of compounds of formula (II) compared to the parental SHC enzyme. Compared to the parental SHC enzyme, SHC enzyme variants can, for example, provide an increased conversion rate of compounds of formula (II) during the first 4 hours, 2 hours, 4 hours, 6 hours, 8 hours, 12 hours, or 24 hours of the reaction. Therefore, the method described herein can have an increased conversion rate of compounds of formula (II) during the first 2 hours, 4 hours, 6 hours, 8 hours, 12 hours, or 24 hours of the reaction compared to the parental SHC enzyme. This can be compared to using two enzymes (i.e., the SHC enzyme variant and the parental SHC enzyme) under the same reaction conditions (e.g., the same pH and temperature), or to using each enzyme under conditions that have been individually defined as optimal for its activity (e.g., optimized pH and temperature), and these conditions can differ from each other.

[0187] The conversion of the compound of formula (II) to the compound of formula (I) can be determined, for example, by the activity assay as described above, and can be calculated as the number of grams of recyclable product per gram of feedstock (which can be calculated as a percentage molar conversion if desired).

[0188] The method for preparing the compound of formula (I) disclosed herein may be carried out at the optimal temperature range or optimal temperature and / or optimal pH range or optimal pH and / or optimal SDS concentration range or optimal SDS concentration of the specific enzyme used, as described in the examples below.

[0189] Nucleic acids and methods for preparing nucleic acids

[0190] The SHC enzymes and enzyme variants described herein can be encoded by nucleic acid sequences. The nucleic acid containing the coding sequence can be, for example, an isolated nucleic acid.

[0191] Therefore, this article provides constructs comprising nucleic acid sequences encoding SHC enzymes or enzyme variants as described herein. As used herein, a “construct” is an artificially generated nucleic acid segment to be transfected into target cells. Constructs may contain nucleic acid sequences encoding SHC enzymes or enzyme variants and gene expression controllers (e.g., promoters).

[0192] This article further provides vectors comprising the constructs described herein. As used herein, a "vector" is a DNA molecule used as a medium to artificially carry foreign genetic material into a cell in which it can be replicated and / or expressed. Vectors can be, for example, plasmids, viral vectors, granules, or artificial chromosomes.

[0193] The terms “construction” and “vector” can overlap, for example, when the construct is a plasmid.

[0194] As used herein, the term "nucleic acid" or "nucleic acid molecule" shall specifically refer to the polynucleotides of this disclosure, which may be DNA, cDNA, genomic DNA, synthetic DNA, or RNA, and may be double-stranded or single-stranded, sense strands and / or antisense strands. The term "nucleic acid" or "nucleic acid molecule" shall be particularly applicable to polynucleotides as used herein, for example as full-length nucleotide sequences or fragments or portions thereof, which encode polypeptides (e.g., enzymes of metabolic pathways) or fragments or portions thereof having enzymatic activity.

[0195] The term also includes individual molecules, such as cDNA, where the corresponding genomic DNA has introns and thus a different sequence; genomic fragments lacking at least one flanking gene; cDNA or genomic DNA fragments produced by polymerase chain reaction (PCR) and lacking at least one flanking gene; restriction fragments lacking at least one flanking gene; DNA fragments encoding non-naturally occurring proteins, such as fusion proteins (e.g., His markers), mutant proteins, or fragments of a given protein; and nucleic acids as degenerate variants of cDNA or naturally occurring nucleic acids. Furthermore, it includes recombinant nucleotide sequences as part of a heterozygous gene (i.e., a gene encoding a non-naturally occurring fusion protein). Fusion proteins may have one or more amino acids (e.g., but not limited to histidine (His)) added to a protein, typically at the N-terminus, but also at the C-terminus or fused within a region of the protein. Such fusion proteins or fusion vectors encoding such proteins are typically used for three purposes: (i) increasing the yield of recombinant proteins; (ii) increasing the solubility of recombinant proteins; and (iii) aiding in the purification of recombinant proteins by providing ligands for affinity purification.

[0196] The term "nucleic acid sequence" also includes codon-optimized sequences suitable for expression in specific microbial host cells (e.g., *E. coli* host cells). As used herein, the term "codon-optimized" refers to a nucleic acid protein-coding sequence that, taking into account its specific codon usage, has been adapted for expression in specific prokaryotic or eukaryotic host cells, particularly bacterial host cells such as *E. coli* host cells, by, for example, replacing one or more, or preferably a large number of, codons with codons more frequently used in bacterial host cell genes (e.g., *E. coli* genes).

[0197] In this regard, the nucleotide sequence or gene encoding the reference amino acid sequence (e.g., SEQ ID NO:1 or SEQ ID NO:13) and its variants / derivatives can be the original nucleotide sequence or gene found in the source, or the gene can be codon-optimized for a selected host organism (e.g., Escherichia coli).

[0198] In another aspect, the nucleic acid sequences disclosed herein are operatively linked to expression control sequences that allow expression in prokaryotic and / or eukaryotic host cells. As used herein, “operatively linked” means incorporated into a genetic construct such that the expression control sequence effectively controls the expression of the coding sequence or the gene of interest. The transcriptional / translational regulatory elements mentioned above include, but are not limited to, inducible and non-inducible, constitutive, cell cycle-regulating, and metabolic-regulating promoters, enhancers, operons, silencers, repressors, and other elements known to those skilled in the art that drive or otherwise regulate gene expression. Such regulatory elements include, but are not limited to, regulatory elements that direct constitutive expression or allow inducible expression, such as the CUP-1 promoter, tet-repressors such as those used in tet-on or tet-off systems, lac systems, and trp systems. As an example, isopropyl β-D-1-galactosylthiopyranoside (IPTG) is an effective inducer of gene expression in concentrations ranging from, for example, 100 μM to 1.0 mM. This compound is a molecular mimic of isolaxose, a lactose metabolite that triggers transcription by the lac operon, and thus it is used to induce gene expression when a gene is under the control of the lac operon. Another example of a regulatory element that induces gene expression is lactose. Similarly, the nucleic acid molecules of this disclosure can form part of a heterozygous gene encoding an additional polypeptide sequence, which, for example, functions as a marker or reporter gene. Examples of markers or reporter genes include β-lactamases, chloramphenicol acetyltransferase (CAT), adenosine deaminase (ADA), glucosyl phosphate transferase (DHFR), hygromycin-β-phosphotransferase (HPH), thymidine kinase (TK), lacZ (encoding β-galactosidase), and xanthine-guanine phosphoribosyltransferase (XGPRT). As with many standard methods associated with the implementation of this disclosure, those skilled in the art will recognize other useful agents, such as additional sequences that can function as markers or reporter genes.

[0199] This document also provides recombinant polynucleotides encoding SHC enzymes or variants thereof, which can be inserted into vectors for gene expression and, optionally, enzyme purification. One type of vector is a plasmid representing a circular double-stranded DNA loop, with an additional DNA segment linked to the circular double-stranded DNA loop. Some vectors can control the expression of genes functionally linked to them. These vectors are called “expression vectors.” Expression vectors commonly used in recombinant DNA technologies are plasmid-type. Typically, expression vectors contain genes for producing wild-type or variant SHC enzymes or as described herein. In this specification, the terms “plasmid” and “vector” are used interchangeably because plasmids are the most commonly used vector type.

[0200] Such vectors may include DNA sequences, including but not limited to DNA sequences naturally occurring in non-host cells, DNA sequences that are not normally transcribed into RNA or translated into protein (“expressed”), and other genes or DNA sequences desired to be introduced into a non-recombinant host. It should be understood that the genome of the recombinant host described herein is typically enhanced by the stable introduction of one or more recombinant genes. However, autonomous or replicating plasmids or vectors may also be used within the scope of this disclosure. Furthermore, this disclosure may be implemented using low copy number (e.g., single copy) or high copy number (as exemplified herein) plasmids or vectors.

[0201] The vectors disclosed herein include plasmids, phage particles, phages, granules, artificial bacterial and yeast chromosomes, knockout or knock-in constructs, synthetic nucleic acid sequences or cartridges, and subgroups can be generated in the form of linear polynucleotides, plasmids, giant plasmids, synthetic or artificial chromosomes (e.g., plant, bacterial, mammalian or yeast artificial chromosomes).

[0202] Preferably, upon vector introduction, the protein encoded by the introduced polynucleotide is expressed intracellularly. The plasmid is typically a standard cloning vector, such as a bacterial multicopy plasmid. The substrate can be incorporated into the same or different plasmids. Typically, at least two different types of plasmids with different types of selection markers are used to allow selection of cells containing at least two types of vectors.

[0203] Typically, bacterial or yeast cells can be transformed with any one or more nucleotide sequences well known in the art. For in vivo recombination, a gene to be recombinated with a chromosome or other gene is used to transform the host using standard transformation techniques. In a suitable embodiment, the construct includes DNA that provides an origin of replication. Those skilled in the art can appropriately select the origin of replication. Depending on the nature of the gene, if the sequence is already present along with a gene that is itself operable as an origin of replication, a supplementary origin of replication may not be necessary.

[0204] Host cells, methods for preparing host cells, and methods for preparing compounds of formula (I) using host cells.

[0205] Recombinant host cells can be used in the methods described herein.

[0206] This article further provides a recombinant host cell comprising the nucleic acid sequence or construct or vector as described herein. This article further provides recombinant host cells that produce the SHC enzyme or enzyme variant as described herein.

[0207] The methods described herein for producing compounds of formula (I) may, for example, involve culturing recombinant host cells as described herein. As used herein, the term "culturing" refers to the process of proliferating living cells to produce SHC enzymes or enzyme variants as described herein, which can be used in methods for producing compounds of formula (I) as described herein.

[0208] When exogenous or heterologous DNA has been introduced into a cell, bacterial or yeast cells can be transformed by this type of DNA. The transformed DNA may or may not integrate, meaning it may be covalently linked to the cell's chromosome. For example, in prokaryotes and yeast, the transformed DNA can be maintained on accessory elements such as plasmids. In eukaryotic cells, stably transfected cells are those in which the transfected DNA has integrated into the chromosome, enabling it to be inherited by daughter cells through chromosomal replication. This stability is demonstrated by the ability of eukaryotic cells to establish cell lines or clones consisting of daughter cell populations containing the transformed DNA.

[0209] Typically, the introduced DNA is not initially present in the host acting as the DNA recipient, but within the scope of this disclosure, a DNA segment is isolated from a given host, and one or more additional copies of that DNA are subsequently introduced into the same host, for example, to enhance the production of a gene product or alter the expression pattern of a gene. In some cases, the introduced DNA will modify or even replace an endogenous gene or DNA sequence, for example, through homologous recombination or site-directed mutagenesis. Suitable recombinant hosts include microorganisms, plant cells, and plants.

[0210] A further feature of this disclosure is the recombinant host. The term "recombinant host" (also referred to as "genetically modified host cell" or "transgenic cell") means a host cell containing a heterologous nucleic acid or its genome that has been amplified by at least one incorporated DNA sequence. The host cells of this invention can be genetically engineered using polynucleotides or vectors as described above.

[0211] Host cells that can be used for the purposes of this disclosure include, but are not limited to, prokaryotic cells, such as bacteria (e.g., *Escherichia coli* and *Bacillus subtilis*), which can be transformed, for example, with recombinant phage DNA, plasmid DNA, bacterial artificial chromosome, or coliform DNA expression vectors containing, for example, the polynucleotide molecules of this disclosure; and simple eukaryotic cells such as yeast (e.g., *Saccharomyces* and *Pichia*), which can be transformed, for example, with recombinant yeast expression vectors containing, for example, the polynucleotide molecules of this disclosure. Depending on the host cell and the corresponding vector used to introduce the polynucleotides of this disclosure, the polynucleotides may be integrated into, for example, chromosomal or mitochondrial DNA, or may be maintained extrachromosomally, for example, in an appendage form, or may only be transiently contained within the cell.

[0212] The term "cell," as used herein, particularly in connection with genetic engineering and the introduction of one or more genes or assembled gene clusters into cells or the production of cells, should be understood to refer to any prokaryotic or eukaryotic cell. Concern is given to prokaryotic and eukaryotic host cells used for the purposes of this disclosure, including bacterial host cells such as *Escherichia coli* or *Bacillus* sp., yeast host cells such as *Saccharomyces cerevisiae*, insect host cells such as *Spodoptera frugiperda*, or human host cells such as *HeLa* and *Jurkat*.

[0213] Specifically, the cells are eukaryotic cells, preferably fungal, mammalian, or plant cells, or prokaryotic cells. Suitable eukaryotic cells include, for example, but not limited to, mammalian cells, yeast cells, or insect cells (including Sf9), amphibian cells (including melanocytes), or worm cells, including cells of the genus *Caenorhabditis* (including *Caenorhabditis elegans*). Suitable mammalian cells include, for example, but not limited to, COS cells (including Cos-1 and Cos-7), CHO cells, HEK293 cells, HEK293T cells, HEK293 T-RexTM cells, or other transfectable eukaryotic cell lines. Suitable bacterial cells include, but are not limited to, *Escherichia coli*.

[0214] Preferably, prokaryotic cells, such as Escherichia coli, Bacillus, Streptomyces, or mammalian cells, such as HeLa cells or Jurkat cells, or plant cells, such as Arabidopsis, can be used.

[0215] Cells can be selected, for example, from prokaryotic, yeast, plant, and / or insect host cells.

[0216] Preferably, the cells are Aspergillus sp. or fungal cells, and more preferably, they can be selected from the genera of yeast, Candida, Kluyveromyces, Hansenula, Schizosaccharomyces, Yarrowia, Pichia pastoris and Aspergillus.

[0217] Preferably, the cells are bacterial cells, such as those selected from the genera *Escherichia*, *Streptomyces*, *Bacillus*, *Pseudomonas*, *Lactobacillus*, and *Lactococcus*. For example, the bacteria may be *Escherichia coli*.

[0218] Preferably, the Escherichia coli host cell is an Escherichia coli host cell recognized by industry and regulatory agencies (including but not limited to Escherichia coli K12 host cells or Escherichia coli BL21 host cells).

[0219] A preferred host cell for use with this disclosure is *Escherichia coli*, which can be recombinantly prepared as described herein. Therefore, the recombinant host can be a recombinant *E. coli* host cell. Libraries exist that provide detailed computer models and other information on *E. coli* mutants, plasmids, metabolism, and other relevant data, allowing for the rational design of various modules to improve product yield.

[0220] In one embodiment, the recombinant Escherichia coli microorganism comprises a nucleotide sequence encoding an SHC enzyme or an enzyme variant (e.g., the microorganism comprises a nucleotide sequence selected from SEQ ID NO: 2, 4, 6, 8, 10, 12 and 14).

[0221] Preferably, the recombinant Escherichia coli microorganism comprises a vector construct as described herein. In another preferred embodiment, the recombinant Escherichia coli microorganism comprises a nucleotide sequence encoding the SHC enzyme and / or enzyme variant disclosed herein.

[0222] Another preferred host cell used in conjunction with this disclosure is brewer's yeast. Libraries exist containing mutants, plasmids, detailed metabolic computer models, and other information for use with brewer's yeast, allowing for the rational design of various modules to improve product yield. Methods for preparing recombinant brewer's yeast microorganisms are known.

[0223] Cell culture can be performed in a conventional manner. The culture medium may contain a carbon source, at least one nitrogen source, and inorganic salts, with vitamins added thereto. The composition of the culture medium may be that commonly used for culturing the microbial species under discussion. The carbon source used in the method of the present invention includes any molecule that can be recombined with host cell metabolism to promote the growth and / or production of the SHC enzyme of interest, which is used to convert the compound of formula (II) into the compound of formula (I). Examples of suitable carbon sources include, but are not limited to, sucrose (e.g., found in molasses), fructose, xylose, glycerol, glucose, cellulose, starch, cellobiose, or other glucose-containing polymers.

[0224] In implementation schemes using yeast as a host, carbon sources such as sucrose, fructose, xylose, ethanol, glycerol, and glucose are suitable. The carbon source can be provided to the host organism in batches or fed-batch throughout the culture period, or alternatively, another energy source, such as protein or protein hydrolysate, can be used.

[0225] The recombinant host cell microorganisms used in the methods of this disclosure can be proliferated in enriched culture media (e.g., LB medium, bacterial tryptone yeast extract medium, nutrient medium, etc.) under pH, temperature, and reaction conditions typically used for microbial proliferation. In one embodiment of this disclosure, culture is performed using a defined basic culture medium such as M9A.

[0226] The components of M9A medium include: 14 g / L KH2PO4, 16 g / L K2HPO4, 1 g / L Na3 citrate·2H2O, 7.5 g / L (NH4)2SO4, 0.25 g / L MgSO4·7H2O, 0.015 g / L CaCl2·2H2O, 5 g / L glucose, and 1.25 g / L yeast extract.

[0227] In another embodiment of this disclosure, a nutrient-rich culture medium, such as LB, is used. The components of LB medium include: 10 g / L tryptone, 5 g / L yeast extract, and 5 g / L NaCl. Other examples of mineral media and M9 mineral media are disclosed, for example, in US6524831B2 and US2003 / 0092143A1.

[0228] Another example of a basic culture medium can be prepared as follows: For a 350 ml culture: Add 307 ml of H2O to 35 ml of citric acid / phosphate stock solution (133 g / L KH2PO4, 40 g / L (NH4)2HPO4, 17 g / L citric acid·H2O, pH adjusted to 6.3), and adjust the pH to 6.8 with 32% NaOH if necessary. After autoclaving, add 0.850 ml of 50% MgSO4, 0.035 ml of trace element solution (see below), 0.035 ml of thiamine solution, and 7 ml of 20% glucose.

[0229] Trace element solution: 50g / l Na2EDTA.2H2O, 20g / l FeSO4.7H2O, 3g / l H3BO3, 0.9g / lMnSO4.2H2O, 1.1g / l CoCl2, 80g / L CuCl2, 240g / l NiSO4.7H2O, 100g / l KI, 1.4g / l (NH4)6Mo7O 24 A deionized aqueous solution of 4H2O and 1 g / L ZnSO4·7H2O.

[0230] Thiamine solution: 2.25 g / L thiamine·HCl deionized water solution.

[0231] MgSO4 solution: 50% (w / v) deionized aqueous solution of MgSO4·7H2O.

[0232] Recombinant microorganisms can be grown in batch, fed-batch, or continuous processes, or combinations thereof. Typically, the recombinant microorganisms are grown in a fermenter at a defined temperature in the presence of a suitable nutrient source (e.g., a carbon source) for a desired time period to produce sufficient SHC enzymes to convert the compound of formula (II) into the compound of formula (I) and produce the desired amount of the compound of formula (I). The recombinant host cells can be cultured in any suitable manner, such as by batch or fed-batch culture.

[0233] As used in this article, the term "batch culture" refers to a culture method in which neither the culture medium is added to nor removed during the culture period.

[0234] As used in this article, the term "feedback" refers to a culture method in which culture medium is added during the culture period but not removed.

[0235] One embodiment of this disclosure provides a method for producing a compound of formula (I) in a suitable cell system, the method comprising generating an SHC enzyme or enzyme variant in the cell system under suitable conditions, feeding a compound of formula (II) into the cell system, converting the compound of formula (II) into a compound of formula (I) using the SHC enzyme or enzyme variant generated in the cell system, collecting the compound of formula (I) from the cell system, and optionally isolating the compound of formula (I) from the system. Expression of other nucleotide sequences may be used to enhance the activity of the cell system used for the biotransformation of the compound of formula (I).

[0236] Another embodiment of this disclosure is a biotransformation method for preparing a compound of formula (I), the method comprising growing a host cell containing a gene encoding an SHC enzyme or an enzyme variant, generating an SHC enzyme or an enzyme variant in the host cell, feeding a compound of formula (II) into the host cell, incubating the host cell under pH, temperature, and solubilizing conditions suitable for promoting the conversion of the compound of formula (II) into the compound of formula (I), and collecting the compound of formula (I). The generation of an SHC enzyme or an enzyme variant in the host cell provides a method for preparing a compound of formula (I) when a compound of formula (II) is added to the host cell under suitable reaction conditions. The conversion rate obtained can be increased by adding more biocatalyst and SDS to the reaction mixture.

[0237] Recombinant host cell microorganisms can be cultured in a variety of ways to provide adequate quantities of cells that have produced SHC enzymes or enzyme variants for subsequent biotransformation steps. Because the microorganisms suitable for biotransformation steps vary widely (e.g., yeast, bacteria, and fungi), culture conditions are naturally tailored to the specific requirements of each species, and these conditions are well-known and documented. Any method known in the art for culturing cells of recombinant host cell microorganisms can be used to produce cells suitable for the subsequent biotransformation steps of this disclosure. Typically, cells are grown to a specific density (measurable as optical density (OD)) to produce sufficient biomass for the biotransformation reaction.

[0238] The chosen culture conditions affect not only the quantity of cells (biomass) obtained, but also the quality of the culture conditions, which in turn affects how the biomass becomes a biocatalyst. Recombinant host cell microorganisms that express SHC enzyme or enzyme variant genes and produce SHC enzymes or enzyme variants are called biocatalysts and are suitable for biotransformation reactions. In some embodiments, the biocatalyst is a recombinant whole cell that produces an SHC enzyme or enzyme variant, or it may be in suspension or immobilized form. In other embodiments, the biocatalyst is a membrane fraction or liquid fraction prepared from recombinant whole cells that produce SHC enzymes or enzyme variants (as disclosed, for example, in Seitz et al. 2012—cited above). Recombinant whole cells that produce SHC enzymes or enzyme variants include whole cells collected from a fermenter (for biotransformation reactions) or cells in a fermenter (which are then used for single-tank reactions). Recombinant whole cells that produce SHC enzymes or enzyme variants may include intact recombinant whole cells and / or cell debris. In either case, the SHC enzyme or enzyme variant is associated with a membrane (such as a cell membrane) in a certain way to receive a substrate (e.g., a compound of formula (II)) and / or interact with a substrate, and this membrane (e.g., a cell membrane) may be part of the whole cell (e.g., a recombinant whole cell). The SHC enzyme or enzyme variant may also be in an immobilized form (e.g., associated with an enzyme carrier) that allows the SHC enzyme or enzyme variant to interact with a substrate (e.g., a compound of formula (II)). The SHC enzyme or enzyme variant may also be used in a soluble form.

[0239] In one implementation, prior to the bioconversion step, the biocatalyst is produced in sufficient quantities (to produce sufficient biomass), harvested and washed (and optionally stored (e.g., freeze-dried)).

[0240] In another implementation, cells are produced in sufficient quantities (to generate enough biocatalyst), and then reaction conditions are adjusted so that the biocatalyst is used for the biotransformation reaction without the need for harvesting and washing. This one-step (or “single-tank”) method is advantageous because it simplifies the process. The culture medium used for growing cells is also suitable for the biotransformation reaction, provided that the reaction conditions are adjusted to promote the biotransformation reaction.

[0241] The optimal pH for cell growth is in the range of 6.0–8.0. The optimal pH for biotransformation depends on the type of SHC enzyme or enzyme variant used in the biotransformation reaction. pH should be adjusted using techniques well-known to those skilled in the art.

[0242] The biotransformation method disclosed herein is carried out under conditions of time, temperature, pH and solubilizer to provide the conversion of the compound of formula (II) to the compound of formula (I).

[0243] For the SHC wild-type enzyme or SHC enzyme variant under consideration, the pH of the reaction mixture can be in the range of 4-8, preferably 4.5-6.5, more preferably 4.5-6.5, and can be maintained by adding a buffer to the reaction mixture. Exemplary buffers for this purpose are citrate buffer or succinate buffer.

[0244] For the wild-type SHC enzyme or SHC enzyme variant under consideration, the temperature is approximately 15°C to approximately 60°C, for example, approximately 15°C to approximately 50°C, or approximately 15°C to approximately 45°C, or approximately 30°C to approximately 60°C, or approximately 35°C to approximately 55°C. During biotransformation, the temperature may be kept constant or may be varied.

[0245] When the ratio of the biocatalyst to the compound of formula (II) is about 2:1, the ratio of [SDS] / [cell] can be in the range of about 10:1 to 20:1, preferably about 15:1 to 18:1, and more preferably about 16:1.

[0246] The method for preparing the compounds of formula (I) disclosed herein can be carried out at optimal temperature, pH and surfactant concentration, which result in optimal activity for each individual SHC enzyme (wild type or variant) under consideration.

[0247] In biotransformation reactions, solubilizers (such as surfactants, detergents, solubility enhancers, water-miscible organic solvents, etc.) may be useful.

[0248] As used herein, the term "surfactant" refers to a component that reduces the surface tension (or interfacial tension) between two liquids or between a liquid and a solid. Surfactants can be used as detergents, wetting agents, emulsifiers, foaming agents, and dispersants. Examples of surfactants include, but are not limited to, Triton X-100, Tween 80, taurine deoxycholate, sodium taurine deoxycholate, sodium dodecyl sulfate (SDS), and / or sodium lauryl sulfate (SLS).

[0249] While Triton X-100 can be used for the partial purification of SHC enzymes or enzyme variants (in soluble or membrane fraction / suspension form), it can also be used for biotransformation reactions (see, for example, the disclosures in Seitz's 2012 PhD paper cited above, and in Neumann and Simon's (1986—cited above) and JP2009060799). SDS can be used as a solubilizer.

[0250] Without being bound by theory, the use of SDS in conjunction with recombinant microbial host cells may be advantageous, as SDS can preferentially interact with the host cell membrane, making it easier for SHC enzymes or enzyme variants (which are membrane-bound enzymes) to access the substrate of compound (II). Furthermore, the inclusion of appropriate levels of SDS in the reaction mixture can improve the properties of the emulsion (e.g., the compound of formula (II) in water) and / or improve the accessibility of the substrate of compound (II) to SHC enzymes within the host. The concentration of solubilizers (e.g., SDS) used in biotransformation reactions is influenced by both biomass and substrate concentration. That is, there is a degree of interdependence between solubilizer (e.g., SDS) concentration, biomass, and substrate concentration. As an example, as the concentration of the substrate of compound (II) increases, sufficient amounts of biocatalyst and solubilizer (e.g., SDS) are required for efficient biotransformation. For example, if the solubilizer (e.g., SDS) concentration is too low, suboptimal transformation of the compound of formula (II) may be observed. On the other hand, for example, if the concentration of the solubilizer (e.g., SDS) is too high, there is a risk that the biocatalyst may be affected by the destruction of intact microbial cells and / or the denaturation / inactivation of SHC enzymes or enzyme variants. The selection of an appropriate concentration of SDS in the context of biomass and substrate must be carefully investigated.

[0251] In some implementations, a biocatalyst to which a substrate of formula (II) is added is used to produce a compound of formula (I).

[0252] The substrate can be added by feeding using known devices (e.g., peristaltic pumps, infusion syringes, etc.). The compounds of formula (II) can be oil-soluble and provided in oil form. Given that the biocatalyst (microbial cells, such as intact recombinant whole cells and / or cell fragments and / or immobilized enzymes) is present in the aqueous phase, the biotransformation reaction can be considered a three-phase system (comprising an aqueous phase, a solid phase, and an oil phase) when the compounds of formula (II) are added to the biotransformation reaction mixture. This is true even in the presence of SDS.

[0253] Fermenters can be used to grow recombinant host cells that produce SHC enzymes or enzyme variant genes and produce active SHC enzymes or enzyme variants in the same fermenter container used to convert compounds of formula (II) into compounds of formula (I) (e.g., mixed with one or more byproducts of formula (III)) to a sufficient biomass concentration suitable for use as a biocatalyst.

[0254] Those skilled in the art will understand that higher cumulative production titers can be achieved by implementing continuous processes, such as product removal, substrate feeding, and biomass addition or (partial) replacement. Preferably, the biotransformation of compound (II) into compound (I) in the presence of a recombinant host cell containing an SHC enzyme or an enzyme variant produces compounds 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67. The yields of compounds of formula (I) of the numbers 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, and 100, expressed as mole percentages and based on the number of moles of compounds of formula (II) used; particularly preferably, the yields are 5 to 100, 10 to 100, and 100, 25 to 100, 30 to 100, 35 to 100, especially 40 to 100, 45 to 100, 50 to 100, 60 to 100, and 70 to 100 mol%.

[0255] The activity of SHC enzymes or enzyme variants is defined by the reaction rate as a percentage ((amount of product / (amount of product + amount of remaining feed)) × 100. Preferably, the bioconversion of compound (II) to compound (I) in the presence of SHC enzyme or enzyme variant produces compounds 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68. The yields of compounds of formula (I) of the numbers 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, and 100, expressed as mole percentages and based on the number of moles of compounds of formula (II) used; particularly preferably, the yields are 5 to 100, 10 to 100, 20 to 100, 25 to 100, 30 to 100, 35 to 100, especially 40 to 100, 45 to 100, 50 to 100, 60 to 100, and 70 to 100.

[0256] In a preferred embodiment of the invention, the yield and / or reaction rate are determined within a defined time period of, for example, 4, 6, 8, 10, 12, 16, 20, 24, 36, 48 or 72 hours, during which the compound of formula (II) is converted into the compound of formula (I) by a recombinant host cell containing a nucleotide sequence encoding an SHC enzyme or enzyme variant and having produced an SHC enzyme or enzyme variant.

[0257] In another embodiment, the reaction is carried out under precisely determined conditions, such as 25°C, 30°C, 40°C, 50°C, or 60°C. In particular, the yield and / or reaction rate are determined by reacting the compound of formula (II) with the SHC enzyme or enzyme variant of the present invention over a time period of 24-72 hours within a temperature range of about 35°C to about 55°C to convert the compound of formula (I).

[0258] In another embodiment of the invention, the recombinant host cell comprising the nucleotide sequence encoding an SHC enzyme variant is characterized in that, under the same conditions, preferably under conditions individually defined as optimal for the activity of the SHC enzyme in question, it exhibits a yield and / or reaction rate of 2-, 3-, 4-, 5-, 6-, 7-, 8-, 9-, 10-, 11-, 12-, 13-, 14-, 15-, 16-, 17-, 18-, 19-, 20-, 21-, 22-, 23-, 24-, 25-, 26-, 27-, 28-, 29-, 30-, 31-, 32-, 33-, 34-, 35-, 36-, 37-, 3 8-, 39-, 40-, 41-, 42-, 43-, 44-, 45-, 46-, 47-, 48-, 49-, 50-, 51-, 52-, 53-, 54-, 55-, 56-, 57-, 58-, 59-, 60-, 61-, 62-, 63-, 64-, 65-, 66-, 67-, 68-, 69-, 70-, 71-, 7 2, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 500, 1000 times or higher. The terminology used herein refers to, but is not limited to, reaction conditions such as pH, temperature, and the concentration of solubilizers (e.g., SDS).

[0259] The successful development of a biotransformation method for preparing compounds of formula (I) from compounds of formula (II) in recombinant Escherichia coli strains containing nucleotide sequences encoding wild-type / reference SHC or SHC variants can provide a low-cost and industrially economical method for the production of compounds of formula (I).

[0260] Throughout this specification and the following claims, unless the context otherwise requires, the word “comprise” and variations thereof, such as “comprises” and “comprising”, shall be understood to imply inclusion of the integer or step or group of integers or steps, but not to exclude any other integer or step or group of integers or steps. The term “comprise” also means “comprising” and “consisting of”, for example, a composition “comprising” X may consist of only X, or may include additional components, such as X+Y. It must also be noted that, as used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural indicators, unless expressly specified otherwise. As an example, the term “gene” or “enzyme” refers to “one or more genes” or “one or more enzymes.”

[0261] It should be understood that this disclosure is not limited to the specific methods, schemes, and reagents described herein, as these can be modified. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure, which will be defined only by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Conventional molecular biology, microbiology, and recombinant DNA techniques within the scope of this disclosure can be employed according to this disclosure.

[0262] The application of this disclosure is not limited to the construction details and component arrangements illustrated in the following description or examples in the accompanying drawings. This disclosure is capable of having other embodiments and can be implemented or carried out in various ways. Furthermore, the wording and terminology used herein are for descriptive purposes and should not be considered limiting. Preferably, the terms used herein are as defined in “A multilingual glossary of biotechnological terms: (IUPAC Recommendations)”, Leuenberger, H.G., Nagel, B., and Kolbl, H. eds. (1995), Helvetica Chimica Acta, CH-4010 Basel, Switzerland.

[0263] Several documents are referenced throughout the text of this specification. Each document cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, GenBank accession number serial number submissions, etc.) is incorporated herein by reference in its entirety, whether preceding or following it.

[0264] The embodiments described herein are examples of this disclosure and are not intended to limit it. Different embodiments of this disclosure have been described in accordance with this disclosure. Many variations and changes can be made to the techniques described and illustrated herein without departing from the spirit and scope of this disclosure. Therefore, it should be understood that these embodiments are merely illustrative and do not limit the scope of this disclosure. Example

[0265] Example 1 - Production of SHC enzyme

[0266] SHC plasmid preparation:

[0267] The gene encoding wild-type or variant squalene-hopaene cyclase (SHC) was inserted into plasmid pET-28a(+), which is under the control of the IPTG-inducible T7 promoter for protein production in *E. coli*. The plasmid was transformed into *E. coli* strain BL21(DE3) using a standard heat shock transformation protocol.

[0268] Culture medium preparation:

[0269] For a 350 ml culture, prepare the default basic culture medium as follows: Add 307 ml of H2O to 35 ml of citric acid / phosphate stock solution (133 g / L KH2PO4, 40 g / L (NH4)2HPO4, 17 g / L citric acid·H2O, pH adjusted to 6.3). Adjust the pH to 6.8 with 32% NaOH if necessary. After autoclaving, add 0.850 ml of 50% MgSO4, 0.035 ml of trace element solution (see below), 0.035 ml of thiamine solution, and 7 ml of 20% glucose.

[0270] Trace element solution: 50g / l Na2EDTA.2H2O, 20g / l FeSO4.7H2O, 3g / l H3BO3, 0.9g / lMnSO4.2H2O, 1.1g / l CoCl2, 80g / L CuCl2, 240g / l NiSO4.7H2O, 100g / l KI, 1.4g / l (NH4)6Mo7O 24 A deionized aqueous solution of 4H2O and 1 g / L ZnSO4·7H2O.

[0271] Thiamine solution: 2.25 g / L thiamine·HCl deionized water solution.

[0272] MgSO4 solution: 50% (w / v) deionized aqueous solution of MgSO4·7H2O.

[0273] SHC enzyme or enzyme variant production (biocatalyst production).

[0274] Small-scale production of biocatalysts (wild-type SHC or SHC variants)

[0275] Inoculate 350 ml of culture (supplemented with 50 μg / ml kanamycin) with a pre-culture of *E. coli* strain BL21(DE3) containing the SHC production plasmid. Allow the cells to grow at a constant stirring speed (250 rpm) until the optical density reaches approximately 0.5 (OD) at 37°C. 650nm ).

[0276] Protein production was then induced by adding IPTG to a concentration of 300 μM, followed by incubation with constant oscillation for 5–6 hours. The resulting biomass was finally collected by centrifugation and washed with, for example, 50 mM Tris-HCl buffer at pH 7.5. The cells were stored as a pellet at 4°C or -20°C until further use. Typically, 2.5 to 4 grams of cells (wet weight) are obtained from 1 liter of culture, regardless of the culture medium used.

[0277] Biocatalyst production in fermenters

[0278] Prepare and run fermentation in a 750 ml InforsHT reactor. Add 168 ml of deionized water to the fermentation vessel. The reaction vessel is equipped with all necessary probes (pO2, pH, sampling, defoamer), a C+N feed, and a sodium hydroxide bottle, and is autoclaved. After autoclaving, add the following components to the reactor:

[0279] 20ml 10x phosphate / citrate buffer

[0280] 14ml 50% glucose

[0281] 0.53 ml MgSO4 solution

[0282] 2ml (NH4)2SO4 solution

[0283] 0.020 ml trace element solution

[0284] 0.400ml thiamine solution

[0285] 0.200 ml kanamycin stock solution

[0286] Operating parameters were set as follows: pH = 6.95, pO2 = 40%, T = 30℃, stirring at 300 rpm. Cascade: rpm setpoint 300, min 300, max 1000; flow rate (L / min) setpoint 0.1, min 0, max 0.6. Defoamer ratio: 1:9.

[0287] The fermenter was inoculated from the seed culture to the OD. 650nm0.4-0.5. The seed culture was grown in LB medium (+kanamycin) at 37°C and 220 rpm for 8 h. Fermentation was initially run in batch mode for 11.5 h, followed by fed-batch feeding with a sterile glucose solution (143 ml H₂O + 35 g glucose) for C+N. After sterilization, the following were added: 17.5 ml (NH₄)₂SO₄ solution, 1.8 ml MgSO₄ solution, 0.018 ml trace element solution, 0.360 ml thiamine solution, and 0.180 ml kanamycin stock solution. The feed was run at a constant flow rate of approximately 4.2 ml / h. Glucose and NH₄ + External measurements were performed to assess the availability of C- and N-sources in the culture. Glucose levels were typically very low.

[0288] The cultures grew for a total of approximately 25 hours, during which they typically reached OD. 650nm 40-45. SHC production was then initiated by adding IPTG to a concentration of 1 mM in the fermenter (either as an IPTG pulse or using an infusion syringe over a 3-4 hour period), setting the temperature to 40°C and pO2 to 20%. SHC production was induced at 40°C for 16 hours. At the end of induction, cells were collected by centrifugation, washed with 0.1 M citrate / sodium citrate buffer (pH 5.4), and stored as a precipitate at 4°C or -20°C until further use.

[0289] Example 2 - GC Analysis

[0290] The sample was extracted with an appropriate volume of tert-butyl methyl ether (MBTE / tBME) to quantify its content in the substrate and reaction products. The solvent fraction was separated from the aqueous phase by centrifugation prior to gas chromatography analysis. 1 μL of the solvent phase was injected (split ratio 10) onto a 30 m x 0.32 mm x 0.25 μm Zebron ZB-Wax column. The column was developed at a constant flow rate (4 mL / min H₂) with a temperature gradient: 100 °C, 15 °C / min to 200 °C, 120 °C / min to 240 °C, 4 min, 240 °C. Inlet temperature: 250 °C, detector temperature: 250 °C. This resulted in the separation of the substrate and product peaks.

[0291] The conversion rate of hydroxyfarnesyacetone is calculated from the peak areas corresponding to the substrate and reaction product using the following formula:

[0292] Conversion rate (%) = 100 x (area) 产物峰 / (area 产物峰 +Area 底物峰 ))

[0293] Example 3 - Screening for cyclization of hydroxyfarethene acetone using SHC enzymes and enzyme variants

[0294] As described in Example 1, a set of SHC enzymes for the cyclization reaction of hydroxyfarethrin acetone was produced in Escherichia coli. The reaction system contained 1 g / L hydroxyfarethrin acetone and cell-to-OD... 650nm 10. Apply the reaction conditions listed in Table 2. An overview of the performance of the SHC enzymes tested under the optimal reaction conditions listed in Table 2 is shown below. Figure 2 middle.

[0295] Table 2. SHC enzymes and reaction conditions.

[0296]

[0297]

[0298] *In citrate / sodium phosphate buffer

[0299] All tested enzymes were able to cyclize hydroxyfarethionine to (+)-amylacetal. The conversion rates of wild-type enzymes ranged from 2% to approximately 90%, and were highest in the case of BmeSHC, such as... Figure 2 As shown. (For example, when used in...) Figure 2 The increased hydroxyfarethionylacetone cyclization was observed in the SHC variants 215G2 SHC, 115A7SHC, 110B8 SHC, 90C7 SHC, SHC#49, SHC#65, and SHC#66 summarized in the study when the mutation was introduced into wild-type AacSHC.

[0300] Example 4 - Cyclohydrazide of farnesylacetone using SHC enzyme

[0301] The cyclization of hydroxyfarylacetone in the reaction was tested using cells producing ZmoSHC1 or BmeSHC enzymes, respectively, in reactions containing 2 and 8 g / L of hydroxyfarylacetone. 600nm 80. The reaction was carried out in 50 mM succinic acid / NaOH buffer at pH 5.2 and incubated at 35 °C for 24 h. GC-FID analysis of the solvent-extracted reaction showed that, at a substrate concentration of 2 g / L, the conversion of hydroxyfarnesyacetone was 52% and 90% for ZmoSHC1 and BmeSHC, respectively. At 8 g / L, the conversion of hydroxyfarnesyacetone using ZmoSHC1 and BmeSHC was 17% and 79%, respectively.

[0302] Example 5 - Cyclohydration of hydroxyfarethion using a variant of squalene-hope cyclase

[0303] The production enzyme SHC variant 215G2, as summarized in Example 1, was used for the cyclization reaction of hydroxyfarethione.

[0304] A typical reaction (150 g total volume) was set up in a 0.75 L Infors fermenter as follows: Hydroxyfarnesyacetone (0.75 g, 2.7 mmol) was added to the reaction vessel, along with 1.95 g of SDS from a 31% (w / w) solution prepared in deionized water. A cell suspension was prepared from *E. coli* cells of the SHC variant of interest by suspending the cells in 0.1 M succinate / NaOH buffer at pH 5.1. After determining the wet weight concentration of the cells in the cell suspension by centrifugation at 10 °C and 17210 g for 10 min, an appropriate volume of cells was added to the reaction vessel to introduce 37.5 g of cells into the reaction. The reaction volume was brought to 150 g with the required amount of reaction buffer at pH 5.1. The reaction was carried out at 35 °C and pH 5.4 with constant stirring (700 rpm). The pH was set to 5.4 using 85% H3PO4. The pH was manually adjusted using 85% phosphate as needed. The reaction was sampled over time (1 ml) and extracted with 5 volumes of MTBE / tBME (5 ml). After clarifying the solvent phase by centrifugation (benchtop centrifuge, 13000 rpm, 2 min), the substrate and product contents of the reaction were determined by GC analysis.

[0305] After approximately 72 hours of reaction, the reaction mixture was extracted five times with 100 mL of MTBE by vigorous shaking, followed by phase separation by centrifugation (6000 g, 10 min, room temperature). The crude extract was filtered through silica gel and the filtrate was concentrated. The residue (2.4 g) was purified by silica gel column chromatography (eluting with a gradient of 7-70% ethyl acetate in n-heptane). (+)-Ambroxanal (the compound of formula (I), 173 mg, 23%, pale yellow solid) and (Z)-5-((5aR,9aR)-6,6,9a-trimethyloctahydrobenzo[b]oxahepta-3(2H)-ylidene)pentan-2-one (the compound of formula (III), 65 mg, 9%, off-white solid) were isolated from it.

[0306] Characterization of (+)-amyral. [α] D 23 = +16.4° (c = 0.42, CHCl3). TLC (silica gel, heptane / EtOAc3:2):R f =0.54.

[0307] 1H-NMR (400MHz, CDCl3)4.32(d,J=6.8Hz,1H),3.37(dd,J=7.1,1.2Hz,1H),1.90(dt,J=12.7,3.4Hz,1H),1.82(dd,J=13.7,4.6Hz,1H),1.49- 1.78(m,8H),1.43-1.47(m,1H),1.42(s,3H),1.10-1.24(m,3H),0.99(dd,J=12.2,2.0Hz,1H),0.90(s,6H),0.83-0.95(m,1H),0.82(s,3H). 13 C-NMR(101MHz, CDCl3):106.0(s),82.6(s),73.5(t),55.7(d),53.3(d),41.8(t),38.7(t),37.3( s),36.2(t),35.9(t),33.6(q),33.1(s),24.3(q),21.7(q),20.0(t),18.3(t),17.4(t),14.6(q).

[0308] EI-MS(70eV):278(M+,<1),263(<1),248(2),236(4),218(36),203(19),190(55),175(43) ,162(11),147(24),137(34),121(42),109(47),95(34),79(38),69(36),55(42),43(100).

[0309] Characterization of (Z)-5-((5aR,9aR)-6,6,9a-trimethyloctahydrobenzo[b]oxapron-3(2H)-ylidene)pentan-2-one

[0310] TLC (silica gel, heptane / EtOAc3:2):R f =0.29.

[0311] 1H-NMR (600MHz, benzene-d6) δppm5.01(br t,J=7.3Hz,1H),4.44(d,J=16.6Hz,1H),4.29(br d,J=16.2Hz,1H),2.47(ddd,J=12.4,7.9,3.8Hz,1H),2.22(ddd,J=12.7,8.6,4.3Hz,1H),2.08(q,J=7.2Hz,2H),),1.94(t, J=1.0Hz,2H),1.73-1.62(m,3H),1.61(s,3H),1.49-1.22(m,4H),1.21-1.14(m,2H),1.19(s,3H),0.87(s,3H),0.72(s,3H). 13 C-NMR (151MHz, benzene-d6) δppm 205.59,143.75,120.82,78.59,62.81,55.39,42.97,42.28,41.65,35.76,35.38,33.45,29.42,26.21,21.56,21.56,20.89,19.64.

[0312] EI-MS(70eV):278(M+,<1),260(<1),245(<1),220(1),141(16),135(15),123(24),109(32),95(25),81(19),69(12),55(13),43(100).

[0313] Example 5A - Cyclohydration of hydroxyfarethione using wild-type BmeSHC

[0314] The production of wild-type BmeSHC, as summarized in Example 1, and its use in the cyclization reaction of hydroxyfarethoxyacetone.

[0315] The reaction (4 mL volume) consisted of 135 g / L hydroxyfarnesyacetone, 221 g / L cells (wet weight), and 0.09% SDS in 0.2 M acetate / sodium acetate buffer, pH 5.2. The reaction mixture was incubated at 45 °C with constant stirring (650 rpm, Radleys Carousel). One reaction served as a control; samples were taken over time, extracted with MTBE, clarified by centrifugation (13000 rpm, 2 min), and the substrate and product concentrations were analyzed by GC. After 50 h of initiation, hydroxyfarnesyacetone conversion was complete or nearly complete (100%), with an expected concentration of approximately 920 mg of (+)-ambroacetal.

[0316] The combined reactants were centrifuged (4500 g, 4 °C, 15 min). The precipitate was recovered and washed three times with 20 mL of deionized water (vigorous shaking + centrifugation), and the aqueous phase was discarded. Finally, the precipitate was resuspended in 15 mL of deionized water and extracted three times with 15–20 mL of MTBE (vigorous shaking + centrifugation). The organic phase was collected and its (+)-ambroeal content was analyzed. The combined organic phase was filtered (silica gel), and the solvent was evaporated under a nitrogen stream to give approximately 720 mg of dry crystalline powder. The residue (720 mg) was purified by silica gel column chromatography, eluting with a gradient of 6–50% MTBE in n-heptane. The fractions containing pure (+)-ambroeal were combined, and the solvent was evaporated. 560 mg of (+)-ambroeal (formula (I)) was isolated (yield approximately 61%).

[0317] Characterization of (+)-amyral. [α] D 25 = +22.6° (c = 0.94, CHCl3). TLC (silica gel, n-heptane / MTBE 3:1):R f =0.55. NMR: The spectral data are consistent with those shown in Example 5.

[0318] Example 6 - Hydroxyfarnesyacetone (see also) Figure 1 )

[0319] Preparation of (4) (E)-6,10-dimethyl-1-((tetrahydro-2H-pyran-2-yl)oxy)undec-5,9-dien-2-one: 10% palladium on carbon (6.0 g, 10% w / w) was added fractionally to a solution of ethyl 4-(benzyloxy)-3-oxobutyrate (1, 60 g, 0.25 mol, 1 equivalent) in ethanol (600 mL) in a 1 L autoclave under a nitrogen atmosphere, and the mixture was stirred at 3 atm hydrogen for 12 h. The reaction mixture was filtered through a C celite bed and washed with a 1:1 mixture of dichloromethane and ethanol. The filtrate was evaporated under vacuum to give a pale yellow residue, which was dissolved in dichloromethane (600 mL). 3,4-dihydro-2H-pyran (42.7 g, 0.51 mol, 2 equivalents) and PPTS (6.3 g, 0.025 mol, 0.1 equivalents) were added to the solution, and the mixture was stirred at room temperature for 16 h. Water (500 mL) was then added, followed by extraction with dichloromethane (2 x 200 mL). The combined organic layers were washed with brine, dried over Na2SO4, and concentrated under vacuum to give a crude product, which was purified by silica gel column chromatography and eluted with ethyl acetate in petroleum ether (5-8%) to give (E)-5,9-dimethyl-2-(2-((tetrahydro-2H-pyran-2-yl)oxy)acetyl)dec-4,8-dienoic acid ethyl ester (2), a pale yellow liquid (45 g, 77%). The product (44 g, 0.19 mol, 1.0 equivalent) was dissolved in THF (440 mL), and potassium carbonate (31.6 g, 0.23 mol, 1.2 equivalent) was added. The suspension was stirred at room temperature for 1 h, and then (E)-1-bromo-3,7-dimethyloctyl-2,6-diene (geranyl bromide, 37.3 g, 0.172 mol, 0.9 equivalent) was added at 0 °C, and the mixture was stirred at room temperature for 16 h. The mixture was filtered, and the filter cake was washed with dichloromethane. The filtrate was concentrated under vacuum to give ethyl (E)-5,9-dimethyl-2-(2-((tetrahydro-2H-pyran-2-yl)oxy)acetyl)dec-4,8-dienoic acid (3, 70 g), which was dissolved in ethanol (500 mL). A solution of KOH (35 g) in water (100 mL) was added, and the mixture was heated to 80 °C for 2 h. The solution was concentrated under vacuum to obtain a residue, which was dissolved in dichloromethane (1 L) and washed with water (2 x 200 mL) and brine. The organic layer was dried over Na2SO4 and concentrated under vacuum to obtain a crude oil, which was purified by silica gel column chromatography using a petroleum ether solution of 5-8% ethyl acetate to give (E)-6,10-dimethyl-1-((tetrahydro-2H-pyran-2-yl)oxy)undec-5,9-dien-2-one (4, 20 g, 36%), as a pale yellow liquid.

[0320] 1b) Preparation of (3-(2-methyl-1,3-dioxolane-2-yl)propyl)triphenylphosphine iodide (8): Triphenylphosphine (66.8 g, 0.25 mol, 1.2 equivalent) was added to a solution of 5-iodopentan-2-one (45 g, 0.21 mol, 1 equivalent) in (200 mL) toluene. The mixture was stirred at 120 °C for 16 h and then cooled to room temperature, at which point a solid precipitate was formed. This precipitate was filtered, washed with ether (100 mL), and dried under vacuum to give (4-oxopentanyl)triphenylphosphine iodide (90 g, 90%) as a light brown solid. The product (40 g, 0.10 mol, 1 equivalent) was suspended in toluene (400 mL), and ethylene glycol (80 mL) and p-toluenesulfonic acid (1 g, 0.0167 mol, 0.16 equivalent) were added. The apparatus was equipped with a Dean-Stark condenser. The mixture was heated to 130 °C for 12 h, then cooled to room temperature and the toluene layer was discarded. The remaining rubbery residue was dissolved in dichloromethane (500 mL) and washed with water (500 mL) and brine (200 mL).

[0321] The organic layer was dried with Na2SO4 and concentrated under vacuum to obtain (3-(2-methyl-1,3-dioxolane-2-yl)propyl)triphenylphosphonium iodide (8, 43 g, 80%), a light brown solid.

[0322] 1c) Preparation of hydroxyfarnesyacetone: A suspension of (3-(2-methyl-1,3-dioxolane-2-yl)propyl)triphenylphosphonium iodide (52.8 g, 0.10 mol, 2 equivalents) in THF (150 mL) was cooled at -78 °C, and then a 1.6 M n-BuLi hexane solution (63.8 mL, 0.10 mol, 2 equivalents) was added. The mixture was stirred at room temperature for 30 min, at which point an orange-yellow suspension was formed, and then cooled to -78 °C. A solution of (E)-6,10-dimethyl-1-((tetrahydro-2H-pyran-2-yl)oxy)undec-5,9-dien-2-one (4,15 g, 0.05 mol, 1 equivalent) in THF (20 mL) was added dropwise. The mixture was stirred continuously at room temperature for 2 h. During this period, the reaction mixture became a yellow suspension. The reaction mixture was quenched with ice water (100 mL) and extracted with ethyl acetate (2 x 100 mL). The organic layer was dried over Na₂SO₄ and concentrated under vacuum to give a crude product, which was ground with hexane (4 x 50 mL), and the precipitated triphenylphosphine oxide was removed by filtration. After removing the solvent under vacuum, 2-(((2Z,5E)-6,10-dimethyl-2-(3-(2-methyl-1,3-dioxolane-2-yl)propylidene)undec-5,9-dien-1-yl)oxy)tetrahydro-2H-pyran (9, 22 g) was given as a pale yellow liquid, which, according to 1H-NMR, still contained trace amounts of triphenylphosphine oxide. The product (22 g, 0.05 mol, 1 equivalent) was dissolved in acetone (200 mL), and a 1.5 N HCl aqueous solution (220 mL) was added at 0 °C. The solution was stirred at room temperature for 16 h, then water (50 mL) was added, and the mixture was extracted with diethyl ether (2 x 100 mL). The combined organic layers were washed with brine, dried over Na₂SO₄, and concentrated under vacuum to give a crude oil, which was purified by silica gel column chromatography, eluted with 10-15% ethyl acetate in petroleum ether, to give (5Z,9E)-6-(hydroxymethyl)-10,14-dimethylpentadecano-5,9,13-trien-2-one (hydroxyfarnesylacetone, 7.5 g, GC / MS purity 81%). Volatile impurities were removed by distillation at 85 °C / 1 mm / Hg. The residue (5.8 g) was subjected to a second silica gel column chromatography, eluted with 10-15% ethyl acetate in petroleum ether, to give hydroxyfarnesylacetone (2.8 g, 19%) as a pale yellow liquid.

[0323] 13C-NMR(d6-DMSO,100MHz):208.4(s),140.2(s),134.7(s),131.1(s),125.0(d),124.6(d),124.6(d) ),58.5(t),43.6(t),39.7(2t),35.0(t),30.1(q),26.7(t),25.9(q),21.9(t),18.0(q),16.2(q).

[0324] sequence list

[0325] SEQ ID NO:1 (Amino acid sequence of wild-type Acidic Cyclocarya spp. SHC (AacSHC))

[0326] MAEQLVEAPAYARTLDRAVEYLLSCQKDEGYWWGPLLSNVTMEAEYVLLCHILDRVDRDRMEKIRRYLLHEQREDGTWALYPGGPPDLDTTIEAYVALKYIGMSRDEEPMQKALRFIQSQGGIESSRVFTRMWLALVGEYPWEKVPMVPPEIMFLGK RMPLNIYEFGSWARATVVALSIVMSRQPVFPLPERARVPELYETDVPPRRRGAKGGGGWIFDALDRALHGYQKLSVHPFRRAAEIRALDWLLERQAGDGSWGGIQPPWFYALIALKILDMTQHPAFIKGWEGLELYGVELDYGGWMFQASISPVWDTG LAVLALRAAGLPADHDRLVKAGEWLLDRQITVPGDWAVKRPNLKPGGFAFQFDNVYYPDVDDTAVVVWALNTLRLPDERRRRDAMTKGFRWIVGMQSSNGGWGAYDVDNTSDLPNHIPFCDFGEVTDPPSEDVTAHVLECFGSFGYDDAWKVIRRAVE YLKREQKPDGSWFGRWGVNYLYGTGAVVSALKAVGIDTREPYIQKALDWVEQHQNPDGGWGEDCRSYEDPAYAGKGASTPSQTAWALMALIAGGRAESEAARRGVQYLVETQRPDGGWDEPYYTGTGFPGDFYLGYTMYRHVFPTLALGRYKQAIERR

[0327] SEQ ID NO:2 (nucleotide sequence encoding wild-type AacSHC)

[0328]

[0329] SEQ ID NO: 3 (Amino acid sequence of AacSHC enzyme variant #65)

[0330] MAEQLVEAPAYARTLDRAVEYLLSCQKDEGYWWGPLLSNVTMEAEYVLLCHILDRVDRDRMEKIRRYLLHEQREDGTWALYPGGPPDLDTTIEAYVALKYIGMSRDEEPMQKALRFIQSQGGIESSRVFTRRWLALVGEYPWEKVPMVPPEIMFLGKRMPLNIYEFGSWARATVVALSIVMSRQPVFPLPERARVPELYETDVPPRRRGAKGGGGWIFDALDRVLHGYQKLSVHPFRRAAEIRALDWLLERQAGDGSWGGIQPPWFYALIALKILDMTQHPAFIKGWEGLELYGVELDYGGWMFQASISPVWDTGLAVLALRAAGLPADHDRLVKAGEWLLDRQITVPGDWAVKRPNLKPGGFAFQFDNVYYPDVDDTAVVVWALNTLRLPDERRRRDAMTKGFRWIVGMQSSNGGWGAYDVDNTSDLPNHTPFCDFGEVTDPPSEDVTAHVLECFGSFGYDDAWKVIRRAVEYLKREQKPDGSWFGRWGVNYLYGTGAVVSALKAVGIDTREPYIQKALDWVEQHQNPDGGWGEDCRSYEDPAYAGKGASTPSQTTWALMALIAGGRAESEAARRGVQYLVETQRPDGGWDEPYYTGTGFPGDFYLGYTMYSHVFPTLALGRYKQAIERR

[0331] SEQ ID NO: 4 (Nucleotide sequence encoding SHC enzyme variant #65)

[0332]

[0333] SEQ ID NO: 5 (Amino acid sequence of AacSHC enzyme variant #66)

[0334] MAEQLVEAPAYARTLDRAVEYLLSCQKDEGYWWGPLLSNVTMEAEYVLLCHILDRVDRDRMEKIRRYLLHEQREDGTWALHPGGPPDLDTTIEAYVALKYIGMSRDEEPMQKALRFIQSQGGIESSRVFTRRWLALVGEYPWEKVPMVPPEIMFLGKRMPLNIYEFGSWARATVVALSIVMSRQPVFPLPERARVPELYETDVPPRRRGAKGGGGWIFDALDRVLHGYQKLSVHPFRRAAEIRALDWLLERQAGDGSWGGIQPPWFYALIALKILDMTQHPAFIKGWEGLELYGVELDYGGWMFQASISPVWDTGLAVLALRAAGLPADHDRLVKAGEWLLDRQITVPGDWAVKRPNLKPGGFAFQFDNVYYPDVDDTAVVVWALNTLRLPDERRRRDAMTKGFRWIVGMQSSNGGWGAYDVDNTSDLPNHTPFCDFGEVTDPPSEDVTAHVLECFGSFGYDDAWKVIRRAVEYLKREQKPDGSWFGRWGVNYLYGTGAVVSALKAVGIDTREPYIQKALDWVEQHQNPDGGWGEDCRSYEDPAYAGKGASTPSQTTWALMALIAGGRAESEAARRGVQYLVETQRPDGGWDEPYYTGTGFPGDFYLGYTMYSHVFPTLALGRYKQAIERR

[0335] SEQ ID NO: 6 (Nucleotide sequence encoding SHC enzyme variant #66)

[0336]

[0337] SEQ ID NO: 7 (Amino acid sequence of AacSHC enzyme variant #90C7)

[0338] MAEQLVEAPAYARTLDRAVEYLLSCQKDEGYWWGPLLSNVTMEAEYVLLCHILDRVDRDRMEKIRRYLLHEQREDGTWALYPGGPPDLDATIEAYVALKYIGMSRDEEPMQKALRFIQSQGGIESSRVFTRRWLALVGEYPWEKVPMVPPEIMFLGKRMPLNIYEFGSWARATVVALSIVMSRQPVFPLPERARVPELYETDVPPRRRGAKGGGGWIFDALDRVLHGYQKLSVHPFRRAAEIRALDWLLERQAGDGSWGGIQPPWFYALIALKILDMTQHPAFIKGWEGLELYGVELDYGGWMFQASISPVWDTGLAVLALRAAGLPADHDRLVKAGEWLLDRQITVPGDWAVKRPNLKPGGFAFQFDNVYYPDVDDTAVVVWALNTLRLPDERRRRDAMTKGFRWIVGMQSSNGGWGAYDVDNTSDLPNHTPFCDFGEVTDPPSEDVTAHVLECFGSFGYDDAWKVIRRAVEYLKREQKPDGSWFGRWGVNYLYGTGAVVSALKAVGIDTREPYIQKALDWVEQHQNPDGGWGEDCRSYEDPAYAGKGASTPSQTAWALMALIAGGRAESEAARRGVQYLVETQRPDGGWDEPYYTGTGFPGDFYLGYTMYSHVFPTLALGRYKQAIERR

[0339] SEQ ID NO: 8 (Nucleotide sequence encoding SHC variant #90C7)

[0340]

[0341] SEQ ID NO: 9 (Amino acid sequence of AacSHC enzyme variant #110B8)

[0342] MAEQLVEAPAYARTLDRAVEYLLSCQKDEGYWWGPLLSNVTMEAEYVLLCHILDRVDRDRMEKIRRYLLHEQREDGTWALHPGGPPDLDTTIEAYVALKYIGMSRDEEPMQKALRFIQSQGGIESSRVFTRRWLALVGEYPWEKVPMVPPEIMFLGKRMPLNIYEFGSWARATVVALSIVMSRQPVFPLPERARVPELYETDVPPRRRGAKGGGGWIFDALDRVLHGYQKLSVHPFRRAAEIRALDWLLERQAGDGSWGGIQPPWFYALIALKILDMTQHPAFIKGWEGLELYGVELDYGGWMFQASISPVWDTGLAVLALRAAGLPADHDRLVKAGEWLLDRQITVPGDWAVKRPNLKPGGFAFQFDNVYYPDVDDTAVVVWALNTLRLPDERRRRDAMTKGFRWIVGMQSSNGGWGAYDVDNTSDLPNLTPFCDFGEVTDPPSEDVTAHVLECFGSFGYDDAWKVIRRAVEYLKREQKPDGSWFGRWGVNYLYGTGAVVSALKAVGIDTREPYIQKALDWVEQHQNPDGGWGEDCRSYEDPAYAGKGASTPSQTTWALMALIAGGRAESEAARRGVQYLVETQRPDGGWDEPYYTGTGFPGDFYLGYTMYRHVFPTLALGRYKQAIERR

[0343] SEQ ID NO: 10 (Nucleotide sequence encoding SHC enzyme variant #110B8)

[0344]

[0345] SEQ ID NO: 11 (amino acid sequence of AacSHC enzyme variant #115A7)

[0346] MAEQLVEAPAYARTLDRAVEYLLSCQKDEGYWWGPLLSNVTMEAEYVLLCHILDRVDRDRMEKIRRYLLHEQREDGTWALYPGGPPDLDTTIEAYVALKYIGMSRDEEPMQKALRFIQSQGGIESSRVFTRRWLALVGEYPWEKVPMVPPEIMFLGKRMPLNIYEFGSWARTTVVALSIVMSRQPVFPLPERARVPELYETDVPPRRRGAKGGGGWIFDALDRVLHGYQKLSVHPFRRAAEIRALDWLLERQAGDGSWGGIQPPWFYALIALKILDKTQHPAFIKGWEGLELYGVELDYGGWMFQASISPVWDTGLAVLALRAAGLPADHDRLVKAGEWLLDRQITVPGDWAVKRPNLKPGGFAFQFDNVYYPDVDDTAVVVWALNTLRLPDERRRRDAMTKGFRWIVGMQSSNGGWGAYDVDNTSDLPNHTPFCDFGEVTDPPSEDVTAHVLECFGSFGYDDAWKVIRRAVEYLKREQKPDGSWFGRWGVNYLYGTGAVVSALKAVGIDTREPYIQKALDWVEQHQNPDGGWGEDCRSYEDPAYAGKGASTPSQTAWALMALIAGGRAESEAARRGVQYLVETQRPDGGWDEPYYTGTGFPGDFYLGYTMYRHVFPTLALGRYKQAIERR

[0347] SEQ ID NO: 12 (nucleotide sequence encoding SHC variant #115A7)

[0348]

[0349] SEQ ID NO: 13 (Amino acid sequence of AacSHC enzyme variant 215G2)

[0350] MAEQLVEAPAYARTLDRAVEYLLSCQKDEGYWWGPLLSNVTMEAEYVLLCHILDRVDRDRMEKIRRYLLHEQREDGTWALYPGGPPDLDTTIEAYVALKYIGMSRDEEPMQKALRFIQSQGGIESSRVFTRRWLALVGEYPWEKVPMVPPEIMFLGKRMPLNIYEFGSWARATVVALSIVMSRQPVFPLPERARVPELYETDVPPRRRGAKGGGGWIFDALDRVLHGYQKLSVHPFRRAAEIRALDWLLERQAGDGSWGGIQPPWFYALIALKILDMTQHPAFIKGWEGLELYGVELDYGGWMFQASISPVWDTGLAVLALRAAGLPADHDRLVKAGEWLLDRQITVPGDWAVKRPNLKPGGFAFQFDNVYYPDVDDTAVVVWALNTLRLPDERRRRDAMTKGFRWIVGMQSSNGGWGAYDVDNTSDLPNHTPFCDFGEVTDPPSEDVTAHVLECFGSFGYDDAWKVIRRAVEYLKREQKPDGSWFGRWGVNYLYGTGAVVSALKAVGIDTREPYIQKALDWVEQHQNPDGGWGEDCRSYEDPAYAGKGASTPSQTAWALMALIAGGRAESEAARRGVQYLVETQRPDGGWDEPYYTGTGFPGDFYLGYTMYRHVFPTLALGRYKQAIERR

[0351] SEQ ID NO: 14 (Nucleotide sequence encoding Aac 215G2 SHC enzyme variant)

[0352]

[0353] SEQ ID NO: 15 (Amino acid sequence of wild-type ZmoSHC1)

[0354] MGIDRMNSLSRLLMKKIFGAEKTSYKPASDTIIGTDTLKRPNRRPEPTAKVDKTIFKTMGNSLNNTLVSACDWLIGQQKPDGHWVGAVESNASMEAEWCLALWFLGLEDHPLRPRLGNALLEMQREDGSWGVYFGAGNGDINATVEAYAALRSLGYSADNPVLKKAAAWIAEKGGLKNIRVFTRYWLALIGEWPWEKTPNLPPEIIWFPDNFVFSIYNFAQWARATMVPIAILSARRPSRPLRPQDRLDELFPEGRARFDYELPKKEGIDLWSQFFRTTDRGLHWVQSNLLKRNSLREAAIRHVLEWIIRHQDADGGWGGIQPPWVYGLMALHGEGYQLYHPVMAKALSALDDPGWRHDRGESSWIQATNSPVWDTMLALMALKDAKAEDRFTPEMDKAADWLLARQVKVKGDWSIKLPDVEPGGWAFEYANDRYPDTDDTAVALIALSSYRDKEEWQKKGVEDAITRGVNWLIAMQSECGGWGAFDKDNNRSILSKIPFCDFGESIDPPSVDVTAHVLEAFGTLGLSRDMPVIQKAIDYVRSEQEAEGAWFGRWGVNYIYGTGAVLPALAAIGEDMTQPYITKACDWLVAHQQEDGGWGESCSSYME

[0355] SEQ ID NO: 16 (Amino acid sequence of wild-type ZmoSHC2)

[0356] MTVSTSSAFHHSPLSDDVEPIIQKATRALLEKQQQDGHWVFELEADATIPAEYILLKHYLGEPEDLEIEAKIGRYLRRIQGEHGGWSLFYGGDLDLSATVKAYFALKMIGDSPDAPHMLRARNEILARGGAMRANVFTRIQLALFGAMSWEHVPQMPVELMLMPEWFPVHINKMAYWARTVLVPLLVLQALKPVARNRRGILVDELFVPDVLPTLQESGDPIWRRFFSALDKVLHKVEPYWPKNMRAKAIHSCVHFVTERLNGEDGLGAIYPAIANSVMMYDALGYPENHPERAIARRAVEKLMVLDGTEDQGDKEVYCQPCLSPIWDTALVAHAMLEVGGDEAEKSAISALSWLKPQQILDVKGDWAWRRPDLRPGGWAFQYRNDYYPDVDDTAVVTMAMDRAAKLSDLHDDFEESKARAMEWTIGMQSDNGGWGAFDANNSYTYLNNIPFADHGALLDPPTVDVSARCVSMMAQAGISITDPKMKAAVDYLLKEQEEDGSWFGRWGVNYIYGTWSALCALNVAALPHDHLAVQKAVAWLKTIQNEDGGWGENCDSYALDYSGYEPMDSTASQTAWALLGLMAVGEANSEAVTKGINWLAQNQDEEGLWKEDYYSGGGFPRVFYLRYHGYSKYFPLWALARYRNLKKANQPIVHYGM

[0357] SEQ ID NO: 17 (Amino acid sequence of wild-type BjaSHC)

[0358] MTVTSSASARATRDPGNYQTALQSTVRAAADWLIANQKPDGHWVGRAESNACMEAQWCLALWFMGLEDHPLRKRLGQSLLDSQRPDGAWQVYFGAPNGDINATVEAYAALRSLGFRDDEPAVRRAREWIEAKGGLRNIRVFTRYWLALIGEWPWEKTPNIPPEVIWFPLWFPFSIYNFAQWARATLMPIAVLSARRPSRPLPPENRLDALFPHGRKAFDYELPVKAGAGGWDRFFRGADKVLHKLQNLGNRLNLGLFRPAATSRVLEWMIRHQDFDGAWGGIQPPWIYGLMALYAEGYPLNHPVLAKGLDALNDPGWRVDVGDATYIQATNSPVWDTILTLLAFDDAGVLGDYPEAVDKAVDWVLQRQVRVPGDWSMKLPHVKPGGWAFEYANNYYPDTDDTAVALIALAPLRHDPKWKAKGIDEAIQLGVDWLIGMQSQGGGWGAFDKDNNQKILTKIPFCDYGEALDPPSVDVTAHIIEAFGKLGISRNHPSMVQALDYIRREQEPSGPWFGRWGVNYVYGTGAVLPALAAIGEDMTQPYIGRACDWLVAHQQADGGWGESCASYMDVSAVGRGTTTASQTAWALMALLAANRPQDKDAIERGCMWLVERQSAGTWDEPEFTGTGFPGYGVGQTIKLNDPALSQRLMQGPELSRAFMLRYGMYRHYFPLMALGRALRPQSHS

[0359] SEQ ID NO: 18 (Amino acid sequence of wild-type TelSHC)

[0360] MPTSLATAIDPKQLQQAIRASQDFLFSQQYAEGYWWAELESNVTMTAEVILLHKIWGTEQRLPLAKAEQYLRNHQRDHGGWELFYGDGGDLSTSVEAYMGLRLLGVPETDPALVKARQFILARGGISKTRIFTKLHLALIGCYDWRGIPSLPPWIMLLPEGSPFTIYEMSSWARSSTVPLLIVMDRKPVYGMDPPITLDELYSEGRANVVWELPRQGDWRDVFIGLDRVFKLFETLNIHPLREQGLKAAEEWVLERQEASGDWGGIIPAMLNSLLALRALDYAVDDPIVQRGMAAVDRFAIETETEYRVQPCVSPVWDTALVMRAMVDSGVAPDHPALVKAGEWLLSKQILDYGDWHIKNKKGRPGGWAFEFENRFYPDVDDTAVVVMALHAVTLPNENLKRRAIERAVAWIASMQCRPGGWAAFDVDNDQDWLNGIPYGDLKAMIDPNTADVTARVLEMVGRCQLAFDRVALDRALAYLRNEQEPEGCWFGRWGVNYLYGTSGVLTALSLVAPRYDRWRIRRAAEWLMQCQNADGGWGETCWSYHDPSLKGKGDSTASQTAWAIIGLLAAGDATGDYATEAIERGIAYLLETQRPDGTWHEDYFTGTGFPCHFYLKYHYYQQHFPLTALGRYARWRNLLAT

[0361] SEQ ID NO: 19 (Amino acid sequence of wild-type ApaSHC1

[0362] MNMASRFSLKKILRSGSDTQGTNVNTLIQSGTSDIVRQKPAPQEPADLSALKAMGNSLTHTLSSACEWLMKQQKPDGHWVGSVGSNASMEAEWCLALWFLGLEDHPLRPRLGKALLEMQRPDGSWGTYYGAGSGDINATVESYAALRSLGYAEDDPAVSKAAAWIISKGGLKNVRVFTRYWLALIGEWPWEKTPNLPPEIIWFPDNFVFSIYNFAQWARATMMPLAILSARRPSRPLRPQDRLDALFPGGRANFDYELPTKEGRDVIADFFRLADKGLHWLQSSFLKRAPSREAAIKYVLEWIIWHQDADGGWGGIQPPWVYGLMALHGEGYQFHHPVMAKALDALNDPGWRHDKGDASWIQATNSPVWDTMLSLMALHDANAEERFTPEMDKALDWLLSRQVRVKGDWSVKLPNTEPGGWAFEYANDRYPDTDDTAVALIAIASCRNRPEWQAKGVEEAIGRGVRWLVAMQSSCGGWGAFDKDNNKSILAKIPFCDFGEALDPPSVDVTAHVLEAFGLLGLPRDLPCIQRGLAYIRKEQDPTGPWFGRWGVNYLYGTGAVLPALAALGEDMTQPYISKACDWLINCQQENGGWGESCASYMEVSSIGHGATTPSQTAWALMGLIAANRPQDYEAIAKGCRYLIDLQEEDGSWNEEEFTGTGFPGYGVGQTIKLDDPAISKRLMQGAELSRAFMLRYDLYRQLFPIIALSRASRLIKLGN

[0363] SEQ ID NO: 20 (Amino acid sequence of wild-type GmoSHC)

[0364] MSPADISTKSSSFQRLDNMLPEAVSSACDWLIDQQKPDGHWVGPVESNACMEAQWCLALWFLGQEDHPLRPRLAQALLEMQREDGSWGIYVGADHGDINTTVEAYAALRSMGYAADMPIMAKSAAWIQQKGGLRNVRVFTRYWLALIGEWPWDKTPNLPPEIIWLPDNFIFSIYNFAQWARATMMPLTILSARRPSRPLLPENRLDGLFPEGRENFDYELPVKGEEDLWGRFFRAADKGLHSLQSFPVRRFVPREAAIRHVIEWIIRHQDADGGWGGIQPPWIYGLMALSVEGYPLHHPVLAKAMDALNDPGWRRDKGDASWIQATNSPVWDTMLAVLALHDAGAEDRYSPQMDKAIGWLLDRQVRVKGDWSIKLPDTEPGGWAFEYANDKYPDTDDTAVALIALAGCRHRPEWRERDIEGAISRGVNWLLAMQSSSGGWGAFDKDNNRSILTKIPFCDFGEALDPPSVDVTAHVLEAFGLLGISRNHPSVQKALAYIRSEQERNGAWFGRWGVNYVYGTGAVLPALAAIGEDMTQPYIVRACDWLMSVQQENGGWGESCASYMDINAVGHGVATASQTAWALIGLLAAKRPKDREAIARGCQFLIERQEDGSWTEEEYTGTGFPGYGVGQAIKLDDPSLPDRLLQGAELSRAFMLRYDLYRQYFPVMALSRARRMMKEDASAAA

[0365] SEQ ID NO: 21 (Amino acid sequence of wild-type BmeSHC)

[0366] MIILLKEVQLEIQRRIAYLRPTQKNDGSFRYCFETGVMPDAFLIMLLRTFDLDKEVLIKQLTERIVSLQNEDGLWTLFDDEEHNLSATIQAYTALLYSGYYQKNDRILRKAERYIIDSGGISRAHFLTRWMLSVNGLYEWPKLFYLPLSLLLVPTYVPLNFYELSTYARIHFVPMMVAGNKKFSLTSRHTPSLSHLDVREQKQESEETTQESRASIFLVDHLKQLASLPSYIHKLGYQAAERYMLERIEKDGTLYSYATSTFFMIYGLLALGYKKDSFVIQKAIDGICSLLSTCSGHVHVENSTSTVWDTALLSYALQEAGVPQQDPMIKGTTRYLKKRQHTKLGDWQFHNPNTAPGGWGFSDINTNNPDLDDTSAAIRALSRRAQTDTDYLESWQRGINWLLSMQNKDGGFAAFEKNTDSILFTYLPLENAKDAATDPATADLTGRVLECLGNFAGMNKSHPSIKAAVKWLFDHQLDNGSWYGRWGVCYIYGTWAAITGLRAVGVSASDPRIIKAINWLKSIQQEDGGFGESCYSASLKKYVPLSFSTPSQTAWALDALMTICPLKDQSVEKGIKFLLNPNLTEQQTHYPTGIGLPGQFYIQYHSYNDIFPLLALAHYAKKHSS

[0367] SEQ ID NO: 22 (amino acid sequence of AacSHC variant #49)

[0368] . sequence list <110> Givaudan Corporation <120> Enzyme-mediated method for preparing ambroxol or ambroxol homologues <130> 31049 PCT <160> twenty two <170> PatentIn version 3.5 <210> 1 <211> 631 <212> PRT <213> wild-type SHC enzymes from *Alicyclobacillus acidocaldarius* <400> 1 Met Ala Glu Gln Leu Val Glu Ala Pro Ala Tyr Ala Arg Thr Leu Asp 1 5 10 15 Arg Ala Val Glu Tyr Leu Leu Ser Cys Gln Lys Asp Glu Gly Tyr Trp 20 25 30 Trp Gly Pro Leu Leu Ser Asn Val Thr Met Glu Ala Glu Tyr Val Leu 35 40 45 Leu Cys His Ile Leu Asp Arg Val Asp Arg Asp Arg Met Glu Lys Ile 50 55 60 Arg Arg Tyr Leu Leu His Glu Gln Arg Glu Asp Gly Thr Trp Ala Leu 65 70 75 80 Tyr Pro Gly Gly Pro Pro Asp Leu Asp Thr Thr Ile Glu Ala Tyr Val 85 90 95 Ala Leu Lys Tyr Ile Gly Met Ser Arg Asp Glu Glu Pro Met Gln Lys 100 105 110 Ala Leu Arg Phe Ile Gln Ser Gln Gly Gly Ile Glu Ser Ser Arg Val 115 120 125 Phe Thr Arg Met Trp Leu Ala Leu Val Gly Glu Tyr Pro Trp Glu Lys 130 135 140 Val Pro Met Val Pro Pro Glu Ile Met Phe Leu Gly Lys Arg Met Pro 145 150 155 160 Leu Asn Ile Tyr Glu Phe Gly Ser Trp Ala Arg Ala Thr Val Val Ala 165 170 175 Leu Ser Ile Val Met Ser Arg Gln Pro Val Phe Pro Leu Pro Glu Arg 180 185 190 Ala Arg Val Pro Glu Leu Tyr Glu Thr Asp Val Pro Pro Arg Arg Arg 195 200 205 Gly Ala Lys Gly Gly Gly Gly Trp Ile Phe Asp Ala Leu Asp Arg Ala 210 215 220 Leu His Gly Tyr Gln Lys Leu Ser Val His Pro Phe Arg Arg Ala Ala 225 230 235 240 Glu Ile Arg Ala Leu Asp Trp Leu Leu Glu Arg Gln Ala Gly Asp Gly 245 250 255 Ser Trp Gly Gly Ile Gln Pro Pro Trp Phe Tyr Ala Leu Ile Ala Leu 260 265 270 Lys Ile Leu Asp Met Thr Gln His Pro Ala Phe Ile Lys Gly Trp Glu 275 280 285 Gly Leu Glu Leu Tyr Gly Val Glu Leu Asp Tyr Gly Gly Trp Met Phe 290 295 300 Gln Ala Ser Ile Ser Pro Val Trp Asp Thr Gly Leu Ala Val Leu Ala 305 310 315 320 Leu Arg Ala Ala Gly Leu Pro Ala Asp His Asp Arg Leu Val Lys Ala 325 330 335 Gly Glu Trp Leu Leu Asp Arg Gln Ile Thr Val Pro Gly Asp Trp Ala 340 345 350 Val Lys Arg Pro Asn Leu Lys Pro Gly Gly Phe Ala Phe Gln Phe Asp 355 360 365 Asn Val Tyr Tyr Pro Asp Val Asp Asp Thr Ala Val Val Val Trp Ala 370 375 380 Leu Asn Thr Leu Arg Leu Pro Asp Glu Arg Arg Arg Arg Asp Ala Met 385 390 395 400 Thr Lys Gly Phe Arg Trp Ile Val Gly Met Gln Ser Ser Asn Gly Gly 405 410 415 Trp Gly Ala Tyr Asp Val Asp Asn Thr Ser Asp Leu Pro Asn His Ile 420 425 430 Pro Phe Cys Asp Phe Gly Glu Val Thr Asp Pro Pro Ser Glu Asp Val 435 440 445 Thr Ala His Val Leu Glu Cys Phe Gly Ser Phe Gly Tyr Asp Asp Ala 450 455 460 Trp Lys Val Ile Arg Arg Ala Val Glu Tyr Leu Lys Arg Glu Gln Lys 465 470 475 480 Pro Asp Gly Ser Trp Phe Gly Arg Trp Gly Val Asn Tyr Leu Tyr Gly 485 490 495 Thr Gly Ala Val Val Ser Ala Leu Lys Ala Val Gly Ile Asp Thr Arg 500 505 510 Glu Pro Tyr Ile Gln Lys Ala Leu Asp Trp Val Glu Gln His Gln Asn 515 520 525 Pro Asp Gly Gly Trp Gly Glu Asp Cys Arg Ser Tyr Glu Asp Pro Ala 530 535 540 Tyr Ala Gly Lys Gly Ala Ser Thr Pro Ser Gln Thr Ala Trp Ala Leu 545 550 555 560 Met Ala Leu Ile Ala Gly Gly Arg Ala Glu Ser Glu Ala Ala Arg Arg 565 570 575 Gly Val Gln Tyr Leu Val Glu Thr Gln Arg Pro Asp Gly Gly Trp Asp 580 585 590 Glu Pro Tyr Tyr Thr Gly Thr Gly Phe Pro Gly Asp Phe Tyr Leu Gly 595 600 605 Tyr Thr Met Tyr Arg His Val Phe Pro Thr Leu Ala Leu Gly Arg Tyr 610 615 620 Lys Gln Ala Ile Glu Arg Arg 625 630 <210> 2 <211> 1896 <212> DNA <213> Alicyclobacillus acidocaldarius wild-type SHC enzyme <400> 2 atggctgagc agttggtgga agcgccggcc tacgcgcgga cgctggatcg cgcggtggag 60 tatctcctct cctgccaaaa ggacgaaggc tactggtggg ggccgcttct gagcaacgtc 120 acgatggaag cggagtacgt cctcttgtgc cacattctcg atcgcgtcga tcgggatcgc 180 atggagaaga tccggcggta cctgttgcac gagcagcgcg aggacggcac gtgggccctg 240 tacccgggtg ggccgccgga cctcgacacg accatcgagg cgtacgtcgc gctcaagtat 300 atcggcatgt cgcgcgacga ggagccgatg cagaaggcgc tccggttcat tcagagccag 360 ggcgggatcg agtcgtcgcg cgtgttcacg cggatgtggc tggcgctggt gggagaatat 420 ccgtgggaga aggtgcccat ggtcccgccg gagatcatgt tcctcggcaa gcgcatgccg 480 ctcaacatct acgagtttgg ctcgtgggct cgggcgaccg tcgtggcgct ctcgattgtg 540 atgagccgcc agccggtgtt cccgctgccc gagcgggcgc gcgtgcccga gctgtacgag 600 accgacgtgc ctccgcgccg gcgcggtgcc aagggagggg gtgggtggat cttcgacgcg 660 ctcgaccggg cgctgcacgg gtatcagaag ctgtcggtgc acccgttccg ccgcgcggcc 720 gagatccgcg ccttggactg gttgctcgag cgccaggccg gagacggcag ctggggcggg 780 attcagccgc cttggtttta cgcgctcatc gcgctcaaga ttctcgacat gacgcagcat 840 ccggcgttca tcaagggctg ggaaggtcta gagctgtacg gcgtggagct ggattacgga ggatggatgt ttcaggcttc catctcgccg gtgtgggaca cgggcctcgc cgtgctcgcg 960 ctgcgcgctg cggggcttcc ggccgatcac gaccgcttgg tcaaggcggg cgagtggctg ttggaccggc agatcacggt tccggggcgac tggggcggtga agccccgaa cctcaagccg 1080. ggcgggttcg cgttccagtt cgacaacgtg tactacccgg acgtggacga cacggccgtc gtggtgtggg cgctcaacac cctgcgcttg ccggacgagc gccgcaggcg ggacgccatg acgaagggat tccgctggat tgtcggcatg cagagctcga acggcggttg gggcgcctac 1260 gacgtcgaca acacgagcga tctcccgac cacatcccgt tctgcgactt cggcgaagtg accgatccgc cgtcagagga cgtcaccgcc cacgtgctcg agtgtttcgg cagcttcggg 1380 tacgatgacg cctggaaggt catccggcgc gcggtggaat atctcaagcg ggagcagaag 1440 ccggacggca gctggttcgg tcgttggggc gtcaattacc tctacggcac gggcgcggtg 1500 gtgtcggcgc tgaaggcggt cgggatcgac acgcgcgagc cgtacattca aaaggcgctc 1560 gactgggtcg agcagcatca gaacccggac ggcggctggg gcgaggactg ccgctcgtac 1620 gaggatccgg cgtacgcggg taagggcgcg agcaccccgt cgcagacggc ctgggcgctg 1680 atggcgctca tcgcgggcgg cagggcggag tccgaggccg cgcgccgcgg cgtgcaatac 1740 ctcgtggaga cgcagcgccc ggacggcggc tgggatgagc cgtactacac cggcacgggc 1800 ttcccagggg atttctacct cggctacacc atgtaccgcc acgtgtttcc gacgctcgcg 1860 ctcggccgct acaagcaagc catcgagcgc aggtga 1896 <210> 3 <211> 630 <212> PRT <213> Artificial sequence <220> <223> SHC enzyme variant #65 <400> 3 Met Ala Glu Gln Leu Val Glu Ala Pro Ala Tyr Ala Arg Thr Leu Asp 1 5 10 15 Arg Ala Val Glu Tyr Leu Leu Ser Cys Gln Lys Asp Glu Gly Tyr Trp 20 25 30 Trp Gly Pro Leu Leu Ser Asn Val Thr Met Glu Ala Glu Tyr Val Leu 35 40 45 Leu Cys His Ile Leu Asp Arg Val Asp Arg Asp Arg Met Glu Lys Ile 50 55 60 Arg Arg Tyr Leu Leu His Glu Gln Arg Glu Asp Gly Thr Trp Ala Leu 65 70 75 80 Tyr Pro Gly Gly Pro Pro Asp Leu Asp Thr Thr Ile Glu Ala Tyr Val 85 90 95 Ala Leu Lys Tyr Ile Gly Met Ser Arg Asp Glu Glu Pro Met Gln Lys 100 105 110 Ala Leu Arg Phe Ile Gln Ser Gln Gly Gly Ile Glu Ser Ser Arg Val 115 120 125 Phe Thr Arg Arg Trp Leu Ala Leu Val Gly Glu Tyr Pro Trp Glu Lys 130 135 140 Val Pro Met Val Pro Pro Glu Ile Met Phe Leu Gly Lys Arg Met Pro 145 150 155 160 Leu Asn Ile Tyr Glu Phe Gly Ser Trp Ala Arg Ala Thr Val Val Ala 165 170 175 Leu Ser Ile Val Met Ser Arg Gln Pro Val Phe Pro Leu Pro Glu Arg 180 185 190 Ala Arg Val Pro Glu Leu Tyr Glu Thr Asp Val Pro Pro Arg Arg Arg 195 200 205 Gly Ala Lys Gly Gly Gly Gly Trp Ile Phe Asp Ala Leu Asp Arg Val 210 215 220 Leu His Gly Tyr Gln Lys Leu Ser Val His Pro Phe Arg Arg Ala Ala 225 230 235 240 Glu Ile Arg Ala Leu Asp Trp Leu Leu Glu Arg Gln Ala Gly Asp Gly 245 250 255 Ser Trp Gly Gly Ile Gln Pro Pro Trp Phe Tyr Ala Leu Ile Ala Leu 260 265 270 Lys Ile Leu Asp Met Thr Gln His Pro Ala Phe Ile Lys Gly Trp Glu 275 280 285 Gly Leu Glu Leu Tyr Gly Val Glu Leu Asp Tyr Gly Gly Trp Met Phe 290 295 300 Gln Ala Ser Ile Ser Pro Val Trp Asp Thr Gly Leu Ala Val Leu Ala 305 310 315 320 Leu Arg Ala Ala Gly Leu Pro Ala Asp His Asp Arg Leu Val Lys Ala 325 330 335 Gly Glu Trp Leu Leu Asp Arg Gln Ile Thr Val Pro Gly Asp Trp Ala 340 345 350 Val Lys Arg Pro Asn Leu Lys Pro Gly Gly Phe Ala Phe Gln Phe Asp 355 360 365 Asn Val Tyr Tyr Pro Asp Val Asp Asp Thr Ala Val Val Val Trp Ala 370 375 380 Leu Asn Thr Leu Arg Leu Pro Asp Glu Arg Arg Arg Arg Asp Ala Met 385 390 395 400 Thr Lys Gly Phe Arg Trp Ile Val Gly Met Gln Ser Ser Asn Gly Gly 405 410 415 Trp Gly Ala Tyr Asp Val Asp Asn Thr Ser Asp Leu Pro Asn His Thr 420 425 430 Pro Phe Cys Asp Phe Gly Glu Val Thr Asp Pro Pro Ser Glu Asp Val 435 440 445 Thr Ala His Val Leu Glu Cys Phe Gly Ser Phe Gly Tyr Asp Asp Ala 450 455 460 Trp Lys Val Ile Arg Arg Ala Val Glu Tyr Leu Lys Arg Glu Gln Lys 465 470 475 480 Pro Asp Gly Ser Trp Phe Gly Arg Trp Gly Val Asn Tyr Leu Tyr Gly 485 490 495 Thr Gly Ala Val Val Ser Ala Leu Lys Ala Val Gly Ile Asp Thr Arg 500 505 510 Glu Pro Tyr Ile Gln Lys Ala Leu Asp Trp Val Glu Gln His Gln Asn 515 520 525 Pro Asp Gly Gly Trp Gly Glu Asp Cys Arg Ser Tyr Glu Asp Pro Ala 530 535 540 Tyr Ala Gly Lys Gly Ala Ser Thr Pro Ser Gln Thr Thr Trp Ala Leu 545 550 555 560 Met Ala Leu Ile Ala Gly Gly Arg Ala Glu Ser Glu Ala Ala Arg Arg 565 570 575 Gly Val Gln Tyr Leu Val Glu Thr Gln Arg Pro Asp Gly Gly Trp Asp 580 585 590 Glu Pro Tyr Tyr Thr Gly Thr Gly Phe Pro Gly Asp Phe Tyr Leu Gly 595 600 605 Tyr Thr Met Tyr Ser His Val Phe Pro Thr Leu Ala Leu Gly Arg Tyr 610 615 620 Lys Gln Ala Ile Glu Arg 625 630 <210> 4 <211> 1896 <212> DNA <213> Artificial Sequence <220> <223> SHC enzyme variant #65 <400> 4 atggctgagc agttggtgga agctccggcc tacgcgcgga cgctggatcg cgcggtggag 60 tatctcctct cctgccaaaa ggacgaaggc tactggtggg ggccgcttct gagcaacgtc 120 acgatggaag cggagtacgt cctcttgtgc cacattctcg atcgcgtcga tcgggatcgc 180 atggagaaga tccggcggta cctgttgcac gagcagcgcg aggacggcac gtgggccctg 240 tacccgggtg ggccgccgga cctcgacacg accatcgagg cgtacgtcgc gctcaagtat 300 atcggcatgt cgcgcgacga ggagccgatg cagaaggcgc tccggttcat tcagagccag 360 ggcgggatcg agtcgtcgcg cgtgttcacg cggaggtggc tggcgctggt gggagaatat 420 ccgtgggaga aggtgcccat ggtcccgccg gagatcatgt tcctcggcaa gcgcatgccg 480 ctcaacatct acgagtttgg ctcgtgggct cgggcgaccg tcgtggcgct ctcgattgtg 540 atgagccgcc agccggtgtt cccgctgccc gagcgggcgc gcgtgcccga gctgtacgag 600 accgacgtgc ctccgcgccg gcgcggtgcc aagggagggg gtgggtggat cttcgacgcg 660 ctcgaccggg tgctgcacgg gtatcagaag ctgtcggtgc acccgttccg ccgcgcggcc 720 gagatccgcg ccttggactg gttgctcgag cgccaggccg gagacggcag ctggggcggg 780 attcagccgc cttggtttta cgcgctcatc gcgctcaaga ttctcgacat gacgcagcat 840 ccggcgttca tcaagggctg ggaaggtcta gagctgtacg gcgtggagct ggattacgga ggatggatgt ttcaggcttc catctcgccg gtgtgggaca cgggcctcgc cgtgctcgcg 960 ctgcgcgctg cggggcttcc ggccgatcac gaccgcttgg tcaaggcggg cgagtggctg ttggaccggc agatcacggt tccggggcgac tggggcggtga agccccgaa cctcaagccg 1080. ggcgggttcg cgttccagtt cgacaacgtg tactacccgg acgtggacga cacggccgtc gtggtgtggg cgctcaacac cctgcgcttg ccggacgagc gccgcaggcg ggacgccatg acgaagggat tccgctggat tgtcggcatg cagagctcga acggcggttg gggcgcctac 1260 gacgtcgaca acacgagcga tctcccgac cacaccccgt tctgcgactt cggcgaagtg accgatccgc cgtcagagga cgtcaccgcc cacgtgctcg agtgtttcgg cagcttcggg 1380 tacgatgacg cctggaaggt catccggcgc gcggtggaat atctcaagcg ggagcagaag 1440 ccggacggca gctggttcgg tcgttggggc gtcaattacc tctacggcac gggcgcggtg 1500 gtgtcggcgc tgaaggcggt cgggatcgac acgcgcgagc cgtacattca aaaggcgctc 1560 gactgggtcg agcagcatca gaacccggac ggcggctggg gcgaggactg ccgctcgtac 1620 gaggatccgg cgtacgcggg taagggcgcg agcaccccgt cgcagacgac ctgggcgctg 1680 atggcgctca tcgcgggcgg cagggcggag tccgaggccg cgcgccgcgg cgtgcaatac 1740 ctcgtggaga cgcagcgccc ggacggcggc tgggatgagc cgtactacac cggcacgggc 1800 ttcccagggg atttctacct cggctacacc atgtacagcc acgtgtttcc gacgctcgcg 1860 ctcggccgct acaagcaagc catcgagcgc aggtga 1896 <210> 5 <211> 631 <212> PRT <213> Artificial sequence <220> <223> SHC enzyme variant #66 <400> 5 Met Ala Glu Gln Leu Val Glu Ala Pro Ala Tyr Ala Arg Thr Leu Asp 1 5 10 15 Arg Ala Val Glu Tyr Leu Leu Ser Cys Gln Lys Asp Glu Gly Tyr Trp 20 25 30 Trp Gly Pro Leu Leu Ser Asn Val Thr Met Glu Ala Glu Tyr Val Leu 35 40 45 Leu Cys His Ile Leu Asp Arg Val Asp Arg Asp Arg Met Glu Lys Ile 50 55 60 Arg Arg Tyr Leu Leu His Glu Gln Arg Glu Asp Gly Thr Trp Ala Leu 65 70 75 80 His Pro Gly Gly Pro Pro Asp Leu Asp Thr Thr Ile Glu Ala Tyr Val 85 90 95 Ala Leu Lys Tyr Ile Gly Met Ser Arg Asp Glu Glu Pro Met Gln Lys 100 105 110 Ala Leu Arg Phe Ile Gln Ser Gln Gly Gly Ile Glu Ser Ser Arg Val 115 120 125 Phe Thr Arg Arg Trp Leu Ala Leu Val Gly Glu Tyr Pro Trp Glu Lys 130 135 140 Val Pro Met Val Pro Pro Glu Ile Met Phe Leu Gly Lys Arg Met Pro 145 150 155 160 Leu Asn Ile Tyr Glu Phe Gly Ser Trp Ala Arg Ala Thr Val Val Ala 165 170 175 Leu Ser Ile Val Met Ser Arg Gln Pro Val Phe Pro Leu Pro Glu Arg 180 185 190 Ala Arg Val Pro Glu Leu Tyr Glu Thr Asp Val Pro Pro Arg Arg Arg 195 200 205 Gly Ala Lys Gly Gly Gly Gly Trp Ile Phe Asp Ala Leu Asp Arg Val 210 215 220 Leu His Gly Tyr Gln Lys Leu Ser Val His Pro Phe Arg Arg Ala Ala 225 230 235 240 Glu Ile Arg Ala Leu Asp Trp Leu Leu Glu Arg Gln Ala Gly Asp Gly 245 250 255 Ser Trp Gly Gly Ile Gln Pro Pro Trp Phe Tyr Ala Leu Ile Ala Leu 260 265 270 Lys Ile Leu Asp Met Thr Gln His Pro Ala Phe Ile Lys Gly Trp Glu 275 280 285 Gly Leu Glu Leu Tyr Gly Val Glu Leu Asp Tyr Gly Gly Trp Met Phe 290 295 300 Gln Ala Ser Ile Ser Pro Val Trp Asp Thr Gly Leu Ala Val Leu Ala 305 310 315 320 Leu Arg Ala Ala Gly Leu Pro Ala Asp His Asp Arg Leu Val Lys Ala 325 330 335 Gly Glu Trp Leu Leu Asp Arg Gln Ile Thr Val Pro Gly Asp Trp Ala 340 345 350 Val Lys Arg Pro Asn Leu Lys Pro Gly Gly Phe Ala Phe Gln Phe Asp 355 360 365 Asn Val Tyr Tyr Pro Asp Val Asp Asp Thr Ala Val Val Val Trp Ala 370 375 380 Leu Asn Thr Leu Arg Leu Pro Asp Glu Arg Arg Arg Arg Asp Ala Met 385 390 395 400 Thr Lys Gly Phe Arg Trp Ile Val Gly Met Gln Ser Ser Asn Gly Gly 405 410 415 Trp Gly Ala Tyr Asp Val Asp Asn Thr Ser Asp Leu Pro Asn His Thr 420 425 430 Pro Phe Cys Asp Phe Gly Glu Val Thr Asp Pro Pro Ser Glu Asp Val 435 440 445 Thr Ala His Val Leu Glu Cys Phe Gly Ser Phe Gly Tyr Asp Asp Ala 450 455 460 Trp Lys Val Ile Arg Arg Ala Val Glu Tyr Leu Lys Arg Glu Gln Lys 465 470 475 480 Pro Asp Gly Ser Trp Phe Gly Arg Trp Gly Val Asn Tyr Leu Tyr Gly 485 490 495 Thr Gly Ala Val Val Ser Ala Leu Lys Ala Val Gly Ile Asp Thr Arg 500 505 510 Glu Pro Tyr Ile Gln Lys Ala Leu Asp Trp Val Glu Gln His Gln Asn 515 520 525 Pro Asp Gly Gly Trp Gly Glu Asp Cys Arg Ser Tyr Glu Asp Pro Ala 530 535 540 Tyr Ala Gly Lys Gly Ala Ser Thr Pro Ser Gln Thr Thr Trp Ala Leu 545 550 555 560 Met Ala Leu Ile Ala Gly Gly Arg Ala Glu Ser Glu Ala Ala Arg Arg 565 570 575 Gly Val Gln Tyr Leu Val Glu Thr Gln Arg Pro Asp Gly Gly Trp Asp 580 585 590 Glu Pro Tyr Tyr Thr Gly Thr Gly Phe Pro Gly Asp Phe Tyr Leu Gly 595 600 605 Tyr Thr Met Tyr Ser His Val Phe Pro Thr Leu Ala Leu Gly Arg Tyr 610 615 620 Lys Gln Ala Ile Glu Arg Arg 625 630 <210> 6 <211> 1896 <212> DNA <213> Artificial sequence <220> <223> SHC enzyme variant #66 <400> 6 atggctgagc agttggtgga agctccggcc tacgcgcgga cgctggatcg cgcggtggag 60 tatctcctct cctgccaaaa ggacgaaggc tactggtggg ggccgcttct gagcaacgtc 120 acgatggaag cggagtacgt cctcttgtgc cacattctcg atcgcgtcga tcgggatcgc 180 atggagaaga tccggcggta cctgttgcac gagcagcgcg aggacggcac gtgggccctg 240 cacccgggtg ggccgccgga cctcgacacg accatcgagg cgtacgtcgc gctcaagtat 300 atcggcatgt cgcgcgacga ggagccgatg cagaaggcgc tccggttcat tcagagccag 360 ggcgggatcg agtcgtcgcg cgtgttcacg cggaggtggc tggcgctggt gggagaatat 420 ccgtgggaga aggtgcccat ggtcccgccg gagatcatgt tcctcggcaa gcgcatgccg 480 ctcaacatct acgagtttgg ctcgtgggct cgggcgaccg tcgtggcgct ctcgattgtg 540 atgagccgcc agccggtgtt cccgctgccc gagcgggcgc gcgtgcccga gctgtacgag 600 accgacgtgc ctccgcgccg gcgcggtgcc aagggagggg gtgggtggat cttcgacgcg 660 ctcgaccggg tgctgcacgg gtatcagaag ctgtcggtgc acccgttccg ccgcgcggcc 720 gagatccgcg ccttggactg gttgctcgag cgccaggccg gagacggcag ctggggcggg 780 attcagccgc cttggtttta cgcgctcatc gcgctcaaga ttctcgacat gacgcagcat 840 ccggcgttca tcaagggctg ggaaggtcta gagctgtacg gcgtggagct ggattacgga ggatggatgt ttcaggcttc catctcgccg gtgtgggaca cgggcctcgc cgtgctcgcg 960 ctgcgcgctg cggggcttcc ggccgatcac gaccgcttgg tcaaggcggg cgagtggctg ttggaccggc agatcacggt tccggggcgac tggggcggtga agccccgaa cctcaagccg 1080. ggcgggttcg cgttccagtt cgacaacgtg tactacccgg acgtggacga cacggccgtc gtggtgtggg cgctcaacac cctgcgcttg ccggacgagc gccgcaggcg ggacgccatg acgaagggat tccgctggat tgtcggcatg cagagctcga acggcggttg gggcgcctac 1260 gacgtcgaca acacgagcga tctcccgac cacaccccgt tctgcgactt cggcgaagtg accgatccgc cgtcagagga cgtcaccgcc cacgtgctcg agtgtttcgg cagcttcggg 1380 tacgatgacg cctggaaggt catccggcgc gcggtggaat atctcaagcg ggagcagaag 1440 ccggacggca gctggttcgg tcgttggggc gtcaattacc tctacggcac gggcgcggtg 1500 gtgtcggcgc tgaaggcggt cgggatcgac acgcgcgagc cgtacattca aaaggcgctc 1560 gactgggtcg agcagcatca gaacccggac ggcggctggg gcgaggactg ccgctcgtac 1620 gaggatccgg cgtacgcggg taagggcgcg agcaccccgt cgcagacgac ctgggcgctg 1680 atggcgctca tcgcgggcgg cagggcggag tccgaggccg cgcgccgcgg cgtgcaatac 1740 ctcgtggaga cgcagcgccc ggacggcggc tgggatgagc cgtactacac cggcacgggc 1800 ttcccagggg atttctacct cggctacacc atgtacagcc acgtgtttcc gacgctcgcg 1860 ctcggccgct acaagcaagc catcgagcgc aggtga 1896 <210> 7 <211> 631 <212> PRT <213> Artificial sequence <220> <223> SHC enzyme variant #90C7 <400> 7 Met Ala Glu Gln Leu Val Glu Ala Pro Ala Tyr Ala Arg Thr Leu Asp 1 5 10 15 Arg Ala Val Glu Tyr Leu Leu Ser Cys Gln Lys Asp Glu Gly Tyr Trp 20 25 30 Trp Gly Pro Leu Leu Ser Asn Val Thr Met Glu Ala Glu Tyr Val Leu 35 40 45 Leu Cys His Ile Leu Asp Arg Val Asp Arg Asp Arg Met Glu Lys Ile 50 55 60 Arg Arg Tyr Leu Leu His Glu Gln Arg Glu Asp Gly Thr Trp Ala Leu 65 70 75 80 Tyr Pro Gly Gly Pro Pro Asp Leu Asp Ala Thr Ile Glu Ala Tyr Val 85 90 95 Ala Leu Lys Tyr Ile Gly Met Ser Arg Asp Glu Glu Pro Met Gln Lys 100 105 110 Ala Leu Arg Phe Ile Gln Ser Gln Gly Gly Ile Glu Ser Ser Arg Val 115 120 125 Phe Thr Arg Arg Trp Leu Ala Leu Val Gly Glu Tyr Pro Trp Glu Lys 130 135 140 Val Pro Met Val Pro Pro Glu Ile Met Phe Leu Gly Lys Arg Met Pro 145 150 155 160 Leu Asn Ile Tyr Glu Phe Gly Ser Trp Ala Arg Ala Thr Val Val Ala 165 170 175 Leu Ser Ile Val Met Ser Arg Gln Pro Val Phe Pro Leu Pro Glu Arg 180 185 190 Ala Arg Val Pro Glu Leu Tyr Glu Thr Asp Val Pro Pro Arg Arg Arg 195 200 205 Gly Ala Lys Gly Gly Gly Gly Trp Ile Phe Asp Ala Leu Asp Arg Val 210 215 220 Leu His Gly Tyr Gln Lys Leu Ser Val His Pro Phe Arg Arg Ala Ala 225 230 235 240 Glu Ile Arg Ala Leu Asp Trp Leu Leu Glu Arg Gln Ala Gly Asp Gly 245 250 255 Ser Trp Gly Gly Ile Gln Pro Pro Trp Phe Tyr Ala Leu Ile Ala Leu 260 265 270 Lys Ile Leu Asp Met Thr Gln His Pro Ala Phe Ile Lys Gly Trp Glu 275 280 285 Gly Leu Glu Leu Tyr Gly Val Glu Leu Asp Tyr Gly Gly Trp Met Phe 290 295 300 Gln Ala Ser Ile Ser Pro Val Trp Asp Thr Gly Leu Ala Val Leu Ala 305 310 315 320 Leu Arg Ala Ala Gly Leu Pro Ala Asp His Asp Arg Leu Val Lys Ala 325 330 335 Gly Glu Trp Leu Leu Asp Arg Gln Ile Thr Val Pro Gly Asp Trp Ala 340 345 350 Val Lys Arg Pro Asn Leu Lys Pro Gly Gly Phe Ala Phe Gln Phe Asp 355 360 365 Asn Val Tyr Tyr Pro Asp Val Asp Asp Thr Ala Val Val Val Trp Ala 370 375 380 Leu Asn Thr Leu Arg Leu Pro Asp Glu Arg Arg Arg Arg Asp Ala Met 385 390 395 400 Thr Lys Gly Phe Arg Trp Ile Val Gly Met Gln Ser Ser Asn Gly Gly 405 410 415 Trp Gly Ala Tyr Asp Val Asp Asn Thr Ser Asp Leu Pro Asn His Thr 420 425 430 Pro Phe Cys Asp Phe Gly Glu Val Thr Asp Pro Pro Ser Glu Asp Val 435 440 445 Thr Ala His Val Leu Glu Cys Phe Gly Ser Phe Gly Tyr Asp Asp Ala 450 455 460 Trp Lys Val Ile Arg Arg Ala Val Glu Tyr Leu Lys Arg Glu Gln Lys 465 470 475 480 Pro Asp Gly Ser Trp Phe Gly Arg Trp Gly Val Asn Tyr Leu Tyr Gly 485 490 495 Thr Gly Ala Val Val Ser Ala Leu Lys Ala Val Gly Ile Asp Thr Arg 500 505 510 Glu Pro Tyr Ile Gln Lys Ala Leu Asp Trp Val Glu Gln His Gln Asn 515 520 525 Pro Asp Gly Gly Trp Gly Glu Asp Cys Arg Ser Tyr Glu Asp Pro Ala 530 535 540 Tyr Ala Gly Lys Gly Ala Ser Thr Pro Ser Gln Thr Ala Trp Ala Leu 545 550 555 560 Met Ala Leu Ile Ala Gly Gly Arg Ala Glu Ser Glu Ala Ala Arg Arg 565 570 575 Gly Val Gln Tyr Leu Val Glu Thr Gln Arg Pro Asp Gly Gly Trp Asp 580 585 590 Glu Pro Tyr Tyr Thr Gly Thr Gly Phe Pro Gly Asp Phe Tyr Leu Gly 595 600 605 Tyr Thr Met Tyr Ser His Val Phe Pro Thr Leu Ala Leu Gly Arg Tyr 610 615 620 Lys Gln Ala Ile Glu Arg Arg 625 630 <210> 8 <211> 1896 <212> DNA <213> Artificial sequence <220> <223> SHC enzyme variant #90C7 <400> 8 atggctgagc agttggtgga agctccggcc tacgcgcgga cgctggatcg cgcggtggag 60 tatctcctct cctgccaaaa ggacgaaggc tactggtggg ggccgcttct gagcaacgtc 120 acgatggaag cggagtacgt cctcttgtgc cacattctcg atcgcgtcga tcgggatcgc 180 atggagaaga tccggcggta cctgttgcac gagcagcgcg aggacggcac gtgggccctg 240 tacccgggtg ggccgccgga cctcgacgcg accatcgagg cgtacgtcgc gctcaagtat 300 atcggcatgt cgcgcgacga ggagccgatg cagaaggcgc tccggttcat tcagagccag 360 ggcgggatcg agtcgtcgcg cgtgttcacg cggaggtggc tggcgctggt gggagaatat 420 ccgtgggaga aggtgcccat ggtcccgccg gagatcatgt tcctcggcaa gcgcatgccg 480 ctcaacatct acgagtttgg ctcgtgggct cgggcgaccg tcgtggcgct ctcgattgtg 540 atgagccgcc agccggtgtt cccgctgccc gagcgggcgc gcgtgcccga gctgtacgag 600 accgacgtgc ctccgcgccg gcgcggtgcc aagggagggg gtgggtggat cttcgacgcg 660 ctcgaccggg tgctgcacgg gtatcagaag ctgtcggtgc acccgttccg ccgcgcggcc 720 gagatccgcg ccttggactg gttgctcgag cgccaggccg gagacggcag ctggggcggg 780 attcagccgc cttggtttta cgcgctcatc gcgctcaaga ttctcgacat gacgcagcat 840 ccggcgttca tcaagggctg ggaaggtcta gagctgtacg gcgtggagct ggattacgga ggatggatgt ttcaggcttc catctcgccg gtgtgggaca cgggcctcgc cgtgctcgcg 960 ctgcgcgctg cggggcttcc ggccgatcac gaccgcttgg tcaaggcggg cgagtggctg ttggaccggc agatcacggt tccggggcgac tggggcggtga agccccgaa cctcaagccg 1080. ggcgggttcg cgttccagtt cgacaacgtg tactacccgg acgtggacga cacggccgtc gtggtgtggg cgctcaacac cctgcgcttg ccggacgagc gccgcaggcg ggacgccatg acgaagggat tccgctggat tgtcggcatg cagagctcga acggcggttg gggcgcctac 1260 gacgtcgaca acacgagcga tctcccgac cacaccccgt tctgcgactt cggcgaagtg accgatccgc cgtcagagga cgtcaccgcc cacgtgctcg agtgtttcgg cagcttcggg 1380 tacgatgacg cctggaaggt catccggcgc gcggtggaat atctcaagcg ggagcagaag 1440 ccggacggca gctggttcgg tcgttggggc gtcaattacc tctacggcac gggcgcggtg 1500 gtgtcggcgc tgaaggcggt cgggatcgac acgcgcgagc cgtacattca aaaggcgctc 1560 gactgggtcg agcagcatca gaacccggac ggcggctggg gcgaggactg ccgctcgtac 1620 gaggatccgg cgtacgcggg taagggcgcg agcaccccgt cgcagacggc ctgggcgctg 1680 atggcgctca tcgcgggcgg cagggcggag tccgaggccg cgcgccgcgg cgtgcaatac 1740 ctcgtggaga cgcagcgccc ggacggcggc tgggatgagc cgtactacac cggcacgggc 1800 ttcccagggg atttctacct cggctacacc atgtacagcc acgtgtttcc gacgctcgcg 1860 ctcggccgct acaagcaagc catcgagcgc aggtga 1896 <210> 9 <211> 631 <212> PRT <213> Artificial sequence <220> <223> SHC enzyme variant #110B8 <400> 9 Met Ala Glu Gln Leu Val Glu Ala Pro Ala Tyr Ala Arg Thr Leu Asp 1 5 10 15 Arg Ala Val Glu Tyr Leu Leu Ser Cys Gln Lys Asp Glu Gly Tyr Trp 20 25 30 Trp Gly Pro Leu Leu Ser Asn Val Thr Met Glu Ala Glu Tyr Val Leu 35 40 45 Leu Cys His Ile Leu Asp Arg Val Asp Arg Asp Arg Met Glu Lys Ile 50 55 60 Arg Arg Tyr Leu Leu His Glu Gln Arg Glu Asp Gly Thr Trp Ala Leu 65 70 75 80 His Pro Gly Gly Pro Pro Asp Leu Asp Thr Thr Ile Glu Ala Tyr Val 85 90 95 Ala Leu Lys Tyr Ile Gly Met Ser Arg Asp Glu Glu Pro Met Gln Lys 100 105 110 Ala Leu Arg Phe Ile Gln Ser Gln Gly Gly Ile Glu Ser Ser Arg Val 115 120 125 Phe Thr Arg Arg Trp Leu Ala Leu Val Gly Glu Tyr Pro Trp Glu Lys 130 135 140 Val Pro Met Val Pro Pro Glu Ile Met Phe Leu Gly Lys Arg Met Pro 145 150 155 160 Leu Asn Ile Tyr Glu Phe Gly Ser Trp Ala Arg Ala Thr Val Val Ala 165 170 175 Leu Ser Ile Val Met Ser Arg Gln Pro Val Phe Pro Leu Pro Glu Arg 180 185 190 Ala Arg Val Pro Glu Leu Tyr Glu Thr Asp Val Pro Pro Arg Arg Arg 195 200 205 Gly Ala Lys Gly Gly Gly Gly Trp Ile Phe Asp Ala Leu Asp Arg Val 210 215 220 Leu His Gly Tyr Gln Lys Leu Ser Val His Pro Phe Arg Arg Ala Ala 225 230 235 240 Glu Ile Arg Ala Leu Asp Trp Leu Leu Glu Arg Gln Ala Gly Asp Gly 245 250 255 Ser Trp Gly Gly Ile Gln Pro Pro Trp Phe Tyr Ala Leu Ile Ala Leu 260 265 270 Lys Ile Leu Asp Met Thr Gln His Pro Ala Phe Ile Lys Gly Trp Glu 275 280 285 Gly Leu Glu Leu Tyr Gly Val Glu Leu Asp Tyr Gly Gly Trp Met Phe 290 295 300 Gln Ala Ser Ile Ser Pro Val Trp Asp Thr Gly Leu Ala Val Leu Ala 305 310 315 320 Leu Arg Ala Ala Gly Leu Pro Ala Asp His Asp Arg Leu Val Lys Ala 325 330 335 Gly Glu Trp Leu Leu Asp Arg Gln Ile Thr Val Pro Gly Asp Trp Ala 340 345 350 Val Lys Arg Pro Asn Leu Lys Pro Gly Gly Phe Ala Phe Gln Phe Asp 355 360 365 Asn Val Tyr Tyr Pro Asp Val Asp Asp Thr Ala Val Val Val Trp Ala 370 375 380 Leu Asn Thr Leu Arg Leu Pro Asp Glu Arg Arg Arg Arg Asp Ala Met 385 390 395 400 Thr Lys Gly Phe Arg Trp Ile Val Gly Met Gln Ser Ser Asn Gly Gly 405 410 415 Trp Gly Ala Tyr Asp Val Asp Asn Thr Ser Asp Leu Pro Asn Leu Thr 420 425 430 Pro Phe Cys Asp Phe Gly Glu Val Thr Asp Pro Pro Ser Glu Asp Val 435 440 445 Thr Ala His Val Leu Glu Cys Phe Gly Ser Phe Gly Tyr Asp Asp Ala 450 455 460 Trp Lys Val Ile Arg Arg Ala Val Glu Tyr Leu Lys Arg Glu Gln Lys 465 470 475 480 Pro Asp Gly Ser Trp Phe Gly Arg Trp Gly Val Asn Tyr Leu Tyr Gly 485 490 495 Thr Gly Ala Val Val Ser Ala Leu Lys Ala Val Gly Ile Asp Thr Arg 500 505 510 Glu Pro Tyr Ile Gln Lys Ala Leu Asp Trp Val Glu Gln His Gln Asn 515 520 525 Pro Asp Gly Gly Trp Gly Glu Asp Cys Arg Ser Tyr Glu Asp Pro Ala 530 535 540 Tyr Ala Gly Lys Gly Ala Ser Thr Pro Ser Gln Thr Thr Trp Ala Leu 545 550 555 560 Met Ala Leu Ile Ala Gly Gly Arg Ala Glu Ser Glu Ala Ala Arg Arg 565 570 575 Gly Val Gln Tyr Leu Val Glu Thr Gln Arg Pro Asp Gly Gly Trp Asp 580 585 590 Glu Pro Tyr Tyr Thr Gly Thr Gly Phe Pro Gly Asp Phe Tyr Leu Gly 595 600 605 Tyr Thr Met Tyr Arg His Val Phe Pro Thr Leu Ala Leu Gly Arg Tyr 610 615 620 Lys Gln Ala Ile Glu Arg Arg 625 630 <210> 10 <211> 1896 <212> DNA <213> Artificial sequence <220> <223> SHC enzyme variant #110B8 <400> 10 atggctgagc agttggtgga agctccggcc tacgcgcgga cgctggatcg cgcggtggag 60 tatctcctct cctgccaaaa ggacgaaggc tactggtggg ggccgcttct gagcaacgtc 120 acgatggaag cggagtacgt cctcttgtgc cacattctcg atcgcgtcga tcgggatcgc 180 atggagaaga tccggcggta cctgttgcac gagcagcgcg aggacggcac gtgggccctg 240 cacccgggtg ggccgccgga cctcgacacg accatcgagg cgtacgtcgc gctcaagtat 300 atcggcatgt cgcgcgacga ggagccgatg cagaaggcgc tccggttcat tcagagccag 360 ggcgggatcg agtcgtcgcg cgtgttcacg cggaggtggc tggcgctggt gggagaatat 420 ccgtgggaga aggtgcccat ggtcccgccg gagatcatgt tcctcggcaa gcgcatgccg 480 ctcaacatct acgagtttgg ctcgtgggct cgggcgaccg tcgtggcgct ctcgattgtg 540 atgagccgcc agccggtgtt cccgctgccc gagcgggcgc gcgtgcccga gctgtacgag 600 accgacgtgc ctccgcgccg gcgcggtgcc aagggagggg gtgggtggat cttcgacgcg 660 ctcgaccggg tgctgcacgg gtatcagaag ctgtcggtgc acccgttccg ccgcgcggcc 720 gagatccgcg ccttggactg gttgctcgag cgccaggccg gagacggcag ctggggcggg 780 attcagccgc cttggtttta cgcgctcatc gcgctcaaga ttctcgacat gacgcagcat 840 ccggcgttca tcaagggctg ggaaggtcta gagctgtacg gcgtggagct ggattacgga ggatggatgt ttcaggcttc catctcgccg gtgtgggaca cgggcctcgc cgtgctcgcg 960 ctgcgcgctg cggggcttcc ggccgatcac gaccgcttgg tcaaggcggg cgagtggctg ttggaccggc agatcacggt tccggggcgac tggggcggtga agccccgaa cctcaagccg 1080. ggcgggttcg cgttccagtt cgacaacgtg tactacccgg acgtggacga cacggccgtc gtggtgtggg cgctcaacac cctgcgcttg ccggacgagc gccgcaggcg ggacgccatg acgaagggat tccgctggat tgtcggcatg cagagctcga acggcggttg gggcgcctac 1260 gacgtcgaca acacgagcga tctcccgac ctcaccccgt tctgcgactt cggcgaagtg accgatccgc cgtcagagga cgtcaccgcc cacgtgctcg agtgtttcgg cagcttcggg 1380 tacgatgacg cctggaaggt catccggcgc gcggtggaat atctcaagcg ggagcagaag 1440 ccggacggca gctggttcgg tcgttggggc gtcaattacc tctacggcac gggcgcggtg 1500 gtgtcggcgc tgaaggcggt cgggatcgac acgcgcgagc cgtacattca aaaggcgctc 1560 gactgggtcg agcagcatca gaacccggac ggcggctggg gcgaggactg ccgctcgtac 1620 gaggatccgg cgtacgcggg taagggcgcg agcaccccgt cgcagacgac ctgggcgctg 1680 atggcgctca tcgcgggcgg cagggcggag tccgaggccg cgcgccgcgg cgtgcaatac 1740 ctcgtggaga cgcagcgccc ggacggcggc tgggatgagc cgtactacac cggcacgggc 1800 ttcccagggg atttctacct cggctacacc atgtaccgcc acgtgtttcc gacgctcgcg 1860 ctcggccgct acaagcaagc catcgagcgc aggtga 1896 <210> 11 <211> 631 <212> PRT <213> Artificial Sequence <220> <223> SHC Enzyme Variant #115A7 <400> 11 Met Ala Glu Gln Leu Val Glu Ala Pro Ala Tyr Ala Arg Thr Leu Asp 1 5 10 15 Arg Ala Val Glu Tyr Leu Leu Ser Cys Gln Lys Asp Glu Gly Tyr Trp 20 25 30 Trp Gly Pro Leu Leu Ser Asn Val Thr Met Glu Ala Glu Tyr Val Leu 35 40 45 Leu Cys His Ile Leu Asp Arg Val Asp Arg Asp Arg Met Glu Lys Ile 50 55 60 Arg Arg Tyr Leu Leu His Glu Gln Arg Glu Asp Gly Thr Trp Ala Leu 65 70 75 80 Tyr Pro Gly Gly Pro Pro Asp Leu Asp Thr Thr Ile Glu Ala Tyr Val 85 90 95 Ala Leu Lys Tyr Ile Gly Met Ser Arg Asp Glu Glu Pro Met Gln Lys 100 105 110 Ala Leu Arg Phe Ile Gln Ser Gln Gly Gly Ile Glu Ser Ser Arg Val 115 120 125 Phe Thr Arg Arg Trp Leu Ala Leu Val Gly Glu Tyr Pro Trp Glu Lys 130 135 140 Val Pro Met Val Pro Pro Glu Ile Met Phe Leu Gly Lys Arg Met Pro 145 150 155 160 Leu Asn Ile Tyr Glu Phe Gly Ser Trp Ala Arg Thr Thr Val Val Ala 165 170 175 Leu Ser Ile Val Met Ser Arg Gln Pro Val Phe Pro Leu Pro Glu Arg 180 185 190 Ala Arg Val Pro Glu Leu Tyr Glu Thr Asp Val Pro Pro Arg Arg Arg 195 200 205 Gly Ala Lys Gly Gly Gly Gly Trp Ile Phe Asp Ala Leu Asp Arg Val 210 215 220 Leu His Gly Tyr Gln Lys Leu Ser Val His Pro Phe Arg Arg Ala Ala 225 230 235 240 Glu Ile Arg Ala Leu Asp Trp Leu Leu Glu Arg Gln Ala Gly Asp Gly 245 250 255 Ser Trp Gly Gly Ile Gln Pro Pro Trp Phe Tyr Ala Leu Ile Ala Leu 260 265 270 Lys Ile Leu Asp Lys Thr Gln His Pro Ala Phe Ile Lys Gly Trp Glu 275 280 285 Gly Leu Glu Leu Tyr Gly Val Glu Leu Asp Tyr Gly Gly Trp Met Phe 290 295 300 Gln Ala Ser Ile Ser Pro Val Trp Asp Thr Gly Leu Ala Val Leu Ala 305 310 315 320 Leu Arg Ala Ala Gly Leu Pro Ala Asp His Asp Arg Leu Val Lys Ala 325 330 335 Gly Glu Trp Leu Leu Asp Arg Gln Ile Thr Val Pro Gly Asp Trp Ala 340 345 350 Val Lys Arg Pro Asn Leu Lys Pro Gly Gly Phe Ala Phe Gln Phe Asp 355 360 365 Asn Val Tyr Tyr Pro Asp Val Asp Asp Thr Ala Val Val Val Trp Ala 370 375 380 Leu Asn Thr Leu Arg Leu Pro Asp Glu Arg Arg Arg Arg Asp Ala Met 385 390 395 400 Thr Lys Gly Phe Arg Trp Ile Val Gly Met Gln Ser Ser Asn Gly Gly 405 410 415 Trp Gly Ala Tyr Asp Val Asp Asn Thr Ser Asp Leu Pro Asn His Thr 420 425 430 Pro Phe Cys Asp Phe Gly Glu Val Thr Asp Pro Pro Ser Glu Asp Val 435 440 445 Thr Ala His Val Leu Glu Cys Phe Gly Ser Phe Gly Tyr Asp Asp Ala 450 455 460 Trp Lys Val Ile Arg Arg Ala Val Glu Tyr Leu Lys Arg Glu Gln Lys 465 470 475 480 Pro Asp Gly Ser Trp Phe Gly Arg Trp Gly Val Asn Tyr Leu Tyr Gly 485 490 495 Thr Gly Ala Val Val Ser Ala Leu Lys Ala Val Gly Ile Asp Thr Arg 500 505 510 Glu Pro Tyr Ile Gln Lys Ala Leu Asp Trp Val Glu Gln His Gln Asn 515 520 525 Pro Asp Gly Gly Trp Gly Glu Asp Cys Arg Ser Tyr Glu Asp Pro Ala 530 535 540 Tyr Ala Gly Lys Gly Ala Ser Thr Pro Ser Gln Thr Ala Trp Ala Leu 545 550 555 560 Met Ala Leu Ile Ala Gly Gly Arg Ala Glu Ser Glu Ala Ala Arg Arg 565 570 575 Gly Val Gln Tyr Leu Val Glu Thr Gln Arg Pro Asp Gly Gly Trp Asp 580 585 590 Glu Pro Tyr Tyr Thr Gly Thr Gly Phe Pro Gly Asp Phe Tyr Leu Gly 595 600 605 Tyr Thr Met Tyr Arg His Val Phe Pro Thr Leu Ala Leu Gly Arg Tyr 610 615 620 Lys Gln Ala Ile Glu Arg Arg 625 630 <210> 12 <211> 1896 <212> DNA <213> Artificial sequence <220> <223> SHC enzyme variant #115A7 <400> 12 atggctgagc agttggtgga agctccggcc tacgcgcgga cgctggatcg cgcggtggag 60 tatctcctct cctgccaaaa ggacgaaggc tactggtggg ggccgcttct gagcaacgtc 120 acgatggaag cggagtacgt cctcttgtgc cacattctcg atcgcgtcga tcgggatcgc 180 atggagaaga tccggcggta cctgttgcac gagcagcgcg aggacggcac gtgggccctg 240 tacccgggtg ggccgccgga cctcgacacg accatcgagg cgtacgtcgc gctcaagtat 300 atcggcatgt cgcgcgacga ggagccgatg cagaaggcgc tccggttcat tcagagccag 360 ggcgggatcg agtcgtcgcg cgtgttcacg cggaggtggc tggcgctggt gggagaatat 420 ccgtgggaga aggtgcccat ggtcccgccg gagatcatgt tcctcggcaa gcgcatgccg 480 ctcaacatct acgagtttgg ctcgtgggct cggacgaccg tcgtggcgct ctcgattgtg 540 atgagccgcc agccggtgtt cccgctgccc gagcgggcgc gcgtgcccga gctgtacgag 600 accgacgtgc ctccgcgccg gcgcggtgcc aagggagggg gtgggtggat cttcgacgcg 660 ctcgaccggg tgctgcacgg gtatcagaag ctgtcggtgc acccgttccg ccgcgcggcc 720 gagatccgcg ccttggactg gttgctcgag cgccaggccg gagacggcag ctggggcggg 780 attcagccgc cttggtttta cgcgctcatc gcgctcaaga ttctcgacaa gacgcagcat 840 ccggcgttca tcaagggctg ggaaggtcta gagctgtacg gcgtggagct ggattacgga ggatggatgt ttcaggcttc catctcgccg gtgtgggaca cgggcctcgc cgtgctcgcg 960 ctgcgcgctg cggggcttcc ggccgatcac gaccgcttgg tcaaggcggg cgagtggctg ttggaccggc agatcacggt tccggggcgac tggggcggtga agccccgaa cctcaagccg 1080. ggcgggttcg cgttccagtt cgacaacgtg tactacccgg acgtggacga cacggccgtc gtggtgtggg cgctcaacac cctgcgcttg ccggacgagc gccgcaggcg ggacgccatg acgaagggat tccgctggat tgtcggcatg cagagctcga acggcggttg gggcgcctac 1260 gacgtcgaca acacgagcga tctcccgac cacaccccgt tctgcgactt cggcgaagtg accgatccgc cgtcagagga cgtcaccgcc cacgtgctcg agtgtttcgg cagcttcggg 1380 tacgatgacg cctggaaggt catccggcgc gcggtggaat atctcaagcg ggagcagaag 1440 ccggacggca gctggttcgg tcgttggggc gtcaattacc tctacggcac gggcgcggtg 1500 gtgtcggcgc tgaaggcggt cgggatcgac acgcgcgagc cgtacattca aaaggcgctc 1560 gactgggtcg agcagcatca gaacccggac ggcggctggg gcgaggactg ccgctcgtac 1620 gaggatccgg cgtacgcggg taagggcgcg agcaccccgt cgcagacggc ctgggcgctg 1680 atggcgctca tcgcgggcgg cagggcggag tccgaggccg cgcgccgcgg cgtgcaatac 1740 ctcgtggaga cgcagcgccc ggacggcggc tgggatgagc cgtactacac cggcacgggc 1800 ttcccagggg atttctacct cggctacacc atgtaccgcc acgtgtttcc gacgctcgcg 1860 ctcggccgct acaagcaagc catcgagcgc aggtga 1896 <210> 13 <211> 631 <212> PRT <213> Artificial Sequence <220> <223> SHC enzyme variant #215G2 <400> 13 Met Ala Glu Gln Leu Val Glu Ala Pro Ala Tyr Ala Arg Thr Leu Asp 1 5 10 15 Arg Ala Val Glu Tyr Leu Leu Ser Cys Gln Lys Asp Glu Gly Tyr Trp 20 25 30 Trp Gly Pro Leu Leu Ser Asn Val Thr Met Glu Ala Glu Tyr Val Leu 35 40 45 Leu Cys His Ile Leu Asp Arg Val Asp Arg Asp Arg Met Glu Lys Ile 50 55 60 Arg Arg Tyr Leu Leu His Glu Gln Arg Glu Asp Gly Thr Trp Ala Leu 65 70 75 80 Tyr Pro Gly Gly Pro Pro Asp Leu Asp Thr Thr Ile Glu Ala Tyr Val 85 90 95 Ala Leu Lys Tyr Ile Gly Met Ser Arg Asp Glu Glu Pro Met Gln Lys 100 105 110 Ala Leu Arg Phe Ile Gln Ser Gln Gly Gly Ile Glu Ser Ser Arg Val 115 120 125 Phe Thr Arg Arg Trp Leu Ala Leu Val Gly Glu Tyr Pro Trp Glu Lys 130 135 140 Val Pro Met Val Pro Pro Glu Ile Met Phe Leu Gly Lys Arg Met Pro 145 150 155 160 Leu Asn Ile Tyr Glu Phe Gly Ser Trp Ala Arg Ala Thr Val Val Ala 165 170 175 Leu Ser Ile Val Met Ser Arg Gln Pro Val Phe Pro Leu Pro Glu Arg 180 185 190 Ala Arg Val Pro Glu Leu Tyr Glu Thr Asp Val Pro Pro Arg Arg Arg 195 200 205 Gly Ala Lys Gly Gly Gly Gly Trp Ile Phe Asp Ala Leu Asp Arg Val 210 215 220 Leu His Gly Tyr Gln Lys Leu Ser Val His Pro Phe Arg Arg Ala Ala 225 230 235 240 Glu Ile Arg Ala Leu Asp Trp Leu Leu Glu Arg Gln Ala Gly Asp Gly 245 250 255 Ser Trp Gly Gly Ile Gln Pro Pro Trp Phe Tyr Ala Leu Ile Ala Leu 260 265 270 Lys Ile Leu Asp Met Thr Gln His Pro Ala Phe Ile Lys Gly Trp Glu 275 280 285 Gly Leu Glu Leu Tyr Gly Val Glu Leu Asp Tyr Gly Gly Trp Met Phe 290 295 300 Gln Ala Ser Ile Ser Pro Val Trp Asp Thr Gly Leu Ala Val Leu Ala 305 310 315 320 Leu Arg Ala Ala Gly Leu Pro Ala Asp His Asp Arg Leu Val Lys Ala 325 330 335 Gly Glu Trp Leu Leu Asp Arg Gln Ile Thr Val Pro Gly Asp Trp Ala 340 345 350 Val Lys Arg Pro Asn Leu Lys Pro Gly Gly Phe Ala Phe Gln Phe Asp 355 360 365 Asn Val Tyr Tyr Pro Asp Val Asp Asp Thr Ala Val Val Val Trp Ala 370 375 380 Leu Asn Thr Leu Arg Leu Pro Asp Glu Arg Arg Arg Arg Asp Ala Met 385 390 395 400 Thr Lys Gly Phe Arg Trp Ile Val Gly Met Gln Ser Ser Asn Gly Gly 405 410 415 Trp Gly Ala Tyr Asp Val Asp Asn Thr Ser Asp Leu Pro Asn His Thr 420 425 430 Pro Phe Cys Asp Phe Gly Glu Val Thr Asp Pro Pro Ser Glu Asp Val 435 440 445 Thr Ala His Val Leu Glu Cys Phe Gly Ser Phe Gly Tyr Asp Asp Ala 450 455 460 Trp Lys Val Ile Arg Arg Ala Val Glu Tyr Leu Lys Arg Glu Gln Lys 465 470 475 480 Pro Asp Gly Ser Trp Phe Gly Arg Trp Gly Val Asn Tyr Leu Tyr Gly 485 490 495 Thr Gly Ala Val Val Ser Ala Leu Lys Ala Val Gly Ile Asp Thr Arg 500 505 510 Glu Pro Tyr Ile Gln Lys Ala Leu Asp Trp Val Glu Gln His Gln Asn 515 520 525 Pro Asp Gly Gly Trp Gly Glu Asp Cys Arg Ser Tyr Glu Asp Pro Ala 530 535 540 Tyr Ala Gly Lys Gly Ala Ser Thr Pro Ser Gln Thr Ala Trp Ala Leu 545 550 555 560 Met Ala Leu Ile Ala Gly Gly Arg Ala Glu Ser Glu Ala Ala Arg Arg 565 570 575 Gly Val Gln Tyr Leu Val Glu Thr Gln Arg Pro Asp Gly Gly Trp Asp 580 585 590 Glu Pro Tyr Tyr Thr Gly Thr Gly Phe Pro Gly Asp Phe Tyr Leu Gly 595 600 605 Tyr Thr Met Tyr Arg His Val Phe Pro Thr Leu Ala Leu Gly Arg Tyr 610 615 620 Lys Gln Ala Ile Glu Arg Arg 625 630 <210> 14 <211> 1896 <212> DNA <213> Artificial Sequence <220> <223> SHC enzyme variant #215G2 <400> 14 atggctgagc agttggtgga agctccggcc tacgcgcgga cgctggatcg cgcggtggag 60 tatctcctct cctgccaaaa ggacgaaggc tactggtggg ggccgcttct gagcaacgtc 120 acgatggaag cggagtacgt cctcttgtgc cacattctcg atcgcgtcga tcgggatcgc 180 atggagaaga tccggcggta cctgttgcac gagcagcgcg aggacggcac gtgggccctg 240 tacccgggtg ggccgccgga cctcgacacg accatcgagg cgtacgtcgc gctcaagtat 300 atcggcatgt cgcgcgacga ggagccgatg cagaaggcgc tccggttcat tcagagccag 360 ggcgggatcg agtcgtcgcg cgtgttcacg cggaggtggc tggcgctggt gggagaatat 420 ccgtgggaga aggtgcccat ggtcccgccg gagatcatgt tcctcggcaa gcgcatgccg 480 ctcaacatct acgagtttgg ctcgtgggct cgggcgaccg tcgtggcgct ctcgattgtg 540 atgagccgcc agccggtgtt cccgctgccc gagcgggcgc gcgtgcccga gctgtacgag 600 accgacgtgc ctccgcgccg gcgcggtgcc aagggagggg gtgggtggat cttcgacgcg 660 ctcgaccggg tgctgcacgg gtatcagaag ctgtcggtgc acccgttccg ccgcgcggcc 720 gagatccgcg ccttggactg gttgctcgag cgccaggccg gagacggcag ctggggcggg 780 attcagccgc cttggtttta cgcgctcatc gcgctcaaga ttctcgacat gacgcagcat 840 ccggcgttca tcaagggctg ggaaggtcta gagctgtacg gcgtggagct ggattacgga ggatggatgt ttcaggcttc catctcgccg gtgtgggaca cgggcctcgc cgtgctcgcg 960 ctgcgcgctg cggggcttcc ggccgatcac gaccgcttgg tcaaggcggg cgagtggctg ttggaccggc agatcacggt tccggggcgac tggggcggtga agccccgaa cctcaagccg 1080. ggcgggttcg cgttccagtt cgacaacgtg tactacccgg acgtggacga cacggccgtc gtggtgtggg cgctcaacac cctgcgcttg ccggacgagc gccgcaggcg ggacgccatg acgaagggat tccgctggat tgtcggcatg cagagctcga acggcggttg gggcgcctac 1260 gacgtcgaca acacgagcga tctcccgac cacaccccgt tctgcgactt cggcgaagtg accgatccgc cgtcagagga cgtcaccgcc cacgtgctcg agtgtttcgg cagcttcggg 1380 tacgatgacg cctggaaggt catccggcgc gcggtggaat atctcaagcg ggagcagaag 1440 ccggacggca gctggttcgg tcgttggggc gtcaattacc tctacggcac gggcgcggtg 1500 gtgtcggcgc tgaaggcggt cgggatcgac acgcgcgagc cgtacattca aaaggcgctc 1560 gactgggtcg agcagcatca gaacccggac ggcggctggg gcgaggactg ccgctcgtac 1620 gaggatccgg cgtacgcggg taagggcgcg agcaccccgt cgcagacggc ctgggcgctg 1680 atggcgctca tcgcgggcgg cagggcggag tccgaggccg cgcgccgcgg cgtgcaatac 1740 ctcgtggaga cgcagcgccc ggacggcggc tgggatgagc cgtactacac cggcacgggc 1800 ttcccagggg atttctacct cggctacacc atgtaccgcc acgtgtttcc gacgctcgcg 1860 ctcggccgct acaagcaagc catcgagcgc aggtga 1896 <210> 15 <211> 608 <212> PRT <213> Zymomonas mobilis wild-type SHC enzyme <400> 15 Met Gly Ile Asp Arg Met Asn Ser Leu Ser Arg Leu Leu Met Lys Lys 1 5 10 15 Ile Phe Gly Ala Glu Lys Thr Ser Tyr Lys Pro Ala Ser Asp Thr Ile 20 25 30 Ile Gly Thr Asp Thr Leu Lys Arg Pro Asn Arg Arg Pro Glu Pro Thr 35 40 45 Ala Lys Val Asp Lys Thr Ile Phe Lys Thr Met Gly Asn Ser Leu Asn 50 55 60 Asn Thr Leu Val Ser Ala Cys Asp Trp Leu Ile Gly Gln Gln Lys Pro 65 70 75 80 Asp Gly His Trp Val Gly Ala Val Glu Ser Asn Ala Ser Met Glu Ala 85 90 95 Glu Trp Cys Leu Ala Leu Trp Phe Leu Gly Leu Glu Asp His Pro Leu 100 105 110 Arg Pro Arg Leu Gly Asn Ala Leu Leu Glu Met Gln Arg Glu Asp Gly 115 120 125 Ser Trp Gly Val Tyr Phe Gly Ala Gly Asn Gly Asp Ile Asn Ala Thr 130 135 140 Val Glu Ala Tyr Ala Ala Leu Arg Ser Leu Gly Tyr Ser Ala Asp Asn 145 150 155 160 Pro Val Leu Lys Lys Ala Ala Ala Trp Ile Ala Glu Lys Gly Gly Leu 165 170 175 Lys Asn Ile Arg Val Phe Thr Arg Tyr Trp Leu Ala Leu Ile Gly Glu 180 185 190 Trp Pro Trp Glu Lys Thr Pro Asn Leu Pro Pro Glu Ile Ile Trp Phe 195 200 205 Pro Asp Asn Phe Val Phe Ser Ile Tyr Asn Phe Ala Gln Trp Ala Arg 210 215 220 Ala Thr Met Val Pro Ile Ala Ile Leu Ser Ala Arg Arg Pro Ser Arg 225 230 235 240 Pro Leu Arg Pro Gln Asp Arg Leu Asp Glu Leu Phe Pro Glu Gly Arg 245 250 255 Ala Arg Phe Asp Tyr Glu Leu Pro Lys Lys Glu Gly Ile Asp Leu Trp 260 265 270 Ser Gln Phe Phe Arg Thr Thr Asp Arg Gly Leu His Trp Val Gln Ser 275 280 285 Asn Leu Leu Lys Arg Asn Ser Leu Arg Glu Ala Ala Ile Arg His Val 290 295 300 Leu Glu Trp Ile Ile Arg His Gln Asp Ala Asp Gly Gly Trp Gly Gly 305 310 315 320 Ile Gln Pro Pro Trp Val Tyr Gly Leu Met Ala Leu His Gly Glu Gly 325 330 335 Tyr Gln Leu Tyr His Pro Val Met Ala Lys Ala Leu Ser Ala Leu Asp 340 345 350 Asp Pro Gly Trp Arg His Asp Arg Gly Glu Ser Ser Trp Ile Gln Ala 355 360 365 Thr Asn Ser Pro Val Trp Asp Thr Met Leu Ala Leu Met Ala Leu Lys 370 375 380 Asp Ala Lys Ala Glu Asp Arg Phe Thr Pro Glu Met Asp Lys Ala Ala 385 390 395 400 Asp Trp Leu Leu Ala Arg Gln Val Lys Val Lys Gly Asp Trp Ser Ile 405 410 415 Lys Leu Pro Asp Val Glu Pro Gly Gly Trp Ala Phe Glu Tyr Ala Asn 420 425 430 Asp Arg Tyr Pro Asp Thr Asp Asp Thr Ala Val Ala Leu Ile Ala Leu 435 440 445 Ser Ser Tyr Arg Asp Lys Glu Glu Trp Gln Lys Lys Gly Val Glu Asp 450 455 460 Ala Ile Thr Arg Gly Val Asn Trp Leu Ile Ala Met Gln Ser Glu Cys 465 470 475 480 Gly Gly Trp Gly Ala Phe Asp Lys Asp Asn Asn Arg Ser Ile Leu Ser 485 490 495 Lys Ile Pro Phe Cys Asp Phe Gly Glu Ser Ile Asp Pro Pro Ser Val 500 505 510 Asp Val Thr Ala His Val Leu Glu Ala Phe Gly Thr Leu Gly Leu Ser 515 520 525 Arg Asp Met Pro Val Ile Gln Lys Ala Ile Asp Tyr Val Arg Ser Glu 530 535 540 Gln Glu Ala Glu Gly Ala Trp Phe Gly Arg Trp Gly Val Asn Tyr Ile 545 550 555 560 Tyr Gly Thr Gly Ala Val Leu Pro Ala Leu Ala Ala Ile Gly Glu Asp 565 570 575 Met Thr Gln Pro Tyr Ile Thr Lys Ala Cys Asp Trp Leu Val Ala His 580 585 590 Gln Gln Glu Asp Gly Gly Trp Gly Glu Ser Cys Ser Ser Tyr Met Glu 595 600 605 <210> 16 <211> 658 <212> PRT <213> Zymomonas mobilis wild-type SHC enzyme <400> 16 Met Thr Val Ser Thr Ser Ser Ala Phe His His Ser Pro Leu Ser Asp 1 5 10 15 Asp Val Glu Pro Ile Ile Gln Lys Ala Thr Arg Ala Leu Leu Glu Lys 20 25 30 Gln Gln Gln Asp Gly His Trp Val Phe Glu Leu Glu Ala Asp Ala Thr 35 40 45 Ile Pro Ala Glu Tyr Ile Leu Leu Lys His Tyr Leu Gly Glu Pro Glu 50 55 60 Asp Leu Glu Ile Glu Ala Lys Ile Gly Arg Tyr Leu Arg Arg Ile Gln 65 70 75 80 Gly Glu His Gly Gly Trp Ser Leu Phe Tyr Gly Gly Asp Leu Asp Leu 85 90 95 Ser Ala Thr Val Lys Ala Tyr Phe Ala Leu Lys Met Ile Gly Asp Ser 100 105 110 Pro Asp Ala Pro His Met Leu Arg Ala Arg Asn Glu Ile Leu Ala Arg 115 120 125 Gly Gly Ala Met Arg Ala Asn Val Phe Thr Arg Ile Gln Leu Ala Leu 130 135 140 Phe Gly Ala Met Ser Trp Glu His Val Pro Gln Met Pro Val Glu Leu 145 150 155 160 Met Leu Met Pro Glu Trp Phe Pro Val His Ile Asn Lys Met Ala Tyr 165 170 175 Trp Ala Arg Thr Val Leu Val Pro Leu Leu Val Leu Gln Ala Leu Lys 180 185 190 Pro Val Ala Arg Asn Arg Arg Gly Ile Leu Val Asp Glu Leu Phe Val 195 200 205 Pro Asp Val Leu Pro Thr Leu Gln Glu Ser Gly Asp Pro Ile Trp Arg 210 215 220 Arg Phe Phe Ser Ala Leu Asp Lys Val Leu His Lys Val Glu Pro Tyr 225 230 235 240 Trp Pro Lys Asn Met Arg Ala Lys Ala Ile His Ser Cys Val His Phe 245 250 255 Val Thr Glu Arg Leu Asn Gly Glu Asp Gly Leu Gly Ala Ile Tyr Pro 260 265 270 Ala Ile Ala Asn Ser Val Met Met Tyr Asp Ala Leu Gly Tyr Pro Glu 275 280 285 Asn His Pro Glu Arg Ala Ile Ala Arg Arg Ala Val Glu Lys Leu Met 290 295 300 Val Leu Asp Gly Thr Glu Asp Gln Gly Asp Lys Glu Val Tyr Cys Gln 305 310 315 320 Pro Cys Leu Ser Pro Ile Trp Asp Thr Ala Leu Val Ala His Ala Met 325 330 335 Leu Glu Val Gly Gly Asp Glu Ala Glu Lys Ser Ala Ile Ser Ala Leu 340 345 350 Ser Trp Leu Lys Pro Gln Gln Ile Leu Asp Val Lys Gly Asp Trp Ala 355 360 365 Trp Arg Arg Pro Asp Leu Arg Pro Gly Gly Trp Ala Phe Gln Tyr Arg 370 375 380 Asn Asp Tyr Tyr Pro Asp Val Asp Asp Thr Ala Val Val Thr Met Ala 385 390 395 400 Met Asp Arg Ala Ala Lys Leu Ser Asp Leu His Asp Asp Phe Glu Glu 405 410 415 Ser Lys Ala Arg Ala Met Glu Trp Thr Ile Gly Met Gln Ser Asp Asn 420 425 430 Gly Gly Trp Gly Ala Phe Asp Ala Asn Asn Ser Tyr Thr Tyr Leu Asn 435 440 445 Asn Ile Pro Phe Ala Asp His Gly Ala Leu Leu Asp Pro Pro Thr Val 450 455 460 Asp Val Ser Ala Arg Cys Val Ser Met Met Ala Gln Ala Gly Ile Ser 465 470 475 480 Ile Thr Asp Pro Lys Met Lys Ala Ala Val Asp Tyr Leu Leu Lys Glu 485 490 495 Gln Glu Glu Asp Gly Ser Trp Phe Gly Arg Trp Gly Val Asn Tyr Ile 500 505 510 Tyr Gly Thr Trp Ser Ala Leu Cys Ala Leu Asn Val Ala Ala Leu Pro 515 520 525 His Asp His Leu Ala Val Gln Lys Ala Val Ala Trp Leu Lys Thr Ile 530 535 540 Gln Asn Glu Asp Gly Gly Trp Gly Glu Asn Cys Asp Ser Tyr Ala Leu 545 550 555 560 Asp Tyr Ser Gly Tyr Glu Pro Met Asp Ser Thr Ala Ser Gln Thr Ala 565 570 575 Trp Ala Leu Leu Gly Leu Met Ala Val Gly Glu Ala Asn Ser Glu Ala 580 585 590 Val Thr Lys Gly Ile Asn Trp Leu Ala Gln Asn Gln Asp Glu Glu Gly 595 600 605 Leu Trp Lys Glu Asp Tyr Tyr Ser Gly Gly Gly Phe Pro Arg Val Phe 610 615 620 Tyr Leu Arg Tyr His Gly Tyr Ser Lys Tyr Phe Pro Leu Trp Ala Leu 625 630 635 640 Ala Arg Tyr Arg Asn Leu Lys Lys Ala Asn Gln Pro Ile Val His Tyr 645 650 655 Gly Met <210> 17 <211> 684 <212> PRT <213> Bradyrhizobium japonicum wild-type SHC enzyme <400> 17 Met Thr Val Thr Ser Ser Ala Ser Ala Arg Ala Thr Arg Asp Pro Gly 1 5 10 15 Asn Tyr Gln Thr Ala Leu Gln Ser Thr Val Arg Ala Ala Ala Asp Trp 20 25 30 Leu Ile Ala Asn Gln Lys Pro Asp Gly His Trp Val Gly Arg Ala Glu 35 40 45 Ser Asn Ala Cys Met Glu Ala Gln Trp Cys Leu Ala Leu Trp Phe Met 50 55 60 Gly Leu Glu Asp His Pro Leu Arg Lys Arg Leu Gly Gln Ser Leu Leu 65 70 75 80 Asp Ser Gln Arg Pro Asp Gly Ala Trp Gln Val Tyr Phe Gly Ala Pro 85 90 95 Asn Gly Asp Ile Asn Ala Thr Val Glu Ala Tyr Ala Ala Leu Arg Ser 100 105 110 Leu Gly Phe Arg Asp Asp Glu Pro Ala Val Arg Arg Ala Arg Glu Trp 115 120 125 Ile Glu Ala Lys Gly Gly Leu Arg Asn Ile Arg Val Phe Thr Arg Tyr 130 135 140 Trp Leu Ala Leu Ile Gly Glu Trp Pro Trp Glu Lys Thr Pro Asn Ile 145 150 155 160 Pro Pro Glu Val Ile Trp Phe Pro Leu Trp Phe Pro Phe Ser Ile Tyr 165 170 175 Asn Phe Ala Gln Trp Ala Arg Ala Thr Leu Met Pro Ile Ala Val Leu 180 185 190 Ser Ala Arg Arg Pro Ser Arg Pro Leu Pro Pro Glu Asn Arg Leu Asp 195 200 205 Ala Leu Phe Pro His Gly Arg Lys Ala Phe Asp Tyr Glu Leu Pro Val 210 215 220 Lys Ala Gly Ala Gly Gly Trp Asp Arg Phe Phe Arg Gly Ala Asp Lys 225 230 235 240 Val Leu His Lys Leu Gln Asn Leu Gly Asn Arg Leu Asn Leu Gly Leu 245 250 255 Phe Arg Pro Ala Ala Thr Ser Arg Val Leu Glu Trp Met Ile Arg His 260 265 270 Gln Asp Phe Asp Gly Ala Trp Gly Gly Ile Gln Pro Pro Trp Ile Tyr 275 280 285 Gly Leu Met Ala Leu Tyr Ala Glu Gly Tyr Pro Leu Asn His Pro Val 290 295 300 Leu Ala Lys Gly Leu Asp Ala Leu Asn Asp Pro Gly Trp Arg Val Asp 305 310 315 320 Val Gly Asp Ala Thr Tyr Ile Gln Ala Thr Asn Ser Pro Val Trp Asp 325 330 335 Thr Ile Leu Thr Leu Leu Ala Phe Asp Asp Ala Gly Val Leu Gly Asp 340 345 350 Tyr Pro Glu Ala Val Asp Lys Ala Val Asp Trp Val Leu Gln Arg Gln 355 360 365 Val Arg Val Pro Gly Asp Trp Ser Met Lys Leu Pro His Val Lys Pro 370 375 380 Gly Gly Trp Ala Phe Glu Tyr Ala Asn Asn Tyr Tyr Pro Asp Thr Asp 385 390 395 400 Asp Thr Ala Val Ala Leu Ile Ala Leu Ala Pro Leu Arg His Asp Pro 405 410 415 Lys Trp Lys Ala Lys Gly Ile Asp Glu Ala Ile Gln Leu Gly Val Asp 420 425 430 Trp Leu Ile Gly Met Gln Ser Gln Gly Gly Gly Trp Gly Ala Phe Asp 435 440 445 Lys Asp Asn Asn Gln Lys Ile Leu Thr Lys Ile Pro Phe Cys Asp Tyr 450 455 460 Gly Glu Ala Leu Asp Pro Pro Ser Val Asp Val Thr Ala His Ile Ile 465 470 475 480 Glu Ala Phe Gly Lys Leu Gly Ile Ser Arg Asn His Pro Ser Met Val 485 490 495 Gln Ala Leu Asp Tyr Ile Arg Arg Glu Gln Glu Pro Ser Gly Pro Trp 500 505 510 Phe Gly Arg Trp Gly Val Asn Tyr Val Tyr Gly Thr Gly Ala Val Leu 515 520 525 Pro Ala Leu Ala Ala Ile Gly Glu Asp Met Thr Gln Pro Tyr Ile Gly 530 535 540 Arg Ala Cys Asp Trp Leu Val Ala His Gln Gln Ala Asp Gly Gly Trp 545 550 555 560 Gly Glu Ser Cys Ala Ser Tyr Met Asp Val Ser Ala Val Gly Arg Gly 565 570 575 Thr Thr Thr Ala Ser Gln Thr Ala Trp Ala Leu Met Ala Leu Leu Ala 580 585 590 Ala Asn Arg Pro Gln Asp Lys Asp Ala Ile Glu Arg Gly Cys Met Trp 595 600 605 Leu Val Glu Arg Gln Ser Ala Gly Thr Trp Asp Glu Pro Glu Phe Thr 610 615 620 Gly Thr Gly Phe Pro Gly Tyr Gly Val Gly Gln Thr Ile Lys Leu Asn 625 630 635 640 Asp Pro Ala Leu Ser Gln Arg Leu Met Gln Gly Pro Glu Leu Ser Arg 645 650 655 Ala Phe Met Leu Arg Tyr Gly Met Tyr Arg His Tyr Phe Pro Leu Met 660 665 670 Ala Leu Gly Arg Ala Leu Arg Pro Gln Ser His Ser 675 680 <210> 18 <211> 642 <212> PRT <213> *Thermosynechococcus elongatus* wild-type SHC enzyme <400> 18 Met Pro Thr Ser Leu Ala Thr Ala Ile Asp Pro Lys Gln Leu Gln Gln 1 5 10 15 Ala Ile Arg Ala Ser Gln Asp Phe Leu Phe Ser Gln Gln Tyr Ala Glu 20 25 30 Gly Tyr Trp Trp Ala Glu Leu Glu Ser Asn Val Thr Met Thr Ala Glu 35 40 45 Val Ile Leu Leu His Lys Ile Trp Gly Thr Glu Gln Arg Leu Pro Leu 50 55 60 Ala Lys Ala Glu Gln Tyr Leu Arg Asn His Gln Arg Asp His Gly Gly 65 70 75 80 Trp Glu Leu Phe Tyr Gly Asp Gly Gly Asp Leu Ser Thr Ser Val Glu 85 90 95 Ala Tyr Met Gly Leu Arg Leu Leu Gly Val Pro Glu Thr Asp Pro Ala 100 105 110 Leu Val Lys Ala Arg Gln Phe Ile Leu Ala Arg Gly Gly Ile Ser Lys 115 120 125 Thr Arg Ile Phe Thr Lys Leu His Leu Ala Leu Ile Gly Cys Tyr Asp 130 135 140 Trp Arg Gly Ile Pro Ser Leu Pro Pro Trp Ile Met Leu Leu Pro Glu 145 150 155 160 Gly Ser Pro Phe Thr Ile Tyr Glu Met Ser Ser Trp Ala Arg Ser Ser 165 170 175 Thr Val Pro Leu Leu Ile Val Met Asp Arg Lys Pro Val Tyr Gly Met 180 185 190 Asp Pro Pro Ile Thr Leu Asp Glu Leu Tyr Ser Glu Gly Arg Ala Asn 195 200 205 Val Val Trp Glu Leu Pro Arg Gln Gly Asp Trp Arg Asp Val Phe Ile 210 215 220 Gly Leu Asp Arg Val Phe Lys Leu Phe Glu Thr Leu Asn Ile His Pro 225 230 235 240 Leu Arg Glu Gln Gly Leu Lys Ala Ala Glu Glu Trp Val Leu Glu Arg 245 250 255 Gln Glu Ala Ser Gly Asp Trp Gly Gly Ile Ile Pro Ala Met Leu Asn 260 265 270 Ser Leu Leu Ala Leu Arg Ala Leu Asp Tyr Ala Val Asp Asp Pro Ile 275 280 285 Val Gln Arg Gly Met Ala Ala Val Asp Arg Phe Ala Ile Glu Thr Glu 290 295 300 Thr Glu Tyr Arg Val Gln Pro Cys Val Ser Pro Val Trp Asp Thr Ala 305 310 315 320 Leu Val Met Arg Ala Met Val Asp Ser Gly Val Ala Pro Asp His Pro 325 330 335 Ala Leu Val Lys Ala Gly Glu Trp Leu Leu Ser Lys Gln Ile Leu Asp 340 345 350 Tyr Gly Asp Trp His Ile Lys Asn Lys Lys Gly Arg Pro Gly Gly Trp 355 360 365 Ala Phe Glu Phe Glu Asn Arg Phe Tyr Pro Asp Val Asp Asp Thr Ala 370 375 380 Val Val Val Met Ala Leu His Ala Val Thr Leu Pro Asn Glu Asn Leu 385 390 395 400 Lys Arg Arg Ala Ile Glu Arg Ala Val Ala Trp Ile Ala Ser Met Gln 405 410 415 Cys Arg Pro Gly Gly Trp Ala Ala Phe Asp Val Asp Asn Asp Gln Asp 420 425 430 Trp Leu Asn Gly Ile Pro Tyr Gly Asp Leu Lys Ala Met Ile Asp Pro 435 440 445 Asn Thr Ala Asp Val Thr Ala Arg Val Leu Glu Met Val Gly Arg Cys 450 455 460 Gln Leu Ala Phe Asp Arg Val Ala Leu Asp Arg Ala Leu Ala Tyr Leu 465 470 475 480 Arg Asn Glu Gln Glu Pro Glu Gly Cys Trp Phe Gly Arg Trp Gly Val 485 490 495 Asn Tyr Leu Tyr Gly Thr Ser Gly Val Leu Thr Ala Leu Ser Leu Val 500 505 510 Ala Pro Arg Tyr Asp Arg Trp Arg Ile Arg Arg Ala Ala Glu Trp Leu 515 520 525 Met Gln Cys Gln Asn Ala Asp Gly Gly Trp Gly Glu Thr Cys Trp Ser 530 535 540 Tyr His Asp Pro Ser Leu Lys Gly Lys Gly Asp Ser Thr Ala Ser Gln 545 550 555 560 Thr Ala Trp Ala Ile Ile Gly Leu Leu Ala Ala Gly Asp Ala Thr Gly 565 570 575 Asp Tyr Ala Thr Glu Ala Ile Glu Arg Gly Ile Ala Tyr Leu Leu Glu 580 585 590 Thr Gln Arg Pro Asp Gly Thr Trp His Glu Asp Tyr Phe Thr Gly Thr 595 600 605 Gly Phe Pro Cys His Phe Tyr Leu Lys Tyr His Tyr Tyr Gln Gln His 610 615 620 Phe Pro Leu Thr Ala Leu Gly Arg Tyr Ala Arg Trp Arg Asn Leu Leu 625 630 635 640 Ala Thr <210> 19 <211> 720 <212> PRT <213> *Acetobacter pasteurianus* wild-type SHC enzyme <400> 19 Met Asn Met Ala Ser Arg Phe Ser Leu Lys Lys Ile Leu Arg Ser Gly 1 5 10 15 Ser Asp Thr Gln Gly Thr Asn Val Asn Thr Leu Ile Gln Ser Gly Thr 20 25 30 Ser Asp Ile Val Arg Gln Lys Pro Ala Pro Gln Glu Pro Ala Asp Leu 35 40 45 Ser Ala Leu Lys Ala Met Gly Asn Ser Leu Thr His Thr Leu Ser Ser 50 55 60 Ala Cys Glu Trp Leu Met Lys Gln Gln Lys Pro Asp Gly His Trp Val 65 70 75 80 Gly Ser Val Gly Ser Asn Ala Ser Met Glu Ala Glu Trp Cys Leu Ala 85 90 95 Leu Trp Phe Leu Gly Leu Glu Asp His Pro Leu Arg Pro Arg Leu Gly 100 105 110 Lys Ala Leu Leu Glu Met Gln Arg Pro Asp Gly Ser Trp Gly Thr Tyr 115 120 125 Tyr Gly Ala Gly Ser Gly Asp Ile Asn Ala Thr Val Glu Ser Tyr Ala 130 135 140 Ala Leu Arg Ser Leu Gly Tyr Ala Glu Asp Asp Pro Ala Val Ser Lys 145 150 155 160 Ala Ala Ala Trp Ile Ile Ser Lys Gly Gly Leu Lys Asn Val Arg Val 165 170 175 Phe Thr Arg Tyr Trp Leu Ala Leu Ile Gly Glu Trp Pro Trp Glu Lys 180 185 190 Thr Pro Asn Leu Pro Pro Glu Ile Ile Trp Phe Pro Asp Asn Phe Val 195 200 205 Phe Ser Ile Tyr Asn Phe Ala Gln Trp Ala Arg Ala Thr Met Met Pro 210 215 220 Leu Ala Ile Leu Ser Ala Arg Arg Pro Ser Arg Pro Leu Arg Pro Gln 225 230 235 240 Asp Arg Leu Asp Ala Leu Phe Pro Gly Gly Arg Ala Asn Phe Asp Tyr 245 250 255 Glu Leu Pro Thr Lys Glu Gly Arg Asp Val Ile Ala Asp Phe Phe Arg 260 265 270 Leu Ala Asp Lys Gly Leu His Trp Leu Gln Ser Ser Phe Leu Lys Arg 275 280 285 Ala Pro Ser Arg Glu Ala Ala Ile Lys Tyr Val Leu Glu Trp Ile Ile 290 295 300 Trp His Gln Asp Ala Asp Gly Gly Trp Gly Gly Ile Gln Pro Pro Trp 305 310 315 320 Val Tyr Gly Leu Met Ala Leu His Gly Glu Gly Tyr Gln Phe His His 325 330 335 Pro Val Met Ala Lys Ala Leu Asp Ala Leu Asn Asp Pro Gly Trp Arg 340 345 350 His Asp Lys Gly Asp Ala Ser Trp Ile Gln Ala Thr Asn Ser Pro Val 355 360 365 Trp Asp Thr Met Leu Ser Leu Met Ala Leu His Asp Ala Asn Ala Glu 370 375 380 Glu Arg Phe Thr Pro Glu Met Asp Lys Ala Leu Asp Trp Leu Leu Ser 385 390 395 400 Arg Gln Val Arg Val Lys Gly Asp Trp Ser Val Lys Leu Pro Asn Thr 405 410 415 Glu Pro Gly Gly Trp Ala Phe Glu Tyr Ala Asn Asp Arg Tyr Pro Asp 420 425 430 Thr Asp Asp Thr Ala Val Ala Leu Ile Ala Ile Ala Ser Cys Arg Asn 435 440 445 Arg Pro Glu Trp Gln Ala Lys Gly Val Glu Glu Ala Ile Gly Arg Gly 450 455 460 Val Arg Trp Leu Val Ala Met Gln Ser Ser Cys Gly Gly Trp Gly Ala 465 470 475 480 Phe Asp Lys Asp Asn Asn Lys Ser Ile Leu Ala Lys Ile Pro Phe Cys 485 490 495 Asp Phe Gly Glu Ala Leu Asp Pro Pro Ser Val Asp Val Thr Ala His 500 505 510 Val Leu Glu Ala Phe Gly Leu Leu Gly Leu Pro Arg Asp Leu Pro Cys 515 520 525 Ile Gln Arg Gly Leu Ala Tyr Ile Arg Lys Glu Gln Asp Pro Thr Gly 530 535 540 Pro Trp Phe Gly Arg Trp Gly Val Asn Tyr Leu Tyr Gly Thr Gly Ala 545 550 555 560 Val Leu Pro Ala Leu Ala Ala Leu Gly Glu Asp Met Thr Gln Pro Tyr 565 570 575 Ile Ser Lys Ala Cys Asp Trp Leu Ile Asn Cys Gln Gln Glu Asn Gly 580 585 590 Gly Trp Gly Glu Ser Cys Ala Ser Tyr Met Glu Val Ser Ser Ile Gly 595 600 605 His Gly Ala Thr Thr Pro Ser Gln Thr Ala Trp Ala Leu Met Gly Leu 610 615 620 Ile Ala Ala Asn Arg Pro Gln Asp Tyr Glu Ala Ile Ala Lys Gly Cys 625 630 635 640 Arg Tyr Leu Ile Asp Leu Gln Glu Glu Asp Gly Ser Trp Asn Glu Glu 645 650 655 Glu Phe Thr Gly Thr Gly Phe Pro Gly Tyr Gly Val Gly Gln Thr Ile 660 665 670 Lys Leu Asp Asp Pro Ala Ile Ser Lys Arg Leu Met Gln Gly Ala Glu 675 680 685 Leu Ser Arg Ala Phe Met Leu Arg Tyr Asp Leu Tyr Arg Gln Leu Phe 690 695 700 Pro Ile Ile Ala Leu Ser Arg Ala Ser Arg Leu Ile Lys Leu Gly Asn 705 710 715 720 <210> 20 <211> 685 <212> PRT <213> Wild-type SHC enzyme from *Gluconobacter morbifer* <400> 20 Met Ser Pro Ala Asp Ile Ser Thr Lys Ser Ser Ser Phe Gln Arg Leu 1 5 10 15 Asp Asn Met Leu Pro Glu Ala Val Ser Ser Ala Cys Asp Trp Leu Ile 20 25 30 Asp Gln Gln Lys Pro Asp Gly His Trp Val Gly Pro Val Glu Ser Asn 35 40 45 Ala Cys Met Glu Ala Gln Trp Cys Leu Ala Leu Trp Phe Leu Gly Gln 50 55 60 Glu Asp His Pro Leu Arg Pro Arg Leu Ala Gln Ala Leu Leu Glu Met 65 70 75 80 Gln Arg Glu Asp Gly Ser Trp Gly Ile Tyr Val Gly Ala Asp His Gly 85 90 95 Asp Ile Asn Thr Thr Val Glu Ala Tyr Ala Ala Leu Arg Ser Met Gly 100 105 110 Tyr Ala Ala Asp Met Pro Ile Met Ala Lys Ser Ala Ala Trp Ile Gln 115 120 125 Gln Lys Gly Gly Leu Arg Asn Val Arg Val Phe Thr Arg Tyr Trp Leu 130 135 140 Ala Leu Ile Gly Glu Trp Pro Trp Asp Lys Thr Pro Asn Leu Pro Pro 145 150 155 160 Glu Ile Ile Trp Leu Pro Asp Asn Phe Ile Phe Ser Ile Tyr Asn Phe 165 170 175 Ala Gln Trp Ala Arg Ala Thr Met Met Pro Leu Thr Ile Leu Ser Ala 180 185 190 Arg Arg Pro Ser Arg Pro Leu Leu Pro Glu Asn Arg Leu Asp Gly Leu 195 200 205 Phe Pro Glu Gly Arg Glu Asn Phe Asp Tyr Glu Leu Pro Val Lys Gly 210 215 220 Glu Glu Asp Leu Trp Gly Arg Phe Phe Arg Ala Ala Asp Lys Gly Leu 225 230 235 240 His Ser Leu Gln Ser Phe Pro Val Arg Arg Phe Val Pro Arg Glu Ala 245 250 255 Ala Ile Arg His Val Ile Glu Trp Ile Ile Arg His Gln Asp Ala Asp 260 265 270 Gly Gly Trp Gly Gly Ile Gln Pro Pro Trp Ile Tyr Gly Leu Met Ala 275 280 285 Leu Ser Val Glu Gly Tyr Pro Leu His His Pro Val Leu Ala Lys Ala 290 295 300 Met Asp Ala Leu Asn Asp Pro Gly Trp Arg Arg Asp Lys Gly Asp Ala 305 310 315 320 Ser Trp Ile Gln Ala Thr Asn Ser Pro Val Trp Asp Thr Met Leu Ala 325 330 335 Val Leu Ala Leu His Asp Ala Gly Ala Glu Asp Arg Tyr Ser Pro Gln 340 345 350 Met Asp Lys Ala Ile Gly Trp Leu Leu Asp Arg Gln Val Arg Val Lys 355 360 365 Gly Asp Trp Ser Ile Lys Leu Pro Asp Thr Glu Pro Gly Gly Trp Ala 370 375 380 Phe Glu Tyr Ala Asn Asp Lys Tyr Pro Asp Thr Asp Asp Thr Ala Val 385 390 395 400 Ala Leu Ile Ala Leu Ala Gly Cys Arg His Arg Pro Glu Trp Arg Glu 405 410 415 Arg Asp Ile Glu Gly Ala Ile Ser Arg Gly Val Asn Trp Leu Leu Ala 420 425 430 Met Gln Ser Ser Ser Gly Gly Trp Gly Ala Phe Asp Lys Asp Asn Asn 435 440 445 Arg Ser Ile Leu Thr Lys Ile Pro Phe Cys Asp Phe Gly Glu Ala Leu 450 455 460 Asp Pro Pro Ser Val Asp Val Thr Ala His Val Leu Glu Ala Phe Gly 465 470 475 480 Leu Leu Gly Ile Ser Arg Asn His Pro Ser Val Gln Lys Ala Leu Ala 485 490 495 Tyr Ile Arg Ser Glu Gln Glu Arg Asn Gly Ala Trp Phe Gly Arg Trp 500 505 510 Gly Val Asn Tyr Val Tyr Gly Thr Gly Ala Val Leu Pro Ala Leu Ala 515 520 525 Ala Ile Gly Glu Asp Met Thr Gln Pro Tyr Ile Val Arg Ala Cys Asp 530 535 540 Trp Leu Met Ser Val Gln Gln Glu Asn Gly Gly Trp Gly Glu Ser Cys 545 550 555 560 Ala Ser Tyr Met Asp Ile Asn Ala Val Gly His Gly Val Ala Thr Ala 565 570 575 Ser Gln Thr Ala Trp Ala Leu Ile Gly Leu Leu Ala Ala Lys Arg Pro 580 585 590 Lys Asp Arg Glu Ala Ile Ala Arg Gly Cys Gln Phe Leu Ile Glu Arg 595 600 605 Gln Glu Asp Gly Ser Trp Thr Glu Glu Glu Tyr Thr Gly Thr Gly Phe 610 615 620 Pro Gly Tyr Gly Val Gly Gln Ala Ile Lys Leu Asp Asp Pro Ser Leu 625 630 635 640 Pro Asp Arg Leu Leu Gln Gly Ala Glu Leu Ser Arg Ala Phe Met Leu 645 650 655 Arg Tyr Asp Leu Tyr Arg Gln Tyr Phe Pro Val Met Ala Leu Ser Arg 660 665 670 Ala Arg Arg Met Met Lys Glu Asp Ala Ser Ala Ala Ala 675 680 685 <210> 21 <211> 625 <212> PRT <213> Wild-type SHC enzyme from *Bacillus megaterium* <400> 21 Met Ile Ile Leu Leu Lys Glu Val Gln Leu Glu Ile Gln Arg Arg Ile 1 5 10 15 Ala Tyr Leu Arg Pro Thr Gln Lys Asn Asp Gly Ser Phe Arg Tyr Cys 20 25 30 Phe Glu Thr Gly Val Met Pro Asp Ala Phe Leu Ile Met Leu Leu Arg 35 40 45 Thr Phe Asp Leu Asp Lys Glu Val Leu Ile Lys Gln Leu Thr Glu Arg 50 55 60 Ile Val Ser Leu Gln Asn Glu Asp Gly Leu Trp Thr Leu Phe Asp Asp 65 70 75 80 Glu Glu His Asn Leu Ser Ala Thr Ile Gln Ala Tyr Thr Ala Leu Leu 85 90 95 Tyr Ser Gly Tyr Tyr Gln Lys Asn Asp Arg Ile Leu Arg Lys Ala Glu 100 105 110 Arg Tyr Ile Ile Asp Ser Gly Gly Ile Ser Arg Ala His Phe Leu Thr 115 120 125 Arg Trp Met Leu Ser Val Asn Gly Leu Tyr Glu Trp Pro Lys Leu Phe 130 135 140 Tyr Leu Pro Leu Ser Leu Leu Leu Val Pro Thr Tyr Val Pro Leu Asn 145 150 155 160 Phe Tyr Glu Leu Ser Thr Tyr Ala Arg Ile His Phe Val Pro Met Met 165 170 175 Val Ala Gly Asn Lys Lys Phe Ser Leu Thr Ser Arg His Thr Pro Ser 180 185 190 Leu Ser His Leu Asp Val Arg Glu Gln Lys Gln Glu Ser Glu Glu Thr 195 200 205 Thr Gln Glu Ser Arg Ala Ser Ile Phe Leu Val Asp His Leu Lys Gln 210 215 220 Leu Ala Ser Leu Pro Ser Tyr Ile His Lys Leu Gly Tyr Gln Ala Ala 225 230 235 240 Glu Arg Tyr Met Leu Glu Arg Ile Glu Lys Asp Gly Thr Leu Tyr Ser 245 250 255 Tyr Ala Thr Ser Thr Phe Phe Met Ile Tyr Gly Leu Leu Ala Leu Gly 260 265 270 Tyr Lys Lys Asp Ser Phe Val Ile Gln Lys Ala Ile Asp Gly Ile Cys 275 280 285 Ser Leu Leu Ser Thr Cys Ser Gly His Val His Val Glu Asn Ser Thr 290 295 300 Ser Thr Val Trp Asp Thr Ala Leu Leu Ser Tyr Ala Leu Gln Glu Ala 305 310 315 320 Gly Val Pro Gln Gln Asp Pro Met Ile Lys Gly Thr Thr Arg Tyr Leu 325 330 335 Lys Lys Arg Gln His Thr Lys Leu Gly Asp Trp Gln Phe His Asn Pro 340 345 350 Asn Thr Ala Pro Gly Gly Trp Gly Phe Ser Asp Ile Asn Thr Asn Asn 355 360 365 Pro Asp Leu Asp Asp Thr Ser Ala Ala Ile Arg Ala Leu Ser Arg Arg 370 375 380 Ala Gln Thr Asp Thr Asp Tyr Leu Glu Ser Trp Gln Arg Gly Ile Asn 385 390 395 400 Trp Leu Leu Ser Met Gln Asn Lys Asp Gly Gly Phe Ala Ala Phe Glu 405 410 415 Lys Asn Thr Asp Ser Ile Leu Phe Thr Tyr Leu Pro Leu Glu Asn Ala 420 425 430 Lys Asp Ala Ala Thr Asp Pro Ala Thr Ala Asp Leu Thr Gly Arg Val 435 440 445 Leu Glu Cys Leu Gly Asn Phe Ala Gly Met Asn Lys Ser His Pro Ser 450 455 460 Ile Lys Ala Ala Val Lys Trp Leu Phe Asp His Gln Leu Asp Asn Gly 465 470 475 480 Ser Trp Tyr Gly Arg Trp Gly Val Cys Tyr Ile Tyr Gly Thr Trp Ala 485 490 495 Ala Ile Thr Gly Leu Arg Ala Val Gly Val Ser Ala Ser Asp Pro Arg 500 505 510 Ile Ile Lys Ala Ile Asn Trp Leu Lys Ser Ile Gln Gln Glu Asp Gly 515 520 525 Gly Phe Gly Glu Ser Cys Tyr Ser Ala Ser Leu Lys Lys Tyr Val Pro 530 535 540 Leu Ser Phe Ser Thr Pro Ser Gln Thr Ala Trp Ala Leu Asp Ala Leu 545 550 555 560 Met Thr Ile Cys Pro Leu Lys Asp Gln Ser Val Glu Lys Gly Ile Lys 565 570 575 Phe Leu Leu Asn Pro Asn Leu Thr Glu Gln Gln Thr His Tyr Pro Thr 580 585 590 Gly Ile Gly Leu Pro Gly Gln Phe Tyr Ile Gln Tyr His Ser Tyr Asn 595 600 605 Asp Ile Phe Pro Leu Leu Ala Leu Ala His Tyr Ala Lys Lys His Ser 610 615 620 Ser 625 <210> 22 <211> 631 <212> PRT <213> SHC enzyme variant #49 <400> 22 Met Ala Glu Gln Leu Val Glu Ala Pro Ala Tyr Ala Arg Thr Leu Asp 1 5 10 15 Arg Ala Val Glu Tyr Leu Leu Ser Cys Gln Lys Asp Glu Gly Tyr Trp 20 25 30 Trp Gly Pro Leu Leu Ser Asn Val Thr Met Glu Ala Glu Tyr Val Leu 35 40 45 Leu Cys His Ile Leu Asp Arg Val Asp Arg Asp Arg Met Glu Lys Ile 50 55 60 Arg Arg Tyr Leu Leu His Glu Gln Arg Glu Asp Gly Thr Trp Ala Leu 65 70 75 80 Tyr Pro Gly Gly Pro Pro Asp Leu Asp Thr Thr Ile Glu Ala Tyr Val 85 90 95 Ala Leu Lys Tyr Ile Gly Met Ser Arg Asp Glu Glu Pro Met Gln Lys 100 105 110 Ala Leu Arg Phe Ile Gln Ser Gln Gly Gly Ile Glu Ser Ser Arg Val 115 120 125 Phe Thr Arg Arg Trp Leu Ala Leu Val Gly Glu Tyr Pro Trp Glu Lys 130 135 140 Val Pro Met Val Pro Pro Glu Ile Met Phe Leu Gly Lys Arg Met Pro 145 150 155 160 Leu Asn Ile Tyr Glu Phe Gly Ser Trp Ala Arg Ala Thr Val Val Ala 165 170 175 Leu Ser Ile Val Met Ser Arg Gln Pro Val Phe Pro Leu Pro Glu Arg 180 185 190 Ala Arg Val Pro Glu Leu Tyr Glu Thr Asp Val Pro Pro Arg Arg Arg 195 200 205 Gly Ala Lys Gly Gly Gly Gly Trp Ile Phe Asp Ala Leu Asp Arg Val 210 215 220 Leu His Gly Tyr Gln Lys Leu Ser Val His Pro Phe Arg Arg Ala Ala 225 230 235 240 Glu Ile Arg Ala Leu Asp Trp Leu Leu Glu Arg Gln Ala Gly Asp Gly 245 250 255 Ser Trp Gly Gly Ile Gln Pro Pro Trp Phe Tyr Ala Leu Ile Ala Leu 260 265 270 Lys Ile Leu Asp Met Thr Gln His Pro Ala Phe Ile Lys Gly Trp Glu 275 280 285 Gly Leu Glu Leu Tyr Gly Val Glu Leu Asp Tyr Gly Gly Trp Met Phe 290 295 300 Gln Ala Ser Ile Ser Pro Val Trp Asp Thr Gly Leu Ala Val Leu Ala 305 310 315 320 Leu Arg Ala Ala Gly Leu Pro Ala Asp His Asp Arg Leu Val Lys Ala 325 330 335 Gly Glu Trp Leu Leu Asp Arg Gln Ile Thr Val Pro Gly Asp Trp Ala 340 345 350 Val Lys Arg Pro Asn Leu Lys Pro Gly Gly Phe Ala Phe Gln Phe Asp 355 360 365 Asn Val Tyr Tyr Pro Asp Val Asp Asp Thr Ala Val Val Val Trp Ala 370 375 380 Leu Asn Thr Leu Arg Leu Pro Asp Glu Arg Arg Arg Arg Asp Ala Met 385 390 395 400 Thr Lys Gly Phe Arg Trp Ile Val Gly Met Gln Ser Ser Asn Gly Gly 405 410 415 Trp Gly Ala Tyr Asp Val Asp Asn Thr Ser Asp Leu Pro Asn Leu Thr 420 425 430 Pro Phe Cys Asp Phe Gly Glu Val Thr Asp Pro Pro Ser Glu Asp Val 435 440 445 Thr Ala His Val Leu Glu Cys Phe Gly Ser Phe Gly Tyr Asp Asp Ala 450 455 460 Trp Lys Val Ile Arg Arg Ala Val Glu Tyr Leu Lys Arg Glu Gln Lys 465 470 475 480 Pro Asp Gly Ser Trp Phe Gly Arg Trp Gly Val Asn Tyr Leu Tyr Gly 485 490 495 Thr Gly Ala Val Val Ser Ala Leu Lys Ala Val Gly Ile Asp Thr Arg 500 505 510 Glu Pro Tyr Ile Gln Lys Ala Leu Asp Trp Val Glu Gln His Gln Asn 515 520 525 Pro Asp Gly Gly Trp Gly Glu Asp Cys Arg Ser Tyr Glu Asp Pro Ala 530 535 540 Tyr Ala Gly Lys Gly Ala Ser Thr Pro Ser Gln Thr Thr Trp Ala Leu 545 550 555 560 Met Ala Leu Ile Ala Gly Gly Arg Ala Glu Ser Glu Ala Ala Arg Arg 565 570 575 Gly Val Gln Tyr Leu Val Glu Thr Gln Arg Pro Asp Gly Gly Trp Asp 580 585 590 Glu Pro Tyr Tyr Thr Gly Thr Gly Phe Pro Gly Asp Phe Tyr Leu Gly 595 600 605 Tyr Thr Met Tyr Arg His Val Phe Pro Thr Leu Ala Leu Gly Arg Tyr 610 615 620 Court of Gln Ala Ile Glu Arg Arg 625 630

Claims

1. A method for preparing compounds of formula (I), The method includes contacting a compound of formula (II) with squalene-hopaene cyclase or a variant of squalene-hopaene cyclase, wherein the squalene-hopaene cyclase has an amino acid sequence of SEQ ID NO: 1, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20 or SEQ ID NO: 21, and wherein the squalene-hopaene cyclase variant has an amino acid sequence that is at least 70% identical to that of the squalene-hopaene cyclase. Where R is methyl or ethyl, and The squalene-hopaene cyclase variants have the amino acid sequences SEQ ID NO: 3, SEQ ID NO: 5, SEQ ID NO: 7, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 13, and SEQ ID NO:

22.

2. The method of claim 1, wherein the method comprises setting the double bond between C-8 and C-9 to be E - Configuration and the double bond between C-4 and C-5 is Z - Compounds of formula (II) with the following configuration are contacted with squalene-hopaene cyclase or a variant of squalene-hopaene cyclase.

3. The method according to claim 1, wherein the compound of formula (III) is prepared as a byproduct. Where R is methyl or ethyl.

4. The method according to claim 3, wherein a compound having the relative configuration shown in formula (IIIa) is prepared as a byproduct. Where R is methyl or ethyl.

5. The method according to any one of claims 1-4, wherein R is methyl.

6. The method according to any one of claims 1-2, wherein the method further comprises purifying the compound of formula (I).

7. A composition comprising a compound of formula (I) and a compound of formula (III): Where R is methyl or ethyl.

8. The composition according to claim 7, wherein R is methyl.

9. Use of the composition according to any one of claims 7 to 8 as a flavoring composition.

10. Use of the composition according to any one of claims 7 to 8 in a flavoring composition.

11. A consumer product comprising the composition according to any one of claims 7 to 8.

Citation Information

Patent Citations

  • Method for producing (-)-ambroxan (r)

    JP2009060799A

  • Enzymes and genes used for producing vanillin

    US20030092143A1

  • Synthetic enzymes for the production of coniferyl alcohol, coniferylaldehyde, ferulic acid, vanillin and vanillic acid and their use

    US6524831B2

  • Biocatalytic production of ambroxan

    WO2010139719A2

  • Method for the biocatalytic cyclisation of terpenes and cyclase mutants which can be used in said method

    WO2012066059A2