Production of mogrol and mogrosides by microorganisms
The recombinant microbial process using engineered enzymes in E. coli and yeast efficiently produces mogrol and mogrol glycosides, overcoming yield and purity limitations of natural extraction, enabling their use as high-intensity sweeteners and flavor enhancers.
Patent Information
- Application Number
- JP2025173313
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-09-30
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-27
AI Technical Summary
The existing methods for producing mogrol and mogrol glycosides, such as mogroside V, face challenges due to low yield, limited purity, and difficulty in sourcing large quantities from the monk fruit, which affects their commercial viability and application in food and beverage industries.
A recombinant microbial process using genetically engineered microbial strains, such as E. coli and yeast, expressing heterologous enzyme pathways to convert isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP) into mogrol and mogrol glycosides, involving enzymes like farnesyl diphosphate synthase, squalene synthase, and uridine diphosphate glycosyltransferases, to enhance production and purity.
This approach allows for the efficient production of mogrol and mogrol glycosides, including mogroside V, with improved yield and purity, addressing the limitations of natural extraction methods and enhancing their use as high-intensity sweeteners and flavor enhancers.
Smart Images

Figure 2026012752000007 
Figure 2026012752000008 
Figure 2026012752000009
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 085,557, filed September 30, 2020, U.S. Provisional Application No. 63 / 075,631, filed September 8, 2020, and U.S. Provisional Application No. 62 / 948,657, filed December 16, 2019, the entire disclosures of which are incorporated herein by reference. [Background technology]
[0002] Mogrosides are unique secondary metabolites derived from triterpenes found in the fruit of Siraitia grosvenorii (also known as monk fruit or monk fruit). Their biosynthesis in the fruit involves numerous sequential glycosylation reactions of the aglycone mogrol. Mogroside fruit extracts are increasingly being used in the food industry as natural, sugarless sweeteners. For example, mogroside V (Mog.V) is approximately 250 times sweeter than sucrose (Kasai et al., Agric Biol Chem (1989)). Furthermore, recent research has revealed additional health benefits of mogrosides (Li et al., Chin J Nat Med (2014)).
[0003] In general, various factors are driving research into mogrosides and monk fruit, as well as the rapidly growing commercial interest, including the surging popularity and demand for natural sweeteners; the difficulty in sourcing large quantities of other promising natural sweeteners derived from the stevia plant, such as rebaudioside M (RebM); the superior taste characteristics of Mog.V compared to other natural and artificial sweetener products on the market; and the medicinal properties of the plant and fruit.
[0004] Purified Mog.V has been approved as a high-intensity sweetener in Japan (Jakinovich et al., Journal of Natural Products (1990)), and its extract has achieved GRAS status in the United States as a non-nutritive sweetener and flavor enhancer (GRAS 522). Extraction of mogrosides from fruit yields products of varying purity, but most have an undesirable aftertaste. Furthermore, due to the low yield of the plant and the special cultivation conditions for this plant, the yield of mogrosides obtained from cultivated fruit is limited. Mogrosides are present at approximately 1% in fresh fruit and approximately 4% in dried fruit (Li HB, et al., 2006). Mog.V is the major component, and its content ranges from 0.5% to 1.4% in dried fruit. Furthermore, due to the difficulty of purification, the purity of Mog.V is limited, and commercial products derived from plant extracts are standardized to approximately 50% Mog.V. Pure Mog.V products are likely to be more commercially successful than blends because they are less likely to off-flavor, are easier to incorporate into products, and have good solubility. Therefore, it would be advantageous to be able to produce sweet-tasting mogroside compounds via biotechnology processes. Summary of the Invention
[0005] In various aspects and embodiments, the present invention provides enzymes (e.g., genetically engineered enzymes), microbial strains, and methods for producing mogrol and mogrol glycosides ("mogrosides") using recombinant microbial processes. In other aspects, the present invention provides methods for producing products, including foods, beverages, and (among other things) sweeteners, that contain mogrol glycosides produced according to the present disclosure.
[0006] In various aspects, the present invention provides a method for producing a microbial strain and a mogrol or mogrol glycoside(s). The present invention provides a method for producing mogrol or a mogrol glycoside. The present invention includes recombinant microbial host cells expressing a heterologous enzyme pathway that catalyzes the conversion of isopentenyl pyrophosphate (IPP) and / or dimethylallyl pyrophosphate (DMAPP) to mogrol or a mogrol glycoside(s). The microbial host cells in various embodiments may be prokaryotic (e.g., E. coli) or eukaryotic (e.g., yeast).
[0007] In various embodiments, the heterologous enzyme pathway comprises recombinantly expressed farnesyl diphosphate synthase (FPPS) and squalene synthase (SQS). In various embodiments, the SQS has an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 2-16, 166, and 167. In some embodiments, the SQS comprises an amino acid sequence at least 70% identical to an SQS (SEQ ID NO: 11) that exhibits high activity in E. coli.
[0008] In some embodiments, host cells express one or more enzymes that produce mogrol from squalene. For example, host cells may express one or more squalene epoxidase (SQE) enzymes, one or more triterpenoid cyclases, epoxide hydrolases (EPHs), one or more cytochrome P450 oxidases (CYP450s), non-heme iron-dependent oxygenases, and cytochrome P450 reductases (CPRs). As shown in Figure 2, heterologous pathways can lead to mogrol via several routes, which may include one or two epoxidations of a core substrate.
[0009] In some embodiments, a heterologous enzyme pathway includes two squalene epoxidase (SQE) enzymes. For example, a heterologous enzyme pathway can include an SQE that produces 2,3-oxidosqualene. In some embodiments, the SQE produces 2,3;22,23-dioxidosqualene, and this conversion can be catalyzed by the same SQE enzyme or an enzyme with at least one amino acid modification that alters the amino acid sequence. For example, a squalene epoxidase enzyme can include at least two SQE enzymes, each of which (independently) comprises an amino acid sequence at least 70% identical to any one of SEQ ID NOs: 17-39, 168-170, and 177-183.
[0010] In some embodiments, at least one SQE comprises an amino acid sequence that is at least 70% identical to SEQ ID NO:39.
[0011] In some embodiments, the host cell comprises two squalene epoxidase enzymes, each comprising an amino acid sequence at least 70% identical to squalene epoxidase (SEQ ID NO: 39). For example, one of the SQE enzymes may have one or more amino acid modifications that enhance specificity and productivity for converting 2,3-oxidosqualene to 2,3;22,23-dioxidosqualene compared to an enzyme having the amino acid sequence of SEQ ID NO: 39. In some embodiments, the amino acid modifications include one or more modifications at the following positions: 35, 133, 163, 254, 283, 380, and 395 of SEQ ID NO: 39. For example, the amino acid at position 35 of SEQ ID NO: 39 may be an arginine (e.g., H35R). The position corresponding to position 133 of SEQ ID NO: 39 may be a glycine (e.g., N133G). The amino acid at position 163 of SEQ ID NO: 39 may be an alanine (e.g., F163A). The amino acid at a position corresponding to position 254 of SEQ ID NO: 39 can be a phenylalanine (e.g., Y254F). The amino acid at a position corresponding to position 283 of SEQ ID NO: 39 can be a leucine (e.g., M283L). The amino acid at a position corresponding to position 380 of SEQ ID NO: 39 can be a leucine (e.g., V280L). The amino acid at a position corresponding to position 395 of SEQ ID NO: 39 can be a tyrosine (e.g., F395Y).
[0012] In various embodiments, the heterologous enzyme pathway includes a triterpene cyclase (TTC) enzyme. In some embodiments, when a microbial cell co-expresses FPPS with SQS, SQE, and one or more triterpene cyclase enzymes, the microbial cell produces 2,3;22,23-dioxidosqualene. 2,3;22,23-dioxidosqualene can be a substrate for downstream enzymes in the heterologous pathway. In some embodiments, the triterpene cyclase (TTC) comprises an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 40-55 and 191-193. In various embodiments, the TTC comprises an amino acid sequence at least 70% identical to the amino acid sequence of SEQ ID NO: 40.
[0013] In various embodiments, the heterologous enzyme pathway contains at least two copies of the TTC enzyme gene, or at least two enzymes with triterpene cyclase activity that convert 22,23-dioxidosqualene to 24,25-epoxycucurbitadienol. In such embodiments, production of cucurbitadienol can be reduced while retaining 24,25-epoxycucurbitadienol as a product. In some embodiments, the heterologous enzyme pathway contains at least one TTC enzyme comprising an amino acid sequence at least 70% identical to one of SEQ ID NO: 191, SEQ ID NO: 192, and SEQ ID NO: 193. For example, co-expression of these enzymes with SgCDS has been demonstrated to improve production of 24,25-epoxycucurbitadienol compared to expression of SgCDS alone.
[0014] In some embodiments, the heterologous enzyme pathway comprises an epoxide hydrolase (EPH). The EPH may comprise an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 56-72, 184-190, and 212. In some embodiments, the EPH may use the substrate 24,25-epoxycucurbitadienol to produce 24,25-dihydroxycucurbitadienol.
[0015] In some embodiments, the heterologous pathway includes at least one EPH that converts 24,25-epoxycucurbitadienol to 24,25-dihydroxycucurbitadienol, wherein the at least one EPH comprises an amino acid sequence at least 70% identical to at least one of: SEQ ID NO:189, SEQ ID NO:58, SEQ ID NO:184, SEQ ID NO:185, SEQ ID NO:187, SEQ ID NO:188, SEQ ID NO:190, and SEQ ID NO:212.
[0016] In some embodiments, the heterologous pathway comprises one or more oxidases active on cucurbitadienol or its oxygenated products as substrates, and can add hydroxyl groups at C11, C24, and C25 (collectively) to produce mogrol. Alternatively, or in addition, the heterologous pathway can comprise one or more oxidases that oxidize C11 of C24,25 dihydroxycucurbitadienol to produce mogrol.
[0017] In some embodiments, at least one oxidase is a cytochrome P450 enzyme. Exemplary cytochrome P450 enzymes include an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 73-91, 171-176, and 194-200.
[0018] In some embodiments, the microbial host cell expresses a heterologous enzyme pathway comprising a P450 enzyme active in oxidizing C24,25 dihydroxycucurbitadienol at C11, thereby producing mogrol. For example, in some embodiments, the cytochrome P450 comprises an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NO:194 and SEQ ID NO:171.
[0019] In various embodiments, the microbial host cell expresses one or more electron transfer proteins selected from cytochrome P450 reductase (CPR), flavodoxin reductase (FPR), and ferredoxin reductase (FDXR) sufficient to regenerate one or more oxidases. Exemplary CPR proteins are provided herein as SEQ ID NOS: 92-99 and 201.
[0020] In some embodiments, the microbial host cell expresses SEQ ID NO: 194 or a derivative thereof, and SEQ ID NO: 98 or a derivative thereof. In some embodiments, the microbial host cell expresses SEQ ID NO: 171 or a derivative thereof, and SEQ ID NO: 201 or a derivative thereof.
[0021] In some embodiments, the heterologous enzyme pathway further comprises one or more uridine diphosphate-dependent glycosyltransferase (UGT) enzymes, thereby producing one or more mogrol glycosides. In some embodiments, the mogrol glycosides may be pentaglycosylated, hexaglycosylated, etc. In other embodiments, the mogrol glycosides have two, three, or four glycosylations. The one or more mogrol glycosides may be selected from Mog.II-E, Mog.III, Mog.III-A1, Mog.III-A2, Mog.III, Mog.IV, MogIV-A, siamenoside, Mog.V, and Mog.VI. In some embodiments, the host cell produces MogV or siamenoside.
[0022] In some embodiments, the host cell expresses a UGT enzyme that catalyzes primary glycosylation of mogrol at the C24 and / or C3 hydroxyl groups, hi some embodiments, the UGT enzyme catalyzes branched glycosylation, such as beta-1,2 and / or beta-1,6 branched glycosylation, at the primary C3 and C24 glucosyl groups.
[0023] In some embodiments, at least one UGT enzyme comprises an amino acid sequence that is at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 116-165, 202-210, 211, and 213-218.
[0024] For example, in some embodiments, the microbial cell expresses at least four UGT enzymes, resulting in glucosylation of mogrol at the C3 hydroxyl group, the C24 hydroxyl group, and further 1,6 glucosylation at the C3 glucosyl group, further 1,6 glucosylation at the C24 glucosyl group, and further 1,2 glucosylation. The product of these glycosylation reactions is Mog.V.
[0025] In some embodiments, at least one UGT enzyme comprises an amino acid sequence having at least 70% sequence identity to one of SEQ ID NOs: 164, 165, 138, 204-211, and 213-218. In some embodiments, the UGT enzyme is engineered to have increased glycosyltransferase productivity compared to the wild-type enzyme.
[0026] In various embodiments, the microbial strain expresses one or more UGT enzymes capable of primary glycosylation at C24 and / or C3 of mogrol. Exemplary UGT enzymes include those comprising an amino acid sequence at least 70% identical to SEQ ID NO: 165, an amino acid sequence at least 70% identical to SEQ ID NO: 146, an amino acid sequence at least 70% identical to SEQ ID NO: 202, an amino acid sequence at least 70% identical to SEQ ID NO: 202, an amino acid sequence at least 70% identical to SEQ ID NO: 129, an amino acid sequence at least 70% identical to SEQ ID NO: 116, an amino acid sequence at least 70% identical to SEQ ID NO: 218, and an amino acid sequence at least 70% identical to SEQ ID NO: 217.
[0027] In various embodiments, the microbial strain expresses one or more UGT enzymes that can catalyze branched glycosylation of one or both primary glycosylations. Such UGT enzymes are summarized in Table 2.
[0028] In some embodiments, the microbial host cell has one or more genetic modifications that increase production of UDP-glucose, a cofactor used by UGT enzymes.
[0029] Mogrol glycoside can be recovered from the microbial culture, for example, it may be recovered from the microbial cells, or in some embodiments, it may be primarily available in the extracellular medium, where it may be recovered or sequestered.
[0030] Other aspects and embodiments of the present invention will become apparent from the following detailed disclosure. [Brief explanation of the drawings]
[0031] [Figure 1] The chemical structures of Mog.V, Mog.VI, Isomog.V, and siamenomeside are shown. The type of glycosylation (e.g., C3 or C24 core glycosylation, and 1-2, 1-4, or 1-6 glycosylation addition) is indicated within each glucose moiety. [Figure 2] The pathway for the production of Mog.V in vivo is shown. The enzymatic conversions required for each step are shown, along with the type of enzyme required. Numbers in parentheses correspond to the chemical structures in Figure 3. Abbreviations: FPP, farnesyl pyrophosphate; SQS, squalene synthase; SQE, squalene epoxidase; TTC, triterpene cyclase; EPH, epoxide hydrolase; CYP450, cytochrome P450 with reductase partner; UGT, uridine diphosphate glycosyltransferase. [Figure 3] Chemical structures of metabolites involved in Mog.V biosynthesis are shown: (1) farnesyl pyrophosphate; (2) squalene; (3) 2,3-oxidosqualene; (4) 2,3;22,23-dioxidosqualene; (5) 24,25-epoxycucurbitadienol; (6) 24,25-dihydroxycucurbitadienol; (7) mogrol; (8) mogroside V; and (9) cucurbitadienol. [Figure 4] The glycosylation pathway toward Mog.V is shown. The bubble structures represent various mogrosides. The white tetracyclic core represents mogrol. The numbers below each structure indicate the specific glycosylated mogroside. The black circles represent C3 or C24 glucosylation. The dark gray vertical circles represent 1,6-glycosylation. The light gray horizontal circles represent 1,2-glucosylation. Abbreviations: Mog, mogrol; sia, siamenomeside. [Figure 5]Figure 1 shows the results of in vivo production of squalene in E. coli using different squalene synthases. Asterisks indicate different plasmid constructs and experiments performed on different days than the other constructs shown in the figure. Legend: (1) SgSQS (SEQ ID NO: 2), (2) AaSQS (SEQ ID NO: 11), (3) EsSQS (SEQ ID NO: 16), (4) ElSQS (SEQ ID NO: 14), (5) FbSQS (SEQ ID NO: 166), (6) BbSQS (SEQ ID NO: 167). [Figure 6]
[0023] Figure 1 shows the results of in vivo production of squalene, 2,3-oxidosqualene, and 2,3;22,23-dioxidosqualene using different squalene epoxidases. Legend: (A) SEQ ID NO:2 and SEQ ID NO:168; (B) SEQ ID NO:11 and SEQ ID NO:168; (C) SEQ ID NO:2 and SEQ ID NO:169; (D) SEQ ID NO:11 and SEQ ID NO:169; (E) SEQ ID NO:2 and SEQ ID NO:170; (F) SEQ ID NO:2 and SEQ ID NO:39; (G) SEQ ID NO:11 and SEQ ID NO:39. [Figure 7] Figure 1 shows the results of in vivo production of cyclized triterpene products. The response correlates with increased expression of the enzymes in E. coli cell lines overexpressing MEP pathway enzymes. Asterisks represent fermentation experiments that were incubated for one-quarter the time of the other experiments. As shown, co-expression of SQS (SEQ ID NO: 11), SQE (SEQ ID NO: 39), and TTC (SEQ ID NO: 40) (lane G) showed significant production of the triterpenoid product, cucurbitadienol. Legend: Product 1 is squalene; product 2 is 2,3-oxidosqualene; product 3 is cucurbitadienol; (A) Expression of SEQ ID NO: 2, (B) Expression of SEQ ID NO: 11, (C) Co-expression of SEQ ID NO: 2 and SEQ ID NO: 17, (D) Co-expression of SEQ ID NO: 2 and SEQ ID NO: 169, (E) Co-expression of SEQ ID NO: 11 and SEQ ID NO: 169; (F) Co-expression of SEQ ID NO: 2, SEQ ID NO: 17, and SEQ ID NO: 40; (G) Co-expression of SEQ ID NO: 11, SEQ ID NO: 39, and SEQ ID NO: 40. [Figure 8]Results are shown for SQE engineered to produce high titers of 2,3,22,23-dioxidosqualene. Expression of SQS (SEQ ID NO: 11), SQE (SEQ ID NO: 39), and TTC (SEQ ID NO: 40) in bacterial artificial chromosomes (BACs) or integrated together results in the production of large amounts of cucurbitadienol. Point mutations in SQE (SEQ ID NO: 39) were screened for complementary to SQE to reduce cucurbitadienol levels, with corresponding increases in 2,3,22,23-dioxidosqualene titers. Two variants are shown in Figure 8. SQE A4 (including H35R, F163A, M283L, V380L, and F395Y substitutions, SEQ ID NO: 203), and SQE C11 (including H35R, N133G, F163A, Y254F, V380L, and F395Y substitutions). [Figure 9] Figure 1 shows the production of 2,3:22,23 dioxidosqualene. Titers are plotted for each strain producing 2,3:22,23 dioxidosqualene. An engineered squalene epoxidase gene, SEQ ID NO:203, was expressed in the strain producing 2,3 oxidosqualene via the squalene epoxidase of SEQ ID NO:39. The strain was incubated for 48 hours before extraction. Lanes: (1) Expression of SQE of SEQ ID NO:39; (2) Expression of SQEs of SEQ ID NO:39 and SEQ ID NO:203. [Figure 10] Co-expression of SQS, SQE, and TTC enzymes is shown. Co-expression of the CDS of SEQ ID NO:40 with SQS (SEQ ID NO:11), SQE (SEQ ID NO:39), and SQE A4 (SEQ ID NO:203) in E. coli resulted in cucurbitadienol and 24,25-epoxycucurbitadienol. E. coli strains co-expressing SQS (SEQ ID NO:11), SQE (SEQ ID NO:39), SQE A4 (SEQ ID NO:203), and CDS (SEQ ID NO:40) with additional TTC produced high levels of 24,25-epoxycucurbitadienol. Legend: TTC1 is SEQ ID NO:92, TTC2 is SEQ ID NO:191, TTC3 is SEQ ID NO:193, and TTC4 is SEQ ID NO:40. [Figure 11]Figure 1 shows the production of cucurbitadienol and 24,25-epoxycucurbitadienol. An E. coli strain producing oxidosqualene and dioxidosqualene was supplemented with a CDS homolog and CAS genes engineered to produce cucurbitadienol. The ratio of 24,25-epoxycucurbitadienol to cucurbitadienol varied from 0.15 for Enzyme 1 (SEQ ID NO: 40) to 0.58 for Enzyme 2 (SEQ ID NO: 192), demonstrating enhanced substrate specificity for the desired 24,25-epoxycucurbitadienol product with Enzyme 2. Enzyme 3 is SEQ ID NO: 219, and Enzyme 4 is SEQ ID NO: 220. [Figure 12]
[0039] Figure 1 shows the screening of EPH enzymes for the hydration of 24,25-epoxycucurbitadienol to produce 24,25-dihydroxycucurbitadienol in an E. coli strain co-expressing SQS (SEQ ID NO: 11), SQE (SEQ ID NO: 39), SQE A4 (SEQ ID NO: 203), and TTC (SEQ ID NO: 40). These fermentation experiments were performed in 96-well plates at 30°C for 72 hours. Legend: EPH1 (SEQ ID NO: 186); EPH2 (SEQ ID NO: 212); EPH3 (SEQ ID NO: 190); EPH4 (SEQ ID NO: 187); EPH5 (SEQ ID NO: 184); EPH6 (SEQ ID NO: 185); EPH7 (SEQ ID NO: 188); EPH8 (SEQ ID NO: 189); and EPH9 (SEQ ID NO: 58). [Figure 13A] Coexpression of SQS, SQE, TTC, EPH, and P450 enzymes to produce mogrol is shown. E. coli strains expressing SEQ ID NOs: 11, 39, and 203 in conjunction with CPR-containing CDS, EPH, and P450 genes produced mogrol and oxomogrol (Figure 13A). These fermentation experiments were performed in 96-well plates at 30°C for 72 hours. Legend: (1) Coexpression of SEQ ID NOs: 40, 58, 194, and 98; (2) Coexpression of SEQ ID NOs: 40, 58, 197, and 98; (3) Coexpression of SEQ ID NOs: 40, 58, 371, and 201. [Figure 13B]Co-expression of SQS, SQE, TTC, EPH, and P450 enzymes to produce mogrol is shown. Mogrol production was verified by LC-QQQ mass spectrometry using spiked authentic standards. [Figure 13C] Co-expression of SQS, SQE, TTC, EPH, and P450 enzymes to produce mogrol is shown. Mogrol production was verified using GC-FID chromatography versus authentic standards. [Figure 14] Screening of cytochrome P450s for the oxidation of the 24,25-dihydroxycucurbitadienol-like molecule cucurbitadienol at C11 is shown. The native anchor P450 enzymes shown are: (1) SEQ ID NO:194, (2) SEQ ID NO:197, (3) SEQ ID NO:171, (4) SEQ ID NO:74, and (5) SEQ ID NO:75. In some cases, the native transmembrane domain was replaced with that of E. coli sohB (anchor 3), E. coli zipA (anchor 2), or bovine 17α (anchor 1) to improve interaction with the E. coli membrane. Each P450 was coexpressed with either CPR SEQ ID NO:98 or CPR (SEQ ID NO:201) to produce 11-hydroxycucurbitadienol. These fermentation experiments were performed in 96-well plates at 30°C for 72 hours. [Figure 15] The formation of the product obtained by oxidation with C11 is shown. [Figure 16A]Figure 1 shows the production of Mog.V using different enzyme combinations. When UGTs of SEQ ID NO:165, SEQ ID NO:146, SEQ ID NO:117, or SEQ ID NO:164 were incubated with the substrate mogrol, a pentaglycosylated product was observed. Strains: (1) expressing SEQ ID NO:165; (2) expressing SEQ ID NO:146; (3) co-expressing SEQ ID NO:165 and SEQ ID NO:146; (4) co-expressing SEQ ID NO:165, SEQ ID NO:146, and SEQ ID NO:117; and (5) co-expressing SEQ ID NO:165, SEQ ID NO:146, SEQ ID NO:117, and SEQ ID NO:164. The mogroside substrate was incubated in Tris buffer containing magnesium chloride, beta-mercaptoethanol, UDP-glucose, a single UGT, and phosphatase. Abbreviations: MogV, mogroside V. [Figure 16B] Figure 1 shows the production of Mog.V using different enzyme combinations. Extracted ion chromatograms (EICs) for the 1285.4 Da (mogroside V + H) reaction product obtained by incubation with Mog.II-E containing SEQ ID NO: 165 and SEQ ID NO: 146 and either Enzyme 1 (SEQ ID NO: 117) or Enzyme 2 (SEQ ID NO: 164). Abbreviations: MogV, mogroside V. [Figure 16C] Figure 1 shows the production of Mog.V using different enzyme combinations. Extracted ion chromatograms (EICs) for reactions of 1285.4 Da (mogroside V + H) containing SEQ ID NO: 165 and SEQ ID NO: 146 and either Enzyme 1 (SEQ ID NO: 117) or Enzyme 2 (SEQ ID NO: 164) incubated with mogrol. Abbreviations: MogV, mogroside V. [Figure 17]Figure 1 shows an in vitro assay demonstrating the conversion of mogroside substrates to more glycosylated products. Mogroside substrates were incubated in Tris buffer containing magnesium chloride, beta-mercaptoethanol, UDP-glucose, a single UGT, and a phosphatase. The panel corresponds to the use of various substrates: (A) mogrol; (B) Mog.IA; (C) Mog.IE; (D) Mog.II-E; (E) Mog.III; (F) Mog.IV-A; (G) Mog.IV; and (H) siamenoside. Enzyme 1 (SEQ ID NO: 165), Enzyme 2 (SEQ ID NO: 146), Enzyme 3 (SEQ ID NO: 116), Enzyme 4 (SEQ ID NO: 117), and Enzyme 5 (SEQ ID NO: 164). [Figure 18] Bioconversion of mogrol to mogroside-IA or mogroside-IIE is demonstrated. In the experiment, a genetically engineered E. coli strain was inoculated with 0.2 mM mogroside at 37°C. Product formation was examined after 48 hours. Values are reported relative to the empty vector control (reported values are those obtained for the detected compound minus the background level detected with the empty vector control). Products were measured by LC / M8-QQQ using authentic standards. Only enzyme 1 demonstrates the formation of mogroside-IIE. Enzymes 1-5 are SEQ ID NOs: 202, 116, 216, 217, and 218, respectively. [Figure 19A] Bioconversion of Mog.IA to Mog.IIE is shown. Genetically engineered E. coli strains (expressing either Enzyme 1, SEQ ID NO:165; Enzyme 2, SEQ ID NO:202; or Enzyme 3, SEQ ID NO:116) were grown at 37°C in fermentation medium containing 0.2 mM Mog.IA. Product formation was measured after 48 hours by LC-MS / MS using authentic standards. Values reported are in excess of empty vector controls. [Figure 19B]Bioconversion of Mog.IE to Mog.IIE is shown. Genetically engineered E. coli strains (expressing either Enzyme 1, SEQ ID NO:165; Enzyme 2, SEQ ID NO:202; or Enzyme 3, SEQ ID NO:116) were grown at 37°C in fermentation medium containing 0.2 mM Mog.IE. Product formation was measured after 48 hours using LC-MS / MS with authentic standards. Values reported are in excess of empty vector controls. [Figure 20]
[0033] Figure 1 shows the production of Mog.III or siamenoside from Mog.II-E by genetically engineered E. coli strains expressing Enzyme 1 (SEQ ID NO: 204), Enzyme 2 (SEQ ID NO: 138), or Enzyme 3 (SEQ ID NO: 206). Strains were grown at 37°C in fermentation medium containing 0.2 mM Mog.IA, and product formation was measured by LC-MS / MS after 48 hours using authentic standards. [Figure 21]
[0039] Figure 2 shows in vitro production of Mog.IIA2 in cells expressing Enzyme 1 (SEQ ID NO: 205). 0.1 mM Mog.IE was added and the reaction was incubated at 37°C for 48 hours. Data were quantified by LCMS / MS using authentic standards of each compound. [Figure 22A] Figure 1 shows production of Mog.V in E. coli. Chromatograms show production of Mog.V from genetically engineered E. coli strains expressing SEQ ID NO:11, SEQ ID NO:39, SEQ ID NO:203, SEQ ID NO:40, SEQ ID NO:189, SEQ ID NO:199, SEQ ID NO:202, SEQ ID NO:165, and SEQ ID NO:122. Strains were incubated at 30°C for 72 hours before extraction. Production of Mog.V was verified by LC-QQQ spectroscopy versus authentic standards. [Figure 22B] 1 shows the production of Mog.V in E. coli. The chromatogram shows the production of Mog.V from biological samples using an authentic standard of spiked Mog.V. [Figure 23] Bioconversion of mogroside-IIE to further glycosylated products is shown, using an engineered version of the UGT enzyme of SEQ ID NO: 164. [Figure 24]The bioconversion of Mog.IA to Mog.IIE is shown, using an engineered version of the UGT enzyme of SEQ ID NO:165. [Figure 25] The bioconversion of Mog.IE to Mog.IIE is shown, using an engineered version of the UGT enzyme of SEQ ID NO:217. [Figure 26] Amino acid alignment of CaUGT_1,6 and SgUGT94_289_3 obtained using Clustal Omega (version CLUSTAL O(1,2,4)). These sequences share 54% amino acid identity. [Figure 27] Amino acid alignment of Homo sapiens squalene synthase (HsSQS) (NCBI accession number NP_004453.3) with AaSQS (SEQ ID NO: 11) obtained using Clustal Omega (version CLUSTAL O(1.2.4)). HsSQS has a published crystal structure (PDB entry: 1EZF). These sequences share 42% amino acid identity. [Figure 28] Amino acid alignment of Homo sapiens squalene epoxidase (HsSQE) (NCBI accession number XP_011515548) with MlSQE (SEQ ID NO: 39) obtained using Clustal Omega (version CLUSTAL O(1.2.4)). HsSQE has a published crystal structure (PDB entry: 6C6N). These sequences share 35% amino acid sequence identity. DETAILED DESCRIPTION OF THE INVENTION
[0032] In various aspects and embodiments, the present invention provides microbial strains and methods for producing mogrol and mogrol glycosides using recombinant microbial processes. In other aspects, the present invention provides methods for producing products, including foods, beverages, and (among other things) sweeteners, that incorporate mogrol glycosides produced according to the methods described herein. In yet other aspects, the present invention provides engineered UGT enzymes that glycosylate secondary metabolite substrates, such as mogrol or mogrosides.
[0033] As used herein, the terms "terpene or triterpene" are used interchangeably with the terms "terpenoid" or "triterpenoid," respectively.
[0034] In various aspects, the present invention provides microbial strains that produce mogrol or its glycoside products, and methods for producing them. The present invention provides recombinant microbial host cells that express heterologous enzyme pathways that catalyze the conversion of isopentenyl pyrophosphate (IPP) and / or dimethylallyl pyrophosphate (DMAPP) to one or more mogrol or mogroside(s).
[0035] Microbial host cells in various embodiments can be prokaryotic or eukaryotic. In some embodiments, the microbial host cell is a bacterium, which can be any selected from the genera Escherichia, Bacillus, Corynebacterium, Rhodobacter, Zymomonas, Vibrio, and Pseudomonas. For example, in some embodiments, the bacterial host cell is a species selected from Escherichia coli, Bacillus subtilis, Corynebacterium glutamicum, Rhodobacter capsulatus, Rhodobacter sphaeroides, Zymomonas mobilis, Vibrio natriegens, or Pseudomonas putida. In some embodiments, the bacterial host cell is E. coli. Alternatively, the microbial cell can be a yeast cell such as, but not limited to, a Saccharomyces, Pichia, or Yarrowia species, for example, Saccharomyces cerevisiae, Pichia pastoris, and Yarrowia lipolytica.
[0036] Microbial cells produce MEP or MVA products that serve as substrates for heterologous enzyme pathways. The MEP (2-C-methyl-D-erythritol 4-phosphate) pathway, also known as the MEP / DOXP (2-C-methyl-D-erythritol 4-phosphate / 1-deoxy-D-xylulose 5-phosphate) pathway, or the mevalonate-independent pathway, refers to a pathway that converts glyceraldehyde-3-phosphate and pyruvate to IPP and DMAPP. Pathways present in bacteria generally involve the following enzymes: 1-deoxy-D-xylulose-5-phosphate synthase (Dxs), 1-deoxy-D-xylulose-5-phosphate reductoisomerase (IspC), 4-diphosphocytidyl-2-C-methyl-D-erythritol synthase (IspD), 4-diphosphocytidyl-2-C-methyl-D-erythritol kinase (IspE), 2C-methyl-D-erythritol 2,4-cyclodiphosphate synthase (IspF), 1-hydroxy-2-methyl-2-(E)-butenyl 4-diphosphate synthase (IspG), and isopentenyl diphosphate isoform synthase (IspG). The MEP pathway includes the action of isoenzyme catalyzed by isoenzyme synthesis enzyme (IspH). The MEP pathway and the genes and enzymes that comprise the MEP pathway are described in U.S. Pat. No. 8,512,988, the entire contents of which are incorporated herein by reference. For example, genes that comprise the MEP pathway include dxs, ispC, ispD, ispE, ispF, ispG, ispH, idi, and ispA. In some embodiments, host cells express or overexpress one or more of dxs, ispC, ispD, ispE, ispF, ispG, ispH, idi, ispA, or modified variants thereof, and enhance production of IPP and DMAPP. In some embodiments, the triterpenoid (e.g., squalene, mogrol, or other intermediate described herein) is produced, at least in part, by metabolic flux through the MEP pathway, and the host cell has at least one additional gene copy of one or more of dxs, ispC, ispD, ispE, ispF, ispG, ispH, idi, ispA, or modified variants thereof.
[0037] The MVA pathway refers to the biosynthetic pathway that converts acetyl-CoA into IPP. The mevalonate pathway present in yeast generally involves the following steps: (a) condensation of two molecules of acetyl-CoA to acetoacetyl-CoA (e.g., by the action of acetoacetyl-CoA thiolase); (b) condensation of acetoacetyl-CoA with acetyl-CoA to form hydroxymethylglutaryl-coenzyme A (HMG-CoA) (e.g., by the action of HMG-CoA synthase (HMGS)); (c) conversion of HMG-CoA to mevalonate (e.g., by the action of HMG-CoA reductase (HMGR)); (d) phosphorylation of mevalonate to mevalonate 5-phosphate (e.g., by the action of mevalonate kinase (MK)); (e) conversion of mevalonate 5-phosphate to mevalonate 5-pyrophosphate (e.g., by the action of phosphomevalonate kinase (PMK)); and (f) conversion of mevalonate to mevalonate 5-pyrophosphate (e.g., by the action of phosphomevalonate kinase (PMK)). The MVA pathway includes an enzyme that catalyzes the conversion of 5-pyrophosphate to isopentenyl pyrophosphate (e.g., by the action of mevalonate pyrophosphate decarboxylase (MPD)). The MVA pathway and the genes and enzymes that comprise it are described in U.S. Pat. No. 7,667,017, the entire contents of which are incorporated herein by reference. In some embodiments, the host cell expresses or overexpresses one or more of acetoacetyl-CoA thiolase, HMGS, HMGR, MK, PMK, and MPD, or modified variants thereof, to enhance production of IPP and DMAPP. In some embodiments, at least a portion of the triterpenoids (e.g., mogrol or squalene) are produced by metabolic flux through the MVA pathway, and the host cell also has at least one additional gene copy of one or more of acetoacetyl-CoA thiolase, HMGS, HMGR, MK, PMK, MPD, or modified variants thereof.
[0038] In some embodiments, the host cell is a bacterial host cell genetically engineered to enhance production of IPP and DMAPP from glucose, as described in U.S. Pat. Nos. 10,480,015 and 10,662,442, the entire contents of which are incorporated herein by reference. For example, in some embodiments, the host cell overexpresses MEP pathway enzymes and balances expression of push / pull carbon flux to IPP and DMAP. In some embodiments, the host cell is genetically engineered to increase the availability or activity of Fe-S cluster proteins, thereby supporting high activity of the Fe-S enzymes IspG and IspH. In some embodiments, the host cell is genetically engineered to overexpress IspG and IspH, thereby increasing carbon flux to the 1-hydroxy-2-methyl-2-(E)-butenyl 4-diphosphate (HMBPP) intermediate, while balancing expression to prevent HMBPP accumulation to an amount that reduces cell growth or viability or inhibits MEP pathway flux and / or terpenoid production. In some embodiments, the host cells exhibit high activity of IspH relative to IspG. In some embodiments, the host cells are genetically engineered to downregulate the ubiquinone biosynthetic pathway, e.g., by reducing the expression or activity of IspB, which utilizes IPP and FPP substrates. do.
[0039] In various embodiments, the heterologous enzyme pathway comprises recombinantly expressed farnesyl diphosphate synthase (FPPS) and squalene synthase (SQS), in which the SQS comprises an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 2-16, 166, and 167.
[0040] By way of example, but not limitation, the FPPS may be Saccharomyces cerevisiae farnesyl pyrophosphate synthase (ScFPPS) (SEQ ID NO: 1), or a modified variant thereof. Modified variants may comprise an amino acid sequence at least 70% identical to SEQ ID NO: 1. For example, the FPPS may comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 1. In some embodiments, the FPPS has an amino acid sequence containing 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, relative to SEQ ID NO: 1, where the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions. Numerous other FPPS enzymes are known in the art and may be used to convert IPP and / or DMAPP to farnesyl diphosphate in accordance with this embodiment.
[0041] In some embodiments, the SQS comprises an amino acid sequence at least 70% identical to SEQ ID NO: 11. For example, the SQS can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 11. In some embodiments, the SQS comprises an amino acid sequence containing 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, with respect to SEQ ID NO: 11, independently selected from amino acid substitutions, deletions, and insertions. Amino acid modifications may be made to enhance enzyme expression or stability in microbial cells or to increase enzyme productivity. As shown in Figure 5, the AaSQS exhibits high activity in E. coli.
[0042] In some embodiments, the SQS comprises an amino acid sequence at least 70% identical to SEQ ID NO:2. For example, the SQS can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:2. In some embodiments, the SQS comprises an amino acid sequence containing 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, with respect to SEQ ID NO:2, independently selected from amino acid substitutions, deletions, and insertions. Amino acid modifications may be made to enhance enzyme expression or stability in microbial cells or to increase enzyme productivity. As shown in Figure 5, SgSQS exhibits high activity in E. coli.
[0043] In some embodiments, the SQS comprises an amino acid sequence at least 70% identical to SEQ ID NO: 14. For example, the SQS can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 14. In some embodiments, the SQS comprises an amino acid sequence comprising 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, with respect to SEQ ID NO: 14, independently selected from amino acid substitutions, deletions, and insertions. Amino acid modifications may be made to enhance enzyme expression or stability in microbial cells or to enhance enzyme productivity. As shown in Figure 5, E1SQS exhibited activity in E. coli.
[0044] In some embodiments, the SQS is an amino acid sequence that is at least 70% identical to SEQ ID NO:16. The SQS may comprise an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 16. In some embodiments, the SQS comprises an amino acid sequence that contains 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, relative to SEQ ID NO: 16, where the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions. Amino acid modifications may be made to enhance enzyme expression or stability in microbial cells or to increase enzyme productivity. As shown in Figure 5, EsSQS exhibited activity in E. coli.
[0045] In some embodiments, the SQS comprises an amino acid sequence at least 70% identical to SEQ ID NO: 166. For example, the SQS can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 166. In some embodiments, the SQS comprises an amino acid sequence containing 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, with respect to SEQ ID NO: 166, independently selected from amino acid substitutions, deletions, and insertions. Amino acid modifications may be made to enhance enzyme expression or stability in microbial cells or to increase enzyme productivity. As shown in Figure 5, FbSQS exhibited activity in E. coli.
[0046] In some embodiments, the SQS comprises an amino acid sequence at least 70% identical to SEQ ID NO: 167. For example, the SQS can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 167. In some embodiments, the SQS comprises an amino acid sequence containing 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, with respect to SEQ ID NO: 167, independently selected from amino acid substitutions, deletions, and insertions. Amino acid modifications may be made to enhance enzyme expression or stability in microbial cells or to increase enzyme productivity. As shown in Figure 5, BbSQS exhibited activity in E. coli.
[0047] Amino acid modifications to SQS enzymes can be derived from available enzyme structures and homology models, such as those described in Aminfar and Tohidfar, "In silico analysis of squalene synthase in the Fabaceae family using bioinformatics tools," J. Genetic Engineer. and Biotech. 16 (2018) 739-747. The publicly available crystal structure of HsSQE (PDB entry: 6C6N) can be used to inform amino acid modifications. An alignment of AaSQS and HsSQS is shown in Figure 27. The enzymes share 42% amino acid identity.
[0048] In some embodiments, the host cell expresses one or more enzymes that produce mogrol from squalene. For example, the host cell may express one or more squalene epoxidase (SQE) enzymes, one or more triterpenoid cyclases, one or more epoxide hydrolase (EPH) enzymes, one or more cytochrome P450 oxidases (CYP45Q), optionally one or more non-heme iron-dependent oxygenases, and one or more cytochrome P450 reductases (CPRs). As shown in Figure 2, heterologous pathways can proceed via several routes to mogrol, which may involve one or two epoxidations of a core substrate. In some embodiments, the pathways proceed via cucurbitadienol and, in some embodiments, do not involve a further epoxidation step. In some embodiments, the cucurbitadienol intermediate is converted to 24,25-epoxycucurbitadienol (5) with one or more epoxidase enzymes (such as that provided herein as SEQ ID NO: 221). In still other embodiments, the pathway is predominantly 2,3;24,25 The process proceeds via dioxidosqualene with negligible or minimal production of the cucurbitadienol intermediate. In some embodiments, one or more of the enzymes SQE, CDS, EPH, CYP450, non-heme iron-dependent oxygenase, flavodoxin reductase (FPR), ferredoxin reductase (FDXR), and CPR are engineered to increase flux to mogrol.
[0049] In some embodiments, the heterologous enzyme pathway comprises two squalene epoxidase (SQE) enzymes. For example, the heterologous enzyme pathway can comprise an SQE that produces 2,3-oxidosqualene (intermediate (3) in Figure 2). In some embodiments, the SQE produces 2,3;22,23-dioxidosqualene (intermediate (4) in Figure 2), and this conversion can be catalyzed by the same SQE enzyme or by enzymes that differ in that their amino acid sequences contain at least one amino acid modification. For example, the squalene epoxidase enzyme can comprise at least two SQE enzymes, each of which (independently) comprises an amino acid sequence at least 70% identical to any one of SEQ ID NOS: 17-39, 168-170, and 177-183. Co-expression of SQE enzymes engineered or screened for substrate specificity for 2,3-oxidosqualene can produce the diepoxy intermediate with low or minimal levels of cucurbitadienol. In these embodiments, the P450 oxygenase enzymes that hydroxylate C24 and C25 of the scaffold can be eliminated.
[0050] In some embodiments, at least one SQE comprises an amino acid sequence that is at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 39. For example, the SQE enzyme can comprise an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 39, where the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0051] As shown in Figure 6, MlSQE exhibits high activity in E. coli when coexpressed with AaSQS, which produces high levels of a single epoxidation product (2,3-oxidosqualene). Therefore, coexpression of AaSQS (or an engineered derivative) with multiple copies of MlSQE engineered as described above offers excellent potential for bioengineering the mogrol pathway. See Figure 9. Amino acid modifications can be made to enhance the expression or stability of the SQE enzyme in microbial cells or to increase the productivity of the enzyme.
[0052] In some embodiments, the host cell comprises two squalene epoxidase enzymes, each comprising an amino acid sequence at least 70% identical to the Methylomonas lenta squalene epoxidase (SEQ ID NO: 39). For example, one of the SQE enzymes has one or more amino acid modifications compared to the amino acid sequence of SEQ ID NO: 39 that improve conversion of 2,3-oxidosqualene to 2,3;22,23 dioxidosqualene and productivity. In some embodiments, the SQE enzyme comprises one or more (or in some embodiments, two, three, four, five, six, or seven) modifications at the following positions: 35, 133, 163, 254, 283, 380, and 395 of SEQ ID NO: 39. For example, the position corresponding to position 35 of SEQ ID NO: 39 may be an arginine or a lysine (e.g., H35R). The amino acid at the position corresponding to position 133 of SEQ ID NO:39 can be glycine, alanine, leucine, isoleucine, or valine (e.g., N133G). The amino acid at the position corresponding to position 163 of SEQ ID NO:39 can be glycine, alanine, leucine, isoleucine, or valine (e.g., F163A). The amino acid at the position corresponding to position 254 of SEQ ID NO:39 can be phenylalanine, alanine, leucine, isoleucine, or valine (e.g., Y254F). The amino acid at the position corresponding to position 283 of SEQ ID NO:39 can be phenylalanine, alanine, leucine, isoleucine, or valine (e.g., Y254F). The amino acid at a position corresponding to position 380 of SEQ ID NO:39 may be alanine, leucine, isoleucine, or valine (e.g., M283L). The amino acid at a position corresponding to position 380 of SEQ ID NO:39 may be alanine, leucine, or glycine (e.g., V280L). The amino acid at a position corresponding to position 395 of SEQ ID NO:39 may be tyrosine, serine, or threonine (e.g., F395Y). Exemplary SQE enzymes in these embodiments are at least 70%, or at least 80%, or at least 90%, or at least 95% identical to SEQ ID NO:39, but include the following set of amino acid substitutions: H35R, F163A, M283L, V380L, F395Y; or H35R, N133G, F163A, Y254F, V380L, and F395Y, and in each case are numbered according to SEQ ID NO:39. For example, the host cell may express an SQE comprising the amino acid sequence of SEQ ID NO: 203 (referred to herein as MlSQE A4).
[0053] In yet another embodiment, the squalene epoxidase comprises an amino acid sequence at least 70% identical to SEQ ID NO: 168. For example, the SQE can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 168. In various embodiments, the SQE has an amino acid sequence containing 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, with respect to SEQ ID NO: 168, independently selected from amino acid substitutions, deletions, and insertions. As shown in FIG. 6, BaESQE exhibited good activity in E. coli. Amino acid modifications may be performed to enhance expression or stability of the enzyme in microbial cells or to increase enzyme productivity.
[0054] In some embodiments, the squalene epoxidase comprises an amino acid sequence at least 70% identical to SEQ ID NO: 169. For example, the SQE can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 169. In various embodiments, the SQE has an amino acid sequence containing 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, with respect to SEQ ID NO: 169, independently selected from amino acid substitutions, deletions, and insertions. As shown in FIG. 6, MsSQE exhibited good activity in E. coli. Amino acid modifications may be performed to enhance expression or stability of the enzyme in microbial cells or to increase enzyme productivity.
[0055] In some embodiments, the squalene epoxidase comprises an amino acid sequence at least 70% identical to SEQ ID NO: 170. For example, the SQE can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 170. In various embodiments, the SQE has an amino acid sequence containing 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, with respect to SEQ ID NO: 170, independently selected from amino acid substitutions, deletions, and insertions. As shown in FIG. 6, MbSQE exhibited good activity in E. coli. Amino acid modifications may be performed to enhance expression or stability of the enzyme in microbial cells or to increase enzyme productivity.
[0056] Amino acid modifications were determined according to Padyana AK, et al., Structure and Inhibition mechanism of the catalytic domain of human squalene epoxidase, Nat.Comm.(2019)Vol.10(97):1-10; or Ruckenstulh et al.,Structure-Function Correlations This can be guided by available enzyme structures and homology models, such as those described in "Two Highly Conserved Motifs in Saccharomyces cerevisiae Squalene Epoxidase," Antimicrob. Agents and Chemo. (2008) Vol. 52(4):1496-1499. Figure 28 shows an alignment of HsSQE with MlSEQ, which is useful for defining the genetic engineering of the enzyme for expression, stability, and productivity in microbial host cells. The two enzymes share 35% identity.
[0057] In various embodiments, the heterologous enzyme pathway includes a triterpene cyclase (TTC). In some embodiments, in which the microbial cell co-expresses FPPS in conjunction with the SQS, SQE, and triterpene cyclase enzymes, the microbial cell produces 2,3;22,23-dioxidosqualene. 2,3;22,23-dioxidosqualene can be a substrate for downstream enzymes in the heterologous pathway. In some embodiments, the triterpene cyclase (TTC) comprises an amino acid sequence at least 70%, or at least 80%, or at least 90%, or at least 95% identical to an amino acid sequence selected from SEQ ID NOs: 40-55, 191-193, and 219-220. In various embodiments, the TTC comprises an amino acid sequence at least 70% identical to the amino acid sequence of SEQ ID NO: 40. In some embodiments, TTC comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 40. For example, TTC has an amino acid sequence that has 1 to 20 amino acid modifications relative to SEQ ID NO: 40, where the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0058] In some embodiments, TTC comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 192. For example, TTC has an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 192, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions. The enzyme defined by SEQ ID NO: 192 exhibits improved specificity for the production of 24,25-epoxycucurbitadienol (Figure 11).
[0059] In various embodiments, the heterologous enzyme pathway includes at least two copies of the TTC enzyme gene, or at least two enzymes with triterpene cyclase activity that convert 22,23-dioxidosqualene to 24,25-epoxycucurbitadienol. In such embodiments, the product can contain 24,25-epoxycucurbitadienol while reducing the production of cucurbitadienol.
[0060] In some embodiments, the heterologous enzyme pathway includes at least one TTC comprising an amino acid sequence at least 70% identical to one of SEQ ID NO:191, SEQ ID NO:192, and SEQ ID NO:193. These enzymes can optionally be co-expressed with SgCDS. These enzymes are notable for the production of 24,25-epoxycucurbitadienol. Figure 10. Thus, in some embodiments, at least one TTC comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to any of SEQ ID NOs:191, 192, and 193. In some embodiments, the TTC has an amino acid sequence with 1 to 20 amino acid modifications relative to one of SEQ ID NOs:191, 192, and 193, where the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0061] To enhance the expression or stability of an enzyme in a microbial cell, or to enhance the productivity of the enzyme, To achieve this, amino acid modifications can be performed. Available enzyme structures and homology models, such as those described in Itkin M., et al., The biosynthetic pathway of the nonsugar, high-intensity sweetener mogroside V from Siraitia grosvenorii, PNAS (2016) Vol 113(47):E7619-E7628, can guide the amino acid modifications. For example, the CDS can be modeled using the structure of human lanosterol synthase (oxidosqualene cyclase) (PDB 1W6K).
[0062] In various embodiments, cucurbitadienol (intermediate 9 in Figure 2) is converted to 24,25-epoxycucurbitadienol (5) by one or more enzymes expressed in a host cell. For example, a heterologous pathway can include an enzyme having at least about 70%, or at least about 80%, or at least about 85%, or at least about 90%, or at least about 95%, or at least about 97%, 98%, or 99% sequence identity to SEQ ID NO:221.
[0063] In some embodiments, the heterologous enzyme pathway comprises at least one epoxide hydrolase (EPH). The EPH may comprise an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 56-72, 184-190, and 212. In some embodiments, the EPH may use the substrate 24,25-epoxycucurbitadienol (intermediate (5) in Figure 2) to produce 24,25-dihydroxycucurbitadienol (intermediate (6) in Figure 2). In some embodiments, the EPH comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 56-72, 184-190, and 212. Thus, in some embodiments, the EPH has an amino acid sequence with 1 to 20 amino acid modifications relative to one of SEQ ID NOs: 56-72, 184-190, and 212, where the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0064] In some embodiments, the heterologous pathway includes at least one EPH enzyme that converts 24,25-epoxycucurbitadienol to 24,25-dihydroxycucurbitadienol, wherein the at least one EPH enzyme comprises an amino acid sequence at least 70% identical to one of: SEQ ID NO: 189, SEQ ID NO: 58, SEQ ID NO: 184, SEQ ID NO: 185, SEQ ID NO: 187, SEQ ID NO: 188, SEQ ID NO: 190, and SEQ ID NO: 212. See Figure 12. In some embodiments, the EPH enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 189, 58, 184, 185, 187, 188, 190, and 212. For example, the EPH has an amino acid sequence with 1 to 20 amino acid modifications, independently selected from amino acid substitutions, deletions, and insertions, relative to one of SEQ ID NOs: 189, 58, 184, 185, 187, 188, 190, and 212. Amino acid modifications may be made to enhance expression or stability of the enzyme in a microbial cell, or to enhance enzyme productivity.
[0065] In some embodiments, the heterologous pathway comprises one or more oxidases active on cucurbitadienol as a substrate or its oxygenated products, which hydroxylate at C11, C24, and C25 (collectively) to produce mogrol (see Figure 2). Alternatively, the heterologous pathway can comprise one or more oxidases that oxidize C11 of C24,25 dihydroxycucurbitadienol to produce mogrol.
[0066] In some embodiments, at least one oxidase is a cytochrome P450 enzyme. Exemplary cytochrome P450 enzymes include an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 73-91, 171-176, and 194-200. In some embodiments, at least one P450 enzyme includes an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 73-91, 171-176, and 194-200. For example, at least one cytochrome P450 enzyme has an amino acid sequence with 1 to 20 amino acid modifications relative to one of SEQ ID NOs: 73-91, 171-176, and 194-200, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0067] In some embodiments, the microbial host cell expresses a heterologous enzyme pathway comprising a P450 enzyme active in oxidizing C24,25-dihydroxycucurbitadienol at C11, thereby producing mogrol. For example, in some embodiments, the cytochrome P450 comprises an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NO: 194 and SEQ ID NO: 171. See Figures 13A-C, 14, and 15. In some embodiments, the microbial host cell expresses a cytochrome P450 enzyme comprising an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 194 and 171. In some embodiments, at least one cytochrome P450 enzyme has an amino acid sequence with 1 to 20 amino acid modifications relative to one of SEQ ID NOs: 194 and 171, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0068] In some embodiments, the cytochrome P450 enzyme has at least a portion of its transmembrane region replaced with a heterologous transmembrane region. For example, particularly in embodiments where the microbial cell is a bacterium, the CYP450 and / or CPR are modified in their entirety as described in US 2018 / 0251738, the entire contents of which are incorporated herein by reference. For example, in some embodiments, the CYP450 enzyme has a deletion of all or a portion of the wild-type P450 N-terminal transmembrane region and the addition of a transmembrane domain derived from an E. coli or bacterial inner membrane, cytoplasmic C-terminal protein. In some embodiments, the transmembrane domain is a single-pass transmembrane domain. In some embodiments, the transmembrane domain is a multiple-pass (e.g., two, three, or more transmembrane helix) transmembrane domain. Exemplary transmembrane domains are derived from E. coli zipA or sohB. Alternatively, the P450 enzyme can use its native transmembrane anchor or the well-known bovine 17α anchor. Please refer to Figure 14.
[0069] In some embodiments, the microbial host cell expresses a non-heme iron oxidase. Exemplary non-heme iron oxidases include an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 100-115. In some embodiments, the non-heme iron oxidase has an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 100-115.
[0070] In various embodiments, the microbial host cell expresses one or more electron transfer proteins selected from cytochrome P450 reductase (CPR), flavodoxin reductase (FPR), and ferredoxin reductase (FDXR) sufficient to regenerate one or more oxidases. Exemplary CPR proteins are provided herein as SEQ ID NOS: 92-99, and 201.
[0071] In some embodiments, the microbial host cell expresses a cytochrome P450 reductase, which may comprise an amino acid sequence at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 92-99, and 201. For example, in some embodiments, the microbial host cell expresses (i.e., has at least 70%, at least 80%, or at least 90% sequence identity to) SEQ ID NO: 194 or a derivative thereof (described above), and SEQ ID NO: 98 or a derivative thereof. In some embodiments, the microbial host cell expresses (i.e., has at least 70%, at least 80%, or at least 90% sequence identity to) SEQ ID NO: 171 or a derivative thereof (described above), and SEQ ID NO: 201 or a derivative thereof.
[0072] In various embodiments, the heterologous enzyme pathway produces mogrol, which may be an intermediate for a downstream enzyme in the heterologous pathway, or in some embodiments, is recovered from the culture. Mogrol may be recovered from the host cell and / or from the culture medium in some embodiments.
[0073] In some embodiments, the heterologous enzyme pathway further comprises one or more uridine diphosphate-dependent glycosyltransferase (UGT) enzymes, thereby producing one or more mogrol glycosides (or "mogrosides"). The mogrol glycosides may be pentaglycosylated, hexaglycosylated, or more (e.g., seven, eight, or nine glycosylations) in some embodiments. In other embodiments, the mogrol glycosides have two, three, or four glycosylations. The one or more mogrol glycosides may be selected from Mog.II-E, Mog.III, Mog.III-A1, Mog.III-A2, Mog.III, Mog.IV, Mog.IV-A, siamenoside, isomog.V, Mog.V, or Mog.VI. In some embodiments, the host cell produces Mog.V or siamenoside.
[0074] In some embodiments, the host cell expresses a UGT enzyme that catalyzes the primary glycosylation of mogrol at the C24 and / or C3 hydroxyl groups. In some embodiments, the UGT enzyme catalyzes branched glycosylation, such as beta-1,2 and / or beta-1,6 branched glycosylation, at the primary C3 and C24 glucosyl groups. UGT enzymes that have been shown to catalyze primary glycosylation at the C24 and / or C3 hydroxyl groups are summarized in Table 1. UGT enzymes that have been shown to catalyze various branched glycosylation reactions are summarized in Table 2.
[0075] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 116-165, 202-210, 211, and 213-218. For example, in some embodiments, the UGT enzyme comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 116-165, 202-210, 211, and 213-218. Thus, at least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications relative to one of SEQ ID NOs: 116-165, 202-210, 211, and 212-218, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0076] For example, in some embodiments, the microbial cell expresses at least four UGT enzymes, resulting in glucosylation of mogrol at the C3 hydroxyl group, the C24 hydroxyl group, and the C3 hydroxyl group. , and further 1,6 glucosylation at the C3 glucosyl group, and further 1,6 glucosylation and 1,2 glucosylation at the C24 glucosyl group. The product of such glycosylation reactions is Mog.V.
[0077] In some embodiments, at least one UGT enzyme comprises an amino acid sequence having at least 70% sequence identity to one of SEQ ID NOs: 164, 165, 138, 204-211, and 213-218.
[0078] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to UGT85C1 (SEQ ID NO: 165). UGT85C1 exhibits primary glycosylation at the C3 and C24 hydroxyl groups. Thus, in some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 165. At least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications with respect to SEQ ID NO: 165, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions. Exemplary amino acid substitutions include substitutions at positions 41 (e.g., L41F or L41Y), 49 (e.g., D49E), and 127 (e.g., C127F or C127Y).
[0079] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 164, which exhibits activity for adding branched glycosylations, both 1-2 and 1-6 branched glycosylations. In various embodiments, at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 164. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence having 1-20 amino acid modifications with respect to SEQ ID NO: 164, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions. Exemplary amino acid substitutions are shown in Table 3. Exemplary amino acid substitutions for SEQ ID NO: 164 include substitutions at one or more positions selected from 150 (e.g., S150F, S150Y), 147 (e.g., T147L, T147V, T147I, and T147A), 207 (e.g., N207K or N207R), 270 (e.g., K270E or K270D), 281 (V281L or V281I), 354 (e.g., L354V or L354I), 13 (e.g., L13F or L13Y), 32 (T32A or T32G or T32L), and 101 (K101A or K101G). Exemplary engineered UGT enzymes for SEQ ID NO: 164 include the amino acid substitutions T147L and N207K.
[0080] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 138, which exhibits activity for catalyzing 1-6 branched glycosylation. In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 138. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications with respect to SEQ ID NO: 138, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0081] In some embodiments, the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 204, which catalyzes 1-6 branched glycosylation, particularly at the C3 primary glycosylation. For example, the at least one UGT enzyme comprises: The at least one UGT enzyme may comprise an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 204. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 204, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0082] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 205, which catalyzes 1-6 branched glycosylation at both C3 and C24 primary glycosylation. For example, at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 205. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 205, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0083] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 206, which catalyzes 1-2 and 1-6 branched glycosylation. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 206. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 206, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0084] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 207, which catalyzes 1-6 branched glycosylation of primary glycosylation. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 207. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications with respect to SEQ ID NO: 207, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0085] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 208, which catalyzes 1-2 and 1-6 branched glycosylation. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 208. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 208, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0086] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:209, which catalyzes 1-6 branched glycosylation of primary glycosylation. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:209. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence having 1-20 amino acid modifications with respect to SEQ ID NO:209, which catalyzes 1-6 branched glycosylation of primary glycosylation. The amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0087] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 210, which catalyzes 1-6 branched glycosylation of the primary glycosylation. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 210. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications with respect to SEQ ID NO: 210, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0088] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 211, which catalyzes 1-2 branched glycosylation of C24 primary glycosylation. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 211. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 211, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0089] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 213, which catalyzes 1-6 branched glycosylation of C24 primary glycosylation. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 213. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 213, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0090] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 214, which catalyzes the primary glucosylation of C24. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 214. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 214, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0091] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 215, which catalyzes 1-6 branched glycosylation at C24. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 215. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO: 215, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0092] In still other embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 146, which is responsible for glucosylation of the C24 hydroxyl group of mogrol or Mog.1E. At least one UGT enzyme comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 146. In some embodiments, at least one UGT enzyme comprises an amino acid sequence with 1-20 or 1-10 amino acid modifications relative to SEQ ID NO: 146, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions. The amino acid modifications may be made to enhance expression or stability of the enzyme in a microbial cell or to increase the productivity of the enzyme for a particular substrate.
[0093] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 202, which catalyzes primary glycosylation at the C3 and C24 hydroxyl groups. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 202. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications relative to SEQ ID NO: 202, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0094] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 218, which catalyzes a primary glycosylation at the C24 hydroxyl group. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 218. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications relative to SEQ ID NO: 218, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0095] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 217, which catalyzes a primary glycosylation at the C24 hydroxyl group. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 217. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications relative to SEQ ID NO: 217, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions. Exemplary amino acid substitutions include substitutions at one or more positions (with respect to SEQ ID NO: 17) selected from 74 (e.g., A74E or A74D), 91 (I91F or I91Y), 101 (e.g., H101P), 241 (e.g., Q241E or Q241D), and 436 (e.g., I436L or I436A). In some embodiments, the UGT enzyme has the following amino acid substitutions with respect to SEQ ID NO: 217: A74E, I91F, and H101P.
[0096] In some embodiments, at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO: 216, which catalyzes a primary glycosylation at the C24 hydroxyl group. For example, at least one UGT enzyme can comprise an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 216. In exemplary embodiments, at least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications with respect to SEQ ID NO: 216, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0097] In some embodiments, the at least one UGT enzyme is selected from the group consisting of SEQ ID NO: 117, SEQ ID NO: 21 0, or SEQ ID NO: 122. For example, the enzyme defined in SEQ ID NO: 117 catalyzes branched glycosylation. In some embodiments, at least one UGT enzyme comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 117, SEQ ID NO: 210, or SEQ ID NO: 122. In some embodiments, at least one UGT enzyme comprises an amino acid sequence having 1 to 20 amino acid modifications with respect to SEQ ID NO: 117, 210, or 122, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
[0098] In some embodiments, the microbial cell expresses at least one UGT enzyme capable of catalyzing the beta-1,2 addition of a glucose molecule to at least a C24 glucosyl group (e.g., of Mog.IVA). Exemplary UGT enzymes according to these embodiments include SEQ ID NO:117, SEQ ID NO:147, SEQ ID NO:148, SEQ ID NO:149, SEQ ID NO:150, or SEQ ID NO:163, or derivatives thereof. Derivatives include enzymes comprising an amino acid sequence at least 70% identical to one or more of SEQ ID NO:117, SEQ ID NO:147, SEQ ID NO:148, SEQ ID NO:149, SEQ ID NO:150, and SEQ ID NO:163. In some embodiments, a UGT enzyme that catalyzes the beta-1,2 addition of a glucose molecule to at least a C24 glucosyl group comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one or more of SEQ ID NO:117, SEQ ID NO:147, SEQ ID NO:148, SEQ ID NO:149, SEQ ID NO:150, and SEQ ID NO:163. In some embodiments, at least one UGT enzyme comprises an amino acid sequence that includes 1 to 20 amino acid modifications, or 1 to 10 amino acid modifications, relative to SEQ ID NO:117, SEQ ID NO:147, SEQ ID NO:148, SEQ ID NO:149, SEQ ID NO:150, and SEQ ID NO:163, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions. The amino acid modifications may be made to enhance expression or stability of the enzyme in a microbial cell or to increase the productivity of the enzyme for a particular substrate.
[0099] In some embodiments, at least one UGT enzyme is a circular permutation of a wild-type UGT enzyme, optionally with amino acid substitutions, deletions, and / or insertions at corresponding positions in the wild-type enzyme. Circular permutations can provide novel and desirable substrate specificities, product profiles, and reaction rates compared to the wild-type enzyme. Circular permutations retain the same basic fold as the parent enzyme but have a different N-terminal position (e.g., "cleavage site"), optionally connecting the original N- and C-termini via a linking sequence. For example, in circular permutations, the N-terminal methionine is located at a site within the protein other than the native N-terminus. UGT circular permutations are described in US 2017 / 0332673, the entire contents of which are incorporated herein by reference. In some embodiments, at least one UTG enzyme is a circular permutation of, but not limited to, SEQ ID NO:146, SEQ ID NO:164, or SEQ ID NO:165, SEQ ID NO:117, SEQ ID NO:147, SEQ ID NO:148, SEQ ID NO:149, SEQ ID NO:150, SEQ ID NO:163, SEQ ID NO:202, SEQ ID NO:216, SEQ ID NO:217, and SEQ ID NO:218. In some embodiments, the circular permutation further comprises one or more amino acid modifications (e.g., amino acid substitutions, deletions, and / or insertions) relative to the parent UGT enzyme. In these embodiments, the circular permutation, when aligned with the corresponding amino acid sequence (i.e., regardless of the novel N-terminus of the circular permutation), has at least about 70%, or at least about 80%, or at least about 90%, or at least about 95%, or at least about 98% identity to the parent enzyme. An exemplary circular permutation for use according to some embodiments is SEQ ID NO:206.
[0100] In some embodiments, the microbial host cell expresses at least three UGT enzymes: mogrol The microbial host cell expresses a first UGT enzyme that catalyzes primary glycosylation at the C24 hydroxyl group, a second UGT enzyme that catalyzes primary glycosylation at the C3 hydroxyl group of mogrol, and a third UGT enzyme that catalyzes one or more branched glycosylation reactions. In some embodiments, the microbial host cell expresses one or two UGT enzymes that catalyze beta-1,2 and / or beta-1,6 branched glycosylation of the C3 and / or C24 primary glycosylation. For example, the UGT enzymes are: SEQ ID NO: 165, or a derivative thereof: SEQ ID NO: 146, or a derivative thereof; SEQ ID NO: 214, or a derivative thereof; SEQ ID NO: 129, or a derivative thereof; SEQ ID NO: 164, or a derivative thereof; SEQ ID NO: 116, or a derivative thereof; SEQ ID NO: 202, or a derivative thereof; SEQ ID NO: 218, or a derivative thereof; SEQ ID NO: 217, or a derivative thereof; SEQ ID NO: 138, or a derivative thereof; SEQ ID NO: 204, or a derivative thereof; SEQ ID NO: 205, or a derivative thereof; SEQ ID NO: 207, or a derivative thereof; SEQ ID NO: 208, or a derivative thereof; SEQ ID NO: 209, or a derivative thereof; SEQ ID NO: 11, or a derivative thereof; SEQ ID NO: 215, or a derivative thereof; SEQ ID NO: 213, or a derivative thereof; SEQ ID NO: 206, or a derivative thereof; SEQ ID NO: 122, or a derivative thereof; and SEQ ID NO: 210, or a derivative thereof; The derivative may comprise three or four UGT enzymes selected from the following: The derivative has sequence identity to the reference enzyme, as described herein.
[0101] In some embodiments, the microbial host cell has one or more genetic modifications that increase production of UDP-glucose, a cofactor used by UGT enzymes, including one or more, or two or more (or all) of: ΔgalE, ΔgalT, ΔgalK, ΔgalM, ΔushA, Δagp, Δpgm, duplication of E. coli galU, expression of Bacillus subtillus UGPA, and expression of Bifidobacterium adolescentis SPL.
[0102] Mogrol glycosides can be recovered from microbial cultures, for example, they can be recovered from microbial cells, or in some embodiments, they can be primarily available in the extracellular medium and recovered or sequestered therein.
[0103] In various embodiments, the reaction is carried out in a microbial cell, and the UGT enzyme is recombinantly expressed in the cell. In some embodiments, mogrol is produced in the cell by a heterologous mogrol synthesis pathway, as described herein. In other embodiments, mogrol or a mogrol glycoside (e.g., monk fruit extract) is provided to the cell for glycosylation. In still other embodiments, the reaction is carried out in vitro using purified UGT enzymes, partially purified UGT enzymes, or recombinant cell lysates.
[0104] As described herein, microbial host cells can be prokaryotic or eukaryotic, and optionally include Escherichia coli, Bacillus subtilis, and the like. In some embodiments, the microbial host cell is a bacterium selected from Saccharomyces, Pichia, or Yarrowia species, such as Saccharomyces cerevisiae, Pichia pastoris, and Yarrowia lipolytica. In some embodiments, the microbial host cell is E. coli.
[0105] Bacterial host cells are cultured to produce triterpenoid products (e.g., mogrosides). In some embodiments, the production stage uses C1, C2, C3, C4, C5, and / or C6 carbon substrates. In exemplary embodiments, the carbon source is glucose, sucrose, fructose, xylose, and / or glycerol. Culture conditions are generally selected from aerobic, microaerobic, and anaerobic.
[0106] In various embodiments, bacterial host cells may be cultured at temperatures between 22°C and 37°C. Commercial biosynthesis in bacteria, such as E. coli, is limited to the temperature stability of overexpressed and / or exogenous enzymes (e.g., plant-derived enzymes), but genetic engineering of recombinant enzymes and maintaining cultures at elevated temperatures can result in higher yields and overall productivity. In some embodiments, culturing is carried out at about 22°C or higher, about 23°C or higher, about 24°C or higher, about 25°C or higher, about 26°C or higher, about 27°C or higher, about 28°C or higher, about 29°C or higher, about 30°C or higher, about 31°C or higher, about 32°C or higher, about 33°C or higher, about 34°C or higher, about 35°C or higher, about 36°C or higher, or about 37°C.
[0107] In some embodiments, the bacterial host cells are further suitable for commercial production on a commercial scale. In some embodiments, the culture size is at least about 100 L, at least about 200 L, at least about 500 L, at least about 1,000 L, or at least about 10,000 L, or at least about 100,000 L, or at least about 500,000 L, or at least about 600,000 L. In certain embodiments, the culture may be performed as a batch culture, a continuous culture, or a semi-continuous culture.
[0108] In various embodiments, the method further includes recovering the product from the cell culture or cell lysate. In some embodiments, the culture produces at least about 100 mg / L, or at least about 200 mg / L, or at least about 500 mg / L, or at least about 1 g / L, or at least about 2 g / L, or at least about 5 g / L, or at least about 10 g / L, or at least about 20 g / L, or at least about 30 g / L, or at least about 40 g / L of the terpenoid or terpenoid glycoside product.
[0109] In some embodiments, indole production (e.g., prenylated indole) is used as a surrogate marker for terpenoid production and / or indole accumulation in the culture is controlled to increase production. For example, in various embodiments, indole accumulation in the culture is controlled to less than about 100 mg / L, less than about 75 mg / L, less than about 50 mg / L, less than about 25 mg / L, or less than about 10 mg / L. Indole accumulation can be controlled by balancing protein expression and activity using the multivariate modular approach described in U.S. Pat. No. 8,927,241 (incorporated herein by reference) and / or is controlled by chemical means.
[0110] Other markers for efficient production of terpenes and terpenoids include the accumulation of DOX or ME in the medium. Generally, strains are engineered to reduce the accumulation of these species, resulting in an accumulation of less than about 5 g / L, or less than about 4 g / L, or less than about 3 g / L in the medium. / L, or less than about 2 g / L, or less than about 1 g / L, or less than about 500 mg / L, or less than about 100 mg / L.
[0111] Optimizing terpene or terpenoid production through manipulation of MEP pathway genes and upstream and downstream pathways is not expected to be a simple linear or additive process. Rather, optimization is achieved by balancing components of the MEP pathway and upstream and downstream pathways through combinatorial analysis. Accumulation of indole (e.g., prenylated indole) and MEP metabolites (e.g., DOX, ME, MEcPP, and / or farnesol) in cultures can be used as surrogate markers to guide this process.
[0112] For example, in some embodiments, the bacterial strain has at least one additional copy (on a plasmid or integrated into the genome) of dxs and idi expressed as an operon / module; or dxs, ispD, ispF, and idi expressed as an operon or module, with additional MEP pathway complementation as described herein to improve MEP carbon. For example, the bacterial strain has additional copies of dxr, and ispG, and / or ispH, optionally, ispE, and / or idi, and can modulate expression of these genes to increase MEP carbon and / or enhance terpene or terpenoid titer. In various embodiments, the bacterial strain has additional copies of at least dxr, ispE, ispG, and ispH, and optionally, an additional copy of idi, and can modulate expression of these genes to increase MEP carbon and / or enhance terpene or terpenoid titer.
[0113] Manipulation of the expression of proteins, such as genes and / or gene modules, can be achieved in a variety of ways. For example, gene or operon expression can be regulated through the selection of promoters of different strengths (e.g., high strength, medium strength, or low strength), such as inducible or constitutive promoters. Some examples of promoters of different strengths include, but are not limited to, Trc, T5, and T7. Furthermore, gene or operon expression can be regulated by manipulating the copy number of the gene or operon in the cell. In some embodiments, gene or operon expression can be regulated by manipulating the order of genes within a module, where genes that are transcribed first are generally expressed at higher levels. In some embodiments, gene or operon expression is regulated by integrating one or more genes or operons into a chromosome.
[0114] Optimization of protein expression can also be achieved by selecting appropriate promoters and ribosome binding sites. In some embodiments, this can include selecting high copy number plasmids, or single copy, low copy, or medium copy number plasmids. The transcription termination step can also be targeted to modulate gene expression through the introduction or elimination of structures such as stem loops.
[0115] Expression vectors containing all the elements necessary for expression are commercially available and known to those skilled in the art. See, for example, Sambrook et al., Molecular Cloning: A See Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, 1989. Heterologous DNA is introduced into cells to genetically engineer the cells. The heterologous DNA is placed under the functional control of transcriptional elements to allow expression of the heterologous DNA within the host cell.
[0116] In contrast to gene complementation, in some embodiments, endogenous genes are edited, which involves modifying the endogenous promoter, ribosomal binding sequence, or other expression control sequences, and / or, in some embodiments, trans-acting factors and / or cis-acting factors in gene regulation. The function factor can be modified. Genome editing can be performed using CRISPR / Cas genome editing technology or similar technology using zinc finger nucleases and TALEN. In some embodiments, endogenous genes are replaced by homologous recombination.
[0117] In some embodiments, gene copy number is controlled to at least partially overexpress the gene. The gene copy number can be easily controlled using a plasmid with various copy numbers, but gene duplication and chromosomal integration can also be used. For example, the process of genetically stable tandem gene duplication is described in US 2011 / 0236927, the entire contents of which are incorporated herein by reference.
[0118] The terpene or terpenoid product can be recovered by any suitable process. For example, the aqueous phase can be recovered for further processing and / or the whole cell biomass can be recovered. The production of the desired product can be measured and / or quantified, for example, by gas chromatography (e.g., GC-MS). The desired product can be produced in a batch system or a continuous bioreactor system.
[0119] The similarity between nucleotide and amino acid sequences, i.e., the percentage of sequence identity, can be determined through sequence alignment. Such alignment can be performed using several algorithms known in the art, such as the mathematical algorithm of Karlin and Altschul (Karlin & Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877), hmmalign (HMMER package, http: / / hmmer.wustl.edu / ), or the CLUSTAL algorithm (Thompson, JD, Higgins, DG & Gibson, TJ (1994) Nucleic Acids Res. 22, 4673-80). The grade of sequence identity (sequence matching) can be calculated, for example, using BLAST, BLAT, or BlastZ (or BlastX). A similar algorithm is incorporated into the BLASTN and BLASTP programs of Altschul et al (1990) J. Mol. Biol. 215:403-410. BLAST polynucleotide searches can be performed with the BLASTN program, score=100, wordlength=12.
[0120] BLAST protein searches can be performed using the BLASTP program, score = 50, word length = 3. To obtain gapped alignments for comparison purposes, gapped BLAST is used as described in Altschul et al (1997) Nucleic Acids Res. 25:3389-3402. When using BLAST and gapped BLAST programs, the default parameters of each program are used. Sequence matching analysis can be supplemented with established homology mapping techniques such as Shuffle-LAGAN (Brudno M., Bioinformatics 2003b, 19 Suppl 1:154-162) or Markov random fields.
[0121] "Conservative substitutions" may be made, for example, on the basis of similarity in polarity, charge, size, solubility, hydrophobicity, hydrophilicity, and / or the amphipathic nature of the amino acid residues involved. The 20 naturally occurring amino acids can be divided into six standard amino acid groups: (1) Hydrophobic: Met, Ala, Val, Leu, Ile; (2) Neutral hydrophilic: Cys, Ser, Thr, Asn, Gln; (3) Acidic: Asp, Glu; (4) basic: His, Lys, Arg; (5) residues that influence chain orientation: Gly, Pro; and (6) Aromatic: Trp, Tyr, Phe.
[0122] As used herein, a "conservative substitution" is defined as the exchange of an amino acid from one of the six standard amino acid groups listed above with another amino acid listed within the same group. For example, an exchange of Asp with Glu maintains a single negative charge in the polypeptide in which such an exchange is made. Furthermore, glycine and proline may be substituted for each other based on their ability to disrupt α-helices. Some preferred conservative substitutions within the six groups listed above are exchanges within the following subgroups: (i) Ala, Val, Leu, and Ile; (ii) Ser and Thr; (iii) Asn and Gln; (iv) Lys and Arg; and (v) Tyr and Phe.
[0123] As used herein, a "non-conservative substitution" is defined as replacing an amino acid with another amino acid listed in a different group of the six standard amino acid groups (1) to (6) described above.
[0124] Modifications of the enzymes described herein can include conservative and / or non-conservative mutations, in some embodiments, substituting or inserting an alanine at position 2 to enhance stability.
[0125] In some embodiments, "rational design" involves the construction of specific mutations in an enzyme. Rational design refers to incorporating knowledge of the enzyme or related enzymes, such as its reaction thermodynamics and kinetics, its three-dimensional structure, its active site(s), substrate(s), and / or enzyme-substrate interactions, into the design of specific mutations. Mutations can be created in an enzyme based on rational design approaches, and the enzymes can then be screened for increased terpene or terpenoid production compared to control levels. In some embodiments, mutations can be rationally designed based on homology modeling. As used herein, "homology modeling" refers to the process of constructing an atomic-resolution model of a protein from its amino acid sequence and the three-dimensional structures of related homologous proteins.
[0126] In another aspect, the present invention provides a method for producing a product comprising mogrol glycoside. The method includes producing mogrol glycoside according to the present disclosure and incorporating the mogrol glycoside into the product. In some embodiments, the mogrol glycoside is siamenoside, Mog.V, Mog.VI, or Isomog.V. In some embodiments, the product is a sweetener composition, a flavor composition, a food, a beverage, a chewing gum, a binder, a pharmaceutical composition, a tobacco product, a dietary supplement composition, or an oral hygiene composition.
[0127] The product may be a sweetener composition containing a blend of artificial and / or natural sweeteners. For example, the composition may further contain one or more of steviol glycosides, aspartame, and neotame. Exemplary steviol glycosides include one or more of RebM, RebB, RebD, RebA, RebE, and RebI.
[0128] Examples of flavors that can be used with the products include, but are not limited to, lime, lemon, orange, fruit, banana, grape, pear, pineapple, mango, almond, cola, cinnamon, sugar, cotton candy, and vanilla flavoring. Examples of other food ingredients include, but are not limited to, flavors, acidulants, and amino acids, colorants, bulking agents, modified starches, gums, binding agents, preservatives, antioxidants, emulsifiers, stabilizers, thickeners, and gelling agents.
[0129] The mogrol glycoside obtained according to the present invention can be incorporated as a high-intensity natural sweetener into foods, beverages, pharmaceutical compositions, cosmetics, chewing gum, tableware, cereals, dairy products, toothpaste, and other oral compositions.
[0130] The mogrol glycoside obtained according to the present invention can be used in combination with various physiologically active substances or functional ingredients. Functional ingredients are generally classified into categories such as carotenoids, dietary fiber, fatty acids, saponins, antioxidants, dietary supplements, flavonoids, isothiocyanates, phenols, plant sterols and stanols (phytosterols and phytostanols), polyols, prebiotics, probiotics, phytoestrogens, soy proteins, sulfides / thiols, amino acids, proteins, vitamins, and minerals. Functional ingredients can also be classified based on their health benefits, such as cardiovascular, cholesterol-lowering, and anti-inflammatory effects.
[0131] The mogrol glycosides obtained according to the present invention can be used as high-intensity sweeteners to produce zero-calorie, low-calorie, or diabetic beverages and foods with improved taste characteristics. They can also be used in beverages, foods, pharmaceuticals, and other products where sugar is not available. Furthermore, highly purified target mogrol glycoside(s), particularly Mog.V, Mog.VI, or Isomog.V, can be used as sweeteners in beverages, foods, and other products specifically for human consumption, as well as in animal diets and feeds with improved properties.
[0132] Examples of products in which mogrol glycoside(s) may be used as sweeteners include alcoholic beverages such as vodka, wine, beer, distilled spirits, and sake; natural fruit juices; soft drinks; carbonated soft drinks; diet drinks; zero-calorie drinks; low-calorie drinks and low-calorie foods; yogurt drinks; instant juices; instant coffee; powdered instant drinks; canned goods; syrups; fermented miso; soy sauce; vinegar; dressings; mayonnaise; ketchup; curry; soups; instant bouillon; powdered soy sauce; powdered vinegar; and biscuits. ; rice crackers; crackers; bread; chocolate; caramel; candy; chewing gum; jelly; pudding; preserved fruit and preserved vegetables; fresh cream; jam; marmalade; flower paste; powdered milk; ice cream; sorbet; bottled vegetables and fruit; canned and boiled pulses; meat and food stews in sweet sauces; agricultural vegetable foods; seafood; ham; sausage; fish ham; fish sausage; kamaboko; fried fish products; dried seafood products; frozen foods; preserved seaweed; preserved meat; tobacco; medicines; and many others.
[0133] Conventional methods such as mixing, kneading, dissolving, pickling, permeation, percolation, sprinkling, spraying, injecting, and other methods may be used in the manufacturing of products such as food, beverages, pharmaceuticals, cosmetics, tableware, and chewing gum.
[0134] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. For example, reference to "a cell" includes a combination of two or more cells, and the like.
[0135] As used herein, the term "about" in connection with a number is generally intended to include values within a range of 10% in either direction (larger or smaller) of that number. [Example]
[0136] Mogroside biosynthesis in fruit involves sequential glycosylation of the aglycone mogrol to the final sweet product, mogroside V (Mog.V), which is approximately 250 times sweeter than sucrose (Kasai et al., Agric Biol Chem (1989)). Mogrosides have also been reported to have health benefits (Li et al., Chin J Nat Med (2014)).
[0137] Interest in mogrosides and monk fruit in general is growing due to a variety of factors, including the burgeoning demand for natural sweeteners, the difficulty in sourcing large quantities of the current lead natural sweetener, rebaudioside M (RebM), derived from the stevia plant; the superior taste characteristics of Mog.V compared to other natural and artificial sweeteners on the market; and the medicinal properties of the plant and fruit.
[0138] Purified Mog.V has been approved as a high-intensity sweetener in Japan (Jakinovich et al., Journal of Natural Products (1990)), and its extract has achieved GRAS status in the United States as a non-nutritive sweetener and flavor enhancer (GRAS 522). Extraction of mogrosides from the fruit has yielded products of varying purity, but most have an undesirable aftertaste. Additionally, the low yield of the plant and the special growing conditions of this plant limit the yield of mogrosides obtained from cultivated fruit. Mogrosides are present at approximately 1% in fresh fruit and approximately 4% in dried fruit. Mog.V is the major component, present at 0.5%–1.4% in dried fruit. Furthermore, due to the difficulty of purification, the purity of Mog.V is limited, and commercial products derived from plant extracts are standardized to approximately 50% Mog.V. A pure Mog.V product is desirable to avoid off-flavor, and Mog.V's good solubility allows for easy formulation into products. Therefore, it would be advantageous to produce sweet-tasting mogroside compounds, including but not limited to Mog.V, via biotechnology processes.
[0139] Figure 1 shows the chemical structures of Mog.V, Mog.VI, Isomog.V, and siamenomeside. Mog.V contains five glycosylations around the mogrol core, including glycosylations at the C3 and C24 hydroxyl groups, followed by 1-2, 1-4, and 1-6 glucosyl additions. These glycosylation reactions are catalyzed by uridine diphosphate-dependent glycosyltransferase enzymes (UGTs).
[0140] Figure 2 shows the in vivo Mog.V production pathway. The enzymatic conversions required for each step are shown, along with the type of enzyme required. The numbers in parentheses correspond to the chemical structures in Figure 3: (1) farnesyl pyrophosphate; (2) squalene; (3) 2,3-oxidosqualene; (4) 2,3;22,23-dioxidosqualene; (5) 24,25-epoxycucurbitadienol; (6) 24,25-dihydroxycucurbitadienol; (7) mogrol; (8) mogroside V; and (9) cucurbitadienol.
[0141] Microbial strains that produce high levels of methylerythritol 4-phosphate (MEP) pathway products in conjunction with heterologous expression of mogrol biosynthetic enzymes and UGT enzymes that direct glucosylation to Mog.V or other desired mogroside compounds can be used to produce mogrosides in a biosynthetic fermentation process, as illustrated in Figure 2. For example, bacteria such as E. coli produce isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP) from glucose, which are then converted to farnesyl diphosphate (FPP) (1) by recombinant farnesyl diphosphate synthase (FPPS). FPP is converted to squalene (2) by a condensation reaction catalyzed by squalene synthase (SQS). Squalene is then converted to 2,3-oxidosqualene (3) by an epoxidation reaction catalyzed by squalene epoxidase (SQE). This pathway can be further epoxidized to 22,23-dioxidosqualene (4), followed by triterpene cyclase-catalyzed cyclization to 24,25-epoxycucurbitadienol (5), and then the remaining epoxy group can be hydrated with epoxide hydrolase to 24,25-dihydroxycucurbitadienol (6). Further P450 oxidase-catalyzed hydroxylation produces mogrol (7).
[0142] Alternatively, this route can involve cyclization of (3) to give cucurbitadienol (9), followed by epoxidation to (5), or multiple hydroxylation of cucurbitadienol to 24,25-dihydroxycucurbitadienol (6), or mogrol (7).
[0143] Figure 4 illustrates the glucosylation pathway leading to Mog.V. Glucosylation of the C3 hydroxyl group generates Mog.IE, or glucosylation of the C24 hydroxyl group generates Mog.I-A1. Glucosylation of Mog.I-A1 at C3 or Mog.I-E1 at C24 generates Mog-IIE. Further 1-6 glucosylation of Mog.II-E at C3 generates Mog.III-A2. Further 1-6 glucosylation of Mog.IIE at C24 generates Mog.III. 1-2 glucosylation of Mog.III-A2 at C24 generates Mog.IV, followed by further 1-6 glycosylation at C24 generates Mog.V. Alternatively, glucosylation can proceed by 1-6 glucosylation at C3 and 1-2 glucosylation at C24 via Mog.III, or 1-6 glucosylation via siamenoside or Mog.IV.
[0144] Biosynthetic enzymes from monk fruit (Siraitia grosvenorii) for mogrol production have been identified (see WO 2016 / 038617 and US 2015 / 0322473, the entire contents of which are incorporated herein by reference), but many of these enzymes lack the productivity or physical properties desirable for overexpression in microbial hosts, particularly for fermentation approaches to this plant that operate at higher temperatures than natural climates. Therefore, there is a need for alternative or engineered enzymes that can improve mogrol production using microbial fermentation and allow mogrol to serve as a substrate for glucosylation to produce Mog.V or other target mogrosides.
[0145] Using an E. coli strain that produces high levels of the MEP pathway products IPP and DMAPP (see US 2018 / 0245103 and US 2018 / 0216137, both incorporated by reference), ScFPPS was overexpressed to screen enzymes for their ability to convert FPP to squalene (SQS activity) and epoxidize squalene to produce 2,3-oxidosqualene (SQE activity). The 2,3-oxidosqualene intermediate can be cyclized by a triterpene cyclase, such as the CDS from Siraitia grosvenorii. As demonstrated in Figure 5, several enzymes were confirmed to have good activity in E. coli. In particular, SEQ ID NO: 11 showed high activity in E. coli at 37°C.
[0146] As shown in Figure 6, co-expression of SQS (SEQ ID NO: 11) and SQE (SEQ ID NO: 39) in E. coli substantially increased the titer of the 2,3-oxidosqualene intermediate. Other SQE enzymes showed activity in E. coli.
[0147] Figure 7 shows the coexpression of SQS, SQE, and TTC enzymes. Coexpression of CDS (or triterpene cyclase, or "TTC") (SEQ ID NO: 40) with SQS (SEQ ID NO: 11) and SQE (SEQ ID NO: 39) resulted in the production of high amounts of the triterpenoid product cucurbitadienol (product 3). These fermentation experiments were carried out at 37°C for 48 to 120 hours. Figure 8 shows the results of genetically engineering SQE to produce high titers of 2,3;22,23-dioxidosqualene. Expression of SQS, SQE, and TTC, whether in a bacterial artificial chromosome (BAC) or integrated form, produced large amounts of cucurbitadienol. Point mutations in SQE (SEQ ID NO: 39) were screened to eliminate SQE (SEQ ID NO: 39), reducing cucurbitadienol levels and correspondingly increasing 2, The titer of 2,3,22,23-dioxidosqualene was increased. Two SQE mutants, SQE A4 and SQE C11, are shown in Figure 8. Supplementing SQE (SEQ ID NO: 39) with a second engineered version with high specificity / activity for 2,3-oxidosqualene can boost the titer of 2,3,22,23-dioxidosqualene, as opposed to cucurbitadienol. This concept is further demonstrated in Figure 9. SQE A4 (SEQ ID NO: 203) was coexpressed with SQE (SEQ ID NO: 39), SQS (SEQ ID NO: 11), and TTC (SEQ ID NO: 40). These fermentation experiments were performed in 96-well plates at 37°C for 48 hours. Titers were plotted for each strain producing 2,3,22,23-dioxidosqualene. As shown in Figure 9, the strain expressing SQE A4 (SEQ ID NO: 203) produced significantly more 2,3,22,23 dioxidosqualene.
[0148] Figure 10 shows the coexpression of SQS, SQE, and TTC enzymes. Coexpression of TTC (SEQ ID NO: 40) with SQS (SEQ ID NO: 11), SQE (SEQ ID NO: 39), and SQE A4 (SEQ ID NO: 203) in E. coli produced cucurbitadienol and 24,25-epoxycucurbitadienol. Candidate enzymes for additional or alternative TTCs include SEQ ID NO: 40, SEQ ID NO: 191, SEQ ID NO: 192, and SEQ ID NO: 193. Each candidate TTC enzyme was expressed in this strain and screened for the production of 24,25-epoxy-cucurbitadienol. These fermentation experiments were performed in 96-well plates at 30°C for 72 hours. The production of 24,25-epoxy-cucurbitadienol was verified by GC-MS spectroscopy. Concentrations were plotted against production of 24,25-epoxy-cucurbitadienol from an E. coli strain expressing SEQ ID NO: 40 as the sole cyclase. As shown in Figure 10, E. coli strains co-expressing SQS (SEQ ID NO: 11), SQE (SEQ ID NO: 39), SQE A4 (SEQ ID NO: 203), and TTC (SEQ ID NO: 40) with additional TTC produced high levels of 24,25-epoxy cucurbitadienol.
[0149] Figure 11 shows the substrate specificity of candidate TTC enzymes for the production of cucurbitadienol and 24,25-epoxycucurbitadienol. Engineered E. coli strains producing oxidosqualene and dioxidosqualene were supplemented with engineered CDS homologs and CAS genes for cucurbitadienol production. Both strains were incubated at 30°C for 72 hours before extraction. The ratio of 24,25-epoxycucurbitadienol to cucurbitadienol varied from 0.15 for Enzyme 1 (SEQ ID NO: 40) to 0.58 for Enzyme 2 (SEQ ID NO: 192), indicating improved substrate specificity for Enzyme 2 toward the desired 24,25-epoxycucurbitadienol product.
[0150] Figure 12 shows the screening of EPH enzymes that hydrate epoxycucurbitadienol to produce 24,25-dihydroxycucurbitadienol in E. coli strains coexpressing SQS (SEQ ID NO:11), SQE (SEQ ID NO:39), SQE A4 (SEQ ID NO:203), and TTC (SEQ ID NO:40). EPH homologs were expressed in the 24,25-epoxycucurbitadienol-producing strains involved in the production of 24,25-dihydroxycucurbitadienol. Candidate EPH enzymes for this reaction include SEQ ID NO:184, SEQ ID NO:185, SEQ ID NO:186, SEQ ID NO:212, SEQ ID NO:187, SEQ ID NO:188, SEQ ID NO:189, SEQ ID NO:189, and SEQ ID NO:190. These fermentation experiments were performed in 96-well plates at 30°C for 72 hours. The production of 24,25-dihydroxycucurbitadienol was verified by GC-MS spectroscopy. The titers for each strain producing 24,25-dihydroxycucurbitadienol were plotted. As shown in Figure 12, E. coli strains expressing EPH were able to produce 24,25-dihydroxycucurbitadienol. In particular, ToEPEI and SgEPEB were able to produce 24,25-dihydroxycucurbitadienol in E. coli. showed high activity.
[0151] Figures 13A-C show that coexpression of SQS, SQE, TTC, EPH, and P450 enzymes produces mogrol. E. coli strains were constructed to express SQS (SEQ ID NO: 11), SQE (SEQ ID NO: 39), SQE A4 (SEQ ID NO: 203), TTC (SEQ ID NO: 40), EPH (SEQ ID NO: 58), and P450s selected from SEQ ID NO: 194, SEQ ID NO: 197, and SEQ ID NO: 171, along with cytochrome P450 reductase (SEQ ID NO: 98 or SEQ ID NO: 201). These fermentation experiments were performed in 96-well plates at 30°C for 72 hours. Mogrol production was verified by LC-QQQ spectroscopy. As shown in Figure 13A, expression of P450s selected from SQS (SEQ ID NO: 11), SQE (SEQ ID NO: 39), SQE A4 (SEQ ID NO: 203), TTC (SEQ ID NO: 40), EPH (SEQ ID NO: 58), and SEQ ID NO: 194, SEQ ID NO: 197, and SEQ ID NO: 171 led to the production of mogrol and oxomogrol. As shown in Figures 13B and 13C, the production of mogrol was verified by LC-QQQ mass spectrometry using spiked authentic standards (Figure 13B) and GC-FID chromatography versus authentic standards (Figure 13C), respectively.
[0152] Figure 14 shows the screening of cytochrome P450s for the oxidation of the 24,25-dihydroxycucurbitadienol-like molecule cucurbitadienol at C11. In most cases, replacing the native transmembrane domain with that from E. coli sohB (SEQ ID NO: 195, SEQ ID NO: 198, and SEQ ID NO: 199), E. coli zipA (SEQ ID NO: 196), or bovine 17α (e.g., SEQ ID NO: 200) improved interaction with the E. coli membrane. Coexpression of each P450 with either SEQ ID NO: 201 or SEQ ID NO: 98 resulted in the production of 11-hydroxycucurbitadienol. These fermentation experiments were performed in 96-well plates at 30°C for 72 hours. The production of 11-hydroxycucurbitadienol was verified by GC-MS. The concentrations of 11-hydroxycucurbitadienol-producing strains were plotted. As shown in Figures 14 and 15, the strains disclosed herein were able to produce 11-hydroxy-cucurbitadienol.
[0153] Mogrol was used as a substrate in an in vitro glucosylation reaction with candidate UGT enzymes to identify candidate enzymes that efficiently glucosylated mogrol to Mog.V. The reaction was carried out in 50 mM Tris-HCl buffer (pH 7.0) containing beta-mercaptoethanol (5 mM), magnesium chloride (400 μM), substrate (200 μM), UDP-glucose (5 mM), and phosphatase (1 U). The results are shown in Figure 16A. Mog.V product was observed when UGT enzymes of SEQ ID NO:165, SEQ ID NO:146, and SEQ ID NO:117 were incubated together. When incubated with UGT enzymes of SEQ ID NO:165, SEQ ID NO:146, and SEQ ID NO:164, a penta-glycosylated product was formed. Figure 16B shows an extracted ion chromatogram (EIC) of 1285.4 Da (mogroside V+H) obtained by incubating reactions containing the enzymes SEQ ID NO:165 + SEQ ID NO:146 and either the enzymes SEQ ID NO:117 (dark gray solid line) or SEQ ID NO:164 (light gray line) with Mog.II-E. Figure 16C shows an extracted ion chromatogram (EIC) of 1285.4 Da (mogroside V+H) obtained by incubating reactions containing the enzymes SEQ ID NO:165 + SEQ ID NO:146 and either the enzymes SEQ ID NO:117 (dark gray solid line) or SEQ ID NO:164 (light gray line) with mogrol.
[0154] Figures 4 and 17 show additional glycosyltransferase activity observed with specific substrates. Co-expression of UGT enzymes can be selected to shift product toward the desired mogroside product.
[0155] Figure 18 shows the bioconversion of mogrol to a mogroside intermediate. Genetically engineered E. coli strains expressing UGT enzymes (see US 2020 / 0087692, incorporated herein by reference in their entirety) were incubated in 96-well plates containing 0.2 mM mogrol. Product formation was determined after 48 hours. Reported values are above empty vector controls. Products were measured by LC-MS / MS using authentic standards. Only Enzyme 1 shows the formation of Mog.IIE. Enzymes 1-5 are SEQ ID NOs: 202, 116, 216, 217, and 218, respectively.
[0156] Figures 19A and 19B show the bioconversion of Mog.IA (Figure 19A) or Mog.IE (Figure 19B) to Mog.IIE. In the experiment, engineered E. coli strains (described above) expressing the UGT enzymes SEQ ID NO:165, SEQ ID NO:202, or SEQ ID NO:116 were incubated in fermentation medium containing 0.2 mM Mog.IA (Figure 19A) or Mog.IE (Figure 19B) at 37°C in a 96-well plate. Product formation was examined after 48 hours. Products were measured by LC-MS / MS with authentic standards. Mog.IIE levels above the empty vector control were calculated. As shown in Figure 19A, SEQ ID NO:165 and SEQ ID NO:202 were able to catalyze the bioconversion of Mog.IA to Mog.IIE. Similarly, as shown in Figure 19B, SEQ ID NO: 165, SEQ ID NO: 202, and SEQ ID NO: 116 were able to catalyze the bioconversion of Mog.IE to Mog.IIE.
[0157] Figure 20 shows the production of Mog.III or siamenoside from Mog.IIE. In the experiment, engineered E. coli strains expressing the UGT enzymes SEQ ID NO:204, SEQ ID NO:138, or SEQ ID NO:206 were grown in fermentation medium containing 0.1 mM Mog.IIE at 37°C for 48 hours. Products were quantified by LCMS / MS with authentic standards for each compound. As shown in Figure 20, all strains were able to catalyze the bioconversion of Mog.IIE to Mog.III. In addition, MbUGT1,2.2 also showed significant production of siamenoside.
[0158] Figure 21 shows the production of Mog.II-A2. 0.1 mM Mog.IE was supplied in vitro. In the experiment, a genetically engineered E. coli strain expressing the UGT enzyme SEQ ID NO:205 was incubated at 37°C for 48 hours. The product was quantified by LC-MS / MS using authentic standards for each compound. As shown in Figure 21, SEQ ID NO:205 can catalyze the bioconversion of Mog.IE to Mog.II-A2.
[0159] The primary glycosylation reactions of mogrol observed at the C3 and C24 hydroxyl groups are summarized in Table 1. Specifically, 0.2 mM mogrol was fed to cells expressing various UGT enzymes. The reactions were incubated at 37°C for 48 hours. Products were quantified by LCMS / MS using authentic standards for each compound. [Table 1]
[0160] The branched glycosylation reactions are summarized in Table 2. 0.2 mM Mog.IIE or Mog.IE was fed to cells expressing various UGT enzymes. The reactions were incubated at 37°C for 48 hours. Products were quantified by LC-MS / MS with authentic standards for each compound. "Indirect" evidence indicates that the substrate was consumed. [Table 2]
[0161] The following enzymes were expressed in an E. coli strain engineered to produce high levels of MEP pathway products: SQS (SEQ ID NO:11), SQE (SEQ ID NO:39), SQE A4 (SEQ ID NO:203), TTC (SEQ ID NO:40), EPH (SEQ ID NO:189), sohB_CppCYP (SEQ ID NO:199), AtUGT73C3 (SEQ ID NO:202), UGT85C1 (SEQ ID NO:165), and UGT94-289-1 (SEQ ID NO:122) to create an exemplary E. coli strain producing Mog.V. Production of Mog.V is demonstrated in Figures 22A and 22B. The strain was incubated at 30°C for 72 hours before extraction. As shown in Figure 22A, production of Mog.V was verified by LC-QQQ spectral analysis versus authentic standards. FIG. 22B is a chromatogram showing the production of Mog.V from biological samples using an authentic standard of spiked Mog.V.
[0162] Using the known structure and primary sequence, the biosynthetic enzymes can be further engineered for expression and activity in microbial cells.
[0163] Figure 26 shows an amino acid alignment of CaUGT_1,6 and SgUGT94_289_3 using Clustal Omega (version CLUSTAL O(1,2,4)). These sequences share 54% amino acid identity. CaUGT_1,6 is predicted to be a beta-D-glucosylcrocetin beta-1,6-glucosyltransferase-like enzyme (XP_027096357.1). With the known UGT structure and primary sequence, CaUGT_1,6 can be further engineered, including circular permutation engineering, for microbial expression and activity.
[0164] Figure 27 shows an amino acid alignment of Homo sapiens squalene synthase (HsSQS) (NCBI accession number NP_004453.3) with AaSQS (SEQ ID NO: 11) using Clustal Omega (version CLUSTAL O(1.2.4)). The crystal structure of HsSQS has been published (PDB entry: 1EZF). These sequences share 42% amino acid identity.
[0165] Figure 28 shows an amino acid alignment of Homo sapiens squalene epoxidase (HsSQE) (NCBI accession number XP_011515548) with MlSQE (SEQ ID NO: 39) using Clustal Omega (version CLUSTAL O(1.2.4)). The crystal structure of HsSQE has been published (PDB entry: 6C6N). These sequences share 35% amino acid identity.
[0166] The UGT enzyme of SEQ ID NO: 164 was engineered for improved glycosylation activity. Various amino acid substitutions were made to the enzyme, as revealed by in silico analysis. The amino acid substitutions in Table 3 below were tested for further glycosylation of mog.IIE. [Table 3-1] [Table 3-2]
[0167] An engineered UGT enzyme based on SEQ ID NO: 164 was prepared with substitutions T147L and N207K. The bioconversion of Mog.IIE to further glycosylated products is shown in Figure 23. In the experiment, an engineered E. coli strain expressing engineered CaUGT_1,6 was inoculated with Mog.IIE substrate at 37°C. Product formation was examined after 48 hours. Products were measured by LC / MS-QQQ using authentic standards.
[0168] The UGT enzyme of SEQ ID NO: 165 was engineered to improve glycosylation activity. The following amino acid substitutions were identified as improving the bioconversion of Mog.IA to Mog.IIE (Table 4): [Table 4]
[0169] Engineered UGT enzymes based on 85C1 were prepared with substitutions L41F, D49E, and C127F. The bioconversion of Mog.IA to Mog.IIE is shown in Figure 24. In the experiment, an engineered E. coli strain expressing engineered 85C1 was inoculated with Mog.IA substrate at 37°C. Product formation was examined after 48 hours. Products were measured by LC / MS-QQQ using authentic standards. Figure 24 shows the fold improvement of the engineered versions compared to the control (85C1).
[0170] The UGT enzyme of SEQ ID NO: 217 (UGT73F24) was engineered to improve glycosylation activity. The following amino acid substitutions were identified that improve the bioconversion of Mog.IE to Mog.IIE using UGT73F24 (Table 5): [Table 5]
[0171] Engineered UGT enzymes based on UGT73F24 were prepared with substitutions A74E, I91F, and H101P. The bioconversion of Mog.IE to Mog.IIE is shown in Figure 25. In the experiment, an engineered E. coli strain expressing engineered UGT73F24 was inoculated with Mog.IE substrate at 37°C. Product formation was examined after 48 hours. Products were measured by LC / MS-QQQ with authentic standards. Figure 25 shows the fold improvement of the engineered version compared to the control (73F24). [Sequence table]
[0172] Farnesyl phosphate synthase (FPPS) Saccharomyces cerevisiae FPPS MASEKEIRRERFLNVFPKLVEELNASLLAYGMPKEACDWYAHSNLYNTPGGKLNRGLSVVDTYAILSNKTVEQLGQEEYEKVAIlgWCIELLQAYFLVADDMMDKSITRRGQPCWYKVPEVGEIAINDAFMLEAAIYKLKSHFRNEKYYIDITELFHEVTFQTELGQLMDLITAP EDKVDLSKFSLKKHSFIVTFKTAYYSFYLPVALAMYVAGITDEKDLKQARDVLIPLGEYFQIQDDYLDCFGTPEQIGKIGTDIQDNKCSWVINKALELASAEQRKTLDENYGKKDSVAEAKCKFNDLKIEQLYHEYEESIAKDLKAKISQVDESRGFKADVLTAFLNKVYKRSK SQS Siraitia grosvenorii SQSa MGSLGAILRHPDDFYPLLKLKMAARHAEKQIPPEPHWGFCYTMHLKVSRSFALVIQQLAPELRNAICIFYLVLRALLDTVEDDTSIQTDIKVPILKAFHCHIYNRDWHFSCGTKDYKVLMDQFHHVSTAFLELGKGYQEAIEDITKRMGAGMAKFICKEVETVDDYDEYCHYVAGLVGLGLSKLFHASLEDLAPDSLNSMGLLLQKT NIIRDYLEDINEIPKSRMFWPREIWGKYADKLEDFKYEEENSVKAVQCLNDLVTNALNHVEDCLKYMSNLRDLSIFRFCAIPQIMAIGTLALCYNNVEVFRGVVKMRRGLTAKVIDRTQTMADVYGAFFDFSVMLKAKVNSSDPNATKTLSRIEAIQKTCEQSSGLLNKRKLYAVKSEPMFNPTLIVLFSLCIILAYLSAKRLPNQPV Siraitia grosvenorii SQSb MGSLGAILRHPDDFYPLLKLKMAARHAEKQIPPEPHWGFCYTMLHKVSRSFALVIQQLAPELRNAICIFYLVLRALDTVEDTSIQTDIKVPILKAFHCHIYNRDWHFSCGTKDYKVLMDQFHHVSTAFLELGKGYQEAIEDITKRMGAGMAKFICKEVETVDDYDEYCHYVAGLVGLGLSKLFHASDLEDLAPDSLNSMGLLLQKT NIIRDYLEDINEIPKSRMFWPREIWGKYADKLEDFKYENSVKAVQCLNDLVTNALNHVEDCLKYMSNLRDLSIFRFCAIPQIMAIGTLALCYNNVEVFRGVVKMRRGLTAKVIDRTQTMADVYGAFFDFSVMLKAKVNNSDPNATKTLSRIEAIQKTCEQSGLLNKRKLYAVKSEPMFNPTLIVILFSLLCIILAYLSAKRLPNQPV Cucumis sativus MGSLGAILKHPDDFYPLLKLKIAARHAEKQIPPEPHWGFCYTMLHKVSRSFALVIQQLKPELRNAVCIFYLVLRALDTVEDTSIQTDIKVPILKAFHCHIYNRDWHFSCGTKDYKVLMDEHFHHVSTAFLELGKGYQEAIEDITKRMGAGMAKFICKEVETVDDYDEYCHYVAGLVGLGLSKLFHAAELEDLAPDSLSNSMGLFLQKT NIIRDYLEDINEIPKSRMFWPREIWGKYADKLEDFKYENSVKAVQCLNDLVTNALNHVEDCLKYMSNLRDLSIFRFCAIPQIMAIGTLALCYNNVEVFRGVVKMRRGLTAKVIDRTKTMADVYGAFFDFSVMLKAKVNSDPNASKTLSRIEAIQKTCKQSGILNRRKLYVVRSEPMFNPAVIVILFSLCCIILAYLSAKRLPAQSV Cucumis melo MGSLGAILKHPDDFYPLLKLKMAARHAEKQIPPESHWGFC YTMLHKVSRSFALVIQQLKPELRNAVCIFYLVLRALDTVEDDTSIQTDIKVPILKAFHCHIYNRDWHFSCGTKDYKVLMDEFHHVSTAFLELGKGYQEAIEDITKRMGAGMAKFICKEVETVDDYDEYCHYVAGLVGLGLSKLFHAAELEDLAPDSLSNSMGLFLQKTNIIRDYLEDINEIPKSRMFW PREIWGKYADKLEDFKYEENSVKAVQCLNDLVTNALNHVEDCLKYMSNLRDLSIFRFCAIPQIMAIGTLALCYNNVEVFRGVVKMRRGLTAKVIDRTKTMADVYGAFFDFSVMLKAKVNSNDPNASKTLSRIEAIQQTCQQSGLMNKRKLYVVRSEPMYNPAVILFSLCCIILAYLSAKRLPAQSV Cucumis melo MGSLGAILKHPDDFYPLLKLKMAARHAEKQIPPESHWGFCYTMLHKVSRSFALVIQQLKPELRNAVCIFYLVLRALDTVEDDTSIQTDIKVPILKAFHCHIYNRDWHFSCGTKDYKVLMDEFHHVSTAFLELGKGYQEAIEDITIKRMGAGMAKFICKEVETVDDYDEYCHYVAGLVGLGLSKLFHAAELEDLAPDSLSNSMGLFLQKTNIIRDYLEDINEIPKSRMWFPREIWGKYADKLEDFKYEENSVKAVQCLNDLVTNALNHVEDCPKYMSNLRDLSIFRFCAIPQIMAIGTLALCYNNVEVFRGVVKMRRGLTLAKIVDRTKTMADVYGAFFDFSVMLKAKVNSNDPNASKTLSRIEAIQQTCQQSGLMNKRKLYVVRSEPMYNPAVIVILFSLCILAYLSAKRLPANQSV Cucurbita moschata MGSLGAILRHPDDIYPLLKLKMAARHAEKQIPPESHWGFCYTMHLKVSRSFALVIQQLKPELRNAVCIFYLVLRALDTVEDTSIQTDIKVPILKAFHCHIYNRDWHFSCGTKDYKVLMDEHFHHVSTAFLELGRGYQEAIEDITKRMGAGMAKFICKEVETVEDYDEYCHYVAGLVGLGLSKLFHASKSENLAPDSLSNSMGLFLQKT NIIRDYLEDINEIPKSRMFWPREIWSKYADKLEDFKYEKNSVKAVQCLNDLVTNALTHVEDCLEYMSNLKDLSIFRFCAIPQIMAIGTLALCYNNVDFRGVVKMRRGLTAKVIYRTKTMDAVYGAFFDFSVMLKAKVNSSDPNASKTLTRIEAIQKTCKQSGLLNKRELYAVRSEPMCNPAAIVVLFSLLCILAYLSAKLLPANQPV Sechium edule MGSLGAILSHPDDLYPLLKLKMAAKHAEKQIPPDPHWGFCFSMLHKVSRSFALVIQQLKPELRNAVCIFYLVLRALDTVEDDTGIHPDIKVPILQAFHCHIYNRDWHFSCGTKHYKVLMDEFHHVSTAFLELGKGYQEAIEDVTERMGAGMAKFICKEVETVDDYDEYCHYVAGLVGLGSKLFHAAELEDLAPDSLSNSMGLFLQKTNIIRDYLEDINEIPKSRMWFPREIWNKYADKLEDFKYEENSVKAVQCLNDLVTNALNHVEDCLKYMSNLKDLSTFRFCAIPQIMAIGTLALCYDNVEVFRGVVKMRRGLTLTAKIIDRTKKIADVYGAFFDFSVMLKAVNSSDPNAAKTLSRIEAIEKTCKESGLLNKRKLYVIRSEPLFNPAVLVILFSLICILLAYLSAKRLPANQPV Panax quinquefolius MGSLGAILKHPDDFYPLLKLKFAARHAEKQIPPEPHWAFCYSMLHKVSRSFGLVIQQLGPQLRDAVCIFYLVLRALDTVEDTSIPTEVKVPILMAFHRHIYDKDWHFSCGTKEYKVLMD EFHHVSNAFLELGSGYQEAIEDITMRMGAGMAKFICKEVETIDDYDEYCHYVAGLVGLGLSKLFHASGAEDLATDSLSNSMGLFLQKTNIIRDYLEDINEIPKSRMFWPRQIWSKYVDKLEDLKYEENSAKAVQCLNDMVTDALVHA EDCLKYMSDLRDPAIFRFCAIPQIMAIGTLALCFNNTQVFRGVVKMRRGLTAKVIDRTKTMSDVYGAFFDFSCLLSKVDNDPNATKTLSRLEAIQKTCKESGTLSKRKSYIIESESGHNSALIAIIIFIILAILYAYLSSSNLLLNKQ Malus domestica MGALSTMLKHPDDIYPLLKLKIASRQIEKQIPAEPHWAFCYTMLQKVSRSFALVIQQLGTELRNAVCLFYLVLRALLDTVEDDTSVADTVKVPILLAFHRHIYDPDWHFACGTNNYKVLMDEFHHVSTAFLELGTGYQEAIEDITKRMGAGMAKFILKEVETIDDEYDEYCHYVAGLVGLGLSKLFHAAGKEDLASDSLSNSMGLFLQ KTNIIRDYLEDINEIPKSRMFWPRQIWSKYVNKLEDLKYEENSEKAVQCLNDMVTNALIHMEDCLKYMAALRDPAIFKFCAIPQIMAIGTLALCYNNIEVFRGVVKMRRGLTAKVIDRTKSMDDVYGAFFDFSSILKSKVDKNDPNATKTLSRVEAVQKLCRDSGALSKRKSYIANREQSYNSTLIVALFIILAIIYAYLSASPRI Artemisia annua MSSLKAVLKHPDDFYPLLKLKMAAKKAEKQIPSQPHWAFSYSMLHKVSRSFALVIQQLNPQLRDAVCIFYLVLRALDTVEDTSIAADIKVPILIAFHKHIYNRDWHFACGTKEYKVLMDQFHHVSTAFLELKRGYQEAIEDITMRMGAGMAKFICKEVETVDDYDEYCHYVAGLVGIGLSKLFHSSGTEILFSDSISNSMGLFLQKTN IIRDYLEDINEIPKSRMFWPREIWSKYVNKLEDLKYEENSEKAVQCLNDMVTNALHIEDCLKYMSQLKDPAIRFCAIPQIMAIGTLALCYNNIEVFRGVVKLRRGLTAKIDRTKTMADVYQAFSDFSDMLKSKVDMHDPNAQTTITRLEAAQKICKDSGTLSNRKSYIVKRESSYSAALLFTILAILYAYLSANRPNKIKFTL Glycine soya MDQRSEDEFYPLLKLKIVARNAEKQIPPEPHWAFCYTMLHKVSRSFALVIQQLGIELRNAVCIFYLVLRALDTVEDDTSIETDVKVPILIAFHRHIYDRDWHFSCGTKEYKVLMGQFHHVSTAFLELGKNYQEAIEDITKRMGAGMAKFICKEVETIDDYDEYCHYVAGLVGLGLSKLFHASGSEDLAPDDLSNSMGLFLQKTN IIRDYLEDINEIPKSRMFWPRQIWSEYVNKLEDLKYEENSVKAVQCLNDMVTNALMHAEDCLTYMAALRDPPIFRFCAIPQIMAIGTLALCYNNIEVFRGVVKMRRGLTAKVIDRTKTMADVYGAFFDFASMLEPKVDKNDPNATKTLSRLEAIQKTCRESGLLSKRKSYIVNDESGYGSTMIVILVIMVSIIFAYLSANHHNS Diospyros kaki MGSLAAMLRHPDVYPLVKLKMAARHAEKQIPPEPHWAFCYTMHLKVSRSFGLVIQQLGTELRNAVCIFYLVLRALDTVEDDTSIATEVKVPILLAFHHHIYDRDWHFSCGTREYKVLMDEFHHVSTAFLELGKGYQEAIEDITMRMGAGMAKFICKEVETIDDEYCHYVAGLVGLGLSKLFHASGLEDLAPDSLNSNS MGLFLQKTNIIRDYLEDINEIPKSRMFWPRQIWSKYVNKLEDLKYEKNSVKSVQCLNDMVTNALIHVDDCLKYMSALRDPAIFRFCAIPQIMAIGTLALCYNNIEVFRGVVKMRRGLTAKVIDQTKTISDVYGAFFDFSCMLKSKVEKNDPNSTKTLSRIEAIQKTCRESGTLSKRKSYILRSKRTHNSTLIFFLVFIILAILFAYLSANRPPINM Euphorbia lathyris MGSLGAILKHPDDFYPLLKLKMAAKHAEKQIPAQPHWGFCYSMLHKVSRSFSLVIQQLGTELRDAVCIFYLVLRALDTVEDDTSIPTDVKVPILIAFHKHIYDPEWHFSCGTKEYKVLMDQIHHLSTAFLELGKSYQEAIEDITKKMGAGMAKFICKEVETVDDYDEYCHYVAGLVGLGLSKLFDASGFEDLAPDDLSNSMGLFL QKTNIIRDYLEDINEIPKSRMFWPRQIWSKYVNKLEDLKYEENSVKAVQCLNDMVTNALIHMDDCLKYMSALRDPAIFRFCAIPQIMAIGTLALCYNNVEVRGVVKMRRGLTAKVIDRTRTMADVYRAFFDFSCMMKSKVDRNDPNAEKTLNRLEAVQKTCKESGLLNKRRSYINESKPYNSTMVILLMIVLAIILAYLSKRAN Camellia oleifera MGSLGAILKHPDDFYPLMKLKMAARRAEKNIPPEPHWGFCYSMLHKVSRSFALVIQQLDTELRNAVCIFYLVLRALDTVEDDTSIATEVKVPILMAFHRHIYDRDWHFSCGTKEYKVLMDEFHHVSTAFSELGRGYQEAIEDITMRMGAGMAKFICKEVETIDDDYDEYCHYVAGLVGLGLSKLFHASGSEDLASDSLSNSMGLFLQVFLTCIKTNIIRDYLEDINEIPKSRMFWPRQIWSKYVNKLEDLKDKENSVKAVECLNDMVTNALIHVEDCLTYMSALRDPSIFRFCAIPQIMAIGTLALCYNNIEVFRGVVKMRRGLTAKVIDRTKTMSDVYGGFFDFSCMLKSKVNKSDPNAMKALSRLEAIQKICRESGTLNKRKSYIIKSEPRYNSTLVFVLFIILAILFAYL Eleutherococcus senticosus MGSLGAILKHPDDFYPLLKLKFAARHAEKQIPPEPHWAFCYSMLHKVSRSFGLVIQQLDAQLRDAVCIFYLVLRALLDTVEDTSIPTEVKVPILMAFHRHIYDKDWHFSCGTKEYKVLMDEFHHVSNAFLELGSGFQEAIEDITMRMGAGMAKFICKEVETIDDYDEYCHYVAGLVGLGLSKLFHASGAEDLATDSLSNSMGLFLQK TNIIRDYLEDINEIPKSRMFWPRQIWSKYVDKLENLKYEENSAKAVQCLNDMVTNALLHAEDCLKYMSNLRDPAIFRFCAIPQIMAIGTLALCFNNIQVFRGVVKMRRGLTAKVIDRTKTMSDVYGAFFDFSCLLKSKVDNNDPNATKTLSRLEAIQKTCKESGTLSKRKSYIIESKSAHNSALIAIIIFIILAILYAYLSSNLPNNQ Flavobacteriales bacteria MLNNSLFSRLEEIPALLKLKLGSKDYYKNNNSETLTCDNLRYCFDTLNKVSRSFATVIKQLPNELGNNVCVFYLILRALDSIEDDMNLPKELKIKLLREFHKKNYESGWNISGVGDKKEHVELLENYDKVIQSFLAIDQKNQLIITDICRKVGAGMANFVKAEIESVEDYNLYCHHVAGLVGIGLSRMFISSGLENDDFLNQDEISNSMGLFLQKTNIVRDYREDLDEGRMFWPKDIWHVYGSKINDFAINPTHDQSVLCLNHMLNNALTHATDCLAYLK HLRNENIFKFCAIPQVMAMATLCKIYSNPDVFIKNVKIRKGLAAKLILNTTSMDEVIKVYKDMLLVIESKISSDNNPVSAETIQLLKQIREYFNDETLIVRKIA Bacteroidetes bacterium (SEQ ID NO: 167) MLNSSLFSRLEEIPALLKLKLGSINNYKNNNSENLTSKNLRYCFDTLNKVSRSFASVIKQLPNELMVNVCLFYLILRALDSIEDDMNLPKDFKINLLREFLDKNYEPGWKISGVGDKKEYVELLENYDKVIQVFLDIDPKNQLIITDICRKMGAGMAHFVEAEINSVKDYNLYCYHVAGLVGIGLSKMFLASGLENCDYLNQEEISSSMGLFLQKTNIVRDYKEDMEENRIFWPKEIWRTYASKFSDFSINPQHETSISCLNHMVNDALGHVIDCLEYLRHLRNENIFKFCAIPQVMAMATLCKVYNNPDVFIKTVKIRKGLAAKLILNTTSMDEVIKVYKGLLLDIENKIPLHNPTSDETLRLIKNIRSYCNNETMVVSKTA Squalene epoxidase Siraitia grosvenorii SQE1 (SEQ ID NO: 17) MVDQCALGWILASALGLVIALCFFVAPRRNHRGVDSKERDECVQSAATTKGECRFNDRDVDVIVVGAGVAGSALAHTLGKDGRRVHVIERDLTEPDRIVGELLQPGGYLKLIELGLQDCVEEIDAQRVYGYALFKDGKNTRLSYPLENFHSDVSGRSFHNGRFIQRMREKAASLPNVRLEQGTVTSLLEEKGTIKGVQYKSKNGEEKTAYAPLTIVCDGCFSNLRRSLCNPMVDVPSYFVGLVLENCELPFANHGHVILGDPSPILFYQISRTEIRCLVDVPGQKVPSIANGEMEKYLKTVVAPQVPPQIYDSFIAAIDKGNIRTMPNRSMPAAPHPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLLKPLKDLSDASTLCKYLESFYTLRKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGIFSNGPVSLLSGLNPRPLSLVLHFFAVAIYGVGRLLLPFPSVKGIWIGARLIYSASGIIFPIIRAEGVRQMFFPATVPAYYRSPPVFKPIV Siraitia grosvenorii SQE2 (Accession No. 18) MVDQCALGWILASVLGAAALYFLFGRKNGGVSNERRHESIKNIATTNGEYKSSNSDGDIIIVGAVAGSALAYTLGKDGRRVHVIERDLTEPDRIVGELLQPGGYLKLTELGLEDCVDDIDAQRVYGYALF KDGKDTRLSYPLEKFHSDVAGRSFHNGRFIQRMREKAASLPKVSLEQGTVTSLLEENGIIKGVQYKTKTGQEMTAYAPLTIVCDGCFSNLRRSLCNPKVDVPSCFVGLVLENCDLPYANHGHVILADPSPI LFYRISSTEIRCLVDVPGQKVPSINSGEMANYLNKNVVAPQIPSQLYDSFVAAIDKGNIRTMPNRSMPADPYPTPGALLMGDAFNMRHPLTGGGMTVALSDVVVLRDLLKPLRDLNDAPTLSKYLEAFYTLR KPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGIFSNGPVSLLSGLNPRISPILLVHFFAVAIYGVGRLLIPFPSPKRVWIGARIISGASAIIFPIIKAEGVRQMFFPATVAAYYRAPRVVKGR Momordica charantia MVDECALGWILAAALGAVIALCLFVAPKTNNQDGGVDSKATPECVQTTNGECRSDGDSDVIIVGAGVAGSALAHTLGKDGRRVHVIERDLTEPDRIVGELLQPGGYLKLIELGLADCVEEIDAQRVYGYALFKDGKNTRLSYPLEKFHSDVSGRSFHNGR FIQRMREKADSLPNVRLEQGTVTSLLEEKGTIKGVQYKSKDGKEKTAYAPLTIVCDGCFSNLRRSLCNPMVDVPSCFVGLVLENCQLPFANHGHVVLGDPSPILFYPISSTEIRCLVDVPGQKVPSISNGEMEKYLKTVVAPQVPPQIYDAFIAAIDKGNIRTMPNRSMPAAPHPTPGALLMG DAFNMRHPLTGGGMTVALSDIVVLRNLLKPLKDLHDAPTLCKYLESFYTLRKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGMFSNGPVSLLSGLNPRPLSLVLHFFAIYGVGRLLFPFPSPKGIWIGARLIYSASGIIFPIIKAEGVRQMFFPATVPAYYRSPPALKPVA Cucurbita maxima MVDYCAFGWILAAVLGLAIALSFFVSPRRNRRGGADSTPRSEGVRSSSTTNGECRSVDGDADVIIVGAGVAGSALAHTLGKDGRLVHVIERDLTEPDRIVGELLQPGGYLKLIELGLQDCVEEIDAQKVYGY ALFKDGKNTQLSYPLEKFQSDVSGRSFHNGRFIQRMREKAASLPNVRLEQGTVTSLLEEKGTIKGVQYKSKNGEEKTAYAPLTIVCDGCFSNLRRRSLCKPVMVDVPSCFVGLVLENCQLPFANHGHVVLGDPS PILFYPISSTEIRCLVDVPGQKIPSISNGEMEKYLKTIVAPQVPPQIHDAFIAIDKGNIRTMPNRSMPAAPQPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLNRLLKPLKDLNDAPTLCKYLESFYTL RKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGIFSNGPVSLLSGLNPRPLSLVLHFFAVAIYGVGRLLLPFPSPKGIWIGARLVYSASGIIFPIIKAEGVRQMFFPATVPAYYRSPPVHKSIA Cucurbita moschata MVDYCAFGWILAAVLGLAIALSFFVSPRRNRRGGADSTPRSEGVRSSSTTNGECRSVDCDADVIIVGAGVAGSALAHTLGKDGRLVHVIERDLTEPDRIVGELLQPGGYLKLIELGLQDCVEEIDAQKVYGY ALFKDGKNTQLSYPLEKFQSDVSGRSFHNGRFIQRMREKAASLPNVRLEQGTVTSLLEEKGTIKGVQYKSKNGEEKTAHPLTIVCDGCFSNLRRRSLCKPVMVDVPSCFVGLVLENCQLPFANHGHVVLGDPS PILFYPISSTEIRCLVDVPGQKVPSISNGEMEKYLKTIVAPQVPPQIHDAFIAIDKGNIRTMPNRSMPAAPQPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLLKPLKDLNDAPTLCKYLESFYTL RKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGIFSNGPVSLLSGLNPRPLSLVLHFFAVAIYGVGRLLLPFPSPKGIWIGARLVYSASGIIFPIIKAEGVRQMFFPATVPAYYRSPPVLKTIA Cucurbita moschata MMVDHCAFAWILDVVLGLVVAVTFFVAAPRRNRRRGGTDSTASKDCVISTAIANGECKPDADAEVIIVGAGVAGSALAYTLGKDGRRVHVIERDLTEPDRIVGEFLQPGGYLKLIELGLGDCVEEIDAQKLYGYALFKDGKNTRVSYPLGNFHSDVSGRSFHNGRFIQRMREKAASLPNV RLEQGTVTSLLETKGTIKGVQYKSKNGEEKTAYAPLTIVCDGCFSNLRRSRSLCKPMVDVPSCFVGLVLENCQLPFANHGHVVLGDPSPILFYPISSTEIRCLVDVPGQKVPSINSGDMEKYLKTVVAPQVPPQIHDAFIAAIEKGNVRTMPNRSMPAAPHPTPGALLMGDAFNMRHPLTGG GMTVALSDIVVLRNLKPLKDLNDASTLCKYLESFYTLRKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGVFSNGPISLLSGLNPRSSLVLHFVAIYGVGRLLLPFPSLKGIWIGARLIYSASGIILPIIKAEGVRQMFFPATVPAYYRSPPVHKPIT Cucumis sativus MVDHCTFGWIFSAFLAFVIAFSFFLSPRKNRRGRGTNSTPRRDCLSSATTNGECRSVDGDADVIIVGAGVAGSALAHTLGKDGRRVHVIERDLTEPDRIVGELLQPGGYLKLIELGLQDCVEEIDAQKVYGYALFKDGKSTRLSYPLENFQSDVSGRSFHNGRFIQRMREKAAFLPNVRLEQGTVTSLLEEKGTITGVQYKSKNGEQKTAYAPLTIVCDGCFSNLRRSLCNPMVDVPSCFVGLVLENCQLPYANLGHVVLGDPSPILFYPISSTEIRCLVDVPGQKVPSISNGEMEKYLKTVVAPQVPPQIHDAFIAAIEKGNIRTMPNRSMPAAPQPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLKPLKDLNDAPTLCKYLESFYTLRKPVASTINTLAGALYKVFCASSDQARKEMRQACFDYLSLGGIFSNGPVSLLSGLNPRPLSLVLHFFAVAIYGVGRLLLPFPSPKGIWIGARLVYSASGIIFPIIKAEGVRQMFFPATVPAYYRTPPVFNS Cucumis melo MVDHCAFGWIFSALLAFPIALSLFLSPWRNRRVRGTDSTPRSASVSSATTNGECRSVDGDADVVVIVGAGVAGSALAHTLGKDGRRVHVIERDLTEPDRIVGELLQPGGYLKLIELGLQDCVEEIDAQKVYGYALFKDGKNTRLSYPLENFHSDVSGRSFHNGRFIQRMREKAASLPNVRLEQGTVTSLLEEKGTITGVQYKSKNGEQKTAYAPLTIVCDGCFSNLRRSLCTPMVDVPSYFVGLVLENCQLPYANLGHVVLGDPSPILFYPISSTEIRCLVDVPGQKVPSISNGEMEKYLKTVVAPQVPPQIHDAFIAAIEKGNIRTMPNRSMPAAPQPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLKPLKDLNDAPTLCKYLESFYTLRKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGIFSNGPVSLLSGLNPRPLSLVLHFFAVAIYGVGRLLLPFPSLKGIWIGARLVYSASGIIFPIIKAEGVRQMFFPATVPAYYRTPPVLNS Cucurbita maxima MMVEHCAYGWILAAVLGLVVAVTFFVAVPRRNRRGGTDSTASKDCVISPAIANGECEPEDADADAVIIVGAGVAGSALAHTLGKDGDRVHVIERDLTEPDRIVGEFLQPGGHLKLIELGLGDCVEEIDAQKLYGYALFKDGKNTRVSYPLGNFHSDVSGRSFHNGRFIQRMREKAASLPNVRLEQGTVTSLLEKKGTIKGVQYKSKNGEEKTAYAPLTIVCDGCFSNLRRSLCKPMVDVPSCFVGLVLENCRLPFANHGHVVLGDPSPILFYPISSTEIRCLVDVPGQKVPSIPNGDMEKYLKTVVAPQVPPQIHDAFIAAIEKGNIRTMPNRSMPAAPHPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLKPLKDLNDAPTLCKYLESYYTLRKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGVFSNGPISLLSGLNPRSCLVLHFFAVAIYGVGRLLLPFPSLKGIWIGARLIYSASGIILPIIKAEGVRQMFFPATVPAYYRSPPVHKPIT Ziziphus jujube MLDQCPLGWILASVLGLFVLCNLIVKNRNSKASLEKRSECVKSIATTNGECRSKSDDVDVIIVGAGVAGSALAHTLGKDGRRLHVIERDLTEPDRIVGELLQPGGYLKLIELGLQDCVEEIDAQRVFGYALFKDGKDTRLSYPLEKFHSDVSGRSFHNGRFIQRMREKSASLPNVRLEQGTVTSLLEEKGTIKGVQYKTKTGQELTAFAPLTIVCDFGCFSNLRRSLCNPKVDVPSCFVGLVLENCELPYANHGHVILADPSPILFYPISSTEVRCLVDVPGQKVPSISNGEMAKYLKSVVAPQIIPPQIYDAFIAAVDKGNIRTMPNRSMAPSFPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRDLLKPLGDLNDAATLCKYLESFYTLRKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGIFSTGPVSLLSGLNPRPLSLVLHFFAVAIYGVGRLLLPFPSPKRIWIGARLISGASGIIFPIKAEGVRQMFFPATVPAYYRAAPVE Morus alba MADPYTMGWILASLLGLFALYYLFVNNKNHREASLQESGSECVKSVAPVKGECRSKNGDADVIIVGAGVAGSALAHTLGKDGRRVHVIERDLAEPDRIVGELLQPGGYLKLIELGLQDCVEEIDSQRWYGY ALFKDGKDTRLSYPLEKFHSDVSGRSFHNGRFIQRMREKAASLPNVQLEQGTVTSLLEENGTIKGVQYKTKTGQELTAYAPLTIVCDGCFSNLRRSLCIPKVDVPSCFVGLVLENCNLPYANHGHVVLADP SPILFYPISTSTEVRCLVDVPGQKVPSISNGEMAKYLKTVVASQIPPQIYDSFVAAVDKGNIRTMPNRSMPAAPHPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRDLLKPLRDLNDSVTLCKYLESFYTLRKPVASTINTLAGALYKVFCASPDQARKEMREACFDYLSLGGVFSEGPVSLLSGLNPRPLSLVCHFFAVAIYGVGRLLLPFPSPKRLWIGARLISGIIFPIIRAEGVRQMFFPAYRAPRPN Juglans regia (JrSQE1) MVDPYALGWSFASVLMGLVALYILVDKKNRSRVSSEARSEGVESVTTTTSGECRLTDGDADVIIVGAGVAGSALAHTLGKDGRRVHVIERDLTPDRIVGELLQPGGYLKLIELGLEDCVEDIDAQRVFGY ALFKDGKNTRLSYPLEKFHSDVSGRSFHNGRFIQRMREKAASLLNVRLEQGTVTSLLEENGTVKGVQYKTKDGNELTAHPLTIVCDGCFSNLRRSLCNPQVDVPSSFVGLVLENCELPYANHGHVILADPS PILFYPISSTEVRCLVDVPGKKVPSIANGEMEKYLKNMVAPQLPPEIYDSFVAAVDRGNIRTPMNRSMPAAPHPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRDLKPLRDLNDAPTLCKYLESFYTL RKPVASTINTLAGALYKVFCASPDRARKEMRQACFDYLSLGGVFSMGPVSLLSGLNPRPLSLVLHFFAVAVYGVRLLVPFPSPSRIWIGARLISGASAIIFPIIKAEGVRQMFFPATVPAYYRAPPVKRDH Cucumis melo MVDQCALGWILASVLGASALYLLFGKKNCGVLNERRRESLKNIATTNGECKSSSNSDGDIIIVGAGVAGSALAYTLAKDGRQVHVIERDLSEPDRIVGELLQPGGYLKLTELGLEDCVDDIDAQRVYGYALFKDGKDTRLSYPLEKFHSDVSGRSFHNGRF IQRMREKAASLPNVRLEQGTVTSLLEENGTIKGVQYKNKSGQEMTAYAPLTIVCDGCFSNLRRSLCNPKVDVPSCFVGLILENCDLPYANHGHVILADPPILFYPISSTEIRCLVDVPGQKVPSINSGEMANYLNKNVVAPQIPPQLYNSFIAAIDKGNIRTMPNRSMPADPYPTPGALLMG DAFNMRHPLTGGGMTVALSDIVVLRDLKPLRDLNDAPTLCKYLEAFYTLRKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGIFSNGPVSLLSGLNPRPLSLVLHFFAIYGVGRLLIPFPSPKRVWIGARLISGASAIIFPIIKAEGVRQMFFPKTVAAYYRAPPVVRER Cucumis sativus MVDQCALGWILASVLGASALYLLFGKKNCGVSNERRRESLKNIATTNGECKSSSNSDGDIIIVGAGVAGSALAYTLAKDGRQVHVIERDLSEPDRIVGELLQPGGYLKLTELGLEDCVDEIDAQRVYGYALF KDGKDTRLSYPLEKFHSDVSGRSFHNGRFIQRMREKAASLPNVRLEQGTVTSLLEENGTIRGVQYKNKSGQEMTAYAPLTIVCDGCFSNLRRSLCNPKVDVPSCFVGLILENCDLPHANHGHVILADPSPI LFYPISSTEIRCLVDVPGQKVPSINSGEMANYLNKNVVAPQIPPQLYNSFIAAIDKGNIRTMPNRSMPADPYPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRDLKPLRDLNDAPTLCKYLEAFYTLR KPVAST INPUT LYKVFCASPDQARKEMRQACFDYLSLGGIFSNGPVSLLSGLNPRPLSLVLHFAVAIYGVGRLLIPFPSPKRVWIGARLISGASAIIFPIIKAEGVRQMFFPKTVAAYYRAPPIVRER Juglans regia (JrSQE2) MVDQYAGLLILASVLGFVVLYNLMAKKNRIRVSSEARTEGVQTVITTTNGECRSIEGDVDVIIVGAGVAGSALAHTLGKDGRKVHVIERDLSEPDRIVGELLQPGYLKLVELGLQDSVEDIDAQRVFGYALFK DGKNTRLSYPLEKFHSDVSGRSFHNGRFIQRMREKAASLPNIRLEQGTVTSLLEENGTIKGVQYKTKDGKELAAHAPLTIVCDGCFSNLRRSLCNPQVDVPSSFVGLVLENCELPYANHGHVVLADPSPPILFYP ISSTEVRCLVDVPGQKVPSISNGEMAKYLKTMVAPQVPPEIYDSFVAAVDRGNIRTPMNRSMPAAPQPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRDLLRPLRDLNDAPTLCKYLESFYTLRKPVASTI NTLAGALYKVFCASPDRARNEMRQACFDYLSLGGVFSTGPVSLLSGLNPRPLSLVLHFFAVAVYGVGRLLVPFPSPSRMWIGARLISGASAIIFPIIKAEGVRQMFFPATVPAYYRAPPVNCQARSLKPDALKGGL Theobroma cacao MADSYVWGWILGSVMTLVALCGVVLKRKGSGISATRTESVKCVSSINGKCRSADGSDADVIIVGAGVAGSALAHTLGKDGRRVHVIERDLTPDRIVGELLQPGGYLKLIELGLEDCVEEIDAQQVFGYALFKDGKHTRLSYPLEKFHSDVSGRSFHNGRFIQRMREKSASLPNVRLEQ GTVTSLLEEKGTIRGVQYKTKDGRELTAFAPLTIVCDGCFSNLRRSLCNPKVDVPSCFVGLVLENCNLPYSNHGHVILADPPILFYPISSTEVRCLVVPGQKVPSIANGEMANYLKTIVAPQVPPEIYNSFVAAVDKGNIRTMPNRSMPAAPYPTPGALLMGDAFNMRHPLTGGGMTV ALSDIVVLRDLLRPLRDLNDAPTLCKYLESFYTLRKPIASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGVFSTGPISLLSGLNPRPVSLVLHFFAVAIYGVGRLLLPFPSPKRIWIGARLISGASGIIFPIIKAEGVRQMFFPATVPAYYRAPPVE Cucurbita moschata (SEQ ID NO: 33) MMVDHCAFAWILDVVLGLVVAVTFFVAAPRRNRRGGTDSTASKDCVISTAIANGECKPDDADAEVIIVGAGVAGSALAYTLGKDGRRVHVIERDLTEPDRIVGEFLQPGGYLKLIELGLGDCVEEIDAQKLY GYALFKDGKNTRVSYPLGNFHSDVSGRSFHNGRFIQRMREKAASLPNVRLEQGTVTSLLETKGTIKGVQYKSKNGEEKTAYAPLTIVCDGCFSNLRSLCKPMVDVPSCFVGLVLENCQLPFANHGHVVLGDP SPILFYPISSTEIRCLVDVPGQKVPSISNGDMEKYLKTVVAPQVPPQIHDAFIAAIEKGNVRTMPNRSMPAAPHPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLLKPLKDLNDASTLCKYLESFYTL RKPVASTINTLAGALYKVFCASPDQARKEMRQACFDYLSLGGVFSNGPISLLSGLNPRPSSLVLHFFAVAIYGVGRLLLPFPSLKGIWIGARLIYSASGIILPIIKAEGVRQMFFPATVPAYYRSPPVHKPIT Phaseolus vulgaris (SEQ ID NO: 34) MLDTYVFGWIICAALSVFVIRNFVFAGKKCCASSETDASMCAENITTAAGECRSSMRDGEFDVLIVGAGVAGSALAYTLGKDGRQVLVIERDLSEPDRIVGELLQPGGYLKLIELGLEDCVDKIDAQQVFG YALFKDGKHIRLSYPLEKFHSDVAGRSFHNGRFIQRMREKAASLPNVRLEQGTVTSLLEKGVIKGVQYKTKDSQELSVCAPFTIVCDGCFSNLRRRSLCDPKVDVPSCFVGLVLENCELPCANHGHVILGE PSPVLFYPISSTEIRCLVDVPGQKVPSISNGEMAKYLKTVIAPQVPHELHNAFIAVDKGSIRTMPNRSMPAAPYPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLLRPLRRDLDAPSLCKYLESF YTLRKPVAST INPUT LYKVFCASSDPARKEMRQACFDYLSLGGQFSEGPISLLSGLNPRLTLVLHFFAVATYGVGRLLLPFPSPKRMWIGLRLISSASGIIMPIIKAEGVRQMFFPATVPAYYRNPPAA Hevea brasiliensis MKMADHYLLGWILASVMGLFAFYYIVYLLVKPEEDNNRRSLPQPRSDFVKTMTATNGECRSDDDSDVDVIIVGAGVAGAALAHTLGKDGRRVHVIERDLTEPDRIVGELLQPGGYLKLIELGLEDCVEEIDAQRVFGYALFKDGKHTQLAYPLEKFHSEVAGRSFHNGRFIQRMREKAASLPSVKLEQGTVTSLLEEKGTIKGVLYKTKTGEELTAFAPLTIVCDGCFSNLRRSLCNPKVDVPSCFVGLVLENCRLPYANNGHVILADPSPILFYPISSSTEVRSLVDVPGQKVPSVSSGEMANYLKNVVAPQVPPEIYDSFVAAVDKGNIRTMPNRSMPASPYPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRDLLKPLRDLHDAPTLCRYLESFYTLRKPVASTINTLAGALYKVFCASPDEARKEMRQACFDYLSLGGVFSTGPVSLLSGLNPRPLSLVLHFFAVAIYGVGRLLLPFPSPHRIWVGARLISGASGIIFPIIKAEGVRQMFFPATVPAYYRAPPIKCN Sorghum bicolor MAAAAAAASGVGFQLIGAAAATLLAAVLVAAVLGRRRRARPQAPLVEAKPAPEGGCAVGDGRTDVIIVGAGVAGSALAYTLGKDGRRVHVIERDLTEPDRIVGELLQPGGYLKLIELGLEDCVEEIDAQRV LGYALFKDGRNTKLAYPLEKFHSDVAGRSFHNGRFIQRMRQKAASLPNVQLEQGTVTSLLEENGTVKGVQYKTKSGEELKAYAPLTIVCDGCFSNLRRALCSPKVDVPSCFVGLVLENCQLPHPNHGHVILA NPSPILFYPISSTEVRCLVDVPGQKVPSIASGEMANYLKTVVAPQIPPEIYDSFIAAIDKGSIRTMPNRSMPAAPHPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLLKPLHNLHDASSLCKYLESF YTLRKPVASTINTLAGALYKVFSASPDQARNEMRQACFDYLSLGGVFSNGPIALLSGLNPRPLSLVAHFFAVAIYWGGRLMLPLPSPKRMWIGARLISGACGIILPIIKAEGVRQMFFPATVPAYYRAAPMGE Zea mays MRKNLEEAGCAVSDGGTTDVIIVGAGVAGSALAYTLGKDGRRVHVIERDLTEPDRIVGELLQPGGYLKLIELGLQDCVEEIDAQRVLGYALFKDGRNTKLAYPLEKFHSDVAGRSFHNGRFIQRMRQKAASLPNVQLEQGTVTSLLEENGTVKGVQYKTKSGEELKAYAPLTIVCDGCFSNLRRALCSPKVDVPSCFVGLVLENCQLPHPNHGHVILANPSPILFYPISTSTEVRCLVDPGQK VPSIATGEMANYLKTVVAPQIPPEIYDSFIAAIDKGSIRTMPNRSMPAAPHPTPPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLLKPLRNLHDASSLCKYLESFYTLRKPVASTINTLAGALYKVFSASPDQARNEMRQACFDYLSLGGVFSNGPIALLSGLNPRPLSLVAHFFAVAIYGVGRLMLPPSPKRMWIGARLISGACGIILPIIKAEGVRQPATVPAYYRAAPTGEKAMFF Medicago sativa (sequence number 38) MDLYNIGWILSSVLSLFALYNLIFSGKRNYHDVNDKVKDSVTSTDAGDIQSEKLNDGADVIIVGAGIAGAALAHTLGKDGGRVRVHIIERDLSEPDRIVGELLQPGGYLKLVELGLQDCVDNIDAQRVFGYALFKDGKHTRLSYPLEKFHSDVSGRSFHNGRFIQRMREKAASLPNVNMEQGTVISLLEEKGTIKGVQYKNKDGQALTAYAPLTIVCDGCFSNLRRSLCNPKVDNPSCFVGLILENCELPCANHGHVILGDPSPILFYPISSTEIRCLVDVPGTKVPSISNGDMTKYLKTTVAPQVPPELYDAFIAAVDKGNIRTMPNRSMPADPRPTPGAVLMGDAFNMRHPLTGGGMTVALSDIVVLRNLKPMRDLNDAPTLCKYLESFYTLRKPVASTINTLAGALYKVFSASPDEARKEMRQACFDYLSLGGLFSEGPISLLSGLNPRPLSLVLHFFAVAVFGVGRLLLPFPSPKRVWIGARLLSGASGIILPIIKAEGIRQMFFPATVPAYYRAPPVNAF Methylomonas lenta MKEEFDICIIGAMGATISAYLAPKGIKIALIDHCYKEKKRIVGELLQPGAVLSLEQMGLSHLLDGFEAQTVKGYALLQGNEKTTIPYPSQHEGIGLHNGRFLQQIRASALENSSVTQIHGKALQLLENERNEIIGVSYRESITSQIKSIYAPLTITSDGFFSNFRAHLSNNQKTVTSYFIGLILKDCEMPFPKGHVF LSGPTPFICYPISDNEVRLLIDFPGEQLPRKNLLQEHLDTNVTPYIPECMRSSYAQAIQEGGFKVMPNHYMAAKPIVRKGAVMLGDALNMRHPLTGGGLTAVFSDIQILSAHLLAMPDFKNTDLIHEKIEAYYRDRKRANANLNILANALYAVMSNDLLKTAVFKYLQCGGANAQESIAVLAGLNRKHFSLIKQFCFLAVFGACNLLQQSISNIPKALKlLKDAFVIIKPLIKNELS Bathymodiolus azoricus Endosymbiont (Accession No. 168) MHTTSEHNDLFDICIVGAGMAGATIATYLAPRGIKIALIDRDYAEKRRIVGELLQPGAVQTLKKMGLEHLLEGFDAQPIYGYALFNKDCEFSIEYNQDKSTNYRGVGLHNGRFLQKIREDALKQPSITQIHGTVSELIEDENHVVTGVKYKEKYTRELKTVNAKLTITSDGFFSSFRKDLTNNVKTVTSFFVGIILKDCELPYPHHGHVFLSAPTPFICYPISSTESRLLIDFPGDQAPKKEAVKHHIENNVIPFLPKEFRLCLDQALRENDYKIMPNHYMPAKPVLKKGVVLLGDALNMRHPITGGGLTAVFNDVYLLSTHLLAMPDFNDTKLIHEKVNLYYNDRYHANTNVNIMANALYGVMSNDLLKQSVFEYLRKGGDNSGGPISLLAGLNRNPTILIKHFFSVALLCLRNLFKAHKMSLTNAFYVIKDAFCIIVPLAINELRPSSFLKKNIHN Methyloprofundus sediment (Accession No. 169) MNTSPEHNDLFDICIVGVGMAGATIAAYLAPRGLKIALIDREYTEKRIVGELLQPGAVQTLKKMGLEHLLEGFDAQPIYGYALFNNDKEFSYSYNSDDSTEYHGVGLHNGRFLQKIREDVFKNETVTQIHGTVSELIEDKKGVVKGVTYREKHTREYKTVKAKLTVTSDGFFSNFRKDLSNNVKTVTSFFIGLVLDCNLPFPNHGHVLFSLAPTPFICYPISSTETRL LIDYPGDKAPKKDEIREHILNKVAPFLFPEEFKECFANAMEDDDFKVMPNHYMPAKPVLKEGAVLLGDALNMRHPLTGGGLTAVFNDVYLLSTHLLAMPDFNDPKLLHEKLELYYQDRYHANTNVNIMANALYGVMSNDLLKQGVFEYLRKGGDNSGGPITLLAGLNRNPTLIKHFFSVAFLCICNLSGNNKMNFTNVFRVMKDAFCIIKPLAVNELRPSSFYKKNIQL Methylomicrobium buryatense MESNFDICIIGAMGAGATIAAYLAPKGINIALIDHCYKEKKRIVGELLQPGAVLSLEQLGLGHLLDGIDAQPVEGYALLQGNEQTTIPYPSPNHGMGLHNGRFLQQIRASALQNSSVTQIQGKALSLLENEQNEIGVNYRDSVSNEIKSIYAPLTITSDGFFSNFRELLSNNEKTVTSYFIGLILKDCEIPVPKHGHVFLSGPTPFICYPISSNEVRLLIDFP GGQFPRKAFLQAHLETNVTPYIPEGMQTSYRHALQEDRLKVMPNHYMAAKKPIRKGAVMLGDALNMRHPLTGGGLTAVFSDIEILSGHLLAMPDFNNNDLIYQKIEAYYRDRQYANALNILANALYGVMSNELLKNSVFKYLQRGGVNAKESIAILAGLNKNHYSLMKQFFFVALFGAYTLVRENITNLPKATKILSDALTIIKPLAKNELSLVGIFSDYFKR Ononis spinosa SQE1 (SEQ ID NO: 177) MVDPYAVGWIICSLTTIVALYNFVFYRQNRSDKTTPTTTENITTATGDCRSLNPNGDVDIVIVGAGVAGSALAYTLGKDGRRVLVIERDLNEPDRIVGELLQPGGYLKLIELGLEDCVEK IDAQQVFGYALFKDGKHTRLSYPLEKFHSDIAGRSFHNGRFIQRMREKAASLPNVQLVQGTVTSLLEENGTIKGVQYKTKDAQELSACAPLTIVCDGCFS NLRRNLCNPKVEVPSCFVGLVLENCELPCANHGHVILGDPSPVLFYPISSTEIRCLVDVPGQKVPSISNGEMAKYLKEVVAPQVPPELHDAFIAAVDKGNI RTMPNRSMPAAPYPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLLKPLRDLNDAPSLCKYLESFYTLRKPVASTINTLAGALYKVFCASPDPARKEMRQACFDYLSLGGLFSEGPVSLLSGLNPRPLSLVLHFFAVAIYGVGRLLLPFPSPKRIWIGVRLIASASGIILPIIKAEGIRQMFFPATVPAYYRTPPAA Ononis spinosa SQE2 (SEQ ID NO: 178) MDLYLLGWILSSVLSLFALYCLVFDGNRSRANAEKQIQRGYSVTTDAGDVKSEKLNGDADVIIVGAGIAGAALAHTLGKDGRRVRVIERDLSEPDRIVGELLQPGGYLKLVELGLADCVDNIDAQKVFGYALFKDGKHTRLSYPLEKFHADVSGRSFHNGRFIQRMREKAASLLNVNLEQGTVTSLLEEKGTIKGVQYKNKDGQELTAYAPLTIVCDGCFSNLRRSLCNPKVDNPSCFVGLVLENCELPCANHGHVILGDPSPILFYPISSTEIRCLVDVPGQKVPSISNGDMTKYLKLTVAPQVPPELYDAFIAAVDKGNIRTMPNKSMPADPCPTPGAVLMGDAFNMRHPLTGGGMTVALSDIVVLRNLLRPLRDLNDAPALCKYLESFYTLRKPVASTINTLAGALYKVFSSSPDQARREMRQACFDYLSLGGLFSEGPISLLSGLNPRPLSLVLHFFAVAVFGVGRLLLPFPSPKRVWIGARLLSAASGIILPIIKAEGIRQMFFPVTVPAYYRAPPTSQE Medicago truncatula SQE1 (SEQ ID NO: 179) MIDPYGFGWITCTLITLAALYNFLFSRKNHSDSTTTENITTATGECRSFNPNGDVDIIIVGAGVAGSALAYTLGKDGRRVLIIERDLNEPDRIVGELLQPGGYLKLIELGLDDCVEKIDAQKVFGYALFKDGKHTRLSYPLEKFHSDIAGRSFHNGRFILRMREKAASLPNVRLEQGTVTSLLEENGTIKGVQYKTKDAQEFSACAPLTIVCDGCFSNLRRSLCNPKVEVPSCFVGLVLENCELPCADHGHVILGDPSPVLFYPISSTEIRCLVDVPGQKVPSISNGEMAKYLKTVVAPQVPPELHAAFIAAVDKGHIRTMPNRSMPADPYPTPGALLMGDAFNMRHPLTGGGMTVALSDIVVLRNLLKPLRDLNDASSLCKYLESFYTLRKPVASTINTLAGALYKVFCASPDPARKEMRQACFDYLSLGGLFSEGPVSLLSGLNPCPLSLVLHFFAVAIYGVGRLLLPFPSPKRLWIGIRLIASASGIILPIIKAEGIRQMFFPATVPAYYRAPPDA Medicago truncatula SQE2 (Accession No. 180) MDLYNIGWILSSVLSLFALYNLIFAGKKNYDVNEKVNQREDSVTSTDAGEIKSDKLNGDADVIIVGAGIAGAALAHTLGKDGRRVHIIERDLSEPDRIVGELLQPGGYLKLVELGLQDCVDNIDAQRVFGYALFKDGKHTRLSYPLEKFHSDVSGRSFHGRFIQRMREKAASLPNVNMEQGTVISLLEEKGTIKGVQYKNKDGQALTAYAPLTIVCDGCFSNLRRSLCNPKVDNPSCFVGLILENCELPCANHGHVILGDPSPILFYPISSTEIRCLVDVPGTKVPSISNGDMTKYLKTTVAPQVPPELYDAFIAAVDKGNIRTMPNRSMPADPRPTPGAVLMGDAFNMRHPLTGGGMTV ALSDIVVLRNLLKPMRDLNDAPTLCKYLESFYTLRKPVASTINTLAGALYKVFSASPDEARKEMRQACFDYLSLGGLFSEGPISLLSGLNPRPLSLVLHFFAVAVFGVGRLLLPFPSPKRVWIGARLLSGASGIILPIIKAEGIRQMFFPATVPAYYRAPPVNAF Hypholoma sublateritium SQE (Sequence No. 181) MSKSRSNYDVIIVGAGIAGCALAHGLSTLSRATPLRIAIVERSLAEPDRIVGELLQPGGVMALQRLGMEGCLEGIDAVKVHGYCVVENGTSVHIPYPGVHEGRSFHHGRFIMKLREAARAARGVELVEATVTELIPREGGKGIAGVRVARKGKDGEEDTTEALGAALVVVADGCFSNFRAAVMGGAAVKPETKSHFVGAILKDARLPIPNHGTVALVKGFGPVLLYQISEHDTRMLVDVKAPLPADLKVCAHILSNIVPQLPAALHLPIQRALDAERLRRMPNSFLPPVEQGATRGAVLVGDAWNMRHPLTGGGMTVALNDVVVLRDLLGSVGDLGDWRQVASTVNILSVALYDLFGADGELQVLRTGCFKYFERGGDCIDGPVSLLSGIAPSPMLLAYHFFSVAFYSIYVIAVGAQNGSAKQVLAVPGALQYPALCVKGLRVFYTACVVFGPLLWTELRW Hypholoma sublateritium SQE2 (Sequence No. 182) MHPTHYDVVIVGAGVAGSSLAHALATLPREKPLQIALIERSFEEPDRIVGELLQPGGVDALKTLKMTSSVEGIDAITVTGYILVESGDMVRIPYPKGKEGRSFHHGRFIMGLRRVALENPNVHPIEATAADLIECPCTGQVIGVRATSKTAPAPSSIDAQQTPPAPFSVYGDLVIVADGCFSNFRNVVMGKAACKATTKSYFVGTILKDAVLPVAGHGTVILPQGSGPVLLYQISEHDTRMLIDIQHPLPSDLRAHILTNILPQLPASIQGVVSDAFTKDRIRRMPNSFLPSVQQGSPLSKKGVILLGDSWNMRHPLTGGGMTVALNDVVYLRSIFASIQNLDDWDEIRYALRHWHWGRKPLSSTINILSGTLYGLFEKDDDDYRALRKGCFKYFQLGGKCIDDPVSLLSGLSPSPLLLSSHFFAVILYAIWVVFTHPRVGSSMSANPADVKRVYDIPSADEYPQLTLKGIRMFSQACGVFLPVLWSEIRWWAPCESS Hypholoma sublateritium SQE3 (SEQ ID NO: 183) MSKSRSNYDVIIVGAGIAGCALAHGLSTLSRATPLRIAIVERSLAEPDRIVGELLQPGGVMALQRLGMEGCLEGIDAVKVHGYCVVENGTSVHIPYPGVHEGRSFHHGRFIMKLREAARAARGVELVEATVTELIPREGGKGIAGVRVARKGKDGEEDTTEALGAALVVVADGCFSNFRAAVMGGAAVKPETKSHFVGAILKDARLPIPNHGTVALVKGFGPVLLYQISEHDTRMLVDVKAPLPADLKAHILSNIVPQLPAALHLPIQRALDAERLRRMPNSFLPPVEQGATRGAVLVGDAWNMRHPLTGGGMTVALNDVVVLRDLLGSVGDLGDWRQVRRALHRWHWDRKPLASTVNILSVALYDLFGADGEELQVLRTGCFKYFERGGDCIDGPVSLLSGIAPSPMLLAYHFFSVAFYSIYVMFAHPQPVAQSKAVGAQNGSAKQVLAVPGALQYPALCVKGLRVFYTACVVFGPLLWTELRWWTAAEASRGRLLVMSLVPLLLLLGAANYGIPGMGLLGVL MlSQE A4 (SEQ ID NO: 203) MAKEEFDICIIGAGMAGATISAYLAPKGIKIALIDRCYKEKKRIVGELLQPGAVLSLEQMGLSHLLDGFEAQTVKGYALL QGNEKTTIPYPSQHEGIGLHNGRFLQQIRASALENSSVTQIHGKALQLLENERNEIIGVSYRESITSQIKSIYAPLTITSDGFASNFRAHLSNNQKTVTSYFIGLILKDCEMPFPKHGHVFLSGPTPFICYPISDNEVRLLIDFPGEQLPRKNLLQEHLDTNVTPYIPECMRSSYAQAI QEGGFKVMPNHYMAAKPIVRKGAVLLGDALNMRHPLTGGGLTAVFSDIQILSAHLLAMPDFKNTDLIHEKIEAYYRDRKRANANLNILANALYAVMSNDLLKTAVFKYLQCGGANAQESIALLAGLNRKHFSLIKQYCFLAVFGACNLLQQSISNIPKALKLLKDAFVIIKPLIKNELS Cucurbitadienol synthase (CDS), triterpene synthase (TTP) Siraitia grosvenorii CDS (SEQ ID NO: 40) MWRLKVGAESVGENDEKWLKSISNHLGRQVWEFCPDAGTQQQLLQVHCARKAFHDDRFHRKQSSDLFITIQYGKEVENGGKTAGVKLKEGEVRKEAVESSLERALSFYSSIQTSDGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYVYNHQNEDGGWGLHIEGPSTMFGSALNYVAL RLLGEDANAGAMPKARAWILDHGGATGITSWGKLWLSVLGVYEWSGNNPLPPEFWLFPYFLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYAVPYHEIDWNKSRNTCAKEDLYYPHPKMQDILWGSLHHVYEPLFTRWPAKRLREKALQTAMQHIHYEDENTRYICLGPVNKVLNLLC CWVEDPYSDAFKLHLQRVHDYLWVAEDGMKMQGYNGSQLWDTAFSIQAIVSTKLVDNYGPTLRKAHDFVKSSQIQQDCPGDPNVWYRHIHKGAWPFSTRDHGWLISDCTAEGLKAALMLSKLPSETVGESLERNRLCDAVNVLLSLQNDNGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTSAT MEALTLFKKLHPGHRTKEIDTAIVRAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNNCLAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPTPLHRAARLLINSQLENGDFPQQEIMGVFNKNCMITYAAYRNIFPIWALGEYCHRVLTE Momordica charantia MWRLKVGAESVGENDEKWVKSISNHLGRQVWEFCPDAGTPQQLLQIEKARKAFQDNRFHRKQTSDLLVSIQCEKGTTNGARVPGTKLKEEGEEVRKEAVKSTLERALSFYSSIQTSDGNWASDLGGPMFLLPGLVIALCVTGALNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIESPSTMFGSALNYVALR LLGEDADGGEGRAMTKARAWILGHGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYFLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPVLSLRKELYTVPYHEIDWNKSRNTCAKEDLYYPHSKMQDILWGSIHHMYEPLFTHWPAKRLREKALKTAMQHIHYEDENTRYICLGPVNKVLNM LCCWVEDPYSEAFKLHLQRVHDYLWVAEDGMKMQGYNGSQLWDTAFSVQAIISTKLVDNYGPTLRKAHDYVKNSQIQQDCPGEPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSETVGEPLERNRLCDAVNVLLSLQNDNGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTS ATMEALALFKKLHPGHRTKEIDTAIARAADFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRAYSNCLAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQGERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYCHRVL TEA Cucurbita maxima MWRLKVGAESVGEKDEKWVKSVSNHLGRQVWEFCADAAADTPHQLLQIQNARNHFHHNRFHRKQSSDLFLAIQYEKEIAKGAKGGAVKVKEGEEVGKEAVKSTLERALGFYSAVQTSDGNWASDLGGPMFLLPGLVIALHVTGVLNSVLSKHHRVEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVAL RLLGEDADGGDGGAMTKARAWILERGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYSLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPKVLSLRQELYTIPYHEIDWNKSRNTCAKEDLYYPHPKMQDILWGSIYHVYEPLFTRWPGKRLREKALQAAMKHIHYEDENSRYICLGPVNKVLNM LCCWVEDPYSDAFKLHLQRVHDYLWVAEDGMRMQGYNGSQLWDTAFSIQAIVATKLVDSYAPTLRKAHDFVKDSQIQEDCPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSTMVGEPLEKNRLCDAVNVLLSLQNDNGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTAA TMEALTLFKKLHPGHRTKEIDTAIGKAANFLEKMQRADGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCLAIRKACEFLLSKELPGGGWGESYLSCQNKVYTNLEGNKPHLVNTAWVLMALIEAGQGERDPAPLHRAARLLMNSQLENGDFVQQEIMGVFNKNCMITYAAYRNIFPIWALGEYCHRVLTE Citrullus colocynthis (CcCDS1) (SEQ ID NO: 43) MWRLKVGAESVGEKEEKWLKSISNHLGRQVWEFCADQPTASPNHLQQIDNARKHFRNNRFHRKQSSDLFLAIQNEKEIANGTKGGGIKVKEEEDVRKETVKNTVERALSFYSAIQTNDGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVAL RLLGEDADGGEGGAMTKARGWILDRGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYCLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYTIPYHEIDWNKSRNTCAKEDLYYPHPKMQDILWGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYICLGPVNKVLNM LCCWVEDPYSDAFKFHLQRVPDYLWIAEDGMRMQGYNGSQLWDTAFSVQAIISTKLIDSFGTTLKKAHDFVKDSQIQQDFPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSKIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTSA TMEALTLFKKLHPGHRTKEIDTAVAKAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYSTCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE Citrullus colocynthis (CcCDS2) (SEQ ID NO: 44) MWRLKVGAESVGEKEEKWLKSISNHLGRQVWEFCAHQPTASPNHLQQIDNARNHFRNNRFHRKQSSDLFLAIQNEKEIANVTKGGGIKVKEEEDVRKETVKNTVERALSFYSAIQTNDGN WASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVALRLLGEDADGGEGGAMTKARSWILDRGGATAITSWGKLWLSVLGVYEWSG NNPLPPEFWLLPYCLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYTIPYHEIDWNRRSRNTCAKEDLYYPHPKMQDILWGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYI CLGPVNKVLNMLCCWVEDPYSDAFKFHLQRVPDYLWVAEDGMRMQGYNGSQLWDTAFSVQAIISTKLIDSFGTTLKKAHDFVKDSQIQQDCPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLM LSKLPSKIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTSATMEALTLFKKLHPGHRTKEIDIAVARAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGL VAAGRTYNSCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE Cucurbita moschata (SEQ ID NO: 45) MWRLKVGAESVGEKDEKWVKSVSNHLGRQVWEFCADAAAAAATPRQLLQIQNARNHFHRNRFHRKQSSDLFLAIQYEKEIAEGGKGGAVKVKEEEEVGKEAVKSTLERALSFYSAVQTSDGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRVEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVA LRLLGEDADGGDDGAMTKARAWILERGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYSLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITKPVLSRQELYTVPYHEIDWNKSRNTCAKEDLYYPHPKMQDILWGSIYHVYEPLFTRWPGKRLREKALQTAMKHIHYEDENSRYICLGPVNKVLN MLCCWVEDPYSDAFKLHLQRVHDYLWVAEDGMRMQGYNGSQLWDTAFSIQAIVATKLVDSFAPTLRKAHDFVKDSQIQEDCPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSTMVGEPLEKNRLCDAVNVLLSLQNDNGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTA ATMEALTLFKKLHPGHRTKEIDTAVGKAANFLEKMQRADGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCLAIRKACEFLLSKELPGGGWGESYLSCQNKVYTNLEGNKPHLVNTAWVLMALIEAGQGERDPAPLHRAARLLMNSQLENGDFVQQEIMGVFNKNCMITYAAYRNIFPIWALGEYCHRVLTE Cucumis sativus (sequence number 46) MWRLKVGKESVGEKEEKWIKSISNHLGRQVWEFCAENDDDDDDEAVIHVVANSSKHLLQQQRRQSSFENARKQFRNNRFHRKQSSDLFLTIQYEKIARNGAKNGGNTKVKEGEDVKKEAVNNTLERALSFYSAIQTSDG NWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRHQEMCRYYYNHQNEDGGWGLHIEGSSTMFGSALNYVALRLLGEDANGGECGAMTKARSWILERGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYSLPFHP GRMWCHCRMVYLPMSYLYGKRFVGPITMHVLSLRKELYTIPYHEIDWNRSRNTCAQEDLYYPHPKMQDILWGSIYHVYEPLFNGWPGRRLREKAMKIAMEHIHYEDENSRYIYLGPVNKVLNMLCCWVEDPYSDAFKFHL QRIPDYLWLAEDGMRMQGYNGSQLWDTAFSIQAILSTKLIDTFGSTLRKAHHFVKHSQIQEDCPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSKIVGEPLEKNRLCDAVNVLLSLQNENGGFAS YELTRSYPWLELINPAETFGDIVIDYSYVECTSATMEALALFKKLHPGHRTKEIDAALAKAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNNCVAIRKACHFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQGERDPAPLHRAARLLINSQLENGDFPQQEIMGVFNKNCMITYAAYRNIFPIWALGEYSHRVLTE Cucumis melo MWRLKVGKESVGEKEEKWIKSISNHLGRQVWEFCSGENENDDDEAIAVANNSASAKFENARNHFRNNRFHRKQSSDLFLAIQCEKEIIRNGAKNEGTTKVKEGEDVKKEAVKNTLERALSFYSAVQTSDGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYYYNHQNEDGGWGLHIEGSSTMFG SALNYVALRLLGEAADGGEHGAMTKARSWILERGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYSLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYTIPYHEIDWNRSRNTCAKEDLYYPHPKMQDILWGSIYHVYEPLFSGWPGKRLREKAMKIAMEHIHYEDENSRYICLGPVN KVLNMLCCWVEDPYSDAFKFHLQRIPDYLWLAEDGMRMQGYNGSQLWDTAFSIQAIISTKLIDTFGPTLRKAHHFVKHSQIQEDCPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSKIVGEPLEKNRLCDAVNVLLSLQNENGGFASYELTRSYPWLELINPAETFGDIVIDYSYVEC TSATMEALALFKKLHPGHRTKEIDAAIAKAANFLENMQKTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNNCVAIRKACNFLLSKELPGGGWGESYLSCQNKVYTNLEGNKPHLVNTAWVMMALIEAGQGERDPAPLHRAARLLINSQLESGDFPQQEIMGVFNKNCMITYAAYRNIFPIWALGEYSHRVLDM Citrullus lanatus subsp. vulgaris DGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVALRLLGEDADGGEGGAMTKARSWILDRGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYCLPFHPGRMWCHCRMVYLPMSYLYG KRFVGPITPIVLSLRKELYTIPYHEIDWNRSRNTCAKEDLYYPHPKMQDILWGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYICLGPVNKVLNMLCCWVEDPYSDAFKFHLQRVPDYLWVAEDGMRMQGYNGSQLWDTAFSVQAIISTKLID SFGTTLKKAHDFVKDSQIQQDCPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSEIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTSATMEALTLFKKLHPGRRTKEIDIAVARAAN FLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE Theobroma cacao MWRLKIGKESVGDNGAWLRSSNDHVGRQVWEFCPESGTPEELSKVEMARQSFSTDRLLKKHSSDLLMRIQYAKENQFVTNFPQVKLKEFEDVKEEATLTTLRRRALNFYSTIQADDGHWPG DYGGPMFLLPGLVITLSVTGALNAVLSKEHQYEMCRYLYNHQNRDGGWGLHIEGPSTMFGTVLNYVTLRLLGEGPEGGQGAVEKACEWILEHGSATAITSWGKMWLSVLGAYEWSGNNPLPPEVWLCPYFLPIHPGRMWCHCRMVYLPMSYLYGKRFVGP ITPIILSLRKELYAVPYHEVDWNKARNTCAKEDLYYPHPLVQDILWASLHYLYEPIFTRWPCKSLREKALRTVMQHIHYEDENTRYICIGPVNKVLNMLSCWVEDPYSESFKLHLPRILDYLWIAEDGMKMQGYNGSQLWDTAFAVQAIISTGLADEYGP ILRKAHDFIKYSQVLEDCPGDLNFWYRHISKGAWPFSTVDHGWPISDCTSEGLKAVLLLSTLPSESVGEPLHMMRLYDAVNVILSLQNVDGGFPTYELTRSYQWLELINPAETFGDIVIDYPYVECTSAAIQALISFKKLFPEHRMEEIENCIGRAVEFI EKIQAADGSWYGSWGVCFTYAGWFGIKGLSAAGRTYNSSNIRKACDFLLSKELATGGWGESYLSCQNKVYTNLEGARPHIVNTSWALLALIEAGQAERDPTPLHRAARILINSQMEDGDFPQEEIMGVFNKNCMISYSAYRNIFPIWALGEYTCRVLRAP Ziziphus jujube MWKLKIGAETVGEGGSDGWLRSVNSHLGRQVWEFHPELGTPEELRQIQDARDAFFNHRFHKQHSSDLLMRIQFAKENPCVANPPQVKVKDTDEVTEESVTTTLRRAINFYSTIQAHDGHWAGDYGGPMFLLPGLVITLSVTGALNAVLSKEHQCEMCRYIYNHQNEDGGWGLHIEGPSTMFGTVLNYVSL RLLGEGAEDGLGTIENARKWILDHGGATAITSWGKMWLSVLGVYEWSGNNPLPPEVWLCPYTLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPTIRSLRKELYTAPYHEIDWNRARNECAKEDLYYPHPLVQDVLWASLHYVYEPIFMRWPAKKLREKALSTVMQHIHYEDENTRYICIGPVNKVLNMML CCWVEDPNSEAFKLHLPRISDYLWIAEDGMKMQGYNGSQLWDTAFAVQAIVSTDLAEEYGPTIRKAHEYIKNSQVLEDCPGDLNFWYRHISKGAWPFSTADHGWPISDCTAEGLKAVLLLLSQLSSETVGDSLDVKRLFNAVNVILSLQNGDGGFATYELTRSYQWLELINPAETFGDIVIDYPYVECTSAA LEALTLFKKSYPGHRREEVENCITNAAMFIENIQAKDGSWYGSWGVCFTYAGWFGIKGLVASGRTYENCPSIRKACDFLLSKELPSGGWGESYLSCQNKVYTNLKDNKPHIVNTAWAMLALIVARQAERDPMPLHRAARILIKSQMHDDGFPQEEIMGVFNKNCMISYAAYRNIFPIWALGEYRLHVLRSL Prunus avium MWKLKIGAETVGEGGYQWLKSVNNHLGRQVWEFNPELGSPEELQRIEDARKAFWDNRFERRHSSDLLMRIQFEKENQCVTNLPQLKVKYEEEVTEEVVKTTLRRAISFYSTIQAHDGHWPGDYGGPMFLLPGLVITLSITGALNDVLSKEHQHEMCRYLYNHQNKDGGWGLHIEGPSTMFGTALNYVTLRLFGEGADDGAMELARKWILDHGGVTKIT SWGKMWLSVLGTYEWSGNNPLPPEVWLCPYSLPHFHPGRMWCRMVYLPMSYLYGKRFVGPITPTIRSLRKELYGVPYHEVDWNQARNLCAKEDLYYPHPMVQDILWASLHYVYEPVFTRWPAKKLRENALQTVMQHIHYEDENTRYICIGPVNKVLNMLCCWAEDPNSDAFKLHLPRIPDYLWVAEDGMKMMQGYNGSQSWDTSFAVQAIISTNLAEEFG PTLRKAHEYIKDSQVLEDCPGDLNFWYRHISKGAWPFSTADHGWPISDCTAEGLKAVLLLSKLPTGTVGESLDMKQLYDAVNVMLSLQNEDGGFATYELTRSYQWLELINPAETFGDIVIDYPYVECTSAAIQALTMFRKLYPGHRREEIESCIARAAKFI EKIQATDGSWYGSWGVCFTYAGWFGIKGLAAAGRTYKDCSSIRKACDFLLSKELPSGGWGESYLSCQNKVYTNLKDNRPHIVHTAWAMLALIGAQAKRDPTPLHRAARVLINSQMENDGFPQKEIMGVFNKNCMISYSAYRNIFPIWAlgEYRCQVLLEAL Brassica napus MWKLKIAEGGSPWLRTTNNHVGRQFWEFDPNLGTPEELAAVEEARKSFRENRFAKKHSSDLLMRLQFSRESLSRPVLPQVNIKDGDDVTEKMVETTLKRGVDFYSTIQASDGHWAGDYGGPMFLLPGLIITLSITGALNTVLSEQHKAEMRRYLHNHQNEDGGWGLHIEGPSTMFGSVLNYVTLRLLGE GPNDGDGAMEKGRDWILNHGGATNITSWGKMWLSVLGAFEWSGNNPLPPEIWLLPYILPIHPGRMWCHCRMVYLPMSYLYGKRFVGPITSTVLSLRKELFTVPYHEVDWNEARNLCAKEDLYYPHPLVQDILWASLHKIVEPVLTRWPGSNLREKALRTTLEHIHYEDENTRYICIGPVNKVLNMLCCWV EDPNSEAFKLHLPRIHDYLWVAEDGMKMQGYNGSQLWDTSFAVQAVLATNFVEEYGPVLKKAHSYVKNSQVSEDCPGDLSYWYRHISKGAWPFSTADHGWPISDCTAEGLKAALLSKVPKEIVGEPVDTKRLYDAVNVIISLQNADGGFATYELTRSYPWLELINPAETFGDIVIDYPYVECTSAAIQA LIAFRKLYPGHRKKEVDECIEKAVKFIESIQESDGSGWYGSWAVCFTYGTWFGVKGLEAAGKTLKNSPTVAKACEFLLSKQLPSGGWGESYLSCQDKVYSNLDGNRSHVVNTAWALLSLIGAGQVEVDQKPLHRAARYLINAQMESGDFPQQEIMGVFNRNCMITYAAYRNIFPIWALGERYRSKVLLQQGE Spinacia oleracea MWKLKIAEGGSPWLRTTNNHVGRQIWEFDPNLGTPEQIREVEEARENFWKNRFEQKHSSDLLMRQFAQENSSNVVLPQVKVKDEDEITEETVATTLRALSYQSTIQAHDGHWPGDYGGPMFLMPGLVIALSVTGALNAVLSKEHQKEMCRYLYNHQNKDGGWGLHIEGHSTMFGTVLTYVTLRLLGE GVDDDGGAMERGRKWTLEHGSATAITSWGKMWLSVLGVFEWAGNMPMPPETWLLPYILPVHPGRMWCHCRMVYLPMSYLYGKRFVGPITTPVLSRRELFDVPYHEIDWDRARNECAKEDLYYPHPLVQDILWASLHKAVEPILMRWPGKKLREKALSTVMEHIHYEDENTRYICIGPVNKVLNMLCCW VEDPNSEAFKLHLPRIPDFLWIAEDGMKMQGYNGSQLWDTTFMVQAILATNLGEEYGGTLRKAHNFIKDSQVREDCPGDLSYWYRHISKGAWPFSTADHGWPISDCTAEGLKAALLLSKVPSDIVGEPLEVKRLYDSVNVLLSLQNGDGGFATYELTRSYPWLELINPAETFGDIVIDYPYVECTSAAI QALVSFKRLYPGHRREEIENCIKKAAKFIEDIQAADGSWYGSWAVCFTYATWFGIKGLVAAGKNYDNCPAIRKACDFLLSKQLSNGGWGESYLSCQNKVYSNIEGNKAHVVNTGWAMLALIGQAKRDPMPLHRAAKVLINSQMPNGDFPQQEIMGVFNRNCMITYYAAYRNIFPTWALGEYRTQVLQK Trigonella foenum graecum MWKLKVAEGGSPWLRTVNNYVGRQVWEFDPNSGSPQELDQIESVRQNFHNNRFSHKSHDDLLMRIQLAKENPMGEVIPCVRVKDVEDVNEESVTTLLRRALNFYSTLQSRDGHWPGDYGGPMFLMPGLVIALSITGALNAVLTDEHQKEMRRYLYNHQNKDGGWGLHIEGPSTMFGVLCYVTLRLLGE GPNDGEGEMEKARDWILEHGGATYITSWGKMWLSVLGVFEWSGNNPLPPEIWLLPYMLPIHPGRMWCHCRMVYLPMSYLYGKRFVGPITTPVLSRKELFTVYHDIDWNQARNLCAKEDLYYPHPLVQDILWASLHKFVEPIFMNWPGKKLREKAVETVMEHVHYEDENTRYICIGPVNKVLNMLCCW VEDPNSEAFKLHLPRIHDFLWIAEDGMKMQGYNGSQLWDTAFAVQAXISTNLIDEFAPTLRKAHTFIKNSQVLEDCPDGLSKWYRHISKGAWPFSTADHGWPISDCTAEGLKAVLLLLSKIGPEIVGEPLDAKGFYDAVNVIISLQNEDGGLATYELTRSYKWLEIINPAETFGDIVIDYTYVECTSAAI QALSTFRKLYPGHRREEIQHCIEKAAAAFIEKIQASDGSWYGSWGVCFTYGTWFGVKGLIAAGKSFSNCLSIRKACDFLLSKQLPSGGWGESYLSCQNKVYSNLESNRSHVVNTGWAMLALIEAEQAKRDPTPLHHAAVCLINSQMENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYRRRHVLQA Ricinus communis MWKLRIAEGSGNPWLRTTNDHIGRQVWEFDSSKIGSPEELSQIENARQNFTKNRFIHKHSSDLLMRIQFSKENPICEVLPQVKVKESEQVTEEKVKITLRRALNYYSSIQADDGHWPGDYGGPMFLMPGLIIALSITGALNAILSEEHKREMCRYLYNHQNRDGGWGLHIEGPSTMFGSVLCYVSLRLL GEGPNEGEGAVERGRNWILKHGGATAITSWGKMWLSVLGAYEWSGNNPLPPEMWLLPYILPVHPGRMWCHCRMVYLPMSYLYGKRFVGPITPTVLSLRKELYTVPYHEIDWNQARNQCAKEDLYYPHPMLQDVLWATLHKFVEPILMHWPGKRLREKAIQTAIEHIHYEDENTRYICIGPVNKVLNMLCC WVEDPNSEAAFKLHLPRLYDYLWLAEDGMKMQGYNGSQLWDTAFAVQAIVSTNLIEEYGPTLKKAHSFIKKMQVLENCPGDLNFWYRHISKGAWPFSTADHGWPISDCTAEGIKALMLLSKIPSEIVGEGLNANRLYDAVNVVLSLQNGDGGFPTYELSRSYSWLEFINPAETFGDIVIDYPYVECTSAAI QALTSFRKSYPEHQREEIECCIKKAAKFMEKIQISDGSWYGSWGVCFTYGTWFGIKGLVAAGKSFGNCSSIRKACDFLLSKQCPSGGWGESYLSCQKKVYSNLEGDRSHVVNTAWAMLSLIDAGQAERDPTPLHRAARYLINAQMENGDFPQQEIMGVFNRNCMITYAAYRDIFPIWALGEYRCRVLKAS Pisum sativum cycloartenol synthase (PsCAS_mut) (SEQ ID NO: 191) MAWKLKVAEGGTPWLRTLNNHVGRQVWEFDPHSGSPQDLDDIETARRNFHDNRFTHKHSDLLMRLQFAKENPMNEVLPKVKVKDVEDVTEEAVATTLRRGLNFYSTIQSHDGHWPGDLGGPMFLMPGLVITLSVTGALNAVLTDEHRKEMRRYNHQNKDGGWGLHIEGPSTMFGSVL CYVTLRLLGEGPNDGEGDMERGRDWILEHGGATYITSWGKMWLSVLGVFEWSGNNPMPPEIWLLPYALPVHPGRMWCHCRMVYLPMSYLYGKRFVGPITPTVLSLRKELFTVPYHDIDWNQARNLCAKEDLYYPHPLVQDILWATLHKFVEPVFMNWPGKKLREKAIKTAIEHIHYEDEN TRYICIGPVNKVLNMLCCWVEPDPNSEAFKLHLPRIYDYLWVAEDGMKMQGYNGSQLWDTAFAAQAIISTNLIDEFGPTLKKAHAFIKNSQVSEDCPGDLSKWYRHISKGAWPFSTADHGWPISDCTAEGLKAVLLLSKIAPEIVGEPLDSKRLYDAVNVILSLQNENGGLATYELTRSYTWLEIINPAETFGDIVIDC PYVECTSAAIQALATFGKLYPGHRREEIQCCIEKAVAFIEKIQASDGSWYGSWGVCFTYGTWFGIKGLIAAGKNFSNCLSIRKACEFLLSKQLPSGGWAESYLSCQNKVYSNLEGNRSHVVNTGWAMLALIEAEQAKRDPTPLHRAAVCLINSQLENGDFPQEEIMGVFNKNCMITYAAYRCIFPIWALGEYRRVLQAC Cucurbita pepo subsp. pepo cycloartenol synthase (CpCAS_mut) (SEQ ID NO: 192) MAWQLKIGADTVPSDPSNAGGWLSLTLNNHVGRQVWHFHPELGSPEDLQQIQQARQHFSDHRFEKKHSADLLMRQFAKENSSSFVNLPQVKVKDKEDVTEEAVTRTLRRAINFYSTIQADDGHWPGDLGGPMFLIPGLVITLSITGALNAVLSTEHQREICRYLYNHQNKDGGWGLHIEGPSTMFGSVLNYV TLRLLGEEAEDGQGAVDKARKWILDHGGAAAITSWGKMWLSVLGVYEWAGNNPLPPELWLPYLLPCHPGRMWCRMVYLPMCYLYGKRFVGPITPIIRSLRKELYLVPYHEVDWNKARNQCAKEDLYYPHPLVQDILWATLHHVYEPLFMHWPAKRLREKALQSVMQHIHYEDENTRYICIGPVNKVLNM LCCWAEDPHSEAFKLHIPRIYDYLWIAEDGMKMQGYNGSQLWDTAFAVQAIISTELAEEYETTLRKAHKYIKDSQVLEDCPGDLQSWYRHISKGAWPFSTADHGWPISDCTAEGLKAVLLLSKLPSEIVGKSIDEQQLYNAVNVILSLQNTDGGFATYELTRSYRWLELMNPAETFGDIVIDYPYVECSSAA IQALAAFKKLYPGHRRDEIDNCIAEAADFIESIQATDGSWYGSWGVCFTYGGWFGIRGLVAAGRRYNNCSSLRKACDFLLSKELAAGGWGESYLSCQNKVYTNIKDDRPHIVNTGWAMLSLIDAGQSERDPPLHRAARVLINSQMEDGDFPQEEIMGVFNKNCMISYSAYRNIFPIWALGEYRSRVLKPLK Zostera marina (ZmCAS_mut) MAWKLKVAEGRDARLRTINGHVGRQIWEFDPDLGTDNERAEVEAVREKFRNNRFEKHSSDLLMRLQLAKENPVSSYLTQVKLEENEDITEEAAVTMTLRRALNFHSSIQSFDGHWAGDLGGPMFLMPGLVISLYITGVLNTVLSEHQREMCRYLYNHQN EDGGWGLHIEGPSTVFGSTLTYITLRLLGENVEDGDGAMEKGRKWILDHGGATYITSWGKMWLSVLGVFDWSGNNPLPPEMWLLPYFLPVHPGRMWCHCRMVYLPMSYLYGKRFVGKITPLVLSLRNEIYTVSYNQIDWNKARNLCAKEDLYYPHPMVQD ILWATLHKFVEPILMHWPGTLLREKALNTTMQHIHYEDESTRYICIGPVNKVLNMLCCWVDDPDSEAFKLHPRISDYLWIAEDGMKCQGYNGSQLWDTAFAVQAYIATNLSDEFGPVLTKAHEYIKNSQVPDDCSGDLSFWYRHISKGAWPFSTGDHGW PISDCTAEGLKASLLLSRISPEVVGKPLNAKRFYDAVNVILSLMNSDGFSFATYELTRSYTWLEMINPAETFGDIVIDYPYVECTSAAIQSLVAFTKLYPGHRREEIDECITKAAKFIESIQKKDGSWYGSWAVCFTYGLWFGIKGLIAAGKTYKNSSAIR KACEFLLSKQLASGGWGESYLSCQDKVYTNLEGNRAHAVNTGWAMLSLIDAGQAERDPSPLHRAARVLINSQMGNGDFPQEEIMGVFNRNCMISYSAYRNIFPIWALGEYRCKVLASKGHE Artemisia annua (AaCASmut) MAWKLKIAEGGDPWLRTTNDHIGRQIWEFDPTLGSVEELAEIEKLRKTFRDNRFEKKHSADLLMRSQFAKENSVSVFPPKVNIKDVEDITEDKVTNVLRRAIGFHSTLQADDGHWPGDLGGPMFLLPGLVITLSITGALNAVLSKEHKREMCRYLYNHQNIDGGWGLHIEGHSTMFGSALNYVTLRLLG EGANDGEGAMEKGRKWILDHGGATAITSWGKFWLSVLGVFEWPGNNPLPPEMWLLPYFLPVHPGRMWCHCRMVYLPMSYLYGKRFVGPITSTVLALRKELFTVPYHDIDWNEARNLCAKEDLYYPHPLIQDVLWATLDKFVEPVLMSWPGKKLREKALRTAMEHIHYEDENTRYICIGPVNKVLNMLCCW VEDPNSEAAFKLHLPRIQDYLWIAEDGMKMQGYNGSQLWDAAFTVQAIMSTNLIEEFGPTLKKGHIFIKKSQVLDNCYGDLDYWYRHISKGAWPFSTADHGWPISDCTAEGLKAALLLSKLPSEIVDEPLDAKRFYDAVNVILSLMNADGSFATYELTRSYSWLELINPAETFGDIVIDYPYVECTSAAIQ ALVAFKRLYPGHRRDEVQGCIDKAAAFLEKIQEADGSWYGSWAVCFTYGTWFGVKGLVAAGKNYSNCSSIRKACNFLLSKQLASGGWGESYLSCVDKVYTNLEGNRSHVVNTGWAMLALIDAEQAKRDPTPLHRAARVLINSQMENGEFPQQEIMGVFNRNCMITYAAYRNIFPIWALGEYRCRVLKVET Citrullus colocynthis (CcCDS2) (SEQ ID NO: 220) MAWRLKVGAESVGEKEEKWLCSISNHLGRQVWEFCAHQPTASPNHLQQIDNARNHFRNRFHRKQSSDLFLAIQNEKEIANVTKGGGIKVKEEEDVRKETVKNTVERALSFYSAIQTNDGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVA LRLLGEDADGGEGGAMTKARSWILDRGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYCLPFHPGRMWCRMVYLPMSYLYGKRFVGPITPIVLSRKELYTIPYHEIDWNRSRNTCAKEDLYYPHPKMQDILWGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYICLGPVNKVLN MLCCWVEDPYSDAFKFHLQRVPDYLWVAEDGMRMQGYNGSQLWDTAFSVQAIISTKLIDSFGTTLKAHDFVKDSQIQQDCPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSKIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTS ATMEALTLFKKLHPGHRTKEIDIAVARAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE エポシブドヒドロラーズ Siraitia grosvenorii EPH1 (SgEPH1) MEKIEHSTIATNGINMHVASAGSGPAVLFLHGFPELWYSWRHQLLYLSSLGYRAIAPDLRGFGDTDAPPSPSSYTAHHIV GDLVGLLDQLGVDQVFLVGDWGAMMAWYFCLFRPDRVKALVNLSVHFTPRNPAISPLDGFRLMLGDDFYVCKFQEPGVAEADFGSVDTATMFKKFLTMRDPRPPIPNGFRSLATPEALPSWLTEEDIDYFAAKFAKTGFTGGFNYYRAIDLTWELTAPWSGSEIKVPTKFIVGDLDLVYHFPGVKEYIHGGGFKKDVPLEEVVVMEGAAHFINQEKADEINSLIYDFIKQF Siraitia grosvenorii EPH2 (SgEPH2) MEKIEHTTISTNGINMHVASIGSGPAVLFLHGFPELWYSWRHQLLFLSSMGYRAIAPDLRGFGDTDAPPSPSSYTAHHIVGDLVGLLDQLGIDQVFLVGHDWGAMMAWYFCLFRPDRVKALVNLSVHFLRRHPSIKFVDGFRALLGDDFYFCQFQEPG VAEADFGSVDVATMLKKFLTMRDPRPPMPIKEKGFRALETPDPPLPAWLTEEDIDYFAGKFRKTGFTGGFNYYRAFNLTWELTAPWSGSEIKVAAKFIVGDLDLVYHFPGAKEYIHGGGFKKDVPLLEEVVVVDGAAHFINQERPAEISSLIYDFIKKF Siraitia grosvenorii EPH3 (SgEPH3) MDQIEHITINTNGIKMHIASVGTGPVVLLLHGFPELWYSWRHQLLYLSSVGYRAIAPDLRGYGDTDSPASPTSYTALHIVGDLVGALDELGIEKVFLVGHDWGAIIAWYFCLFRPDRIKALVNLSVQFIPRNPAIPFIEGFRTAFGDDFYMCRFQVPG EAEEDFASIDTAQLFKTSLCNRSSAPPCLPKEIGFRAIPPPENPLSWLTEEDINYYAAAKFKQTGFTGALNYYRAFDLTWELTAPWTGAQIQPVVKFIVGDSDLTYHFPGAKEYIHNGGFKKDVPLLEEVVVVKDACHFINQERPQEINAHIHDFINKF Momordica charantia (SEQ ID NO: 59) MEKIEHSTIAANGITIHVAVAVGSGPAVLLLHGFPELWYSWRHQLLFLASKGYRAIAPDLRGFGDSDAPPSPSSYTPLHIVGDLVALLDHLGIDLVFLVGHDWGAMMAWHFCLLRPDRVKALVNLSVHFMPRNPAMSPLDGMRLLLGDDFYVCRFQEP GAAEADFGSVDTATMMKKFLTMRDPRPPIIPNGFRSLETPQALPPWLTEEDIDYFAAKFAKTGFTGGFNYYRAIGRTWELTAPWTGSKIKVPAKFIVGDLDMVYHLPDAKEYIHGGGFKEDVPLLEEVVVIEGAAHFINQEKPDEISSLIYDFIKKF Cucurbita moschata (SEQ ID NO: 60) MEKIEHSTIATNGINMHVASIGSGPPVLFLHGFPELWYSWRHQLLFLASKGFRAIAPDLRGFGDSDVPPSPSSYTPFHIIGDLIGLDHLGIEQVFLVGHDWGAMMAWYFCLFRPDRVKALVNLSVHYNPRNPAISPLSRTRQFLGDDFYICKFQTP GVAEADFGSVDTATMMKKFLTIRDPSPPIIPNGFKTLKTPETLPSWLTEEDIDYFASKFTKTGFTGGFNYYRAIEQTWELTGPWSGAKIKVPTKYVVGDVDMVYHLPGAKQYIHGGGFKKDVPLLEEVVVMEGAAHFINQEKADEISAHIYDFIIKF Cucurbita maxima (SEQ ID NO: 61) MENIEHTIVPTNGINMHIASIGSGPAVLFLHGFPELWYSWRHQLLFLASNGFRAIAPDLRGFGDTDVPPSPSSYTAHHIVGDLIGLLDHLGIDRVFLVGHDWGAMMAWYFCLFRPDRVRALVNLSVHYLHRHPSIKFVDGFRAFLGDDFYFCQFQEPGVAEADFGSVDTATMLKKFLTMRDPRPPMIPKEKGFRALETPD PLPSWLTEEDVDYFASKFSKTGFTGGFNYYRAFDLSWELTAPWSGSQVKVPAKFIVGDLDLVYHFPGAKEYIHGGRFKEDVPFLEEVVVIEGAAHFINQERADEISSLIYEFINKF Prunus persica (SEQ ID NO: 62) MEKIEHTTVSTNGINMHIASIGTGPVVLFLHGFPELWYSWRHQLLSLSSLGYRCIAPDLRGFGDTDAPPSPASYSALHIVGDLIGLLDHLGIDQVFLVGHDWGAVIAWWFCLFRPDRVKALVNMSVAFSPRNPKRKPVDGFRALFGDDYYICRFQEPG EIEKEFAGYDTTSIMKKFLTGRSPKPPCLPKELGLRAWKTPETLPPWLSEEDLNYFASKFSKTGFVGGLNYYRALNLTWELTGPWTGLQVKVPVKFIVGDLDITYHIPGVKNYIHNGGFKRDVPFLQEVVVIEDGAHFINQERPDEISRHVYDFIQKF Morus notabilis (SEQ ID NO: 63) MEKIEHSTVHTNGINMHVASVGTGPAILFLHGFPELWYSWRHQMISLSSLGYRCIAPDLRGYGDTDAPPSPTSYTSLHIVGDLVGLIDHLVIEKLFLVGHDWGAMIAWYFCLFRPDRIKALVNLSVPFFPRNPKINFVDGFRAELGDDFYICRFQEPG ESEADFSSDTVAVFRRILANRDPKPPLIPKEIGFRGVYEDPVALPSWLTEDDINHFANKFNETGFTGGLNYYRALNLTWELTAAWTGARVQVPTKFIMGDLDLVYYFPGMKEYILNGGFKRDVPLLQELVIIEGAAHFINQEKPDEISSHIHHFIQKF Ricinus communis (SEQ ID NO: 64) MEKIEHTTVATNGINMHVAAIGTGPEILFLHGFPELWYSWRHQLLSLSSRGYRCIAPDLRGYGDTDAPESLTGYTALHIVGDLIGLLDSMGIEQVFLVGHDWGAMMAWYLCMFRPDRIKALVNTSVAYMSRNPQLKSLELFRTVYGDDYYVCRFQEPG GAEEDFAQVDTAKLIRSVFTSRDPNPPIVPKEIGFRSLPDPPSLPSWLSEEDVNYYADKFNKKGFTGGLNYYRNIDQNWELTAPWDGLQIKVPVKFVIGDLDLTYHFPGIKDYIHNGGFKQVVPLLQEVVVMEGVAHFINQEKPEEISEHIYDFIKKF Citrus unshiu (SEQ ID NO: 65) MEKIEHTTVGTNGINMHVASIGTGPVVLFIHGFPELWYSWRNQLLYLSSRGYRAIAPDLRGYGDTDAPPSVTSYTALHLVGDLIGLLDKLGIHQVFLVGHDWGALIAWYFCLFRPDRVKALVNMSVPFPPRNPAVRPLNNFRAVYGDDYYICRFQEPG EIEEEFAQIDTARLMKKFLCLRIAKPLCIPKDTGLSTVPDPSALPSWLSEEDVNYYASKFNQKGFTGPVNYYRCSDLNWELMAPWTGVQLEVPVKFIVGDQDLVYNNKGMKEYIHNGGFKKYVPYLQEVVVMEGVAHFINQEKAEEVGAHIYEFIKKF Hevea brasiliensis (SEQ ID NO: 66) MEKIEHITVFTNGINMHIASIGTGPEILFLHGFPELWYSWRHQLLSLSSLGYRCIAPDLRGYGDTDAPQSVNQYTVLHIVGDLVGLLDSLGIQQVFLVGHDWGAFIAWYFCIFRPDRIKALVNTSVAFMPRNPQVKPLDGLRSMFGDDYYICQFQKPG KAEEDFAQVNTAKLIKLLFTSRDPRPPHFLKEVGLKALQDPPSQQSWLTEEDVNFYAAKFNQKGFRGGLNYYQNINMNWELAAAWTGVQIKVPVKFIIGDLDLTYHFPGIKEYIHNGGFKKDVPLLQDVVVMEGVAHFLNQEKPEEVSKHIYDFIKKF Handroanthus impetiginosus (SEQ ID NO: 67) MDKIQHKIIQTNGINIHVAEIGDGPAVLFLHGFPELWYSW RHQMLFLSSRGYRAIAPDLRGYGDSDAPPCATSYTAFHIIGDLVGLLDAMGLDRVFLVGHDWGAVMAWYFCLLRPDRIKALVNLSVVFQPRNPKRKPVESMRAKLGDDYYICRFQEPGEAEEEFARVDTARLIKKLLT TRNPAPPRLPKEVGFGCLPHKPITMPSWLSEEDVQYYAAKFNQKGFTGGLNYYRAMDLSWELAAPWTGVQIKVPVKFIVGDLDITYNTPGVKEYIHKGRFKQHVPFLQELVILEGVAHFLNQEKPDEINQHIYDFIHKF Camelina sativa (SEQ ID NO: 68) MEKIEHTTVSTNGINMHVASIGSGPVILFLHGFPDLWYSWRHQLLSFAALGYRAIAPDLRGYGDSDAPPSPESYTILHIVGDLVGLLDSLGVDRVFLVGHDWGAIVAWWLCMIRPDRVKALVNTSVVFNPRNPSVKPVDKFRDLFGDDYYVCRFQETGEIEE DFAQVDTKKLITRFFVSRNRPPCIPKSVGFRGLPDPPSLPAWLTEQDVSFYGDKFSQKGFTGGLNYYRAMNLSWELTAPWAGLQIKVPVKFIVGDLDITYNIPGTKEYIHGGGLKKHVPFLQEVVVMEGVGHFLQQEKPDEVTDHIYGFFEKFRTRETSSL Coffea canephora MDKIQHRQVPVNGINLHVAEIGDGPAILFLHGFPELWYSWRHQLLLSLSAKGYRALAPDLRGYGDSDAPPSPSNYTALHIVGDLVGLLDSLGLDRVFLVGHDWGAVMAWYFCLLRPDRIKALVNMSVVFTPRNPKRKPLEAMRARFGDDYYICRFQEPG EAEEEFARVDTARIIKKFLTSRRPGPLCVPKEVGFGGSPHNPIQLPSWLSEDDVNYFASKFSQKGFTGGLNYYRAMDLNWELTAPWTGLQIKVPVKFIVGDLDVTFTTPGVKEYIQKGGFKRDVPFLQELVVMEGVAHFVNQEKPEEVSAHIYDFIQKF Punica granatum MEKIQHTTVRTNGINMHVATAGSGPDSILFVHGFPELWYWTWRHQMVSLAALGYRTIAPDLRGYGDTDAPPSHESYTAHFIVGDLVGLLDSMGIEKVFLVGHDWGAAIAWYFCLFRPDRIKALVNMSVVFHPRNPNRKPVDGLRAILGDDYYICRFQAPG EIEEDFARADTANIIKFFLVSRNPRPPQIPKEGFSCLANSRQMDLPSWLSEEDINYYASKFSEKGFTGGLNYYRVMNLNWELTAPFTGLQIKVPAKFMVGDLDITYNTPGTGKEFIHNGGLKKHVPFLQEVVVMEGVAHFINQEKPEEVTAHIYDFIKKF Arabidopsis lyrata subsp. lyrata (SEQ ID NO: 71)MEKIEHTTVSTNGINMHVASIGSGPVILFLHGFPDLWYSWRHQLLSFAALGYRAIAPDLRGYGDSDAPPSRESYTILHIVGDLVGLLNSLGVDRVFLVGHDWGAIVAWWLCMIRPDRVNALVNTSVVFNPRNPSVKPVDAFRALFGDDYYICRFQEPGEIEEDFAQVDTKKLITRFFISRNPRPPCIPKSVGFRGLPDPPSLPAWLTEEDVSFYGDKFSQKGFTGGLNYYRALNLSWELTAPWAGLQIKVPVKFIVGDLDITYNIPGTKEYIHEGGLKKHVPFLQEVVVLEGVGHFLHQEKPDEITDHIYGFFKKFRTRETASL Rhinolophus sinicus MDKIEHTTVSTNGINMHVASIGSGPVILFLHGFPDLWYSWRHQLLSFAGGLGYRAIAPDLRGYGDSDSPPSHESYTILHIVGDLVGLLDSLGVDRVFLVGHDWGAVVAWWLCMIRPDRVNALVNTSVVFNPRNPSVKPVDAFKALFGEDYYVCRFQEPGEI EEDFAQVDTKKLINRFFTSRNPRPCIPKTLGFRGLPDPPALPAWLTEQDVSFYADKFSQKGFTGGLNYYRAMNLSWELTAPWAGLQIKVPVKFIVGDLDITYNIPGTKEYIHEGGLKKHVPFLQEVVVMEGVGHFLHQEKPDEVTDHIYGFFKKF Gossypium raimondii (GrEPH) (SEQ ID NO: 184) MAEKIEHTTVTTNGIKMHVASIGSGPILFLHGFPELWYTWRHQLLSLSSLGYRCVAPDLRGYGDSDAPPSPESYTVFHIVGDLVGLLDALGVDKVFLVGHDWGAMIAWNFCLFRPDRIKALVNLSIPYHPRNPKVKTVDGYRALFGDDFYICRFQVP GEAEAHFAQMDTAKVMKKFLTTRDPNPPCIPRETGLKALPDPPALPSWLSEDEINYFATKFSQKGFTGGLNYYRAMNLNWELMAPWTGLQIQVPVKFIVGDLDITYHIPGVKEYLQNGGFKKNVPFLQELVVMEGVAHFINQEKPQEISMHIYDFIKKF Gossypium hirsutum (GhEPH) (SEQ ID NO: 185) MAEKIEHTTVTTNGIKMHVASIGSGPILFLHGFPELWYTWRRQLLSLSSLGYRCVAPDLRGYGDSDAPPSPESYTVFHVVGDLVGLLDALGVDKVFLVGHDWGAMIAWNFCLFRPDRIKALVNLSVPYHPRNPKVKTVDGYRALFGDDFYICRFQVP GEAEAHFAQMDTAKVLKKFLTTRDPNPPCIPKETGLKALPDPPALPSWLSEDEINYFATKFNQKGFTGGLNYYRAMNLNWELMAPWTGLQIQVPVKFIVGDLDITYHIPGVKEYLQNGGFKKNVPFLQELVVMEGVAHFINQEKPQEISMHIYDFVKKF Siraitia grosnevorii (SgEPH4) (SEQ ID NO: 186) MAENIEHTTVQTNGIKMHVAAIGTGPPVLLLHGFPELWYSWRHQLLYLSSAGYRAIAPDLRGYGDTDAPPSPSSYTALHIVGDLVGLLDVLGIEKVFLIGHDWGAIIAWYFCLFRPDRIKALVNLSVQFFPRNPTTPFVKGFRAVLGDQFYMVRFQEP GKAEEEFASVDIREFFKNVLSNRDPQAPYLPNEVKFEGVPPPALAPWLTPEDIDVYADKFAETGFTGGLNYYRAFDRTWELTAPWTGARIGVPVKFIVGDLDLTYHFPGAQKYIHGEGFKKAVPGLEEVVVMEDTSHFINQERPHEINSHIHDFFSKFC Cucumis melo (CmEPH1) (SEQ ID NO: 187) MADKIQHSTISTNGINIHFASIGSGPVVLFLHGFPELWYSWRHQLLFLASKGFRAIAPDLRGFGDSDAPPSPSSYTPHHIVGDLIGLLDHLGIDQVFLVGHDWGAMMAWYFCLFRPDRVKALVNLSVHYTPRNPAGSPLAVTRRYLGDDFYICKFQEP GVAEADFGSVDTATMMKKFLTMRDPRPAIIPNGFKTLLETPEILPSWLTEEDIEYFASKFSKTGFTGGFNYYRALDITWELTGPWSRAQIKVPTKFIVGDLDLVYNFPGAKEYIHGGGFKKDVPLLEDVVVVIEGAAHFINQEKPDEISSLIYDFITKF Cucumis melo (CmEPH2) (SEQ ID NO: 188) MAEKIEHTTIPTNGINMHVASIGSGPAVLFLHGFPQLWYSWRHQLLFLASKFRALAPDLRGFGDTAPPSPSSYTFLHIIGDLIGLLDHLGLEKVFLVGHDWGAMIAWYFCLFRPDRVKALVNLSVYYIKRHPSISFVDGFRAVAGDNFYICQFQEAGVAE ADFGRVDTATMMKKFMGMRDPEAPLIFTKEKGFSSMETDPDPLPCWLTEEDIDFFATKFSKTGFTGFNYYRALNLSWELTAAWNGSKIEVPVKFIVGDLDLVYHFPGAKQYIHGGEFKKDVPFLEEVVVIKDAAHFIHQEKPHQINSLIYHFINKFSTSTSPPA Trema orientale (ToEPH) MAEKIEHTTINTNGVNLHVASIGTGPAVLFLHGFPELWYSWRHQMLALSSLGYRAIAPDLRGYGDSDAPPSPESYSSLHIVGDLVGLIDQLGIDQIFLVGHDWGAVIAWQFCLFRPDRVKALVNMSVPFRPRHPTRKPIETFRALFGDDYYVCRFQAP GEVEEDFASDDTANLLKKFYGGRNPRPCVPKEIGFKGLKAPELPSWLSEEDLNYFAEKFNQRGFTGGLNYYRALDLTWELTAAWTGVQVKVPTKFIVGDLDITYHIPGAKEYINEGGLKKDVPYLQEVVVMEGVAHFVNQEKAEEVSAHIHDFIKKF Arachis hypogaea (AhEPH) MAEKIEHTWVNTNGIKMHVASIGSGPAVLFLHGFPELWYSWRHQLLSLSAQGYRCIAPDLRGYGDTDAPPSPSSYSALHIVSDLVGLLDALRIDQVFLVGHDWGAAMAWYFCLFRPDRIKALVNMSVVFRPRNPKWKPLQSLRAMLGDDYYICRFQKPG EAEEEFARAGTSRIIKTFLVSRDPRPPCVPKEIGFGGSPNLQLALPSWLTEEDVNYYASKFDQKGFTGGLNYYRAIDLTWELTAPWTGVQIKVPVKFIVGDLDVTYNTPGVKEYIHGGGFKKEVPFLQELVVMEGVAHFINQERPDEISAHIHDFIKKF Mycobacterium tuberculosis (MtEPH) (SEQ ID NO: 212) MASQVHRILNCRGTRIHAVADSPPDQQGPLVVLLHGFPESWYSWRHQIPALAGAGYRVVAIDQRGYGRSSKYRVQKAYRIKELVGDVVGVLDSYGAEQAFVVGHDWGAPVAWTFAWLHPDRCAGVVGISVPFAGRGVIGLPGSPFGERRPSDYHLELAGPGRVWYQDYFAVQDGIITE IEEDLRGWLLGLTYTVSGEGMMAATKAAVDAGVDLESMDPIDVIRAGPLCMAEGARLKDAFVYPETMPAWFTEADLDFYTGEFERSGFGGPLSFYHNIDNDWHDLADQQGKPLTPALFIGGQYDVGTIWGAQAIERAHEVMPNYRGTHMIADVGHWIQQEAPEETNRLLLDFLGGLRP Cytochrome P450 Siraitia grosvenorii CYP87D18 (SEQ ID NO: 73) MWTVVLGLATLFVAYYIHWINKWRDSKFNGVLPPGTMGLPLIGETIQLSRPSDSLDVHPFIQKKVERYGPIFKTCLAGRPVVSADAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGLIHKYIRSITLNHFGAEALRERFLPFIEASSMEALHSWSTQPSVEVKNASALMVFRTSVNKMFGEDAKKLSGNIPGKFTKLLGGFLSLPLNFPGTTYHKCLKDMKEIQKKLR EVVDDRLANVGPDFEDFLGQAFKDKESKFISEEFIIQLLFSISFASFESISTTLTLILKLLDEHPEVVKELEVEHEAIRKARADPDGPITWEEYKSMTFTLQVINETLRLGSVTPALLRKTVKDLQVKGKIIPEGWTIMLVTASRHRDPKVYKDPHIFNPWRWKDLDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILCTKYRWTKLGGGTIARAHILSFEDGLHVKFTPKE Cucumis melo MWTILLGLATLAIAYYIHWVNKWKDSKFNGVLPPGTMGLPLIGETIQLSRPSDSLDVHPFIQSKVKRYGPIFKTCLAGRPVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGLIHKYIRSITLNHFGAESLRERFLPRIEESARETLHYWSTQPSVEVKESAAAMVFRTSIVKMFSEDSSKLLTAGLTKKFTGLGGFLTLPLNVPGTTYHKCIKDMKEIQKKLKDIL EERLAKGVSIDEDFLGQAIKDKESQQFISEEFIIQLLFSISFASFESISTTLTLILNFLADHPDVAKELEAEHEAIRKARADPDGPITWEEYKSMNFTLNVICETLRLGSVTPALLRKTTKEIQIKGYTIPEGWTVMLVTASRHDPEVYKDPDTFNPRWKELDSITIQRNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWRKLKGGKIARAHILRFEDGLYVNFTPKE Cucurbita maxima MWTIVVGLATLAVAYYIHWINKWKDSKFNGVLPPGTMGLPLIGETLQLSRPSDSLDVHPFIKKKVKRYGSIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLCYWATQPSVEVKDSAAVMVFRTSMVKMVSKDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYNKCMKDMKEIQKKLREILEGRLASGAGSDEDFLGQAVKDKGSQKFISDDFIIQLLFSISFASFESISTTLTLILNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARAHILFEDGLHMKFTPKE Cucumis sativus MWTILLGLATLAIAYYIHWVNKWKDSKFNGVLPPGTMGLPLIGETIQLSRPSDSLDVHPFIQRKVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGLIHKYIRSITLNHFGAESLRERFLPRIEESARETLHYWSTQTSVEVKESAAAMVFRTSIVKMFSEDSSKLLTEGLTKKFTGGFLTLPLNLPGTTYHKCIKDMKQIQKKLK DILEERLAKGVKIDEDFFLGQAIKDKESQQFISEEFIIQLLFSISFASFESISTTLTLILNFLADHPDVVKELEAEHEAIRKARADPDGPITWEEYKSMNFTLNVICETLLRLGSVTPALLRKTTKEIQIKGYTIPEGWTVMLVTASRRHDPEVYKDPTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWRKLKGGKIARAHILRFEDGLYVNFTPKE Cucurbita moschata MWAIVVGLATLAVAYYIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGSIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTSMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYNKCMKDMKEIQKKLR EILEGRLASGAGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARAHILSFEDGLHVKFTPKE Prunus avium (SEQ ID NO: 78) MWTLVGLSLVALLVIYFTHWIIKWRNPKCNGVLPPGSMGLPLIGETLNLIIPSYSLDLHPFIKKRLQRYGPIFRTSLAGRPVVVTADPEFNNYIFQQEGRMVELWYLDTFSKIFVHEGDSKTNAIGMVHKYVRSIFLNHFGAERLKEKLLPQIEEFVNKS LCAWSSKASVEVKHAGSVMVFNFSAKQMISYDAEKSSDDLSEKYTKIIDGLMSFPLNIPGTAYYNCSKHQKNVTTMLRDMLKERRISPETRRGDFLDQLSIDMEKEKFLSEDFSVQLVFGGLFATFESISAVIALAFSLLADHPSVVEELTAEHEAIL KNRENPNSSITWDEYKSMTFTLQVINEILRLGNVAPGLLRRALKDIPVKGFTIPEGWTIMVVTSALQLSPNTFEDPLEFNPWRWKDLDSYAVSKNFMPFGGGMRQCAGAEYSRVFLATFLHVLVTKYRWTTIKAARIARNPILGFGDGIHIKFEEKKT Populus trichocarpa (SEQ ID NO: 79) MWAIGLVVVALVVIYYTHMIFKWRSPKIEGVLPPGSMGWPLIGETLQFISPGKSLDLHPFVKKRMEKYGPIFKTSLVGRPIIVSTDYEMNKYILQHEGTLVELWYLDSFAKFFALEGETRVNAIGTVHKYLRSITLNHFGVESLKESLLPKIEDMLHTNLAKWASQGPVDVKQVISVMVFNFTANKIFGYDAENSKEKLSENYTKILNSFISLPLNIPGTSFHKCMQDREKMLKMLK DTLMERLNDPSKRRGDFLDQAIDDMKTEKFLTEDFIPQLMFGILFASFESMSTTLTLTFKFLTENPRVVEELRAEHEAIVKKRENPNSRLTWEEYRSMTFTQMVVNETLRISNIPPGLFRKALKDFQVKGYTVPAGWTVMLVTPATQLNPDTFKDPVTFNPWRWQELDQVTISKNFMPFGGGTRQCAGAEYSKLVLSTFLHILVTNYSFTKIRGGDVSRTPIISFGDGIHIKFTARA Prunus persica MWTLVGLSLVGLLVIYFTHWIIKWRNPKCNGVLPPGSMGLPFIGETNLNLIIPSYSLDLHPFIKKRLQRYGPIFRTSLAGRQVVVTADPEFNNYLFQQEGRMVELWYLDTFSKIFVHEGESKTNAVGMVHKYVRSIFLNHFGAERLKEKLLPQIEEFVNKSLCAWSSKASVEVKHAGSVMVFNFSAKQMISYDAEKSSDDLSEKYTKIIDGLMSFPLNIPGTAYNCLKHQKNVTTMLR DMLKERQISPETRRGDFLDQISIDMEKEFLSEDFSVQLVFGGLFATFESISAVLALAFSLLAEHPSVVEELTAEHEAILKNRENLNSSLTWDEYKSMTFTLQVINEILRLGNVAPGLLRRALKDIPVKGFTIPEGWTIMVVTSALQLSPNTFEDPLEFNPWRWKDLDSYAVSKNFMPFGGGMRQCAGAEYSRVFLATFLHVLVTKYRWTTIKAARNPILGFGDGIHIKFEEKKT Populus euphratica MWTFVLCVVAVLVVYYTHWINKWRNPTCNGVLPPGSMGLPIIGETLELIIPSYSLDLHPFIKKRIQRYGPIFRTNILGRPAVVSADPEINSYIFQNEGKLVEMWYMDTFSKLFAQSGESRTNAFGIIHKYARSLTLTHFGSESLKERLLPQVENIVSKSLQMWSSDASVDVKPAVSIMVCDFTAKQLFGYDAENSSDKISEKFTKVIDAFMSLPLNIPGTTYHKCLKDKDSTLSILRNTLKERMNSPAESRGGDFLDQIIADMDKEKFLETEDFTVNLIFGILFASFESISAALTLSLKLIGDHPSVLEELTVEHEAILKNRENPDSPLTWAEYNSMTFSLQVINETLRLGNVAPGLLRRALQDMQVKGYTIPAGWVIMVVNSALHLNPATFKDPLEFNPWRWKDFDSYAVSKNLMPFGGGRRQCAGSEFTKLFMAIFLHKLVTKYRWNIIKQGNIGRNPILGGFGDGIHISFSPKDI Juglans regia MWKVGLCVVGVVIVVWFTRWINKWRNPKCNGILPPGSMGPPLIGESLQLIIPSYSLDLHPFIKKRVQRYGPIFRTSVVGQP MVVSTDVEFNHYLAKQEGRLVHFWYLDSFAEIFNLEDENAISAVGLIHKYGRSIVLNHFGTDSLKKTLLSQIEEIVNKTLQTWSSLPSVEVKHAASVMAFDLTAKQCFGYDVENSAVKMSEKFLYTLDSLISFPFNIPGTVYHKCLKDKKEVLNMLRNIVKERMNSPEKYRGDFLDQITADMNKESFLTQDFIVYLLYGLLFASFESISASLSLTLKLLAEHPAVLQQLTAEHEAILKNRDNPNSSLTWDEYKSMTFTFQVINEALRLGNVAPGLLRRALKDIEFKGYTIPAGWTIMLANSAIQLNPNTYEDPLAFNPWRWQDLDPQIVSKNFMPFGGGIRQCAGAEYSKTFLATFLVTKYRWTKVKGGKMARNPILWFADGIHINFALKHN Pyrus x bretschneideri MWDVVGLSFVALLVIYLTYWITQWKNPKCNGVLPPGSMGLPLIGETNLNLLIPSYSLDLHPFIRKRLERYGPIFRTSLAGKPVLVSADPEFNNYVLKQEGRMVEFWYLDTFSKIFMQEGGNGTNQIGVIHKYARSIFLNHFGAECIKEKLLTQIEGSINKHLRAWSNQESVEVKKAGSMALNFCAEHMIGYDAETATENLGEIYHRVFQGLISFPNLNVPGTAYHNCLKIHKKATTMLR AMLRERRSSPEKRRGDFLDQIIDDLDQEKFLSEDFCIHLIFGGLFAIFESISTVLTLFFSLLADHPAVLQELTAEHEALLKNREDPNSALTWDEYKSMTFTLQVINETLRLVNTAPGLLRRALKDIPVKGYTIPAGWTILLVTPALHLTSNTFKDHLEFNPWRWKDLDSLVISKNFMPFGSGLRQCAGAEFSRAYLSTFLHVLVTKYRWTTIKGARISRRPMLTFGDGAHIKFSEKKN Morus notabilis MWNTICLSVVGLVVIWISNWIRRWRNPKCNGVLPPGSMGFPLIGETLPLIIPTYSLDHPFIKNRLQRYGSIFRTSIVGRPVVISADPEFNNFLFQQEGSLVELYYLDTFSKIFVHEGVSRTNEFGVVHKYIRSIFLNHFGAERLKEKLLPEIEQMVNKTLSAWSTQASVEVKHAASVLVLDFSAKQIISYDAKKSSSESLSLETYTRIIQGFMSFPLNIPGTAYNQCVKDQKKIIAMLRDMLKERRASPETNRGDFLDQISKDMDKEKFLSEDFVVQLIFGGLFATFESVSAVLALGFMLLSEHPSVLEEMIAEHETILKNREHPNSLLAWGEYKSMTFTLQVINETLRLGNVAPGLLRKALKDIRVKGFTIPKGWAIMMVTSALQLSPSTFKNPLEFNPWRWKDLDSLVISKNFMPFGRGMRQCAGAEYSRAFMATFFHVLLTKYRWTIKVGNVSRNPILRFGNGIHIKFSKKN Jatropha curcas (JcP450.1) MWIIGLCFASLLVIYCTHFFYKWRNPKCKGVLPPGSMGLPIIGETLQLIIPSYSLDHHPFIQKRIQRYGPIFRTNLVGRPVIVSADPEVNQYIFQQEGNSVEMWYLDAYAKIFQLDGESRLSAVGRVHKYIRSITLNNFGIENLKENLLPQIQDLVNQSLQKWSNKASVDVKQAASVMVFNLTAKQMFSYGVEKNSSEEMTEKFTGIFNSLPLNIPGTTYHKCLKDREAMLKML RDTLKQRLSSPDTHRGDFLDQAIDDMDTEKFLTGDCIPQLIFGILLAGFETTATTLTLAFKFLAEHPLVLEELTAEHEKILSKRENLESPLTWDEYKSMTFTHHVINETLRLANFLPGLLRKALKDIQVKNYTIPAGWTIMVVKSAMQLNPEIYKDPLAFNPWRWKDLDSYTVSKNFMPFGGGSRQCAGADYSKLFMTIFHLHVLVTKYRWRKIKGGDIARPILGFGDGLHIEVSAKN Hevea brasiliensis MLTVVLLLVGFFIIYTYWISKWRNPNCNGVLPPGSMPLIGETLQLLIPSYSLDLHPFIKKRIHRYGPIFRSNLAGRPVIVSADPEFNYILSQEGRSVEIWYLDTFSKLFRQQGESRTNVAGYVHKYLRGAFLSQIGSENLREKLLLHIQDMVNRTLCSWSNQESVEVKHSASLAVCDFTAKVLFGYDAEKSPDNLSETFTRFVEGLISFPLNIPRTAYRQCLQDRQKALSILK NVLTDRRNSVENYRGDVLDLLLNDMGKEKFLTEDFICLIMLGGLFASFESISTITTLLLKLFSAHPEVVQELEAEHEKILVSRHGSDSLSITWDEYKSMTFTHQVINETLRLGNVAPGLLRRAIKDVQFKGYTIPSGWTIMMVTSAQQVNPEVYKDPLVFNPWRWKDFDSITVSKNFTPFGGGTRQVCGAEYSRLTLSLFIHLLVTKYRWTKIKEGEIRRAPFSEPMLGFGDGIHFKKE Jatropha curcas (JcP450.2) MKRAIYICLARITKQGLSLIEMLMTELLFGAFFIIFLTYWINRWRNPKCNGVLPPGSMGLPLLLGETLQLLIPRYSLDLHPFIRKRIQRYGPIFRSNVAGRPIVFTADPELNHYIFIQERRLVEL WYMDTFSNLFVLDGESRPTGATGYIHKYMRGLFLTHFGAERLKDKLLHQIQELIHTTLQSWCKQPTIEVKHAASAVICDFSAKFLFGYEAEKSPFNMSERFAKFAESLVSFPLNIPGTAYHQSLE DREKVMKLLKNVLRERRNSTKKSEEDVLKQILDDMEKENFITDDFIIQILFGALFAISESIPMTIALLVKFLSAQPSVVEELTAEHEEILKNKKEKGLDSSITWEDYKSMTFTLQVINETLRIA NVAPGLLRRTLRDIHYKGYTIPAGWTIMVLTSSRHMNPEIYKDPVEFNPWRWKDLDSQTISKNFTPFGGGTRQCAGAEYSRAFISMFLHVLVTKYRWKNVKEGKICRGPILRIEDGIHIKLYEKH Chenopodium quinoa (SEQ ID NO: 88) MWPTMGLYVATIVAICFILLELKRRNSREKQVVLPPGSKGFPLIGETLQLLVPSYSLDLPSFIRTRIQRYGPIFKTRLVGRPVVMSADPGFNRYIVQQEGKSVEMWYLDTFSKLFAQDGEARTTAAGLVHKYLRNLTLSHFGSESLRVNLLPHLESLVRNTLGWSSKDTIDVKESALTMTIEFVAKQLFGYDSDKSKEKIGEKFGNISQGLFSLPLNIPGTTYHSCLKSQREVMDMMRTALKDRLTTPESYRGDFLDHALKDLSTEKFLSEEFILQIMFGLLFASSESTSMTLTLVLKLLSENPHVLKELEAEHERIIKNKESPDSPLTWAEVKSMTFTLQVINESLRLGINVSLGILRRTLKDIEINGYTIPAGWTIMLVTSACQYNSDIYKDPLTFNPWRWKEMQPDVIAKNFMPFGGGTRQCAGAEFAKVLMTIFLHNLVTNYRWEKIKGGEIVRTPILGFRNALRVKLTKKN Spinacia oleracea MVLLPGSKGFPFIGETLQLLLPSYSLDLPSFIRTRIQRYGPIFQTRLVGRPVVVSADPGFNRYIVQQEGKMVEMWYLDTFSKIFAQQGEGRTNAAGLVHKYLRNITFTHFGSQTLRDKLLPHLEILVRKTLHGWTSQESIDVKEAALTMTIEFVAKQLFGYDSDKSKERIGDKFANISQGLLSFPNLIPGTTYHSCLKSQ REVMDMMRKTLKERLASPDTCQGDFLDHALKDLNTDKFLTEDFILQIMFGLLFASSESTSITLTLILKFLSENPHVLEELEVEHERILKNRESPDSPLTWAEVKSMTFTLQVINESLRLGNVSLGLLRRTLKDIEINGYTIPAGWTIMLVTSACQYNSDVYKDPLTFNPWRWKEMQPDVIAKNFMPFGGGTRQCAGAEFA KVLMTIFLHVLVTTYRWEKIKGGEIIRTPILGFRNGLHVKLIKKARLS Manihot esculenta MEMWSVWLYIISLIIIATHWIYRWRNPKCNGKLPPGSMGIPFIGETIQFLIPSKSLDVPNFIKKRMNKYGPLFRTNLVGRPVIVSSDPDFNYYLLQREGKLVERWYMDSFSKLLHHDVTQIIIKHGSIHKYLRNLVLGHFGPEPLKDKLLPQLESAISQRLQDWSKQPSIEAKSASSAMIFDFTAKILFSYEPEKSGENIGEIFNSFLQGLMSIPLNIPGTAFHRCLKNQKRAIQMI TEILKERRSNPEIHKGDFLDQIVEDMKKDSFWTEEFAIYMMFGLLLASFETISTSLALAIIFLTDNPPVVQKLTEEHEAILKARENRDSGLSWKEYKSLSYTHQVVNESLRLASVAPGILRRAITDIQVDGYTIPKGWTIMVVPAAVQLNPNTFEDPLVFNPSRWEDMGAVAMAKNFIAFGGGSRSCAGAEFSRVLMSVFVHVFVTNYRWTKIGGDMVRSPALGFGNGFHIRVSEKQL Olea europaea var. sylvestris MAALDLSTVGYLIVGLLTVYITHWIYKWRNPKCNGVLPPGSMGLPLIGETIQLVIPNASLDLPPFIKKRMKRYGPIFRTNVAGRPVIITADPEFNHFLLRQDGKLVDTWSMDTFAEVFDQASQSSRKYTRHLTLNHFGVEALREKLLPQMEDMVRTTLSNWSSQESVEVKSASVTMAIDYAARQIYSGNLENAPLKISDLFRDLVDGLMSFPINIPGTAHHRCLQTHKKVREMMKDIVKTRLEEPERQYGDMLDHMIEDMKKESFLDEDFIVQLMFGLFFVTSDSISTTLALAFKLLAEHPLVLEELTAEHEAILKKREKSESHLTWNDYKSMTFTLQVINEVLRLGNIAPGFFRRALQDIPVNGYTIPSGWVIMIATAGLHLNSNQFEDPLKFNPWRWKVCKVSSVIAKCFMPFGSGMKQCAGAEYSRVLLATFIHVLTTKYRWAIVKGGKIVRSPIIRFPDGFHYKIIEKTN Cucurbita pepo subsp. pepo (Accession No. 171) MWAIVVGLATLAVAYYIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTE WLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTSMVKMVSEDSSKLLTGGLTKKFTGLLGGGFLTLPINVPGTTYNKCMKDMKEIQKKLR EILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPAL LRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARAHILSFEDGLHVKFTPKE Capsella rubella CYP705A38 (SEQ ID NO: 172) MATLMTIDLQNCFIFIILSLLCYYLLFKKQKGSRAGCVLPPSPPSLPIIGHLHLLLSNLTHKSLQNISTKFGSFLYLRVVNLPIVLVSSPSVAYEIYKTHDVNVSSRVATSLGDSLFLGSSGFITAPYGDYWKFMKKMVATKLLRPQAIEQSRGGRAEEL QMFYENLLDKAMKKESIEVSKEAMKLTNNIICRMSMGRSCSDENGEAERVRELLVKSTALTKKIFFANMFPRIPLFKKEIMGVSSEFDDLLERLLVEHEERVEEHENKDMMDLLLEAYRDENAEYKISRKQIKSLFVEIFLGGTTDTSAQTVQWILAELIN KPNILERIREEIDSVVGKSRLMKETDLPNLPYLQATVKEGLRMHPPSPLLVRTFQESCEVKGFYMPEKTMLVINVYALMRDPDTWEDPNEFKPERFLLSSRSRQEDEKEQGMMKYLPFGAGRRGCPGSNLAYLFVGIAVGVMVQCFDWKIKEDKVNMEETTAGMNLAMAHPFKCTPVVRNDPLTLNLENPSS Brassica rapa CYP705A37v2 (Accession No. 173) MIVDFQNCSIFILLCFFTFLCYSVFFFFKKTNDLGPSPPSLPIIGHLHHFLSGLPHKAFQKISTKYGPLLHLHIFSFPIVLVSSPTMAHEIFTTHDLNISSRNTPAIDESLLFGPSGFTVAPYGDYVKFIKKLLATKLLRPRAIEKSRGVRAEELKQFYLKVQDKALKKESIEIGKETMKFTNNMICRMSIGRSFSEENGEVETLRELIIKSFALSKQILFVNVLRRPLEMLGLMSLFKKDIMDVSRGFDELLERVLAEHEEKREEDQDMDMMDLLLEACRDENAEYKITRNQIKSLFVEIFLGGTDTSAHTTQWTMAELVNNPNILGRLRDEIDLVVGKERLIQETDLPNLPYLQAVVKEGLRLHPPAPLLVRMFDKKCVIKDFFKVPEKTTLVVNVYGVMRDPDSWEDPNEFKPERFLTSKQEEDKVLKYLPFAAGRRGCPATNVGYIFVGISIGMMVQCFDWSIKEKVSMEEVYAGMSLSMAHPPTCTPVSRLSL Siraitia grosvenorii (Accession No. 174) MDFFSAFLLLLLTVLILLQIRTRRRNLPPSPPSLPIIGHHLLLKRPIHRNFHKIAAEYGPIFSLRFGSRLAVIVSSLDIAEECFTKNDLIFANRPRLLISKHLGYNCTTMATSPYGDHWRNLRRLAAIEIFSTARLNSSLSIRKDEIQRLLLKLHSGSSGEFTKVELKTMFSELAFNALMRIVAGKRYYGDEVSDEEEAREFRGLMEEISLHGGASHWVDFMPILKWIGGGGFEKSLVRLTKRTDKFMQAL IEERRNKKVLERKNSLLDRLLELQASEPEYYTDQIIKGLVLVLLRAGTDTSAVTLNWAMAQLLNNPELLAKAKAELDTKIGQDRPVDEPDLPNLSYLQAIVSETLRLHPAAPMLLSHYSSDCTVAGYDIPRGTLLVNAWAIHRDPKLWDDPTSFRPERFLGAANELQSKKLIAFGLGRRSCPGDTMALRFVGLTLGLLIQCYQWKKCGDEKVDMGEGGITIHKAKPLEAMCKARPAMYKLLLNALDKI Camelina sativa MATMMIFDFQNCFIFIILCFVSLLCYTILFKKQESSRTGCVLPPSPPSLPIIGHHLLLLSSLTHKSLHNISSKGFGPFLYLRVVNLPIVLVSSASVAYEIYKTQDVNVSSRVATSLGDSFLGLGSSGFIT APYGDYWKFMKKMVATKLLRPQAIEQSRGGRAEELQGLYENLLDKAMKKESIEISKEAMKFTNNIICRMSMGRSCSDENGEAEIVRELLVKSTALTKKIFFANMFPRIPLFKKEIMGVSNQFDELLERL LVEERVEEHENKDMMDLLLEAFRDEHAEYKISRKQIKSLFVEIFLGGTDTSAQTVQWIMAELINKPSIIEKIREEIDSVVGKTRLIKETDLPKLPYLQVVVKEGLRMHPPSPLVVRTFQESCEVKGFYMPEKTMLVINVYALMRDPESWEDPNEFKPERFLPSSKSRQDEEKEQGLKYLPFGAGRRGCPGSNLAYLFVGLAVGVMVQCFDWKIKEDKVNMEETTAGMNLAMAHPFKCTPVVRIDPLTFNLKSPSP Raphanus sativus (sequence number 176) MAPMTIDFQTCFIFILLSFFSFFCYFFFFKKTNDLGPSPPSLPIIGHLHHFLSVLPHKAFQQISTKYGPLLHLRIFSFPI VLVSSATMAYEIFTTHDLNISSRNAPAIDESLVFGSSGFIVSPYGDYVKFIKKLLATKLLRPRAIEKSRGVRAEELKQFYLKLHDKALKKESIEIGNETMKFTNNMICGMSRSCSEENGEETVRGLINKSFALSRKILFVNVLRRPLEKLGLLSLFKKDILDVSNRFDELLERILLEHEEKPEEEQDMDMMDLLLEASRDENAEYKITRNQIKALFVEIFMGGTDTSAHTTQWTMAELVNNPNSLEKLRDEIDMVVGKSRLIQETDLPNLPYLQAVVKEGLRLHPPAPLLVRMFEKKCVIKDFFNVPEKTTLVVNLYGVMRDPDSWEDPNEFKPERFLTSKQEEEKTLKYLPFAAGRGCPATNVAYIFVGISIGMMVQCFDWSIKDKVSMEEVYAGMSLSMAHPPKFTPVSRLSL Cucumis sativus (CsCYP87D20) MAWTILLGLATLAIAYYIHWVNKWKDSKFNGVLPPGTMGLPLIGETIQLSRPSDSLDVHPFIQRKVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGLIHKYIRSITLNHFGAESLRERFLPRIEESARETLHYWSTQTSVEVKESAAAMVFRTSIVKMFSEDSSKLLTEGLTKKFTGLLGGFLTLPLNLPGTTYHKCIKDMKQIQKKLKDILEERLAKGVKIDEDFLGQAIKDKESQQFISEEFIIQLLFSISFASFESISTTLTLILNFLADHPDVVKELEAEHEAIRKARADPDGPITWEEYKSMNTLNVICETLRLGSVTPALLRKTTKEIQIKGYTIPEGWTVMLVTASRHDPEVYKDPDTFNPRWKELDSITIQKNFMPFGGLRHCAGAEYSKVYLCTFLHILFTKYRWRKLKGGKIARAHILRFEDGLYVNFTPKE Cucumis sativus (sohB_CsCYP87D20) MALLSEYGLFLAKIVTVVLAIAAIAAIHWVNKWKDSKFNGVLPPGTMGLPLIGETIQLSRPSDSLDVHPFIQRKVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGLIHKYIRSITLNHFGAESLRERFLPRIEESARETLHYWSTQTSVEVKESAAAMVFRTSIVKMFSEDSSKLLTEGLTKKFTGLGGFLTLPLNLPGTTYHKCIKDMKQIQKKLKDILEERLAKGVKIDEDFLGQAIKDKESQQFISEEFIIQLLFSISFASFESISTTLTLILNFLADHPDVVKELEAEHEAIRKARADPDGPITWEEYKSMNTLNVICETLRLGSVTPALLRKTTKEIQIKGYTIPEGWTVMLVTASRHDPEVYKDPDTFNPRWKELDSITIQKNFMPFGGLRHCAGAEYSKVYLCTFLHILFTKYRWRKLKGGKIARAHILRFEDGLYVNFTPKE Cucumis sativus (zipA_CsCYP87D20) MAQDLRLILIIVGAIAIIALLVHGFHWVNKWKDSKFNGVLPPGTMGLPLIGETIQLSRPSDSLDVHPFIQRKVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGLIHKYIRSITLNHFGAESLRERFLPRIEESARETLHYWSTQTSVEVKESAAAMVFRTSIVKMFSEDSSKLLTEGLTKKFTGLLGGFLTLPLNLPGTTYHKCIKDMKQIQKKLKDILEERLAKGVKIDEDFLGQAIKDKESQQFISEEFIIQLLFSISFASFESISTTLTLILNFLADHPDVVKELEAEHEAIRKARADPDGPITWEEYKSMNFTLNVICETLRLGSVT PALLRKTTKEIQIKGYTIPEGWTVMLVTASRHRDPEVYKDPTFNPWRWKELDSITIQKNFMPPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWRKLKGGKIARAHILRFEDGLYVNFTPKE Cucumis sativus (CsCYP87D20_mut) MAWTILLGLATLAIAYYIHWVNKWKDSKFNGVLPPGTMGLPLIGETIQFSRPSDSLDVHPFIQRKVKRYGPIFKTCIAGRPVVSTDAEFNHYIMLQEGRAVEMWYLDTFSKFLGLDTEWLKALGLIHKYIRSITLNHFGAESLRERFLPRIEESARETLHYWSTQTSVEVKESAAAMVFRTSIVKMFSEDSSKLLTEGLTKKFTGLGGFLTLPLNLPGTTYHKCIKDMKQIQKKLKDILEERLAKGVKIDEDFLGQAIKDKESQQFISEEFIIQLLFSISFASFASISTTLTLILNFLADHPDVVKELEAEHEAIRKARADPDGPITWEEYKSMNTLNVICETLRLGSVTPALLRKTTKEIQIKGYTIPEGWTVMLVTASRHDPEVYKDPDTFNPRWKELDSITIQKNFMPFGGLRHCAGAEYSKVYLCTFLHILFTKYRWRKLKGGKIARALILRFEDGLYVNFTPKE Cucumis sativus (sohB_CsCYP87D20_mut) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWVNKWKDSKFNGVLPPGTMGLPLIGETIQFSRPSDSLDVHPFIQRKVKRYGPIFKTCIAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTFSKFLGLDTEWLKALGLIHKYIRSITLNHFGAESLRERFLPRIEESARETLHYWSTQTSVEVKESAAAMVFRTSIVKMFSEDSSKLLTEGLTKKFTGLLGGFLTLPLNLPGTTYHKCIKDMKQIQKKLKDILEERLAKGVKIDEDFLGQAIKDKESQQFISEEFIIQLLFSISFASFASISTTLTLILNFLADHPDVVKELEAEHEAIRKARADPDGPITWEEYKSMNFTLNVICETLRLGSVTPALLRKTTKEIQIKGYTIPEGWTVMLVTASRHRDPEVYKDPDTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWRKLKGGKIARALILRFEDGLYVNFTPKE Cucurbita pepo subsp. pepo (sohB_CppCYP)(SEQ ID NO: 199) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTSMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYNKCMKDMKE IQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLTVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARAHILSFEDGLHVKFTPKE Cucurbita pepo subsp. pepo (17アルファ_CppCYP) MALLLAVFHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTSMVKMVSEDSSKLLTGGLTKKFTGLGGFLTLPINVPGTTYNKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTTLLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGLRHCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARAHILFEDGLHVKFTPKE Siraitia grosvenorii (CYP1798) MEMSSSVAATISIWMVVVCIVGVGWRVVNWVWLRPKKLEKRLREQGLAGNSYRLLFGDLKERAAMEEQANSKPINFSHDIGPRVFPSMYKTIQNYGKNSYMWLGPYPRVHIMDPQQLKTVFTLVYDIQKPNLN PLIKFLLDGIVTHEGEKWAKHRKIINPAFHLEKLKDMIPAFFHSCNEIVNEWERLISKEGSCELDVMPYLQNLAADAISRTAFGSSYEEGKMIFQLLKELTDLVVKVAFGVYIPGWRFLPTKSNNKMKEINNRKI KSLLLGIINKRQKAMEEGEAGQSDLLGILMESNSNEIQGEGNNKEDGMSIEDVIEECKVFYIGGQETTARLLIWTMILLSSHTEWQERARTEVLKVFGNKKPDFDGLSRLKVVTMILNEVLRLYPPASMLTRII QKETRVGKLTLPAGVILIMPIILIHRDHDLWGEDANEFKPERFSKGVSKAAKVQPAFFPFGWGPRICMGQNFAMIEAKMALSLILQRFSFELSSSYVHAPTVVFTTQPQHGAHIVLRKL Cytochrome P450 Reductase Stevia rebaudiana (SrCPR1) (SEQ ID NO: 92) MAQSDSVKVSPFDLVSAAMNGKAMEKLNASESEDPTTLPALKMLVENRELLTLFTTSFAVLIGCLVFLMWRRSSSKKLVQDPVPQVIVVKKKEKESEVDDGKKKVSIFYGTQTGTAEGFAKALVEEAKVRYEKTSFKVIDLDDYAADDDEYEEKLKKESLAFFFLATYGDGEPTDNAANFYKWFTEGDDKGEWLKKLQYGVFGLGNRQYEHFNKIAIVVDDKLTEMGAKRLVPVGLGDDDQCIEDDFTAWKELVWPELDQLLRDEDDTSVTTPYTAAVLEYRVVYHDKPADSYAEDQTHTNGHVVHDAQHPSRSNVAFKKELHTSQSDRSCTHLEFDISHTGLSYETGDHVGVYSENLSEVVDEALKLLGLSPDTYFSVHADKEDGTPIGGASLPPPFPPCTLRDALTRYADVLSSPKKVALLALAAHASDPSEADRLKFLASPAGKDEYAQWIVANQRSLLEVMQSFPSAKPPLGVFFAAVAPRLQPRYYSISSSPKMSPNRIHVTCALVYETTPAGRIHRGLCSTWMKNAVPLTESPDCSQASIFVRTSNFRLPVDPKVPVIMIGPGTGLAPFRGFLQERLALKESGTELGSSIFFFGCRNRKVDFIYEDELNNFVETGALSELIVAFSREGTAKEYVQHKMSQKASDIWKLLSEGAYLYVCGDAKGMAKDVHRTLHTIVQEQGSLDSSKAELYVKNLQMSGRYLRDVW Arabidopsis thaliana CPR1 (AtCPR1)(SEQ ID NO: 93) MATSALYASDLFKQLKSIMGTDSLSDDVVLVIATTSLALVAGFVVLLWKKTTADRSGELKPLMIPKSLMAKDEDDDLDLG SGKTRVSIFFGTQTGTAEGFAKALSEEIKARYEKAAVKVIDLDDYAADDQYEEKLKKETLAFFCVATYGDGEPTDNAARFYKWFTEENERDIKLQQLAYGVFALGNRQYEHFNKIGIVLDEELCKKGAKRLIEVGLGDDDQSIEDDFNAWKESLWSELDKLLKDEDDKSVATPYTAVIPEYRVVTHDPRFTTQKSMESNVANGNTTIDIHHPCRVDVAVQKELHTHESDRSCIHLEFDISRTGITYETGDHVGVYAENHVEIVEEAGKLLGHSLDLVFSIHADKEDGSPLESAVPPFPGPCTLG TGLARYADLLNPPRKSALVALAAAYATEPSEAEKLKHLTSPDGKDEYSQWIVASQRSLLEVMAAFPSAKPPLGVAIFAIAPRLQPRYYSISSSPRLAPSRVHVTSALVYGPTPTGRIHKGVCSTWMKNAVPAEKSHECSGAPIRASNFKLPS NPSTPIVMVGPGTGLAPFRGFLQERMALKEDGEELGSSLLFFGCRNRQMDFIYEDELNNFVDQGVISELIMAFSREGAQKEYVQHKMMEKAAQVWDLIKEEGYLYVCGDAKGMARDVHRTLHTIVQEQEGVSSSEAEAIVKKLQTEGRYLRDVW Arabidopsis thaliana CPR2 (AtCPR2) MASSSSSSSTMIDLMAAIIKGEPVIVSDPANASAYESVAAELSSMLIENRQFAMIVTTSIAVLIGCIVMLVWRRSGSGNSKRVEPLKPLVIKPREEIDDGRKKKVTIFFGTQTGTAEGFAKALGEEAKARYEKTRFKIVDLDDYAAADDDEYEEKLKKEDVAFFFLATYGDGEPTDNAARFYKWFTEGNDRGEWLKNLKYGVFGLGNRQYEHFNKVAKVVDDILVEQGAQRLVQVGLGDDDQCIEDDFTAWREALWPELDTILREEGDTVATPYTAAVLEYRVSIHDSEDAKFNDINMANGNGYTVFDAQHPYKANVAVKRELHTPESDRSCIHLEFDIAGSGLTYETGDHVGVL CDNLSETVDEALRLLDMSPDTYFSLHAEKEDGTPISSSLPPFPPCNLRTALTRYACLLSSPKKSALVALAAHASDPTEAERLKHLASPAGKDEYSKWVVESQRSLLEVMAEFPSAKPPLGVFFAGVAPRLQPRFYSISSSPKIAETRIHVTCALVYEKMPTGRIHKGVCSTWMKNAV PYEKSENCSSAPIFVRQSNFKLPDSKVPIIMIGPGTGLAPFRGFLQERLALVESGVELGPSVLFFGCNRRRMDFIYEEELQRFVESGALAELSVAFSREGPTKEYVQHKMMDKASDIWNMISQGAYLYVCGDAKGMARDVHRSLHTIAQEQGSMDSTKAEGFVKNLQTSGRYLRDVW Arabidopsis thaliana (AtCPR3) MASSSSSSSTSMIDLMAAIIKGEPVIVSDPANASAYESVAAELSSMLIENRQFAMIVTTSIAVLIGCIVMLVWRRSGSGNSKRVEPLKPLVIKPREEEIDDGRKKVTIFFGTQTGTAEGFAKALGEEAKARYEKTRFKIVDLDDYAADDDEYEEKLKKEDVAFFFLATYGDGEPTDNAARFYKWFTEGNDRGEWLKNLKYGVLGNRQYEHFNKVVDDILVEQGAQRLVQVGLGDDDQCIEDDFTAWREALWPELD TILREEGDTAVATPYTAAVLEYRVSIHDSEDAKFNDITLANGNGYTVFDAQHPYKAVNAVKRELHTPESDRSCIHLEFDIAGSGLTMKLGDHHVGVLCDNLSETVDEALRLLDMSPDTYFSLHAEKEDGTPISSLPPPFPPCNLRTALTRYACLLSSPKKSALVALAAHASDPTEAERLKHLASPAGKDEYSKWVVESQRSLLEVMAEFPSAKPPLGVFFAGVAPRLQPRFYSISSSPKIAETRIHVTCALVYEKMPTGR IHKGVCSTWMKNAVPYEKSEKLFLGRPIFVRQSNFKLPSDSKVPIIIMIGPGTGLAPFRGFLQERLALVESGVELGPSVLFFGCNRRMDFIYEEELQRFVESGALAELSVAFSREGPTKEYVQHKMMDKASDIWNMISQGAYLYVCGDAKGMARDVHRSLHTIAQEQGSMDSTKAEGFVKNLQTSGRYLDVW Stevia rebaudiana CPR2 (SrCPR2) MAQSESVEASTIDLMTAVLKDTVIDTANASDNGDSKMPPALAMMFEIRDLLLILTTSVAVLVGCFVVLVWKRSSGKKSKGELEPPKIVVPKRRLEQEVDDGKKKVTIFFGTQTGTAEGFAKALFEEAKARYEKAAFKVIDLDDYAADLDEYAEKLKKETYAFFFLATYGDGEPTDNA AKFYKWFTEGDEKGVWLQKLQYGVFGLGNRQYEHFNKIGIVVDDGLTEQGAKRIVPVGLGDDDQSIEDDFSAWKELVWPELDLLRDEDDKAAATPYTAAIPEYRVVFHDKPDAFSDDHTQTNGHAVHDAQHPCRSNVAVKKELHTPESDRSCTHLEFDISHTGLSYETGDHVGVYCE NLIEVVEEAGKLLGLSTDTYFSLHIDNEDGSPLGGPSLQPPFPPCTLRKALTNYADLLSSPKKSTLLALAAHASDPTEADRLRFLASREGKDEYAEWVVANQRSLLEVMEAFPSARPPLGVFFAAVAPRLQPRYYSISSSPKMEPNRIHVTCALVYEKTPAGRHIHKGICSTWMKNAV PLTESQDCSWAPIFVRTSNFRLPIDPKVPVIMIGPTGLAPFRGFLQERLALKESGTELGSSILFFGCRNRKVDYIYENELNNFVENGALSELDVAFSRDGPTKEYVQHKMTQKASEIWNMLSEGAYLYVCGDAKGMAKDVHRTLHTIVQEQGSLDSSKAELYVKNLQMSGRYLRDVW Stevia rebaudiana CPR3 (SrCPR3) MAQSNSVKISPLDLVTALFSGKVLDTSNASESGESAMLPTIAMIMINRELMLLTSSAVLIGCVVVLVWRRSSTKKSALEPPVIVVPKRVQEEEVDDGKKKVTVFFGTQTGTAEGFAKALVEEAKARYAKAVFKVIDLDDYAADDDEYEEKLKKSLAFFLATYGDGEPTDNAAR FYKWFTEGDAKGEWNLQYGVFGLGNRQYEHFNKIAKVVDDGLVEQGAKRLVPVGLGDDDQCIEDDFTAWKELVWPELDQLLRDEDDTTVATPYTAAVAEYRVVFHEKPDALSEDYSYTNGHAVHDAQHPCRSNVAVKKELHSPESDRSCTHLEFDISNTGLSYETGDHVVYCEN LSEVVNDAERLVGLPPDTYFSIHTDSEDGSPLGGASLPPPFPPCTLRKALTCYADVLSSPKKSALLALAAHATDPSEADRLKFLASPAGKDEYSQWIVASQRSLLEVMEAFPSAKPSLGVFFASVAPRLQPRYYSISSSPKMAPDRIHVTCALVYEKTPAGRIHKGVCSTWMKNAVPMTESQDCSWAPIYVRTSNFRLPSDPKPVPVIMIGPGTGLAPFRGFLQERLALKEAGTDLGLSILFFGCRNRKVDFIYENELNNFVETGALSELIVAFSREGPTKEYVQHKMSEKASIWNLLSEGAYLYVCGDAKGMAKDVHRTLHTIVQEQGSLDSSKAELYVKNLQMSGRYLRDVW Artemisia annua CPR (AaCPR) MAQSTTSVKLSPFDLMTALLNGKVSFDTSNTSDTNIPLAVFMENRELLMILTTSVAVLIGCVVVLVWRRSSSAAKKAAESPVIVVPKKVTEDEVDDGRKKVTVFFGTQTGTAEGFAKALVEEAKARYEKAVFKVIDLDDYAAEDDEYEEEKLKKESLAFFFLATYGDGEPTDNAARFYKWFTEGEEKGEWLDKLQYAVFGLGNRQYEHFNKIAKVVDEKLVEQGAKRLVPVGMGDDDQCIE DDFTAWKELVPELDQLLRDEDDTSVATPYTAAVAEYRVVFHDKPETYDQDQLTNGHAVHDAQHPCRSNVAVKKELHSPLSDRSCTHLEFDISNTGLSYETGDHVGVYVENLSEVDEAEKLIGLPPHTYFSVHADNEDGTPLGGASLPPPFPPCTLRKALASYADVLSSPKKSALLALAAHATDSTEADRLKFLASPAGKDEYAQWIVASHRSLLEVMEAFPSAKPPLGVFFASVAPRLQPRYYSISSSPRFAPNRIHVTCALVYEQTPSGRVHKGVCSTWMKNAVPMTESQDCSWAPIYVRTSNFRLPSDPKVPVIMIGPGTGLAPFRGFLQERLAQKEAGTELGTAILFFGCRNRKVDFIYEDELNNFVETGALSELVTAFSREGATKEYVQHKMTQKASDIWNLLSEGAYLYVCGDAKGMAKDVHRTLHTIVQEQGSLDSSKAELYVKNLQMAGRYLRDVW CPR (PgCPR) MAQSSSGSMSPFDFMTAIIKGKMEPSNASLGAAGEVTAMILDNRELVMILTTSIAVLIGCVVVFIWRRSSSQTPTAVQPLKPLLAKETESEVDDGKQKVTIFFGTQTGTAEGFAKALADEAKARYDKVTFKVVDLDDYAADDEEYEEKLKKETLAFFFLATYGDGEPTDNAARFYK WFLEGKERGEWLQNLKFGVFGLGNRQYEHFNKIAIVVDEILAEQGGKRLISVGLGDDDQCIEDDFTAWRESLWPELDQLLRDEDDTTVSTPYTAAVLEYRVVFHDPADAPTLEKSYSNANGHSVVDAQHPLRANVAVRRELHTPASDRSCTHLEFDISGTGIAYETGDHVGVYCENL AETVEEALELLGLSPDTYFSVHADKEDGTPLSGSSLPPPFPPCTLRTALTLHADLLSSPKKSALLALAAHASDPTEADRLRHLASPAGKDEYAQWIVASQRSLLEVMAEFPSAKPPLGVFFASVAPRLQPRYYSISSSPRIAPSRIHVTCALVYEKTPTGRVHKGVCSTWMKNSVP SEKSDECSWAPIFVRQSNFKLPADAKVPIIIMIGPGTGLAPFRGFLQERLALKEAGTELGPSILFFGCRNSKMDYIYEDELDNFVQNGALSELVLAFSREGPTKEYVQHKMMEKASDIWNLISQGAYLYVCGDAKGMARDVHRTLHTIAQEQGSLDSSKAESMVKNLQMSGRYLRDVW Camptotheca acuminate CaCPR (SEQ ID NO: 201) MAQSSSVKVSTFDLMSAILRGRSMDQTNVSFESGESPALAMLIENRELVMILTTSVAVLIGCFVVLLWRRSSGKSGKVTEPPKPLMVKTEPEPEVDDGKKKVSIFYGTQTGTAEGFAKALAEEAKVRYEKASFKVIDLDDYAADDEEYEEKLKKETLTFFFLATYGDGEPTDNAARFYKWFMEGKERGDWLKNLHYGVFGLGNRQYEHFNRIAKVVDDTIAEQGGKRLIPVGLGDDDQCIEDDFAAWRELLWPELDQLLQDEDGTTVATPYTAAVLEYRVVFHDSPDASLLDKSFSKSNGHAVHDAQHPCRANVAVRRELHTPASDRSCTHLEFDISGTGLVYETGDHVGVYCENLIEVVEEAEMLLGLSPDTFFSIHTDKEDGTPLSGSSLPPPFPPCTLRRALTQYADLLSSPKKSSLLALAAHCSDPSEADRLRHLASPSGKDEYAQWVVASQRSLLEVMAEFPSAKPPIGAFFAGVAPRLQPRYYSISSSPRMAPSRIHVTCALVFEKTPVGRIHKGVCSTWMKNAVPLDESRDCSWAPIFVRQSNFKLPADTKVPVLMIGPGTGLAPFRGFLQERLALKEAGAELGPAILFFGCRNRQMDYIYEDELNNFVETGALSELIVAFSREGPKKEYVQHKMMEKASDIWNMISQEGYIYVCGDAKGMARDVHRTLHTIVQEQGSLDSSKTESMVKNLQMNGRYLRDVW Non-heme iron oxidase Acetobacter pasteurianus subsp. ascendens (ApGA2ox)(SEQ ID NO: 100) MSVSKTTETFTSIPVIDISKLYSSDLAERKAVAEKGDAARNIGFLYISGHNVSADLIEGVRKAARDFFAEPFEKKMEYYIGTSATHKGFVPEGEEVYSAGRPDHKEAFDIGYEVPANHPLVQAGTPLLGPNNWPDIPGFRSAAEAYRTVFDLGRTLFRGFALALGLNESYFDTVANFPPSKLRMIHYPYDADAQDAPGIGAHTDYECFTILLADKPGLEVMNGNGDWIDAPPIGAFVVNIGDMLEVMTAGEFVATAHRVRKVSEERYSFPLFYACDYHTQIRPLPAFAKKIDASYETITIGEHMWAQALQTYQYLVKKVEKGELKLPKGARKTATFGHFKRNSAA Cucurbita maxima (CmGA2ox) (sequence number 101) MAAASSFSAAFYSGIPLIDLSAPDAKQLIVKACEELGFFKVVKHGVPMELISSLESESTKFFSLPLSEKQRAGPPSPFGYGNKQIGRNGDVGWVEYLLLNTHLESNSDGFLSMFGQDPQKLRSAVNDYISAVRNMAGEILELMAEGLKIQQRNVFSKLVM DEQSDSVFRVNHYPPCPDLQALKGTNMIGFGEHTDPQIISVLRSNNTSGFQISLADGNWISVPPDHSSFFINVGDSLQVMTNRFKSVKHRVLTNNSSKSRVSMIYFGGPPLSEKIAPLASLMQGEERSLYKEFTWFEYKRSAYNSRLADNRLVPFERIAAS Dendrobium catenatum (DcGA3ox) MPSLSKEHFDLYSAFHVPETHAWSSSLHDHPIAGDGATIPVIDSPDAASMVGGACRSWGVFYATSHGIPADLHQVESHARRLFSLPLHRKLQTAPRDGSLSGYGRPPISAFFPKLMWSEGFTLAGHDDHLAVTSQLSPFDSLSFCEVMEAYRKEMKKLAGRLFRLILSLGL EEEEMGQVGPLKELSQAADAIQLNSYPTCPEPERAIGMAAHTDSAFLTVLHQTDGAGGLQVLRDQDESGSARWVDVLRPDCLVVNVGDLLHILSNGRFKSVRHRAVVNRADHRISAAYFIGPAHMKVGSITKLVDMRTGPMYRPVTWPEYLGIRTRLFDKALDSVKFQEKELEKD Cucurbita maxima (CmGA3ox)(SEQ ID NO:103) MATTIADVFKSFPVHIPAHKNLDFDSLHELPDSYAWIQPDSFPSPTHKHHNSILDSDSDSVPLIDLSLPNAAALIGNAFRSWGAFQVINHGVPISLLQSIESSADTLFSLPPSHKLKAARTPDGISGYGLVRISSFFPKRMWSEGFTIVGSPLDHFRQLWPHDYHKHCEIVEEYDREMRSLCGRLMWLGLGELGITRDDMKWAGPDGDFKTSPAATQFNSYPVCPDPDRAMGLGPHTDTSLLTIVYQSNTRGLQVLREGKRWVTVEPVAGGLVVQVGDLLHILTNGLYPSALHQAVVNRTRKRLSVAYVFGPPESAEISPLKKLLGPTQPPLYRPVTWTEYLGKKAEHFNNALSTVRLCAPITGLLDVNDHSRVKVG Cucurbita maxima (CmGA20ox) (sequence number 104) MHVVTSTPEARHDGAPLVFDASVLRHQHNIPKQFIWPDEEKPAATCPELEVPLIDLSFLSGEKDAAAEAVRLVGEACEKHGFFLVVNHGVDRKLIGEAHKYMDEFFELPLSQKQSAQRKAGEHCGYASSFTGRFSSKLPWKETLSFRFAADESLNNLVLHYLNDKLGDQFAKFGRVYQDYCEAMSGLSLGIMELLGKSLGVEEQCFKNFFKDNDSIMRLNFYPPCQKPHLTLGTGPHCD PTSLTILHQDQVGGLQVFVDNQWRLITPNFDAFVVNIGDTFMALSNGRYKSCLHRAVVNSERTRKSLAFFLCPRNDKVVRPPRELVDTQNPRRYPDFTWSMLLRFTQTHYRADMKTLEAFSAWLQQEQQEQQEQQFNI Agapanthus praecox subsp. orientalis (ApoGA20ox) MVLQPFVFDAALLRDEHNIPTQFIWPEEDKPSPDASEELILPFIDLKAFLSGDPDSPFQVSKQVGEACESLGAFQVTNHGIDFDLLEEAHSCIQKFFSMPLCEKQRALRKAGESYGYASSFTGRFCSKLPWKETLSFRYSSSSSDIVQNYFVRTLGEEFRHFGEVYQKYCESMSKLSLMI MEVLGLSLGVGRMHFREFFEGNDSTMRLNYYPPCKKPDLTLTGPHCDPTSLTILHQDDVSGLQVFTGGKWLTVRPKTDAFVVNIGDTFTALSNGRYKSCLHRAVVNSKTARKSLAFFLCPAMNKIVRPPRELVDIDHPRAYPDFTWSALLEFTQKHYRADMQTLNEFSKYILQAQGTLHK Arabidopsis thaliana (AtF3H) MAPGTLTELAGESKLNSKFVRDEDERPKVAYNVFSDEIPVISLAGIDDVDGKRGEICRQIVEACENWGIFQVVDHGVDTNLVADMTRLARDFFALPPEDKLRFDMSGGKKGGFIVSSHLQGEAVQDWREIVTYFSYPVRNRDYSRWPDKPEGWVKVTEEYSERLMSLACKLLEVLSEAM GLEKESLTNACVDMDQKIVVNYYPKCPQPDLTLGLKRHTDPGTITLLLQDQVGGLQATRDNGKTWITVQPVEGAFVVNLGDHGHFLSNGRFKNADHQAVVNSNSSRLSIATFQNPAPDATVYPLKVREGEKAILEEPITFAEMYKRKMGRDLELARLKKLAKEERDHKEVDKPVDQIFA Chrysosplenium americanum (CaF6H) QEKTLNSRFVARDEDSLERPKVSAIYNGSFDEIPVLISLAGIDMTGAGTDAAARRSEICRKIVEACEDWGIFGEIDDDHGKRAEICDKIVKACEDWGVFQPDEKLESVMSAAKKGDFVVDHGVDAEVISQWTTFAKPTSHTQFETTTRDFPNKPEGWKATTEQYSRTLMGLACKLLGVISEAMGLEKEALTKACVDMDQKVVVNYYPKCPQPDLTLGLKRHTDPGTITLLQDQVGGLQATRDGGKTWITVQPVKDNGWILLHIGDSNGHRHGHFLSNGRFKSHQAYRYRRPTRGSPTFGTKVSNYPPCPEQSLVRPPAGRPYGRALNALDAKKLASAKQQLESAAILLISELAVAYIILAILPSSEIIAEEGYL Datura stramonium (DsH6H) MATFVSNWSTNVSESFIAPLEKRAEKDVALGNDVPIIDLQQDHLLIVQQITKACQDFGLFQVINHGVPEKLMVEAMEVYKEFFALPAEEKEKFQPKGEPAKFELPLEQKAKLYVEGERRCNEEFLYWKDTLAHGCYPLHEELLNSWPEKPPTYRDVIAKYSVEVRKLTMRILDYICEGLGLKLGYFDNELTQIQMLLANYYPSCPDPSSTIGSGGHYDGNLITLLQQDLVGLQQLIVKDDKWIAVEPIPTAFVVNLGLTLKVMSNEKFEGSIHRVVTHPTRNRISIGTLIGTPDYSCTIEPIKELLSQENPPLYKPYYAKFAEIYLSDKSDYDAGVKPYKINQFPN Arabidopsis thaliana (AtH6DH) MENHTTMKVSSLNCIDLANDDLNHSVVSLKQACLDCGFFY VINHGISEEFMDDVFEQSKKLFALPLEEKMKVLRNEKHRGYTPVLDELLDPKNQINGDHKEGYYIGIEVPKDDPHWDKPFYGPNPWPDADVLPGWRETMEKYHQEALRVSMAIARLLALALDLDVGYFDRTEMLKGPIATMRLL RYQGISDPSKGIYACGAHSDFGMMTLLATDGVMGLQICKDKNAMPQKWEYVPPIKGAFIVNLGDMLERWSNGFFKSTLHRVLGNGQERYSIPFFVEPNHDCLVECLPTCKSESELPKYPPIKCSTYLTQRYEETHANLSIYHQQT Solanum lycopersicum (SlF35H) (sequence number 110) MALRINELFVAAIIYIIVHIIISKLITTVRERGRRLPLPPGPTGWPVIGALPLLGSMPHVALAKMAKKYGPIMYLKVGTCGMVVASTPNAAKAFLKTLDINFSNRPPNAGATHLAYNAQDMVFAPYG PRWKLLRKLSNLHMLGGKALENWANVRANELGHMLKSMFDASQDGECVVIADVLTFAMANMIGQVMLSKRVFVEKGVEVNEFKNMVVELMTVAGYFNIGDFIPKLAWMDIQGIEKGMKNLHKKFDDLL TKMFDEHEATSNERKENPDFLDVVMANRDNSEGERLSTTNIKALLLNLFTAGTDTSSSVIEWALAEMMKNPKIFEKAQQEMDQVIGKNRRLIESDIPNLPYLRAICKETFRKHPSTPLNLPRVSSEPC TVDGYYIPKNTRLSVNIWAIGRDPDVWENPLEFTPERFLSGKNAKIEPRGNDFELIPFGAGRRICAGTRMGIVMVEYILGTLVHSFDWKLPNNVIDINMEESFGLALQKAVPLEAMVTPRLSLDVYRC D4H (SEQ ID NO: 111) MPKSWPIVISSHSFCFLPNSEQERKMKDLNFHAATLSEEESLRELKAFDETKAGVKGIVDTGITKIPRIFIDQPKNLDRISVCRGKSDIKIPVINLNGLSSNSEIRREIVEKIGEASEKYGFFQIVNHGIPQDVMDKMVDGVRKFHEQDDQIKRQYYSRDRFNKNFLYSSNYVLIPGIACNWRDTMECIMNSNQPDPQEF PDVCRDILMKYSNYVRNLGLILFELLSEALGLKPNHLEEMDCAEGLILLGHYYPACPQPELTFGTSKHSDSGFLTILMQDQIGGLQILLENQWIDVPFIPGALVINIADLLQLITNDKFKSVEHRVLANKVGPRISVAVAFGIKTQTQEGVSPRLYGPIKELISEENPPIYKEVTVKDFITIRFAKRFDDSSSLSPFRLNN Catharanthus roseus (CrD4H-like) (SEQ ID NO: 112) MKELNNSEEELKAFDDTKAGVKALVDSGITEIPRIFLDHPTNLDQISSKDREPKFKKNIPVIDLDGISTNSEIRREIVEKIREASEKWGFFQIVNHGIPQEVMDDMIVGIRRFHEQDNEIKKQFYTRDRTKSFRYTSNFVLNPKIACNWRDTFECTMAPHQPNPQDLPDICRDIMMKYISYTRNLGLTLFELLSEALGLKSNRLKDMHCDEGVELVGHYYPACQPELTLGTSKHTDTGFLTMLQQDQIGGLQVLYENHQWVDVPFIPGALIINIGDFLQIISNDKFKSAPHRVLANKNGPRISTASVFMPNFLESAEVRLYGPIKELLSEENPPIYEQITAKDYVTVQFSRGLDGDSFLSPFMLNKDNMEK Zea mays (ZmBX6) MAPTTATKDDSGYGDERRRELQAFDDTKLGVKGLVDSGVKSIPSIFHHPPEALSDIISPAPLPSSPPSGAAIPVVDLSVTRREDLVEQVRHAAGTVGFFWLVNHGVAEELMGGMLRGVRQFNEGPVEAKQALYSRDLARNLRFASNFDLFKAAADWRDT LFCEVAPNPPPREELPEPLRNVMLEYGAAVTKLARFVFELLSESLGMPSDHLYEMECMQNLNVVCQYYPPPEPHRTVGVKRHTDPGFFTILLQDGMGGLQVRLGNNGQSGGCWVDIAPRPGALMVNIGDLLQLVTNDRFSVEHRVPANKSSDTARVSVASFFNTDVRRSERMYGPIPDPSKPPLYRSVRARDFIAKFNTIGLDGRALDHFRL Hordeum vulgare subsp. vulgare (HvIDS2) MAKVMNLTPVHASSIPDSFLLPADRLHPATTDVSLPIIDMSRGRDEVRQAILDSGKEYGFIQVVNHGISEPMLHEMYAVCHEFFDMPAEDKAEFFSEDRSERNKLFCGSAFETLGEKYWIDVLELLYPLPSGDTKDWPHKPQMLREVVGNYTSLARGVAMEILRLLCEG LGLRPDFFVGDISGGRVVVDINYYPPSPNPSRTLGLPPHCDRDLMTVLLPGAVPGLEIAYKGGWIKVQPVPNSLVINFGLQLEVVTNGYLKAVEHRAATNFAEPRLSVASFIVPADDCVVGPAEEFVSEDNPPRYRTLTVGEFKRKHNVVNLDSSINQIININNNQKGI Hordeum vulgare subsp. vulgare (HvIDS3) (SEQ ID NO: 115) MENILHATPAPVSLPESFVFASDKVPPATKAVVSLPIIDLSCGRDEVRRSILEAGKELGFFQVVNHGVSKQVMRDMEGMCEQFFHLPAADKASLYSEERHKPNRLFSGATYDTGGEKYWRDCLRLACPFPVDDSINEWPDTPKGLRDVIEKFTSQTRDVGKELLRLLCE GMGIRADYFEGDLSGGNVILNINHYPSCPNPDKALGQPPHCDRNLITLLPGAVNGLEVSYKGDWIKVDPAPNAFVVNFGQQLEVVTNGLLKSIEHRAMTNSALARTSVATFIMPTQECLIGPAKEFLSKENPPCYRTTMFRDFMRIYNVVKLGSSLNLTTNLKNVQKEI Uridine diphosphate-dependent glycosyltransferases (UGTs) Siraitia grosvenorii UGT720-269-1 (SEQ ID NO: 116) MEDRNAMDMSRIKYRPQPLRPASMVQPRVLLFPFPALGHVKPFLSLAELLSDAGIDVVFLSTEYNHRRISNTEALASRFPTLHFETIPDGLPPNESRALADGPLYFSMREGTKPRFRQLIQSLNDGRWPITCIITDIMLSSPIEVAEEFGIPVIAFCPCSARYLSIHFFIPKLVEEGQIPYADDDPIGEIQGVPLFEGLLRRNHLPGSWSDKSADISFSHGLINQTLAAGRASALILNTFDELEAPFLTHLSSIFNKIYTIGPLHALSKSRLGDSSSSASALSGFWKEDRACMSWLDCQPPRSVVFVSFGSTMKMKADELREFWYGLVSSGKPFLCVLRSDVVSGGEAAELIEQMAEEEGAGGKLGMVVEWAAQEKVLSHPAVGGFLTHCGWNSTVESIAAGVPMMCWPILGDQPSNATWIDRVWKIGVERNNREWDRLTVEKMVRALMEGQKRVEIQRSMEKLSKLANEKVVRGINLHPTISLKKDTPTTSEHPRHEFENMRGMNYEMLVGNAIKSPTLTKK Siraitia grosvenorii UGT94-289-3 (Accession No. 117) MTIFFSVEILVLGIAEFAAIAMDAAQQGDTTTILMLPWLGYGHLSAFLELAKSLSRRNFHIYFCSTSVNLDAIKPKLPSSFSDSIQFVELHLPSSPEFPPHLHTTNGLPPTLMPALHQAF SMAAQHFESILQTLAPHLLIYDSLQPWAPRVASSLKIPAINFNTTGVFVISQGLHPIHYPHSKFPFSEFVLHNHWKAMYSTADGASTERTRKRGEAFLYCLHASCSVILINSFRELEGKYMDYLSVLLNKKVVPVGPLVYEPNQDGEDEGYSSIKNWLDKKEPSSTVFVSFGSEYFPSKEEMEEIAHGLEASEVNFIWVVRFPQGDNTSGIEDALPKGFLERAGERGMVVKGWAPQAKILKHWSTGGFVSHCGWNSVMESMMFGVPIIGVPMHVDQPFNAGLVEEAGVGVEAKRDPDGKIQRDEVAKLIKEVVVEKTREDVRKKAREMSEILRSKGEEKFDEMVAEISLLLKI Siraitia grosvenorii UGT74-345-2 (Accession No. 118) MDETTVNGGRRASDVVVFAFPRHGHMSPMLQFSKRLVSKGLRVTFLITTSATESLRLNLPPSSSLDLQVISDVPESNDIATLEGYLRSFKATVSKTLADFIDGIGNPPKFIVYDSVMPWVQEVARGRGLDAAPFFTQSSAVNHILNHVYGGSLSIPAPENTAVSLPSMPVLQAEDLPAFPDDPEVVMNFMTSQFSNFQDAKWIFFNTFDQLECKKQSQVVNWMADRWPIKTVGPTIPSAYLDDGRLEDDRAFGLNLLKPEDGKNTRQWQWLDSKDTASVLYISFGSLAILQEEQVKELAYFLKDTNLSFLWVLRDSELQKLPHNFVQETSHRGLVVNWCSQLQVLSHRAVSCFVTHCGWNSTLEALSLGVPMVAIPQWVDQTTNAKFVADVWRVGVRVKKKDERIVTKEELEASIRQVVQGEGRNEFKHNAIKWKKLAKEAVDEGGSSDKNIEEFVKTIA Siraitia grosvenorii UGT75-281-2 (Accession No. 119) MGDNGDGGEKKELKENVKKGKELGRQAIGEGYINPSLQLARRLISLGVNVTFATTVLAGRRMKNKTHQTATTPGLSFATFSDGFDDETLKPNGDLTHYFSELRRCGSESLTHLITSAANEGRPITFVIYSLLLSWAADIASTYDIPSALFFAQPATVLALYFYYFHGYGDTICSKLQDPSSYIELPGLPLLTSQDMPSFFSPSGPHAFILPPMREQAEFLGRQSQPKVLVNTFDALEADALRAIDKLKMLAIGPLIPSALLGGNDSSDASFCGDLFQVSSEDYIEWLNSKPDSSVVYISVGSICVLSDEQEDELVHALLNSGHTFLWVKRSKENNEGVKQETDEEKLKKLEEQGKMVSWCRQVEVLKHPALGCFLTHCGWNSTIESLVSGLPVVAFPQQIDQATNAKLIEDVWKTGVRVKANTEGIVEREEIRRCLDLVMGSRDGQKEEIERNAKKWKELARQAIGEGGSSDSNLKTFLWEIDLEI Siraitia grosvenorii UGT720-269-4 (Accession No. 120) MAEQAHDLLHVLLFPFPAEGHIKPFLCLAELLCNAGFHVTFLNTDYNHRRLHNLHLLAARFPSLHFESISDGLPPDQPRDILDPKFFISICQVTKPLFRELLLSYKRISSVQTGRPPITCVITDVIFRFPIDVAEELDIPVFSFCTFSARFMFLYFWIPKLIEDGQLPYPNGNINQKLYGVAPEAEGLLRCKDLPGHWAFADELKDDQLNFVDQTTASSRSSGLILNTFDDLEAPFLGRLSTIFKKIYAVGPIHSLLNSHHCGLWKEDHSCLAWLDSRAAKSVVFVSFGSLVKITSRQLMEFWHGLLNSGKSFLFVLRSDVVEGDDEKQVVKEIYETKAEGKWLVVGWAPQEKVLAHEAVGGFLTHSGWNSILESIAAGVPMISCPKIGDQSSNCTWISK VWKIGLEMEDRYDRVSVETMVRSIMEQEGEKMQKTIAELAKQAKYKVSKDGTSYQNLECLIQDIKKLNQIEGFINNPNFSDLLRV Siraitia grosvenorii UGT94-289-2 (Accession No. 121) MDAQQGHTTTILMLPWVGYGHLLPFLELAKSLSRRKLFHIYFCSTSVSLDAIKPKLPPSISSDDSIQLVELRLPSSPELPPHLHTTNGLPSHLMPALHQAFVMAAQHFQVILQTLAPHLLIYDILQPWAPQVASSLNIPAINFSTTGASMLSRTLHPTHYPSSKFPISEFVLHNHWRAMYTTADGALTEEGHKIEETLANCLHTSCGVVLVNSFRELETKYIDYLSVLLNKKVVPVGPLVYEPNQEGEDEGYSSIKNWLDKKEPSSTVFVSFGTEYFPSKEEMEEIAYGLELSEVNFIWVLRFPQGDSTSTIEDALPKGFLERAGERAMVVKGWAPQAKILKHWSTGGLVSHCGWNSMMEGMMFGVPIIAVPMHLDQPFNAGLVEEAGVGVEAKRDSDGKIQREEVAKSIKEVVIEKTREDVRKKAREMDTKHGPTYFSRSKVSSFGRLYKINRPTTLTVGRFWSKQIKMKRE<000104�>Siraitia grosvenorii UGT94-289-1 (Accession No. 122) It should be noted that in the original text, seems to be misspelled as <000104�> in the provided content. I translated it as as it's likely a typo in the original. If this is incorrect, please let me know.MDAQRGHTTTILMFPWLGYGHLSAFLELAKSLSRRNFHIYFCSTSVNLDAIKPKLPSSSSDSIQLVELCLPSSPDQLPPHLHTTNALPPHLMPTLHQAFSMAAQHFAAILHTLAPHLLIYDSFQPWAPQLASSLNIPAINFNTTGASVLTRMLHATHYPSSKPFISEFVLHDYWKAMYSAAGGAVTKKDHKIGETLANCLHASCSVILINSFRELEEKYMDYLSV LLNKKVVPVGPLVYEPNQDGEDEGYSSIKNWLDKKEPSSTVFVSFGSEYFPSKEEMEEIAHGLASEVHFIWVVRFPQGDNTSAIEDALPKGFLERVGERGMVVKGWAPQAKILKHWSTGGFVSHCGWNSVMESMMFGVPIIGVPMHLDQPFNAGLAEEAGVGVEAKRDPDGKIQRDEVAKLIKEVVVEKTREDVRKKAREMSEILRSKGEMDEMVAAISLFLKI Momordica charantia 1 (McUGT1) (SEQ ID NO:123) MAQPQTQARVLVFPYPTVGHIKPFLSLAELLADGGLDVVFLSTEYNHRRIPNLEALASRFPTLHFDTIPDGLPIDKPRVIIGGELYTSMRDGVKQRLRQVLQSYNDGSSPITCVICDVMLSGPIEAAEEELGIPVVTFCPYSARYLCAHFVMPKLIEEGQIPFTDGNLAGEIQGVPLFGGLLRRDHLPGFWFVKSLSDEVWSHAFLNQTLAVGRTSALIINTLDELEAPFLAHLSST FDKIYPIGPLDALSKSRLGDSSSSSTVLTAFWKEDQACMSWLDSQPPKSVIFVSFGSTMRMTADKLVEFWGHGLVNSGTRFLCVLRSDIVEGGGAADLIKQVGETGNGIVVEWAAQEKVLAHRAVGGFLTHCGWNSTMESIAAGVPMMCWQIYGDQMINATWIGKVWKIGERDDKWDRSTVEKMIKELMEGEGAIQRSMEKFSKLANDKVVKGGTSFENLELIVEYLKKLKPSN Momordica charantia 2 (McUGT2) (SEQ ID NO:124) MAQPRVLLFPFPAMGHVKPFLSLAELLSDAGVEVVFLSTEYNHRRIPDIGALAARFPTLHFETIPDGLPPDQPRVLADGHLYFSMLDGTKPRFRQLIQSLNGNPRPITCIINDVMLSSPIEVAEEFGIPVIAFCPCSARFLSVHFFMPNFIEEAQIPYTDENPMGKIEATVFEGLLRRKDLPGLWCAKSSNISFSHRFI NQTIAAGRASLILNTFDELESPFLNHLSSIFPKIYCIGPLNALSRSRLGKSSSSSLASAGFWKEDQAYMSWLESQPPRSVIFVSFGSTMKMEAWKLAEFWYGLVNSGSPFLFVFRPDVINSGDAAEVMEGR GRGMVVEWASQEKVLAHPAVGGFLTHCGWNSTVESIVAGVPMMCCPIVADQLSNATWIHKVWKIGIEGDEKWDRSTVEMMIKELMESQKGTEIRTSIEMLSKLANEKVVKGGTSLNNFELLVEDIKTLRRPYT Momordica charantia 3 (McUGT3) (SEQ ID NO:125) MEQSDSNSDDHQHHVLLFPPFAKGHIKPFLCLAQLLCGAGLQVTFLNTDHNHRRIDDRHRRLLATQFPMLHFKSISDGLPDHPRDLLDGCLIASMRRVTESLFRQLLLSYNGYGNGTNNVSNSGRRPPISCVITDVIFSFPVEVAEELGIPVFSFATFSARFLYFWIPKLIQEGQLPFPDGKTNQELYPGAEGIIRCKDLPGSWSVEAVAKNDPMNFVKQTLASSRSSGLILNTFE DLEAPFVTHLSNTFDKIYTIGPIHSLLGTSHCGLWKEDYACLAWLDARPRKSVVFVSFGSLVKTTSRELMELWHGLVSSGKSFLLVLRSDVVEGEDEEQVVKEILESNGEGKWLVVGWAPQEEVLAHEAIGGFLTHSGWNSTMESIAAGVPMVCWPKIGDQPSNCTWVSRVWKVGLEMEERYDRSTVAMARSMMEQEGKEMERRIAELAKRVKYRVGKDGESYRNLESLIRDIKTSSN Momordica charantia 4 (McUGT4) (SEQ ID NO:126) MDAHQQAEHTTTILMLPWVGYGHLTAYLELAKALSRRNFHIYYCSTPVNIESIKKPLTIPCSSIQFVELHLPSSDDLPPNLHTTNGLPSHLMPTLHQAFSAAAPLFEEILQTLCPHLLIYDSLQPWAPKIASSLKIPALNFNTSGVSVIAQALHAIHHPDSKFPLSDFILHNYWKSTYTTADGGASEKTRRAREAFLYCLNSSGNAILINTFRELEGEYIDYLSL LLNKKVIPIGPLVYEPNQDEDQDEEYRSIKNWLDKKEPCSTVFVSFGSEYFPSNEEMEEIAPGLEESGANFIWVVRFPKLENRNGIIEGGLERAGERGMVIKEWAPQARILRHGSIGGFVSHCGWNSVMESIICGVPVIGVPMRVDQPYNAGLVEEAGVGVEAKRDPDGKIQRHEVSKLIKQVVVEKTRDDVRKKVAQMSEILRRKGDEKIDEMVALISLLPKG Momordica charantia 5 (McUGT5) (SEQ ID NO:127) MDARQQAEHTTTILMLPWVGYGHLSAYLELAKALSRRNFHIYYCSTPVNIESIKKPKLTIPCSSIQFVELHLPFSDDLPPNLHTTNGLPSHLMPALHQAFSAAAPLFEAILQTLCPHLLIYDSLQPWAPQIASSLKIPALNFNTTGVSVIARALHTIHHPDSKFPLSEIVLHNYWKATHATADGANPEKFRRDLEALLCCLHSSCNAILINTFRELEGEYIDYLSL LLNKKVTPIGPLVYEPNQDEEQDEEYRSIKNWLDKKEPYSTIFVSFGSEYFPSNEEMEEIARGLEESGANFIWVVRFHKLENGNGITEEGLLERAGERGMVIQGWAPQARILRHGSIGGFVSHCGWNSVMESIICGVPVIGVPMGLDQPYNAGLVEEAGVGVEAKRDPDGKIQRHEVSKLIKQVVVEKTRDDVRKKVAQMSEILRKGDEKIDEMVALISLLKG Cucumis sativus MGLSPTDHVLLFPPAKGHIKPFFCLAHLLCNAGLRVTFLSTEHHHQKLHNLTHLAAQIPSLHFQSISDGLSLDHPRNLL DGQLFKSMPQVTKPLFRQLLLSYKDGTSPITCVITDLILRFPMDVAQELDIPVFCFSTFSARFLFLYFSIPKLLEDGQIPYPEGNSNQVLHGIPGAEGLLRCKDLPGYWSVEAVANYNPMNFVNQTIATSKSHGLILNTFDELEVPFITNLSKIYKKVYTIGPIHSLLKKSVQTQYEFWKEDHSCLAWLDSQPPRSVMFVSGFGSIVKLKSSQLKEFWNGLVDSGKAFLLVLRSDALVEETGEEDEKQKELVIKEIMETKEEGRWVIVNWAPQEKVLEHKAIGGFLTHSGWNSTLESVAVGVPMVSWPQIGDQPSNATWLSKVWKIGVEMEDSYDRSTVESKVRSIMEHEDKKMENAIVELAKRVDDRVSKEGTSYQNLQRLIEDIEGFKLN Cucurbita maxima 1 (CmaUGT1)(SEQ ID NO:129) MELSHTHHVLLFPPAKGHIKPFFSLAQLLCNAGLRVTFLNTDHHHRRIHDLNRLAAQLPTLHFDSVSDGLPPDEPRNVFDGKLYESIRQVTSSLFRELLVSYNNGTSSGRPPITCVITDVMFRFPIDIAEELGIPVFTFSTFSARFLIFWIPKLLEDGQLRYPEQELHGVPGAEGLIRWKDLPGFWSVEDVADWDPMNFVNQTLATSRSSGLILNTFDELEAPFLTSL SKIYKKIYSLGPINSLLKNFQSQPQYNLWKEDHSCMAWLDSQPRKSVVFVSFGSVVKLTSRQLMEFWNGLVNSGMPFLLVLRSDVIEAGEEVVREIMERKAEGRWVIVSWAPQEEVLAHDAVGGFLTHSGWNSTLESLAAGVPMISWPQIGDQTSNSTWISKVWRIGLQLEDGFDSSIETMVRSIMDQTMEKTVAELAERAKNRASKNGTSYRNFQTLIQDITNIIETHI Cucurbita maxima 2 (CmaUGT2)(SEQ ID NO:130) MDAQKAVDTPPTTVLMLPWIGYGHLSAYLELAKALSRRNFHVYFCSTPVNLDSIKPNLIPPPSSSIQFVDLHLPSSPELPPHLHTTNGLPSHLKPTLHQAFSAAAQHFEAILQTLSPHLLIYDSLQPWAPRIASSLNIPAINFNTTAVSIIAHALHSVHYPDSKFPFSDFVLHDYWKAKYTTADGATSEKIRRGAEAFLYCLNASCDVLVNSFRELEGEYMDYLSVL LKKKVVSVGPLVYEPSEGEEDEEYWRIKKWLDEKEALSTVLVSFGSEYFPSKEEMEEIAHGLEESEANFIWVVRFPKGEESCRGIEALPKGFVERAGERAMVVKKWAPQGKILKHGSIGGFVSHCGWNSVLESIRFGPVIGVPMHLDQPYNAGLLEEAGIGVEAKRDADGKIQRDQVASLIKRVVVEKTREDIWKTVREMREVLRRRRDDMIDEMVAEISVVLKI Cucurbita maxima 3 (CmaUGT3)(SEQ ID NO:131) MSSNFLKISIPFGRLRDSALNCSVFHCKLHLAIAIAMDAQQAANKSPTATTIFMLPWAGYGHLSAYLELAKALSTRNFHIYFCSTPVSLASIKPRLIPSCSSIQFVELHLPSSDEFPPHLHT TNGLPSRLVPTFHQAFSEAAQTFEAFLQTLRPHLLIYDSLQPWAPRIASSLNIPAINFFTAGAFAVSHVLRAFHYPDSQFPSSDFVLHSRWKIKNTTAESPTQAKLPKIGEAIGYCLNASRGV ILTNSFRELEGKYIDYLSVILKKRVFPIPLVYQPNQDEEDEDYSRIKNWLDRKEASSTVLVSFGSEFFLSKEETEAIAHGLEQSEANFIWGIRFPKGAKKNAIEEALPEGFLERAGGRAMVV EEWVPQGKILKHGSIGGFVSHCGWNSAMESIVCGVPIIGIPMQVDQPFNAGILEEAGVGVEAKRDSDGKIQRDEVAKLIKEVVVERTREDIRNKLEKINEILRSRREEKLDELATEISLLSRN Cucurbita moschata 1 (CmoUGT1) (SEQ ID NO: 132) MELSPTHHLLLFPFPAKGHIKPFFSLAQLLCNAGARVTFLNTDHHHRRIHDLDRLAAQLPTLHFDSVSDGLPPDESRNVFDGKLYESIRQVTSSLFRELLVSYNNGTSSGRPPITCVITDCMFRFPIDIAEELGIPVFTFSTFSARFLFFWIPKLLEDGQLRYPEQELHGVPGAEGLIRCKDLPGFLSDEDVAHWKPINFVNQILATSRSSGLILNTFDELEAPFLTSL SKIYKKIYSLGPINSLLKNFQSQPQYNLWKEDHSCMAWLDSQPPKSVVFVSFGSVVKLTNRQLVEFWNGLVNSGKPFLLVLRSDVIEAGEEVVRENMERKAEGRWMIVSWAPQEEVLAHDAVGGFLTHSGWNSTLESLAAGVPMISWTQIGDQTSNSTWVSKVWRIGLQLEDGFDSFTIETMVRSVMDQTMETVAELAKNRASKNGTSYNFQTLIQDITNIIETHI Cucurbita moschata 2 (CmoUGT2) (SEQ ID NO:133) MDAQKAVDTPPTTTVLMLPWIGYGHLSAYLELAKALSRRNFHVYFCSTPVNLDSIKPNLIPPPSIQFVDLHLPSSPELPPHLHTTNGLPSHLKPTLHQAFSAAAQHFEAILQTLSPHLLIYDSLQPWAPRIASSLNIPAINFNTTAVSIIAHALHSVHYPDSKFPFSDFVLHDYWKAKYTTADGATSEKTRRGVEAFLYCLNASCVLVNSFRELEGEYMDYLSVLL KKKVVSVGPLVYEPSEGEEDEEYWRIKKWLDEKEALSTVLVSFGSEYFPPPKEEMEEIAHGLEESEANFIWVVRFPKGESSSRGIEEALPKGFVERAGERAMVVKKWAPQGKILKHGSIGGFVSHCGWNSVLESIRFGVPVIGAPMHLDQPYNAGLLEEAGIGVEAKRDADGKIQRDQVASLIKQVVVEKTREDIWKKVREMREVLRRRRDDDMMIDEMVAVISVVLKI Cucurbita moschata 3 (CmoUGT3) (SEQ ID NO:134) MDAQQAANKSPTASTIFMLPWVGYGHLSAYLELAKALSTRNFHVYFCSTPVSLASIKPRLIPSCSSIQFVELHLPSSDEFPPHLHTTNGLPAHLVPTIHQAFAAAAQTFEAFLQTLRPHLLIYDSLQPWAPRIASSLNIPAINFFTAGAFAVSHVLRAFHYPDSQFPSSDFVLHSRWKIKNTTAESPTQVKIPKIGEAIGYCLNASRGVILTNSFRELEGKYIDYLS VILKKRVLPIGPLVYQPNQDEEDEDYSRIKNWLDRKEASSTVLVSFGSEFFLSKEETEAIAHGLEQSEANFIWGIRFPKGAKKNAIEELPEGFLERVGGRAMVVEEWVPQGKILKHGNIGGFVSHCGWNSAMESIMCGVPVIGIPMQVDQPFNAGILEEAGVGVEAKRDSDGKIQRDEVAKLIKEVVVERTREDIRNKLEEINEILRTRREEKLDELATEISLLCKN Prunus persica MAMKQPHVIIFPFPLQGHMKPLLCLAELLCHAGLHVTYVNTHHNHQRLANRQALSTHFPTLHFESISDGLPEDDPRTLNSQLLIALKTSIRPHFRELLKTISLKAESNDTLVPPPSCIMTDGLVTFAFDVAEEGLPILSFNVPCPRYLWTCLCLKPKLIENGQLPFQDDDMNVEITGVPGMEGLLHRQDLPGFCRVKQAD HPSLQFAINETQTLKRASALILDTVYELDAPCISHMALMFPKIYTLGPLHALLNSQIGDMSRGLASHGSLWKSDLNCMTWLDSQPSKSIIYVSFGTLVHLTRAQVIEFWYGLVNSGHPFLWVMRSDITSGDHQIPAELENGTKERGCIVDWVSQEEVLAHKSVGGFLTHSGWNSTLESIVAGLPMICWPKLGDHYIISST VCRQWKIGLQLNENCDRSNIESMVQTLMGSKREEIQSSMDAISKLSRDSVAEGGSSHNNLEQLIEYIRNLQHQN Theobroma cacao MRQPHVLVLFPPAQGHIKPMLCLAELLCQAGLRVTFLNTHHSHRRLNNLQDLSTRFPTLHFESVSDGLPEDHPRNLVHFMHLVHSIKNVTKPLLRDLLTSLSLKTDPPVSCIADGILSFAIDVAEELQIKVIIFRTISSSCCLWSYLVCVPKLIQQGELQFSDSDMGQKVSSVPEMKGSLRLHDRPYSFGLKQLEDPNFQFFVSETQAMTRASAVIFNTFDSLEAPVLSQMIPL LPKVYTIGPHLARKARLGDLSQHSSFNGNLREADHNCITWLDSQPLRSVVYVSFGSHVVLTSEELLEFWHGLVNSGKRFLWVLRPDIIAGEKDHNQIIAREPDLGTKEKGLLVDWAPQEEVLAHPSVGGFLTHCGWNSTLESMVAGVPMLCWPKLPDQLVNSSCVSEVWKIGLDLKDMCDRSTVEKMVRALMEDRREEVMRSVDGISKLARESVSHGGSSSNLEMLIQELET Corchorus capsularis MDSKQKKMSVLMFPWLAYGHISPFLELAKKLSKRNFHTFFFFSTPINLNSIKSKLSKPYAQSIQFVELHLPSLPDLPPHYHTTNGLPPHLMNTLKKAFDMSSLQFSKILKTLNPDLLVYDFIQPWAPLLALSNKIPAVHFLCTSAAMSSFSVHAFKKPCEDFPFPNIYVHGNFMNAKFNNMENCSSDDSISDQDRVLQCFERSTKIILVKTFEELEGKFMDYLSVLLNKKIV PTGPLTQDPNEDEGDDDERTKLLLEWLNKKSSKSSTVFVSFGSEYFLSKEEEREEIAYGLELSKVNFIWVIRFPLGENKTNLEEALPQGFLQRVSERGLVVENWAPQAKILQHSSIGGFVSHCGWSSVMESLKFGVPIIAIPMHLDQPLNARLVVDVVGGLEVIRNHGSLEREEIAKLIKEVVLGNGNNDGEIVRRKAREMSNHIKKKGDMDELVELMLICKMKPNSCHLS Ziziphus jujube MMERQRSIKVLMFPWLAHGHISPFLELAKRLTDRNFQIYFCSTPVNLTSVKPKLSQKYSSIKLVELHLPSLPDLPPHYHTTNGLALNLIPTLKKAFDMSSSFSTILSTIKPDLLIYDFLQPWAPQLASCMNIPAVNFLSAGASMVSFVLHSIKYNGDDHDDEFLTTELHLSDSMEAKFAEMTESSPDEHIDRAVTCLERSNSLILIKSFRELEGKYLDYLSLSFAKKVVPIGPLVAQDTNPEDDSMDIINWLDKKEKSTVFVSFGSEYYLTNEEMEEIAYGLELSKVNFIWVVRFPLGQKMAVEEALPKGFLERVGEKGVVEDWAPQMKILGHSSIGGFVSHCGWSSLMESLKLGVPIIAMPMQLDQPINAKLVERSGVGLEVKRDKNGRIEREYLAKVIREIVVEKARQDIEKKAREMSNIITEKGEEIDNVVEELAKLCGM Vitis vinifera MDARQSDGISVLMFPWLAHGHISPFLQLAKKLSKRNFSIYFCSTPVNLDPIKGKLSESYSLSIQLVKLHLPSLPELPPQYHTTNGLPPHLMPTLKMAFDMASPNFSNILKTLHPDLLIYDFLQPWAPAASSLNIPAVQFLSTGATLQSFLAHRHRKPGI EFPFQEIHLPDYEIGRLNRFLEPSAGRISDRDRANQCLERSSRFSLIKTFREIAKYLDYVSDLTKKKMVTVGPLLQDPEDEDEATDIVEWLNKKCEASAVFVSFGSEYFVSKEEEMEIAHGLELSNVDFIWVVRFPMGEKIRLEDALPPGFLHRLGDRG MVVEGWAPQRKILGHSSIGGFVSHCGWSSVMEGMKFGVPIIAMPMHLDQPINAKLVEAVGVGREVKRDENRKLEREEIAKVIKEVVGEKNGENVRRKARELSTLRKKGDEEIDVVVEELKQLCSY Juglans regia MDTARKRIRVVMLPWLAHGHISPFLELSKKLAKRNFHIYFCSTPVNLSIKPKLSGKYSRSIQLVELHLPSLPELPPQYHTTKGLPPPHLNATLKRAFDMAGPHFSNILKTLSPDLLIYDFLQPWAPAIAASQNIPAINFLSTGAAMTSFVLHAMKKPGDEFPFPEIHLDECMKTRFVDLPEDHSPSDDHNHISDKDRALKCFERSSGFVMMKTFEELEGKYINFLSHLMQKKIVPVGPLVQNPVRGDHEKAKTLEWLDKRKQSSAVFVSFGTEYFLSKEEMEEIAYGLELSNVNFIWVVRFPEGEKVKLEEALPEGFLQRVGEKGMVVEGWAPQAKILMHPSIGGFVSHCGWSSVMESIDFGVPIVAIPMQLDQPVNAKVVEQAGVGVEVKRDRDGKLEREEVATVIREVVMGNIGESVRKKEREMRDNIRKKGEEKMDGVAQELVQLYGNGIKNV Hevea brasiliensis METLQRRKISVLMFPWLAHGHLSPFELSKKLNRFHVYFCSTPVNLDSIKPKLSAEYSFSIQLVELHLPSSPELPLHYHTTNGLPPHLMKNLKNAFDMASSSFFNILKTLKPDLLIYDFIQPWAPALASSLNIPAVNFLCTSMAMSCFGLHLNNQEAKFPFPGIYPRDYMRMKVFGALESSSNDIKDGERAGRCMDQSFHLILAKTFRELEGKYIDYLSVKLMKKI VVPGPLVQDPIFEDDEKIMDHHQVIKWLEKKERLSTVFVSFGTEYFLSTEEMEIAYGLELSKAHFIWVVRFPTGEKINLEESLPKRYLERVQERGKIVEGWAPQQKILRHSSIGGFVSHCGWSSIMESMKFGVPIIAMPMNLDQPVNSRIVEDAGVGIEVRRNKSGELEREEIAKTIRKVVVEKDGKNVSRKAREMSDTIRKKGEEEIDGVVDELLQLCDVKTNYLQ Manihot esculenta MATAQTRKISVLMFPWLAHGHLSPFLELSKKLANRNFHVYFCSTPVNLDSIKPKLSPEYHFSIQFVELHLPSSPELPSHYHTTNGLPPHLMKTLKKAFDMASSSFFNILKTLNPDLLIYDFLQPWAPALASSLNIPAVNLCSSMAMSCFGLNLNKNKEIKFLFPEIYPRDYMEMKLFRVFESSSNQIKDGERAGRCIDQSFHVILAKTFRELEGKYYVSVKCNKKIV PVGPLVEDTIHEDDEKTMDHHHHHHDEVIKWLEKKERSTTVFVSFGSEYFLSKEEMEEIAHGLELSKVNFIWVVRFPKGECINLEESLPEGYLERIQERGKIVEGWAPQRKILGHSSIGGFVSHCGWSSIMESMKLGVPIIAMPMNLDQPINSRIVEAAGVGIEVSRNQSGELEREEMAKTIRKVVVEREGVYVRRKAREMSDVLRKKGEEIDGVDELVQLCDMKTNYL Cephalotus follicularis MDLKRRSIRVLMLPWLAHGHISPFLELAKKLTNRNFLIYFCSTPINLNSIKPKLSSKYSFSIQLVELHLPSLPELPPHYHTTNGLPLHLMNTLKTAFDMASPFLNILKTLKPDLLICDHLQPWAPSLASSLNIPAIIFPTNSAIMMAFSLHHAKNPGEEFPFPSININDDMVKSINFLHSASNGLTDMDRVLQCLERSSNTMLLKTFRQLEAKYVDYSSALLKKIVLAGPLVQVPDNE DEKIEIIKWLDSRGQSSTVFVSFGSEYFLSKEEREDIAHGLELSKVNFIWVVRFPVGEKVKLEEALPNGFAERIGERGLVVEGWAPQAMILSHSSIGGFVSHCGWSSMMESKFGVPIIAMPMHIDQPLNARLVEDVGVGLEIKRNKDGRFEREELARVIKEVLVYKNGDAVRSKAREMSEHIKKNGDQEIDGVADALVKLCEMKTNSLNQD Stevia rebaudiana UGT74G1 (SEQ ID NO: 144) MAEQQKIKKSPHVLLIPFPLQGHINPFIQFGKRLISKGVKTTLVTTIHTLNSTLNHSNTTTTSIEIQAISDGCDEGGFMSAGESYLETFKQVGSKSLADLIKKLQSEGTTIDAIIYDSMTEWVLDVAIEFGIDGGSFFTQACVVNSLYYHVHKGLISLPLGETVSVPGFPVLQRWETPLILQNHEQIQSPWSQMLFGQFANIDQARWVFTNSFYKLEEEVIEWTRKIWNLKVIGPTLPSMYLDKRLDDDKDNGFNLYKANHHECMNWLDDKPKESVVYVAFGSLVKHGPEQVEEITRALIDSDVNFLWVIKHKEEGKLPENLSEVIKTGKGLIVAWCKQLDVLAHESVGCFVTHCGFNSTLEAISLGVPVVAMPQFSDQTTNAKLLDEILGVGVRVKADENGIVRRGNLASCIKMIMEEERGVIIRKNAVKWKDLAKVAVHEGGSSDNDIVEFVSELIKA Stevia rebaudiana UGT76G1 (SEQ ID NO: 145) MENKTETTVRRRRRIILFPVPFQGHINPILQLANVLYSKGFSITIFHTNFNKPKTSNYPHFTFRFILDNDPQDERISNLPTHGPLAGMRIPIINEHGADELRRELELLMLASEEDEEVSCLITDALWYFAQSVADSLNLRRLVLMTSSLFNFHAHVSLPQFDELGYLDPDDKTRLEEQASGFPMLKVKDIKSAYSNWQILKEILGKMIKQTKASSGVIWNSFKELEESELETVIREIPAPSFLIPLPKHLTASSSSLLDHDRTVFQWLDQQPPSSVLYVSFGSTSEVDEKDFLEIARGLVDSKQSFLWVVRPGFVKGSTWVEPLPDGFLGERGRIVKWVPQQEVLAHGAIGAFWTHSGWNSTLESVCEGVPMIFSDFGLDQPLNARYMSDVLKVGVYLENGWERGEIANAIRRVMVDEEGEYIRQNARVLKQKADVSLMKGGSSYESLESLVSYISSL Stevia rebaudiana UGT85C2 (Accession No. 146) MDAMATTEKKPHVIFIPFPAQSHIKAMLKLAQLLHHKGLQITFVNTDFIHNQFLESSGPHCLDGAPGFRFETIPDGVSHSPEASIPIRESLLRSIETNFLDRFIDLVTKLPDPPTCIISDGFLSVFTIDAAKKLGIPVMMYWTLAACGFMGFYHIHSLIEKGFAPLKDASYLTNGYLDTVIDWVPGMEGIRLKDFPLDWSTDLNDKVLMFTTEAPQRSHKVSHHIFHTFDELEPSIIKTLSLRYNHIYTIGPLQLLLDQIPEEKKQTGITSLHGYSLVKEEPECFQWLQSKEPNSVVYVNFGSTTVMSLEDMTEFGWGLANSNHYFLWIIRSNLVIGENAVLPPELEEHIKKRGFIASWCSQEKVLKHPSVGGFLTHCGWGSTIESLSAGVPMICWPYSWDQLTNCRYICKEWEVGLEMGTKVKRDEVKRLVQELMGEGGHKMRNKAKDWKEKARIAIAPNGSSSLNIDKMVKEITVLARN Stevia rebaudiana UGT91D1 (Accession No. 147) MYNVTYHQNSKAMATSDSIVDDRKQLHVATFPWLAFGHILPFLQLSKLIAEKGHKVSFLSTTRNIQRLSSHISPLINVVQLTLPRVQELPEDAEATTDVHPEDIQYLKKAVDGLQPEVTR FLEQHSPDWIIYDFTHYWLPSIAASLGISRAYFCVITPWTIAYLAPSSDAMINDSDGRTTVEDLTTPPKWFPFPTKVCWRKHDLARMEPYEAPGISDGYRMGMVFKGSDCLLFKCYHEFGTQWLPLLETLHQVPVVPVGLLPPEIPGDEKDETWVSIKKWLDGKQKGSVVYVALGSEALVSQTEVVELALGLELSGLPFVWAYRKPKGPAKSDSVELPDGFVERTRDRGLVWTSWAPQLRILSHESVCGFLTHCGSGSIVEGLMFGHPLIMLPIFCDQPLNARLLEDKQVGIEIPRNEEDGCLTKESVARSLRSVVVENEGEIYKANARALSKIYNDTKVEKEYVSQFVDYLEKNARAVAIDHES Stevia rebaudiana UGT91D2 (Accession No. 148) MATSDSIVDDRKQLHVATFPWLAFGHILPYLQLSKLIAEKGHKVSFLSTTRNIQRLSSHISPLINVVQLTLPRVQELPEDAEATTDVHPEDIPYLKKASDGLQPEVTRFLEQHSPDWIIYDYTHYWLPSIAASLGISRAHFSVTTPWAIAYMGPSADAMINGSDGRTTVEDLTTPPKWFPFPTKVCWRKHDLARLVPYKAPGISDGYRMGLVLKGSDCLLSKCYHEFGTQWLPLLETLHQVPVVPVGLLPPEVPGDEKDETWVSIKKWLDGKQKGSVVYVALGSEVLVSQTEVVELALGLELSGLPFVWAYRKPKGPAKSDSVELPDGFVERTRDRGLVWTSWAPQLRILSHESVCGFLTHCGSGSIVEGLMFGHPLIMLPIFGDQPLNARLLEDKQVGIEIPRNEEDGCLTKESVARSLRSVVVEKEGEIYKANARELSKIYNDTKVEKEYVSQFVDYLEKNTRAVAIDHES Stevia rebaudiana UGT91D2e (Accession No. 149) MATSDSIVDDRKQLHVATFPWLAFGHILPYLQLSKLIAEKGHKVSFLSTTRNIQRLSSHISPLINVVQLTLPRVQELPEDAEATTDVHPEDIPYLKKASDGLQPEVTRFLEQHSPDWI IYDYTHYWLPSIAASLGISRAHFSVTTPWAIAYMGPSADAMINGSDGRTTVEDLTTPPKWFPFPTKVCWRKHDLARLVPYKAPGISDGYRMGLVLKGSDCLLSKCYHEFGTQWLPLLE TLHQVPVVPVGLLPPEIPGDEKDETWVSIKKWLDGKQKGSVVYVALGSEVLVSQTEVVELALGLELSGLPFVWAYRKPKGPAKSDSVELPDGFVERTRDRGLVWTSWAPQLRILSHES VCGFLTHCGSGSIVEGLMFGHPLIMLPIFGDQPLNARLLEDKQVGIEIPRNEEDGCLTKESVARSLRSVVVEKEGEIYKANARELSKIYNDTKVEKEYVSQFVDYLEKNARAVAIDHES OsUGT1-2 (SEQ ID NO: 150) MDSGYSSSYAAAAGMHVVICPWLAFGHLLPCLDLAQRLASRGHRVSFVSTPRNISRLPPVRPALAPLVAFVALPLPRVEGLPDGAESTNDVPHDRPDMVELHRRAFDGLAAPFSE FLGTACADWVIVDVFHHWAAAAALEHKVPCAMMLLGSAHMIASIADRRLERAETESPAAAGQGRPAAAPTFEVARMKLIRTKGSSGMSLAERFSLTLSRSSLVVGRSCVEFEPETV PLLSTLRGKPITFLGLMPPLHEGRREDGEDATVRWLDAQPAKSVVYVALGSEVPLGVEKVHELALGLELAGTRFLWALRKPTGVSDADLLPAGFEERTRGRGVVATRWVPQMSIL AHAAVGAFLTHCGWNSTIEGLMFGHPPLIMLPIFGDQGPNARLIEAKNAGLQVARNDGDGSFDREGVAAAIRAVAVEEESSKVFQAKAKKLQEIVADMACHERYIDGFIQQLRSYKD Arabidopsis thaliana AAN72025.1 (SEQ ID NO: 151) MGSISEMVFETCPSPNPIHVMLVSFQGQGHVNPLLRLGKLIASKGLLVTFVTTELWGKKMRQANKIVDGELKPVGSGSIRFEFFDEEWAEDDDRRADFSLYIAHLESVGIREVSKLVRRYEEANEPVSCLINNPFIPWVCHVAEEFNIPCAVLWVQSCACFSAYYHYQDGSVSFPTETEPELDVKLPCVPVLKNDEIPSFLHPSSRFTGFRQAILGQFKNLSKSFCVLIDSFDSLEREVIDYMSSLCPVKTVGPLFKVARTVTSDVSGDICKSTDKCLEWLDSRPKSSVVYISFGTVAYLKQEQIEEIAHGVLKSGLSFLWVIRPPPHDLKVETHVLPQELKESSAKGKGMIVDWCPQEQVLSHPSVACFVTHCGWNSTMESLSSGVPVVCCPQWGDQVTDAVYLIDVFKTGVRLGRGATEERVVPREEVAEKLLEATVGEKAEELRKNALKWKAEAEAAVAPGGSSDKNFREFVEKLGAGVTKTKDNGY Arabidopsis thaliana AAF87256.1 (SEQ ID NO: 152) MGSHVAQKQHVVCVPYPAQGHINPMMKVAKLLYAKGFHITFVNTVYNHNRLLRSRGPNAVDGLPSFRFESIPDGLPETDVDVTQDIPTLCESTMKHCLAPFKELLRQINARDDVPPVSCIVSDGCMSFTLDAAEELGVPEVLFWTTSACGFLAYLYYYRFIEKGLSPIKDESYLTKEHLDTKIDWIPSMKNLRLKDIPSFIRTTNPDDIMLNFIIREADRAKRASAIILNTFDDLEHDVIQSMKSIVPPVYSIGPLHLLEKQESGEYSEIGRTGSNLWREETECLDWLNTKARNSVVYVNFGSITVLSAKQLVEFAWGLAATGKEFLWVIRPDLVAGDEAMVPPEFLTATADRRMLASWCPQEKVLSHPAIGGFLTHCGWNSTLESLCGGVPMVCWPFFAEQQTNCKFSRDEWEVGIEIGGDVKREEVEAVVRELMDEEKGKNMREKAEEWRRLANEATEHKHGSSKLNFEMLVNKVLLGE Columba livia ClUGT1 (SEQ ID NO: 153) MIHCGKKHICAFVTCILISASILMYSWKDPQLQNNITRKIFQATSALPASQLCRGKPAQNVITALEDNRTFIISPYFDDRESKVTRVIGIVHHEDVKQLYCWFCCQPDGKIYVARAKIDVHSDRFGFPYGAADIVCLEPENCNPTHVSIHQSPHANIDQLPSFKIKNRKSETFSVDFTVCISAMFGNYNNVLQFIQSVEMYKILGVQKVVIYKNNCSQLMEKVLKFYMEEGTVEIIPWPINSHLKVSTKWHFSMDAKDIGYYGQITALNDCIYRNMQRSKFVVLNDADEIILPLKHLDWKAMMSSLQEQNPGAGIFLFENHIFPKTVSTPVFNISSWNRVPGVNILQHVHREPDRKEVFNPKKMIIDPRQVVQTSVHSVLRAYGNSVNVPADVALVYHCRVPLQEELPRESLIRDTALWRYNSSLITNVNKVLHQTVL Haemophilus ducreyi LgtF Q9L875 (Accession No. 154) MPTLTVAMIVKNEAQDLAECLKTVDGWVDEIVIVDSGSTDDTLKIATQFNAKVYVNSDWQGFGPQRQFAQQYVTSDYVLWLDADERVTPELKASILQAVQHNQKNTVYKVSRLSEIFGKEIRYSGWYPDYVVRLYPTYLAKYGDELVHEKVHYPADSRVEKLQGDLLHFTYKNIHHYLVKSASYAKAWAMQRAKAGKKASLLDGVTHAIACFLKMYLFKAGFLDGKQGFLLAVLSAHSTFVKYADLWDRTRS Neisseria gonorrhoeae Q5F735 (Accession No. 155) MKKVSVLIVAKNEANHIRECIESCRFDKEVIVIDDHSADNTAEIAEGLGAKVFRRHLNGDFGAQKTFAIEQAGGEWVFLI DADERCTPELSDEISKIVRTGDYAAYFVERRNLFPNHPATHGAMRPDSVCRLMPKKGGSVQGKVHETVQTPYPERRLKHFMYHYTYDNWEQYFNKFNKYTSISAEKYREQGKPVSFVRDIILRPIWGFFKIYILNKGFLDGKMGWIMSVNHSYYTMIKYVKLYYLYKSGGKF Rhizobium meliloti (strain 1021) ExoM P33695 (SEQ ID NO: 156) MPNETLHIDIGVCTYRRPELAETLRSLAAMNVPERARLRVIVADNDAEPSARALVEGLRPEMPFDILYVHCPHSNISIARNCCLDNSTGDFLAFLDDDETVSGDWLTRLLETARTTGAAAVLGPVRAHYGPTAPRWMRSGDFHSTLPVWAKGEI RTGYTCNALLRRDAASLLGRRFKLSLGKSGGEDTDFFTGMHCAGGTIAFSPEAWVHEPVPENRASLAWLAKRRFRSGQTHGRLLAEKAHGLRQAWNIALAGAKSGFCATAAVLCFPSAARRNRFALRAVLHAGVISGLLGLKEIEQYGAREVTSA Rhizobium radiobacter Q44418 (SEQ ID NO: 157) MCRCGRAVRSRPVCRPGQLVVRRSPRPRSRNHSRCRPLRLSVFPRPHRRVRHHCQRDLRWEPGRWIAVRWKAARSHRRFRRCPFPRQLVWPVRERHRDAGDRRNQRERRRRDAYHEISEPKFRTRKRTESFWMNKAITVIVWLLVSLCVLAIITMPVSLQTHLVATAISLILLATIKSFNGQGAWRLVALGFGTAIVLRYVYWRTTSTLPPVNQLENFIPGFLLYLAEMYSVVMLGLSLVIVSMPLPSRKTRPGSPDYRPTVDVFVPSYNEDAELLANTLAAAKNMDYPADRFTVWLLDDGGSVQKRNAANIVEAQAAQRRHEELKKLCEDLDVRYLTRERNVHAKAGNLNNGLAHSTGELVTVFDADHAPARDFLLETVGYFDEDPRLFLVQTPHFFVNPDPIERNLRTFETMPSENEMFYGIIQRGLDKWNGAFFCGSAAVLRREALQDSDGFSGVSITEDCETALALHSRGWNSVYVDKPLIAGLQPATFASFIGQRSRWAQGMMQILIFRQPLFKRGLSFTQRLCYMSSTLFWLFPFPRTIFLFAPLFYLFFDLQIFVASGGEFLAYTAAYMLVNLMMQNYLYGSFRWPWISELYEYVQTVHLLPAVVSVIFNPGKPTFKVTAKDESIAEARLSEISRPFFVIFALLLVAMAFAVWRIYSEPYKADVTLVVGGWNLLNLIFAGCALGVVSERGDKSASRRITVKRRCEVQLGGSDTWVPASIDNVSVHGLLINIFDSATNIEKGATAIVKVKPHSEGVPETMPLNVVRTVRGEGFVSIGCTFSPQRAVDHRLIADLIFANSEQWSEFQRVRRKKPGLIRGTAIFLAIALFQTQRGLYYLVRARRPAPKSAKPVGAVK Streptococcus agalactiae cpsI O87183 (Accession No. 158) MIKKIEKDLISVIVPIYNVEDYLVECIESLIVQTYRNIEILLINDGSTDNCATIAKEFSERDCRVIYIEKSNGGLSEARNYGIYHSKGKYLTFVDSDDKVSSDYIANLYNAIQKHDSSIAIGGYLEFYERHNSIRNYEYLDKVIPVEEALLNMYDIKTYGSIFITAWGKLFHKSIFNDLEFALNKYHEDEFFNYKAYLKANSITYIDKPLYHYRIRVGSIMNNSDNVIIARKKLDVLSALDERIKLITSLRKYSVFLQKTEIFYVNQYFRTKKFLKQQSVMFKEDNYIDAYRMYGRLLRKVKLVDKLKLIKNRFF Streptococcus pneumoniae cps3S Q54611 (Accession No. 159) MYTFILMLLDFFQNHDFHFFMLFFVFILIRWAVIYFHAVRYKSYSCSVSDEKLFSSVIIPVVDEPLNLFESVLNRISRHKPSEIIVVINGPKNERLVKLCHDFNEKLENNMTPIQCYYTPVPGKRNAIRVGLEHVDSQSDITVLVDSDTVWTPRTLSELLKPFVCDKKIGGVTTRQKILDPERNLVTMFANLLEEIRAEGTMKAMSVTGKVGCLPGRTIAFRNIVERVYTKFIEETFMGFHKEVSDDRSLTNLTLKKGYKTVMQDTSVVYTDAPTSWKKFIRQQLRWAEGSQYNNLKMTPWMIRNAPLMFFIYFTDMILPMLLISFGVNIFLLKILNITTIVYTASWWEIILYVLLGMIFSFGGRNFKAMSRMKWYYVFLIPVFIIVLSIIMCPIRLLGLMRCSDDLGWGTRNLTE MbUGTc13 (Accession No. 160) MADAMATTEKKPHVIFIPFPAQSHIKAMLKLAQLLHHKGLQITFVNTDFIHNQFLESSGPHCLDGAPGFRFETIPDGVSHSPEASIPIRESLLRSIETNFLDRFIDLVTKLPDPPTCIISDGFLSVFTIDAAKKLGIPVMMYWTLAACGFMGFYHIHSLIEKGFAPLKDASYLTNGYLDTVIDWVPGMEGIRLKDFPLDWSTDLNDKVLMFTTEATQRSHKVSHHIFHTFDELEPSIIKTLSLRYNHIYTIGPLQLLLDQIPEEKKQTGITSLHGYSLVKEEPECFQWLQSKEPNSVVYVNFGSTTVMSLEDMTEFGWGLANSNHYFLWIIRSNLVIGENAVLPPELEEHIKKRGFIASWCSQEKVLKHPSVGGFLTHCGWGSTIESLSAGVPMICWPYSWDQLTNCRYICKEWEVGLEMGTKVKRDEVKRLVQELMGEGGHKMRNKAKDWKEKARIAIAPNGSSSLNIDKMVKEITVLARN MbUGTc19 (Accession No. 161) MANHHECMNWLDDKPKESVVYVAFGSLVKHGPEQVEEITRALIDSDVNFLWVIKHKEEGKLPENLSEVIKTGKGLIVAWCKQLDVLAHESVGCFVTHCGFNSTLEAISLGVPVVAMPQFSDQTTNAKLLDEILGVGVRVKADENGIVRRGNLASCIKMIMEEERGVIIRKNAVKWKDLAKVAVHEGGSSDNDIVEFVSELIKAGSGEQQKIKKSPHVLLIPFPLQGHINPFIQFGKRLISKGVKTTLVTTIHTLNSTLNHSNTTTTSIEIQAISDGCDEGGFMSAGESYLETFKQVGSKSLADLIKKLQSEGTTIDAIIYDSMTEWVLDVAIEFGIDGGSFFTQACVVNSLYYHVHKGLISLPLGETVSVPGFPVLQRWETPLILQNHEQIQSPWSQMLFGQFANIDQARWVFTNSFYKLEEEVIEWTRKIWNLKVIGPTLPSMYLDKRLDDDKDNGFNLYKA MbUGT1-3 (SEQ ID NO: 162) MENKTETTVRRRRRIILFPVPFQGHINPILQLANVLYSKGFSITIFHTNFNKPKTSNYPHFTFRFILDNDPQDERISNLPTHGPLAGMRIPIINEHGADELRRELELLML ASEEDEEVSCLITDALWYFAQSVADSLNLRRLVLMTSSLFNFHAHVSLPQFDELGYLDPDDKTRLEEQASGFPMLKVKDIKSAYSNWQILKEILGKMIKQTKASSGVIWN SFKELEESELETVIREIPAPSFLIPLPKHLTASSSSLLDHDRTVFQWLDQQPPSSVLYVSFGSTSEVDEKDFLEIARGLVDSKQSFLWVVRPGFVKGSTWVEPLPDGFLG ERGRIVKWVPQQEVLAHGAIGAFWTHSGWNSTLESVCEGVPMIFSDFGLDQPLNARYMSDVLKVGVYLENGWERGEIANAIRRVMVDEEGEYIRQNARVLKQKADVSLMK GGSSYESLESLVSYISSL MbUGT1-2 (SEQ ID NO: 163) MATKGSSGMSLAERFWLTLSRSSLVVGRSCVEFEPETVPLLSTLRGKPITFLGLMPPLHEGRREDGEDATVRWLDAQPAKSVVYVALGSEVPLGVEKVHELALGLELAGTRFLWALRKPTGVSDADLLPAGFEERTRGVVATRWVPQMSILAHAAVGAFLTHCGWNSTIEGLMFGHPLIMLPIFGDQGPNARLIEAKNAGLQVRNDDGGSDFDREGVAAAIRAVAEE SSKVFQAKAKKLQEIVADMACHERYIDGFIQQLRSYKDDSGYSSSYAAAAGMHVVICPWLAFGHLLPCLDLAQRLASRGHRVSFVSTPRNISRLPPVRPLAPLVAFVALPLPRVEGLPDGAESTNDVPHDRPDMVELHRRAFDGLAAPFSEFLGTACADWVIVDVFHHWAAAAAALEHKVPCAMMLLGSAEMISIADERLEHAETESPAAAGQGRPAAAPTFEVARMKLIR Coffea arabica (CaUGT_1,6) MAENHATFNVLMLPWLAHGHVSPYLELAKKLTARNFNVYLCSSPATLSSVRSKLTEKFSQSIHLVELHLPKLPELPAEYHTTNGLPPHLMPTLKDAFDMAKPNFCNVLKSLKPDLLIYDLLQPWAPEAASAFNIPAVVFISSSATMTSFGLHFFKNPGTKYPYGNAIFIRDYESVFVENLTRRDRDTYRVINCMERSSKIILIKGFNEIEGKYFDYFSCLTGKKVV PVGPLVQDPVLDDEDCRIMQWLNKKEKGSTVFVSFGSEYFLSKKDMEEIAHGLEVSNVDFIWVVRFPKGENIVIEETLPKGFFERVGERGLVVNGWAPQAKILTHPNVGGFVSHCGWNSVMESMKFGLPIIAMPMHLDQPINARLIEEVGAGVEVLRDSKGKLHRERMAETINKVMKEASGESVRKKARELQEKLELKGDEEIDVVKELVQLCATKNKRNGLHYY Stevia rebaudiana UGT85C1 MADQMAKIDEKKPHVVFIFPPAQSHIKCMLKLARILHQKGLYITFINTDTNHERLVASGGTQWLENAPGFWFKTVPDGFGSAKDDGVKPTDALRELMDYLKTNFFDLFLDLVLKLEVPATCIICDGCMTFANTIRAAEKNLIPVILFWTMAACGFMAFYQAKVLKEKEIVPVKDETYLTNGYLDMEIDWIPGMKRIRLRDLPEFILATKQNYFAFEFLFETAQLADKVSHMIIHTFEELEAS LVSEIKSIFPNVYTIGPLQLLLNKITQKETNNDSYSLWKEEPECVEWLNSKEPNSVVYVNFGSLAVMSLQDLVEFGWGLVNSNHYFLWIIRANLIDGKPAVMPQELKEAMNEKGFVGSWCSQEEVLNHPAVGGFLTHCGWGSIIESLSAGVPMLGWPSIGDQRANCRQMCKEWEVGMEIGKNVKRDEVEKLVRLMEGLEGERMRKKALEWKKSATLATCCNGSSLDVEKLANEIKKLSRN Arabidopsis thaliana AtUGT73C3 MATEKTHQFHPSLHFVLFFPMAQGHMIPMIDARLLAQRGVTITIVTTPHNAARFKNVLNRAIESGLAINILHVKFPYQEFGLPEGKENIDSLDSTELMVPFFKAVNLLEDPVMKLMEEMKPRPSCLISDWCLPYTSIIAKNFNIPCIVFHGMGCFNLLC MHVLRRNLEILENVKSDEEYFLVPSFPDRVEFTKLQLPVKANASGDWKEIMDEMVKAEYTSYGVIVNTFQELEPPYVKDYKEAMDGKVWSIGPVSLCNKAGADKAERGSKAAIDQDECLQWLDSKEEGSVLYVCLGSICNLPLSQLKEGLGLLEESRRSF IWVIRGSEKYKELFEWMLESGFEERIKERGLLIKGWAPQVLILSHPSVGGFLTHCGWNSTLEGITSGIPLITWPLFGDQFCNQKLVVQVLKAGVSAGVEEVMKWGEEDKIGVLVDKEGVKKAVEELMGDSDDAKERRRRVKELGELAHKAVEKGGSSHSNITLLLQDIMQLAQFKN Hordeum vulgare subsp. Vulgare HvUGT_B1 (Accession No. 204) MAQAESERMRVVMFPWLAHGHINPYLELAKRLIASASGDHHLDVVVHLVSTPANLAPLAHHQTDRLRLVELHLPSLPDLPPALHTTKGLPARLMPVLKRACDLAAPRFGALLDELCPDILVYDFIQPWAPLEAEARGVPAFHFATCGAAATAFFIHCLKTDRPPSAFPFESISLGGVDEDAKYTALVTVREDSTALVAERDRLPLSLERSSGFVAVKSSADIERKYMEYLSQLLGKEIIPTGPLLVDSGGSEEQRDGGRIMRWLDGEEPGSVVFVSFGSEYFMSEHQMAQMARGLELSGVPFLWVVRFPNAEDDARGAARSMPPGFEPELGLVVEGWAPQRRILSHPSCGAFLTHCGWSSVLESMAAGVPMVALPLHIDQPLNANLAVELGAAAARVKQERFGEFTAEEVARAVRAAVKGKEGEAARRRARELQEVVARNNGNDGQIATLLQRMARLCGKDQAVPN Hordeum vulgare subsp. Vulgare HvUGT_B3 (Accession No. 205) MAEANDGGKMHVVMLPWLAFGHVLPFTEFAKRVARQGHRVTLLSAPRNTRRLIDIPPGLAGLIRVVHVPLPRVDGLPEHAEATIDLPSDHLRPCLRRAFDAAFERELSRLLQEEAKPDWVLVDYASYWAPTAAARHGVPCAFLSLFGAAALSFFGTPETLLLGIGRHAKTEPAHLTVVPEYVPFPTTVAYRGYEARELFEPGMVPDSGVSEGYRFAKTIEGCQLVGIRSSSEFEPEWL RLLGELYRKPVIPVGLFPPAPQDDVAGHEATLRWLDGQAPSSVVYAAFGSEVKLTGAQLQRIALGLEASGLPFIWAFRAPTSTETGAASGGLPEGFEERLAGRVCRGWVPQVKFLAHASVGGFLTHAGWNSIAEGLAHGVRLVLLPLVFEQGLNARNIVDKNIGVEVARDEQDGSFAAGDIAAALRRVMVEDEGEGFGAKVKELAKVFGDDEVNDQCVREFLMHLSDHSKKNQGQD MbUGT1,2.2(sequence no.206) MATKGSSGMSLAERFWLTLSRSSLVVGRSCVEFEPETVPLLSTLRGKPITFLGLMPPLHEGRREDGEDATVRWLDAQPAKSVVYVALGSEVPLGVEKVHELALGLELAGTRFLWALRKPTGVSDADLLPAGFEERTRGRGVVATRWVPQMSILAHAAVGAFLTHCGWNSTIEGLMFGHPLIMLPIFGDQGPNARLIEAKNAGLQVARNDDGGSDFDREGVAAAIRAVAVEEE SSKVFQAKAKKLQEIVADMACHERYIDGFIQQLRSYKDDSGYSSSYAAAAGMHVVICPWLAFGHLLPCLDLAQRLASRGHRVSFVSTPRNISRLPPVRPALAPLVAFVALPLPRVEGLPDGAESTNDVPHDRPDMVELHRRAFDGLAAPFSEFLGTACADWVIVDVFHHWAAAAAALEHKVPCAMMLLGSAEMIASIADERLEHAETESPAAAGQGRPAAAPTFEVARMKLIR Coffea canephora (CcUGT_1,6) (207) MAENHATFNVLMLPWLAHGHVSPYLELAMKLTARNFNVYLCSSPATLSSVRSKLTEKFSQSIHLVELHLPKLPELPAEYHTTNGLPPHLMPTLKDAFDMAKPNFCNVLKSLKPDLLIYDL LQPWAPEAASAFNIPAVVFISSSATMTSFGLHFFKNPGTKYPYGNTIFYRDYESVFVENLKKRDRDTYRVVNCMERSSKIILIKGFKEIEGKYFDYFSCLTGKKVVPVGPLVQDPVLDDEDCRIMQWLNKKEKGSTVFVSFGSEYFLSKEDMEEIAHGLELSNVDF IWVVRFPKGENIVIEETLPKGFFERVGERGLVVNGWAPQAKILTHPNVGGFVSHCGWNSVMESMKFGLPIVAMPMHLDQPINARLIEEVGAGVEVLRDSKGKLHRERMAETINKVTKEASGEPARKKARELQEKLELKGDEEIDDVVKELVQLCATKNKRNGLHCYN Coffea eugenioides (CeUGT_1,6) (208) MAENHATFNVLMLPWLAHGHVSPYLELAKKLTARNFNVYLCSSPATLSSVRSKLTEKFSQSIHLVELHLPKLPELPAEYHTTNGLPPHLMPTLKDAFDMAEPNFCNVLKSLKPDLLIYDLLQPWAPEAASAFNIPAVVFISSSATMTSFGLHFFKNPGTKYPYGNTIFYRDYESVFVENLKRRDRDTYRVVNCMERSSKIILIKGFKEIEGKYFDYFSCLTGKKVVPVGPLVQDPVLDDEDCRIMQWLNKKEKGSTVFVSFGSEYFLSKEDMEEIAHGLELSNVDFIWVVRFPKGENIVIEETLPKGFFERVGERGLVVNGWAPQAKILTHPNVGGFVSHCGWNSVMESMKFGLPIIAMPMHLDQPINARLIEEVGAGVEVLRDSKGKLHRERMAETINKVTKEASGESVRKKARELQEKLELKGDEEIDDVVKELVQLCATKNKRNGLHYN Coffea eugenioides (CeUGT_1,6.2) (209) MAENHATFNVLMLPWLAHGHVSPYLELAKKLTARNFNVYLCSSPATLSSVRSKLTEKFSQSIHLVELHLPKLPELPAEYHTTNGLPPHLMPTLKDAFDMAKPNFCNVLKSLKPDLLIYDLLQPWAPEAASAFNIPAVVFISSSATMTSFGLHFFKNPGTKYPYGNAIFYRDYESVFVENLTRRDRDTYRVINCMERSSKIILIKGFNEIEGKYFDYFSCLTGKKVVPVGPLVQDPVLDDEDCEIMQWLNKKEKVSTVFVSFGSEYFLSKKDMEEIAHGLELSNVDFIWVVRFPKGENIVIEETLPKGFFERVGERGLVVNGWAPQAKILTHPNVGGFVSHCGWNSVMESMKFGLPIIAMPMHLDQPINARLIEEVGAGVEVLRDSKGKLHRERMAETINKVMKEASGESVRKKARELQEKMDLKGDEEIDDVVKELVQLCATKNKRNGLHYY Siraitia grosvenorii (SgUGT94-289-3.2) (210) MADAAQQGDTTTILMLPWLGYGHLSAFLELAKSLSRRNFHIYFCSTSVNLDAIKPKLPSSFDSSIQFVELHLPSSPEFPPHLHTTNGLPPTLMPALHQAFSMAAQHFESILQTLAPHLLIYDSLQPWAPRVASSLKIPAINNFNTTGVFVISQGLHPIHYPHSKFPFSEFVLHNHWKAMYSTADGASTERTRKRGEAFLYCLHASCSVILINSFRLEGKYMDYLSV LLNKKVVPVGPLVYEPNQDGEDEGYSSIKNWLDKKEPSSTVFVSFGSEYFPSKEEMEEIAHGLEASEVNFIWVVRFPQGDNTSGIEDALPKGFLERAGERGMVVKGWAPQAKILKHWSTGGFVSHCGWNSVMESMMFGVPIIGVPMHVDQPFNAGLVEEAGVGVEAKRDPDGKIQRDEVAKLIKEVVVEKTREDVRKKAREMSEILSKGEEKFDEMVAEISLLLKI Oryza sativa (OsJUGT_1,6) MAQAEERRLRVLMFPWLAHGHINPYLELATRLTTTSSSQIDVVVHLVSTPVNLAAVAHRRTDRISLVELHLPELPGLPPALHTTKHLPPRLMPALKRACDLAAPAFGALLDELSPDVVLYDFIQPWAPLEAAARGVPAVHFSTCSAAATAFFLHFLDGGGGGGGRGAFPFEAISLGGAEEDARYTMLTCRDDGTALLPKGERLPLSFARSSEFVAVKTCVEIESKYMDYLSKLVGKE IIPCGPLLVDSGDVSAGSEADGVMRWLDGQEPGSVVLVSFGSEYFMTEKQLAEMARGLELSGAAFVWVVRFPQQSPDGDEDDHGAAAARMPPGFAPARGLVVEGWAPQRRVLSHRSCGAFLTHCGWSSVMESMSAGVPMVALPLHIDQPVGANLAAELGVAARVRQERFGEFEAEEVARAVRAVMRGGEALRRATELREVVARRDAECDEQIGALLHRMARLCGKGTGRAQLGH Panax ginseng (PsUGT94_B1) MADNQNGRISIALLPFLAHGHISPFFELAKQLAKRNCNVFLCSTPINLSSIKDKDSSASIKLVELHLPSSPDLPPHYHTTNGLPSHLMLPLRNAFETAGPTFSEILKTLNPDLLIYDFNPSWAPEIASSHNIPAVYFLTTAAASSSIGLHAFKNPGEKYPFPDFYDNSNITPEPPSADMKLLHDFIACFERSCDIILIKSFRELEGKYIDLLSLSDKTL VPVGPLVQDPMGHNEDPKTEQIINWLDKRAESTVVFVCFGSEYFLSNEELEEVAIGLEISTVNFIWAVRLIEGEKKGILPEGFVQRVGDRGLVVEGWAPQARILGHSSTGGFVSHCGWSSIAESMKFGVPVIAMARHLDQPLNGKLAAEVGVGMEVVRDENGKYKREGIAEVIRKVVVEKSGEVIRRKARELSEKMKEKGEQEIDRALEELVQICKKKDEQ Stevia rebaudiana (SrUGT73E1, optionally with a His tag) (SEQ ID NO: 214) MAHHHHHHVGTGSNDDDDKSPDPNWASTSELVFIPSPGAGHLPPTVELAKLLLHRDQRLSVTIIVMNLWLGPKHNTEARPCVPSLRFVDIPCDESTMALISPNTFISAFVEHHKPRVRDIVRGI IESDSVRLAGFVLDMFCMPMSDVANEFGVPSYNYFTSGAATLGLMFHLQWKRDHEGYDATELKNSDTELSVPSYVNPVPAKVLPEVVLDKEGGSKMFLDLAERIRESKGIIVNSCQAIERHALEY LSSNNNGIPPVFPVGPILNLENKKDDAKTDEIMRWLNEQPESSVVFLCFGSMGSFNEKQVKEIAVAIERSGHRFLWSLRRPTPKEKIEFPKEYENLEEVLPEGFLKRTSSIGKVIGWAPQMAVLS HPSVGGFVSHCGWNSTLESMWCGVPMAAWPLYAEQTLNAFLLVVELGLAAEIRMDYRTDTKAGYDGGMEVTVEEIEEDGIRKLMSDGEIRNKVKDVKEKSRAAVVEGGSSYASIGKFIEHVSNVTI Oryza sativa (OsUGT1-2) (SEQ ID NO: 215) MADSGYSSSYAAAAGMHVVICPWLAFGHLLPCLDLAQRLASRGHRVSFVSTPRNISRLPPVRPALAPLVAFVALPLPRVEGLPDGAESTNDVPHDRPDMVELHRRAFDGLAAPFSEFLGTACADWVIVDVFHHWAAAAALEHKVPCAMMLLGSAHMIASIADRRLERAETESPAAAGQGR PAAAPTFEVARMKLIRTKGSSGMSLAERFSLTLSRSSLVVGRSCVEFEPETVPLLSTLRGKPITFLGLMPPLHEGRREDGEDATVRWLDAQPAKSVVYVALGSEVPLGVEKVHELALGLELAGTRFLWALRKPTGVSDADLLPAGFEERTRGRGVVATRWVPQMSILAHAAVGAFLTHCG WNSTIEGLMFGHPLIMLPIFGDQGPNARLIEAKNAGLQVARNDGDGSFDREGVAAIRAVAVEEESSKVFQAKAKKLQEIVADMACHERYIDGFIQQLRSYKD Camelina sativa (XP_010516905.1) MASEKTLQVHPPLHFVLFPFMAQGHMIPMVDIARLLAQRGATTVIVTTRYNAGRFENVLSRAVESGLPINIVHVKFPYEEVGLPKGKENIDSLDSMELMVPFFKAVNMLQDPVVKLMEEMESRPSIISDLLPYTSKIAKKFNIPKIVFHGISCFCLCVHVLRRNELITNLKSDKEYFLVPSFPDRVEFTKPQVTVETNASGDWKEFLDEMVEAEDTSYGVIINTFEELEPAYVKDYKDARAGNVWSIGPVSLCNKAGVDKAERGNKATIDQDECLKWLDSKEEGSVLYVCLGSICNLPLVQLKELGLGLEESQRPFIWVIRGWEKYNELSEWMVESGFEERIRERGLLIRWAPQVLILSHPSVGGFLTHCGWNSTVEGITSGVPLITWPLFGDQFCNQTLVVQVLKAGVSVGVEEVMKWGEEEKIGVLVDKEGVKKAVEDLMGESDDAKERTKRVKELGGLAHKAVEEGGSSHSNITLFQDIRQVQVQSV Glycyrrhiza uralensis (UGT73F24) (SEQ ID NO: 217) MADVAEEQPLKIYFIPYLAAGHMIPLCDIATLFASRGHHVTIITTPSNAQTLRESHHFRVQTIQFPSQEVGLPAGVQNLTAVTNLDDSYKIYHATMLLRKHIEDFVERDPDCIVADFLFPWVDDVATKLHIPRLVFTNGFTLFTICAMESHKAHPLPVDAASGSFVIPDFPHHVTINSTPPKRTKEFVDPLLTEAFKSHGFLINSFVELDGEECVEHYERITGGHKAWHLGPAFLVHRTAQDRGEKSVVSTQECLSWLDSKRDNSVLYICFGTICYFPDKQLYEIASAIEASGHEFIWVVPEKRGNADESEEEKEKWLPKGFEERNNGKKGMIIRGWAPQVAILGHPAVGGFLTHCGWNSTVEAVSAGVPMITWPVHSDQYFNEKLITQVRGIGVEVGAEEWIVTAFRETEKLVGRDRIERAVRRVMDGGGDEAVQIRRRARELGEMARQAVQEGGSSHTNTALINDLKRWRDSKQLN Glycyrrhiza uralensis (UGT73C33) (SEQ ID NO:218) MAVFQANQPHFVLFPLMAQGHIIPMIDIARLLAQRGAIVTIFTPKNASRFTSVLSRAVSSGLQIRLVHLHFPSKEAGLPEGCENLDMVASHDMICNIFQAIRMLQKQAEELFETLTPKPSCIISDFCIPWTTQVAEKHHIPRISFHGFSCFCLHCMLKIHTSKVLEGITSEYFTVPGIPDQIQVTKQQVPGPMIDEMKEFGEQMRDAEIRSYGVIINTFEELEKAYVNDYKKERNGKVWCIGPVSLCNKDGLDKAQRGNKASISEHHCLEWLDLQQPNSVIYVCLGSLCNLTPPQLMELALGLEATKRPFIWVIREGNKFEELEKWISEEGFEERIKGRGLIIRGWAPQVLILSHPSIGGFLTHCGWNSTLEGVTAGVPMVTWPLFADQFLNEKLVTQVLRIGVSLGVDVPLKWGEEEKVGVQVKKEGIEKAICMVMDEGEESKERRERAKELSEMAKRAVEKDGSSHLNMTMLIQDIMQQSSSKVET
Claims
1. 1. A method for producing mogrol or mogroside, comprising: The present invention provides a recombinant microbial host cell expressing a heterologous enzyme pathway that catalyzes the conversion of isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP) to mogrol or mogroside, the pathway comprising: (A) at least two squalene epoxidase enzymes (SQEs) that convert squalene to 2,3;22,23 dioxidosqualene; (B) at least one triterpene cyclase enzyme that converts 22,23-dioxidosqualene to 24,25-epoxycucurbitadienol, wherein the at least one triterpene cyclase enzyme comprises an amino acid sequence that is at least 70% identical to one of SEQ ID NO:191, SEQ ID NO:192, and SEQ ID NO:193; (C) at least one epoxide hydrolase that converts 24,25-epoxycucurbitadienol to 24,25-dihydroxycucurbitadienol, wherein the at least one epoxide hydrolase comprises an amino acid sequence at least 70% identical to any one of SEQ ID NOs: 189, 58, 184, 185, 187, 188, 190, and 212; (D) a cytochrome P450 enzyme comprising an amino acid sequence having at least 70% sequence identity to an amino acid sequence selected from SEQ ID NO: 194 and SEQ ID NO: 171; and (E) at least one uridine diphosphate-dependent glycosyltransferase (UGT) enzyme comprising an amino acid sequence having at least 70% sequence identity to any one of SEQ ID NOs: 164, 165, 138, 204-211, 213-218; providing the recombinant microbial host cell comprising at least one of: and culturing the host cell under conditions to produce mogrol or mogroside.
2. 2. The method of claim 1, wherein the at least one squalene epoxidase comprises an amino acid sequence at least 70% identical to any one of SEQ ID NOs: 17-39, 168-170, and 177-183.
3. 3. The method of claim 2, wherein the at least one squalene epoxidase comprises an amino acid sequence at least 70% identical to SEQ ID NO:
39.
4. 4. The method of claim 3, wherein the at least one SQE comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
39.
5. 4. The method of claim 3, wherein the SQE comprises an amino acid sequence having 1 to 20 amino acid modifications relative to SEQ ID NO: 39, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
6. 4. The method of claim 3, wherein the host cell comprises two squalene epoxidase enzymes, each comprising an amino acid sequence that is at least 70% identical to SEQ ID NO:
39.
7. 7. The method of claim 6, wherein one of the SQE enzymes has one or more amino acid modifications that increase its specificity or productivity for converting 2,3-oxidosqualene to 2,3;22,23 dioxidosqualene compared to an enzyme having the amino acid sequence of SEQ ID NO:
39.
8. The amino acid modifications to the squalene epoxidase are at one or more of the following positions: 35, 133, 163, 254, 283, 380, and 395 of SEQ ID NO:
39. The method of claim 6 or 7, comprising the above modifications.
9. 9. The method of claim 8, wherein the amino acid modifications to the squalene epoxidase comprise two, three, four, five, or six amino acid modifications selected from substitutions at positions corresponding to positions 35, 133, 163, 254, 283, 380, and 395 of SEQ ID NO:
39.
10. The amino acid modification is The amino acid at the position corresponding to position 35 of SEQ ID NO: 39 is arginine or lysine; The amino acid at the position corresponding to position 133 of SEQ ID NO: 39 is glycine, alanine, leucine, isoleucine, or valine; The amino acid at the position corresponding to position 163 of SEQ ID NO: 39 is glycine, alanine, leucine, isoleucine, or valine; The amino acid at the position corresponding to position 254 of SEQ ID NO: 39 is phenylalanine, alanine, leucine, isoleucine, or valine; The amino acid at the position corresponding to position 283 of SEQ ID NO: 39 is alanine, leucine, isoleucine, or valine; The amino acid at the position corresponding to position 380 of SEQ ID NO: 39 is alanine, leucine, or glycine; and The amino acid at the position corresponding to position 395 of SEQ ID NO: 39 is tyrosine, serine, or threonine. The method according to claim 8 or 9, wherein the compound is selected from the group consisting of:
11. 11. The method of claim 10, wherein the squalene epoxidase comprises the amino acid substitutions numbered according to SEQ ID NO: 39: H35R, F163A, M283L, V380L, and F395Y.
12. 11. The method of claim 10, wherein the squalene epoxidase comprises the amino acid substitutions numbered according to SEQ ID NO: 39: H35R, N133G, F163A, Y254F, V380L, and F395Y.
13. 13. The method of any one of claims 1 to 12, wherein the heterologous enzyme pathway further comprises squalene synthase (SQS).
14. 14. The method of claim 13, wherein the SQS comprises an amino acid sequence that is at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 2-16, 166, and 167.
15. 15. The method of claim 14, wherein the SQS comprises an amino acid sequence that is at least 70% identical to SEQ ID NO:
11.
16. 16. The method of claim 15, wherein the SQS comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
11.
17. 16. The method of claim 15, wherein the SQS comprises an amino acid sequence having 1 to 20 amino acid modifications relative to SEQ ID NO: 11, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
18. The SQS is selected from the group consisting of SEQ ID NO: 2, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 166, and SEQ ID NO:
167.
15. The method of claim 14, comprising an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to sequence number 167.
19. 19. The method of any one of claims 1 to 18, wherein the heterologous enzyme pathway comprises at least one triterpene cyclase (TTC).
20. 20. The method of claim 19, wherein the at least one TTC comprises an amino acid sequence that is at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 40-55, 191-193, and 219-220.
21. 21. The method of claim 20, wherein the heterologous enzyme pathway comprises at least two enzymes that have triterpene cyclase activity and that convert 22,23-dioxidosqualene to 24,25-epoxycucurbitadienol.
22. 22. The method of claim 20 or 21, wherein the TTC comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of SEQ ID NO:
40.
23. 23. The method of claim 22, wherein the TTC comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO:
40.
24. 24. The method of claim 23, wherein the TTC comprises an amino acid sequence having 1 to 20 amino acid modifications relative to SEQ ID NO: 40, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
25. 25. The method of any one of claims 19 to 24, wherein the heterologous enzyme pathway comprises at least one TTC, wherein the TTC comprises an amino acid sequence that is at least 70% identical to one of SEQ ID NO:191, SEQ ID NO:192, and SEQ ID NO:
193.
26. 26. The method of claim 25, wherein at least one TTC is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NO: 191, SEQ ID NO: 192, and SEQ ID NO:
193.
27. 27. The method of claim 26, wherein the TTC comprises an amino acid sequence having 1 to 20 amino acid modifications relative to one of SEQ ID NOs: 191, 192, and 193, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
28. 28. The method of any one of claims 1 to 27, wherein the heterologous pathway comprises an enzyme that converts cucurbitadienol to 24,25-epoxycucurbitadienol.
29. 29. The method of claim 28, wherein the enzyme that converts cucurbitadienol to 24,25-epoxycucurbitadienol comprises an amino acid sequence having at least about 70% sequence identity to SEQ ID NO:
221.
30. 30. The method of any one of claims 1 to 29, wherein the heterologous enzyme pathway comprises an epoxide hydrolase (EPH).
31. The EPH is an amino acid sequence selected from SEQ ID NOs: 56-72, 184-190, and 212.
31. The method of claim 30, wherein the amino acid sequence is at least 70% identical to the amino acid sequence of
32. 32. The method of claim 31 , wherein the EPH comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 56-72, 184-190, and 212.
33. 33. The method of claim 32, wherein the EPH comprises an amino acid sequence having 1 to 20 amino acid modifications relative to one of SEQ ID NOs: 56-172, 184-190, and 212, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
34. 32. The method of claim 31 , wherein the heterologous pathway comprises at least one EPH that converts 24,25-epoxycucurbitadienol to 24,25-dihydroxycucurbitadienol, and the at least one EPH comprises an amino acid sequence at least 70% identical to one of SEQ ID NOs: 189, 58, 184, 185, 187, 188, 190, and 212.
35. 35. The method of claim 34, wherein the EPH comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 189, 58, 184, 185, 187, 188, 190, and 212.
36. 36. The method of claim 35, wherein the EPH comprises an amino acid sequence having 1 to 20 amino acid modifications relative to one of the amino acids of SEQ ID NOs: 189, 58, 184, 185, 187, 188, 190, and 212, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
37. 37. The method of any one of claims 1 to 36, wherein the heterologous pathway comprises one or more oxidases that oxidize C11 of C24,25 dihydroxycucurbitadienol to produce mogrol.
38. 38. The method of claim 37, wherein the at least one oxidase is a cytochrome P450 enzyme.
39. 39. The method of claim 38, wherein the at least one cytochrome P450 enzyme comprises an amino acid sequence that is at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 73-91, 171-176, and 194-200.
40. 40. The method of claim 39, wherein the at least one cytochrome P450 enzyme comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 73-91, 171-176, and 194-200.
41. 41. The method of claim 40, wherein the at least one cytochrome P450 enzyme comprises an amino acid sequence having from 1 to 20 amino acid modifications relative to one of SEQ ID NOs: 73-91, 171-176, and 194-200, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
42. The cytochrome P450 is an amino acid sequence selected from SEQ ID NO: 194 and SEQ ID NO:
171.
42. The method of any one of claims 37 to 41, comprising an amino acid sequence that is at least 70% identical to the amino acid sequence of claim 37.
43. 43. The method of claim 42, wherein the cytochrome P450 enzyme comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 194 and 171.
44. 44. The method of claim 43, wherein the at least one cytochrome P450 enzyme comprises an amino acid sequence of one of SEQ ID NOs: 194 and 171 with 1 to 20 amino acid modifications, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
45. 45. The method of any one of claims 42 to 44, wherein the cytochrome P450 enzyme has at least a portion of its transmembrane domain replaced with a heterologous transmembrane domain.
46. 38. The method of claim 37, wherein the at least one oxidase is a non-heme iron oxidase.
47. 47. The method of claim 46, wherein the non-heme iron oxidase comprises an amino acid sequence that is at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 100-115.
48. 48. The method of any one of claims 37 to 47, wherein the microbial host cell expresses one or more electron transfer proteins selected from cytochrome P450 reductase (CPR), flavodoxin reductase (FPR), and ferredoxin reductase (FDXR) sufficient to regenerate one or more oxidases.
49. 49. The method of claim 48, wherein the microbial host cell expresses a cytochrome P450 reductase comprising an amino acid sequence that is at least 70% identical to one of SEQ ID NOs: 92-99 and 201.
50. 50. The method of claim 49, wherein the cytochrome P450 reductase comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to at least one of SEQ ID NOs: 92-99 and 201.
51. 50. The method of claim 49, wherein the microbial host cell expresses SEQ ID NO: 194 or a derivative thereof and SEQ ID NO: 98 or a derivative thereof.
52. 50. The method of claim 49, wherein the microbial host cell expresses SEQ ID NO: 171 or a derivative thereof and SEQ ID NO: 201 or a derivative thereof.
53. 53. The method of any one of claims 1 to 52, wherein the heterologous enzyme pathway comprises one or more uridine diphosphate-dependent glycosyltransferase (UGT) enzymes, thereby producing one or more mogrol glycosides.
54. 54. The method of claim 53, wherein the one or more mogrol glycosides are selected from Mog.II-E, Mog.III, Mog.III-A1, Mog.III-A2, Mog.III, Mog.IV, MogIV-A, siamenoside, Mog.V, and Mog.VI.
55. 54. The method of claim 53, wherein the one or more mogrol glycosides include Mog. VI, Isomog. V, and Mog. V.
56. 54. The method of claim 53, wherein the host cell produces Mog. V or siamenoside.
57. 57. The method of any one of claims 53-56, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 116-165, 202-210, 211, 213-218.
58. 58. The method of claim 57, wherein the at least one UGT enzyme comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs:116-165, 202-210, 211, 213-218.
59. 59. The method of claim 58, wherein the at least one UGT enzyme comprises an amino acid sequence having from 1 to 20 amino acid modifications relative to one of SEQ ID NOs:116-165, 202-210, 211, 213-218, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
60. 60. The method of any one of claims 53-59, wherein the at least one uridine diphosphate-dependent glycosylation (UGT) enzyme comprises an amino acid sequence having at least 70% sequence identity to one of SEQ ID NOs: 164, 165, 138, 204-211, and 213-218.
61. 61. The method of Claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:
165.
62. 62. The method of claim 61, wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:165, or comprises an amino acid sequence that has from 1 to 20 amino acid modifications relative to SEQ ID NO:165, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
63. 63. The method of claim 62, wherein the at least one UGT enzyme comprises a substitution at one or more of positions 41, 49, and 127 relative to SEQ ID NO: 165, and optionally the one or more substitutions comprise L41F, D49E, and C127F.
64. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:
164.
65. 65. The method of claim 64, wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 164, or comprises an amino acid sequence that has from 1 to 20 amino acid modifications relative to SEQ ID NO: 164, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
66. The at least one UGT enzyme comprises one or more substitutions shown in Table 3 relative to SEQ ID NO:164, and optionally, S150F, T147L, N2 66. The method of claim 65, wherein the antibody has one or more amino acid substitutions selected from: O5K, K270E, V281L, L354V, L13F, T32A, and K101A.
67. 60. The method of Claim 59, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:
138.
68. 68. The method of claim 67, wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO: 138, or comprises an amino acid sequence that has from 1 to 20 amino acid modifications relative to SEQ ID NO: 138, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
69. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:
204.
70. 70. The method of claim 69, wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:204, or comprises an amino acid sequence that has from 1 to 20 amino acid modifications relative to SEQ ID NO:204, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
71. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:
205.
72. 72. The method of claim 71, wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:205, or comprises an amino acid sequence that has from 1 to 20 amino acid modifications relative to SEQ ID NO:205, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
73. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:
206.
74. 74. The method of claim 73, wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:206, or comprises an amino acid sequence that has from 1 to 20 amino acid modifications relative to SEQ ID NO:206, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
75. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:207, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
207.
76. The at least one UGT enzyme is at least 70% identical to SEQ ID NO:
208.
61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
208.
77. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:209, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
209.
78. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:210, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
210.
79. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:211, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
211.
80. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:213, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
213.
81. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:214, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
214.
82. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:215, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
215.
83. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:218, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
218.
84. The at least one UGT enzyme comprises an amino acid sequence that is at least 70% identical to SEQ ID NO:217, or the at least one UGT enzyme comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 100% identical to SEQ ID NO:
217.
61. The method of claim 60, comprising an amino acid sequence that is at least 95%, or at least 98%, or at least 99% identical to said sequence, and optionally having one or more amino acid substitutions selected from A74E, I91F, H101P, Q241E, and I436L.
85. 61. The method of claim 60, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:216, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
216.
86. 86. The method of any one of claims 60-85, wherein the at least one UGT enzyme further comprises an amino acid sequence at least 70% identical to SEQ ID NO:
146.
87. 87. The method of claim 86, wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:146, or the at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO:146, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
88. 88. The method of any one of claims 60-87, wherein the at least one UGT enzyme further comprises an amino acid sequence at least 70% identical to SEQ ID NO:
202.
89. 89. The method of claim 88, wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:202, or the at least one UGT enzyme comprises an amino acid sequence with 1 to 20 amino acid modifications relative to SEQ ID NO:202, wherein the amino acid modifications are independently selected from amino acid substitutions, deletions, and insertions.
90. 90. The method of any one of claims 60-89, wherein the at least one UGT enzyme is a circular permutation of a wild-type UGT enzyme, or a derivative thereof.
91. 91. The method of any one of claims 60-90, wherein the microbial host cell expresses at least three UGT enzymes: a first UGT enzyme that catalyzes a primary glycosylation at the C24 hydroxyl group of mogrol, a second UGT enzyme that catalyzes a primary glycosylation at the C3 hydroxyl group of mogrol, and a third UGT enzyme that catalyzes one or more branched glycosylation reactions.
92. 92. The method of claim 91 , wherein the microbial host cell expresses one or two UGT enzymes that catalyze beta 1,2 and / or beta 1,6 branched glycosylation of C3 and / or C24 primary glycosylation.
93. The UGT enzyme is SEQ ID NO: 165, or a derivative thereof: SEQ ID NO: 146, or a derivative thereof; SEQ ID NO: 214, or a derivative thereof; SEQ ID NO: 129, or a derivative thereof; SEQ ID NO: 164, or a derivative thereof; SEQ ID NO: 116, or a derivative thereof; SEQ ID NO: 202, or a derivative thereof; SEQ ID NO: 218, or a derivative thereof; SEQ ID NO: 217, or a derivative thereof; SEQ ID NO: 138, or a derivative thereof; SEQ ID NO: 204, or a derivative thereof; SEQ ID NO: 205, or a derivative thereof; SEQ ID NO: 207, or a derivative thereof; SEQ ID NO: 208, or a derivative thereof; SEQ ID NO: 209, or a derivative thereof; SEQ ID NO: 11, or a derivative thereof; SEQ ID NO: 215, or a derivative thereof; SEQ ID NO: 213, or a derivative thereof; SEQ ID NO: 206, or a derivative thereof; SEQ ID NO: 122, or a derivative thereof; and SEQ ID NO: 210, or a derivative thereof; 58. The method of any one of claims 53 to 57, comprising three or four UGT enzymes selected from:
94. The microbial host cell may be prokaryotic or eukaryotic, and is optionally a bacterium selected from Escherichia coli, Bacillus subtilis, Corynebacterium glutamicum, Rhodobacter capsulatus, Rhodobacter sphaeroides, Zymomonas mobilis, Vibrio natriegens, or Pseudomonas putida, or a yeast selected from a species of Saccharomyces, Pichia, or Yarrowia, and optionally a species of Saccharomyces cerevisiae, Pichia The method of any one of claims 1 to 93, wherein the selected strains are P. pastoris and Yarrowia lipolytica.
95. 95. The method of claim 94, wherein the microbial host cell is E. coli.
96. 96. The method of claim 94 or 95, wherein the microbial host cell is a bacterium that produces increased MEP pathway products.
97. 97. The method of any one of claims 1 to 96, wherein the heterologous enzyme pathway comprises farnesyl diphosphate synthase (FPPS).
98. 98. The method of any one of claims 1 to 97, wherein the microbial host cell comprises one or more genetic modifications that increase UDP-glucose production or utilization.
99. 99. The method of claim 98, wherein the microbial host cell is a bacterial cell having one or more genetic modifications selected from: AgalE, AgalT, AgalK, AgalM, AushA, Aagp, Apgm, duplication or overexpression of E. coli GALU, expression of Bacillus subtillus UGPA, and expression of Bifidobacterium adolescentis SPL.
100. 100. The method of any one of claims 1 to 99, wherein the mogrol glycoside product is recovered from the extracellular medium.
101. 1. A method for producing a product containing mogrol glycoside, comprising: Producing mogrol glycoside according to any one of claims 1 to 100, and and incorporating glulol glycoside into the product.
102. 102. The method of claim 101, wherein the product is a sweetener composition, a flavoring composition, a food, a beverage, a chewing gum, a binder, a pharmaceutical composition, a tobacco product, a dietary supplement composition, or an oral hygiene composition.
103. 103. The method of claim 101 or 102, wherein the product further comprises one or more of steviol glycosides, aspartame, and neotame.
104. 104. The method of claim 103, wherein the steviol glycosides include one or more of RebM, RebB, RebD, RebA, RebE, and RebI.
105. 1. A microbial host cell expressing a heterologous enzyme pathway that catalyzes the conversion of isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP) to mogrol or mogroside, the pathway comprising: (A) at least two squalene epoxidase enzymes (SQEs) that convert squalene to 2,3;22,23 dioxidosqualene; (B) at least one triterpene cyclase enzyme that converts 22,23-dioxidosqualene to 24,25-epoxycucurbitadienol, wherein the at least one triterpene cyclase enzyme comprises an amino acid sequence at least 70% identical to one of SEQ ID NO:191, SEQ ID NO:192, and SEQ ID NO:193; (C) at least one epoxide hydrolase that converts 24,25-epoxycucurbitadienol to 24,25-dihydroxycucurbitadienol, wherein the at least one epoxide hydrolase comprises an amino acid sequence at least 70% identical to any one of SEQ ID NOs: 189, 58, 184, 185, 187, 188, 190, and 212; (D) a cytochrome P450 enzyme comprising an amino acid sequence having at least 70% sequence identity to an amino acid sequence selected from SEQ ID NO: 194 and SEQ ID NO: 171; and (E) at least one uridine diphosphate-dependent glycosyltransferase (UGT) enzyme comprising an amino acid sequence having at least 70% sequence identity to any one of SEQ ID NOs: 164, 165, 138, 204-211, 213-218; The microbial host cell comprising at least one of:
106. 106. The microbial host cell of claim 105, wherein the at least one squalene epoxidase comprises an amino acid sequence at least 70% identical to any one of SEQ ID NOs: 17-39, 168-170, 177-183.
107. 107. The microbial host cell of claim 106, wherein the at least one squalene epoxidase comprises an amino acid sequence at least 70% identical to SEQ ID NO:
39.
108. The microbial host cell of claim 107, wherein the at least one SQE comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
39.
109. 109. The microbial host cell of claim 108, wherein the host cell comprises two squalene epoxidase enzymes, each comprising an amino acid sequence that is at least 70% identical to SEQ ID NO:
39.
110. One of the SQE enzymes has a 2,3- 110. The microbial host cell of claim 109, having one or more amino acid modifications that enhance the specificity or productivity of converting oxidosqualene to 2,3;22,23 dioxidosqualene.
111. 111. The microbial host cell of claim 110, wherein the amino acid modifications to the squalene epoxidase comprise one or more modifications at the following positions: 35, 133, 163, 254, 283, 380, and 395 of SEQ ID NO:
39.
112. 112. The microbial host cell of claim 111, wherein the squalene epoxidase comprises amino acid substitutions numbered according to SEQ ID NO: 39, namely H35R, F163A, M283L, V380L, and F395Y; or amino acid substitutions numbered according to SEQ ID NO: 39, namely H35R, N133G, F163A, Y254F, V380L, and F395Y.
113. 113. The microbial host cell of any one of claims 105 to 112, wherein the heterologous enzyme pathway further comprises a squalene synthase (SQS).
114. The microbial host cell of claim 113, wherein the SQS comprises an amino acid sequence that is at least 70% identical to SEQ ID NO:11, or wherein the SQS comprises a sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
11.
115. 115. The microbial host cell of any one of claims 105 to 114, wherein the heterologous enzyme pathway comprises at least one triterpene cyclase (TTC).
116. 116. The microbial host cell of claim 115, wherein the heterologous enzyme pathway comprises at least two enzymes that have triterpene cyclase activity and that convert 22,23-dioxidosqualene to 24,25-epoxycucurbitadienol.
117. A microbial host cell described in claim 115 or 116, wherein the TTC comprises an amino acid sequence that is at least 70% identical to the amino acid sequence of SEQ ID NO: 40, or wherein the TTC comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
40.
118. 118. The microbial host cell of any one of claims 115 to 117, wherein the heterologous enzyme pathway comprises at least one TTC comprising an amino acid sequence that is at least 70% identical to one of SEQ ID NO:191, SEQ ID NO:192, SEQ ID NO:193, or at least one TTC comprising an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NO:191, SEQ ID NO:192, SEQ ID NO:
193.
119. 119. The microbial host cell of any one of claims 105-118, wherein the heterologous pathway comprises an enzyme that converts cucurbitadienol to 24,25-epoxycucurbitadienol.
120. 120. The microbial host cell of claim 119, wherein the enzyme that converts cucurbitadienol to 24,25-epoxycucurbitadienol comprises an amino acid sequence having at least about 70% sequence identity to SEQ ID NO:
221.
121. 121. The microbial host cell of any one of claims 105 to 120, wherein the heterologous enzyme pathway comprises an epoxide hydrolase (EPH).
122. 122. The microbial host cell of claim 121, wherein the heterologous pathway comprises at least one EPH that converts 24,25-epoxycucurbitadienol to 24,25-dihydroxycucurbitadienol, and the at least one EPH comprises an amino acid sequence at least 70% identical to one of SEQ ID NO:189, SEQ ID NO:58, SEQ ID NO:184, SEQ ID NO:185, SEQ ID NO:187, SEQ ID NO:188, SEQ ID NO:190, and SEQ ID NO:
212.
123. 123. The microbial host cell of claim 122, wherein the EPH comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 189, 58, 184, 185, 187, 188, 190, and 212.
124. 124. The microbial host cell of any one of claims 105-123, wherein the heterologous pathway comprises one or more oxidases that oxidize C11 of C24,25 dihydroxycucurbitadienol to produce mogrol.
125. 125. The microbial host cell of claim 124, wherein the at least one oxidase is a cytochrome P450 enzyme.
126. 126. The microbial host cell of claim 124 or 125, wherein the cytochrome P450 comprises an amino acid sequence that is at least 70% identical to an amino acid sequence selected from SEQ ID NO: 194 and SEQ ID NO: 171, or wherein the cytochrome P450 enzyme comprises an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to one of SEQ ID NOs: 194 and 171.
127. 127. The microbial host cell of claim 125 or 126, wherein the cytochrome P450 enzyme has at least a portion of its transmembrane domain replaced with a heterologous transmembrane domain.
128. 128. The microbial host cell of any one of claims 124 to 127, wherein the microbial host cell expresses one or more electron transfer proteins selected from cytochrome P450 reductase (CPR), flavodoxin reductase (FPR), and ferredoxin reductase (FDXR) sufficient to regenerate one or more oxidases.
129. 129. The microbial host cell of claim 128, wherein the microbial host cell expresses SEQ ID NO: 194 or a derivative thereof and SEQ ID NO: 98 or a derivative thereof.
130. 130. The microbial host cell of claim 129, wherein the microbial host cell expresses SEQ ID NO: 171 or a derivative thereof and SEQ ID NO: 201 or a derivative thereof.
131. 131. The microbial host cell of any one of claims 105-130, wherein the heterologous enzyme pathway comprises one or more uridine diphosphate-dependent glycosyltransferase (UGT) enzymes, thereby producing one or more mogrol glycosides.
132. 132. The microbial host cell of claim 131, wherein the host cell produces one or more mogrol glycosides selected from Mog.II-E, Mog.III, Mog.III-A1, Mog.III-A2, Mog.III, Mog.IV, MogIV-A, siamenoside, Mog.V, and Mog.VI.
133. 133. The microbial host cell of claim 132, wherein the host cell produces MogV or siamenoside.
134. 134. The microbial host cell of any one of claims 105-133, wherein the at least one uridine diphosphate-dependent glycosyltransferase (UGT) enzyme comprises an amino acid sequence at least 70% identical to an amino acid sequence selected from SEQ ID NOs: 164, 165, 138, 204-211, and 213-218.
135. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to Stevia rebaudiana UGT85C1 (SEQ ID NO:165), or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
165.
136. 136. The microbial host cell of claim 135, wherein the UGT enzyme has an amino acid substitution at one or more positions selected from 41, 49, and 127 relative to SEQ ID NO: 165, optionally including one or more of L41F, D49E, C127F.
137. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to a Coffea arabica UGT (SEQ ID NO:164), or the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
164.
138. 138. The microbial host cell of claim 137, wherein the UGT enzyme has one or more amino acid substitutions as set forth in Table 3 relative to SEQ ID NO: 164, and optionally includes one or more of S150F, T147L, N207K, K270E, V281L, L354V, L13F, T32A, and K101A.
139. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:138, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
138.
140. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:204, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
204.
141. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:205, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
205.
142. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:206, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
206.
143. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:207, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
207.
144. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:208, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
208.
145. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:209, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
209.
146. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:210, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
210.
147. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:211, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
211.
148. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:213, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
213.
149. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:214, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
214.
150. The at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:215, or the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:
215.
135. The microbial host cell of claim 134, comprising an amino acid sequence that is at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to
151. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:218, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
218.
152. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:217, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:217, and the UGT optionally has an amino acid substitution selected from A74E, I91F, H101P, Q241E, and I436L.
153. 135. The microbial host cell of claim 134, wherein the at least one UGT enzyme comprises an amino acid sequence at least 70% identical to SEQ ID NO:216, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
216.
154. 154. The microbial host cell of any one of claims 134-153, wherein the at least one UGT enzyme further comprises an amino acid sequence at least 70% identical to SEQ ID NO:146, or the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
146.
155. 155. The microbial host cell of any one of claims 134-154, wherein the at least one UGT enzyme further comprises an amino acid sequence at least 70% identical to SEQ ID NO:202, or wherein the at least one UGT enzyme comprises an amino acid sequence at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identical to SEQ ID NO:
202.
156. 156. The microbial host cell of any one of claims 134-155, wherein the microbial host cell expresses at least three UGT enzymes: a first UGT enzyme that catalyzes a primary glycosylation at the C24 hydroxyl group of mogrol, a second UGT enzyme that catalyzes a primary glycosylation at the C3 hydroxyl group of mogrol, and a third UGT enzyme that catalyzes one or more branched glycosylation reactions.
157. 157. The microbial host cell of claim 156, wherein the microbial host cell expresses one or two UGT enzymes that catalyze beta 1,2 and / or beta 1,6 branched glycosylation of C3 and / or C24 primary glycosylation.
158. The UGT enzyme is SEQ ID NO: 165, or a derivative thereof: SEQ ID NO: 146, or a derivative thereof; SEQ ID NO: 214, or a derivative thereof; SEQ ID NO: 129, or a derivative thereof; SEQ ID NO: 164, or a derivative thereof; SEQ ID NO: 116, or a derivative thereof; SEQ ID NO: 202, or a derivative thereof; SEQ ID NO: 218, or a derivative thereof; SEQ ID NO: 217, or a derivative thereof; SEQ ID NO: 138, or a derivative thereof; SEQ ID NO: 204, or a derivative thereof; SEQ ID NO: 205, or a derivative thereof; SEQ ID NO: 207, or a derivative thereof; SEQ ID NO: 208, or a derivative thereof; SEQ ID NO: 209, or a derivative thereof; SEQ ID NO: 11, or a derivative thereof; SEQ ID NO: 215, or a derivative thereof; SEQ ID NO: 213, or a derivative thereof; SEQ ID NO: 206, or a derivative thereof; SEQ ID NO: 122, or a derivative thereof; and SEQ ID NO: 210, or a derivative thereof; comprising three or four UGT enzymes selected from 158. The microbial host cell of claim 157.
159. The microbial host cell is prokaryotic or eukaryotic, and is optionally a bacterium selected from Escherichia coli, Bacillus subtilis, Corynebacterium glutamicum, Rhodobacter capsulatus, Rhodobacter sphaeroides, Zymomonas mobilis, Vibrio natriegens, or Pseudomonas putida, or a yeast selected from a species of Saccharomyces, Pichia, or Yarrowia, and is optionally a species of Saccharomyces cerevisiae, Pichia 159. The microbial host cell of any one of claims 105 to 158, which is P. pastoris, and Yarrowia lipolytica.
160. 160. The microbial host cell of claim 159, wherein the microbial host cell is E. coli.
161. 161. The microbial host cell of claim 159 or 160, wherein the microbial host cell is a bacterium that produces increased MEP pathway products.
162. 162. The microbial host cell of any one of claims 105-161, wherein the heterologous enzyme pathway comprises farnesyl diphosphate synthase (FPPS).
163. 163. The microbial host cell of any one of claims 105-162, wherein the microbial host cell comprises one or more genetic modifications that increase UDP-glucose production or utilization.
164. 164. The method of claim 163, wherein the microbial host cell is a bacterial cell having one or more genetic modifications selected from: ΔgalE, ΔgalT, ΔgalK, ΔgalM, ΔushA, Δagp, Δpgm, duplication or overexpression of E. coli GALU, expression of Bacillus subtillus UGPA, and expression of Bifidobacterium adolescentis SPL.
165. 1. A UGT enzyme, or a host cell expressing said UGT enzyme, wherein the UGT enzyme comprises an amino acid sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 97% sequence identity to SEQ ID NO: 165, and wherein the UGT enzyme has one or more amino acid substitutions selected from L41F, D49E, and C127F with respect to SEQ ID NO:
165.
166. 166. The UGT enzyme or host cell of claim 165, wherein the UGT enzyme comprises the amino acid substitutions L41F, D49F, and C127F relative to SEQ ID NO:
165.
167. 1. A UGT enzyme, or a host cell expressing said UGT enzyme, wherein the UGT enzyme comprises an amino acid sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 97% sequence identity to SEQ ID NO: 164, and wherein the UGT enzyme has one or more amino acid substitutions selected from Table 3.
168. 168. The UGT enzyme or host cell of claim 167, wherein the UGT enzyme has one or more substitutions selected from S150F, T147L, N207K, K270E, V281L, L354V, L13F, T32A, and K101A relative to SEQ ID NO:
164.
169. 169. The UGT enzyme of claim 168, comprising the amino acid substitutions T147L and N207K relative to SEQ ID NO:
164.
170. 1. A UGT enzyme, or a host cell expressing said UGT enzyme, wherein the UGT enzyme comprises an amino acid sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 97% sequence identity to SEQ ID NO:217, and wherein the UGT enzyme has one or more amino acid substitutions selected from A74E, I91F, H101P, Q241E, and I436L relative to SEQ ID NO:
217.
171. 171. The UGT enzyme or host cell of claim 170, comprising the amino acid substitutions A74E, I91F, and H101P relative to SEQ ID NO:217.