Enzymes, host cells, and methods for producing mogrosides
Patent Information
- Application Number
- CN202480082208.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-06
- Filing Date
- 2024-11-06
- Publication Date
- 2026-08-28
Smart Images

Figure CN122663291A_ABST
Abstract
Description
priority
[0001] This application claims the benefit and priority of U.S. Application No. 63 / 596,330, filed November 6, 2023, which is incorporated herein by reference in its entirety.
[0002] sequence list This application contains a sequence list, which has been submitted in XML format through the Patent Centre. The contents of the XML file named “MAN-039PC_107590-5038_Sequence_Listing” are incorporated herein by reference in their entirety. The file was created on November 5, 2024, and is 123,589 bytes in size. Background Technology
[0003] Mogrosides are derived from the cucurbitacin found in the mogroside of the Cucurbitaceae family (Sira Siraitia grosvenorii Mogroside V (also known as monk fruit or Luo Han Guo) is a specialized secondary metabolite derived from triterpenes, found in the fruit of the fruit. Its biosynthesis in the fruit involves a series of successive glycosylations of the glycoside aglycone mogroside. The food industry is increasingly using mogroside fruit extracts as natural non-sugar food sweeteners. For example, mogroside V (Mog.V) has a sweetening capacity approximately 250 times that of sucrose (Kasai et al.). Agric Biol Chem (1989)). Furthermore, recent studies have revealed additional health benefits of mogrosides (Li et al., 1989). Chin J Nat Med (2014)).
[0004] Several factors are driving a surge in research and commercialization interest in mogrosides and mogrosides in general, including, for example, the explosive growth in popularity and demand for natural sweeteners; the difficulty in scaling up the production of other promising natural sweeteners (such as rebaudioside M (RebM)) from stevia plants; the superior taste performance of Mog.V relative to other natural and artificial sweetener products on the market; and the medicinal potential of the plant and fruit.
[0005] Purified Mog.V has been approved in Japan as a high-intensity sweetener (Jakinovich et al., Journal of Natural Products(1990) and the extract has been granted GRAS status (GRAS 522) in the United States as a non-nutritive sweetener and flavor enhancer. Extraction of mogrosides from fruit can produce products of varying purities and is often accompanied by undesirable aftertastes. Furthermore, the yield of mogrosides in cultivated fruit is limited due to low plant yields and the specific cultivation requirements of the plant. Mogrosides are present in fresh fruit at approximately 1% and in dried fruit at approximately 4% (Li HB, et al., 2006). Mog.V is the major component, present in dried fruit at a concentration of 0.5% to 1.4%. Moreover, purification difficulties limit the purity of Mog.V, and commercial products derived from plant extracts are standardized to approximately 50% Mog.V. Pure Mog.V products are highly likely to achieve greater commercial success than blends because they are less likely to have off-flavors, will be easier to formulate into products, and have good solubility potential. Therefore, it is advantageous to be able to produce sweet mogroside compounds via biotechnological processes. Attached Figure Description
[0006] Figure 1A The biosynthetic pathway for the production of Mog.V in vivo is shown. The enzymatic conversions required at each step and the types of enzymes needed are indicated. Undesirable reactions that produce off-pathway chemicals (e.g., cucurbitacinol, hydrolyzed 2,3;22,23-squalene dioxide, and glycosylated 2,4,25-dihydroxycucurbitacinol) are shown with X. Figure 1B This diagram illustrates the glycosylation pathway to Mog.V. Bubble structures represent different mogrosides. The white tetracyclic core represents mogroside. The number below each structure indicates the specific glycosylated mogroside. Black circles represent C3 or C24 glycosylation. Dark gray vertical circles represent 1,6-glycosylation. Light gray horizontal circles represent 1,2-glycosylation. Abbreviations: Mog, mogroside; Sia, symmenidine.
[0007] Figures 2A to 2E This study demonstrates the increased yield of 2,3;22,23-squalene dioxide by bacterial strains that produce squalene and express squalene cyclooxygenase (SQE) derivatives. Figures 2A to 2C This study demonstrates the improved yield of 2,3;22,23-squalene dioxide by a bacterial strain that produces squalene and expresses a squalene cyclooxygenase (SQE) derivative, obtained through successive rounds of engineering. Figure 2A The results show the fold increase in the production of 2,3;22,23-squalene dioxide by the bacterial strain expressing SQE_2 (SEQ ID NO: 3) compared to the bacterial strain expressing SQE_1 (SEQ ID NO: 2). Figure 2BThe results show the fold increase in 2,3;22,23-squalene dioxide production by bacterial strains expressing SQE_3 compared to strains expressing SQE_2. Figure 2C The results show that the bacterial strain expressing SQE_4 (SEQ ID NO: 4) increased the yield of 2,3;22,23-squalene dioxide by a factor of 1 compared to the strain expressing SQE_3. Figure 2D and Figure 2E This study demonstrates that a bacterial strain that produces squalene and expresses a derivative of squalene epoxidase (McSQE_0; SEQ ID NO: 77), obtained through successive rounds of engineering, has improved the yield of 2,3;22,23-squalene dioxide. Figure 2D The results show that the bacterial strain expressing McSQE_1 (SEQ ID NO: 78) increased the yield of 2,3;22,23-squalene dioxide by a factor of 1 compared to the bacterial strain expressing McSQE_0. Figure 2E The results show the fold increase in 2,3;22,23-squalene production of bacterial strains expressing McSQE_2 (SEQ ID NO:79) compared to bacterial strains expressing McSQE_1.
[0008] Figure 3 This study demonstrates the increased yield of 24,25-dihydroxy-cucurbitadienol achieved through the engineering of an epoxide hydrolase (EPH). The bacterial strain producing 24,25-epoxide-cucurbitadienol and expressing the engineered form EPH_1 (SEQ ID NO: 6) showed a fold increase in 24,25-dihydroxy-cucurbitadienol yield compared to the same strain expressing EPH_0 (SEQ ID NO: 5).
[0009] Figures 4A to 4E This study demonstrates the increased yield of 24,25-epoxy-cucurbitacinol by a bacterial strain that produces a derivative of 2,3;22,23-squalene dioxide and expresses cucurbitacinol synthase (CDS) after successive rounds of engineering. Figure 4A The results show the fold increase in 24,25-epoxy-cucurbitadienol production by the bacterial strain expressing ECDS_1 (SEQ ID NO: 8) compared to the strain expressing ECDS_0 (SEQ ID NO: 7). Figure 4B The results show that the bacterial strain expressing ECDS_2 (SEQ ID NO: 9) increased the yield of 24,25-epoxy-cucurbitadienol and the extra-pathway compound cucurbitadienol by a factor of 1 compared to the strain expressing ECDS_1. Figure 4CThe results show the fold increase in 24,25-epoxy-cucurbitadienol production by the strain expressing ECDS_3 (SEQ ID NO: 10) compared to the strain expressing ECDS_2. Figure 4D The results show that the strain expressing ECDS_4 (SEQ ID NO: 11) increased the yield of 24,25-dihydroxy-cucurbitadienol by a factor of 10 compared to the strain expressing ECDS_3. Figure 4E The results show the fold increase in 24,25-dihydroxy-cucurbitadienol production by the strain expressing ECDS_5 (SEQ ID NO: 12) compared to the strain expressing ECDS_4.
[0010] Figure 5 The comparison of CcCDS (SEQ ID NO: 7) with homologs SgCDS (SEQ ID NO: 68), CpCAS (SEQ ID NO: 69), PsCAS (SEQ ID NO: 70), and AsCAS (SEQ ID NO: 71) is shown. As described in Table 11, some sites used for mutagenesis are highlighted.
[0011] Figure 6 The study showed that bacterial strains expressing AaSQS (SEQ ID NO: 13), SQE_3, ECDS_4, and EPH_1 exhibited an improvement in the 24,25-dihydroxycucurbitadienol:cucurbitadienol ratio compared to bacterial strains expressing AaSQS, SQE_2, ECDS_3, and EPH_1.
[0012] Figures 7A to 7T This study demonstrates the increased mogroside production of a bacterial strain that produces 24,25-dihydroxycucurbitadienol and expresses a derivative of cytochrome P450 enzyme, obtained through engineered processes involving successive rounds of hydroxylation of the C11 of 24,25-dihydroxycucurbitadienol. Figure 7A The results show the fold increase in mogroside production by the bacterial strain expressing sohB_CppCYP87A3 (SEQ ID NO: 15) compared to the bacterial strain expressing CppCYP87A3 (SEQ ID NO: 14). Figure 7B The results show that the bacterial strain expressing sohB_C11CYP_1 (SEQ ID NO: 16) increased mogroside production by a factor of 1 compared to the strain expressing the parental enzyme sohB_CppCYP87A3. Figure 7C The results show the fold increase in mogroside production of the strain expressing the engineered enzyme sohB_C11CYP_2 (SEQ ID NO: 17) compared to the strain expressing the parental enzyme sohB_C11CYP_1. Figure 7DThe results show the fold increase in mogroside production of the strain expressing the engineered enzyme sohB_C11CYP_3 (SEQ ID NO: 18) compared to the strain expressing the parental enzyme sohB_C11CYP_2. Figure 7E The results show the fold increase in mogrool production of the strain expressing the engineered enzyme sohB_C11CYP_4 (SEQ ID NO: 19) compared to the strain expressing the parental enzyme sohB_C11CYP_3. Figure 7F The results show the fold increase in mogroside production of the strain expressing the engineered enzyme sohB_C11CYP_5 (SEQ ID NO: 20) compared to the strain expressing the parental enzyme sohB_C11CYP_4. Figure 7G The results show that bacterial strains expressing the engineered enzyme sohB_C11CYP_6 (SEQ ID NO: 21) increased the yields of mogrool, 11-hydroxy-cucurbitadienol, and oxo-mogrool compared to strains expressing the parental enzyme sohB_C11CYP_5. Figure 7H The results show the fold increase in the production of mogroside and 11-oxomogroside by the strain expressing SohB_C11CYP_8 (SEQ ID NO: 45) compared to the strain expressing SohB_C11CYP_6. Figure 7I The results show the fold increase in the production of mogroside and 11-oxomogroside by the strain expressing SohB_C11CYP_9 (SEQ ID NO: 46) compared to the strain expressing SohB_C11CYP_8. Figure 7J The results show the fold increase in mogroside production of the strain expressing SohB_C11CYP_10 (SEQ ID NO: 47) compared to the strain expressing SohB_C11CYP_9. Figure 7K The results show the fold increase in mogroside production of the strain expressing SohB_C11CYP_11 (SEQ ID NO: 80) compared to the strain expressing SohB_C11CYP_10. Figure 7L The results show the fold increase in mogroside production of the strain expressing SohB_C11CYP_12 (SEQ ID NO: 81) compared to the strain expressing SohB_C11CYP_11. Figure 7M The results show the fold increase in the ratio of mogroside to deoxymogroside and the yield of 11-oxomogroside by the strain expressing SohB_C11CYP_13 (SEQ ID NO: 82) compared to the strain expressing SohB_C11CYP_12. Figure 7NThe results show the fold increase in mogroside production of the strain expressing SohB_C11CYP_14 (SEQ ID NO: 83) compared to the strain expressing SohB_C11CYP_13. Figure 7O The results show the fold increase in mogroside production of the strain expressing SohB_C11CYP_15 (SEQ ID NO: 84) compared to the strain expressing SohB_C11CYP_14. Figure 7P The results show the fold increase in mogroside production by the strain expressing n20_C11CYP_15 (SEQ ID NO: 85) compared to the strain expressing SohB_C11CYP_15. Figure 7Q The results show the fold increase in mogroside production of the strain expressing n20_C11CYP_16 (SEQ ID NO: 86) compared to the strain expressing n20_C11CYP_15. Figure 7R The results show the fold increase in mogroside production of the strain expressing n20_C11CYP_17 (SEQ ID NO: 87) compared to the strain expressing n20_C11CYP_16. Figure 7S The results show the fold increase in mogroside production of the strain expressing n20_C11CYP_18 (SEQ ID NO: 88) compared to the strain expressing n20_C11CYP_17. Figure 7T The results show the fold increase in mogroside production of the strain expressing n20_C11CYP_19 (SEQ ID NO: 89) compared to the strain expressing n20_C11CYP_18.
[0013] Figures 8A to 8D This study demonstrates an enhanced mogroside production pathway achieved by a bacterial strain developed through multiple rounds of engineering, which expresses a biosynthetic pathway producing mogroside and a derivative of cytochrome P450 reductase. This pathway supports the use of engineered cytochrome P450 enzymes for cytochrome P450-based hydroxylation of 24,25-dihydroxycucurbitadienol at C11. Figure 8A The results show the fold increase in the production of mogroside, including MogCPR1_1 (SEQ ID NO: 23), compared to strains expressing MogCPR1_0 (SEQ ID NO: 22). Figure 8B The results show the fold increase in mogroside production by the bacterial strain expressing MogCPR2_0 (SEQ ID NO: 24) compared to the strain expressing MogCPR1_0. Figure 8CThe results show the fold increase in mogroside production of the strain expressing MogCPR2_1 (SEQ ID NO: 25) compared to the strain expressing MogCPR2_0. Figure 8D The results show the fold increase in mogroside production of the strain expressing MogCPR2_2 / SohB_C11CYP_6 (SEQ ID NO: 26) compared to the strain expressing MogCPR2_0.
[0014] Figures 9A to 9F This study demonstrates the increased C3 glycosylation of mogroside after successive rounds of engineering of a uridine diphosphate-dependent glycosyltransferase (UGT) for C3 glycosylation. Figure 9A The figure shows the fold increase in the yield of MogIIE produced from MogIA substrates in an enzyme assay using MogUGTC3_1 (SEQ ID NO: 28) compared to MogUGTc3_0 (SEQ ID NO: 27). Figure 9B The results show the fold increase in MogIE production by the bacterial strain that produces mogroside and expresses MogUGTC3_2 (SEQ ID NO: 29) compared to the strain expressing MogUGTC3_1. Figure 9C The results show that, compared with strains expressing MogUGTC3_2, strains expressing MogUGTC3_3 (SEQ ID NO: 30) showed a fold increase in MogIE production relative to the glycosylation of products outside the dihydroxycucurbitadiol pathway. Figure 9D The results show the fold increase in MogIE yield resulting from the expression of CmUGTc3_1 compared to CmUGTc3_0 in the mogroside-producing strain. Figure 9E The study shows the fold increase in MogIE and deoxy-MogIE yields resulting from the expression of the CmUGTc3_1 variant in mogroside-producing strains, and demonstrates enhanced specificity for mogrosides relative to deoxy-mogrosides. Figure 9F The results show the fold increase in MogIE and deoxy-MogIE yields resulting from the expression of CmUGTc3_3 compared to CmUGTc3_2 in the mogroside-producing strains.
[0015] Figures 10A to 10E This study demonstrates the increased C24 glycosylation of mogrool after successive rounds of engineering of a uridine diphosphate-dependent glycosyltransferase (UGT) for C24 glycosylation. Figure 10AThe figure shows the fold increase in the yield of MogIIE produced from MogIE substrates in an enzyme assay using MogUGTC24_1 (SEQ ID NO: 32) compared to MogUGTc24_0 (SEQ ID NO: 31). Figure 10B The figure shows the fold increase in the yield of MogIIE produced from MogIE substrates in enzyme assays using MogUGTC24_2 (SEQ ID NO: 33) compared to MogUGTC24_1. Figure 10C The figure shows the fold increase in the yield of MogIIE produced from MogIE substrates in enzyme assays using MogUGTC24_3 (SEQ ID NO: 34) compared to MogUGTC24_2. Figure 10D The results show the fold increase in MogIA yield resulting from expression of MogUGTc24_4 compared to MogUGTc24_3 in the mogroside-producing strains. Figure 10E The results show the fold increase in MogIA and deoxy-MogIA yields in mogroside-producing strains resulting from expression of MogUGTc24_5.
[0016] Figures 11A to 11D This study demonstrates the increased MogV production following successive rounds of engineering of uridine diphosphate-dependent glycosyltransferase (UGT). Figure 11A The figure shows the fold increase in MogV production in an enzymatic assay of the conversion of MogIIE to MogV by the MogUGT1216_1 enzyme (SEQ ID NO: 37) compared to UGT94-289-1 enzyme (SEQ ID NO: 35) or MogUGT1216_0 (SEQ ID NO: 36). Figure 11B The results show the fold increase in the yields of MogIII, MogIV-A, sialoside, and MogV in enzyme assays of MogUGT1216_1 converted to MogIIE as a double mutant (L14V-L266Q) compared to the L14V monosubstituted form (Mut1) or the L266Q monosubstituted form (Mut16). Figure 11C The figure shows the fold increase in MogV production in an enzymatic assay of the conversion of MogIIE to MogV via MogUGT1216_2 (SEQ ID NO: 38) compared to MogUGT1216_1. Figure 11D The figure shows the fold increase in MogV production in an enzymatic assay of the conversion of MogIIE to MogV via MogUGT1216_3 (SEQ ID NO: 90) compared to MogUGT1216_2.
[0017] Figure 12 It shows the process of passing through coffee trees ( Coffea arabica The enzyme assay of the conversion of MogIIE to MogV by uridine diphosphate-dependent glycosyltransferase 1,6 (MogUGT16_0, SEQ ID NO: 39) and its engineered derivative MogUGT16_1 (SEQ ID NO: 40) showed an increase in MogV production.
[0018] Figures 13A to 13C This study demonstrates the increased yield of 2,3;22,23-squalene dioxide after expressing PntAB (a membrane-bound transhydrogenase that produces NADPH in E. coli) and undergoing multiple rounds of engineering. Figure 13A It shows the expression without. pntAB Compared to strains that express SQS and SQE and overexpress [the other strain], [the other strain] expresses SQS and SQE. pntAB The bacterial strains (SEQ ID NO: 41 and 42) showed a fold increase in the production of 2,3;22,23-squalene dioxide. Figure 13B It shows the expression of SQS, SQE, ECDS, and EPH, but not the expression of SQS, SQE, ECDS, and EPH. pntAB Compared to other strains, this strain expressed SQS, SQE, ECDS, and EPH and overexpressed them. pntAB The bacterial strain increased the yield of 24,25-epoxycucurbitadienol by a factor of 1. Figure 13C The results show that overexpression of native *E. coli* compared to strains expressing only SQS and SQE significantly improves the efficacy of strains. fldA Fold changes in squalene, 2,3-squalene oxide, and 2,3:22,23-squalene dioxide in bacterial strains that expressed SQS and SQE, either the heterologous Cv.fdx gene (SEQ ID NO: 43) or SQS and SQE.
[0019] Figures 14A to 14C It shows that in the by ybhGFSR In strains with reduced expression of ABC-type transport proteins encoded by operons, the production of mogroside and 24,24-dihydroxycucurbitadienol was increased. Figure 14A This demonstrates the complete metabolic pathway for the production of mogroside. ybhGFSR + Compared to the control strain (CTL), it expressed the complete metabolic pathway for the production of mogrool and lacked [the other pathway]. ybhGFSR The levels of 2,3;22,23-squalene dioxide (OOSQ) and 24,24-dihydroxycucurbitadienol + mogroside in the operon strains. Figure 14B This demonstrates the complete metabolic pathway for the production of mogroside. ybhGFSR +Compared to the control strain (CTL), it expresses the complete metabolic pathway for the production of mogrool and has a sub-efficacy. ybhG The levels of 2,3;22,23-squalene dioxide (OOSQ) and 24,24-dihydroxycucurbitadienol + mogroside in the mutant strain; the sub-effect ybhG mutation was constructed by replacing the ribosome binding site in the chromosome with a form weaker than the natural sequence. Figure 14C This demonstrates the complete metabolic pathway for the production of mogroside. ybhGFSR + Compared to the control strain (CTL), this strain expressed the complete metabolic pathway for the production of mogrool and lacked a chromosome. ybhF Genes are expressed via plasmid overexpression. ybhGFSR The levels of 24,24-dihydroxycucurbitadienol + mogroside (top) and 22,3;22,23-squalene dioxide (OOSQ, bottom) in the operon strains. Detailed Implementation
[0020] This disclosure is based on the development of engineered enzymes and microbial cells that improve the production of mogrosides or mongholic acid from squalene microorganisms compared to known processes by increasing the yield of various pathway intermediates. In various respects, the engineered enzymes disclosed herein increase the yield of pathway intermediates and reduce the formation of extra-pathway metabolites by increasing substrate preference for preferred pathway intermediates or decreasing substrate preference for undesirable pathway intermediates or extra-pathway products, thereby directing the metabolic resources of microbial cells toward the production of pathway intermediates while reducing the formation of undesirable extra-pathway metabolites.
[0021] Pathway intermediates and certain extra-pathway metabolites are shown in Figure 1A and Figure 1B . Figure 1A The in vivo biosynthetic pathways leading to mogroside and MogV are shown. This pathway involves the condensation of farnesyl pyrophosphate to squalene, catalyzed by squalene synthase (SQS); the sequential epoxidation of squalene to 2,3-squalene oxide and 2,3;22,23-squalene dioxide, both catalyzed by squalene epoxidase (SQE); the cyclization of 2,3;22,23-squalene dioxide to 24,25-epoxycucurbitadienol, catalyzed by epoxycucurbitadienol synthase (ECDS); the hydration of 24,25-epoxycucurbitadienol to 24,25-dihydroxycucurbitadienol, catalyzed by epoxide hydrolase (EPH); and the hydroxylation of 24,25-dihydroxycucurbitadienol to mogroside, catalyzed by cytochrome P450 and cytochrome P450 reductase active at the C11 position of 24,25-dihydroxycucurbitadienol. Figure 1A ). Figure 1BThe glycosylation pathway to Mog.V is illustrated. The conversion of mogroside to Mog.V involves C3-glucosylation, C24-glucosylation, 1,6-glucosylation, and 1,2-glucosylation reactions, which are catalyzed by uridine diphosphate-dependent glycosyltransferases (UGTs, see [link to article]). Figure 1A and Figure 1B ).
[0022] Figure 1A The intermediates shown in the pathway are farnesyl pyrophosphate, squalene; 2,3-squalene oxide, 2,3;22,23-squalene dioxide, 24,25-epoxycucurbitadienol, 24,25-dihydroxycucurbitadienol, and mogroside; and Figure 1B The mogrosides shown are referred to in this article as "on-pathway products." Figure 1A As shown, enzymes such as CDS, EPH, and UGT catalyze the conversion of intra-pathway intermediates into undesirable products, including cucurbitacinol, hydrolyzed 2,3;22,23-squalene dioxide, and glycosides of 2,4,25-dihydroxycucurbitacinol. These undesirable products are referred to herein as “extra-pathway” products. Bacterial cells expressing squalene synthase (SQS), squalene epoxidase (SQE), cucurbitacinol synthase (CDS), and epoxide hydrolase (EPH) typically produce approximately 90% or more (by weight) of extra-pathway products, along with relatively low amounts of mogroside. This document discloses a pathway comprising engineered enzymes that increase the yield of intra-pathway products. Relative to Figure 1A The total amount of in-pathway and out-of-pathway products shown is 100%. Bacterial cells containing the engineered enzymes disclosed herein produce more than about 40% (by weight), or more than about 50%, or more than about 60%, or more than about 70%, or more than about 80% of the in-pathway products. Bacterial cells containing the engineered enzymes disclosed herein produce less than about 50% (by weight), or less than about 40%, or less than about 30%, or less than about 20% of the out-of-pathway products (and specifically the out-of-pathway products cucurbitadienol, hydrolyzed 2,3;22,23-squalene dioxide, and glycosides of 2,4,25-dihydroxycucurbitadienol).
[0023] Therefore, in various aspects and embodiments, this disclosure provides enzymes, microbial strains, and methods for preparing mogrosides and mogrosides using recombinant microbial processes. In other aspects, the invention provides methods for preparing products, including foods, beverages, and sweeteners (and other products), by incorporating mogrosides produced according to the methods described herein. Still in other aspects, the invention provides engineered UGT enzymes for glycosylation of secondary metabolite substrates, such as mogrosides or mogrosides.
[0024] As used herein, the terms “terpene” or “triterpene” are used interchangeably with the terms “terpene compounds” or “triterpene compounds”, respectively.
[0025] Therefore, in one aspect, this disclosure provides a method for producing mogroside or mogroside, the method comprising: providing microbial cells that produce squalene and express a biosynthetic pathway for converting squalene into said mogroside or mogroside in microbial cells, and culturing said microbial cells under conditions suitable for producing said mogroside or mogroside. The culture produces a squalene derivative, wherein the extra-pathway products cucurbitacinol, hydrolyzed 2,3;22,23-squalene dioxide, and glycosylated 24,25-dihydroxycucurbitacinol are less than about 40% by weight. In some embodiments, the culture produces at least 50% (by weight) or at least 60% of the in-pathway products farnesyl pyrophosphate, squalene, 2,3-squalene oxide, 2,3;22,23-squalene dioxide, 24,25-epoxycucurbitadienol, 24,25-dihydroxycucurbitadienol, mogroside, and mogroside; wherein the total amount of such in-pathway products and the extra-pathway products cucurbitadienol, hydrolyzed glycosides of 2,3;22,23-squalene dioxide and 24,25-dihydroxycucurbitadienol is considered 100%.
[0026] In various aspects, this disclosure provides a microbial cell that produces squalene and expresses a biosynthetic pathway for converting squalene into mogroside or mogroside. When the microbial cell is cultured under conditions suitable for producing said mogroside or mogroside, the microbial cell produces a squalene derivative wherein the extra-pathway products cucurbitadienol, hydrolyzed 2,3;22,23-squalene dioxide, and glycosylated 2,4,25-dihydroxycucurbitadienol are less than 40% by weight, less than 30% by weight, or less than 20% by weight. In some embodiments, the microbial cells produce at least 50% or at least 60% or at least 70% by weight of the in-pathway products farnesyl pyrophosphate, squalene, 2,3-squalene oxide, 2,3;22,23-squalene dioxide, 24,25-epoxycucurbitadienol, 24,25-dihydroxycucurbitadienol, mogroside, and mogroside; wherein the total amount of such in-pathway products and the extra-pathway products cucurbitadienol, hydrolyzed glycosides of 2,3;22,23-squalene dioxide and 24,25-dihydroxycucurbitadienol is considered 100%.
[0027] In some embodiments, the biosynthetic pathway comprises: squalene synthase (SQS); squalene epoxidase (SQE) for the production of 2,3;22,23-squalene dioxide from squalene; epoxidized cucurbitacinol synthase (ECDS) for the production of 2,3;22,23-epoxidized cucurbitacinol from 2,3;22,23-squalene dioxide; epoxide hydrolase (EPH) for the production of 24,25-dihydroxy-cucurbitacinol from 24,25-epoxidized cucurbitacinol; and a cytochrome P450 enzyme (e.g., P450 hydroxylase) for the production of mogroside from 24,25-dihydroxy-cucurbitacinol. In some embodiments, the microbial cells express a P450 reductase chaperone for regenerating P450 hydroxylase (see [link to relevant documentation]). Figure 1A ).
[0028] In some embodiments, the microbial cells express one or more uridine diphosphate-dependent glycosyltransferases (UGTs) that glycosylate the C-3 and C-24 positions of mogrosides to produce one or more mogrosides from mogrosides, including but not limited to Mog V. In some embodiments, the microbial cells express a UGT enzyme (UGTc3) that produces C-3 mogroside glycosides from mogrosides. In some embodiments, the microbial cells express a UGT enzyme (UGTc24) that produces C-24 mogroside glycosides from mogrosides. In some embodiments, the microbial cells express a UGT enzyme (UGT1216) that produces 1,6- and / or 1,2-branched glycosides of C-3 and / or C-24 mogrosides (see also...). Figure 1B ).
[0029] According to various aspects and embodiments, this disclosure provides engineered enzymes for biosynthetic pathways, including SQE, ECDS, EPH, cytochrome P450 hydroxylase, P450 reductase, UGTc3, UGTc24, and UGT1216 enzymes. These engineered enzymes are described in detail below.
[0030] In various aspects and embodiments, this disclosure provides a squalene epoxidase comprising an amino acid sequence having at least 80%, or at least 85%, at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 2, and comprising one or more mutations or mutation groups listed in Tables 1, 2, and 3. In some embodiments, the squalene epoxidase comprises at least two, at least three, at least four, at least five, or at least six or more mutations or mutation groups listed in Tables 1, 2, and / or 3. In some embodiments, the squalene epoxidase comprises at least two, at least three, at least four, or at least five mutations or mutation groups from Tables 1, 2, and / or 3, showing that it provides at least a 1.2-fold increase in 2,3-squalene dioxide titer and / or at least a 1.2-fold increase in 2,3;22,23-squalene dioxide titer compared to the parent enzyme. In some embodiments, microbial cells expressing the squalene epoxidase disclosed herein provide an increased yield of 2,3;22,23-squalene dioxide compared to the parent enzyme.
[0031] In some embodiments, the squalene cyclooxygenase, relative to SEQ ID NO: 2, comprises one or more substitutions at positions selected from S116, N134, I145, A164, A169, and A400. In some embodiments, S116 is substituted with a charged amino acid selected from Glu and Asp. In some embodiments, N134 is substituted with a nonpolar amino acid selected from Gly, Ala, Ile, Leu, and Val. In some embodiments, I145 is substituted with a charged amino acid selected from His, Arg, and Lys. In some embodiments, A164 is substituted with a polar amino acid selected from Asn and Gln. In some embodiments, A169 is substituted with a charged amino acid selected from Arg, Lys, Glu, Asp, or His. In some embodiments, A400 is substituted with a nonpolar amino acid selected from Val, Leu, and Ile. In some embodiments, the squalene epoxidase relative to SEQ ID NO: 2 comprises one or more mutations selected from S116E, N134G, I145H, A164N, A169R, and A400V. In some embodiments, the squalene epoxidase relative to SEQ ID NO: 2 comprises at least two, at least three, at least four, at least five, or all six mutations selected from S116E, N134G, I145H, A164N, A169R, and A400V.
[0032] In various embodiments, the SQE comprises the amino acid sequence of SEQ ID NO: 3 or SEQ ID NO: 4.
[0033] In various aspects and embodiments, this disclosure provides a squalene epoxidase comprising an amino acid sequence having at least 80%, or at least 85%, at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 77, and comprising one or more mutations or groups of mutations listed in Tables 4 and / or 5.
[0034] In some embodiments, the squalene epoxidase comprises one or more substitutions at D41, M315, and / or S366 relative to SEQ ID NO: 77. In some embodiments, S366 is substituted with a charged amino acid selected from Arg, Lys, and His. In some embodiments, D41 is substituted with a positively charged amino acid selected from His, Lys, and Arg. In some embodiments, M315 is substituted with a nonpolar amino acid selected from Leu, Ile, and Val. In some embodiments, the squalene epoxidase comprises one or more mutations selected from S366R, D41H, and M315L relative to SEQ ID NO: 77.
[0035] In various embodiments, the SQE comprises the amino acid sequence of SEQ ID NO: 78 or SEQ ID NO: 79.
[0036] The intermediate 2,3-squalene dioxide may be redirected to an undesirable cyclization, leading to the production of the extra-pathogen product cucurbitadienol (see [link to original text]). Figure 1A The generation of the off-pathogen cucurbitacinol may lead to increased impurities, increased production costs, and reduced overall productivity of the strain and process. Not wishing to be bound by theory, the squalene epoxidase disclosed herein is believed to catalyze two consecutive epoxidation reactions at a rate limiting the use of 2,3-squalene dioxide substrates for undesirable cyclization reactions. Therefore, in some embodiments, the squalene epoxidase disclosed herein catalyzes two epoxidation reactions: the first catalyzes the epoxidation of squalene to produce 2,3-squalene dioxide, and the second catalyzes the epoxidation of 2,3-squalene dioxide to produce 2,3;22,23-squalene dioxide. Therefore, in some embodiments, culturing microbial cells expressing the squalene epoxidase and ECDS disclosed herein produces at least 50% by weight, or at least 60% by weight, or at least 70% by weight of the in-pathway products 2,3-squalene oxide, 2,3;22,23-squalene dioxide, and 24,25-epoxycucurbita diol; wherein the total amount of these in-pathway products and the extra-pathway product cucurbita diol is considered 100%. In some embodiments, culturing microbial cells expressing the squalene epoxidase disclosed herein produces less than 40% (by weight), or less than 30%, or less than 20% cucurbita diol.
[0037] Further amino acid modifications of SQE can be guided by available enzyme structures and homology models, including those by Padyana AK et al. Structure and inhibition mechanism of the catalytic domain of human squalene epoxidase , Nat. Comm (2019) Vol. 10(97): 1-10; or Ruckenstulh et al., Structure-Function Correlations of Two Highly Conserved Motifs in Saccharomyces cerevisiae Squalene Epoxidase , Antimicrobial agents and Chemo. Those described in (2008) Vol. 52(4): 1496-1499.
[0038] Cucurbitadienol synthase catalyzes two cyclization reactions: the cyclization of the intra-path intermediate 2,3-squalene dioxide to the extra-path product cucurbitadienol; and the cyclization of the intra-path intermediate 2,3;22,23-squalene dioxide to the intra-path intermediate 24,25-epoxy-cucurbitadienol (see [link to relevant documentation]). Figure 1A This document discloses an engineered cucurbitacin synthase that exhibits enhanced specificity for 2,3- and 22,23-squalene substrates compared to 2,3-squalene dioxide substrates. The engineered cucurbitacin synthase is referred to herein as 24,25-epoxycucurbitacin synthase (“ECDS”). In some embodiments, microbial cells expressing the ECDS disclosed herein produce fewer extra-pathway product cucurbitacinols compared to microbial cells expressing non-engineered cucurbitacin synthases. In some embodiments, microbial cells expressing the ECDS disclosed herein produce more intra-pathway product 24,25-epoxycucurbitacinols compared to microbial cells expressing non-engineered cucurbitacin synthases.
[0039] In various aspects and embodiments, this disclosure provides a 24,25-epoxycucurbitadienol synthase (ECDS) comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 7, and comprising one or more mutations or mutant groups listed in Tables 7, 8, 9, and 10. The ECDS enzyme produces 24,25-epoxycucurbitadienol from 2,3;22,23-squalene dioxide. In some embodiments, the ECDS enzyme exhibits substrate preference for 2,3;22,23-squalene dioxide relative to 2,3-oxide, thereby reducing the production of the extra-pathway product cucurbitadienol. In some embodiments, the ECDS enzyme comprises the following amino acid sequence, wherein the amino acid sequence has at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% sequence identity with respect to SEQ ID NO: 7, and has one or more mutations relative to SEQ ID NO: 7 that increase enzyme productivity or increase substrate preference for 2,3;22,23-squalene dioxide relative to 2,3-squalene oxide. In some embodiments, the ECDS enzyme comprises one or more mutations or groups of mutations listed in Tables 7, 8, 9, and / or 10. In some embodiments, the ECDS enzyme comprises at least one, at least two, at least three, at least four, or at least five mutations listed in Tables 7, 8, 9, and / or 10, which increase the 24,25-epoxycucurbitadienol titer by at least 1.2-fold and / or reduce cucurbitadienol production compared to the corresponding parent enzyme (as shown in the table).
[0040] In some embodiments, ECDS contains one or more mutations at one or more positions selected from S24, C35, D50, N121, L245, W331, Q401, I490, I553, and A556 relative to SEQ ID NO: 7. In some embodiments, S24 is substituted with a polar amino acid selected from Asn and Gln. In some embodiments, C35 is substituted with a charged amino acid selected from Asp and Glu. In some embodiments, D50 is substituted with Glu. In some embodiments, N121 is substituted with His, Lys, or Arg. In some embodiments, L245 is substituted with Ile or Val. In some embodiments, W331 is substituted with Ser, Thr, or Tyr. In some embodiments, Q401 is substituted with Ala, Gly, or Val. In some embodiments, I490 is substituted with Val, Leu, or Ala. In some embodiments, I553 is substituted with Met. In some embodiments, A556 is substituted with a polar amino acid selected from Ser, Thr, and Tyr. In some embodiments, ECDS comprises, relative to SEQ ID NO: 7, at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or all ten mutations selected from 24N, 35D, 50E, 121H, 245I, 331S, 401A, 490V, 553M, and 556S. Exemplary ECDS enzymes according to various embodiments are disclosed herein as SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, and SEQ ID NO: 12.
[0041] In some embodiments, the ECDS enzyme contains one or more mutations in the substrate entry channel that improve substrate selectivity relative to 2,3-squalene oxide for 2,3;22,23-squalene dioxide. This entry channel is shared with homologous enzymes (see [link to homologous enzyme]). Figure 5Amino acid modifications located at positions in the substrate entry channel liner can alter substrate selectivity. Exemplary positions include amino acids 245, 331, 553, and 556 relative to SEQ ID NO: 7, which are close to the membrane carrying ECDS. Indeed, mutations 245I, 331S, and 553M (relative to SEQ ID NO: 7) significantly increase the production of 24,25-epoxycucurbitadienol and 24,25-dihydroxycucurbitadienol while reducing the production of cucurbitadienol (see Table 10). Therefore, in some embodiments, the ECDS enzyme is an enzyme having at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% sequence identity to SEQ ID NO: 7, but with mutations at positions corresponding to SEQ ID NO: 7, 245, 331, 553, and 556. In some embodiments, the ECDS comprises the amino acid sequence of SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, or SEQ ID NO: 71 (or an enzyme having at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% sequence identity with it), and contains one or more mutations (e.g., 1, 2, 3, or 4 mutations) at positions 245, 331, 553, and 556 corresponding to SEQ ID NO: 7. See Table 11. For example, the amino acid corresponding to position 245 of SEQ ID NO: 7 may be a hydrophobic amino acid, such as Ile, Val, Ala, Met, or Phe; the amino acid corresponding to position 331 of SEQ ID NO: 7 may be a polar amino acid, such as, but not limited to, Ser or Thr; the amino acid corresponding to position 553 of SEQ ID NO: 7 may be a hydrophobic amino acid, such as Ile, Val, Ala, Met, or Phe; and the amino acid corresponding to position 556 of SEQ ID NO: 7 may be a low-polarity amino acid, such as Ser or Thr. Alternatively or additionally, ECDS may contain mutations at positions corresponding to positions 50, 490, 121, and / or 401 relative to SEQ ID NO: 7. See Table 11. For example, it is believed that amino acid I490 of SEQ ID NO: 7 constitutes part of the active site, and therefore a mutation at this position may affect substrate preference and / or productivity. In some embodiments, the amino acid corresponding to position 490 of SEQ ID NO: 7 is a hydrophobic amino acid, such as Val, Leu, or Ala. See Table 11. Other inactive site locations that may be important for controlling substrate preference include positions 50, 121, and 401 in SEQ ID NO: 7.In one embodiment, the position 50 corresponding to SEQ ID NO: 7 is mutated to an amino acid selected from Asp, His, Gln, Asn, and Glu. See Table 11. In another embodiment, the position 121 corresponding to SEQ ID NO: 7 is mutated to an amino acid selected from Asp, His, Gln, Asn, and Glu. See Table 11. In yet another embodiment, the position 401 corresponding to SEQ ID NO: 7 is mutated to an amino acid selected from Ala, Asp, His, Gln, Asn, Glu, and Pro. See Table 11.
[0042] Further amino acid modifications of ECDS enzymes can be guided by available enzyme structures and homology models, including those described in the Swiss model Q6BE24 and the AlphaFold model K7NBZ9.
[0043] Epoxide hydrolase (EPH) catalyzes two hydration reactions: the hydration of the intra-pathway intermediate 2,3;22,23-squalene dioxide to the extra-pathway product 2,3;22,23-squalene dioxide; and the hydration of the intra-pathway intermediate 24,25-epoxy-cucurbitadienol to the intra-pathway intermediate 24,25-dihydroxy-cucurbitadienol (see [link]). Figure 1A This document discloses engineered EPH enzymes that exhibit enhanced specificity for 24,25-dihydroxy-cucurbitadiene compared to 2,3;22,23-squalene dioxide. Therefore, in some embodiments, when the total amount of the substrates 2,3;22,23-squalene dioxide and 24,25-epoxy-cucurbitadiene, as well as the products 2,3;22,23-squalene dioxide and 24,25-dihydroxy-cucurbitadiene, is considered 100%, microbial cells expressing the disclosed EPH (as well as the SQE and ECDS enzymes) produce less than 30% or less than 20% of 2,3;22,23-squalene dioxide. In some embodiments, microbial cells expressing the disclosed EPH (as well as the SQE and ECDS enzymes) produce less extra-pathway product 2,3;22,23-squalene dioxide compared to microbial cells expressing non-engineered EPH. In some implementations, culturing microbial cells expressing the EPH disclosed herein produces more of the pathway product 24,25-dihydroxy-cucurbitadienol compared to culturing microbial cells expressing non-engineered EPH.
[0044] In various aspects and embodiments, this disclosure provides an epoxide hydrolase (EPH) for the production of 24,25-dihydroxy-cucurbita-dienol from 24,25-epoxycucurbita-dienol. In some embodiments, the EPH enzyme has a substrate preference for 24,25-epoxycucurbita-dienol relative to 2,3;22,23-squalene dioxide and thus produces less 2,3;22,23-squalene dioxide hydrolyzed by extra-pathway products. In some embodiments, the EPH enzyme comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% sequence identity relative to SEQ ID NO: 5. In some embodiments, the EPH has one or more mutations relative to SEQ ID NO: 5 that increase enzyme productivity or substrate preference for 24,25-epoxycucurbita-dienol relative to 2,3;22,23-squalene dioxide, and said mutations are selected from Table 6. In some embodiments, the EPH enzyme contains at least 2, at least 3, at least 4, or at least 5 amino acid substitutions listed in Table 6 relative to the amino acid sequence of SEQ ID NO: 5.
[0045] In some embodiments, the epoxide hydrolase contains one or more mutations relative to SEQ ID NO: 5 at one or more positions selected from S68, Q128, G141, T144, A145, E163, E191, L262, C295, and N299. In some embodiments, S68 is substituted with a nonpolar amino acid selected from Ala, Gly, and Val. In some embodiments, Q128 is substituted with a nonpolar amino acid selected from Ala, Pro, Gly, and Val. In some embodiments, G141 is substituted with a polar amino acid selected from Ser, Thr, and Tyr. In some embodiments, T144 is substituted with an amino acid selected from Gln, Asn, Ala, Gly, and Val. In some embodiments, A145 is substituted with a nonpolar amino acid selected from Val, Ile, and Leu. In some embodiments, E163 is substituted with a nonpolar amino acid selected from Ala, Gly, and Val. In some embodiments, E191 is substituted with a nonpolar amino acid selected from Gly, Ala, and Val. In some embodiments, L262 is substituted with Met. In some embodiments, C295 is substituted with a nonpolar amino acid selected from Gly, Ala, and Val. In some embodiments, N299 is substituted with Gln. In some embodiments, the epoxide hydrolase has at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or all ten mutations selected from S68A, Q128A, G141S, T144Q, A145V, E163A, E191G, L262M, C295G, and N299Q relative to SEQ ID NO: 5. An exemplary EPH enzyme according to this disclosure comprises the amino acid sequence of SEQ ID NO: 6.
[0046] Further amino acid modifications of EPH enzymes can be guided by available enzyme structures and homology models, including those described in the following literature: Mitusińska, K, et al. Structure-function relationship between soluble epoxide hydrolases structure and their tunnel network , Comput Struct Biotechnol J. 2022; 20: 193–205; Zou et al., Structure of Aspergillus niger epoxide hydrolase at 1.8 Å resolution: implications for the structure and function of the mammalian microsomal class of epoxide hydrolases , Structure 8(2): 111-122 (2000); Schulz et al., The crystal structure of mycobacterial epoxide hydrolase A , Scientific Reports 10: 16539 (2020); Gomez et al., Human soluble epoxide hydrolase: Structural basis of inhibition by 4-(3- cyclohexylureido)-carboxylic acids ,Protein Sci. 15(1): 58-64 (2006); Biswal et al., The molecular structure of epoxide hydrolase B from Mycobacterium tuberculosis and its complex with a urea-based inhibitor , J Mol Biol 381: 897-912 (2008); Nardini et al., The X-ray structure of epoxide hydrolase from Agrobacterium radiobacter AD1. An enzyme to detoxify harmful epoxides , J Biol Chem 274(21):14579-86 (1999).
[0047] In various aspects and embodiments, this disclosure provides a cytochrome P450 enzyme for producing mogroside from 24,25-dihydroxycucurbita diol. In some embodiments, the cytochrome P450 enzyme selectively hydroxylates the C11 of 24,25-dihydroxycucurbita diol. In some embodiments, the cytochrome P450 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 14 or SEQ ID NO: 15.
[0048] The functional expression of plant cytochrome P450 in heterologous hosts has inherent limitations attributable to the heterologous host. For example, bacterial systems exhibit translational incompatibility of the membrane signaling modules of P450 enzymes due to the absence of the endoplasmic reticulum. Furthermore, the expression and function of cytochrome P450 enzymes become restricted in heterologous hosts. In various aspects and embodiments, this disclosure provides engineered cytochrome P450 enzymes that efficiently express and selectively hydroxylate the C11 of 24,25-dihydroxycucurbitadienol in microbial hosts, including bacterial hosts.
[0049] In some embodiments, such as those disclosed in U.S. Patent No. 10,774,314, the cytochrome P450 enzyme has all or part of its native transmembrane domain replaced by a transmembrane domain derived from a C-terminal protein from the E. coli endomembrane cytoplasm, which is incorporated herein by reference in its entirety. In some embodiments, the E. coli endomembrane cytoplasmic C-terminal protein is selected from ycgG, yhcB, zipA, waaA, sohB, djlA, ycgG, lpxK, ypfN, ypfN, yhhM, and FliO. In some embodiments, the E. coli endomembrane cytoplasmic C-terminal protein is sohB or FliO.
[0050] In various embodiments, the P450 enzyme has a partial or complete deletion or truncation of its native transmembrane domain. Typically, the deletion or truncation is about the first 15 to 75 amino acids, and the desired length can be determined using prediction tools known in the art. For example, the deletion / truncation of the native transmembrane region can be about 15 to about 30 amino acids. The native P450 transmembrane region is replaced by a membrane-anchoring sequence derived from an *E. coli* protein, such as sohB or FliO, whose C-terminus is located in the cytoplasm. For example, the P450 enzyme described herein with SEQ ID NO: 15 corresponds to CYP87A3 (zucchini (… Cucurbita pepo The subspecies *pepo* has a portion (18 amino acids) of its natural transmembrane domain replaced by the transmembrane domain of *E. coli* sohB.
[0051] Cytochrome P450 enzymes can exhibit broad substrate specificity or low substrate selectivity. See, for example, Kim et al. Genetic engineering, purification, crystallization and preliminary X-ray diffraction of cytochrome P450 p-coumarate-3-hydroxylase (C3H), the Arabidopsis membrane protein , Protein Expression and Purification 79(1): 149-155 (2011); Yamaguchi et al. Cytochrome P450 CYP71AT96 catalyses the final step of herbivore-induced phenylacetonitrile biosynthesis in the giant knotweed, Fallopia sachalinensis , Plant Molecular Biology 91: 229-239 (2016). The broad substrate specificity and low substrate selectivity of cytochrome P450 enzymes lead to wasteful extra-pathway reactions and reduced levels of desired molecule oxidation. This document discloses engineered cytochrome P450 enzymes with enhanced specificity for the C-11 position of 24,25-dihydroxycucurbita dienol. In some embodiments, the engineered cytochrome P450 enzymes disclosed herein exhibit enhanced substrate selectivity for the C-11 position of 24,25-dihydroxycucurbita dienol compared to the C-11 position of cucurbita dienol, thereby allowing for the production of higher levels of mogrosides compared to 11-hydroxycucurbita dienol (see, for example...). Figure 7G and Figure 7I ).
[0052] Therefore, in some embodiments, when the total amount of the substrates cucurbitacinol and 24,25-dihydroxycucurbitacinol, and the products mogroside, 11-hydroxycucurbitacinol, and 11-oxo-mogrosideol is considered 100%, microbial cells expressing the engineered cytochrome P450 (and SQE, ECDS, and EPH enzymes disclosed herein) produce less than 40%, less than 30%, or less than 20% of 11-hydroxycucurbitacinol and 11-oxo-mogrosideol. In some embodiments, when the total amount of the products mogrosideol, 11-hydroxycucurbitacinol, and 11-oxo-mogrosideol is considered 100%, microbial cells expressing the engineered cytochrome P450 (and SQE, ECDS, and EPH enzymes disclosed herein) produce more than 60%, more than 70%, or more than 80% of mogrosideol. In some embodiments, microbial cells expressing the engineered cytochrome P450 disclosed herein produce fewer extra-pathway products, 11-hydroxycucurbita diol and 11-oxo-mogroside, compared to microbial cells expressing non-engineered cytochrome P450 enzymes. In some embodiments, microbial cells expressing the engineered cytochrome P450 disclosed herein produce more mogroside compared to microbial cells expressing non-engineered cucurbita diol synthase.
[0053] In various embodiments, the cytochrome P450 enzyme comprises an S193T substitution relative to the enzyme of SEQ ID NO: 15, which is provided herein as SEQ ID NO: 16. In various embodiments, the cytochrome P450 enzyme comprises one or more mutations or groups of mutations (e.g., at least 2, 3, 4, or 5 mutations or groups of mutations) listed in Tables 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, and / or 29 relative to the corresponding parent enzyme. In various embodiments, the P450 enzyme comprises at least two, at least three, at least four, or at least five mutants or mutant groups listed in Tables 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, and / or 29, which increase the mogrool titer by at least 1.5 times relative to the corresponding parent enzyme.
[0054] In some embodiments, the cytochrome P450 enzyme, relative to SEQ ID NO: 15, comprises substitutions at one or more positions selected from A12, I27, S63, K75, V76, N102, G125, T128, W130, L131, K132, T192, S193, L218, G219, T223, V228, T231, T232, Y233, N234, K246, F354, K369, H433, F452, V465, A468, and T482. In some embodiments, S63 is substituted with Asn, Phe, Gln, Tyr, or Trp. In some embodiments, A12 is substituted with Trp, Tyr, or Phe. In some embodiments, I27 is substituted with Gly or Ala. In some embodiments, S63 is substituted with Asn, Phe, Gln, Tyr, or Trp. In some embodiments, K75 is substituted with Arg. In some embodiments, V76 is substituted with Met, Leu, or Ile. In some embodiments, N102 is substituted with His, Lys, or Arg. In some embodiments, G125 is substituted with an amino acid selected from Ala, Leu, Ile, and Val. In some embodiments, T128 is substituted with an amino acid selected from Gly, Ser, Pro, and Ala. In some embodiments, W130 is substituted with an amino acid selected from Asn, Gly, Ser, Leu, Gln, Ala, Val, and Thr. In some embodiments, L131 is substituted with an amino acid selected from Val, Ile, Pro, Thr, and Ser. In some embodiments, K132 is substituted with an amino acid selected from Asn, Ser, Thr, and Gln. In some embodiments, T192 is substituted with Leu, Ile, or Val. In some embodiments, S193 is substituted with Thr. In some embodiments, L218 is substituted with Ile, Val, and Ala. In some embodiments, G219 is substituted with an amino acid selected from Ala, Gln, Arg, Val, Asn, His, and Lys. In some embodiments, T223 is substituted with Ser. In some embodiments, V228 is substituted with Ile or Leu. In some embodiments, T231 is substituted with Phe, Trp, or Tyr. In some embodiments, T232 is substituted with Ala, Gly, or Val. In some embodiments, Y233 is substituted with Phe. In some embodiments, N234 is substituted with His, Arg, or Lys. In some embodiments, K246 is substituted with Val, Ile, or Leu. In some embodiments, F354 is substituted with Tyr or Trp. In some embodiments, K369 is substituted with Arg, Lys, or His. In some embodiments, H433 is substituted with Phe, Trp, or Tyr.In some embodiments, F452 is substituted with Ile, Val, or Leu. In some embodiments, V465 is substituted with an amino acid selected from Ile, Met, or Leu. In some embodiments, A468 is substituted with an amino acid selected from Thr, Ser, Asn, and Gln. In some embodiments, T482 is substituted with Ser. In some embodiments, the cytochrome P450 enzyme comprises at least one, or at least one, selected from A12W, I27G, S63F, S63N, K75R, V76M, N102H, G125A, T128G, W130N, L131I, K132N, T192L, S193T, L218I, G219A, T223S, V228I, T231F, T232A, Y233F, N234H, K246V, F354Y, K369R, H433F, F452I, V465I, A468T, and T482S. Two, or at least three, or at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten, or at least eleven, or at least twelve, or at least thirteen, or at least fourteen, or at least fifteen, or at least sixteen, or at least seventeen, or at least eighteen, or at least nineteen, or at least twenty, or at least twenty-one, or at least twenty-two, or at least twenty-three, or at least twenty-four, or at least twenty-five, or at least twenty-six, or at least twenty-seven, or at least twenty-eight, or at least twenty-nine, or all thirty substitutions, each of which is relative to SEQ ID NO: 15.
[0055] In some embodiments, the cytochrome P450 enzyme comprises, relative to SEQ ID NO: 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, or at least about 50 N-terminal deletions of amino acids. In some embodiments, the cytochrome P450 enzyme comprises deletions selected from L3 to H29 relative to SEQ ID NO: 15. In some embodiments, the cytochrome P450 enzyme comprises an amino acid sequence having at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 85. In some embodiments, the cytochrome P450 enzyme comprises one or more mutations or groups of mutations listed in Tables 26, 27, 28, and / or 29.
[0056] In some embodiments, the cytochrome P450 enzyme contains substitutions at one or more positions selected from K8, D9, S10, N13, V15, I45, K46, K47, M49, K50, R51, I191, A192, F406, E411, Y412, and S413 relative to SEQ ID NO: 85. In some embodiments, K8 is substituted with Arg. In some embodiments, D9 is substituted with Asn or Gln. In some embodiments, S10 is substituted with Arg, Lys, or His. In some embodiments, N13 is substituted with Lys, Arg, or His. In some embodiments, V15 is substituted with Lys, Arg, or His. In some embodiments, I45 is substituted with Val, Ile, or Leu. In some embodiments, K46 is substituted with Arg, Lys, or His. In some embodiments, K47 is substituted with Glu or Asp. In some embodiments, M49 is replaced by Val, Ile, or Leu. In some embodiments, K50 is replaced by Glu or Asp. In some embodiments, R51 is replaced by Lys. In some embodiments, I191 is replaced by Leu or Val. In some embodiments, A192 is replaced by Lys, Arg, or His. In some embodiments, F406 is replaced by Tyr or Trp. In some embodiments, E411 is replaced by Asp. In some embodiments, Y412 is replaced by Phe, Leu, Ile, or Trp. In some embodiments, S413 is replaced by Ala, Gly, or Val. In some embodiments, the cytochrome P450 enzyme comprises, relative to SEQ ID NO:85, at least one, or at least two, or at least three, or at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten, or at least eleven, or at least twelve, or at least thirteen, or at least fourteen, or at least fifteen, or at least eleven, or at least 116, or all seventeen of the following substitutions: K8R, D9N, S10R, N13K and V15K, I45V, K46R, K47E, M49V, K50E, R51K, F406Y, E411D, Y412F, S413A, I191L, and A192K.
[0057] According to embodiments of this disclosure, the cytochrome P450 enzyme comprises the amino acid sequence of SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83 or SEQ ID NO: 84.
[0058] In various embodiments, the cytochrome P450 enzyme comprises, relative to SEQ ID NO: 15, an N-terminal deletion of at least about 20, or at least about 25, or at least about 30, or at least about 35, or at least about 40, or at least about 45, or at least about 50 amino acids. In some embodiments, the cytochrome P450 enzyme comprises a deletion selected from L3 to H29 relative to SEQ ID NO: 15. In some embodiments, the cytochrome P450 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 85, and comprises one or more mutations or groups of mutations listed in Tables 26, 27, 28, and / or 29.
[0059] In some embodiments, the cytochrome P450 enzyme contains substitutions at one or more positions selected from K8, D9, S10, N13, V15, I45, K46, K47, M49, K50, R51, I191, A192, F406, E411, Y412, and S413 relative to SEQ ID NO: 85. In some embodiments, K8 is substituted with Arg. In some embodiments, D9 is substituted with Asn or Gln. In some embodiments, S10 is substituted with Arg, Lys, or His. In some embodiments, N13 is substituted with Lys, Arg, or His. In some embodiments, V15 is substituted with Lys, Arg, or His. In some embodiments, I45 is substituted with Val, Ile, or Leu. In some embodiments, K46 is substituted with Arg, Lys, or His. In some embodiments, K47 is substituted with Glu or Asp. In some embodiments, M49 is replaced by Val, Ile, or Leu. In some embodiments, K50 is replaced by Glu or Asp. In some embodiments, R51 is replaced by Lys. In some embodiments, I191 is replaced by Leu or Val. In some embodiments, A192 is replaced by Lys, Arg, or His. In some embodiments, F406 is replaced by Tyr or Trp. In some embodiments, E411 is replaced by Asp. In some embodiments, Y412 is replaced by Phe, Leu, Ile, or Trp. In some embodiments, S413 is replaced by Ala, Gly, or Val. In some embodiments, the cytochrome P450 enzyme, relative to SEQ ID NO:85, comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen, at least fifteen, at least sixteen, or all seventy-seven substitutions selected from K8R, D9N, S10R, N13K and V15K, I45V, K46R, K47E, M49V, K50E, R51K, F406Y, E411D, Y412F, S413A, I191L and A192K. According to embodiments of this disclosure, the cytochrome P450 enzyme comprises the amino acid sequence of SEQ ID NO: 86, or SEQ ID NO: 87, or SEQ ID NO: 88, or SEQ ID NO: 89.
[0060] In various aspects and embodiments, this disclosure provides a cytochrome P450 reductase capable of reducing the P450 hydroxylase of any embodiment disclosed herein. In some embodiments, the P450 reductase comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with SEQ ID NO:22 or SEQ ID NO:24.
[0061] In some embodiments, the cytochrome P450 reductase comprises replacing all or a portion (e.g., 15 to 75 amino acids) of the native transmembrane domain with a transmembrane domain derived from a C-terminal protein of the E. coli endomembrane cytoplasm (see U.S. Patent No. 10,774,314, which is incorporated herein by reference in its entirety). In some embodiments, the E. coli endomembrane cytoplasmic C-terminal protein is selected from YcgG, YhcB, ZipA, WaaA, SohB, DjlA, LpxK, YpfN, and YhhM. An exemplary P450 reductase (based on SEQ ID NO: 22 of the CPR enzyme) is shown in Table 30, having 67 amino acids truncated from the native transmembrane domain and replaced with a transmembrane sequence (21 to 30 amino acids) derived from a C-terminal protein of the E. coli endomembrane cytoplasm. In some embodiments, the transmembrane sequence is derived from the E. coli protein YcgG. In some embodiments, about 50 to about 80 amino acids, or about 65 to about 78 amino acids (e.g., 65 to 75 amino acids) from the N-terminus of SEQ ID NO: 22 are deleted, and 21 to 28 amino acids from the transmembrane domain of an E. coli inner membrane cytoplasmic C-terminal protein (e.g., YcgG) are inserted. In some embodiments, the native N-terminus of ID NO: 22 (as described) is replaced by 20 to 30 amino acids from the N-terminus of YcgG. In some embodiments, the native N-terminus of ID NO: 22 is replaced by about 25 or 26 amino acids from the N-terminus of YcgG. In some embodiments, the engineered P450 reductase having the native N-terminus of SEQ ID NO: 22 replaced by the N-terminus of YcgG further comprises one or more mutations or mutant groups listed in Table 31 relative to the corresponding parent enzyme. In some embodiments, the CPR enzyme comprises at least two, at least three, at least four, or at least five mutants or mutant groups shown in Table 31, which increase mogroside yield by at least 1.2-fold relative to the corresponding parent enzyme. Exemplary CPR enzymes according to this disclosure include those comprising the amino acid sequence of SEQ ID NO: 23.
[0062] In other embodiments, the P450 reductase (based on SEQ ID NO: 24, the CPR enzyme) is shown in Table 32, having approximately 75 amino acids truncated from the native transmembrane domain and replaced by a transmembrane sequence (21 to 30 amino acids) from the C-terminal protein of the E. coli inner membrane cytoplasm. In some embodiments, the transmembrane sequence is derived from the E. coli protein YcgG. In some embodiments, the native N-terminus (approximately 75 amino acids) of SEQ ID NO: 24 is replaced by approximately 25 amino acids from the N-terminus of YcgG. In some embodiments, the engineered P450 reductase having the native N-terminus of SEQ ID NO: 24 replaced by the N-terminus of YcgG further comprises one or more mutations or mutant groups listed in Table 33 relative to the corresponding parent enzyme. In some embodiments, the CPR enzyme comprises at least two, at least three, at least four, or at least five mutations or mutant groups shown in Table 33, which increase mogroside yield. Exemplary CPR enzymes according to this disclosure include those comprising the amino acid sequences of SEQ ID NO: 25 and SEQ ID NO: 26. In some embodiments, the cytochrome P450 reductase comprises a substitution of the L90 residue relative to SEQ ID NO: 24, optionally substituted with Pro.
[0063] In various aspects and embodiments, this disclosure provides a uridine diphosphate-dependent glycosyltransferase (UGT) for the production of mogrosides from mogrosides. Exemplary mogrosides that can be produced are shown in... Figure 1B This includes Mog 1A, Mog.1E, Mog. IIA1, Mog. IIE, Mog. IIA2, Mog. IIIA1, Mog. III, Mog. IIIA2, sarmancoside, Mog. IVA, Mog. IV, and Mog. V. In some embodiments, mogroside is Mog. V. Other mogrosides, such as iso-Mog. V, can be produced. In some embodiments, the UGT enzyme produces at least 2 times more of the desired mogroside (e.g., Mog. V) than the glycosylated product of 24,25-dihydroxycucurbitadienol. In some embodiments, the UGT enzyme produces at least 3 times, or at least 5 times, or at least 10 times more of the desired mogroside (e.g., Mog. V) than the glycosylated product of 24,25-dihydroxycucurbitadienol.
[0064] Urate diphosphate-dependent glycosyltransferases (UGTs) can exhibit substrate diversity. See, for example, Song et al. Functional Characterization and Substrate Promiscuity of UGT71 Glycosyltransferases from Strawberry ( Fragaria x ananassa ) . Plant and Cell Physiology , 56: 2478-2493 (2015); Bönisch et al., Activity based profiling of a A physiologic aglycone library reveals sugar acceptor promiscuity of family 1UDP-glucosyltransferases from Vitis vinifera . Plant Physiology , 166: 23-39(2014); Bönisch et al., A UDP-glucose: Monoterpenol glucosyltransferase adds to the chemical diversity of the grapevine metabolome . Plant Physiology , 165:561-581 (2014); Hefner et al., Arbutin synthase, a novel member of the NRD1β glycosyltransferase family, is a unique multifunctional enzyme converting various natural products and xenobiotics , Bioorganic and Medicinal Chemistry 10: 1731-1741 (2002). The substrate diversity of UGT enzymes leads to wasteful extra-pathway reactions and reduced levels of desired molecule oxidation. This paper discloses engineered cytochrome UGT enzymes with enhanced specificity for mogrool at the C-3 or C-24 position or for 1,6 and 1,2 branched glycosylation.
[0065] In some aspects and embodiments, this disclosure provides engineered UGT enzymes that exhibit enhanced substrate selectivity for the C-3 position of mogroside, compared to positions other than, for example, the C-3 position of 24,25-dihydroxycucurbitadienol or the C-3 position of mogroside. Such engineered enzymes are referred to herein as UGTc3. Therefore, in some embodiments, the UGTc3 enzymes disclosed herein provide higher levels of Mog1E yield compared to glucosylated dihydroxycucurbitadienol and deoxy-MogIE byproducts (see, for example...). Figure 9C , Figure 9E and Figure 9F In some embodiments, when total glucosylation products are considered 100%, the UGTc3 enzyme disclosed herein uses an equimolar mixture comprising mogroside, deoxy-mogroside, and dihydroxycucurbita periol as substrates to produce less than 60%, or less than 50%, or less than 40%, or less than 30%, or less than 20% of glucosylated dihydroxycucurbita periol and deoxy-MogIE byproducts. In some embodiments, compared to non-engineered UGT enzymes, the UGTc3 enzyme disclosed herein uses a mixture comprising mogroside, deoxy-mogroside, and dihydroxycucurbita periol to produce fewer extra-pathway products of glucosylated dihydroxycucurbita periol and deoxy-MogIE. In some embodiments, compared to culturing microbial cells expressing non-engineered UGT enzymes, the UGTc3 enzyme disclosed herein uses a mixture comprising mogroside, deoxy-mogroside, and dihydroxycucurbita periol to produce more Mog1E.
[0066] In various aspects and embodiments, this disclosure provides a uridine diphosphate-dependent glycosyltransferase (UGT) (UGTc3) for the production of C-3 mogroside from mogroside. In some embodiments, the UGTc3 enzyme comprises an amino acid sequence having at least 70% identity, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequences of SEQ ID NO: 27, SEQ ID NO: 52, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 61, SEQ ID NO: 62, SEQ ID NO: 63 and SEQ ID NO: 64 (see Table 34). In some embodiments, the UGTc3 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 27. In some embodiments, the UGT enzyme converts Mog. 1A to Mog. IIE and mogroside to Mog. IE. In some embodiments, the UGT enzyme comprises one or more mutations relative to SEQ ID NO: 27, which result in higher specificity and / or activity for glycosylation at the C-3 position (e.g., in the case of mogroside and / or Mog. IE substrates) compared to the UGT enzyme having the amino acid sequence of SEQ ID NO: 27. In some embodiments, one or more mutations (e.g., at least two, at least three, at least four, or at least five mutations) are selected from Tables 35 and / or 36. In some embodiments, at least two, three, four, or five mutations selected from Tables 35 and 36 are shown to increase the Mog. IIE titer or Mog. IE titer by at least 1.2-fold or at least 1.3-fold. In various embodiments, UGT enzymes with C3 activity exhibit specificity for reducing 24,25-dihydroxycucurbitadienol and its glycosylated derivatives.
[0067] In some embodiments, the UGTc3 enzyme, relative to SEQ ID NO: 27, comprises one or more mutations at one or more positions selected from L41, D49, T74, C127, F303, and A307. In some embodiments, L41 is substituted with an aromatic amino acid selected from Phe, Tyr, and Trp. In some embodiments, D49 is substituted with Glu. In some embodiments, T74 is substituted with Met. In some embodiments, C127 is substituted with an aromatic amino acid selected from Phe, Tyr, and Trp. In some embodiments, F303 is substituted with Tyr. In some embodiments, A307 is substituted with an amino acid selected from Ile, Leu, and Val. In some embodiments, the UGTc3 enzyme, relative to SEQ ID NO: 27, comprises at least one, at least two, at least three, at least four, at least five, or all six of the following substitutions: L41F, D49E, T74M, C127F, F303Y, and A307I. According to this disclosure, an exemplary engineered UGT enzyme having activity at the C3 site of mogroside comprises the amino acid sequence of SEQ ID NO:28, SEQ ID NO:29, or SEQ ID NO:30.
[0068] In various aspects and embodiments, this disclosure provides a UGT enzyme (UGTc3) for producing C-3 mogroside from mogroside. In some embodiments, the UGTc3 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 60 or SEQ ID NO: 65. The enzyme defined by SEQ ID NO: 65 is a circular variant of the enzyme defined by SEQ ID NO: 60. Circular variants of UGT are described in U.S. Patent No. 10,463,062, which is incorporated herein by reference in its entirety. In some embodiments, the UGT enzyme comprises one or more mutations relative to SEQ ID NO: 60 or SEQ ID NO: 65, which result in higher specificity and / or activity for glycosylation at the C-3 position compared to a UGT enzyme having the amino acid sequence of SEQ ID NO: 60 or SEQ ID NO: 65. In some embodiments, one or more mutations (e.g., at least 2, at least 3, at least 4, or at least 5 mutations) relative to SEQ ID NO: 60 are selected from Table 37. In some embodiments, at least 1, or at least 2, or at least 3, or at least 4, or at least 5 mutations are selected from Table 37, and the mutations show that they increase the Mog IE titer by at least 1.4-fold or at least 1.5-fold. In some embodiments, the mutation is a circular arrangement mutation having the cleavage sites shown in Table 38 and optionally having a linker sequence. In some embodiments, one or more mutations or groups of mutations (e.g., at least 2, at least 3, at least 4, or at least 5 mutations or groups of mutations) relative to SEQ ID NO: 65 are selected from Table 39. In some embodiments, at least 1, or at least 2, or at least 3, or at least 4, or at least 5 mutations or groups of mutations are selected from Table 39, and the mutations show that they increase the Mog IE titer by at least 1.4-fold or at least 1.5-fold.
[0069] In some embodiments, the UGTc3 enzyme is a circular variant relative to SEQ ID NO: 60. In some embodiments, the circular variant has a cleavage site corresponding to amino acids 150 to 180, or 160 to 180, or 165 to 180 of SEQ ID NO: 60 (e.g., in some embodiments, the cleavage site may be located at position 169 or 177). In some embodiments, the circular variant further comprises a linker sequence located between amino acids corresponding to the N-terminal and C-terminal residues of SEQ ID NO: 60. For example, in various embodiments, the linker may have a length of 2 to 25 amino acids and may be a rigid or flexible linker. A flexible linker may consist primarily or entirely of Ser and Gly amino acids. More rigid linkers may be constructed from larger amino acids, and examples are shown in Table 38.
[0070] In some embodiments, the UGTc3 enzyme relative to SEQ ID NO: 65 includes one or more mutations at one or more positions selected from H40, I216, S241, and L352. In some embodiments, H40 is substituted with Asp, Gly, Gln, Ala, or Val. In some embodiments, I216 is substituted with Ala, Gly, or Val. In some embodiments, S241 is substituted with Gly, Ala, or Val. In some embodiments, L352 is substituted with Met. In some embodiments, the UGTc3 enzyme relative to SEQ ID NO: 65 includes at least one, at least two, at least three, or all four of the following substitutions: H40N, I216A, S241G, and L352M. Exemplary enzymes according to these embodiments are provided herein as SEQ ID NO: 66 and SEQ ID NO: 67.
[0071] In some aspects and embodiments, this disclosure provides engineered UGT enzymes that exhibit enhanced substrate selectivity for the C-24 position of mogroside compared to positions other than, for example, the C-24 position of 24,25-dihydroxycucurbitadienol or the C-24 position of mogroside. Such engineered enzymes are referred to herein as UGTc24. Therefore, in some embodiments, the UGTc24 enzymes disclosed herein provide higher levels of Mog1A production compared to deoxy-MogIA byproducts (see, for example...). Figure 10EIn some embodiments, when total glucosylation products are considered 100%, the UGTc24 enzyme disclosed herein uses a mixture comprising equal amounts of mogromectin and deoxy-mogromectin (in molar amounts) as a substrate, producing less than 40%, or less than 30%, or less than 20% of the deoxy-MogIA byproduct. In some embodiments, compared to non-engineered UGT enzymes, the UGTc24 enzyme disclosed herein uses a mixture comprising mogromectin and deoxy-mogromectin as a substrate, producing less of the extra-pathway product deoxy-MogIA. In some embodiments, compared to culturing microbial cells expressing non-engineered UGT enzymes, the UGTc24 enzyme disclosed herein uses a mixture comprising mogromectin and deoxy-mogromectin as a substrate, producing more Mog1A.
[0072] In various aspects and embodiments, this disclosure provides a UGT enzyme (UGTc24) for producing C-24 mogroside from mogroside. In some embodiments, the UGTc24 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with an amino acid sequence selected from SEQ ID NO: 31, SEQ ID NO: 50, SEQ ID NO: 51, SEQ ID NO: 63, and SEQ ID NO: 64 (see Table 34). In some embodiments, the UGTc24 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 31. In some embodiments, the UGTc24 enzyme comprises one or more mutations that result in higher specificity and / or activity for glycosylation at the C-24 position (e.g., including but not limited to the conversion of Mog. 1E to Mog. 1IE and / or the conversion of mogroside to Mog. 1A) compared to the UGT enzyme having the amino acid sequence of SEQ ID NO: 31. In some embodiments, the UGTc24 enzyme comprises one or more mutations or mutation groups (e.g., at least 2, 3, 4, or 5 mutations or mutation groups) selected from Tables 40, 41, 42, 43, 44, and / or 45 (in each case relative to the parent enzyme). In some embodiments, the UGTc24 enzyme comprises one or more mutations or mutation groups (e.g., at least 2, 3, 4, or 5 mutations or mutation groups) selected from Tables 40, 41, 42, 43, 44, and / or 45, showing that the one or more mutations or mutation groups increase the Mog. IIE titer or Mog. IA titer by at least 1.5-fold or at least 2-fold. In some embodiments, at least one, at least two, at least three, at least four, or at least five mutations or mutation groups increase the specificity for conversion to Mog. IA relative to deoxy-Mog. IA (e.g., see Tables 43 and 44).
[0073] In some embodiments, the UGTc24 enzyme contains one or more mutations relative to SEQ ID NO: 31 at positions selected from L18, A74, S88, I91, T95, H101, T184, F187, P190, T273, N306, M332, and Y381. In some embodiments, L18 is substituted with Met. In some embodiments, I63 is substituted with Val, Ala, or Gly. In some embodiments, A74 is substituted with an acidic amino acid selected from Glu and Asp. In some embodiments, S88 is substituted with an amino acid selected from Ala, Gly, Leu, Val, and Ile. In some embodiments, I91 is substituted with an aromatic amino acid selected from Phe, Tyr, and Trp. In some embodiments, T95 is substituted with an amino acid selected from Ala, Gly, and Val. In some embodiments, H101 is substituted with Pro. In some embodiments, T184 is substituted with an amino acid selected from Phe, Tyr, and Trp. In some embodiments, F187Y is substituted with Tyr or Trp. In some embodiments, P190 is substituted with an amino acid selected from Glu and Asp. In some embodiments, T273 is substituted with Ser or is deleted. In some embodiments, N306 is substituted with Lys, His, or Arg. In some embodiments, M332 is substituted with an amino acid selected from Leu, Ile, Val, and Ala. In some embodiments, Y381 is substituted with an amino acid selected from Phe and Trp. In some embodiments, the UGTc24 enzyme, relative to SEQ ID NO: 31, comprises at least one, or at least two, or at least three, or at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten, or at least eleven, or at least twelve, or at least thirteen, or all fourteen mutations selected from L18M, I63V, A74E, S88A, I91F, T95A, H101P, T184F, F187Y, P190E, T273S, N306K, M332L, and Y381F. Exemplary enzymes according to embodiments of this disclosure comprise amino acid sequences selected from SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 48, and SEQ ID NO: 49.
[0074] In various aspects and embodiments, this disclosure provides a uridine diphosphate-dependent glycosyltransferase (UGT) (UGT1216) that produces 1,6- and / or 1,2-branched glycosides of C-3 and / or C-24 glycosides of mogroside. In some embodiments, UGT1216 is a circular variant of the enzyme represented by SEQ ID NO: 35 or a sequence having at least 70% or at least 90% sequence identity therewith. In some embodiments, the circular variant has a cleavage site at positions 180 to 195 of the sequence corresponding to SEQ ID NO: 35 or a sequence having at least 70% sequence identity therewith. In some embodiments, the circular variant has a cleavage site at positions 185 to 190 of the sequence corresponding to SEQ ID NO: 35 or a sequence having at least 70% sequence identity therewith. In some embodiments, the circular variant has a cleavage site at position 187 of the sequence corresponding to SEQ ID NO: 35 or a sequence having at least 70% sequence identity therewith.
[0075] In some embodiments, the UGT1216 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 36. In some embodiments, compared to the UGT enzyme having the amino acid sequence of SEQ ID NO: 36, the UGT1216 enzyme has one or more mutations, said mutations resulting in higher specificity and / or activity for 1,6 and / or 1,2 glycosylation of mogrosides C-3 and / or C-24. In some embodiments, compared to the UGT enzyme having the amino acid sequence of SEQ ID NO: 36, the UGT1216 enzyme has one or more mutations, said mutations resulting in the production of more MogV from MogIIE. In some embodiments, the UGT1216 enzyme has one or more mutations or mutation groups selected from Tables 46, 47, 48, and / or 49 relative to the corresponding parent enzyme (e.g., at least 2, 3, 4, or 5 mutations or mutation groups). In some embodiments, the UGT1216 enzyme has one or more mutations or mutation groups selected from Tables 46, 47, 48, and / or 49 relative to the corresponding parent enzyme (e.g., at least 2, 3, 4, or 5 mutations or mutation groups), and said mutations or mutation groups show that the Mog. V titer is increased by at least 1.5 or at least 2.0.
[0076] In some embodiments, the UGT1216 enzyme, relative to SEQ ID NO: 36, includes mutations at one or more positions selected from L14, W160, L266, and H390. In some embodiments, L14 is substituted with a nonpolar amino acid selected from Val, Ile, Ala, and Gly. In some embodiments, W160 is substituted with Pro, Ser, or Thr. In some embodiments, L266 is substituted with a polar amino acid selected from Gln and Asn. In some embodiments, H390 is substituted with an amino acid selected from Asp and Glu. In some embodiments, the UGT1216 enzyme, relative to SEQ ID NO: 36, includes at least one, at least two, at least three, or all four mutations selected from L14V, H19E, W160P, L266Q, T281S, S318K, and H390D. The exemplary UGT1216 enzyme according to this disclosure comprises an amino acid sequence selected from SEQ ID NO: 37, SEQ ID NO: 38 and SEQ ID NO: 90.
[0077] In various aspects and embodiments, this disclosure provides a uridine diphosphate-dependent glycosyltransferase (UGT) (UGT1216) for producing 1,6- and / or 1,2-branched glycosides of mogrosides (C-3 and / or C-24 glycosides). In some embodiments, the UGT1216 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 39. In some embodiments, the UGT1216 enzyme has one or more mutations compared to the UGT enzyme having the amino acid sequence of SEQ ID NO: 39, said mutations resulting in higher specificity and / or activity for the 1,6- and 1,2-branched glycosylation of mogrosides (C-3 and / or C-24 glycosides). In some embodiments, the UGT1216 enzyme has one or more mutations compared to the UGT enzyme having the amino acid sequence of SEQ ID NO: 39, said one or more mutations resulting in the production of more Mog. IVA and / or more Mog. V from Mog. IIE. In some embodiments, the UGT1216 enzyme has one or more mutations or mutation groups (e.g., at least 2, at least 3, at least 4, or at least 5 mutations or mutation groups) listed in Table 50 relative to SEQ ID NO: 39. In some embodiments, the enzyme has one or more mutations or mutation groups (e.g., at least 2, at least 3, at least 4, or at least 5 mutations or mutation groups) listed in Table 50 relative to SEQ ID NO: 39, and said one or more mutations or mutation groups show providing at least 5-fold Mog. V titers.
[0078] In some embodiments, the UGT1216 enzyme contains mutations at positions selected from T147 and N207, relative to SEQ ID NO: 39. In some embodiments, T147 is substituted with a nonpolar amino acid selected from Leu, Ile, and Val. In some embodiments, N207 is substituted with a basic amino acid selected from Lys, Arg, and His. In some embodiments, the UGT1216 enzyme contains T147L and / or N207K substitutions, relative to SEQ ID NO: 39.
[0079] In various aspects, this disclosure provides a uridine diphosphate-dependent glycosyltransferase (UGT) (UGT1216) for producing 1,6- and / or 1,2-branched glycosides of C-3 and / or C-24 glycosides of mogroside, wherein the UGT1216 enzyme comprises an amino acid sequence having at least 80%, or at least 85%, at least 90%, or at least 95%, or at least 97%, or at least 98% identity with the amino acid sequence of SEQ ID NO: 38, and comprises a T147L and / or N207K mutation relative to SEQ ID NO: 39. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 40.
[0080] In various aspects, this disclosure provides a microbial cell that produces squalene and expresses a biosynthetic pathway for converting squalene into mogroside or mogroside, wherein said microbial cell expresses one or more engineered enzymes of any aspect or embodiment disclosed herein.
[0081] In various embodiments, the microbial cells further express heterologous farnesyl diphosphate synthase (FPPS) and squalene synthase (SQS) to convert metabolic precursors from the MEP or MVA pathway into squalene. SQS and FPPS enzymes are known in the art. Exemplary SQS and FPPS enzymes (as well as other enzymes for the production of mogrosides and mongholic acid in microbial host cells) are described in US 2021 / 0032669, which is incorporated herein by reference in its entirety.
[0082] The microbial host cell in each embodiment can be prokaryotic or eukaryotic. In some embodiments, the microbial host cell is a bacterium, and it may optionally be selected from species of the genus *Escherichia* (…). Escherichia spp ), Bacillus species ( Bacillus spp ), species of Corynebacterium ( Corynebacterium spp ), species of the genus Rhodophyta ( Rhodobacter spp ), species of the genus *Fermentomonas* ( Zymomonas spp ), species of Vibrio ( Vibrio spp ) and species of the genus Pseudomonas ( Pseudomonas sppFor example, in some implementations, the bacterial host cell is selected from *Escherichia coli* (E. coli). Escherichia coli Bacillus subtilis ( Bacillus subtilis ), Corynebacterium glutamicum ( Corynebacterium glutamicum ), capsular red bacteria ( Rhodobacter capsulatus ), Rhodopsycetes ( Rhodobacter sphaeroides ), biogenic fermentation monoclonal bacteria ( Zymomonas mobilis ), sodium-dependent Vibrio ( Vibrio natriegens ) or Pseudomonas putida ( Pseudomonas putida The species is *Escherichia coli*. In some embodiments, the bacterial host cell is *Escherichia coli*. Alternatively, the microbial cell may be a yeast cell, such as, but not limited to, *Yeast*. Saccharomyces ), Pichia pastoris ( Pichia ) or Yersinia genus ( Yarrowia ) species, including brewer's yeast ( Saccharomyces cerevisiae Pichia pastoris () Pichia pastoris ) and Yersinia lipophila ( Yarrowia lipolytica ).
[0083] Surprisingly, ybhGFSR Overexpression of the operon reduces the production of mogroside and its precursor 24,25-dihydroxycucurbitadienol. On the other hand, ybhGFSR Sub-effective or null mutations in the operon (e.g., reducing gene expression or activity in the ybhGFSR operon) increase the production of mogroside and / or its direct precursor 24,25-dihydroxycucurbitadienol. It has been found that... ybhGFSR Overexpression of the operon increases the production of 2,3;22,23-squalene dioxide, which is an early intermediate in the synthesis of mogrool. ybhGFSR Decreased expression of transporter proteins increases intracellular (int.) 2,3;22,23-squalene dioxide and enhances conversion to downstream products 24,25-dihydroxycucurbitadienol and mogroside.
[0084] Therefore, in various respects, this disclosure relates to a microbial cell for producing mogroside or mogroside, wherein said microbial cell has ybhGFSRThe expression or activity of one or more genes of the operon is reduced, and the microbial cell expresses a biosynthetic pathway that converts squalene into mogroside or mogroside. In various embodiments, the expression of one or more genes (e.g., ybhG or F) of the ybhGFSR operon is modified by replacing or modifying the natural promoter, modifying the RBS, or modifying the gene to increase the rate of mRNA degradation. In various embodiments, one or more ybhGFSR operon genes are engineered to include sub-effect mutations, thereby reducing activity. In some embodiments, one or more genes of the ybhGFSR operon are disrupted (e.g., deleted) (e.g., ybhF).
[0085] The YbhGFSR pump is a member of the ATP-binding cassette (ABC) transporter superfamily. ABC transporters, also known as ABC efflux pumps, are ubiquitous in all organisms because they use energy from the hydrolysis of ATP to ADP to actively transport various molecules (such as ions, xenobiotics, drugs, sugars, amino acids, lipids, and oligopeptides) across lipid membranes. For example, it is estimated that *E. coli* possesses 79 ABC proteins, including 57 ABC pumps, making it the largest family of paralogous proteins. Linton and Higgins, The Escherichia coli ATP-binding cassette (ABC) proteins , Mol. Microbiol ., 28: 5-13 (1998).
[0086] In various aspects and embodiments, this disclosure provides a microbial cell for producing mogroside or mogroside, wherein the microbial cell carries one or more gene modifications that reduce the efflux of intermediates in mogroside synthesis, thereby increasing the efficiency of mogroside synthesis compared to cells without said gene modifications. In some embodiments, the microbial cell contains mutations that reduce the amount or activity of transport proteins. In some embodiments, the transport protein is an ABC transport protein. In some embodiments, the transport protein is YbhGFSR or its ortholog. In some embodiments, the microbial cell expresses a biosynthetic pathway that converts squalene into mogroside or mogroside. In some embodiments, the microbial cell contains... ybhG , ybhF , ybhS and ybhR One or more of the sub-effective mutations or null mutations in the middle.
[0087] In some implementations, microbial cells are contained within ybhG , ybhF , ybhS and ybhR One or more invalid mutations in the [missing information]. In some implementations, invalid mutations are [missing information]. ybhG ,ybhF , ybhS and ybhR One or more of the motifs may be completely or partially missing. In some implementations, the null mutation is generated by mutating one or more of Walker A, Q ring, and Walker B, as well as the characteristic motifs of the D ring and the transformed motif. ybhF ATPase ineffective mutant.
[0088] In some implementations, microbial cells are contained within ybhG , ybhF , ybhS and ybhR One or more sub-effect mutations in the gene. In some embodiments, a sub-effect mutation is a mutation that reduces the expression of one or more genes, such as a promoter mutation, and / or one or more mutations selected from mutated ribosome binding sites, the addition of one or more rare codons, and in... ybhG , ybhF , ybhS and ybhR One or more of the RNA secondary structures are mutated. In some embodiments, the RBS of ybhG is modified to reduce expression, and / or to disrupt ybhF expression (e.g., ybhF is deleted). In some embodiments, strategies including but not limited to the following are used to reduce... ybhG , ybhF , ybhS and ybhR The expression and / or activity of one or more of them: ybhGFSR The promoter is mutated to a weaker promoter than the natural promoter; the ribosome binding site (RBS) is mutated to... ybhG , ybhF , ybhS and ybhR One or more of the mRNAs have one or more RBSs that are weaker than the natural RBSs; making ybhG , ybhF , ybhS and ybhR One or more mutations in the protein are used to include one or more protease recognition sites, thereby increasing the degradation rate of the YbhGFSR transporter; ybhG , ybhF , ybhS and ybhR One or more mutations in the mRNA are used to include secondary structures in its mRNA, thereby reducing the translation rate; ybhG , ybhF , ybhS and ybhR One or more mutations in the formula to include rare codons to reduce the translation rate; or a combination thereof.
[0089] Microbial cells will produce MEP or MVA products, which serve as substrates for heterologous enzyme pathways. The MEP (2-C-methyl-D-erythritol 4-phosphate) pathway, also known as the MEP / DOXP (2-C-methyl-D-erythritol 4-phosphate / 1-deoxy-D-xylulose 5-phosphate) pathway, or the mevalonate-free pathway or mevalonate-independent pathway, refers to the pathway that converts glyceraldehyde-3-phosphate and pyruvate into IPP and DMAPP. The pathway present in bacteria typically involves the action of the following enzymes: 1-deoxy-D-xyulose-5-phosphate synthase (Dxs), 1-deoxy-D-xyulose-5-phosphate reductase (IspC), cytidine-2-C-methyl-D-erythritol synthase (IspD), cytidine-2-C-methyl-D-erythritol kinase (IspE), 2-C-methyl-D-erythritol 2,4-cyclic diphosphate synthase (IspF), 1-hydroxy-2-methyl-2-(E)-butenyl-4-bisphosphate synthase (IspG), and isopentenyl diphosphate isomerase (IspH). The MEP pathway, along with the genes and enzymes constituting it, is described in US 8,512,988, which is incorporated herein by reference in its entirety. For example, genes constituting the MEP pathway include dxs, ispC, ispD, ispE, ispF, ispG, ispH, idi, and ispA. In some embodiments, the host cell expresses or overexpresses one or more of dxs, ispC, ispD, ispE, ispF, ispG, ispH, idi, ispA, or modified variants thereof, which leads to increased production of IPP and DMAPP. In some embodiments, the triterpenoid skeleton is at least partially produced by metabolic flux through the MEP pathway, and said host cell has at least one additional copy of one or more of dxs, ispC, ispD, ispE, ispF, ispG, ispH, idi, ispA, or modified variants thereof.
[0090] The MVA pathway refers to the biosynthetic pathway that converts acetyl-CoA into IPP. The mevalonate pathway, which is present in yeast, typically involves enzymes that catalyze the following steps: (a) condensing two acetyl-CoA molecules to acetoacetyl-CoA (e.g., by the action of acetoacetyl-CoA thiolase); (b) condensing acetoacetyl-CoA with acetyl-CoA to form hydroxymethylglutaryl-CoA (HMG-CoA) (e.g., by the action of HMG-CoA synthase (HMGS)); (c) converting HMG-CoA to mevalonate (e.g., by the action of HMG-CoA reductase (HMGR)); (d) phosphorylating mevalonate to mevalonate 5-phosphate (e.g., by the action of mevalonate kinase (MK)); (e) converting mevalonate 5-phosphate to mevalonate 5-pyrophosphate (e.g., by the action of phosphate mevalonate kinase (PMK)); and (f) converting mevalonate 5-pyrophosphate to isopentenyl pyrophosphate (e.g., by the action of mevalonate pyrophosphate decarboxylase (MPD)). The MVA pathway, as well as the genes and enzymes constituting the MVA pathway, are described in US 7,667,017, which is incorporated herein by reference in its entirety. In some embodiments, the host cell expresses or overexpresses one or more of acetyl-CoA thiolase, HMGS, HMGR, MK, PMK, and MPD, or modified variants thereof, resulting in increased production of IPP and DMAPP. In some embodiments, the triterpenoid backbone is at least partially generated by metabolic flux through the MVA pathway, and said host cell has at least one additional copy of one or more of acetyl-CoA thiolase, HMGS, HMGR, MK, PMK, MPD, or modified variants thereof.
[0091] In some embodiments, the host cell is a bacterial host cell engineered to increase the production of IPP and DMAPP from glucose, as described in US 10,480,015 and US 10,662,442, the contents of which are incorporated herein by reference in their entirety. For example, in some embodiments, the host cell overexpresses MEP pathway enzymes and has balanced expression to push / pull carbon flux to IPP and DMAP. In some embodiments, the host cell is engineered to increase the utilization or activity of Fe-S cluster proteins to support higher activity of IspG and IspH as Fe-S enzymes. In some embodiments, the host cell is engineered to overexpress IspG and IspH to provide more carbon flux to the 1-hydroxy-2-methyl-2-(E)-butenyl-4-bisphosphate (HMBPP) intermediate, but balanced expression prevents HMBPP from accumulating in amounts that reduce cell growth or viability, or in amounts that inhibit MEP pathway flux and / or terpene compound production. In some embodiments, host cells exhibit higher IspH activity relative to IspG. In some embodiments, host cells are engineered to downregulate the ubiquinone biosynthetic pathway, for example, by reducing the expression or activity of IspB using IPP and FPP substrates.
[0092] In other embodiments, microbial cells expressing the mogroside or mogroside biosynthesis pathway co-express the isoprenol utilization pathway, as described in US 2019 / 0367950, which is incorporated herein by reference in its entirety. Such cells can bypass the endogenous MEP pathway by generating IPP and DMAPP precursors from isoprenol and / or isoprenol substrates provided to the culture.
[0093] In some embodiments, the microbial cells contain one or more gene modifications that enhance electron supply and transfer for the production of squalene and mogroside via biosynthetic pathways. In some embodiments, the microbial cells express at least one recombinant electron carrier. Such electron carriers support the function of various enzymes disclosed herein, including enzymes of the MEP pathway. For example, the functional expression of plant cytochrome P450 in *Escherichia coli* has inherent limitations attributable to the bacterial platform (such as the absence of electron transfer mechanisms and cytochrome P450 reductase, and translational incompatibility of the membrane signaling module of P450 enzymes due to the lack of endoplasmic reticulum). Therefore, in some embodiments, the microbial cells overexpress at least one endogenous electron carrier, such as one or more oxidoreductases, including oxidoreductases that oxidize pyruvate and / or cause ferredoxin reduction. In various embodiments, the microbial strains comprise overexpression or complementation of one or more of flavin reductin (fldA), flavin reductase, ferredoxin (fdx), and ferredoxin reductase. In some implementations, the microbial cells express heterologous *Escherichia coli* fldA (SEQ ID NO: 43) or heterologous *Cv. fdx* (from *Heterochromis vinifera*). Allochromatium vinosum The strain expresses one or more non-natural fdx and / or fldA homologs (SEQ ID NO: 44), or an enzyme having at least 80% or 90% or 95% or 97% sequence identity with it. In some embodiments, the strain expresses one or more non-natural fdx and / or fldA homologs. For example, the fdx homolog may be selected from Hm. fdx1 (Moderately Thermophilic Sunbacterium) Heliobacterium modesticaldum )), Pa.fdx (Pseudomonas aeruginosa ( Pseudomonas aeruginosa Cv. fdx (Heterochilus acetoneus), Ca. fdx (Fusobacterium acetobutyricum) Clostridium acetobutylicum )), Cp.fdx (Clostridium pasteurellum ( Clostridium pasteurianum )), Ev2.fdx (Saphoshnikov's exosulfuron-lactonium spirospira) Ectothiorhodospira shaposhnikovii Pp1.fdx (Pseudomonas putida) and Pp2.fdx ( (Pseudomonas putida). In some embodiments, fldA homologs include those selected from Ec. fldA (Escherichia coli), Ac. fldA2 (Azotobacter chroococcus), and others. Azotobacter chroococcum )), Av.fldA2 (brown azotocin ( Azotobacter vinelandii One or more of Bs.fldA (Bacillus subtilis). See WO2018140778, which is incorporated herein by reference in its entirety.
[0094] In some embodiments, microbial cells are engineered to increase the metabolic supply of NADP by expressing or overexpressing alternative or exogenous NADPH biosynthetic pathways. In some embodiments, the microbial cells express bacterial... pntAB Or a homolog, or an ortholog, or a variant thereof. In some embodiments, the microbial cell expresses bacteria. pntAB Or a highly active variant of a homolog or its direct homolog.
[0095] In various embodiments, mogroside is produced via a heterologous enzyme pathway. The mogroside may be an intermediate of a downstream enzyme in the heterologous pathway, or in some embodiments it may be recovered from the culture. In some embodiments, mogroside may be recovered from host cells and / or from the culture medium.
[0096] In some embodiments, the heterologous enzyme pathway further comprises one or more uridine diphosphate-dependent glycosyltransferases (UGTs) to produce one or more mogrosides (or "monkhorinosides"). In some embodiments, the mogrosides may be pentasylated, hexasylated, or more (e.g., 7, 8, or 9 glycosylations). In other embodiments, the mogrosides have two, three, or four glucosylations. One or more mogrosides may be selected from Mog.II-E, Mog.III, Mog.III-A1, Mog.III-A2, Mog.III, Mog.IV, Mog.IV-A, sarcosinoside, isomog.V, Mog.V, or Mog.VI. In some embodiments, the host cell produces Mog.V or sarcosinoside.
[0097] In some embodiments, the host cell expresses a UGT enzyme that catalyzes primary glycosylation of mogroside at the C24 and / or C3 hydroxyl groups, and one or more branched glycosylations, such as β1,2 and / or β1,6 branched glycosylation at the primary C3 and C24 glucosyl groups. In some embodiments, at least one UGT enzyme comprises an amino acid sequence having at least 70% identity with an amino acid sequence selected from SEQ ID NO: 27-40 and 48-67. For example, in some embodiments, the UGT enzyme comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% identity with one of SEQ ID NO: 27-40 and 48-67. Thus, at least one UGT enzyme comprises an amino acid sequence modified with 1 to 20 amino acids relative to one of SEQ ID NO: 27-40 and 48-67, said amino acid modification being independently selected from amino acid substitution, deletion, and insertion.
[0098] For example, in some embodiments, the microbial cells express at least four UGT enzymes, leading to glucosylation of mogroside at the C3 hydroxyl group and the C24 hydroxyl group, as well as further 1,6-glucosylation at the C3 glucosyl group, and further 1,6-glucosylation and further 1,2-glucosylation at the C24 glucosyl group. The product of such glucosylation reactions is Mog.V.
[0099] In some embodiments, the microbial host cell has one or more gene modifications that increase the production of UDP-glucose (a cofactor used by the UGT enzyme). These gene modifications may include one or more, or two or more (or all) of the following: ΔgalE, ΔgalT, ΔgalK, ΔgalM, ΔushA, Δagp, Δpgm, replication of Escherichia coli galU, expression of Bacillus subtilis UGPA, and expression of Bifidobacterium adolescentis SPL.
[0100] In various embodiments, the reaction is carried out in microbial cells, and the UGT enzyme is recombinantly expressed in the cells. In some embodiments, mogroside is produced in cells via a heterologous mogroside synthesis pathway, as described herein. In other embodiments, mogroside or mogroside glycosides (e.g., mogroside extract) are supplied to cells for glycosylation. In still other embodiments, the reaction is carried out in vitro using purified UGT enzyme, partially purified UGT enzyme, or recombinant cell lysate.
[0101] Bacterial host cells are cultured to produce triterpenoid products (e.g., mogrosides). In some embodiments, the production phase employs carbon substrates such as C1, C2, C3, C4, C5, and / or C6 carbon substrates. In exemplary embodiments, the carbon source is glucose, sucrose, fructose, xylose, and / or glycerol. Culture conditions are typically selected from aerobic, microaerobic, and anaerobic conditions.
[0102] In various embodiments, the bacterial host cells can be cultured at temperatures between 22°C and 37°C. While commercial biosynthesis in bacteria such as *Escherichia coli* may be limited by the temperature at which overexpressed and / or exogenous enzymes (e.g., plant-derived enzymes) stabilize, recombinant enzymes can be engineered to allow cultures to be maintained at higher temperatures, resulting in higher yields and higher overall productivity. In some embodiments, cultures are carried out at approximately 22°C or higher, approximately 23°C or higher, approximately 24°C or higher, approximately 25°C or higher, approximately 26°C or higher, approximately 27°C or higher, approximately 28°C or higher, approximately 29°C or higher, approximately 30°C or higher, approximately 31°C or higher, approximately 32°C or higher, approximately 33°C or higher, approximately 34°C or higher, approximately 35°C or higher, approximately 36°C or higher, or approximately 37°C.
[0103] In some embodiments, the bacterial host cells are also suitable for commercial production on a commercial scale. In some embodiments, the culture size is at least about 100 L, at least about 200 L, at least about 500 L, at least about 1,000 L, or at least about 10,000 L, or at least about 100,000 L, or at least about 500,000 L, or at least about 600,000 L. In one embodiment, the culture can be carried out in batch, continuous, or semi-continuous form.
[0104] In various embodiments, the method further includes recovering the product from the cell culture or from the cell lysate. In some embodiments, the culture produces at least about 100 mg / L, or at least about 200 mg / L, or at least about 500 mg / L, or at least about 1 g / L, or at least about 2 g / L, or at least about 5 g / L, or at least about 10 g / L, or at least about 20 g / L, or at least about 30 g / L, or at least about 40 g / L of terpenoid or terpene glycoside product.
[0105] In some embodiments, the production of indole (including isopentenylated indole) is used as a substitute marker for terpene production and / or to control the accumulation of indole in the culture to increase production. For example, in various embodiments, the accumulation of indole in the culture is controlled to be below about 100 mg / L, or below about 75 mg / L, or below about 50 mg / L, or below about 25 mg / L, or below about 10 mg / L. This can be controlled by balancing protein expression and activity using a multivariate modular approach as described in U.S. Patent No. 8,927,241 (which is incorporated herein by reference in its entirety), and / or by chemical means.
[0106] Other markers for efficient production of terpenes and terpenoids include the accumulation of DOX or ME in the culture medium. Typically, bacterial strains can be engineered to accumulate fewer of these chemicals in the culture, less than about 5 g / L, or less than about 4 g / L, or less than about 3 g / L, or less than about 2 g / L, or less than about 1 g / L, or less than about 500 mg / L, or less than about 100 mg / L.
[0107] Optimization of terpene or terpene production by manipulating MEP pathway genes and upstream and downstream pathways is not expected to be a simple linear or additive process. Instead, optimization is expected through combinatorial analysis, balancing the components of the MEP pathway and its upstream and downstream pathways. The accumulation of indoles (including isopentenylated indoles) and MEP metabolites (e.g., DOX, ME, MEcPP, and / or farnesol) in cultures can serve as alternative markers to guide this process.
[0108] For example, in some embodiments, the bacterial strain has at least one additional copy (on a plasmid or integrated into the genome) of dxs and idi, represented as operons / modules; or dxs, ispD, ispF, and idi, represented as operons or modules, and has other MEP pathway complementarities described herein to improve MEP carbon. For example, the bacterial strain may have another copy of dxr and ispG and / or ispH, optionally another copy of ispE and / or idi, and the expression of these genes is modulated to increase MEP carbon and / or improve terpene or terpene titers. In various embodiments, the bacterial strain has at least another copy of dxr, ispE, ispG, and ispH, optionally another copy of idi, and the expression of these genes is modulated to increase MEP carbon and / or improve terpene or terpene titers.
[0109] Manipulation of gene and / or protein (including gene modules) expression can be achieved through various methods. For example, gene or operon expression can be regulated by selecting promoters of varying strengths (e.g., strong, medium, or weak), such as inducible or constitutive promoters. Several non-limiting examples of promoters of varying strengths include Trc, T5, and T7. Additionally, gene or operon expression can be regulated by manipulating the copy number of the gene or operon in the cell. In some embodiments, gene or operon expression can be regulated by manipulating the order of genes within a module, where genes transcribed first are typically expressed at higher levels. In some embodiments, gene or operon expression is regulated by integrating one or more genes or operons into a chromosome.
[0110] Optimization of protein expression can also be achieved by selecting appropriate promoters and ribosome binding sites. In some implementations, this may include selecting plasmids with high copy numbers, or plasmids with single, low, or medium copy numbers. Gene expression can also be regulated by targeting transcription termination steps by introducing or eliminating structures such as stem-loops.
[0111] Expression vectors containing all the essential elements for expression are commercially available and are known to those skilled in the art. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., ColdSpring Harbor Laboratory Press, 1989. Cellular genetic engineering is achieved by introducing heterologous DNA into the cell. Heterologous DNA is placed under the operative control of transcriptional elements to allow its expression in the host cell.
[0112] In some implementations, endogenous genes are edited, which is the opposite of gene complementation. Editing can modify endogenous promoters, ribosome-binding sequences, or other expression control sequences, and / or, in some implementations, modify trans- and / or cis-acting factors in gene regulation. Genome editing can be performed using CRISPR / Cas genome editing technology, or similar technologies such as zinc finger nucleases and TALEN. In some implementations, endogenous genes are replaced by homologous recombination.
[0113] In some implementations, gene overexpression is achieved, at least in part, by controlling the gene copy number. While gene copy number can be conveniently controlled using plasmids with different copy numbers, gene replication and chromosome integration can also be employed. For example, a method for stable tandem gene replication is described in US 2011 / 0236927, which is incorporated herein by reference in its entirety.
[0114] Mogrosides can be recovered from microbial cultures. For example, mogrosides can be recovered from microbial cells, or in some embodiments, they are primarily present in extracellular media from which they can be recovered or isolated. In some embodiments, mogrosides are recovered essentially as described in WO 2022 / 115527, which is incorporated herein by reference in its entirety. For example, recovery may include adjusting the pH of the culture or medium to below about pH 5 or above about pH 10, raising the temperature to at least about 50°C, optionally adding one or more glycoside solubilizers; followed by removal of biomass. By adjusting the properties of the culture or medium prior to biomass removal, products with desired qualities can be produced, including: high purity of the glycoside product, appealing color, easy solubility, odorlessness, and / or high recovery yield. For example, initial pH and temperature adjustments of the culture can alter the fluid properties of the fermentation broth and improve the efficiency of disc separators used for biomass removal. Furthermore, the solubility of the glycoside product, and therefore the yield, can be significantly improved by pH and / or temperature adjustments and / or the addition of glycoside solubilizers.
[0115] In some embodiments, recovery includes the addition of one or more glycoside solubility enhancers. Exemplary solubility enhancers include chemical agents (including organic acids and polymers) having alcohol functional groups and / or polar agents (including those having ether, ester, aldehyde, and ketone functional groups), and include, but are not limited to, glycerol, 1,3-propanediol, polyvinyl alcohol, polyethylene glycol, etc. Other exemplary solubilizers include organic acids, sugars, and polysaccharides. Other solubility enhancers are described in US 2020 / 0268026, which is incorporated herein by reference in its entirety. Improved glycoside solubility facilitates the removal of biomass without significant loss of product. Typically, solubility enhancers can be added to the harvested culture material in the range of about 0.1 wt% to about 2 wt%, such as from about 0.1 wt% to about 1 wt% (e.g., about 0.5 wt%).
[0116] Subsequently, the biomass is removed by centrifugation to prepare a clarified fermentation broth. An exemplary process for removing biomass employs a disc centrifuge to separate the liquid and solid phases. The clarified fermentation broth (liquid phase) is recovered for further processing to purify the glycoside products. The separated biomass (solid phase) can be reprocessed for further glycoside product recovery, or alternatively disposed of as waste.
[0117] In some embodiments, the glycosides are crystallized from the clarified fermentation broth. In some embodiments, the method includes one, two, or three crystallization steps. In some embodiments, the glycoside product is purified from the clarified fermentation broth using one or more processes selected from filtration, ion exchange, activated carbon, bentonite, affinity chromatography, and digestion, said processes optionally performed before crystallization and / or recrystallization. These processes can be selected to achieve high product purity, appealing color, easy solubility, odorless properties, and high recovery yield. In some embodiments, the method employs affinity chromatography, such as using one or more of styrene-divinylbenzene adsorption resins, strongly acidic cation exchange resins, weakly acidic cation exchange resins, strongly basic anion exchange resins, weakly basic anion exchange resins, and hydrophobically interacting resins. In other embodiments, the method employs simulated moving bed chromatography as described in U.S. Patent 10,213,707, which is incorporated herein by reference in its entirety. In other embodiments, the recovery method is non-chromatographic (i.e., without chromatographic steps), providing a significant cost advantage. For example, a recovery method after biomass removal can consist essentially of filtration and crystallization steps, or filtration and crystallization steps alone. In some embodiments, the recovery method uses an organic solvent (e.g., ethanol), but in other embodiments, the method uses an aqueous solvent entirely. In some embodiments, two crystallization steps are employed.
[0118] In some embodiments, the recovery method will include one or more tangential flow filtration (TFF) steps. For example, a TFF with a filter having a pore size of about 5 kD can remove endotoxins, large proteins, and other cellular debris, while also improving the solubility of the final powder product. In some embodiments, the glycoside product is purified by tangential flow filtration, optionally having a membrane pore size of about 5 kD, prior to initial crystallization. A TFF with a filter having a pore size of about 0.5 kD can also be used downstream to remove small molecule impurities and salts, and / or concentrate the mother liquor for recrystallization. In some embodiments, a TFF with a pore size of about 0.5 kD is used prior to recrystallization.
[0119] In each case, the crystallization step may include one or more stages of static crystallization, stirred crystallization, and evaporative crystallization. For example, the crystallization step may include a static stage followed by a stirred stage to control the crystal morphology. The static stage may grow large crystals with highly crystalline domains. The crystallization process may include seeding crystals, or in some embodiments, seeding crystals may not be involved (i.e., crystals form spontaneously). In various embodiments, the crystallization solvent comprises water or water / ethanol. Exemplary crystallization solvents include water, optionally containing about 5% to about 50% ethanol by volume, or about 25% to about 50% ethanol by volume (e.g., about 30% to about 40% ethanol by volume). In some embodiments, after seeding crystals during the static stage, the stirred stage rapidly grows the crystals and increases the degree of amorphous domains. Using this method, the resulting crystals can have better final solubility and higher purity of the glycoside product, and can be more easily recovered and washed.
[0120] In various embodiments, prior to recrystallization, the glycoside product is redissolved in a solvent (such as, but not limited to, water and / or ethanol), which may involve one or more of the following: lowering the pH of the solvent and the glycoside product solution or suspension to below about pH 5 or raising the pH of the solution or suspension to above about pH 9, heating to at least about 50°C, and adding one or more glycoside solubilizers. Target values for pH, temperature, and glycoside solubilizer concentration can alternatively be used for biomass removal. For example, the pH of the glycoside product solution or suspension may be adjusted to a range of about 2 to about 5. In other embodiments, the pH is adjusted to an alkaline pH range, such as in the range of about 9 to about 12, or in the range of about 9.5 or about 10 to about 12. According to known methods, the pH can be adjusted by adding or titrating an organic or inorganic acid or hydroxide ion. In other embodiments, recrystallization is carried out at a pH of about 4 to about 12. Alternatively or additionally, the temperature of the solution or suspension is adjusted to a temperature between about 50°C and about 90°C, such as about 50°C to about 80°C. Exemplary recrystallization solvents include water, optionally containing about 5% to about 50% ethanol by volume, or about 25% to about 50% ethanol by volume (e.g., about 30% to about 40% ethanol by volume). Alternatively or additionally, as described, a solubility enhancer may be added to the solution / suspension in the range of about 0.1 wt% to about 2 wt%, such as in the range of about 0.1 wt% to about 1 wt% (e.g., about 0.5 wt%). Exemplary solubility enhancers include glycerol.
[0121] In some embodiments, after crystallization, the resulting crystals are separated, for example, using a basket centrifuge or belt filter, to separate the glycoside wet cake (e.g., mogroside wet cake). The washing in the basket centrifuge step can be done with water, or alternatively with other rinsing solutions (e.g., cooled water / ethanol). In some embodiments, the cake is dissolved and recrystallized. The wet cake from the recrystallization can then be dried, optionally using a belt dryer, paddle dryer, or spray dryer. The dried cake can be ground and packaged.
[0122] Prior to recrystallization, the glycoside solution can be filtered to remove impurities. In some embodiments, the filter may be an approximately 0.2-micron filter. Alternatively, other pore sizes may be used, such as approximately 0.45-micron filters and approximately 1.2-micron filters. Depending on the embodiment, the filter material can be selected to further remove impurities, such as through adsorption. For example, hydrophilic materials such as polyethersulfone (PES) have significant advantages over more hydrophobic materials such as polypropylene. Other exemplary hydrophilic filter materials include nylon, cellulose acetate, cellulose nitrate, and conventionally hydrophobic materials that have been functionalized to produce hydrophilic materials (such as PTFE or PVDF coated with fluoroalkyl-terminated polyethylene glycol).
[0123] The similarity of nucleotide and amino acid sequences, i.e., the percentage of sequence identity, can be determined by sequence alignment. Such alignments can be performed using several algorithms known in the art, such as the mathematical algorithm of Karlin and Altschul (Karlin & Altschul (1993) Proc. Natl. Acad. Sci. USA 90: 5873-5877), hmmalign (HMMER software package, http: / / hmmer.wustl.edu / ), or the CLUSTAL algorithm (Thompson, JD, Higgins, DG & Gibson, TJ (1994) Nucleic Acids Res. 22, 4673-80). Sequence identity (sequence matching) levels can be calculated using, for example, BLAST, BLAT, or BlastZ (or BlastX). Similar algorithms are incorporated into the BLASTN and BLASTP programs of Altschul et al. (1990) J. Mol. Biol. 215: 403-410. BLAST polynucleotide search can be performed using the BLASTN program with a score of 100 and a word length of 12.
[0124] BLAST protein searches can be performed using the BLASTP program with a score of 50 and a word length of 3. To obtain gap alignments for comparison, GappedBLAST can be used as described by Altschul et al. (1997) Nucleic Acids Res. 25: 3389-3402. When using BLAST and GappedBLAST programs, the default parameters for each program should be used. Sequence matching analysis can be supplemented by established homology mapping techniques such as Shuffle-LAGAN (Brudno M., Bioinformatics 2003b, 19 Supplement 1:154-162) or Markov random fields.
[0125] "Conservative substitution" can be based, for example, on the similarity of the polarity, charge, size, solubility, hydrophobicity, hydrophilicity, and / or amphiphilic properties of the amino acid residues involved. The 20 natural amino acids can be divided into the following six standard amino acid groups: (1) Hydrophobicity: Met, Ala, Val, Leu, Ile; (2) Neutral hydrophilicity: Cys, Ser, Thr; Asn, Gln; (3) Acidic: Asp, Glu; (4) Alkaline: His, Lys, Arg; (5) Residues affecting chain orientation: Gly, Pro; and (6) Aromatics: Trp, Tyr, Phe.
[0126] As used herein, “conservative substitution” is defined as the exchange of an amino acid for another amino acid listed in the same group of the six standard amino acid groups shown above. For example, the exchange of Asp with Glu in such a modified polypeptide retains a negative charge. Additionally, glycine and proline can substitute for each other based on their ability to disrupt the α-helix. Some preferred conserved substitutions within the six groups above are exchanges within the following subgroups: (i) Ala, Val, Leu, and Ile; (ii) Ser and Thr; (iii) Asn and Gln; (iv) Lys and Arg; (v) Tyr and Phe.
[0127] As used herein, “non-conservative substitution” is defined as the exchange of an amino acid for another amino acid listed in different groups of the six standard amino acid groups (1) to (6) shown above.
[0128] As described herein, enzyme modifications may include conserved and / or non-conserved mutations. In some embodiments, an alanine residue is substituted or inserted at position 2 to increase stability.
[0129] In some implementations, “rational design” involves constructing specific mutations in an enzyme. Rational design refers to incorporating knowledge about the enzyme or related enzymes, such as their reaction thermodynamics and kinetics, their three-dimensional structure, their active sites, their substrates, and / or enzyme-substrate interactions, into the design of a specific mutation. Based on rational design methods, mutations can be generated in the enzyme, and then screening can be performed for increased yields of terpenes or terpenoids relative to control levels. In some implementations, mutations can be rationally designed based on homology modeling. As used herein, “homology modeling” refers to the process of constructing an atomic resolution model of a protein based on its amino acid sequence and the three-dimensional structures of related homologous proteins.
[0130] In other aspects, the present invention provides a method for preparing a product comprising mogroside. The method includes generating mogroside according to the present disclosure and incorporating the mogroside into a product. In some embodiments, the mogroside is symmonoside, Mog.V, Mog.VI, or Isomog.V. In some embodiments, the product is a sweetener composition, flavoring composition, food, beverage, chewing gum, texture agent, pharmaceutical composition, tobacco product, nutritional composition, or oral hygiene composition.
[0131] The product may be a sweetener composition comprising a blend of artificial and / or natural sweeteners. For example, the composition may also contain one or more of steviol glycosides, aspartame, and neotame. Exemplary steviol glycosides include one or more of RebM, RebB, RebD, RebA, RebE, and RebI.
[0132] Non-limiting examples of flavoring agents that can be used in combination with the product include lime, lemon, orange, fruit, banana, grape, pear, pineapple, mango, bitter almond, kola nut, cinnamon, sugar, marshmallow, and vanilla flavorings. Other non-limiting examples of food ingredients include flavoring agents, acidifiers and amino acids, coloring agents, extenders, modified starches, gums, conditioning agents, preservatives, antioxidants, emulsifiers, stabilizers, thickeners, and gelling agents.
[0133] The mogrosides obtained according to the present invention can be incorporated as high-intensity natural sweeteners into foods, beverages, pharmaceutical compositions, cosmetics, chewing gum, tabletop products, cereals, dairy products, toothpaste, and other oral compositions.
[0134] The mogrosides obtained according to this invention can be used in combination with various physiologically active substances or functional ingredients. Functional ingredients are generally classified into the following categories: carotenoids, dietary fiber, fatty acids, saponins, antioxidants, nutritional foods, flavonoids, isothiocyanates, phenols, phytosterols and sterols (phytosterols and phytosterols); polyols; prebiotics, probiotics; phytoestrogens; soy protein; sulfides / thiols; amino acids; proteins; vitamins; and minerals. Functional ingredients can also be classified based on their health benefits, such as cardiovascular, cholesterol-lowering, and anti-inflammatory benefits.
[0135] The mogrosides obtained according to the present invention can be used as high-intensity sweeteners to produce zero-calorie, low-calorie, or diabetic beverages and foods with improved flavor characteristics. They can also be used in beverages, foods, pharmaceuticals, and other products where sugar cannot be used. Furthermore, highly purified target mogrosides, particularly Mog.V, Mog.VI, or Isomog.V, can be used not only as sweeteners in beverages, foods, and other products intended for human consumption, but also in animal feed and forage with improved characteristics.
[0136] Examples of products in which mogrosides can be used as sweeteners include, but are not limited to: alcoholic beverages such as vodka, wine, beer, spirits, and sake; natural fruit juices; refreshing beverages; carbonated soft drinks; meal replacement drinks; zero-calorie beverages; low-calorie beverages and foods; yogurt drinks; instant fruit juices; instant coffee; powdered instant beverages; canned products; syrups; fermented soybean pastes; soy sauce; vinegar; sauces; mayonnaise; ketchup; curry; soups; instant broth; powdered soy sauce; powdered vinegar; and various biscuits. Dried; rice cakes; crispy biscuits; bread; chocolate; caramel; candy; chewing gum; jelly; pudding; pickled fruits and vegetables; whipped cream; jam; orange marmalade; flower jam; milk powder; ice cream; fruit juice sorbet; bottled vegetables and fruits; canned and boiled beans; meat and sweet sauce cooked foods; agricultural products and vegetable foods; seafood; ham; sausage; fish ham; fish sausage; surimi; fried fish products; dried seafood; frozen foods; pickled seaweed; cured meat; tobacco; medicinal products; and many other products.
[0137] In the manufacturing process of products such as food, beverages, pharmaceuticals, cosmetics, tabletop products and chewing gum, conventional methods such as mixing, kneading, dissolving, pickling, permeation, filtration, spraying, atomizing, filling and other methods can be used.
[0138] Unless otherwise expressly stated, as used in this specification and the appended claims, the singular forms “a” and “the” include plural indicators. For example, reference to “a cell” includes a combination of two or more cells, etc.
[0139] As used in this article, the term “about” when referring to numbers is generally understood to include numbers that fall within 10% of either direction (greater or less).
[0140] Any aspect or implementation disclosed herein may be combined with any aspect or implementation disclosed herein.
[0141] Example Figure 1A and Figure 1BThe in vivo biosynthetic pathways to mogrool and MogV are shown, including the condensation of farnesyl pyrophosphate to squalene catalyzed by squalene synthase (SQS); the epoxidation of squalene to 2,3-squalene oxide and 2,3;22,23-squalene dioxide catalyzed by squalene epoxidase (SQE); and the cyclization of 2,3;22,23-squalene dioxide to 24,25-epoxy Cucurbitadienol; the synthesis of 24,25-dihydroxycucurbitadienol from 24,25-epoxycucurbitadienol catalyzed by epoxide hydrolase (EPH); the hydroxylation of 24,25-dihydroxycucurbitadienol to mogroside catalyzed by cytochrome P450 and cytochrome P450 reductase active at the C11 position of 24,25-dihydroxycucurbitadienol; and the synthesis of mogroside from uridine diphosphate-dependent glycosyltransferase (UGT, see...) Figure 1A and Figure 1B Catalyzes the C3, C24, 1,6-glycosylation, and 1,2-glycosylation of mogroside.
[0142] like Figure 1A As shown, in addition to catalyzing the desired conversion, enzymes such as CDS, EPH, and UGT can also divert intra-pathway products to undesirable extra-pathway products, such as cucurbitacinol, hydrolyzed glycosides of 2,3;22,23-squalene dioxide, and 24,25-dihydroxycucurbitacinol. Bacterial cells expressing unengineered SQS, SQE, CDS, and EPH typically produce approximately 90% or more of the extra-pathway products, along with relatively low amounts of mogroside. This paper discloses engineered enzymes that exhibit higher substrate specificity and / or activity to divert the production of intra-pathway products.
[0143] Example 1: Engineering of squalene cyclooxygenase (SQE) to increase the yield of 2,3;22,23-squalene dioxide Downstream enzymes were screened for their ability to convert squalene into in-pathway products using *E. coli* strains that produce high levels of the MEP pathway products IPP and DMAPP (see US 2018 / 0245103 and US 2018 / 0216137, which are incorporated herein by reference) and express ScFPPS and squalene synthase (e.g., SEQ ID NO:13). See also Figure 1A and 1B .
[0144] SQE_1 (SEQ ID NO: 2) is a type of *Mortimes methylmonas* (…). Methylomonas lentaThe SQE (SEQ ID NO:1) derivative has mutations H36R, F164A, M284L, V381L, and F396Y relative to SEQ ID NO:1. This derivative is disclosed in WO 2021 / 126960, the entire contents of which are incorporated herein by reference. Further improvements to SQE_1 are made for the production of 2,3;22,23-squalene dioxide.
[0145] In the first round, *E. coli* strains that produced squalene and expressed high levels of the SQE_1 mutant derivative were screened. For screening, fermentation was performed in 96-well plates at 37 °C for 72 h. The yields of 2,3-squalene oxide and 2,3;22,23-squalene dioxide were quantified by gas chromatography coupled with flame ionization detection (GC-FID). The effect of amino acid substitutions was evaluated by comparing the titers of 2,3-squalene oxide and 2,3;22,23-squalene dioxide produced during fermentation by bacterial strains expressing a given SQE derivative with those expressing SQE_1. Table 1 shows fifty groups of amino acid substitutions that showed beneficial changes in the titers of 2,3-squalene oxide and / or 2,3;22,23-squalene dioxide compared to SQE_1: Table 1: SQE Engineering Round 1
[0146] Based on the above results, a combined mutation was generated. SQE_2 (SEQ ID NO: 3) with mutations S116E, N134G, I145H, A169R, and A400V was further investigated and compared with SQE_1. Squalene-producing bacterial strains expressing either SQE_1 or SQE_2 were fermented in 96-well plates at 37°C for 72 hours. The yield of 2,3;22,23-squalene dioxide was quantified by GC-FID and normalized to the level of 2,3;22,23-squalene dioxide produced by the SQE_1-expressing bacterial strain. Figure 2A As shown, compared with bacterial strains expressing SQE_1, bacterial strains expressing SQE_2 produced approximately 5-fold higher levels of 2,3;22,23-squalene dioxide.
[0147] In the second round, *E. coli* strains expressing further SQE mutant derivatives were screened. Some of these mutant derivatives included mutant clusters (see, for example, several rows in Table 2 showing various mutations). For screening, bacterial strains producing squalene and expressing various SQE_2 derivatives were fermented in 96-well plates at 37°C for 72 h. The yields of 2,3-squalene oxide and 2,3;22,23-squalene dioxide were quantified by GC-FID. The effect of amino acid substitution was evaluated by comparing the titers of 2,3-squalene oxide and 2,3;22,23-squalene dioxide produced during fermentation by bacterial strains expressing a given SQE derivative with those expressing SQE_2. Table 2 shows fourteen mutants that exhibited beneficial changes in the titers of 2,3-squalene oxide and / or 2,3;22,23-squalene dioxide compared to SQE_2: Table 2: SQE Engineering Round 2
[0148] The SQE_2 sequence was re-encoded without any further alteration to the amino acid sequence. The resulting sequence was designated SQE_3. Squalene-producing bacterial strains expressing either SQE_2 or SQE_3 were fermented in 96-well plates at 37°C for 72 hours. The yield of 2,3;22,23-squalene dioxide was quantified by GC-FID and normalized to the level of 2,3;22,23-squalene dioxide produced by the SQE_2-expressing bacterial strains. Figure 2B As shown, bacterial strains expressing SQE_3 produced more than 2-fold higher levels of 2,3;22,23-squalene dioxide compared to bacterial strains expressing SQE_2.
[0149] In the third round, *E. coli* strains expressing large amounts of SQE mutant derivatives were screened. Some of these mutant derivatives included mutant clusters (see, for example, several rows in Table 3 for various mutations). For screening, fermentation was performed in 96-well plates at 37°C for 72 h. The yields of 2,3-squalene oxide and 2,3;22,23-squalene dioxide were quantified by GC-FID. The effect of amino acid substitution was evaluated by comparing the titers of 2,3-squalene oxide and 2,3;22,23-squalene dioxide produced during fermentation by bacterial strains expressing a given SQE derivative with those expressing SQE_3. Table 3 shows eight mutants that exhibited beneficial changes in the titers of 2,3-squalene oxide and / or 2,3;22,23-squalene dioxide compared to SQE_3: Table 3: SQE Engineering Round 3
[0150] Based on these results, SQE_4 (SEQ ID NO:4), which contains the A164N mutation in addition to the mutation present in SQE_3, was further compared with SQE_3. Bacterial strains expressing either SQE_3 or SQE_4 were fermented in 96-well plates at 37°C for 72 hours. The yield of 2,3;22,23-squalene dioxide was quantified by GC-FID and normalized to the level of 2,3;22,23-squalene dioxide produced by the bacterial strain expressing SQE_3. Figure 2C As shown, bacterial strains expressing SQE_4 produced increased levels of 2,3;22,23-squalene dioxide compared to bacterial strains expressing SQE_3.
[0151] Example 2: Engineering of McSQE to increase the yield of 3,22,23-squalene dioxide In the first round, strains that produce squalene and express capsular methylcocci were screened. Methylococcus capsulatus E. coli strains expressing mutant derivatives of SQE (McSQE_0, SEQ ID NO: 77) were used for screening. Fermentation was performed in 96-well plates at 37°C for 72 hours. The yield of 2,3;22,23-squalene dioxide was quantified by up-conversion liquid chromatography with diode array detection (UPLC-DAD). The effect of amino acid substitution was evaluated by comparing the titers of 2,3;22,23-squalene dioxide produced during fermentation by bacterial strains expressing a given SQE derivative with those expressing McSQE_0. Table 4 shows 10 mutants that exhibited beneficial changes in the titer of 2,3;22,23-squalene dioxide compared to McSQE_0: Table 4: McSQE Engineering Phase 1
[0152] Further studies were conducted comparing McSQE_1 (SEQ ID NO: 78), which has the S366R mutation relative to SEQ ID NO: 77, with McSQE_0. Squalene-producing bacterial strains expressing either McSQE_0 or McSQE_1 were fermented in 96-well plates at 37°C for 72 hours. The yield of 2,3;22,23-squalene dioxide was quantified by UPLC-DAD and normalized to the level of 2,3;22,23-squalene dioxide produced by bacterial strains expressing McSQE_0. Figure 2D As shown, compared with bacterial strains expressing McSQE_0, bacterial strains expressing McSQE_1 produced approximately 4-fold higher levels of 2,3;22,23-squalene dioxide.
[0153] In the second round of screening, *E. coli* strains that produced squalene and expressed a McSQE_1 mutant derivative were selected. McSQE_2 (SEQ ID NO: 79), which, in addition to the mutations in McSQE_1, also possessed the mutations D41H and M315L relative to SEQ ID NO: 77, was further compared with McSQE_1. Squalene-producing bacterial strains expressing either McSQE_1 or McSQE_2 were fermented in 96-well plates at 37°C for 72 hours. The yield of 2,3;22,23-squalene dioxide was quantified by UPLC-DAD and normalized to the level of 2,3;22,23-squalene dioxide produced by bacterial strains expressing McSQE_0. Figure 2E As shown, compared with bacterial strains expressing McSQE_1, bacterial strains expressing McSQE_1 produced approximately 1.4 times more 2,3;22,23-squalene dioxide.
[0154] In the third round, *E. coli* strains that produced squalene and expressed mutant derivatives of McSQE_2 were screened. For screening, fermentation was performed in 96-well plates at 37°C for 72 h. The yield of 2,3;22,23-squalene dioxide was quantified by UPLC-DAD. The effect of amino acid substitution was evaluated by comparing the titers of 2,3;22,23-squalene dioxide produced during fermentation by bacterial strains expressing a given SQE derivative with those expressing McSQE_2. Table 5 shows the mutants that showed beneficial changes in the titer of 2,3;22,23-squalene dioxide compared to McSQE_2: Table 5: McSQE Engineering Phase 1
[0155] These results demonstrate that, compared to McSQE_2, the aforementioned mutant derivatives further increased the titer of 2,3;22,23-squalene dioxide.
[0156] Example 3: Engineering of epoxide hydrolase (EPH) to increase the yield of 24,25-dihydroxycucurbitadienol Expressions were filtered Monk fruit An Escherichia coli strain with an amino acid-substituted derivative of EPH3 (EPH_0; SEQ ID NO: 5). Monk fruitEPH3 is disclosed in WO 2019169027, which is incorporated herein by reference in its entirety. Screening was performed using bacterial strains that produced 24,25-epoxycucurbitadienol and expressed EPH_0 variants. Briefly, fermentation was carried out in 96-well plates at 37°C for 72 hours. The yield of 24,25-dihydroxy-cucurbitadienol was quantified by gas chromatography coupled with flame ionization detection (GC-FID). The effect of amino acid substitutions was evaluated by comparing the titers of 24,25-dihydroxy-cucurbitadienol produced during fermentation by bacterial strains expressing a given EPH_0 derivative with those expressing EPH_0. Table 6 shows twenty-two amino acid substitutions that showed beneficial changes in the titer of 24,25-dihydroxy-cucurbitadienol compared to EPH_0: Table 6: EPH Engineering
[0157] Based on these results, EPH_1 variants with the following mutations relative to EPH_0 were constructed: S68A, Q128A, G141S, T144Q, A145V, E163A, E191G, L262M, C295G, and N299Q. EPH_1 (SEQ ID NO: 6) was further investigated. Bacterial strains producing 24,25-epoxycucurbitadienol and expressing either EPH_0 or EPH_1 were fermented in 96-well plates at 37°C for 72 hours. The yield of 24,25-dihydroxycucurbitadienol was quantified by GC-FID. Figure 3 As shown, compared with bacterial strains expressing EPH_0, bacterial strains expressing EPH_1 produced approximately 3.5 times more 24,25-dihydroxy-cucurbitadienol.
[0158] Example 4: Engineering of 24,25-epoxycucurbitacin synthase (ECDS) to improve the efficiency of 24,25-epoxycucurbitacin synthase. Enol production Escherichia coli strains expressing mutant derivatives of the drug cucurbitacin synthase 2 (CcCDS_0 or ECDS_0; SEQ ID NO: 7) were screened. Cucurbitacin synthase 2 is disclosed in WO 2021 / 126960, which is incorporated herein by reference in its entirety. Briefly, bacterial strains producing 2,3;22,23-squalene dioxide and expressing ECDS_0 or its derivatives were fermented in 96-well plates at 37°C for 72 h. The yield of 24,25-epoxy-cucurbitacinol was quantified by GC-FID. The effect of amino acid substitution was evaluated by comparing the titers of cucurbitacinol and 24,25-epoxy-cucurbitacinol produced during fermentation by bacterial strains expressing a given ECDS_0 derivative with those expressing ECDS_0. Table 7 shows fourteen amino acid substitutions that exhibited beneficial changes in the titer of 24,25-epoxy-cucurbitadienol compared to ECDS_0: Table 7: ECDS Engineering Phase 1
[0159] Based on these results, ECDS_1 (SEQ ID NO:8) carrying S24N, C35D, and A556S substitutions was further developed and compared with ECDS_0. Fermentation was carried out in 96-well plates at 37°C for 72 hours using bacterial strains that produce 2,3;22,23-squalene dioxide and express ECDS_0 or ECDS_1. The yield of 2,25-epoxycucurbitadienol was quantified by GC-FID. Figure 4A As shown, compared with bacterial strains expressing ECDS_0, bacterial strains expressing ECDS_1 produced approximately 6-fold higher levels of 24,25-epoxycucurbitadienol.
[0160] As discussed above, using bacterial strains that produce 2,3;22,23-squalene dioxide and express mutant derivatives of ECDS_1, *E. coli* strains expressing further mutant derivatives of ECDS_1 were constructed and screened. Based on these results, ECDS_2 (SEQ ID NO: 9), which carries I490V and I553M substitutions in addition to those present in ECDS_1, was further compared with ECDS_1. Bacterial strains that produce 2,3;22,23-squalene dioxide and express ECDS_1 or ECDS_2 were fermented in 96-well plates at 37°C for 72 hours. The yields of cucurbitadienol and 2,4,25-epoxy-cucurbitadienol were quantified by GC-FID. Figure 4BAs shown, compared with bacterial strains expressing ECDS_1, bacterial strains expressing ECDS_2 produced approximately 3.5-fold higher levels of 24,25-epoxy-cucurbitadienol, while the increase in cucurbitadienol levels was less than approximately 2-fold. These results particularly demonstrate that ECDS_2 exhibits increased activity or specificity for 2,3;22,23-squalene dioxide substrates compared to 2,3-squalene oxide substrates.
[0161] Using bacterial strains that produce 2,3;22,23-squalene dioxide and express ECDS_2 mutant derivatives, *E. coli* strains expressing further mutants of ECDS_2 were screened. Based on these results, ECDS_3 (SEQ ID NO: 10), carrying D50E, N121H, and Q401A substitutions in addition to those present in ECDS_2, was constructed and further studied in comparison with ECDS_2. Fermentation was performed in 96-well plates at 37°C for 72 hours using bacterial strains that produce 2,3;22,23-squalene dioxide and express ECDS_2 or ECDS_3. The yield of 24,25-epoxy-cucurbitadienol was quantified by GC-FID. Figure 4C As shown, compared with bacterial strains expressing ECDS_2, bacterial strains expressing ECDS_3 produced more than 2 times higher levels of 24,25-epoxy-cucurbitadienol.
[0162] Mutant derivatives of ECDS_3 that produced 2,3;22,23-squalene dioxide and expressed ECDS_3, as well as downstream EPH strains of *E. coli*, were screened. Briefly, fermentation was carried out in 96-well plates at 37°C for 72 hours. The yields of cucurbitadienol, 24,25-epoxy-cucurbitadienol, and 24,25-dihydroxycucurbitadienol were quantified using ultra-high performance liquid chromatography with diode array detection (UPLC-DAD). The effect of amino acid substitutions was evaluated by comparing the titers of cucurbitadienol, 24,25-epoxy-cucurbitadienol, and 24,25-dihydroxycucurbitadienol produced during fermentation by bacterial strains expressing a given ECDS_3 derivative with those expressing ECDS_3. Table 8 shows sixteen amino acid substitutions that showed beneficial changes in the titers of 24,25-epoxy-cucurbitadienol and 24,25-dihydroxycucurbitadienol compared to ECDS_3. Table 8: ECDS Engineering Phase 4
[0163] Based on these results, ECDS_4 (SEQ ID NO: 11), which carries L245I substitution in addition to the substitutions present in ECDS_3, was constructed and further investigated by comparing it with ECDS_3. Bacterial strains producing 2,3;22,23-squalene dioxide and expressing ECDS_3 or ECDS_4 (except EPH) were fermented in 96-well plates at 37°C for 72 hours. The yield of 2,25-dihydroxycucurbitadienol was quantified by UPLC-DAD. Figure 4D As shown, compared with bacterial strains expressing ECDS_3, bacterial strains expressing ECDS_4 produced approximately 1.4 times more 24,25-dihydroxy-cucurbitadienol.
[0164] E. coli strains that produced mutant derivatives of 2,3;22,23-squalene dioxide and expressed ECDS_4 (as well as EPH) were screened. Some of these mutant derivatives included double mutants (e.g., double mutants shown in several rows of Table 9). Briefly, fermentation was carried out in 96-well plates at 37°C for 72 h. The yields of cucurbitacinol, 24,25-epoxy-cucurbitacinol, and 24,25-dihydroxycucurbitacinol were quantified by UPLC-DAD. The effect of amino acid substitution was evaluated by comparing the titers of cucurbitacinol, 24,25-epoxy-cucurbitacinol, and 24,25-dihydroxycucurbitacinol produced during fermentation by bacterial strains expressing a given ECDS_4 mutant with those expressing ECDS_4. Table 9 shows twenty-three mutants that showed beneficial changes in the titer of 24,25-dihydroxycucurbitacinol compared to ECDS_4: Table 9: ECDS Engineering Phase 5
[0165] Based on these results, ECDS_5 (SEQ ID NO: 12), which carries a W331S substitution in addition to the substitutions present in ECDS_4, was further investigated and compared with ECDS_4. Bacterial strains expressing ECDS_4 or ECDS_5 (except EPH) were fermented in 96-well plates at 37°C for 72 hours. The yield of 24,25-dihydroxycucurbitadienol was quantified by UPLC-DAD. Figure 4E As shown, compared with bacterial strains expressing ECDS_4, bacterial strains expressing ECDS_5 produced approximately 1.25 times more 24,25-dihydroxy-cucurbitadienol.
[0166] A molecular model of cucurbitacin synthase 2 was developed. Not wishing to be bound by theory, the enzyme may include an entry channel located at the base of the cucurbitacin synthase structure, near its surface interacting with the cell membrane (cytoplasmic side), through which squalene substrates enter the active site. Interestingly, the mutations I553M, L245I, and W331S, identified in rounds 2, 4, and 5 of ECDS engineering, respectively, altered the topology of the entry channel. To directly compare the effects of these mutations on substrate specificity (and therefore product specificity), the production of 24,25-dihydroxy-cucurbitacinol and 24,25-epoxy-cucurbitacinol (both derived via ECDS cyclization of 2,3;22,23-squalene dioxide) was compared with the production of cucurbitacinol (derived via ECDS cyclization of 2,3-squalene dioxide). The ECDS variants of 24,25-epoxycucurbitacin synthase (ECDS) disclosed in this paper were compared with ECDS_0. In short, bacterial strains producing 2,3;22,23-squalene dioxide and expressing ECDS_0 or its derivatives with I553M, L245I, and W331S substitutions were fermented in 96-well plates at 37°C for 72 hours. The yields of cucurbitacinol, 24,25-epoxy-cucurbitacinol, and 24,25-dihydroxycucurbitacinol were quantified using UPLC-DAD. The effect of amino acid substitutions was evaluated by comparing the titers of cucurbitacinol or the total titers of 24,25-epoxy-cucurbitacinol + 24,25-dihydroxycucurbitacinol produced during fermentation by bacterial strains expressing a given ECDS derivative with those expressing ECDS_0. Table 10 shows the results: Table 10: Effects of mutations in the substrate entry channel of cucurbita dienol synthase.
[0167] As shown above, each mutation resulted in an increased yield of 24,25-dihydroxy-cucurbitadienol + 24,25-epoxy-cucurbitadienol derived from 2,3;22,23-squalene dioxide, and / or a decreased yield of cucurbitadienol derived from 2,3-squalene dioxide. These results particularly suggest that mutations in the squalene channel of cucurbitadienol synthase (CDS) provide improved substrate specificity, increasing the production of 24,25-dihydroxy-cucurbitadienol and 24,25-epoxy-cucurbitadienol (an intermediate in the production of mogroside), and reducing the production of the undesirable extra-pathway byproduct cucurbitadienol.
[0168] Figure 5The comparison of CcCDS (SEQ ID NO: 7) with homologs SgCDS (SEQ ID NO: 68), CpCAS (SEQ ID NO: 69), PsCAS (SEQ ID NO: 70), and AsCAS (SEQ ID NO: 71) is shown. Positions corresponding to D50, N121, L245, W331, Q401, I490, I553, and A556 of SEQ ID NO: 7 are highlighted. These positions are useful for mutagenesis, potentially affecting the putative substrate entry channel (e.g., due to membrane proximity), the active site, or other positions that have been shown to influence activity or substrate specificity.
[0169] Table 11: ECDS, locations used for mutagenesis that affect productivity or substrate specificity.
[0170] Table 12: Percentage identity of ECDS enzymes with each other.
[0171] Example 5: Synergistic effect of engineered squalene epoxidase (SQE) and 24,25-epoxycucurbitacin synthase (ECDS) Same effect Two *E. coli* strains expressing different engineered forms of squalene epoxidase (SQE) and 24,25-epoxycucurbitadienol synthase (ECDS) were constructed. These two strains expressed (1) *Artemisia annua* (… Artemisia annua (1) Squalene synthase (AaSQS; SEQ ID NO: 13), SQE_3, ECDS_4, and EPH_1; and (2) AaSQS, SQE_2, ECDS_3, and EPH_1. The yields of cucurbitacinol and 24,25-dihydroxycucurbitacinol from these strains were analyzed. Briefly, fermentation was carried out in 96-well plates at 37°C for 72 hours. The yields of cucurbitacinol and 24,25-dihydroxycucurbitacinol were quantified using UPLC-DAD. The ratio of cucurbitacinol to 24,25-dihydroxycucurbitacinol produced by these strains was determined and plotted. Figure 6 As shown, strains expressing AaSQS, SQE_2, ECDS_3, and EPH_1 produced less than half the amount of 24,25-dihydroxycucurbitadienol as cucurbitadienol. On the other hand, strains expressing AaSQS, SQE_3, ECDS_4, and EPH_1 produced more than twice the amount of 24,25-dihydroxycucurbitadienol as cucurbitadienol. Figure 6 Therefore, compared to the former strain, the latter strain is able to produce more than 4 times more 24,25-dihydroxy-cucurbitadiol relative to cucurbitadiol. Figure 6Given that strains expressing SQE_3 produce approximately twice as much 2,3;22,23-squalene dioxide compared to strains expressing SQE_2 (see...), Figure 2B Furthermore, strains expressing ECDS_4 alone produced approximately 1.4 times more 24,25-dihydroxy-cucurbitadienol compared to strains expressing ECDS_3 (see [link to ECDS_4]). Figure 4D These results demonstrate the synergistic effect of SQE and ECDS engineering.
[0172] Example 6: Engineering for C11-specific cytochrome P450 enzymes For zucchini subspecies pepo The cytochrome P450 enzyme CYP87A3 (CppCYP87A3, SEQ ID NO: 14) was engineered. As a first step, the native transmembrane domain of CppCYP87A3 was replaced with a transmembrane domain from *E. coli* sohB to produce sohB_CppCYP87A3 (SEQ ID NO: 15). See U.S. Patent 10,774,314, which is incorporated herein by reference in its entirety. The effects of these cytochrome P450 enzymes on 11-hydroxycucurbita-dienol production were compared in cucurbita-dienol-producing *E. coli* strains. Briefly, fermentation was carried out in 96-well plates at 37°C for 72 hours. 11-hydroxycucurbita-dienol yields were quantified using UPLC-DAD. Figure 7A As shown, the strain expressing sohB_CppCYP87A3 produces approximately 1.75 times more 11-hydroxycucurbitadienol than the strain expressing CppCYP87A3.
[0173] Escherichia coli strains expressing a mutant derivative of sohB_CppCYP87A3 were screened among bacterial strains producing 24,25-dihydroxycucurbitadienol. Based on these results, further studies were conducted comparing sohB_C11CYP_1 (SEQ ID NO: 16) carrying an S193T substitution with sohB_CppCYP87A3. In short, bacterial strains producing 24,25-dihydroxycucurbitadienol and expressing sohB_CppCYP87A3 or sohB_C11CYP_1 were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by GC-FID. Figure 7B As shown, bacterial strains expressing sohB_C11CYP_1 produced approximately 5-fold higher levels of mogroside compared to bacterial strains expressing sohB_CppCYP87A3. These results particularly demonstrate that sohB_C11CYP_1 has improved specificity or activity for hydroxylation at the C-11 position.
[0174] E. coli strains expressing mutant derivatives of SohB_C11CYP_1 were screened. In short, bacterial strains producing 24,25-dihydroxycucurbitadienol and expressing the sohB_C11CYP_1 derivative were fermented in 96-well plates at 37°C for 72 h. The yield of mogromaturin was quantified by GC-FID. The effect of amino acid substitutions in SohB_C11CYP_1 was evaluated by comparing the mogromaturin titers produced by bacterial strains expressing a given sohB_C11CYP_1 derivative with those expressing sohB_C11CYP_1. Table 13 shows sixty-four amino acid substitutions that showed beneficial changes in mogromaturin titers compared to sohB_C11CYP_1: Table 13: C11CYP Engineering Phase 2
[0175] Based on these results, SohB_C11CYP_2 (SEQ ID NO: 17), which carries substitutions of W130N, N234H, and A468T in addition to those present in SohB_C11CYP_1, was further investigated and compared with sohB_CppCYP87A3. Bacterial strains producing 24,25-dihydroxycucurbitadienol and expressing sohB_CppCYP87A3 or SohB_C11CYP_2 were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by GC-FID. Figure 7C As shown, compared with bacterial strains expressing sohB_CppCYP87A3, bacterial strains expressing SohB_C11CYP_2 produced approximately 2 times more mogroside.
[0176] E. coli strains expressing mutant derivatives of SohB_C11CYP_2 were screened. In short, bacterial strains producing 24,25-dihydroxycucurbitadienol and expressing either sohB_C11CYP_2 or a sohB_C11CYP_2 derivative were fermented in 96-well plates at 37°C for 72 h. The yield of mogromaturin was quantified by GC-FID. The effect of amino acid substitutions in SohB_C11CYP_2 was evaluated by comparing the mogromaturin titers produced during fermentation between bacterial strains expressing a given sohB_C11CYP_2 derivative and those expressing sohB_C11CYP_2. Table 14 shows fifty-three amino acid substitutions that showed beneficial changes in mogromaturin titers compared to sohB_C11CYP_2: Table 14: C11CYP Engineering Phase 3
[0177] Based on these results, SohB_C11CYP_3 (SEQ ID NO: 18), which carries an H433F substitution in addition to the substitutions present in SohB_C11CYP_2, was further compared with SohB_C11CYP_2. Bacterial strains producing 24,25-dihydroxy-cucurbitadienol and expressing SohB_C11CYP_3 or SohB_C11CYP_2 were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by GC-FID. Figure 7D As shown, compared with bacterial strains expressing SohB_C11CYP_2, bacterial strains expressing SohB_C11CYP_3 produced more than 2 times the level of mogroside.
[0178] E. coli strains expressing mutant derivatives of SohB_C11CYP_3 were screened. Some of the mutant derivatives included mutants with mutant clusters (see, for example, the mutant clusters shown in some rows of Table 15). Briefly, bacterial strains producing 24,25-dihydroxycucurbitadienol and expressing either SohB_C11CYP_3 or a SohB_C11CYP_3 derivative were fermented in 96-well plates at 37°C for 72 h. The yield of mogromaturin was quantified by GC-FID. The effect of amino acid substitutions in SohB_C11CYP_3 was evaluated by comparing the mogromaturin titers produced by bacterial strains expressing a given SohB_C11CYP_3 derivative with those expressing SohB_C11CYP_3. Table 15 shows thirty-two mutants that exhibited beneficial changes in mogromaturin titers compared to SohB_C11CYP_3: Table 15: C11CYP Engineering Phase 4
[0179] Based on these results, SohB_C11CYP_4 (SEQ ID NO: 19), which carries G125A and T128G substitutions in addition to those present in SohB_C11CYP_3, was further compared with SohB_C11CYP_3. Bacterial strains producing 24,25-dihydroxy-cucurbitadienol and expressing SohB_C11CYP_4 or SohB_C11CYP_3 were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by GC-FID. Figure 7EAs shown, compared with bacterial strains expressing SohB_C11CYP_3, bacterial strains expressing SohB_C11CYP_4 produced more than 1.6 times higher levels of mogroside.
[0180] E. coli strains expressing mutant derivatives of SohB_C11CYP_4, recoded forms, or derivatives with alternative N-terminal anchors were screened. Some of the mutant derivatives included mutants with mutant clusters (see, for example, mutant clusters shown in some rows of Table 16). Briefly, bacterial strains producing 24,25-dihydroxycucurbitadienol and expressing SohB_C11CYP_4 or SohB_C11CYP_4 derivatives were fermented in 96-well plates at 37°C for 72 h. The yield of mogromaturin was quantified by GC-FID. The effect of the SohB_C11CYP_4 mutant was evaluated by comparing the mogromaturin titers produced during fermentation by bacterial strains expressing a given SohB_C11CYP_4 derivative with those expressing SohB_C11CYP_4. Table 16 shows fourteen derivatives that exhibited beneficial changes in mogrool titers compared to SohB_C11CYP_4: Table 16: C11CYP Engineering Phase 5
[0181] Based on these results, SohB_C11CYP_5 (SEQ ID NO: 20), which carries an L131V substitution in addition to the substitutions present in SohB_C11CYP_4, was further compared with SohB_C11CYP_4. Bacterial strains producing 24,25-dihydroxy-cucurbitadienol and expressing SohB_C11CYP_5 or SohB_C11CYP_4 were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by GC-FID. Figure 7F As shown, compared with bacterial strains expressing SohB_C11CYP_4, bacterial strains expressing SohB_C11CYP_5 produced more than 1.3 times higher levels of mogroside.
[0182] E. coli strains expressing mutant derivatives of SohB_C11CYP_5 were screened to improve substrate specificity. Based on these data, SohB_C11CYP_6 (SEQ ID NO: 21), which carries substitutions of V131L, K132N, and V465I in addition to those present in SohB_C11CYP_5, was further compared with SohB_C11CYP_5. Bacterial strains producing 24,25-dihydroxy-cucurbitadienol and expressing SohB_C11CYP_6 or SohB_C11CYP_5 were fermented in 96-well plates at 37°C for 72 hours. The yields of mogroside, 11-hydroxy-cucurbitadienol, and oxo-mogroside were quantified by GC-FID. Figure 7G As shown, bacterial strains expressing SohB_C11CYP_6 produced lower levels of 11-hydroxy-cucurbitadienol and oxo-mogroside compared to bacterial strains expressing SohB_C11CYP_5, without affecting mogroside levels. These results particularly indicate that SohB_C11CYP_6 has improved substrate specificity (e.g., better differentiation of cucurbitadienol than 2,4,25-dihydroxy-cucurbitadienol) compared to SohB_C11CYP_5.
[0183] E. coli strains expressing mutant derivatives of SohB_C11CYP_6 were screened. In short, bacterial strains expressing either MogCPR2_0 or MogCPR2_2 and producing mogromaturin via SohB_C11CYP_6 derivatives were fermented in 96-well plates at 37°C for 72 h. The yield of mogromaturin was quantified by GC-FID. The effect of amino acid substitutions in SohB_C11CYP_6 was evaluated by comparing the resulting mogromaturin titers. Table 17 shows nine groups of amino acid substitutions that showed beneficial changes in mogromaturin titers compared to SohB_C11CYP_6: Table 17: C11CYP Engineering Phase 6
[0184] E. coli strains expressing additional mutant derivatives of SohB_C11CYP_6 were screened. In short, bacterial strains expressing either MogCPR2_0 or MogCPR2_2 and producing mogromaturin via SohB_C11CYP_6 derivatives were fermented in 96-well plates at 37°C for 72 h. The yields of mogromaturin and 11-oxomogromaturin were quantified by GC-FID. The effect of amino acid substitutions in SohB_C11CYP_6 was evaluated by comparing the titers of the produced mogromaturin and 11-oxomogromaturin. Table 18 shows eight groups of amino acid substitutions that showed beneficial changes in mogromaturin titers compared to SohB_C11CYP_6: Table 18: C11CYP Engineering Phase 7
[0185] Based on these and other results, SohB_C11CYP_8 (SEQ ID NO:45) with L218I and G219A substitutions was further investigated and compared with SohB_C11CYP_6. Bacterial strains expressing SohB_C11CYP_6 or SohB_C11CYP_8 to produce mogromaturin and expressing MogCPR2_0 or MogCPR2_2 were fermented in 96-well plates at 37°C for 72 hours. The yields of mogromaturin and 11-oxomogromaturin were quantified by GC-FID. Figure 7H As shown, compared with bacterial strains expressing SohB_C11CYP_6, bacterial strains expressing SohB_C11CYP_8 produced approximately 1.2 times more mogroside and approximately 2.3 times more 11-oxomogroside.
[0186] E. coli strains expressing mutant derivatives of SohB_C11CYP_8 were screened. In short, bacterial strains expressing either MogCPR2_1 or MogCPR2_1 that produce mogrool were fermented in 96-well plates at 37°C for 72 h. The yield of mogrool was quantified by GC-FID. The effect of amino acid substitutions in SohB_C11CYP_8 was evaluated by comparing the resulting mogrool titers. Table 19 shows seventeen groups of amino acid substitutions that showed beneficial changes in mogrool titers compared to SohB_C11CYP_8: Table 19: C11CYP Engineering Phase 8
[0187] Based on these results, SohB_C11CYP_9 (SEQ ID NO: 46) with N102H and S63F substitutions was further investigated by comparing it with SohB_C11CYP_8. Bacterial strains expressing SohB_C11CYP_8 or SohB_C11CYP_9 to produce mogromaturin and expressing MogCPR2 were fermented in 96-well plates at 37°C for 72 hours. The yields of mogromaturin and 11-oxomogromaturin were quantified by UPLC-DAD. Figure 7I The bacterial strain expressing SohB_C11CYP_9 produced approximately 1.25 times more mogroside and 0.25 times less 11-oxomogroside compared to the bacterial strain expressing SohB_C11CYP_8.
[0188] E. coli strains expressing mutant derivatives of SohB_C11CYP_9 were screened. In short, a bacterial strain expressing the SohB_C11CYP_9 derivative to produce mogromaturin and expressing MogCPR2_ was fermented in 96-well plates at 37°C for 72 h. The yield of mogromaturin was quantified by UPLC-DAD. The effect of amino acid substitutions in SohB_C11CYP_9 was evaluated by comparing the resulting mogromaturin titers. Table 20 shows twenty-three groups of amino acid substitutions that showed beneficial changes in mogromaturin titers compared to SohB_C11CYP_9: Table 20: C11CYP Engineering Round 9
[0189] Based on these results, SohB_C11CYP_10 (SEQ ID NO:47) with F63N, K75R, and F354Y substitutions was further investigated and compared with SohB_C11CYP_9. Bacterial strains expressing SohB_C11CYP_9 or SohB_C11CYP_10 to produce mogroside and expressing MogCPR2 were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by UPLC-DAD. Figure 7J As shown, compared with bacterial strains expressing SohB_C11CYP_9, bacterial strains expressing SohB_C11CYP_10 produced approximately 1.4 times more mogroside.
[0190] To investigate the increased amounts of mogrosides and their derivatives compared to compounds derived from 2,25-dihydroxycucurbitacin and deoxymogroside, *E. coli* strains expressing mutant derivatives of SohB_C11CYP_10 were screened. In short, bacterial strains expressing mogrosides via SohB_C11CYP_10 derivatives, MogCPR2, and two UGTs (MogUGTc24_3 and MogUGTc3_2) were fermented in 96-well plates at 37°C for 72 h. The yields of mogrosides and their derivatives were quantified by liquid chromatography-mass spectrometry (LC-MS). The effect of amino acid substitutions in SohB_C11CYP_10 was evaluated by comparing the titers of mogrosides and their derivatives. Table 21 shows the twelve amino acid substitutions that show beneficial changes in mogrool titers compared to SohB_C11CYP_10: Table 21: C11CYP Engineering Round 10
[0191] Based on these results, a further study was conducted comparing SohB_C11CYP_11 (SEQ ID NO: 80), which, in addition to the mutation in SohB_C11CYP_10, also possesses a K369R substitution, with SohB_C11CYP_10. Bacterial strains expressing SohB_C11CYP_11 or SohB_C11CYP_11 to produce mogroside and expressing MogCPR2, as well as UGT MogUGTc24_3 and MogUGTc3_2, were fermented in 96-well plates at 37°C for 72 hours. The yields of mogroside and deoxymogroside were quantified by LC-MS. Figure 7K As shown, compared with bacterial strains expressing SohB_C11CYP_10, bacterial strains expressing SohB_C11CYP_11 produced approximately 1.6 times more mogroside.
[0192] To investigate the increased amounts of mogrosides and mogroside-derived mogrosides compared to compounds derived from 2,4,25-dihydroxycucurbitadienol and deoxymogroside, *E. coli* strains expressing mutant derivatives of SohB_C11CYP_11 were screened. In short, bacterial strains expressing mogrosides via SohB_C11CYP_11 derivatives, MogCPR2, and UGTMogUGTc24_3 and MogUGTc3_2 were fermented in 96-well plates at 37°C for 72 h. The production of mogrosides and deoxymogrosides was quantified by liquid chromatography-mass spectrometry (LC-MS). The effect of amino acid substitutions in SohB_C11CYP_11 was evaluated by comparing the titers of mogrosides and deoxymogrosides. Table 22 shows eleven groups of amino acid substitutions that showed beneficial changes in mogroside titers compared to SohB_C11CYP_10: Table 22: C11CYP Engineering Round 11
[0193] Based on these results, SohB_C11CYP_12 (SEQ ID NO: 81), which, in addition to the mutation in SohB_C11CYP_11, also possesses T223S and V228I substitutions, was further compared with SohB_C11CYP_11. Bacterial strains expressing mogrosides (SohB_C11CYP_11) or SohB_C11CYP_12, and expressing MogCPR2, as well as UGT MogUGTc24_3 and MogUGTc3_2, were fermented in 96-well plates at 37°C for 72 hours. The yields of mogrosides and deoxymogrosides were quantified by LC-MS. Figure 7L As shown, compared with bacterial strains expressing SohB_C11CYP_11, bacterial strains expressing SohB_C11CYP_12 produced approximately 1.3 times more mogroside.
[0194] To investigate the increased amounts of mogrosides and mogrosides derived from mogrosides relative to compounds derived from 2,2,5-dihydroxycucurbitacin and deoxymogroside, *E. coli* strains expressing mutant derivatives of SohB_C11CYP_12 were screened. In short, bacterial strains expressing mogrosides via SohB_C11CYP_12 derivatives, MogCPR2, and UGTMogUGTc24_3 and MogUGTc3_2 were fermented in 96-well plates at 37°C for 72 h. The yields of mogrosides and deoxymogrosides were quantified by liquid chromatography-mass spectrometry (LC-MS). The effect of amino acid substitutions in SohB_C11CYP_12 was evaluated by comparing the titers of mogrosides and deoxymogrosides. Table 23 shows six groups of amino acid substitutions that showed beneficial changes in mogroside titers compared to SohB_C11CYP_10: Table 23: C11CYP Engineering Phase 12
[0195] Based on these results, SohB_C11CYP_13 (SEQ ID NO: 82), which, in addition to the mutation in SohB_C11CYP_12, also possesses T231F, T232A, and Y233F substitutions, was further compared with SohB_C11CYP_12. Bacterial strains expressing mogrosides (SohB_C11CYP_12) or SohB_C11CYP_13, and expressing MogCPR2, as well as UGT MogUGTc24_3 and MogUGTc3_2, were fermented in 96-well plates at 37°C for 72 hours. The yields of mogrosides, deoxymogrosides, and 11-oxomogrosides were quantified by LC-MS. Figure 7M As shown, compared with bacterial strains expressing SohB_C11CYP_12, bacterial strains expressing SohB_C11CYP_13 produced approximately 60% less 11-oxomonasoside.
[0196] To investigate the increased amounts of mogrosides and mogrosides derived from mogrosides relative to compounds derived from 2,2,5-dihydroxycucurbitacin and deoxymogrosides, *E. coli* strains expressing mutant derivatives of SohB_C11CYP_13 were screened. In short, bacterial strains expressing mogrosides via SohB_C11CYP_13 derivatives, MogCPR2, and UGTMogUGTc24_3 and MogUGTc3_2 were fermented in 96-well plates at 37°C for 72 h. The yields of mogrosides and deoxymogrosides were quantified by liquid chromatography-mass spectrometry (LC-MS). The effect of amino acid substitutions in SohB_C11CYP_13 was evaluated by comparing the titers of mogrosides and deoxymogrosides. Table 24 shows twenty-three groups of amino acid substitutions that showed beneficial changes in mogroside titers compared to SohB_C11CYP_10: Table 24: C11CYP Engineering Phase 13
[0197] Based on these results, SohB_C11CYP_14 (SEQ ID NO: 83), which, in addition to the mutation in SohB_C11CYP_13, also possesses A12W, F452I, and T192L substitutions, was further investigated and compared with SohB_C11CYP_13. Using bacterial strains that produce mogroside by expressing SohB_C11CYP_13 or SohB_C11CYP_14 and expressing MogCPR2, as well as UGT MogUGTc24_3 and MogUGTc3_2, fermentation was carried out in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by LC-MS. Figure 7N As shown, compared with bacterial strains expressing SohB_C11CYP_13, bacterial strains expressing SohB_C11CYP_14 produced approximately 1.7 times more mogroside.
[0198] To investigate the increased amounts of mogrosides and mogrosides derived from mogrosides relative to compounds derived from 24,25-dihydroxycucurbitadienol and deoxymogrosides, *E. coli* strains expressing mutant derivatives of SohB_C11CYP_14 were screened. In short, bacterial strains expressing mogrosides, MogCPR2, UGTMogUGTc24_3, and MogUGTc3_2, were fermented in 96-well plates at 37°C for 72 h. The yields of mogrosides and deoxymogrosides were quantified by liquid chromatography-mass spectrometry (LC-MS). The effect of amino acid substitutions in SohB_C11CYP_14 was evaluated by comparing the titers of mogrosides and deoxymogrosides. Table 25 shows twenty-four groups of amino acid substitutions that showed beneficial changes in mogroside titers compared to SohB_C11CYP_10: Table 25: C11CYP Engineering Round 14
[0199] Based on these results, SohB_C11CYP_15 (SEQ ID NO: 84), which, in addition to the mutation in SohB_C11CYP_14, also possesses I27G, K246V, T482S, and V76M substitutions, was further compared with SohB_C11CYP_14. Bacterial strains expressing SohB_C11CYP_14 or SohB_C11CYP_15 to produce mogroside and expressing MogCPR2, as well as UGTMogUGTc24_3 and MogUGTc3_2, were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by LC-MS. Figure 7O As shown, compared with bacterial strains expressing SohB_C11CYP_14, bacterial strains expressing SohB_C11CYP_15 produced approximately 1.4 times more mogroside.
[0200] A variant of SohB_C11CYP_15, n20_C11CYP_15 (SEQ ID NO: 85), with deletions of amino acids L3 to H29, was constructed. This mutant showed improved solubility compared to SohB_C11CYP_15. Further investigation was conducted by comparing bacterial strains producing the SohB_C11CYP_15 mutant with SohB_C11CYP_15. Fermentation was performed in 96-well plates at 37°C for 72 hours using bacterial strains expressing either SohB_C11CYP_15 or n20_C11CYP_15 to produce mogroside and expressing MogCPR2, as well as UGT MogUGTc24_3 and MogUGTc3_2. The yield of mogroside was quantified by LC-MS. Figure 7P As shown, bacterial strains expressing n20_C11CYP_15 produced similar levels of mogroside compared to bacterial strains expressing SohB_C11CYP_15.
[0201] To investigate the increased amounts of mogrosides and mogrosides derived from mogrosides relative to compounds derived from 2,25-dihydroxycucurbitacin and deoxymogrosides, *E. coli* strains expressing mutant derivatives of n20_C11CYP_15 were screened. In short, bacterial strains expressing mogrosides, MogCPR2, UGTMogUGTc24_3, and MogUGTc3_2, which produce mogrosides via n20_C11CYP_15 derivatives, were fermented in 96-well plates at 37°C for 72 h. The yields of mogrosides and deoxymogrosides were quantified by liquid chromatography-mass spectrometry (LC-MS). The effect of amino acid substitutions in n20_C11CYP_15 was evaluated by comparing the titers of mogrosides and deoxymogrosides. Table 26 shows four amino acid substitutions relative to SEQ ID NO: 85, which show beneficial changes in mogrool titers compared to n20_C11CYP_15: Table 26: C11CYP Engineering Round 15
[0202] Based on these results, further studies were conducted comparing n20_C11CYP_16 (SEQ ID NO: 86), which, in addition to the mutation in n20_C11CYP_15, also possesses K8R, D9N, S10R, N13K, and V15K substitutions relative to SEQ ID NO: 85, with n20_C11CYP_15. Bacterial strains expressing either n20_C11CYP_15 or n20_C11CYP_16 to produce mogroside and expressing MogCPR2, as well as UGT MogUGTc24_3 and MogUGTc3_2, were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by LC-MS. Figure 7Q As shown, compared with bacterial strains expressing n20_C11CYP_15, bacterial strains expressing n20_C11CYP_16 produced approximately 2.25 times more mogroside.
[0203] To investigate the increased amounts of mogroside and its derivatives relative to compounds derived from 2,25-dihydroxycucurbitacin and deoxymogroside, *E. coli* strains expressing mutant derivatives of n20_C11CYP_16 were screened. In short, bacterial strains expressing mogroside, MogCPR2, UGTMogUGTc24_3, and MogUGTc3_2, which produce mogroside via n20_C11CYP_16 derivatives, were fermented in 96-well plates at 37°C for 72 h. The yields of mogroside and deoxymogroside were quantified by liquid chromatography-mass spectrometry (LC-MS). The effect of amino acid substitutions in n20_C11CYP_16 was evaluated by comparing the titers of mogroside and deoxymogroside. Table 27 shows five groups of amino acid substitutions that showed beneficial changes in mogroside titers compared to n20_C11CYP_16.
[0204] Table 27: C11CYP Engineering Round 16
[0205] Based on these results, n20_C11CYP_17 (SEQ ID NO: 87), which, in addition to the mutation in n20_C11CYP_16, also has I45V, K46R, K47E, M49V, K50E, and R51K substitutions relative to SEQ ID NO: 85, was further compared with n20_C11CYP_16. Bacterial strains expressing either n20_C11CYP_16 or n20_C11CYP_17 to produce mogroside and expressing MogCPR2, as well as UGT MogUGTc24_3 and MogUGTc3_2, were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified by LC-MS. Figure 7R As shown, compared with bacterial strains expressing n20_C11CYP_16, bacterial strains expressing n20_C11CYP_17 produced approximately 1.4 times more mogroside.
[0206] To investigate the increased amounts of mogroside and its derivatives relative to compounds derived from 2,25-dihydroxycucurbitacin and deoxymogroside, *E. coli* strains expressing mutant derivatives of n20_C11CYP_17 were screened. In short, bacterial strains expressing mogroside, MogCPR2, UGTMogUGTc24_3, and MogUGTc3_2, which produce mogroside via n20_C11CYP_17 derivatives, were fermented in 96-well plates at 37°C for 72 h. The yields of mogroside and deoxymogroside were quantified by liquid chromatography-mass spectrometry (LC-MS). The effect of amino acid substitutions in n20_C11CYP_17 was evaluated by comparing the titers of mogroside and deoxymogroside. Table 28 shows five groups of amino acid substitutions that showed beneficial changes in mogroside titers compared to n20_C11CYP_16.
[0207] Table 28: C11CYP Engineering Round 17
[0208] Based on these results, further investigations were conducted comparing n20_C11CYP_18 (SEQ ID NO: 88), which, in addition to the mutation in n20_C11CYP_17, also possesses I191L and A192K substitutions relative to SEQ ID NO: 85, with n20_C11CYP_17. Bacterial strains expressing mogromaturin via n20_C11CYP_17 or n20_C11CYP_18 and expressing MogCPR2, as well as UGT MogUGTc24_3 and MogUGTc3_2, were fermented in 96-well plates at 37°C for 72 hours. The yield of mogromaturin was quantified by LC-MS. Figure 7S As shown, compared with bacterial strains expressing n20_C11CYP_17, bacterial strains expressing n20_C11CYP_18 produced approximately 1.35 times more mogroside.
[0209] To investigate the increased amounts of mogrosides and mogrosides derived from mogrosides relative to compounds derived from 24,25-dihydroxycucurbita diol and deoxymogrosides, *E. coli* strains expressing mutant derivatives of n20_C11CYP_18 were screened. In short, bacterial strains expressing mogrosides, MogCPR2, and UGT MogUGTc24_3 and MogUGTc3_2 were fermented in 96-well plates at 37°C for 72 hours. The yields of mogrosides and deoxymogrosides were quantified by liquid chromatography-mass spectrometry (LC-MS). The effect of amino acid substitutions in n20_C11CYP_18 was evaluated by comparing the titers of mogrosides and deoxymogrosides. Table 29 shows twenty-three amino acid substitutions that exhibit beneficial changes in mogrool titers compared to n20_C11CYP_16.
[0210] Table 29: C11CYP Engineering Phase 18
[0211] Based on these results, n20_C11CYP_19 (SEQ ID NO: 89), which, in addition to the mutation in n20_C11CYP_17, also has the substitutions of F406Y, E411D, Y412F, and S413A relative to SEQ ID NO: 85, was further compared with n20_C11CYP_17. Bacterial strains expressing mogromaturin via n20_C11CYP_17 or n20_C11CYP_18 and expressing MogCPR2, as well as UGT MogUGTc24_3 and MogUGTc3_2, were fermented in 96-well plates at 37°C for 72 hours. The yield of mogromaturin was quantified by LC-MS. Figure 7T As shown, bacterial strains expressing n20_C11CYP_18 produce approximately 1.5 times more [products / products]. Example 7: Engineering of cytochrome P450 reductase right pepo subspecies of zucchini Cytochrome P450 reductase (CppCPR4, also referred to herein as MogCPR1_0, SEQ ID NO: 22) was engineered. As a first step, the native transmembrane domain of CppCPR4 was replaced with transmembrane domains derived from various *E. coli* membrane proteins. In this process, CppCPR4 was truncated, removing the N-terminal 67 amino acids. The effects of these transmembrane domains on mogromol production were compared in *E. coli* strains expressing SohB_C11CYP_4. Briefly, fermentation was carried out in 96-well plates at 37°C for 72 h. Mogromol yield was quantified by UPLC-DAD. As shown in Table 30, strains expressing CppCPR4 with transmembrane domains derived from various *E. coli* membrane proteins produced more mogromol compared to strains expressing CppCPR4 itself. U.S. Patent No. 10,774,314 describes E. coli membrane proteins and the use of their N-terminal transmembrane domains or portions thereof as membrane anchors for cytochrome P450 enzymes and reductases, which is incorporated herein by reference in its entirety.
[0212] Table 30: C11CPR Engineering Round 1
[0213] E. coli strains expressing mutant derivatives of CppCPR4 were screened. In short, bacterial strains producing mogrool and expressing CppCPR4 or CppCPR4 derivatives were fermented in 96-well plates at 37°C for 72 hours. Mogrool yield was quantified using UPLC-DAD. The effect of amino acid substitutions in CppCPR4 was evaluated by comparing the mogrool titers produced during fermentation between bacterial strains expressing a given CppCPR4 derivative and those expressing CppCPR4. Table 31 shows eighteen mutants that exhibited beneficial changes in mogrool titers compared to CppCPR4: Table 31: C11CPR Engineering Round 2
[0214] Based on these results in Tables 30 and 31, the deletion of 67 amino acids at the N-terminus and the E. coli-based... ycgG The transmembrane domains of the 25 N-terminal amino acids (i.e., n25_ ycgG Further studies were conducted comparing MogCPR1_1 (SEQ ID NO: 23) with MogCPR1_0 (CppCPR4). Fermentation was performed in 96-well plates at 37°C for 72 hours using bacterial strains that produce mogroside and express either MogCPR1_0 or MogCPR1_1. The yield of mogroside was quantified using UPLC-DAD. Figure 8A As shown, compared with bacterial strains expressing MogCPR1_0, bacterial strains expressing MogCPR1_1 produced more than 1.25 times higher levels of mogroside.
[0215] Green beans ( Phaseolus vulgaris Cytochrome P450 reductase (PvCPR, also referred to herein as MogCPR2_0, SEQ ID NO: 24) was compared with MogCPR1_0. Fermentation was performed in 96-well plates at 37°C for 72 hours using bacterial strains that produced mogroside and expressed MogCPR1_0 or MogCPR2_0. The yield of mogroside was quantified by UPLC-DAD. Figure 8B As shown, compared with bacterial strains expressing MogCPR1_0, bacterial strains expressing MogCPR2_0 produced more than 1.4 times higher levels of mogroside.
[0216] Replace the native transmembrane domain of MogCPR2_0 with one from E. coli. ycgGThe transmembrane domain of MogCPR2_0 was truncated during the process, removing the N-terminal 75 amino acids. The effects of these transmembrane domains with various point mutations on mogromol production were compared in *E. coli* strains expressing SohB_C11CYP_6. Briefly, fermentation was carried out for 72 hours at 37°C in 96-well plates. Mogromol yield was quantified using UPLC-DAD. As shown in Table 32, strains expressing MogCPR2_0 with transmembrane domains derived from *E. coli* ycgG* and various point mutations produced more mogromol compared to strains expressing MogCPR2_0.
[0217] Table 32: C11CPR Engineering Round 3
[0218] Based on these results, the deletion of 75 amino acids at the N-terminus and the E. coli-based ycgG The transmembrane domain of the N-terminal 25 amino acids (i.e., n25_) ycgG Further studies were conducted comparing MogCPR2_1 (SEQ ID NO: 25) with MogCPR2_0. Fermentation was performed in 96-well plates at 37°C for 72 hours using bacterial strains expressing SohB_C11CYP_6 to produce mogroside and expressing either MogCPR2_0 or MogCPR2_1. The yield of mogroside was quantified using UPLC-DAD. Figure 8C As shown, compared with bacterial strains expressing MogCPR2_0, bacterial strains expressing MogCPR2_1 produced increased levels of mogroside.
[0219] E. coli strains expressing mutant derivatives of MogCPR2_1 were screened. In short, bacterial strains that produce mogroside by expressing SohB_C11CYP_6 and express MogCPR2_1 or MogCPR2_1 derivatives were fermented in 96-well plates at 37°C for 72 h. The yield of mogroside was quantified by UPLC-DAD. The effect of amino acid substitutions in MogCPR2_1 was evaluated by comparing the resulting mogroside titers. Table 33 shows nine amino acid substitutions that showed beneficial changes in mogroside titers compared to MogCPR2_1: Table 33: C11CPR Engineering Round 4
[0220] Based on these results, MogCPR2_2 (SEQ ID NO: 26) with L90P substitution was further investigated and compared with MogCPR2_0. Bacterial strains expressing SohB_C11CYP_6 to produce mogroside and expressing either MogCPR2_0 or MogCPR2_2 were fermented in 96-well plates at 37°C for 72 hours. The yield of mogroside was quantified using UPLC-DAD. Figure 8D As shown, compared with bacterial strains expressing MogCPR2_0, bacterial strains expressing MogCPR2_2 produced approximately 1.25 times more mogroside.
[0221] Example 8: Engineering of a C3-specific uridine diphosphate-dependent glycosyltransferase (UGT) Various UGT homologs were screened for O-glycosylation activity at the C3 or C24 positions of mogroside. In short, the enzyme assay for converting mogroside to MogIE or MogIA was performed in 96-well plates at 37°C for 48 hours. The yields of MogIE or MogIA were quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). Table 34 shows the results: Table 34: Screening of various UGT homologs for the O-glycosylation activity of mogroside C3 or C24
[0222] Stevia (Stevia rebaudiana) UGT85C1 (MogUGTc3_0, SEQ ID NO: 27) was engineered. Mutant derivatives of MogUGTc3_0 were screened for an enzyme assay to convert MogIA to MogIIE. Assays were performed in 96-well plates at 37°C for 48 hours. The yield of MogIIE was quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effect of amino acid substitutions was evaluated by comparing the yield of MogIIE produced by a given amino acid substitution in MogUGTc3_0 with that produced by MogUGTc3_0. Table 35 shows three amino acid substitutions that showed beneficial changes in the conversion of MogIA to MogIIE compared to MogUGTc3_0: Table 35: MogUGTc3 Engineering Round 1
[0223] Based on these results, MogUGTc3_1 (SEQ ID NO: 28) with L41F, D49E, and C127F substitutions was further investigated and compared with MogUGTc3_0. Enzyme assays using MogUGTc3_0 or MogUGTc3_1 and MogIA substrate were performed in 96-well plates at 37°C for 48 hours. MogIIE yield was quantified using UPLC-QQQ. Figure 9A As shown, the MogUGTc3_1 enzyme produces approximately 2 times more MogIIE levels compared to the MogUGTc3_0 enzyme.
[0224] E. coli strains producing mogroside and expressing a mutant derivative of MogUGTc3_1 were screened. Briefly, fermentation was carried out in 96-well plates at 37°C for 72 hours. The yield of MogIE was quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effect of amino acid substitutions was evaluated by comparing the titers of MogIE produced during fermentation between bacterial strains expressing a given MogUGTc3_1 derivative and bacterial strains producing mogroside and expressing MogUGTc3_1. Table 36 shows thirty-three amino acid substitutions that showed beneficial changes in MogIE titers compared to MogUGTc3_1: Table 36: MogUGTc3 Engineering Round 2
[0225] Based on these results, MogUGTc3_2 (SEQ ID NO: 29) with A307I substitution was further compared with MogUGTc3_1. Bacterial strains expressing either MogUGTc3_1 or MogUGTc3_2 were fermented in 96-well plates at 37°C for 72 hours. The yield of MogIE was quantified using UPLC-QQQ. Figure 9B As shown, compared with bacterial strains expressing MogUGTc3_1, bacterial strains expressing MogUGTc3_2 produced approximately 1.6 times more MogIE levels.
[0226] To improve substrate selectivity, *E. coli* strains expressing mutant derivatives of MogUGTc3_2 were screened. These studies yielded MogUGTc3_3 (SEQ ID NO: 30) with T74M and F303Y substitutions. Further studies compared MogUGTc3_3 with MogUGTc3_2 were conducted. Fermentation was performed in 96-well plates at 37°C for 72 hours using bacterial strains that produced mogroside and expressed either MogUGTc3_2 or MogUGTc3_3. The yields of MogIE and two glucosylated 24,25-dihydroxycucurbitadienol derivatives (undesired byproducts) were quantified using UPLC-QQQ. Figure 9C As shown, compared with bacterial strains expressing MogUGTc3_2, bacterial strains expressing MogUGTc3_3 produced lower levels of glucosylated 24,25-dihydroxycucurbitadienol derivatives and approximately 1.25-fold increased MogIE levels. These results particularly indicate that MogUGTc3_3 exhibits improved substrate specificity (e.g., distinguishing mogroside from 24,25-dihydroxycucurbitadienol) compared to MogUGTc3_2.
[0227] For cantaloupe ( Cucumis melo CmUGTc3 (XP_008445481.1, SEQ ID NO: 60, “CmUGTc3_0”) was engineered. Mutant derivatives of CmUGTc3_0 were screened for an enzymatic assay to convert mogromectin to MogIE. Assays were performed in 96-well plates at 37°C for 72 hours. MogIE yield was quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effect of amino acid substitutions was evaluated by comparing the MogIE yield of a given CmUGTc3_0 variant with that of CmUGTc3_0. Table 37 shows thirty-four amino acid substitutions and one cyclic variant (CP16) that showed beneficial changes in the conversion of mogromectin to MogIE compared to CmUGTc3_0. Table 37: CmUGTc3_0 Engineering Round 1
[0228] Table 38: Circular Arrangement Variant Constructs
[0229] Further investigations were conducted by comparing CmUGTc3_1 (SEQ ID NO: 65), a circular variant of CmUGTc3_0 with a cleavage site at position 169, with CmUGTc3_0. Compared to the wild-type enzyme, the circular variant retains the same basic folding as the parent enzyme but has a different N-terminal position (e.g., “cleavage site”), where the original N-terminus and C-terminus are optionally linked by an additional linker sequence, optionally connected by a linker sequence. For example, in the circular variant, the N-terminal methionine is located at a site other than the native N-terminus in the protein. Enzyme assays were performed in 96-well plates at 37°C for 72 h using either CmUGTc3_0 or CmUGTc3_1 with mogroside substrate. MogIE yield data were quantified by UPLC-QQQ. Figure 9D As shown, the CmUGTc3_1 enzyme produced more than 1.5 times more MogIE levels compared to the CmUGTc3_0 enzyme.
[0230] Mutant derivatives of CmUGTc3_1 were screened for an enzymatic assay to convert mogroside to MogIE. The assay was performed in 96-well plates at 37°C for 72 hours. MogIE production was quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effect of amino acid substitutions was evaluated by comparing the MogIE production of a given CmUGTc3_1 variant with that of CmUGTc3_1. Table 39 shows thirty-seven variants that exhibited beneficial changes in the conversion of mogroside to MogIE compared to CmUGTc3_1. Table 39: CmUGTc3_0 Engineering Round 2
[0231] Based on these results, further studies were conducted comparing CmUGTc3_2 (SEQ ID NO:66), which has an L352M substitution relative to CmUGTc3_1, with CmUGTc3_1 and another CmUGTc3_1 derivative with an L448V substitution. Enzyme assays were performed in 96-well plates at 37°C for 72 hours using CmUGTc3_2, the L448V derivative of CmUGTc3_1, or CmUGTc3_1 with mogroside as a substrate. The yields of MogIE and deoxy-MogIE were quantified using UPLC-QQQ. Figure 9EAs shown, compared to the CmUGTc3_1 enzyme, the CmUGTc3_2 enzyme produced approximately 1.2-fold more MogIE and decreased deoxy-MogIE levels by approximately 0.9-fold. Compared to the CmUGTc3_1 enzyme, the L448V derivative of CmUGTc3_1 produced approximately 1.3-fold more MogIE and approximately 1.2-fold more deoxy-MogIE levels.
[0232] Based on the above results, CmUGTc3_3 (SEQ ID NO: 67), which has H40N, I216A, S241G, and M352L substitutions relative to CmUGTc3_1, was further compared with CmUGTc3_2. Enzyme assays were performed in 96-well plates at 37°C for 72 hours using either CmUGTc3_3 or CmUGTc3_2 with mogroside as a substrate. The yields of MogIE and deoxy-MogIE were quantified using UPLC-QQQ. Figure 9F As shown, compared with the CmUGTc3_2 enzyme, the CmUGTc3_3 enzyme produced approximately 1.6 times more MogIE and approximately 1.1 times more deoxy-MogIE.
[0233] Example 9: Engineering of a C24-specific uridine diphosphate-dependent glycosyltransferase (UGT) Licorice was screened ( Glycyrrhiza uralensis The mutant derivative of (UGT73F24) (MogUGTc24_0, SEQ ID NO:31) was used. In short, the enzyme assay for converting MogIE to MogIIE was performed in 96-well plates at 37°C for 48 hours. The yield of MogIIE was quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effect of amino acid substitutions was evaluated by comparing the yield of MogIIE produced by a given MogUGTc24_0 derivative with that produced by MogUGTc24_0. Table 40 shows five amino acid substitutions that showed beneficial changes in MogIIE titers compared to MogUGTc24_0: Table 40: MogUGTc24 Engineering Phase 1
[0234] Based on these results, MogUGTc24_1 (SEQ ID NO: 32) with A74E, I91F, and H101P substitutions was further investigated and compared with MogUGTc24_1. Enzymatic assays were performed in 96-well plates at 37°C for 48 hours to convert MogIE to MogIIE using either MogUGTc24_1 or MogUGTc24_0. The yield of MogIIE was quantified using UPLC-QQQ. Figure 10AAs shown, the MogUGTc24_1 enzyme produces approximately 3.5 times more MogIIE levels compared to the MogUGTc24_0 enzyme.
[0235] Mutant derivatives of MogUGTc24_1 were screened. In short, the enzyme assay for converting MogIE to MogIIE was performed in 96-well plates at 37°C for 48 hours. The yield of MogIIE was quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effect of amino acid substitutions was evaluated by comparing the conversion of MogIE to MogIIE by a given MogUGTc24_1 derivative with that of MogUGTc24_1. Table 41 shows ninety-seven amino acid substitutions that showed beneficial changes in MogIIE titers compared to MogUGTc24_0: Table 41: MogUGTc24 Engineering Round 2
[0236] Based on these results, MogUGTc24_2 (SEQ ID NO: 33), which has substitutions of T184F, Y381F, T95A, P190E, and L18M in addition to those present in MogUGTc24_1, was further compared with MogUGTc24_1. Enzymatic assays were performed on the conversion of MogIE to MogIIE using either MogUGTc24_2 or MogUGTc24_1 in 96-well plates at 37°C for 48 hours. The yield of MogIIE was quantified using UPLC-QQQ. Figure 10B As shown, the MogUGTc24_2 enzyme produces approximately 4.5 times more MogIIE levels compared to the MogUGTc24_1 enzyme.
[0237] E. coli strains expressing mutant derivatives of MogUGTc24_2 were screened. In short, the mogroside-producing strains were fermented in 96-well plates at 37°C for 72 hours. The yield of MogIA was quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effect of amino acid substitution was evaluated by comparing the titers of MogIA produced during fermentation by bacterial strains expressing a given MogUGTc24_2 derivative with those expressing MogUGTc24_2. Table 42 shows twenty-nine mutants that exhibited beneficial changes in MogIA titers compared to MogUGTc24_2: Table 42: MogUGTc24 Engineering Phase 3
[0238] Based on these results, further studies were conducted comparing MogUGTc24_3 (SEQ ID NO: 34) with substitutions of S88A, T273S, N306K, and M332L, in addition to those present in MogUGTc24_2. Fermentation was performed in 96-well plates at 37°C for 72 hours using bacterial strains that produce mogroside and express MogUGTc24_3 or MogUGTc24_2. The yield of MogIA was quantified using UPLC-QQQ. Figure 10C As shown, compared with bacterial strains expressing MogUGTc24_2, bacterial strains expressing MogUGTc24_3 produced approximately 4 times more MogIA levels.
[0239] E. coli strains expressing mutant derivatives of MogUGTc24_3 were screened. In short, the mogroside-producing strains were fermented in 96-well plates at 37°C for 72 hours. The yield of MogIA was quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effect of amino acid substitution was evaluated by comparing the titers of MogIA produced during fermentation by bacterial strains expressing a given MogUGTc24_3 derivative with those expressing MogUGTc24_3. Table 43 shows ten mutants that exhibited beneficial changes in MogIA titers compared to MogUGTc24_3: Table 43: MogUGTc24 Engineering Round 4
[0240] Based on these and other results, MogUGTc24_4 (SEQ ID NO: 48), which has an F187Y substitution in addition to the substitutions present in MogUGTc24_3, was further compared with MogUGTc24_3. Fermentation was carried out in 96-well plates at 37°C for 72 hours using bacterial strains that produce mogroside and express MogUGTc24_3 or MogUGTc24_4. The yield of MogIA was quantified by UPLC-QQQ. Figure 10D As shown, compared with bacterial strains expressing MogUGTc24_3, bacterial strains expressing MogUGTc24_4 produced approximately 1.2 times more MogIA.
[0241] E. coli strains expressing mutant derivatives of MogUGTc24_4 were screened. In short, the mogroside-producing strains were fermented in 96-well plates at 37°C for 72 hours. The yields of MogIA and deoxy-MogIA were quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effects of amino acid substitution were evaluated by comparing the titers of MogIA and deoxy-MogIA produced during fermentation and by calculating the improvement in MogIA transformation specificity of bacterial strains expressing a given MogUGTc24_4 derivative compared to bacterial strains expressing MogUGTc24_4. Table 44 shows twenty-three mutants that showed beneficial changes in MogIA titers compared to MogUGTc24_4, several of which also showed improved specificity. Table 44: MogUGTc24 Engineering Round 5
[0242] Based on these results, further studies were conducted comparing MogUGTc24_5 (SEQ ID NO: 49), which, in addition to the substitutions present in MogUGTc24_4, also has an I63V substitution, with MogUGTc24_4. Fermentation was performed in 96-well plates at 37°C for 72 hours using bacterial strains that produce mogroside and express either MogUGTc24_5 or MogUGTc24_4. The yields of MogIA and deoxyMogA were quantified using UPLC-QQQ. Figure 10E As shown, compared with bacterial strains expressing MogUGTc24_4, bacterial strains expressing MogUGTc24_5 produced approximately 1.4 times more MogIA levels, while deoxy MogA levels increased only approximately 1.1 times.
[0243] E. coli strains expressing mutant derivatives of MogUGTc24_5 were screened. In short, the mogroside-producing strains were fermented in 96-well plates at 37°C for 72 hours. The yields of MogIA and deoxy-MogIA were quantified by ultra-high performance liquid chromatography-triple quadrupole mass spectrometry (UPLC-QQQ). The effect of the mutation was evaluated by comparing the titers of MogIA and deoxy-MogIA produced during fermentation and by calculating the improvement in MogIA transformation specificity of bacterial strains expressing a given MogUGTc24_5 derivative compared to bacterial strains expressing MogUGTc24_5. Table 45 shows the mutants that showed beneficial changes in MogIA titers compared to MogUGTc24_5: Table 45: MogUGTc24 Engineering Round 6
[0244] Example 10: Uridine diphosphate-dependent glycosyltransferase (UGT) for 1,6 and 1,2 branched glycosylation Engineering Filtered Monk fruit Derivatives of UGT94-289 (SEQ ID NO: 35), including its cyclic variant (MogUGT1216_0, SEQ ID NO: 36) having a cleavage site at V187. U.S. Patent No. 10,463,062 discloses a strategy for deriving cyclic variants of UGT enzymes, which is incorporated herein by reference in its entirety. Enzymatic assays for converting MogIIE to MogV were performed for 48 hours in 96-well plates at 37°C using MogUGT1216_0 and its variants. MogV was quantified using UPLC-QQQ. Table 46 shows numerous derivatives that exhibit beneficial changes in MogV titers compared to MogUGT1216_0: Table 46: First Round of Engineering for Uridine Diphosphate-Dependent Glycosyltransferases (UGTs) for 1,6 and 1,2 Branched Glycosylation
[0245] Based on these results, MogUGT1216_1 (SEQ ID NO: 37), which, in addition to the circular arrangement mutation at V187 of SgUGT94_289_1, also possesses L14V and L266Q substitutions, was further investigated and compared with SgUGT94_289_1 and MogUGT1216_0. Enzymatic assays were performed in 96-well plates at 37°C for 48 hours to convert MogIIE to MogV using MogUGT1216_0, SgUGT94_289_1, and MogUGT1216_1. MogV data were quantified using UPLC-QQQ. Figure 11A As shown, SgUGT94_289_1 produces approximately 10-fold more MogV levels compared to the cyclic variant MogUGT1216_0. Interestingly, MogUGT1216_1 produces approximately 1.4-fold more MogV levels compared to SgUGT94_289_1. Figure 11A ).
[0246] The effects of L14V and L266Q mutations in MogUGT1216_1 were investigated. To this end, MogUGT1216_0 derivatives with L14V and L266Q single mutants were created and compared with MogUGT1216_0 (without mutation). Enzymatic assays were performed in 96-well plates at 37°C for 48 hours to convert MogIIIIE to various mogrosides. Data on different mogrosides were quantified and plotted using UPLC-QQQ, and the amounts of MogIII, MogIV-A, ciprofloxacin, and MogV produced were normalized to the productivity of MogUGT1216_0. The L14V (Mut1) and L266Q (Mut16) single mutants produced approximately 2-fold and approximately 5-fold increases in MogV, respectively. Figure 11B In comparison, such as Figure 11B As shown, compared to MogUGT1216_0, the L14V / L266Q double mutant produced approximately 18-fold increased MogV levels. The L14V single mutant and L266Q single mutant also produced significantly higher levels of ciprofloxacin (…). Figure 11B ).
[0247] Derivatives of MogUGT1216_1 were screened using an enzyme assay to convert MogIVA to MogV. Enzyme assays were performed on MogUGT1216_1 and its variants in 96-well plates at 37°C for 48 hours. MogV was quantified by UPLC-QQQ. Table 47 shows the derivatives that showed beneficial changes in MogV titers compared to MogUGT1216_1: Table 47: Round 2 Engineering of Uridine Diphosphate-Dependent Glycosyltransferases (UGTs) for 1,6 and 1,2 Branched Glycosylation
[0248] Based on these results, MogUGT1216_2 (SEQ ID NO: 38), which contains W160P and H390D substitutions in addition to the mutation present in MogUGT1216_1, was further compared with MogUGT1216_1. Enzymatic assays were performed in 96-well plates at 37°C for 48 hours to convert MogIIE to MogV using both MogUGT1216_0 and MogUGT1216_1. MogV data were quantified using UPLC-QQQ. Figure 11C As shown, MogUGT1216_2 (SEQ ID NO: 38) produces approximately 19 times more MogV levels compared to MogUGT1216_1.
[0249] Derivatives of MogUGT1216_2 were screened from Mog IIE producing strains. The production of downstream mogrosides, including MogV, was quantified, and the total UDP-glucose consumption was determined based on the amount of product produced and the total number of glycosylations catalyzed by the enzyme. Measurements were performed in 96-well plates at 37°C for 72 hours. The conversion of MogIIE to MogV was quantified using UPLC-QQQ. Table 48 shows the derivatives that showed beneficial changes in MogV titers compared to MogUGT1216_2: Table 48: Round 3 Engineering of Uridine Diphosphate-Dependent Glycosyltransferases (UGTs) for 1,6 and 1,2 Branching Glycosylation
[0250] Based on these results, MogUGT1216_3 (SEQ ID NO: 90), which contains H19E, T281S, and S318K substitutions in addition to the mutations present in MogUGT1216_2, was further compared with MogUGT1216_2. Enzymatic assays for the conversion of MogIIE to MogV were performed in 96-well plates at 37°C for 48 hours using MogUGT1216_0 and MogUGT1216_1. MogV data were quantified using UPLC-QQQ. Figure 11D As shown, MogUGT1216_3 (SEQ ID NO: 90) produces approximately 1.6 times more MogV levels compared to MogUGT1216_2.
[0251] Derivatives of MogUGT1216_3 were screened from Mog IIE producing strains. The production of downstream mogrosides, including MogV, was quantified, and the total UDP-glucose consumption was determined based on the amount of product produced and the total number of glycosylations catalyzed by the enzyme. Measurements were performed in 96-well plates at 37°C for 72 hours. The conversion of MogIIE to MogV was quantified using UPLC-QQQ. Table 49 shows derivatives that showed beneficial changes in MogV titers compared to MogUGT1216_2: Table 49: Round 4 of Engineering of Uridine Diphosphate-Dependent Glycosyltransferases (UGTs) for 1,6 and 1,2 Branched Glycosylation
[0252] These results indicate that, compared with MogUGT1216_2, the above mutant derivatives further improved the titer of MogV.
[0253] Example 11: Engineering of coffee tree uridine diphosphate-dependent glycosyltransferase Constructed Coffee TreeDerivatives of uridine diphosphate-dependent glycosyltransferase 1,6 (MogUGT16_0, SEQ ID NO: 39) were screened using an enzyme assay to convert MogIIE to MogV. Enzyme assays were performed on MogUGT16_0 and its variants in 96-well plates at 37°C for 48 hours. MogV was quantified by UPLC-QQQ. Table 50 shows derivatives that showed beneficial changes in MogV titers compared to MogUGT16_0: Table 50: Coffee Tree Round 1 of engineering uridine diphosphate-dependent glycosyltransferases
[0254] Based on these results, MogUGT16_1 (SEQ ID NO: 40) with T147L and N207K substitutions was further investigated and compared with MogUGT16_0. Enzymatic assays for the conversion of MogIIE to MogV were performed in 96-well plates at 37°C for 48 hours using both MogUGT16_0 and MogUGT16_1. MogV data were quantified using UPLC-QQQ. Figure 12 As shown, compared with MogUGT16_0, MogUGT16_1 produces more than 40 times higher levels of MogV.
[0255] Example 12: Strategies to improve the production of 2,3;22,23-squalene dioxide precursor PntAB is a membrane-bound transhydrogenase that produces NADPH in *E. coli*. Since each of the two oxygenation reactions catalyzed by squalene epoxidase requires NADPH as a cofactor (Figure 1), an overexpression of natural *E. coli* was constructed to provide additional NADPH for these reactions. pntA-pntB A strain of 2,3:22,23-squalene dioxide was developed (through gene complementation), and the levels of 2,3:22,23-squalene dioxide produced by this strain were analyzed. In short, fermentation was carried out in 96-well plates at 37°C for 72 hours. The yield of 2,3:22,23-squalene dioxide was quantified by GC-FID. Figure 13A As shown, and not expressed pntAB Compared to strains, overexpression pntAB The strains (SEQ ID NO: 41 and 42) produced approximately 1.4 times more 2,3;22,23-squalene dioxide. These results particularly suggest that the additional NADPH produced by PntAB enabled more squalene to be converted to 2,3-squalene dioxide, and then subsequently to 2,3;22,23-squalene dioxide.
[0256] The next problem to be solved is how to co-express enzymes that catalyze the epoxidation of squalene to 2,3;22,23-squalene dioxide and catalyze the cyclization of 2,3;22,23-squalene dioxide to 24,25-epoxycucurbitadienol, and overexpress native *E. coli*. pntA-pntB Whether the gene increases the yield of 24,25-epoxycucurbitadienol was investigated. Therefore, native *E. coli* strains expressing SQS, SQE, ECDS, and EPH, with or without expression, were constructed and studied. pntA-pntB The *E. coli* strain was used. In short, fermentation was carried out in 96-well plates at 37°C for 72 hours. The yield of 24,25-epoxycucurbitadienol was quantified by GC-FID (24,25-dihydroxycucurbitadienol was measured as a surrogate indicator). Figure 13B As shown, the display and the expression are not shown. pntAB Compared to strains, expression pntAB The strain produced approximately 1.9 times more 24,25-epoxycucurbitadienol. These results particularly indicate that the additional NADPH produced by PntAB enabled more squalene to be converted to 2,3-squalene oxide, and then subsequently to 2,3;22,23-squalene dioxide.
[0257] To further explore the role of supplemental reducing power, strains containing genome integration cassettes were constructed from strains expressing SQS and SQE, and these cassettes were used to overexpress native *Escherichia coli*. fldA (SEQ ID NO: 43) or heterogeneous Cv.fdx (from Wine-colored heterochromatic bacteria The fdx homolog (SEQ ID NO: 44). These strains were compared with Escherichia coli that overexpressed SQS and SQE but did not express them. fldA or Cy.fdx A comparative study was conducted with wild-type strains. In short, fermentation was carried out in 96-well plates at 37°C for 72 hours. The yields of squalene, 2,3-squalene oxide, and 2,3:22,23-squalene dioxide were quantified using GC-FID. Figure 13C As shown, compared with strains that overexpress only SQS and SQE, overexpressing native Escherichia coli... fldA The strains produced approximately 4-fold increased levels of 2,3:22,23-squalene dioxide. Similarly, strains overexpressing heterologous Cv. fdx also produced more 2,3-squalene dioxide and 2,3:22,23-squalene dioxide compared to strains overexpressing only SQS and SQE. These results particularly suggest that providing additional electron carriers (such as ferredoxin or flavin-redoxin) can lead to additional NADPH availability, enabling more squalene to be converted to 2,3-squalene dioxide, and then subsequently to 2,3;22,23-squalene dioxide.
[0258] Example 13: In a strain with reduced expression of YbhGFSR ABC type transporter protein, mogroside and 24,24- Production of dihydroxycucurbitadienol ABC pumps (such as those made by...) ybhGFSR Overexpression of the operon-encoded ABC pump can be used to increase the production of compounds in microbial cells. Therefore, derivatives of strains expressing the complete metabolic pathway for the production of mogroside were constructed and tested.
[0259] For these experiments, parental strains expressing the complete metabolic pathway for mogrool production were modified and evaluated. To evaluate... ybhGFSR The role of the operon is to construct a deletion from the chromosome. ybhGFSR The strain of the operon, and its relationship with the parent. ybhGFSR + The strains were evaluated by comparison. The strains were grown in 96-well plates at 37°C for 72 hours. Samples of the entire culture, including cells and culture supernatant, were extracted and analyzed. The yields of 2,3:22,23-squalene dioxide, 24,25-dihydroxycucurbitacinol, and mogroside were quantified by LC-DAD and normalized to the production of the parent strains. Surprisingly, as... Figure 14A As shown, compared with the parent strain, ybhGFSR The operon deletion strains exhibited increased production of 24,25-dihydroxycucurbitadienol + mogroside. ybhGFSR The operon deletion strain exhibited a reduced yield of 2,3;22,23-squalene dioxide compared to the parental strain. Figure 14A ).
[0260] To further evaluate this phenomenon, the parental strain expressing the complete metabolic pathway for the production of mogroside was modified to produce a sub-efficacy. ybhG Alleles. In short, alleles are genes that carry the genes from chromosomes. ybhG The ribosome binding site is replaced with a weaker form than the natural sequence to produce ybhG Sub-effect alleles. Those with sub-effect... ybhG Allele strains and parents ybhGFSR + The strains were evaluated through comparison. The strains were grown in 96-well plates at 37°C for 72 hours. Samples of the entire culture, including cells and culture supernatant, were extracted and analyzed. The yields of 2,3:22,23-squalene dioxide, 24,25-dihydroxycucurbitacinol, and mogroside were quantified by LC-DAD and normalized to the production levels of the parent strains. Figure 14B As shown, compared with the parent strain, it carries ybhG Strains carrying the hypoepivalent allele exhibited increased production of 24,25-dihydroxycucurbitadienol + mogroside. Compared to the parental strain, those carrying the hypoepivalent allele showed increased production. ybhGThe strains with the hypoepivalent allele also exhibited reduced production of 2,3;22,23-squalene dioxide. Figure 14B ).
[0261] In order to understand ybhGFSR The effect of operon overexpression: Modifying the parental strain expressing the complete metabolic pathway for mogrool production to overexpress... ybhGFSR Operon. In short, a deletion from a chromosome. ybhF Genes, and overexpression from plasmids ybhGFSR Operator. Overexpression ybhGFSR The strain of the operon and the parent ybhGFSR + The strains were evaluated through comparison. The strains were grown in 96-well plates at 37°C for 72 hours. Samples of the entire culture, including cells and culture supernatant, were extracted and analyzed. The yields of 2,3:22,23-squalene dioxide, 24,25-dihydroxycucurbitacinol, and mogroside were quantified by LC-DAD and normalized to the production levels of the parent strains. Figure 14C As shown in the figure above, compared with the parent strain, overexpression ybhGFSR The operon strain exhibited decreased production of 24,25-dihydroxycucurbitadienol + mogroside, while overall production of 2,3:22,23-squalene dioxide, 24,25-dihydroxycucurbitadienol, and mogroside was increased. This confirmed that overexpression compared to the parental strain... ybhGFSR The operon strains exhibited a significantly increased yield of 2,3;22,23-squalene dioxide. Figure 14C (See image below)
[0262] These results prove that, ybhGFSR Sub-effective or null mutations in the operon increase the yield of mogroside and / or its direct precursor 24,25-dihydroxycucurbitadienol, while decreasing the yield of 2,3;22,23-squalene dioxide. These results also demonstrate that overexpression of the ybhGFSR operon decreases the yield of mogroside and / or its direct precursor 24,25-dihydroxycucurbitadienol, while increasing the yield of 2,3;22,23-squalene dioxide. To avoid being bound by theory, it is hypothesized that the YbhGFSRABC type transporter increases the efflux of 2,3;22,23-squalene dioxide, thereby preventing its conversion to downstream products. ybhGFSR Decreased expression of transporter proteins increases intracellular (int.) 2,3;22,23-squalene dioxide and enhances conversion to downstream products 24,25-dihydroxycucurbitadienol and mogroside.
[0263] sequence Squalene epoxidase SQE (SEQ ID NO: 1) MAKEEFDICIIGAGMAGATISAYLAPKGIKIALIDHCYKEKKRIVGELLQPGAVLSLEQMGLSHLLDGFEAQTVKGYALLQGNEKTTIPYPSQHEGIGLHNGRFLQQIRASALENSSVTQIHGKALQLLENERNEIIGVSYRESITSQIKSIYAPLTITSDGFFSNFRAHLSNNQKTVTSYFIGLILKDCEMPFPKHGHVFLSGPTPFICYPISDNEVRLLIDFPGEQLPRKNLLQEHLDTNVTPYIPECMRSSYAQAIQEGGFKVMPNHYMAAKPIVRKGAVMLGDALNMRHPLTGGGLTAVFSDIQILSAHLLAMPDFKNTDLIHEKIEAYYRDRKRANANLNILANALYAVMSNDLLKTAVFKYLQCGGANAQESIAVLAGLNRKHFSLIKQFCFLAVFGACNLLQQSISNIPKALKLLKDAFVIIKPLIKNELS SQE_1 (SEQ ID NO: 2) MAKEEFDICIIGAGMAGATISAYLAPKGIKIALIDRCYKEKKRIVGELLQPGAVLSLEQMGLSHLLDGFEAQTVKGYALLQGNEKTTIPYPSQHEGIGLHNGRFLQQIRASALENSSVTQIHGKALQLLENERNEIIGVSYRESITSQIKSIYAPLTITSDGFASNFRAHLSNNQKTVTSYFIGLILKDCEMPFPKHGHVFLSGPTPFICYPISDNEVRLLIDFPGEQLPRKNLLQEHLDTNVTPYIPECMRSSYAQAIQEGGFKVMPNHYMAAKPIVRKGAVLLGDALNMRHPLTGGGLTAVFSDIQILSAHLLAMPDFKNTDLIHEKIEAYYRDRKRANANLNILANALYAVMSNDLLKTAVFKYLQCGGANAQESIALLAGLNRKHFSLIKQYCFLAVFGACNLLQQSISNIPKALKLLKDAFVIIKPLIKNELS SQE_2 (SEQ ID NO: 3) MAKEEFDICIIGAGMAGATISAYLAPKGIKIALIDRCYKEKKRIVGELLQPGAVLSLEQMGLSHLLDGFEAQTVKGYALLQGNEKTTIPYPSQHEGIGLHNGRFLQQIRASALENESVTQIHGKALQLLENERGEIIGVSYRESHTSQIKSIYAPLTITSDGFASNFRRHLSNNQKTVTSYFIGLILKDCEMPFPKHGHVFLSGPTPFICYPISDNEVRLLIDFPGEQLPRKNLLQEHLDTNVTPYIPECMRSSYAQAIQEGGFKVMPNHYMAAKPIVRKGAVLLGDALNMRHPLTGGGLTAVFSDIQILSAHLLAMPDFKNTDLIHEKIEAYYRDRKRANANLNILANALYAVMSNDLLKTAVFKYLQCGGANAQESIALLAGLNRKHFSLIKQYCFLVVFGACNLLQQSISNIPKALKLLKDAFVIIKPLIKNELS SQE_4 (SEQ ID NO: 4) MAKEEFDICIIGAGMAGATISAYLAPKGIKIALIDRCYKEKKRIVGELLQPGAVLSLEQMGLSHLLDGFEAQTVKGYALLQGNEKTTIPYPSQHEGIGLHNGRFLQQIRASALENESVTQIHGKALQLLENERGEIIGVSYRESHTSQIKSIYAPLTITSDGFNSNFRRHLSNNQKTVTSYFIGLILKDCEMPFPKHGHVFLSGPTPFICYPISDNEVRLLIDFPGEQLPRKNLLQEHLDTNVTPYIPECMRSSYAQAIQEGGFKVMPNHYMAAKPIVRKGAVLLGDALNMRHPLTGGGLTAVFSDIQILSAHLLAMPDFKNTDLIHEKIEAYYRDRKRANANLNILANALYAVMSNDLLKTAVFKYLQCGGANAQESIALLAGLNRKHFSLIKQYCFLVVFGACNLLQQSISNIPKALKLLKDAFVIIKPLIKNELS Capsular methylcoccus SQE (McSQE_0) (SEQ ID NO: 77) MASSIELGEWDVLIAGGSVAGSAAAAALSGLGLRVLIVEPDPDPGRRLAGELIHPPGIDGLLELGLIHDDVPQGSVVNGFAIFPFNDGEGAPATLLPYGEIHGRQRCGRVIEHTLLKSHLLETVRGFERVSVWLGARVTGMEHEDGKGYVATVTHEGTETRMEVRLIIGADGPMSQLRKMVGISHETQRYSGMIGLEVEDTHLPNPGYGNIFLNPAGVSYAYGIGGGRARVMFEVLKGADSKESIRDHLRLFPAPFRGDIEAVLAQGKPLAAANYCIVPEASVKANVALVGDARGCCHPLTASGITAAVKDAFVMRDALQATGLNFEAALKRYSVQCGRLQLTRRTLAEELREAFLAQTPEAELLSQCIFSYWRNSPKGRQASMALLSTLDSSIFSLASQYTLVGLQAFRLLPQWLGAKMGGDWFRGVAQLVSKSLKFQQDALNQALRAK McSQE_1 (SEQ ID NO: 78) MASSIELGEWDVLIAGGSVAGSAAAAALSGLGLRVLIVEPDPDPGRRLAGELIHPPGIDGLLELGLIHDDVPQGSVVNGFAIFPFNDGEGAPATLLPYGEIHGRQRCGRVIEHTLLKSHLLETVRGFERVSVWLGARVTGMEHEDGKGYVATVTHEGTETRMEVRLIIGADGPMSQLRKMVGISHETQRYSGMIGLEVEDTHLPNPGYGNIFLNPAGVSYAYGIGGGRARVMFEVLKGADSKESIRDHLRLFPAPFRGDIEAVLAQGKPLAAANYCIVPEASVKANVALVGDARGCCHPLTASGITAAVKDAFVMRDALQATGLNFEAALKRYSVQCGRLQLTRRTLAEELREAFLAQTPEAELLRQCIFSYWRNSPKGRQASMALLSTLDSSIFSLASQYTLVGLQAFRLLPQWLGAKMGGDWFRGVAQLVSKSLKFQQDALNQALRAK McSQE_2 (SEQ ID NO: 79) MASSIELGEWDVLIAGGSVAGSAAAAALSGLGLRVLIVEPHPDPGRRLAGELIHPPGIDGLLELGLIHDDVPQGSVVNGFAIFPFNDGEGAPATLLPYGEIHGRQRCGRVIEHTLLKSHLLETVRGFERVSVWLGARVTGMEHEDGKGYVATVTHEGTETRMEVRLIIGADGPMSQLRKMVGISHETQRYSGMIGLEVEDTHLPNPGYGNIFLNPAGVSYAYGIGGGRARVMFEVLKGADSKESIRDHLRLFPAPFRGDIEAVLAQGKPLAAANYCIVPEASVKANVALVGDARGCCHPLTASGITAAVKDAFVLRDALQATGLNFEAALKRYSVQCGRLQLTRRTLAEELREAFLAQTPEAELLRQCIFSYWRNSPKGRQASMALLSTLDSSIFSLASQYTLVGLQAFRLLPQWLGAKMGGDWFRGVAQLVSKSLKFQQDALNQALRAK Epoxide hydrolase Monk fruit EPH3 (EPH_0) (SEQ ID NO: 5) MADQIEHITINTNGIKMHIASVGTGPVVLLLHGFPELWYSWRHQLLYLSSVGYRAIAPDLRGYGDTDSPASPTSYTALHIVGDLVGALDELGIEKVFLVGHDWGAIIAWYFCLFRPDRIKALVNLSVQFIPRNPAIPFIEGFRTAFGDDFYMCRFQVPGEAEEDFASIDTAQLFKTSLCNRSSAPPCLPKEIGFRAIPPPENLPSWLTEEDINYYAAKFKQTGFTGALNYYRAFDLTWELTAPWTGAQIQVPVKFIVGDSDLTYHFPGAKEYIHNGGFKKDVPLLEEVVVVKDACHFINQERPQEINAHIHDFINKF EPH_1 (SEQ ID NO: 6) MADQIEHITINTNGIKMHIASVGTGPVVLLLHGFPELWYSWRHQLLYLSSVGYRAIAPDLRGYGDTDAPASPTSYTALHIVGDLVGALDELGIEKVFLVGHDWGAIIAWYFCLFRPDRIKALVNLSVAFIPRNPAIPFIESFRQVFGDDFYMCRFQVPGEAEADFASIDTAQLFKTSLCNRSSAPPCLPKGIGFRAIPPPENLPSWLTEEDINYYAAKFKQTGFTGALNYYRAFDLTWELTAPWTGAQIQVPVKFIVGDSDMTYHFPGAKEYIHNGGFKKDVPLLEEVVVVKDAGHFIQQERPQEINAHIHDFINKF Cucurbita dienol synthase (CDS) Citrullus colocynthis (CcCDS2; ECDS_0) (SEQ ID NO: 7) MAWRLKVGAESVGEKEEKWLKSISNHLGRQVWEFCAHQPTASPNHLQQIDNARNHFRNNRFHRKQSSDLFLAIQNEKEIANVTKGGGIKVKEEEDVRKETVKNTVERALSFYSAIQTNDGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVALRLLGEDADGGEGGAMTKARSWILDRGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYCLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYTIPYHEIDWNRSRNTCAKEDLYYPHPKMQDILWGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYICLGPVNKVLNMLCCWVEDPYSDAFKFHLQRVPDYLWVAEDGMRMQGYNGSQLWDTAFSVQAIISTKLIDSFGTTLKKAHDFVKDSQIQQDCPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSKIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTSATMEALTLFKKLHPGHRTKEIDIAVARAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE ECDS_1 (SEQ ID NO: 8) MAWRLKVGAESVGEKEEKWLKSINNHLGRQVWEFDAHQPTASPNHLQQIDNARNHFRNNRFHRKQSSDLFLAIQNEKEIANVTKGGGIKVKEEEDVRKETVKNTVERALSFYSAIQTNDGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVALRLLGEDADGGEGGAMTKARSWILDRGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYCLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYTIPYHEIDWNRSRNTCAKEDLYYPHPKMQDILWGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYICLGPVNKVLNMLCCWVEDPYSDAFKFHLQRVPDYLWVAEDGMRMQGYNGSQLWDTAFSVQAIISTKLIDSFGTTLKKAHDFVKDSQIQQDCPGDPNVWFRHIHKGAWPFSTRDHGWLISDCTAEGLKASLMLSKLPSKIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELINPSETFGDIVIDYPYVECTSATMEALTLFKKLHPGHRTKEIDIAVARAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE ECDS_2 (SEQ ID NO: 9) MAWRLKVGAESVGEKEEKWLKSINNHLGRQVWEFDAHQPTASPNHLQQIDNARNHFRNNRFHRKQSSDLFLAIQNEKEIANVTKGGGIKVKEEEDVRKETVKNTVERALSFYSAIQTNDGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVALRLLGEDADGGEGGAMTKARSWILDRGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYCLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYTIPYHEIDWNRSRNTCAKEDLYYPHPKMQDILWGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYICLGPVNKVLNMLCCWVEDPYSDAFKFHLQRVPDYLWVAEDGMRMQGYNGSQLWDTAFSVQAIISTKLIDSFGTTLKKAHDFVKDSQIQQDCPGDPNVWFRHIHKGAWPFSTRDHGWLVSDCTAEGLKASLMLSKLPSKIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELMNPSETFGDIVIDYPYVECTSATMEALTLFKKLHPGHRTKEIDIAVARAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE ECDS_3 (SEQ ID NO: 10) MAWRLKVGAESVGEKEEKWLKSINNHLGRQVWEFDAHQPTASPNHLQQIENARNHFRNNRFHRKQSSDLFLAIQNEKEIANVTKGGGIKVKEEEDVRKETVKNTVERALSFYSAIQTNDGHWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVALRLLGEDADGGEGGAMTKARSWILDRGGATAITSWGKLWLSVLGVYEWSGNNPLPPEFWLLPYCLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYTIPYHEIDWNRSRNTCAKEDLYYPHPKMQDILWGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYICLGPVNKVLNMLCCWVEDPYSDAFKFHLARVPDYLWVAEDGMRMQGYNGSQLWDTAFSVQAIISTKLIDSFGTTLKKAHDFVKDSQIQQDCPGDPNVWFRHIHKGAWPFSTRDHGWLVSDCTAEGLKASLMLSKLPSKIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELMNPSETFGDIVIDYPYVECTSATMEALTLFKKLHPGHRTKEIDIAVARAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE ECDS_4 (SEQ ID NO: 11) MAWRLKVGAESVGEKEEKWLKSINNHLGRQVWEFDAHQPTASPNHLQQIENARNHFRNNRFHRKQSSDLFLAIQNEKEIANVTKGGGIKVKEEEDVRKETVKNTVERALSFYSAIQTNDGHWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVALRLLGEDADGGEGGAMTKARSWILDRGGATAITSWGKLWLSVLGVYEWSGNNPIPPEFWLLPYCLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYTIPYHEIDWNRSRNTCAKEDLYYPHPKMQDILWGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYICLGPVNKVLNMLCCWVEDPYSDAFKFHLARVPDYLWVAEDGMRMQGYNGSQLWDTAFSVQAIISTKLIDSFGTTLKKAHDFVKDSQIQQDCPGDPNVWFRHIHKGAWPFSTRDHGWLVSDCTAEGLKASLMLSKLPSKIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELMNPSETFGDIVIDYPYVECTSATMEALTLFKKLHPGHRTKEIDIAVARAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE ECDS_5 (SEQ ID NO: 12) MAWRLKVGAESVGEKEEKWLKSINNHLGRQVWEFDAHQPTASPNHLQQIENARNHFRNNRFHRKQSSDLFLAIQNEKEIANVTKGGGIKVKEEEDVRKETVKNTVERALSFYSAIQTNDGHWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYLYNHQNEDGGWGLHIEGTSTMFGSALNYVALRLLGEDADGGEGGAMTKARSWILDRGGATAITSWGKLWLSVLGVYEWSGNNPIPPEFWLLPYCLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYTIPYHEIDWNRSRNTCAKEDLYYPHPKMQDILSGSIYHLYEPLFTRWPGKRLREKALQMAMKHIHYEDENSRYICLGPVNKVLNMLCCWVEDPYSDAFKFHLARVPDYLWVAEDGMRMQGYNGSQLWDTAFSVQAIISTKLIDSFGTTLKKAHDFVKDSQIQQDCPGDPNVWFRHIHKGAWPFSTRDHGWLVSDCTAEGLKASLMLSKLPSKIVGEPLEKSRLCDAVNVLLSLQNENGGFASYELTRSYPWLELMNPSETFGDIVIDYPYVECTSATMEALTLFKKLHPGHRTKEIDIAVARAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNSCVAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPAPLHRAARLLINSQLENGDFPQEEIMGVFNKNCMITYAAYRNIFPIWALGEYFHRVLTE SgCDS (SEQ ID NO: 68) MAWRLKVGAESVGENDEKWLKSISNHLGRQVWEFCPDAGTQQQLLQVHKARKAFHDDRFHRKQSSDLFITIQYGKEVENGGKTAGVKLKEGEEVRKEAVESSLERALSFYSSIQTSDGNWASDLGGPMFLLPGLVIALYVTGVLNSVLSKHHRQEMCRYVYNHQNEDGGWGLHIEGPSTMFGSALNYVALRLLGEDANAGAMPKARAWILDHGGATGITSWGKLWLSVLGVYEWSGNNPLPPEFWLFPYFLPFHPGRMWCHCRMVYLPMSYLYGKRFVGPITPIVLSLRKELYAVPYHEIDWNKSRNTCAKEDLYYPHPKMQDILWGSLHHVYEPLFTRWPAKRLREKALQTAMQHIHYEDENTRYICLGPVNKVLNLLCCWVEDPYSDAFKLHLQRVHDYLWVAEDGMKMQGYNGSQLWDTAFSIQAIVSTKLVDNYGPTLRKAHDFVKSSQIQQDCPGDPNVWYRHIHKGAWPFSTRDHGWLISDCTAEGLKAALMLSKLPSETVGESLERNRLCDAVNVLLSLQNDNGGFASYELTRSYPWLELINPAETFGDIVIDYPYVECTSATMEALTLFKKLHPGHRTKEIDTAIVRAANFLENMQRTDGSWYGCWGVCFTYAGWFGIKGLVAAGRTYNNCLAIRKACDFLLSKELPGGGWGESYLSCQNKVYTNLEGNRPHLVNTAWVLMALIEAGQAERDPTPLHRAARLLINSQLENGDFPQQEIMGVFNKNCMITYAAYRNIFPIWALGEYCHRVLTE PsCAS_mut (SEQ ID NO: 69) MWKLKVAEGGTPWLRTLNNHVGRQVWEFDPHSGSPQDLDDIETARRNFHDNRFTHKHSDDLLMRLQFAKENPMNEVLPKVKVKDVEDVTEEAVATTLRRGLNFYSTIQSHDGHWPGDLGGPMFLMPGLVITLSVTGALNAVLTDEHRKEMRRYLYNHQNKDGGWGLHIEGPSTMFGSVLCYVTLRLLGEGPNDGEGDMERGRDWILEHGGATYITSWGKMWLSVLGVFEWSGNNPMPPEIWLLPYALPVHPGRMWCHCRMVYLPMSYLYGKRFVGPITPTVLSLRKELFTVPYHDIDWNQARNLCAKEDLYYPHPLVQDILWATLHKFVEPVFMNWPGKKLREKAIKTAIEHIHYEDENTRYICIGPVNKVLNMLCCWVEDPNSEAFKLHLPRIYDYLWVAEDGMKMQGYNGSQLWDTAFAAQAIISTNLIDEFGPTLKKAHAFIKNSQVSEDCPGDLSKWYRHISKGAWPFSTADHGWPISDCTAEGLKAVLLLSKIAPEIVGEPLDSKRLYDAVNVILSLQNENGGLATYELTRSYTWLEIINPAETFGDIVIDCPYVECTSAAIQALATFGKLYPGHRREEIQCCIEKAVAFIEKIQASDGSWYGSWGVCFTYGTWFGIKGLIAAGKNFSNCLSIRKACEFLLSKQLPSGGWAESYLSCQNKVYSNLEGNRSHVVNTGWAMLALIEAEQAKRDPTPLHRAAVCLINSQLENGDFPQEEIMGVFNKNCMITYAAYRCIFPIWALGEYRRVLQAC CpCAS_mut (SEQ ID NO: 70) MAWQLKIGADTVPSDPSNAGGWLSTLNNHVGRQVWHFHPELGSPEDLQQIQQARQHFSDHRFEKKHSADLLMRMQFAKENSSFVNLPQVKVKDKEDVTEEAVTRTLRRAINFYSTIQADDGHWPGDLGGPMFLIPGLVITLSITGALNAVLSTEHQREICRYLYNHQNKDGGWGLHIEGPSTMFGSVLNYVTLRLLGEEAEDGQGAVDKARKWILDHGGAAAITSWGKMWLSVLGVYEWAGNNPLPPELWLLPYLLPCHPGRMWCHCRMVYLPMCYLYGKRFVGPITPIIRSLRKELYLVPYHEVDWNKARNQCAKEDLYYPHPLVQDILWATLHHVYEPLFMHWPAKRLREKALQSVMQHIHYEDENTRYICIGPVNKVLNMLCCWAEDPHSEAFKLHIPRIYDYLWIAEDGMKMQGYNGSQLWDTAFAVQAIISTELAEEYETTLRKAHKYIKDSQVLEDCPGDLQSWYRHISKGAWPFSTADHGWPISDCTAEGLKAVLLLSKLPSEIVGKSIDEQQLYNAVNVILSLQNTDGGFATYELTRSYRWLELMNPAETFGDIVIDYPYVECSSAAIQALAAFKKLYPGHRRDEIDNCIAEAADFIESIQATDGSWYGSWGVCFTYGGWFGIRGLVAAGRRYNNCSSLRKACDFLLSKELAAGGWGESYLSCQNKVYTNIKDDRPHIVNTGWAMLSLIDAGQSERDPTPLHRAARVLINSQMEDGDFPQEEIMGVFNKNCMISYSAYRNIFPIWALGEYRSRVLKPLK AaCAS_mut (SEQ ID NO: 71) MWKLKIAEGGDPWLRTTNDHIGRQIWEFDPTLGSVEELAEIEKLRKTFRDNRFEKKHSADLLMRSQFAKENSVSVFPPKVNIKDVEDITEDKVTNVLRRAIGFHSTLQADDGHWPGDLGGPMFLLPGLVITLSITGALNAVLSKEHKREMCRYLYNHQNIDGGWGLHIEGHSTMFGSALNYVTLRLLGEGANDGEGAMEKGRKWILDHGGATAITSWGKFWLSVLGVFEWPGNNPLPPEMWLLPYFLPVHPGRMWCHCRMVYLPMSYLYGKRFVGPITSTVLALRKELFTVPYHDIDWNEARNLCAKEDLYYPHPLIQDVLWATLDKFVEPVLMSWPGKKLREKALRTAMEHIHYEDENTRYICIGPVNKVLNMLCCWVEDPNSEAFKLHLPRIQDYLWIAEDGMKMQGYNGSQLWDAAFTVQAIMSTNLIEEFGPTLKKGHIFIKKSQVLDNCYGDLDYWYRHISKGAWPFSTADHGWPISDCTAEGLKAALLLSKLPSEIVDEPLDAKRFYDAVNVILSLMNADGSFATYELTRSYSWLELINPAETFGDIVIDYPYVECTSAAIQALVAFKRLYPGHRRDEVQGCIDKAAAFLEKIQEADGSWYGSWAVCFTYGTWFGVKGLVAAGKNYSNCSSIRKACNFLLSKQLASGGWGESYLSCVDKVYTNLEGNRSHVVNTGWAMLALIDAEQAKRDPTPLHRAARVLINSQMENGEFPQQEIMGVFNRNCMITYAAYRNIFPIWALGEYRCRVLKVET Squalene synthase (SQS) Artemisia annua SQS (SEQ ID NO: 13) MASSLKAVLKHPDDFYPLLKLKMAAKKAEKQIPSQPHWAFSYSMLHKVSRSFALVIQQLNPQLRDAVCIFYLVLRALDTVEDDTSIAADIKVPILIAFHKHIYNRDWHFACGTKEYKVLMDQFHHVSTAFLELKRGYQEAIEDITMRMGAGMAKFICKEVETVDDYDEYCHYVAGLVGIGLSKLFHSSGTEILFSDSISNSMGLFLQKTNIIRDYLEDINEIPKSRMFWPREIWSKYVNKLEDLKYEENSEKAVQCLNDMVTNALIHIEDCLKYMSQLKDPAIFRFCAIPQIMAIGTLALCYNNIEVFRGVVKLRRGLTAKVIDRTKTMADVYQAFSDFSDMLKSKVDMHDPNAQTTITRLEAAQKICKDSGTLSNRKSYIVKRESSYSAALLALLFTILAILYAYLSANRPNKIKFTL Cytochrome P450 pepo subspecies of zucchini Cytochrome P450 enzyme CYP87A3 (SEQ ID NO: 14) MAWAIVVGLATLAVAYYIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTSMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYNKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARAHILSFEDGLHVKFTPKE sohB_CppCYP87A3 (SEQ ID NO: 15) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTSMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYNKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARAHILSFEDGLHVKFTPKE sohB_C11CYP_1 (SEQ ID NO: 16) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTEWLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYNKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARAHILSFEDGLHVKFTPKE SohB_C11CYP_2 (SEQ ID NO: 17) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTENLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRHCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARTHILSFEDGLHVKFTPKE sohB_C11CYP_3 (SEQ ID NO: 18) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFGLDTENLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARTHILSFEDGLHVKFTPKE- sohB_C11CYP_4 (SEQ ID NO: 19) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFALDGENLKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARTHILSFEDGLHVKFTPKE sohB_C11CYP_5 (SEQ ID NO: 20) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFALDGENVKALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKVARTHILSFEDGLHVKFTPKE sohB_C11CYP_6 (SEQ ID NO: 21) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLLGGFLTLPINVPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKIARTHILSFEDGLHVKFTPKE sohB_C11CYP_8 (SEQ ID NO: 45) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPSDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNNYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLTLPINVPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKIARTHILSFEDGLHVKFTPKE sohB_C11CYP_9 (SEQ ID NO: 46) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPFDSLDVHPFIKKKVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLTLPINVPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIFETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKIARTHILSFEDGLHVKFTPKE sohB_C11CYP_10 (SEQ ID NO: 47) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPNDSLDVHPFIKKRVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLTLPINVPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRKTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKIARTHILSFEDGLHVKFTPKE sohB_C11CYP_11 (SEQ ID NO: 80) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPNDSLDVHPFIKKRVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLTLPINVPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKIARTHILSFEDGLHVKFTPKE sohB_C11CYP_12 (SEQ ID NO: 81) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPNDSLDVHPFIKKRVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLSLPINIPGTTYHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKIARTHILSFEDGLHVKFTPKE sohB_C11CYP_13 (SEQ ID NO: 82) MALLSEYGLFLAKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPNDSLDVHPFIKKRVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRTTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLSLPINIPGFAFHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILFTKYRWTKLKGGKIARTHILSFEDGLHVKFTPKE sohB_C11CYP_14 (SEQ ID NO: 83) MALLSEYGLFLWKIVTVVLAIAAIAAIIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPNDSLDVHPFIKKRVKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRLTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLSLPINIPGFAFHKCMKDMKEIQKKLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILITKYRWTKLKGGKIARTHILSFEDGLHVKFTPKE sohB_C11CYP_15 (SEQ ID NO: 84) MALLSEYGLFLWKIVTVVLAIAAIAAGIHWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPNDSLDVHPFIKKRMKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRLTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLSLPINIPGFAFHKCMKDMKEIQKVLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILITKYRWTKLKGGKIARTHILSFEDGLHVKFSPKE n20_C11CYP_15 (SEQ ID NO: 85) MAWINKWKDSKFNGVLPPGTMGLPLVGETLQLARPNDSLDVHPFIKKRMKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRLTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLSLPINIPGFAFHKCMKDMKEIQKVLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILITKYRWTKLKGGKIARTHILSFEDGLHVKFSPKE n20_C11CYP_16 (SEQ ID NO: 86) MAWINKWRNRKFKGKLPPGTMGLPLVGETLQLARPNDSLDVHPFIKKRMKRYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRLTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLSLPINIPGFAFHKCMKDMKEIQKVLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILITKYRWTKLKGGKIARTHILSFEDGLHVKFSPKE n20_C11CYP_17 (SEQ ID NO: 87) MAWINKWRNRKFKGKLPPGTMGLPLVGETLQLARPNDSLDVHPFVRERVEKYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRLTMVKMVSEDSSKLLTGGLTKKFTGLIAGFLSLPINIPGFAFHKCMKDMKEIQKVLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILITKYRWTKLKGGKIARTHILSFEDGLHVKFSPKE n20_C11CYP_18 (SEQ ID NO: 88) MAWINKWRNRKFKGKLPPGTMGLPLVGETLQLARPNDSLDVHPFVRERVEKYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRLTMVKMVSEDSSKLLTGGLTKKFTGLLKGFLSLPINIPGFAFHKCMKDMKEIQKVLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRFCAGAEYSKVYLCTFLHILITKYRWTKLKGGKIARTHILSFEDGLHVKFSPKE n20_C11CYP_19 (SEQ ID NO: 89) MAWINKWRNRKFKGKLPPGTMGLPLVGETLQLARPNDSLDVHPFVRERVEKYGPIFKTCLAGRPVVVSTDAEFNHYIMLQEGRAVEMWYLDTLSKFFALDGENLNALGFIHKYIRSITLNHFGAESLRERFLPRIEESAKETLRYWATQPSVEVKDSAAVMVFRLTMVKMVSEDSSKLLTGGLTKKFTGLLKGFLSLPINIPGFAFHKCMKDMKEIQKVLREILEGRLASGGGSDEDFLGQAIKDKGSQQFISDDFIIQLLFSISFASFESISTTLTLVLNYLADHPDVVKELEAEHEAIRNARADPDGPITWEEYKSMTFTLHVIYETLRLGSVTPALLRRTTKELQINGYTIPEGWTVMLVTASRHRDPAVYKDPHTFNPWRWKELDSITIQKNFMPFGGGLRYCAGADFAKVYLCTFLHILITKYRWTKLKGGKIARTHILSFEDGLHVKFSPKE Cytochrome P450 reductase pepo subspecies of zucchini Cytochrome P450 reductase CppCPR4 or MogCPR1_0 (SEQ ID NO: 22) MAQSESRSMKVSPLELMSAIIRKAMDPSQDSSESVREVATLILENREFVMILTTSIAVLIGCVVVLVWKRSSDQKAKSFEPPKQLIVKKPEPEVDDGKKKVTVFFGTQTGTAEGFAKALAEEAKARYEKATFRVVDLDDYAADDDEYEEKLKKETLAIFFLATYGDGEPTDNAARFYKWFSEGKEKGEWISNLQYAVFGLGNRQYEHFNKIAKVVDEQLAEQGGKRLVPVGLGDDDQCIEDDFSAWREALWPELDKLLREEDEFTTVSTPYTAAVLEYRVVFYDAADVSGGDKKWAFANGHAVYDIQHPCRANVAVRKELHTSASDRSCTHLEFDISGTGLTYETGDHVGVFCE NLDEVVEEAIRLIGLSPETYFSIHTDKEDGTPLSGSSLPPPFAPCTLRTALTQYADLLSSPKKSALVALAAHASDPAEADRLRHLSLPGKDEYSQWIVASQRSLLEVMAEFPSARPPLGVFFAAVAPRLQPRYYISSSSPRMAPSRIHVTCALVYDKTPTGRHIHKGLCSTWMKNAI PLEESQACSWAPIYVRQSNFKLPTDSKLPIIMIGPGTGLAPFRGFLQERLALKESGVELGHSILFFGCRNRKMDYIYEDELNNFVETGALSELIFASREGPSKEYVQHKMVEKASDIWNLLSQGAYYYVCGDAKGMARDVHRTLNHIVQEQGSLDSSKAESMVKNLQMTGRYLRDVW MogCPR1_1 (SEQ ID NO: 23) MANTLIPILVAICLFITGVAILNIQWKRSSDQKAKSFEPPKQLIVKKPEPEVDDGKKKVTVFFGTQTGTAEGFAKALAEEAKARYEKATFRVVDLDDYAADDDEYEEKLKKETLAIFFLATYGDGEPTDNAARFYKWFSEGKEKGEWISNLQYAVFGLGNRQYEHFNKIAKVVDEQLAEQGGKRLVPVGLGDDDQCIEDDFSAWREALWPELDKLLREEDEFTTVSTPYTAAVLEYRVVFYDAADVSGGDKKWAFANGHAVYDIQHPCRANVAVRKELHTSASDRSCTHLEFDISGTGLTYETGDHVGVFCENLDEVVEEAIRLIGLSPETYFSIHTDKEDGTPLSGSSLPPPFAPCTLRTALTQYADLLSSPKKSALVALAAHASDPAEADRLRHLSLPAGKDEYSQWIVASQRSLLEVMAEFPSARPPLGVFFAAVAPRLQPRYYSISSSPRMAPSRIHVTCALVYDKTPTGRIHKGLCSTWMKNAIPLEESQACSWAPIYVRQSNFKLPTDSKLPIIMIGPGTGLAPFRGFLQERLALKESGVELGHSILFFGCRNRKMDYIYEDELNNFVETGALSELIVAFSREGPSKEYVQHKMVEKASDIWNLLSQGAYIYVCGDAKGMARDVHRTLHNIVQEQGSLDSSKAESMVKNLQMTGRYLRDVW kidney bean Cytochrome P450 reductase (PvCPR or MogCPR2_0) (SEQ ID NO: 24) MAQSSSSSSSSSSSSMSPFDLMAAIIKGKKVLDPSNVSSDSSVSEVANIIFENREFVMILTTSIAVLIGCVVVLIWRRSSGQKVKPVEPLKPLTVKEPEVEVDDGKQKVTIFFGTQTGTAEGFAKALADEAKARYEKAKFRVVDLDDYAADDDEYEEKLKKESLALFFLATYGDGEPTDNAARFYKWFTEGKERGEWLQNLKYGVFGLGNRQYEHFNKVAKVVDDTLIEQGAKRLVPVGLGDDDQCIEDDFTAWREMLWPELDQLLRDEDDSTTVSTPYTAAISEYRVVFYDPADAPLEDKSWGNANGHAVHDAQHPCRSNVAVRKELHTPQSDRSCTHLEFDIAGTGLSYETGDHVG VYCENLIETVEEALKLLGLSPDTYFSHISDKEDGTPLGGSSLPPTFPPCTLRTALTKYADLLSSPKKSALLALAAHASDPTEADRLRYLASPAGKDEYAQWIVASQRSLLEVMAEFPSAKPPLGVFFAAVVPRLQPRYYSISSSPRMAPSRIHVTCALVYEKTPAGRHIHKGLCSTWMKN CVPLEKSSDCSWAPIFVRQSNFKLPTDPEVPVIMIGPGTGLAPFRGFLQERFAMKEDGVELGPSILFFGCRNRQMDYYEDELNNFVQSGALSELVVAFSREGPTKEYVQHKMMEKASDIWNMISQGGYLYVCGDAKGMARDVHRTLHTIVQEQGSLDSSKAESMVKNLQMTGRYLRDVW MogCPR2_1 (SEQ ID NO: 25) MANTLIPILVAICLFITGVAILNIQWRRSSGQKVKPVEPLKPLTVKEPEVEVDDGKQKVTIFFGTQTGTAEGFAKALADEAKARYEKAKFRVVDLDDYAADDDEYEEKLKKSALLFLATYGDGEPTDNAARFYKWFTEGKERGEWLQNLKYGVFGGLNRQYEHF NKVAKVVDDTLIEQGAKRLVVPGLGDDDQCIEDDFTAWREMLWPELDQLLRDEDDSTTVSTPYTAAISEYRVVFYDPADAPLEDKSWGNANGHAVHDAQHPCRSNVAVRKELHTPQSDRSCTHLEFDIAGTGLSYETGDHVGVYCENLIETVEEALKLLGLSPDTYF SIHSDKEDGTPLGGSSLPPTFPPCTLRTKYADLLSSPKKSALLALAAHASDPTEADRLRYLASPAGKDEYAQWIVASQRSLLEVMAEFPPSAKPPLGVFFAAVVPRLQPRYYISSSPRMAPSRIHVTCALVYEKTPAGRIHKGLCSTWMKNCVPLEKSSDCSWA PIFVRQSNFKLPTDPEVPVIMIGPGGTGLAPFRGFLQERFAMKEDGVELGPSILFFGCRNRQMDYYEDELNNFVQSGALSELVVAFSREGPTKEYVQHKMMEKASDIWNMISQGGYLYVCGDAKGMARDVHRTLHTIVQEQGSLDSSKAESMVKNLQMTGRYLRDVW MogCPR2_2 (SEQ ID NO: 26) MAQSSSSSSSSSSSSMSPFDLMAAIIKGKKVLDPSNVSSDSSVSEVANIIFENREFVMILTTSIAVLIGCVVVLIWRRSSGQKVKPVEPPKPLTVKEPEVEVDDGKQKVTIFFGTQTGTAEGFAKALADEAKARYEKAKFRVVDLDDYAADDDEYEEKLKKESLALFFLATYGDGEPTDNAARFYKWFTEGKERGEWLQNLKYGVFGLGNRQYEHFNKVAKVVDDTLIEQGAKRLVPVGLGDDDQCIEDDFTAWREMLWPELDQLLRDEDDSTTVSTPYTAAISEYRVVFYDPADAPLEDKSWGNANGHAVHDAQHPCRSNVAVRKELHTPQSDRSCTHLEFDIAGTGLSYETGDHVG VYCENLIETVEEALKLLGLSPDTYFSHISDKEDGTPLGGSSLPPTFPPCTLRTALTKYADLLSSPKKSALLALAAHASDPTEADRLRYLASPAGKDEYAQWIVASQRSLLEVMAEFPSAKPPLGVFFAAVVPRLQPRYYSISSSPRMAPSRIHVTCALVYEKTPAGRHIHKGLCSTWMKN CVPLEKSSDCSWAPIFVRQSNFKLPTDPEVPVIMIGPGTGLAPFRGFLQERFAMKEDGVELGPSILFFGCRNRQMDYYEDELNNFVQSGALSELVVAFSREGPTKEYVQHKMMEKASDIWNMISQGGYLYVCGDAKGMARDVHRTLHTIVQEQGSLDSSKAESMVKNLQMTGRYLRDVW Urate diphosphate-dependent glycosyltransferase (UGT) AtUGT73C3 (SEQ ID NO: 50) MATEKTHQFHPSLHFVLFPFMAQGHMIPMIDIARLLAQRGVTITIVTTPHNAARFKNVLNRAIESGLAINILHVKFPYQEFGLPEGKENIDSLDSTELMVPFFKAVNLLEDPVMKLMEEMKPRPSCLISDWCLPYTSIIAKNFNIPKIVFHGMGCFNLLCMHVLRRNLEILENVKSDEEYFLVPSFPDRVEFTKLQLPVKANASGDWKEIMDEMVKAEYTSYGVIVNTFQELEPPYVKDYKEAMDGKVWSIGPVSLCNKAGADKAERGSKAAIDQDECLQWLDSKEEGSVLYVCLGSICNLPLSQLKELGLGLEESRRSFIWVIRGSEKYKELFEWMLESGFEERIKERGLLIKGWAPQVLILSHPSVGGFLTHCGWNSTLEGITSGIPLITWPLFGDQFCNQKLVVQVLKAGVSAGVEEVMKWGEEDKIGVLVDKEGVKKAVEELMGDSDDAKERRRRVKELGELAHKAVEKGGSSHSNITLLLQDIMQLAQFKN SgUGT720_269_1_E (SEQ ID NO: 51) MAVQPRVLLFPFPALGHVKPFLSLAELLSDAGIDVVFLSTEYNHRRISNTEALASRFPTLHFETIPDGLPPNESRALADGPLYFSMREGTKPRFRQLIQSLNDGRWPITCIITDIMLSSPIEVAEEFGIPVIAFCPCSARYLSIHFFIPKLVEEGQIPYADDDPIGEIQGVPLFEGLLRRNHLPGSWSDKSADISFSHGLINQTLAAGRASALILNTFDELEAPFLTHLSSIFNKIYTIGPLHALSKSRLGDSSSSASALSGFWKEDRACMSWLDCQPPRSVVFVSFGSTMKMKADELREFWYGLVSSGKPFLCVLRSDVVSGGEAAELIEQMAEEEGAGGKLGMVVEWAAQEKVLSHPAVGGFLTHCGWNSTVESIAAGVPMMCWPILGDQPSNATWIDRVWKIGVERNNREWDRLTVEKMVRALMEGQKRVEIQRSMEKLSKLANEKVVRGGLSFDNLEVLVEDIKKLKPYKF CmOUGT (SEQ ID NO: 52) MAELSHTHHVLLFPFPAKGHIKPFFSLAQLLCNAGLRVTFLNTDHHHRRIHDLNRLAAQLPTLHFDSVSDGLPPDEPRNVFDGKLYESIRQVTSSLFRELLVSYNNGTSSGRPPITCVITDVMFRFPIDIAEELGIPVFTFSTFSARFLFLIFWIPKLLEDGQLRYPEQELHGVPGAEGLIRWKDLPGFWSVEDVADWDPMNFVNQTLATSRSSGLILNTFDELEAPFLTSLSKIYKKIYSLGPINSLLKNFQSQPQYNLWKEDHSCMAWLDSQPRKSVVFVSFGSVVKLTSRQLMEFWNGLVNSGMPFLLVLRSDVIEAGEEVVREIMERKAEGRWVIVSWAPQEEVLAHDAVGGFLTHSGWNSTLESLAAGVPMISWPQIGDQTSNSTWISKVWRIGLQLEDGFDSSTIETMVRSIMDQTMEKTVAELAERAKNRASKNGTSYRNFQTLIQDITNIIETHI SgUGT720_269_4_Manus (SEQ ID NO: 53) MAEQAHDLLHVLLFPFPAEGHIKPFLCLAELLCNAGFHVTFLNDTDYNHRRLHNLLAARFPSLHFESISDGLPPDQPRDILDPKFFISICQVTKPLFRELLLSYKRISSVQTGRPPITCVITDVIFRFPIDVAEELDIPVFSFCTFSARFMFLYFWIPKLIEDGQLPYPNGNINQKLYGVAPEAEGLLRCKDLPGHWAFADELKDDQLNFVDQTTASSRSSGLILNTFDDLEAPFLGRLSTIFKKIYAVGPIHSLLNSHHCGLWKEDHSCLAWLDSRAAKSVVFVSSFGSLVKITSRQLMEFWHGLLNSGKSFLFVLRSDVVEGDDEKQVVKEIYETKAEGKWLVVGWAPQEKVLAHEAVGGFLTHSGWNSILESIAAGVPMISCPKIGDQSSNCTWISKVWKIGLEMEDRYDRVSVETMVRSIMEQEGEKMQKTIAELAKQAKYKVSKDGTSYQNLECLIQDIKKLNQIEGFINNPNFSDLLRV SgUGT720_269_4_E (SEQ ID NO: 54) MAEQAHDLLHVLLFPYPAKGHIKPFLCLAELLCNAGLNVTFLNTDYNHRRLHNLHLLAACFPSLHFESISDGLQPDQPRDILDPKFYISICQVTKPLFRELLLSYKRTSSVQTGRPPITCVITDVIFRFPIDVAEELDIPVFSFCTFSARFMFLYFWIPKLIEDGQLPYPNGNINQKLYGVAPEAEGLLRCKDLPGHWAFADELKDDQLNFVDQTTASLRSSGLILNTFDDLEAPFLGRLSTIFKKIYAVGPIHALLNSHHCGLWKEDHSCLAWLDSRAARSVVFVSFGSLVKITSRQLMEFWHGLLNSGTSFLFVLRSDVVEGDGEKQVVKEIYETKAEGKWLVVGWAPQEKVLAHEAVGGFLTHSGWNSILESIAAGVPMISCPKIGDQSSNCTWISKVWKIGLEMEDQYDRATVEAMVRSIMKHEGEKIQKTIAELAKRAKYKVSKDGTSYRNLEILIEDIKKIKPN SgUGT74_345_2_I (SEQ ID NO: 55) MADETTVNGGRRASDVVVFAFPRHGHMSPMLQFSKRLVSKGLRVTFLITTSATESLRLNLPPSSSLDLQVISDVPESNDIATLEGYLRSFKATVSKTLADFIDGIGNPPKFIVYDSVMPWVQEVARGRGLDAAPFFTQSSAVNHILNHVYGGSLSIPAPENTAVSLPSMPVLQAEDLPAFPDDPEVVMNFMTSQFSNFQDAKWIFFNTFDQLECKVVNWMADRWPIKTVGPTIPSAYLDDGRLEDDRAFGLNLLKPEDGKNTRQWQWLDSKDTASVLYISFGSLAILQEEQVKELAYFLKDTNLSFLWVLRDSELQKLPHNFVQETSHRGLVVNWCSQLQVLSHRAVSCFVTHCGWNSTLEALSLGVPMVAIPQWVDQTTNAKFVADVWRVGVRVKKKDERIVTKEELEASIRQVVQGEGRNEFKHNAIKWKKLAKEAVDEGGSSDKNIEEFVKTIA XP_023543158.1 (SEQ ID NO: 56) MAELSHTHHVLLFPFPAKGHIKPFFSLAQLLCNAGLRVTFLNTDHHHRRIHDLDRLAAQLPTLHFDSVSDGLPPDEPRNVFDGKLYESIRQVTSSLFRDLLVSYNNGTSSGRPPITCVITDCMFRFPIDIAEELGIPVFTFSTFSARFLFLFFWIPKLLEDGQLRYPEQELHGVPGAEGLIRCKDLPGFLSDEDVAHWKPMNFVNQILATSRSSGLILNTFDELEAPFLTSLSKIYKKIYSLGPINSLLKNFQSQPQYNLWKEDHSCMAWLDSQPRKSVVFVSFGSVVKLTSRQLMEFWNGLVNSGKPFLLVLRSDVTEAGEEVVREIMERKAEGRWVIVNWAPQEEVLAHDAVGGFLTHSGWNSTLESLAAGVPMISWPQIGDQTSNSTWVSKVWRIGLQLEDGFDSSTIETMVRSIMDQKMEKTVAELAERAKNRASKNGTSYRNFQTLIQDITNIIETHI XP_022151514.1 (SEQ ID NO: 57) MANSDDHQHHVLLFPFPAKGHIKPFLCLAQLLCGAGLQVTFLNTDHNHRRIDDRHRRLLATQFPMLHFKSISDGLPPDHPRDLLDGKLIASMRRVTESLFRQLLLSYNGYGNGTNNVSNSGRRPPISCVITDVIFSFPVEVAEELGIPVFSFATFSARFLFLYFWIPKLIQEGQLPFPDGKTNQELYGVPGAEGIIRCKDLPGSWSVEAVAKNDPMNFVKQTLASSRSSGLILNTFEDLEAPFVTHLSNTFDKIYTIGPIHSLLGTSHCGLWKEDYACLAWLDARPRKSVVFVSFGSLVKTTSRELMELWHGLVSSGKSFLLVLRSDVVEGEDEEQVVKEILESNGEGKWLVVGWAPQEEVLAHEAIGGFLTHSGWNSTMESIAAGVPMVCWPKIGDQPSNCTWVSRVWKVGLEMEERYDRSTVARMARSMMEQEGKEMERRIAELAKRVKYRVGKDGESYRNLESLIRDIKITKSSN XP_023764969.1 (SEQ ID NO: 58) MAGDAVPIIGQKQPHVVFVPFPAQSHIKCMLKLARILHHNGLHVTFINTHSNQKRLVKSNGILGLDKAPGFQFKTVPDGLSSATDDGVEHTQTMAELWTYLGANFLGSFLDVVSGLEIPVTCIICDGFMTYTNTIHAAEKLNIPIILFWTMAASGFMGFYQVKVLAEKGILPLKDEIYLTNGYLDMEIDWIPGMEGIRLKELPEIKLFTKHDNPAFKFLLETAQLAHKVSHMIIHTFEELEASLIKELKSIFPNVYSVGPLELLLNQITEKETNKSLCNGYSLWKEEPECVQWLQSKEPNSVVYVNFGSIAVMSLQDLLEFGWGLVNSKHEFLWIIRTDLVDGKPVVLPQELEDAMKGKGFVASWCSQEEVLNHSSVGGFLTHGGWGSIIESLSAGVPMICWPVSGDQQTNCRQMCKEWGVGMEISRNVKRDEVEKLVKELMEGMEGKRMRKKALEWKKVAEKATGSNGSSWIDAEKLANQIVKLSTKFPTV XP_043637218.1 (SEQ ID NO: 59) MANTMTRVDEKQPHVVFIPFPAQSHIKCMLKLAELLHHKGIHITFINTRSNHKKLVESGGTHDFKDIPSFQFRTVSDCPSGAEDDGVEPVTPTLVEVWMYLADNFFDSFLDIVSGLESTPATCIICDGFMTYTNMITAAEKLKIPIILFWTMAACGFMGFYQAKVLTDKGLLPLQDESYLTNGYLEMEIDWIPGMEGCRLKDLPEDMLVTKVDDPGYRYLLETAQAAHNVSYIIMHTFEELEARLVNEIKSIFPNVYTVGPLQLLLNQIKEKENKKTAFNCYSLWNEEPECIQWLESKEPSSVVYVNFGSLAVMSLEDLIEFGWGLVNSDYEFLWIIRADLVKGNPVVLPKELEEEIGKKGFLAKWCSQEEVLNHPSVGGFLTHGGWGSTIESLSAGLPMICWPSIGDQRANCRQMCKQWQVGMEIGKNVKRDEVEKLVKLLMDGLEGKRMKEKARHWKKMAEKATSSNGSSSIDVEKLANQIINFP XP_008445481.1 (CmUGTc3_0) (SEQ ID NO: 60) MAEMTAANGGGERIKQSHVIVFPFPRHGHMSPMLQFSKRLISKGLLLTFLITSSASQSLTINIPPSPSFHFKIISDLPESDDVATLDAYLRSFRAAVTKSLSNFIDEVLTSSSNEEVPPTLIVYDSVMPWVQSVAAERGLDSAPFFTESAAVNHLLHLVYGGSLSIPPPDNVVVSLPSEIVLQPEDLPSFPDDPEVVLDFMTSQFSHLENVKWIFINTFDRLESKVVNWMAKTLPIKTVGPTIPSAYLDGRLEKDKAYGLNVSKSNNGKCPIKWLDSKETASVIYISFGSLVILSEEQVKELTNLLRDTDFSFLWVLRESEMVKLPKNFVQDTSDRGLIVNWCCQLQVLSHKAVSCFVTHCGWNSTLEALSLGVPMVAIPQWIDQTTNAKFVADVWRVGVRVKKNEKSVAIKEELEASIRKIVVQGNGTNEFKQNAIKWKNLAKEAVDERGSSDKNIEEFVQALV XP_022132196.1 (SEQ ID NO: 72) MADNVVVSLPSEIVLQPEDLPSFPDDPEVVLDFMTSQFSHLENVKWIFINTFDRLESKVVNWMAKTLPIKTVGPTIPSAYLDGRLEKDKAYGLNVSKSNNGKCPIKWLDSKETASVIYISFGSLVILSEEQVKELTNLLRDTDFSFLWVLRESEMVKLPKNFVQDTSDRGLIVNWCCQLQVLSHKAVSCFVTHCGWNSTLEALSLGVPMVAIPQWIDQTTNAKFVADVWRVGVRVKKNEKSVAIKEELEASIRKIVVQGNGTNEFKQNAIKWKNLAKEAVDERGSSDKNIEEFVQALVAANGGGERGKQSHVIVFPFPRHGHMSPMLQFSKRLISKGLLLTFLITSSASQSLTINIPPSPSFHFKIISDLPESDDVATLDAYLRSFRAAVTKSLSNFIDEVLTSSSNEEVPPTLIVYDSVMPWVQSVAAERGLDSAPFFTESAAVNHLLHLVYGGSLSIPPP XP_022132196.1 (SEQ ID NO: 61) MAEKATANGGRRSSHVLLFAYPMHGHMSPMLQFAKRLASKGLLVTFLTTSSVTESLQIDLPPSYPIHLRFISDFHTEVIETLKQRHEAFAAAVSRSLGEFLDGALINGDHPPRLMVFDSVMPWAMEVARSRGLEAAPFFTESAAVNHILNQVYEGSLSIPAPENAAVSIPSLPNLEAEDLPYFPSVIREVTLEFMTRQFSSFKDAKWIFINTFDQLEPQIVNWMGERWPIKTVGPTVPSAYLDGRLEKDKTYGLKRQKPEDGRAVEWLDSKETASVVYISFGSLVMLAEKQVKELTNFLTESGLPFLWVLRESEMEKLPENFIQETSGKGLVVNWCSQLEVLSHKAVGCFVTHGGWNSTLEALSSGVPMVAVPQWIDQTTNAKFIADVWEIGVRVKLNEKHEIATKDELEASIRQVIEGREKKNSIKWRKLAKEAVDEGGSSDKNIEDFAKTIM XP_023538293.1 (SEQ ID NO: 62) MASNTTLNGGRRSSHVVLFAYPKQGHLSPMLQFAKRLASKGLRITFLTTTSTTKSLEIDLPASYQIDLRFISDVRTEPILSLKDEHESFEAVVSKSFGDFIDGTLRSSGFDPPRFVIFDSVMPWAMDVARVRGIDSAPFFTESCVVNHILNQVYEGSFSIPPVENVAAGISIPPLPVLQTEDLPYFSYEPELVLKFMTDQFSSFKNAKWIFVNTFDQLEMKVVNWMTQKWPIKTIGPNIPSAYLDGRLKDDKTYGLNHQNLNNCKIFQWLDSKEIASVIYLSFGSLVILPEEQVNELARFFKDTNFSFLWVLRESEQEKLPNNFVQQTSHKGLVVKWCCQLQVLSHKAVSCFVTHCGWNSTIEALSLGVPMVAVPQWIDQTTNAKFVADVWKVGARVKMNDKGIATKLELESIRHVSQGYRQNEIKQNSIKLRNLAKEAMDEGGSSDKNIEQFVKELD UGT74AC1 (SEQ ID NO: 63) MAEKGDTHILVFPFPSQGHINPLLQLSKRLIAKGIKVSLVTTLHVSNHLQLQGAYSNSVKIEVISDGSEDRLETDTMRQTLDRFRQKMTKNLEDFLQKAMVSSNPPKFILYDSTMPWVLEVAKEFGLDRAPFYTQSCALNSINYHVLHGQLKLPPETPTISLPSMPLLRPSDLPAYDFDPASTDTIIDLLTSQYSNIQDANLLFCNTFDKLEGEIIQWMETLGRPVKTVGPTVPSAYLDKRVENDKHYGLSLFKPNEDCLKWLDSKPGSVLYVSYGSLVEMGEEQLKELALGIKETGKFFLWVVRDTEAEKLPPNFVESVAEKGLVVSWCSQLEVLAHPSVGCFFTHCGWNSTLEALCLGVPVVAFPQWADQVTNAKFLEDVWKVGKRNEQRLASKEEVRSCIWEVMEGERASEFKSNSMEWKKWAKEAVDEGGSSSDKNIEEFVAMLKQT UGT74AC1_M7 (SEQ ID NO: 64) MAEKGDTHILVFPFPAQGHINPLLQLSKHLIAKGIKVSLVTTLHVSNRMQLQGAYSNSVKIEVISDGSEDRLETDTLRQYLDRFRQKMTKNLEDFLQKAMVSSNPPKFIIYDSTMPWVLEVAKEFGLDRAPFYTQSCALNSINYHVLHGQLKLPPETPTISLPSMPLLRPSDLPAYDFDPASTDTIIDLLTSQYSNIQDANLLFCNTFDKLEGEIIQWMETLGRPVKTVGPTVPSAYLDKRVENDKHYGLSLFKPNEDCLKWLDSKPGSVLYVSYGSLVEMGEEQLKELALGIKETGKFFLWVVRDTEAEKLPPNFVESVAEKGLVVSWCSQLEVLAHPSVGCFFTHCGWNSTLEALCLGVPVVAFPQWADQVTNAKFLEDVWKVGKRNEQRLASKEEVRSCIWEVMEGERASEFKSNSMEWKKWAKEAVDEGGSSSDKNIEEFVAMLKQT Stevia UGT85C1 MogUGTc3_0 (SEQ ID NO: 27) MADQMAKIDEKKPHVVFIFPPAQSHIKCMLKLARILHQKGLYITFINTDTNHERLVASGGTQWLENAPGFWFKTVPDGFGSAKDDGVKPTDALRELMDYLKTNFFDLFLDLVLKLEVPATCIICDGCMTFANTIRAAEKNLIPVILFWTMAACGFMAFYQAKVLKEIVPVKDETYLTNGYLDMEIDWIPGMKRIRLRDLPEFILATKQNYFEFLFETAQLADKVSHMIIHTFEELEAS LVSEIKSIFPNVYTIGPLQLLLNKITQKETNNDSYSLWKEEPECVEWLNSKEPNSVVYVNFGSLAVMSLQDLVEFGWGLVNSNHYFLWIIRANLIDGKPAVMPQELKEAMNEKGFVGSWCSQEEVLNHPAVGGFLTHCGWGSIIESLSAGVPMLGWPSIGDQRANCRQMCKEWEVGMEIGKNVKRDEVEKLVRLMEGLEGERMRKKALEWKKSATLATCCNGSSLDVEKLANEIKKLSRN MogUGTC3_1 (SEQ ID NO: 28) MADQMAKIDEKKPHVVFIFPPAQSHIKCMLKLARILHQKGFYITFINTETNHERLVASGGTQWLENAPGFWFKTVPDGFGSAKDDGVKPTDALRELMDYLKTNFFDLFLDLVLKLEVPATCIICDGFMTFANTIRAAEKNLIPVILFWTMAACGFMAFYQAKVLKEIVPVKDETYLTNGYLDMEIDWIPGMKRIRLRDLPEFILATKQNYFEFLFETAQLADKVSHMIIHTFEELEAS LVSEIKSIFPNVYTIGPLQLLLNKITQKETNNDSYSLWKEEPECVEWLNSKEPNSVVYVNFGSLAVMSLQDLVEFGWGLVNSNHYFLWIIRANLIDGKPAVMPQELKEAMNEKGFVGSWCSQEEVLNHPAVGGFLTHCGWGSIIESLSAGVPMLGWPSIGDQRANCRQMCKEWEVGMEIGKNVKRDEVEKLVRLMEGLEGERMRKKALEWKKSATLATCCNGSSLDVEKLANEIKKLSRN MogUGTc3_2 (SEQ ID NO: 29) MADQMAKIDEKKPHVVFIFPPAQSHIKCMLKLARILHQKGFYITFINTETNHERLVASGGTQWLENAPGFWFKTVPDGFGSAKDDGVKPTDALRELMDYLKTNFFDLFLDLVLKLEVPATCIICDGFMTFANTIRAAEKNLIPVILFWTMAACGFMAFYQAKVLKEIVPVKDETYLTNGYLDMEIDWIPGMKRIRLRDLPEFILATKQNYFEFLFETAQLADKVSHMIIHTFEELEAS LVSEIKSIFPNVYTIGPLQLLLNKITQKETNNDSYSLWKEEPECVEWLNSKEPNSVVYVNFGSLIVMSLQDLVEFGWGLVNSNHYFLWIIRANLIDGKPAVMPQELKEAMNEKGFVGSWCSQEEVLNHPAVGGFLTHCGWGSIIESLSAGVPMLGWPSIGDQRANCRQMCKEWEVGMEIGKNVKRDEVEKLVRLMEGLEGERMRKKALEWKKSATLATCCNGSSLDVEKLANEIKKLSRN MogUGTc3_3 (SEQ ID NO: 30) MADQMAKIDEKKPHVVFIPFPAQSHIKCMLKLARILHQKGFYITFINTETNHERLVASGGTQWLENAPGFWFMVPDGGFGSAKDDGVKPTDALRELMDYLKTNFFDLFLDLVLKLEVPATCIICDGFMTFANTIRAAEKLNIPVILFWTMAACGFMAFYQAKVLKEKEIVPVKDETYLTNGYLDMEIDWIPGMKRIRLRDLPEFILATKQNYFAFEFLFETAQLADKVSHMIIHTFEELEASLVSEIKSIFPNVYTIGPLQLLLNKITQKETNNDSYSLWKEEPECVEWLNSKEPNSVVYVNYGSLIVMSLQDLVEFGWGLVNSNHYFLWIIRANLIDGKPAVMPQELKEAMNEKGFVGSWCSQEEVLNHPAVGGFLTHCGWGSIIESLSAGVPMLGWPSIGDQRANCRQMCKEWEVGMEIGKNVKRDEVEKLVRMLMEGLEGERMRKKALEWKKSATLATCCNGSSSLDVEKLANEIKKLSRN MogUGTC24_0 (SEQ ID NO: 31) MADVAEEQPLKIYFIPYLAAGHMIPLCDIATLFASRGHHVTIITTPSNAQTLRESHHFRVQTIQFPSQEVGLPAGVQNLTAVTNLDDSYKIYHATMLLRKHIEDFVERDPDCIVADFLFPWVDDVATKLHIPRLVFTNGFTLFTICAMESHKAHPLPVDAASGSFVIPDFPHHVTINSTPPKRTKEFVDPLLTEAFKSHGFLINSFVELDGEECVEHYERITGGHKAWHLGPAFLVHRTAQDRGEKSVVSTQECLSWLDSKRDNSVLYICFGTICYFPDKQLYEIASAIEASGHEFIWVVPEKRGNADESEEEKEKWLPKGFEERNNGKKGMIIRGWAPQVAILGHPAVGGFLTHCGWNSTVEAVSAGVPMITWPVHSDQYFNEKLITQVRGIGVEVGAEEWIVTAFRETEKLVGRDRIERAVRRVMDGGGDEAVQIRRRARELGEMARQAVQEGGSSHTNTALINDLKRWRDSKQLN MogUGTC24_1 (SEQ ID NO: 32) MADVAEEQPLKIYFIPYLAAGHMIPLCDIATLFASRGHHVTIITTPSNAQTLRESHHFRVQTIQFPSQEVGLPEGVQNLTAVTNLDDSYKFYHATMLLRKPIEDFVERDPDCIVADFLFPWVDDVATKLHIPRLVFNGFTLFTICAMESHKAHPLPVDAASGSFVIPDFPHHVTINSTPPKRTKEFVDPLLTEAFKSHGFLINSFVELDGEECVEHYERITGGHKAWHLGPAFLVHRTAQDRGEKSVVSTQECLSWLDSKRDNSVLYICFGTICYFPDKQLYEIASAIEASGHEFIWVVPEKRGNADESEEEKEKWLPKGFEERNNGKKGMIIRGWAPQVAILGHPAVGGFLTHCGWNSTVEAVSAGVPMITWPVHSDQYFNEKLITQVRGIGVEVGAEEWIVTAFRETEKLVGRDRIERAVRRVMDGGGDEAVQIRRRARELGEMARQAVQEGGSSHTNTALINDLKRWRDSKQLN MogUGTC24_2 (SEQ ID NO: 33) MADVAEEQPLKIYFIPYMAAGHMIPLCDIATLFASRGHHVTIITTPSNAQTLRESHHFRVQTIQFPSQEVGLPEGVQNLTAVTNLDDSYKFYHAAMLLRKPIEDFVERDPDCIVADFLFPWVDDVATKLHIPRLVFNGFTLFTICAMESHKAHPLPVDAASGSFVIPDFPHHVTINSTPPKRFKEFVDELLTEAFKSHGFLINSFVELDGEECVEHYERITGGHKAWHLGPAFLVHRTAQDRGEKSVVSTQECLSWLDSKRDNSVLYICFGTICYFPDKQLYEIASAIEASGHEFIWVVPEKRGNADESEEEKEKWLPKGFEERNNGKKGMIIRGWAPQVAILGHPAVGGFLTHCGWNSTVEAVSAGVPMITWPVHSDQFFNEKLITQVRGIGVEVGAEEWIVTAFRETEKLVGRDRIERAVRRVMDGGGDEAVQIRRRARELGEMARQAVQEGGSSHTNTALINDLKRWRDSKQLN MogUGTC24_3 (SEQ ID NO: 34) MADVAEEQPLKIYFIPYMAAGHMIPLCDIATLFASRGHVTIITTPSNAQTLRESHHFRVQTIQFPSQEVGLPEGVQNLTAVTNLDDAYKFYHAAMLLRKPIEDFVERDPDCIVADFLFPWVDDVATKLHIPRLVFNGFTLFTICAMESHKAHPLPVDAASGSFVIPDFPHHVTINSTPPKRFKEFVDELLTEAFKSHGFLINSFVELDGEECVEHYERITGGHKAWHLGPAFLVHRTAQDRGEKSVVSTQECLSWLDSKRDNSVLYICFGSICYFPDKQLYEIASAIEASGHEFIWVVPEKRGKADESEEEKEKWLPKGFEERNNGKKGLIIRGWAPQVAILGHPAVGGFLTHCGWNSTVEAVSAGVPMITWPVHSDQFFNEKLITQVRGIGVEVGAEEWIVTAFRETEKLVGRDRIERAVRRVMDGGGDEAVQIRRRARELGEMARQAVQEGGSSHTNTALINDLKRWRDSKQLN MogUGTC24_4 (SEQ ID NO: 48) MADVAEEQPLKIYFIPYMAAGHMIPLCDIATLFASRGHHVTIITTPSNAQTLRESHHFRVQTIQFPSQEVGLPEGVQNLTAVTNLDDAYKFYHAAMLLRKPIEDFVERDPDCIVADFLFPWVDDVATKLHIPRLVFNGFTLFTICAMESHKAHPLPVDAASGSFVIPDFPHHVTINSTPPKRFKEYVDELLTEAFKSHGFLINSFVELDGEECVEHYERITGGHKAWHLGPAFLVHRTAQDRGEKSVVSTQECLSWLDSKRDNSVLYICFGSICYFPDKQLYEIASAIEASGHEFIWVVPEKRGKADESEEEKEKWLPKGFEERNNGKKGLIIRGWAPQVAILGHPAVGGFLTHCGWNSTVEAVSAGVPMITWPVHSDQFFNEKLITQVRGIGVEVGAEEWIVTAFRETEKLVGRDRIERAVRRVMDGGGDEAVQIRRRARELGEMARQAVQEGGSSHTNTALINDLKRWRDSKQLN MogUGTC24_5 (SEQ ID NO: 49) MADVAEEQPLKIYFIPYMAAGHMIPLCDIATLFASRGHHVTIITTPSNAQTLRESHHFRVQTVQFPSQEVGLPEGVQNLTAVTNLDDAYKFYHAAMLLRKPIEDFVERDPPDCIVADFLFPWVDDVATKLHIPRLVFNGFTLFTICAMESHKAHPLPVDAASGSFVIPDFPHHVTINSTPPKRFKEYVDELLTEAFKSHGFLINSFVELDGEECVEHYERITGGHKAWHLGPAFLVHRTAQDRGEKSVVSTQECLSWLDSKRDNSVLYICFGSICYFPDKQLYEIASAIEASGHEFIWVVPEKRGKADESEEEKEKWLPKGFEERNNGKKGLIIRGWAPQVAILGHPAVGGFLTHCGWNSTVEAVSAGVPMITWPVHSDQFFNEKLITQVRGIGVEVGAEEWIVTAFRETEKLVGRDRIERAVRRVMDGGDEAVQIRRRARELGEMARQAVQEGGSSHTNLTALINDLKRWRDSKQLN Monk fruit UGT94-289,SEQ ID NO: 35) MADAQRGHTTTILMFPWLGYGHLSAFLELAKSLSRRNFHIYFCSTSVNLDAIKPKLPSSSSSDSIQLVELCLPSSPDQLPPHLHTTNALPPHLMPTLHQAFSMAAQHFAAILHTLAPHLLIYDSFQPWAPQLASSLNIPAINFNTTGASVLTRMLHATHYPSSKFPISEFVLHDYWKAMYSAAGGAVTKKDHKIGETLANCLHASCSVILINSFRELEEKYMDYLSVLLNKKVVPVGPLVYEPNQDGEDEGYSSIKNWLDKKEPSSTVFVSFGSEYFPSKEEMEEIAHGLEASEVHFIWVVRFPQGDNTSAIEDALPKGFLERVGERGMVVKGWAPQAKILKHWSTGGFVSHCGWNSVMESMMFGVPIIGVPMHLDQPFNAGLAEEAGVGVEAKRDPDGKIQRDEVAKLIKEVVVEKTREDVRKKAREMSEILRSKGEEKMDEMVAAISLFLKI MogUGT1216_0, SEQ ID NO: 36) MAVTKKDHKIGETLANCLHASCSVILINSFRELEEKYMDYLSVLLNKKVVPVGPLVYEPNQDGEDEGYSSIKNWLDKKEPSSTVFVSFGSEYFPSKEEMEEIAHGLEASEVHFIWVVRFPQGDNTSAIEDALPKGFLERVGERGMVVKGWAPQAKILKHWSTGGFVSHCGWNSVMESMMFGVPIIGVPMHLDQPFNAGLAEEAGVGVEAKRDPDGKIQRDEVAKLIKEVVVEKTREDVRKKAREMSEILRSKGEEKMDEMVAAISLFLKIGGGSDAQRGHTTTILMFPWLGYGHLSAFLELAKSLSRRNFHIYFCSTSVNLDAIKPKLPSSSSSDSIQLVELCLPSSPDQLPPHLHTTNALPPHLMPTLHQAFSMAAQHFAAILHTLAPHLLIYDSFQPWAPQLASSLNIPAINFNTTGASVLTRMLHATHYPSSKFPISEFVLHDYWKAMYSAAGGA MogUGT1216_1 (SEQ ID NO: 37) MAVTKKDHKIGETVANCLHASCSVILINSFRELEEKYMDYLSVLLNKKVVPVGPLVYEPNQDGEDEGYSSIKNWLDKKEPSSTVFVSFGSEYFPSKEEMEEIAHGLEASEVHFIWVVRFPQGDNTSAIEDALPKGFLERVGERGMVVKGWAPQAKILKHWSTGGFVSHCGWNSVMESMMFGVPIIGVPMHLDQPFNAGLAEEAGVGVEAKRDPDGKIQRDEVAKLIKEVVVEKTREDVRKKAREMSEILRSKGEEKMDEMVAAISQFLKIGGGSDAQRGHTTTILMFPWLGYGHLSAFLELAKSLSRRNFHIYFCSTSVNLDAIKPKLPSSSSSDSIQLVELCLPSSPDQLPPHLHTTNALPPHLMPTLHQAFSMAAQHFAAILHTLAPHLLIYDSFQPWAPQLASSLNIPAINFNTTGASVLTRMLHATHYPSSKFPISEFVLHDYWKAMYSAAGGA MogUGT1216_2 (SEQ ID NO: 38) MAVTKKDHKIGETVANCLHASCSVILINSFRELEEKYMDYLSVLLNKKVVPVGPLVYEPNQDGEDEGYSSIKNWLDKKEPSSTVFVSFGSEYFPSKEEMEEIAHGLEASEVHFIWVVRFPQGDNTSAIEDALPKGFLERVGERGMVVKGWAPQAKILKHPSTGGFVSHCGWNSVMESMMFGVPIIGVPMHLDQPFNAGLAEEAGVGVEAKRDPDGKIQRDEVAKLIKEVVVEKTREDVRKKAREMSEILRSKGEEKMDEMVAAISQFLKIGGGSDAQRGHTTTILMFPWLGYGHLSAFLELAKSLSRRNFHIYFCSTSVNLDAIKPKLPSSSSSDSIQLVELCLPSSPDQLPPHLHTTNALPPHLMPTLHQAFSMAAQHFAAILHTLAPDLLIYDSFQPWAPQLASSLNIPAINFNTTGASVLTRMLHATHYPSSKFPISEFVLHDYWKAMYSAAGGA MogUGT1216_3 (SEQ ID NO: 90) MAVTKKDHKIGETVANCLEASCSVILINSFRELEEKYMDYLSVLLNKKVVPVGPLVYEPNQDGEDEGYSSIKNWLDKKEPSSTVFVSFGSEYFPSKEEMEEIAHGLEASEVHFIWVVRFPQGDNTSAIEDALPKGFLERVGERGMVVKGWAPQAKILKHPSTGGFVSHCGWNSVMESMMFGVPIIGVPMHLDQPFNAGLAEEAGVGVEAKRDPDGKIQRDEVAKLIKEVVVEKTREDVRKKAREMSEILRSKGEEKMDEMVAAISQFLKIGGGSDAQRGHSTTILMFPWLGYGHLSAFLELAKSLSRRNFHIYFCSTKVNLDAIKPKLPSSSSSDSIQLVELCLPSSPDQLPPHLHTTNALPPHLMPTLHQAFSMAAQHFAAILHTLAPDLLIYDSFQPWAPQLASSLNIPAINFNTTGASVLTRMLHATHYPSSKFPISEFVLHDYWKAMYSAAGGA Coffee Tree Uridine diphosphate-dependent glycosyltransferase 1,6 MogUGT16_0 (SEQ ID NO: 39) MAENHATFNVLMLPWLAHGHVSPYLELAKKLTARNFNVYLCSSPATLSSVRSKLTEKFSQSIHLVELHLPKLPELPAEYHTTNGLPPHLMPTLKDAFDMAKPNFCNVLKSLKPDLLIYDLLQPWAPEAASAFNIPAVVFISSSATMTSFGLHFFKNPGTKYPYGNAIFYRDYESVFVENLTRRDRDTYRVINCMERSSKIILIKGFNEIEGKYFDYFSCLTGKKVVPVGPLVQDPVLDDEDCRIMQWLNKKEKGSTVFVSFGSEYFLSKKDMEEIAHGLEVSNVDFIWVVRFPKGENIVIEETLPKGFFERVGERGLVVNGWAPQAKILTHPNVGGFVSHCGWNSVMESMKFGLPIIAMPMHLDQPINARLIEEVGAGVEVLRDSKGKLHRERMAETINKVMKEASGESVRKKARELQEKLELKGDEEIDDVVKELVQLCATKNKRNGLHYY MogUGT16_1 (SEQ ID NO: 40) MAENHATFNVLMLPWLAHGHVSPYLELAKKLTARNFNVYLCSSPATLSSVRSKLTEKFSQSIHLVELHLPKLPELPAEYHTTNGLPPHLMPTLKDAFDMAKPNFCNVLKSLKPDLLIYDLLQPWAPEAASAFNIPAVVFISSSATMLSFGLHFFKNPGTKYPYGNAIFYRDYESVFVENLTRRDRDTYRVINCMERSSKIILIKGFKEIEGKYFDYFSCLTGKKVVPVGPLVQDPVLDDEDCRIMQWLNKKEKGSTVFVSFGSEYFLSKKDMEEIAHGLEVSNVDFIWVVRFPKGENIVIEETLPKGFFERVGERGLVVNGWAPQAKILTHPNVGGFVSHCGWNSVMESMKFGLPIIAMPMHLDQPINARLIEEVGAGVEVLRDSKGKLHRERMAETINKVMKEASGESVRKKARELQEKLELKGDEEIDDVVKELVQLCATKNKRNGLHYY CmUGTc3_1 (SEQ ID NO: 65) MADNVVVSLPSEIVLQPEDLPSFPDDPEVVLDFMTSQFSHLENVKWIFINTFDRLESKVVNWMAKTLPIKTVGPTIPSAYLDGRLEKDKAYGLNVSKSNNGKCPIKWLDSKETASVIYISFGSLVILSEEQVKELTNLLRDTDFSFLWVLRESEMVKLPKNFVQDTSDRGLIVNWCCQLQVLSHKAVSCFVTHCGWNSTLEALSLGVPMVAIPQWIDQTTNAKFVADVWRVGVRVKKNEKSVAIKEELEASIRKIVVQGNGTNEFKQNAIKWKNLAKEAVDERGSSDKNIEEFVQALVAANGGGERGKQSHVIVFPFPRHGHMSPMLQFSKRLISKGLLLTFLITSSASQSLTINIPPSPSFHFKIISDLPESDDVATLDAYLRSFRAAVTKSLSNFIDEVLTSSSNEEVPPTLIVYDSVMPWVQSVAAERGLDSAPFFTESAAVNHLLHLVYGGSLSIPPP CmUGTc3_2 (SEQ ID NO: 66) MADNVVVSLPSEIVLQPEDLPSFPDDPEVVLDFMTSQFSHLENVKWIFINTFDRLESKVVNWMAKTLPIKTVGPTIPSAYLDGRLEKDKAYGLNVSKSNNGKCPIKWLDSKETASVIYISFGSLVILSEEQVKELTNLLRDTDFSFLWVLRESEMVKLPKNFVQDTSDRGLIVNWCCQLQVLSHKAVSCFVTHCGWNSTLEALSLGVPMVAIPQWIDQTTNAKFVADVWRVGVRVKKNEKSVAIKEELEASIRKIVVQGNGTNEFKQNAIKWKNLAKEAVDERGSSDKNIEEFVQALVAANGGGERGKQSHVIVFPFPRHGHMSPMLQFSKRLISKGLLLTFLITSSASQSMTINIPPSPSFHFKIISDLPESDDVATLDAYLRSFRAAVTKSLSNFIDEVLTSSSNEEVPPTLIVYDSVMPWVQSVAAERGLDSAPFFTESAAVNHLLHLVYGGSLSIPPP CmUGTc3_3 (SEQ ID NO: 67) MADNVVVSLPSEIVLQPEDLPSFPDDPEVVLDFMTSQFSNLENVKWIFINTFDRLESKVVNWMAKTLPIKTVGPTIPSAYLDGRLEKDKAYGLNVSKSNNGKCPIKWLDSKETASVIYISFGSLVILSEEQVKELTNLLRDTDFSFLWVLRESEMVKLPKNFVQDTSDRGLIVNWCCQLQVLSHKAVSCFVTHCGWNSTLEALSLGVPMVAIPQWADQTTNAKFVADVWRVGVRVKKNEKGVAIKEELEASIRKIVVQGNGTNEFKQNAIKWKNLAKEAVDERGSSDKNIEEFVQALVAANGGGERGKQSHVIVFPFPRHGHMSPMLQFSKRLISKGLLLTFLITSSASQSLTINIPPSPSFHFKIISDLPESDDVATLDAYLRSFRAAVTKSLSNFIDEVLTSSSNEEVPPTLIVYDSVMPWVQSVAAERGLDSAPFFTESAAVNHLLHLVYGGSLSIPPP CmUGTc3_CP7 (SEQ ID NO: 73) MADNVVVSLPSEIVLQPEDLPSFPDDPEVVLDFMTSQFSHLENVKWIFINTFDRLESKVVNWMAKTLPIKTVGPTIPSAYLDGRLEKDKAYGLNVSKSNNGKCPIKWLDSKETASVIYISFGSLVILSEEQVKELTNLLRDTDFSFLWVLRESEMVKLPKNFVQDTSDRGLIVNWCCQLQVLSHKAVSCFVTHCGWNSTLEALSLGVPMVAIPQWIDQTTNAKFVADVWRVGVRVKKNEKSVAIKEELEASIRKIVVQGNGTNEFKQNAIKWKNLAKEAVDERGSSDKNIEEFVQALVEMTAANGGGERIKQSHVIVFPFPRHGHMSPMLQFSKRLISKGLLLTFLITSSASQSLTINIPPSPSFHFKIISDLPESDDVATLDAYLRSFRAAVTKSLSNFIDEVLTSSSNEEVPPTLIVYDSVMPWVQSVAAERGLDSAPFFTESAAVNHLLHLVYGGSLSIPPP CmUGTc3_CP9 (SEQ ID NO: 74) MASEIVLQPEDLPSFPDDPEVVLDFMTSQFSHLENVKWIFINTFDRLESKVVNWMAKTLPIKTVGPTIPSAYLDGRLEKDKAYGLNVSKSNNGKCPIKWLDSKETASVIYISFGSLVILSEEQVKELTNLLRDTDFSFLWVLRESEMVKLPKNFVQDTSDRGLIVNWCCQLQVLSHKAVSCFVTHCGWNSTLEALSLGVPMVAIPQWIDQTTNAKFVADVWRVGVRVKKNEKSVAIKEELEASIRKIVVQGNGTNEFKQNAIKWKNLAKEAVDERGSSDKNIEEFVQALVEMTAANGGGERIKQSHVIVFPFPRHGHMSPMLQFSKRLISKGLLLTFLITSSASQSLTINIPPSPSFHFKIISDLPESDDVATLDAYLRSFRAAVTKSLSNFIDEVLTSSSNEEVPPTLIVYDSVMPWVQSVAAERGLDSAPFFTESAAVNHLLHLVYGGSLSIPPPDNVVVSLP CmUGTc3_CP10 (SEQ ID NO: 75) MASEIVLQPEDLPSFPDDPEVVLDFMTSQFSHLENVKWIFINTFDRLESKVVNWMAKTLPIKTVGPTIPSAYLDGRLEKDKAYGLNVSKSNNGKCPIKWLDSKETASVIYISFGSLVILSEEQVKELTNLLRDTDFSFLWVLRESEMVKLPKNFVQDTSDRGLIVNWCCQLQVLSHKAVSCFVTHCGWNSTLEALSLGVPMVAIPQWIDQTTNAKFVADVWRVGVRVKKNEKSVAIKEELEASIRKIVVQGNGTNEFKQNAIKWKNLAKEAVDERGSSDKNIEEFVQALVAANGGGERGKQSHVIVFPFPRHGHMSPMLQFSKRLISKGLLLTFLITSSASQSLTINIPPSPSFHFKIISDLPESDDVATLDAYLRSFRAAVTKSLSNFIDEVLTSSSNEEVPPTLIVYDSVMPWVQSVAAERGLDSAPFFTESAAVNHLLHLVYGGSLSIPPPDNVVVSLP CmUGTc3_CP16 (SEQ ID NO: 76) MARLEKDKAYGLNVSKSNNGKCPIKWLDSKETASVIYISFGSLVILSEEQVKELTNLLRDTDFSFLWVLRESEMVKLPKNFVQDTSDRGLIVNWCCQLQVLSHKAVSCFVTHCGWNSTLEALSLGVPMVAIPQWIDQTTNAKFVADVWRVGVRVKKNEKSVAIKEELEASIRKIVVQGNGTNEFKQNAIKWKNLAKEAVDERGSSDKNIEEFVQALVAANGGGERGKQSHVIVFPFPRHGHMSPMLQFSKRLISKGLLLTFLITSSASQSLTINIPPSPSFHFKIISDLPESDDVATLDAYLRSFRAAVTKSLSNFIDEVLTSSSNEEVPPTLIVYDSVMPWVQSVAAERGLDSAPFFTESAAVNHLLHLVYGGSLSIPPPDNVVVSLPSEIVLQPEDLPSFPDDPEVVLDFMTSQFSHLENVKWIFINTFDRLESKVVNWMAKTLPIKTVGPTIPSAYLDG Other sequences pntA (SEQ ID NO: 41) MRIGIPRERLTNETRVAATPKTVEQLLKLGFTVAVESGAGQLASFDDKAFVQAGAEIVEGNSVWQSEIILKVNAPLDDEIALLNPGTTLVSFIWPAQNPELMQKLAERNVTVMAMDSVPRISRAQSLDALSSMANIAGYRAIVEAAHEFGRFFTGQITAAGKVPPAKVMVIGAGVAGLAAIGAANSLGAIVRAFDTRPEVKEQVQSMGAEFLELDFKEEAGSGDGYAKVMSDAFIKAEMELFAAQAKEVDIIVTTALIPGKPAPKLITREMVDSMKAGSVIVDLAAQNGGNCEYTVPGEIFTTENGVKVIGYTDLPGRLPTQSSQLYGTNLVNLLKLLCKEKDGNITVDFDDVVIRGVTVIRAGEITWPAPPIQVSAQPQAAQKAAPEVKTEEKCTCSPWRKYALMALAIILFGWMASVAPKEFLGHFTVFALACVVGYYVVWNVSHALHTPLMSVTNAISGIIVVGALLQIGQGGWVSFLSFIAVLIASINIFGGFTVTQRMLKMFRKN pntB (SEQ ID NO: 42) MSGGLVTAAYIVAAILFIFSLAGLSKHETSRQGNNFGIAGMAIALIATIFGPDTGNVGWILLAMVIGGAIGIRLAKKVEMTEMPELVAILHSFVGLAAVLVGFNSYLHHDAGMAPILVNIHLTEVFLGIFIGAVTFTGSVVAFGKLCGKISSKPLMLPNRHKMNLAALVVSFLLLIVFVRTDSVGLQVLALLIMTAIALVFGWHLVASIGGADMPVVVSMLNSYSGWAAAAAGFMLSNDLLIVTGALVGSSGAILSYIMCKAMNRSFISVIAGGFGTDGSSTGDDQEVGEHREITAEETAELLKNSHSVIITPGYGMAVAQAQYPVAEITEKLRARGINVRFGIHPVAGRLPGHMNVLLAEAKVPYDIVLEMDEINDDFADTDTVLVIGANDTVNPAAQDDPKSPIAGMPVLEVWKAQNVIVFKRSMNTGYAGVQNPLFFKENTHMLFGDAKASVDAILKAL Escherichia coli fldA (SEQ ID NO: 43) MAITGIFFGSDTGNTENIAKMIQKQLGKDVADVHDIAKSSKEDLEAYDILLLGIPTWYYGEAQCDWDDFFPTLEEIDFNGKLVALFGCGDQEDYAEYFCDALGTIRDIIEPRGATIVGHWPTAGYHFEASKGLADDDHFVGLAIDEDRQPELTAERVEKWVKQISEELHLDEILNA Wine-colored heterochromatobacter Cv. fdx (SEQ ID NO: 44) MALMITDECINCDVCEPECPNGAISQGDETYVIEPSLCTECVGHYETSQCVEVCPVDCIIKDPSHEETEDELRAKYERITGEG
Claims
1. A method for producing mogroside or mogroside, the method comprising: Provides microbial cells that produce squalene and express a biosynthetic pathway for converting squalene into mogroside or mogroside. The microbial cells are cultured under conditions suitable for producing the mogroside or mogroside, wherein the culture produces a squalene derivative, wherein the extra-pathogen product cucurbitadienol, hydrolyzed 2,3;22,23-squalene dioxide and glycosylated 24,25-dihydroxycucurbitadienol are less than about 40 by weight.
2. The method of claim 1, wherein the biosynthetic pathway comprises squalene epoxidase (SQE), which generates 2,3;22,23-squalene dioxide from squalene.
3. The method of claim 2, wherein the squalene epoxidase comprises an amino acid sequence having at least 80%, or at least 85%, at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 2, and comprises one or more mutations or groups of mutations listed in Tables 1, 2, and / or 3.
4. The method of claim 3, wherein the squalene epoxidase comprises one or more substitutions relative to SEQ ID NO: 2 at positions selected from the following, optionally wherein: (i) S116 is replaced by charged amino acids selected from Glu and Asp; (ii) N134 is replaced by a nonpolar amino acid selected from Gly, Ala, Ile, Leu and Val; (iii) I145 is replaced by charged amino acids selected from His, Arg and Lys; (iv) A164 is replaced by a polar amino acid selected from Asn and Gln. (v) A169 is replaced by a charged amino acid selected from Arg, Lys, Glu, Asp, or His; and / or (vi) A400 is replaced by nonpolar amino acids selected from Val, Leu and Ile.
5. The method of claim 4, wherein the squalene epoxidase comprises one or more mutations selected from S116E, N134G, I145H, A164N, A169R and A400V relative to SEQ ID NO:
2.
6. The method of claim 2, wherein the squalene epoxidase comprises an amino acid sequence having at least 80%, or at least 85%, at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 77, and comprises one or more mutations or groups of mutations listed in Tables 4 and / or 5.
7. The method of claim 6, wherein the squalene epoxidase comprises one or more substitutions relative to SEQ ID NO: 77 at positions selected from the following, optionally wherein: (i) S366 is replaced by charged amino acids selected from Arg, Lys and His; (ii) D41 is replaced by positively charged amino acids selected from His, Lys and Arg; (iii) M315 is replaced by nonpolar amino acids selected from Leu, Ile and Val.
8. The method of claim 7, wherein the squalene epoxidase comprises one or more mutations selected from S366R, D41H and M315L relative to SEQ ID NO:
77.
9. The method according to any one of claims 1 to 8, wherein the biosynthetic pathway comprises 24,25-epoxycucurbitadienol synthase (ECDS), which generates 24,25-epoxycucurbitadienol from 2,3;22,23-squalene dioxide.
10. The method of claim 9, wherein the ECDS enzyme has a substrate preference for 2,3;22,23-squalene dioxide relative to 2,3-squalene oxide.
11. The method of claim 10, wherein the ECDS enzyme comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% sequence identity with respect to SEQ ID NO: 7, and having one or more mutations relative to SEQ ID NO: 7, said one or more mutations increasing enzyme productivity or substrate preference for 2,3;22,23-squalene dioxide relative to 2,3-squalene oxide.
12. The method of claim 11, wherein the ECDS enzyme comprises one or more mutations or mutation groups listed in Tables 7, 8, 9 and / or 10.
13. The method of claim 11 or claim 12, wherein the ECDS comprises one or more mutations relative to SEQ ID NO:7 at one or more positions selected from S24, C35, D50, N121, L245, W331, Q401, I490, I553 and A556, optionally wherein: (i) S24 is replaced by a polar amino acid selected from Asn and Gln; (ii) C35 is replaced by charged amino acids selected from Asp and Glu; (iii) D50 is replaced by Glu; (iv) N121 was replaced by His, Lys and Arg. (v) L245 was replaced by Ile or Val; (vi) W331 was replaced by Ser, Thr or Tyr; (vii) Q401 was replaced by Ala, Gly or Val; (viii) I490 was replaced by Val, Leu, or Ala; (ix) The I553 was replaced by the Met; and / or (x) A556 is replaced by polar amino acids selected from Ser, Thr and Tyr.
14. The method of claim 13, wherein the ECDS comprises one or more mutations selected from S24N, C35D, D50E, N121H, L245I, W331S, Q401A, I490V, I553M and A556S relative to SEQ ID NO:
7.
15. The method according to any one of claims 9 to 14, wherein the ECDS enzyme comprises one or more mutations in a substrate entry channel that improves substrate selectivity for 2,3;22,23-squalene relative to 2,3-squalene oxide.
16. The method according to any one of claims 1 to 15, wherein the biosynthetic pathway comprises epoxide hydrolase (EPH) that generates 24,25-dihydroxy-cucurbitadiol from 24,25-epoxycucurbitadiol.
17. The method of claim 16, wherein the EPH enzyme has a substrate preference for 24,25-epoxycucurbitadienol relative to 2,3;22,23-squalene dioxide.
18. The method of claim 17, wherein the EPH enzyme comprises the following amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% sequence identity with respect to SEQ ID NO:
5.
19. The method of claim 18, wherein the EPH has one or more mutations relative to SEQ ID NO: 5, the one or more mutations increasing enzyme productivity or substrate preference for 24,25-epoxycucurbitadienol relative to 2,3;22,23-squalene dioxide, and wherein one or more of the mutations are selected from Table 6.
20. The method of claim 19, wherein the epoxide hydrolase comprises one or more mutations relative to SEQ ID NO: 5 at one or more positions selected from S68, Q128, G141, T144, A145, E163, E191, L262, C295, N299, optionally wherein: (i) S68 is replaced by nonpolar amino acids selected from Ala, Gly and Val; (ii) Q128 is replaced by a nonpolar amino acid selected from Ala, Pro, Gly, and Val; (iii) G141 is replaced by polar amino acids selected from Ser, Thr and Tyr; (iv) T144 is replaced by an amino acid selected from Gln, Asn, Ala, Gly and Val; (v) A145 is replaced by nonpolar amino acids selected from Val, Ile, and Leu; (vi) E163 is replaced by nonpolar amino acids selected from Ala, Gly, and Val; (vii) E191 is replaced by nonpolar amino acids selected from Gly, Ala and Val; (viii) L262 was replaced by Met; (ix) C295 is replaced by a nonpolar amino acid selected from Gly, Ala, and Val; and / or (x) N299 is replaced by Gln.
21. The method of claim 20, wherein the epoxide hydrolase comprises one or more mutations selected from S68A, Q128A, G141S, T144Q, A145V, E163A, E191G, L262M, C295G and N299Q relative to SEQ ID NO:
5.
22. The method according to any one of claims 1 to 21, wherein the biosynthetic pathway comprises a cytochrome P450 enzyme that produces mogroside from 24,25-dihydroxycucurbitadienol.
23. The method of claim 22, wherein the cytochrome P450 enzyme selectively hydroxylates the C11 of 24,25-dihydroxycucurbitadienol.
24. The method of claim 23, wherein the cytochrome P450 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 14 or SEQ ID NO:
15.
25. The method of claim 24, wherein the cytochrome P450 enzyme has all or part of the natural transmembrane domain replaced by a transmembrane domain from a C-terminal protein of the Escherichia coli inner membrane cytoplasm, wherein the Escherichia coli protein is optionally SohB or FliO.
26. The method of claim 24 or 25, wherein the cytochrome P450 enzyme comprises a mutation or mutation group listed in Tables 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28 and / or 29.
27. The method according to any one of claims 24 to 26, wherein the cytochrome P450 enzyme comprises a substitution relative to SEQ ID NO: 15 at one or more positions selected from A12, I27, S63, K75, V76, N102, G125, T128, W130, L131, K132, T192, S193, L218, G219, T223, V228, T231, T232, Y233, N234, K246, F354, K369, H433, F452, V465, A468, and T482, wherein the one or more substitutions are optionally selected from: (i) A12 is replaced by Trp, Tyr or Phe; (ii) I27 was replaced by Gly or Ala; (iii) S63 is replaced by Asn, Phe, Gln, Tyr or Trp; (iv) The K75 was replaced by the Arg; (v) V76 was replaced by Met, Leu, or Ile; (vi) N102 is replaced by His, Lys or Arg; (vii) G125 is replaced by amino acids selected from Ala, Leu, Ile and Val; (viii) T128 is replaced by an amino acid selected from Gly, Ser, Pro and Ala; (ix) W130 is replaced by amino acids selected from Asn, Gly, Ser, Leu, Gln, Ala, Val and Thr; (x) L131 is replaced by amino acids selected from Val, Ile, Pro, Thr and Ser; (xi) K132 is replaced by amino acids selected from Asn, Ser, Thr and Gln; (xii) T192 was replaced by Leu, Ile or Val; (xiii) S193 was replaced by Thr; (xiv) L218 was replaced by Ile, Val and Ala; (xv) G219 is replaced by amino acids selected from Ala, Gln, Arg, Val, Asn, His, and Lys; (xvi) T223 was replaced by Ser; (xvii) V228 was replaced by Ile or Leu; (xviii) T231 is replaced by Phe, Trp or Tyr; (xix) T232 was replaced by Ala, Gly, or Val; (xx) Y233 was replaced by Phe; (xxi) N234 is replaced by His, Arg or Lys; (xxii) K246 is replaced by Val, Ile or Leu; (xxiii) F354 was replaced by Tyr or Trp; (xxiv) K369 was replaced by Arg, Lys or His; (xxv) H433 is replaced by Phe, Trp or Tyr; (xxvi) F452 was replaced by Ile, Val, or Leu; (xxvii) V465 is replaced by an amino acid selected from Ile, Met or Leu; (xxviii) A468 is replaced by an amino acid selected from Thr, Ser, Asn, and Gln; and (xxix) T482 was replaced by Ser.
28. The method of claim 27, wherein the cytochrome P450 enzyme comprises one or more substitutions selected from A12W, I27G, S63F, S63N, K75R, V76M, N102H, G125A, T128G, W130N, L131I, K132N, T192L, S193T, L218I, G219A, T223S, V228I, T231F, T232A, Y233F, N234H, K246V, F354Y, K369R, H433F, F452I, V465I, A468T, and T482S relative to SEQ ID NO:
15.
29. The method according to any one of claims 24 to 28, wherein the cytochrome P450 enzyme comprises an N-terminal deletion of at least about 20, or at least about 25, or at least about 30, or at least about 35, or at least about 40, or at least about 45, or at least about 50 amino acids relative to SEQ ID NO:
15.
30. The method of claim 29, wherein the cytochrome P450 enzyme comprises a deletion selected from the L3 to H29 deletions relative to SEQ ID NO:
15.
31. The method of claim 30, wherein the cytochrome P450 enzyme comprises the following amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 85, and comprises one or more mutations or groups of mutations listed in Tables 26, 27, 28 and / or 29.
32. The method of claim 31, wherein the cytochrome P450 enzyme comprises a substitution at one or more positions selected from K8, D9, S10, N13, V15, I45, K46, K47, M49, K50, R51, I191, A192, F406, E411, Y412, and S413 relative to SEQ ID NO:85, wherein the one or more substitutions are optionally selected from: (i) K8 is replaced by Arg; (ii) D9 is replaced by Asn or Gln; (iii) S10 is replaced by Arg, Lys or His; (iv) N13 is replaced by Lys, Arg or His; (v) V15 is replaced by Lys, Arg or His; (vi) I45 was replaced by Val, Ile, or Leu; (vii) K46 was replaced by Arg, Lys or His; (viii) K47 is replaced by Glu or Asp; (ix) M49 was replaced by Val, Ile, or Leu; (x) K50 is replaced by Glu or Asp; (xi) R51 was replaced by Lys; (xii) I191 was replaced by Leu or Val; (xiii) A192 is replaced by Lys, Arg or His; (xiv) F406 was replaced by Tyr or Trp; (xv) E411 was replaced by Asp; (xvi) Y412 was replaced by Phe, Leu, Ile, or Trp; and (xvii) S413 is replaced by Ala, Gly or Val.
33. The method of claim 32, wherein the cytochrome P450 enzyme comprises one or more substitutions relative to SEQ ID NO:85 selected from K8R, D9N, S10R, N13K and V15K, I45V, K46R, K47E, M49V, K50E, R51K, F406Y, E411D, Y412F, S413A, I191L and A192K.
34. The method according to any one of claims 22 to 33, wherein the microbial cells express P450 reductase.
35. The method of claim 34, wherein the P450 reductase comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with SEQ ID NO: 22 or SEQ ID NO:
24.
36. The method of claim 35, wherein the cytochrome P450 reductase comprises replacing all or part of the native transmembrane domain with a transmembrane domain from a C-terminal protein of the inner membrane cytoplasm of *E. coli*, wherein the *E. coli* protein is selected from YcgG, yhcB, zipA, waaA, sohB, djlA, lpxK, ypfN, and yhhM.
37. The method of claim 36, wherein the *E. coli* protein is YcgG .
38. The method according to claim 36 or 37, wherein about 60 to about 80 amino acids at the N-terminus of SEQ ID NO: 22 or SEQ ID NO: 24 are deleted, and optionally about 65 to about 78 amino acids at the N-terminus of SEQ ID NO: 22 or SEQ ID NO: 24 are deleted.
39. The method according to claim 37 or 38, wherein: The natural N-terminus of SEQ ID NO: 22 or SEQ ID NO: 24 is derived from YcgG The N-terminus of the amino acid is replaced by 20 to 30 amino acids, and optionally replaced by amino acids from... YcgG The 25 amino acid substitutions at the N-terminus.
40. The method according to any one of claims 37 to 39, wherein the cytochrome P450 reductase comprises one or more mutations relative to SEQ ID NO: 22 listed in Table 31, or comprises one or more mutations relative to SEQ ID NO: 24 listed in Table 32 or 33, and optionally a substitution of the L90 residue relative to SEQ ID NO:
24.
41. The method according to any one of claims 1 to 40, wherein the biosynthetic pathway comprises one or more uridine diphosphate-dependent glycosyltransferases (UGTs) that produce one or more mogrosides from mogrool, wherein the one or more mogrosides are selected from: Mog 1A, Mog. 1E, Mog. IIA1, Mog. IIE, Mog. IIA2, Mog. IIIA1, Mog. III, Mog. IIIA2, sarcosinolate, Mog. IVA, Mog. IV, Mog. V and iso-Mog. V, and optionally wherein the mogroside is substantially Mog. V.
42. The method of claim 41, wherein the MogV produced by the UGT enzyme is at least twice as much as the glycosylated product of 2,25-dihydroxycucurbitadienol.
43. The method of claim 42, wherein the MogV produced by the UGT enzyme is at least 3 times, at least 5 times, or at least 10 times more than the glycosylated product of 24,25-dihydroxycucurbitadienol.
44. The method according to any one of claims 1 to 43, wherein the biosynthetic pathway comprises a UGT enzyme (UGTc3) for producing C-3 mogroside from mogroside.
45. The method of claim 44, wherein the UGTc3 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 27, wherein the UGT enzyme optionally comprises one or more mutations relative to SEQ ID NO: 27, the one or more mutations resulting in higher specificity and / or activity for glycosylation at the C-3 position compared to a UGT enzyme having the amino acid sequence of SEQ ID NO:
27.
46. The method of claim 45, wherein one or more mutations are selected from Table 35 and / or Table 36.
47. The method of claim 46, wherein the UGTc3 enzyme comprises one or more mutations relative to SEQ ID NO: 27 at one or more positions selected from L41, D49, T74, C127, F303, and A307, optionally wherein the mutation is selected from: (i) L41 is replaced by aromatic amino acids selected from Phe, Tyr and Trp; (ii) D49 was replaced by Glu; (iii) The T74 was replaced by the Met; (iv) C127 is replaced by an aromatic amino acid selected from Phe, Tyr and Trp; (v) The F303 was replaced by the Tyr; and (vi) A307 is replaced by an amino acid selected from Ile, Leu and Val.
48. The method of claim 47, wherein the UGTc3 enzyme comprises one or more of the following substitutions relative to SEQ ID NO: 27: L41F, D49E, T74M, C127F, F303Y, and A307I.
49. The method of claim 44, wherein the UGTc3 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 60 or SEQ ID NO: 65, wherein the UGTc3 enzyme optionally comprises one or more mutations relative to SEQ ID NO: 60 or SEQ ID NO: 65, wherein the one or more mutations result in higher specificity and / or activity for glycosylation at the C-3 position compared to a UGT enzyme having the amino acid sequence of SEQ ID NO: 60 or SEQ ID NO:
65.
50. The method of claim 49, wherein one or more mutations are selected from Table 37 relative to SEQ ID NO: 60, and / or Table 39 relative to SEQ ID NO:
65.
51. The method of claim 49 or claim 50, wherein the UGTc3 enzyme is a cyclic variant relative to SEQ ID NO:
60.
52. The method of claim 51, wherein the cyclic variant has a cleavage site corresponding to amino acids 150 to 180, or amino acids 155 to 180, or amino acids 160 to 180, or amino acids 165 to 180, or amino acids 169 to 180 of SEQ ID NO: 60, and optionally has a linker sequence located between amino acids corresponding to the N-terminal and C-terminal residues of SEQ ID NO: 60, and wherein the linker optionally has a length of 2 to 25 amino acids.
53. The method of claim 52, wherein the UGTc3 enzyme comprises one or more mutations relative to SEQ ID NO: 65 at one or more positions selected from H40, I216, S241, and L352, optionally wherein the mutations are selected from: (i) H40 is replaced by Asp, Gly, Gln, Ala or Val; (ii) I216 is replaced by Ala, Gly, or Val; (iii) S241 is replaced by Gly, Ala, or Val; and (iv) L352 was replaced by Met.
54. The method of claim 53, wherein the UGTc3 enzyme comprises one or more of the following substitutions relative to SEQ ID NO: 65: H40N, I216A, S241G, and L352M.
55. The method according to any one of claims 1 to 54, wherein the biosynthetic pathway comprises a UGT enzyme (UGTc24) for producing C-24 mogroside from mogroside.
56. The method of claim 55, wherein the UGTc24 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 31, wherein the UGTc24 enzyme optionally comprises one or more mutations that result in higher specificity and / or activity for glycosylation at the C-24 position compared to a UGT enzyme having the amino acid sequence of SEQ ID NO:
31.
57. The method of claim 56, wherein one or more mutations are selected from Tables 40, 41, 42, 43, 44 and / or 45.
58. The method according to claim 56 or 57, wherein the UGTc24 enzyme comprises one or more mutations relative to SEQ ID NO: 31 at positions selected from L18, A74, S88, I91, T95, H101, T184, F187, P190, T273, N306, M332, and Y381, wherein the one or more mutations are optionally selected from: (i) L18 was replaced by Met; (ii) I63 was replaced by Val, Ala or Gly; (iii) A74 is replaced by an acidic amino acid selected from Glu and Asp; (iv) S88 is replaced by an amino acid selected from Ala, Gly, Leu, Val and Ile; (v) I91 is replaced by aromatic amino acids selected from Phe, Tyr and Trp; (vi) T95 is replaced by amino acids selected from Ala, Gly and Val; (vii) H101 was replaced by Pro; (viii) T184 is replaced by an amino acid selected from Phe, Tyr and Trp; (ix) The F187 was replaced by the Tyr or Trp. (x) P190 is replaced by amino acids selected from Glu and Asp; (xi) T273 is replaced by Ser or is missing; (xii) N306 is replaced by Lys, His or Arg; (xiii) M332 is replaced by an amino acid selected from Leu, Ile, Val, and Ala; and (xiv) Y381 is replaced by amino acids selected from Phe and Trp.
59. The method of claim 58, wherein the UGTc24 enzyme comprises one or more mutations selected from L18M, I63V, A74E, S88A, I91F, T95A, H101P, T184F, F187Y, P190E, T273S, N306K, M332L and Y381F relative to SEQ ID NO:
31.
60. The method according to any one of claims 1 to 59, wherein the biosynthetic pathway comprises a uridine diphosphate-dependent glycosyltransferase (UGT) (UGT1216) that produces 1,6- and / or 1,2-branched glycosides of mogroside C-3 and / or C-24 glycosides.
61. The method of claim 60, wherein the UGT1216 is a circular variant of the enzyme represented by SEQ ID NO: 35 or a sequence having at least 70% identity with it.
62. The method of claim 61, wherein the cyclic variant has a cleavage site at positions 180 to 195 of amino acids corresponding to or having at least 70% identity with SEQ ID NO: 35, and Optionally, a cleavage site is provided at positions 185 to 190 corresponding to SEQ ID NO: 35 or a sequence having at least 70% identity with it, and Optionally, a cleavage site is provided at position 187 of SEQ ID NO: 35 or a sequence having at least 70% identity with it.
63. The method of claim 62, wherein the UGT1216 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 36, wherein the UGT1216 enzyme has one or more mutations that result in higher specificity and / or activity for 1,6 and / or 1,2 glycosylation of mogrosides C-3 and / or C-24 glycosides compared to a UGT enzyme having the amino acid sequence of SEQ ID NO:
36.
64. The method of claim 63, wherein the UGT1216 enzyme has one or more mutations that result in the production of more MogV from MogIIE compared to a UGT enzyme having the amino acid sequence of SEQ ID NO:
36.
65. The method of claim 64, wherein one or more mutations or groups of mutations of UGT1216 are selected from Tables 46, 47, 48 and / or 49.
66. The method according to claim 64 or 65, wherein the UGT1216 enzyme comprises a mutation relative to SEQ ID NO: 36 at one or more positions selected from L14, H19, W160, L266, T281, S318, and H390, optionally wherein: (i) L14 is replaced by nonpolar amino acids selected from Val, Ile, Ala and Gly; (ii) H19 is replaced by a charged amino acid selected from Glu, Lys, Asp and Glu; (iii) The W160 was replaced by Pro, Ser, or Thr; (iv) L266 is replaced by polar amino acids selected from Gln and Asn; (v) T281 was replaced by Ser; (vi) S318 is replaced by an amino acid selected from Lys, Pro, and Arg; and (vii) H390 is replaced by amino acids selected from Asp and Glu.
67. The method of claim 66, wherein the UGT1216 enzyme comprises one or more substitutions selected from L14V, H19E, W160P, L266Q, T281S, S318K and H390D relative to SEQ ID NO:
36.
68. The method of claim 67, wherein the UGT1216 enzyme comprises mutant substitutions of L14V, W160P, L266Q and H390D relative to SEQ ID NO:
36.
69. The method of claim 62, wherein the UGT1216 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 39, wherein the UGT1216 enzyme has one or more mutations that result in higher specificity and / or activity for 1,6 and 1,2 branched glycosylation of mogrosides C-3 and / or C-24 glycosides compared to a UGT enzyme having the amino acid sequence of SEQ ID NO:
39.
70. The method of claim 69, wherein the UGT1216 enzyme has one or more mutations that, compared with the UGT enzyme having the amino acid sequence of SEQ ID NO: 39, result in the production of more MogIVA from MogIIE or more MogV from MogIIE.
71. The method of claim 70, wherein the UGT1216 enzyme comprises one or more mutations listed in Table 50.
72. The method according to claim 70 or 71, wherein the UGT enzyme comprises a mutation relative to SEQ ID NO: 39 at a position selected from T147 and N207, optionally wherein: (i) T147 is replaced by a nonpolar amino acid selected from Leu, Ile, and Val; and / or (ii) N207 is replaced by a basic amino acid selected from Lys, Arg and His.
73. The method of claim 72, wherein the UGT enzyme comprises a T147L and / or N207K substitution relative to SEQ ID NO:
39.
74. The method according to any one of claims 1 to 73, wherein the microbial cells express squalene synthase (SQS).
75. The method of claim 74, wherein the microbial cells express farnesyl pyrophosphate (FPP) synthase.
76. The method according to claim 74 or 75, wherein the microbial cells overexpress one or more enzymes of the MEP or MVA pathway.
77. The method of claim 76, wherein the microbial cell is a bacterium, and optionally selected from species of Escherichia, Bacillus, Corynebacterium, Rhodophyton, Fermentomonas, Vibrio, and Pseudomonas, and optionally selected from species of Escherichia coli, Bacillus subtilis, Corynebacterium glutamicum, Rhodophyton capsulatum, Rhodophyton floccosum, Fermentomonas motilityis, Vibrio natriureticus, and Pseudomonas putida.
78. The method according to claim 77, wherein the microbial cell is a yeast cell, optionally selected from species of the genera *Saccharomyces*, *Pichia*, or *Yersinia*, and optionally *Saccharomyces cerevisiae*, *Pichia*, and *Yersinia lipolytica*.
79. The method according to any one of claims 1 to 78, wherein the microbial cell expresses recombinant... pntAB Or its variants.
80. The method according to any one of claims 1 to 79, wherein the microbial cell expresses at least one recombinant electron carrier.
81. The method according to claim 80, wherein the at least one recombinant electron carrier is recombinant ferroredoxin (fdx) or recombinant flavin redoxin (Fld).
82. The method according to any one of claims 1 to 81, wherein the microbial cells produce more than about 40%, or more than about 50%, or more than about 60%, or more than about 70%, or more than about 80% of the in-path products 2,3-squalene oxide, 2,3;22,23-squalene dioxide, 24,25-epoxycucurbitadienol, 24,25-dihydroxycucurbitadienol, mogroside, and mogroside.
83. The method according to any one of claims 1 to 82, wherein the microbial cells produce less than about 35%, or less than about 30%, or less than about 25%, or less than about 20% of the extra-pathway product cucurbitadienol, hydrolyzed 2,3;22,23-squalene dioxide, and glycosylated 24,25-dihydroxycucurbitadienol.
84. The method according to any one of claims 1 to 83, wherein the microbial cell contains a mutation that reduces the amount or activity of the transport protein.
85. The method of claim 84, wherein the transporter is an ABC transporter.
86. The method according to claim 84 or claim 85, wherein the transporter protein is YbhGFSR or its ortholog.
87. The method of claim 86, wherein the microbial cells are... ybhG , ybhF , ybhS and ybhR One or more of them contain sub-effective mutations or null mutations.
88. The method of claim 87, wherein the invalid mutation is ybhG , ybhF , ybhS and ybhR One or more of them are completely or partially missing.
89. The method of claim 87, wherein the meta-effect mutation is a promoter mutation, and / or a mutation selected from the group consisting of: mutated ribosome binding sites, in ybhG , ybhF , ybhS and ybhR One or more rare codons are added to one or more of the RNA secondary structures, and one or more mutations are added.
90. A method for producing mogroside or mogroside, the method comprising: Provides microbial cells that produce squalene and express a biosynthetic pathway for converting squalene into mogroside or mogroside. The microbial cells contain one or more gene modifications that reduce the efflux of intermediates in mogroside synthesis, optionally wherein the intermediate is 2,3;22,23-squalene dioxide. Compared to cells without the aforementioned gene modification, the microbial cells exhibited increased mogroside synthesis.
91. The method of claim 90, wherein the gene modification comprises a mutation that reduces the amount or activity of the transporter protein.
92. The method according to claim 90 or claim 91, wherein the transporter protein is YbhGFSR or its ortholog.
93. The method according to claim 92, wherein the microbial cells are... ybhG , ybhF , ybhS and ybhR One or more of them contain sub-effective mutations or null mutations.
94. The method of claim 93, wherein the invalid mutation is ybhG , ybhF , ybhS and ybhR One or more of them are completely or partially missing.
95. The method of claim 93, wherein the meta-effect mutation is a promoter mutation, and / or a mutation selected from the group consisting of: mutated ribosome binding sites; ybhG , ybhF , ybhS and ybhR One or more rare codons are added to one or more of the RNA secondary structures, and one or more mutations are added.
96. The method according to any one of claims 90 to 95, wherein the microbial cell is Escherichia coli.
97. The method according to any one of claims 90 to 96, wherein the biosynthetic pathway comprises one or more enzymes disclosed herein.
98. A microbial cell that produces squalene and expresses a biosynthetic pathway for converting squalene into mogroside or mogroside, wherein when said microbial cell is cultured under conditions suitable for producing said mogroside or mogroside, said microbial cell produces squalene derivatives, wherein the extra-pathway products cucurbitadienol, hydrolyzed 2,3;22,23-squalene dioxide, and glycosylated 24,25-dihydroxycucurbitadienol are less than about 40 by weight.
99. The microbial cell of claim 98, wherein the biosynthetic pathway comprises squalene epoxidase (SQE), the squalene epoxidase producing 2,3;22,23-squalene dioxide from squalene.
100. The microbial cell of claim 99, wherein the squalene epoxidase comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 2, and comprises one or more mutations or groups of mutations listed in Tables 1, 2, and 3.
101. The microbial cell of claim 100, wherein the squalene epoxidase comprises one or more substitutions relative to SEQ ID NO: 2 at positions selected from the following, optionally wherein: (i) S116 is replaced by charged amino acids selected from Glu and Asp; (ii) N134 is replaced by a nonpolar amino acid selected from Gly, Ala, Ile, Leu and Val; (iii) I145 is replaced by charged amino acids selected from His, Arg and Lys; (iv) A164 is replaced by a polar amino acid selected from Asn and Gln. (v) A169 is replaced by a charged amino acid selected from Arg, Lys, Glu, Asp, or His; and / or (vi) A400 is replaced by nonpolar amino acids selected from Val, Leu and Ile.
102. The microbial cell of claim 101, wherein the squalene epoxidase comprises one or more mutations selected from S116E, N134G, I145H, A164N, A169R and A400V relative to SEQ ID NO:
2.
103. The microbial cell of claim 99, wherein the squalene epoxidase comprises an amino acid sequence having at least 80%, or at least 85%, at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 77, and comprises one or more mutations or groups of mutations listed in Tables 4 and / or 5.
104. The microbial cell of claim 103, wherein the squalene epoxidase comprises one or more substitutions relative to SEQ ID NO: 77 at positions selected from the following, optionally wherein: (i) S366 is replaced by charged amino acids selected from Arg, Lys and His; (ii) D41 is replaced by positively charged amino acids selected from His, Lys and Arg; (iii) M315 is replaced by nonpolar amino acids selected from Leu, Ile and Val.
105. The microbial cell of claim 104, wherein the squalene epoxidase comprises one or more mutations selected from S366R, D41H and M315L relative to SEQ ID NO:
77.
106. The microbial cell according to any one of claims 98 to 105, wherein the biosynthetic pathway comprises 24,25-epoxycucurbitadienol synthase (ECDS), said 24,25-epoxycucurbitadienol synthase producing 24,25-epoxycucurbitadienol from 2,3;22,23-squalene dioxide.
107. The microbial cell of claim 106, wherein the ECDS enzyme has a substrate preference for 2,3;22,23-squalene dioxide relative to 2,3-squalene oxide.
108. The microbial cell of claim 107, wherein the ECDS enzyme comprises the following amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% sequence identity with respect to SEQ ID NO: 7, and having one or more mutations relative to SEQ ID NO: 7, said one or more mutations increasing enzyme productivity or substrate preference for 2,3;22,23-squalene dioxide relative to 2,3-squalene oxide.
109. The microbial cell of claim 108, wherein the ECDS enzyme comprises one or more mutations or mutation groups listed in Tables 7, 8, 9 and / or 10.
110. The microbial cell according to claim 108 or claim 109, wherein the ECDS comprises one or more mutations relative to SEQ ID NO: 7 at one or more positions selected from S24, C35, D50, N121, L245, W331, Q401, I490, I553 and A556, optionally wherein: (i) S24 is replaced by a polar amino acid selected from Asn and Gln; (ii) C35 is replaced by charged amino acids selected from Asp and Glu; (iii) D50 is replaced by Glu; (iv) N121 was replaced by His, Lys and Arg. (v) L245 was replaced by Ile or Val; (vi) W331 was replaced by Ser, Thr or Tyr; (vii) Q401 was replaced by Ala, Gly or Val; (viii) I490 was replaced by Val, Leu, or Ala; (ix) The I553 was replaced by the Met; and / or (x) A556 is replaced by polar amino acids selected from Ser, Thr and Tyr.
111. The microbial cell of claim 110, wherein the ECDS comprises one or more mutations relative to SEQ ID NO: 7 selected from S24N, C35D, D50E, N121H, L245I, W331S, Q401A, I490V, I553M and A556S.
112. The microbial cell according to any one of claims 106 to 111, wherein the ECDS enzyme comprises one or more mutations in a substrate entry channel that improves substrate selectivity for 2,3;22,23-squalene relative to 2,3-squalene oxide.
113. The microbial cell according to any one of claims 98 to 112, wherein the biosynthetic pathway comprises epoxide hydrolase (EPH) that produces 24,25-dihydroxy-cucurbitadiol from 24,25-epoxycucurbitadiol.
114. The microbial cell of claim 113, wherein the EPH enzyme has a substrate preference for 24,25-epoxycucurbitadienol relative to 2,3;22,23-squalene dioxide.
115. The microbial cell of claim 114, wherein the EPH enzyme comprises the following amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% sequence identity with respect to SEQ ID NO: 5, optionally wherein the EPH has one or more mutations relative to SEQ ID NO: 5, the one or more mutations increasing enzyme productivity or substrate preference for 24,25-epoxycucurbitadienol relative to 2,3;22,23-squalene dioxide, and wherein one or more of the mutations are selected from Table 6.
116. The microbial cell of claim 115, wherein the epoxide hydrolase comprises one or more mutations relative to SEQ ID NO: 5 at one or more positions selected from S68, Q128, G141, T144, A145, E163, E191, L262, C295, N299, optionally wherein: (i) S68 is replaced by nonpolar amino acids selected from Ala, Gly and Val; (ii) Q128 is replaced by a nonpolar amino acid selected from Ala, Pro, Gly, and Val; (iii) G141 is replaced by polar amino acids selected from Ser, Thr and Tyr; (iv) T144 is replaced by an amino acid selected from Gln, Asn, Ala, Gly and Val; (v) A145 is replaced by nonpolar amino acids selected from Val, Ile, and Leu; (vi) E163 is replaced by nonpolar amino acids selected from Ala, Gly, and Val; (vii) E191 is replaced by nonpolar amino acids selected from Gly, Ala and Val; (viii) L262 was replaced by Met; (ix) C295 is replaced by a nonpolar amino acid selected from Gly, Ala, and Val; and / or (x) N299 is replaced by Gln.
117. The microbial cell of claim 116, wherein the epoxide hydrolase comprises one or more mutations selected from S68A, Q128A, G141S, T144Q, A145V, E163A, E191G, L262M, C295G and N299Q relative to SEQ ID NO:
5.
118. The microbial cell according to any one of claims 98 to 117, wherein the biosynthetic pathway comprises a cytochrome P450 enzyme that produces mogroside from 24,25-dihydroxycucurbitadienol.
119. The microbial cell of claim 118, wherein the cytochrome P450 enzyme selectively hydroxylates the C11 of 24,25-dihydroxycucurbitadienol, optionally wherein the cytochrome P450 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 14 or SEQ ID NO:
15.
120. The microbial cell of claim 119, wherein the cytochrome P450 enzyme has all or part of the natural transmembrane domain replaced by a transmembrane domain derived from a C-terminal protein of the Escherichia coli inner membrane cytoplasm, wherein the Escherichia coli protein is optionally SohB or FliO.
121. The microbial cell of claim 119 or 120, wherein the cytochrome P450 enzyme comprises a mutation or mutation group listed in Tables 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28 and / or 29.
122. The microbial cell according to any one of claims 118 to 121, wherein the cytochrome P450 enzyme comprises a substitution relative to SEQ ID NO: 15 at one or more positions selected from A12, I27, S63, K75, V76, N102, G125, T128, W130, L131, K132, T192, S193, L218, G219, T223, V228, T231, T232, Y233, N234, K246, F354, K369, H433, F452, V465, A468, and T482, wherein the one or more substitutions are optionally selected from: (i) A12 is replaced by Trp, Tyr or Phe; (ii) I27 was replaced by Gly or Ala; (iii) S63 is replaced by Asn, Phe, Gln, Tyr or Trp; (iv) The K75 was replaced by the Arg; (v) V76 was replaced by Met, Leu, or Ile; (vi) N102 is replaced by His, Lys or Arg; (vii) G125 is replaced by amino acids selected from Ala, Leu, Ile and Val; (viii) T128 is replaced by an amino acid selected from Gly, Ser, Pro and Ala; (ix) W130 is replaced by amino acids selected from Asn, Gly, Ser, Leu, Gln, Ala, Val and Thr; (x) L131 is replaced by amino acids selected from Val, Ile, Pro, Thr and Ser; (xi) K132 is replaced by amino acids selected from Asn, Ser, Thr and Gln; (xii) T192 was replaced by Leu, Ile or Val; (xiii) S193 was replaced by Thr; (xiv) L218 was replaced by Ile, Val and Ala; (xv) G219 is replaced by amino acids selected from Ala, Gln, Arg, Val, Asn, His, and Lys; (xvi) T223 was replaced by Ser; (xvii) V228 was replaced by Ile or Leu; (xviii) T231 is replaced by Phe, Trp or Tyr; (xix) T232 was replaced by Ala, Gly, or Val; (xx) Y233 was replaced by Phe; (xxi) N234 is replaced by His, Arg or Lys; (xxii) K246 is replaced by Val, Ile or Leu; (xxiii) F354 was replaced by Tyr or Trp; (xxiv) K369 was replaced by Arg, Lys or His; (xxv) H433 is replaced by Phe, Trp or Tyr; (xxvi) F452 was replaced by Ile, Val, or Leu; (xxvii) V465 is replaced by an amino acid selected from Ile, Met or Leu; (xxviii) A468 is replaced by an amino acid selected from Thr, Ser, Asn, and Gln; and (xxix) T482 was replaced by Ser.
123. The microbial cell of claim 122, wherein the cytochrome P450 enzyme comprises one or more substitutions relative to SEQ ID NO: 15 selected from A12W, I27G, S63F, S63N, K75R, V76M, N102H, G125A, T128G, W130N, L131I, K132N, T192L, S193T, L218I, G219A, T223S, V228I, T231F, T232A, Y233F, N234H, K246V, F354Y, K369R, H433F, F452I, V465I, A468T, and T482S.
124. The microbial cell according to any one of claims 118 to 123, wherein the cytochrome P450 enzyme comprises an N-terminal deletion of at least about 20, or at least about 25, or at least about 30, or at least about 35, or at least about 40, or at least about 45, or at least about 50 amino acids relative to SEQ ID NO:
15.
125. The microbial cell of claim 123, wherein the cytochrome P450 enzyme comprises a deletion selected from the L3 to H29 deletions relative to SEQ ID NO: 15, optionally wherein the cytochrome P450 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 85, and comprises one or more mutations or groups of mutations listed in Tables 26, 27, 28 and / or 29.
126. The microbial cell of claim 125, wherein the cytochrome P450 enzyme comprises a substitution relative to SEQ ID NO: 85 at one or more positions selected from K8, D9, S10, N13, V15, I45, K46, K47, M49, K50, R51, I191, A192, F406, E411, Y412, and S413, wherein the one or more substitutions are optionally selected from: (i) K8 is replaced by Arg; (ii) D9 is replaced by Asn or Gln; (iii) S10 is replaced by Arg, Lys or His; (iv) N13 is replaced by Lys, Arg or His; (v) V15 is replaced by Lys, Arg or His; (vi) I45 was replaced by Val, Ile, or Leu; (vii) K46 was replaced by Arg, Lys or His; (viii) K47 is replaced by Glu or Asp; (ix) M49 was replaced by Val, Ile, or Leu; (x) K50 is replaced by Glu or Asp; (xi) R51 was replaced by Lys; (xii) I191 was replaced by Leu or Val; (xiii) A192 is replaced by Lys, Arg or His; (xiv) F406 was replaced by Tyr or Trp; (xv) E411 was replaced by Asp; (xvi) Y412 was replaced by Phe, Leu, Ile, or Trp; and (xvii) S413 is replaced by Ala, Gly or Val.
127. The microbial cell of claim 126, wherein the cytochrome P450 enzyme comprises one or more substitutions relative to SEQ ID NO: 85 selected from K8R, D9N, S10R, N13K and V15K, I45V, K46R, K47E, M49V, K50E, R51K, F406Y, E411D, Y412F, S413A, I191L and A192K.
128. The microbial cell according to any one of claims 118 to 127, wherein the microbial cell expresses P450 reductase.
129. The microbial cell of claim 128, wherein the P450 reductase comprises an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with SEQ ID NO:22 or SEQ ID NO:
24.
130. The microbial cell of claim 129, wherein the cytochrome P450 reductase comprises replacing all or part of the native transmembrane domain with a transmembrane domain derived from a C-terminal protein of the inner membrane cytoplasm of *Escherichia coli*, wherein the *E. coli* protein is selected from YcgG, yhcB, zipA, waaA, sohB, djlA, lpxK, ypfN, and yhhM.
131. The microbial cell according to claim 130, wherein the *E. coli* protein is... YcgG .
132. The microbial cell according to claim 130 or 131, wherein about 60 to about 80 amino acids at the N-terminus of SEQ ID NO: 22 or SEQ ID NO: 24 are deleted, and optionally about 65 to about 78 amino acids at the N-terminus of SEQ ID NO: 22 or SEQ ID NO: 24 are deleted.
133. The microbial cell according to claim 132, wherein: The natural N-terminus of SEQ ID NO: 22 or SEQ ID NO: 24 is derived from YcgG The N-terminus of the amino acid is replaced by 20 to 30 amino acids, and optionally by amino acids from the N-terminus of the amino acid. YcgG The 25 amino acid substitutions at the N-terminus.
134. The microbial cell according to any one of claims 128 to 133, wherein the cytochrome P450 reductase comprises one or more mutations relative to those listed in Table 31 of SEQ ID NO: 22, or comprises one or more mutations relative to those listed in Table 32 or 33 of SEQ ID NO: 24, and optionally a substitution of the L90 residue relative to SEQ ID NO:
24.
135. The microbial cell according to any one of claims 98 to 134, wherein the biosynthetic pathway comprises one or more uridine diphosphate-dependent glycosyltransferases (UGTs) that generate MogV from mogroside.
136. The microbial cell of claim 135, wherein the UGT enzyme produces at least twice as much MogV as the glycosylated product of 2,25-dihydroxycucurbitadienol.
137. The microbial cell of claim 136, wherein the MogV produced by the UGT enzyme is at least 3 times, at least 5 times, or at least 10 times more than the glycosylated product of 24,25-dihydroxycucurbitadienol.
138. The microbial cell according to any one of claims 135 to 137, wherein the biosynthetic pathway comprises a UGT enzyme (UGTc3) for producing C-3 mogroside from mogroside.
139. The microbial cell of claim 138, wherein the UGTc3 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 27, wherein the UGT enzyme optionally comprises one or more mutations relative to SEQ ID NO: 27, the one or more mutations resulting in higher specificity and / or activity for glycosylation at the C-3 position compared to a UGT enzyme having the amino acid sequence of SEQ ID NO:
27.
140. The microbial cell of claim 139, wherein one or more mutations are selected from Table 35 and / or Table 36.
141. The microbial cell of claim 140, wherein the UGTc3 enzyme comprises one or more mutations relative to SEQ ID NO:27 at one or more positions selected from L41, D49, T74, C127, F303, and A307, optionally wherein the mutation is selected from: (i) L41 is replaced by aromatic amino acids selected from Phe, Tyr and Trp; (ii) D49 was replaced by Glu; (iii) The T74 was replaced by the Met; (iv) C127 is replaced by an aromatic amino acid selected from Phe, Tyr and Trp; (v) The F303 was replaced by the Tyr; and (vi) A307 is replaced by an amino acid selected from Ile, Leu and Val.
142. The microbial cell of claim 141, wherein the UGTc3 enzyme comprises one or more of the following substitutions relative to SEQ ID NO:27: L41F, D49E, T74M, C127F, F303Y, and A307I.
143. The microbial cell of claim 142, wherein the UGTc3 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 60 or SEQ ID NO: 65, wherein the UGT enzyme optionally comprises one or more mutations relative to SEQ ID NO: 60 or SEQ ID NO: 65, wherein the one or more mutations result in higher specificity and / or activity for glycosylation at the C-3 position compared to a UGT enzyme having the amino acid sequence of SEQ ID NO: 60 or SEQ ID NO: 65, wherein optionally one or more mutations are selected from Table 24 relative to SEQ ID NO: 60, and / or Table 25 relative to SEQ ID NO:
65.
144. The microbial cell of claim 143, wherein the UGTc3 enzyme is a circular variant relative to SEQ ID NO:
60.
145. The microbial cell of claim 144, wherein the circular variant has a cleavage site corresponding to amino acids 150 to 170, or 155 to 175, or 160 to 185, or 165 to 190, or between 165 and 175 of SEQ ID NO: 60, and optionally has a linker sequence located between amino acids corresponding to the N-terminal and C-terminal residues of SEQ ID NO: 60, and wherein the linker optionally has a length of 2 to 25 amino acids.
146. The microbial cell of claim 145, wherein the UGTc3 enzyme comprises one or more mutations relative to SEQ ID NO:65 at one or more positions selected from H40, I216, S241, and L352, optionally wherein the mutation is selected from: (i) H40 is replaced by Asp, Gly, Gln, Ala or Val; (ii) I216 is replaced by Ala, Gly, or Val; (iii) S241 is replaced by Gly, Ala, or Val; and (iv) L352 was replaced by Met.
147. The microbial cell of claim 146, wherein the UGTc3 enzyme comprises one or more of the following substitutions relative to SEQ ID NO:65: H40N, I216A, S241G, and L352M.
148. The microbial cell according to any one of claims 135 to 147, wherein the biosynthetic pathway comprises a UGT enzyme (UGTc24) that produces C-24 mogroside from mogroside.
149. The microbial cell of claim 148, wherein the UGTc24 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 31, wherein the UGTc24 enzyme optionally comprises one or more mutations that result in higher specificity and / or activity for glycosylation at the C-24 position compared to a UGT enzyme having the amino acid sequence of SEQ ID NO: 31, wherein optionally one or more mutations are selected from Tables 40, 41, 42, 43, 44, and / or 45.
150. The microbial cell of claim 149, wherein the UGTc24 enzyme comprises one or more mutations relative to SEQ ID NO:31 at positions selected from L18, A74, S88, I91, T95, H101, T184, P190, T273, N306, M332, and Y381, wherein the one or more mutations are optionally selected from: (i) L18 was replaced by Met; (ii) I63 was replaced by Val, Ala or Gly; (iii) A74 is replaced by an acidic amino acid selected from Glu and Asp; (iv) S88 is replaced by an amino acid selected from Ala, Gly, Leu, Val and Ile; (v) I91 is replaced by aromatic amino acids selected from Phe, Tyr and Trp; (vi) T95 is replaced by amino acids selected from Ala, Gly and Val; (vii) H101 was replaced by Pro; (viii) T184 is replaced by an amino acid selected from Phe, Tyr and Trp; (ix) The F187 was replaced by the Tyr or Trp. (x) P190 is replaced by amino acids selected from Glu and Asp; (xi) T273 is replaced by Ser or is missing; (xii) N306 is replaced by Lys, His or Arg; (xiii) M332 is replaced by an amino acid selected from Leu, Ile, Val, and Ala; and (xiv) Y381 is replaced by amino acids selected from Phe and Trp.
151. The microbial cell of claim 150, wherein the UGTc24 enzyme comprises one or more mutations selected from L18M, I63V, A74E, S88A, I91F, T95A, H101P, T184F, F187Y, P190E, T273S, N306K, M332L and Y381F relative to SEQ ID NO:
31.
152. The microbial cell according to any one of claims 135 to 151, wherein the biosynthetic pathway comprises a uridine diphosphate-dependent glycosyltransferase (UGT) (UGT1216) that produces 1,6- and / or 1,2-branched glycosides of mogroside C-3 and / or C-24 glycosides.
153. The microbial cell of claim 152, wherein the UGT1216 is a circular variant of the enzyme represented by SEQ ID NO: 35 or a sequence having at least 70% identity with it.
154. The microbial cell of claim 153, wherein the circular arrangement variant has a cleavage site at positions 180 to 195 of amino acids corresponding to or having at least 70% identity with SEQ ID NO: 35, and Optionally, a cleavage site is provided at positions 185 to 190 corresponding to SEQ ID NO: 35 or a sequence having at least 70% identity with it, and Optionally, a cleavage site is provided at position 187 of SEQ ID NO: 35 or a sequence having at least 70% identity with it.
155. The microbial cell of claim 154, wherein the UGT1216 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 36, wherein the UGT1216 enzyme has one or more mutations that result in higher specificity and / or activity for 1,6 and / or 1,2 glycosylation of mogrosides C-3 and / or C-24 glycosides compared to a UGT enzyme having the amino acid sequence of SEQ ID NO:
36.
156. The microbial cell of claim 155, wherein the UGT1216 enzyme has one or more mutations that result in the production of more MogV from MogIIE compared to a UGT enzyme having the amino acid sequence of SEQ ID NO: 36, optionally wherein one or more mutations or groups of mutations of the UGT1216 are selected from Tables 46, 47, 48 and / or 49.
157. The microbial cell according to claim 155 or 156, wherein the UGT1216 enzyme comprises a mutation relative to SEQ ID NO: 36 at one or more sites selected from L14, H19, W160, L266, T281, S318, and H390, optionally wherein: (i) L14 is replaced by nonpolar amino acids selected from Val, Ile, Ala and Gly; (ii) H19 is replaced by a charged amino acid selected from Glu, Lys, Asp and Glu; (iii) The W160 was replaced by Pro, Ser, or Thr; (iv) L266 is replaced by polar amino acids selected from Gln and Asn; (v) T281 was replaced by Ser; (vi) S318 is replaced by an amino acid selected from Lys, Pro, and Arg; and (vii) H390 is replaced by amino acids selected from Asp and Glu.
158. The microbial cell of claim 157, wherein the UGT1216 enzyme comprises one or more substitutions selected from L14V, H19E, W160P, L266Q, T281S, S318K and H390D relative to SEQ ID NO:
36.
159. The microbial cell of claim 158, wherein the UGT1216 enzyme comprises mutant substitutions of L14V, W160P, L266Q and H390D relative to SEQ ID NO:
36.
160. The microbial cell of claim 154, wherein the UGT1216 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 30, wherein the UGT1216 enzyme has one or more mutations that result in higher specificity and / or activity for 1,6 and 1,2 branched glycosylation of mogrosides C-3 and / or C-24 glycosides compared to a UGT enzyme having the amino acid sequence of SEQ ID NO:
39.
161. The microbial cell of claim 160, wherein the UGT1216 enzyme has one or more mutations that, compared with the UGT enzyme having the amino acid sequence of SEQ ID NO: 39, result in the production of more MogIVA from MogIIE or more MogV from MogIIE.
162. The microbial cell of claim 161, wherein the UGT1216 enzyme comprises one or more mutations listed in Table 50.
163. The microbial cell of claim 161 or 162, wherein the UGT enzyme comprises a mutation relative to SEQ ID NO: 39 at a position selected from T147 and N207, optionally wherein: (i) T147 is replaced by a nonpolar amino acid selected from Leu, Ile, and Val; and / or (ii) N207 is replaced by a basic amino acid selected from Lys, Arg and His.
164. The microbial cell of claim 163, wherein the UGT enzyme comprises a T147L and / or N207K substitution relative to SEQ ID NO:
39.
165. The microbial cell according to any one of claims 98 to 164, wherein the microbial cell expresses squalene synthase (SQS).
166. The microbial cell of claim 165, wherein the microbial cell expresses farnesyl pyrophosphate (FPP) synthase.
167. The microbial cell according to claim 165 or 166, wherein the microbial cell overexpresses one or more enzymes of the MEP or MVA pathway.
168. The microbial cell of claim 167, wherein the microbial cell is a bacterium, and optionally selected from species of Escherichia, Bacillus, Corynebacterium, Rhodophyton, Fermentomonas, Vibrio, and Pseudomonas, and optionally selected from species of Escherichia coli, Bacillus subtilis, Corynebacterium glutamicum, Rhodophyton capsulatum, Rhodophyton floccosum, Fermentomonas motilityis, Vibrio natriureticus, and Pseudomonas putida.
169. The microbial cell according to claim 168, wherein the microbial cell is a yeast cell, optionally selected from species of the genera *Saccharomyces*, *Pichia*, or *Yersinia*, and optionally *Saccharomyces cerevisiae*, *Pichia*, and *Yersinia lipolytica*.
170. The microbial cell according to any one of claims 98 to 169, wherein the microbial cell expresses recombinant... pntAB Or its variants.
171. The microbial cell according to any one of claims 98 to 170, wherein the microbial cell expresses at least one recombinant electron carrier.
172. The microbial cell according to claim 171, wherein the at least one recombinant electron carrier is recombinant ferroredoxin (fdx) or recombinant flavin redoxin (Fld).
173. The microbial cell according to any one of claims 98 to 172, wherein when the microbial cell is cultured under conditions suitable for producing the mogroside or mogroside, the microbial cell produces more than about 40%, or more than about 50%, or more than about 60%, or more than about 70%, or more than about 80% of the in-path products 2,3-squalene oxide, 2,3;22,23-squalene dioxide, 24,25-epoxycucurbitadienol, 24,25-dihydroxycucurbitadienol, mogroside, and mogroside.
174. The microbial cell according to any one of claims 98 to 173, wherein when the microbial cell is cultured under conditions suitable for producing the mogroside or mogroside, the microbial cell produces less than about 35%, or less than about 30%, or less than about 25%, or less than about 20% of the extra-pathway product cucurbitadienol, hydrolyzed 2,3;22,23-squalene dioxide, and glycosylated 24,25-dihydroxycucurbitadienol.
175. The microbial cell according to any one of claims 98 to 174, wherein the microbial cell comprises a mutation that reduces the amount or activity of the transport protein.
176. The microbial cell of claim 175, wherein the transporter protein is an ABC transporter protein.
177. The microbial cell according to claim 175 or claim 176, wherein the transporter protein is YbhGFSR or its ortholog.
178. The microbial cell according to claim 177, wherein the microbial cell is in ybhG , ybhF , ybhS and ybhR One or more of them contain sub-effective mutations or null mutations.
179. The microbial cell according to claim 178, wherein the invalid mutation is ybhG , ybhF , ybhS and ybhR One or more of them are completely or partially missing.
180. The microbial cell of claim 178, wherein the sub-effect mutation is a promoter mutation, and / or a mutation selected from the group consisting of: mutated ribosome binding sites; ybhG , ybhF , ybhS and ybhR One or more rare codons are added to one or more of the RNA secondary structures, and one or more mutations are added.
181. A microbial cell for producing mogroside or mogroside, said microbial cell expressing a biosynthetic pathway that converts squalene into said mogroside or mogroside. The microbial cells contain one or more gene modifications that reduce the efflux of intermediates in mogroside synthesis, optionally wherein the intermediate is 2,3;22,23-squalene dioxide. Compared to cells without the aforementioned gene modification, the microbial cells exhibited increased mogroside synthesis.
182. The microbial cell of claim 181, wherein the genetic modification comprises a mutation that reduces the amount or activity of the transporter protein.
183. The microbial cell according to claim 181 or claim 182, wherein the transporter protein is YbhGFSR or its ortholog.
184. The microbial cell according to claim 183, wherein the microbial cell is in ybhG , ybhF , ybhS and ybhR One or more of them contain sub-effective mutations or null mutations.
185. The microbial cell according to claim 184, wherein the invalid mutation is ybhG , ybhF , ybhS and ybhR One or more of them are completely or partially missing.
186. The microbial cell of claim 184, wherein the meta-effect mutation is a promoter mutation, and / or a mutation selected from the group consisting of: mutated ribosome binding sites; ybhG , ybhF , ybhS and ybhR One or more rare codons are added to one or more of the RNA secondary structures, and one or more mutations are added.
187. The microbial cell according to any one of claims 181 to 186, wherein the microbial cell is Escherichia coli.
188. The microbial cell according to any one of claims 181 to 187, wherein the biosynthetic pathway comprises one or more enzymes disclosed herein.
189. A squalene epoxidase comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 2, and comprising one or more mutations or groups of mutations listed in Tables 1, 2 and 3.
190. The squalene epoxidase according to claim 189, wherein the amino acid sequence relative to SEQ ID NO: 2 comprises at least two, or at least three, or at least four, or at least five, or at least six, or more mutations or mutation groups listed in Tables 1, 2 and / or 3.
191. The squalene epoxidase according to claim 189, wherein the squalene epoxidase comprises one or more substitutions selected from the group consisting of the amino acid sequence of SEQ ID NO: 2: (i) S116 is replaced by charged amino acids selected from Glu and Asp; (ii) N134 is replaced by a nonpolar amino acid selected from Gly, Ala, Ile, Leu and Val; (iii) I145 is replaced by charged amino acids selected from His, Arg and Lys; (iv) A164 is replaced by a polar amino acid selected from Asn and Gln. (v) A169 is replaced by a charged amino acid selected from Arg, Lys, Glu, Asp, or His; and / or (vi) A400 is replaced by nonpolar amino acids selected from Val, Leu and Ile.
192. The squalene epoxidase according to claim 191, wherein the squalene epoxidase comprises the substitutions S116E, N134G, I145H, A164N, A169R and A400V relative to SEQ ID NO:
2.
193. A squalene epoxidase comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 77, and comprising one or more mutations or groups of mutations listed in Tables 4 and / or 5.
194. The squalene epoxidase according to claim 193, wherein the squalene epoxidase comprises one or more substitutions relative to SEQ ID NO: 77 at positions selected from the following, optionally wherein: (i) S366 is replaced by charged amino acids selected from Arg, Lys and His; (ii) D41 is replaced by positively charged amino acids selected from His, Lys and Arg; (iii) M315 is replaced by nonpolar amino acids selected from Leu, Ile and Val.
195. The squalene epoxidase according to claim 194, wherein the squalene epoxidase comprises one or more mutations selected from S366R, D41H and M315L relative to SEQ ID NO:
77.
196. A 24,25-epoxycucurbitadienol synthase (ECDS) comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 7, and comprising one or more mutations or groups of mutations listed in Tables 7, 8, 9 and 10.
197. The ECDS of claim 196, comprising at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or more mutations or mutation groups listed in Tables 7, 8, 9 and / or 10 relative to SEQ ID NO:
7.
198. The ECDS of claim 197, wherein the ECDS comprises one or more substitutions selected from: (i) S24 is replaced by a polar amino acid selected from Asn and Gln; (ii) C35 is replaced by charged amino acids selected from Asp and Glu; (iii) D50 is replaced by Glu; (iv) N121 was replaced by His, Lys and Arg. (v) L245 was replaced by Ile or Val; (vi) W331 was replaced by Ser, Thr or Tyr; (vii) Q401 was replaced by Ala, Gly or Val; (viii) I490 was replaced by Val, Leu, or Ala; (ix) The I553 was replaced by the Met; and / or (x) A556 is replaced by polar amino acids selected from Ser, Thr and Tyr.
199. The ECDS of claim 198, wherein the ECDS comprises substitutions of S24N, C35D, D50E, N121H, L245I, W331S, Q401A, I490V, I553M and A556S relative to SEQ ID NO:
7.
200. An epoxide hydrolase comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 5, and comprising one or more mutations or mutant groups listed in Table 6.
201. The epoxide hydrolase of claim 200, comprising at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or more mutations or mutation groups listed in Table 6 relative to SEQ ID NO:
5.
202. The epoxide hydrolase of claim 201, wherein the epoxide hydrolase comprises one or more substitutions selected from: (i) S68 is replaced by nonpolar amino acids selected from Ala, Gly and Val; (ii) Q128 is replaced by a nonpolar amino acid selected from Ala, Pro, Gly, and Val; (iii) G141 is replaced by polar amino acids selected from Ser, Thr and Tyr; (iv) T144 is replaced by an amino acid selected from Gln, Asn, Ala, Gly and Val; (v) A145 is replaced by nonpolar amino acids selected from Val, Ile, and Leu; (vi) E163 is replaced by nonpolar amino acids selected from Ala, Gly, and Val; (vii) E191 is replaced by nonpolar amino acids selected from Gly, Ala and Val; (viii) L262 was replaced by Met; (ix) C295 is replaced by a nonpolar amino acid selected from Gly, Ala, and Val; and / or (x) N299 is replaced by Gln.
203. The epoxide hydrolase according to claim 202, wherein the epoxide hydrolase comprises the substitutions S68A, Q128A, G141S, T144Q, A145V, E163A, E191G, L262M, C295G and N299Q relative to SEQ ID NO:
5.
204. A cytochrome P450 comprising an amino acid sequence having at least 80%, or at least 85%, at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 14 or 15, and comprising one or more mutations or groups of mutations listed in Tables 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28 and / or 29.
205. The cytochrome P450 of claim 204, wherein the cytochrome P450 enzyme has all or part of the natural transmembrane domain replaced by a transmembrane domain from a C-terminal protein of the Escherichia coli inner membrane cytoplasm, wherein the Escherichia coli protein is optionally SohB or FliO, wherein optionally about 15 to about 20 amino acids of the cytochrome P450 enzyme are replaced by about 25 to about 30 amino acids of SohB or about 36 to about 42 amino acids of FliO.
206. The cytochrome P450 of claim 205, comprising at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or more mutations or mutation groups listed in Tables 10, 11, 12, 13, 14, 15, 16 and / or 17.
207. The cytochrome P450 of claim 206, wherein the cytochrome P450 enzyme comprises a substitution relative to SEQ ID NO: 15 at one or more positions selected from A12, I27, S63, K75, V76, N102, G125, T128, W130, L131, K132, T192, S193, L218, G219, T223, V228, T231, T232, Y233, N234, K246, F354, K369, H433, F452, V465, A468, and T482, wherein the one or more substitutions are optionally selected from: (i) A12 is replaced by Trp, Tyr or Phe; (ii) I27 was replaced by Gly or Ala; (iii) S63 is replaced by Asn, Phe, Gln, Tyr or Trp; (iv) The K75 was replaced by the Arg; (v) V76 was replaced by Met, Leu, or Ile; (vi) N102 is replaced by His, Lys or Arg; (vii) G125 is replaced by amino acids selected from Ala, Leu, Ile and Val; (viii) T128 is replaced by an amino acid selected from Gly, Ser, Pro and Ala; (ix) W130 is replaced by amino acids selected from Asn, Gly, Ser, Leu, Gln, Ala, Val and Thr; (x) L131 is replaced by amino acids selected from Val, Ile, Pro, Thr and Ser; (xi) K132 is replaced by amino acids selected from Asn, Ser, Thr and Gln; (xii) T192 was replaced by Leu, Ile or Val; (xiii) S193 was replaced by Thr; (xiv) L218 was replaced by Ile, Val and Ala; (xv) G219 is replaced by amino acids selected from Ala, Gln, Arg, Val, Asn, His, and Lys; (xvi) T223 was replaced by Ser; (xvii) V228 was replaced by Ile or Leu; (xviii) T231 is replaced by Phe, Trp or Tyr; (xix) T232 was replaced by Ala, Gly, or Val; (xx) Y233 was replaced by Phe; (xxi) N234 is replaced by His, Arg or Lys; (xxii) K246 is replaced by Val, Ile or Leu; (xxiii) F354 was replaced by Tyr or Trp; (xxiv) K369 was replaced by Arg, Lys or His; (xxv) H433 is replaced by Phe, Trp or Tyr; (xxvi) F452 was replaced by Ile, Val, or Leu; (xxvii) V465 is replaced by an amino acid selected from Ile, Met or Leu; (xxviii) A468 is replaced by an amino acid selected from Thr, Ser, Asn, and Gln; and (xxix) T482 was replaced by Ser.
208. The cytochrome P450 of claim 207, wherein the cytochrome P450 comprises, relative to SEQ ID NO: 14 or 15, the substitutions of A12W, I27G, S63F, S63N, K75R, V76M, N102H, G125A, T128G, W130N, L131I, K132N, T192L, S193T, L218I, G219A, T223S, V228I, T231F, T232A, Y233F, N234H, K246V, F354Y, K369R, H433F, F452I, V465I, A468T, and T482S.
209. A cytochrome P450 enzyme comprising an N-terminal deletion of at least about 20, or at least about 25, or at least about 30, or at least about 35, or at least about 40, or at least about 45, or at least about 50 amino acids relative to SEQ ID NO:
15.
210. The cytochrome P450 enzyme of claim 209, wherein the cytochrome P450 enzyme comprises a deletion selected from L3 to H29 relative to SEQ ID NO:
15.
211. The cytochrome P450 enzyme of claim 210, wherein the cytochrome P450 enzyme comprises an amino acid sequence having at least 70%, or at least 80%, or at least 90%, or at least 95%, or at least 97%, or at least 98%, or at least 99% identity with the amino acid sequence of SEQ ID NO: 85, and comprises one or more mutations or groups of mutations listed in Tables 26, 27, 28 and / or 29.
212. The cytochrome P450 enzyme of claim 211, wherein the cytochrome P450 enzyme comprises a substitution at one or more positions selected from K8, D9, S10, N13, V15, I45, K46, K47, M49, K50, R51, I191, A192, F406, E411, Y412 and S413 relative to SEQ ID NO: 85, wherein the one or more substitutions are optionally selected from: (i) K8 is replaced by Arg; (ii) D9 is replaced by Asn or Gln; (iii) S10 is replaced by Arg, Lys or His; (iv) N13 is replaced by Lys, Arg or His; (v) V15 is replaced by Lys, Arg or His; (vi) I45 was replaced by Val, Ile, or Leu; (vii) K46 was replaced by Arg, Lys or His; (viii) K47 is replaced by Glu or Asp; (ix) M49 was replaced by Val, Ile, or Leu; (x) K50 is replaced by Glu or Asp; (xi) R51 was replaced by Lys; (xii) I191 was replaced by Leu or Val; (xiii) A192 is replaced by Lys, Arg or His; (xiv) F406 was replaced by Tyr or Trp; (xv) E411 was replaced by Asp; (xvi) Y412 was replaced by Phe, Leu, Ile, or Trp; and (xvii) S413 is replaced by Ala, Gly or Val.
213. The cytochrome P450 enzyme of claim 212, wherein the cytochrome P450 enzyme comprises one or more substitutions selected from the group consisting of K8R, D9N, S10R, N13K and V15K, I45V, K46R, K47E, M49V, K50E, R51K, F406Y, E411D, Y412F, S413A, I191L and A192K, relative to SEQ ID NO:
85.
214. A cytochrome P450 reductase comprising an amino acid sequence having at least 80%, or at least 85%, at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 22 or SEQ ID NO: 24, and comprising one or more mutations or groups of mutations listed in Table 19 relative to SEQ ID NO: 22, or in Tables 20 and / or 21 relative to SEQ ID NO:
24.
215. The cytochrome P450 reductase of claim 214, wherein the cytochrome P450 reductase comprises replacing all or part of the native transmembrane domain with a transmembrane domain derived from a C-terminal protein of the inner membrane cytoplasm of *E. coli*, wherein the *E. coli* protein is selected from YcgG, yhcB, zipA, waaA, sohB, djlA, lpxK, ypfN, and yhhM.
216. The cytochrome P450 reductase of claim 215, comprising at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or more mutations or mutation groups listed in Tables 18, 19, 20, and 21.
217. The cytochrome P450 reductase of claim 216, wherein the cytochrome P450 reductase comprises a substitution of the L90 residue relative to SEQ ID NO:
24.
218. The cytochrome P450 reductase of claim 217, wherein the cytochrome P450 reductase comprises a substituted L90P relative to SEQ ID NO:
24.
219. A uridine diphosphate-dependent glycosyltransferase (UGTc3) capable of producing C-3 mogroside from mogroside, comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 27, and comprising one or more mutations or mutant groups listed in Tables 23 and / or 24.
220. The UGTc3 enzyme of claim 219, comprising at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, or more mutations or mutation groups listed in Tables 35 and / or 36.
221. The UGTc3 enzyme of claim 220, wherein the UGTc3 enzyme comprises one or more substitutions selected from the group consisting of, relative to SEQ ID NO: 27: (i) L41 is replaced by aromatic amino acids selected from Phe, Tyr and Trp; (ii) D49 was replaced by Glu; (iii) The T74 was replaced by the Met; (iv) C127 is replaced by an aromatic amino acid selected from Phe, Tyr and Trp; (v) The F303 was replaced by the Tyr; and (vi) A307 is replaced by an amino acid selected from Ile, Leu and Val.
222. The UGTc3 enzyme of claim 221, wherein the UGTc3 enzyme comprises substituted L41F, D49E, T74M, C127F, F303Y and A307I relative to SEQ ID NO:
27.
223. A uridine diphosphate-dependent glycosyltransferase (UGTc3) capable of producing C-3 mogroside from mogroside, comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 65, SEQ ID NO: 66, or SEQ ID NO:
67.
224. A uridine diphosphate-dependent glycosyltransferase (UGTc3) capable of producing C-3 mogroside from mogroside, comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 60, and comprising one or more mutations or groups of mutations listed in Table 37 relative to SEQ ID NO: 60, and / or in Table 39 relative to SEQ ID NO:
65.
225. A uridine diphosphate-dependent glycosyltransferase (UGTc3) capable of producing C-3 mogroside from mogroside, wherein the UGTc3 enzyme is a cyclic variant relative to SEQ ID NO:
60.
226. The uridine diphosphate-dependent glycosyltransferase of claim 225, wherein the cyclic variant has a cleavage site corresponding to amino acids 150 to 170, or 155 to 175, or 160 to 185, or 165 to 190, or 165 to 175 of SEQ ID NO: 60, and optionally has a linker sequence located between amino acids corresponding to the N-terminal and C-terminal residues of SEQ ID NO: 60, and wherein the linker optionally has a length of 2 to 25 amino acids.
227. The uridine diphosphate-dependent glycosyltransferase of claim 225, wherein the UGTc3 enzyme comprises one or more mutations relative to SEQ ID NO: 65 at one or more positions selected from H40, I216, S241, and L352, optionally wherein the mutation is selected from: (i) H40 is replaced by Asp, Gly, Gln, Ala or Val; (ii) I216 is replaced by Ala, Gly, or Val; (iii) S241 is replaced by Gly, Ala, or Val; and (iv) L352 was replaced by Met.
228. The uridine diphosphate-dependent glycosyltransferase of claim 227, wherein the UGTc3 enzyme comprises one or more of the following substitutions relative to SEQ ID NO: 65: H40N, I216A, S241G, and L352M.
229. A uridine diphosphate-dependent glycosyltransferase (UGTc24) capable of producing C-24 mogroside from mogroside, comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 31, and comprising one or more mutations or mutant groups listed in Tables 27, 28, 29, 30, 31 and 32.
230. The UGTc24 enzyme of claim 229, comprising at least two, or at least three, or at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten, or at least eleven, or at least twelve, or at least thirteen or more mutations or mutation groups listed in Tables 40, 41, 42, 43, 44 and / or 45.
231. The UGTc24 enzyme of claim 230, wherein the UGTc24 enzyme comprises one or more substitutions selected from: (i) L18 was replaced by Met; (ii) I63 was replaced by Val, Ala or Gly; (iii) A74 is replaced by an acidic amino acid selected from Glu and Asp; (iv) S88 is replaced by an amino acid selected from Ala, Gly, Leu, Val and Ile; (v) I91 is replaced by aromatic amino acids selected from Phe, Tyr and Trp; (vi) T95 is replaced by amino acids selected from Ala, Gly and Val; (vii) H101 was replaced by Pro; (viii) T184 is replaced by an amino acid selected from Phe, Tyr and Trp; (ix) The F187 was replaced by the Tyr or Trp. (x) P190 is replaced by amino acids selected from Glu and Asp; (xi) T273 is replaced by Ser or is missing; (xii) N306 is replaced by Lys, His or Arg; (xiii) M332 is replaced by an amino acid selected from Leu, Ile, Val, and Ala; and (xiv) Y381 is replaced by amino acids selected from Phe and Trp.
232. The UGTc24 enzyme according to claim 231, wherein the UGTc24 enzyme comprises the substitutions L18M, A74E, S88A, I91F, T95A, H101P, T184F, P190E, T273S, N306K, M332L and Y381F relative to SEQ ID NO:
31.
233. A uridine diphosphate-dependent glycosyltransferase (UGT1216) capable of producing 1,6 and / or 1,2-branched glycosides of mogroside C-3 and / or C-24 glycosides, comprising an amino acid sequence having at least 80%, or at least 85%, at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 35 or SEQ ID NO: 36, and comprising one or more mutations or mutant groups listed in Tables 33 and / or 34.
234. The UGT1216 of claim 233, wherein the UGT1216 is a circular variant of the enzyme represented by SEQ ID NO: 35 or a sequence having at least 70% identity with it.
235. The UGT1216 of claim 234, wherein the cyclic variant has a cleavage site at positions 180 to 195 of amino acids corresponding to or having at least 70% identity with SEQ ID NO:35, and Optionally, a cleavage site is provided at positions 185 to 190 corresponding to SEQ ID NO: 35 or a sequence having at least 70% identity with it, and Optionally, a cleavage site is provided at position 187 of SEQ ID NO: 35 or a sequence having at least 70% identity with it.
236. The UGT1216 enzyme of claim 235, comprising at least two, at least three, at least four, or more mutations or mutation groups listed in Tables 33 and / or 34.
237. The UGT1216 enzyme according to claim 236, wherein the UGT1216 enzyme comprises one or more substitutions selected from: (i) L14 is replaced by nonpolar amino acids selected from Val, Ile, Ala and Gly; (ii) The W160 was replaced by Pro, Ser, or Thr; (iii) L266 is replaced by a polar amino acid selected from Gln and Asn; and (iv) H390 is replaced by amino acids selected from Asp and Glu.
238. The UGT1216 enzyme of claim 237, wherein the UGT1216 enzyme comprises the substitutions L14V, W160P, L266Q and H390D relative to SEQ ID NO:
36.
239. A uridine diphosphate-dependent glycosyltransferase (UGT1216) capable of producing 1,6 and / or 1,2-branched glycosides of mogroside C-3 and / or C-24 glycosides, comprising an amino acid sequence having at least 80%, or at least 85%, or at least 90%, or at least 95% identity with the amino acid sequence of SEQ ID NO: 39, and comprising one or more mutations or mutant groups listed in Table 50.
240. The UGT1216 enzyme of claim 239, comprising at least two or more mutations or mutant groups listed in Table 50.
241. The UGT1216 enzyme of claim 240, wherein the UGT1216 enzyme comprises one or more substitutions selected from: (i) T147 is replaced by a nonpolar amino acid selected from Leu, Ile, and Val; and / or (ii) N207 is replaced by a basic amino acid selected from Lys, Arg and His.
242. The UGT1216 enzyme according to claim 241, wherein the UGT1216 enzyme comprises a substitution of T147L and / or N207K relative to SEQ ID NO:
39.
243. A method for producing mogroside comprising C3-glucosylated mogroside, the method comprising contacting mogroside with a uridine diphosphate-dependent glycosyltransferase (UGT) according to any one of claims 233 to 236.
244. A method for producing mogrosides comprising one or more of C24-glucosylated mogrosides, the method comprising contacting mogrosides with a uridine diphosphate-dependent glycosyltransferase (UGT) according to any one of claims 206 to 208.
245. A method for producing 1,6 and 1,2-glucosylated mogrosides comprising C3- or C24-glucosylated mogroside, the method comprising contacting mogroside, C3-glucosylated mogroside, or C24-glucosylated mogroside with a uridine diphosphate-dependent glycosyltransferase (UGT) according to any one of claims 209 to 218.
Citation Information
Patent Citations
Continuous process for purification of steviol glycosides from stevia leaves using simulated moving bed chromatography
US10213707B2
Microbial production of steviol glycosides
US10463062B2
Metabolic engineering for microbial production of terpenoid products
US10480015B2
Metabolic engineering for microbial production of terpenoid products
US10662442B2
Increasing productivity of E. coli host cells that functionally express P450 enzymes
US10774314B2