Recombinant production of steviol glycosides

Recombinant hosts are engineered to produce steviol glycosides like rebaudioside D, addressing variability and impurities in stevia extracts, enhancing sweetness and reducing plant-derived contaminants.

JP7803999B2Active Publication Date: 2026-01-21ダンスター ファーマント エージー
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024074215
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-02-27
Filing Date
2024-05-01
Publication Date
2026-01-21
Estimated Expiration
2032-08-08

AI Technical Summary

Technical Problem

Existing stevia extracts vary in composition and quality, with rebaudioside D being present in low amounts, leading to inconsistent flavor characteristics and impurities that affect the sweetness and taste of sweeteners.

Method used

Development of recombinant hosts, such as microorganisms and plant cells, engineered to produce steviol glycosides like rebaudioside D through the expression of biosynthetic genes, particularly UDP glycosyltransferases, to enhance production and reduce impurities.

Benefits of technology

The recombinant hosts produce steviol glycosides with increased rebaudioside D content, improving sweetness quality and reducing plant-derived impurities, resulting in higher-quality sweeteners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007803999000022
    Figure 0007803999000022
  • Figure 0007803999000023
    Figure 0007803999000023
  • Figure 0007803999000024
    Figure 0007803999000024
Patent Text Reader

Abstract

To provide recombinant microorganisms, plants, and plant cells that have been engineered to express recombinant genes encoding UDP-glycosyltransferases (UGTs).SOLUTION: Such microorganisms, plants, or plant cells can produce steviol glycosides, e.g., Rebaudioside A and / or Rebaudioside D, which can be used as natural sweeteners in food products and dietary supplements.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. patent application Ser. No. 61 / 521,084, filed Aug. 8, 2011, U.S. patent application Ser. No. 61 / 521,203, filed Aug. 8, 2011, U.S. patent application Ser. No. 61 / 521,051, filed Aug. 8, 2011, U.S. patent application Ser. No. 61 / 523,487, filed Aug. 15, 2011, U.S. patent application Ser. No. 61 / 567,929, filed Dec. 7, 2011, and U.S. patent application Ser. No. 61 / 603,639, filed Feb. 27, 2012, all of which are incorporated herein by reference in their entireties.

[0002] Technical Field The present disclosure relates to the recombinant production of steviol glycosides. In particular, the present disclosure relates to the production of steviol glycosides, such as rebaudioside D, by recombinant hosts, such as recombinant microorganisms, plants, or plant cells. The present disclosure also provides compositions containing steviol glycosides. The present disclosure also relates to tools and methods for producing terpenoids by modulating the biosynthesis of terpenoid precursors in the squalene pathway. [Background technology]

[0003] Sweeteners are known as the most commonly used ingredients in the food, beverage, and confectionery industries. They can be incorporated into the final product during production or used alone. When properly diluted, they can serve as tabletop sweeteners or as a sugar substitute in home baking. Sweeteners include natural sweeteners such as sucrose, high-fructose corn syrup, molasses, maple syrup, and honey, as well as artificial sweeteners such as aspartame, saccharin, and sucralose. Stevia extract is a natural sweetener that can be isolated and extracted from the perennial shrub Stevia rebaudiana. Stevia is commonly cultivated in South America and Asia for the commercial production of stevia extract. Stevia extract, purified to various degrees, is used commercially as a high-intensity sweetener in foods or blends, or as a tabletop sweetener on its own.

[0004] Extracts of the stevia plant contain rebaudiosides and other steviol glycosides that contribute sweetness, although the amount of each glycoside often varies with different production batches. Existing commercial products contain mostly rebaudioside A with smaller amounts of other glycosides, such as rebaudioside C, D, and F. Stevia extracts may also contain impurities, such as plant-derived compounds that contribute to off-flavors. These off-flavors may be more or less problematic depending on the food system or application of choice. Potential impurities include pigments, lipids, proteins, phenols, sugars, spathulenol and other sesquiterpenes, labdane diterpenes, monoterpenes, decanoic acid, 8,11,14-eicosatrienoic acid, 2-methyloctadecane, pentacosane, octacosane, tetracosane, octadecanol, stigmasterol, β-sitosterol, α- and β-amyrin, lupeol, β-amyrin acetate, pentacyclic triterpenes, centaureidin, quercetin, epi-α-cadinol, caryophyllene and its derivatives, β-pinene, β-sitosterol, and gibberellins. Summary of the Invention

[0005] overview Provided herein is a recombinant host, such as a microorganism, plant, or plant cell, that contains one or more biosynthetic genes whose expression results in the production of a steviol glycoside, such as rebaudioside A, rebaudioside C, rebaudioside D, rebaudioside E, rebaudioside F, or dulcoside A. In particular, the uridine 5'-diphospho (UDP) glycosyltransferase described herein, EUGTl1, alone or in combination with one or more other UDP glycosyltransferases, such as UGT74G1, UGT76G1, UGT85C2, and UGT91D2e, enables the production and accumulation of rebaudioside D in a recombinant host or using an in vitro system. As described herein, EUGTl1 has potent 1,2-19-O-glucose glycosylation activity, which is a key step in rebaudioside D production.

[0006] Stevioside and rebaudioside A are typically the predominant compounds in commercially produced stevia extracts. Stevioside has been reported to have a more bitter and less sweet taste than rebaudioside A. The composition of stevia extracts can vary from lot to lot depending on the soil and climate in which the plant is grown. Depending on the source plant, climatic conditions, and extraction process, the amount of rebaudioside A in commercial preparations has been reported to vary from 20 to 97% of the total steviol glycoside content. Other steviol glycosides are present in varying amounts in stevia extracts. For example, rebaudioside B is typically present at less than 1-2%, while rebaudioside C can be present at levels as high as 7-15%. Rebaudioside D is typically present at levels below 2% of the total steviol glycosides, and rebaudioside F is typically present at levels below 3.5%. The amount of minor steviol glycosides affects the flavor characteristics of stevia extracts. Furthermore, rebaudioside D and other highly glycosylated steviol glycosides are believed to be higher quality sweeteners than rebaudioside A. Thus, the recombinant hosts and methods described herein are particularly useful for producing steviol glycoside compositions with increased amounts of rebaudioside D for use, for example, as non-caloric sweeteners with functional and sensory properties superior to many high-potency sweeteners.

[0007] In one aspect, this document features a recombinant host that includes a recombinant gene encoding a polypeptide having at least 80% identity to the amino acid sequence set forth in SEQ ID NO:152. This document also features a recombinant host that includes a recombinant gene encoding a polypeptide capable of transferring a second sugar moiety to the C-2' of the 19-O-glucose of rubusoside.This document also features a recombinant host that includes a recombinant gene encoding a polypeptide capable of transferring a second sugar moiety to the C-2' of the 19-O-glucose of stevioside. In another aspect, this document also features a recombinant host that includes a recombinant gene encoding a polypeptide having the ability to transfer a second sugar moiety to the C-2' of the 19-O-glucose of rubusoside and the C-2' of the 13-O-glucose of rubusoside.

[0008] This document also features a recombinant host that includes a recombinant gene encoding a polypeptide having the ability to transfer a second sugar moiety to the C-2' of a 19-O-glucose of rebaudioside A to produce rebaudioside D, wherein the catalytic rate of the polypeptide is at least 20 times faster (e.g., 25 times or 30 times faster) than a 91D2e polypeptide having the amino acid sequence set forth in SEQ ID NO:5, when the reaction is performed under corresponding conditions. In any of the recombinant hosts described herein, the polypeptide can have at least 85% sequence identity (e.g., 90%, 95%, 98%, or 99% sequence identity) to the amino acid sequence set forth in SEQ ID NO: 152. The polypeptide can have the amino acid sequence set forth in SEQ ID NO: 152. Any of the hosts described herein can further comprise a recombinant gene encoding a UGT85C polypeptide having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 3. The UGT85C polypeptide can include one or more amino acid substitutions at residues 9, 10, 13, 15, 21, 27, 60, 65, 71, 87, 91, 220, 243, 270, 289, 298, 334, 336, 350, 368, 389, 394, 397, 418, 420, 440, 441, 444, and 471 of SEQ ID NO: 3.

[0009] Any of the hosts described herein can further comprise a recombinant gene encoding a UGT76G polypeptide having at least 90% identity to the amino acid sequence set forth in SEQ ID NO: 7. The UGT76G polypeptide can include one or more amino acid substitutions at residues 29, 74, 87, 91, 116, 123, 125, 126, 130, 145, 192, 193, 194, 196, 198, 199, 200, 203, 204, 205, 206, 207, 208, 266, 273, 274, 284, 285, 291, 330, 331, and 346 of SEQ ID NO:7. Any of the hosts described herein can further include a gene (e.g., a recombinant gene) encoding a UGT74G1 polypeptide. Any of the hosts described herein can further include a gene (e.g., a recombinant gene) encoding a functional UGT91D2 polypeptide. The UGT91D2 polypeptide can have at least 80% sequence identity to the amino acid sequence set forth in SEQ ID NO: 5. The UGT91D2 polypeptide can have a mutation at positions 206, 207, or 343 of SEQ ID NO: 5. The UGT91D2 polypeptide can also have a mutation at positions 211 and 286 of SEQ ID NO: 5 (e.g., L211M and V286A, referred to as UGT91D2e-b). The UGT91D2 polypeptide can have the amino acid sequence set forth in SEQ ID NO: 5, 10, 12, 76, 78, or 95.

[0010] Any of the hosts described herein may further include one or more of the following: (i) the gene encoding geranylgeranyl diphosphate synthase; (ii) a gene encoding a bifunctional copalyl diphosphate synthase and kaurene synthase, or a gene encoding a copalyl diphosphate synthase and a gene encoding a kaurene synthase; (iii) genes encoding kaurene oxidase; (iv) a gene encoding a steviol synthase. Each of the genes (i), (ii), (iii), and (iv) can be a recombinant gene. Any of the hosts described herein may further include one or more of the following: (v) a gene encoding a truncated HMG-CoA; (vi) genes encoding CPRs; (vii) a gene encoding rhamnose synthase; (viii) a gene encoding UDP-glucose dehydrogenase; (ix) a gene encoding UDP-glucuronic acid decarboxylase. At least one of the genes (i), (ii), (iii), (iv), (v), (vi), (vii), (viii), or (ix) can be a recombinant gene.

[0011] The geranylgeranyl diphosphate synthase can have greater than 90% sequence identity with one of the amino acid sequences set forth in SEQ ID NOs: 121-128. The copalyl diphosphate synthase can have greater than 90% sequence identity with one of the amino acid sequences set forth in SEQ ID NOs: 129-131. The kaurene synthase can have greater than 90% sequence identity with one of the amino acid sequences set forth in SEQ ID NOs: 132-135. The kaurene oxidase can have greater than 90% sequence identity with one of the amino acid sequences set forth in SEQ ID NOs: 138-141. The steviol synthase can have greater than 90% sequence identity with one of the amino acid sequences set forth in SEQ ID NOs: 142-146. Any recombinant host, when cultured under conditions in which each gene is expressed, is capable of producing at least one steviol glycoside, which can be selected from the group consisting of rubusoside, rebaudioside A, rebaudioside B, rebaudioside C, rebaudioside D, rebaudioside E, rebaudioside F, dulcoside A, stevioside, steviol-19-O-glucoside, steviol-13-O-glucoside, steviol-1,2-bioside, steviol-1,3-bioside, 1,3-stevioside, and other rhamnosylated or xylosylated intermediates. Steviol glycosides (e.g., rebaudioside D) can accumulate to at least 1 mg / liter (e.g., at least 10 mg / liter, 20 mg / liter, 100 mg / liter, 200 mg / liter, 300 mg / liter, 400 mg / liter, 500 mg / liter, 600 mg / liter, or 700 mg / liter, or more) of culture medium when cultured under the above conditions.

[0012] This document also features a method for producing steviol glycosides. The method includes growing any of the hosts described herein in a culture medium under conditions in which the genes are expressed; and recovering the steviol glycoside produced by the host. The growing step can include inducing expression of one or more genes. The steviol glycoside can be a 13-O-1,2-diglycosylated and / or 19-O-1,2-diglycosylated steviol glycoside (e.g., stevioside, steviol 1,2bioside, rebaudioside D, or rebaudioside E). For example, the steviol glycoside can be rebaudioside D or rebaudioside E. Other examples of steviol glycosides include rebaudioside A, rebaudioside B, rebaudioside C, rebaudioside F, and dulcoside A. This document also features a recombinant host. The host includes (i) a gene encoding UGT74G1; (ii) a gene encoding UGT85C2; (iii) a gene encoding UGT76G1; (iv) a gene encoding a glycosyltransferase capable of transferring a second sugar moiety to the C-2' of the 19-O-glucose of rubusoside or stevioside; and (v) optionally, a gene encoding UGT91D2e, wherein at least one of the genes is a recombinant gene. In some embodiments, each gene is a recombinant gene. The host is capable of producing at least one steviol glycoside (e.g., rebaudioside D) when cultured under conditions in which each of the genes (e.g., recombinant genes) is expressed. The host can further comprise (a) a gene encoding a bifunctional copalyl diphosphate synthase and kaurene synthase, or a gene encoding a copalyl diphosphate synthase and a gene encoding a kaurene synthase; (b) a gene encoding a kaurene oxidase; (c) a gene encoding a steviol synthase; and (d) a gene encoding a geranylgeranyl diphosphate synthase.

[0013] This document also features steviol glycoside compositions produced by any of the hosts described herein, which have reduced levels of stevia plant-derived impurities compared to stevia extract. In yet another aspect, this document features a steviol glycoside composition produced by any of the hosts described herein, the composition having a steviol glycoside composition enriched in rebaudioside D relative to the steviol glycoside composition of a wild-type Stevia plant. In yet another aspect, this document features a method for producing a steviol glycoside composition. The method includes growing a host described herein in a culture medium under conditions in which each gene is expressed; and recovering the steviol glycoside composition produced by the host (e.g., a microorganism). The composition is enriched for rebaudioside A, rebaudioside B, rebaudioside C, rebaudioside D, rebaudioside E, rebaudioside F, or dulcoside A compared to the steviol glycoside composition of a wild-type stevia plant. The steviol glycoside composition produced by the host (e.g., a microorganism) has reduced levels of stevia plant-derived impurities compared to a stevia extract.

[0014] This document also features a method for transferring a second sugar moiety to the C-2' of a 19-O-glucose or the C-2' of a 13-O-glucose in a steviol glycoside. The method includes contacting a steviol glycoside with a EUGT11 polypeptide described herein or a UGT91D2 polypeptide described herein (e.g., UGT91D2e-b) and a UDP-sugar under suitable reaction conditions for transferring the second sugar moiety to the steviol glycoside. The steviol glycoside can be rubusoside, where the second sugar moiety is glucose, and stevioside is produced upon transfer of the second glucose moiety. The steviol glycoside can be stevioside, where the second sugar moiety is glucose, and rebaudioside E is produced upon transfer of the second glucose moiety. The steviol glycoside can be rebaudioside A, where rebaudioside D is produced upon transfer of the second glucose moiety. In another aspect of the improved downstream steviol glycoside pathway as disclosed herein, materials and methods are provided for the recombinant production of sucrose synthase, materials and methods for increasing UDP-glucose production in a host, particularly for the purpose of enhancing glycosylation reactions in cells to increase the availability of UDP-glucose in vivo, and methods for reducing UDP concentrations in cells are provided.

[0015] This document also provides a recombinant host comprising one or more exogenous nucleic acids encoding a sucrose transporter and a sucrose synthase, wherein expression of the one or more exogenous nucleic acids having a glucosyltransferase results in elevated UDP-glucose levels in the host. Optionally, the one or more exogenous nucleic acids comprise a SUS1 sequence. Optionally, the SUS1 sequence is from Coffea arabica or encodes a functional homolog of the sucrose synthase encoded by the SUS1 sequence of Coffea arabica, although as described herein, SUS from Arabidopsis thaliana or Stevia rebaudiana can also be used. In the recombinant host of the present invention, the one or more exogenous nucleic acids can comprise a sequence encoding a polypeptide having the sequence set forth in SEQ ID NO: 180 or an amino acid sequence at least 90% identical thereto, and optionally, the one or more exogenous nucleic acids comprise a SUC1 sequence. In one embodiment, the SUC1 sequence is from Arabidopsis thaliana, or the SUC1 sequence encodes a functional homolog of the sucrose transporter encoded by the Arabidopsis thaliana SUC1 sequence. In the recombinant host, the one or more exogenous nucleic acids can comprise a sequence encoding a polypeptide having the sequence set forth in SEQ ID NO: 179, or an amino acid sequence at least 90% identical thereto. The recombinant host has a reduced ability to degrade exogenous sucrose compared to a corresponding host lacking the one or more exogenous nucleic acids.

[0016] The recombinant host may be a microorganism, for example a yeast such as Saccharomyces cerevisiae. Alternatively, the microorganism is Escherichia coli. In an alternative embodiment, the recombinant host is a plant or plant cell. The present invention also provides a method for increasing the level of UDP-glucose and decreasing the level of UDP in a cell, the method comprising expressing a recombinant sucrose synthase sequence and a recombinant sucrose transporter sequence in a cell in a medium containing sucrose, wherein the cell is deficient in sucrose degradation. The present invention further provides a method for enhancing glycosylation in a cell, comprising expressing a recombinant sucrose synthase sequence and a recombinant sucrose transporter sequence in a cell in a medium containing sucrose, wherein the expression results in a decrease in UDP levels in the cell and an increase in UDP-glucose levels in the cell, thereby increasing glycosylation in the cell.

[0017] In either method for increasing the level of UDP-glucose or promoting glycosylation, the cell can produce vanillin glucoside, resulting in increased vanillin glucoside production by the cell, or can produce steviol glucoside, resulting in increased steviol glucoside production by the cell. Optionally, the SUS1 sequence is an A. thaliana, S. rebaudiana, or Coffea arabica SUS1 sequence (see, e.g., Figure 17, SEQ ID NOS:175-177), or a sequence encoding a functional homolog of the sucrose synthase encoded by the A. thaliana, S. rebaudiana, or Coffea arabica SUS1 sequence. The recombinant sucrose synthase optionally comprises a nucleic acid encoding a polypeptide having the sequence set forth in SEQ ID NO: 180 or an amino acid sequence at least 90% identical thereto, wherein optionally the recombinant sucrose transporter sequence is a SUC1 sequence, or wherein optionally the SUC1 sequence is an Arabidopsis thaliana SUC1 sequence or a sequence encoding a functional homolog of a sucrose transporter encoded by an Arabidopsis thaliana SUC1 sequence, or wherein optionally the recombinant sucrose transporter sequence comprises a nucleic acid encoding a polypeptide having the sequence set forth in SEQ ID NO: 179 or an amino acid sequence at least 90% identical thereto. In either method, the host can be a microorganism, for example, a yeast such as Saccharomyces cerevisiae, or the host can be Escherichia coli, or the host can be a plant cell.

[0018] Also provided herein is a recombinant host, such as a microorganism, containing one or more biosynthetic genes whose expression results in the production of diterpenoids. Such genes include a gene encoding ent-copalyl diphosphate synthase (CDPS) (EC 5.5.1.13), a gene encoding ent-kaurene synthase, a gene encoding ent-kaurene oxidase, or a gene encoding steviol synthase. At least one of these genes is a recombinant gene. The host can also be a plant cell. Expression of these gene(s) in a stevia plant can result in increased levels of steviol glycosides in the plant. In some embodiments, the recombinant host further contains multiple copies of a recombinant gene encoding a CDPS polypeptide (EC 5.5.1.13) lacking a chloroplast transit peptide sequence. The CDPS polypeptide can have at least 90%, 95%, 99%, or 100% identity to the truncated CDPS amino acid sequence shown in Figure 14. The host can further comprise multiple copies of a recombinant gene encoding a KAH polypeptide, e.g., a KAH polypeptide having at least 90%, 95%, 99%, or 100% identity to the KAH amino acid sequence shown in Figure 12. The host can further comprise one or more of the following: (i) a gene encoding geranylgeranyl diphosphate synthase; (ii) a gene encoding ent-kaurene oxidase; and (iii) a gene encoding ent-kaurene synthase. The host can further comprise one or more of the following: (iv) a gene encoding a truncated HMG-CoA; (v) a gene encoding a CPR; (vi) a gene encoding a rhamnose synthase; (vii) a gene encoding a UDP-glucose dehydrogenase; and (viii) a gene encoding a UDP-glucuronic acid decarboxylase. For example, two or more exogenous CPRs can be present. Expression of one or more of such genes can be inducible.At least one of the genes (i), (ii), (iii), (iv), (v), (vi), (vii), or (viii) can be a recombinant gene, and in some cases, each of the genes (i), (ii), (iii), (iv), (v), (vi), (vii), or (viii) is a recombinant gene. The geranylgeranyl diphosphate synthase can have greater than 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 127; the kaurene oxidase can have greater than 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 138; the CPR can have greater than 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 168; the CPR can have greater than 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 170, and the kaurene synthase can have greater than 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 156.

[0019] In one aspect, this document features an isolated nucleic acid encoding a polypeptide having the amino acid sequence set forth in SEQ ID NO:5, wherein the polypeptide includes substitutions at positions 211 and 286 of SEQ ID NO:5. For example, the polypeptide can include a methionine at position 211 and an alanine at position 286. In one aspect, this document features an isolated nucleic acid encoding a polypeptide having at least 80% identity (e.g., at least 85%, 90%, 95%, or 99% identity) to the amino acid sequence set forth in Figure 12C (SEQ ID NO: 164). The polypeptide can have the amino acid sequence set forth in Figure 12C. In another aspect, this document features a nucleic acid construct including a regulatory region operably linked to a nucleic acid encoding a polypeptide having at least 80% identity (e.g., at least 85%, 90%, 95%, or 99% identity) to the amino acid sequence set forth in Figure 12C (SEQ ID NO: 164). The polypeptide can have the amino acid sequence set forth in Figure 12C.

[0020] This document also features a recombinant host including a recombinant gene (e.g., multiple copies of the recombinant gene) encoding a KAH polypeptide having at least 80% identity (e.g., at least 85%, 90%, 95%, or 99% identity) to the amino acid sequence set forth in FIG. 12C. The polypeptide can have the amino acid sequence set forth in FIG. 12C. The host can be a microorganism such as yeast (e.g., Saccharomyces cerevisiae) or Escherichia coli. The host can be a plant or plant cell (e.g., a stevia, Physcomitrella, or tobacco plant or plant cell). The stevia plant or plant cell is a Stevia rebaudiana plant or plant cell. The recombinant host can produce steviol when cultured under conditions in which each gene is expressed. The recombinant host can further include a gene encoding a UGT74G1 polypeptide, a gene encoding a UGT85C2 polypeptide, a gene encoding a UGT76G1 polypeptide, a gene encoding a UGT91D2 polypeptide, and / or a gene encoding a EUGT11 polypeptide. When cultured under conditions in which each gene is expressed, the host can produce at least one steviol glycoside. The steviol glycoside can be steviol-13-O-glucoside, steviol-19-O-glucoside, rubusoside, rebaudioside A, rebaudioside B, rebaudioside C, rebaudioside D, rebaudioside E, rebaudioside F, and / or dulcoside A.The recombinant host may further comprise one or more of the following: a gene encoding deoxyxylulose 5-phosphate synthase (DXS); a gene encoding D-1-deoxyxylulose 5-phosphate reductoisomerase (DXR); a gene encoding 4-diphosphocytidyl-2-C-methyl-D-erythritol synthase (CMS); a gene encoding 4-diphosphocytidyl-2-C-methyl-D-erythritol kinase (CMK); a gene encoding 4-diphosphocytidyl-2-C-methyl-D-erythritol 2,4-cyclodiphosphate synthase (MCS); a gene encoding 1-hydroxy-2-methyl-2(E)-butenyl 4-diphosphate synthase (HDS); and a gene encoding 1-hydroxy-2-methyl-2(E)-butenyl 4-diphosphate reductase (HDR). The recombinant host can further include one or more of the following: a gene encoding an acetoacetyl-CoA thiolase; a gene encoding a truncated HMG-CoA reductase; a gene encoding a mevalonate kinase; a gene encoding a phosphomevalonate kinase; and a gene encoding a mevalonate pyrophosphate decarboxylase. In another aspect, this document features a recombinant host further including a gene encoding an ent-kaurene oxidase (EC 4.2.3.19); and / or a gene encoding a gibberellin 20-oxidase (EC 1.14.11.12). When cultured under conditions in which each gene is expressed, the host produces the gibberellin GA3.

[0021] This document also features an isolated nucleic acid encoding a CPR polypeptide having at least 80% sequence identity (e.g., at least 85%, 90%, 95%, or 99% sequence identity) to the S. rebaudiana CPR amino acid sequence set forth in Figure 13. In some embodiments, the polypeptide has the S. rebaudiana CPR amino acid sequence set forth in Figure 13 (SEQ ID NOs: 169 and 170). In any of the hosts described herein, expression of one or more genes is inducible. In any of the hosts described herein, one or more genes encoding endogenous phosphatases can be deleted or disrupted to reduce endogenous phosphatase activity. For example, the yeast genes DPP1 and / or LPP1 can be disrupted or deleted to reduce the breakdown of farnesyl pyrophosphate (FPP) to farnesol and the breakdown of geranylgeranyl pyrophosphate (GGPP) to geranylgeraniol (GGOH). In another embodiment, as described herein, ERG9 can be modified as defined below, resulting in reduced production of squalene synthase (SQS) and the accumulation of terpenoid precursors, which may or may not be secreted into the culture medium and can then be used as substrates for enzymes capable of metabolizing the terpenoid precursors into desired terpenoids.

[0022] Thus, a main aspect of the present invention is a cell comprising a nucleic acid sequence, said nucleic acid comprising: i) the promoter sequence, which is ii) operably linked to a heterologous insert sequence, which iii) operably linked to an open reading frame, which iv) operably linked to a transcription termination signal; wherein the heterologous insert sequence comprises the general formula (I): -X1-X2-X3-X4-X5- wherein X2 comprises at least four consecutive nucleotides that are complementary to and form a hairpin secondary structure element with at least four consecutive nucleotides of X4; and wherein X3 is optional and, if present, contains an unpaired nucleotide involved in the formation of a hairpin loop between X2 and X4; and wherein X1 and X5 independently and optionally comprise one or more nucleotides, and wherein the open reading frame, when expressed, encodes a polypeptide sequence having at least about 70% identity to a squalene synthase (EC 2.5.1.21) or a biologically active fragment thereof, wherein the fragment has at least 70% sequence identity with the squalene synthase within an overlap of at least 100 amino acids. Regarding the cells.

[0023] The cells of the invention are useful for increasing the yield of industrially interesting terpenoids. Accordingly, another aspect of the invention relates to a method for producing terpenoid compounds synthesized via the squalene pathway in a cell culture, the method comprising the steps of: (a) providing a cell as defined herein above; (b) culturing the cells of step (a); (c) recovering the terpenoid product compounds. By providing a cell containing the genetically engineered construct defined herein above, the accumulation of terpenoid precursors is enhanced (see, for example, Figure 20).

[0010] Accordingly, in another aspect, the present invention relates to a method for producing a terpenoid from a terpenoid precursor selected from the group consisting of farnesyl pyrophosphate (FPP), isopentenyl pyrophosphate (IPP), dimethylallyl pyrophosphate (DMAPP), geranyl pyrophosphate (GPP), and / or geranylgeranyl pyrophosphate (GGPP), said method comprising: (a) contacting the precursor with an enzyme of the squalene synthase pathway; (b) recovering the terpenoid product.

[0024] The present invention may operate, at least in part, by sterically hindering the binding of ribosomes to RNA, thus reducing translation of squalene synthase. Accordingly, one aspect of the present invention relates to a method for reducing the translation rate of functional squalene synthase (EC 2.5.1.21), said method comprising: (a) providing a cell as defined herein above; (b) culturing the cells of step (a). Similarly, in another aspect, the present invention relates to a method for reducing the conversion of farnesyl-pp to squalene, said method comprising: (a) providing a cell as defined herein above; (b) culturing the cells of step (a).

[0025] As shown in Figure 20, knockdown of ERG9 results in an increase in the precursor to squalene synthase. Thus, in one aspect, the present invention relates to a method for enhancing the accumulation of a compound selected from the group consisting of farnesyl pyrophosphate, isopentenyl pyrophosphate, dimethylallyl pyrophosphate, geranyl pyrophosphate, and geranylgeranyl pyrophosphate, comprising the steps of: (a) providing a cell as defined herein above; (b) culturing the cells of step (a). In one aspect, the present invention relates to the production of geranylgeranyl pyrophosphate (GGPP) and other terpenoids that can be prepared from geranylgeranyl pyrophosphate (GGPP).

[0026] In this aspect of the invention, the reduced production of squalene synthase (SQS) described above can be combined with increased activity of geranylgeranyl pyrophosphate synthase (GGPPS), which converts FPP to geranylgeranyl pyrophosphate (GGPP), resulting in increased production of GGPP. Thus, in one aspect, the present invention relates to a microbial cell comprising a nucleic acid sequence, said nucleic acid comprising: i) the promoter sequence, which is ii) operably linked to a heterologous insert sequence, which iii) operably linked to an open reading frame, which iv) operably linked to a transcription termination signal; wherein the heterologous insert sequence and open reading frame are as defined herein; wherein said microbial cell further comprises a heterologous nucleic acid encoding GGPPS operably linked to a nucleic acid sequence that directs expression of GGPPS in said cell. Furthermore, this document relates to methods for producing steviol or steviol glycosides, wherein the methods comprise the use of any one of the microbial cells described above.

[0027] All hosts described herein can be microorganisms (e.g., yeast such as Saccharomyces cerevisiae, or Escherichia coli), or plants or plant cells (e.g., Stevia such as Stevia rebaudiana, Physcomitrella, or tobacco plants or plant cells). Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used to practice the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. Furthermore, the materials, methods, and examples are illustrative only and are not intended to be limiting. Other features and advantages of the present invention will be apparent from the following detailed description. Applicants reserve the right to alternatively claim any disclosed inventions using the transitional phrases "comprising," "consisting essentially of," or "consisting of," in accordance with standard practice in patent law. [Brief explanation of the drawings]

[0028] [Figure 1] FIG. 1 is a diagram of the chemical structures of various steviol glycosides. [Figure 2A] FIG. 2A shows a representative pathway for the biosynthesis of steviol glycosides from steviol. [Figure 2B] FIG. 2B shows a representative pathway for the biosynthesis of steviol glycosides from steviol. [Figure 2C] FIG. 2C shows a representative pathway for the biosynthesis of steviol glycosides from steviol. [Figure 2D] FIG. 2D shows a representative pathway for the biosynthesis of steviol glycosides from steviol. [Figure 3] Figure 3 is a schematic diagram of the 19-O-1,2-diglycosylation reaction by EUGT11 and UGT91D2e. Numbers are the average signal intensities of the substrate or product of the reaction from liquid chromatography-mass spectrometry (LC-MS) chromatograms. [Figure 4] Figure 4 shows LC-MS chromatograms depicting the production of rebaudioside D (RebD) from rebaudioside A (RebA) using in vivo transcribed or translated UGT91D2e (SEQ ID NO: 5) (left panel) or EUGT11 (SEQ ID NO: 152) (right panel). The LC-MS was configured to detect specific masses corresponding to steviol + 5 glucoses (e.g., RebD), steviol + 4 glucoses (e.g., RebA), etc. Each "lane" is scaled according to the largest peak.

[0029] [Figure 5] FIG. 5 is an LC-MS chromatogram showing the conversion of rubusoside to stevioside and compounds “2” and “3” (RebE) by UGT91D2e (left panel) and EUGT11 (right panel). [Figure 6] FIG. 6 shows an alignment of the amino acid sequence of EUGT11 (SEQ ID NO: 152, top line) and the amino acid sequence of UGT91D2e (SEQ ID NO: 5, bottom line). [Figure 7]FIG. 7 shows the amino acid sequence of EUGT11 (SEQ ID NO: 152), the nucleotide sequence encoding EUGT11 (SEQ ID NO: 153), and the nucleotide sequence encoding EUGT11 that has been codon-optimized for expression in yeast (SEQ ID NO: 154). [Figure 8] Figure 8 shows an alignment of the predicted secondary structures of UGT91D2e with those of UGT85H2 and UGT71G1. Secondary structure predictions were performed by submitting the amino acid sequences of the three UGTs to NetSurfP ver. 1.1 - Protein Surface Accessibility and Secondary Structure Predictions, available on the World Wide Web at cbs.dtu.dk / services / NetSurfP / . This predicted the presence and location of α-helices, β-sheets, and coils in the protein. These were subsequently labeled as shown for UGT91D2e. For example, the first N-terminal β-sheet was labeled Nβ1. The y-axis represents the certainty of the prediction; higher values ​​indicate greater certainty; the x-axis indicates the amino acid position. Although the primary sequence identity between these UGTs is very low, the secondary structures show a very high degree of conservation.

[0030] [Figure 9] FIG. 9 shows an alignment of the amino acid sequences of UGT91D1 and UGT91D2e (SEQ ID NO: 5). [Figure 10] 10 is a bar graph of the activity of double amino acid substitution mutants of UGT91D2e. The black bars represent stevioside production, and the white bars represent 1,2-bioside production. [Figure 11] Figure 11 is a schematic diagram of the regeneration of UDP-glucose for the biosynthesis of steviol glycosides. SUS = sucrose synthase, Steviol = steviol or steviol glycoside substrate, UGT = UDP glycosyltransferase. [Figure 12A] FIG. 12A is a diagram showing the nucleotide sequence encoding Stevia rebaudiana KAH (SEQ ID NO: 163), which is referred to herein as SrKAHe1. [Figure 12B] FIG. 12B contains the nucleotide sequence encoding Stevia rebaudiana KAHe1 (SEQ ID NO: 165), which has been codon-optimized for expression in yeast. [Figure 12C] FIG. 12C is the amino acid sequence of Stevia rebaudiana KAHe1 (SEQ ID NO: 164).

[0031] [Figure 13A] Figure 13A shows the amino acid sequences of CPR polypeptides from S. cerevisiae (encoded by the NCP1 gene) (SEQ ID NO: 166), A. thaliana (encoded by ATR1 and ATR2) (SEQ ID NOs: 148 and 168), and S. rebaudiana (encoded by CPR7 and CPR8) (SEQ ID NOs: 169 and 170). [Figure 13B] FIG. 13B contains the ATR1 nucleotide sequence (accession number CAA23011) codon-optimized for expression in yeast (SEQ ID NO: 171); the ATR2 nucleotide sequence codon-optimized for expression in yeast (SEQ ID NO: 172); the Stevia rebaudiana CPR7 nucleotide sequence (SEQ ID NO: 173); and the Stevia rebaudiana CPR8 nucleotide sequence (SEQ ID NO: 174). [Figure 14A] 14A depicts the nucleotide sequence (SEQ ID NO:157) encoding the CDPS polypeptide (SEQ ID NO:158) from Zea mays. The bold and underlined sequence can be deleted to remove the sequence encoding the chloroplast transit sequence. [Figure 14B] 14B is the amino acid sequence (SEQ ID NO: 158) encoding the CDPS polypeptide from Zea mays. The bold and underlined sequence can be deleted to remove the sequence encoding the chloroplast transit sequence.

[0032] [Figure 15A]FIG. 15A contains a codon-optimized nucleotide sequence (SEQ ID NO: 161) encoding the bifunctional CDPS-KS polypeptide (SEQ ID NO: 162) from Gibberella fujikuroi. [Figure 15B] FIG. 15B contains the amino acid sequence encoding the bifunctional CDPS-KS polypeptide from Gibberella fujikuroi (SEQ ID NO: 162). [Figure 16] 16 shows a graph of the growth of two strains of S. cerevisiae: enhanced EFSC1972 (designated T2) and enhanced EFSC1972 with overexpression of Arabidopsis thaliana kaurene synthase (KS-5) (designated T7, squares). Numbers on the y-axis are OD values ​​of the cell cultures, and numbers on the x-axis represent hours of growth in synthetic-based medium at 30° C. [Figure 17] FIG. 17 shows the nucleic acid sequences encoding A. thaliana, S. rebaudiana (contig10573 selection_ORF S11E, where the mutation changing S11 to glutamic acid (E) is shown in bold lowercase), and coffee (Coffea arabica) sucrose synthases (SEQ ID NOs: 175, 176, and 177, respectively).

[0033] [Figure 18] Figure 18 shows a bar graph of rebD production in permeabilized S. cerevisiae transformed with EUGT11 or an empty plasmid ("Empty"). Cells were grown to logarithmic phase, washed in PBS buffer, and then treated with Triton X-100 (0.3% or 0.5% in PBS) for 30 minutes at 30°C. After permeabilization, cells were washed in PBS and resuspended in a reaction mix containing 100 μM RebA and 300 μM UDP-glucose. The reaction proceeded for 20 hours at 30°C. [Figure 19A]FIG. 19A contains the amino acid sequence of A. thaliana UDP-glycosyltransferase UGT72E2 (SEQ ID NO: 178). [Figure 19B] FIG. 19B shows the amino acid sequence of the sucrose transporter SUC1 from A. thaliana (SEQ ID NO: 179). [Figure 19C] FIG. 19C shows the amino acid sequence of sucrose synthase from coffee (SEQ ID NO: 180). [Figure 20] FIG. 20 is a schematic diagram of the isoprenoid pathway in yeast, showing the location of ERG9. [Figure 21] FIG. 21 shows the nucleotide sequences of the Saccharomyces cerevisiae Cyc1 promoter (SEQ ID NO: 185) and the Saccharomyces cerevisiae Kex2 promoter (SEQ ID NO: 186).

[0034] [Figure 22] Figure 22 is a schematic diagram of a PCR product containing two regions, HR1 and HR2, that are homologous to portions of genomic sequence within the ERG9 promoter or the 5' end of the ERG9 open reading frame (ORF). Also present on the PCR product is an antibiotic marker, NatR, that can be embedded between two Lox sites (L) for subsequent excision with Cre recombinase. The PCR product can further contain a promoter, such as wild-type ScKex2 or wild-type ScCyc1, which can further contain a heterologous insert, such as a hairpin (SEQ ID NOs: 181-184), at its 3' end (see Figure 23). [Figure 23]Figure 23 shows a schematic of the promoter and ORF with a hairpin stem-loop immediately upstream of the translation start site (arrow), as well as an alignment of a portion of the wild-type S. cerevisiae CYC1 promoter sequence with the initial ATG of the ERG9 OPR without a heterologous insert (SEQ ID NO: 187) and with four different heterologous inserts (SEQ ID NOs: 188-191). 75% represents a construct containing the ScCyc1 promoter followed by SEQ ID NO: 184 (SEQ ID NO: 191); 50% represents a construct containing the ScCyc1 promoter followed by SEQ ID NO: 183 (SEQ ID NO: 190); 20% represents a construct containing the ScCyc1 promoter followed by SEQ ID NO: 182 (SEQ ID NO: 189); and 5% represents a construct containing the ScCyc1 promoter followed by SEQ ID NO: 181 (SEQ ID NO: 188).

[0035] [Figure 24] 24 is a bar graph showing amorphadiene produced in yeast strains with different promoter constructs inserted in front of the ERG9 gene in the host genome. CTRL-ADS indicates the unmodified control strain; ERG9-CYC1-100% indicates a construct containing the ScCyc1 promoter and no insert; ERG9-CYC1-50% indicates a construct containing the ScCyc1 promoter followed by SEQ ID NO: 183 (SEQ ID NO: 190); ERG9-CYC1-20% indicates a construct containing the ScCyc1 promoter followed by SEQ ID NO: 182 (SEQ ID NO: 189); ERG9-CYC1-5% indicates a construct containing the ScCyc1 promoter followed by SEQ ID NO: 181 (SEQ ID NO: 188); and ERG9-KEX2-100% indicates a construct containing the ScKex2 promoter. [Figure 25]FIG. 25 contains the amino acid sequences of squalene synthase polypeptides from Saccharomyces cerevisiae, Schizosaccharomyces pombe, Yarrowia lipolytica, Candida glabrata, Ashbya gossypii, Cyberlindnera jadinii, Candida albicans, Saccharomyces cerevisiae, Homo sapiens, Mus musculus, and Rattus norvegicus (SEQ ID NOs: 192-202), and the amino acid sequences of geranylgeranyl diphosphate synthase (GGPPS) from Aspergilus nidulans and S. cerevisiae (SEQ ID NOs: 203 and 167). [Figure 26] FIG. 26 is a bar graph showing the accumulation of geranylgeraniol (GGOH) after 72 hours in the ERG9-CYC1-5% and ERG9-KEX2 strains. [Figure 27] 27 is a representative chromatograph showing the conversion of rubusoside to a xylosylated intermediate for RebF production by UGT91D2e and EUGT11. Like reference symbols in the various drawings indicate like elements. Detailed Description of the Invention

[0036] This document is based on the discovery that recombinant hosts, such as plant cells, plants, or microorganisms, can be developed that express polypeptides useful for the biosynthesis of steviol glycosides, such as rebaudioside A, rebaudioside C, rebaudioside D, rebaudioside E, rebaudioside F, or dulcoside A. The recombinant hosts described herein are particularly useful for the production of rebaudioside D. Such hosts can express one or more uridine 5'-diphospho (UDP) glycosyltransferases suitable for the production of steviol glycosides. Expression of these biosynthetic polypeptides in various microbial chassis allows for the consistent and reproducible production of steviol glycosides from energy and carbon sources, such as sugars, glycerol, CO2, H2, and sunlight. The proportion of each steviol glycoside produced by a recombinant host can be tailored by incorporating preselected biosynthetic enzymes into the host and expressing them at appropriate levels, resulting in the production of sweetener compositions with consistent taste profiles. Furthermore, the concentrations of steviol glycosides produced by the recombinant host are expected to be higher than the levels produced in the Stevia plant, which will improve the efficiency of downstream purification. Such sweetener compositions contain little or no plant-based impurities compared to the amount of impurities present in Stevia extracts.

[0037] At least one of the genes is a recombinant gene, with the particular recombinant gene(s) depending on the species or strain selected for use. Additional genes or biosynthetic modules can be included to increase the yield of steviol glycosides, improve the efficiency with which energy and carbon sources are converted to steviol and its glycosides, and / or enhance productivity from cell cultures or plants. Such additional biosynthetic modules include genes involved in the synthesis of the terpenoid precursors, isopentenyl diphosphate and dimethylallyl diphosphate. Additional biosynthetic modules include terpene synthase and terpene cyclase genes, such as genes encoding geranylgeranyl diphosphate synthase and copalyl diphosphate synthase; these genes can be endogenous or recombinant.

[0038] 1. Steviol and Steviol Glycoside Biosynthesis Polypeptides A. Steviol Biosynthesis Polypeptides The chemical structures of some of the compounds found in stevia extracts, including the diterpene steviol and various steviol glycosides, are shown in Figure 1. CAS numbers are listed in Table A below. See also Steviol Glycosides Chemical and Technical Assessment 69th JECFA, prepared by Harriet Wallin, Food Agric. Org. (2007). [Table 1]

[0039] It has been discovered that expression of certain genes in a host, such as a microorganism, confers on the host the ability to synthesize steviol glycosides. As explained in more detail below, one or more of such genes may be naturally present in the host. Typically, however, one or more of such genes are recombinant genes that have been transformed into a host that does not naturally possess them. The biochemical pathway for producing steviol involves the formation of geranylgeranyl diphosphate, cyclization to (-)copalyl diphosphate, followed by oxidation and hydroxylation to form steviol. Thus, conversion of geranylgeranyl diphosphate to steviol in recombinant microorganisms involves expression of a gene encoding kaurene synthase (KS), a gene encoding kaurene oxidase (KO), and a gene encoding steviol synthase (KAH). Steviol synthase is also known as kaurenoic acid 13-hydroxylase.

[0040] Suitable KS polypeptides are known. For example, suitable KS enzymes include those produced by Stevia rebaudiana, Zea mays, Populus trichocarpa, and Arabidopsis thaliana. See Table 1 and SEQ ID NOS: 132-135 and 156. Nucleotide sequences encoding these polypeptides are set forth in SEQ ID NOS: 40-47 and 155. The nucleotide sequences set forth in SEQ ID NOS: 40-43 have been modified for expression in yeast, while the nucleotide sequences set forth in SEQ ID NOS: 44-47 are from the source organisms from which the KS polypeptides were identified. [Table 2]

[0041] Suitable KO polypeptides are known. For example, suitable KO enzymes include those produced by Stevia rebaudiana, Arabidopsis thaliana, Gibberella fujikoroi, and Trametes versicolor. See Table 2 and SEQ ID NOs: 138-141. Nucleotide sequences encoding these polypeptides are set forth in SEQ ID NOs: 52-59. The nucleotide sequences set forth in SEQ ID NOs: 52-59 were modified for expression in yeast. The nucleotide sequences set forth in SEQ ID NOs: 56-59 are from the source organisms from which the KO polypeptides were identified. [Table 3]

[0042] Suitable KAH polypeptides are known. For example, suitable KAH enzymes include those produced by Stevia rebaudiana, Arabidopsis thaliana, Vitis vinifera, and Medicago trunculata. See Table 3 and SEQ ID NOS: 142-146; U.S. Patent Publication No. 2008-0271205; U.S. Patent Publication No. 2008-0064063, and GenBank Accession No. GI 189098312. The steviol synthase from Arabidopsis thaliana is classified as CYP714A2. Nucleotide sequences encoding these KAH enzymes are set forth in SEQ ID NOS: 60-69. The nucleotide sequences set forth in SEQ ID NOS: 60-64 have been modified for expression in yeast, while the nucleotide sequences from the source organisms from which the polypeptides were identified are set forth in SEQ ID NOS: 65-69. [Table 4]

[0043] Furthermore, the KAH polypeptide from Stevia rebaudiana identified herein is particularly useful in recombinant hosts. The nucleotide sequence (SEQ ID NO: 163) encoding the S. rebaudiana KAH (SrKAHe1) (SEQ ID NO: 164) is shown in Figure 12A. The nucleotide sequence (SEQ ID NO: 165) encoding the S. rebaudiana KAH codon-optimized for expression in yeast is shown in Figure 12B. The amino acid sequence of the S. rebaudiana KAH is shown in Figure 12C. When expressed in S. cerevisiae, the S. rebaudiana KAH exhibits significantly higher steviol synthase activity than the Arabidopsis thaliana ent-kaurenoic acid hydroxylase described by Yamaguchi et al. (U.S. Patent Publication No. 2008 / 0271205 A1). The S. rebaudiana KAH polypeptide shown in Figure 12C has less than 20% identity to the KAH from US Patent Publication No. 2008 / 0271205 and less than 35% identity to the KAH from US Patent Publication No. 2008 / 0064063.

[0044] In some embodiments, the recombinant microorganism contains a recombinant gene encoding a KO and / or KAH polypeptide. Such a microorganism also typically contains a recombinant gene encoding a cytochrome P450 reductase (CPR) polypeptide, since certain combinations of KO and / or KAH polypeptides require expression of an exogenous CPR polypeptide. In particular, the activity of KO and / or KAH polypeptides of plant origin can be significantly enhanced by including a recombinant gene encoding an exogenous CPR polypeptide. Suitable CPR polypeptides are known. For example, suitable CPR enzymes include those produced by Stevia rebaudiana and Arabidopsis thaliana. See, e.g., Table 4 and SEQ ID NOs: 147 and 148. Nucleotide sequences encoding these polypeptides are set forth in SEQ ID NOs: 70, 71, 73, and 74. The nucleotide sequences set forth in SEQ ID NOs: 70-72 were modified for expression in yeast. Nucleotide sequences from the source organisms from which the polypeptides were identified are set forth in SEQ ID NOs: 73-75. [Table 5]

[0045] For example, steviol synthase encoded by SrKAHe1 is activated by the S. cerevisiae CPR encoded by the gene NCP1 (YHR042W). Better activation of steviol synthase encoded by SrKAHe1 is observed when coexpressed with the Arabidopsis thaliana CPR encoded by the gene ATR2 or the S. rebaudiana CPR encoded by the gene CPR8. Figure 13A contains the amino acid sequences of S. cerevisiae, A. thaliana (derived from the ATR1 and ATR2 genes), and S. rebaudiana CPR polypeptides (derived from the CPR7 and CPR8 genes) (SEQ ID NOS: 166-170). Figure 13B contains the nucleotide sequences encoding the A. thaliana and S. cerevisiae CPR polypeptides (SEQ ID NOS: 171-174).

[0046] For example, the yeast gene DPP1 and / or the yeast gene LPP1 can be disrupted or deleted to reduce the degradation of farnesyl pyrophosphate (FPP) to farnesol and reduce the degradation of geranylgeranyl diphosphate (GGPP) to geranylgeraniol (GGOH). Alternatively, promoter or enhancer elements of endogenous genes encoding phosphatases can be altered to modify the expression of the proteins they encode. Homologous recombination can be used to disrupt endogenous genes. For example, a "gene replacement" vector can be constructed in such a way that it contains a selectable marker gene. The selectable marker gene can be operably linked at both the 5' and 3' ends to a portion of the gene of sufficient length to mediate homologous recombination. The selectable marker can be any of a number of genes that complement auxotrophies of the host cell, confer antibiotic resistance, or cause a color change. The linearized DNA fragment of the gene replacement vector is then introduced into cells using methods well known in the art (see below). The integration of the linear fragment into the genome and the disruption of the gene can be determined based on the selection marker, and can be confirmed, for example, by Southern blot analysis.After its use in selection, the selection marker can be removed from the genome of the host cell, for example, by the Cre-loxP system (see, for example, Gossen et al. (2002) Ann. Rev. Genetics 36:153-173 and U.S. Patent Publication No. 20060014264).Alternatively, the gene replacement vector can be constructed in such a way that it contains a part of the gene to be disrupted, where this part lacks any promoter sequence of the endogenous gene and does not encode the coding sequence of the gene or encodes an inactive fragment thereof. An "inactive fragment" is a fragment of a gene that encodes a protein having, for example, less than about 10% (e.g., less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, less than about 1%, or 0%) of the activity of the protein produced from the full-length coding sequence of the gene.Such a portion of the gene is inserted into a vector in such a way that a known promoter sequence is not operably linked to the gene sequence, but a stop codon and a transcription termination sequence are operably linked to the portion of the gene sequence. This vector can then be linearized at the portion of the gene sequence and transformed into a cell. By the method of single homologous recombination, this linearized vector is then integrated into the endogenous counterpart of the gene. Expression of these genes in a recombinant microorganism results in the conversion of geranylgeranyl diphosphate to steviol.

[0047] B. Steviol Glycoside Biosynthesis Polypeptides The recombinant hosts described herein are capable of converting steviol to steviol glycosides. Such hosts (e.g., microorganisms) also contain genes encoding one or more UDP glycosyltransferases, also known as UGTs. UGTs transfer monosaccharide units from activated nucleotide sugars to an acceptor moiety, in this case the -OH or -COOH moiety of steviol or a steviol derivative. UGTs have been classified into families and subfamilies based on sequence homology. Li et al. J. Biol. Chem. 276:4338-4343 (2001).

[0048] B.1 Rubusoside Biosynthetic Polypeptides The biosynthesis of rubusoside involves glycosylation of the 13-OH and 19-COOH of steviol. See Figure 2A. Conversion of steviol to rubusoside in a recombinant host, such as a microorganism, can be achieved by expression of genes encoding UGT85C2 and UGT74G1, which transfer a glucose unit to the 13-OH or 19-COOH of steviol, respectively. A suitable UGT85C2 functions as a uridine 5'-diphosphoglucosyl:steviol 13-OH transferase and a uridine 5'-diphosphoglucosyl:steviol-19-O-glycoside 13-OH transferase. Functional UGT85C2 polypeptides can also catalyze glucosyltransferase reactions utilizing steviol glycoside substrates other than steviol and steviol-19-O-glucoside. Suitable UGT74G1 polypeptides function as uridine 5'-diphosphoglucosyl:steviol 19-COOH transferases and uridine 5'-diphosphoglucosyl:steviol-13-O-glycoside 19-COOH transferases. Functional UGT74G1 polypeptides can also catalyze glycosyltransferase reactions that utilize steviol glycoside substrates other than steviol and steviol-13-O-glucoside, or that transfer sugar moieties from acceptors other than uridine diphosphate glucose.

[0049] A recombinant microorganism expressing a functional UGT74G1 and a functional UGT85C2 can make rubusoside and both steviol monosides (i.e., steviol-13-O-monoglucoside and steviol-19-O-monoglucoside) when steviol is used as a feedstock in the culture medium. One or more of these genes may be naturally present in the host. Typically, however, such genes are recombinant genes transformed into a host (e.g., a microorganism) that does not naturally possess them. As used herein, the term recombinant host is intended to refer to a host whose genome has been augmented by at least one integrated DNA sequence. Such DNA sequences include, but are not limited to, non-naturally occurring genes, DNA sequences that are not normally transcribed into RNA or translated into protein ("expressed"), and other genes or DNA sequences that one desires to introduce into a non-recombinant host. Typically, it is understood that the genome of a recombinant host described herein is augmented by the stable introduction of one or more recombinant genes. Generally, the introduced DNA is not naturally present in the host that is the recipient of the DNA, but it is within the scope of the present invention to isolate a DNA segment from a given host and subsequently introduce one or more additional copies of that DNA into the same host, e.g., to enhance production of a gene's product or alter the gene's expression pattern. In some instances, the introduced DNA also modifies or replaces an endogenous gene or DNA sequence, e.g., by homologous recombination or site-directed mutagenesis. Suitable recombinant hosts include microorganisms, plant cells, and plants.

[0050] The term "recombinant gene" refers to a gene or DNA sequence that is introduced into a recipient host, regardless of whether an identical or similar gene or DNA sequence may already be present in such host. In this context, "introduced" or "augmented" are known in the art to mean introduced or augmented by the hand of man. Thus, a recombinant gene may be a DNA sequence from another species, or it may be a DNA sequence derived from or present in the same species, but which has been incorporated into the host by recombinant methods to form a recombinant host. It is understood that a recombinant gene introduced into a host may be identical to a DNA sequence normally present in the host being transformed, and that is introduced to provide one or more additional copies of DNA, thereby allowing overexpression or modified expression of the gene product of that DNA. Suitable UGT74G1 and UGT85C2 polypeptides include those produced by Stevia rebaudiana. Genes encoding functional UGT74G1 and UGT85C2 polypeptides from Stevia are reported in Richman, et al. Plant J. 41: 56-67 (2005). The amino acid sequences of the S. rebaudiana UGT74G1 and UGT85C2 polypeptides are set forth in SEQ ID NOs: 1 and 3, respectively. Nucleotide sequences encoding UGT74G1 and UGT85C2 optimized for expression in yeast are set forth in SEQ ID NOs: 2 and 4, respectively. DNA2.0 codon-optimized sequences for UGTs 85C2, 91D2e, 74G1, and 76G1 are set forth in SEQ ID NOs: 82, 84, 83, and 85, respectively. See also the variants of UGT85C2 and UGT74G1 described in the "Functional Homologs" section below. For example, UGT85C2 polypeptides containing substitutions at positions 65, 71, 270, 289, and 389 can be used (e.g., A65S, E71Q, T270M, Q289H, and A389V).

[0051] In some embodiments, the recombinant host is a microorganism. The recombinant microorganism can be grown on a medium containing steviol to produce rubusoside. In other embodiments, however, the recombinant microorganism expresses one or more recombinant genes involved in steviol biosynthesis, such as a CDPS gene, a KS gene, a KO gene, and / or a KAH gene. Suitable CDPS polypeptides are known. For example, suitable CDPS enzymes include those produced by Stevia rebaudiana, Streptomyces clavuligerus, Bradyrhizobium japonicum, Zea mays, and Arabidopsis. See, e.g., Table 5 and SEQ ID NOS: 129-131, 158, and 160. Nucleotide sequences encoding these polypeptides are set forth in SEQ ID NOS: 34-39, 157, and 159. The nucleotide sequences set forth in SEQ ID NOS: 34-36 were modified for expression in yeast. Nucleotide sequences from the source organisms from which the polypeptides were identified were identified and are set forth in SEQ ID NOS: 37-39.

[0052] In some embodiments, a CDPS polypeptide lacking the chloroplast transit peptide at the amino terminus of the unmodified polypeptide can be used. For example, the first 150 nucleotides from the 5' end of the Zea mays CDPS coding sequence shown in Figure 14 (SEQ ID NO: 157) can be removed. This removes the amino-terminal 50 residues of the amino acid sequence shown in Figure 14 (SEQ ID NO: 158), which encodes the chloroplast transit peptide. The truncated CDPS gene can be fitted with a new ATG translation start site and operably linked to a promoter, typically a constitutive or high-expression promoter. When multiple copies of the truncated coding sequence are introduced into a microbial expression system, expression of the CDPS polypeptide from the promoter results in increased carbon flux directed toward ent-kaurene biosynthesis. [Table 6]

[0053] CDPS-KS bifunctional proteins (SEQ ID NOS:136 and 137) may also be used. The nucleotide sequence encoding the CDPS-KS bifunctional enzyme shown in Table 6 was modified for expression in yeast (see SEQ ID NOS:48 and 49). The nucleotide sequence from the source organism in which the polypeptide was originally identified is set forth in SEQ ID NOS:50 and 51. The nucleotide sequence encoding the Gibberella fujikuroi bifunctional CDPS-KS enzyme was modified for expression in yeast (see Figure 15A, SEQ ID NO:161). [Table 7]

[0054] Thus, a microorganism containing UGT74G1 and UGT85C2 as well as CDPS, KS, KO, and KAH genes can produce both steviolmonoside and rubusoside without the need to use steviol as a feedstock. In some embodiments, the recombinant microorganism further expresses a recombinant gene encoding geranylgeranyl diphosphate synthase (GGPPS). Suitable GGPPS polypeptides are known. For example, suitable GGPPS enzymes include those produced by Stevia rebaudiana, Gibberella fujikuroi, Mus musculus, Thalassiosira pseudonana, Streptomyces clavuligerus, Sulfulobus acidocaldarius, Synechococcus species, and Arabidopsis thaliana. See Table 7 and SEQ ID NOs: 121-128. Nucleotide sequences encoding these polypeptides are set forth in SEQ ID NOs: 18-33. The nucleotide sequences set forth in SEQ ID NOs: 18-25 were modified for expression in yeast, while the nucleotide sequences from the source organisms from which the polypeptides were identified are set forth in SEQ ID NOs: 23-26.

[0055] [Table 8] In some embodiments, the recombinant microorganism can further express recombinant genes involved in diterpene biosynthesis or production of terpenoid precursors, such as genes in the methylerythritol 4-phosphate (MEP) pathway or genes in the mevalonate (MEV) pathway described below, have reduced phosphatase activity, and / or express sucrose synthase (SUS), as described herein.

[0056] B.2 Biosynthetic Polypeptides of Rebaudioside A, Rebaudioside D, and Rebaudioside E The biosynthesis of rebaudioside A involves the glucosylation of the aglycone steviol. Specifically, rebaudioside A is formed by glucosylation of the 13-OH of steviol to form 13-O-steviolmonoside, glucosylation of the C-2' of the 13-O-glucose of steviolmonoside to form steviol-1,2-bioside, glucosylation of the C-19 carboxyl of steviol-1,2-bioside to form stevioside, and glucosylation of the C-3' of the C-13-O-glucose of stevioside. The order in which each glucosylation reaction occurs can vary. See Figure 2A. The biosynthesis of rebaudioside E and / or rebaudioside D involves glucosylation of the aglycone steviol. Specifically, rebaudioside E is formed by glucosylation of the 13-OH of steviol to form steviol-13-O-glucoside, glucosylation of the C-2' of the 13-O-glucose of steviol-13-O-glucoside to form steviol-1,2-bioside, glucosylation of the C-19 carboxyl of the 1,2-bioside to form 1,2-stevioside, and glucosylation of the C-2' of the 19-O-glucose of 1,2-stevioside to form rebaudioside E. Rebaudioside D is formed by glucosylation of the C-3' of the C-13-O-glucose of rebaudioside E. The order in which each glycosylation reaction occurs can be varied. For example, glucosylation of the C-2' of 19-O-glucose may be the final step in the pathway, where rebaudioside A is an intermediate in the pathway. See Figure 2C.

[0057] It has been discovered that conversion of steviol to rebaudioside A, rebaudioside D, and / or rebaudioside E in a recombinant host can be achieved by expressing the following functional UGTs: EUGT11, 74G1, 85C2, and 76G1, and optionally 91D2. Thus, a recombinant microorganism expressing a combination of these four or five UGTs can make rebaudioside A and rebaudioside D when steviol is used as a feedstock. Typically, one or more of these genes are recombinant genes transformed into a microorganism that does not naturally possess them. It has also been discovered that a UGT, referred to herein as SM12UGT, can substitute for UGT91D2. In some embodiments, fewer than five UGTs (e.g., 1, 2, 3, or 4) are expressed in a host. For example, a recombinant microorganism expressing a functional EUGT11 can make rebaudioside D when rebaudioside A is used as a feedstock. A recombinant microorganism expressing two functional UGTs, EUGT11 and 76G1, and optionally a functional 91D12, can make rebaudioside D when rubusoside or 1,2-stevioside is used as a feedstock. As another alternative, a recombinant microorganism expressing three functional UGTs, EUGT11, 74G1, 76G1, and optionally 91D12, can make rebaudioside D when fed the monoside steviol-13-O-glucoside in the medium. Similarly, conversion of steviol-19-O-glucoside to rebaudioside D in a recombinant microorganism can be achieved by expression of genes encoding the UGTs EUGT11, 85C2, 76G1, and optionally 91D12 when steviol-19-O-glucoside is provided in the culture medium. Typically, one or more of these genes are recombinant genes transformed into a host that does not naturally possess them.

[0058] Suitable UGT74G1 and UGT85C2 polypeptides include those described above. A suitable UGT74G1 adds a glucose moiety to the C-3' of the C-13-O-glucose of an acceptor molecule, a steviol-1,2 glycoside. Thus, UGT74G1 functions, for example, as a uridine 5'-diphosphoglucosyl:steviol-13-O-1,2 glucoside C-3' glucosyltransferase and a uridine 5'-diphosphoglucosyl:steviol-19-O-glucose, 13-O-1,2 bioside C-3' glucosyltransferase. Functional UGT74G1 polypeptides can also catalyze glucosyltransferase reactions utilizing steviol glycoside substrates containing sugars other than glucose, such as steviol rhamnoside and steviol xyloside. See Figures 2A, 2B, 2C, and 2D. Suitable UGT76G1 polypeptides include those produced by S. rebaudiana and reported in Richman, et al. Plant J. 41: 56-67 (2005). The amino acid sequence of the S. rebaudiana UGT76G1 polypeptide is set forth in SEQ ID NO: 7. The nucleotide sequence encoding the UGT76G1 polypeptide of SEQ ID NO: 7 has been optimized for expression in yeast and is set forth in SEQ ID NO: 8. See also the UGT76G1 variants described in the "Functional Homologs" section.

[0059] A suitable EUGT11 or UGT91D2 polypeptide functions as a uridine 5'-diphosphoglucosyl:steviol-13-O-glucoside transferase (also called steviol-13-monoglucoside 1,2-glucosylase), transferring a glucose moiety to the C-2' of the 13-O-glucose of the acceptor molecule, steviol-13-O-glucoside. A suitable EUGT11 or UGT91D2 polypeptide also functions as a uridine 5'-diphosphoglucosyl:rubusoside transferase, transferring a glucose moiety to the C-2' of the 13-O-glucose of the acceptor molecule rubusoside to produce stevioside. EUGT11 polypeptides also transfer a glucose moiety to the C-2' of the 19-O-glucose of the acceptor molecule rubusoside to produce 19-O-1,2-diglycosylated rubusoside (compound 2 in Figure 3).

[0060] Functional EUGT11 or UGT91D2 polypeptides can also catalyze reactions that utilize steviol glycoside substrates other than steviol-13-O-glucoside and rubusoside. For example, a functional EUGT11 polypeptide can use stevioside as a substrate and transfer a glucose moiety to the C-2' of the 19-O-glucose residue to produce rebaudioside E (see compound 3 in Figure 3). A functional EUGT11 or UGT91D2 polypeptide can also use rebaudioside A as a substrate and transfer a glucose moiety to the C-2' of the 19-O-glucose residue of rebaudioside A to produce rebaudioside D. As shown in the examples, EUGT11 can convert rebaudioside A to rebaudioside D at a rate at least 20 times faster (e.g., at least 25 times or at least 30 times faster) than the rate of the corresponding UGT91D2e (SEQ ID NO: 5) when the reactions are carried out under similar conditions, i.e., similar time, temperature, purity, and substrate concentration. Thus, EUGT11 produces larger amounts of RebD than UGT91D2e when incubated under similar conditions. Furthermore, functional EUGT11 exhibits significant C-2' 19-O-diglycosylation activity using rubusoside or stevioside as substrates, whereas UGT91D2e has no detectable diglycosylation activity using these substrates. Thus, functional EUGT11 can be distinguished from UGT91D2e by their different steviol glycoside substrate specificities. Figure 3 provides a schematic diagram of the 19-O-1,2 diglycosylation reactions by EUGT11 and UGT91D2e.

[0061] Functional EUGT11 or UGT91D2 polypeptides typically do not transfer a glucose moiety to steviol compounds having a 1,3-linked glucose at the C-13 position, i.e., transfer of a glucose moiety to steviol 1,3-bioside and 1,3-stevioside does not occur. Functional EUGT11 and UGT91D2 polypeptides can transfer sugar moieties from donors other than uridine diphosphate glucose. For example, a functional EUGT11 or UGT91D2 polypeptide can act as a uridine 5'-diphospho D-xylosyl:steviol-13-O-glucoside transferase, transferring a xylose moiety to the C-2' of the 13-O-glucose of an acceptor molecule, steviol-13-O-glucoside. As another example, a functional EUGT11 or UGT91D2 polypeptide can act as a uridine 5'-diphospho L-rhamnosyl:steviol-13-O-glucoside transferase, transferring a rhamnose moiety to the C-2' of the 13-O-glucose of an acceptor molecule, steviol-13-O-glucoside. Suitable EUGT11 polypeptides are described herein and can include the EUGT11 polypeptide from Oryza sativa (GenBank Accession No. AC133334). For example, the EUGT11 polypeptide can have an amino acid sequence having at least 70% sequence identity (e.g., at least 75, 80, 85, 90, 95, 96, 97, 98, or 99% sequence identity) to the amino acid sequence set forth in SEQ ID NO: 152 (see FIG. 7). The nucleotide sequence encoding the amino acid sequence of SEQ ID NO: 152 is set forth in SEQ ID NO: 153. SEQ ID NO: 154 is a nucleotide sequence encoding the polypeptide of SEQ ID NO: 152 that has been codon-optimized for expression in yeast.

[0062] Suitable functional UGT91D2 polypeptides include those disclosed herein, e.g., the polypeptides designated UGT91D2e and UGT91D2m. The amino acid sequence of an exemplary UGT91D2e polypeptide from Stevia rebaudiana is set forth in SEQ ID NO:5. SEQ ID NO:6 is a nucleotide sequence encoding the polypeptide of SEQ ID NO:5, codon-optimized for expression in yeast. The S. rebaudiana nucleotide sequence encoding the polypeptide of SEQ ID NO:5 is set forth in SEQ ID NO:9. The amino acid sequences of exemplary UGT91D2m polypeptides from S. rebaudiana are set forth in SEQ ID NOs:10 and 12, which are encoded by the nucleic acid sequences set forth in SEQ ID NOs:11 and 13, respectively. Additionally, UGT91D2 variants containing substitutions at amino acid residues 206, 207, and 343 of SEQ ID NO:5 can be used. For example, the amino acid sequence set forth in SEQ ID NO:95, which has the following mutations relative to wild-type UGT92D2e (SEQ ID NO:5): G206R, Y207C, and W343R, can be used. Additionally, one can use a UGT91D2 variant that includes substitutions at amino acid residues 211 and 286. For example, a UGT91D2 variant can include a substitution of methionine for leucine at position 211 and an alanine for valine at position 286 of SEQ ID NO: 5 (UGT91D2e-b).

[0063] As noted above, the UGT referred to herein as SM12UGT can replace UGT91D2. Suitable functional SM12UGT polypeptides include those produced by Ipomoea purpurea (Japanese morning glory) and described in Morita et al. Plant J. 42, 353-363 (2005). The amino acid sequence encoding the I. purpurea IP3GGT polypeptide is set forth in SEQ ID NO: 76. SEQ ID NO: 77 is a nucleotide sequence encoding the polypeptide of SEQ ID NO: 76, codon-optimized for expression in yeast. Another suitable SM12UGT polypeptide is a Bp94B1 polypeptide with an R25S mutation. See Osmani et al. Plant Phys. 148: 1295-1308 (2008) and Sawada et al. J. Biol. Chem. 280:899-906 (2005). The amino acid sequence of the Bellis perennis (red daisy) UGT94B1 polypeptide is set forth in SEQ ID NO: 78. SEQ ID NO: 79 is a nucleotide sequence encoding the polypeptide of SEQ ID NO: 78 that has been codon-optimized for expression in yeast.

[0064] In some embodiments, the recombinant microorganism is grown on a medium containing steviol-13-O-glucoside or steviol-19-O-glucoside to produce rebaudioside A and / or rebaudioside D. In such embodiments, the microorganism contains and expresses genes encoding a functional EUGT11, a functional UGT74G1, a functional UGT85C2, a functional UGT76G1, and optionally a functional UGT91D2, and is capable of accumulating rebaudioside A and rebaudioside D when steviol, one or both of the steviolmonosides, or rubusoside is used as a feedstock. In another embodiment, the recombinant microorganism is grown on medium containing rubusoside to produce rebaudioside A and / or rebaudioside D. In such embodiments, the microorganism contains and expresses genes encoding a functional EUGT11, a functional UGT76G1, and optionally a functional UGT91D2, and is capable of producing rebaudioside A and / or rebaudioside D when rubusoside is used as a feedstock.

[0065] In another embodiment, the recombinant microorganism expresses one or more genes involved in steviol biosynthesis, e.g., a CDPS gene, a KS gene, a KO gene, and / or a KAH gene. Thus, for example, a microorganism containing a CDPS gene, a KS gene, a KO gene, and a KAH gene, in addition to EUGT11, UGT74G1, UGT85C2, UGT76G1, and optionally a functional UGT91D2 (e.g., UGT91D2e), can produce rebaudioside A, rebaudioside D, and / or rebaudioside E without the need to include steviol in the culture medium. In some embodiments, the recombinant host further contains and expresses a recombinant GGPPS gene to increase levels of the diterpene precursor geranylgeranyl diphosphate and increase flux through the steviol biosynthetic pathway. In some embodiments, the recombinant host further contains a construct that suppresses expression of non-steviol pathways that consume geranylgeranyl diphosphate, ent-kaurenoic acid, or farnesyl pyrophosphate, thereby increasing flux through the steviol and steviol glycoside biosynthetic pathways. For example, flux to sterol production pathways, such as ergosterol, can be reduced by downregulating the ERG9 gene. See the ERG9 section below and Examples 24-25. In cells that produce gibberellins, gibberellin synthesis can be downregulated to increase flux of ent-kaurenoic acid to steviol. In carotenoid-producing organisms, flux to steviol can be increased by downregulating one or more carotenoid biosynthetic genes. In some embodiments, the recombinant microorganism can further express recombinant genes involved in diterpene biosynthesis or production of terpenoid precursors, such as genes in the MEP or MEV pathways described below, have reduced phosphatase activity, and / or express a SUS as described herein.

[0066] Those skilled in the art will recognize that by adjusting the relative expression levels of different UGT genes, a recombinant host can be tailored to specifically produce steviol glycoside products in desired proportions. Transcriptional regulation of steviol biosynthesis genes and steviol glycoside biosynthesis genes can be achieved by a combination of transcriptional activation and repression, using techniques well known to those skilled in the art. In in vitro reactions, those skilled in the art will recognize that adding various levels of combinations of UGT enzymes, or conditions that affect the relative activity of different UGT combinations, will direct synthesis to desired proportions of each steviol glycoside. Those skilled in the art will recognize that a higher proportion of rebaudioside D or E, or more efficient conversion to rebaudioside D or E, can be obtained with a diglycosylation enzyme that has higher activity for the 19-O-glucoside reaction compared to the 13-O-glucoside reaction (substrates rebaudioside A and stevioside).

[0067] In some embodiments, a recombinant host such as a microorganism produces a rebaudioside D-enriched steviol glycoside composition having greater than at least 3% rebaudioside D by weight of total steviol glycosides, e.g., at least 4% rebaudioside D, at least 5% rebaudioside D, 10-20% rebaudioside D, 20-30% rebaudioside D, 30-40% rebaudioside D, 40-50% rebaudioside D, 50-60% rebaudioside D, 60-70% rebaudioside D, or 70-80% rebaudioside D. In some embodiments, a recombinant host such as a microorganism produces a steviol glycoside composition having at least 90% rebaudioside D, e.g., 90-99% rebaudioside D. Other steviol glycosides present include those shown in Figure 2C, e.g., steviol monosides, steviol glucobiosides, rebaudioside A, rebaudioside E, and stevioside. In some embodiments, the rebaudioside D-enriched composition produced by a host (e.g., a microorganism) is further purified, and the rebaudioside D or rebaudioside E thus purified can then be mixed with other steviol glycosides, flavors, or sweeteners to obtain a desired flavor system or sweetener composition. For example, a rebaudioside D-enriched composition produced by a recombinant host can be combined with a rebaudioside A-, C-, or F-enriched composition produced by a different recombinant host, where rebaudioside A, F, or C is purified from a Stevia extract or produced in vitro.

[0068] In some embodiments, rebaudioside A, rebaudioside D, rebaudioside B, steviol monoglucoside, steviol-1,2-bioside, rubusoside, stevioside, or rebaudioside E can be produced using in vitro methods while providing appropriate UDP-sugars and / or a cell-free system for regeneration of UDP-sugars. See, e.g., Jewett MC, et al. Molecular Systems Biology, Vol. 4, article 220 (2008); Masada S et al. FEBS Letters, Vol. 581, 2562-2566 (2007). In some embodiments, sucrose and sucrose synthase can be provided in the reaction vessel to regenerate UDP-glucose from UDP generated during the glycosylation reaction. See FIG. 11. The sucrose synthase can be derived from any suitable organism. For example, a sucrose synthase coding sequence from Arabidopsis thaliana, Stevia rebaudiana or Coffea arabica can be cloned into an expression plasmid under the control of a suitable promoter and expressed in a host, such as a microorganism or a plant.

[0069] Conversions requiring multiple reactions can be carried out simultaneously or stepwise. For example, rebaudioside D can be produced from rebaudioside A, commercially available as an enriched extract or produced via biosynthesis, by adding stoichiometric or excess amounts of UDP-glucose and EUGT11. Alternatively, rebaudioside D can be produced from a steviol glycoside extract enriched for stevioside and rebaudioside A using EUGT11 and a suitable UGT76G1 enzyme. In some embodiments, a phosphatase is used to remove secondary products and improve reaction yields. UGTs and other enzymes for in vitro reactions can be provided in soluble or immobilized forms.

[0070] In some embodiments, rebaudioside A, rebaudioside D, or rebaudioside E can be produced using whole cells fed with raw materials containing precursor molecules such as steviol and / or steviol glycosides, including a mixture of steviol glycosides derived from plant extracts. The raw materials can be fed during or after cell growth. The whole cells can be in suspension or immobilized. The whole cells can be entrapped in beads, such as calcium alginate beads or sodium alginate beads. The whole cells can be connected to a hollow fiber tubular reactor system. The whole cells can be concentrated and encapsulated in a membrane reactor system. The whole cells can be in fermentation broth or reaction buffer. In some embodiments, a permeabilizing agent is utilized for efficient transfer of substrates into the cells. In some embodiments, the cells are permeabilized with a solvent such as toluene or a detergent such as Triton-X or Tween. In some embodiments, the cells are permeabilized with a detergent, such as a cationic detergent such as cetyltrimethylammonium bromide (CTAB). In some embodiments, cells are permeabilized by electroporation or periodic mechanical shock, such as a slight osmotic shock. The cells can contain one recombinant UGT or multiple recombinant UGTs. For example, the cells can contain UGT 76G1 and EUGT11, such that a mixture of stevioside and RebA is efficiently converted to RebD. In some embodiments, the whole cells are host cells described in Section IIIA. In some embodiments, the whole cells are Gram-negative bacteria, such as E. coli. In some embodiments, the whole cells are Gram-positive bacteria, such as Bacillus. In some embodiments, the whole cells are fungal species, such as Aspergillus, or yeast, such as Saccharomyces. In some embodiments, the term "whole-cell biocatalysis" is used to refer to a process in which whole cells are grown (e.g., in a medium and optionally permeabilized) as described above, and a substrate, such as rebA or stevioside, is provided and converted to an end product using enzymes from the cells. The cells may or may not be viable, and may or may not grow during the bioconversion reaction.In contrast, in fermentation, cells are cultured in a growth medium, supplied with carbon and an energy source such as glucose, and an end product is produced by living cells.

[0071] B.3 Biosynthetic polypeptides of dulcoside A and rebaudioside C The biosynthesis of rebaudioside C and / or dulcoside A involves glucosylation and rhamnosylation of the aglycone steviol. Specifically, dulcoside A can be formed by glucosylation of the 13-OH of steviol to form steviol-13-O-glucoside, rhamnosylation of the C-2' of the 13-O-glucoside of steviol-13-O-glucoside to form the 1,2 rhamnobioside, and glucosylation of the C-19 carboxyl of the 1,2 rhamnobioside. Rebaudioside C can be formed by glucosylation of the C-3' of the C-13-O-glucose of dulcoside A. The order in which each glycosylation reaction occurs can vary. See Figure 2B.

[0072] It has been discovered that conversion of steviol to dulcoside A in a recombinant host can be achieved by expression of a gene(s) encoding the following functional UGTs: 85C2, EUGT11 and / or 91D2e, and 74G1. Thus, recombinant microorganisms expressing these three or four UGTs and rhamnose synthase can make dulcoside A when provided with steviol in the medium. Alternatively, recombinant microorganisms expressing two UGTs, EUGT11 and 74G1, and rhamnose synthase can produce dulcoside A when provided with the monosides, steviol-13-O-glucoside or steviol-19-O-glucoside, in the medium. Similarly, conversion of steviol to rebaudioside C in a recombinant microorganism can be achieved by expression of a gene(s) encoding UGTs 85C2, EUGT11, 74G1, 76G1, optionally 91D2, and rhamnose synthase when feeding steviol; by expression of genes encoding UGTs EUGT11, and / or 91D2, 74G1, and 76G1, and rhamnose synthase when feeding steviol-13-O-glucoside; by expression of genes encoding UGTs 85C2, EUGT11, and / or 91D2e, 76G1, and rhamnose synthase when feeding steviol-19-O-glucoside; or by expression of genes encoding UGTs EUGT11, and / or 91D2e, 76G1, and rhamnose synthase when feeding rubusoside. Typically, one or more of these genes are recombinant genes that have been transformed into a microorganism that does not naturally possess them.

[0073] Suitable EUGT11, UGT91D2, UGT74G1, UGT76G1, and UGT85C2 polypeptides include the functional UGT polypeptides described herein. Rhamnose synthetase increases the amount of UDP-rhamnose donor for rhamnosylation of the steviol compound acceptor. Suitable rhamnose synthetases include those made by Arabidopsis thaliana, such as the product of the A. thaliana RHM2 gene. In some embodiments, the UGT79B3 polypeptide replaces the UGT91D2 polypeptide. Suitable UGT79B3 polypeptides include those made by Arabidopsis thaliana, which are capable of rhamnosylating steviol-13-O-monoside in vitro. A. thaliana UGT79B3 can rhamnosylate glucosylated compounds to form 1,2-rhamnosides. The amino acid sequence of Arabidopsis thaliana UGT79B3 is set forth in SEQ ID NO: 150. The nucleotide sequence encoding the amino acid sequence of SEQ ID NO: 150 is set forth in SEQ ID NO: 151.

[0074] In some embodiments, rebaudioside C can be produced using in vitro methods by providing an appropriate UDP-sugar and / or a cell-free system for regenerating the UDP-sugar. See, for example, Jewett MC, Calhoun KA, Voloshin A, Wuu JJ, and Swartz JR, "An integrated cell-free metabolic platform for protein production and synthetic biology," in Molecular Systems Biology, 4, article 220 (2008) and Masada S et al., FEBS Letters, Vol. 581, 2562-2566 (2007). In some embodiments, sucrose and sucrose synthase can be provided in the reaction vessel to regenerate UDP-glucose from UDP during the glycosylation reaction. See FIG. 11. The sucrose synthase can be derived from any suitable organism. For example, a sucrose synthase coding sequence from Arabidopsis thaliana, Stevia rebaudiana, or Coffea arabica can be cloned into an expression plasmid under the control of a suitable promoter and expressed in a host (e.g., a microorganism or a plant). In some embodiments, the RHM2 enzyme (rhamnose synthase) can be supplied with NADPH to produce UDP-rhamnose from UDP-glucose.

[0075] Reactions can be carried out simultaneously or stepwise. For example, rebaudioside C can be produced by adding stoichiometric amounts of UDP-rhamnose and EUGT11, followed by UGT76G1 and a stoichiometric or excess amount of UDP-glucose. In some embodiments, phosphatases are used to remove secondary products and improve reaction yields. UGTs and other enzymes for in vitro reactions can be provided in soluble or immobilized forms. In some embodiments, rebaudioside C, dulcoside A, or other steviol rhamnosides can be produced using whole cells as described above. The cells can contain one recombinant UGT or multiple recombinant UGTs. For example, the cells can contain UGT76G1 and EUGT11 so that a mixture of stevioside and RebA is efficiently converted to RebD. In some embodiments, the whole cells are host cells described in Section IIIA. In other embodiments, the recombinant host expresses one or more genes involved in steviol biosynthesis, such as a CDPS gene, a KS gene, a KO gene, and / or a KAH gene. Thus, for example, a microorganism containing a CDPS gene, a KS gene, a KO gene, and a KAH gene in addition to a UGT85C2, a UGT74G1, a EUGT11 gene, optionally a UGT91D2e gene, and a UGT76G1 gene can produce rebaudioside C without the need to include steviol in the culture medium. Additionally, the recombinant host typically expresses an endogenous or recombinant gene encoding a rhamnose synthase. Such genes are useful for providing increased amounts of UDP-rhamnose donor for rhamnosylation of steviol compound acceptors. Suitable rhamnose synthases include those made by Arabidopsis thaliana, such as the product of the A. thaliana RHM2 gene.

[0076] Those skilled in the art will recognize that by modulating the relative expression levels of different UGT genes and the availability of UDP-rhamnose, a recombinant host can be tailored to specifically produce steviol glycoside products in desired proportions. Transcriptional regulation of steviol biosynthesis genes and steviol glycoside biosynthesis genes can be achieved by a combination of transcriptional activation and repression, using techniques well known to those skilled in the art. In in vitro reactions, those skilled in the art will recognize that adding various levels of combinations of UGT enzymes, or under conditions that affect the relative activity of combinations of different UGT enzymes, will direct synthesis to desired proportions of each steviol glycoside. In some embodiments, the recombinant host further contains and expresses a recombinant GGPPS gene to increase levels of the diterpene precursor geranylgeranyl diphosphate and thereby increase flux through the rebaudioside A biosynthetic pathway. In some embodiments, the recombinant host further contains a construct that suppresses or reduces expression of non-steviol pathways that consume geranylgeranyl diphosphate, ent-kaurenoic acid, or farnesyl pyrophosphate, thereby increasing flux through the steviol and steviol glycoside biosynthetic pathways. For example, flux to sterol production pathways, such as ergosterol, can be reduced by downregulating the ERG9 gene. See the ERG9 section below and Examples 24-25. In cells that produce gibberellins, gibberellin synthesis can be downregulated to increase flux of ent-kaurenoic acid to steviol. In carotenoid-producing organisms, flux to steviol can be increased by downregulating one or more carotenoid biosynthetic genes.

[0077] In some embodiments, the recombinant host further contains and expresses recombinant genes involved in diterpene biosynthesis or production of terpenoid precursors, e.g., genes in the MEP or MEV pathway, has reduced phosphatase activity, and / or expresses a SUS as described herein. In some embodiments, a recombinant host such as a microorganism produces a steviol glycoside composition having greater than at least 15% rebaudioside C relative to total steviol glycosides, e.g., at least 20% rebaudioside C, 30-40% rebaudioside C, 40-50% rebaudioside C, 50-60% rebaudioside C, 60-70% rebaudioside C, 70-80% rebaudioside C, or 80-90% rebaudioside C. In some embodiments, a recombinant host such as a microorganism produces a steviol glycoside composition having at least 90% rebaudioside C, e.g., 90-99% rebaudioside C. Other steviol glycosides present may include those shown in Figures 2A and B, such as steviol monosides, steviol glucobiosides, steviol rhamnobiosides, rebaudioside A, and dulcoside A. In some embodiments, the rebaudioside C-enriched composition produced by the host is further purified, and the rebaudioside C or dulcoside A thus purified is then mixed with other steviol glycosides, flavors, or sweeteners to obtain a desired flavor system or sweetening composition. For example, a rebaudioside C-enriched composition produced by a recombinant microorganism can be combined with a rebaudioside A, F, or D-enriched composition produced by a different recombinant microorganism, with rebaudioside A, F, or D purified from a Stevia extract, or with rebaudioside A, F, or D produced in vitro.

[0078] B.4 Rebaudioside F Biosynthetic Polypeptides The biosynthesis of rebaudioside F involves glucosylation and xylosylation of the aglycone steviol. Specifically, rebaudioside F is formed by glucosylation of the 13-OH of steviol to form steviol-13-O-glucoside, xylosylation of the C-2' of the 13-O-glucose of steviol-13-O-glucoside to form steviol-1,2-xylobioside, glucosylation of the C-19 carboxyl of the 1,2-xylobioside to form 1,2-stevioxyloside, and glucosylation of the C-3' of the C-13-O-glucose of 1,2-stevioxyloside to form rebaudioside F. The order in which each glycosylation reaction occurs can be varied. See Figure 2D. It has been discovered that conversion of steviol to rebaudioside F in a recombinant host can be achieved by expressing genes encoding the following functional UGTs: 85C2, EUGT11, and / or 91D2e, 74G1, and 76G1, along with endogenous or recombinantly expressed UDP-glucose dehydrogenase and UDP-glucuronic acid decarboxylase. Thus, a recombinant microorganism expressing these four or five UGTs, along with endogenous or recombinant UDP-glucose dehydrogenase and UDP-glucuronic acid decarboxylase, can make rebaudioside F when fed with steviol in the medium. Alternatively, a recombinant microorganism expressing two functional UGTs, EUGT11 or 91D2e and 76G1, can make rebaudioside F when fed with rubusoside in the medium. As another alternative, a recombinant microorganism expressing functional UGT76G1 can make rebaudioside F when fed 1,2-steviorhamnoside. As another alternative, a recombinant microorganism expressing 76G1, EUGT11, and / or 91D2e, 76G1 can make rebaudioside F when fed the monoside, steviol-13-O-glucoside, in the medium. Similarly, conversion of steviol-19-O-glucoside to rebaudioside F in a recombinant microorganism can be achieved by expression of genes encoding the UGTs 85C2, EUGT11, and / or 91D2e, and 76G1 when fed steviol-19-O-glucoside. Typically, one or more of these genes are recombinant genes transformed into a host that does not naturally possess them.

[0079] Suitable EUGT11, UGT91D2, UGT74G1, UGT76G1, and UGT85C2 polypeptides include the functional UGT polypeptides described herein. In some embodiments, as described above, the UGT79B3 polypeptide replaces UGT91. UDP-glucose dehydrogenase and UDP-glucuronic acid decarboxylase increase the amount of UDP-xylose donor for xylosylation of the steviol compound acceptor. Suitable UDP-glucose dehydrogenase and UDP-glucuronic acid decarboxylase polypeptides include those produced by Arabidopsis thaliana or Cryptococcus neoformans. For example, suitable UDP-glucose dehydrogenase and UDP-glucuronic acid decarboxylase polypeptides can be encoded by the UGD1 gene and UXS3 gene, respectively, of A. thaliana. Oka and Jigami, FEBS J. 273:2645-2657 (2006). In some embodiments, rebaudioside F can be produced using in vitro methods, providing the appropriate UDP-sugar and / or a cell-free system for regeneration of the UDP-sugar. See, e.g., Jewett MC, et al. Molecular Systems Biology , Vol. 4, article 220 (2008);Masada S et al. FEBS Letters, Vol. 581, 2562-2566 (2007). In some embodiments, sucrose and sucrose synthase may be provided in the reactor to regenerate UDP-glucose from UDP during the glycosylation reaction. See FIG. 11. The sucrose synthase can be derived from any suitable organism. For example, the sucrose synthase coding sequence from Arabidopsis thaliana, Stevia rebaudiana, or Coffea arabica can be cloned into an expression plasmid under the control of a suitable promoter and expressed in a host (e.g., a microorganism or a plant). In some embodiments, UDP-xylose can be produced from UDP-glucose by providing appropriate enzymes, such as the Arabidopsis thaliana UGD1 (UDP-glucose dehydrogenase) and UXS3 (UDP-glucuronic acid decarboxylase) enzymes, along with the NAD+ cofactor.

[0080] Reactions can be performed simultaneously or stepwise. For example, rebaudioside F can be produced from rubusoside by adding stoichiometric amounts of UDP-xylose and EUGT11, followed by UGT76G1 and an excess or stoichiometric amount of UDP-glucose. In some embodiments, phosphatases are used to remove secondary products and improve reaction yields. UGTs and other enzymes for in vitro reactions can be provided in soluble or immobilized form. In some embodiments, rebaudioside F or other steviol xylosides can be produced using whole cells as described above. For example, the cells can contain UGT76G1 and EUGT11 so that a mixture of stevioside and RebA is efficiently converted to RebD. In some embodiments, the whole cells are host cells described in Section IIIA. In other embodiments, the recombinant host expresses one or more genes involved in steviol biosynthesis, such as a CDPS gene, a KS gene, a KO gene, and / or a KAH gene. Thus, for example, a microorganism containing a CDPS gene, a KS gene, a KO gene, and a KAH gene in addition to EUGT11, UGT85C2, UGT74G1, optionally a UGT91D2e gene, and a UGT76G1 gene can produce rebaudioside F without the need to include steviol in the culture medium. Additionally, the recombinant host typically expresses endogenous or recombinant genes encoding UDP-glucose dehydrogenase and UDP-glucuronic acid decarboxylase. Such genes are useful for providing increased amounts of UDP-xylose donor for xylosylation of the steviol compound acceptor. Suitable UDP-glucose dehydrogenases and UDP-glucuronic acid decarboxylases include those made by Arabidopsis thaliana or Cryptococcus neoformans. For example, suitable UDP-glucose dehydrogenase and UDP-glucuronic acid decarboxylase polypeptides can be encoded by the UGD1 and UXS3 genes of A. thaliana, respectively. See Oka and Jigami, FEBS J. 273:2645-2657 (2006).

[0081] Those skilled in the art will recognize that by modulating the relative expression levels of different UGT genes and the availability of UDP-xylose, a recombinant host can be tailored to specifically produce steviol glycoside products in desired proportions. Transcriptional regulation of steviol biosynthetic genes can be achieved by a combination of transcriptional activation and repression, using techniques well known to those skilled in the art. In in vitro reactions, those skilled in the art will recognize that adding various levels of combinations of UGT enzymes, or conditions that affect the relative activity of combinations of different UGT enzymes, will direct synthesis to desired proportions of each steviol glycoside. In some embodiments, the recombinant host further contains and expresses a recombinant GGPPS gene to increase levels of the diterpene precursor geranylgeranyl diphosphate and thereby increase flux through the rebaudioside A biosynthetic pathway. In some embodiments, the recombinant host further contains a construct that suppresses expression of non-steviol pathways that consume geranylgeranyl diphosphate, ent-kaurenoic acid, or farnesyl pyrophosphate, thereby increasing flux through the steviol and steviol glycoside biosynthetic pathways. For example, flux to sterol production pathways, such as ergosterol, can be reduced by downregulating the ERG9 gene. See the ERG9 section below and Examples 24-25. In cells that produce gibberellins, gibberellin synthesis can be downregulated to increase flux of ent-kaurenoic acid to steviol. In carotenoid-producing organisms, flux to steviol can be increased by downregulating one or more carotenoid biosynthetic genes. In some embodiments, the recombinant host further contains and expresses recombinant genes involved in diterpene biosynthesis, such as genes in the MEP pathway described below.

[0082] In some embodiments, a recombinant host such as a microorganism produces rebaudioside F-enriched steviol glycoside compositions having greater than at least 4% rebaudioside F by weight of total steviol glycosides, e.g., at least 5% rebaudioside F, at least 6% rebaudioside F, 10-20% rebaudioside F, 20-30% rebaudioside F, 30-40% rebaudioside F, 40-50% rebaudioside F, 50-60% rebaudioside F, 60-70% rebaudioside F, or 70-80% rebaudioside F. In some embodiments, a recombinant host such as a microorganism produces steviol glycoside compositions having at least 90% rebaudioside F, e.g., 90-99% rebaudioside F. Other steviol glycosides present may include those shown in Figures 2A and D, such as steviolmonosides, steviolglucobiosides, steviolxylobiosides, rebaudioside A, stevioxyloside, rubusoside, and stevioside. In some embodiments, the rebaudioside F-enriched composition produced by a host is mixed with other steviol glycosides, flavors, or sweeteners to achieve a desired flavor system or sweetening composition. For example, a rebaudioside F-enriched composition produced by a recombinant microorganism can be combined with a rebaudioside A, C, or D-enriched composition produced by a different recombinant microorganism, with rebaudioside A, C, or D purified from a Stevia extract, or with rebaudioside A, C, or D produced in vitro.

[0083] C. Other Polypeptides Genes for additional polypeptides whose expression facilitates more efficient or larger-scale production of steviol or steviol glycosides can also be introduced into a recombinant host. For example, a recombinant microorganism, plant, or plant cell can also contain one or more genes encoding geranylgeranyl diphosphate synthase (also called GGPPS or GGDPS). As another example, a recombinant host can contain one or more genes encoding rhamnose synthase, or one or more genes encoding UDP-glucose dehydrogenase and / or UDP-glucuronic acid decarboxylase. As another example, a recombinant host can also contain one or more genes encoding cytochrome P450 reductase (CPR). Expression of recombinant CPR facilitates the cycling of NADP+ to regenerate NADPH, which is utilized as a cofactor for terpenoid biosynthesis. Other methods can also be used to regenerate NADHP levels. In situations where NADPH becomes limiting, strains can be further modified to include an exogenous transhydrogenase gene. For example, see Sauer et al. J. Biol. Chem. 279: 6613-6619 (2004). Other methods for reducing or otherwise altering the ratio of NADH / NADPH so that desired cofactor levels are increased are well known to those of skill in the art. As another example, a recombinant host can contain one or more genes encoding one or more enzymes in the MEP pathway or mevalonate pathway. Such genes are useful because they can increase carbon flux to the diterpene biosynthetic pathway, producing geranylgeranyl diphosphate from isopentenyl diphosphate and dimethylallyl diphosphate through the pathway. The geranylgeranyl diphosphate thus produced can be directed toward steviol and steviol glycoside biosynthesis by expression of steviol biosynthesis polypeptides and steviol glycoside biosynthesis polypeptides.

[0084] As another example, a recombinant host can contain one or more genes encoding sucrose synthase and, optionally, a sucrose uptake gene. The sucrose synthase reaction can be used to increase the UDP-glucose pool in a fermentation host or in a whole-cell bioconversion process. This regenerates UDP-glucose from the UDP and sucrose produced during glycosylation, allowing for efficient glycosylation. In some organisms, disruption of endogenous invertase is advantageous to prevent sucrose degradation. For example, the S. cerevisiae SUC2 invertase can be disrupted. The sucrose synthase (SUS) can be derived from any suitable organism. For example, but not limited to, a sucrose synthase coding sequence from Arabidopsis thaliana, Stevia rebaudiana, or Coffea arabica can be cloned into an expression plasmid under the control of a suitable promoter and expressed in a host (e.g., a microorganism or a plant). Sucrose synthase can be expressed in such strains in combination with a sucrose transporter (e.g., the A. thaliana SUC1 transporter or a functional homolog thereof) and one or more UGTs (e.g., UGT85C2, UGT74G1, UGT76G1, and one or more of UGT91D2e, EUGT11, or a functional homolog thereof). Culturing the host in a medium containing sucrose can promote the production of not only UDP-glucose, but also one or more glucosides (e.g., steviol glycosides). Additionally, as described herein, the recombinant host may have reduced phosphatase activity.

[0085] C.1 MEP Biosynthetic Polypeptides In some embodiments, the recombinant host contains one or more genes encoding enzymes involved in the methylerythritol 4-phosphate (MEP) pathway for isoprenoid biosynthesis, including deoxyxylulose 5-phosphate synthase (DXS), D-1-deoxyxylulose 5-phosphate reductoisomerase (DXR), 4-diphosphocytidyl-2-C-methyl-D-erythritol synthase (CMS), 4-diphosphocytidyl-2-C-methyl-D-erythritol kinase (CMK), 4-diphosphocytidyl-2-C-methyl-D-erythritol 2,4-cyclodiphosphate synthase (MCS), 1-hydroxy-2-methyl-2(E)-butenyl 4-diphosphate synthase (HDS), and 1-hydroxy-2-methyl-2(E)-butenyl 4-diphosphate synthase (HDR). One or more of the DXS gene, DXR gene, CMS gene, CMK gene, MCS gene, HDS gene, and / or HDR gene can be incorporated into a recombinant microorganism. Plant Phys. 130: 1079-1089 (2002). Suitable genes encoding DXS, DXR, CMS, CMK, MCS, HDS and / or HDR polypeptides include those produced by E. coli, Arabidopsis thaliana, and Synechococcus leopoliensis. Nucleotide sequences encoding DXR polypeptides are described, for example, in U.S. Patent No. 7,335,815.

[0086] C.2 Mevalonate Biosynthetic Polypeptides In some embodiments, the recombinant host contains one or more genes encoding enzymes involved in the mevalonate pathway for isoprenoid biosynthesis. Genes suitable for transformation into the host encode enzymes in the mevalonate pathway, such as a truncated 3-hydroxy-3-methyl-glutaryl (HMG) CoA reductase (tHMG), and / or a gene encoding mevalonate kinase (MK), and / or a gene encoding phosphomevalonate kinase (PMK), and / or a gene encoding mevalonate pyrophosphate decarboxylase (MPPD). Thus, one or more HMG-CoA reductase genes, MK genes, PMK genes, and / or MPPD genes can be incorporated into a recombinant host, such as a microorganism. Suitable genes encoding mevalonate pathway polypeptides are known. For example, suitable polypeptides include those produced by E. coli, Paracoccus denitrificans, Saccharomyces cerevisiae, Arabidopsis thaliana, Kitasatospora griseola, Homo sapiens, Drosophila melanogaster, Gallus gallus, Streptomyces sp. KO-3988, Nicotiana attenuata, Kitasatospora griseola, Hevea brasiliensis, Enterococcus faecium, and Haematococcus pluvialis. See, e.g., Table 8 and U.S. Patent Nos. 7,183,089, 5,460,949, and 5,306,862. [Table 9]

[0087] C.3 Sucrose synthase polypeptides Sucrose synthase (SUS) can be used as a tool for producing UDP-sugars. SUS (EC 2.4.1.13) catalyzes the formation of UDP-glucose and fructose from sucrose and UDP (Figure 11). Therefore, UDP produced by the UGT reaction can be converted to UDP-glucose in the presence of sucrose. See, for example, Chen et al. (2001) J. Am. Chem. Soc. 123:8866-8867; Shao et al. (2003) Appl. Env. Microbiol. 69:5238-5242; Masada et al. (2007) FEBS Lett. 581:2562-2566; and Son et al. (2009) J. Microbiol. Biotechnol. 19:709-712. Sucrose synthase can be used to generate UDP-glucose and remove UDP, facilitating the efficient glycosylation of compounds in various systems. For example, yeast lacking the ability to utilize sucrose can be grown on sucrose by introducing a sucrose transporter and SUS. For example, Saccharomyces cerevisiae does not have an efficient sucrose uptake system and relies on extracellular SUC2 for sucrose utilization. The combination of disrupting the endogenous S. cerevisiae SUC2 invertase and expressing recombinant SUS resulted in a yeast strain capable of metabolizing intracellular but not extracellular sucrose (Riesmeier et al. ((1992) EMBO J. 11:4705-4713)). This strain was used to isolate sucrose transporters by transforming a cDNA expression library and selecting transformants that had gained the ability to uptake sucrose.

[0088] As described herein, combined expression of recombinant sucrose synthase and sucrose transporter in vivo can result in increased availability of UDP-glucose and removal of unwanted UDP. For example, functional expression of recombinant sucrose synthase, sucrose transporter, and glycosyltransferases, combined with knockout of the native sucrose degradation system (SUC2 in the case of S. cerevisiae), can be used to produce cells capable of producing increased amounts of glycosylated compounds, such as steviol glycosides. This higher glycosylation capacity is due, at least, to (a) a greater ability to produce UDP-glucose in a more energy-efficient manner and (b) removal of UDP from the growth medium (since UDP can inhibit glycosylation reactions). The sucrose synthase can be derived from any suitable organism. For example, but not limited to, a sucrose synthase coding sequence from Arabidopsis thaliana, Stevia rebaudiana, or Coffea arabica (see, e.g., Figures 19A-19C, SEQ ID NOs: 178, 179, and 180) can be cloned into an expression plasmid under the control of a suitable promoter and expressed in a host (e.g., a microorganism or a plant). As described in the Examples herein, a SUS coding sequence can be expressed in a SUC2 (sucrose hydrolase)-deficient S. cerevisiae strain to prevent the yeast from degrading extracellular sucrose. In such a strain, the sucrose synthase can be expressed in combination with a sucrose transporter (e.g., the A. thaliana SUC1 transporter or a functional homolog thereof) and one or more UGTs (e.g., UGT85C2, UGT74G1, UGT76G1, EUGT11, and UGT91D2e, or one or more functional homologs thereof). Culturing a host in a medium containing sucrose can promote the production of not only UDP-glucose but also one or more glucosides (e.g., steviol glucoside). Note that in some cases, sucrose synthases and sucrose transporters can be expressed along with UGTs in host cells that are also recombinant for the production of a particular compound (e.g., steviol).

[0089] C.4 Regulation of ERG9 activity The purpose of the present disclosure is to produce terpenoids based on the concept of increasing the accumulation of terpenoid precursors in the squalene pathway.Non-limiting examples of terpenoids include hemiterpenoids, 1 isoprene unit (5 carbons); monoterpenoids, 2 isoprene units (10C); sesquiterpenoids, 3 isoprene units (15C); diterpenoids, 4 isoprene units (20C) (e.g., ginkgolides); triterpenoids, 6 isoprene units (30C); tetraterpenoids, 8 isoprene units (40C) (e.g., carotenoids); and polyterpenoids with multiple isoprene units. Hemiterpenoids include isoprene, prenol, and isovaleric acid.Monoterpenoids include geranyl pyrophosphate, eucalyptol, limonene, and pinene.Sesquiterpenoids include farnesyl pyrophosphate, artemisinin, and bisabolol.Diterpenoids include geranylgeranyl pyrophosphate, steviol, retinol, retinal, phytol, taxol, forskolin, and aphidicolin.Triterpenoids include squalene and lanosterol.Tetraterpenoids include lycopene and carotene. Terpenes are hydrocarbons derived from the combination of several isoprene units. Terpenoids can be considered terpene derivatives. The term terpene is sometimes used broadly to include terpenoids. Similar to terpenes, terpenoids can also be classified according to the number of isoprene units used. The present invention focuses on terpenoids, and in particular, terpenoids derived from the precursors farnesyl pyrophosphate (FPP), isopentenyl pyrophosphate (IPP), dimethylallyl pyrophosphate (DMAPP), geranyl pyrophosphate (GPP), and / or geranylgeranyl pyrophosphate (GGPP) via the squalene pathway.

[0090] By terpenoids, the following are understood: terpenoids of the hemiterpenoid class, such as, but not limited to, isoprene, prenol, and isovaleric acid; terpenoids of the monoterpenoid class, such as, but not limited to, geranyl pyrophosphate, eucalyptol, limonene, and pinene; terpenoids of the sesquiterpenoid class, such as, but not limited to, farnesyl pyrophosphate, artemisinin, and bisabolol; terpenoids of the diterpenoid class, such as, but not limited to, geranylgeranyl pyrophosphate, steviol, retinol, retinal, phytol, taxol, forskolin, and aphidicolin; terpenoids of the triterpenoid class, such as, but not limited to, lanosterol; and terpenoids of the tetraterpenoid class, such as, but not limited to, lycopene and carotene. In one aspect, the present invention relates to the production of terpenoids biosynthesized from geranylgeranyl pyrophosphate (GGPP). In particular, such terpenoid may be steviol. In one aspect, the present invention relates to the production of terpenoids biosynthesized from geranylgeranyl pyrophosphate (GGPP). In particular, such terpenoid may be steviol.

[0091] cell The present invention relates to a cell, such as any of the hosts described in Section III, modified to contain the construct shown in Figure 22. Thus, in a main aspect, the present invention relates to a cell comprising a nucleic acid sequence, the nucleic acid sequence being: i) the promoter sequence, which is ii) operably linked to a heterologous insert sequence, which iii) operably linked to an open reading frame, which iv) operably linked to a transcription termination signal; wherein the heterologous insert sequence comprises the general formula (I): -X1-X2-X3-X4-X5- wherein X2 comprises at least four consecutive nucleotides that are complementary to and form a hairpin secondary structure element with at least four consecutive nucleotides of X4; and wherein X3 is optional and, if present, comprises nucleotides involved in the formation of a hairpin loop between X2 and X4; and wherein X1 and X5 independently and optionally comprise one or more nucleotides, and wherein the open reading frame, when expressed, encodes a squalene synthase (EC 2.5.1.21), e.g., a polypeptide sequence having at least 70% identity to a squalene synthase (EC 2.5.1.21) or a biologically active fragment thereof, wherein the fragment has at least 70% sequence identity with the squalene synthase within an overlap of at least 100 amino acids. The present invention relates to the cell, which has:

[0092] In addition to the above-described nucleic acids comprising heterologous insert sequences, the cells may also contain one or more additional heterologous nucleic acid sequences (e.g., nucleic acids encoding any of the steviol and steviol glycoside biosynthetic polypeptides of Section I. In one preferred embodiment, the cells contain a heterologous nucleic acid encoding GGPPS operably linked to a nucleic acid sequence that directs expression of GGPPS in the cell. Heterologous Insert Sequence The heterologous insert sequence can adapt the secondary structure elements of the hairpin to the hairpin loop. The hairpin portion comprises complementary and hybridizing portions X2 and X4. The X2 and X4 portions are adjacent to portion X3, which contains the nucleotides that form the loop, i.e., the hairpin loop. The term "complementary" is understood by those skilled in the art to mean two sequences compared to each other, counting nucleotide by nucleotide from the 5' end to the 3' end or in reverse. The heterologous insert sequence is long enough to complete a hairpin but short enough to allow limited translation of an ORF located in frame immediately after the 3' end of the heterologous insert sequence. Thus, in one embodiment, the heterologous insert sequence comprises 10 to 50 nucleotides, preferably 10 to 30 nucleotides, more preferably 15 to 25 nucleotides, even more preferably 17 to 22 nucleotides, more preferably 18 to 21 nucleotides, more preferably 18 to 20 nucleotides, and more preferably 19 nucleotides.

[0093] X2 and X4 can independently consist of any suitable number of nucleotides, provided that a contiguous sequence of at least four nucleotides of X2 is complementary to a contiguous sequence of at least four nucleotides of X4. In a preferred embodiment, X2 and X4 consist of the same number of nucleotides. X2 may, for example, consist of nucleotides in the range of 4 to 25 (such as the range of 4 to 20), for example in the range of 4 to 15 (such as the range of 6 to 12), for example in the range of 8 to 12 (such as the range of 9 to 11). X4 may consist of, for example, 4 to 25 nucleotides (such as 4 to 20 nucleotides), for example, 4 to 15 nucleotides (such as 6 to 12 nucleotides), for example, 8 to 12 nucleotides (such as 9 to 11 nucleotides). In a preferred embodiment, X2 consists of a nucleotide sequence complementary to the nucleotide sequence of X4, ie, all nucleotides of X2 are preferably complementary to the nucleotide sequence of X4. In a preferred embodiment, X4 consists of a nucleotide sequence complementary to the nucleotide sequence of X2, i.e., all nucleotides of X4 are preferably complementary to the nucleotide sequence of X2. Highly preferably, X2 and X4 consist of the same number of nucleotides, wherein X2 is complementary to X4 over the entire length of X2 and X4.

[0094] X3 may be absent, i.e., X3 may consist of zero nucleotides. X3 may also consist of 1 to 5 nucleotides, such as 1 to 3 nucleotides. X1 may be absent, i.e., X1 may consist of zero nucleotides. X1 may also consist of nucleotides in the range of 1 to 25 (such as in the range of 1 to 20), for example in the range of 1 to 15 (such as in the range of 1 to 10), for example in the range of 1 to 5 (such as in the range of 1 to 3). X5 may be absent, i.e., X5 may consist of zero nucleotides. X5 may also consist of 1 to 5 nucleotides, for example 1 to 3 nucleotides. The sequence may be any suitable sequence that meets the requirements defined herein above.Thus, the heterologous insert sequence may comprise a sequence selected from the group consisting of SEQ ID NO:181, SEQ ID NO:182, SEQ ID NO:183, and SEQ ID NO:184.In a preferred embodiment, the insert sequence is selected from the group consisting of SEQ ID NO:181, SEQ ID NO:182, SEQ ID NO:183, and SEQ ID NO:184.

[0095] Squalene synthase Squalene synthase (SQS) is the first committed enzyme in the biosynthetic pathway leading to the production of sterols. It catalyzes the synthesis of squalene from farnesyl pyrophosphate via the intermediate presqualene pyrophosphate. This enzyme is a key branch point enzyme in terpenoid / isoprenoid biosynthesis and is thought to regulate the flux of isoprene intermediates through the sterol pathway. The enzyme is sometimes referred to as farnesyl diphosphate farnesyltransferase (FDFT1). The mechanism of SQS is the conversion of two units of farnesyl pyrophosphate into squalene. SQS is believed to be a eukaryotic or higher organism enzyme, although at least one prokaryotic organism has been shown to have a functionally similar enzyme. From a structural and mechanical standpoint, squalene synthase most closely resembles phytoene synthase, which in many plants plays a key role in the elaboration of phytoene, the precursor to many carotenoid compounds.

[0096] A high level of sequence identity indicates the likelihood that the first sequence is derived from the second sequence. Amino acid sequence identity requires identical amino acid sequences between two aligned sequences. Thus, a candidate sequence sharing 70% amino acid identity with a reference sequence requires that, following alignment, 70% of the amino acids in the candidate sequence are identical to the corresponding amino acids in the reference sequence. Identity can be determined with the aid of computer analysis, for example, but not limited to, the ClustalW computer alignment program, as described in Section D. Using this program with default settings, the mature (biologically active) portions of the query and reference polypeptides are aligned. The number of completely conserved residues is counted and divided by the length of the reference sequence polypeptide. The ClustalW algorithm may also be used to align nucleotide sequences. Sequence identity can be calculated in a manner similar to that shown for amino acid sequences. In one important embodiment the cell of the invention, upon expression as defined herein, has a squalene synthase activity of at least 75%, such as at least 76%, for example at least 77%, such as at least 78%, for example at least 79%, such as at least 80%, for example at least 81%, such as at least 82%, for example at least 83%, such as at least 84%, for example at least 85%, such as at least 86%, for example at least 87%, such as at least 88%, for example at least 89%, such as at least 90%, for example at least 91%, for example at least 92%, for example at least 93%, such as at least 94%, for example at least 95%, such as at least 96%, for example at least 97%, such as at least 98%, for example at least 99%, such as at least 100%, for example at least 101%, such as at least 102%, for example at least 103%, such as at least 104%, for example at least 105%, such as at least 106%, for example at least 107%, such as at least 108%, for example at least 109%, such as at least 1109%, such as at least 111%, The present invention also includes a nucleic acid sequence encoding a squalene synthase which is at least 85%, such as at least 86%, for example at least 87%, such as at least 88%, for example at least 89%, such as at least 90%, for example at least 91%, such as at least 92%, for example at least 93%, such as at least 94%, for example at least 95%, such as at least 96%, for example at least 97%, such as at least 98%, for example at least 99%, such as at least 99.5%, for example at least 99.6%, such as at least 99.7%, for example at least 99.8%, such as at least 99.9%, such as 100% identical.

[0097] promoter A promoter is a region of DNA that promotes transcription of a specific gene. Promoters are located near the genes they regulate, on the same strand, typically upstream (toward the 5' region of the sense strand). For transcription to occur, an RNA-synthesizing enzyme known as RNA polymerase must attach to the DNA near the gene. Promoters contain specific DNA sequences and response elements that provide secure initial binding sites for RNA polymerase and proteins called transcription factors that recruit RNA polymerase. These transcription factors have specific activator or repressor sequences of corresponding nucleotides that attach to specific promoters and regulate gene expression. In bacteria, promoters are recognized by RNA polymerase and associated sigma factors, which are then often delivered to the promoter DNA by an activator protein that binds to its own nearby DNA-binding site. In eukaryotes, the process is more complex, with at least seven different factors required for RNA polymerase II to bind to the promoter. Promoters represent a critical element that can work in concert with other regulatory regions (enhancers, silencers, boundary elements / insulators) to direct the level of transcription of a given gene. A promoter is usually immediately adjacent to the open reading frame (ORF) of interest, and its position in the promoter is specified relative to the transcription start site, where transcription of RNA begins for a particular gene (i.e., upstream positions are negative numbers measured from -1, e.g., -100 is a position 100 base pairs upstream).

[0098] Promoter element Core promoter – the minimal part of the promoter required to properly initiate transcription O Transcription start site (TSS) Approximately -35 bp upstream and / or downstream of the start site o RNA polymerase binding site RNA polymerase I: transcribes the gene encoding ribosomal RNA RNA polymerase II: transcribes genes encoding messenger RNA and certain small nuclear RNAs RNA polymerase III: transcribes genes that encode tRNAs and other small RNAs O General transcription factor binding sites Proximal promoter – the proximal sequence upstream of a gene that tends to contain the primary regulatory elements O Approximately -250 bp upstream of the start site Specific transcription factor binding sites Distal promoter - the distal sequence upstream of a gene that may contain additional regulatory elements that often have a weaker effect than the proximal promoter Everything further upstream (except enhancers or other regulatory regions whose effects are position / orientation independent) Specific transcription factor binding sites

[0099] Prokaryotic promoters In prokaryotes, a promoter consists of two short sequences upstream of the transcription start site at positions -10 and -35. Sigma factors not only help enhance the binding of RNAP to the promoter, but also help target RNAP to specific genes for transcription. The sequence at -10 is called the Pribnow box or -10 element and is usually composed of the six nucleotides TATAAT. The Pribnow box is essential for initiating transcription in prokaryotes. Another sequence at -35 (the -35 element) usually consists of the seven nucleotides TTGACAT, and its presence allows for very high transcription rates.

[0100] Both of the above consensus sequences, while conserved on average, are not found intact in most promoters. On average, only three of the six base pairs in each consensus sequence are found in any given promoter. To date, no promoters have been identified with intact consensus sequences at both the -10 and -35 positions; artificial promoters with complete conservation of the -10 / -35 hexamers have been found to promote RNA chain initiation with very high efficiency. Some promoters contain a UP element (consensus sequence 5'-AAAWWTWTTTTNNNAAANNN-3'; W = A or T; N = any base) centered at -50; the presence of the -35 element does not appear to be important for transcription from UP element-containing promoters.

[0101] Eukaryotic promoters Eukaryotic promoters are usually located upstream of the gene (ORF) and can contain regulatory elements several kilobases (kb) from the transcription start site. In eukaryotes, the transcription complex can induce DNA to fold back on itself, allowing regulatory sequences to be located far from the actual transcription site. Many eukaryotic promoters contain a TATA box (sequence TATAAA), which in turn binds the TATA-binding protein, which aids in the formation of the RNA polymerase transcription complex. The TATA box is usually located very close to the transcription start site (often within 50 bases). The cells of the present invention comprise a nucleic acid sequence comprising a promoter sequence. The promoter sequence is not limiting for the present invention and can be any promoter suitable for the host cell of choice. In one embodiment of the invention, the promoter is a constitutive or an inducible promoter. In further embodiments of the invention, the promoter is selected from the group consisting of endogenous promoter, PGK-1, GPD1, PGK1, ADH1, ADH2, PYK1, TPI1, PDC1, TEF1, TEF2, FBA1, GAL1-10, CUP1, MET2, MET14, MET25, CYC1, GAL1-S, GAL1-L, TEF1, ADH1, CAG, CMV, human UbiC, RSV, EF-1α, SV40, Mt1, Tet-On, Tet-Off, Mo-MLV-LTR, Mx1, progesterone, RU486, and rapamycin-inducible promoter.

[0102] Post-transcriptional regulation Post-transcriptional regulation is the control of gene expression at the RNA level, and therefore between the transcription and translation of the gene. The first example of regulation is in transcription (transcriptional regulation), where genes are differentially transcribed depending on chromatin structure and the activity of transcription factors. Once produced, the stability and distribution of different transcripts are regulated (post-transcriptional regulation) by RNA-binding proteins (RBPs), which control various steps and rates of transcript evolution, such as alternative splicing, nuclear degradation (exosome), processing, nuclear export (three alternative pathways), sequestration in DCP2 bodies for storage or degradation, and finally translation. These proteins achieve these events by binding to specific sequences or secondary structures of the transcript, usually RNA recognition motifs (RRMs) in the 5' and 3' UTRs of the transcript. Regulating capping, splicing, addition of poly(A) tails, sequence-specific rates of nuclear export, and in some contexts sequestration of RNA transcripts occurs in eukaryotes but not in prokaryotes. This regulation is the result of proteins or transcripts that can then be regulated to have affinity for specific sequences.

[0103] Capping Capping changes the 5'-end of the mRNA to a 3'-end with a 5'-5' linkage, which protects the mRNA from 5' exonucleases that degrade foreign RNA. The cap also aids in ribosome binding. Splicing Splicing removes introns, non-coding regions transcribed into RNA, so that mRNA can make proteins. Cells do this by attaching spliceosomes to either side of the intron, looping the intron into a joining circle, and then cleaving it. The two ends of the exon are then joined together.

[0104] Polyadenylation Polyadenylation is the addition of a poly(A) tail to the 3' end of a construct, i.e., the poly(A) tail is composed of multiple adenosine monophosphates. The poly(A) sequence acts as a buffer against 3' exonucleases, thus extending the half-life of mRNA. Furthermore, a long poly(A) tail can increase translation. In this way, the poly(A) tail can be used to further regulate the translation of the construct of the present invention to achieve an optimal translation rate. In eukaryotes, polyadenylation is part of the process that generates mature messenger RNA (mRNA) for translation. The poly(A) tail is also important for mRNA nuclear export, translation, and stability. In one embodiment, the nucleic acid sequence of the cell of the invention as defined herein above further comprises a polyadenylation / polyadenylation sequence, preferably the 5' end of said polyadenylation / polyadenylation sequence being operably linked to the 3' end of an open reading frame, such as an open reading frame encoding squalene synthase.

[0105] RNA editing RNA editing is a process that results in sequence changes in RNA molecules and is catalyzed by enzymes. These enzymes include adenosine deaminase acting on RNA (ADAR) enzymes, which convert specific adenosine residues in mRNA molecules to inosine by hydrolytic deamination. Three ADAR enzymes, ADAR1, ADAR2, and ADAR3, have been cloned, but only the first two subtypes have been shown to have RNA editing activity. Many mRNAs are susceptible to RNA editing, including the glutamate receptor subunits GluR2, GluR3, GluR4, GluR5, and GluR6 (components of AMPA and kainate receptors), the serotonin 2C receptor, the GABA-alpha3 receptor subunit, the tryptophan hydroxylase enzyme TPH2, hepatitis delta virus, and over 16% of microRNAs. In addition to ADAR enzymes, CDAR enzymes exist, which convert cytosine to uracil in specific RNA molecules. These enzymes, called "APOBECs," have a genetic locus at 22q13, a region close to the chromosomal deletion that occurs in palatocardiofacial syndrome (22q11) and that has been linked to psychiatric disorders. RNA editing has been widely investigated in relation to infectious diseases because the editing process alters viral function.

[0106] Post-transcriptional regulatory elements The use of a posttranscriptional regulatory element (PRE) is often necessary to obtain a vector with sufficient performance for a specific application. Schambach et al. in Gene Ther. (2006) 13(7):641-5 reported that the introduction of a posttranscriptional regulatory element (PRE) from woodchuck hepatitis virus (WHV) into the 3' untranslated region of retroviral and lentiviral gene transfer vectors enhances both titer and transgene expression. The enhancing activity of a PRE depends on the exact composition of its sequence and the context of the vector and the cells into which it is introduced. Thus, the use of a PRE such as the woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) may be useful in preparing the cells of the invention when using gene therapy approaches. Thus, in one embodiment, the nucleic acid sequence of the cell defined herein further comprises a post-transcriptional regulatory element. In a further embodiment, the post-transcriptional regulatory element is a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE).

[0107] Terminal repeat To insert gene sequences into host DNA, viruses often use DNA sequences that are repeated thousands of times, so-called repeats, or terminal repeats, including long terminal repeats (LTRs) and inverted terminal repeats (ITRs), where the repeat sequences can be both 5' and 3' terminal repeats. ITRs assist in the formation of concatemers in the nucleus after the single-stranded vector DNA is converted into double-stranded DNA by the host cell DNA polymerase complex. ITR sequences can be derived from viral vectors, such as AAVs, such as AAV2. In one embodiment, the nucleic acid sequence of the cell or vector defined herein comprises a 5' terminal repeat and a 3' terminal repeat. In one embodiment, said 5' and 3' terminal repeats are selected from inverted terminal repeats (ITRs) and long terminal repeats (LTRs). In one embodiment, the 5' and 3' terminal repeats are AAV inverted terminal repeats (ITRs).

[0108] Geranylgeranyl pyrophosphate synthase In a preferred embodiment, the microbial cells of the present invention may contain a heterologous nucleic acid sequence encoding geranylgeranyl pyrophosphate synthase (GGPPS). See, for example, Table 7. GGPPS is an enzyme that catalyzes the chemical reaction that converts one molecule of farnesyl pyrophosphate (FPP) into one molecule of geranylgeranyl pyrophosphate (GGPP). Genes encoding GGPPS can be found, for example, in organisms that contain the mevalonate pathway. The GGPPS used in the present invention may be any useful enzyme capable of catalyzing the conversion of a farnesyl pyrophosphate (FPP) molecule to a geranylgeranyl pyrophosphate (GGPP) molecule. In particular, the GGPPS used in the present invention may be any enzyme capable of catalyzing the following reaction: (2E,6E)-Farnesyl diphosphate + isopentenyl diphosphate → diphosphate + geranylgeranyl diphosphate. Preferably, the GGPPS for use with the present invention are enzymes classified under EC 2.5.1.29. The GGPPS may be GGPPS from various sources, such as bacteria, fungi, or mammals. The GGPPS may be any type of GGPPS, such as GGPPS-1, GGPPS-2, GGPPS-3, or GGPPS-4. The GGPPS may be wild-type GGPPS or a functional homolog thereof.

[0109] For example, the GGPPS may be GGPPS-1 of S. acidicaldarius (SEQ ID NO: 126), GGPPS-2 of A. nidulans (SEQ ID NO: 203), GGPPS-3 of S. cerevisiae (SEQ ID NO: 167), GGPPS-4 of M. musculus (SEQ ID NO: 123), or a functional homologue of any of the foregoing. The heterologous nucleic acid encoding GGPPS may be any nucleic acid sequence encoding GGPPS. Thus, in embodiments of the invention in which GGPPS is a wild-type protein, the nucleic acid sequence may be, for example, a wild-type cDNA sequence encoding the protein. However, it will often be the case that the heterologous nucleic acid is a nucleic acid sequence encoding any particular GGPPS, where the nucleic acid has been codon-optimized for a particular microbial cell. For example, if the microbial cell is S. cerevisiae, the nucleic acid encoding GGPPS is preferably codon-optimized for optimal expression in S. cerevisiae.

[0110] A functional homologue of GGPPS is preferably a protein that has the above activity and shares at least 70% amino acid identity with the sequence of the reference GGPPS. Methods for determining sequence identity are described herein above in the section "Squalene synthase" and in Section D. In one embodiment the cell, such as a microorganism, of the invention comprises a nucleic acid sequence encoding a GGPPS or a functional homologue thereof, wherein said functional homologue is selected from the group consisting of SEQ ID NO: 123, SEQ ID NO: 126, SEQ ID NO: 167 and SEQ ID NO: 203 and a GGPPS selected from the group consisting of SEQ ID NO: 128, SEQ ID NO: 129, SEQ ID NO: 230, SEQ ID NO: 231, SEQ ID NO: 232, SEQ ID NO: 233, SEQ ID NO: 234, SEQ ID NO: 235, SEQ ID NO: 236, SEQ ID NO: 237, SEQ ID NO: 238, SEQ ID NO: 239, SEQ ID NO: 240, SEQ ID NO: 241, SEQ ID NO: 242, SEQ ID NO: 243, SEQ ID NO: 244, SEQ ID NO: 245, SEQ ID NO: 246, SEQ ID NO: 247, SEQ ID NO: 248, SEQ ID NO: 249, SEQ ID NO: 250, SEQ ID NO: 251, SEQ ID NO: 252, SEQ ID NO: 253, SEQ ID NO: 254, SEQ ID NO: 255, SEQ ID NO: 256, SEQ ID NO: 257, SEQ ID NO: 258, SEQ ID NO: 259, SEQ ID NO: 260, SEQ ID NO: 261, SEQ ID NO: 262, SEQ ID NO: 263, SEQ ID NO: 264, SEQ ID NO: 265, SEQ ID NO: 266, SEQ ID NO: 267, SEQ ID NO: 268, SEQ ID NO: For example, at least 86%, such as at least 87%, for example at least 88%, such as at least 89%, for example at least 90%, such as at least 91%, for example at least 92%, such as at least 93%, for example at least 94%, such as at least 95%, for example at least 96%, such as at least 97%, for example at least 98%, such as at least 99%, for example at least 99.5%, such as at least 99.6%, for example at least 99.7%, such as at least 99.8%, for example at least 99.9%, such as 100% identical. The heterologous nucleic acid sequence encoding GGPPS is generally operably linked to a nucleic acid sequence that directs the expression of GGPPS in the microbial cell. The nucleic acid sequence that directs the expression of GGPPS in the microbial cell may be a promoter sequence, and preferably, the promoter sequence is selected depending on the specific microbial cell. The promoter may, for example, be any of the promoters described in the section "Promoters" herein above.

[0111] vector A vector is a DNA molecule used as a vehicle to transfer foreign genetic material into another cell. The main types of vectors are plasmids, viruses, cosmids, and artificial chromosomes. Common to all engineered vectors are an origin of replication, a multiple cloning site, and a selectable marker. The vector itself is generally a DNA sequence composed of an insert (transgene) and a larger sequence that serves as the "backbone" of the vector. The purpose of a vector to transfer genetic information to another cell is typically to isolate, multiply, or express the insert in the target cell. Vectors called expression vectors (expression constructs) are specifically intended for the expression of a transgene in target cells and generally have a promoter sequence that drives transgene expression. Simple vectors called transcription vectors can only be transcribed, not translated: they can be replicated in target cells, but unlike expression vectors, they are not translated. Transcription vectors are used to amplify their inserts. The insertion of a vector into a target cell is usually called transformation for bacterial cells and transfection for eukaryotic cells, although the insertion of a viral vector is often called transduction.

[0112] Plasmid A plasmid vector is a double-stranded, generally circular, DNA sequence capable of autonomous replication within a host cell. A plasmid vector minimally consists of an origin of replication, allowing semi-independent replication of the plasmid within the host, and a transgene insert. Modern plasmids generally have multiple functions, notably a "multiple cloning site" containing nucleotide overhangs for insertion of the insert, and multiple restriction enzyme consensus sites on either side of the insert. In the case of a plasmid used as a transcription vector, incubation of bacteria with the plasmid generates hundreds or thousands of copies of the vector within the bacteria within a few hours. The vector can then be extracted from the bacteria, and the multiple cloning site can be cleaved with a restriction enzyme to excise the hundredfold or thousandfold amplified insert. These plasmid transcription vectors characteristically lack critical sequences encoding polyadenylation and translation termination sequences in the translated mRNA, making protein expression from the transcription vector impossible. Plasmids can be conjugative / transmissible and non-conjugative: - Conjugative: mediates the transfer of DNA via conjugation and therefore spreads rapidly within a population of bacterial cells; e.g., F plasmid, many R and some col plasmids. - Non-conjugative: do not transmit DNA via conjugation, e.g., many R and col plasmids.

[0113] viral vectors Viral vectors are generally genetically engineered viruses that carry modified viral DNA or RNA that has been made non-infectious, but contain a viral promoter and a transgene, thereby enabling the viral promoter to translate the transgene.However, because viral vectors often lack infectious sequences, they require a helper virus or packaging line for large-scale transfection.Viral vectors are often designed for permanent integration of inserts into the host genome, so after integration of the transgene, they leave a unique genetic marker in the host genome.For example, retroviruses leave a characteristic retroviral integration pattern after insertion, which can be detected and indicates that the viral vector has been integrated into the host genome. In one aspect, the present invention also relates to viral vectors capable of transfecting host cells, such as culturable cells, e.g., yeast cells or any other suitable eukaryotic cells, which can then be transfected with a nucleic acid comprising a heterologous insert sequence as described herein. The viral vector may be any suitable viral vector, such as a viral vector selected from the group consisting of vectors derived from the retroviridae family, including lentivirus, HIV, SIV, FIV, EAIV, CIV.

[0114] The viral vector may also be selected from the group consisting of alphavirus, adenovirus, adeno-associated virus, baculovirus, HSV, coronavirus, bovine papilloma virus, Mo-MLV and adeno-associated virus. In aspects of the invention in which the microbial cell comprises a heterologous nucleic acid encoding GGPPS, the heterologous nucleic acid may be located on the vector containing the nucleic acid encoding squalene synthase, or the heterologous nucleic acid encoding GGPPS may be located on a different vector, which may be contained in any of the vectors described herein above. In embodiments of the present invention in which the microbial cell comprises a heterologous nucleic acid encoding HMCR, the heterologous nucleic acid may be located on a vector containing the nucleic acid encoding squalene synthase, or the heterologous nucleic acid encoding HMCR may be located on a different vector. The heterologous nucleic acid encoding HMCR may be contained in any of the vectors described herein above. It is within the scope of the present invention that the heterologous nucleic acid encoding GGPPS and the heterologous nucleic acid encoding HMCR may be located on the same or separate vectors.

[0115] Transcription Transcription is a necessary element for all vectors: a prerequisite for a vector is to amplify an insert (although expression vectors later also drive translation of the amplifying insert). Therefore, stable expression is determined by stable transcription, which generally depends on the promoter within the vector. However, expression vectors have different expression patterns: constitutive (consistent expression) or inducible (expression only under specific conditions or chemicals). This expression is based on different promoter activity, not post-transcriptional activity. Therefore, these two different types of expression vectors depend on different types of promoters. Viral promoters are often used in plasmids and viral vectors for constitutive expression because they usually ensure constant transcription in many cell lines and types. Inducible expression depends on the promoter responding to the inducing condition: for example, the mouse mammary tumor virus promoter initiates transcription only after application of dexamethasone, and the Drosophila heat shock promoter is initiated only after high temperature. Transcription is the synthesis of mRNA. Genetic information is copied from DNA to RNA.

[0116] Expression Expression vectors require sequences encoding, for example, polyadenylation tails (see herein above): they create polyadenylation tails at the ends of the transcribed pre-mRNA that protect the mRNA from exonucleases and ensure transcription and translation termination: they stabilize mRNA production. Minimum UTR length: Because UTRs contain certain features that may interfere with transcription or translation, optimal expression vectors encode the shortest or no UTRs. Kozak sequence: The vector must encode a Kozak sequence in the mRNA that assembles ribosomes for mRNA translation. The above conditions are necessary for expression vectors in eukaryotes but not in prokaryotes.

[0117] Modern vectors may include additional features beyond the transgene insert and backbone, such as promoters (discussed above), genetic markers that allow for confirmation, for example, that the vector has integrated into the host's genomic DNA, antibiotic resistance genes for antibiotic selection, and affinity tags for purification. In one embodiment, the cells of the invention comprise a nucleic acid sequence incorporated into a vector, such as an expression vector. In one embodiment, the vector is selected from the group consisting of a plasmid vector, a cosmid, an artificial chromosome, and a viral vector. The plasmid vector should be able to be maintained and replicated in bacteria, fungi and yeast. The present invention also relates to cells containing the plasmid and cosmid vectors, as well as artificial chromosome vectors. The important factors are that the vector is functional and that it contains at least a nucleic acid sequence that contains a heterologous insert sequence as described herein. In one embodiment, the vector is functional in fungal and mammalian cells. In one aspect, the present invention relates to a cell transformed or transduced with a vector as defined herein.

[0118] Terpenoid production method As described herein above, the cells (e.g., recombinant host cells) of the invention are useful for increasing the yield of industrially relevant terpenoids. The cells of the invention can therefore be used in a variety of setups to increase the accumulation of terpenoid precursors and therefore the yield of terpenoid products obtained from the enzymatic conversion of said (upstream) terpenoid precursors. Thus, in one aspect, the present invention relates to a method for producing terpenoid compounds synthesized via the squalene pathway in a cell culture, the method comprising the steps of: (a) providing a cell as defined herein above; (b) culturing the cells of step (a); (c) recovering the terpenoid product compounds. By providing cells containing the genetically engineered construct defined herein above, the accumulation of terpenoid precursors is enhanced (Figure 20). Thus, in another aspect, the present invention relates to a method for producing terpenoids from terpenoid precursors selected from the group consisting of farnesyl pyrophosphate (FPP), isopentenyl pyrophosphate (IPP), dimethylallyl pyrophosphate (DMAPP), geranyl pyrophosphate (GPP) and / or geranylgeranyl pyrophosphate (GGPP), the method comprising: (a) contacting the precursor with an enzyme in the squalene synthase pathway; (b) recovering the terpenoid product; Includes:

[0119] In one embodiment, the terpenoid (product) of the method defined herein above is selected from the group consisting of hemiterpenoids, monoterpenes, sesquiterpenoids, diterpenoids, sesterpenes, triterpenoids, tetraterpenoids, and polyterpenoids. In further embodiments, the terpenoid is selected from the group consisting of farnesyl phosphate, farnesol, geranylgeranyl, geranylgeraniol, isoprene, prenol, isovaleric acid, geranyl pyrophosphate, eucalyptol, limonene, pinene, farnesyl pyrophosphate, artemisinin, bisabolol, geranylgeranyl pyrophosphate, retinol, retinal, phytol, taxol, forskolin, aphidicolin, lanosterol, lycopene, and carotene. The terpenoid product can be used as a starting point in additional purification processes. Thus, in one embodiment, the method further comprises dephosphorylating the farnesyl phosphate to produce farnesol. The enzyme or enzymes used in the process for preparing a target product terpenoid compound are preferably "downstream" enzymes of terpenoid precursors such as farnesyl pyrophosphate, isopentenyl pyrophosphate, dimethylallyl pyrophosphate, geranyl pyrophosphate, and geranylgeranyl pyrophosphate, for example, enzymes downstream of the terpenoid precursors farnesyl pyrophosphate, isopentenyl pyrophosphate, dimethylallyl pyrophosphate, geranyl pyrophosphate, and geranylgeranyl pyrophosphate shown in the squalene synthesis pathway of Figure 20. The enzyme used in the process for preparing a target product terpenoid, which is based on the accumulation of a precursor according to the present invention, can therefore be selected from the group consisting of dimethylallyltransferase (EC 2.5.1.1), isoprene synthase (EC 4.2.3.27), and geranyltransferase (EC 2.5.1.10).

[0120] The present invention can operate by at least in part by sterically hindering ribosome binding to RNA, thereby reducing translation of squalene synthase. Thus, in one aspect, the present invention relates to a method for reducing the translation rate of functional squalene synthase (EC 2.5.1.21), said method comprising: (a) providing a cell as defined herein above; (b) Culturing the cells of (a). Similarly, in another aspect, the present invention relates to a method for reducing the conversion of farnesyl-pp to squalene, said method comprising: (a) providing a cell as defined herein above; (b) Culturing the cells of (a). As shown in Figure 20, knockdown of ERG9 results in an increase in the precursor of squalene synthase. Thus, in one aspect, the present invention relates to a method for enhancing the accumulation of a compound selected from the group consisting of farnesyl pyrophosphate, isopentenyl pyrophosphate, dimethylallyl pyrophosphate, geranyl pyrophosphate, and geranylgeranyl pyrophosphate, the method comprising the steps of: (a) providing a cell as defined herein above, and (b) Culturing the cells of (a).

[0121] In one aspect, the method of the present invention as defined herein above further comprises recovering the compounds farnesyl pyrophosphate, isopentenyl pyrophosphate, dimethylallyl pyrophosphate, geranyl pyrophosphate, and geranylgeranyl pyrophosphate. The recovered compounds may be used in further processes to produce desired terpenoid product compounds. The further processes may be carried out in the same cell culture as the process carried out and defined herein above, for example, the accumulation of terpenoid precursors by the cells of the present invention. Alternatively, the recovered precursors may be added to another cell culture or a cell-free system to produce the desired product. Although precursors are intermediates, primarily stable intermediates, specific endogenous production of terpenoid products may occur based on terpenoid precursor substrates. Cells of the invention may also have additional genetic modifications that allow for both the accumulation of terpenoid precursors (constructs of the cells of the invention) and the performance of all or substantially all of the subsequent biosynthetic processes to the desired terpenoid product.

[0122] Therefore, in one embodiment, the method of the present invention further comprises recovering compounds synthesized via the squalene pathway, which compounds are derived from the farnesyl pyrophosphate, isopentenyl pyrophosphate, dimethylallyl pyrophosphate, geranyl pyrophosphate, and / or geranylgeranyl pyrophosphate. Sometimes, it may be advantageous to include a squalene synthase inhibitor when culturing the cells of the present invention. Chemical inhibition of squalene synthase, for example, with lapaquistat, is known in the art and has been investigated as a method for lowering cholesterol levels, for example, in preventing cardiovascular disease. It has also been suggested that variants of this enzyme may be involved in genetic associations with hypercholesterolemia. Other squalene synthase inhibitors include zaragozic acid and RPR 107393. Thus, in one aspect, the culturing step of the method(s) defined herein above is carried out in the presence of a squalene synthase inhibitor. The cells of the present invention may further be genetically engineered to further enhance production of certain key terpenoid precursors. In one embodiment, the cells are further genetically engineered to enhance the activity and / or overexpress one or more enzymes selected from the group consisting of phosphomevalonate kinase (EC 2.7.4.2), diphosphomevalonate decarboxylase (EC 4.1.1.33), 4-hydroxy-3-methylbut-2-en-1-yl diphosphate synthase (EC 1.17.7.1), 4-hydroxy-3-methylbut-2-enyl diphosphate reductase (EC 1.17.1.2), isopentenyl diphosphate delta-isomerase 1 (EC 5.3.3.2), short-chain Z-isoprenyl diphosphate synthase (EC 2.5.1.68), dimethylallyltransferase (EC 2.5.1.1), geranyltransferase (EC 2.5.1.10), and geranylgeranyl pyrophosphate synthase (EC 2.5.1.29).

[0123] As described hereinabove, in one embodiment of the present invention, a microbial cell contains both a nucleic acid encoding a squalene synthase as described hereinabove and a heterologous nucleic acid encoding GGPPS. Such a microbial cell is particularly useful for the preparation of GGPP and terpenoids in the biosynthesis of which GGPP is an intermediate. Thus, in one aspect, the present invention relates to a method for preparing GGPP, said method comprising the steps of: a. providing a microbial cell comprising a nucleic acid sequence, wherein the nucleic acid i) the promoter sequence, which is ii) operably linked to a heterologous insert sequence, which iii) operably linked to an open reading frame, which iv) operably linked to a transcription termination signal; wherein said heterologous insert sequence and open reading frame are as defined herein above; wherein said microbial cell further comprises a heterologous nucleic acid encoding GGPPS operably linked to a nucleic acid sequence that directs expression of GGPPS in said cell; Cultivating microbial cells of ba; c. Recovering GGPP.

[0124] In another aspect, the present invention relates to a method for preparing a terpenoid of which GGPP is an intermediate in the biosynthetic pathway, the method comprising the steps of: a. providing a microbial cell, wherein the microbial cell comprises a nucleic acid sequence, the nucleic acid comprising: i) the promoter sequence, which is ii) operably linked to a heterologous insert sequence, which iii) operably linked to an open reading frame, which iv) operably linked to a transcription termination signal; wherein said heterologous insert sequence and open reading frame are as defined herein above; wherein said microbial cell further comprises a heterologous nucleic acid encoding GGPPS operably linked to a nucleic acid sequence that directs expression of GGPPS in said cell; Cultivating microbial cells of ba; and

[0125] c. Recovering terpenoids; wherein said terpenoid may be any terpenoid described above in the section "Terpenoids" that has GGPP as its biosynthetic intermediate; and said microbial cell may be any microbial cell described herein above in the section "Cells"; and said promoter may be any promoter, such as any promoter described herein above in the section "Promoters"; and said heterologous insert sequence may be any heterologous insert sequence described herein above in the section "Heterologous Insert Sequences"; and said open reading frame encodes a squalene synthase, which may be any squalene synthase described herein above in the section "Squalene Synthases"; and said GGPPS may be any GGPPS described herein above in the section "Geranylgeranyl Pyrophosphate Synthases". In this embodiment, the microbial cell may optionally contain one or more additional heterologous nucleic acids encoding one or more enzymes involved in the biosynthetic pathway of the terpenoid.

[0126] In one particular aspect, the present invention relates to a method for preparing steviol, wherein the method comprises the steps of: a. providing a microbial cell, wherein the microbial cell comprises a nucleic acid sequence, the nucleic acid comprising: i) the promoter sequence, which is ii) operably linked to a heterologous insert sequence, which iii) operably linked to an open reading frame, which iv) operably linked to a transcription termination signal; wherein said heterologous insert sequence and open reading frame are as defined herein above; wherein said microbial cell further comprises a heterologous nucleic acid encoding GGPPS operably linked to a nucleic acid sequence that directs expression of GGPPS in said cell; Cultivating microbial cells of ba;

[0127] c. Recovering steviol; wherein said microbial cell may be any microbial cell described herein above in the section "Cells"; and said promoter may be any promoter, such as any promoter described herein above in the section "Promoters"; and said heterologous insert sequence may be any heterologous insert sequence described herein above in the section "Heterologous Insert Sequences"; and said open reading frame encodes a squalene synthase, which may be any squalene synthase described herein above in the section "Squalene Synthases"; and said GGPPS may be any GGPPS described herein above in the section "Geranylgeranyl Pyrophosphate Synthases". In this embodiment, the microbial cell may optionally contain one or more additional heterologous nucleic acids encoding one or more enzymes involved in the biosynthetic pathway of steviol.

[0128] In another aspect, the present invention relates to a method for preparing a terpenoid of which GGPP is an intermediate in the biosynthetic pathway, the method comprising the steps of: a. providing a microbial cell, wherein the microbial cell comprises a nucleic acid sequence, the nucleic acid comprising: i) the promoter sequence, which is ii) operably linked to a heterologous insert sequence, which iii) operably linked to an open reading frame, which iv) operably linked to a transcription termination signal; wherein said heterologous insert sequence and open reading frame are as defined herein above; wherein the microbial cell further comprises a heterologous nucleic acid encoding GGPPS operably linked to a nucleic acid sequence that directs expression of GGPPS in the cell; and wherein said microbial cell further comprises a heterologous nucleic acid encoding HMCR operably linked to a nucleic acid sequence that directs expression of HMCR in said cell; Cultivating microbial cells of ba;

[0129] c. Recovering terpenoids; wherein said terpenoid may be any terpenoid described herein above in the section "Terpenoids" that has GGPP as an intermediate in its biosynthesis; wherein said microbial cell may be any microbial cell described herein above in the section "Cells"; and said promoter may be any promoter, such as any promoter described herein above in the section "Promoters"; and said heterologous insert sequence may be any heterologous insert sequence described herein above in the section "Heterologous Insert Sequences"; and said open reading frame encodes a squalene synthase, which may be any squalene synthase described herein above in the section "Squalene Synthases"; and said GGPPS may be any GGPPS described herein above in the section "Geranylgeranyl Pyrophosphate Synthases"; and said HMCR may be any HMCR described herein above in the section "HMCRs". In this embodiment, the microbial cell may optionally contain one or more additional heterologous nucleic acids encoding one or more enzymes involved in the biosynthetic pathway of the terpenoid.

[0130] In certain aspects, the present invention relates to a method for preparing steviol, wherein the method comprises the steps of: a. providing a microbial cell, wherein the microbial cell comprises a nucleic acid sequence, the nucleic acid comprising: i) the promoter sequence, which is ii) operably linked to a heterologous insert sequence, which iii) operably linked to an open reading frame, which iv) operably linked to a transcription termination signal; wherein said heterologous insert sequence and open reading frame are as defined herein above; wherein said microbial cell further comprises a heterologous nucleic acid encoding GGPPS operably linked to a nucleic acid sequence that directs expression of GGPPS in said cell; Cultivating microbial cells of ba;

[0131] c. Recovering steviol; wherein said microbial cell may be any microbial cell described herein above in the section "Cells"; and said promoter may be any promoter, such as any promoter described herein above in the section "Promoters"; and said heterologous insert sequence may be any heterologous insert sequence described herein above in the section "Heterologous Insert Sequences"; and said open reading frame encodes a squalene synthase, which may be any squalene synthase described herein above in the section "Squalene Synthase"; and said GGPPS may be any GGPPS described herein above in the section "Geranylgeranyl Pyrophosphate Synthase", and said HMCR may be any HMCR described herein above in the section "HMCR". In this embodiment, the microbial cell may optionally contain one or more additional heterologous nucleic acids encoding one or more enzymes involved in the biosynthetic pathway of steviol.

[0132] In one embodiment, the cells are further genetically engineered to enhance the activity of and / or overexpress one or more enzymes selected from the group consisting of acetoacetyl-CoA thiolase, HMG-CoA reductase or its catalytic domain, HMG-CoA synthase, mevalonate kinase, phosphomevalonate kinase, phosphomevalonate decarboxylase, isopentenyl pyrophosphate isomerase, farnesyl pyrophosphate synthase, D-1-deoxyxylulose 5-phosphate synthase, and 1-deoxy-D-xylulose 5-phosphate reductoisomerase and farnesyl pyrophosphate synthase. In one embodiment of the methods of the invention, the cell comprises a mutation in the ERG9 open reading frame. In another embodiment of the methods of the present invention, the cell comprises an ERG9[delta]::HIS3 deletion / insertion allele. In yet another embodiment, the step of recovering the compound of the method of the present invention further comprises purifying the compound from the cell culture medium.

[0133] D. Functional homolog Functional homologs of the above polypeptides are also suitable for use in producing steviol or steviol glycosides in a recombinant host. A functional homolog is a polypeptide that has sequence similarity to a reference polypeptide and performs one or more of the biochemical or physiological functions of the reference polypeptide. The functional homolog and the reference polypeptide may be naturally occurring polypeptides, and the sequence similarity may be due to convergent or divergent evolutionary events. As such, functional homologs are sometimes referred to in the literature as homologs, or orthologs, or paralogs, etc. A variant of a naturally occurring functional homolog, e.g., a polypeptide encoded by a variant of a wild-type coding sequence, may itself be a functional homolog. Functional homologs can also be generated through site-directed mutagenesis of the coding sequence of the polypeptide or by combining domains from the coding sequences of different naturally occurring polypeptides ("domain swapping"). Techniques for modifying genes encoding the functional UGT polypeptides described herein are well known and include, among others, directed evolution, site-directed mutagenesis, and random mutagenesis techniques, and are useful for increasing the specific activity of a polypeptide, altering substrate specificity, changing expression levels, changing subcellular localization, or modifying polypeptide:polypeptide interactions in a desired manner. Such modified polypeptides are considered functional homologs. The term "functional homolog" is sometimes applied to nucleic acids encoding functionally homologous polypeptides.

[0134] Functional homologs can be identified by analyzing alignments of nucleotide and polypeptide sequences. For example, a database of nucleotide or polypeptide sequences can be queried to identify homologs of steviol or steviol glycoside biosynthesis polypeptides. Sequence analysis includes BLAST, reciprocal BLAST, or PSI-BLAST analysis of non-redundant databases using the GGPPS, CDPS, KS, KO, or KAH amino acid sequence as a reference sequence. In some instances, the amino acid sequence is deduced from the nucleotide sequence. Polypeptides in the database with 40% or more sequence identity are candidates for further evaluation for suitability as steviol or steviol glycoside biosynthesis peptides. Amino acid sequence similarity allows for conservative amino acid substitutions, such as the substitution of one hydrophobic residue for another or one polar residue for another. If desired, manual inspection of such candidates can be performed to narrow the number of candidates for further evaluation. Manual inspection can be performed by selecting candidates that appear to have domains present in steviol biosynthesis polypeptides, such as conserved functional domains.

[0135] Conserved regions can be identified by locating regions within the primary amino acid sequence of a steviol or steviol glycoside biosynthetic polypeptide that are repetitive, form certain secondary structures (e.g., helices and beta-sheets), establish positively or negatively charged domains, or represent protein motifs or domains. See, for example, the Pfam website, which describes consensus sequences for various protein motifs or domains, on the World Wide Web at sanger.ac.uk / Software / Pfam / and pfam.janelia.org / . The information contained in the Pfam database is described in Sonnhammer et al., Nucl. Acids Res. , 26:320-322 (1998);Sonnhammer et al., Proteins, 28:405-420 (1997); and Bateman et al., Nucl. Acids Res. , 27:260-262 (1999). Conserved regions can also be determined by aligning sequences of identical or related polypeptides from closely related species. Closely related species are preferably from the same family. In some embodiments, alignment of sequences from two different species is appropriate. Typically, polypeptides that exhibit at least about 40% amino acid sequence identity are useful for identifying conserved regions. Conserved regions of related polypeptides exhibit at least 45% amino acid sequence identity (e.g., at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% amino acid sequence identity). In some embodiments, conserved regions exhibit at least 92%, 94%, 96%, 98%, or 99% amino acid sequence identity.

[0136] For example, polypeptides suitable for producing steviol glycosides in a recombinant host include functional homologs of EUGT11, UGT91D2e, UGT91D2m, UGT85C, and UGT76G. Such homologs have greater than 90% (e.g., at least 95% or 99%) sequence identity to the amino acid sequence of EUGT11 (SEQ ID NO:152), UGT91D2e (SEQ ID NO:5), UGT91D2m (SEQ ID NO:10), UGT85C (SEQ ID NO:3), or UGT76G (SEQ ID NO:7). Variants of EUGT11, UGT91D, UGT85C, and UGT76G polypeptides typically have no more than 10 amino acid substitutions in the primary amino acid sequence, such as no more than 7 amino acid substitutions, 5 or conservative amino acid substitutions, or 1 to 5 amino acid substitutions. However, in some embodiments, variants of EUGT11, UGT91D, UGT85C, and UGT76G polypeptides can have 10 or more amino acid substitutions (e.g., 10, 15, 20, 25, 30, 35, 10-20, 10-35, 20-30, or 25-35 amino acid substitutions). Substitutions can be conservative or, in some embodiments, non-conservative. Non-limiting examples of non-conservative changes within a UGT91D2e polypeptide include glycine to arginine and tryptophan to arginine. Non-limiting examples of non-conservative substitutions within a UGT76G polypeptide include valine to glutamic acid, glycine to glutamic acid, glutamine to alanine, and serine to proline. Non-limiting examples of changes to a UGT85C polypeptide include a histidine to aspartic acid, a proline to serine, a lysine to threonine, and a threonine to arginine change.

[0137] In some embodiments, useful UGT91D2 homologs can have amino acid substitutions (e.g., conservative amino acid substitutions) in regions of the polypeptide outside of predicted loops, e.g., residues 20-26, 39-43, 88-95, 121-124, 142-158, 185-198, and 203-214 are predicted loops in the N-terminal domain of SEQ ID NO: 5, and residues 381-386 are predicted loops in the C-terminal domain of SEQ ID NO: 5. For example, useful UGT91D2 homologs can include at least one amino acid substitution at residues 1-19, 27-38, 44-87, 96-120, 125-141, 159-184, 199-202, 215-380, or 387-473 of SEQ ID NO: 5. In some embodiments, a UGT91D2 homolog can have an amino acid substitution at one or more residues selected from the group consisting of residues 30, 93, 99, 122, 140, 142, 148, 153, 156, 195, 196, 199, 206, 207, 211, 221, 286, 343, 427, and 438 of SEQ ID NO: 5. For example, a UGT91D2 functional homolog can have an amino acid substitution at one or more of residues 206, 207, and 343, such as an arginine at residue 206, a cysteine ​​at residue 207, and an arginine at residue 343 of SEQ ID NO: 5. See SEQ ID NO: 95.Other functional homologs of UGT91D2 can have one or more of the following: a tyrosine or phenylalanine at residue 30, a proline or glutamine at residue 93, a serine or valine at residue 99, a tyrosine or phenylalanine at residue 122, a histidine or tyrosine at residue 140, a serine or cysteine ​​at residue 142, an alanine or threonine at residue 148, a methionine at residue 152, an alanine at residue 153, an alanine at residue 156, or of SEQ ID NO:5. serine, glycine at residue 162, leucine or methionine at residue 195, glutamic acid at residue 196, lysine or glutamic acid at residue 199, leucine or methionine at residue 211, leucine at residue 213, serine or phenylalanine at residue 221, valine or isoleucine at residue 253, valine or alanine at residue 286, lysine or asparagine at residue 427, alanine at residue 438, and either alanine or threonine at residue 462. In another embodiment, the UGT91D2 functional homolog comprises a methionine at residue 211 and an alanine at residue 286.

[0138] In some embodiments, useful UGT85C homologs can have one or more amino acid substitutions at residues 9, 10, 13, 15, 21, 27, 60, 65, 71, 87, 91, 220, 243, 270, 289, 298, 334, 336, 350, 368, 389, 394, 397, 418, 420, 440, 441, 444, and 471 of SEQ ID NO:3. Non-limiting examples of useful UGT85C homologs include polypeptides having (relative to SEQ ID NO: 3) a substitution at residue 65 (e.g., serine at residue 65), a substitution at residue 65 in combination with residue 15 (leucine at residue 15), 270 (e.g., methionine, arginine, or alanine at residue 270), 418 (e.g., valine at residue 418), 440 (e.g., aspartic acid at residue 440), or 441 (e.g., asparagine at residue 441); residues 13 (e.g., phenylalanine at residue 13), 15, 60 (e.g., aspartic acid at residue 60), 270, 289 (e.g., histidine at residue 289), and 418; residues 13, 60, and and 270; substitutions at residues 60 and 87 (e.g., phenylalanine at residue 87); substitutions at residues 65, 71 (e.g., glutamine at residue 71), 220 (e.g., threonine at residue 220), 243 (e.g., tryptophan at residue 243), and 270; substitutions at residues 65, 71, 220, 243, 270, and 442; substitutions at residues 65, 71, 220, 389 (e.g., valine at residue 389), and 394 (e.g., valine at residue 394); substitutions at residues 65, 71, 270, 289; substitutions at residues 220, 243, 270, and 334 (e.g., serine at residue 334); or substitutions at residues 270 and 289. The following amino acid mutations did not result in loss of activity in the 85C2 polypeptide: V13F, F15L, H60D, A65S, E71Q, I87F, K220T, R243W, T270M, T270R, Q289H, L334S, A389V, I394V, P397S, E418V, G440D, and H441N.Additional mutations found in active clones include K9E, K10R, Q21H, M27V, L91P, Y298C, K350T, H368R, G420R, L431P, R444G, and M471T. In some embodiments, UGT85C2 contains substitutions at positions 65 (e.g., serine), 71 (glutamine), 270 (methionine), 289 (histidine), and 389 (valine).

[0139] Stevia rebaudiana UGTs 74G1, 76G1, and 91D2e, which have an N-terminal in-frame fusion of the first 158 ​​amino acids of the human MDM2 protein, and Stevia rebaudiana UGT85C2, which has an N-terminal in-frame fusion of four repeats of the synthetic PMI peptide (4×TSFAEYWNLLSP, SEQ ID NO:86), are shown in SEQ ID NOs:90, 88, 94, and 92, respectively; for the nucleotide sequences encoding the fusion proteins, see SEQ ID NOs:89, 92, 93, and 95. In some embodiments, useful UGT76G homologs can have one or more amino acid substitutions at residues 29, 74, 87, 91, 116, 123, 125, 126, 130, 145, 192, 193, 194, 196, 198, 199, 200, 203, 204, 205, 206, 207, 208, 266, 273, 274, 284, 285, 291, 330, 331, and 346 of SEQ ID NO:7. Non-limiting examples of useful UGT76G homologs include substitutions (with respect to SEQ ID NO: 7) at residues 74, 87, 91, 116, 123, 125, 126, 130, 145, 192, 193, 194, 196, 198, 199, 200, 203, 204, 205, 206, 207, 208, and 291; or at residues 74, 87, 91, 116, 123, 125, 126, 130, 145, 192, 193, 194, 196, 198, 199, 200, 203, 204, 205, 206, 207, 208, 266, 273, 274, 284, 285, 291, 330, 331, and 346. See Table 9.

[0140] [Table 10] Methods for modifying the substrate specificity of, for example, EUGT11 or UGT91D2e are known to those skilled in the art and include, but are not limited to, site-directed / rational mutagenesis approaches, random directed evolution approaches, and combinations of random mutagenesis / saturation techniques performed near the active site of the enzyme. See, e.g., Sarah A. Osmani, et al. Phytochemistry 70 (2009) 325-347.

[0141] Candidate sequences typically have a length of 80% to 200% of the length of the reference sequence, for example, 82, 85, 87, 89, 90, 93, 95, 97, 99, 100, 105, 110, 115, 120, 130, 140, 150, 160, 170, 180, 190, or 200% of the length of the reference sequence. Functional homolog polypeptides typically have a length of 95% to 105% of the length of the reference sequence, for example, 90, 93, 95, 97, 99, 100, 105, 110, 115, or 120% of the length of the reference sequence, or any range of lengths therebetween. The percent identity of any candidate nucleic acid or polypeptide to a reference nucleic acid or polypeptide can be determined as follows. A reference sequence (e.g., a nucleic acid sequence or an amino acid sequence) is aligned to one or more candidate sequences using the computer program ClustalW (version 1.83, default parameters), which allows alignment of nucleic acid or polypeptide sequences over their entire length (global alignment). Chenna et al., Nucleic Acids Res., 31(13):3497-500 (2003).

[0142] ClustalW calculates the best match between a reference sequence and one or more candidate sequences and aligns them so that identities, similarities, and differences can be determined. Gaps of one or more residues can be inserted into the reference sequence, the candidate sequences, or both to maximize sequence alignment. For fast pairwise alignment of nucleic acid sequences, the following default parameters are used: word size: 2; window size: 4; scoring method: percentage; number of top diagonals: 4; and gap penalty: 5. For multiple alignment of nucleic acid sequences, the following parameters are used: gap opening penalty: 10.0; gap extension penalty: 5.0; and weight transition: yes. For fast pairwise alignment of protein sequences, the following parameters are used: word size: 1; window size: 5; scoring method: percentage; number of top diagonals: 5; and gap penalty: 3. For multiple alignment of protein sequences, the following parameters are used: weight matrix: blosum; gap opening penalty: 10.0; gap extension penalty: 0.05; hydrophilic gaps: on; hydrophilic residues: Gly, Pro, Ser, Asn, Asp, Gln, Glu, Arg, and Lys; residue-specific gap penalties: on. The output of ClustalW is a sequence alignment that reflects the relationships between sequences. ClustalW can be run, for example, at the Baylor College of Medicine Search Launcher site on the World Wide Web (searchlauncher.bcm.tmc.edu / multi-align / multi-align.html) and at the European Bioinformatics Institute site on the World Wide Web (ebi.ac.uk / clustalw).

[0143] To determine the percent identity of a candidate nucleic acid or amino acid sequence to a reference sequence, the sequences are aligned using ClustalW, the number of identical matches in the alignment is divided by the length of the reference sequence, and the result is multiplied by 100. Note that percent identity values ​​can be rounded to the nearest tenth. For example, 78.11, 78.12, 78.13, and 78.14 are rounded down to 78.1, and 78.15, 78.16, 78.17, 78.18, and 78.19 are rounded up to 78.2. It is understood that a functional UGT may contain additional amino acids that are not involved in glycosylation or other enzymatic activities performed by the enzyme, and therefore such polypeptides may be longer than they would otherwise be. For example, a EUGT11 polypeptide can include a purification tag (e.g., a HIS tag or a GST tag), a chloroplast transit peptide, a mitochondrial transit peptide, an amyloplast peptide, a signal peptide, or a secretion tag added to the amino- or carboxy-terminus. In some embodiments, a EUGT11 polypeptide includes an amino acid sequence that functions as a reporter, such as green fluorescent protein or yellow fluorescent protein.

[0144] II. Steviol and Steviol Glycoside Biosynthesis Nucleic Acids A recombinant gene encoding a polypeptide described herein comprises a coding sequence for that polypeptide operably linked in a sense orientation to one or more regulatory regions suitable for expressing the polypeptide. Because many microorganisms are capable of expressing multiple gene products from polycistronic mRNAs, expression of multiple polypeptides, if desired, can be under the control of a single regulatory region for those microorganisms. A coding sequence and a regulatory region are considered operably linked when they are positioned such that the regulatory region is effective to regulate the transcription or translation of the sequence. Typically, the translation initiation site of the translational reading frame of the coding sequence is located between 1 and about 50 nucleotides downstream of the regulatory region for a monocistronic gene.

[0145] In many cases, the coding sequence of a polypeptide described herein is identified in a species other than the recombinant host, i.e., is a heterologous nucleic acid. Thus, when the recombinant host is a microorganism, the coding sequence can be derived from another prokaryotic or eukaryotic microorganism, a plant, or an animal. In some cases, however, the coding sequence is a sequence native to the host and reintroduced into that organism. Native sequences can often be distinguished from naturally occurring sequences by the presence of non-native sequences linked to the exogenous nucleic acid, for example, by non-native regulatory sequences flanking the native sequence in the recombinant nucleic acid construct. In addition, stably transformed exogenous nucleic acids are typically integrated at a location other than the location where the native sequence is found.

[0146] A "regulatory region" refers to a nucleic acid having nucleotide sequences that influence the initiation and rate of transcription or translation, and the stability and / or mobility of the transcription or translation product. Regulatory regions include, but are not limited to, promoter sequences, enhancer sequences, response elements, protein recognition sites, inducible elements, protein binding sequences, 5' and 3' untranslated regions (UTRs), transcription initiation sites, termination sequences, polyadenylation sequences, introns, and combinations thereof. Regulatory regions typically include at least a core (basal) promoter. Regulatory regions may also include at least one control element, such as an enhancer sequence, upstream element, or upstream activation region (UAR). A regulatory region is operably linked to a coding sequence by positioning the regulatory region and the coding sequence such that the regulatory region is effective to regulate the transcription or translation of the sequence. For example, to operably link a coding sequence to a promoter sequence, the translation initiation site of the translational reading frame of the coding sequence is typically positioned between 1 and approximately 50 nucleotides downstream of the promoter. However, the regulatory region may be located up to about 5,000 nucleotides upstream of the translation start site or up to about 2,000 nucleotides upstream of the transcription start site.

[0147] The selection of the regulatory region to be included depends on several factors, including but not limited to efficiency, selectivity, inducibility, desired expression level, and preferential expression at a particular culture stage. It is a routine matter for those skilled in the art to regulate the expression of a coding sequence by appropriately selecting and positioning a regulatory region relative to the coding sequence. It is understood that more than one regulatory region may be present, such as introns, enhancers, upstream activation regions, transcription terminators, and inducible elements. One or more genes can be combined in a recombinant nucleic acid construct as a "module" useful for discrete aspects of steviol and / or steviol glycoside production. Combining multiple genes within a module, particularly a polycistronic module, facilitates the use of the module in various species. For example, a steviol biosynthesis gene cluster, or a UGT gene cluster, can be combined within a polycistronic module, allowing the module to be introduced into various species after inserting appropriate regulatory regions. As another example, a UGT gene cluster can be combined such that each UGT coding sequence is operably linked to a separate regulatory region to form a UGT module. Such a module can be used in species where monocistronic expression is necessary or desired. In addition to genes useful for steviol or steviol glycoside production, recombinant constructs typically also contain an origin of replication and one or more selectable markers for maintenance of the construct in the appropriate species.

[0148] Due to the degeneracy of the genetic code, it is understood that numerous nucleic acids can encode a particular polypeptide; i.e., for many amino acids, there is more than one nucleotide triplet that functions as a codon for that amino acid. Therefore, the codons in the coding sequence for a given polypeptide can be altered to obtain optimal expression in a particular host (e.g., a microorganism) using an appropriate codon bias table for that host. SEQ ID NOS: 18-25, 34-36, 40-43, 48-49, 52-55, 60-64, 70-72, and 154 represent nucleotide sequences encoding specific enzymes for the biosynthesis of steviol and steviol glycosides, modified for increased expression in yeast. Like isolated nucleic acids, these modified sequences can exist as purified molecules or can be incorporated into vectors or viruses used to construct modules of recombinant nucleic acid constructs.

[0149] In some cases, inhibiting one or more functions of endogenous polypeptides is desirable to divert metabolic intermediates toward steviol or steviol glycoside biosynthesis. For example, it may be desirable to downregulate sterol synthesis in a yeast strain, e.g., downregulating squalene epoxidase, to further increase steviol or steviol glycoside production. As another example, it may be desirable to inhibit the degradative function of a particular endogenous gene product, e.g., a glycohydrolase that removes glucose moieties from secondary metabolites or a phosphatase, as described herein. As another example, expression of a membrane transporter involved in the transport of steviol glycosides can be suppressed, such that secretion of glycosylated stevioside is inhibited. Such modulation can be beneficial to suppress steviol glycoside secretion for a desired time during microbial culture, thereby increasing the yield of the glycoside product(s) at harvest. In such cases, a nucleic acid that inhibits expression of a polypeptide or gene product may be included in a recombinant construct transformed into the strain. Alternatively, mutagenesis can be used to generate mutations in the gene whose function it is desired to disrupt.

[0150] III.Host A. Microorganisms Many prokaryotic and eukaryotic organisms are suitable for use in constructing the recombinant microorganisms described herein, including gram-negative bacteria, yeast, and fungi. The species and strain selected for use as a steviol or steviol glycoside production strain are first analyzed to determine which production genes are endogenous to the strain and which genes are absent. Genes for which no endogenous counterpart is present in the strain are assembled into one or more recombinant constructs, which are then transformed into the strain to provide the missing function(s). Exemplary prokaryotic and eukaryotic species are described in more detail below. However, it is understood that other species may be suitable. For example, a suitable species may be a genus selected from the following group: Agaricus, Aspergillus, Bacillus, Candida, Corynebacterium, Escherichia, Fusarium / Gibberella, Kluyveromyces, Laetiporus, Lentinus, Phaffia, Phanerochaete, Pichia, Physcomitrella, Rhodoturula, Saccharomyces, Schizosaccharomyces, Sphaceloma, Xanthophyllomyces, and Yarrowia. Representative species of such genera include Lentinus tigrinus, Laetiporus sulphureus, Phanerochaete chrysosporium, Pichia pastoris, Physcomitrella patens, Rhodoturula glutinis 32, Rhodoturula mucilaginosa, Phaffia rhodozyma UBV-AX, Xanthophyllomyces dendrorhous, Fusarium fujikuro / Gibberella fujikuroi, Candida utilis, and Yarrowia lipolytica. In some embodiments, the microorganism can be an ascomycete such as Gibberella fujikuroi, Kluyveromyces lactis, Schizosaccharomyces pombe, Aspergillus niger, or Saccharomyces cerevisiae. In some embodiments, the microorganism can be a prokaryote, such as Escherichia coli, Rhodobacter sphaeroides, or Rhodobacter capsulatus. It is understood that certain microorganisms can be used to screen and test genes of interest in a high-throughput manner, while other microorganisms with desirable productivity or growth characteristics can be used for large-scale production of steviol glycosides.

[0151] Saccharomyces cerevisiae Saccharomyces cerevisiae is a widely used chassis organism in synthetic biology and can be used as a recombinant microbial platform. Libraries of mutants, plasmids, detailed computer models of metabolism, and other information are available for S. cerevisiae, allowing for the rational design of various modules to enhance product yield. Methods for generating recombinant microorganisms are known. The steviol biosynthetic gene cluster can be expressed in yeast using any of a number of known promoters. Strains that overproduce terpenes are known and can be used to increase the amount of geranylgeranyl diphosphate available for the production of steviol and steviol glycosides.

[0152] Aspergillus species Aspergillus species, such as A. oryzae, A. niger, and A. sojae, are widely used in food production and can also be used as recombinant microbial platforms. Nucleotide sequences are available for the genomes of A. nidulans, A. fumigatus, A. oryzae, A. clavatus, A. flavus, A. niger, and A. terreus, enabling the rational design and modification of endogenous pathways for enhanced flux and increased product yield. Metabolic models have been developed for Aspergillus, as well as transcriptomic and proteomic studies. A. niger is cultivated for the industrial production of many food ingredients, such as citric acid and gluconic acid; therefore, species such as A. niger are generally suitable for the production of food ingredients such as steviol and steviol glycosides. Escherichia coli Escherichia coli is another platform organism widely used in synthetic biology and can also be used as a recombinant microbial platform. Similar to Saccharomyces, there are libraries of mutants, plasmids, detailed computer models of metabolism, and other information available for E. coli, allowing for the rational design of various modules to enhance product yield. Similar methods to those described above for Saccharomyces can be used to generate recombinant E. coli microorganisms.

[0153] Agaricus, Gibberella, and Phanerochaete species Agaricus, Gibberella, and Phanerochaete species may be useful because they are known to produce large amounts of gibberellins in culture. Therefore, the terpene precursors for producing large amounts of steviol and steviol glycosides are already produced by endogenous genes. Therefore, modules containing recombinant genes for steviol or steviol glycoside biosynthetic polypeptides can be introduced into species of these genera without the need to introduce mevalonate or MEP pathway genes. Arxula adeninivorans (Blastobotrys adeninivorans) Arxula adeninivorans is a dimorphic yeast with unusual biochemical properties (it grows as a budding yeast, like baker's yeast, up to temperatures of 42°C, and grows in a filamentous form above this threshold). It can grow on a wide range of substrates and assimilate nitrate. This has been successfully applied to generate strains capable of producing natural plastics and to develop biosensors for estrogens in environmental samples.

[0154] Yarrowia lipolytica Yarrowia lipolytica is a dimorphic yeast (see Arxula adeninivorans) that can grow on a wide range of substrates. It has great potential for industrial applications, but no commercially available recombinant products exist yet. Rhodobacter species Rhodobacter can be used as a recombinant microbial platform. Similar to E. coli, there are libraries of mutants available, as well as suitable plasmid vectors, allowing for the rational design of various modules to enhance product yield. The isoprenoid pathway has been engineered in membranous bacterial species of Rhodobacter to enhance carotenoid and CoQ10 production. See U.S. Patent Publication Nos. 20050003474 and 20040078846. Recombinant Rhodobacter microorganisms can be generated using methods similar to those described above for E. coli. Candida boidinii Candida boidinii is a methylotrophic yeast (capable of growing on methanol). Like other methylotrophic species, such as Hansenula polymorpha and Pichia pastoris, it provides an excellent platform for the production of heterologous proteins. Yields of secreted foreign proteins in the multigram range have been reported. The computational method IPRO recently predicted a mutation that experimentally switched the cofactor specificity of the Candida boidinii xylose reductase from NADPH to NADH.

[0155] Hansenula polymorpha(Pichia angusta) Hansenula polymorpha is another methylotrophic yeast (see Candida boidinii). It is also capable of growth on a wide range of other substrates; it is heat tolerant and can assimilate nitrate (see also Kluyveromyces lactis). It has also been applied in the production of hepatitis B vaccines, insulin and interferon alpha-2a for the treatment of hepatitis C, as well as a range of technical enzymes. Kluyveromyces lactis Kluyveromyces lactis is a yeast regularly applied in the production of kefir. It is able to grow on several sugars, most importantly the lactose present in milk and whey. It has been successfully applied for cheese production, in particular for the production of chymosin, an enzyme normally present in calf stomach. Production is carried out in fermenters at a 40,000 L scale. Pichia pastoris Pichia pastoris is a methylotrophic yeast (see Candida boidinii and Hansenula polymorpha). It provides an efficient platform for the production of foreign proteins. Platform elements are available as kits and are used worldwide in academia for protein production. Strains have been engineered to produce complex human glycans (yeast glycans are similar, but not identical, to those found in humans). Physcomitrella species When grown in suspension culture, Physcomitrella mosses have properties similar to those of yeast or other fungal cultures, making this genus an important cell type for the production of plant secondary metabolites that can be difficult to produce in other cell types.

[0156] B. Plant cells or plants In some embodiments, the nucleic acids and polypeptides described herein are introduced into plants or plant cells to increase overall steviol glycoside production or to enrich the production of a particular steviol glycoside relative to others. Thus, the host can be a plant or plant cell containing at least one recombinant gene described herein. A plant or plant cell can be transformed by incorporating the recombinant gene into its genome, i.e., stably transformed. Stably transformed cells typically retain the introduced nucleic acid with each cell division. A plant or plant cell can also be transiently transformed, such that the recombinant gene is not integrated into its genome. Transiently transformed cells typically lose all or a portion of the introduced nucleic acid with each cell division, and the introduced nucleic acid becomes undetectable in daughter cells after a sufficient number of cell divisions. Both transiently transformed and stably transformed transgenic plants and plant cells can be useful in the methods described herein.

[0157] The transgenic plant cells used in the methods described herein can constitute part or all of a whole plant. Such plants can be grown in a grow box, a greenhouse, or in the field in a manner appropriate to the species under consideration. Transgenic plants can be cultivated as needed for specific purposes, such as to introduce recombinant nucleic acids into other lines, to transfer recombinant nucleic acids to other species, or for further selection of other desirable traits. Alternatively, transgenic plants can be vegetatively propagated for species suitable for such techniques. As used herein, transgenic plants also refer to the progeny of the initial transgenic plant, provided that the progeny inherit the transgene. Seeds produced by transgenic plants can be grown and then self-pollinated (or cross-pollinated and then self-pollinated) to obtain homozygous seeds of the nucleic acid construct. Transgenic plants can be grown in suspension culture, or in tissue or organ culture. For purposes of the present invention, solid and / or liquid tissue culture techniques can be used. When using solid media, the transgenic plant cells can be placed directly on the medium or on a filter, which can then be placed in contact with the medium. When using liquid media, the transgenic plant cells can be placed in a suspension device, for example, on a porous membrane in contact with the liquid medium.

[0158] When transiently transformed plant cells are used, a reporter sequence encoding a reporter polypeptide having reporter activity can be included in the transformation procedure, and assays for reporter activity or expression can be performed at an appropriate time after transformation. A suitable time for performing the assay is generally about 1 to 21 days after transformation, e.g., about 1 to 14 days, about 1 to 7 days, or about 1 to 3 days. The use of transient assays is particularly convenient for rapid analysis in different species or for confirming the expression of a heterologous polypeptide whose expression has not previously been confirmed in a particular recipient cell. Techniques for introducing nucleic acids into monocotyledonous and dicotyledonous plants are well known in the art and include, but are not limited to, Agrobacterium-mediated transformation, viral vector-mediated transformation, electroporation and particle gun transformation, U.S. Patent Nos. 5,538,880, 5,204,253, 6,329,571, and 6,013,863. When cells or cultured tissues are used as recipient tissues for transformation, plants can be regenerated from the transformed cultures, if desired, by techniques well known to those skilled in the art.

[0159] A population of transgenic plants can be screened and / or selected for members of the population that possess a trait or phenotype conferred by expression of the transgene. For example, a population of progeny from a single transformation event can be screened for plants that have the desired level of expression of a steviol or steviol glycoside biosynthesis polypeptide or nucleic acid. Physical and biochemical methods can be used to identify expression levels. These include Southern analysis or PCR amplification for the detection of polynucleotides; Northern blots, S1 RNase protection, primer extension, or RT-PCR amplification for the detection of RNA transcripts; enzyme assays for detecting enzyme or ribozyme activity of polypeptides and polynucleotides; and protein gel electrophoresis, Western blots, immunoprecipitation, and enzyme-linked immunoassays for detecting polypeptides. Other techniques, such as in situ hybridization, enzyme staining, and immunostaining, can also be used to detect the presence or expression of polypeptides and / or nucleic acids. Methods for performing all of the referenced techniques are known. Alternatively, a population of plants containing independent transformation events can be screened for plants with a desired trait, such as steviol glycoside production or regulated steviol glycoside biosynthesis. Selection and / or screening can be carried out over one or more generations and / or in two or more geographic locations. In some cases, transgenic plants can be grown and selected under conditions that induce the desired phenotype or other conditions required to produce the desired phenotype in the transgenic plants. Selection and / or screening can also be applied during a particular developmental stage at which the phenotype is expected to be exhibited by the plant. Selection and / or screening can be carried out to select transgenic plants with a statistically significant difference in steviol glycoside levels compared to control plants lacking the transgene.

[0160] The nucleic acids, recombinant genes, and constructs described herein can be used to transform many monocotyledonous and dicotyledonous plants and plant cell lines. Non-limiting examples of suitable monocotyledonous plants include crop plants such as rice, rye, sorghum, millet, wheat, maize, and barley. Plants can also be non-cereal monocotyledonous plants such as asparagus, banana, or onion. Plants can also be dicotyledonous plants such as stevia (Stevia rebaudiana), soybean, cotton, sunflower, pea, geranium, spinach, or tobacco. In some cases, plants can contain precursor pathways for phenylphosphate production, such as the mevalonate pathway, which is typically found in the cytoplasm and mitochondria. Non-mevalonate pathways are more frequently found in plant plastids [Dubey, et al., 2003]. J. Biosci. 28 637-646]. One skilled in the art can target expression of steviol glycoside biosynthesis polypeptides to appropriate organelles by use of leader sequences so that steviol glycoside biosynthesis occurs at the desired site in the plant cell. Optionally, one skilled in the art can use an appropriate promoter to direct synthesis, for example, to the leaves of the plant. Expression can also occur in tissue culture, such as callus culture or hairy root culture, if desired.

[0161] In one embodiment, one or more nucleic acids or polypeptides described herein are introduced into Stevia (e.g., Stevia rebaudiana) to increase overall steviol glycoside biosynthesis or to selectively enrich the overall steviol glycoside composition for one or more particular steviol glycosides (e.g., rebaudioside D). For example, one or more recombinant genes can be introduced into Stevia to express a EUGT11 enzyme (e.g., SEQ ID NO: 152 or a functional homolog thereof), alone or in combination with one or more of the following: a UGT91D enzyme, such as UGT91D2e (e.g., SEQ ID NO: 5 or a functional homolog thereof), UGT91D2m (e.g., SEQ ID NO: 10); a UGT85C enzyme, such as a variant described in the "Functional Homologues" section; a UGT76G1 enzyme, such as a variant described in the "Functional Homologues" section; or a UGT74G1 enzyme. Nucleic acid constructs typically include a suitable promoter (e.g., a 35S, e35S, or ssRUBISCO promoter) operably linked to a nucleic acid encoding a UGT polypeptide. The nucleic acid can be introduced into Stevia by Agrobacterium-mediated transformation; electroporation-mediated gene transfer into protoplasts; or particle bombardment. See Singh, et al., Compendium of Transgenic Crop Plants: Transgenic Sugar, Tuber, and Fiber, Edited by Chittaranjan Kole and Timothy C. Hall, Blackwell Publishing Ltd. (2008), pp. 97-115. For particle bombardment of Stevia leaf-derived callus, the parameters are as follows: 6 cm distance, 1000 psi He pressure, gold particles, and one bombardment.

[0162] Stevia plants can be regenerated by somatic embryogenesis as described above in Singh et al., 2008. Specifically, leaf segments (approximately 1–2 cm long) can be removed from 5–6-week-old in vitro-grown plants and incubated (adaxial side down) on MS medium supplemented with vitamin B5, 30 g sucrose, and 3 g Gelrite. 2,4-Dichlorophenoxyacetic acid (2,4-D) can be used in combination with 6-benzyladenine (BA), kinetin (KN), or zeatin. Proembryogenic masses appear after 8 weeks of subculture. Somatic embryos appear on the surface of the culture within 2–3 weeks of subculture. Embryos can be matured in medium containing BA in combination with 2,4-D, α-naphthaleneacetic acid (NAA), or indolebutyric acid (IBA). Mature somatic embryos that germinate and form plantlets can be excised from the callus. After the plantlets reach 3-4 weeks, they are transferred to pots with vermiculite and grown in growth boxes for 6-8 weeks for acclimatization before being transferred to a greenhouse. In one embodiment, steviol glycosides are produced in rice. Rice and maize can be easily transformed using techniques such as Agrobacterium-mediated transformation. Binary vector systems are commonly used for Agrobacterium-mediated foreign gene transfer into monocotyledonous plants. See, for example, U.S. Patent Nos. 6,215,051 and 6,329,571. In a binary vector system, one vector contains a T-DNA region containing a gene of interest (e.g., a UGT described herein), and the other vector is a disarmed Ti plasmid containing a vir region. Cointegration and mobilization vectors can also be used. The type and pretreatment of tissue to be transformed, the Agrobacterium strain used, the duration of inoculation, and prevention of Agrobacterium-induced overgrowth and necrosis can be easily adjusted by those skilled in the art. Rice immature embryo cells can be prepared for Agrobacterium-mediated transformation using binary vectors. The culture medium used is supplemented with phenolic compounds. Alternatively, transformation can be performed in plants using vacuum infiltration. See, for example, WO 2000037663, WO 2000063400, and WO 2001012828.

[0163] IV. Methods for Producing Steviol Glycosides The recombinant hosts described herein can be used in methods for producing steviol or steviol glycosides. For example, when the recombinant host is a microorganism, the method can include growing the recombinant microorganism in a culture medium under conditions in which the steviol and / or steviol glycoside biosynthetic genes are expressed. The recombinant microorganism can be grown in a fed-batch or continuous process. Typically, the recombinant microorganism is grown in a fermentor at a defined temperature(s) for a desired period of time. Depending on the particular microorganism used in the method, other recombinant genes, such as isopentenyl biosynthetic genes and terpene synthase and cyclase genes, can also be present and expressed. Substrate and intermediate levels, such as isopentenyl diphosphate, dimethylallyl diphosphate, geranylgeranyl diphosphate, kaurene, and kaurenoic acid, can be determined by extracting samples from the culture medium for analysis according to published methods. After the recombinant microorganism has been grown in culture for a desired period of time, steviol and / or one or more steviol glycosides can then be recovered from the culture using various techniques known in the art. In some embodiments, a permeabilizing agent can be added to aid in the flow of feedstock into the host and product out. When the recombinant host is a plant or plant cell, steviol or steviol glycosides can be extracted from the plant tissue using various techniques known in the art. For example, a crude lysate of the cultured microorganism or plant tissue can be centrifuged to obtain a supernatant. The resulting supernatant can then be applied to a chromatography column, such as a C-18 column, washed with water to remove hydrophilic compounds, followed by elution of the compound(s) of interest with a solvent such as methanol. The compound(s) can then be further purified by preparative HPLC. See also WO 2009 / 140394.

[0164] The amount of steviol glycoside (e.g., rebaudioside D) produced can be about 1 mg / L to about 1500 mg / L, e.g., about 1 to about 10 mg / L, about 3 to about 10 mg / L, about 5 to about 20 mg / L, about 10 to about 50 mg / L, about 10 to about 100 mg / L, about 25 to about 500 mg / L, about 100 to about 1500 mg / L, or about 200 to about 1000 mg / L. Generally, longer culture times lead to larger amounts of product. Thus, the recombinant microorganism can be cultured for 1 to 7 days, 1 to 5 days, 3 to 5 days, about 3 days, about 4 days, or about 5 days. It is understood that the various genes and modules described herein can be present in two or more recombinant microorganisms rather than one microorganism. When multiple recombinant microorganisms are used, they can be grown in mixed cultures to produce steviol and / or steviol glycosides. For example, a first microorganism can contain one or more biosynthetic genes for steviol production, while a second microorganism contains steviol glycoside biosynthetic genes. Alternatively, two or more microorganisms can each be grown in separate culture media, and the product of the first culture medium, e.g., steviol, can be introduced into a second culture medium, which is converted to a subsequent intermediate or to an end product, such as rebaudioside A. The product produced by the second or final microorganism is then recovered. It will also be understood that in some embodiments, recombinant microorganisms are grown using nutrient sources other than culture media and systems other than fermentors.

[0165] Steviol glycosides do not necessarily perform equally well in different food systems. Therefore, it is desirable to have the ability to direct synthesis to a selected steviol glycoside composition. The recombinant hosts described herein can produce compositions selectively enriched for specific steviol glycosides (e.g., rebaudioside D) and with a consistent taste profile. Thus, the recombinant microorganisms, plants, and plant cells described herein can facilitate the production of compositions tailored to meet a desired sweetness profile for a given food product and with a consistent proportion of each steviol glycoside from batch to batch. The microorganisms described herein do not produce undesirable plant by-products found in stevia extracts. Thus, the steviol glycoside compositions produced by the recombinant microorganisms described herein are distinct from compositions derived from stevia plants.

[0166] V.Food The steviol glycosides obtained by the methods disclosed herein can be used to prepare food products, dietary supplements, and sweetener compositions. For example, substantially pure steviol or steviol glycosides, such as rebaudioside A or rebaudioside D, can be included in food products such as ice cream, carbonated beverages, fruit juice, yogurt, baked goods, chewing gum, hard and soft candies, and sauces. Substantially pure steviol or steviol glycosides can also be included in non-food products, such as pharmaceutical products, medicinal products, dietary supplements, and nutritional supplements. Substantially pure steviol or steviol glycosides can also be included in animal feed products for both the agricultural and pet industries. Alternatively, a mixture of steviol and / or steviol glycosides can be produced by culturing recombinant microorganisms separately or growing different plants / plant cells, each producing a specific steviol or steviol glycoside, recovering steviol or steviol glycosides in substantially pure form from each microorganism or plant / plant cell, and then combining the compounds to obtain a mixture containing each compound in the desired ratio. The recombinant microorganisms, plants, and plant cells described herein allow for a more precise and consistent mixture compared to current stevia products. In another alternative, substantially pure steviol or steviol glycosides can be incorporated into foods along with other sweeteners, such as saccharin, dextrose, sucralose, fructose, erythritol, aspartame, sucralose, monatin, or acesulfame potassium. The weight ratio of steviol or steviol glycosides to other sweeteners can be varied as needed to achieve a favorable taste in the final food product. See, for example, U.S. Patent Publication No. 2007 / 128311. In some embodiments, steviol or steviol glycosides may be provided with a flavoring (e.g., citrus) as a flavor modulator.For example, rebaudioside C can be used as a sweetness enhancer or sweetness modulator, especially for carbohydrate-based sweeteners, allowing for the reduction of the amount of sugar in foods.

[0167] Compositions produced by the recombinant microorganisms, plants, or plant cells described herein can be incorporated into foods. For example, steviol glycoside compositions produced by recombinant microorganisms, plants, or plant cells can be incorporated into foods in amounts ranging from about 20 mg steviol glycoside / kg food to about 1800 mg steviol glycoside / kg food on a dry weight basis, depending on the steviol glycoside and the type of food. For example, steviol glycoside compositions produced by recombinant microorganisms, plants, or plant cells can be incorporated into desserts, frozen treats (e.g., ice cream), dairy products (e.g., yogurt), or beverages (e.g., carbonated beverages) so that the food has up to 500 mg steviol glycoside / kg food on a dry weight basis. Steviol glycoside compositions produced by recombinant microorganisms, plants, or plant cells can be incorporated into baked goods (e.g., biscuits) so that the food has up to 300 mg steviol glycoside / kg food on a dry weight basis. Steviol glycoside compositions produced by recombinant microorganisms, plants, or plant cells can be incorporated into sauces (e.g., chocolate syrup) or vegetable products (e.g., pickles) so that the food product has up to 1000 mg steviol glycoside / kg food on a dry weight basis. Steviol glycoside compositions produced by recombinant microorganisms, plants, or plant cells can be incorporated into bread so that the food product has up to 160 mg steviol glycoside / kg food on a dry weight basis. Steviol glycoside compositions produced by recombinant microorganisms, plants, or plant cells can be incorporated into hard or soft candies so that the food product has up to 1600 mg steviol glycoside / kg food on a dry weight basis. Steviol glycoside compositions produced by recombinant microorganisms, plants, or plant cells can be incorporated into processed fruit products (e.g., fruit juices, fruit fillings, jams, and jellies) so that the food product has up to 1000 mg steviol glycoside / kg food on a dry weight basis.

[0168] For example, such a steviol glycoside composition can have 90-99% rebaudioside A and an undetectable amount of stevia plant-derived impurities and can be incorporated into a food product at 25-1600 mg / kg on a dry weight basis, e.g., 100-500 mg / kg, 25-100 mg / kg, 250-1000 mg / kg, 50-500 mg / kg, or 500-1000 mg / kg. Such steviol glycoside compositions can be rebaudioside B-enriched compositions having greater than 3% rebaudioside B and can be incorporated into foods such that the amount of rebaudioside B in the product is 25-1600 mg / kg, e.g., 100-500 mg / kg, 25-100 mg / kg, 250-1000 mg / kg, 50-500 mg / kg, or 500-1000 mg / kg, on a dry weight basis. Typically, rebaudioside B-enriched compositions have undetectable amounts of stevia plant-derived impurities. Such steviol glycoside compositions can be rebaudioside C-enriched compositions having greater than 15% rebaudioside C and can be incorporated into foods such that the amount of rebaudioside C in the product is 20-600 mg / kg, e.g., 100-600 mg / kg, 20-100 mg / kg, 20-95 mg / kg, 20-250 mg / kg, 50-75 mg / kg, or 50-95 mg / kg, on a dry weight basis. Typically, rebaudioside C-enriched compositions have an undetectable amount of stevia plant-derived impurities.

[0169] Such steviol glycoside compositions can be rebaudioside D-enriched compositions having greater than 3% rebaudioside D and can be incorporated into foods such that the amount of rebaudioside D in the product is 25-1600 mg / kg, e.g., 100-500 mg / kg, 25-100 mg / kg, 250-1000 mg / kg, 50-500 mg / kg, or 500-1000 mg / kg, on a dry weight basis. Typically, rebaudioside D-enriched compositions have undetectable amounts of stevia plant-derived impurities. Such steviol glycoside compositions can be rebaudioside E-enriched compositions having greater than 3% rebaudioside E and can be incorporated into foods such that the amount of rebaudioside E in the product is 25-1600 mg / kg, e.g., 100-500 mg / kg, 25-100 mg / kg, 250-1000 mg / kg, 50-500 mg / kg, or 500-1000 mg / kg, on a dry weight basis. Typically, rebaudioside E-enriched compositions have undetectable amounts of stevia plant-derived impurities. Such steviol glycoside compositions can be rebaudioside F-enriched compositions having greater than 4% rebaudioside F and can be incorporated into foods such that the amount of rebaudioside F in the product is 25-1000 mg / kg, e.g., 100-600 mg / kg, 25-100 mg / kg, 25-95 mg / kg, 50-75 mg / kg, or 50-95 mg / kg, on a dry weight basis. Typically, rebaudioside F-enriched compositions have undetectable amounts of stevia plant-derived impurities.

[0170] Such steviol glycoside compositions can be dulcoside A-enriched compositions having greater than 4% dulcoside A and can be incorporated into foods such that the amount of dulcoside A in the product is 25-1000 mg / kg, e.g., 100-600 mg / kg, 25-100 mg / kg, 25-95 mg / kg, 50-75 mg / kg, or 50-95 mg / kg, on a dry weight basis. Typically, dulcoside A-enriched compositions have undetectable amounts of stevia plant-derived impurities. Such steviol glycoside compositions can be enriched for rubusoside xylosylated at either of two positions, i.e., 13-O-glucose or 19-O-glucose. Such compositions can have greater than 4% xylosylated rubusoside compounds and can be incorporated into foods such that the amount of xylosylated rubusoside compounds in the product is 25-1000 mg / kg, e.g., 100-600 mg / kg, 25-100 mg / kg, 25-95 mg / kg, 50-75 mg / kg, or 50-95 mg / kg, on a dry weight basis. Typically, xylosylated rubusoside-enriched compositions have undetectable amounts of stevia plant-derived impurities.

[0171] Such steviol glycoside compositions can be enriched for compounds rhamnosylated at either of two positions, i.e., 13-O-glucose or 19-O-glucose, or for compounds containing one rhamnose and multiple glucoses (e.g., steviol 13-O-1,3-diglycoside-1,2-rhamnoside). Such compositions can have greater than 4% rhamnosylated compounds and can be incorporated into foods such that the amount of rhamnosylated compounds in the product is 25-1000 mg / kg, e.g., 100-600 mg / kg, 25-100 mg / kg, 25-95 mg / kg, 50-75 mg / kg, or 50-95 mg / kg, on a dry weight basis. Typically, compositions enriched for rhamnosylated compounds have undetectable amounts of stevia plant-derived impurities. In some embodiments, substantially pure steviol or steviol glycosides are incorporated into tabletop sweeteners or "cup-for-cup" products. Such products are typically diluted to an appropriate sweetness level with one or more bulking agents, such as maltodextrin, known to those of skill in the art. Steviol glycoside compositions enriched for rebaudioside A, rebaudioside C, rebaudioside D, rebaudioside E, rebaudioside F, dulcoside A, or rhamnosylated or xylosylated compounds can be packaged in sachets for tabletop use, e.g., at 10,000-30,000 mg steviol glycoside / kg product on a dry weight basis.

[0172] In some aspects, the present disclosure relates to the following: 1. A recombinant host cell comprising a nucleic acid sequence, the nucleic acid comprising a heterologous insert sequence operably linked to an open reading frame, wherein the heterologous insert sequence has the general formula (I): -X1-X2-X3-X4-X5- wherein X2 comprises at least four consecutive nucleotides complementary to at least four consecutive nucleotides of X4; wherein X3 contains zero nucleotides or one or more nucleotides that form a hairpin loop; wherein X1 and X5 each independently consist of zero nucleotides or one or more nucleotides, and wherein the open reading frame encodes squalene synthase (EC 2.5.1.21). The recombinant host cell comprising: 2. The recombinant cell of claim 1, wherein the nucleic acid comprises, in 5' to 3' order, a promoter sequence operably linked to a heterologous insert sequence operably linked to an open reading frame, wherein the heterologous insert sequence and open reading frame are as defined in claim 1.

[0173] 3. A cell comprising a nucleic acid sequence, the nucleic acid comprising: i) the promoter sequence, which is ii) operably linked to a heterologous insert sequence, which iii) operably linked to an open reading frame, which iv) operably linked to a transcription termination signal; wherein the heterologous insert sequence comprises the general formula (I): -X1-X2-X3-X4-X5- wherein X2 comprises at least four consecutive nucleic acids that are complementary to and form a hairpin secondary structure element with at least four consecutive nucleic acids of X4; and wherein X3 contains an unpaired nucleic acid, thereby forming a hairpin loop between X2 and X4; and wherein X1 and X5 independently and optionally comprise one or more nucleic acids; and wherein the open reading frame, when expressed, encodes a polypeptide sequence having at least about 70% identity to a squalene synthase (EC 2.5.1.21) or a biologically active fragment thereof, wherein the fragment has at least about 70% sequence identity with the squalene synthase within an overlap of at least 100 amino acids. The cell comprising:

[0174] 4. The cell according to any one of items 1 to 3, wherein the heterologous insert sequence comprises 10 to 50 nucleotides, preferably 10 to 30 nucleotides, more preferably 15 to 25 nucleotides, more preferably 17 to 22 nucleotides, more preferably 18 to 21 nucleotides, more preferably 18 to 20 nucleotides, more preferably 19 nucleotides. 5. The cell according to any one of items 1 to 4, wherein X2 and X4 consist of the same number of nucleotides. 6. The cell according to any one of items 1 to 5, wherein all X2s consist of 4 to 25 nucleotides, for example in the range of 4 to 20 nucleotides, for example in the range of 4 to 15 nucleotides, for example in the range of 6 to 12 nucleotides, for example in the range of 8 to 12 nucleotides, for example in the range of 9 to 11 nucleotides. 7. The cell according to any one of items 1 to 6, wherein all X4s consist of 4 to 25 nucleotides, for example in the range of 4 to 20 nucleotides, for example in the range of 4 to 15 nucleotides, for example in the range of 6 to 12 nucleotides, for example in the range of 8 to 12 nucleotides, for example in the range of 9 to 11 nucleotides. 8. The cell according to any one of items 1 to 7, wherein X2 consists of a nucleotide sequence complementary to the nucleotide sequence of X4. 9. The cell according to any one of items 1 to 8, wherein X4 consists of a nucleotide sequence complementary to the nucleotide sequence of X2.

[0175] 10. The cell according to any one of items 1 to 9, wherein X3 is absent, i.e. X3 consists of zero nucleotides. 11. The cell according to any one of items 1 to 9, wherein X3 consists of 1 to 5 nucleotides, for example 1 to 3 nucleotides. 12. The cell according to any one of items 1 to 11, wherein X1 is absent, i.e. X1 consists of zero nucleotides. 13. The cell according to any one of items 1 to 11, wherein X1 consists of nucleotides in the range of 1 to 25, such as in the range of 1 to 20, for example in the range of 1 to 15, such as in the range of 1 to 10, for example in the range of 1 to 5, for example in the range of 1 to 3. 14. The cell according to any one of items 1 to 13, wherein X5 is absent, i.e. X5 consists of zero nucleotides. 15. The cell according to any one of items 1 to 13, wherein X5 consists of 1 to 5 nucleotides, for example 1 to 3 nucleotides. 16. The cell of any one of items 1 to 15, wherein the heterologous insert sequence comprises a sequence selected from the group consisting of SEQ ID NO: 181, SEQ ID NO: 182, SEQ ID NO: 183, and SEQ ID NO: 184. 17. The cell of any one of items 1 to 16, wherein the heterologous insert sequence is selected from the group consisting of SEQ ID NO: 181, SEQ ID NO: 182, SEQ ID NO: 183, and SEQ ID NO: 184.

[0176] 18. The cell according to any one of items 1 to 17, wherein the squalene synthase is at least 75% identical, such as at least 80%, such as at least 85%, for example at least 87%, such as at least 90%, for example at least 91%, such as at least 92%, for example at least 93%, such as at least 94%, for example at least 95%, such as at least 96%, for example at least 97%, such as at least 98%, for example at least 99%, such as 100% identical to a squalene synthase selected from the group consisting of SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 201, SEQ ID NO: 202. 19. The cell of any one of items 1 to 18, wherein the promoter is a constitutive promoter or an inducible promoter. 20. The cell of any one of items 1 to 19, wherein the promoter is selected from the group consisting of endogenous promoter, GPD1, PGK1, ADH1, ADH2, PYK1, TPI1, PDC1, TEF1, TEF2, FBA1, GAL1-10, CUP1, MET2, MET14, MET25, CYC1, GAL1-S, GAL1-L, TEF1, ADH1, CAG, CMV, human UbiC, RSV, EF-1α, SV40, Mt1, Tet-On, Tet-Off, Mo-MLV-LTR, Mx1, progesterone, RU486, and rapamycin-inducible promoter. 21. The cell of any one of items 1 to 20, wherein the nucleic acid sequence further comprises a polyadenylation sequence.

[0177] 22. The cell of item 21, wherein the 5' end of the polyadenylation sequence is operably linked to the 3' end of the nucleic acid of item 1. 23. The cell of any one of items 1 to 22, wherein the nucleic acid sequence further comprises a post-transcriptional regulatory element. 24. The cell according to item 23, wherein the post-transcriptional regulatory element is a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE). 25. The cell of any one of items 1 to 24, wherein the nucleic acid comprises a 5' terminal repeat and a 3' terminal repeat. 26. The cell of item 25, wherein the 5' and 3' terminal repeats are selected from inverted terminal repeats (ITRs) and long terminal repeats (LTRs). 27. The cell of any one of items 1 to 26, wherein the nucleic acid sequence is incorporated into a vector. 28. The cell according to item 27, wherein the vector is an expression vector. 29. The cell according to item 27, wherein the vector is selected from the group consisting of a plasmid vector, a cosmid, an artificial chromosome, and a viral vector.

[0178] 30. The cell according to item 29, wherein the plasmid vector can be maintained and replicated in bacteria, fungi, and yeast. 31. The cell according to item 29, wherein the viral vector is selected from the group consisting of vectors derived from the Retroviridae family, including lentivirus, HIV, SIV, FIV, EAIV, and CIV. 32. The cell according to item 31, wherein the viral vector is selected from the group consisting of alphavirus, adenovirus, adeno-associated virus, baculovirus, HSV, coronavirus, bovine papillomavirus, Mo-MLV and adeno-associated virus. 33. The cell according to any one of items 27 to 32, wherein the vector is functional in mammalian cells. 34. The cell according to any one of items 1 to 33, wherein the cell is transformed or transduced with a vector according to any one of items 27 to 33. 35. The cell of any one of items 1 to 34, wherein the cell is a eukaryotic cell. 36. The cell according to any one of items 1 to 34, wherein the cell is a prokaryotic cell.

[0179] 37. The cell according to item 35, wherein the cell is selected from the group consisting of fungal cells, such as yeast and Aspergillus; microalgae, such as Chlorella and Prototheca; plant cells; and mammalian cells, such as human, cat, pig, monkey, dog, murine, rat, mouse, and rabbit cells. 38. The cell according to item 37, wherein the yeast is selected from the group consisting of Saccharomyces cerevisiae, Schizosaccharomyces pombe, Yarrowia lipolytica, Candida glabrata, Ashbya gossypii, Cyberlindnera jadinii, and Candida albicans. 39. The cell according to item 37, wherein the cell is selected from the group consisting of CHO, CHO-K1, HEI193T, HEK293, COS, PC12, HiB5, RN33b, and BHK cells. 40. The cell according to item 36, wherein the cell is E. coli, Corynebacterium, Bacillus, Pseudomonas, or Streptomyces. 41. The cell of any one of items 35 to 40, wherein the prokaryotic or fungal cell is genetically engineered to express at least some of the enzymes of the mevalonate-independent pathway.

[0180] 42. The cell of any one of items 1 to 41, wherein the cell further comprises a heterologous nucleic acid encoding GGPPS operably linked to a nucleic acid sequence that directs expression of GGPPS in the cell. 43. The cell according to item 42, wherein the GGPPS is selected from the group consisting of SEQ ID NO: 126, SEQ ID NO: 123, SEQ ID NO: 203, SEQ ID NO: 167, and functional homologs thereof having at least 75% sequence identity to any of the foregoing. 43. A method for producing a terpenoid compound synthesized via the squalene pathway, the method comprising the steps of: (a) providing the cell according to any one of items 1 to 42; (b) culturing the cells of step (a); (c) recovering the terpenoid product compounds; The method comprising:

[0181] 44. A method for producing terpenoids derived from terpenoid precursors selected from the group consisting of farnesyl pyrophosphate (FPP), isopentenyl pyrophosphate (IPP), dimethylallyl pyrophosphate (DMAPP), geranyl pyrophosphate (GPP), and / or geranylgeranyl pyrophosphate (GGPP), comprising: (a) contacting the precursor with an enzyme of the squalene synthase pathway; (b) recovering the terpenoid product; The method comprising: 45. The method according to item 44 or 45, wherein the terpenoid product is selected from the group consisting of hemiterpenoids, monoterpenes, sesquiterpenoids, diterpenoids, sesterpenes, triterpenoids, tetraterpenoids, and polyterpenoids. 46. ​​The method according to item 44, wherein the terpenoid is selected from the group consisting of farnesyl phosphate, farnesol, geranylgeranyl, geranylgeraniol, isoprene, prenol, isovaleric acid, geranyl pyrophosphate, eucalyptol, limonene, pinene, farnesyl pyrophosphate, artemisinin, bisabolol, geranylgeranyl pyrophosphate, retinol, retinal, phytol, taxol, forskolin, aphidicolin, lanosterol, lycopene, and carotene.

[0182] 47. The method of claim 46, wherein the method further comprises dephosphorylating farnesyl phosphate to produce farnesol. 48. The method of item 44, wherein the enzyme of the squalene synthase pathway is selected from the group consisting of dimethylallyltransferase (EC 2.5.1.1), isoprene synthase (EC 4.2.3.27), and geranyltranstransferase (EC 2.5.1.10). 49. A method for reducing the translation rate of functional squalene synthase (EC 2.5.1.21), comprising: (a) providing the cell according to any one of items 1 to 42; (b) culturing the cells of (a); The method comprising: 50. A method for reducing the conversion of farnesyl-pp to squalene, the method comprising: (d) providing the cell according to any one of items 1 to 42; (e) culturing the cells of (d); The method comprising: 51. A method for enhancing accumulation of a compound selected from the group consisting of farnesyl pyrophosphate, isopentenyl pyrophosphate, dimethylallyl pyrophosphate, geranyl pyrophosphate, and geranylgeranyl pyrophosphate, comprising the steps of: (a) providing a cell according to any one of items 1 to 42, and (b) culturing the cells of step (a). The method comprising:

[0183] 52. The method of claim 51, further comprising recovering the farnesyl pyrophosphate, isopentenyl pyrophosphate, dimethylallyl pyrophosphate, geranyl pyrophosphate, or geranylgeranyl pyrophosphate compound. 53. The method of item 51 or 52, further comprising recovering a compound synthesized via the squalene pathway, wherein the compound is derived from farnesyl pyrophosphate, isopentenyl pyrophosphate, dimethylallyl pyrophosphate, geranyl pyrophosphate, and / or geranylgeranyl pyrophosphate. 54. The method according to any one of items 43 to 53, wherein the step of culturing the cells is carried out in the presence of a squalene synthase inhibitor. 55. The cells further express phosphomevalonate kinase (EC 2.7.4.2), diphosphomevalonate decarboxylase (EC 4.1.1.33), 4-hydroxy-3-methylbut-2-en-1-yl diphosphate synthase (EC 1.17.7.1), 4-hydroxy-3-methylbut-2-enyl diphosphate reductase (EC 1.17.1.2), isopentenyl diphosphate δ-isomerase 1 (EC 5.3.3.2), short-chain Z-isoprenyl diphosphate synthase (EC 2.5.1.68), dimethylallyltransferase (EC 2.5.1.1), geranyltransferase (EC 2.5.1.10), and geranylgeranyl pyrophosphate synthase (EC 55. The method according to any one of items 43 to 54, wherein the cell culture medium is genetically engineered to enhance the activity of and / or overexpress one or more enzymes selected from the group consisting of: 2.5.1.29).

[0184] 56. The method of any one of paragraphs 43 to 55, wherein the cells are further genetically engineered to enhance the activity of and / or overexpress one or more enzymes selected from the group consisting of acetoacetyl-CoA thiolase, HMG-CoA reductase or its catalytic domain, HMG-CoA synthase, mevalonate kinase, phosphomevalonate kinase, phosphomevalonate decarboxylase, isopentenyl pyrophosphate isomerase, farnesyl pyrophosphate synthase, D-1-deoxyxylulose 5-phosphate synthase, and 1-deoxy-D-xylulose 5-phosphate reductoisomerase and farnesyl pyrophosphate synthase. 57. The method of any one of items 43 to 56, wherein the cell comprises a mutation in the ERG9 open reading frame. 58. The method of any one of items 43 to 57, wherein the cell comprises an ERG9[delta]::HIS3 deletion / insertion allele. 59. The method of any one of items 43 to 58, wherein recovering the compound comprises purifying the compound from the cell culture medium.

[0185] VI. Examples The present invention is further described in the following examples, which do not limit the scope of the invention as described in the claims. In the examples described herein, unless otherwise specified, the following LC-MS methods were used for the analysis of steviol glycosides and steviol pathway intermediates. 1) Analysis of steviol glycosides LC-MS analysis was performed using an Agilent 1200 Series HPLC system (Agilent Technologies, Wilmington, DE, USA) equipped with a Phenomenex® kinetex C18 column (150 × 2.1 mm, 2.6 μm particles, 100 Å pore size) connected to a TSQ Quantum Access (ThermoFisher Scientific) triple quadrupole mass spectrometer with a heated electrospray ionization (HESI) source. Elution was performed using a mobile phase of eluent B (MeCN with 0.1% formic acid) and eluent A (water with 0.1% formic acid) with a gradient increasing from 10 to 40% B from 0.0 to 1.0 min, from 40 to 50% B from 1.0 to 6.5 min, and from 50 to 100% B from 6.5 to 7.0 min, with a final wash and re-equilibration. The flow rate was 0.4 ml / min, and the column temperature was 30° C. Detection of steviol glycosides was performed using SIM (Single Ion Monitoring) in positive mode with the following m / z traces:

[0186] [Table 11] Steviol glycoside levels were quantified by comparison with a calibration curve obtained using certified reference materials from the LGC standard. For example, standard solutions of 0.5–100 μM rebaudioside A were commonly used to generate a calibration curve.

[0187] 2) Analysis of steviol and ent-kaurenoic acid LC-MS analysis of steviol and ent-kaurenoic acid was performed using the above system. For separation, a Thermo Science Hypersil Gold column (C-18, 3 μm, 100 × 2.1 mm) was used, with 20 mM aqueous ammonium acetate as eluent A and acetonitrile as eluent B. The gradient conditions were 20 to 55% B from 0.0 to 1.0 min, 55 to 100% B from 1.0 to 7.0 min, and final wash and re-equilibration. The flow rate was 0.5 ml / min, and the column temperature was 30°C. Steviol and ent-kaurenoic acid were detected using SIM (Single Ion Monitoring) in negative mode using the following m / z traces: [Table 12]

[0188] 3) HPLC quantification of UDP-glucose For the quantification of UDP-glucose, an Agilent 1200 Series HPLC system was used with a Waters XBridgeBEH Amide (2.5 μM, 3.0 × 50 mm) column. Eluent A was 10 mM aqueous ammonium acetate (pH 9.0), and eluent B was acetonitrile. The gradient conditions were: 0.0 min to 0.5 min: hold at 95% B, 0.5 min to 4.5 min: decrease from 95 to 50% B, 4.5 min to 6.8 min: hold at 50% B, and finally re-equilibration to 95% B. The flow rate was 0.9 ml / min, and the column temperature was 20°C. UDP-glucose was analyzed by UV spectroscopy. 262nm The detection was based on the absorbance of the sample. The amount of UDP-glucose was quantified by comparison with a calibration curve obtained using commercially available standards (eg, from Sigma Aldrich).

[0189] Example 1 - Identifying EUGT11 Fifteen genes were tested for RebA 1,2-glycosylation activity. See Table 10. [Table 13]

[0190] These genes were transcribed and translated in vitro, and the resulting UGTs were incubated with RebA and UDP-glucose. After incubation, the reactions were analyzed by LC-MS. The reaction mixture containing EUGT11 (rice, AC133334, SEQ ID NO: 152) was shown to convert significant amounts of RebA to RebD. See the LC-MS chromatograms in Figure 4. As shown in the left panel of Figure 4, UGT91D2e produced trace amounts of RebD when RebA was used as the feedstock. As shown in the right panel of Figure 4, EUGT11 produced significant amounts of RebD when RebA was used as the feedstock. Preliminary quantification of the amount of RebD produced indicated that EUGT11 was approximately 30 times more efficient than UGT91D2e at converting RebA to RebD. To further characterize EUGT11 and quantitatively compare it to UGT91D2e, the nucleotide sequence encoding EUGT11 (SEQ ID NO: 153, non-codon optimized, FIG. 7) was cloned into two E. coli expression vectors, one containing an N-terminal HIS tag and one containing an N-terminal GST tag. EUGT11 was expressed and purified using both systems. Incubation of the purified enzyme with UDP-glucose and RebA produced RebD.

[0191] Example 2 - Identification of the EUGT11 reaction EUGT11 was produced by in vitro transcription and translation and incubated with various substrates in the RebD pathway. Similar experiments were performed using in vitro transcribed and translated UGT91D2e. Figure 3 shows a schematic diagram of the 19-O-1,2-diglycosylation reaction performed by EUGT11 and UGT91D2e. Compounds 1–3 were identified solely by mass and estimated retention time. The numbers in Figure 3 are the average peak heights of the indicated steviol glycosides obtained from LC-MS chromatograms. While not quantitative, they can be used to compare the activities of the two enzymes. EUGT11 and UGT91D2e were unable to use steviol as a substrate. Both enzymes were able to convert steviol 19-O-monoglucoside (SMG) to compound 1, with EUGT11 being approximately 10-fold more efficient than UGT91D2e at converting 19-SMG to compound 1.

[0192] Both enzymes were able to convert rubusoside to stevioside with comparable activity, but only EUGT11 was able to convert rubusoside to compounds 2 and 3 (RebE). See Figure 5. The left panel of Figure 5 shows LC-MS chromatograms of the conversion of rubusoside to stevioside. The right panel of Figure 5 shows chromatograms of the conversion of rubusoside to stevioside, compounds 2, and 3 (RebE). The conversion of rubusoside to compound 3 requires two sequential 1,2-O-glycosylations at the 19 and 13 positions of steviol. UGT91D2e was able to produce trace amounts of compound 3 (RebE) in one experiment, while EUGT11 produced significant amounts of compound 3. Both enzymes were able to convert RebA to RebD, but EUGT11 was approximately 30-fold better at converting RebA to RebD. Overall, EUGT11 appears to produce more product than UGT91D2e in all reactions (with similar time, concentration, temperature, and enzyme purity) except for the conversion of rubusoside to stevioside.

[0193] Example 3 - Expression of EUGT11 in yeast The nucleotide sequence encoding EUGT11 was codon-optimized (SEQ ID NO: 154) and transformed into yeast along with nucleic acids encoding all four UGTs (UGT91D2e, UGT74G1, UGT76G1, and UGT85C2). The resulting yeast strain was grown in medium containing steviol, and accumulated steviol glycosides were analyzed by LC-MS. EUGT11 was required for the production of RebD. In other experiments, production of RebD was observed with UGT91D2e, UGT74G1, UGT76G1, and UGT85C2.

[0194] Example 4 - UGT activity in 19-O-1,2-diglycosylated steviol glycosides The 19-O-1,2-diglycosylated steviol glycosides produced by EUGT11 require further glycosylation for conversion to RebD. The following experiments were performed to determine whether other UGTs could use these intermediates as substrates. In one experiment, compound 1 was produced in vitro from 19-SMG by either EUGT11 or UGT91D2e in the presence of UDP-glucose. After boiling the sample, UGT85C2 and UDP-glucose were added. The sample was analyzed by LC-MS, and compound 2 was detected. This experiment demonstrated that UGT85C2 can use compound 1 as a substrate. In another experiment, compound 2 was incubated with UGT91D2e and UDP-glucose. The reaction was analyzed by LC-MS. UGT91D2e was unable to convert compound 2 to compound 3 (RebE). Incubating compound 2 with EUGT11 and UDP-glucose resulted in the production of compound 3. UGT76G1 was able to use RebE as a substrate for RebD production. This indicates that 19-O-1,2-diglycosylation of steviol glycosides can occur at any time during RebD production because downstream enzymes can metabolize the 19-O-1,2-diglycosylated intermediates.

[0195] Example 5 - Sequence comparison of EUGT11 and UGT91D2e The amino acid sequence of EUGT11 (SEQ ID NO: 152, FIG. 7) and the amino acid sequence of UGT91D2e (SEQ ID NO: 5) were aligned using the FASTA algorithm (Pearson and Lipman, Proc. Natl. Acad. Sci., 85:2444-2448 (1998)). See FIG. 6. EUGT11 and UGT91D2e are 42.7% identical over 457 amino acids.

[0196] Example 6 - Modulation of the 19-1,2-diglycosylation activity of UGT91D2e Crystal structures are available for many UGTs. In general, the N-terminal half of a UGT is primarily involved in substrate binding, while the C-terminal half is involved in binding the UDP-sugar donor. Modeling the secondary structure of UGT91D2e onto the secondary structures of crystallized UGTs revealed a conserved pattern of secondary structure despite the highly divergent primary sequences, as shown in Figure 8. Crystal structures of UGT71G1 and UGT85H2 have been reported (e.g., H. Shao et al., The Plant Cell November 2005 vol. 17 no 11 3141-3154 and L. Li et al., J Mol Biol. 2007 370(5):951-63). The known loops, α-helices, and β-sheets are shown for UGT91D2e in Figure 8. Although the homology of these UGTs at the primary structure level is fairly low, the secondary structures appear to be conserved, allowing predictions about the locations of amino acids involved in substrate binding to UGT91D2e based on the locations of such amino acids in UGT85H2 and UGT71G1.

[0197] Superimposition of the region commonly involved in substrate binding onto UGT91D2e showed that it closely matches the 22 amino acid difference from UGT91D1 (GenBank Accession No. Protein Accession No. AAR06918, GI:37993665). UGT91D1 is highly expressed in Stevia and is thought to be a functional UGT. However, its substrate is not steviol glycoside. This suggests that UGT91D1 has a different substrate, defined by 22 amino acids, than UGT91D2e. Figure 9 shows an alignment of the amino acid sequences of UGT91D1 and UGT91D2e. The boxes represent the regions reported to be involved in substrate binding. The amino acids highlighted in dark gray indicate the 22 amino acid differences between UGT91D1 and UGT91D2e. Asterisks indicate amino acids that have been shown to be involved in substrate binding in UGTs whose crystal structures have been determined (multiple asterisks under a particular amino acid indicate that substrate binding has been shown in more than one UGT whose structure has been determined). There is a strong correlation between the 22 amino acid differences between the two UGT91s, which are regions known to be involved in substrate binding and are the amino acids that are actually involved in substrate binding in the UGTs whose crystal structures have been determined. This suggests that the 22 amino acid differences between the two UGT91s are involved in substrate binding.

[0198] All 22 modified 91D2e variants were expressed in the XJb autolytic E. coli strain from the pGEX-4T1 vector. To assess the enzyme activity, we performed two substrate feeding experiments: in vitro and in vivo. While most mutants had lower activity than the wild-type, five mutants showed increased activity. This was recapitulated by in vitro transcription and translation (IVT). C583A, C631A, and T857C had approximately three-fold higher stevioside-forming activity than wild-type UGT91D2e, while C662t and A1313C had approximately two-fold higher stevioside-forming activity (nucleotide numbering). These modifications result in amino acid mutations corresponding to L195M, L211M, V286A; and S221F and E438A, respectively. The increased activity varied depending on the substrate, with C583A and C631A showing an almost 10-fold increase using 13-SMG as the substrate and an approximately 3-fold increase using rubusoside as the substrate, while T857C showed a 3-fold increase using either 13-SMG or rubusoside as the substrate. To investigate whether these mutations are additive, a range of double mutants were created and analyzed for activity (Figure 10). In this particular experiment, higher wild-type levels of activity were observed than in the previous four experiments; however, the relative activities of the mutants remained the same. Because rubusoside accumulates in many S. cerevisiae strains expressing four UGTs (UGT74G1, UGT85C2, UGT76G1, and UGT91D2e), stevioside-forming activity may be more important for increasing steviol glycoside production. Thus, the double mutant C631A / T857C (nucleotide numbering) may be useful. This mutant, designated UGT91D2e-b, contains the amino acid modifications L211M and V286A. The experiment was replicated in vitro using S. cerevisiae-expressed UGT91D2e mutants.

[0199] To improve the 19-1,2-diglycosylation activity of UGT91D2e, we performed a directed saturation mutagenesis screen of UGT91D2e for the 22 amino acid differences between UGT91D2e and UGT91D1. Using site-saturation mutagenesis from GeneArt's® (Life Technologies, Carlsbad, CA), we obtained a library containing each mutation. The library was cloned into the BamHI and NotI sites of the pGEX4T1 bacterial expression plasmid, which expresses mutant forms of 91D2e as GST fusion proteins, to obtain a new library (Lib#116). Lib#116 was transformed into the XJbAutolysis E. coli strain (ZymoResearch, Orange, CA) to generate approximately 1600 clones containing 418 predicted mutations (i.e., 22 positions with 19 different amino acids at each position). Other plasmids expressing GST-tagged versions of 91D2e (EPCS1314), 91D2e-b (EPSC1888), or EUGT11 (EPSC1744) as well as empty pGEX4T1 (PSB12) were transformed similarly.

[0200] LC-MS screening To analyze approximately 1600 mutant clones of UGT91D2e, E. coli transformants were grown overnight at 30°C in 1 ml of NZCYM containing ampicillin (100 mg / L) and chloramphenicol (33 mg / L) in a 96-well format. The next day, 150 μl of each culture was inoculated into 3 ml of NZCYM containing ampicillin (100 mg / L), chloramphenicol (33 mg / L), arabinose 3 mM, IPTG 0.1 mM, and ethanol 2% (v / v) in a 24-well format and incubated at 20°C and 200 rpm for ∼20 hours. The next day, cells were spun down, and the pellet was resuspended in 100 μl of lysis buffer containing 10 mM Tris-HCl, pH 8, 5 mM MgCl2, 1 mM CaCl2, and Complete Mini Protease Inhibitor EDTA-free (3 tablets / 100 ml) (Hoffmann-La Roche, Basel, Switzerland) and frozen at -80°C for at least 15 minutes to facilitate cell lysis. The pellet was thawed at room temperature, and 50 μl of DNase mix (1 μl of 1.4 mg / ml DNase (~80,000 / ml) in H2O, 1.2 μl of 500 mM MgCl2, and 47.8 μl of 4x PBS buffer) was added to each well. The plate was shaken at 500 rpm for 5 minutes at room temperature to allow for degradation of genomic DNA. Plates were spun down at 4000 rpm for 30 minutes at 4°C, and 6 μl of the lysate was used in UGT in vitro reactions using rubusoside or rebaudioside A as substrates, as described for GST-91D2e-b. In each case, the resulting compounds, stevioside or rebaudioside D (rebD), were measured by LC-MS. Results were analyzed in comparison with stevioside or rebD produced by lysates expressing the corresponding controls (91D2e, 91D2e-b, EUGT11, and empty plasmid). Clones showing similar or higher activity than those expressing 91D2e-b were selected as primary hits.

[0201] Half of the 1,600 clones and corresponding controls were assayed for their ability to glycosylate rubusoside and rebaudioside A. Stevioside and RebD were quantified by LC-MS. Under the conditions used, lysates from clones expressing native UGT91D2e exhibited near-background activity for both substrates (approximately 0.5 μM stevioside and 1 μM RebD), whereas clones expressing UGT91D2e-b consistently showed improved product formation (>10 μM stevioside; >1.5 μM RebD). Clones expressing EUGT11 consistently exhibited high levels of activity, particularly with RebA as a substrate. In screening, the cutoff for considering clones as primary hits was generally 1.5 μM for both products, but in some cases was adjusted for each independent assay.

[0202] Example 7—EUGT11 homolog A Blastp search of the NCBInr database using the EUGT11 protein sequence revealed approximately 79 potential UGT homologs from 14 plant species (one of which, Stevia UGT91D1, is approximately 67% identical to EUGT11 in the conserved UGT region but less than 45% overall). Homologs with 90% or greater identity in the conserved region were identified from maize, soybean, Arabidopsis, grapevine, and Sorghum. The overall homology of full-length EUGT11 homologs ranged from 28 to 68% at the amino acid level. RNA was extracted from plant material according to the method described by Iandolino et al. (Iandolino et al., Plant Mol Biol Reporter 22, 269-278, 2004), using the RNeasy Plant mini Kit (Qiagen) according to the manufacturer's instructions, or using the Fast RNA Pro Green Kit (MP Biomedicals) according to the manufacturer's instructions. cDNA was generated using the AffinityScript QPCR cDNA Synthesis Kit (Agilent) according to the manufacturer's instructions. Genomic DNA was extracted using the FastDNA Kit (MP Biomedicals) according to the manufacturer's instructions. PCR was performed on the cDNA using either Dream Taq polymerase (Fermentas) or Phusion polymerase (New England Biolabs) and a series of primers designed to amplify homologs.

[0203] PCR reactions were analyzed by electrophoresis in a SyberSafe-containing agarose-TAE gel. DNA was visualized by UV irradiation on a transilluminator. Bands of the correct size were excised, purified using spin columns according to the manufacturer's specifications, and cloned into TOPO-Zero Blunt (for products generated with Phusion polymerase) or TOPO-TA (for products generated with DreamTaq). TOPO vectors containing the PCR products were transformed into E. coli DH5Bα and plated on LB-agar plates containing the appropriate selection antibiotic. DNA was extracted from surviving colonies and sequenced. Genes with the correct sequence were excised by restriction digestion with SbfI and AscI, cloned into a similarly digested IVT8 vector, and transformed into E. coli. PCR was performed on all cloned genes to amplify the gene and flanking regions required for in vitro transcription and translation. Proteins were produced by in vitro transcription and translation from PCR products using the Promega L5540, TNT T7 Quick for PCR DNA Kit according to the manufacturer's instructions. 35 S-methionine incorporation was assessed followed by separation by SDS-PAGE and visualization on a Typhoon fluorescent imager.

[0204] Activity assays were set up: 20% (by volume) of each in vitro reaction, 0.1 M rubusoside or RebA, 5% DMSO, 100 mM Tris-HCl, pH 7.0, 0.01 units of fast alkaline phosphatase (Fermentas), and 0.3 mM UDP-glucose (final concentrations). After 1 hour of incubation at 30°C, samples were analyzed by LC-MS for stevioside and RebD production as described above. UGT91D2e and UGT91D2e-b (double mutants described in Example 6) were used as positive controls, along with EUGT11. Under the initial assay conditions, clone P64B (see Table 11) produced trace amounts of product using rubusoside and RebA. Table 11 shows the percent identity at the amino acid level compared to EUGT11 for the full-length UGTs, which ranged from 28 to 58%. A high amount of homology (96-100%) was observed over short stretches of sequence, which may indicate highly conserved domains of plant UGTs. [Table 14-1] [Table 14-2]

[0205] Example 8 - Cell-free biocatalytic production of Reb-D The cell-free approach is an in vitro system in which RebA, stevioside, or a mixture of steviol glycosides is enzymatically converted to RebD. This system requires a stoichiometric amount of UDP-glucose, and therefore regeneration of UDP-glucose from UDP and sucrose using sucrose synthase can be used. Furthermore, sucrose synthase removes UDP generated during the reaction, which improves conversion to glycosylated products due to the reduced product inhibition observed for glycosylation reactions. See WO 2011 / 153378.

[0206] Enzyme expression and purification UGT91D2e-b (described in Example 6) and EUGT11 are key enzymes that catalyze the glycosylation of RebA to produce RebD. While these UGTs were expressed in bacteria (E. coli), those skilled in the art will understand that such proteins can also be prepared using different methods and hosts (e.g., other bacteria such as Bacillus species, yeast such as Pichia species or Saccharomyces species, other fungi (e.g., Aspergillus), or other organisms). For example, the proteins can be produced by in vitro transcription and translation or by protein synthesis. The UGT91D2e-b and EUGT11 genes were cloned into pET30a or pGEX4T1 plasmids. The resulting vectors were transformed into the XJb(DE3) autolytic E. coli strain (ZymoResearch, Orange, CA). E. coli transformants were first grown overnight at 30°C in NZCYM medium, then induced with 3 mM arabinose and 0.1 mM IPTG and further incubated overnight at 30°C. The corresponding fusion proteins were purified by affinity chromatography using 6HIS- or GST-tags and standard methods. Those skilled in the art will understand that other protein purification methods, such as gel filtration or other chromatographic techniques, can also be used, for example, in conjunction with ammonium sulfate precipitation / crystallization or fractionation. While EUGT11 was expressed well using the initial conditions, UGT91D2e-b required some modifications to the basic protocol to increase protein solubility, including lowering the temperature of overnight expression from 30 °C to 20 °C and adding 2% ethanol to the expression medium. Typically, 2-4 mg / L of soluble GST-EUGT11 and 400-800 µg / L of GST-UGT91D2e-b were purified with this method.

[0207] EUGT11 stability Reactions were performed to explore the stability of EUGT11 under various RebA to RebD reaction conditions. The substrate was omitted from the reaction mixture, and EUGT11 was preincubated for various times. Following preincubation of the enzyme in 100 mM Tris-HCl buffer, the substrate (100 μM RebA) and other reaction components (300 μM UDP-glucose and 10 U / mL alkaline phosphatase (Fermentas / Thermo Fisher, Waltham, MA)) were added (0, 1, 4, or 24 h after the start of incubation). The reaction was then allowed to proceed for 20 h, after which the reaction was stopped and the formation of RebD product was measured. The experiment was repeated at different temperatures: 30°C, 32.7°C, 35.8°C, and 37°C. When the enzyme was preincubated at 37°C, the activity of EUGT11 rapidly decreased, reaching approximately half of its activity after 1 hour and almost no activity after 4 hours. At 30°C, the activity did not decrease significantly after 4 hours, and approximately one-third of the activity remained after 24 hours, suggesting that EUGT11 is thermolabile.

[0208] To assess the thermal stability of EUGT11 and compare it to other UGTs in the steviol glycosylation pathway, the denaturation temperature of the protein was determined using differential scanning calorimetry (DSC). D The use of DSC thermograms to estimate T is described, for example, by E. Freire in Methods in Molecular Biology 1995, Vol. 40, 191-218. DSC was performed with 6HIS-purified EUGT11, and the apparent T D whereas, when GST-purified 91D2e-b was used, the measured T D The T was 79°C. For reference, the T was measured using 6HIS-purified UGT74G1, UGT76G1, and UGT85C2. DThe temperature was 86°C in all cases. One skilled in the art can improve protein stability by enzyme immobilization or by adding thermoprotectants to the reaction. Non-limiting examples of thermoprotectants include trehalose, glycerol, ammonium sulfate, betaine, trimethylamine oxide, and proteins.

[0209] Enzyme kinetics A series of experiments was performed to determine the kinetic parameters of EUGT11 and 91D2e-b. For both enzymes, 100 μM RebA, 300 μM UDP-glucose, and 10 U / mL alkaline phosphatase (Fermentas / Thermo Fisher, Waltham, MA) were used in the reactions. For EUGT11, the reaction was carried out at 37°C using 100 mM Tris-HCl, pH 7, and 2% enzyme. For 91D2e-b, the reaction was carried out at 30°C using 20 mM Hepes-NaOH, pH 7.6, and 20% (by volume) enzyme. The initial velocity (V0) was calculated in the linear range of the product versus time plot. To first investigate the linear range, initial time courses were performed for each enzyme. EUGT11 was tested at initial concentrations of 100 μM RebA and 300 μM UDP-glucose at 37°C for 48 h. UGT91D2e-b was tested at initial concentrations of 200 μM RebA and 600 μM UDP-glucose at 37°C for 24 h. Based on these range-finding studies, we determined that the first 10 min for EUGT11 and the first 20 min for UGT91D2e-b were linear ranges for product formation, and therefore the initial rates for each reaction were calculated within these intervals. For EUGT11, the RebA concentrations tested were 30 μM, 50 μM, 100 μM, 200 μM, 300 μM, and 500 μM. The UDP-glucose concentration was always three times the RebA concentration, and incubations were performed at 37°C. The Michaelis-Menten curve was constructed by plotting the calculated V as a function of substrate concentration. The Lineweaver-Burk diagram was obtained by plotting the reciprocal of V versus the reciprocal of [S], where y = 339.85x + 1.8644; R 2=0.9759.

[0210] V max and K. M Parameters were determined from curve-fitted Lineweaver-Burk data calculated from the x and y intercepts: (x = 0, y = 1 / V max ) and (y=0, x=-1 / K M ). Additionally, the same parameters were calculated by nonlinear least-squares regression using the Solver function in Excel. The results for EUGT11 and RebA obtained by both methods are shown in Table 12, along with all the kinetic parameters for this example. The results from both the nonlinear least-squares fit and the Lineweaver-Burk plot are shown in Table 12. K cat is V max The calculation is based on dividing by the approximate amount of protein in the assay.

[0211] [Table 15]

[0212] Similar kinetic analyses were performed to examine the effects of UDP-glucose concentration and the affinity of EUGT11 for UDP-glucose on the glycosylation reaction. EUGT11 was incubated with increasing amounts of UDP-glucose (20 μM, 50 μM, 100 μM, and 200 μM) while maintaining an excess of RebA (500 μM). Kinetic parameters were calculated as described above and are shown in Table 12. For UGT91D2e-b, the RebA concentrations tested were 50 μM, 100 μM, 200 μM, 300 μM, 400 μM, and 500 μM. The UDP-glucose concentration was always three times the RebA concentration, and incubations were performed at 30°C under the reaction conditions described above for UGT91D2e-b. Kinetic parameters were calculated as described above, and the resulting kinetic parameters are shown in Table 12. Additionally, the kinetic parameters of UGT91D2e-b toward UDP-glucose were also determined. UGT91D2e-b was incubated with increasing amounts of UDP-glucose (30 μM, 50 μM, 100 μM, and 200 μM), while maintaining an excess of RebA (1500 μM). Incubations were performed at 30°C under the optimal conditions for UGT91D2e-b. Kinetic parameters were calculated as described above, and the results are shown in Table 12. Comparison of the kinetic parameters for EUGT11 and 91D2e-b reveals that 91D2e-b has a lower K cat and lower affinity for RebA (higher K M ), but the K M It was concluded that UGT91D2e-b has a lower catalytic rate (K cat ) and the strength of the enzyme-substrate bond (K M ), we can obtain a measure of catalytic efficiency, K cat / K M has a lower value for

[0213] Determine the limiting factor in the reaction Under the conditions described above for EUGT11, approximately 25% of the applied RebA was converted to RebD. The limiting factor in these conditions was either the enzyme, UDP-glucose, or RebA. An experiment was designed to distinguish between these possibilities. A standard assay procedure was performed for 4 hours. Next, either extra RebA substrate, extra enzyme, extra UDP-glucose, or extra enzyme and UDP-glucose was added. Addition of extra enzyme resulted in a relative increase in conversion of approximately 50%. Addition of extra RebA or UDP-glucose alone did not significantly increase conversion, but simultaneous addition of enzyme and UDP-glucose increased conversion by approximately 2-fold. Experiments were conducted to examine the limitations of adding a bolus of UDP-glucose and fresh enzyme in the conversion of RebA to RebD. Additional enzyme or enzyme and UDP-glucose were added after 1, 6, 24, and 28 hours. Conversions of over 70% were achieved when both extra EUGT11 and UDP-glucose were added. Other components did not significantly affect conversion. This indicates that EUGT11 is the primary limiting factor in the reaction, but UDP-glucose is also limiting. Because UDP-glucose is present at a threefold higher concentration than RebA, this suggests that UDP-glucose may be somewhat unstable in the reaction mixture, at least in the presence of EUGT11. Alternatively, as discussed below, EUGT11 may metabolize UDP-glucose.

[0214] Inhibition test Experiments were conducted to determine whether factors such as sucrose, fructose, UDP, product (RebD), and impurities in low-purity stevia extract raw materials inhibit the extent of conversion of steviol glycoside substrates to RebD. In a standard reaction mixture, excess amounts of potential inhibitors (sucrose, fructose, UDP, RebD, or a commercial blend of steviol glycosides (Stevia, Steviva Brands, Inc., Portland, OR)) were added. After incubation, RebD production was quantified. Addition of 500 μg / ml of commercial stevia mix (approximately 60% 1,2-stevioside, 30% RebA, 5% rubusoside, 2% 1,2-bioside, less than 1% RebD, RebC, etc., as assessed by LC-MS) was not found to be inhibitory, but did increase overall RebD production (from approximately 30 μM without addition to approximately 60 μM) by the blend, far exceeding the initial RebD (approximately 5 μM). Among the molecules tested, only UDP was shown to have an inhibitory effect on RebD production at the concentration used (500 μM), as measured by LC-MS. Less than 7 μM RebD was produced. This inhibition can be alleviated in in vitro or in vivo reactions for RebD production by including a UDP recycling system in UDP-glucose, either with yeast or by adding SUS (sucrose synthase enzyme) along with sucrose. Furthermore, when working with lower amounts of UDP-glucose (300 μM), the addition of alkaline phosphatase to remove UDP-G did not increase the amount of RebD produced by in vitro glycosylation, suggesting that the produced UDP may not be inhibitory at these concentrations.

[0215] RebA vs. Crude Steviol Glycoside Mix In some experiments, a crude steviol glycoside mix was used as the source of RebA instead of purified RebA. Because such a crude steviol glycoside mix contains a high proportion of stevioside along with RebA, UGT76G1 was included in the reaction. In vitro reactions were carried out as described above using 0.5 g / L Steviva® mix as substrate and enzyme (UGT76G1 and / or EUGT11) and incubated at 30°C. The presence of steviol glycosides was analyzed by LC-MS. When only UGT76G1 was added to the reaction, stevioside was converted very efficiently to RebA. An unknown pentaglycoside (with a retention time peak at 4.02 min) was also detected. When only EUGT11 was added to the reaction, large amounts of RebE, RebA, RebD, and an unknown steviol pentaglycoside (with a retention time peak at 3.15 min) were found. When both EUGT11 and UGT76G1 were added to the reaction, the stevioside peak decreased and was almost completely converted to RebA and RebD. A trace amount of the unknown steviol pentaglycoside (peak at 4.02 min) was also present. RebE was not detected, nor was the second unknown steviol pentaglycoside (peak at 3.15 min). These results suggest that the use of stevia extract as a substrate for in vitro production of RebD is possible when EUGT11 and UGT76G1 are used in combination.

[0216] Nonspecific UDP-glucose metabolism To determine whether EUGT11 can metabolize UDP-glucose independently of the conversion of RebA to RebD, GST-purified EUGT11 was incubated in the presence or absence of the RebA substrate, and UDP-glucose utilization was measured as UDP release using the TR-FRET Transcreener® kit (BellBrook Labs). The Transcreener® kit is based on a tracer molecule conjugated to an antibody. The tracer molecule is sensitively and quantitatively displaced by UDP or ADP. The FP kit contains an Alexa633 tracer conjugated to an antibody. The tracer is displaced by UDP / ADP. The displaced tracer is free to rotate, resulting in a decrease in fluorescence polarization. Therefore, UDP production is proportional to the decrease in polarization. The FI kit contains a quenched Alexa594 tracer conjugated to an antibody, which is coupled to an IRDye® QC-1 quencher. The tracer is displaced by UDP / ADP, which prevents the displaced tracer from being quenched, resulting in a positive increase in fluorescence intensity. Therefore, UDP production is proportional to the increase in fluorescence. The TR-FRET kit contains a HiLyte647 tracer bound to an antibody-Tb complex. Excitation of the terbium complex in the UV range (approximately 330 nm) results in energy transfer to the tracer and, after a time delay, emission at a higher wavelength (665 nm). The tracer is displaced by UDP / ADP, causing a decrease in TR-FRET. The measured UDP-glucose was observed to be identical, independent of the presence of the RebA substrate. No UDP release was detected in the absence of the enzyme, indicating nonspecific degradation of UDP-glucose by EUGT11. Nevertheless, RebD continued to be produced when RebA was added, suggesting that EUGT11 preferentially catalyzes RebA glycosylation over nonspecific UDP-glucose degradation.

[0217] Experiments were set up to explore the destination of glucose molecules in the absence of RebA or other obvious glycosylation substrates. One common factor in all previous reactions was the presence of Tris buffer and / or trace amounts of glutathione, both of which contain potential glycosylation sites. The effect of these molecules on nonspecific UDP-glucose consumption was assayed in in vitro reactions using GST-purified EUGT11 (with glutathione) and HIS-purified enzyme (without glutathione) in the presence or absence of RebA. UDP-glucose utilization was measured as UDP-release using the TR-FRETTranscreener® kit. UDP release occurred in all cases and was independent of the presence of RebA. UDP release was slower when using the HIS-purified enzyme, but the overall catalytic activity of the enzyme in converting RebA to RebD was also lower, suggesting a reduced amount of active, soluble enzyme present in the assay. Thus, UDP-glucose metabolism by EUGT11 appears to be independent of the presence of substrate under the conditions tested and independent of the presence of glutathione in the reaction. To test the effect of Tris on the metabolism of UDP-glucose by EUGT11, GST-EUGT11 was purified using Tris- or PBS-based buffers for elution, yielding similar amounts of protein in both cases. Tris- and PBS-purified enzymes were used in in vitro reactions using Tris and HEPES as buffers, respectively, in the presence or absence of RebA, in a manner similar to that described above. Under both conditions, UDP release was identical in reactions with or without RebA added, indicating that the metabolism of UDP-glucose by EUGT11 is independent of the presence of both RebA and Tris in the reaction. This suggests that the detected UDP release is an artifact somehow resulting from the properties of EUGT11, or alternatively, that EUGT11 is capable of hydrolyzing UDP-glucose. EUGT11 remains efficient in preferentially converting RebA to RebD, and the loss of UDP-G can be compensated for by the addition of a sucrose synthase recycling system, as described below.

[0218] RebA solubility The solubility of RebA determines the concentration that can be used for both whole-cell and cell-free approaches. Several different aqueous solutions of RebA were prepared and left at room temperature for several days. After 24 hours of storage, RebA precipitated at concentrations of 50 mM and above. 25 mM RebA began to precipitate after 4–5 days, while concentrations below 10 mM remained in solution even when stored at 4°C. RebD solubility The solubility of RebD was evaluated by making several different aqueous solutions of RebD and incubating them at 30°C for 72 hours. RebD was initially found to be soluble in water at concentrations of 1 mM or less, while concentrations of 0.5 mM or less were found to be stable for longer periods. Those skilled in the art will recognize that solubility can be affected by any number of conditions, such as pH, temperature, or different matrices.

[0219] sucrose synthase Sucrose synthase (SUS) is used to regenerate UDP-glucose from UDP and sucrose (Figure 11) for the glycosylation of other small molecules (Masada Sayaka et al. FEBS Letters 581 (2007) 2562-2566). Three SUS1 genes from A. thaliana, S. rebaudiana, and coffee (Coffea arabica) were cloned into the pGEX4T1 E. coli expression vector (see Figure 17 for sequences). Using a method similar to that described for EUGT11, approximately 0.8 mg / L of GST-AtSUS1 (A. thaliana SUS1) was purified. Initial expression of CaSUS1 (Coffea arabica SUS1) and SrSUS1 (S. rebaudiana SUS1) followed by GST purification did not produce significant amounts of protein, but Western blot analysis confirmed the presence of GST-SrSUS1. When GST-SrSUS1 was expressed at 20°C in the presence of 2% ethanol, approximately 50 μg / L of the enzyme was produced.

[0220] Experiments were performed to evaluate the UDP-glucose regenerating activity of purified GST-AtSUS1 and GST-SrSUS1. In vitro assays were performed in 100 mM Tris-HCl, pH 7.5, and 1 mM UDP (final concentration). ∼2.4 μg of purified GST-AtSUS1, ∼0.15 μg of GST-SrSUS1, or ∼1.5 μg of commercially available BSA (New England Biolabs, Ipswich, MA) was also added. Reactions were performed in the presence or absence of ∼200 mM sucrose and incubated at 37°C for 24 h. Product UDP-glucose was measured by HPLC as described in the analytical section. AtSUS1 produced ∼0.8 mM UDP-glucose in the presence of sucrose. No UDP-glucose was observed with SrSUS1 or the negative control (BSA). The lack of activity observed with SrSUS1 may be explained by the low quality and concentration of the purified enzyme. It was concluded that the production of UDP-glucose by AtSUS1 is sucrose dependent, and therefore AtSUS1 can be used in coupled reactions to regenerate UDP-glucose for use by EUGT1 or other UGTs for the glycosylation of small molecules (Figure 11, above). As shown in Figure 11, AtSUS catalyzes the formation of UDP-glucose and fructose from sucrose and UDP. This UDP-glucose can then be used by EUGT11 to glycosylate RebA to produce RebD. In vitro assays were performed as described above, adding ~200 mM sucrose, 1 mM UDP, 100 μM RebA, ~1.6 μg of purified GST-AtSUS1, and ~0.8 μg of GST-EUGT11. The formation of the product, RebD, was assessed by LC-MS. When AtSUS, EUGT11, sucrose, and UDP were mixed with RebA, 81 ± 5 μM RebD was formed. The reaction was dependent on the presence of AtSUS, EUGT11, and sucrose. The conversion rate was similar to that observed previously using exogenously provided UDP-glucose. This indicates that AtSUS can be used to regenerate UDP-glucose for EUGT11-mediated RebD formation.

[0221] Example 9: Whole-cell biocatalytic production of RebD In this example, several parameters were considered to support the use of whole-cell biocatalytic systems to produce RebD from RebA or other steviol glycosides. The ability of the feedstock to cross the cell membrane and the availability of UDP-glucose are two such factors. Permeabilization agents and different cell types were explored to determine which system would be most beneficial for RebD production.

[0222] Permeabilizing Agents Several different permeabilization agents have previously been shown to enable intracellular enzymatic conversion of a variety of compounds that are normally unable to cross the cell membrane (Chow and Palecek, Biotechnol Prog. 2004 Mar-Apr;20(2):449-56). In some cases, the approach resembles partial lysis of the cells, which in yeast often relies on the removal of the cell membrane with detergents and the encapsulation of enzymes within the remaining cell wall, which is permeable to small molecules. Common to these methods is exposure to a permeabilization agent followed by pelleting of the cells by a centrifugation step before the addition of substrate. For example, regarding yeast permeability, see Flores et al., Enzyme Microb. Technol., 16, pp. 340-346 (1994);Presecki & Vasic-Racki, Biotechnology Letters, 27, pp. 1835-1839 (2005);Yu et al., J Ind Microbiol Biotechnol, 34, 151-156 (2007);Chow and Palecek, Cells. Biotechol. Prog., 20, pp. 449-456 (2004);Fernandez et al., Journal of Bacteriology, 152, pp. 1255-1264 (1982);Kondo et al., Enzyme and microbial technology, 27, pp. 806-811 (2000);Abraham and Bhat, J. See Ind Microbiol Biotechnol, 35, pp. 799-804 (2008); Liu et al.,: Journal of bioscience and bioengineering, 89, pp. 554-558 (2000); and Gietz and Schiestl, Nature Protocols, 2, pp. 31-34 (2007).For bacterial permeabilization, see Naglak and Wang, Biotechnology and Bioengineering, 39, pp. 732-740 (1991); Alakomi et al., Applied and environmental Microbiology, 66, pp. 2001-2005 (2000); and Fowler and Zabin, Journal of bacteriology, 92, pp. 353-357 (1966). As described in this example, it was determined that de novo UDP-glucose biosynthesis could be maintained if the cells remained viable.

[0223] Experiments were performed to establish permeabilization conditions in E. coli and yeast. Growing cells (S. cerevisiae or E. coli) were treated with different concentrations / combinations of permeabilizing agents: toluene, chloroform, and ethanol for S. cerevisiae, and guanidine, lactate, DMSO, and / or Triton X-100 for E. coli. The tolerance of both model organisms to high concentrations of RebA and other potential substrates was also evaluated. Permeability was measured by the amount of RebD produced by EUGT11-expressing organisms after incubation in RebA-containing medium (feeding experiment). Enzyme activity was monitored by lysing cells before and after exposure to the permeabilizing agents and analyzing the activity of the released UGTs in an in vitro assay. In yeast, none of the permeabilization conditions tested resulted in an increase in RebD above the detected background (i.e., impurity RebD levels present in the RebA stock used for feeding), indicating that under the conditions tested, yeast cells remain impermeable to RebA and / or the solvent-induced reduction in cell viability also results in a reduction in EUGT11 activity.

[0224] In E. coli, none of the conditions tested resulted in cell permeabilization and subsequent production of RebD above background levels. Detectable levels of RebD were measured when lysates from strains expressing EUGT11 were used in in vitro reactions (data not shown), indicating that the EUGT11 enzyme was present and active after all permeabilization treatments (although activity levels varied). Permeabilization treatments had little or no effect on cell viability, except for cultures treated with 0.2 M guanidine and 0.5% Triton X-100, which significantly reduced viability. S. cerevisiae was also subjected to permeabilization assays using Triton X-100, N-lauryl sarcosine (LS), or lithium acetate plus polyethylene glycol (LiAc + PEG), which do not allow further cell growth. Under these conditions, cells are rendered non-viable by permeabilization, which completely removes the plasma membrane while retaining the cell wall as a barrier to keep enzymes and gDNA inside. In this method, UDP-glucose can be supplemented or recycled as described above. The advantage of permeabilization over purely in vitro approaches is that it eliminates the need to individually produce and isolate individual enzymes.

[0225] N-laurylsarcosine treatment resulted in the inactivation of EUGT11, and only a slight increase in RebD was detected when LiAc / PEG was applied (data not shown). Treatment with 0.3% or 0.5% Triton X-100, however, increased the amount of RebD above background levels while maintaining EUGT11 activity (see Figure 18). For the Triton X-100 assay, overnight cultures were washed three times with PBS buffer. An OD of 6 units was obtained. 600 The corresponding cells were resuspended in PBS containing 0.3% or 0.5% Triton X-100, respectively. The treated cells were vortexed and incubated at 30°C for 30 minutes. After treatment, the cells were washed in PBS buffer. Five units of OD were obtained as described for GST-EUGT11. 600Cells corresponding to 0.6 units of OD were used in an in vitro assay. 600 were resuspended in reaction buffer as described for the LS-treated samples and incubated overnight at 30° C. An untreated sample was used as a control. Lysates from transformants expressing EUGT11 were able to convert some RebA to RebD (measured at 8–50 μM in the reaction) when cells were untreated or after treatment with LiAC / PEG or Triton X-100. However, no RebD was measured in the lysates of cell pellets treated with LS. Permeabilized but unlysed cells were able to produce some RebD (measured at 1.4–1.5 μM) when treated with 0.3% or 0.5% Triton X-100 (Figure 18), whereas no RebD was found in samples treated with LS or LiAC / PEG. These results demonstrate that RebD can be biocatalytically produced from RebA when whole cells are used and Triton X-100 is used as the permeabilizing agent.

[0226] Example 10: Assessment of codon-optimized UGT sequences Optimal coding sequences for UGTs 91d2e, 74G1, 76G1, and 85C2 were designed and synthesized for yeast expression using two methodologies provided by GeneArt (Regensburg, Germany) (SEQ ID NOS: 6, 2, 8, and 4, respectively) or DNA2.0 (Menlo Park, CA) (SEQ ID NOS: 84, 83, 85, and 82, respectively). The amino acid sequences of UGTs 91d2e, 74G1, 76G1, and 85C2 (SEQ ID NOS: 5, 1, 7, and 3, respectively) were not changed. Wild-type, DNA2.0, and GeneArt sequences were assayed for in vitro activity to compare reactivity with substrates in the steviol glycoside pathway. UGTs were inserted into high-copy (2µ) vectors and expressed from a strong constitutive promoter (GPD1) (vectors P423-GPD, P424-GPD, P425-GPD, and P426-GPD). Plasmids were individually transformed into the universal Watchmaker strain EFSC301 (described in Example 3 of WO 2011 / 153378), and assays were performed using cell lysates prepared from equal amounts of cells (8 OD units). For the enzymatic reactions, 6 μL of each cell lysate was incubated in a 30 μL reaction with 0.25 mM steviol (final concentration) to test UGT74G1 and UGT85C2 clones, and with 0.25 mM 13-SMG (13SMG) (final concentration) to test 76G1UGT and 91D2eUGT. Assays were performed at 30° C. for 24 hours. Prior to LC-MS analysis, 1 volume of 100% DMSO was added to each reaction, samples were centrifuged at 16,000 g, and the supernatants were analyzed.

[0227] Lysates expressing the GeneArt-optimized genes provided high levels of UGT activity under the conditions tested. When expressed as a percentage of wild-type yeast, GeneArt lysates showed activity equivalent to wild-type for UGT74G1, 170% for UGT76G1, 340% for UGT85C2, and 130% for UGT91D2e. When expressed in S. cerevisiae, UGT85C2 can improve the overall flux and productivity of cells for RebA and RebD production. Further experiments were conducted to determine whether codon-optimized UGT85C2 could reduce 19-SMG accumulation and increase the production of rubusoside and highly glycosylated steviol glycosides. The production of 19-SMG and rubusoside was analyzed in steviol-feeding experiments using S. cerevisiae strain BY4741 expressing wild-type UGT74G1 and codon-optimized UGT85C2 from a high-copy (2μ) vector under a strong constitutive promoter (GPD1) (vectors P426-GPD and P423-GPD, respectively). For total glycoside levels, whole culture samples (without cell removal) were taken and boiled in an equal volume of DMSO. The reported intracellular concentrations were obtained by pelleting cells and resuspending them in 50% DMSO relative to the volume of the original culture sample, followed by boiling. "Total" glycoside levels and normalized intracellular levels were then measured using LC-MS. Using wild-type UGT74G1 and wild-type UGT85C2, a total of approximately 13.5 μM of rubusoside was produced, with a maximum normalized intracellular concentration of approximately 7 μM. In contrast, using wild-type UGT74G1 and codon-optimized UGT85C2, a maximum of 26 μM of rubusoside was produced, or approximately twice that produced using wild-type UGT85C2. Furthermore, the maximum normalized intracellular concentration of rubusoside was 13 μM, again approximately twice that produced using wild-type UGT85C2. The intracellular concentration of 19-SMG was significantly reduced from a maximum of 35 μM using wild-type UGT85C2 to 19 μM using codon-optimized UGT85C2. As a result, approximately 10 μM less total 19-SMG was measured with codon-optimized UGT85C2. This indicates that more 19-SMG is converted to rubusoside and confirms that wild-type UGT85C2 is the bottleneck.

[0228] During diversity screening, another UGT85C2 homolog was discovered during cDNA cloning in Stevia rebaudiana. This homolog has the following combination of conserved amino acid polymorphisms (amino acid numbering for wild-type S. rebaudiana UGT85C, coding sequence set forth in accession number AY345978.1): A65S, E71Q, T270M, Q289H, and A389V. This clone, designated UGT85C2 D37, was expressed through in vitro transcription / translation of PCR products (TNT® T7 Quick for PCR DNA kit, Promega). The expression product was assayed for glycosylation activity as described in WO / 2011 / 153378, except that steviol (0.5 mM) was used as the glycosyl acceptor and the assay was incubated for 24 hours. Compared to the wild-type UGT85C2 control assay, the D37 enzyme appears to have approximately 30% higher glycosylation activity.

[0229] Example 11: Identification of a novel S. rebaudiana KAH A partial sequence (GenBank accession number BG521726) was identified in the Stevia rebaudiana EST database with some homology to Stevia KAH. The partial sequence was blasted against raw Stevia rebaudiana pyrosequencing reads using CLC Main Workbench software. Reads that overlapped the ends of the partial sequence were identified and used to increase the length of the partial sequence. This was done several times until the sequence encompassed both the start and stop codons. The complete sequence was analyzed for possible nucleotide substitutions and frameshift mutations by blasting the complete sequence against the raw pyrosequencing reads. The resulting sequence was designated SrKAHe1. See Figure 12. The activity of the KAH encoded by SrKAHe1 was evaluated in vivo in the S. cerevisiae background strain CEN.PK 111-61A, which expresses genes encoding the enzymes that comprise the entire biosynthetic pathway from the yeast secondary metabolites isopentenyl pyrophosphate (IPP) and farnesyl pyrophosphate (FPP) to steviol-19-O-monoside, except for steviol synthase, which converts ent-kaurenoic acid to steviol.

[0230] Briefly, S. cerevisiae strain CEN.PK 111-61A was modified to express the following from chromosomally integrated gene copies, with transcription driven by the TPI1 and GPD1 yeast promoters: Aspergillus nidulans GGPPS, a 150 nt truncated Zea mays CDPS (with a new start codon; see below), S. rebaudiana KA, S. rebaudiana KO, and Stevia rebaudiana UGT74G1. The CEN.PK 111-61A yeast strain expressing all these genes was designated EFSC2386. Thus, strain EFSC2386 contains the following integrated genes: Aspergillus nidulans geranylgeranyl pyrophosphate synthase (GGPPS); Zea mays ent-copalyl diphosphate synthase (CDPS); Stevia rebaudiana ent-kaurene synthase (KS); Stevia rebaudiana ent-kaurene oxidase (KO); and Stevia rebaudiana UGT74G1, combined with the pathway from IPP and FPP to steviol-19-O-monoside, but without steviol synthase (KAH). Expression of different steviol synthases (from episomal expression plasmids) was tested in combination with expression of various CPRs (from episomal expression plasmids) in strain EFSC2386, and production of steviol-19-O-monoside was detected by LC-MS analysis of culture sample extracts. Nucleic acids encoding CPRs were inserted into the multiple cloning site of the p426 GPD basic plasmid, while nucleic acids encoding steviol synthases were inserted into the multiple cloning site of the P415 TEF basic plasmid (p4XX basic plasmid series by Mumberg et al., Gene 156 (1995), 119-122). Production of steviol-19-O-monoside occurs in the presence of functional steviol synthase.

[0231] The KAHs expressed from episomal expression plasmids in strain EFSC2386 were: "indKAH" (Kumar et al., Accession No. DQ398871; Reeja et al., Accession No. EU722415); "KAH1" (S. rebaudiana steviol synthase from Brandle et al., U.S. Patent Publication No. 2008 / 0064063 A1); "KAH3" (A. thaliana steviol synthase from Yamaguchi et al., U.S. Patent Publication No. 2008 / 0271205 A1); "SrKAHe1" (S. rebaudiana steviol synthase cloned from S. rebaudiana acDNA, as described above); and "DNA2.0.SrKAHe1" (codon-optimized sequence (DNA2.0) encoding S. rebaudiana steviol synthase, see Figure 12B). The CPRs expressed from episomal expression plasmids in strain EFSC2386 were: "CPR1" (S. rebaudiana NADPH-dependent P450 reductase (Kumar et al., accession number DQ269454); "ATR1" (A. thaliana CPR, accession number CAA23011, see also Figure 13); "ATR2" (A. thaliana CPR, accession number CAA46815, see also Figure 13); "CPR7" (S. rebaudiana CPR, see also Figure 13; CPR7 is similar to "CPR1"); "CPR8" (S. rebaudiana CPR, similar to Artemisia annua CPR, see Figure 13), and "CPR4" (S. cerevisiae NCP1 (accession number YHR042W, see also Figure 13). Table 13 shows the levels of steviol-19-O-monoside (μM) in strain EFSC2386 with various combinations of steviol synthase and CPR.

[0232] [Table 16-1] [Table 16-2]

[0233] When expressed in S. cerevisiae, only the steviol synthases encoded by KAH3 and SrKAHe1 were active. The DNA2.0 codon-optimized SrKAHe1 sequence encoding steviol synthase resulted in approximately one order of magnitude higher accumulation levels of steviol-19-O-monoside compared to the codon-optimized KAH3 when coexpressed with each of the optimized CPRs. In the experiments presented in this example, the combination of KAH1 and ATR2 CPRs did not result in the production of steviol-19-O-monoside.

[0234] Example 12 - Pairing of CPR and KO The CEN.PK S. cerevisiae strain EFSC2386 and CPR referred to in this example are described in Example 11 ("Identification of S. rebaudiana KAH"). EFSC2386 contains the following integrated genes: Aspergillus nidulans geranylgeranyl pyrophosphate synthase (GGPPS); Zea maysent-copalyl diphosphate synthase (CDPS); Stevia rebaudianaent-kaurene synthase (KS); and Stevia rebaudianaent-kaurene oxidase (KO). This strain produces ent-kaurenoic acid, as detected by LC-MS analysis. A collection of cytochrome P450 reductases (CPRs) was expressed in strain EFSC2386 and tested: "CPR1" (S. rebaudiana NADPH-dependent cytochrome P450 reductase, Kumar et al., accession number DQ269454); "ATR1" (A. thaliana CPR, accession number CAA23011), "ATR2" (A. thaliana CPR, accession number CAA46815), "CPR7" (S. rebaudiana CPR, CPR7 is similar to "CPR1"), "CPR8" (S. rebaudiana CPR, similar to Artemisia annua CPR; and "CPR4" (S. cerevisiae NCP1, accession number YHR042W).

[0235] Overexpression of the endogenous native CPR of S. cerevisiae (referred to as CPR4 in Table 14), and in particular overexpression of one of the A. thaliana CPRs, namely ATR2, provides good activation of the Stevia rebaudiana kaurene oxidase (the latter referred to as KO1 in Table 14), resulting in increased accumulation of ent-kaurenoic acid. See Table 14, which shows the area under the curve (AUC) of the ent-kaurenoic acid peak in the LC-MS chromatograms. KO1 is an ent-kaurenoic acid-producing yeast control strain without additional overexpression of a CPR. [Table 17]

[0236] Example 13 - Evaluation of KS-5 and KS-1 in the steviol pathway Yeast strain EFSC1972 is an S. cerevisiae strain of CEN.PK 111-61A that harbors a biosynthetic pathway from IPP / FPP to rubusoside, expressed by integrated gene copies encoding: Aspergillus nidulans GGPPS (internal name GGPPS-10), Stevia rebaudiana KS (KS1, SEQ ID NO: 133), Arabidopsis thaliana KAH (KAH-3, SEQ ID NO: 144), Stevia rebaudiana KO (KO1, SEQ ID NO: 138), Stevia rebaudiana CPR (CPR-1, SEQ ID NO: 147), full-length Zea mays CDPS (CDPS-5, SEQ ID NO: 158), Stevia rebaudiana UGT74G1 (SEQ ID NO: 1), and Stevia rebaudiana UGT85C2 (SEQ ID NO: 3). Furthermore, EFSC1972 has down-regulation of ERG9 gene expression due to replacement of the endogenous promoter with the copper-inducible promoter CUP1.

[0237] When EFSC1972 was transformed with a CEN / ARS-based plasmid expressing Stevia rebaudiana SrKAHe1 from the TEF1 promoter and cotransformed with a 2µ-based plasmid expressing a truncated version of Synechococcus species GGPPS (GGPPS-7) and Zea mays CDPS (truncated CDPS-5) from the GPD promoter, the result was a growth-impaired S. cerevisiae producer of rubusoside (and 19-SMG). This strain is referred to as "enhanced EFSC1972" in the following text. To determine whether the slow growth rate was caused by the accumulation of the toxic pathway intermediate ent-copalyl diphosphate, a collection of kaurene synthase (KS) genes was expressed in the "enhanced EFSC1972" strain, and subsequent growth and steviol glycoside production were assessed. Expression of the A. thaliana KS (KS5) resulted in improved growth and steviol glycoside production in the "enhanced EFSC1972" strain. See Figure 16. A similar positive effect on growth cannot be achieved by additional overexpression of Stevia rebaudiana kaurene synthase (KS-1) in the enhanced EFSC1972.

[0238] Example 14 - Yeast strain EFSC1859 The Saccharomyces cerevisiae strain EFSC1859 contains the coding sequences for GGPPS-10, CDPS-5, KS-1, KO-1, KAH-3, CPR-1, and UGT74G1, which are integrated into the genome and expressed from the strong constitutive GPD1 and TPI promoters (see Table 15). Additionally, the endogenous promoter of the yeast ERG9 gene was replaced with the copper-inducible promoter CUP1 to downregulate the ERG9 squalene synthase. In standard yeast growth medium, the ERG9 gene is transcribed at very low levels due to the low copper concentration in such medium. Reduced ergosterol production in this strain results in an increased amount of isoprene units available for isoprenoid biosynthesis. The EFSC1859 strain also expresses UGT85C2 from a 2 micron multicopy vector using the GPD1 promoter. EFSC1859 produces rubusoside and steviol 19-O-glycosides. Zea mays CDPS DNA was expressed from a 2 micron multicopy plasmid using the GPD promoter in the presence or absence of a chloroplast signal peptide. The nucleotide and amino acid sequences of Zea mays CDPS are shown in Figure 14. The chloroplast signal peptide is encoded by nucleotides 1-150, which corresponds to residues 1-50 of the amino acid sequence.

[0239] [Table 18] The full-length CDPS plasmid in EFSC1859+ maize and the truncated CDPS plasmid in EFSC+ maize were grown in selective yeast medium with 4% glucose. Rubusoside and 19-SMG production were measured by LC-MS to estimate production levels. Removal of the plastid leader sequence did not appear to increase steviol glycoside production compared to the wild-type sequence, demonstrating that the CDPS transit peptide can be removed without causing a loss of steviol glycoside biosynthesis.

[0240] Example 15 - Yeast strain EFSC1923 The Saccharomyces cerevisiae strain CEN.PK 111-61A was modified to produce steviol glycosides by the introduction of steviol glycoside pathway enzymes from various organisms. The modified strain was designated EFSC1923. Strain EFSC1923 contains the following: an Aspergillus nidulans GGPP synthase gene expression cassette integrated into the S. cerevisiae PRP5-YBR238C intergenic region; a Zea mays full-length CDPS and Stevia rebaudiana CPR gene expression cassette integrated into the MPT5-YGL176C intergenic region; a Stevia rebaudiana kaurene synthase and CDPS-1 gene expression cassette integrated into the ECM3-YOR093C intergenic region; an Arabidopsis thaliana KAH and Stevia rebaudiana KO gene expression cassette integrated into the KIN1-INO2 intergenic region; a Stevia rebaudiana UGT74G1 gene expression cassette integrated into the MGA1-YGR250C intergenic region; and a Stevia rebaudiana UGT85C2 gene expression cassette integrated by replacing the ORF of the TRP1 gene (see Table 15). In addition, the endogenous promoter of the yeast ERG9 gene was replaced with the copper-inducible promoter CUP1. Strain EFSC1923 produced approximately 5 μM of the steviol glycoside, steviol 19-O-monoside, on selective yeast medium with 4% glucose.

[0241] Example 16 - Expression of a truncated maize CDPS in yeast strain EFSC1923 The 5'-terminal 150 nucleotides (SEQ ID NO: 157, see Figure 14) of the Zea mays CDP synthase coding sequence shown in Table 15 were deleted, a new translation initiation ATG was inserted into the remainder of the coding sequence, and the truncated sequence was operably linked to the GPD1 promoter in the multicopy plasmid p423GPD of Saccharomyces cerevisiae EFSC1923. Plasmid p423GPD was prepared as described by Mumberg, D et al. Gene , 156:119-122 (1995). EFSC1923 and EFSC1923 plus p423GPD-ZmtCDPS were grown for 96 hours in selective yeast medium containing 4% glucose. Under these conditions, the amount of steviol 19-O-monoside produced by EFSC1923 plus p423GPD-ZmtCDPS (truncated Zea mays CDPS) was approximately 2.5-fold greater than that produced by EFSC1923 without the plasmid. The Arabidopsis thaliana KAH coding sequence from Table 15 was inserted into a multicopy plasmid designated p426GPD under the control of the GPD1 promoter. Plasmid p426GPD was described in Mumberg, D et al. Gene , 156: 119-122 (1995). No significant difference was observed between the amount of steviol 19-O-monoside produced by EFSC1923+p426GPD-AtKAH and EFSC1923 without the plasmid.

[0242] EFSC1923 was transformed with both p423GPD-ZmtCDPS and p426 p426GPD-AtKAH. Surprisingly, under these conditions, the amount of steviol 19-O-monoside produced by EFSC1923 carrying both plasmids (i.e., the truncated Zea mays CDPS and Arabidopsis KAH) was more than six-fold greater than the amount produced by EFSC1923 alone. A bifunctional CDPS-KS from Gibberella fujikuroi (NCBI accession number Q9UVY5.1, FIG. 15 ) was cloned and compared with truncated CDPS-5. The bifunctional Gibberella CDPS-KS was cloned into a 2μ plasmid with a GPD promoter and transformed into EFSC1923 with a plasmid expressing Arabidopsis thaliana KAH-3 from a 2μ-based plasmid driven by a GPD promoter. In shake flask experiments, this bifunctional CDPS-KS was approximately 5.8-fold more active in producing steviol 19-O-monoside than strain EFSC1923 carrying only KAH-3. However, under the conditions tested, it was found to be less optimal than the combination of KAH-3 and truncated CDPS. Therefore, additional strains were constructed using KS-5 and truncated CDPS.

[0243] Example 17 - Toxicity of intermediates The effects of geranylgeranyl pyrophosphate (GGPP), ent-copalyl diphosphate (CDP), or ent-kaurene production on S. cerevisiae vitality were examined in the laboratory S. cerevisiae strain CEN.PK background by expressing Synechococcus sp. GGPP alone (GGPP production), GGPP together with a 50-amino acid N-terminally truncated Zea mays CDPS (see Example 16) (CDP production), or GGPP, truncated CDPS, and Arabidopsis thaliana kaurene synthase (KS5) together (ent-kaurene production). The genes were expressed from a 2μ plasmid, with the GPD promoter driving transcription of the truncated CDPS and KS5, while transcription of GGPPS was driven by the ADH1 promoter. Growth of S. cerevisiae CEN.PK transformed with various combinations of these plasmids (GGPP alone; GGPP + truncated CDPS; or GGPP + truncated CDPS + KS5) or with no gene insert was observed. GGPP production, and especially CDP production, was toxic to S. cerevisiae when produced as an end product. Interestingly, ent-kaurene did not appear to be toxic to yeast at the levels produced in this experiment.

[0244] Example 18 - Disruption of endogenous phosphatase activity The yeast genes DPP1 and LPP1 encode phosphatases that degrade FPP and GGPP to farnesol and geranylgeraniol, respectively. The DPP1-encoding gene was deleted in strain EFSC1923 (described in Example 15) to determine whether this affected steviol glycoside production. When this dpp1 mutant strain was further transformed with a plasmid expressing Zea mays CDPS lacking a chloroplast transport sequence (Example 16), both large and small transformants emerged. The "large colony" type strain produced ~40% more 19-SMG under the test conditions compared to the "small colony" type strain and the DPP1-free strain. These results indicate that deletion of DPP1 can have a positive effect on steviol glycoside production, and therefore, prenyl pyrophosphate degradation in yeast may negatively impact steviol glycoside production.

[0245] Example 19 - Construction of a genetically stable yeast reporter strain producing vanillin glucoside from glucose with a disrupted SUC2 gene A yeast strain producing vanillin glucoside from glucose was generated essentially as described in Brochado et al. ((2010) Microbial Cell Factories 9:84-98) (strain VG4), except that an expression cassette containing the E. coli EntD PPTase controlled by the yeast TPI1 promoter was additionally integrated into the ECM3 interlocus region of the yeast genome (as described in Hansen et al. (2009) Appl. Environ. Microbiol. 75(9):2765-2774), SUC2 was disrupted by replacing the coding sequence with a MET15 expression cassette, and LEU2 was disrupted by replacing the coding sequence with a Tn5ble expression cassette, conferring resistance to phleomycin. The resulting yeast strain is designated V28. This strain also encodes a recombinant A. thaliana UDP-glycosyltransferase (UGT72E2, GenBank accession number Q9LVR1) having the amino acid sequence shown in Figure 19 (SEQ ID NO: 178).

[0246] Example 20 - Expression of sucrose transporter and sucrose synthase in yeast already biosynthesizing vanillin glucoside The sucrose transporter SUC1 from Arabidopsis thaliana was isolated by PCR amplification from cDNA prepared from A. thaliana using a proofreading PCR polymerase. The resulting PCR fragment was transferred by restriction digestion with SpeI and EcoRI and inserted into the corresponding position of the low-copy-number yeast expression vector P416-TEF (a CEN-ARS-based vector), which allows gene expression from the strong TEF promoter. The resulting plasmid was named pVAN192. The sequence of the encoded sucrose transporter is shown in Figure 19B (GenBank accession number AEE35247, SEQ ID NO: 179). The sucrose synthase SUS1 from Coffea arabica (accession number CAJ32596) was isolated by PCR amplification from cDNA prepared from C. arabica using a proofreading PCR polymerase. The PCR fragment was transferred by restriction digestion with SpeI and SalI and inserted into the corresponding position of the high-copy-number yeast expression vector p425-GPD (a 2µm-based vector), which allows gene expression from the strong GPD promoter. The resulting plasmid was designated pMUS55. The sequence of the encoded sucrose synthase is shown in Figure 19C (GenBank accession number CAJ32596; SEQ ID NO: 180).

[0247] pVAN192 and pMUS55 were genetically transformed into yeast strain V28 using a lithium acetate ...

Claims

1. An in vitro method for producing a target steviol glycoside or a target steviol glycoside composition, comprising adding to a reaction mixture a precursor steviol glycoside having 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose, and / or a mixture thereof, a first recombinant polypeptide capable of beta-1,2 glycosylation at the C2' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside, and one or more UDP-glucoses, thereby producing the target steviol glycoside or the target steviol glycoside composition; The first recombinant polypeptide has at least 90% identity to the amino acid sequence set forth in SEQ ID NO: 152; The method.

2. An in vitro method for producing a target steviol glycoside or a target steviol glycoside composition, comprising: precursor steviol glycosides having 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose, and / or mixtures thereof; A first recombinant polypeptide capable of beta-1,2 glycosylation at the C2' of a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose of a precursor steviol glycoside, wherein the first recombinant polypeptide has at least 90% identity to the amino acid sequence set forth in SEQ ID NO:152; and Polypeptides (a) to (e): (a) a polypeptide capable of rhamnosylation of steviol 13-O-monoside; (b) a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-19 carboxyl group; (c) a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group; (d) a polypeptide capable of beta-1,3 glycosylation at the C3' of a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose of a precursor steviol glycoside; and / or (e) a second polypeptide capable of beta-1,2 glycosylation at the C2' of the 13-O-glucose, the 19-O-glucose, or both the 13-O-glucose and the 19-O-glucose of the precursor steviol glycoside. one or more of: wherein at least one of the polypeptides (a) to (e) is a recombinant polypeptide; and one or more UDP-sugars, the one or more UDP-sugars comprising UDP-glucose, UDP-rhamnose, and / or UDP-xylose; to a reaction mixture, thereby producing a target steviol glycoside or a target steviol glycoside composition. The method.

3. 3. The method of claim 2, (a) a polypeptide capable of rhamnosylation of steviol-13-O-monoside comprises a polypeptide having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 150; (b) polypeptides capable of glycosylating steviol or a precursor steviol glycoside at its C-19 carboxyl group include polypeptides having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO:1; (c) a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group, (i) a polypeptide having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO:3; or (ii) a polypeptide having one or more amino acid substitutions of K9E, K10R, V13F, F15L, Q21H, M27V, H60D, A65S, E71Q, I87F, L91P, K220T, R243W, T270M, T270R, Q289H, Y298C, L334S, K350T, H368R, A389V, I394V, P397S, E418V, G420R, L431P, G440D, H441N, R444G, and M471T in the amino acid sequence set forth in SEQ ID NO:

3. Includes; (d) a polypeptide capable of beta-1,3 glycosylation at the C3' of a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose of a precursor steviol glycoside, (i) a polypeptide having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 7; or (ii) a polypeptide having one or more amino acid substitutions of M29I, V74E, V87G, L91P, G116E, A123T, Q125A, I126L, T130A, V145M, C192S, S193A, F194Y, M196N, K198Q, K199I, Y200L, Y203I, F204L, E205G, N206K, I207M, T208I, P266Q, S273P, R274S, G284T, T285S, L330V, G331A, and L346I in the amino acid sequence set forth in SEQ ID NO:

7. and / or (e) a second polypeptide capable of beta-1,2 glycosylation at the C2' of a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose of a precursor steviol glycoside; (i) a polypeptide having at least 90% sequence identity with the amino acid sequence set forth in SEQ ID NO:5; (ii) a polypeptide having an arginine at residue 206, a cysteine ​​at residue 207, and an arginine at residue 343 in the amino acid sequence set forth in SEQ ID NO:5; (iii) in the amino acid sequence set forth in SEQ ID NO: 5, tyrosine or phenylalanine at residue 30, proline or glutamine at residue 93, serine or valine at residue 99, tyrosine or phenylalanine at residue 122, histidine or tyrosine at residue 140, serine or cysteine ​​at residue 142, alanine or threonine at residue 148, methionine at residue 152, alanine at residue 153, alanine or serine at residue 156, and glutamine at residue 162 a polypeptide having a lysine, a leucine or methionine at residue 195, a glutamic acid at residue 196, a lysine or glutamic acid at residue 199, a leucine or methionine at residue 211, a leucine at residue 213, a serine or phenylalanine at residue 221, a valine or isoleucine at residue 253, a valine or alanine at residue 286, a lysine or asparagine at residue 427, an alanine at residue 438, and an alanine or threonine at residue 462; (iv) a polypeptide having a methionine at residue 211 and an alanine at residue 286 in the amino acid sequence set forth in SEQ ID NO:5; or (v) a polypeptide having an amino acid sequence set forth in SEQ ID NO: 10, 12, 76, 78 or 95 Including, The method.

4. 4. The method of claim 2 or 3, wherein the target steviol glycoside comprises steviol-13-O-glucoside, steviol-1,2-bioside, steviol-1,3-bioside, steviol-19-O-glucoside, stevioside, 1,3 stevioside, rubusoside, rebaudioside A, rebaudioside B, rebaudioside C, rebaudioside D, rebaudioside E, rebaudioside F, dulcoside A, steviol monoglucoside, steviol rhamnoside, or steviol xyloside.

5. 5. The method of any one of claims 2-4, wherein the target steviol glycoside composition comprises steviol-13-O-glucoside, steviol-1,2-bioside, steviol-1,3-bioside, steviol-19-O-glucoside, stevioside, 1,3 stevioside, rubusoside, rebaudioside A, rebaudioside B, rebaudioside C, rebaudioside D, rebaudioside E, rebaudioside F, dulcoside A, steviol monoglucosides, steviol rhamnosides, and / or steviol xylosides, or combinations thereof.

6. 3. The method of claim 2, (a) the precursor steviol glycoside is steviol-13-O-glucoside, steviol-19-O-glucoside, rubusoside, steviol-1,2-bioside, stevioside, or rebaudioside B, and / or a mixture thereof; one or more UDP-sugars is UDP-glucose; Steviol, precursor steviol glycosides, and / or mixtures thereof are A polypeptide capable of beta 1,2 glycosylation at the C2' of a precursor steviol glycoside, a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 152; and one or more of a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group, a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-19 carboxyl group, a second polypeptide capable of beta-1,2 glycosylation of the C2' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside, and a polypeptide capable of beta-1,3 glycosylation of the C3' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside. has been contacted with; and Rebaudioside A is produced upon the transfer of one or more sugar moieties from one or more UDP-glucoses to steviol, precursor steviol glycosides, and / or mixtures thereof; or (b) the precursor steviol glycoside is steviol-13-O-glucoside or steviol-1,2-bioside, and / or a mixture thereof; one or more UDP-sugars is UDP-glucose; Steviol, precursor steviol glycosides, and / or mixtures thereof are A polypeptide capable of beta 1,2 glycosylation at the C2' of a precursor steviol glycoside, a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 152; and one or more of a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group, a second polypeptide capable of beta-1,2 glycosylation of the C2' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside, and a polypeptide capable of beta-1,3 glycosylation of the C3' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside. has been contacted with; and Rebaudioside B is produced upon the transfer of one or more sugar moieties from one or more UDP-glucoses to steviol, precursor steviol glycosides, and / or mixtures thereof; or (c) the precursor steviol glycoside is steviol-13-O-glucoside, steviol-19-O-glucoside, rubusoside, steviol-1,2-bioside, or stevioside, and / or mixtures thereof; one or more UDP-sugars is UDP-glucose; Steviol, precursor steviol glycosides, and / or mixtures thereof are A polypeptide capable of beta 1,2 glycosylation at the C2' of a precursor steviol glycoside, a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 152; and one or more of a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group, a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-19 carboxyl group, a second polypeptide capable of beta-1,2 glycosylation of the C2' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside, and a polypeptide capable of beta-1,3 glycosylation of the C3' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside. has been contacted with; and Rebaudioside E is produced upon the transfer of one or more sugar moieties from one or more UDP-glucoses to steviol, precursor steviol glycosides, and / or mixtures thereof; or (d) the precursor steviol glycoside is steviol-13-O-glucoside, steviol-19-O-glucoside, rubusoside, steviol-1,2-bioside, stevioside, rebaudioside A, rebaudioside B, or rebaudioside E, and / or a mixture thereof; one or more UDP-sugars is UDP-glucose; Steviol, precursor steviol glycosides, and / or mixtures thereof are A polypeptide capable of beta 1,2 glycosylation at the C2' of a precursor steviol glycoside, a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 152; and one or more of a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group, a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-19 carboxyl group, a second polypeptide capable of beta-1,2 glycosylation of the C2' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside, and a polypeptide capable of beta-1,3 glycosylation of the C3' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside. has been contacted with; and Rebaudioside D is produced upon the transfer of one or more sugar moieties from one or more UDP-glucoses to steviol, precursor steviol glycosides, and / or mixtures thereof; or (e) the precursor steviol glycoside is steviol-13-O-glucoside, steviol-19-O-glucoside, rubusoside, or steviol-1,2-rhamnobioside, and / or mixtures thereof; the one or more UDP-sugars are UDP-glucose and UDP-rhamnose; Steviol, precursor steviol glycosides, and / or mixtures thereof are A polypeptide capable of beta 1,2 glycosylation at the C2' of a precursor steviol glycoside, a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 152; and one or more of a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group, a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-19 carboxyl group, a second polypeptide capable of beta-1,2 glycosylation of the C2' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside, and a polypeptide capable of rhamnosylation of steviol 13-O-monoside. has been contacted with; and Dulcoside A is produced upon the transfer of one or more sugar moieties from one or more of UDP-glucose and UDP-rhamnose to steviol, precursor steviol glycosides, and / or mixtures thereof; or (f) the precursor steviol glycoside is steviol-13-O-glucoside, steviol-19-O-glucoside, rubusoside, steviol-1,2-rhamnobioside, or dulcoside A, and / or mixtures thereof; the one or more UDP-sugars are UDP-glucose and UDP-rhamnose; Steviol, precursor steviol glycosides, and / or mixtures thereof are A polypeptide capable of beta 1,2 glycosylation at the C2' of a precursor steviol glycoside, a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 152; and one or more of a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group, a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-19 carboxyl group, a second polypeptide capable of beta-1,2 glycosylation at the C2' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside, a polypeptide capable of rhamnosylation of steviol 13-O-monoside, and a polypeptide capable of beta-1,3 glycosylation at the C3' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside. has been contacted with; and Rebaudioside C is produced upon the transfer of one or more sugar moieties from one or more of UDP-glucose and UDP-rhamnose to steviol, precursor steviol glycosides, and / or mixtures thereof; or (g) the precursor steviol glycoside is steviol-13-O-glucoside, steviol-19-O-glucoside, rubusoside, 1,2-steviolxyloside, or steviol-1,2-xylobioside, and / or mixtures thereof; the one or more UDP-sugars are UDP-glucose and UDP-xylose; Steviol, precursor steviol glycosides, and / or mixtures thereof are A polypeptide capable of beta 1,2 glycosylation at the C2' of a precursor steviol glycoside, a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 152; and one or more of: a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group; a polypeptide capable of glycosylation of steviol or a precursor steviol glycoside at its C-19 carboxyl group; a second polypeptide capable of beta-1,2 glycosylation of the C2' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside; a polypeptide capable of beta-1,3 glycosylation of the C3' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside; a UDP-glucose dehydrogenase (UGD1) polypeptide; and / or a UDP-glucuronic acid decarboxylase (UXS3) polypeptide. has been contacted with; and Rebaudioside F is produced upon the transfer of one or more sugar moieties from one or more of UDP-glucose and UDP-xylose to steviol, precursor steviol glycosides, and / or mixtures thereof; or (h) the precursor steviol glycoside is steviol-13-O-glucoside; one or more UDP-sugars is UDP-glucose; Steviol, precursor steviol glycosides, and / or mixtures thereof are A polypeptide capable of beta 1,2 glycosylation at the C2' of a precursor steviol glycoside, a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 152; and one or more of: a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group; a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-19 carboxyl group; and a polypeptide capable of beta-1,2 glycosylation of the C2' of a precursor steviol glycoside at 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose. has been contacted with; and Steviol-1,2-bioside is produced upon the transfer of one or more sugar moieties from one or more UDP-glucoses to steviol, precursor steviol glycosides, and / or mixtures thereof; or (i) the precursor steviol glycoside is steviol-13-O-glucoside, steviol-19-O-glucoside, steviol-1,2-bioside, or rubusoside, and / or mixtures thereof; one or more UDP-sugars is UDP-glucose; Steviol, precursor steviol glycosides, and / or mixtures thereof are A polypeptide capable of beta 1,2 glycosylation at the C2' of a precursor steviol glycoside, a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO: 152; and one or more of a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-13 hydroxyl group, a polypeptide capable of glycosylating steviol or a precursor steviol glycoside at its C-19 carboxyl group, a second polypeptide capable of beta-1,2 glycosylation of the C2' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside, and a polypeptide capable of beta-1,3 glycosylation of the C3' of 13-O-glucose, 19-O-glucose, or both 13-O-glucose and 19-O-glucose of a precursor steviol glycoside. has been contacted with; and Stevioside is produced upon the transfer of one or more sugar moieties from one or more UDP-glucoses to steviol, precursor steviol glycosides, and / or mixtures thereof. The method.

7. The method of any one of claims 2 to 6, further comprising adding a cell-free system for the regeneration of one or more UDP-sugars.

8. 8. The method of any one of claims 1 to 7, further comprising isolating the target steviol glycoside or target steviol glycoside composition produced from the reaction mixture, the isolating step comprises separating a liquid phase of the reaction mixture from a solid phase of the reaction mixture to obtain a supernatant containing the target steviol glycoside or target steviol glycoside composition produced; and (a) contacting the supernatant with one or more adsorption resins to obtain at least a portion of the target steviol glycoside or target steviol glycoside composition produced; or (b) contacting the supernatant with one or more ion exchange or reverse phase chromatography columns to obtain at least a portion of the target steviol glycoside or target steviol glycoside composition produced; or (c) crystallizing or extracting the target steviol glycoside or target steviol glycoside composition produced; This allows the target steviol glycoside or target steviol glycoside composition produced to be isolated. The method comprising:

9. 9. The method of any one of claims 1 to 8, further comprising recovering the target steviol glycoside or target steviol glycoside composition produced from the reaction mixture, wherein the recovered target steviol glycoside or target steviol glycoside composition has reduced levels of stevia plant-derived components compared to a steviol glycoside composition obtained from a plant-derived stevia extract.

10. (a) a recombinant polypeptide capable of beta 1,2 glycosylation at the C2' of a 13-O-glucose, a 19-O-glucose, or both a 13-O-glucose and a 19-O-glucose of a precursor steviol glycoside, and having at least 90% sequence identity to the amino acid sequence set forth in SEQ ID NO:152; (b) a target steviol glycoside or a target steviol glycoside composition, wherein the target steviol glycoside comprises steviol-1,2-bioside, stevioside, steviol-19-O-1,2-diglucoside, 19-O-1,2-diglycosylated rubusoside, rebaudioside D, or rebaudioside E; (c) UDP-glucose; and (d) Reaction buffer and / or salts A reaction mixture comprising:

Citation Information

Patent Citations

  • Preparation of beta-1,3-glycosyl stevioside

    JP1983149697A

  • plant fatty acid hydroxylase gene

    JP2001519164A

  • Steviol synthetic enzyme gene and method for producing steviol

    JP2008237110A

  • Sequence-determined DNA fragments and corresponding polypeptides encoded thereby

    US20060150283A1

  • Compositions and methods for producing steviol and steviol glycosides

    US20080064063A1