Host cells for producing triterpene glycosides and precursors and derivatives thereof

Host cells engineered with heterologous polynucleotides enhance the production of triterpene glycosides and derivatives, addressing inefficiencies in existing production methods and facilitating their use in vaccine adjuvants.

WO2026090314A1PCT designated stage Publication Date: 2026-04-30GINKGO BIOWORKS INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GINKGO BIOWORKS INC
Filing Date
2025-10-22
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Production of triterpene glycosides, such as QS-7 and QS-21, is complex and inefficient.

Method used

Development of host cells comprising heterologous polynucleotides encoding polypeptides involved in biosynthetic pathways for triterpene glycosides, including beta-amyrin synthase, cytochrome P450 oxidases, and UDP-sugar transferases, to enhance the production of triterpene glycosides and their derivatives.

Benefits of technology

The host cells efficiently produce triterpene glycosides and derivatives, such as quillaic acid and QS-21, improving the production process and enabling the use of these compounds in vaccine adjuvants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052103_30042026_PF_FP_ABST
    Figure US2025052103_30042026_PF_FP_ABST
Patent Text Reader

Abstract

Described in this application are host cells comprising heterologous polynucleotides encoding enzymes for producing triterpene glycosides (e.g., prosapogenin) and precursors or derivatives thereof. Methods of culturing the host cells are also included herein. Also described are methods of producing a product of interest, such as triterpene glycosides (e.g., prosapogenin) and precursors or derivatives thereof, using the host cells described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] HOST CELLS FOR PRODUCING TRITERPENE GLYCOSIDES AND PRECURSORS AND DERIVATIVES THEREOF RELATED APPLICATION

[0002] This application claims the benefit under 35 U.S.C. § 119(e) of United States Provisional Application Number 63 / 711,104, filed October 23, 2024, entitled HOST CELLS FOR PRODUCING TRITERPENE GLYCOSIDES AND PRECURSORS AND DERIVATIVES THEREOF, the entire disclosure of which is hereby incorporated by reference in its entirety.

[0003] FEDERALLY SPONSORED RESEARCH

[0004] This invention was made with government support under MCDC2205-014 awarded by the Medical CBRN Defense Consortium. The government has certain rights in the invention. As detailed below, aspects of this invention were made without government support.

[0005] REFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0006] The contents of the electronic sequence listing (G091970125WO00-SEQ-AXW.xml; Size: 6,383,552 bytes; and Date of Creation: October 22, 2025) are herein incorporated by reference in their entirety.

[0007] FIELD

[0008] The present disclosure relates to host cells comprising heterologous polynucleotides encoding polypeptides in biosynthetic pathways related to production of triterpene glycosides, such as QS-7 and QS-21 and / or precursors and derivatives of triterpene glycosides, and to the production and uses of the triterpene glycosides, precursors and derivatives.

[0009] BACKGROUND

[0010] Saponins such as QS-7 and QS-21 are triterpene glycoside natural products isolated from soap bark trees (Quillaja saponaria). QS-7 and QS-21 comprise a triterpene core (quillaic acid) which is glycosylated at the triterpene C3 oxygen and C28 carboxylate positions. Saponins are a key component in vaccine adjuvants, which make vaccines more effective, modifying and broadening the immune response, and enabling the use of a lower vaccine dosage. QS-21 is currently used in the shingles and malaria vaccines, Shingrix and Mosquirix, respectively.

[0011] Production of triterpene glycosides, such as saponins, is complex and inefficient.

[0012] SUMMARY

[0013] Aspects of the present disclosure relate, at least in part, to host cells comprising heterologous polynucleotides encoding polypeptides related to producing triterpene glycosides (e.g., prosapogenin).

[0014] In some embodiments, a host cell comprises one or more of:

[0015] a) a heterologous polynucleotide encoding a beta-amyrin synthase (BAS) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 59-314;

[0016] b) a heterologous polynucleotide encoding a cytochrome P450 C16C28 oxidase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315-318;

[0017] c) a heterologous polynucleotide encoding a cytochrome P450 C28 oxidase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315 and 318-320;

[0018] d) a heterologous polynucleotide encoding a cytochrome P450 C23 oxidase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 321-335;

[0019] e) a heterologous polynucleotide encoding a cytochrome P450 reductase (CPR) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 336-377; f) a heterologous polynucleotide encoding a membrane steroid binding protein (MSBP) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ IDNOs: 378-416;

[0020] g) a heterologous polynucleotide encoding a UDP-GlcA transferase (GlcAT) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 417-441;

[0021] h) a heterologous polynucleotide encoding a UDP -galactose transferase (GalT) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 442-537;

[0022] i) a heterologous polynucleotide encoding a UDP -xylose transferase (XylT) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 484, 485, 491, 494, 502, 504, 517, and 538-572;

[0023] j) a heterologous polynucleotide encoding a UDP-glucose dehydrogenase (UGD) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 46-56, and 712-1142;

[0024] k) a heterologous polynucleotide encoding a UDP-xylose synthase (UXS) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 573-711;

[0025] l) a heterologous polynucleotide encoding a phosphotransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1239-1257;

[0026] m) a heterologous polynucleotide encoding a UDP -glycosyltransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1258-1283;

[0027] n) a heterologous polynucleotide encoding a methyltransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1284-1289;

[0028] o) a heterologous polynucleotide encoding a sulfotransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1290-1307;

[0029] p) a heterologous polynucleotide encoding an acyltransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1308-1454;

[0030] q) a heterologous polynucleotide encoding a glycosyl hydrolase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1455-1472;

[0031] r) a heterologous polynucleotide encoding an esterase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1473-1499;

[0032] In some embodiments, the host cell is capable of producing one or more of: 2,3-oxidosqualene; beta-amyrin; oleanolic acid; echinocystic acid; hederagenin; gypsogenin; gypsogenic acid; 16a-OH-hederagenin; quillaic acid (QA); 16a-OH-gypsogenic acid; 3-O-{p-D-glucopyranosiduronic acid}-quillaic acid (QA-C3-GlcA); 3-O-{p-D-galactopyranosyl-(l->2)-p-D-glucopyranosiduronic acid}-quillaic acid (QA-C3-GlcA-Gal); 3-O-{p-D-xylopyranosyl-(l- >3)-[p-D-galactopyranosyl-(l->2)]-p-D-glucopyranosiduronic acid}-quillaic acid (QA-C3-GlcA-Gal-Xyl,QA-TriX, prosapogenin). In various chemical names: X or Xyl is xylose, Gal is galactose, and GlcA is glucuronic acid. In some embodiments, the host cell is a yeast cell, a plant cell, a bacterial cell, or a filamentous fungi cell.

[0033] Aspects of the present disclosure provide methods comprising the step(s) of culturing a host cell under conditions such that it expresses gene(s) encoded by the heterologous polynucleotide(s).

[0034] Aspects of the present disclosure provide host cells comprising one or more heterologous polynucleotides collectively encoding: a) a beta-amyrin synthase (BAS); b) a cytochrome P450 C16C28 oxidase (or separate P450 C16 and C28 oxidases); c) a cytochrome P450 C23 oxidase; d) a cytochrome P450 reductase (CPR); and e) a membrane steroid binding protein (MSBP). In some embodiments, the host cell is capable of producing 2,3-oxidosqualene.

[0035] In some embodiments: a) the beta-amyrin synthase (BAS) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 59-314; b) the cytochrome P450 C16C28 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315-318; c) the cytochrome P450 C28 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315 and 318-320; d) the cytochrome P450 C23 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 321-335; e) the cytochrome P450 reductase (CPR) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 336-377; f) the membrane steroid binding protein (MSBP) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 378-416; or g) any combination of two or more of a) to f).

[0036] In some embodiments: a) the beta-amyrin synthase (BAS) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 59-314; b) the cytochrome P450 C16C28 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315-318; c) the cytochrome P450 C28 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315 and 318-320; d) the cytochrome P450 C23 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 321-335; e) the cytochrome P450 reductase (CPR) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 336-377; or f) the membrane steroid binding protein (MSBP) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 378-416; or g) any combination of two or more of a) to f). In some embodiments, the host cell is a yeast cell, a plant cell, or a filamentous fungi cell.

[0037] In some embodiments, the host cell comprises one or more genetic modifications such that the host cell has increased mevalonate flux relative to a host cell that lacks the one or more genetic modifications. In some embodiments, the one or more genetic modifications comprise a heterologous polynucleotide(s) encoding one or more genes associated with mevalonate flux. In some embodiments, the genes associated with mevalonate flux comprise erg 10, ergl3, hmgl, ergl2, erg8, ergl9, idil, and / or erg20. In some embodiments, the host cell overexpresses erglO, ergl3, hmgl, ergl2, erg8, ergl9, idil, and / or erg20. Aspects of the present disclosure provide methods of producing quillaic acid (QA), the method comprising culturing any one of the host cells of the present disclosure under conditions such that the host cell expresses the beta-amyrin synthase (BAS), the cytochrome P450 C16C28 oxidase, the cytochrome P450 C23 oxidase, the cytochrome P450 reductase (CPR), and the membrane steroid binding protein (MSBP), wherein the host cell is capable of producing 2,3-oxidosqualene.

[0038] Aspects of the present disclosure provide host cells comprising a heterologous polynucleotide encoding a UDP-glucose dehydrogenase. In some embodiments, the host cell is capable of producing UDP-glucose (UDP-Glu). In some embodiments, the UDP-glucose dehydrogenase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 46-56, and 712-1142. In some embodiments, the host cell is a yeast cell, a plant cell, or a filamentous fungi cell.

[0039] Aspects of the present disclosure provide methods of producing UDP-glucuronic acid (UDP-GlcA), the method comprising culturing any of the host cells of the present disclosure under conditions such that the host cell expresses the UDP-glucose dehydrogenase, wherein the host cell is capable of producing UDP-glucose (UDP-Glu).

[0040] Aspects of the present disclosure provide host cells comprising a heterologous polynucleotide encoding a UDP-xylose synthase. In some embodiments, the host cell is capable of producing UDP-glucuronic acid (UDP-GlcA). In some embodiments, the UDP-xylose synthase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 573-711. In some embodiments, the host cell is a yeast cell, a plant cell, a bacterial cell, or a filamentous fungi cell.

[0041] Aspects of the present disclosure provide methods of producing UDP-xylose (UDP-Xyl) in a host cell, the method comprising culturing the any of the host cells of the present disclosure under conditions such that the host cell expresses the UDP-xylose synthase, wherein the host cell is capable of producing UDP-glucuronic acid (UDP-GlcA).

[0042] Aspects of the present disclosure provide host cells comprising one or more heterologous polynucleotides collectively encoding: a) a UDP-GlcA transferase (GlcAT); b) a UDP -galactose transferase (GalT); and c) a UDP -xylose transferase (XylT). In some embodiments, the host cell is capable of producing quillaic acid (QA), UDP-glucuronic acid (UDP-GlcA), UDP-galactose (UDP-Gal), and / or UDP-xylose (UDP-Xyl). In some embodiments, a) the UDP-GlcA transferase (GlcAT) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 417-441; b) the UDP-galactose transferase (GalT) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 442-537; c) the UDP-xylose transferase (XylT) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 484, 485, 491, 494, 502, 504, 517, and 538-572; or d) any combination thereof In some embodiments, the host cell is a yeast cell, a plant cell, or a filamentous fungi cell.

[0043] Aspects of the present disclosure provide methods of producing QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin), the method comprising culturing any of the host cells of the present disclosure under conditions such the host cell expresses the UDP-GlcA transferase (GlcAT), the UDP-galactose transferase (GalT), and the UDP-xylose transferase (XylT), wherein the host cell is capable of producing quillaic acid (QA), UDP-glucuronic acid (UDP-GlcA), UDP-galactose (UDP-Gal), and UDP-xylose (UDP-Xyl).

[0044] Aspects of the present disclosure provide host cells comprising one or more of: a) a heterologous polynucleotide encoding a phosphotransferase; b) a heterologous polynucleotide encoding a UDP -glycosyltransferase; c) a heterologous polynucleotide encoding a methyltransferase; d) a heterologous polynucleotide encoding a sulfotransferase; e) a heterologous polynucleotide encoding an acyltransferase; f) a heterologous polynucleotide encoding a glycosyl hydrolase; g) a heterologous polynucleotide encoding an esterase. In some embodiments, the host cell is capable of producing QS-21. In some embodiments, a) the phosphotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1239-1257; b) the UDP-glycosyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1258-1283; and / or c) the methyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1284-1289; d) the sulfotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1290-1307; e) the acyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1308-1454; f) the glycoside hydrolase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1455-1472; and / or g) the esterase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1473-1499. In some embodiments, the host cell is a yeast cell, a plant cell, a bacterial cell, or a filamentous fungi cell.

[0045] Aspects of the present disclosure provide methods of producing a QS-21 derivative, the method comprising culturing any of the host cell of the present disclosure under conditions such the host cell expresses the phosphotransferase, the UDP-glycosyltransferase, the methyltransferase, the sulfotransferase, the acyltransferase, the glycosyl hydrolase, and / or the esterase, wherein the host cell is capable of producing QS-21.

[0046] Aspects of the present disclosure provide methods of producing a QS-21 derivative, the method comprising contacting QS-21 with a phosphotransferase, a UDP-glycosyltransferase, a methyltransferase, a sulfotransferase, an acyltransferase, a glycosyl hydrolase, and / or an esterase. In some embodiments, a) the phosphotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1239-1257; b) the UDP-glycosyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1258-1283; c) the methyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1284-1289; d) the sulfotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1290-1307; e) the acyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1308-1454; f) the glycosyl hydrolase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1455-1472; and / or g) the esterase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1473-1499. In some embodiments, the phosphotransferase, the UDP-glycosyltransferase, the methyltransferase, the sulfotransferase, the acyltransferase, the glycosyl hydrolase, and / or the esterase is a purified enzyme. In some embodiments, the method comprises contacting QS-21 with a cell lysate, wherein the cell lysate comprises the phosphotransferase, the UDP-glycosyltransferase, the methyltransferase, the sulfotransferase, the acyltransferase, the glycosyl hydrolase, and / or the esterase. In some embodiments, the cell lysate is produced by lysing a population of host cells, wherein the host cells comprise one or more of: a) a heterologous polynucleotide encoding the phosphotransferase; b) a heterologous polynucleotide encoding the UDP-glycosyltransferase; c) a heterologous polynucleotide encoding the methyltransferase; d) a heterologous polynucleotide encoding the sulfotransferase; e) a heterologous polynucleotide encoding the acyltransferase; f) a heterologous polynucleotide encoding the glycosyl hydrolase; and g) a heterologous polynucleotide encoding the esterase.

[0047] Aspects of the present disclosure provide adjuvants comprising a triterpene glycoside produced by a host cell described herein. In some embodiments, the triterpene glycoside is quillaic acid (QA), 16a-OH-gypsogenic acid, QA-C3-GlcA; QA-C3-GlcA-Gal, or QA-C3-GlcA-Gal-Xyl (QA-TriX) (prosapogenin), wherein X or Xyl is xylose, Gal is galactose, and GlcA is glucuronic acid.

[0048] Aspects of the present disclosure provide therapeutic compositions comprising a therapeutically relevant agent and a triterpene glycoside produced by a host cell described herein. In some embodiments, the triterpene glycoside is quillaic acid (Q A), 16a-0H-gypsogenic acid, QA-C3-GlcA; QA-C3-GlcA-Gal, or QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin).

[0049] Aspects of the present disclosure provide compositions comprising phosphorylated QS-21, glycosylated QS-21, methylated QS-21, sulfated QS-21, acetylated QS-21, malonylated QS-21, benzoylated QS-21, esterase hydrolyzed QS-21, and glycoside hydrolase hydrolyzed QS-21.

[0050] BRIEF DESCRIPTION OF DRAWINGS

[0051] The accompanying drawings are not intended to be drawn to scale. The drawings are illustrative only and are not required for enablement of the disclosure. For purposes of clarity, not every component may be labeled in every drawing. In the drawings:

[0052] FIG. 1 shows a 2,3-oxidosqualene to prosapogenin biosynthetic pathway, including product intermediates, including quillaic acid (QA).

[0053] FIG. 2 shows UDP-sugar biosynthetic pathways relevant to the 2,3-oxidosqualene to prosapogenin biosynthetic pathway.

[0054] FIGs. 3A-3B show plots of liquid chromatography mass spectrometry (LCMS) chromatograms and mass spectra for the monophosphorylated product generated by the reaction of QS-21 with ATP and a phosphotransferase.

[0055] FIG. 3A shows a chromatographic peak at RT = 11.17 min (shaded) corresponding to the precursor ion mono-phosphorylated QS-21 isomer 1 at m / z 1033.9429 ([M-2H]2"), detected predominantly in the doubly charged state. This m / z value is consistent with the calculated mass for mono-phosphorylated QS-21 [M-H] ", corresponding to the addition of a single phosphate moiety to QS-21. FIG. 3B shows chromatographic peak at RT = 11.41 min (shaded) corresponding to the precursor ion mono-phosphorylated QS-21 isomer 2 at m / z 1033.9427 ([M-2H]2), detected predominantly in the doubly charged state. This m / z value is consistent with the calculated mass for mono-phosphorylated QS-21 [M-H] ", corresponding to the addition of a single phosphate moiety to QS-21.

[0056] FIGs. 4A-4G show plots of liquid chromatography mass spectrometry (LCMS) chromatograms and mass spectra for the mono-glycosylated product generated by the reaction of QS-21 with a UDP-sugar and a UDP -glycosyltransferase.

[0057] FIG. 4A shows a chromatographic peak at RT = 10.91 min. (shaded) corresponding to the precursor ion mono-glucosylated QS-21 isomer 1 at m / z 1074.9824 ([M-2H]2'), detected predominantly in the doubly charged state. This m / z value correspond to the parent ion with m / z 2150.9770, consistent with the predicted m / z for mono-glucosylated QS-21 [M-H]" (i.e., QS-21 + one glucose sugar).

[0058] FIG. 4B shows a chromatographic peak at RT = 11.05 min. (shaded) corresponding to a precursor ion mono-glucosylated QS-21 isomer 2 at m / z 1074.9818 ([M-2H]2'), detected predominantly in the doubly charged state. This m / z value correspond to the parent ion with m / z 2150.9770, consistent with the predicted m / z for mono-glucosylated QS-21 [M-H]" (i.e., QS-21 + one glucose sugar).

[0059] FIG. 4C shows a chromatographic peak at RT = 11.12 min. (shaded) corresponding to a precursor ion mono-glucosylated QS-21 isomer 3 at m / z 1074.9814 ([M-2H]2'), detected predominantly in the doubly charged state. This m / z value correspond to the parent ion with m / z 2150.9770, consistent with the predicted m / z for mono-glucosylated QS-21 [M-H]" (i.e., QS-21 + one glucose sugar).

[0060] FIG. 4D shows a chromatographic peak at RT = 11.20 min. (shaded) corresponding to a precursor ion mono-glucosylated QS-21 isomer 4 at m / z 1074.9818 ([M-2H]2'), detected predominantly in the doubly charged state. This m / z value correspond to the parent ion with m / z 2150.9770, consistent with the predicted m / z for mono-glucosylated QS-21 [M-H]" (i.e., QS-21 + one glucose sugar).

[0061] FIG. 4E shows a chromatographic peak at RT = 11.28 min. (shaded) corresponding to a precursor ion mono-glucosylated QS-21 isomer 5 at m / z 1074.9829 ([M-2H]2'), detected predominantly in the doubly charged state. This m / z value corresponds to the parent ion with m / z 2150.9770, consistent with the predicted m / z for mono-glucosylated QS-21 [M-H] (i.e., QS-21 + one glucose sugar).

[0062] FIG. 4F shows a chromatographic peak at RT = 11.44 min. (shaded) corresponding to a precursor ion mono-glucosylated QS-21 isomer 6 - at m / z 1074.9821 ([M-2H]2'), detected predominantly in the doubly charged state. This m / z value corresponds to the parent ion with m / z 2150.9770, consistent with the predicted m / z for mono-glucosylated QS-21 [M-H] (i.e., QS-21 + one glucose sugar).

[0063] FIG. 4G shows a chromatographic peak at RT = 11.61 min. (shaded) corresponding to a precursor ion mono-glucosylated QS-21 isomer 7 at m / z 1074.9833 ([M-2H]2'), detected predominantly in the doubly charged state. This m / z value corresponds to the parent ion with m / z 2150.9770, consistent with the predicted m / z for mono-glucosylated QS-21 [M-H] (i.e., QS-21 + one glucose sugar).

[0064] FIGs. 5A-5D show plots of liquid chromatography mass spectrometry (LCMS) chromatograms and mass spectra for the mono- and di-methylated products generated by the reaction of QS-21 with S-Adenosyl methionine and a methyltransferase.

[0065] FIG. 5A shows a chromatographic peak at RT = 11.84 min. (shaded) corresponding to a parent ion mono-methylated QS-21 isomer 1 with m / z 2001.9341, consistent with the predicted m / z for mono-methylated QS-21 [M-H] (i.e., QS-21 + one methyl group).

[0066] FIG. 5B shows a chromatographic peak at RT = 11.44 min. (shaded) corresponding to a parent ion mono-methylated QS-21 isomer 2 with m / z 2001.9349, consistent with the predicted m / z for mono-methylated QS-21 [M-H] (i.e., QS-21 + one methyl group).

[0067] FIG. 5C shows a chromatographic peak at RT = 11.94 min. (shaded) corresponding to a parent ion di-methylated QS-21 isomer 1 with m / z 2015.9517, consistent with the predicted m / z for di-methylated QS-21 [M-H] (i.e., QS-21 + two methyl groups).

[0068] FIG. 5D shows a chromatographic peak at RT = 11.52 min. (shaded) corresponding to a parent ion di-methylated QS-21 isomer 2 with m / z 2015.9510, consistent with the predicted m / z for di-methylated QS-21 [M-H] (i.e., QS-21 + two methyl groups).

[0069] FIGs. 6A-6D show plots of liquid chromatography mass spectrometry (LCMS) chromatograms and mass spectra for the mono-sulfated product generated by the reaction of QS-21 with 3'-phosphoadenosine-5'-phosphosulfate and a methyltransferase.

[0070] FIG. 6A shows a chromatographic peak at RT = 11.27 min (shaded) corresponding to a precursor ion mono-sulfated QS-21 isomer 1 with m / z 1033.9351 ([M-2H]2), detected predominantly in the doubly charged state, consistent with the calculated mass for monosulfated QS-21 [M-H] (i.e., QS-21 + one sulfate moiety).

[0071] FIG. 6B shows a chromatographic peak at RT = 11.14 min (shaded) corresponding to a precursor ion mono-sulfated QS-21 isomer 2 with m / z 1033.9347 ([M-2H]2), detected predominantly in the doubly charged state, consistent with the calculated mass for monosulfated QS-21.

[0072] FIG. 6C shows a chromatographic peak at RT = 11.52 min (shaded) corresponding to a precursor ion mono-sulfated QS-21 isomer 3 with m / z 1033.9343 ([M-2H]2), detected predominantly in the doubly charged state, consistent with the calculated mass for monosulfated QS-21 [M-H] (i.e., QS-21 + one sulfate moiety).

[0073] FIG. 6D shows a chromatographic peak at RT = 11.63 min (shaded) corresponding to a precursor ion mono-sulfated QS-21 isomer 4 with m / z 1033.9353 ([M-2H]2), detected predominantly in the doubly charged state, consistent with the calculated mass for monosulfated QS-21 [M-H] (i.e., QS-21 + one sulfate moiety).

[0074] FIGs. 7A-7F show plots of liquid chromatography mass spectrometry (LCMS) and sample-prep LCMS chromatograms and mass spectra for the mono-acetylated, malonylated, and benzoylated product generated by the reaction of QS-21 with acetyl-CoA and an acyltransferase.

[0075] FIG. 7A shows a chromatographic peak at RT = 11.70 min. (shaded) corresponding to a parent ion mono-acetylated QS-21 isomer 1 with m / z 2029.9371, consistent with the predicted m / z for mono-acetylated QS-21 [M-H] (i.e., QS-21 + one acetyl group).

[0076] FIG. 7B shows a chromatographic peak at RT = 11.83 min. (shaded) corresponding to a parent ion mono-acetylated QS-21 isomer 2 with m / z 2029.9371, consistent with the predicted m / z for mono-acetylated QS-21 [M-H] (i.e., QS-21 + one acetyl group).

[0077] FIG. 7C shows a chromatographic peak at RT = 12.14 min. (shaded) corresponding to a parent ion mono-acetylated QS-21 isomer 3 with m / z 2029.9355, consistent with the predicted m / z for mono-acetylated QS-21 [M-H] (i.e., QS-21 + one acetyl group).

[0078] FIG. 7D shows a chromatographic peak at RT = 12.23 min. (shaded) corresponding to a parent ion mono-acetylated QS-21 isomer 4 with m / z 2029.9354, consistent with the predicted m / z for mono-acetylated QS-21 [M-H] (i.e., QS-21 + one acetyl group).

[0079] FIG. 7E shows a sample-prep LC-MS spectra corresponding to a precursor ion mono-malonylated QS-21 product at m / z 1036.4542 ([M-2H]2) detected predominantly in the doubly charged state, consistent with the predicted m / z for mono-malonylated QS-21 [M-2H]2, corresponding to the addition of a single malonyl group to QS-21.

[0080] FIG. 7F shows a sample-prep LC-MS spectra corresponding to a precursor ion mono-benzoylated QS-21 product at m / z 2091.9316 ([M-H]Q, consistent with the predicted m / z for mono-benzoylated QS-21 [M-H] (i.e., QS-21 + one benzoyl group).

[0081] FIGs. 8A-8B show plots of sample-prep LC-MS mass spectra for the hydrolyzed QS-21 product generated by the reaction of QS-21 with malonyl-CoA and a glycosyl hydrolase.

[0082] FIG. 8A shows a sample-prep LC-MS spectra showing a precursor ion of GH deglycosylated QS-21 product 1 with m / z 1517.7900, consistent with the predicted m / z for QS-21 minus the C3 branching sugars.

[0083] FIG. 8B shows a sample-prep LC-MS spectra showing a precursor ion of GH deglycosylated QS-21 product 2 with m / z 1385.7460, consistent with the predicted m / z for QS-21 minus the C3 branching sugars and a pentose.

[0084] FIG. 9 shows the chemical structure of QS-21.

[0085] DETAILED DESCRIPTION

[0086] Triterpene glycosides, such as QS-7 and QS-21, are saponins that can be derived from Quillaja saponaria. These compounds, and precursors and derivatives of these compounds, can be used as adjuvants in therapeutic compositions (such as vaccines). Alternatively, these saponins, and precursors and derivatives thereof, are useful for other therapeutic applications. Cell culture-based production of saponins could allow for lowered cost of therapeutics involving saponins. However, production of saponins involves a complex biosynthetic pathway.

[0087] Synthesis of Triterpene Glycosides

[0088] Aspects of the disclosure provide methods and compositions (e.g., host cells) for synthesis of triterpene glycosides (e.g., prosapogenin), and / or precursors and / or derivatives thereof.

[0089] FIG. 1 shows a biosynthetic pathway for the production of QA and prosapogenin (and related intermediate products), starting with 2,3-oxidosqualene. Structures for QA and prosapogenin (and related intermediate products) are provided. Relevant enzymes are identified and further discussed herein. FIG. 2 shows biosynthetic pathways for the production of UDP-sugars relevant to the prosapogenin biosynthetic pathway. Structures for the UDP-sugars (and related intermediate products) are provided. Relevant enzymes are identified and further discussed herein.

[0090] Synthesis of Quillaic Acid (QA), and Related Intermediates

[0091] To produce QA (and related intermediates), 2,3-oxidosqualene is first contacted with a beta-amyrin synthase, generating beta-amyrin.

[0092] In some embodiments, beta-amyrin is contacted with a C28 oxidase to generate oleanolic acid.

[0093] In some embodiments, the beta-amyrin is contacted with a C16C28 oxidase to generate echinocystic acid. In some embodiments, the echinocystic acid is then contacted with a C23 oxidase to generate QA.

[0094] In some embodiments, the oleanolic acid is contacted with a C23 oxidase to generate hederagenin, gypsogenin, and / or gypsogenic acid.

[0095] In some aspects, host cells are disclosed herein which comprise a heterologous polynucleotide encoding an enzyme of the QA biosynthetic pathway. In some embodiments, the host cell comprises one or more heterologous polynucleotides collectively encoding: a) a beta-amyrin synthase (BAS); b) a cytochrome P450 C16C28 oxidase (or separate P450 C16 and C28 oxidases); c) a cytochrome P450 C23 oxidase; d) a cytochrome P450 reductase (CPR); and e) a membrane steroid binding protein (MSBP).

[0096] Beta-amyrin synthase (BAS)

[0097] As disclosed herein, a beta-amyrin synthase (BAS) is an enzyme that produces beta-amyrin. In some aspects, a BAS uses 2,3-oxidosqualene as a substrate to produce beta-amyrin. In some embodiments, a BAS comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 59-314.

[0098] Cytochrome P450 (CYP) C16C28 oxidase

[0099] As disclosed herein, a C16C28 oxidase is an enzyme that oxidizes both C16 and C28 carbons, for example, to a hydroxyl group. In some aspects, a C16C28 oxidase: oxidizes the C16 and C28 carbons of 2,3-oxidosqualene to generate echinocystic acid. In some embodiments, a C16C28 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ IDNOs: 315-318.

[0100] Cytochrome P450 (CYP) C28 oxidase

[0101] As disclosed herein, a CYPC28 oxidase (C28 oxidase) is an enzyme that oxidizes a C28 carbon, for example, to a hydroxyl group. In some aspects, a C28 oxidase oxidizes the C28 carbon of beta-amyrin to generate oleanolic acid. In some embodiments, a C28 oxidase is also a C16 oxidase, otherwise referred to as a C16C28 oxidase. In some embodiments, a C28 oxidase (and / or C16C28 oxidase) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315 and 318-320.

[0102] Cytochrome P450 (CYP) C23 oxidase

[0103] As disclosed herein, a CYPC23 oxidase (C23 oxidase) is an enzyme that oxidizes a C23 carbon, for example, to an aldehyde group. In some aspects, a C23 oxidase oxidizes: the C23 carbon of oleanolic acid to generate hederagenin, gysogenin, and / or gypsogenic acid; and / or the C23 carbon of echinocystic acid to generate QA. In some embodiments, a C23 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 321-335.

[0104] Cytochrome P450 reductase (CPR)

[0105] As disclosed herein, a cytochrome P450 reductase (CPR) is an enzyme that acts as redox partner for one or more of: a C16 oxidase, a C28 oxidase, a C16C28 oxidase, and a C23 oxidase. In some aspects, a CPR acts as a redox partner for each of a C16 oxidase, a C28 oxidase, a C16C28 oxidase, and a C23 oxidase. In some embodiments, a CPR comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 336-377.

[0106] Membrane Steroid Binding Protein (MSBP)

[0107] As disclosed herein, a membrane steroid binding protein (MSBP) is a protein that acts as scaffold protein for one or more of: a Cl 6 oxidase, a C28 oxidase, a C16C28 oxidase, a C23 oxidase, and a CPR. In some aspects, an MSBP acts as a scaffold protein for each of a C16 oxidase, a C28 oxidase, a C16C28 oxidase, a C23 oxidase, and a CPR. In some embodiments, a MSBP comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 378-416.

[0108] Synthesis of QA-TriX (prosapogenin), and Related Intermediates

[0109] To produce QA-TriX (prosapogenin) (and related intermediates), QA is first contacted with a UDP-glucuronic acid transferase to generate QA-C3-GlcA.

[0110] In some embodiments, QA-C3-GlcA is contacted with a UDP -galactose transferase (GalT) to generate QA-C3-GlcA-Gal.

[0111] In some embodiments, QA-C3-GlcA-Gal is contacted with a UDP -xylose transferase (XylT) to generate prosapogenin.

[0112] In some aspects, host cells are disclosed herein which comprise a heterologous polynucleotide encoding an enzyme of the prosapogenin biosynthetic pathway. In some embodiments, the host cell comprises one or more heterologous polynucleotides collectively encoding: a) a UDP-GlcA transferase (GlcAT), b) a UDP-galactose transferase (GalT), and c) a UDP -xylose transferase (XylT). In some embodiments, the host cell further comprises one or more heterologous polynucleotides collectively encoding: a) a beta-amyrin synthase (BAS); b) a cytochrome P450 C16 oxidase; c) a cytochrome P450 C28 oxidase; d) a cytochrome P450 C23 oxidase; e) a cytochrome P450 reductase (CPR); and f) a membrane steroid binding protein (MSBP). In some embodiments, the host cell produces UDP-sugars relevant (e.g., UDP-GlcA, UDP-Gal, and / or UDP-Xyl) to the prosapogenin production shown in FIG. 1. In some embodiments, the host cell comprises one or more heterologous polynucleotides encoding a UDP-Glc transferase, a UDP-galactose transferase, and / or a UDP-xylose transferase. UDP-GlcA transferase (GlcAT)

[0113] As disclosed herein, a UDP-GlcA transferase (GlcAT) is an enzyme that attaches a glucuronic acid residue to a target molecule. In some aspects, GlcAT uses UDP-GlcA as a substrate and attaches a glucuronic acid residue to the C3 position of QA to form QA-C3-GlcA . In some embodiments, a GlcAT comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 417-441.

[0114] UDP-galactose transferase (GalT)

[0115] As disclosed herein, a UDP-Galactose transferase (GalT) is an enzyme that attaches a galactose residue to a target molecule. In some aspects, a GalT uses UDP-Gal as a substrate and attaches a galactose residue to QA-C3-GlcA to form QA-C3-GlcA-Gal. In some embodiments, a GalT comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 442-537.

[0116] UDP-xylose transferase (XylT)

[0117] As disclosed herein, a UDP-Xylose transferase (XylT) is an enzyme that attaches a xylose residue to a target molecule. In some aspects, a XylT uses UDP-Xylose as a substrate and attaches a xylose residue to QA-C3-GlcA-Gal to form QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin). In some embodiments, a XylT comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ IDNOs: 484, 485, 491, 494, 502, 504, 517, and 538-572.

[0118] Synthesis of UDP-glucuronic acid

[0119] To produce UDP-glucuronic acid, UDP-glucose is contacted with a UDP-glucose dehydrogenase. In some aspects, host cells are disclosed herein which comprise a heterologous polynucleotide encoding an enzyme of the UDP-glucuronic acid biosynthetic pathway. In some embodiments, the host cell comprises one or more heterologous polynucleotides encoding a UDP-glucose dehydrogenase.

[0120] UDP-glucose dehydrogenase (UGD)

[0121] As disclosed herein, a UDP-glucose dehydrogenase (UGD) is an oxidoreductase, which catalyzes the oxidation of UDP-glucose. In some aspects, a UGD converts UDP-glucose to UDP-GlcA. In some embodiments, a UGD comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 46-56, and 712-1142.

[0122] Synthesis of UDP-xylose

[0123] To produce UDP-xylose, UDP-glucuronic acid is contacted with a UDP-xylose synthase. In some aspects, host cells are disclosed herein which comprise a heterologous polynucleotide encoding an enzyme of the UDP-xylose biosynthetic pathway. In some embodiments, the host cell comprises one or more heterologous polynucleotides encoding a UDP-xylose synthase.

[0124] UDP-xylose synthase (UXS)

[0125] As disclosed herein, a UDP-xylose synthase (UXS) is a synthase that generates UDP-xylose. In some aspects, a UXS converts UDP-GlcA to UDP-Xyl. In some embodiments, a UXS comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 573-711.

[0126] Modifying QS-21 and QS-21 Derivatives

[0127] To produce modified QS-21 (and QS-21 derivatives), QS-21 (FIG. 9) is contacted with one or more of a phosphotransferase, a UDP-glycosyltransferase, a methyltransferase, a sulfotransferase, an acyltransferase, a glycoside hydrolase, and an esterase. In some aspects, host cells are disclosed herein which comprise a heterologous polynucleotide encoding an enzyme for modifying QS-21. In some embodiments, the host cell comprises one or more heterologous polynucleotides collectively encoding one or more of: a phosphotransferase, a UDP-glycosyltransferase, a methyltransferase, a sulfotransferase, an acyltransferase, a glycosyl hydrolase, and an esterase.

[0128] Phosphotransferase

[0129] As disclosed herein, a phosphotransferase is an enzyme that transfers a phosphoryl group from a donor molecule to a target molecule. In some aspects, a phosphotransferase uses ATP as a donor molecule and transfers and attaches one or more phosphoryl groups to QS-21. In some embodiments, a phosphotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1239-1257.

[0130] UDP-Glycosyltransferase

[0131] As disclosed herein, a UDP-glycosyltransferase is an enzyme that transfers a sugar from a donor molecule to a target molecule. In some aspects, a UDP-glycosyltransferase uses UDP-glucose as a donor molecule and transfers and attaches one or more glycosyl groups to QS-21. In some embodiments, a UDP-glycosyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1258-1283.

[0132] Methyltransferase

[0133] As disclosed herein, a methyltransferase is an enzyme that transfers a methyl group from donor molecules to target molecules. In some aspects, a methyltransferase uses S-adenosyl methionine as a donor molecule and transfers and attaches one or more methyl groups to QS-21. In some embodiments, a methyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1284-1289. Sulfotransferase

[0134] As disclosed herein, a sulfotransferase is an enzyme that transfers a sulfonyl group from a donor molecule to a target molecule. In some embodiments, a sulfotransferase uses 3'-phosphoadenosine-5'-phosphosulfate as a donor molecule and transfers and attaches one or more sulfonyl groups to QS-21. In some embodiments, a sulfotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1290-1307.

[0135] Acyltransferase

[0136] As disclosed herein, an acetyltransferase is an enzyme that transfers an acyl group from a donor molecule to a target molecule. In some aspects, an acyltransferase uses acetyl-coA, malonyl-coA, or benzoyl-coA as a donor molecule and transfers and attaches one or more acyl groups to QS-21. In some embodiments, a methyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1308-1454.

[0137] Glycosyl hydrolase

[0138] As disclosed herein, a glycosyl hydrolase is an enzyme that cleaves a glycosidic bond. In some aspects, a glycosyl hydrolase cleaves one or more glycosidic bonds in QS-21. In some embodiments, a glycosyl hydrolase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1455-1472.

[0139] Esterase

[0140] As disclosed herein, an esterase is an enzyme that cleaves an ester bond. In some aspects, an esterase cleaves one or more ester bonds in QS-21. In some embodiments, an esterase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1473-1499.

[0141] In some embodiments, the host cells are capable of producing QS-21. In some embodiments, QS-21 is produced via a method which involves producing each of the various intermediates. In some embodiments, an intermediate in a pathway for production of QS-21 is produced using an enzymatic step (e.g., as described herein) or is produced via chemical synthesis. In some embodiments, QS-21 can be purified directly from the Quillaja saponaria.

[0142] Polynucleotides

[0143] The term “heterologous” with respect to a polynucleotide, such as a polynucleotide comprising a gene, is used interchangeably with the term “exogenous” and the term “recombinant” and refers to: a polynucleotide that has been artificially supplied to a biological system such as a cell; a polynucleotide that has been modified within a biological system; or a polynucleotide whose expression or regulation has been manipulated within a biological system. A heterologous polynucleotide that is introduced into or expressed in a host cell may be a synthetic polynucleotide, a polynucleotide that comes from a different organism or species from the host cell, or a polynucleotide that results from modification or selective editing within the host cell of a polynucleotide that is endogenous to the host cell. A polynucleotide comprising a sequence that is endogenous to a host cell also may be considered heterologous when it is, for example: situated non-naturally in the host cell; expressed recombinantly in the host cell, either stably or transiently; present in a copy number that differs from the naturally occurring copy number within the host cell; or expressed in a non-natural way or at a non-natural level within the host cell, such as through manipulation of regulatory regions that control expression of the polynucleotide. In some embodiments, a heterologous polynucleotide is a polynucleotide that comprises a sequence endogenous to a host cell but whose expression is driven by a promoter that does not naturally regulate expression of the polynucleotide. In other embodiments, a heterologous polynucleotide is a polynucleotide that comprises a sequence endogenous to the host cell and whose expression is driven by a promoter that does naturally regulate expression of the polynucleotide, but the promoter driving its expression or another regulatory region regulating its expression has been modified. In some embodiments, the promoter is recombinantly activated or repressed. For example, gene-editing techniques may be used to regulate expression of a polynucleotide in a cell, including an endogenous polynucleotide, from a promoter, including an endogenous promoter. See, e.g., Chavez etal., Nat Methods. 2016 Jul; 13(7): 563-567. A heterologous polynucleotide may comprise a wild-type sequence or a mutant sequence as compared with a reference polynucleotide sequence.

[0144] A polynucleotide encoding a polypeptide associated with the disclosure may be incorporated into any appropriate vector through any method known in the art. For example, the vector may be an expression vector, including but not limited to a viral vector (e.g., a lentiviral, retroviral, adenoviral, or adeno-associated viral vector), any vector suitable for transient expression, any vector suitable for constitutive expression, or any vector suitable for inducible expression (e.g., a galactose-inducible or doxycycline-inducible vector). The vector may be a cloning vector, such as a plasmid, fosmid, phagemid, virus genome or artificial chromosome.

[0145] As used in this application, the term "expression vector" or "expression construct" refers to a nucleic acid construct, generated recombinantly or synthetically, with a series of specified nucleic acid elements that permit transcription of a particular polynucleotide in a host cell, such as a yeast cell or bacterial cell. In some embodiments, a polynucleotide associated with the disclosure is inserted into an expression vector or expression construct such that it is operably joined to regulatory sequences and, in some embodiments, expressed as an RNA transcript. In some embodiments, the expression vector or expression construct contains one or more markers, such as a selectable marker, to identify cells transformed or transfected with the expression vector or expression construct. A polynucleotide encoding a polypeptide associated with the disclosure is “operably joined” or “operably linked” to a regulatory sequence when the polynucleotide and the regulatory sequence are covalently linked, and the expression or transcription of the polynucleotide is under the influence or control of the regulatory sequence.

[0146] In some embodiments, a polynucleotide encoding any of the polypeptides described in this application is under the control of regulatory sequences (e.g., enhancer sequences). In some embodiments, a polynucleotide (e.g., a polynucleotide comprising a gene) is expressed under the control of a promoter. In some embodiments, the promoter is a native promoter, corresponding to the promoter of the gene in its endogenous context. In other embodiments, the promoter is not the native promoter of the gene, e.g., the promoter is different from the promoter of the gene in its endogenous context.

[0147] In some embodiments, the promoter is a eukaryotic promoter. Non-limiting examples of eukaryotic promoters include TDH3, PGK1, PKC1, PDC1, TEF1, TEF2, RPL18B, SSA1, TDH2, PYK1,TPI1 GALI, GAL10, GAL7, GALS, GAL2, MET3, MET25, HXT3, HXT7, ACT1, ADH1, ADH2, CUP1-1, EN02, and S0D1, as would be known to one of ordinary skill in the art (see, e.g., Addgene website: blog. addgene. org / plasmids-101-the-promoter-regi on). In some embodiments, the promoter is a prokaryotic promoter (e.g., bacteriophage or bacterial promoter). Non-limiting examples of bacteriophage promoters include Plslcon, T3, T7, SP6, and PL. Non-limiting examples of bacterial promoters include Pbad, PmgrB, Ptrc2, Plac / ara, Ptac, and Pm.

[0148] In some embodiments, the promoter is an inducible promoter. As used in this application, an “inducible promoter” is a promoter controlled by the presence or absence of a molecule. Non-limiting examples of inducible promoters include chemically regulated promoters and physically regulated promoters. For chemically regulated promoters, the transcriptional activity can be regulated by one or more compounds, such as alcohol, an antibiotic such as tetracycline, a carbon source such as galactose, a steroid, a metal, or other compounds. For physically regulated promoters, transcriptional activity can be regulated by a phenomenon such as light or temperature. Non-limiting examples of tetracycline-regulated promoters include anhydrotetracycline (aTc)-responsive promoters and other tetracyclineresponsive promoter systems (e.g., a tetracycline repressor protein (tetR), a tetracycline operator sequence (tetO) and a tetracycline transactivator fusion protein (tTA)). Non-limiting examples of steroid-regulated promoters include promoters based on the rat glucocorticoid receptor, human estrogen receptor, moth ecdysone receptors, and promoters from the steroid / retinoid / thyroid receptor superfamily. Non-limiting examples of metal-regulated promoters include promoters derived from metallothionein (proteins that bind and sequester metal ions) genes. Non-limiting examples of pathogenesis-regulated promoters include promoters induced by salicylic acid, ethylene or benzothiadi azole (BTH). Non-limiting examples of temperature / heat-inducible promoters include heat shock promoters. Non-limiting examples of light-regulated promoters include light responsive promoters from plant cells. In certain embodiments, the inducible promoter is a galactose-inducible promoter. In some embodiments, the inducible promoter is induced by one or more physiological conditions (e.g., pH, temperature, radiation, osmotic pressure, saline gradients, cell surface binding, or concentration of one or more extrinsic or intrinsic inducing agents). Non-limiting examples of an extrinsic inducer or inducing agent include amino acids and amino acid analogs, saccharides and polysaccharides, nucleic acids, protein transcriptional activators and repressors, cytokines, toxins, petroleum-based compounds, metal containing compounds, salts, ions, enzyme substrate analogs, hormones or any combination thereof.

[0149] In some embodiments, the promoter is a constitutive promoter. As used in this application, a “constitutive promoter” refers to an unregulated promoter that allows continuous transcription of a gene. Non-limiting examples of a constitutive promoter include TDH3, PGK1, PKC1, PDC1, TEF1, TEF2, RPL18B, SSA1, TDH2, PYK1, TPI1, HXT3, HXT7, ACT1, ADH1, ADH2, EN02, and SOD1.

[0150] Other inducible promoters or constitutive promoters known to one of ordinary skill in the art are also contemplated.

[0151] In some embodiments, introduction of a polynucleotide, such as a polynucleotide encoding a polypeptide associated with the disclosure, into a host cell results in genomic integration of the polynucleotide. In some embodiments, a host cell (e.g., a yeast cell or other eukaryotic cell) comprises at least 1 copy, at least 2 copies, at least 3 copies, at least 4 copies, at least 5 copies, at least 6 copies, at least 7 copies, at least 8 copies, at least 9 copies, at least 10 copies, at least 11 copies, at least 12 copies, at least 13 copies, at least 14 copies, at least 15 copies, at least 16 copies, at least 17 copies, at least 18 copies, at least 19 copies, at least 20 copies, at least 21 copies, at least 22 copies, at least 23 copies, at least 24 copies, at least 25 copies, at least 26 copies, at least 27 copies, at least 28 copies, at least 29 copies, at least 30 copies, at least 31 copies, at least 32 copies, at least 33 copies, at least 34 copies, at least 35 copies, at least 36 copies, at least 37 copies, at least 38 copies, at least 39 copies, at least 40 copies, at least 41 copies, at least 42 copies, at least 43 copies, at least 44 copies, at least 45 copies, at least 46 copies, at least 47 copies, at least 48 copies, at least 49 copies, at least 50 copies, at least 60 copies, at least 70 copies, at least 80 copies, at least 90 copies, at least 100 copies, or more, including any values in between, of a polynucleotide sequence, such as a polynucleotide sequence encoding any of the polypeptides described in this application, in its genome. Said copies may be inserted into the same locus or into different loci of a host cell of the disclosure.

[0152] In some embodiments, the sequence of a polynucleotide (e.g., a polynucleotide comprising a gene) is codon optimized. Codon optimization may increase expression of a gene by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%, including all values in between) relative to a reference sequence that is not codon-optimized.

[0153] In some embodiments, a heterologous polynucleotide described herein encodes a protein provided in Table 22 (where corresponding example polynucleotides are provided in Table 33) and comprises a sequence having a least 80%, at least 85%, at least 90%, at least 95%, or 100% with a sequence provided in Table 22.

[0154] Percent Identity

[0155] Aspects of the disclosure relate to variant polypeptides for producing a triterpene glycoside (e.g., prosapogenin) or a precursor or derivative thereof. As used in this disclosure, a "variant" polynucleotide refers to a polynucleotide that differs from a reference polynucleotide by one or more nucleotides in its sequence. As used in this disclosure, a "variant" polypeptide refers to a polypeptide that differs from a reference polypeptide by one or more amino acids in its sequence.

[0156] A variant may share at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with a reference sequence, including all values in between.

[0157] Unless otherwise noted, the term “sequence identity” refers to the relatedness of the sequences of two polypeptides or polynucleotides when the sequences are aligned, and the term “percent identity” refers to the percentage of residues (amino acids or nucleotides) that are identical when two polypeptide or polynucleotide sequences are aligned. In some embodiments, sequence identity and / or percent identity is determined across the entire length of a sequence, while in other embodiments, sequence identity and / or percent identity is determined over a region of a sequence.

[0158] Percent identity of polypeptide or polynucleotide sequences can be calculated by any of the methods known to one of ordinary skill in the art. For example, percent identity can be determined using the algorithm of Karlin and Altschul Proc. Natl. Acad. Sci. USA 87:2264-68, 1990, modified as in Karlin and Altschul Proc. Natl. Acad. Sci. USA 90:5873-77, 1993. Such an algorithm is incorporated into the NBLAST® and XBLAST® programs (version 2.0) of Altschul et al., J. Mol. Biol. 215:403-10, 1990. BLAST® protein searches can be performed, for example, with the XBLAST program, score=50, wordlength=3. Where gaps exist between two sequences, Gapped BLAST® can be utilized, for example, as described in Altschul et al., Nucleic Acids Res.

[0159] 25(17):3389-3402, 1997. When utilizing BLAST® and Gapped BLAST® programs, the default parameters of the respective programs e.g., XBLAST® and NBLAST®) can be used, or the parameters can be adjusted appropriately as would be understood by one of ordinary skill in the art.

[0160] A second example of a local alignment technique is based on the Smith-Waterman algorithm (Smith, T.F. & Waterman, M.S. (1981) J. Mol. Biol. 147:195-197). An example of a global alignment technique is the Needleman-Wunsch algorithm (Needleman, S.B. & Wunsch, C.D. (1970) J. Mol. Biol. 48:443-453), which is based on dynamic programming. A further example of a global alignment technique is the Fast Optimal Global Sequence Alignment Algorithm (FOGSAA).

[0161] In some embodiments, the identity of two polypeptide sequences is determined by aligning the amino acid sequences of the polypeptides, calculating the number of identical amino acids, and dividing by the length of one of the polypeptide sequences. In some embodiments, the identity of two polynucleotide sequences is determined by aligning the nucleic acid sequences of the polynucleotides, calculating the number of identical nucleic acids and dividing by the length of one of the polynucleotide sequences.

[0162] For multiple sequence alignments, computer programs including Clustal Omega (Sievers et al., Mol SystBiol. 2011 Oct 11 ;7: 539) may be used.

[0163] In some embodiments, sequence identity is determined using the algorithm of Karlin and Altschul Proc. Natl. Acad. Sci. USA 87:2264-68, 1990, modified as in Karlin and Altschul Proc. Natl. Acad. Sci. USA 90:5873-77, 1993 (e.g., BLAST®, NBLAST®, XBLAST® or Gapped BLAST® programs, using default parameters of the respective programs).

[0164] In some embodiments, the sequence identity of two amino acid or polynucleotide sequences is determined using the Smith-Waterman algorithm (Smith, T.F. & Waterman, M.S. (1981) J. Mol. Biol. 147:195-197) or the Needleman-Wunsch algorithm (Needleman, S.B. & Wunsch, C.D. (1970) J. Mol. Biol. 48:443-453).

[0165] In some embodiments, the sequence identity of two amino acid or polynucleotide sequences is determined using a Fast Optimal Global Sequence Alignment Algorithm

[0166] (FOGSAA). In some embodiments, the sequence identity of two amino acid or polynucleotide sequences is determined using Clustal Omega (Sievers etal.,Mol SystBiol. 2011 Oct 11;7:539).

[0167] As used in this application, a residue (such as a nucleic acid residue or an amino acid residue) in sequence “X” is referred to as corresponding to a position or residue (such as a nucleic acid residue or an amino acid residue) “Z” in a different sequence “Y” when the residue in sequence “X” is at the counterpart position of “Z” in sequence “Y” when sequences X and Y are aligned using sequence alignment tools known in the art.

[0168] Variant sequences may be homologous sequences. As used in this application, homologous sequences are sequences (e.g., nucleic acid or amino acid sequences) that share a certain percent identity (e.g., at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% percent identity, including all values in between) and may be paralogous sequences, orthologous sequences, or sequences arising from convergent evolution. Paralogous sequences arise from duplication of a gene within a genome of a species, while orthologous sequences diverge after a speciation event. Two different species may have evolved independently but may each comprise a sequence that shares a certain percent identity with a sequence from the other species as a result of convergent evolution.

[0169] In some embodiments, a polypeptide variant (e.g., a BAS variant or C16C28 oxidase variant) comprises the same or substantially similar secondary structure (e.g., alpha helix, beta sheet) as that of a reference polypeptide (e.g., a reference BAS or a reference C16C28 oxidase). In some embodiments, a polypeptide variant (e.g., a BAS variant or a C16C28 oxidase variant) comprises the same or substantially similar tertiary structure as that of a reference polypeptide (e.g., a reference BAS or a reference C16C28 oxidase). As a non-limiting example, a variant polypeptide may have low primary sequence identity (e.g., less than 80%, less than 75%, or less than 70% sequence identity) compared to a reference polypeptide, but comprises one or more secondary structures (e.g., including but not limited to loops, alpha helices, or beta sheets) or the same tertiary structure as that of a reference polypeptide. For example, a loop may be located between a beta sheet and an alpha helix, between two alpha helices, or between two beta sheets. Homology modeling may be used to compare two or more tertiary structures. In some embodiments, an algorithm that determines the percent identity between a sequence of interest and a reference sequence described in this application accounts for the presence of circular permutation between the sequences. The presence of circular permutation may be detected using any method known in the art, including, for example, RASPODOM (Weiner etal., Bioinformatics. 2005 Apr l;21(7):932-7). In some embodiments, the presence of circulation permutation is corrected for (e.g, the domains in at least one sequence are rearranged) prior to calculation of the percent identity between a sequence of interest and a sequence described in this application.

[0170] Functional variants of polypeptides disclosed in this application are also encompassed by the present disclosure. For example, functional variants may bind one or more of the same substrates (e.g, 2,3-oxidosqualene) or produce one or more of the same products (e.g., phenol or phenyl acetate). Functional variants may be identified using any method known in the art. For example, the algorithm of Karlin and Altschul Proc. Natl. Acad. Sci. USA 87:2264-68, 1990 described above may be used to identify homologous proteins.

[0171] Host Cells

[0172] Any of the polynucleotides or polypeptides of the disclosure may be expressed in a host cell. As used in this application, the term “host cell” refers to a cell that can be used to express a polynucleotide, such as a polynucleotide that encodes a polypeptide used in production of triterpene glycosides (e.g., prosapogenin).

[0173] Any suitable host cell may be used to produce any of the polypeptides, including a BAS, a P450 oxidase or any other polypeptide disclosed in this application, including eukaryotic cells or prokaryotic cells. Suitable host cells include, but are not limited to, fungal cells (e.g., yeast cells), bacterial cells (e.g., E. coli cells), algal cells, plant cells, insect cells, and animal cells, including mammalian cells.

[0174] Suitable yeast host cells include, but are not limited to, Candida, Escherichia, Hansenula, Saccharomyces e.g., S. cerevisiae), Schizosaccharomyces, Pichia, Kluyveromyces (e.g., K. lactis), and Yarrowia (e.g., Y. lipolylica). In some embodiments, the yeast cell is Hansenula polymorpha, Saccharomyces cerevisiae, Saccharomyces carlsber gensis, Saccharomyces diastaticus, Saccharomyces norbensis, Saccharomyces kluyveri, Schizosaccharomyces pombe, Pichia fmlandica, Pichia trehalophila, Pichia kodamae, Pichia membranaefaciens, Pichia opuntiae, Pichia thermotolerans, Pichia salictaria, Pichia quercuum, Pichia pijperi, Pichia stipitis, Pichia methanolica, Pichia angusta, Komagataella phaffii, Komagataella pastoris, Kluyveromyces lactis, Candida albicans, or Yarrowia lipolytica.

[0175] In some embodiments, the yeast strain is an industrial polyploid yeast strain. Other nonlimiting examples of fungal cells include cells obtained from Aspergillus spp., Penicillium spp., Fusarium spp., Rhizopus spp., Acremonium spp., Neurospora spp., Sordaria spp., Magnaporthe spp., Allomyces spp., Ustilago spp., Botrytis spp., Myrothecium spp., and Trichoderma spp.

[0176] In certain embodiments, the host cell is an algal cell such as, Chlamydomonas (e.g., C. Reinhard / ii) and Phormidium (P. sp. ATCC29409).

[0177] In other embodiments, the host cell is a prokaryotic cell. Suitable prokaryotic cells include gram positive, gram negative, and gram-variable bacterial cells. The host cell may be a species of, but not limited to: Agrobacterium, Alicyclobacillus, Anabaena, Anacystis, Acinetobacter , Acidothermus, Arthrobacter, Azotobacter, Bacillus, Bifidobacterium, Brevibacterium, Butyrivibrio, Buchnera, Campestris, Campylobacter, Clostridium, Corynebacterium, Chromatium, Coprococcus, Escherichia, Enterococcus, Enterobacter , Erwinia, Fusobacterium, Faecalibacterium, Francisella, Flavobacterium, Geobacillus, Haemophilus, Helicobacter, Klebsiella, Lactobacillus, Lactococcus, Ilyobacter, Micrococcus, Microbacterium, Mesorhizobium, Methylobacterium, Methylobacterium, Mycobacterium, Neisseria, Pantoea, Pseudomonas, Prochlorococcus, Rhodobacter, Rhodopseudomonas, Rhodopseudomonas, Roseburia, Rhodospirillum, Rhodococcus, Scenedesmus, Streptomyces, Streptococcus, Synecoccus, Saccharomonospora, Saccharopolyspora, Staphylococcus, Serratia, Salmonella, Shigella, Thermoanaerobacterium, Tropheryma, Tularensis, Temecula, Thermosynechococcus, Thermococcus, Ureaplasma, Xanthomonas, Xylella, Yersinia, and Zymomonas.

[0178] In some embodiments, the bacterial host cell is of the Agrobacterium species (e.g., A. radiobacter, A. rhizogenes, A. rubi), the Arthrobacter species (e.g., A. aurescens, A. citreus, A. globformis, A. hydrocarboglutamicus, A. mysorens, A. nicotianae, A. paraffineus, A. protophonniae, A. roseoparaffinus, A. sulfureus, A. ureafaciens), or the Bacillus species (e.g., B. thuringiensis, B. anthracis, B. megaterium, B. subtilis, B. lentus, B. circulans, B. pumilus, B. lautus, B. coagulans, B. brevis, B. firmus, B. alkaophius, B. licheniformis, B. clausii, B. stearothermophilus, B. halodurans, B. amyloliquefaciens . In particular embodiments, the host cell is an industrial Bacillus strain including but not limited to B. subtilis, B. pumilus, B. licheniformis, B. megaterium, B. clausii, B. stearothermophilus and B. amyloliquefaciens. In some embodiments, the host cell is an industrial Clostridium species (e.g, C. acetobutylicum, C. tetani E88, C. lituseburense, C. saccharobutylicum, C. perfringens, C. beijerinckii). In some embodiments, the host cell is an industrial Corynebacterium species (e.g., C. glutamicum, C. acetoacidophilum). In some embodiments, the host cell is an industrial Escherichia species (e.g., E. colt). In some embodiments, the host cell is an industrial Erwinia species e.g., E. uredovora, E. carotovora, E. ananas, E. herbicola, E. punctata, E. terreus). In some embodiments, the host cell is an industrial Pantoea species e.g., P. citrea, P. agglomerans). In some embodiments, the host cell is an industrial Pseudomonas species, e.g., P. putida, P. aeruginosa, P. mevalonii). In some embodiments, the host cell is an

[0179] industrial Streptococcus species e.g., S. equisimiles, S. pyogenes, S. uberis). In some embodiments, the host cell is an industrial Streptomyces species e.g., S. ambofaciens, S. achromogenes, S. avermitilis, S. coelicolor, S. aureofaciens, S. aureus, S. fimgicidicus, S. griseus, S. lividans). In some embodiments, the host cell is an industrial Zymomonas species e.g., Z. mobilis, Z. lipolytica).

[0180] The present disclosure is also suitable for use with a variety of animal cell types, including mammalian cells, for example, human (including 293, HeLa, WI38, PER.C6 and Bowes melanoma cells), mouse (including 3T3, NSO, NS1, Sp2 / 0), hamster (CHO, BHK), monkey (COS, FRhL, Vero), and hybridoma cell lines.

[0181] The present disclosure is also suitable for use with a variety of plant cell types.

[0182] The term “cell,” as used in this application, may refer to a single cell or a population of cells, such as a population of cells belonging to the same cell line or strain. Use of the singular term “cell” should not be construed to refer explicitly to a single cell rather than a population of cells.

[0183] The host cell may comprise genetic modifications relative to a wild-type counterpart. As a non-limiting example, a host cell e.g., S. cerevisiae or Y. lipolytica may be modified to introduce or increase expression of one or more genes related to mevalonate flux. In some embodiments, the host cell comprises one or more genetic modifications such that the host cell has increased mevalonate flux relative to a host cell that lacks the one or more genetic modifications. In some embodiments, the genetic modification comprises one or more heterologous polynucleotides expressing one or more genes associated with mevalonate flux. Genes associated with mevalonate flux include, but are not limited to erg 10, ergl3, hmgl, erg 12, erg8, ergl9, idil, and erg20. In some embodiments, a host cell is modified to reduce or inactivate at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 genes.

[0184] In some embodiments, a host cell is modified to reduce or inactivate 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 genes.

[0185] Reduction of gene expression and / or gene inactivation may be achieved through any suitable method, including but not limited to deletion of the gene, introduction of a point mutation into the gene, truncation of the gene, introduction of an insertion into the gene, introduction of a tag or fusion into the gene, or selective editing of the gene. For example, polymerase chain reaction (PCR)-based methods may be used (see, e.g., Gardner et al., Methods Mol Biol. 2014;1205:45-78) or well-known gene-editing techniques may be used. As a nonlimiting example, genes may be deleted through gene replacement (e.g., with a marker, including a selection marker). A gene may also be truncated through the use of a transposon system (see, e.g., Poussu et al., Nucleic Acids Res. 2005; 33(12): el04).

[0186] A vector or polynucleotide encoding any of the polypeptides described in this application may be introduced into a suitable host cell using any method known in the art. Non-limiting examples of yeast transformation protocols are described in Gietz et al., yeast transformation can be conducted by the LiAc / SS Carrier DNA / PEG method. Methods Mol Biol. 2006; 313 : 107-20, which is incorporated by reference in its entirety. Host cells may be cultured under any suitable conditions as would be understood by one of ordinary skill in the art. For example, any media, temperature, and incubation conditions known in the art may be used. For host cells carrying an inducible vector, cells may be cultured with an appropriate inducible agent to promote expression.

[0187] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a BAS. In some embodiments, a heterologous polynucleotide encoding a BAS comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, identical to any one of SEQ ID NOs: 1500-1755. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a BAS that comprises the amino acid sequence of any one of SEQ ID NOs: 59-314.

[0188] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a C16C28 oxidase. In some embodiment, a heterologous polynucleotide encoding a C16C28 oxidase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, identical to any one of SEQ ID NOs: 1756-1759. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a C16C28 oxidase that comprises the amino acid sequence of any one of SEQ ID NOs: 315-318.

[0189] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a C28 oxidase. In some embodiments, a heterologous polynucleotide encoding a C28 oxidase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, identical to any one of SEQ ID NOs: 1756 and 1759-1761. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a C28 oxidase that comprises the amino acid sequence of any one of SEQ ID NOs: 315 and 318-320.

[0190] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a C23 oxidase. In some embodiment, a heterologous polynucleotide encoding a C23 oxidase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 1762-1776. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a C23 oxidase that comprises the amino acid sequence of any one of SEQ ID NOs: 321-335.

[0191] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a CPR. In some embodiments, a heterologous polynucleotide encoding a CPR comprises a sequence that is at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 1777-1818. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a CPR that comprises the amino acid sequence of any one of SEQ ID NOs: 336-377.

[0192] In some embodiments, a host cell comprises a heterologous polynucleotide encoding an MSBP. In some embodiments, a heterologous polynucleotide encoding a MSBP comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 1819-1857. In some embodiments, a host cell comprises a heterologous polynucleotide encoding an MSBP that comprises the amino acid sequence of any one of SEQ ID NOs: 378-416.

[0193] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a GlcAT. In some embodiments, a heterologous polynucleotide encoding a GlcAT comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 1858-1882. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a GlcAT that comprises the amino acid sequence of any one of SEQ ID NO: 417-441.

[0194] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a GalT. In some embodiments, a heterologous polynucleotide encoding a GalT comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 1883-1978. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a GalT that comprises the amino acid sequence of any one of SEQ ID NO: 442-537.

[0195] In some embodiments, a host cell comprises a heterologous polynucleotide encoding an XylT. In some embodiments, a heterologous polynucleotide encoding a XylT comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ IDNOs: 1925, 1926, 1932, 1935, 1943, 1945, 1958, and 1979-2013. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a XylT that comprises the amino acid sequence of any one of SEQ ID NOs: 484, 485, 491, 494, 502, 504, 517, and 538-572.

[0196] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a UGD. In some embodiments, a heterologous polynucleotide encoding an UGD comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 2153-2594. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a UGD that comprises the amino acid sequence of any one of SEQ ID NOs: 46-56, and 712-1142.

[0197] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a UXS. In some embodiments, a heterologous polynucleotide encoding an UXS comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 2014-2152. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a UXS that comprises the amino acid sequence of any one of SEQ ID NOs: 573-711.

[0198] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a phosphotransferase. In some embodiments, a heterologous polynucleotide encoding a phosphotransferase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 2595-2613. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a phosphotransferase that comprises the amino acid sequence of any one of SEQ ID NOs: 1239-1257.

[0199] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a UDP-glycosyltransf erase. In some embodiments, a heterologous polynucleotide encoding an UDP-glycosyltransferase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 2614-2639. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a UDP-glycosyltransferase that comprises the amino acid sequence of any one of SEQ ID NOs: 1258-1283.

[0200] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a methyltransferase. In some embodiments, a heterologous polynucleotide encoding a methyltransferase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 2640-2645. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a methyltransferase that comprises the amino acid sequence of any one of SEQ ID NOs: 1284-1289.

[0201] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a sulfotransferase. In some embodiments, a heterologous polynucleotide encoding a sulfotransferase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 2646-2663. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a sulfotransferase that comprises the amino acid sequence of any one of SEQ ID NOs: 1290-1307.

[0202] In some embodiments, a host cell comprises a heterologous polynucleotide encoding an acyltransferase. In some embodiments, a heterologous polynucleotide encoding an acyltransferase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 2664-2810. In some embodiments, a host cell comprises a heterologous polynucleotide encoding an acyltransferase that comprises the amino acid sequence of any one of SEQ ID NOs: 1308-1454.

[0203] In some embodiments, a host cell comprises a heterologous polynucleotide encoding a glycosyl hydrolase. In some embodiments, a heterologous polynucleotide encoding a glycosyl hydrolase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 2811-2828. In some embodiments, a host cell comprises a heterologous polynucleotide encoding a glycosyl hydrolase that comprises the amino acid sequence of any one of SEQ ID NOs: 1455-1472.

[0204] In some embodiments, a host cell comprises a heterologous polynucleotide encoding an esterase. In some embodiments, a heterologous polynucleotide encoding an esterase comprises a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 2829-2855. In some embodiments, a host cell comprises a heterologous polynucleotide encoding an esterase that comprises the amino acid sequence of any one of SEQ ID NOs: 1473-1499. In some embodiments, a host cell comprises one or more heterologous polynucleotides collectively encoding:

[0205] a) a BAS comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 59-314;

[0206] b) a C16C28 oxidase comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 315-318;

[0207] c) a C23 oxidase comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 321-335;

[0208] d) a CPR comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 336-377; and

[0209] e) an MSBP comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 378-416.

[0210] In some embodiments, the cell is capable of producing 2,3-oxidosqualene.

[0211] In some embodiments, a host cell comprises one or more heterologous polynucleotides collectively encoding:

[0212] a) a GlcAT comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 417-441;

[0213] b) a GalT comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 442-537;

[0214] c) a XylT comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 484, 485, 491, 494, 502, 504, 517, and 538-572;

[0215] d) a UGD comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 46-56, and 712-1142; and

[0216] e) a UXS comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 573-711.

[0217] In some embodiments, the host cell is capable of producing QA.

[0218] In some embodiments, a host cell comprises one or more heterologous polynucleotides collectively encoding:

[0219] a) a phosphotransferase comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 1239-1257;

[0220] b) an UDP-glycosyltransferase comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 1258-1283;

[0221] c) a methyltransferase comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 1284-1289;

[0222] d) a sulfotransferase comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 1290- 1307;

[0223] e) an acyltransferase comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 1308- 1454;

[0224] f) a glycoside hydrolase comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 1455- 1472; and

[0225] g) an esterase comprising a sequence that is at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence of any one of SEQ ID NOs: 1473-1499.

[0226] Methods

[0227] Aspects of the disclosure provide, at least in part, methods for producing a bioproduct using host cells associated with the disclosure. In some embodiments, cells associated with the disclosure are used to make QA, prosapogenin, or a precursor or derivative of any of the foregoing. In some embodiments, a method for producing QA, prosapogenin, or a precursor or derivative thereof, comprises culturing a host cell of the present disclosure.

[0228] Any of the cells disclosed in this application can be cultured in media of any type (rich or minimal) and any composition prior to, during, and / or after contact and / or integration of a nucleic acid. The conditions of the culture or culturing process can be optimized through routine experimentation as would be understood by one of ordinary skill in the art. In some embodiments, the selected media is supplemented with various components. In some embodiments, the concentration and amount of a supplemental component is optimized. In some embodiments, other aspects of the media and growth conditions (e.g., pH, temperature, etc.) are optimized through routine experimentation. In some embodiments, the frequency that the media is supplemented with one or more supplemental components, and the amount of time that the cell is cultured, is optimized.

[0229] Culturing of the cells described in this application can be performed in culture vessels known and used in the art. In some embodiments, an aerated reaction vessel (e.g., a stirred tank reactor) is used to culture the cells. In some embodiments, a bioreactor or fermenter is used to culture the cells. Thus, in some embodiments, the cells are used in fermentation. As used in this application, the terms “bioreactor” and “fermentor” are interchangeably used and refer to an enclosure, or partial enclosure, in which a biological, biochemical and / or chemical reaction takes place, involving a living organism, part of a living organism, or purified proteins. Any type of bioreactor or fermenter known in the art may be compatible with aspects of the disclosure, including large or industrial scale bioreactors such as those with volumes in the range of liters or hundreds or thousands of liters or more.

[0230] In some embodiments, fermentation processes are operated in continuous, semi-continuous or non-continuous modes. Non-limiting examples of operation modes are batch, fed batch, extended batch, repetitive batch, draw / fill, rotating-wall, spinning flask, and / or perfusion mode of operation. In some embodiments, a bioreactor allows continuous or semi-continuous replenishment of the substrate stock, for example a carbohydrate source and / or continuous or semi-continuous separation of the product, from the bioreactor.

[0231] In some embodiments, the method involves batch fermentation (e.g., shake flask fermentation). General considerations for batch fermentation (e.g., shake flask fermentation) include the level of oxygen and glucose. For example, batch fermentation (e.g., shake flask fermentation) may be oxygen and glucose limited, so in some embodiments, the capability of a strain to perform in a well-designed fed-batch fermentation is underestimated. Also, the final product (e.g., QA, prosapogenin, or precursor or derivative thereof) may display some differences from the substrate (e.g., QA, prosapogenin, or precursor or derivative thereof) in terms of solubility, toxicity, cellular accumulation and secretion and in some embodiments can have different fermentation kinetics.

[0232] In some embodiments, the method further comprises isolating QA, prosapogenin, or a precursor or derivative thereof from the cell culture. Triterpene glycosides produced by any of the host cells disclosed in this application may be identified and extracted using any method known in the art. Mass spectrometry (e.g., LC-MS, GC-MS) is a non-limiting example of a method for identification and may be used to help identify a compound of interest.

[0233] Aspects of the present disclosure provide, at least in part, methods of modifying QS-21 in vitro. In some embodiments, QS-21 is modified with a purified enzyme. In some embodiments, the phosphotransferase, the UDP-glycosyltransferase, the methyltransferase, the sulfotransferase, the acyltransferase, the glycosyl hydrolase, and / or the esterase for modifying QS-21 is a purified enzyme.

[0234] In some embodiments, QS-21 is modified by contact with a cell lysate. In some embodiments, the cell lysate comprises a phosphotransferase, a UDP-glycosyltransferase, a methyltransferase, a sulfotransferase, an acyltransferase, a glycosyl hydrolase, and / or an esterase. In some embodiments, the cell lysate is produced by lysing a population of host cells, wherein the host cells comprise one or more of: a) a heterologous polynucleotide encoding a phosphotransferase; b) a heterologous polynucleotide encoding a UDP-glycosyltransferase; c) a heterologous polynucleotide encoding a methyltransferase; d) a heterologous polynucleotide encoding a sulfotransferase; e) a heterologous polynucleotide encoding an acyltransferase; f) a heterologous polynucleotide encoding a glycosyl hydrolase; and g) a heterologous polynucleotide encoding an esterase. Derivatives of triterpene glycosides

[0235] Any triterpene glycoside disclosed herein can be used to produce a derivative thereof. Derivatives of triterpene glycosides can be produced from triterpene glycosides via any method, including but not limited to enzymatic and chemical methods, and / or any method known in the art. In some embodiments, any host cell or enzyme or method described herein can be used to produce a triterpene glycoside, which can in turn be used to produce a derivative of the triterpene glycoside.

[0236] In some embodiments, a derivative of a triterpene glycoside (including but not limited to QS-7 and QS-21) can be produced from a triterpene glycoside by removal of, addition of, and / or substitution of at least one atom (or at least one functional group) with a moiety such as: a sugar, a phosphate (e.g., to produce a phosphorylated triterpene glycoside such as a monophosphorylated triterpene glycoside), an acyl group (e.g. acetyl group, malonyl group, benzoyl group) (e.g. to produce an acetylated triterpene glycoside such as a mono-acetylated triterpene glycoside), a functional group comprising a sulfur atom (e.g., in the form of a sulfate) (e.g., to form a sulfated triterpene glycoside such as a mono-sulfated triterpene glycoside), an methyl group (e.g. to produce an methylated triterpene glycoside such as a mono-methylated triterpene glycoside),. In various embodiments, a derivative of a triterpene glycoside can comprise an additional sugar, a phosphate, an acetyl group, a sulfur-comprising group (e.g., a sulfate), a lipophilic moiety, a lipophobic moiety, a fatty acid, an alkyl group (e.g., which may be straight, branched, multiply-branched and / or comprise a circular moiety), an alkenyl group, an alkynyl group, an amine, an amide or other nitrogen-comprising group (e.g., -NH-), an alcohol or other oxygen-comprising group (e.g., -O, -CO-, etc.), a thiol, an aldehyde, a diethyl ether, a ketone, a dimethyl sulfide, a carboxylic acid, a carbonyl group, an acetyl chloride, an acetic acid group, a haloalkane or other halide-comprising group, an acyl halide, an ether, a nitrile, an ester, a protein, an antibody, and / or a nucleic acid.

[0237] In some embodiments, a derivative of a triterpene glycoside such as QS-7 or QS-21 can be another known triterpene glycoside. In some embodiments, a derivative of a triterpene glycoside is an enantiomerically triterpene glycoside (e.g., an enantiomerically pure QS-7 or QS-21).

[0238] Various other known triterpene glycosides and derivatives of various triterpene glycosides are described in the art, for example: Avilov et al. 2000 J. Nat. Products. 63: 65-71; Kalinin et al. 2021 Nat. Prod. Comm. 16: 1-24; Kalinin et al. 1996 Toxicon 34: 475; and Perrone et al. 2007 J. Nat. Prod. 70: 584.

[0239] Compositions

[0240] Aspects of the disclosure provide, at least in part, compositions comprising phosphorylated QS-21, glycosylated QS-21, methylated QS-21, sulfated QS-21, acylated QS-21, glycoside hydrolase deglycosylated QS-21, or esterase hydrolyzed QS-21.

[0241] Aspects of the disclosure provide, at least in part, compositions comprising a monophosphorylated QS-21, a mono-glucosylated QS-21, a mono-methylated QS-21, a di-methylated QS-21, a mono-sulfated QS-21, a mono-acetylated QS-21, a di-acetylated QS-21, a mono-malonylated QS-21, a mono-benzoylated QS-21, a glycoside hydrolase deglycosylated QS-21, or an esterase hydrolyzed QS-21.

[0242] In some aspects, the disclosure provides an adjuvant comprising a mono-phosphorylated QS-21, a mono-glucosylated QS-21, a mono-methylated QS-21, a di-methylated QS-21, a mono-sulfated QS-21, a mono-acetylated QS-21, a di-acetylated QS-21, a mono-malonylated QS-21, a mono-benzoylated QS-21, a glycoside hydrolase deglycosylated QS-21, or an esterase hydrolyzed QS-21.

[0243] Aspects of the disclosure provide, at least in part, compositions comprising a quillaic acid (QA), a 16a-OH-gypsogenic acid, a QA-C3-GlcA, a QA-C3-GlcA-Gal, a QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin), a mono-phosphorylated QS-21, a mono-glucosylated QS-21, a mono-methylated QS-21, a di-methylated QS-21, a mono-sulfated QS-21, a mono-acetylated QS-21, a di-acetylated QS-21, a mono-malonylated QS-21, a mono-benzoylated QS-21, a glycoside hydrolase deglycosylated QS-21, or an esterase hydrolyzed QS-21 produced by a host cell described herein.

[0244] In some aspects, the disclosure provides a precursor to an adjuvant comprising a triterpene glycoside produced by a host cell described herein. In some embodiments, the triterpene glycoside is quillaic acid (QA), 16a-OH-gypsogenic acid, QA-C3-GlcA; QA-C3-GlcA-Gal, or QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin) produced by a host cell described herein. In some embodiments, a precursor to an adjuvant is glycosylated and acetylated (once or multiple times) to form an adjuvant. In some embodiments, QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin) is glycosylated and acetylated to form an adjuvant. In some aspects, the disclosure provides therapeutic compositions comprising a therapeutically relevant agent and an adjuvant, wherein the adjuvant is produced from a precursor, wherein the precursor is a triterpene glycoside produced by a host cell described herein. In some embodiments, the triterpene glycoside is quillaic acid (QA), 16a-0H-gypsogenic acid, QA-C3-GlcA; QA-C3-GlcA-Gal, or QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin) produced by a host cell described herein.

[0245] The phraseology and terminology used in this application is for the purpose of description and should not be regarded as limiting. The use of terms such as “including,” “comprising,” “having,” “containing,” “involving,” and / or variations thereof in this application, is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.

[0246] The present invention is further illustrated by the following Examples, which in no way should be construed as further limiting. The entire contents of all of the references (including literature references, issued patents, published patent applications, and co-pending patent applications) cited throughout this application are hereby expressly incorporated by reference.

[0247] EXAMPLES

[0248] Example 1: Functional expression of quillaic acid (QA) biosynthesis pathway genes in engineered Saccharomyces cerevisiae host cells

[0249] This example describes an engineered S. cerevisiae strain capable of producing quillaic acid (QA). The QA biosynthetic pathway is composed of a P-amyrin synthase (BAS) and multiple cytochrome P450 (CYP) oxidases (C16, C23, and C28 oxidases), and may include accessory proteins not directly involved in catalytic activities, such as a cytochrome P450 reductase (CPR) and / or a membrane steroid binding protein (MSBP). CPR and MSBP are accessory proteins. The CPR is a redox partner of cytochrome P450 oxidases. The MSBP is believed to reorganize the oxidases in the endoplasmic reticulum membrane. A P-amyrin synthase (BAS) may be used to produce P-amyrin by converting 2,3-oxidosqualene to P-amyrin. Multiple cytochrome P450 (CYP) oxidases (P450 C16, P450 C23, P450 C16C28, and / or P450 C28 oxidases) can be used to oxidize the C16 and C28 carbons of P-amyrin to hydroxyl groups, and oxidize the C23 carbon of P-amyrin to an aldehyde group. Further, a cytochrome P450 reductase (CPR) can act as a redox partner for P450 oxidases (Cl 6, C23, C16C28, and / or C28 oxidases). A membrane steroid binding protein (MSBP) may serve as a scaffold protein that physically interacts with the P450 oxidases (C16, C23, C16C28, and / or C28 oxidases) and the CPR.

[0250] Specific homologs of these genes were used in this Example: A BAS from Quillaja saponaria (QsbAS; SEQ ID NO: 1), a MSBP from Siraitia grosvenorii (CB5; SEQ ID NO: 3) (see, e.g., WO2022192688A1), a CPR from Arabidopsis thaliana (AtATR2; SEQ ID NO: 9), a C23 oxidase from Psammosilene tunicoides (CYP72A567; SEQ ID NO: 12) (see, e.g., Li et al. (2021) ACS Synth. Biol., 10: 1874-81), a C16C28 oxidase from Psammosilene tunicoides (CYP716A262; SEQ ID NO: 41) (see, e.g., Li et al. (2021)).

[0251] All heterologous genes were codon-optimized for expression in S. cerevisiae and placed under the control of galactose-inducible promoters.

[0252] The expression constructs were transformed into a S. cerevisiae host strain with increased mevalonate flux by overexpressing the native S. cerevisiae erg 10, ergl3, hmgl, erg 12, erg8, ergl9, idil and erg20. Single colonies resulting from transformation were inoculated in preculture media (10 g / L yeast extract, 20 g / L bacto peptone, 2% glucose) and grown in a shaking incubator at 30°C for 48 hours at 1000 rpm. After 48 hours, 20 pL aliquot of preculture was added to each well of a 96-well deepwell plate along with 500 pL of assay media (10 g / L yeast extract, 20 g / L bacto peptone, 2% galactose, 2% raffinose, 6 g / L MgSCU, Bird’s trace metals, and Bird’s trace vitamins). The plate was incubated at 30°C at 1,000 rpm for 3 days. After 3 days, 100 pL of production cultures was diluted with 900 pL of 80% methanol and analyzed for quillaic acid titer by LC-MS.

[0253] Table 1. Quillaic acid titers from expressing QA biosynthetic pathway in S. cerevisiae.

[0254] Expressed Genes QA [mg / L]

[0255] Parent chassis 0

[0256] QsbAS 0

[0257] QsbAS, CB5 0

[0258] QsbAS, AtATR2, CYP72A567, CYP16A262 2.90

[0259] QsbAS, CB5, AtATR2, CYP72A567, CYP16A262 26.59

[0260]

[0261] In Examples 3-8, libraries were screened for different enzymes that had the same (or similar) activities. Example 2: Production of UDP-Sugars (Glucuronic acid, Xylose, Rhamnose, and Fucose) Non-native to Yeast

[0262] This example describes an engineered S. cerevisiae strain capable of producing uracil diphosphate sugar (UDP-sugar), in particular, producing UDP-D-glucuronic acid (UDP-GlcA) and UDP-D-xylose (UDP-Xyl).

[0263] The UDP-GlcA biosynthetic pathway is composed of a UDP-glucose dehydrogenase (UGD) selected from the list below. Previous examples of heterologous UDP-sugar production in yeast used plant derived UGD (such as Arabidopsis thaliana UDP-glucose dehydrogenase AtUGD). Specifically, AtUGD was used as the seed sequence to search in public database for prokaryote derived UGD and tested the following homologs for UDP-GlcA production in yeast: a UGD from Bacteroidales bacterium (BbacUGD; SEQ ID NO: 46), a UGD from Blautia luti (BlutUGD; SEQ ID NO: 47), a UGD from Escherichia coli (EcolUGD; SEQ ID NO: 48), a UGD from Firmicutes bacterium (FbacUGD; SEQ ID NO: 50), a UGD from Holdemania fdiformis (HfilUGD; SEQ ID NO: 51), a UGD from Lachnospiraceae bacterium (LbacUGD; SEQ ID NO: 52).

[0264] The UDP-Xyl biosynthetic pathway is composed of a UDP -xylose synthase (UXS), and a UXS from Arabidopsis thaliana (AtUXS; SEQ ID NO: 57) was used.

[0265] All heterologous genes were codon-optimized for expression in S. cerevisiae and placed behind a galactose-inducible promoter. The expression constructs were transformed in QA producing S. cerevisiae strain and was cultivated using procedure as described in Example 1. After 3 days cultivation, cell pellets from 500 pL production cultures were diluted with 150 pL of 100% methanol and analyzed for UDP-sugar titer by LC-MS.

[0266] Table 2. UDP-GlcA and UDP-Xyl titers from expressing UGD and UXS in S. cerevisiae Expressed Genes UDP-GlcA [mg / L] UDP-Xyl [mg / L] Parent Chassis 0 0

[0267] BbacUGD, AtUXS 0.20 66.13

[0268] BlutUGD, AtUXS 0.23 48.45

[0269] EcolUGD, AtUXS 0.17 37.57

[0270] FbacUGD, AtUXS 0.17 45.11

[0271]

[0272] HfilUGD, AtUXS 0.19 35.54

[0273] LbacUGD, AtUXS 0.23 93.61

[0274]

[0275] Example 3: Generation and Screening of BAS Library

[0276] This example describes the identification of BAS variants that were functional in the described QA biosynthesis pathway in Example 1 for quillaic acid (QA) production.

[0277] To identify BAS variants that were functional in the described QA biosynthesis pathway for quillaic acid production, a metagenomic library of approximately 800 variants was generated based on the QsbAS (SEQ ID NO: 1) and the Glycyrrhiza glabra BAS (GgbAS; SEQ ID NO: 2) sequences.

[0278] The library was transformed, along with the remaining QA biosynthetic genes listed in Example 1, into mevalonate production strain as described in Example 1.

[0279] To initiate cell growth in preparation for screening, glycerol stocks of the BAS variant transformants were thawed at room temperature. Then, 50 pL of preculture media (10 g / L yeast extract, 20 g / L bacto peptone, 2% glucose) was added to each well of a 384-well plate, and 2 pL aliquots of glycerol stock was added to each well; the plate was incubated at 30°C at 1,200 rpm and 80% humidity for 1 day. In the next step, 2 pL aliquots of the preculture sample were transferred from wells of the first plate to wells of a second plate each containing 500 pL of preculture media; the plate was incubated at 30°C at 1,000 rpm and 80% humidity for 1 day. Next, 20 pL aliquots of the preculture 2 sample was transferred wells of the second plate to wells of a third plate each containing 35 pL of production media (10 g / L yeast extract, 20 g / L bacto peptone, 2% galactose, 2% raffinose, 6 g / L MgSCU, Bird’s trace metals, and Bird’s trace vitamins); the plate was incubated at 30°C at 1,000 rpm and 80% humidity for 2 days. After 2 days, 10 pL of production cultures was diluted with 40 pL of 80% methanol, the samples were analyzed for quillaic acid titer by LC-MS.

[0280] Table 3. Relative QA titers from BAS Library Screen

[0281] BAS Candidates QA Production Relative to Positive Control (%)

[0282] QsbAS (SEQ ID NO: 1) 1.00

[0283] 6603218 (SEQ ID NO: 59) 1.46

[0284]

[0285] 6603201 (SEQ ID NO: 60) 1.41 6603220 (SEQ ID NO: 61) 1.41 6603210 (SEQ ID NO: 62) 1.40 6603249 (SEQ ID NO: 63) 1.40 6603119 (SEQ ID NO: 64) 1.34 6603063 (SEQ ID NO: 65) 1.31 6603225 (SEQ ID NO: 66) 1.31 6603131 (SEQ ID NO: 67) 1.30 6603159 (SEQ ID NO: 68) 1.30 6603188 (SEQ ID NO: 69) 1.30 6603105 (SEQ ID NO: 70) 1.24 6603121 (SEQ ID NO: 71) 1.22 6603257 (SEQ ID NO: 72) 1.21 6603180 (SEQ ID NO: 73) 1.21 6603093 (SEQ ID NO: 74) 1.20 6603212 (SEQ ID NO: 75) 1.19 6603040 (SEQ ID NO: 76) 1.19 6603080 (SEQ ID NO: 77) 1.16 6603107 (SEQ ID NO: 78) 1.16 6603048 (SEQ ID NO: 79) 1.15 6603056 (SEQ ID NO: 80) 1.14 6603195 (SEQ ID NO: 81) 1.14 6603155 (SEQ ID NO: 82) 1.13 6603233 (SEQ ID NO: 83) 1.12 6603169 (SEQ ID NO: 84) 1.12 6603127 (SEQ ID NO: 85) 1.09 6603331 (SEQ ID NO: 86) 1.09

[0286]

[0287] 6603192 (SEQ ID NO: 87) 1.09 6603039 (SEQ ID NO: 88) 1.07 6603194 (SEQ ID NO: 89) 1.07 6603089 (SEQ ID NO: 90) 1.06 6603353 (SEQ ID NO: 91) 1.06 6603112 (SEQ ID NO: 92) 1.05 6603060 (SEQ ID NO: 93) 1.03 6603204 (SEQ ID NO: 94) 1.03 6603054 (SEQ ID NO: 95) 1.01 6603049 (SEQ ID NO: 96) 1.00 6603232 (SEQ ID NO: 97) 0.99 6603096 (SEQ ID NO: 98) 0.99 6603070 (SEQ ID NO: 99) 0.99 6603237 (SEQ ID NO: 100) 0.99 6603083 (SEQ ID NO: 101) 0.99 6603146 (SEQ ID NO: 102) 0.99 6603081 (SEQ ID NO: 103) 0.99 6603227 (SEQ ID NO: 104) 0.98 6603052 (SEQ ID NO: 105) 0.97 6603196 (SEQ ID NO: 106) 0.97 6603161 (SEQ ID NO: 107) 0.97 6603345 (SEQ ID NO: 108) 0.97 6603115 (SEQ ID NO: 109) 0.96 6603260 (SEQ ID NO: 110) 0.94 6603158 (SEQ ID NO: 111) 0.94 6603064 (SEQ ID NO: 112) 0.93 6603128 (SEQ ID NO: 113) 0.93

[0288]

[0289] 6603164 (SEQ ID NO: 114) 0.92 6603071 (SEQ ID NO: 115) 0.92 6603156 (SEQ ID NO: 116) 0.91 6603130 (SEQ ID NO: 117) 0.91 6603067 (SEQ ID NO: 118) 0.89 6603167 (SEQ ID NO: 119) 0.88 6603219 (SEQ ID NO: 120) 0.87 6603176 (SEQ ID NO: 121) 0.87 6603066 (SEQ ID NO: 122) 0.86 6603123 (SEQ ID NO: 123) 0.84 6603189 (SEQ ID NO: 124) 0.83 6603055 (SEQ ID NO: 125) 0.83 6603150 (SEQ ID NO: 126) 0.82 6603254 (SEQ ID NO: 127) 0.79 6603114 (SEQ ID NO: 128) 0.79 6603148 (SEQ ID NO: 129) 0.78 6603166 (SEQ ID NO: 130) 0.77 6603206 (SEQ ID NO: 131) 0.77 6603094 (SEQ ID NO: 132) 0.73 6603216 (SEQ ID NO: 133) 0.72 6603163 (SEQ ID NO: 134) 0.71 6603157 (SEQ ID NO: 135) 0.71 6603172 (SEQ ID NO: 136) 0.68 6603241 (SEQ ID NO: 137) 0.68 6603269 (SEQ ID NO: 138) 0.68 6603414 (SEQ ID NO: 139) 0.67 6603050 (SEQ ID NO: 140) 0.66

[0290]

[0291] 6603207 (SEQ IDNO: 141) 0.65 6603072 (SEQ ID NO: 142) 0.64 6603252 (SEQ ID NO: 143) 0.63 6603117 (SEQ ID NO: 144) 0.62 6603240 (SEQ ID NO: 145) 0.61 6603116 (SEQ ID NO: 146) 0.61 6603099 (SEQ ID NO: 147) 0.61 6603129 (SEQ IDNO: 148) 0.60 6603250 (SEQ ID NO: 149) 0.59 6603247 (SEQ ID NO: 150) 0.59 6603058 (SEQ ID NO: 151) 0.56 6603182 (SEQ IDNO: 152) 0.55 6603186 (SEQ IDNO: 153) 0.53 6603062 (SEQ ID NO: 154) 0.52 6603208 (SEQ IDNO: 155) 0.52 6603179 (SEQ IDNO: 156) 0.52 6603217 (SEQ ID NO: 157) 0.49 6603197 (SEQ IDNO: 158) 0.48 6603302 (SEQ IDNO: 159) 0.46 6603076 (SEQ ID NO: 160) 0.46 6603057 (SEQ ID NO: 161) 0.46 6603177 (SEQ IDNO: 162) 0.46 6603059 (SEQ ID NO: 163) 0.45 6603113 (SEQ IDNO: 164) 0.45 6603069 (SEQ ID NO: 165) 0.45 6603174 (SEQ IDNO: 166) 0.43 6603160 (SEQ IDNO: 167) 0.43

[0292]

[0293] 6603320 (SEQ IDNO: 168) 0.42 6603223 (SEQ ID NO: 169) 0.41 6603077 (SEQ ID NO: 170) 0.41 6603782 (SEQ IDNO: 171) 0.41 6603092 (SEQ ID NO: 172) 0.40 6603140 (SEQ IDNO: 173) 0.40 6603122 (SEQ IDNO: 174) 0.38 6603224 (SEQ ID NO: 175) 0.38 6603279 (SEQ ID NO: 176) 0.38 6603796 (SEQ ID NO: 177) 0.36 6603051 (SEQ IDNO: 178) 0.36 6603079 (SEQ ID NO: 179) 0.35 6603754 (SEQ IDNO: 180) 0.34 6603209 (SEQ IDNO: 181) 0.34 6603065 (SEQ IDNO: 182) 0.34 6603723 (SEQ IDNO: 183) 0.34 6603103 (SEQ IDNO: 184) 0.33 6603187 (SEQ IDNO: 185) 0.33 6603409 (SEQ ID NO: 186) 0.31 6603397 (SEQ IDNO: 187) 0.29 6603786 (SEQ IDNO: 188) 0.29 6603416 (SEQ ID NO: 189) 0.28 6603748 (SEQ ID NO: 190) 0.28 6603749 (SEQ IDNO: 191) 0.27 6603770 (SEQ ID NO: 192) 0.27 6603459 (SEQ ID NO: 193) 0.27 6603488 (SEQ ID NO: 194) 0.26

[0294]

[0295] 6603810 (SEQ ID NO: 195) 0.25 6603061 (SEQ ID NO: 196) 0.24 6603313 (SEQ ID NO: 197) 0.23 6603774 (SEQ ID NO: 198) 0.23 6603118 (SEQ ID NO: 199) 0.23 6603415 (SEQ ID NO: 200) 0.23 6603746 (SEQ ID NO: 201) 0.23 6603494 (SEQ ID NO: 202) 0.22 6603795 (SEQ ID NO: 203) 0.22 6603759 (SEQ ID NO: 204) 0.22 6603135 (SEQ ID NO: 205) 0.22 6603787 (SEQ ID NO: 206) 0.21 6603282 (SEQ ID NO: 207) 0.21 6603477 (SEQ ID NO: 208) 0.21 6603090 (SEQ ID NO: 209) 0.20 6603088 (SEQ ID NO: 210) 0.20 6603826 (SEQ ID NO: 211) 0.20 6603781 (SEQ ID NO: 212) 0.19 6603333 (SEQ ID NO: 213) 0.18 6603513 (SEQ ID NO: 214) 0.18 6603463 (SEQ ID NO: 215) 0.18 6603470 (SEQ ID NO: 216) 0.18 6603803 (SEQ ID NO: 217) 0.18 6603680 (SEQ ID NO: 218) 0.17 6603426 (SEQ ID NO: 219) 0.17 6603799 (SEQ ID NO: 220) 0.17 6603047 (SEQ ID NO: 221) 0.17

[0296]

[0297] 6603398 (SEQ ID NO: 222) 0.16 6603456 (SEQ ID NO: 223) 0.16 6603731 (SEQ ID NO: 224) 0.16 6603831 (SEQ ID NO: 225) 0.16 6603823 (SEQ ID NO: 226) 0.16 6603504 (SEQ ID NO: 227) 0.15 6603747 (SEQ ID NO: 228) 0.15 6603429 (SEQ ID NO: 229) 0.15 6603716 (SEQ ID NO: 230) 0.14 6603773 (SEQ ID NO: 231) 0.14 6603503 (SEQ ID NO: 232) 0.14 6603521 (SEQ ID NO: 233) 0.14 6603779 (SEQ ID NO: 234) 0.13 6603310 (SEQ ID NO: 235) 0.13 6603097 (SEQ ID NO: 236) 0.13 6603563 (SEQ ID NO: 237) 0.13 6603519 (SEQ ID NO: 238) 0.13 6603454 (SEQ ID NO: 239) 0.13 6603347 (SEQ ID NO: 240) 0.12 6603489 (SEQ ID NO: 241) 0.12 6603287 (SEQ ID NO: 242) 0.12 6603412 (SEQ ID NO: 243) 0.12 6603410 (SEQ ID NO: 244) 0.12 6603413 (SEQ ID NO: 245) 0.12 6603551 (SEQ ID NO: 246) 0.12 6603790 (SEQ ID NO: 247) 0.12 6603213 (SEQ ID NO: 248) 0.12

[0298]

[0299] 6603423 (SEQ ID NO: 249) 0.11 6603335 (SEQ ID NO: 250) 0.11 6603673 (SEQ ID NO: 251) 0.11 6603775 (SEQ ID NO: 252) 0.11 6603403 (SEQ ID NO: 253) 0.11 6603075 (SEQ ID NO: 254) 0.10 6603570 (SEQ ID NO: 255) 0.10 6603211 (SEQ ID NO: 256) 0.10 6603451 (SEQ ID NO: 257) 0.10 6603490 (SEQ ID NO: 258) 0.09 6603458 (SEQ ID NO: 259) 0.09 6603383 (SEQ ID NO: 260) 0.08 6603473 (SEQ ID NO: 261) 0.08 6603450 (SEQ ID NO: 262) 0.08 6603137 (SEQ ID NO: 263) 0.08 6603438 (SEQ ID NO: 264) 0.07 6603838 (SEQ ID NO: 265) 0.07 6603315 (SEQ ID NO: 266) 0.07 6603154 (SEQ ID NO: 267) 0.07 6603288 (SEQ ID NO: 268) 0.07 6603399 (SEQ ID NO: 269) 0.06 6603085 (SEQ ID NO: 270) 0.06 6603453 (SEQ ID NO: 271) 0.06 6603325 (SEQ ID NO: 272) 0.05 6603327 (SEQ ID NO: 273) 0.05 6603695 (SEQ ID NO: 274) 0.05 6603231 (SEQ ID NO: 275) 0.05

[0300]

[0301] 6603340 (SEQ ID NO: 276) 0.05 6603149 (SEQ ID NO: 277) 0.05 6603702 (SEQ ID NO: 278) 0.05 6603406 (SEQ ID NO: 279) 0.05 6603692 (SEQ ID NO: 280) 0.05 6603181 (SEQ ID NO: 281) 0.05 6603396 (SEQ ID NO: 282) 0.04 6603743 (SEQ ID NO: 283) 0.04 6603511 (SEQ ID NO: 284) 0.04 6603703 (SEQ ID NO: 285) 0.04 6603144 (SEQ ID NO: 286) 0.04 6603430 (SEQ ID NO: 287) 0.04 6603664 (SEQ ID NO: 288) 0.04 6603717 (SEQ ID NO: 289) 0.04 6603303 (SEQ ID NO: 290) 0.04 6603449 (SEQ ID NO: 291) 0.04 6603706 (SEQ ID NO: 292) 0.03 6603446 (SEQ ID NO: 293) 0.03 6603263 (SEQ ID NO: 294) 0.03 6603500 (SEQ ID NO: 295) 0.03 6603441 (SEQ ID NO: 296) 0.03 6603527 (SEQ ID NO: 297) 0.03 6603153 (SEQ ID NO: 298) 0.03 6603087 (SEQ ID NO: 299) 0.03 6603173 (SEQ ID NO: 300) 0.03 6603425 (SEQ ID NO: 301) 0.03 6603676 (SEQ ID NO: 302) 0.03

[0302]

[0303] 6603704 (SEQ ID NO: 303) 0.02

[0304] 6603259 (SEQ ID NO: 304) 0.02

[0305] 6603465 (SEQ ID NO: 305) 0.02

[0306] 6603322 (SEQ ID NO: 306) 0.02

[0307] 6603281 (SEQ ID NO: 307) 0.02

[0308] 6603724 (SEQ ID NO: 308) 0.02

[0309] 6603684 (SEQ ID NO: 309) 0.02

[0310] 6603401 (SEQ ID NO: 310) 0.02

[0311] 6603497 (SEQ ID NO: 311) 0.02

[0312] 6603143 (SEQ ID NO: 312) 0.02

[0313] 6603251 (SEQ ID NO: 313) 0.02

[0314] 6603280 (SEQ ID NO: 314) 0.02

[0315]

[0316] Example 4: Generation and Screening of Cl 6C28 oxidase Library

[0317] This example describes the identification of C16C28 oxidase variants that were functional in the described QA biosynthesis pathway in Example 1 for quillaic acid (QA) production.

[0318] To identify C16C28 oxidase variants that are functional in the described QA biosynthesis pathway for quillaic acid production, a metagenomic library of approximately 2400 variants was generated based on the Quillaja saponaria C28 oxidase (CYP716A224; SEQ ID NO: 16), the Arabidopsis thaliana C28 oxidase (CYP716A1; SEQ ID NO: 17), the Medicago truncatula C28 oxidase (CYP716A12; SEQ ID NO: 18), the two Vitis vinifera C28 oxidase (CYP716A15; SEQ ID NO: 19; and CYP716A17; SEQ ID NO: 20), the two Solanum lycopersicum C28 oxidase (CYP716A44; SEQ ID NO: 21; and CYP716A46; SEQ ID NO: 22), the Panax ginseng C28 oxidase (CYP716A52v2; SEQ ID NO: 23), the Maesa lanceolata C28 oxidase (CYP716A75; SEQ ID NO: 24), the two Chenopodium quinoa C28 oxidase (CYP716A78; SEQ ID NO: 25; and CYP716A79; SEQ ID NO: 26), the two Barbarea vulgaris C28 oxidase (CYP716A80; SEQ ID NO: 27; and CYP716A81; SEQ ID NO: 28), the two Centella asiatica C28 oxidase (CYP716A83; SEQ ID NO: 29; and CYP716A86; SEQ ID NO: 30), the Cathar anthus roseus C28 oxidase (CYP716A154; SEQ ID NO: 31), the Aquilegia coerulea C28 oxidase (CYP716A110; SEQ ID NO: 32), the two Platycodon grandiflorus C28 oxidase (CYP716A140; SEQ ID NO: 33; and CYP716A141; SEQ ID NO: 34), the Glycyrrhiza uralensis C28 oxidase (CYP716A179; SEQ ID NO: 35), the two Ocimum basilicum C28 oxidase (CYP716A252; SEQ ID NO: 36; and CYP716A253; SEQ ID NO: 37), the Quillaja saponaria C16 oxidase (CYP716; SEQ ID NO: 38), the Maesa lanceolata C16 oxidase (CYP87D16; SEQ ID NO: 39), the Bupleurum falcatum C16 oxidase (CYP716Y1; SEQ ID NO: 40), and the Psammosdene tunicoides C16C28 oxidase (CYP716A262; SEQ ID NO: 41) sequences.

[0319] The library was transformed, along with the remaining QA biosynthetic genes listed in Example 1, into mevalonate production strain as described in Example 1. The QA production titer of the transformants were screened based on procedure described in Example 3.

[0320] Table 4. Relative QA titers from C16C28 Oxidase Library Screen

[0321] C16C28 Oxidase Candidates QA Production Relative to Positive Control (%)

[0322] CYP716A262 (SEQ ID NO: 41) 1.00

[0323] 6684675 (SEQ ID NO: 315) 0.32

[0324] 6683513 (SEQ ID NO: 316) 0.04

[0325] 6684095 (SEQ ID NO: 317) 0.02

[0326] 6683085 (SEQ ID NO: 318) 0.01

[0327]

[0328] Example 5: Generation and Screening of C28 oxidase Library

[0329] This example describes the identification of C28 oxidase variants that are functional in the described QA biosynthesis pathway in Example 1 for quillaic acid (QA) production.

[0330] To identify C28 oxidase variants that were functional in the described QA biosynthesis pathway for quillaic acid production, a metagenomic library of approximately 2400 variants as described in Example 4 was prepared.

[0331] The library was transformed, along with QsbAS, CB5, CPR, CYP716, and CYP72A567, into mevalonate production strain as described in Example 1.

[0332] To initiate cell growth in preparation for screening, glycerol stocks of the BAS variant transformants were thawed at room temperature. Next, 500 pL of preculture media (10 g / L yeast extract, 20 g / L bacto peptone, 2% glucose) was added to each well of a 96-well plate, and 5 pL aliquots of glycerol stock was added to each well; the plate was incubated at 30°C at 1,200 rpm and 80% humidity for 1 day. Then, 20 pL aliquots of the preculture 2 sample was transferred wells of the second plate to wells of a separate 96-well plate each containing 500 pL of growth media (10 g / L yeast extract, 20 g / L bacto peptone, 2% glucose, 6 g / L MgSC , Bird’s trace metals, and Bird’s trace vitamins); the plate was incubated at 30°C at 1,000 rpm and 80% humidity for 1 days. The next day, galactose and raffinose were added to each well to final concentration of 2% per component, and growth was continued at 30°C at 1,200 rpm and 80% humidity for 3 days. After 3 days, 10 pL of production cultures was diluted with 40 pL of 80% methanol, the samples were analyzed for quillaic acid titer by LC-MS.

[0333] Table 5. Relative QA titres from C28 Oxidase Library Screen

[0334] C28 Oxidase Candidates QA Production Relative to Positive Control (%)

[0335] CYP716A262 (SEQ ID NO: 41) 1.00

[0336] 6684675 (SEQ ID NO: 315) 0.57

[0337] 6683085 (SEQ ID NO: 318) 0.04

[0338] 6683651 (SEQ ID NO: 319) 0.01

[0339] 6683029 (SEQ ID NO: 320) 0.01

[0340]

[0341] Example 6: Generation and Screening of C23 oxidase Library

[0342] This example describes the identification of C23 oxidase variants that were functional in the described QA biosynthesis pathway in Example 1 for quillaic acid (QA) production.

[0343] To identify C23 oxidase variants that were functional in the described QA biosynthesis pathway for quillaic acid production, a metagenomic library of approximately 800 variants was generated based on the Psammosilene tunicoides C28 oxidase (CYP72A567; SEQ ID NO: 12), the Hibiscus syriacus C28 oxidase (CYP714; SEQ ID NO: 13), the Medicago truncatula C28 oxidase (CYP72A68; SEQ ID NO: 14), and the Ilex asprella C28 oxidase (CYP714E19; SEQ ID NO: 15) sequences.

[0344] The library was transformed, along with the remaining QA biosynthetic genes listed in Example 1, into mevalonate production strain as described in Example 1. The library transformants were cultivated and screened using the procedure as described in Example 5. Table 6. Relative QA titers from C23 Oxidase Library Screen

[0345] C23 Oxidase Candidates QA Production Relative to Positive Control (%)

[0346] CYP72A567 (SEQ ID NO: 12) 0.99

[0347] 6663110 (SEQ ID NO: 321) 0.67

[0348] 6662992 (SEQ ID NO: 322) 0.60

[0349] 6663040 (SEQ ID NO: 323) 0.53

[0350] 6663019 (SEQ ID NO: 324) 0.49

[0351] 6663035 (SEQ ID NO: 325) 0.44

[0352] 6663112 (SEQ ID NO: 326) 0.41

[0353] 6662638 (SEQ ID NO: 327) 0.40

[0354] 6662582 (SEQ ID NO: 328) 0.29

[0355] 6662946 (SEQ ID NO: 329) 0.26

[0356] 6663113 (SEQ ID NO: 330) 0.26

[0357] 6662793 (SEQ ID NO: 331) 0.26

[0358] 6662909 (SEQ ID NO: 332) 0.23

[0359] 6662709 (SEQ ID NO: 333) 0.22

[0360] 6663281 (SEQ ID NO: 334) 0.20

[0361] 6663114 (SEQ ID NO: 335) 0.17

[0362]

[0363] Example 7: Generation and Screening of CPR Library

[0364] This example describes the identification of CPR variants that are functional in the described QA biosynthesis pathway in Example 1 for quillaic acid production.

[0365] To identify CPR variants that are functional in the described QA biosynthesis pathway for quillaic acid production, a metagenomic library of approximately 500 variants was generated based on the two Arabidopsis thaliana CPR (AtATRl; SEQ ID NO: 9; and AtATR2; SEQ ID NO: 10), and the Lotus japonicus CPR (LjCPR; SEQ ID NO: 11) sequences. The library was transformed, along with the remaining QA biosynthetic genes listed in Example 1, into mevalonate production strain as described in Example 1. The library transformants were cultivated and screened using procedure as described in Example 5.

[0366] Table 7. Relative QA titers from CPR Library Screen

[0367] CPR Candidates QA Production Relative to Positive Control (%)

[0368] AtATRl (SEQ ID NO: 9) 1.00

[0369] 8230548 (SEQ ID NO: 336) 1.49

[0370] 8230634 (SEQ ID NO: 337) 1.38

[0371] 8230529 (SEQ ID NO: 338) 1.28

[0372] 8230604 (SEQ ID NO: 339) 1.25

[0373] 8230601 (SEQ ID NO: 340) 1.20

[0374] 8230617 (SEQ ID NO: 341) 1.22

[0375] 8230673 (SEQ ID NO: 342) 1.20

[0376] 8230565 (SEQ ID NO: 343) 1.23

[0377] 8230450 (SEQ ID NO: 344) 1.19

[0378] 8230642 (SEQ ID NO: 345) 1.13

[0379] 8230567 (SEQ ID NO: 346) 1.24

[0380] 8230540 (SEQ ID NO: 347) 1.20

[0381] 8230680 (SEQ ID NO: 348) 1.21

[0382] 8230542 (SEQ ID NO: 349) 1.13

[0383] 8230675 (SEQ ID NO: 350) 1.16

[0384] 8230651 (SEQ ID NO: 351) 1.12

[0385] 8230568 (SEQ ID NO: 352) 1.15

[0386] 8230550 (SEQ ID NO: 353) 1.14

[0387]

[0388] 8230685 (SEQ ID NO: 354) 1.13 8230488 (SEQ ID NO: 355) 1.09 8230649 (SEQ ID NO: 356) 1.13 8230638 (SEQ ID NO: 357) 1.08 8230677 (SEQ ID NO: 358) 1.14 8230598 (SEQ ID NO: 359) 1.06 8230549 (SEQ ID NO: 360) 1.08 8230457 (SEQ ID NO: 361) 1.06 8230575 (SEQ ID NO: 362) 1.04 8230461 (SEQ ID NO: 363) 1.08 8230676 (SEQ ID NO: 364) 1.06 8230582 (SEQ ID NO: 365) 1.09 8230629 (SEQ ID NO: 366) 1.10 8230472 (SEQ ID NO: 367) 1.01 8230538 (SEQ ID NO: 368) 1.01 8230449 (SEQ ID NO: 369) 1.01 8230454 (SEQ ID NO: 370) 1.05 8230645 (SEQ ID NO: 371) 0.92 8230607 (SEQ ID NO: 372) 0.96 8230606 (SEQ ID NO: 373) 0.78 8230586 (SEQ ID NO: 374) 0.53 8230455 (SEQ ID NO: 375) 0.48 8230912 (SEQ ID NO: 376) 0.44 8230469 (SEQ ID NO: 377) 0.02

[0389]

[0390] Example 8: Generation and Screening of MSBP Library

[0391] This example describes the identification of MSBP variants that are functional in the described QA biosynthesis pathway in Example 1 for quillaic acid production.

[0392] To identify MSBP variants that are functional in the described QA biosynthesis pathway for quillaic acid production, a metagenomic library of approximately 500 variants was generated based on the two Arabidopsis thaliana MSBP (AtMSBPl; SEQ ID NO: 4; and AtMSBP2; SEQ ID NO: 5), the Quillaja saponaria MSBP (QsMSBPl; SEQ ID NO: 6), and the two Saponaria vaccaria MSBP (SvMSBPl; SEQ ID NO: 7; an SvMSBP2; SEQ ID NO: 8) sequences.

[0393] The library was transformed, along with the remaining QA biosynthetic genes listed in Example 1, into mevalonate production strain as described in Example 1. The library transformants were cultivated and screened using procedure as described in Example 5.

[0394] Table 8. Relative QA titers from MSBP Library Screen

[0395] MSBP Candidates QA Production Relative to Positive Control (%)

[0396] CB5 (SEQ ID NO: 3) 1.00

[0397] 8232784 (SEQ ID NO: 378) 1.50

[0398] 8232692 (SEQ ID NO: 379) 1.13

[0399] 8232788 (SEQ ID NO: 380) 1.17

[0400] 8232724 (SEQ ID NO: 381) 1.14

[0401] 8232901 (SEQ ID NO: 382) 1.06

[0402] 8232826 (SEQ ID NO: 383) 1.07

[0403] 8232695 (SEQ ID NO: 384) 1.14

[0404] 8232748 (SEQ ID NO: 385) 1.02

[0405] 8232736 (SEQ ID NO: 386) 1.02

[0406] 8232846 (SEQ ID NO: 387) 1.01

[0407] 8232728 (SEQ ID NO: 388) 1.01

[0408] 8232686 (SEQ ID NO: 389) 1.06

[0409]

[0410] 8232738 (SEQ ID NO: 390) 1.00 8232701 (SEQ ID NO: 391) 0.97 8232842 (SEQ ID NO: 392) 0.96 8232751 (SEQ ID NO: 393) 1.01 8232776 (SEQ ID NO: 394) 0.95 8232848 (SEQ ID NO: 395) 0.92 8232763 (SEQ ID NO: 396) 0.95 8232713 (SEQ ID NO: 397) 0.95 8232704 (SEQ ID NO: 398) 0.97 8232712 (SEQ ID NO: 399) 0.97 8232654 (SEQ ID NO: 400) 0.92 8232765 (SEQ ID NO: 401) 0.97 8232829 (SEQ ID NO: 402) 0.92 8232660 (SEQ ID NO: 403) 0.88 8232815 (SEQ ID NO: 404) 0.90 8232657 (SEQ ID NO: 405) 0.88 8232786 (SEQ ID NO: 406) 0.90 8232768 (SEQ ID NO: 407) 0.85 8232730 (SEQ ID NO: 408) 0.89 8233072 (SEQ ID NO: 409) 0.82 8232650 (SEQ ID NO: 410) 0.81 8232711 (SEQ ID NO: 411) 0.85 8232716 (SEQ ID NO: 412) 0.73 8232750 (SEQ ID NO: 413) 0.66 8233078 (SEQ ID NO: 414) 0.65

[0411]

[0412] 8233055 (SEQ ID NO: 415) 0.26

[0413] 8233003 (SEQ ID NO: 416) 0.23

[0414]

[0415] Example 9: Functional expression of prosapogenin biosynthesis in engineered Saccharomyces cerevisiae host cells

[0416] This example describes an engineered S. cerevisiae strain capable of producing prosapogenin based on the quillaic acid and UDP-sugar production strains as described in Example 1 and 2, respectively.

[0417] The prosapogenin biosynthetic pathway was composed of, in addition to QA and UDP-sugar biosynthetic genes, a glucuronosyltransferase (an enzyme which is capable of mediating an enzymatic step of converting QA to QA with a GlcA residue attached at the C3 position (QA-C3-GlcA), designated “GlcAT”), a galactosyltransferase (an enzyme which is capable of mediating an enzymatic step of transferring UDP-Gal and attaching a galactose residue to QA-C3-GlcA to form QA-C3-GlcA-Gal, designated “GalT”), and a xylosyltransf erase (an enzyme which is capable of mediating an enzymatic step of transferring UDP-Xyl and attaching a xylose residue to QA-C3-GlcA-Gal to form QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin), designated “XylT”) (see FIG. 1). It should be noted that any enzyme capable of mediating a GlcAT, GalT or XylT reaction might be able to perform additional or alternative enzymatic reactions. Specifically, a GlcAT from Quillaja saponaria (QsCslG2; SEQ ID NO: 43), a GalT from Quillaja saponaria (QsGalT; SEQ ID NO: 44), and a XylT from Quillaja saponaria (QsC3XylT; SEQ ID NO: 45) were used for the reactions.

[0418] All heterologous genes were codon-optimized for expression in S. cerevisiae and expressed under a galactose-inducible promoter. The expression constructs were transformed in the UDP-sugar producing S. cerevisiae strain (Example 2) and was cultivated using procedure as described in Example 1. The production culture was assayed for prosapogenin production using LC-MS.

[0419] Table 9. Prosapogenin titers from expressing GlcAT, GalT, and XylT in S. cerevisiae Expressed Genes Prosapogenin [mg / L]

[0420] Parent Chassis 0

[0421] GlcAT, GalT, XylT 42.19

[0422]

[0423] Examples 10-12 describe screening libraries for different enzymes that have the same (or similar) activities.

[0424] Example 10: Generation and Screening of GlcAT Library

[0425] This example describes the identification of GlcAT variants that are functional in the described prosapogenin biosynthesis pathway in Example 9 for prosapogenin production.

[0426] To identify GlcAT variants that are functional in the described prosapogenin biosynthesis pathway for quillaic acid production, a metagenomic library of approximately 1000 variants was generated based on the two Quillaja saponaria GlcATs (QsCslGl; SEQ ID NO: 42; and QsCslG2; SEQ ID NO: 43) sequences.

[0427] The metagenomic library is screened for activity in the prosapogenin biosynthesis pathway. The library was transformed, along with the remaining prosapogenin biosynthetic genes listed in Example 9, into QA production strain as described in Example 1. The library transformants were cultivated and screened using procedure as described in Example 5.

[0428] Table 10. Relative Prosapogenin titers from GlcAT Library Screen

[0429] GlcAT Candidates Prosapogenin Production Relative to Positive Control (%)

[0430] QsCslG2 (SEQ ID NO: 43) 0.98

[0431] 6799784 (SEQ ID NO: 417) 1.23

[0432] 6800772 (SEQ ID NO: 418) 1.20

[0433] 6799916 (SEQ ID NO: 419) 1.19

[0434] 6799769 (SEQ ID NO: 420) 1.15

[0435] 6799913 (SEQ ID NO: 421) 1.14

[0436] 6799803 (SEQ ID NO: 422) 1.14

[0437] 6799899 (SEQ ID NO: 423) 0.89

[0438] 6799933 (SEQ ID NO: 424) 0.89

[0439] 6799859 (SEQ ID NO: 425) 0.87

[0440] 6799892 (SEQ ID NO: 426) 0.78

[0441]

[0442] 6800774 (SEQ ID NO: 427) 0.77

[0443] 6800777 (SEQ ID NO: 428) 0.71

[0444] 6799823 (SEQ ID NO: 429) 0.64

[0445] 6799879 (SEQ ID NO: 430) 0.63

[0446] 6799964 (SEQ ID NO: 431) 0.62

[0447] 6800508 (SEQ ID NO: 432) 0.59

[0448] 6799904 (SEQ ID NO: 433) 0.57

[0449] 6799972 (SEQ ID NO: 434) 0.56

[0450] 6799946 (SEQ ID NO: 435) 0.54

[0451] 6799815 (SEQ ID NO: 436) 0.54

[0452] 6799824 (SEQ ID NO: 437) 0.53

[0453] 6799841 (SEQ ID NO: 438) 0.53

[0454] 6800551 (SEQ ID NO: 439) 0.51

[0455] 6800262 (SEQ ID NO: 440) 0.36

[0456] 6800310 (SEQ ID NO: 441) 0.20

[0457]

[0458] Example 11: Generation and Screening of GalT Library

[0459] This example describes the identification of GalT variants that are functional in the described prosapogenin biosynthesis pathway in Example 9 for prosapogenin production.

[0460] To identify GalT variants that are functional in the described prosapogenin biosynthesis pathway for quillaic acid production, a metagenomic library of approximately 2000 variants was generated based on the Quillaja saponaria GalT (QsGalT; SEQ ID NO: 44) and the Quillaja saponaria XylT (QsC3XylT; SEQ ID NO: 45) sequences.

[0461] The metagenomic library is screened for activity in the prosapogenin biosynthesis pathway. The library was transformed, along with the remaining prosapogenin biosynthetic genes listed in Example 9, into QA production strain as described in Example 1. The library transformants were cultivated and screened using procedure as described in Example 5.

[0462] Table 11. Relative Prosapogenin titers from GalT Library Screen GalT Candidates Prosapogenin Production Relative to Positive Control (%)

[0463] QsGalT (SEQ IDNO: 44) 1.02

[0464] 6797741 (SEQ ID NO: 442) 2.32

[0465] 6798826 (SEQ ID NO: 443) 1.41

[0466] 6797736 (SEQ ID NO: 444) 1.27

[0467] 6798307 (SEQ ID NO: 445) 1.17

[0468] 6797867 (SEQ ID NO: 446) 1.16

[0469] 6797724 (SEQ ID NO: 447) 1.16

[0470] 6797538 (SEQ ID NO: 448) 1.14

[0471] 6798048 (SEQ ID NO: 449) 1.10

[0472] 6797863 (SEQ ID NO: 450) 1.10

[0473] 6797902 (SEQ ID NO: 451) 1.09

[0474] 6797826 (SEQ ID NO: 452) 1.09

[0475] 6797833 (SEQ ID NO: 453) 1.07

[0476] 6797537 (SEQ ID NO: 454) 1.04

[0477] 6797747 (SEQ ID NO: 455) 1.03

[0478] 6797958 (SEQ ID NO: 456) 1.02

[0479] 6798853 (SEQ ID NO: 457) 0.99

[0480] 6797518 (SEQ ID NO: 458) 0.94

[0481] 6797733 (SEQ ID NO: 459) 0.89

[0482] 6797584 (SEQ ID NO: 460) 0.89

[0483] 6798043 (SEQ ID NO: 461) 0.83

[0484] 6798350 (SEQ ID NO: 462) 0.80

[0485] 6797859 (SEQ ID NO: 463) 0.80

[0486]

[0487] 6797585 (SEQ ID NO: 464) 0.77 6797579 (SEQ ID NO: 465) 0.77 6798298 (SEQ ID NO: 466) 0.77 6797515 (SEQ ID NO: 467) 0.77 6798291 (SEQ ID NO: 468) 0.75 6798313 (SEQ ID NO: 469) 0.75 6797982 (SEQ ID NO: 470) 0.74 6798912 (SEQ ID NO: 471) 0.72 6798296 (SEQ ID NO: 472) 0.71 6798239 (SEQ ID NO: 473) 0.67 6797564 (SEQ ID NO: 474) 0.65 6798414 (SEQ ID NO: 475) 0.64 6798317 (SEQ ID NO: 476) 0.62 6798244 (SEQ ID NO: 477) 0.58 6799141 (SEQ ID NO: 478) 0.55 6797897 (SEQ ID NO: 479) 0.54 6798092 (SEQ ID NO: 480) 0.52 6798109 (SEQ ID NO: 481) 0.52 6797524 (SEQ ID NO: 482) 0.51 6797729 (SEQ ID NO: 483) 0.50 6797725 (SEQ ID NO: 484) 0.50 6797810 (SEQ ID NO: 485) 0.49 6797462 (SEQ ID NO: 486) 0.48 6798320 (SEQ ID NO: 487) 0.47 6797638 (SEQ ID NO: 488) 0.43

[0488]

[0489] 6798477 (SEQ ID NO: 489) 0.41 6797753 (SEQ ID NO: 490) 0.41 6798172 (SEQ ID NO: 491) 0.40 6797453 (SEQ ID NO: 492) 0.40 6797578 (SEQ ID NO: 493) 0.40 6798661 (SEQ ID NO: 494) 0.39 6798405 (SEQ ID NO: 495) 0.39 6799406 (SEQ ID NO: 496) 0.38 6799387 (SEQ ID NO: 497) 0.36 6798910 (SEQ ID NO: 498) 0.36 6798822 (SEQ ID NO: 499) 0.33 6799044 (SEQ ID NO: 500) 0.33 6798341 (SEQ ID NO: 501) 0.32 6798125 (SEQ ID NO: 502) 0.32 6798866 (SEQ ID NO: 503) 0.30 6798532 (SEQ ID NO: 504) 0.30 6797534 (SEQ ID NO: 505) 0.29 6797868 (SEQ ID NO: 506) 0.29 6799072 (SEQ ID NO: 507) 0.29 6798008 (SEQ ID NO: 508) 0.28 6797821 (SEQ ID NO: 509) 0.28 6798525 (SEQ ID NO: 510) 0.27 6797744 (SEQ ID NO: 511) 0.27 6799392 (SEQ ID NO: 512) 0.27 6798218 (SEQ ID NO: 513) 0.26

[0490]

[0491] 6797679 (SEQ ID NO: 514) 0.25 6797448 (SEQ ID NO: 515) 0.24 6799359 (SEQ ID NO: 516) 0.24 6799003 (SEQ ID NO: 517) 0.24 6798976 (SEQ ID NO: 518) 0.23 6799033 (SEQ ID NO: 519) 0.23 6798628 (SEQ ID NO: 520) 0.23 6797467 (SEQ ID NO: 521) 0.23 6797641 (SEQ ID NO: 522) 0.22 6798312 (SEQ ID NO: 523) 0.22 6798705 (SEQ ID NO: 524) 0.22 6798528 (SEQ ID NO: 525) 0.21 6798383 (SEQ ID NO: 526) 0.21 6797642 (SEQ ID NO: 527) 0.20 6798958 (SEQ ID NO: 528) 0.20 6799202 (SEQ ID NO: 529) 0.19 6798733 (SEQ ID NO: 530) 0.19 6797553 (SEQ ID NO: 531) 0.19 6797917 (SEQ ID NO: 532) 0.19 6798995 (SEQ ID NO: 533) 0.18 6798983 (SEQ ID NO: 534) 0.18 6797899 (SEQ ID NO: 535) 0.15 6798401 (SEQ ID NO: 536) 0.15 6799402 (SEQ ID NO: 537) 0.10

[0492]

[0493] Example 12: Generation and Screening ofXylT Library

[0494] This example describes the identification ofXylT variants that are functional in the described prosapogenin biosynthesis pathway in Example 9 for prosapogenin production.

[0495] To identify XylT variants that are functional in the described prosapogenin biosynthesis pathway for quillaic acid production, a metagenomic library of approximately 2000 variants as described in Example 11.

[0496] The metagenomic library is screened for activity in the prosapogenin biosynthesis pathway. The library was transformed, along with the remaining prosapogenin biosynthetic genes listed in Example 9, into QA production strain as described in Example 1. The library transformants were cultivated and screened using procedure as described in Example 5.

[0497] Table 12. Relative Prosapogenin titers from XylT Library Screen

[0498] XylT Candidates Prosapogenin Production Relative to Positive Control (%)

[0499] QsC3XylT (SEQ ID NO: 45) 1.03

[0500] 6797740 (SEQ ID NO: 538) 1.20

[0501] 6797690 (SEQ ID NO: 539) 1.17

[0502] 6797723 (SEQ ID NO: 540) 1.17

[0503] 6797885 (SEQ ID NO: 541) 1.03

[0504] 6798560 (SEQ ID NO: 542) 1.02

[0505] 6797688 (SEQ ID NO: 543) 1.01

[0506] 6798292 (SEQ ID NO: 544) 1.00

[0507] 6798353 (SEQ ID NO: 545) 0.95

[0508] 6797554 (SEQ ID NO: 546) 0.94

[0509] 6798391 (SEQ ID NO: 547) 0.91

[0510] 6799186 (SEQ ID NO: 548) 0.85

[0511] 6799241 (SEQ ID NO: 549) 0.83

[0512] 6799098 (SEQ ID NO: 550) 0.80

[0513]

[0514] 6797810 (SEQ ID NO: 485) 0.78 6798588 (SEQ ID NO: 551) 0.70 6798823 (SEQ ID NO: 552) 0.57 6798661 (SEQ ID NO: 494) 0.57 6799346 (SEQ ID NO: 553) 0.55 6798608 (SEQ ID NO: 554) 0.49 6798416 (SEQ ID NO: 555) 0.48 6798172 (SEQ ID NO: 491) 0.43 6799132 (SEQ ID NO: 556) 0.43 6799004 (SEQ ID NO: 557) 0.39 6798785 (SEQ ID NO: 558) 0.38 6799173 (SEQ ID NO: 559) 0.38 6798687 (SEQ ID NO: 560) 0.35 6799248 (SEQ ID NO: 561) 0.34 6799003 (SEQ ID NO: 517) 0.34 6798715 (SEQ ID NO: 562) 0.33 6798125 (SEQ ID NO: 502) 0.33 6797913 (SEQ ID NO: 563) 0.30 6797589 (SEQ ID NO: 564) 0.29 6798352 (SEQ ID NO: 565) 0.27 6799290 (SEQ ID NO: 566) 0.26 6799313 (SEQ ID NO: 567) 0.24 6798534 (SEQ ID NO: 568) 0.23 6797834 (SEQ ID NO: 569) 0.20 6797725 (SEQ ID NO: 484) 0.19

[0515]

[0516] 6797664 (SEQ ID NO: 570) 0.18

[0517] 6798699 (SEQ ID NO: 571) 0.16

[0518] 6799177 (SEQ ID NO: 572) 0.15

[0519] 6798532 (SEQ ID NO: 504) 0.15

[0520]

[0521] Example 13: Generation and Screening ofUGD Library

[0522] This example describes the identification ofUGD variants that were functional in the described UDP-GlcA biosynthesis pathway in Example 2.

[0523] To identify UGD variants that were functional in the described UGD-GlcA biosynthesis pathway for prosapogenin production, a metagenomic library of approximately 500 variants was generated based on the Bacteroidales bacterium UGD (BbacUGD; SEQ ID NO: 46), the Blautia luti UGD (BlutUGD; SEQ ID NO: 47), two Escherichia coli UGD (EcolUGD; SEQ ID NO: 48; and EcolUGD2; SEQ ID NO: 49), the Firmicutes bacterium UGD (FbacUGD; SEQ ID NO: 50), the Holdemania filiformis UGD (HfilUGD; SEQ ID NO: 51), the Lachnospiraceae bacterium UGD (LbacUGD; SEQ ID NO: 52), the Lactococcus garvieae UGD (LgarUGD; SEQ ID NO: 53), the Streptococcus iniae UGD (SiniUGD; SEQ ID NO: 54), the Streptococcus equi UGD (SequUGD; SEQ ID NO: 55), and the Streptococcus pyogenes UGD (SpyoUGD; SEQ ID NO: 56) sequences.

[0524] The library was transformed, along with the remaining prosapogenin biosynthetic genes listed in Example 9, into QA production strain as described in Example 1. The library transformants were cultivated and screened using procedure as described in Example 5.

[0525] Table 13. Relative Prosapogenin titers from UGD Library Screen

[0526] UGD Candidates Prosapogenin Production Relative to Positive Control (%)

[0527] LbacUGD (SEQ ID NO: 52) 1.00

[0528] 7750758 (SEQ ID NO: 712) 1.55

[0529] 7750967 (SEQ ID NO: 713) 1.50

[0530] 7750846 (SEQ ID NO: 714) 1.49

[0531] 7750778 (SEQ ID NO: 715) 1.49

[0532]

[0533] 7750864 (SEQ ID NO: 716) 1.49 7750992 (SEQ ID NO: 717) 1.43 7750448 (SEQ ID NO: 718) 1.43 7750982 (SEQ ID NO: 719) 1.43 7750339 (SEQ ID NO: 720) 1.41 7750863 (SEQ ID NO: 721) 1.38 7751028 (SEQ ID NO: 722) 1.37 7751063 (SEQ ID NO: 723) 1.35 7750874 (SEQ ID NO: 724) 1.33 7751101 (SEQ ID NO: 725) 1.32 7751021 (SEQ ID NO: 726) 1.31 7750984 (SEQ ID NO: 727) 1.31 7750925 (SEQ ID NO: 728) 1.31 7750908 (SEQ ID NO: 729) 1.30 7750440 (SEQ ID NO: 730) 1.28 7750873 (SEQ ID NO: 731) 1.28 7750402 (SEQ ID NO: 732) 1.27 7750425 (SEQ ID NO: 733) 1.27 7750334 (SEQ ID NO: 734) 1.26 7750418 (SEQ ID NO: 735) 1.25 7750851 (SEQ ID NO: 736) 1.24 7750686 (SEQ ID NO: 737) 1.23 7750831 (SEQ ID NO: 738) 1.23 7750986 (SEQ ID NO: 739) 1.23 7750579 (SEQ ID NO: 740) 1.22 7750600 (SEQ ID NO: 741) 1.22 7750829 (SEQ ID NO: 742) 1.22

[0534]

[0535] 7750483 (SEQ ID NO: 743) 1.22 7750941 (SEQ ID NO: 744) 1.21 7750310 (SEQ ID NO: 745) 1.20 7750647 (SEQ ID NO: 746) 1.20 7750856 (SEQ ID NO: 747) 1.19 7751067 (SEQ ID NO: 748) 1.19 7750364 (SEQ ID NO: 749) 1.18 7750481 (SEQ ID NO: 750) 1.18 7750821 (SEQ ID NO: 751) 1.17 7750532 (SEQ ID NO: 752) 1.17 7750939 (SEQ ID NO: 753) 1.17 7750861 (SEQ ID NO: 754) 1.17 7750953 (SEQ ID NO: 755) 1.17 7750895 (SEQ ID NO: 756) 1.16 7750929 (SEQ ID NO: 757) 1.16 7750776 (SEQ ID NO: 758) 1.16 7750502 (SEQ ID NO: 759) 1.15 7750911 (SEQ ID NO: 760) 1.15 7750880 (SEQ ID NO: 761) 1.15 7750470 (SEQ ID NO: 762) 1.15 7750345 (SEQ ID NO: 763) 1.15 7750823 (SEQ ID NO: 764) 1.15 7750476 (SEQ ID NO: 765) 1.14 7750342 (SEQ ID NO: 766) 1.14 7750438 (SEQ ID NO: 767) 1.14 7750572 (SEQ ID NO: 768) 1.14 7750330 (SEQ ID NO: 769) 1.14

[0536]

[0537] 7751034 (SEQ ID NO: 770) 1.14 7750979 (SEQ ID NO: 771) 1.13 7750868 (SEQ ID NO: 772) 1.13 7750858 (SEQ ID NO: 773) 1.12 7750315 (SEQ ID NO: 774) 1.12 7750844 (SEQ ID NO: 775) 1.12 7750377 (SEQ ID NO: 776) 1.11 7750486 (SEQ ID NO: 777) 1.11 7750878 (SEQ ID NO: 778) 1.11 7750389 (SEQ ID NO: 779) 1.11 7751041 (SEQ ID NO: 780) 1.10 7750422 (SEQ ID NO: 781) 1.10 7750824 (SEQ ID NO: 782) 1.09 7750304 (SEQ ID NO: 783) 1.09 7750574 (SEQ ID NO: 784) 1.09 7750741 (SEQ ID NO: 785) 1.09 7750749 (SEQ ID NO: 786) 1.09 7750305 (SEQ ID NO: 787) 1.09 7751098 (SEQ ID NO: 788) 1.09 7750367 (SEQ ID NO: 789) 1.09 7750759 (SEQ ID NO: 790) 1.08 7750784 (SEQ ID NO: 791) 1.08 7750473 (SEQ ID NO: 792) 1.08 7750399 (SEQ ID NO: 793) 1.08 7750404 (SEQ ID NO: 794) 1.08 7750347 (SEQ ID NO: 795) 1.08 7750584 (SEQ ID NO: 796) 1.08

[0538]

[0539] 7750332 (SEQ ID NO: 797) 1.07 7750407 (SEQ ID NO: 798) 1.07 7750927 (SEQ ID NO: 799) 1.07 7750374 (SEQ ID NO: 800) 1.07 7750324 (SEQ ID NO: 801) 1.07 7750786 (SEQ ID NO: 802) 1.07 7750871 (SEQ ID NO: 803) 1.07 7750327 (SEQ ID NO: 804) 1.06 7750545 (SEQ ID NO: 805) 1.06 7750370 (SEQ ID NO: 806) 1.06 7750609 (SEQ ID NO: 807) 1.05 7750488 (SEQ ID NO: 808) 1.05 7750552 (SEQ ID NO: 809) 1.05 7750592 (SEQ ID NO: 810) 1.05 7751094 (SEQ ID NO: 811) 1.05 7750352 (SEQ ID NO: 812) 1.05 7750916 (SEQ ID NO: 813) 1.04 7750325 (SEQ ID NO: 814) 1.04 7751026 (SEQ ID NO: 815) 1.04 7750335 (SEQ ID NO: 816) 1.04 7750337 (SEQ ID NO: 817) 1.04 7750555 (SEQ ID NO: 818) 1.04 7750322 (SEQ ID NO: 819) 1.03 7750713 (SEQ ID NO: 820) 1.03 7750951 (SEQ ID NO: 821) 1.03 7750692 (SEQ ID NO: 822) 1.03 7750617 (SEQ ID NO: 823) 1.03

[0540]

[0541] 7750523 (SEQ ID NO: 824) 1.03 7750839 (SEQ ID NO: 825) 1.03 7750743 (SEQ ID NO: 826) 1.03 7750405 (SEQ ID NO: 827) 1.03 7750543 (SEQ ID NO: 828) 1.02 7750360 (SEQ ID NO: 829) 1.02 7750733 (SEQ ID NO: 830) 1.02 7750734 (SEQ ID NO: 831) 1.02 7750350 (SEQ ID NO: 832) 1.02 7750427 (SEQ ID NO: 833) 1.01 7750349 (SEQ ID NO: 834) 1.01 7750672 (SEQ ID NO: 835) 1.01 7750409 (SEQ ID NO: 836) 1.01 7750736 (SEQ ID NO: 837) 1.01 7750753 (SEQ ID NO: 838) 1.01 7751019 (SEQ ID NO: 839) 1.00 7751065 (SEQ ID NO: 840) 1.00 7750684 (SEQ ID NO: 841) 1.00 7750669 (SEQ ID NO: 842) 1.00 7750954 (SEQ ID NO: 843) 1.00 7751003 (SEQ ID NO: 844) 1.00 7751084 (SEQ ID NO: 845) 1.00 7750575 (SEQ ID NO: 846) 1.00 7750934 (SEQ ID NO: 847) 0.99 7750644 (SEQ ID NO: 848) 0.99 7750989 (SEQ ID NO: 849) 0.99 7750975 (SEQ ID NO: 850) 0.99

[0542]

[0543] 7751053 (SEQ IDNO: 851) 0.99 7750529 (SEQ ID NO: 852) 0.99 7750961 (SEQ IDNO: 853) 0.98 7750897 (SEQ ID NO: 854) 0.98 7750657 (SEQ ID NO: 855) 0.98 7751024 (SEQ ID NO: 856) 0.98 7750451 (SEQ IDNO: 857) 0.98 7750930 (SEQ ID NO: 858) 0.98 7750314 (SEQ ID NO: 859) 0.98 7750357 (SEQ ID NO: 860) 0.98 7750738 (SEQ ID NO: 861) 0.98 7750475 (SEQ ID NO: 862) 0.98 7750340 (SEQ ID NO: 863) 0.97 7750355 (SEQ ID NO: 864) 0.97 7750968 (SEQ ID NO: 865) 0.96 7750666 (SEQ ID NO: 866) 0.96 7750682 (SEQ ID NO: 867) 0.96 7750420 (SEQ ID NO: 868) 0.96 7750577 (SEQ ID NO: 869) 0.96 7750530 (SEQ ID NO: 870) 0.96 7750390 (SEQ IDNO: 871) 0.96 7750811 (SEQ IDNO: 872) 0.96 7750909 (SEQ ID NO: 873) 0.96 7750788 (SEQ ID NO: 874) 0.95 7750614 (SEQ ID NO: 875) 0.95 7750460 (SEQ ID NO: 876) 0.95 7750649 (SEQ ID NO: 877) 0.95

[0544]

[0545] 7750921 (SEQ ID NO: 878) 0.95 7750836 (SEQ ID NO: 879) 0.95 7750913 (SEQ ID NO: 880) 0.95 7750480 (SEQ ID NO: 881) 0.94 7750801 (SEQ ID NO: 882) 0.94 7750505 (SEQ ID NO: 883) 0.94 7750468 (SEQ ID NO: 884) 0.94 7750694 (SEQ ID NO: 885) 0.94 7750918 (SEQ ID NO: 886) 0.94 7750510 (SEQ ID NO: 887) 0.94 7750803 (SEQ ID NO: 888) 0.94 7751070 (SEQ ID NO: 889) 0.94 7750691 (SEQ ID NO: 890) 0.93 7750319 (SEQ ID NO: 891) 0.93 7750703 (SEQ ID NO: 892) 0.93 7750503 (SEQ ID NO: 893) 0.93 7750659 (SEQ ID NO: 894) 0.93 7750848 (SEQ ID NO: 895) 0.93 7750620 (SEQ ID NO: 896) 0.93 7750562 (SEQ ID NO: 897) 0.93 7750466 (SEQ ID NO: 898) 0.93 7750520 (SEQ ID NO: 899) 0.93 7750949 (SEQ ID NO: 900) 0.93 7750796 (SEQ ID NO: 901) 0.93 7750965 (SEQ ID NO: 902) 0.93 7751087 (SEQ ID NO: 903) 0.93 7750443 (SEQ ID NO: 904) 0.92

[0546]

[0547] 7750731 (SEQ IDNO: 905) 0.92 7751029 (SEQ ID NO: 906) 0.92 7750605 (SEQ ID NO: 907) 0.92 7750582 (SEQ ID NO: 908) 0.92 7750569 (SEQ ID NO: 909) 0.92 7750369 (SEQ ID NO: 910) 0.92 7750518 (SEQ ID NO: 911) 0.92 7750450 (SEQ ID NO: 912) 0.92 7751125 (SEQ IDNO: 913) 0.92 7750804 (SEQ ID NO: 914) 0.92 7750714 (SEQ ID NO: 915) 0.91 7750744 (SEQ ID NO: 916) 0.91 7750765 (SEQ ID NO: 917) 0.91 7750392 (SEQ IDNO: 918) 0.91 7750886 (SEQ IDNO: 919) 0.91 7750727 (SEQ ID NO: 920) 0.91 7750522 (SEQ ID NO: 921) 0.91 7750789 (SEQ ID NO: 922) 0.91 7750365 (SEQ ID NO: 923) 0.91 7750838 (SEQ ID NO: 924) 0.90 7750372 (SEQ ID NO: 925) 0.90 7751115 (SEQ IDNO: 926) 0.90 7750309 (SEQ ID NO: 927) 0.90 7750998 (SEQ ID NO: 928) 0.90 7750843 (SEQ ID NO: 929) 0.90 7750890 (SEQ ID NO: 930) 0.90 7750567 (SEQ IDNO: 931) 0.90

[0548]

[0549] 7750932 (SEQ ID NO: 932) 0.90 7750974 (SEQ ID NO: 933) 0.90 7750894 (SEQ ID NO: 934) 0.90 7750841 (SEQ ID NO: 935) 0.90 7750814 (SEQ ID NO: 936) 0.90 7750379 (SEQ ID NO: 937) 0.89 7750642 (SEQ ID NO: 938) 0.89 7751072 (SEQ ID NO: 939) 0.89 7750550 (SEQ ID NO: 940) 0.89 7750687 (SEQ ID NO: 941) 0.89 7750708 (SEQ ID NO: 942) 0.89 7750394 (SEQ ID NO: 943) 0.89 7750513 (SEQ ID NO: 944) 0.89 7750508 (SEQ ID NO: 945) 0.89 7751023 (SEQ ID NO: 946) 0.89 7751036 (SEQ ID NO: 947) 0.89 7750906 (SEQ ID NO: 948) 0.88 7750354 (SEQ ID NO: 949) 0.88 7750629 (SEQ ID NO: 950) 0.88 7750866 (SEQ ID NO: 951) 0.88 7750553 (SEQ ID NO: 952) 0.88 7750442 (SEQ ID NO: 953) 0.88 7750699 (SEQ ID NO: 954) 0.88 7750761 (SEQ ID NO: 955) 0.88 7750548 (SEQ ID NO: 956) 0.88 7750960 (SEQ ID NO: 957) 0.88 7750580 (SEQ ID NO: 958) 0.88

[0550]

[0551] 7750981 (SEQ IDNO: 959) 0.88 7750570 (SEQ ID NO: 960) 0.88 7751031 (SEQ IDNO: 961) 0.88 7750763 (SEQ ID NO: 962) 0.88 7750816 (SEQ ID NO: 963) 0.88 7750343 (SEQ ID NO: 964) 0.87 7750854 (SEQ ID NO: 965) 0.87 7750458 (SEQ ID NO: 966) 0.87 7750320 (SEQ ID NO: 967) 0.87 7750437 (SEQ ID NO: 968) 0.87 7750674 (SEQ ID NO: 969) 0.87 7751074 (SEQ ID NO: 970) 0.87 7750994 (SEQ ID NO: 971) 0.87 7750794 (SEQ ID NO: 972) 0.86 7750947 (SEQ ID NO: 973) 0.86 7750538 (SEQ ID NO: 974) 0.86 7750779 (SEQ ID NO: 975) 0.86 7750542 (SEQ ID NO: 976) 0.86 7750809 (SEQ ID NO: 977) 0.86 7750359 (SEQ ID NO: 978) 0.86 7750775 (SEQ ID NO: 979) 0.86 7750497 (SEQ ID NO: 980) 0.86 7750689 (SEQ IDNO: 981) 0.86 7750991 (SEQ ID NO: 982) 0.86 7750455 (SEQ ID NO: 983) 0.86 7750718 (SEQ ID NO: 984) 0.85 7750709 (SEQ ID NO: 985) 0.85

[0552]

[0553] 7750640 (SEQ ID NO: 986) 0.85 7750540 (SEQ ID NO: 987) 0.85 7750490 (SEQ ID NO: 988) 0.85 7750317 (SEQ ID NO: 989) 0.85 7751058 (SEQ ID NO: 990) 0.85 7750885 (SEQ ID NO: 991) 0.84 7750533 (SEQ ID NO: 992) 0.84 7750602 (SEQ ID NO: 993) 0.84 7750739 (SEQ ID NO: 994) 0.84 7750701 (SEQ ID NO: 995) 0.83 7750923 (SEQ ID NO: 996) 0.83 7750585 (SEQ ID NO: 997) 0.83 7750876 (SEQ ID NO: 998) 0.83 7750944 (SEQ ID NO: 999) 0.83 7750646 (SEQ ID NO: 1000) 0.83 7750937 (SEQ ID NO: 1001) 0.83 7750946 (SEQ ID NO: 1002) 0.83 7750719 (SEQ ID NO: 1003) 0.82 7750465 (SEQ ID NO: 1004) 0.82 7750972 (SEQ ID NO: 1005) 0.82 7750507 (SEQ ID NO: 1006) 0.82 7750512 (SEQ ID NO: 1007) 0.81 7750485 (SEQ ID NO: 1008) 0.81 7750770 (SEQ ID NO: 1009) 0.81 7750812 (SEQ ID NO: 1010) 0.81 7750869 (SEQ ID NO: 1011) 0.80 7750833 (SEQ ID NO: 1012) 0.80

[0554]

[0555] 7750751 (SEQ IDNO: 1013) 0.79 7751096 (SEQ IDNO: 1014) 0.79 7750892 (SEQ IDNO: 1015) 0.79 7750547 (SEQ ID NO: 1016) 0.79 7750853 (SEQ IDNO: 1017) 0.79 7750942 (SEQ ID NO: 1018) 0.78 7750559 (SEQ ID NO: 1019) 0.78 7750706 (SEQ ID NO: 1020) 0.78 7750627 (SEQ ID NO: 1021) 0.78 7751006 (SEQ IDNO: 1022) 0.78 7750599 (SEQ ID NO: 1023) 0.78 7750651 (SEQ IDNO: 1024) 0.78 7750859 (SEQ ID NO: 1025) 0.78 7751080 (SEQ IDNO: 1026) 0.77 7750711 (SEQ IDNO: 1027) 0.77 7751068 (SEQ ID NO: 1028) 0.77 7750808 (SEQ ID NO: 1029) 0.77 7750624 (SEQ ID NO: 1030) 0.77 7750671 (SEQ IDNO: 1031) 0.77 7750307 (SEQ ID NO: 1032) 0.76 7750901 (SEQ IDNO: 1033) 0.76 7751016 (SEQ ID NO: 1034) 0.76 7750447 (SEQ ID NO: 1035) 0.76 7750915 (SEQ ID NO: 1036) 0.76 7751086 (SEQ IDNO: 1037) 0.76 7751082 (SEQ IDNO: 1038) 0.76 7750883 (SEQ ID NO: 1039) 0.75

[0556]

[0557] 7750590 (SEQ ID NO: 1040) 0.75 7751039 (SEQ ID NO: 1041) 0.75 7750748 (SEQ ID NO: 1042) 0.75 7750724 (SEQ ID NO: 1043) 0.75 7750622 (SEQ ID NO: 1044) 0.75 7750722 (SEQ ID NO: 1045) 0.75 7750977 (SEQ ID NO: 1046) 0.75 7751014 (SEQ ID NO: 1047) 0.75 7750312 (SEQ ID NO: 1048) 0.74 7750302 (SEQ ID NO: 1049) 0.74 7750721 (SEQ ID NO: 1050) 0.74 7750604 (SEQ ID NO: 1051) 0.74 7750799 (SEQ ID NO: 1052) 0.74 7750625 (SEQ ID NO: 1053) 0.74 7750500 (SEQ ID NO: 1054) 0.73 7750834 (SEQ ID NO: 1055) 0.73 7751062 (SEQ ID NO: 1056) 0.73 7750634 (SEQ ID NO: 1057) 0.73 7750329 (SEQ ID NO: 1058) 0.73 7750457 (SEQ ID NO: 1059) 0.73 7750445 (SEQ ID NO: 1060) 0.73 7750956 (SEQ ID NO: 1061) 0.73 7750935 (SEQ ID NO: 1062) 0.72 7750375 (SEQ ID NO: 1063) 0.72 7750849 (SEQ ID NO: 1064) 0.72 7751108 (SEQ ID NO: 1065) 0.72 7750899 (SEQ ID NO: 1066) 0.71

[0558]

[0559] 7750667 (SEQ ID NO: 1067) 0.71 7750493 (SEQ ID NO: 1068) 0.71 7750610 (SEQ ID NO: 1069) 0.71 7750564 (SEQ ID NO: 1070) 0.71 7750987 (SEQ ID NO: 1071) 0.71 7750888 (SEQ ID NO: 1072) 0.71 7750920 (SEQ ID NO: 1073) 0.70 7750677 (SEQ ID NO: 1074) 0.70 7750999 (SEQ ID NO: 1075) 0.70 7750768 (SEQ ID NO: 1076) 0.70 7750565 (SEQ ID NO: 1077) 0.70 7750612 (SEQ ID NO: 1078) 0.69 7750806 (SEQ ID NO: 1079) 0.69 7750902 (SEQ ID NO: 1080) 0.69 7750828 (SEQ ID NO: 1081) 0.69 7750904 (SEQ ID NO: 1082) 0.69 7750589 (SEQ ID NO: 1083) 0.69 7750463 (SEQ ID NO: 1084) 0.68 7750756 (SEQ ID NO: 1085) 0.67 7750716 (SEQ ID NO: 1086) 0.67 7751033 (SEQ ID NO: 1087) 0.65 7751051 (SEQ ID NO: 1088) 0.65 7750679 (SEQ ID NO: 1089) 0.65 7751018 (SEQ ID NO: 1090) 0.65 7750754 (SEQ ID NO: 1091) 0.65 7751060 (SEQ ID NO: 1092) 0.65 7750632 (SEQ ID NO: 1093) 0.65

[0560]

[0561] 7750635 (SEQ ID NO: 1094) 0.65 7750478 (SEQ ID NO: 1095) 0.64 7750791 (SEQ ID NO: 1096) 0.64 7751046 (SEQ ID NO: 1097) 0.64 7750639 (SEQ ID NO: 1098) 0.64 7751100 (SEQ ID NO: 1099) 0.62 7750766 (SEQ ID NO: 1100) 0.62 7750607 (SEQ ID NO: 1101) 0.62 7750453 (SEQ ID NO: 1102) 0.60 7750471 (SEQ ID NO: 1103) 0.60 7751111 (SEQ ID NO: 1104) 0.60 7751119 (SEQ ID NO: 1105) 0.60 7750773 (SEQ ID NO: 1106) 0.59 7750595 (SEQ ID NO: 1107) 0.59 7750525 (SEQ ID NO: 1108) 0.59 7750362 (SEQ ID NO: 1109) 0.58 7750696 (SEQ ID NO: 1110) 0.58 7750826 (SEQ ID NO: 1111) 0.58 7751127 (SEQ ID NO: 1112) 0.58 7750560 (SEQ ID NO: 1113) 0.58 7750970 (SEQ ID NO: 1114) 0.58 7750681 (SEQ ID NO: 1115) 0.58 7750704 (SEQ ID NO: 1116) 0.57 7750495 (SEQ ID NO: 1117) 0.57 7750395 (SEQ ID NO: 1118) 0.56 7750535 (SEQ ID NO: 1119) 0.56 7750515 (SEQ ID NO: 1120) 0.56

[0562]

[0563] 7750637 (SEQ ID NO: 1121) 0.55

[0564] 7750387 (SEQ ID NO: 1122) 0.54

[0565] 7750958 (SEQ ID NO: 1123) 0.54

[0566] 7750587 (SEQ ID NO: 1124) 0.54

[0567] 7750537 (SEQ ID NO: 1125) 0.53

[0568] 7750498 (SEQ ID NO: 1126) 0.53

[0569] 7751055 (SEQ ID NO: 1127) 0.52

[0570] 7750798 (SEQ ID NO: 1128) 0.52

[0571] 7750557 (SEQ ID NO: 1129) 0.52

[0572] 7750594 (SEQ ID NO: 1130) 0.49

[0573] 7750491 (SEQ ID NO: 1131) 0.49

[0574] 7750517 (SEQ ID NO: 1132) 0.48

[0575] 7751113 (SEQ ID NO: 1133) 0.47

[0576] 7751106 (SEQ ID NO: 1134) 0.45

[0577] 7751001 (SEQ ID NO: 1135) 0.44

[0578] 7750963 (SEQ ID NO: 1136) 0.44

[0579] 7750819 (SEQ ID NO: 1137) 0.43

[0580] 7750996 (SEQ ID NO: 1138) 0.42

[0581] 7751117 (SEQ ID NO: 1139) 0.41

[0582] 7750818 (SEQ ID NO: 1140) 0.39

[0583] 7750781 (SEQ ID NO: 1141) 0.39

[0584] 7750423 (SEQ ID NO: 1142) 0.38

[0585]

[0586] Example 14: Generation and Screening of UXS Library

[0587] This example describes the identification of UXS variants that were functional in the described UDP-Xyl biosynthesis pathway in Example 2.

[0588] To identify UXS variants that were functional in the described UGD-Xyl biosynthesis pathway for prosapogenin production, a metagenomic library of approximately 500 variants was generated based on the Arabidopsis thaliana UXS (AtUXS; SEQ ID NO: 57) and the Quillaja saponaria UXS (QtUXS; SEQ ID NO: 58) sequences.

[0589] The library was transformed, along with the remaining prosapogenin biosynthetic genes listed in Example 9, into QA production strain as described in Example 1. The library transformants were cultivated and screened using procedure as described in Example 5.

[0590] Table 14. Relative Prosapogenin titers from UXS Library Screen

[0591] UXS Candidates Prosapogenin Production Relative to Positive Control (%)

[0592] AtUXS (SEQ ID NO: 57) 1.00

[0593] 8123775 (SEQ ID NO: 573) 1.63

[0594] 8123777 (SEQ ID NO: 574) 1.50

[0595] 8123645 (SEQ ID NO: 575) 1.26

[0596] 8123879 (SEQ ID NO: 576) 1.25

[0597] 8123691 (SEQ ID NO: 577) 1.24

[0598] 8123678 (SEQ ID NO: 578) 1.18

[0599] 8123579 (SEQ ID NO: 579) 1.18

[0600] 8123636 (SEQ ID NO: 580) 1.14

[0601] 8123648 (SEQ ID NO: 581) 1.14

[0602] 8123634 (SEQ ID NO: 582) 1.13

[0603] 8123739 (SEQ ID NO: 583) 1.10

[0604] 8123749 (SEQ ID NO: 584) 1.08

[0605] 8123746 (SEQ ID NO: 585) 1.08

[0606] 8123742 (SEQ ID NO: 586) 1.06

[0607] 8123760 (SEQ ID NO: 587) 1.05

[0608] 8123502 (SEQ ID NO: 588) 1.05

[0609] 8123708 (SEQ ID NO: 589) 1.05

[0610]

[0611] 8123644 (SEQ ID NO: 590) 1.04 8123611 (SEQ ID NO: 591) 1.04 8123460 (SEQ ID NO: 592) 1.04 8123667 (SEQ ID NO: 593) 1.02 8123531 (SEQ ID NO: 594) 1.00 8123637 (SEQ ID NO: 595) 1.00 8123664 (SEQ ID NO: 596) 1.00 8123493 (SEQ ID NO: 597) 0.99 8123622 (SEQ ID NO: 598) 0.97 8123674 (SEQ ID NO: 599) 0.97 8123705 (SEQ ID NO: 600) 0.95 8123614 (SEQ ID NO: 601) 0.94 8123554 (SEQ ID NO: 602) 0.94 8123658 (SEQ ID NO: 603) 0.93 8123505 (SEQ ID NO: 604) 0.93 8123736 (SEQ ID NO: 605) 0.92 8123520 (SEQ ID NO: 606) 0.92 8123613 (SEQ ID NO: 607) 0.92 8123563 (SEQ ID NO: 608) 0.91 8123532 (SEQ ID NO: 609) 0.91 8123517 (SEQ ID NO: 610) 0.90 8123572 (SEQ ID NO: 611) 0.90 8123647 (SEQ ID NO: 612) 0.90 8123719 (SEQ ID NO: 613) 0.89 8123712 (SEQ ID NO: 614) 0.88

[0612]

[0613] 8123724 (SEQ ID NO: 615) 0.88 8123482 (SEQ ID NO: 616) 0.87 8123665 (SEQ ID NO: 617) 0.87 8123675 (SEQ ID NO: 618) 0.87 8123509 (SEQ ID NO: 619) 0.87 8123625 (SEQ ID NO: 620) 0.87 8123542 (SEQ ID NO: 621) 0.87 8123657 (SEQ ID NO: 622) 0.86 8123628 (SEQ ID NO: 623) 0.86 8123524 (SEQ ID NO: 624) 0.85 8123461 (SEQ ID NO: 625) 0.85 8123566 (SEQ ID NO: 626) 0.84 8123635 (SEQ ID NO: 627) 0.84 8123518 (SEQ ID NO: 628) 0.83 8123591 (SEQ ID NO: 629) 0.81 8123646 (SEQ ID NO: 630) 0.80 8123528 (SEQ ID NO: 631) 0.80 8123567 (SEQ ID NO: 632) 0.80 8123668 (SEQ ID NO: 633) 0.80 8123602 (SEQ ID NO: 634) 0.78 8123643 (SEQ ID NO: 635) 0.77 8123656 (SEQ ID NO: 636) 0.76 8123696 (SEQ ID NO: 637) 0.76 8123539 (SEQ ID NO: 638) 0.72 8123606 (SEQ ID NO: 639) 0.72

[0614]

[0615] 8123683 (SEQ IDNO: 640) 0.71 8123639 (SEQ ID NO: 641) 0.71 8123618 (SEQ IDNO: 642) 0.70 8123697 (SEQ ID NO: 643) 0.70 8123491 (SEQ IDNO: 644) 0.67 8123526 (SEQ IDNO: 645) 0.67 8123546 (SEQ ID NO: 646) 0.65 8123650 (SEQ ID NO: 647) 0.65 8123463 (SEQ IDNO: 648) 0.65 8123741 (SEQ IDNO: 649) 0.63 8123573 (SEQ IDNO: 650) 0.63 8123603 (SEQ IDNO: 651) 0.61 8123462 (SEQ ID NO: 652) 0.58 8123688 (SEQ ID NO: 653) 0.56 8123459 (SEQ ID NO: 654) 0.55 8123580 (SEQ IDNO: 655) 0.53 8123544 (SEQ IDNO: 656) 0.53 8123670 (SEQ IDNO: 657) 0.51 8123536 (SEQ ID NO: 658) 0.51 8123743 (SEQ IDNO: 659) 0.51 8123629 (SEQ ID NO: 660) 0.50 8123761 (SEQ IDNO: 661) 0.49 8123720 (SEQ ID NO: 662) 0.49 8123600 (SEQ IDNO: 663) 0.47 8123624 (SEQ ID NO: 664) 0.46

[0616]

[0617] 8123713 (SEQ ID NO: 665) 0.46 8123673 (SEQ ID NO: 666) 0.44 8123550 (SEQ ID NO: 667) 0.44 8123555 (SEQ ID NO: 668) 0.44 8123716 (SEQ ID NO: 669) 0.42 8123797 (SEQ ID NO: 670) 0.42 8123677 (SEQ ID NO: 671) 0.42 8123666 (SEQ ID NO: 672) 0.41 8123576 (SEQ ID NO: 673) 0.41 8123605 (SEQ ID NO: 674) 0.38 8123693 (SEQ ID NO: 675) 0.38 8123521 (SEQ ID NO: 676) 0.38 8123616 (SEQ ID NO: 677) 0.37 8123704 (SEQ ID NO: 678) 0.36 8123565 (SEQ ID NO: 679) 0.36 8123485 (SEQ ID NO: 680) 0.36 8123694 (SEQ ID NO: 681) 0.35 8123679 (SEQ ID NO: 682) 0.34 8123802 (SEQ ID NO: 683) 0.34 8123574 (SEQ ID NO: 684) 0.34 8123740 (SEQ ID NO: 685) 0.33 8123699 (SEQ ID NO: 686) 0.33 8123700 (SEQ ID NO: 687) 0.31 8123835 (SEQ ID NO: 688) 0.30 8123589 (SEQ ID NO: 689) 0.29

[0618]

[0619] 8123789 (SEQ ID NO: 690) 0.28

[0620] 8123794 (SEQ ID NO: 691) 0.27

[0621] 8123698 (SEQ ID NO: 692) 0.27

[0622] 8123553 (SEQ ID NO: 693) 0.26

[0623] 8123859 (SEQ ID NO: 694) 0.26

[0624] 8123799 (SEQ ID NO: 695) 0.25

[0625] 8123684 (SEQ ID NO: 696) 0.22

[0626] 8123709 (SEQ ID NO: 697) 0.22

[0627] 8123558 (SEQ ID NO: 698) 0.22

[0628] 8123710 (SEQ ID NO: 699) 0.18

[0629] 8123552 (SEQ ID NO: 700) 0.16

[0630] 8123549 (SEQ ID NO: 701) 0.13

[0631] 8123785 (SEQ ID NO: 702) 0.13

[0632] 8123718 (SEQ ID NO: 703) 0.13

[0633] 8123727 (SEQ ID NO: 704) 0.12

[0634] 8123557 (SEQ ID NO: 705) 0.11

[0635] 8123560 (SEQ ID NO: 706) 0.11

[0636] 8123792 (SEQ ID NO: 707) 0.10

[0637] 8123795 (SEQ ID NO: 708) 0.10

[0638] 8123689 (SEQ ID NO: 709) 0.09

[0639] 8123642 (SEQ ID NO: 710) 0.04

[0640] 8123619 (SEQ ID NO: 711) 0.04

[0641]

[0642] Example 15: Generation and Screening of QS-21 Derivatization Enzymes

[0643] This example describes the identification of enzymes capable of modifying QS-21. QS-21 is a commercially available product (Desert King International, Chula Vista, CA; desertking.

[0644]

[0645] S-' . Examples 15.1 and 15.2 were performed without the use of any government funding.

[0646] 15.1 Phosphotransferase mediated production of mono-phosphorylated QS-21

[0647] This example describes the identification of phosphotransferase variants that are functional in phosphorylating QS-21.

[0648] To identify phosphotransferase variants that were capable of phosphorylating QS-21, a metagenomic library of approximately 500 variants was generated based on Bombyx mori phosphotransferase (Q0PCR8; SEQ ID NO: 1143), a Candidates Entotheonella phosphotransferase (A0A068PCD5; SEQ ID NO: 1144), and Psilocybe cubensis phosphotransferase (P0DPA8; SEQ ID NO: 1145) sequences.

[0649] The library was codon-optimized for expression in Escherichia coli and placed behind the T7 lac promoter and was transformed into the E. coli KRX host strain.

[0650] To initiate cell growth in preparation for screening, glycerol stocks of the phosphotransferase variant transformants were thawed at room temperature. Then, 500 pL of autoinduction media (10 g / L tryptone, 5 g / L yeast extract, 13.52 mg / L Ferric Chloride, 2.94 mg / L Calcium Chloride, 1.98 mg / L Manganese Chloride, 2.88 mg / L Zinc Sulfate, 476 pg / L Cobalt Chloride, 341 pg / L Cupric Chloride, 455 pg / L Nickel Chloride, 489 pg / L Sodium Molybdate, 346 pg / L Sodium Selenite, 124 pg / L Boric Acid, 120 mg / L Magnesium Sulfate, 3.3 g / L Ammonium Sulfate, 6.8 g / L Monobasic Potassium Phosphate, 15.7 g / L Disodium phosphate, 5 g / L Glycerol, 0.5 g / L Glucose, 2 g / L Lactose, 0.5 mg / L Rhamnose and 100 mg / L Carbenicillin) was added to each well of a 96-well 2mL deep plate, and 5pL aliquots of glycerol stock was added to each well. The plate was incubated at 30°C at 1,000 rpm and 80% humidity for 18 hours. Next, 60 pL of the cultures were compressed to a 384w plate and centrifuged at 4°C, 4000g for 10 minutes, after which the culture supernatant was discarded and cell pellet was stored in -80°C, overnight. The cell pellet was thawed at room temperature. Then, 60 pL of lysis buffer (EDTA-free protease inhibitors, 20 mM Tris-HCl (pH 6.8), 0.5x BugBuster, 0.1 mL / L Benzoase, 0.5 mL / L Lysozyme) was added to each cell pellet; the plate was incubated at 30°C and 1,000 rpm for 30 minutes followed by 10 minutes centrifugation at 4000rpm.

[0651] To screen for phosphotransferase activity, the E. coli lysate was screened using QS-21 as substrate. Next, 46 pL of reaction buffer (11 mM Tris-HCl (pH 6.8), 110 pM QS-21, 275 pM ATP, and 1.1 mM magnesium chloride) was added to each well of a 384-well plate along with 4 pL of E. coli lysate supernatant; the plate was sealed and then incubated at room temperature for 18 hours. A new 384-well plate was filled with 110 pL of extraction buffer (50% Methanol, 5 mg / L hederacoside C) along with 12 pL reaction sample and centrifuge. The formation of phosphorylated QS-21 in the reaction supernatant was assayed with Echo-MS, sample-prep LC-MS, orLC-MS.

[0652] The screen identified 19 phosphotransferase enzymes (see Table 15) where upon reaction with QS-21 and ATP in vitro, formed 2 product peaks that correspond to a precursor ion at m / z 1033.942 ([M-2H]2), detected predominantly in the doubly charged state.

[0653] The m / z value of one product is consistent with the calculated mass (m / z 1033.942; see FIGs. 3A-3B) for a mono-phosphorylated QS-21 [M-2H]2", and is believed to be QS-21 with an addition of a single phosphate group. The different retention times between the 2 isomers are believed to be different sites of phosphorylation, although the specific phosphorylation sites were not determined.

[0654] Table 15. QS-21 Phosphtransferases from library screen

[0655] QS-21 Phosphotransferase QS-21 Derivatives Profile

[0656] 6003250 (SEQ ID NO: 1254) Monophosphorylated QS-21 Isomer 1

[0657] 6003270 (SEQ ID NO: 1257) Monophosphorylated QS-21 Isomer 1,

[0658] Monophosphorylated QS-21 Isomer 2

[0659] 6003289 (SEQ ID NO: 1258) Monophosphorylated QS-21 Isomer 1

[0660] 6003352 (SEQ ID NO: 1264) Monophosphorylated QS-21 Isomer 1

[0661] 6003380 (SEQ ID NO: 1267) Monophosphorylated QS-21 Isomer 1,

[0662] Monophosphorylated QS-21 Isomer 2

[0663] 6003386 (SEQ ID NO: 1269) Monophosphorylated QS-21 Isomer 1,

[0664] Monophosphorylated QS-21 Isomer 2

[0665] 6003408 (SEQ ID NO: 1272) Monophosphorylated QS-21 Isomer 1

[0666] 6003416 (SEQ ID NO: 1274) Monophosphorylated QS-21 Isomer 1,

[0667] Monophosphorylated QS-21 Isomer 2

[0668] 6003422 (SEQ ID NO: 1275) Monophosphorylated QS-21 Isomer 1,

[0669] Monophosphorylated QS-21 Isomer 2

[0670]

[0671] 6003432 (SEQ ID NO: 1277) Monophosphorylated QS-21 Isomer 1

[0672] 6003439 (SEQ ID NO: 1280) Monophosphorylated QS-21 Isomer 1,

[0673] Monophosphorylated QS-21 Isomer 2

[0674] 6003444 (SEQ ID NO: 1282) Monophosphorylated QS-21 Isomer 1,

[0675] Monophosphorylated QS-21 Isomer 2

[0676] 6003461 (SEQ ID NO: 1283) Monophosphorylated QS-21 Isomer 2

[0677] 6003476 (SEQ ID NO: 1285) Monophosphorylated QS-21 Isomer 1,

[0678] Monophosphorylated QS-21 Isomer 2

[0679] 6003486 (SEQ ID NO: 1286) Monophosphorylated QS-21 Isomer 1,

[0680] Monophosphorylated QS-21 Isomer 2

[0681] 6003488 (SEQ ID NO: 1287) Monophosphorylated QS-21 Isomer 1,

[0682] Monophosphorylated QS-21 Isomer 2

[0683] 6003535 (SEQ ID NO: 1293) Monophosphorylated QS-21 Isomer 1,

[0684] Monophosphorylated QS-21 Isomer 2

[0685] 6003536 (SEQ ID NO: 1294) Monophosphorylated QS-21 Isomer 1,

[0686] Monophosphorylated QS-21 Isomer 2

[0687] 6003572 (SEQ ID NO: 1296) Monophosphorylated QS-21 Isomer 1,

[0688] Monophosphorylated QS-21 Isomer 2

[0689]

[0690] 15.2 UDP-Glycosyltransferase mediated production of mono-glucosylated QS-21

[0691] This example describes the identification of UDP -glycosyltransferase variants that were functional in glycosylating QS-21.

[0692] To identify UDP-glycosyltransferase variants that were capable of glycosylating QS-21, a metagenomic library of approximately 1500 variants was generated based on three Spinacia oleracea UDP-glycosyltransferases (KNA07536.1; SEQ ID NO: 1146; and KNA23861.1; SEQ ID NO: 1147; and KNA13786.1; SEQ ID NO: 1148) and four Quillaja saponaria UDP-glycosyltransferases (QsGalT; SEQ ID NO: 44; and Qs_2015879; SEQ ID NO: 1149; andKAJ7951238.1; SEQ ID NO: 1150; and QsC3XylT; SEQ ID NO: 45) sequences.

[0693] The library was codon-optimized for expression in Escherichia coli and placed behind the T7 lac promoter and was transformed into the E. coli BL21-AI host strain. The cell cultivation and lysis procedure were as described in Example 15.1, with addition of 0.25 mg / L arabinose instead of 0.5 mg / L rhamnose.

[0694] To screen for UDP-glycosyltransf erase activity, the E. coli lysate was screened using QS-21 as substrate. 46 pL of reaction buffer (11 mM Tris-HCl (pH 6.8), 110 pM QS-21, 1.1 mM UDP -D-glucose, 11 units / mL shrimp alkaline phosphatase (rSAP) and 1.1 mM magnesium chloride) was added to each well of a 384-well plate along with 4 pL of E. coli lysate supernatant. The reaction and extraction procedures were as described in Example 15.1.

[0695] The screen identified 26 UDP-glycosyltransferase enzymes (see Table 16) where upon reaction with QS-21 and UDP -D-Glucose in vitro, formed 7 product peaks with a precursor ion at m / z 1074.98 ([M-2H]2’).

[0696] This m / z value of one product is consistent with the calculated mass (m / z 1074.98; see FIGs. 4A-4G) for a mono-glucosylated QS-2 1[M-H] ", and is believed to be QS-21 with an addition of a single glucose sugar. The different retention times between the 7 isomers are believed to be different sites of glycosylation, although the specific glycosylation sites were not determined.

[0697] Table 16. QS-21 LTDP-glycosyltransferase from library screen

[0698] QS-21 UDP-glucosyltransferase QS-21 Derivatives Profile

[0699] 6065558(SEQ ID NO: 1298) Monoglycosylated QS-21 isomer 1,

[0700] Monoglycosylated QS-21 isomer 4

[0701] 6065585(SEQ ID NO: 1299) Monoglycosylated QS-21 isomer 1,

[0702] Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 5, Monoglycosylated QS-21 isomer 6

[0703] 6066001(SEQ ID NO: 1300) Monoglycosylated QS-21 isomer 5,

[0704] Monoglycosylated QS-21 isomer 6, Monoglycosylated QS-21 isomer 7

[0705] 6066132(SEQ ID NO: 1301) Monoglycosylated QS-21 isomer 1,

[0706] Monoglycosylated QS-21 isomer 4

[0707] 6066240(SEQ ID NO: 1302) Monoglycosylated QS-21 isomer 1,

[0708] Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 5, Monoglycosylated QS-21 isomer 6

[0709]

[0710] 6066256(SEQ ID NO: 1303) Monoglycosylated QS-21 isomer 1,

[0711] Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 6 6066258(SEQ ID NO: 1304) Monoglycosylated QS-21 isomer 1,

[0712] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 6 6066274(SEQ ID NO: 1305) Monoglycosylated QS-21 isomer 1,

[0713] Monoglycosylated QS-21 isomer 5, Monoglycosylated QS-21 isomer 6 6066304(SEQ ID NO: 1306) Monoglycosylated QS-21 isomer 1,

[0714] Monoglycosylated QS-21 isomer 5, Monoglycosylated QS-21 isomer 6 6066319(SEQ ID NO: 1307) Monoglycosylated QS-21 isomer 1,

[0715] Monoglycosylated QS-21 isomer 2, Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 6 6066335(SEQ ID NO: 1308) Monoglycosylated QS-21 isomer 4,

[0716] Monoglycosylated QS-21 isomer 6 6066336(SEQ ID NO: 1309) Monoglycosylated QS-21 isomer 1,

[0717] Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 6 6066338(SEQ ID NO: 1310) Monoglycosylated QS-21 isomer 2,

[0718] Monoglycosylated QS-21 isomer 5, Monoglycosylated QS-21 isomer 6 6066380(SEQ ID NO: 1311) Monoglycosylated QS-21 isomer 2,

[0719] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 5, Monoglycosylated QS-21 isomer 6 6066420(SEQ ID NO: 1312) Monoglycosylated QS-21 isomer 3,

[0720] Monoglycosylated QS-21 isomer 5 6066506(SEQ ID NO: 1313) Monoglycosylated QS-21 isomer 2,

[0721] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 6

[0722]

[0723] 6066550(SEQ ID NO: 1314) Monoglycosylated QS-21 isomer 2,

[0724] Monoglycosylated QS-21 isomer 6 606663 O(SEQ ID NO: 1315) Monoglycosylated QS-21 isomer 1,

[0725] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 6 6066673 (SEQ ID NO: 1316) Monoglycosylated QS-21 isomer 3,

[0726] Monoglycosylated QS-21 isomer 6 6066692(SEQ ID NO: 1317) Monoglycosylated QS-21 isomer 1,

[0727] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 5, Monoglycosylated QS-21 isomer 6 6066773 (SEQ ID NO: 1318) Monoglycosylated QS-21 isomer 2,

[0728] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 5, Monoglycosylated QS-21 isomer 6 6066827(SEQ ID NO: 1319) Monoglycosylated QS-21 isomer 1,

[0729] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 6 6066861(SEQ ID NO: 1320) Monoglycosylated QS-21 isomer 1,

[0730] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 6, Monoglycosylated QS-21 isomer 7 6066867(SEQ ID NO: 1321) Monoglycosylated QS-21 isomer 1,

[0731] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 5, Monoglycosylated QS-21 isomer 6 6066901(SEQ ID NO: 1322) Monoglycosylated QS-21 isomer 3,

[0732] Monoglycosylated QS-21 isomer 6 6066910(SEQ ID NO: 1323) Monoglycosylated QS-21 isomer 1,

[0733] Monoglycosylated QS-21 isomer 3, Monoglycosylated QS-21 isomer 4, Monoglycosylated QS-21 isomer 6

[0734]

[0735] 15.3 Methyltransferase mediated production of mono-methylated QS-21

[0736] This example describes the identification of methyltransferase variants that were functional in methylating QS-21.

[0737] To identify methyltransferase variants that were capable of methylating QS-21, a metagenomic library of approximately 1000 variants was generated based on a Streptomyces niveus methyltransferase (Q9L9F2; SEQ ID NO: 1151), a Streptosporangium amethystogenes methyltransferase (M5ABB9; SEQ ID NO: 1152), three Streptomyces olivaceus methyltransferases (Q9AJU2; SEQ ID NO: 1153; Q9AJU1; SEQ ID NO: 1154; and Q9AJU0; SEQ ID NO: 1155), two Streptomyces fradiae methyltransferases (Q9S4D5; SEQ ID NO: 1156; Q9ZHQ4; SEQ ID NO: 1157), Streptomyces sp. SCSIO 01127 methyltransferase (M9T242; SEQ ID NO: 1158), a Streptomyces sp. Al(2016) methyltransferase (A0A172MB36; SEQ ID NO: 1159), a Streptomyces sp. KCTC 0041BP methyltransferase (Q331Q6; SEQ ID NO: 1160), and two Streptomyces kanamyceticus methyltransferase (Q6L726; SEQ ID NO: 1161; and Q65CE6; SEQ ID NO: 1162) sequences.

[0738] The library was codon-optimized for expression in Escherichia coli and placed behind the T7 lac promoter and was transformed into the E. coli KRX host strain. The cell cultivation and lysis procedure were as described in Example 15.1.

[0739] To characterize methyltransferase activity, the E. coli lysate was screened using QS-21 as substrate. Then, 46 pL of reaction mix (11 mM Tris-HCl (pH 6.8), 1.1 mM magnesium chloride, 220 mM S-adenosyl methionine (SAM)) was added to a new 384-well plate along with 4 pL lysate supernatant; the plate was incubated at room temperature for 18 hours. The reaction and extraction procedures were as described in Example 15.1.

[0740] The screen identified 6 methyltransferase enzymes (see Table 17) where upon reaction with QS-21 and SAM in vitro, formed 2 product peaks with a parent ion with m / z 2001.934 and 2015.951 ([M-H]’).

[0741] The m / z value of one product is consistent with the predicted mass (m / z 2001.934; see FIGs. 5A-5B) for a mono-methylated QS-21 [M-H] ", and is believed to be QS-21 with addition of a single methyl group. The different retention times between the 2 isomers are believed to be different sites of methylation, although the specific methylation site was not determined.

[0742] The m / z value of the second product is consistent with the predicted mass (m / z 2015.951; see FIGs. 5C-5D) for a di-methylated QS-21 [M-H]", and is believed to be QS-21 with addition of two methyl groups. The different retention times between the 2 isomers are believed to be different sites of methylation, although the specific methylation sites were not determined.

[0743] Table 17. QS-21 Methyltransferase from library screen

[0744] QS-21 Methytransferase QS-21 Derivatives Profile

[0745] 6304698 (SEQ ID NO: 1324) Mono-methylated QS-21 Isomer 1, mono-methylated QS- 21 isomer 2, di-methylated QS-21 isomer 1, di-methylated QS-21 isomer 2

[0746] 6304708 (SEQ ID NO: 1325) Mono-methylated QS-21 isomer 1, mono-methylated QS- 21 isomer 2

[0747] 6304768 (SEQ ID NO: 1326) Mono-methylated QS-21 isomer (Isomer type not determined)

[0748] 6305140 (SEQ ID NO: 1327) Mono-methylated QS-21 isomer (Isomer type not determined)

[0749] 6305213 (SEQ ID NO: 1328) Mono-methylated QS-21 isomer (Isomer type not determined)

[0750] 6305488 (SEQ ID NO: 1329) Mono-methylated QS-21 isomer 1, mono-methylated QS- 21 isomer 2

[0751]

[0752] 15.4 Sulfotransferase mediated production of mono-sulfated QS-21

[0753] This example describes the identification of sulfotransferase variants that were functional in sulfating QS-21.

[0754] To identify sulfotransferase variants that were capable of sulfating QS-21, a metagenomic library of approximately 500 variants was generated based on a. Bombyx mori sulfotransferase (H9JMT9; SEQ ID NO: 1163), Bombyx mori sulfotransferase (Q3LFP9; SEQ ID NO: 1164), Bombyx mori sulfotransferase (A0PCF9; SEQ ID NO: 1165), Bombyx mori sulfotransferase (XP 004932587.1; SEQ ID NO: 1166), Bombyx mori sulfotransferase (XP_037872720.1; SEQ ID NO: 1167), Tribolium castaneum sulfotransferase (MW664928.1; SEQ ID NO: 1168), Aedes aegypti sulfotransferase (QI 76K0; SEQ ID NO: 1169), Anopheles gambiae sulfotransferase (Q7PXJ0; SEQ ID NO: 1170), Drosophila melanogaster sulfotransferase (A1Z9J8; SEQ ID NO: 1171), Spodoptera frugiperda sulfotransferase (Q26490; SEQ ID NO: 1172), Ixodes scapularis sulfotransferase (DQ066225.1; SEQ ID NO: 1173), and Ixodes scapularis sulfotransferase (DQ066226.1; SEQ ID NO: 1174) sequences.

[0755] The library was codon-optimized for expression in Escherichia coli and placed behind the T7 lac promoter and was transformed into the E. coli KRX host strain. The cell cultivation and lysis procedure were as described in Example 15.1.

[0756] To characterize sulfotransferase activity, the E. coli lysate was screened using QS-21 as substrate. Then, 45 pL of reaction mix (11 mM Tris-HCl (pH 6.8), 110 pM 3'-phosphoadenosine-5'-phosphosulfate (PAPS)) was added to a new 384-well plate along with 30 pL lysate supernatant; the plate was incubated at 37°C for 18 hours. A 384-well plate was filled with 70 pL of extraction buffer (50% Methanol, 5 pM hederacoside C) along with 48 pL reaction sample; the plate was centrifuged at 3000 rpm for 5 min. The formation of sulfated QS-21 was determined by assaying the reaction supernatant with Echo-MS, sample-prep LC-MS, or LC-MS.

[0757] The screen identified 18 sulfotransferases (see Table 18) where upon reaction with QS-21 and PAPS in vitro, formed 4 new product peaks with a precursor ion m / z 1033.93 ([M-2H]2' ), detected predominantly in the doubly charged state.

[0758] The m / z values for the product are consistent with the calculated mass (m / z 1033.93; see FIG. 6A-6D) for mono-sulfated QS-21 [M-H] ", and is believed to be QS-21 with addition of a single sulfate group. The different retention times between the 4 isomers are believed to be different sites of sulfation, although the specific sulfation site was not determined.

[0759] Table 18. QS-21 Sulotransferase from library screen

[0760] QS-21 Sulfotransferase QS-21 Derivatives Profile

[0761] 7750528 (SEQ ID NO: 1330) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0762] Isomer 3, mono-sulfated QS-21 Isomer 4

[0763] 7750887 (SEQ ID NO: 1331) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0764] Isomer 4

[0765] 7750643 (SEQ ID NO: 1332) Mono-sulfated QS-21 Isomer 1, Mono-sulfated QS-21

[0766] Isomer 4

[0767] 7750601 (SEQ ID NO: 1333) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0768] Isomer 2, mono-sulfated QS-21 Isomer 4

[0769]

[0770] 7750591 (SEQ IDNO: 1334) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0771] Isomer 2, mono-sulfated QS-21 Isomer 3, mono-sulfated QS-21 Isomer 4

[0772] 7751995 (SEQ IDNO: 1335) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0773] Isomer 3, mono-sulfated QS-21 Isomer 4

[0774] 7751150 (SEQ IDNO: 1336) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0775] Isomer 3, mono-sulfated QS-21 Isomer 4

[0776] 7751153 (SEQ IDNO: 1337) Mono-sulfated QS-21 (Isomer not determined) 7750971 (SEQ IDNO: 1338) Mono-sulfated QS-21 (Isomer not determined) 7750504 (SEQ ID NO: 1339) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0777] Isomer 2, mono-sulfated QS-21 Isomer 3

[0778] 7750494 (SEQ ID NO: 1340) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0779] Isomer 2, mono-sulfated QS-21 Isomer 3, mono-sulfated QS-21 Isomer 4

[0780] 7750524 (SEQ ID NO: 1341) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0781] Isomer 3, mono-sulfated QS-21 Isomer 4

[0782] 7750650 (SEQ ID NO: 1342) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0783] Isomer 2, mono-sulfated QS-21 Isomer 3, mono-sulfated QS-21 Isomer 4

[0784] 7751180 (SEQ ID NO: 1343) Mono-sulfated QS-21 (Isomer not determined) 7750875 (SEQ ID NO: 1344) Mono-sulfated QS-21 Isomer 2, mono-sulfated QS-21

[0785] Isomer 4

[0786] 7751204 (SEQ ID NO: 1345) Mono-sulfated QS-21 (Isomer not determined) 7750630 (SEQ ID NO: 1346) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0787] Isomer 4

[0788] 7750278 (SEQ ID NO: 1347) Mono-sulfated QS-21 Isomer 1, mono-sulfated QS-21

[0789] Isomer 2, mono-sulfated QS-21 Isomer 3, mono-sulfated QS-21 Isomer 4

[0790]

[0791] 15.5 Acyltransferase mediated production of acetylated, malonylated, and benzoylated QS- 21

[0792] This example describes the identification of acyltransferase variants that were functional in acylating QS-21.

[0793] To identify acyltransferase variants that were capable of acylating QS-21, a metagenomic library of approximately 1500 variants was generated based on the Spinacia oleracea acyltransferase (A0A0K9QYV8; SEQ ID NO: 1175), two Salvia splendens acyltransferases (Q6TXD2; SEQ ID NO: 1176; Q8W1W9; SEQ ID NO: 1177), three Chrysanthemum morifolium acyltransferases (Q6WB12; SEQ ID NO: 1178; Q6WB13; SEQ ID NO: 1179;

[0794] A4PHY4; SEQ ID NO: 1180), Gentiana triflora acyltransferase (Q9ZWR8; SEQ ID NO: 1181), Dahlia pinnata acyltransferase (Q8GSN8; SEQ ID NO: 1182), Nicotiana tabacum acyltransferase (Q589 Y0; SEQ ID NO: 1183), Verbena hybrida acyltransferase (Q6RFS6; SEQ ID NO: 1184), Lamium purpureum acyltransferase (Q6RFS5; SEQ ID NO: 1185), Perilla frutescens acyltransferase (Q9MBC1; SEQ ID NO: 1186), six Arabidopsis thaliana acyltransferases (Q9LJB4; SEQ ID NO: 1187; Q9ZWB4; SEQ ID NO: 1188; A0A2H1ZEA8; SEQ ID NO: 1189; Q940Z5; SEQ ID NO: 1190; Q9LRQ8; SEQ ID NO: 1191; Q9SV07; SEQ ID NO: 1192), Glycine max acyltransferase (A7BIC9; SEQ ID NO: 1193), Ax Medicago truncatula acyltransferases (B4Y0U0; SEQ ID NO: 1194; B4Y0U1; SEQ ID NO: 1195;

[0795] B4Y0U2; SEQ ID NO: 1196; F4ZG53; SEQ ID NO: 1197; F4ZG54; SEQ ID NO: 1198;

[0796] F4ZG55; SEQ ID NO: 1199), two Crocosmia x crocosmiiflora acyltransferases (A0A2Z5CVQ4; SEQ ID NO: 1200; A0A2Z5CVK5; SEQ ID NO: 1201), five Taxus cuspidata acyltransferases (Q9FPW3; SEQ ID NO: 1202; Q9M6F0; SEQ ID NO: 1203; Q9M6E2; SEQ ID NO: 1204; Q8H2B5; SEQ ID NO: 1205; Q8LL69; SEQ ID NO: 1206), two Lavandula x intermedia acyltransferases (A0A0K0LBP0; SEQ ID NO: 1207; A0A0K0LCG5; SEQ ID NO: 1208), and Celastrus angulatus acyltransferase (A0A7D5UIR6; SEQ ID NO: 1209) sequences.

[0797] The library was codon-optimized for expression in Escherichia coli and placed behind the T7 lac promoter and was transformed into the E. coli KRX host strain. The cell cultivation and lysis procedure were as described in Example 15.1.

[0798] To characterize acyltransferase activity, the / . coli lysate was screened using QS-21 as substrate. Then, 46 pL of reaction mix (11 mM Tris-HCl (pH 6.8), 1.1 mM magnesium chloride, 1.43 mM acetyl-coA, malonyl-coA, and benzoyl-coA) was added to a new 384-well plate along with 4 pL lysate supernatant; the plate was incubated at room temperature for 18 hours. The reaction and extraction procedures were as described in Example 15.1.

[0799] The screen identified 147 acyltransferases (see Table 19). Some of the identified acyltransferases, where upon reaction with QS-21 and acetyl-coA in vitro, formed 4 new monoacetylated product peaks with a parent ion m / z 2029.93 ([M-H] ) and a di-acetylated product peak with a parent ion m / z 2071.9375 ([M-H] ). Some of the identified acyltransferases, where upon reaction with QS-21 and malonyl-coA in vitro, formed mono-malonylated product peaks with a parent ion m / z 1036.4542 ([M-2H]2"), detected predominantly in the doubly charged state. Some of the identified acyltransferases, where upon reaction with QS-21 and benzoyl-coA in vitro, formed mono-benzoylated product peaks with a parent ion m / z 2091.9316 ([M-H]").

[0800] The m / z value of one product is consistent with the calculated mass (m / z 2071.9375; see FIGs. 7A-7D) for mono-acetylated QS-21 [M-H] ", and is believed to be QS-21 with addition of one acetyl group. The different retention times between the 4 isomers are believed to be different sites of acetylation, although the specific acylation site was not determined.

[0801] The m / z value of one product is consistent with the calculated mass (m / z 2071.9375) for di-acetylated QS-21 [M-H] , and is believed to be QS-21 with addition of two acetyl groups. The site of acylation or product isomers were not determined.

[0802] The m / z value of one product is consistent with the calculated mass (m / z 1036.4542 ([M-2H]2"; see FIG. 7E) for mono-malonylated QS-21 [M-H] , and is believed to be QS-21 with addition of one malonyl groups. The site of acylation or product isomers were not determined.

[0803] The m / z value of one product is consistent with the calculated mass (m / z2091.9316 ([M-H]_; see FIG. 7F) for mono-benzoylated QS-21 [M-H] , and is believed to be QS-21 with addition of one benzoyl groups. The site of acylation or product isomers were not determined.

[0804] Table 19. QS-21 Acyltransferases from library screen

[0805] QS-21 Acyltransferase QS-21 Derivatives Profile

[0806] 6292883 (SEQ ID NO: 1348) Mono-malonyl QS-21 (Isomer not determined) 6292885 (SEQ ID NO: 1349) Mono-acetyl QS-21 isomer 2, mono-acetyl QS-21 isomer 4, mono-malonyl QS-21

[0807] 6292890 (SEQ ID NO: 1350) Mono-acetyl QS-21 (Isomer not determined)

[0808] 6292906 (SEQ ID NO: 1351) Mono-malonyl QS-21

[0809]

[0810] 6292913 (SEQ ID NO: 1352) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0811] 6292924 (SEQ ID NO: 1353) Mono-benzoylated QS-21

[0812] 6292926 (SEQ ID NO: 1354) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0813] 6292940 (SEQ ID NO: 1355) Mono-acetyl QS-21 (Isomer not determined) 6292946 (SEQ ID NO: 1356) Mono-benzoylated QS-21 (Isomer not determined) 6292947 (SEQ ID NO: 1357) Mono-acetyl QS-21 (Isomer not determined) 6292958 (SEQ ID NO: 1358) Mono-acetyl QS-21 (Isomer not determined) 6292989 (SEQ ID NO: 1359) Mono-acetyl QS-21 isomer 4

[0814] 6293010 (SEQ ID NO: 1360) Mono-acetyl QS-21 isomer 1, mono-acetyl QS-21 isomer 2, mono-acetyl QS-21 isomer 3, mono-acetyl QS-21 isomer 4, mono-malonyl QS-21

[0815] 6293025 (SEQ ID NO: 1361) Mono-acetyl QS-21 (Isomer not determined) 6293036 (SEQ ID NO: 1362) Mono-benzoylated QS-21 (Isomer not determined) 6293037 (SEQ ID NO: 1363) Di-acetylated QS-21 (Isomer not determined) 6293063 (SEQ ID NO: 1364) Di-acetylated QS-21, mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0816] 6293077 (SEQ ID NO: 1365) Mono-acetyl QS-21 (Isomer not determined) 6293089 (SEQ ID NO: 1366) Mono-malonyl QS-21 (Isomer not determined) 6293092 (SEQ ID NO: 1368) Mono-acetyl QS-21 (Isomer not determined) 6293093 (SEQ ID NO: 1369) Mono-acetyl QS-21 (Isomer not determined) 6293103 (SEQ ID NO: 1370) Mono-benzoylated QS-21 (Isomer not determined) 6293104 (SEQ ID NO: 1371) Mono-acetyl QS-21 (Isomer not determined) 6293105 (SEQ ID NO: 1372) Mono-acetyl QS-2, mono-malonyl QS-21 (Isomer not determined)

[0817] 6293137 (SEQ ID NO: 1374) Mono-acetyl QS-21 (Isomer not determined)

[0818]

[0819] 6293169 (SEQ ID NO: 1375) Mono-acetyl QS-21 isomer 3, mono-acetyl QS-21 isomer 4, mono-malonyl QS-21

[0820] 6293173 (SEQ ID NO: 1376) Mono-benzoylated QS-21 (Isomer not determined) 6293175 (SEQ ID NO: 1377) Mono-acetyl QS-21 (Isomer not determined) 6293181 (SEQ ID NO: 1378) Mono-benzoylated QS-21 (Isomer not determined) 6293188 (SEQ ID NO: 1379) Mono-acetyl QS-21 (Isomer not determined) 6293202 (SEQ ID NO: 1381) Mono-malonyl QS-21 (Isomer not determined) 6293213 (SEQ ID NO: 1382) Mono-acetyl QS-21, mono-benzoylated QS-21, monomalonyl QS-21 (Isomer not determined)

[0821] 6293217 (SEQ ID NO: 1383) Mono-acetyl QS-21 isomer 4, mono-malonyl QS-21 6293219 (SEQ ID NO: 1384) Mono-acetyl QS-21 (Isomer not determined) 6293221 (SEQ ID NO: 1385) Mono-acetyl QS-21, mono-benzoyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0822] 6293223 (SEQ ID NO: 1386) Mono-acetyl QS-21 (Isomer not determined) 6293225 (SEQ ID NO: 1387) Mono-acetyl QS-21, mono-benzoyl QS-21 (Isomer not determined)

[0823] 6293263 (SEQ ID NO: 1389) Mono-acetyl QS-21 (Isomer not determined) 6293265 (SEQ ID NO: 1390) Mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0824] 6293270 (SEQ ID NO: 1391) Mono-benzoyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0825] 6293275 (SEQ ID NO: 1392) Mono-acetyl QS-21 (Isomer not determined) 6293296 (SEQ ID NO: 1394) Di-acetyl QS-21, mono-acetyl QS-21, mono-malonyl QS- 21 (Isomer not determined)

[0826] 6293300 (SEQ ID NO: 1395) Di-acetyl QS-21 (Isomer not determined)

[0827] 6293312 (SEQ ID NO: 1396) Mono-malonyl QS-21 (Isomer not determined) 6293313 (SEQ ID NO: 1397) Mono-acetyl QS-21, mono-benzoylated QS-21, monomalonyl QS-21 (Isomer not determined)

[0828]

[0829] 6293330 (SEQ ID NO: 1398) Mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0830] 6293332 (SEQ ID NO: 1399) Mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0831] 6293333 (SEQ ID NO: 1400) Mono-benzoylated QS-21 (Isomer not determined) 6293343 (SEQ ID NO: 1401) Mono-malonyl QS-21 (Isomer not determined) 6293359 (SEQ ID NO: 1402) Mono-benzoylated QS-21 (Isomer not determined) 6293360 (SEQ ID NO: 1403) Mono-acetyl QS-21 (Isomer not determined) 6293366 (SEQ ID NO: 1404) Mono-acetyl QS-21, mono-benzoylated QS-21, monomalonyl QS-21 (Isomer not determined)

[0832] 6293392 (SEQ ID NO: 1405) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0833] 6293394 (SEQ ID NO: 1406) Mono-acetyl QS-21, mono-benzoylated QS-21, monomalonyl QS-21 (Isomer not determined)

[0834] 6293409 (SEQ ID NO: 1407) Mono-acetyl QS-21 (Isomer not determined) 6293424 (SEQ ID NO: 1408) Mono-acetyl QS-21 (Isomer not determined) 6293442 (SEQ ID NO: 1410) Mono-malonyl QS-21 (Isomer not determined) 6293448 (SEQ ID NO: 1411) Mono-acetyl QS-21 (Isomer not determined) 6293470 (SEQ ID NO: 1412) Mono-malonyl QS-21 (Isomer not determined) 6293481 (SEQ ID NO: 1413) Mono-acetyl QS-21, mono-benzoylated QS-21, monomalonyl QS-21 (Isomer not determined)

[0835] 6293486 (SEQ ID NO: 1414) Mono-malonyl QS-21 isomer

[0836] 6293496 (SEQ ID NO: 1415) Mono-malonyl QS-21 (Isomer not determined) 6293507 (SEQ ID NO: 1416) Di-acetylated QS-21 (Isomer not determined) 6293513 (SEQ ID NO: 1417) Mono-acetyl QS-21 (Isomer not determined) 6293518 (SEQ ID NO: 1418) Mono-acetyl QS-21 (Isomer not determined) 6293528 (SEQ ID NO: 1419) Mono-benzoylated QS-21 (Isomer not determined)

[0837]

[0838] 6293529 (SEQ ID NO: 1420) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0839] 6293532 (SEQ ID NO: 1421) Mono-acetyl QS-21 isomer 1, mono-acetyl QS-21 isomer 2, mono-acetyl QS-21 isomer 4

[0840] 6293552 (SEQ ID NO: 1422) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0841] 6293574 (SEQ ID NO: 1425) Mono-acetyl QS-21 isomer 1, mono-acetyl QS-21 isomer 2, mono-acetyl QS-21 isomer 4

[0842] 6293585 (SEQ ID NO: 1426) Mono-benzoylated QS-21, mono-malonyl QS-21 (Isomer not determined)

[0843] 6293595 (SEQ ID NO: 1427) Mono-acetyl QS-21 (Isomer not determined) 6293600 (SEQ ID NO: 1428) Mono-acetyl QS-21 (Isomer not determined) 6293601 (SEQ ID NO: 1429) Mono-acetyl QS-21 (Isomer not determined) 6293604 (SEQ ID NO: 1430) Di-acetylated QS-21, mono-acetyl QS-21, mono- benzoylated QS-21 (Isomer not determined)

[0844] 6293625 (SEQ ID NO: 1431) Mono-acetyl QS-21 (Isomer not determined) 6293641 (SEQ ID NO: 1432) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0845] 6293646 (SEQ ID NO: 1433) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0846] 6293648 (SEQ ID NO: 1434) Mono-malonyl QS-21

[0847] 6293650 (SEQ ID NO: 1435) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0848] 6293663 (SEQ ID NO: 1436) Mono-benzoylated QS-21 (Isomer not determined) 6293678 (SEQ ID NO: 1437) Mono-benzoylated QS-21 (Isomer not determined) 6293681 (SEQ ID NO: 1438) Mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0849] 6293684 (SEQ ID NO: 1439) Di-acetylated QS-21 (Isomer not determined) 6293686 (SEQ ID NO: 1440) Mono-benzoylated QS-21 (Isomer not determined)

[0850]

[0851] 6293687 (SEQ ID NO: 1441) Di-acetylated QS-21 (Isomer not determined) 6293693 (SEQ ID NO: 1442) Di-acetylated QS-21 (Isomer not determined) 6293699 (SEQ ID NO: 1443) Mono-malonyl QS-21 (Isomer not determined) 6293700 (SEQ ID NO: 1444) Mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0852] 6293703 (SEQ ID NO: 1445) Mono-malonyl QS-21 (Isomer not determined) 6293720 (SEQ ID NO: 1446) Di-acetylated QS-21, mono-acetyl QS-21 (Isomer not determined)

[0853] 6293734 (SEQ ID NO: 1447) Mono-malonyl QS-21 (Isomer not determined) 6293742 (SEQ ID NO: 1448) Mono-benzoylated QS-21 (Isomer not determined) 6293743 (SEQ ID NO: 1449) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0854] 6293754 (SEQ ID NO: 1450) Mono-benzoylated QS-21 (Isomer not determined) 6293769 (SEQ ID NO: 1451) Mono-malonyl QS-21 (Isomer not determined) 6293783 (SEQ ID NO: 1452) Mono-acetyl QS-21 (Isomer not determined) 6293789 (SEQ ID NO: 1453) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0855] 6293797 (SEQ ID NO: 1454) Mono-acetyl QS-21 (Isomer not determined) 6293801 (SEQ ID NO: 1455) Mono-acetyl QS-21 (Isomer not determined) 6293819 (SEQ ID NO: 1457) Mono-malonyl QS-21

[0856] 6293836 (SEQ ID NO: 1459) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0857] 6293838 (SEQ ID NO: 1460) Mono-acetyl QS-21 (Isomer not determined) 6293845 (SEQ ID NO: 1461) Mono-acetyl QS-21 (Isomer not determined) 6293873 (SEQ ID NO: 1462) Mono-acetyl QS-21 (Isomer not determined) 6293911 (SEQ ID NO: 1465) Mono-malonyl QS-21 (Isomer not determined) 6293913 (SEQ ID NO: 1466) Di-acetylated QS-21, mono-acetylated QS-21 (Isomer not

[0858]

[0859] determined)

[0860] 6293920 (SEQ ID NO: 1467) Di-acetylated QS-21 (Isomer not determined) 6293933 (SEQ ID NO: 1468) Di-acetylated QS-21 (Isomer not determined) 6293969 (SEQ ID NO: 1469) Mono-benzoylated QS-21 (Isomer not determined) 6293974 (SEQ ID NO: 1470) Mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0861] 6293975 (SEQ ID NO: 1471) Mono-benzoylated QS-21 (Isomer not determined) 6294000 (SEQ ID NO: 1472) Mono-acetyl QS-21 (Isomer not determined) 6294006 (SEQ ID NO: 1473) Mono-benzoylated QS-21 (Isomer not determined) 6294008 (SEQ ID NO: 1474) Mono-acetyl QS-21 (Isomer not determined) 6294036 (SEQ ID NO: 1475) Mono-malonyl QS-21

[0862] 6294053 (SEQ ID NO: 1476) Mono-acetyl QS-21, mono-benzoylated QS-21 (Isomer not determined)

[0863] 6294068 (SEQ ID NO: 1477) Mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0864] 6294071 (SEQ ID NO: 1479) Mono-acetyl QS-21 (Isomer not determined) 6294100 (SEQ ID NO: 1480) Mono-malonyl QS-21

[0865] 6294102 (SEQ ID NO: 1481) Mono-malonyl QS-21 (Isomer not determined) 6294104 (SEQ ID NO: 1482) Mono-malonyl QS-21 (Isomer not determined) 6294106 (SEQ ID NO: 1484) Mono-malonyl QS-21 (Isomer not determined) 6294107 (SEQ ID NO: 1485) Mono-acetyl QS-21 (Isomer not determined)

[0866] 6294110 (SEQ ID NO: 1486) Mono-malonyl QS-21 (Isomer not determined) 6294116 (SEQ ID NO: 1487) Mono-malonyl QS-21 (Isomer not determined) 6294121 (SEQ ID NO: 1488) Di-acetylated QS-21 (Isomer not determined) 6294127 (SEQ ID NO: 1489) Mono-acetyl QS-21 isomer 2

[0867] 6294142 (SEQ ID NO: 1491) Mono-malonyl QS-21

[0868]

[0869] 6294158 (SEQ ID NO: 1492) Mono-benzoylated QS-21 (Isomer not determined) 6294164 (SEQ ID NO: 1493) Mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0870] 6294177 (SEQ ID NO: 1495) Mono-acetyl QS-21 (Isomer not determined) 6294186 (SEQ ID NO: 1496) Mono-benzoylated QS-21 (Isomer not determined) 6294187 (SEQ ID NO: 1497) Mono-acetyl QS-21 isomer 4, mono-malonyl QS-21 isomer

[0871] 6294218 (SEQ ID NO: 1498) Diacetyl QS-21 (Isomer not determined)

[0872] 6294223 (SEQ ID NO: 1499) Mono-acetyl QS-21 isomer 1, mono-acetyl QS-21 isomer 4, mono-malonyl QS-21

[0873] 6294226 (SEQ ID NO: 1500) Diacetyl QS-21 (Isomer not determined)

[0874] 6294234 (SEQ ID NO: 1501) Mono-acetyl QS-21, mono-malonyl QS-21 (Isomer not determined)

[0875] 6294243 (SEQ ID NO: 1502) Mono-acetyl QS-21 (Isomer not determined) 6294260 (SEQ ID NO: 1504) Mono-acetyl QS-21 (Isomer not determined) 6294276 (SEQ ID NO: 1505) Di-acetylated QS-21 (Isomer not determined) 6294294 (SEQ ID NO: 1507) Mono-benzoylated QS-21 (Isomer not determined) 6294337 (SEQ ID NO: 1508) Mono-acetyl QS-21 (Isomer not determined) 6294364 (SEQ ID NO: 1509) Mono-malonyl QS-21

[0876] 6294367 (SEQ ID NO: 1510) Mono-acetyl QS-21 (Isomer not determined) 6294371 (SEQ ID NO: 1511) Di-acetylated QS-21 (Isomer not determined) 6294387 (SEQ ID NO: 1512) Mono-acetyl QS-21, mono-benzoylated QS-21, monomalonyl QS-21 (Isomer not determined)

[0877]

[0878] 15.6 Glycoside hydrolase mediated production of QS-21 sugar hydrolysis products This example describes the identification of glycoside hydrolase variants that were functional in hydrolyzing QS-21. To identify glucoside hydrolase variants that were capable of hydrolyzing QS-21, a metagenomic library of approximately 1000 variants was generated based on the Cladosporium fulvum glycoside hydrolase (CfToml; SEQ ID NO: 1210) Clavibacter michiganensis glycoside hydrolase (Q7X3X6; SEQ ID NO: 1211), two Fusarium graminearum glycoside hydrolases (EYB27127; SEQ ID NO: 1212; 093976; SEQ ID NO: 1213), Streptomyces scabiei glycoside hydrolase (C9ZB10; SEQ ID NO: 1214), Septaria lycopersici glycoside hydrolase (Q99324; SEQ ID NO: 1215), Aspergillus oryzae glycoside hydrolase (D9J2M5; SEQ ID NO: 1216;

[0879] Q2WGL5; SEQ ID NO: 1217), Fusarium neocosmosporiellum glycoside hydrolase (Q76BW2; SEQ ID NO: 1218), and Eupenicillium brefeldianum glycoside hydrolase (Q2WGL4; SEQ ID NO: 1219) sequences.

[0880] The library was codon-optimized for expression in Escherichia coli and placed behind the T7 lac promoter and was transformed into the E. coli KRX host strain. The cell cultivation and lysis procedure were as described in Example 15.1.

[0881] To characterize glycoside hydrolase activity, the / / . coli lysate was screened using QS-21 as substrate. Then, 46 pL of reaction mix (11 mM Tris-HCl (pH 6.8), 1.1 mM magnesium chloride) was added to a new 384-well plate along with 4 pL lysate supernatant; the plate was incubated at room temperature for 18 hours. The reaction and extraction procedures were as described in Example 15.1.

[0882] The screen identified 18 glycoside hydrolases (see Table 20) where upon reaction with QS-21 in vitro, formed 2 new product peaks with a parent ion m / z 1517.7800 and 1385.7469 ([M-H] ).

[0883] The m / z value of one product is consistent with the calculated mass (m / z 1517.7800; see FIG. 8 A) for a deglycosylated QS-21 [M-H] ", and is believed to be QS-21 from the C3 branching sugars been removed.

[0884] The m / z value of the second product is consistent with the calculated mass (m / z 1385.7469; see FIG. 8B) for a different deglycosylated QS-21 [M-H]", and is believed to be QS-21 from the C3 branching sugars been removed and furthermore an additional pentose has been removed. Table 20. QS-21 Glycosyl hydrolases from library screen

[0885] QS-21 Glycosyl hydrolase QS-21 Derivatives Profile

[0886] 6311352 (SEQ ID NO: 1514) GH Deglycosylated QS-21 Product 1

[0887] 6311430 (SEQ ID NO: 1515) GH Deglycosylated QS-21 Product 2

[0888] 6311497 (SEQ ID NO: 1516) GH Deglycosylated QS-21 Product 1

[0889] 6311604 (SEQ ID NO: 1517) GH Deglycosylated QS-21 Product 2

[0890] 6311627 (SEQ ID NO: 1518) GH Deglycosylated QS-21 Product 1

[0891] 6311632 (SEQ ID NO: 1519) GH Deglycosylated QS-21 Product 1,

[0892] GH Deglycosylated QS-21 Product 2

[0893] 6311670 (SEQ ID NO: 1520) GH Deglycosylated QS-21 Product 2

[0894] 6311689 (SEQ ID NO: 1521) GH Deglycosylated QS-21 Product 2

[0895] 6311773 (SEQ ID NO: 1522) GH Deglycosylated QS-21 Product 1

[0896] 6311809 (SEQ ID NO: 1523) GH Deglycosylated QS-21 Product 1

[0897] 6311922 (SEQ ID NO: 1524) GH Deglycosylated QS-21 Product 1

[0898] 6311974 (SEQ ID NO: 1525) GH Deglycosylated QS-21 Product 1

[0899] 6312035 (SEQ ID NO: 1526) GH Deglycosylated QS-21 Product 1

[0900] 6312057 (SEQ ID NO: 1527) GH Deglycosylated QS-21 Product 2

[0901] 6312136 (SEQ ID NO: 1528) GH Deglycosylated QS-21 Product 2

[0902] 6312161 (SEQ ID NO: 1529) GH Deglycosylated QS-21 Product 2

[0903] 6312163 (SEQ ID NO: 1530) GH Deglycosylated QS-21 Product 1,

[0904] GH Deglycosylated QS-21 Product 2

[0905] 6312310 (SEQ ID NO: 1531) GH Deglycosylated QS-21 Product 2

[0906] 6311352 (SEQ ID NO: 1514) GH Deglycosylated QS-21 Product 1

[0907]

[0908] 15.7 Esterase mediated production of deacetylated QS-21

[0909] This example describes the identification of esterase variants that were functional in hydrolyzing QS-21. To identify esterase variants that were capable of hydrolyzing QS-21, a metagenomic library of approximately 500 variants was generated based on three Primus persica esterases (XP_007200443.1; SEQ IDNO: 1220; XP_007201832.1; SEQ IDNO: 1221; XP_007199906.1; SEQ ID NO: 1222), Solanum lycopersicum esterase (K7SGP9; SEQ ID NO: 1223), Spodoptera exigua esterase (AEJ38206.1; SEQ ID NO: 1224), two Plutella xylostella esterases (PxylCCE016a; SEQ ID NO: 1225; PxylCCE016c; SEQ ID NO: 1226), uncultured bacterial esterase (A0A2P1NSC8; SEQ ID NO: 1227), Klebsiella sp. esterase (Q52NW7; SEQ ID NO: 1228), Sphingobium faniae esterase (ACY01919.1; SEQ ID NO: 1229), Helicoverpa armigera esterase (ADF43460.1; SEQ ID NO: 1230), three Culex quinquefasciatus esterases (B0XFI1; SEQ ID NO: 1231; B0XFI2; SEQ ID NO: 1232; B0XFI3; SEQ ID NO: 1233), four Liposcelis bostrychophila esterases (ACI16653.1; SEQ ID NO: 1234; ACI16654.1; SEQ ID NO: 1235; ANG60749.1; SEQ ID NO: 1236; ANG60750.1; SEQ ID NO: 1237), and Sphingobium faniae esterase (D0VUS3; SEQ ID NO: 1238) sequences.

[0910] The library was codon-optimized for expression in Escherichia coli and placed behind the T7 lac promoter and was transformed into the E. coli KRX host strain. The cell cultivation and lysis procedure were as described in Example 15.1.

[0911] To characterize esterase activity, the A. coli lysate was screened using QS-21 as substrate. Then, 46 pL of reaction mix (11 mM Tris-HCl (pH 6.8), 1.1 mM magnesium chloride) was added to a new 384-well plate along with 4 pL lysate supernatant; the plate was incubated at room temperature for 18 hours. The reaction and extraction procedures were as described in Example 15.1.

[0912] The screen identified 27 esterases (see Table 21) where upon reaction with QS-21 in vitro, formed a new product peaks with a parent ion m / z 1511.6542 ([M-H]").

[0913] This m / z value of the product is consistent with the calculated mass for hydrolyzed QS-21 [M-H] ", corresponding to the removal of a single acyl chain group to QS-21.

[0914] Table 21. QS-21 Esterases from library screen

[0915] QS-21 Esterase QS-21 Derivatives Profile

[0916] 6380822 (SEQ ID NO: 1532) Esterase hydrolyzed QS-21 product 1

[0917] 6380838 (SEQ ID NO: 1533) Esterase hydrolyzed QS-21 product 1

[0918] 6380940 (SEQ ID NO: 1534) Esterase hydrolyzed QS-21 product 1

[0919]

[0920] 6380957 (SEQ ID NO: 1535) Esterase hydrolyzed QS-21 product 1 6380964 (SEQ ID NO: 1536) Esterase hydrolyzed QS-21 product 1 6380968 (SEQ ID NO: 1537) Esterase hydrolyzed QS-21 product 1 6380989 (SEQ ID NO: 1538) Esterase hydrolyzed QS-21 product 1 6381006 (SEQ ID NO: 1539) Esterase hydrolyzed QS-21 product 1 6381008 (SEQ ID NO: 1540) Esterase hydrolyzed QS-21 product 1 6381050 (SEQ ID NO: 1541) Esterase hydrolyzed QS-21 product 1 6381059 (SEQ ID NO: 1542) Esterase hydrolyzed QS-21 product 1 6381072 (SEQ ID NO: 1543) Esterase hydrolyzed QS-21 product 1 6381075 (SEQ ID NO: 1544) Esterase hydrolyzed QS-21 product 1 6381080 (SEQ ID NO: 1545) Esterase hydrolyzed QS-21 product 1 6381081 (SEQ ID NO: 1546) Esterase hydrolyzed QS-21 product 1 6381157 (SEQ ID NO: 1547) Esterase hydrolyzed QS-21 product 1 6381160 (SEQ ID NO: 1548) Esterase hydrolyzed QS-21 product 1 6381172 (SEQ ID NO: 1549) Esterase hydrolyzed QS-21 product 1 6381178 (SEQ ID NO: 1550) Esterase hydrolyzed QS-21 product 1 6381196 (SEQ ID NO: 1551) Esterase hydrolyzed QS-21 product 1 6381198 (SEQ ID NO: 1552) Esterase hydrolyzed QS-21 product 1 6381235 (SEQ ID NO: 1553) Esterase hydrolyzed QS-21 product 1 6381241 (SEQ ID NO: 1554) Esterase hydrolyzed QS-21 product 1 6381248 (SEQ ID NO: 1555) Esterase hydrolyzed QS-21 product 1 6381250 (SEQ ID NO: 1556) Esterase hydrolyzed QS-21 product 1 6381264 (SEQ ID NO: 1557) Esterase hydrolyzed QS-21 product 1 6381294 (SEQ ID NO: 1558) Esterase hydrolyzed QS-21 product 1

[0921]

[0922] EQUIVALENTS

[0923] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described in this application. Such equivalents are intended to be encompassed by the following claims.

Claims

CLAIMSWhat is claimed is:

1. A host cell comprising one or more of:a) a heterologous polynucleotide encoding a beta-amyrin synthase (BAS) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 59-314;b) a heterologous polynucleotide encoding a cytochrome P450 C16C28 oxidase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315-318;c) a heterologous polynucleotide encoding a cytochrome P450 C28 oxidase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315 and 318-320;d) a heterologous polynucleotide encoding a cytochrome P450 C23 oxidase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 321-335;e) a heterologous polynucleotide encoding a cytochrome P450 reductase (CPR) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 336-377;f) a heterologous polynucleotide encoding a membrane steroid binding protein (MSBP) that comprises an amino acid sequence having at least 70% identity, at least 75%identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ IDNOs: 378-416;g) a heterologous polynucleotide encoding a UDP-GlcA transferase (GlcAT) that comprises an amino acid sequence having at least 70% identity, 80% identity, 90% identity, 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 417-441;h) a heterologous polynucleotide encoding a UDP -galactose transferase (GalT) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 442-537;i) a heterologous polynucleotide encoding a UDP -xylose transferase (XylT) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 484, 485, 491, 494, 502, 504, 517, and 538-572;j) a heterologous polynucleotide encoding a UDP-glucose dehydrogenase (UGD) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 46-56, and 712-1142;k) a heterologous polynucleotide encoding a UDP-xylose synthase (UXS) that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 573-711;l) a heterologous polynucleotide encoding a phosphotransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1239-1257;m) a heterologous polynucleotide encoding a UDP -glycosyltransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1258-1283;n) a heterologous polynucleotide encoding a methyltransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1284-1289;o) a heterologous polynucleotide encoding a sulfotransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1290-1307;p) a heterologous polynucleotide encoding an acyltransferase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1308-1454;q) a heterologous polynucleotide encoding a glycosyl hydrolase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1455-1472; and / or r) a heterologous polynucleotide encoding an esterase that comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1473-1499.

2. The host cell of claim 1, wherein the host cell is capable of producing one or more of: 2,3-oxidosqualene; beta-amyrin; oleanolic acid; echinocystic acid; hederagenin; gypsogenin; gypsogenic acid; 16a-OH-hederagenin; quillaic acid (QA); 16a-OH-gypsogenic acid; QA-C3-GlcA; QA-C3-GlcA-Gal; and QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin).

3. The host cell of claim 1 or claim 2, wherein the host cell is a yeast cell, a plant cell, a bacterial cell, or a filamentous fungi cell.

4. A method comprising culturing the host cell of any one of claims 1-3 under conditions such that it expresses gene(s) encoded by the heterologous polynucleotide(s).

5. A host cell comprising one or more heterologous polynucleotides collectively encoding:a) a beta-amyrin synthase (BAS);b) a cytochrome P450 C 16 oxidase;c) a cytochrome P450 C28 oxidase;d) a cytochrome P450 C23 oxidase;e) a cytochrome P450 reductase (CPR); andf) a membrane steroid binding protein (MSBP);optionally wherein the host cell is capable of producing 2,3-oxidosqualene.

6. The host cell of claim 5, wherein:a) the beta-amyrin synthase (BAS) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 59-314;b) the cytochrome P450 C16C28 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315-318;c) the cytochrome P450 C23 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 321-335;d) the cytochrome P450 reductase (CPR) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 336-377;e) the membrane steroid binding protein (MSBP) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 378-416; orf) any combination thereof.

7. The host cell of claim 5 or claim 6, wherein:a) the beta-amyrin synthase (BAS) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 59-314;b) the cytochrome P450 C16C28 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 315-318;c) the cytochrome P450 C23 oxidase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 321-335;d) the cytochrome P450 reductase (CPR) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 336-377; ande) the membrane steroid binding protein (MSBP) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 378-416.

8. The host cell of any one of claims 5-7, wherein the host cell is a yeast cell, a plant cell, a bacterial cell, or a filamentous fungi cell.

9. The host cell of any one of claims 5-8, wherein the host cell comprises one or more genetic modifications such that the host cell has increased mevalonate flux relative to a host cell that lacks the one or more genetic modifications.

10. The host cell of claim 9, wherein the one or more genetic modifications comprise a heterologous polynucleotide(s) encoding one or more genes associated with mevalonate flux.

11. The host cell of claim 9 or claim 10, wherein the genes associated with mevalonate flux comprise ERG10, ERG13, HMG1, ERG12, ERG8, ERG19, IDI1, and / or ERG20.

12. The host cell of any one of claims 5-11, wherein the host cell overexpresses ERG10, ERG13, HMG1, ERG12, ERG8, ERG19, IDI1, and ERG20.

13. A method of producing quillaic acid (QA), the method comprising culturing the host cell of any one of claims 5-12 under conditions such that the host cell expresses the beta-amyrin synthase (BAS), the cytochrome P450 C16 oxidase, the cytochrome P450 C28 oxidase, the cytochrome P450 C23 oxidase, the cytochrome P450 reductase (CPR), and the membrane steroid binding protein (MSBP), wherein the host cell is capable of producing 2,3-oxidosqualene.

14. A host cell comprising a heterologous polynucleotide encoding a UDP-glucose dehydrogenase; optionally wherein the host cell is capable of producing UDP-glucose (UDP-Glu).

15. The host cell of claim 14, wherein the UDP-glucose dehydrogenase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 46-56, and 712-1142.

16. The host cell of claim 14 or claim 15, wherein the host cell is a yeast cell, a plant cell, a bacterial cell, or a filamentous fungi cell.

17. A method comprising culturing the host cell of any one of claims 14-16.

18. A method of producing UDP-glucuronic acid (UDP-GlcA), the method comprising culturing the host cell of any one of claims 14-16 under conditions such that the host cell expresses the UDP -glucose dehydrogenase, wherein the host cell is capable of producing UDP-glucose (UDP-Glu).

19. A host cell comprising a heterologous polynucleotide encoding a UDP-xylose synthase; optionally wherein the host cell is capable of producing UDP-glucuronic acid (UDP-GlcA).

20. The host cell of claim 19, wherein the UDP-xylose synthase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 573-711.

21. The host cell of claim 19 or claim 20, wherein the host cell is a yeast cell, a plant cell, a bacterial cell, or a filamentous fungi cell.

22. A method comprising culturing the host cell of any one of claims 19-21.

23. A method of producing UDP-xylose (UDP-Xyl) in a host cell, the method comprising culturing the host cell of any one of claims 19-21 under conditions such that the host cell expresses the UDP-xylose synthase, wherein the host cell is capable of producing UDP-glucuronic acid (UDP-GlcA).

24. A host cell comprising one or more heterologous polynucleotides collectively encoding:a) a UDP-glcA transferase (GlcAT);b) a UDP -galactose transferase (GalT); andc) a UDP-xylose transferase (XylT);optionally wherein the host cell is capable of producing quillaic acid (QA), UDP-glucuronic acid (UDP-GlcA), UDP -galactose (UDP-Gal), and / or UDP-xylose (UDP-Xyl).

25. The host cell of claim 24, wherein:a) the UDP-glcA transferase (GlcAT) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 417-441;b) the UDP -galactose transferase (GalT) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 442-537;c) the UDP -xylose transferase (XylT) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of any one of SEQ ID NOs: 484, 485, 491, 494, 502, 504, 517, and 538-572; ord) any combination thereof.

26. The host cell of claim 24 or claim 25, wherein:a) the UDP-glcA transferase (GlcAT) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 417-441;b) the UDP -galactose transferase (GalT) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NO: 442-537; andc) the UDP -xylose transferase (XylT) comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 484, 485, 491, 494, 502, 504, 517, and 538-572.

27. The host cell of any one of claims 24-26, wherein the host cell is a yeast cell, a plant cell, a bacterial cell, or a filamentous fungi cell.

28. A method comprising culturing the host cell of any one of claims 24-27.

29. A method of producing QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin), the method comprising culturing the host cell of any one of claims 24-27 under conditions such the host cell expresses the UDP-glcA transferase (GlcAT), the UDP-galactose transferase (GalT), and the UDP -xylose transferase (XylT), wherein the host cell is capable of producing quillaic acid (QA), UDP-glucuronic acid (UDP-GlcA), UDP-galactose (UDP-Gal), and UDP -xylose (UDP-Xyl).

30. A host cell comprising one or more of:a) a heterologous polynucleotide encoding a phosphotransferase;b) a heterologous polynucleotide encoding a UDP-glycosyltransferase;c) a heterologous polynucleotide encoding a methyltransferase;d) a heterologous polynucleotide encoding a sulfotransferase;e) a heterologous polynucleotide encoding an acyltransferase;f) a heterologous polynucleotide encoding a glycosyl hydrolase;g) a heterologous polynucleotide encoding an esterase; optionally wherein the host cell is capable of producing QS-21.

31. The host cell of claim 30, wherein:a) the phosphotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1239-1257;b) the UDP-glycosyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1258-1283;c) the methyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1284-1289.d) the sulfotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1290-1307;e) the acyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1308-1454;f) the glycosyl hydrolase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1455-1472; and / org) the esterase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1473-1499.

32. The host cell of claim 30 or claim 31, wherein the host cell is a yeast cell, a plant cell, a bacterial cell, or a filamentous fungi cell.

33. A method comprising culturing the host cell of any one of claims 30-32.

34. A method of producing a QS-21 derivative, the method comprising culturing the host cell of any one of claims 30-32 under conditions such the host cell expresses the sulfotransferase, the phosphotransferase, the UDP-glycosyltransferase, the acyltransferase, the methyltransferase, the glycosyl hydrolase, the esterase, the carbamoyl transferase, the 2-oxoglutarate-dependent dioxygenase, and / or the 4-keto-reductase, wherein the host cell is capable of producing QS-21.

35. A method of producing a QS-21 derivative, the method comprising contacting QS-21 with a sulfotransferase, a phosphotransferase, a UDP-glycosyltransferase, an acyltransferase, a methyltransferase, a glycosyl hydrolase, and / or an esterase.

36. The method of claim 35, wherein:a) the phosphotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1239-1257;b) the UDP -glycosyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1258-1283;c) the methyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1284-1289.d) the sulfotransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1290-1307;e) the acyltransferase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1308-1454;f) the glycosyl hydrolase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1455-1472; and / org) the esterase comprises an amino acid sequence having at least 70% identity, at least 75% identity, at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, at least 98% identity, at least 99% identity, or 100% identity with the sequence of any one of SEQ ID NOs: 1473-1499.

37. The method of claim 35 or claim 36, wherein the sulfotransferase, the phosphotransferase, the UDP -glycosyltransferase, the acyltransferase, the methyltransferase, the glycosyl hydrolase, and / or the esterase is a purified enzyme.

38. The method of claim 35 or claim 36, wherein the method comprises contacting QS-21 with a cell lysate, wherein the cell lysate comprises the sulfotransferase, the phosphotransferase, the UDP -glycosyltransferase, the acyltransferase, the methyltransferase, the glycosyl hydrolase, and / or the esterase.

39. The method of claim 38, wherein the cell lysate is produced by lysing a population of host cells, wherein the host cells comprise one or more of:a) a heterologous polynucleotide encoding the phosphotransferase;b) a heterologous polynucleotide encoding the UDP-glycosyltransferase;c) a heterologous polynucleotide encoding the methyltransferase;d) a heterologous polynucleotide encoding the sulfotransferase;e) a heterologous polynucleotide encoding the acyltransferase;f) a heterologous polynucleotide encoding the glycosyl hydrolase; andg) a heterologous polynucleotide encoding the esterase.

40. An adjuvant comprising a triterpene glycoside, optionally wherein the triterpene glycoside is quillaic acid (QA), 16a-OH-gypsogenic acid, QA-C3-GlcA; QA-C3-GlcA-Gal, QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin), QA-TriX-F, QA-TriX-FR, QA-TriX-FXR, QA-TriX-FRXX (DS1), or QA-TriX-FRXX-RG-Ac (QS-7), wherein the triterpene glycoside is produced by the host cell of any one of claims 1-3, 5-12, 14-16, 19-21, 24-27, and 30-32.

41. A therapeutic composition comprising a therapeutically relevant agent and a triterpene glycoside, optionally wherein the triterpene glycoside is quillaic acid (QA), 16a-OH-gypsogenic acid, QA-C3-GlcA; QA-C3-GlcA-Gal, QA-C3-GlcA-Gal-Xyl (QA-TriX, prosapogenin), QA-TriX-F, QA-TriX-FR, QA-TriX-FXR, QA-TriX-FRXX (DS1), or QA-TriX-FRXX-RG-Ac (QS-7), wherein the triterpene glycoside is produced by the host cell of any one of claims 1-3, 5-12, 14-16, 19-21, 24-27, and 30-32.

42. A composition comprising phosphorylated QS-21, glycosylated QS-21, methylated QS-21, sulfated QS-21, acylated QS-21, esterase hydrolyzed QS-21, and / or glycoside hydrolase hydrolyzed QS-21.

43. The composition of claim 42, wherein the phosphorylated QS-21 is mono-phosphorylated QS-21, the glycosylated QS-21 is mono-glucosylated QS-21, the methylated QS-21 is mono-methylated or di-methylated QS-21, the sulfated QS-21 is mono-sulfated QS-21, or the acylated QS-21 is mono-acetylated QS-21, de-acetylated QS-21, mono-malonylated QS-21, or mono-benzoylated QS-21.

Citation Information

Patent Citations

  • Prunella vulgaris 2, 3-oxysqualene cyclase OSC protein as well as coding gene and application thereof

    CN118109447A

  • Metabolic engineering

    WO2019122259A1

  • Saponin production in yeast

    WO2023122801A2