Biosynthetic enzyme

Characterizing genes for triterpenoid saponin biosynthesis in Saponaria officinalis enables the production of QA and its derivatives in alternative hosts, addressing the lack of understanding in biosynthetic pathways and facilitating large-scale production for various applications.

JP2025521790APending Publication Date: 2025-07-10PLANT BIOSCIENCE LIMITED
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024577115
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-29
Filing Date
2023-06-27
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The biosynthetic pathway of complex triterpenoid saponins, such as saponariosides A and B, is not well understood, limiting the ability to metabolically engineer alternative host systems for large-scale production.

Method used

Identification and characterization of genes involved in the biosynthesis of triterpenoid saponins in Saponaria officinalis, including enzymes for the conversion of 2,3-oxidosqualene to quillaic acid and its glycosylation, enabling the production of QA and its glycosylated products through a series of enzymatic steps.

Benefits of technology

Facilitates the production of triterpenoids like QA and saponarioside B in heterologous hosts, potentially allowing for large-scale synthesis and utilization in cosmetics, dietary supplements, and bioremediation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521790000007
    Figure 2025521790000007
  • Figure 2025521790000008
    Figure 2025521790000008
  • Figure 2025521790000009
    Figure 2025521790000009
Patent Text Reader

Abstract

The present invention relates to a method for producing triterpenoids using one or more of the following polypeptides: (i) Saponaria officinalis synthase (SobAS), (ii) S. officinalis C28 oxidase (SoC28), (iii) S. officinalis C28C16 oxidase (SoC28C16), (iv) S. officinalis C23 oxidase (SoC23), (v) S. officinalis QA 3-O-glucuronosyltransferase (“SoCSL”), (vi) S. officinalis QA-GlcA galactosyltransferase (“SoC3Gal”), (vii) S. officinalis QA-GlcA-Gal xylosyltransferase (“SoC3Xyl”), (viii) S. officinalis QA-Tri fucosyltransferase (“SoC28Fu”), (ix) S. officinalis QA-TriF rhamnosyltransferase (“SoC28Rha”), (x) S. officinalis QA-TriFR xylosyltransferase (“SoC28Xyl1”), (xi) S. officinalis QA-TriFRX xylosyltransferase (“SoC28Xyl2”), (xii) S. officinalis QA-TriFRXX quinovosyltransferase (“SoGH1”), and (xiii) S. officinalis QA-TriF(Q)RXX acetyltransferase (“SoBAHD1”). Methods, host cells, isolated polypeptides, nucleic acids, and plants are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Field The present invention relates to the biosynthesis of complex triterpenoid saponins and intermediates such as quillaic acid, and genes and polypeptides involved in this biosynthesis.

Background Art

[0002] Background Saponaria officinalis (Caryophyllaceae), commonly known as soapwort, is a perennial flowering plant native to Europe and Asia and has traditionally been used as a raw material for soap [1]. The well-known cleaning properties of soapwort are due to the high content of amphiphilic saponins contained in the plant extract. Ancient Greeks, Romans, and Egyptians used soapwort extracts for cleaning and laundering clothes, and later, the first American settlers brought soapwort plants from Europe to North America for household use [2]. In addition to its cleaning action, soapwort extracts have been used in folk remedies to treat symptoms such as syphilis, gout, rheumatism, and jaundice [3]. Soapwort extracts are also used in the production of tahini halva, a traditional dessert in the Middle East, and thus play an important role in Middle Eastern culture [4]. Even today, soapwort extracts are used in cosmetics, dietary supplements, and phytopharmaceutical products [5]. Furthermore, the saponin layer of soapwort extracts is being studied for its potential use in bioremediation, food surfactants, antifungal activity, and immunotoxic activity [6-9].

[0003] S. officinalis is a rich source of saponins with various aglycone cores such as chirayic acid, diosgenin, and diosgeninic acid. The main saponins contained in the soapwort extract have been reported as saponariosides A and B (SpA, SpB) [1]. SpA and SpB have similar chemical structures. Both consist of the chirayic acid aglycone, a C-30 triterpenoid, and the C-3 position is decorated with a branched trisaccharide and the C28 position with a linear tetrasaccharide. The first sugar of the tetrasaccharide chain, β-D-fucose, is linked to β-D-quinovose with an acetyl group attached. The only chemical difference between SpA and SpB is that β-D-xylose is added to the quinovose moiety on the C28 sugar chain of SpA.

[0004] Interestingly, QS-21, a triterpenoid saponin discovered from Quillaja saponaria, is chemically very similar to SpA and SpB (Figure 1). QS-21 is a complex triterpenoid saponin synthesized by the Chilean tree Quillaja saponaria (order Fabales). Biochemically, QS-21 contains the quillaic acid skeleton of a C-30 triterpenoid. This scaffold is decorated with a branched trisaccharide at the C-3 position and a linear tetrasaccharide at the C-28 position. The terminal sugar of the tetrasaccharide is either β-D-apiose or β-D-xylose. Finally, the β-D-fucose sugar within the tetrasaccharide is also characterized by a C-18 acyl chain glycosylated with an arabinose sugar. QS-21 is a potent immunostimulant that can enhance the antibody response and boost specific T cell responses, offering the potential for an important adjuvant (Del Giudice et al. Seminars in Immunology, 2018. 39: p. 14-21; Marciani, D.J. Trends in Pharmacological Sciences, 2018. 39(6): p. 573-585). The AS01 adjuvant is a liposomal formulation of QS-21 and 3-O-desacyl-4'-monophosphoryl lipid A and is currently approved as part of GlaxoSmithKline's "Shingrix" vaccine against herpes zoster and "Mosquirix" against malaria (Del Giudice et al., supra).

[0005] Despite the promising commercial potential of saponariosides and their intermediates, nothing is known about their biosynthetic pathway. The biosynthesis of saponariosides can conceptually be divided into two stages: (i) the biosynthesis of the quillaic acid core and (ii) the decoration of quillaic acid (Figure 2). However, the order may be different in the actual plant, and further details are unknown. Many plant natural products are present in low amounts in plants, and chemical synthesis is often impossible due to their complex chemical structures. Knowledge of the biosynthetic pathway may enable metabolic engineering in alternative host systems and potentially lead to the large-scale production of the desired compounds. SUMMARY OF THE INVENTION

[0006] Summary The inventors have identified genes involved in the biosynthesis of complex triterpenoid saponins in Saponaria officinalis and characterized them. These include genes encoding enzymes involved in the biosynthesis of QA and glycosyltransferases involved in the glycosylation of QA. The expression of one or more of these genes would be useful for the production of QA and glycosylated products of QA.

[0007] The first aspect of the present invention provides a method for producing a triterpenoid comprising one or more of the following; (i) contacting 2,3-oxidosqualene (OS) with Saponaria officinalis synthase (SobAS) comprising an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 8 such that the OS is converted to β-amyrin; (ii) any of the following: (a) contacting β-amyrin with a SoC28 oxidase polypeptide comprising an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 2; oxidizing the C28 position of the β-amyrin to a carboxylic acid to produce oleanolic acid, and contacting the oleanolic acid with a SoC28C16 oxidase polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 4, and oxidizing the C16 alcohol of the oleanolic acid to produce echinocystic acid; or (b) contacting β-amyrin with a SoC28C16 oxidase polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 4, oxidizing the C28 carboxylic acid of the β-amyrin, and oxidizing the C16 position of the β-amyrin to an alcohol, thereby producing echinocystic acid; (iii) contacting echinocystic acid with a SoC23 oxidase polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 6 such that the C-23 position of the echinocystic acid is oxidized to an aldehyde, thereby producing quillaic acid (QA); (iv) contacting QA with a Saponaria officinalis QA 3-O-glucuronosyltransferase (「SoCSL」) polypeptide comprising an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 10 so that the QA is converted to QA-GlcA; (v) contacting QA-GlcA with a Saponaria officinalis QA-GlcA galactosyltransferase (「SoC3Gal」) polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 12 so that the QA-GlcA is converted to QA-GlcA-Gal; (vi) contacting QA-GlcA-Gal with a Saponaria officinalis QA-GlcA-Gal xylosyltransferase (「SoC3Xyl」) polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 14 so that the QA-GlcA-Gal is converted to QA-Tri (QA-GlcA-Gal-Xyl); (vii) contacting QA-Tri with a Saponaria officinalis QA-Tri fucosyltransferase (「SoC28Fu」) polypeptide comprising an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 16 so that the QA-Tri is converted to QA-TriF; (viii) contacting QA-TriF with a Saponaria officinalis QA-TriF rhamnosyltransferase (「SoC28Rha」) polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 18 so that the QA-TriF is converted to QA-TriFR; (ix) contacting QA-TriFR with a Saponaria officinalis QA-TriFR xylosyltransferase (「SoC28Xyl1」) polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 20 so that the QA-TriFR is converted to QA-TriFRX; and / or Contacting QA-TriFRX with a Saponaria officinalis QA-TriFRX xylosyltransferase ("SoC28Xyl2") polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 22 such that the QA-TriFRX is converted to QA-TriFRXX; Contacting QA-TriFRXX with a Saponaria officinalis QA-TriFRXX quinovosyltransferase ("SoGH1") polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 34 such that the QA-TriFRXX is converted to QA-TriF(Q)RXX; and / or Contacting QA-TriF(Q)RXX with a Saponaria officinalis QA-TriF(Q)RXX acetyltransferase ("SoBAHD1") polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 36 such that the QA-TriF(Q)RXX is converted to saponarioside B (SpB).

[0008] The method of the first aspect can include the following; Contacting β-amyrin with a Saponaria officinalis C28 oxidase (SoC28 oxidase) to oxidize the C28 position to a carboxylic acid to form oleanolic acid, wherein the amino acid sequence of the SoC28 oxidase has at least 80% sequence identity with SEQ ID NO: 2; Contacting oleanolic acid with a Saponaria officinalis C28C16 oxidase (SoC28C16 oxidase) to oxidize the C16 position of oleanolic acid to an alcohol to form echinocystic acid, wherein the amino acid sequence of the SoC28C16 oxidase has at least 50% sequence identity with SEQ ID NO: 4; and A step of contacting echinocystic acid with Saponaria officinalis C-23 oxidase (SoC23 oxidase) to oxidize the C16 position of echinocystic acid to an aldehyde to produce kirenolic acid (QA), wherein the amino acid sequence of SoC23 oxidase has at least 50% sequence identity with SEQ ID NO: 6.

[0009] The method of the first aspect can include the following; A step of contacting β-amyrin with Saponaria officinalis C28C16 oxidase (SoC28C16 oxidase) to oxidize the C28 position to a carboxylic acid and the C16 position to an alcohol to form echinocystic acid, wherein the amino acid sequence of SoC2816 oxidase has at least 50% sequence identity with SEQ ID NO: 4; and A step of contacting echinocystic acid with Saponaria officinalis C-23 oxidase (SoC23 oxidase) to oxidize the C16 position of echinocystic acid to an aldehyde to form kirenolic acid (QA), wherein the amino acid sequence of SoC23 oxidase has at least 50% sequence identity with SEQ ID NO: 6.

[0010] The method of the first aspect can include or further include the following; A step of contacting QA with Saponaria officinalis QA 3-O-glucuronosyltransferase ("SoCSL") to covalently bond D-glucuronic acid ("GlcA") to the 3-O position of kirenolic acid to form 3-O-{β-D-glucopyranosiduronic acid}-kirenolic acid ("QA-GlcA" or "QA-mono"), wherein the amino acid sequence of SoCSL has at least 60% sequence identity with SEQ ID NO: 10; Contacting QA-GlcA with Saponaria officinalis QA-GlcA galactosyltransferase ("SoC3Gal") and covalently binding D-galactose ("Gal") to QA-GlcA via a β-1→2 linkage to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-oleanolic acid ("QA-GlcA-Gal or QA-Di"), wherein the amino acid sequence of SoC3Ga has at least 50% sequence identity with SEQ ID NO: 12; and Contacting QA-GlcA-Gal with Saponaria officinalis QA-GlcA-Gal xylosyltransferase ("SoC3Xyl") and covalently binding D-xylose ("Xyl") to QA-GlcA-Gal via a 1,3 linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-oleanolic acid ("QA-Tri" or "QA-GlcA-Gal-Xyl"), wherein the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14.

[0011] The method of the first aspect can include or further include the following; Contacting 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-oleanolic acid (QA-Tri) with Saponaria officinalis QA-Tri fucosyltransferase ("SoC28Fu") and binding fucose to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF), wherein the amino acid sequence of SoC28Fu has at least 60% sequence identity with SEQ ID NO: 16; Contacting QA-TriF with Saponaria officinalis QA-TriF rhamnosyltransferase ("SoC28Rha") to covalently attach rhamnose to QA-TriF via a 1,2-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFR), where the amino acid sequence of SoC28Rha has at least 50% sequence identity with SEQ ID NO: 18; where the amino acid sequence of SoC28Rha has at least 50% sequence identity with SEQ ID NO: 18; Contacting QA-TriFR with Saponaria officinalis QA-TriFR xylosyltransferase ("SoC28Xyl1") to covalently attach xylose to QA-TriFR via a 1,4-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→4)-α-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRX), where the amino acid sequence of SoC28Xyl1 has at least 50% sequence identity with SEQ ID NO: 20; Contacting QA-TriFRX with Saponaria officinalis QA-TriFRX-xylosyltransferase ("SoC28Xyl2") to covalently attach xylose to QA-TriFRX via a 1,3-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRXX), where the amino acid sequence of SoC28Xyl2 has at least 50% sequence identity with SEQ ID NO: 22.

[0012] The method of the first aspect can include or further include the following; contacting QA-TriFRXX with Saponaria officinalis QA-TriFRXX quinovosyltransferase (“SoGH1”) to covalently attach quinovose to QA-TriFRXX via a 1,4 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)]-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX), wherein the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34; and / or contacting QA-TriF(Q)RXX with Saponaria officinalis QA-TriF(Q)RXX acetyltransferase (“SoBAHD1”) to covalently attach an acetyl group to QA-TriF(Q)RXX to form saponarioside B, wherein the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36.

[0013] The second aspect of the present invention provides a method for converting a host from a phenotype incapable of performing triterpenoid biosynthesis from β-amyrin to a phenotype capable of performing said triterpenoid biosynthesis, the method including the following; a step of expressing a heterologous nucleic acid in the host or one or more of its cells following a step of introducing the nucleic acid into the host or one of its ancestors, wherein the heterologous nucleic acid encodes one or more or all of the following polypeptides: (i) SoC28 oxidase (“SoC28 oxidase”) capable of oxidizing β-amyrin at the C28 position to a carboxylic acid; said SoC28 oxidase has at least 80% sequence identity with SEQ ID NO: 2; (ii) SoC28C16 oxidase (referred to as "SoC28C16 oxidase") that can oxidize β - amyrin to carboxylic acid at the C28 position and to alcohol at the C16 position; the C28C16 oxidase has at least 50% sequence identity with SEQ ID NO: 4; and (iii) SoC23 oxidase (referred to as "SoC23 oxidase") that can oxidize echinocystic acid to aldehyde at the C - 23 position, and the SoC23 oxidase has at least 50% sequence identity with SEQ ID NO: 6; (iv) Saponaria officinalis QA 3 - O glucuronosyltransferase (referred to as "SoCSL") that can bind D - glucuronic acid (referred to as "GlcA") to the 3 - O position of quillaic acid to form 3 - O - {β - D - glucopyranosiduronic acid} - quillaic acid (referred to as "QA - mono" or "QA - GlcA"); the SoCSL has at least 60% sequence identity with SEQ ID NO: 10; (v) Bind D - galactose (referred to as "Gal") to QA - GlcA via a β - 1→2 bond to form 3 - O - {[β - D - galactopyranosyl - (1→2)] - β - D - glucopyranosiduronic acid} - quillaic acid (referred to as "QA - Di" or "QA - GlcA - Gal"); here, the amino acid sequence of SoC3Gal has at least 50% sequence identity with SEQ ID NO: 12; (vi) Bind D - xylose (referred to as "xylose") to QA - GlcA - Gal via a 1,3 bond to form 1,3 - O - {β - D - xylopyranosyl - (1→3) - [β - D - galactopyranosyl - (1→2)] - β - D - glucopyranosiduronic acid} - quillaic acid (referred to as "QA - GlcA - Gal - Xyl" or "QA - Tri"); here, the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14; (vii) Attach fucose ("Fuc") to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosyl ester}-kirayaic acid (QA-TriF); wherein the amino acid sequence of SoC28Fu has at least 50% sequence identity with SEQ ID NO: 16; (viii) Attach rhamnose ("Rha") to QA-TriF via a 1,2 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFR); the amino acid sequence of said SoC28Rha has at least 50% sequence identity with SEQ ID NO: 18; (ix) Attach D-xylose ("Xyl") to QA-TriFR via a 1,4 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→4)-α-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFRX); wherein the amino acid sequence of C28Xyl1 has at least 50% sequence identity with SEQ ID NO: 20; (x) Attach D-xylose ("Xyl") to QA-TriFRX via a 1,3 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-l-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFRXX); wherein the amino acid sequence of C28Xyl2 has at least 50% sequence identity with SEQ ID NO: 22; (xi) QA-TrFRXX binds to QA-TrFRXX via a 1,4 linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-dD-quinovopyranosyl-(1→4)]-β-DD-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX); wherein the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34; and / or (xii) Saponaria officinalis QA-TriF(Q)RXX acetyltransferase (「SoBAHD1」) for binding an acetyl group to QA-TriF(Q)RXX to form saponarioside B; wherein the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36.

[0014] The heterologous nucleic acid in the method of the second aspect may encode the following polypeptides; (i) SoC28 oxidase capable of oxidizing β-amyrin to carboxylic acid at the C28 position to form oleanolic acid; said SoC28 oxidase has at least 80% sequence identity with SEQ ID NO: 2; (ii) SoC28C16 oxidase (「C28C16 oxidase」) capable of oxidizing β-amyrin to carboxylic acid at the C28 position and to alcohol at the C16 position to form echinocystic acid; said SoC28C16 oxidase has at least 50% sequence identity with SEQ ID NO: 4; and (iii) SoC23 oxidase capable of oxidizing echinocystic acid to aldehyde at the C-23 position to form oleanolic acid (QA), said SoC23 oxidase has at least 50% sequence identity with SEQ ID NO: 6, SoC23 oxidase.

[0015] The heterologous nucleic acid in the method of the second aspect may further encode the following polypeptides; (iv) Saponaria officinalis QA 3-O glucuronosyltransferase (「SoCSL」) for forming 3-O-{β-D-glucopyranosiduronic acid}-oleanolic acid (「QA-GlcA」) by binding D-glucuronic acid (「GlcA」) to the 3-O position of oleanolic acid; here, the amino acid sequence of SoCSL has at least 60% sequence identity with SEQ ID NO: 10; (v) Binding D-galactose (「Gal」) to QA-GlcA via a β-1→2 bond to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-oleanolic acid (「QA-GlcA-Gal」); here, the amino acid sequence of SoC3Gal has at least 50% sequence identity with SEQ ID NO: 12; and (vi) Binding D-xylose (「xylose」) to QA-GlcA-Gal via a 1,3 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-oleanolic acid (「QA-GlcA-Gal-Xyl」 or 「QA-Tri」); here, the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14.

[0016] The heterologous nucleic acid in the method of the second aspect may further encode the following polypeptide; (vii) SoC28Fu for binding fucose (「Fuc」) to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF); said SoC28Fu has at least 60% sequence identity with SEQ ID NO: 16; (viii) Attach rhamnose ("Rha") to QA-TriF via a 1,2 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{α-L-rhamnopyranosyl-(1→2)-Β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFR); wherein the SoC28Rha has at least 50% sequence identity with SEQ ID NO: 18; (ix) Attach D-xylose ("Xyl") to QA-TriFR via a 1,4 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→4)-α-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFRX); wherein the amino acid sequence of SoC28Xyl1 has at least 50% sequence identity with SEQ ID NO: 20; and (x) Attach D-xylose ("Xyl") to QA-TriFRX via a 1,3 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-l-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFRXX); wherein the amino acid sequence of SoC28Xyl2 has at least 80% sequence identity with SEQ ID NO: 22.

[0017] The heterologous nucleic acid in the method of the second aspect may further encode the following polypeptide; (xi) Quinovose (Q) is coupled to QA-TrFRXX via a 1,4 linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX) using Saponaria officinalis QA-TriFRXX xylosyltransferase (「SoGH1」); here, the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34), and / or (xii) Saponaria officinalis QA-TriF(Q)RXX acetyltransferase (「SoBAHD1」) for coupling an acetyl group to QA-TriF(Q)RXX to form saponarioside B; here, the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36.

[0018] A third aspect of the present invention provides a host cell comprising, or transformed with, a heterologous nucleic acid comprising each of a plurality of nucleotide sequences encoding a polypeptide having triterpenoid biosynthetic activity, wherein the plurality of nucleotide sequences encode one or more of the following polypeptides: (i) SoC28 oxidase (「SoC28 oxidase」) capable of oxidizing β-amyrin to a carboxylic acid at the C28 position; said SoC28 oxidase has at least 80% sequence identity with SEQ ID NO: 2; (ii) SoC28C16 oxidase (「C28C16 oxidase」) capable of oxidizing β-amyrin to a carboxylic acid at the C28 position and to an alcohol at the C16 position; said C28C16 oxidase has at least 50% sequence identity with SEQ ID NO: 4; and (iii) SoC23 oxidase (referred to as "SoC23 oxidase") which can oxidize echinocystic acid to an aldehyde at the C-23 position, wherein the SoC23 oxidase has at least 50% sequence identity with SEQ ID NO: 6; (iv) Saponaria officinalis QA 3-O glucuronosyltransferase (referred to as "SoCSL") for binding D-glucuronic acid ("GlcA") to the 3-O position of kiragic acid to form 3-O-{β-D-glucopyranosiduronic acid}-kiragic acid ("QA-GlcA"); the SoQA-GlcT has at least 60% sequence identity with SEQ ID NO: 10; (v) Binding D-galactose ("Gal") to QA-GlcA via a β-1→2 bond to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-kiragic acid ("QA-GlcA-Gal"); here, the amino acid sequence of SoC3Gal has at least 50% sequence identity with SEQ ID NO: 12; and (vi) Binding D-xylose ("xylose") to QA-GlcA-Gal via a 1,3 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-kiragic acid ("QA-GlcA-Gal-Xyl" or "QA-Tri"); here, the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14; (vii) SoC28Fu for binding fucose ("Fuc") to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-fucopyranosyl ester}-kiragic acid (QA-TriF); the SoC28Fu has at least 50% sequence identity with SEQ ID NO: 16; (viii) Attach rhamnose ("Rha") to QA-TriF via a 1,2 linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFR); the SoC28Rha has at least 50% sequence identity with SEQ ID NO: 18; (ix) Attach D-xylose ("Xyl") to QA-TriFR via a 1,4 linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFRX); wherein the amino acid sequence of SoC28Xyl1 has at least 50% sequence identity with SEQ ID NO: 20; (x) Attach D-xylose ("Xyl") to QA-TriFRX via a 1,3 linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFRXX); wherein the amino acid sequence of SoC28Xyl2 has at least 50% sequence identity with SEQ ID NO: 22; (xi)Saponaria officinalis QA-TriFRXX xylosyltransferase (「SoGH1」) for conjugating quinovose (Q) to QA-TrFRXX via a 1,4-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX); wherein the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34; and / or (xii)Saponaria officinalis QA-TriF(Q)RXX acetyltransferase (「SoBAHD1」) for conjugating an acetyl group to QA-TriF(Q)RXX to form saponarioside B; wherein the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36.

[0019] The plurality of nucleotide sequences in the host cell of the third aspect may encode the following polypeptides; (i) SoC28 oxidase capable of oxidizing β-amyrin to carboxylic acid at the C28 position to form oleanolic acid; said SoC28 oxidase having at least 80% sequence identity with SEQ ID NO: 2; (ii) SoC28C16 oxidase (「C28C16 oxidase」) capable of oxidizing β-amyrin to carboxylic acid and / or alcohol at the C16 position at the C28 position; and (iii) SoC23 oxidase capable of oxidizing echinocystic acid to aldehyde at the C-23 position to form oleanolic acid (QA), said SoC23 oxidase having at least 50% sequence identity with SEQ ID NO: 6.

[0020] The plurality of nucleotide sequences in the host cell of the third aspect may further encode the following polypeptides; (iv) Saponaria officinalis QA 3-O glucuronosyltransferase (「SoCSL」) for forming 3-O-{β-D-glucopyranosiduronic acid}-oleanolic acid (「QA-GlcA」) by binding D-glucuronic acid (「GlcA」) to the 3-O position of oleanolic acid; said SoQA-GlcT has at least 60% sequence identity with SEQ ID NO: 10; (v) Binding D-galactose (「Gal」) to QA-GlcA via a β-1→2 bond to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-oleanolic acid (「QA-GlcA-Gal」); wherein the amino acid sequence of QA-GlcA-Gal has at least 50% sequence identity with SEQ ID NO: 12; and (vi) Binding D-xylose (「xylose」) to QA-GlcA-Gal via a 1,3 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-oleanolic acid (「QA-GlcA-Gal-Xyl」QA-Tri); wherein the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14.

[0021] The plurality of nucleotide sequences in the host cell of the third aspect may further encode the following polypeptides; (vii) SoC28Fu for binding fucose (「Fuc」) to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF); said SoC28Fu has at least 60% sequence identity with SEQ ID NO: 16; (viii) Link rhamnose ("Rha") to QA-TriF via a 1,2 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFR); wherein said SoC28Rha has at least 50% sequence identity with SEQ ID NO: 18; (ix) Link D-xylose ("Xyl") to QA-TriFR via a 1,4 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRX); wherein the amino acid sequence of SoC28Xyl1 has at least 50% sequence identity with SEQ ID NO: 20; and (x) Link D-xylose ("Xyl") to QA-TriFRX via a 1,3 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRXX); wherein the amino acid sequence of SoC28Xyl2 has at least 80% sequence identity with SEQ ID NO: 22.

[0022] The plurality of nucleotide sequences in the host cell of the third aspect may further encode the following polypeptides; (xi) Quinovose (Q) is attached to QA-TrFRXX via a 1,4-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX) using Saponaria officinalis QA-TriFRXX xylosyltransferase (「SoGH1」); wherein the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34, and / or (xii) Saponaria officinalis QA-TriF(Q)RXX acetyltransferase (「SoBAHD1」) for attaching an acetyl group to QA-TriF(Q)RXX to form saponarioside B, wherein the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36.

[0023] A fourth aspect of the present invention provides a method for producing a host cell comprising transforming or transfecting a host cell with a heterologous nucleic acid comprising the plurality of nucleotide sequences described in the second and third aspects.

[0024] A fifth aspect provides a process for producing a transgenic plant: (a) performing the method of the fourth aspect, wherein the host cell is a plant cell, and (b) regenerating a plant from the transformed plant cell.

[0025] A sixth aspect provides a transgenic plant obtainable by the method of the fifth aspect, or a clone of said transgenic plant, or a transgenic plant that is a self-propagating or hybrid offspring or other offspring of the transgenic plant, Here, the expression of the heterologous nucleic acid confers an increased ability to carry out triterpenoid biosynthesis compared to the wild-type plant corresponding to the transgenic plant.

[0026] The seventh aspect provides a method for producing triterpenoids in a heterologous host. This method includes culturing the host cell described in the third aspect and purifying the triterpenoid therefrom.

[0027] The eighth aspect provides a method for producing triterpenoids in a heterologous host, which includes growing the plant of the sixth aspect, then harvesting it, and purifying the triterpenoid therefrom.

[0028] The triterpenoids of the seventh and eighth aspects can be QA or glycosylated QA, such as QA-Tri, QA-TriFRXX or QA-F(Q-Ac)RXX, or an intermediate or derivative thereof.

[0029] SobAS, SoC28 oxidase, SoC23 oxidase, SoC28C16 oxidase, SoCSL, SoC3Gal, SoC3Xyl, SoC28Fu, SoC28Rha, SoC28Xyl1, SoC28Xyl2, SoGH1 and SoBAHD1 of the first to eighth aspects can be obtained from or derived from Saponaria officinalis.

[0030] Other aspects and embodiments of the present invention will be described in more detail below.

Brief Description of the Drawings

[0031]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

[0032] DETAILED DESCRIPTION The present invention relates to the production of triterpenoids such as saponarioside and its intermediates, using biosynthetic enzymes encoded by genes newly characterized or identified from the soapwort plant (Saponaria officinalis) and its variants. These enzymes include β-amyrin synthase (bAS; SobAS; SEQ ID NO: 8), SoC28 oxidase (SoC28; SEQ ID NO: 2), SoC23 oxidase (SoC23; SEQ ID NO: 4), C28C16 oxidase (SoC28C16; SEQ ID NO: 6), QA 3-O-glucuronosyltransferase (SoQA-GlcAT; SoCSL; SEQ ID NO: 10), QA-GlcA galactosyltransferase (SoC3Gal; SEQ ID NO: 12), QA-GlcA-Gal xylosyltransferase (SoQA-RXylT; SoC3Xyl; SEQ ID NO: 14), QA-Tri fucosyltransferase (QATriFuT; SoC28F; SEQ ID NO: 16), QA-TriF rhamnosyltransferase (QA-TriFR; SoC28Rha; SEQ ID NO: 18), QA-TriFR xylosyltransferase (SoQA-TriFRXylT; SoC28Xyl1; SEQ ID NO: 20), QA-TriFRX xylosyltransferase (SoQA-TriFRXXylT; SoC28Xyl2; SEQ ID NO: 22), QA-TriFRXX quinovosyltransferase (SoGH1; SEQ ID NO: 34) and / or QA-TriF(Q)RXX acetyltransferase (SoBAHD1; SEQ ID NO: 36).

[0033] Each gene, polypeptide sequence and nucleotide sequence described herein is optionally obtained from or derived from S. officinalis.

[0034] The gene polypeptide sequences and nucleotide sequences described herein may be useful for the production of cyclic triterpenes such as β-amyrin, oleanolic acid, echinocystic acid, quillaic acid (“QA”), and glycosylated forms of QA such as saponarioside, QS-7, QS-21, and analogs and intermediates of these glycosylated forms of QA.

[0035] In some embodiments, one, two, three, four, or more of the genes described herein may be useful in the production of quillaic acid (QA). QA is a derivative of β-amyrin, a simple triterpene, which is synthesized by cyclization of the universal linear precursor 2,3-oxidosqualene (OS) by oxidosqualene cyclase (OSC). The β-amyrin skeleton is further oxidized at the C16, C-23, and C28 positions with alcohol, aldehyde, and carboxylic acid, respectively, to form QA. Although the proposed linear biosynthetic pathway is shown in FIG. 15, the three oxidation reactions may occur in different orders via the corresponding intermediates.

[0036] In a preferred embodiment, QA can be produced from OS using a gene encoding a biosynthase as shown below

[0037] 2,3-Oxidosqualene (OS) may be converted to β-amyrin using Saponaria officinalis synthase (SobAS). SobAS may have the amino acid sequence of SEQ ID NO: 8, or may be a variant or fragment thereof.

[0038] Alternatively, 2,3-oxidosqualene (OS) may be converted to β-amyrin by an endogenous enzyme in the host cell.

[0039] The C28 position of β-amyrin may be oxidized to carboxylic acid using SoC28 oxidase (SoC28) to produce oleanolic acid. SoC28 oxidase may have the amino acid sequence of SEQ ID NO: 2, or may be a variant or fragment thereof.

[0040] Next, the C16 position of oleanolic acid can be oxidized to alcohol using Saponaria officinalis C28C16 oxidase (SoC28C16) to produce echinocystic acid. C28C16 oxidase has the amino acid sequence of SEQ ID NO: 4, or is a variant or fragment thereof.

[0041] Alternatively, Saponaria officinalis C28C16 oxidase (SoC28C16) may be used to oxidize the C28 position of β-amyrin to carboxylic acid and the C16 position to alcohol to produce echinocystic acid. SoC28C16 oxidase may have the amino acid sequence of SEQ ID NO: 4, or may be a variant or fragment thereof.

[0042] The C-23 position of echinocystic acid can be oxidized to aldehyde to produce QA using SoC23 oxidase (SoC23). SoC23 oxidase may have the amino acid sequence of SEQ ID NO: 6, or may be a variant or fragment thereof.

[0043] In some embodiments, the genes described herein may be useful for the glycosylation of the C3 position of QA.

[0044] Glycosylation is initiated with a Β-D-glucuronic acid (GlcA) residue attached to the 3-O position of QA. The GlcA residue binds to D-galactose (Gal) via a β-1→2 bond and to D-xylose (Xyl) via a β-1,3 bond.

[0045] In a preferred embodiment, QA or the C28 glycosylated form of QA can be glycosylated at the 3-O position using the gene encoding the crude synthase defined below.

[0046] D-glucuronic acid (“GlcA”) can be transferred to the 3-O position of quillaic acid to form 3-O-{β-D-glucopyranosiduronic acid}-quillaic acid (“QA-GlcA” or “QA-Mono”) using Saponaria officinalis QA 3-O glucuronosyltransferase (“SoQA-GlcAT;SoCSL”). SoCSL may have the amino acid sequence of SEQ ID NO: 10, or may be a variant or fragment thereof.

[0047] D-Galactose ("Gal") is transferred to QA mono (QA-GlcA) via a β-1→2 linkage and Saponaria officinalis QA-GlcA galactosyltransferase ("SoQA-GalT" or "SoC3Gal") is used to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-oleanolic acid ("QA-GlcA-Gal" or QA-Di). SoC3Gal may have the amino acid sequence of SEQ ID NO: 12, or may be a variant or fragment thereof.

[0048] D-Xylose ("Xyl") is transferred to QA-Di via a 1→3 linkage and may form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)-β-D-glucopyranosiduronic acid]-oleanolic acid ("QA-GlcA-[Gal]-xyl" or QA-Tri). QA-XylT may have the amino acid sequence of SEQ ID NO: 14, or may be a variant or fragment thereof.

[0049] In some embodiments, the genes described herein may be useful for glycosylation at the C28 position of QA or the glycosylation of the C-3 glycosylated form of QA.

[0050] Saponaria officinalis QA-Tri fucosyltransferase ("SoQA-TriFuT" or "SoC28Fu") is used to transfer D-fucose ("Fuc") to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF). SoC28Fu may have the amino acid sequence of SEQ ID NO: 16, or may be a variant or fragment thereof.

[0051] L-Rhamnose ("Rhap") is transferred to QA-TriF via a β-1→2 bond and can form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFR) using Saponaria officinalis QA-TriF rhamnosyltransferase ("SoQA-TriFRhaT" or "SoC28Rha"). SoC28Rha can have the amino acid sequence of SEQ ID NO: 18, or can be a variant or fragment thereof.

[0052] D-Xylose ("Xyl") is transferred to QA-TriFR via a 1→4 bond and can form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRX) using Saponaria officinalis QA-TriFR xylosyltransferase ("SoQA-TriFRXylT" or SoC28Xyl1). SoC28Xyl1 can have the amino acid sequence of SEQ ID NO: 20, or can be a variant or fragment thereof.

[0053] D-Xylose (“xylose”) is transferred to QA-TriFRX via a 1→3 linkage and can be used with Saponaria officinalis QA-TriFRX xylosyltransferase (“SoQA-TriFRXXylT” or SoC28Xyl2) to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRXX). SoQA-TriFRXXylT can have the amino acid sequence of SEQ ID NO: 22 or can be a variant or fragment thereof.

[0054] The quinovosyl group of QA-TriF(Q)RXX is acetylated and can be used with the Saponaria officinalis QA-TriF(Q)RXX acetyltransferase (“SoBAHD1”) polypeptide to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-2)-[β-D-4-O-acetylquinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (SpB). SoBAHD1 can have the amino acid sequence of SEQ ID NO: 36 or can be a variant or fragment thereof

[0055] In preferred embodiments, the methods described herein will optionally involve using one or more of these newly characterized triterpenoid biosynthetic nucleic acids (e.g., one, two, three or more such nucleic acids) in combination with manipulation of other genes that affect QA or glycosylated QA biosynthesis known in the art.

[0056] These newly characterized triterpenoid biosynthetic amino acid and nucleotide sequences (SEQ ID NOs: 1-22, and 33-36) from Saponaria officinalis, as well as variants of these sequences and methods of using them, form aspects of the invention in and of themselves. Any one of these sequences or variants can be used to alter the QA or glycosylated QA content of a plant, as disclosed herein. For example, a variant nucleic acid can include a sequence encoding a variant polypeptide that shares the relevant biological activity of the native polypeptide, as discussed above. Examples include variants of any of SEQ ID NOs: 1-22 and 33-36. For the sake of brevity, in the context of the present invention, particularly in the methods and uses described herein, the polypeptide or nucleotide sequences of SEQ ID NOs: 1-22 and 33-36, as well as variants thereof described herein, may be referred to herein as "triterpenoid biosynthetic sequences", such as triterpenoid biosynthetic genes and triterpenoid biosynthetic polypeptides.

[0057] Provided herein is a Saponaria officinalis synthase (SobAS) polypeptide having the amino acid sequence of SEQ ID NO: 8 or a variant thereof, such as an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 8. Also provided herein are a nucleic acid encoding the SobAS polypeptide having the nucleotide sequence of SEQ ID NO: 7 or a variant thereof, such as a nucleotide sequence having at least 80% sequence identity with SEQ ID NO: 7; and a vector comprising the nucleic acid. The SobAS polypeptide can cyclize the universal linear precursor 2,3-oxidosqualene (OS) to triterpenes.

[0058] In addition, in the present specification, there is provided a SoC28 oxidase polypeptide having the amino acid sequence of SEQ ID NO: 2 or a variant thereof, for example, an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 2. Further, in the present specification, there are provided a nucleic acid encoding the SoC28 oxidase polypeptide having the nucleotide sequence of SEQ ID NO: 1 or a variant thereof, for example, a nucleotide sequence having at least 80% sequence identity with SEQ ID NO: 1; and a vector containing the nucleic acid. The SoC28 oxidase polypeptide can oxidize β-amyrin at the C28 position to form a carboxylic acid, oleanolic acid.

[0059] In addition, provided in the present specification is a SoC28C16 oxidase polypeptide having the amino acid sequence of SEQ ID NO: 4 or a variant thereof, for example, an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 4. Further, in the present specification, there are provided a nucleic acid encoding the SoC28C16 oxidase polypeptide having the nucleotide sequence of SEQ ID NO: 3 or a variant thereof, for example, a nucleotide sequence having at least 50% sequence identity with SEQ ID NO: 3; and a vector containing the nucleic acid. The SoC28C16 oxidase polypeptide can oxidize β-amyrin to an alcohol at the C16 position and to a carboxylic acid at the C28 position to form echinocystic acid.

[0060] In addition, in the present specification, there is provided a SoC23 oxidase polypeptide having the amino acid sequence of SEQ ID NO: 6 or a variant thereof, for example, an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 6. Further, in the present specification, there are provided a nucleic acid encoding the SoC23 oxidase polypeptide having the nucleotide sequence of SEQ ID NO: 5 or a variant thereof, for example, a nucleotide sequence having at least 50% sequence identity with SEQ ID NO: 5; and a vector containing the nucleic acid. The SoC23 oxidase can oxidize echinocystic acid at the C-23 position to convert it into an aldehyde that forms QA.

[0061] Also provided herein is a Saponaria officinalis QA 3-O-glucuronosyltransferase ("SoCSL") polypeptide having the amino acid sequence of SEQ ID NO: 10 or a variant thereof, such as an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 10. Also provided herein are a nucleic acid encoding the QA 3-O-glucuronosyltransferase ("SoCSL") polypeptide having the nucleotide sequence of SEQ ID NO: 9 or a variant thereof, such as a nucleotide sequence having at least 60% sequence identity with SEQ ID NO: 9; and a vector containing the nucleic acid. SoCSL can bind D-glucuronic acid ("GlcA") to the 3-O position of quillaic acid to form 3-O-{β-D-glucopyranosiduronic acid}-quillaic acid ("QA-GlcA" or "QA-Mono").

[0062] Also provided herein is a Saponaria officinalis QA-GlcA galactosyltransferase ("SoC3Gal") polypeptide having the amino acid sequence of SEQ ID NO: 12 or a variant thereof, such as an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 12. Also provided herein are a nucleic acid encoding the Saponaria officinalis QA-GlcA galactosyltransferase ("SoC3Gal") polypeptide having the nucleotide sequence of SEQ ID NO: 11 or a variant thereof, such as a nucleotide sequence having at least 80% sequence identity with SEQ ID NO: 11; and a vector containing the nucleic acid. SoC3Gal can bind D-galactose ("Gal") to QA-GlcA via a β-1→2 linkage to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-quillaic acid ("QA-Di" or "QA-GlcA-Gal").

[0063] Also provided herein is a Saponaria officinalis QA-GlcA-Gal xylosyltransferase (SoC3Xyl) polypeptide having the amino acid sequence of SEQ ID NO: 14 or a variant thereof, for example, an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 14. Also provided herein are a nucleic acid encoding the Saponaria officinalis QA-GlcA-Gal xylosyltransferase (SoC3Xyl) polypeptide having the nucleotide sequence of SEQ ID NO: 13 or a variant thereof, for example, a nucleotide sequence having at least 80% sequence identity with SEQ ID NO: 13; and a vector containing the nucleic acid. SoC3Xyl can bind D-xylose (Xyl) to QA-GlcA-Gal via a 1,3-linkage to form 1,3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-kirayaic acid (QA-GlcAGalXyl or QA-Tri).

[0064] Also provided herein is a Saponaria officinalis QA-Tri fucosyltransferase (SoC28Fu) polypeptide having the amino acid sequence of SEQ ID NO: 16 or a variant thereof, for example, an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 16. Also provided herein are a nucleic acid encoding the Saponaria officinalis QA-Tri fucosyltransferase (SoC28Fu) polypeptide having the nucleotide sequence of SEQ ID NO: 15 or a variant thereof, for example, a nucleotide sequence having at least 60% sequence identity with SEQ ID NO: 15; and a vector containing the nucleic acid. SoC28Fu can bind fucose (Fuc) to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-fucopyranosyl ester}-kirayaic acid (QA-TriF).

[0065] Also provided herein is a Saponaria officinalis QA-TriF rhamnosyltransferase ("SoC28Rha") polypeptide having the amino acid sequence of SEQ ID NO: 18 or a variant thereof, such as an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 18. Further provided herein are a nucleic acid encoding the Saponaria officinalis QA-TriF rhamnosyltransferase ("SoC28Rha") polypeptide having the nucleotide sequence of SEQ ID NO: 17 or a variant thereof, such as a nucleotide sequence having at least 50% sequence identity with SEQ ID NO: 17; and a vector containing the nucleic acid. SoC28Rha can bind rhamnose ("Rha") to QA-TriF via a 1,2 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{α-L-rhamnopyranosyl-(1→2)-Β-D-fucopyranosyl ester}-gypsogenic acid (QA-TriFR).

[0066] Also provided herein is a Saponaria officinalis QA-TriFR xylosyltransferase ("SoC28Xyl1") polypeptide having the amino acid sequence of SEQ ID NO: 20 or a variant thereof, e.g., an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 20. Also provided herein are a nucleic acid encoding the Saponaria officinalis QA-TriFR xylosyltransferase ("SoC28Xyl1") polypeptide having the nucleotide sequence of SEQ ID NO: 19 or a variant thereof, e.g., a nucleotide sequence having at least 50% sequence identity with SEQ ID NO: 19; and a vector comprising the nucleic acid. SoC28Xyl1 attaches D-xylose ("xylose") to QA-TriFR via a 1,4-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kiragic acid (QA-TriFRX).

[0067] Also provided herein is a Saponaria officinalis QA-TriFRX xylosyltransferase (“SoC28Xyl2”) polypeptide having the amino acid sequence of SEQ ID NO: 22 or a variant thereof, such as an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 22. Further provided herein are a nucleic acid encoding the Saponaria officinalis QA-TriFRX xylosyltransferase (“SoC28Xyl2”) polypeptide having the nucleotide sequence of SEQ ID NO: 21 or a variant thereof, such as a nucleotide sequence having at least 50% sequence identity with SEQ ID NO: 21; and a vector comprising the nucleic acid. SoC28Xyl2 attaches D-xylose (“Xyl”) to QA-TriFRX via a 1,3 linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kie acid (QA-TriFRXX).

[0068] Also provided herein is a Saponaria officinalis QA-TriFRXX quinovosyltransferase ("SoGH1") polypeptide having the amino acid sequence of SEQ ID NO: 34 or a variant thereof, e.g., an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 34. Further provided herein are a nucleic acid encoding the Saponaria officinalis QA-TriFRXX quinovosyltransferase ("SoGH1") polypeptide having the nucleotide sequence of SEQ ID NO: 33 or a variant thereof, e.g., a nucleotide sequence having at least 50% sequence identity with SEQ ID NO: 33; and a vector comprising the nucleic acid. SoGH1 can bind D-quinovose ("Q") to QA-TriFRXX via a 1,4-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-quinic acid (QA-TriF(Q)RXX).

[0069] Also provided herein is a Saponaria officinalis QA-TriF(Q)RXX acetyltransferase ("SoBAHD1") polypeptide having the amino acid sequence of SEQ ID NO: 36 or a variant thereof, for example, an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 36. Further provided herein are a nucleic acid encoding the Saponaria officinalis QA-TriF(Q)RXX acetyltransferase ("SoBAHD1") polypeptide having the nucleotide sequence of SEQ ID NO: 35 or a variant thereof, for example, a nucleotide sequence having at least 50% sequence identity with SEQ ID NO: 35; and a vector comprising the nucleic acid. SoBAHD1 can acetylate QA-TriF(Q)RXX to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-4-O-acetylquinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-gypsogenic acid (SpB).

[0070] Any one of the amino acid sequences described herein that are variants of the reference sequence, for example, the peptide, polypeptide or protein sequences described herein, such as any one of SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 34 or 36, may have one or more amino acid residues modified relative to the reference sequence. For example, up to 50 amino acid residues may be modified relative to the reference sequence, preferably up to 45, up to 40, up to 30, up to 20, up to 15, up to 10, up to 5, or up to 3, 2 or 1. For example, variants described herein may include sequences of the reference sequence with up to 50, up to 45, up to 40, up to 30, up to 20, up to 15, up to 10, up to 5, up to 3, up to 2 or 1 amino acid residues mutated.

[0071] The amino acid residues in the reference sequence may be altered or mutated by insertion, deletion, or substitution, preferably substitution with a different amino acid residue. Such changes can be caused by one or more of addition, insertion, deletion, or substitution of one or more nucleotides in the encoding nucleic acid.

[0072] Any one of the nucleotide sequences described herein that are variants of the reference sequence, such as SEQ ID NO: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 33, or 35, may have one or more nucleotides modified relative to the reference sequence. For example, 50 or fewer nucleotides may be modified relative to the reference sequence, preferably 45 or fewer, 40 or fewer, 30 or fewer, 20 or fewer, 15 or fewer, 10 or fewer, 5 or fewer, 3 or fewer, 2 or fewer, or 1 or fewer. For example, the variants described herein can include the nucleotide sequences of the reference sequence with 50 or fewer, 45 or fewer, 40 or fewer, 30 or fewer, 20 or fewer, 15 or fewer, 10 or fewer, 5 or fewer, 3 or fewer, 2, or 1 nucleotide mutations.

[0073] The peptides, polypeptides, or proteins described herein that are variants of a reference sequence such as the above amino acid sequence or nucleotide sequence, or the nucleotide sequences described herein, may share at least 50% sequence identity, at least 55%, at least 60%, at least 65%, at least 70%, at least about 80%, at least 90%, at least 95%, at least 98%, or at least 99% sequence identity with the reference sequence. For example, variants of the proteins described herein can include amino acid sequences having at least 50% sequence identity with the reference amino acid sequence, at least 55%, at least 60%, at least 65%, at least 70%, at least about 80%, at least 90%, at least 95%, at least 98%, or at least 99% sequence identity with the reference amino acid sequence, such as one or more of SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 34, and 36.

[0074] The variants of the nucleic acids described herein have a nucleotide sequence having at least 50% sequence identity, at least 55%, at least 60%, at least 65%, at least 70%, at least about 80%, at least 90%, at least 95%, at least 98% or at least 99% sequence identity with a reference amino acid sequence, for example, may include one or more of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 33 and 35.

[0075] Variants of different variant triterpenoid biosynthetic sequences may share different levels of sequence identity with their respective reference sequences. Combinations of variant triterpenoid biosynthetic sequences having all levels of sequence identity disclosed above are encompassed by the present invention.

[0076] The identity of sequences is generally defined with reference to the algorithm GAP (Wisconsin GCG package, Accelerys Inc, San Diego USA). GAP uses the Needleman and Wunsch algorithm to align two complete sequences, maximizing the number of matches and minimizing the number of gaps. Usually, the default parameters are used, with a gap creation penalty = 12 and a gap extension penalty = 4. The use of GAP is preferred, but other algorithms may be used. For example, BLAST (Altschul et al. Biol:405-410), FASTA (which uses the method of Pearson and Lipman (1988) PNAS USA 85: 2444-2448), or the Smith-Waterman algorithm (Smith and Waterman (1981) J. Mol Biol. 147: 195-197), or the TBLASTN program of Altschul et al. (1990) supra generally employ default parameters. In particular, the psi-Blast algorithm may be used (Nucl. Acids Res. (1997) 25 3389-3402). Sequence identity and similarity can also be determined using Genomequest™ software (Gene-IT, Worcester MA USA).

[0077] Sequence comparisons are preferably made over the entire length of the relevant sequences described herein.

[0078] Variant polypeptides may share the relevant biological activities of the reference polypeptide. Variant nucleic acids may encode the relevant variant polypeptides. In this context, the "biological activity" of the polypeptides described herein is the ability to catalyze each of the reactions shown in Figure 15 and described above. The relevant biological activities can be assayed based on the reactions shown in Figure 15 in vitro. Alternatively, as described in the Examples, the activities in vivo, i.e., by introducing a plurality of heterologous constructs for generating each product into a host, can be assayed, which can be assayed by, for example, LC-MS.

[0079] Preferred variations are as follows: (i) Naturally occurring nucleic acids such as alleles (including polymorphisms or mutations at one or more bases) or pseudoalleles (which can occur at loci closely linked to the biosynthetic genes described herein). Also included are paralogs, isogenes, or other homologous genes belonging to the same family as the biosynthetic genes described herein, for example, sharing a clade or subclade. Also included are orthologs or homologous genes from other plant species (i.e., plants other than S. officinalis). Homology can be at the nucleotide sequence level and / or the amino acid sequence level, as described below. (ii) Artificial nucleic acids that can be prepared by those skilled in the art in light of the present disclosure. Such derivatives can be prepared, for example, by site-directed or random mutagenesis, or by direct synthesis. Preferably, the mutant nucleic acid is generated directly or indirectly (e.g., via one or more amplification or replication steps) from the original nucleic acid having all or part of the sequence of the biosynthetic gene described herein.

[0080] Variants include nucleic acids corresponding to the above nucleic acids but extended at the 3' or 5' end.

[0081] A method for producing a variant triterpenoid biosynthetic nucleic acid may include modifying any of the genes described herein, such as one or more of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 33, and 35.

[0082] The changes may be desirable for many reasons. For example, introduction or removal of restriction endonuclease sites, change of codon usage, etc. This is particularly desirable when the gene is expressed in an alternative host, such as a microbial host like yeast. Methods for codon-optimizing genes for this purpose are known in the art (see, for example, Elena, Claudia, et al. "Expression of codon optimized genes in microbial systems: current industrial applications and perspectives." Frontiers in microbiology 5 (2014)). The sequences described herein that include codon modifications to maximize yeast expression represent embodiments of the present invention.

[0083] Alternatively, the sequence changes can produce derivatives by one or more (e.g., several) methods of addition, insertion, deletion, or substitution of one or more nucleotides in the nucleic acid, resulting in addition, insertion, deletion, or substitution of one or more (e.g., several) amino acids in the encoded polypeptide.

[0084] Such changes can modify cleavage sites of the encoded polypeptide; sites required for post-translational modifications such as motifs for phosphorylation of the encoded polypeptide. If isolation from a microbial system is desired, a leader sequence or other targeting sequence (e.g., a membrane or Golgi localization sequence) can be added to the expressed protein to determine its location after expression.

[0085] Other desirable mutations are random mutagenesis or site-directed mutagenesis to change the activity (e.g., specificity) or stability of the encoded polypeptide. That is, substituting hydrophobic residues such as isoleucine, valine, leucine, methionine, etc. with another residue, or substituting arginine with lysine, glutamic acid with aspartic acid, glutamine with asparagine. As is well known to those skilled in the art, even when changing the primary structure of a polypeptide by conservative substitution, the activity of the peptide may not change significantly. This is because the side chain of the inserted amino acid in the sequence can form similar bonds and contacts as the side chain of the substituted amino acid. This is the same even when the substitution is in a region important for determining the conformation of the peptide. Also included are variants with non-conservative substitutions. As is well known to those skilled in the art, substitutions in regions not important for determining the three-dimensional structure of a peptide may not significantly change the three-dimensional structure of the peptide and thus may not have a significant impact on activity. In regions important for determining the three-dimensional structure and activity of a peptide, such changes may confer advantageous properties on the polypeptide. In fact, changes such as those described above may confer slightly advantageous properties on the peptide, such as changes in stability or specificity.

[0086] In some embodiments, the variant nucleotide sequence encoding the So polypeptide can be obtained by a method comprising: (a) providing a preparation of nucleic acids from, for example, plant cells. The test nucleic acids can be provided from the cells as genomic DNA, cDNA, RNA, or a mixture thereof, preferably as a library in a suitable vector. When using genomic DNA, the probe can be used to identify the untranscribed region of the gene (e.g., promoter, etc.) as described below. (b) providing a nucleic acid molecule that is a probe or primer as described above. (c) contacting the nucleic acids in the preparation with the nucleic acid molecule under hybridization conditions of the nucleic acid molecule to any of the genes or homologous genes in the preparation, and (d) Identifying the gene or homolog, if present, by hybridization with the nucleic acid molecule. Binding of the probe to the target nucleic acid (e.g., DNA) can be measured using any of a variety of techniques freely available to those skilled in the art. For example, the probe can be labeled radioactively, fluorescently, or enzymatically. Other methods without using a label for the probe include amplification using PCR (see below), RNase cleavage, allele-specific oligonucleotide probing, etc. After confirming successful hybridization, the hybridized nucleic acid is isolated, which includes one or more steps of PCR or amplification of the vector in a suitable host.

[0087] Preliminary experiments can be performed by conducting hybridization under low stringency conditions. For probing, more stringent conditions are preferred such that there is a simple pattern of a small number of hybridizations that can be further investigated and identified as positive.

[0088] For example, hybridization can be performed using a hybridization solution containing the following, according to the method of Sambrook et al. (below): 5X SSC (where "SSC" = 0.15 M sodium chloride; 0.15 M sodium citrate; pH 7), 5X Denhardt's reagent, 0.5 - 1.0% SDS, 100 μg / ml denatured fragmented salmon sperm DNA, 0.05% sodium pyrophosphate, and up to 50% formamide. Hybridization is performed at 37 - 42 °C for at least 6 hours. After hybridization, the filter is washed as follows: (1) in 2X SSC and 1% SDS for 5 minutes at room temperature, (2) in 2X SSC and 0.1% SDS for 15 minutes at room temperature, (3) in 1X SSC and 1% SDS at 37 °C for 30 minutes to 1 hour, (4) in 1X SSC and 1% SDS at 42 - 65 °C for 2 hours, with the solution changed every 30 minutes.

[0089] One common formula for calculating the stringency conditions necessary to achieve hybridization between nucleic acid molecules of a specified sequence homology is (Sambrook et al., 1989): T m = 81.5°C + 16.6 Log [Na+] + 0.41 (%G+C) - 0.63 (% formamide) - 600 / #bp (duplex).

[0090] As an explanation of the above formula, when [Na+] = [0.368], 50% formamide is used, the GC content is 42%, and the average probe size is 200 bases, T m is 57°C. The T m of a DNA duplex decreases by 1 - 1.5°C for every 1% decrease in homology. Therefore, targets with approximately 75% or more sequence identity will be observed at a hybridization temperature of 42°C. Such sequences are considered to be substantially homologous to the nucleic acid sequences of the present invention.

[0091] It is well known in the art to gradually increase the stringency of hybridization until only a few positive clones remain. Other suitable conditions are, for example, for detecting sequences that are approximately 80 - 90% identical, hybridize overnight at 42°C in 0.25 M Na2HPO4, pH 7.2, 6.5% SDS, 10% dextran sulfate, and wash finally at 55°C in 0.1X SSC, 0.1% SDS. For detecting sequences that are approximately 90% or more identical, the conditions of hybridizing overnight at 65°C in 0.25 M Na2HPO4, pH 7.2, 6.5% SDS, 10% dextran sulfate and washing finally at 60°C in 0.1X SSC, 0.1% SDS are suitable.

[0092] In a further embodiment, hybridization to a variant of a triterpenoid biosynthetic nucleic acid molecule can be determined or identified indirectly, for example, using a nucleic acid amplification reaction, particularly the polymerase chain reaction (PCR). In PCR, since two primers are required to specifically amplify the target nucleic acid, preferably two nucleic acid molecules having sequences characteristic of the triterpenoid biosynthetic gene are employed. If RACE PCR is used, only one such primer may be sufficient (see "PCR protocols; A Guide to Methods and Applications", Eds. Innis et al, Academic Press, New York, (1990)).

[0093] Accordingly, methods including the use of PCR in obtaining the mutant triterpenoid biosynthetic nucleic acids described herein may include the following: (a) providing a preparation of plant nucleic acids, for example, a preparation from seeds or other suitable tissues or organs, (b) providing a pair of nucleic acid molecule primers useful for PCR (i.e., suitable for PCR), at least one of said primers being a primer directed to a triterpenoid biosynthetic sequence as described above, (c) contacting the nucleic acids in said preparation with said primers under PCR performance conditions, (d) performing PCR and determining the presence or absence of the amplified PCR product. The presence of the amplified PCR product may indicate the identification of a variant.

[0094] In all of the above cases, if necessary, the clones or fragments identified by the search can be extended. For example, if they are suspected to be incomplete, the original DNA source (such as a clone library, mRNA preparation, etc.) can be reexamined to isolate the missing part. For example, sequences, probes or primers based on the part already obtained can be used to identify other clones containing overlapping sequences.

[0095] The methods described herein can utilize fragments of the triterpenoid biosynthetic genes described herein, such as one or more of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 33, and 35; or fragments of variants of these genes. Also provided are fragments of the full-length polypeptides disclosed herein, particularly the production and use of the active portions thereof. The “active portion” of a polypeptide means a peptide that is less than the full-length polypeptide but retains its essential biological activity, for example, in relation to the production of QA or the glycosylation of QA.

[0096] Fragments of the full-length reference triterpenoid biosynthetic polypeptide sequences such as SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 34, or 36 are contiguous amino acid sequences from the full-length protein sequence and consist of at least one fewer amino acid than the full-length protein sequence. For example, the fragment may lack 10 or more, 20 or more, 50 or more, 100 or more amino acid sequences relative to the full-length sequence. Preferably, the fragment shares the relevant biological activity of the full-length reference polypeptide.

[0097] In some embodiments, the fragment of the polypeptide can contain one or more epitopes useful for raising antibodies against any of the amino acid sequences disclosed herein. Preferred epitopes are those to which an antibody can specifically bind, which means binding to the polypeptide or a fragment thereof with an affinity that is at least about 1000-fold greater than that for other polypeptides.

[0098] Purified proteins (polypeptides, enzymes), or fragments, variants, derivatives or mutants thereof, such as those recombinantly produced by expression from nucleic acids encoding triterpenoid biosynthetic nucleic acids, form one aspect of the invention.

[0099] Such purified polypeptides can be used to generate antibodies using standard techniques in the art. Polypeptides comprising antibodies and antigen-binding fragments of antibodies can also be used to identify homologs from other species, as described further below.

[0100] Methods for producing antibodies include immunizing a mammal (e.g., human, mouse, rat, rabbit, horse, goat, sheep or monkey) with the protein or a fragment thereof. Antibodies can be obtained from the immunized animal using any of a variety of techniques known in the art, and preferably can be screened using binding of the antibody to the antigen of interest. For example, Western blotting techniques or immunoprecipitation can be used (Armitage et al, 1992, Nature 357: 80-82). The antibodies can be polyclonal or monoclonal.

[0101] As an alternative or supplement to immunization of a mammal, antibodies having appropriate binding specificity can be obtained from a recombinant production library of expressed immunoglobulin variable domains. For example, λ bacteriophage or filamentous bacteriophage having functional immunoglobulin binding domains on their surface are used; see, for example, WO92 / 01047.

[0102] Antibodies raised against a polypeptide or peptide can be used to identify and / or isolate homologous polypeptides and the genes encoding them.

[0103] Antibodies can be modified in a variety of ways. In fact, the term "antibody" should be construed to include any specific binding substance having a binding domain with the required specificity. Thus, the term covers antibody fragments, derivatives, functional equivalents, homologs, including any polypeptide comprising an immunoglobulin binding domain, whether natural or synthetic.

[0104] Mevalonic acid (MVA) is an important intermediate in triterpenoid synthesis. Therefore, in order to maximize the yield of triterpenoids such as QA, it is considered desirable to express the rate-limiting MVA pathway genes in the host.

[0105] 3-Hydroxy-3-methylglutaryl-CoA reductase (HMGR) is considered to be the rate-limiting enzyme of the MVA pathway.

[0106] The use of a recombinant feedback-insensitive truncated form of HMGR (tHMGR) has been demonstrated to increase the triterpene (β-amyrin) content during transient expression in N. benthamiana [Reed, J., et al. Metab Eng, 2017. 42: p. 185-193].

[0107] In some embodiments, a heterologous HMGR (e.g., a feedback-insensitive HMGR) can be used with the triterpenoid biosynthetic genes described herein.

[0108] Examples of sequences encoding HMGR or polypeptide sequences include SEQ ID NOs: 23-26, or variants or fragments thereof. Variants may be homologs, alleles, or artificial derivatives, etc., as discussed in relation to the biosynthetic genes or polypeptides described above. For example, an HMGR native to the host utilized may be preferred - for example, yeast HMGR in a yeast host. HMGR genes are known in the art and can be appropriately selected in light of the present disclosure.

[0109] It has also been reported that squalene synthase (SQS) may be the rate-limiting step [Reed et al supra].

[0110] In some embodiments, a heterologous SQS can be used with the biosynthetic genes described herein and optionally with the HMGR described herein

[0111] Examples of SQS coding sequences or polypeptide sequences include SEQ ID NOs: 27 and 28, or variants or fragments thereof. Variants may be homologs, alleles, or artificial derivatives, etc., as discussed in relation to the biosynthetic genes or polypeptides described above. For example, an SQS native to the host to be utilized is preferred, such as yeast SQS in a yeast host. SQS genes are known in the art and can be appropriately selected in light of the present disclosure.

[0112] When using a specific host (e.g., yeast), it may be desirable to introduce additional genes to improve the flux of biosynthetic production. Examples include one or more plant cytochrome P450 reductases (CPRs) that function as the redox partner of the introduced P450. In some embodiments, a heterologous cytochrome P450 reductase such as AtATR2 (Arabidopsis thaliana cytochrome P450 reductase 2) can be used together with the biosynthetic polypeptides and genes described herein. Examples of the sequence or polypeptide sequence encoding AtATR2 include SEQ ID NOs: 29 and 30, or variants or fragments thereof. Variants may be homologs, alleles, or artificial derivatives, etc., as discussed in relation to the biosynthetic polypeptides and genes described above.

[0113] In some embodiments, the heterologous nucleic acids described herein may further encode one or more of the following polypeptides: (i) HMG-CoA reductase (HMGR) and / or (ii) squalene synthase (SQS). HMGR or SQS can be arbitrarily selected from the respective polypeptides of SEQ ID NOs: 24, 26, and 28, or variants or fragments of any of the polypeptides, or encoded by the respective polynucleotides of SEQ ID NOs: 23, 25, and 27, or variants or fragments of any of the polynucleotides.

[0114] Nucleic acids include cDNA, RNA, genomic DNA, modified nucleic acids or nucleic acid analogs (such as peptide nucleic acids). When a DNA sequence is identified, for example with reference to a figure, the RNA equivalent in which U appears instead of T is included, unless otherwise required by the context. Nucleic acids can contain multiple nucleic acid molecules. The nucleic acid molecules according to the present invention are isolated and / or purified from their natural environment, provided in a substantially pure or homogeneous form, or free of or substantially free of other nucleic acids of the originating species, and can be provided in double-stranded or single-stranded form. As used herein, the term "isolated" encompasses all of these possibilities. Nucleic acid molecules may be wholly or partially synthetic. In particular, they may be recombinants in which nucleic acid sequences that do not occur together (are not contiguous) in nature are artificially joined by ligation or other methods. Nucleic acids may contain, consist of, or consist essentially of any of the sequences described hereinafter.

[0115] The complement of a nucleic acid as described herein means the complementary sequence of a nucleic acid or the nucleotide sequence composed of the nucleic acid. Optionally, the complementary sequence is the full length compared to the reference nucleotide sequence.

[0116] As used herein, the term "heterologous" is used in a broad sense to indicate that the sequence of the gene / nucleotide in question (e.g., encoding a biosynthetically modified polypeptide) has been introduced into the host or the said cell of its ancestor by genetic engineering, i.e., by artificial intervention. A heterologous nucleic acid with respect to a host cell is not naturally present in a cell of that type, variety or species. Thus, a heterologous nucleic acid can include a coding sequence of a particular type of plant cell or species or variety of plant, or a coding sequence derived therefrom, placed within a plant cell of a different type or species or variety. A further possibility is that the nucleic acid sequence is placed within a cell in which it or a homolog is naturally found, but the nucleic acid sequence is operably linked to one or more regulatory sequences, such as a promoter sequence, for the control of expression, and is linked and / or adjacent to a nucleic acid that is not naturally present in the cell, or in a plant cell of that type or species or variety.

[0117] As used in this context, "transformation" means that the nucleotide sequence of a heterologous nucleic acid changes one or more properties of a cell, and thus the phenotype, with respect to the ability to biosynthesize triterpenoids such as QA or glycosylated QA, for example QATri, QATriFRXX, QATriF(Q)RXX or SpB. Such transformation may be transient or stable.

[0118] "Unable to perform biosynthesis" means that the host before transformation does not naturally produce or is considered not to produce a detectable or recoverable level of product under the normal metabolic conditions of the host. After application of the present invention, a detectable or recoverable level of product can be produced.

[0119] The nucleotide sequence information provided herein can be used for the design of probes and primers for probing or amplification. The oligonucleotides used for probing or PCR may be about 30 nucleotides or less in length (e.g., 18, 21 or 24). Generally, specific primers are at least 14 nucleotides in length. To obtain optimal specificity and cost-effectiveness, primers 16-24 nucleotides in length are preferred. Those skilled in the art are proficient in the design of primers for use in processes such as PCR. If necessary, probing can be performed using the entire restriction enzyme fragment of the gene disclosed herein. If necessary, small mutations can be introduced into the sequence to create "consensus" or "degenerate" primers.

[0120] Standard Southern blotting techniques can be used for probing. For example, DNA is extracted from cells and digested with different restriction enzymes. The restriction enzyme fragments are then separated by electrophoresis on an agarose gel, denatured, and transferred to a nitrocellulose filter. A labeled probe can be hybridized to the single-stranded DNA fragments on the filter to determine binding. The DNA for probing can be prepared from an RNA preparation from the cells. Probing can also be performed using so-called "nucleic acid chips" (see review by Marshall & Hodgson (1998) Nature Biotechnology 16: 27-31).

[0121] The methods described herein may employ co-infiltration of multiple Agrobacterium tumefaciens strains, each having one or more of the triterpenoid biosynthetic genes described above, for coordinated expression in the biosynthetic pathway described above.

[0122] In some embodiments, at least two or three different Agrobacterium tumefaciens strains are co-infiltrated, for example, each having a triterpenoid biosynthetic nucleic acid.

[0123] The genes may be present from transient expression vectors

[0124] Vectors (typically binary vectors) for use as described herein may typically contain an expression cassette comprising: (i) a promoter, optionally linked to: (ii) an enhancer sequence derived from the RNA-2 genomic segment of a bipartite RNA virus, wherein the target start site in the RNA-2 genomic segment is mutated; (iii) a nucleic acid sequence encoding one or more of the biosynthetic genes described above; (iv) a terminator sequence; and optionally (v) a 3'UTR located upstream of the terminator sequence.

[0125] Further examples of vectors and expression systems suitable for use as described herein are set forth below.

[0126] The above triterpenoid biosynthesis, recombinant, preferably replicable vectors may be included in or in the form of such vectors. Examples of vectors include, in particular, double-stranded or single-stranded linear or circular plasmids, cosmids, phages or Agrobacterium binary vectors, which may or may not be self-transmissible or mobilizable and which may transform prokaryotic or eukaryotic hosts by integration into the cell genome or may exist episomally (e.g., autonomously replicating plasmids having an origin of replication).

[0127] Suitable expression vectors include binary vectors for transient expression mediated by Agrobacterium tumefaciens (see, for example, Bevan et al Nucl Acid Res 984 Nov 26; 12(22): 8711-872). As is well known to those skilled in the art, the "binary vector" system includes: (a) border sequences that enable the transfer of a desired nucleotide sequence into the plant cell genome; (b) the desired nucleotide sequence itself, generally including an expression cassette operably linked to (i) a plant-active promoter, (ii) a target sequence, and optionally an enhancer. The desired nucleotide sequence is located between the border sequences and can be inserted into the plant genome under appropriate conditions. The binary vector system generally requires other sequences (derived from A. tumefaciens) for integration. This is generally achieved by using so-called "agroinfiltration", which uses transient transformation via Agrobacterium. Briefly, this technique is based on the property that Agrobacterium tumefaciens transfers a part of its DNA ("T-DNA") into the host cell and integrates it into the nuclear DNA. T-DNA is defined by left and right border sequences approximately 21-23 nucleotides in length. Infiltration can be achieved, for example, by syringe (leaves) or vacuum (whole plant). In the present invention, the border sequences generally surround the desired nucleotide sequence (T-DNA), and one or more vectors are introduced into the plant material by agroinfiltration.

[0128] Other suitable expression systems can utilize a system called the "Hyper-Translatable Cowpea Mosaic Virus ("CPMV-HT") system". Suitable vectors based on the pEAQ-HT expression plasmid for use in the CPMV-HT system are well known in the art (e.g., WO2009 / 087391; Sainsbury et al (2009) Plant Biotechnol J 7(7): 682-693).

[0129] Generally, one of ordinary skill in the art is fully capable of constructing vectors and designing protocols for recombinant gene expression (e.g., expressing a heterologous nucleic acid in a host or one or more cells of a host). Suitable vectors can be selected or constructed to contain appropriate regulatory sequences, including promoter sequences, terminator fragments, polyadenylation sequences, enhancer sequences, marker genes, and other appropriate sequences. For further details, see, for example, Molecular Cloning: a Laboratory Manual: 2nd edition, Sambrook et al, 1989, Cold Spring Harbor Laboratory Press or Current Protocols in Molecular Biology, Second Edition, Ausubel et al. eds., John Wiley & Sons, 1992.

[0130] Specifically, a shuttle vector is included. A shuttle vector means a DNA vehicle that is capable of replicating in two different host organisms, either naturally or by design, and can be selected from actinomycetes and related species, bacteria, eukaryotes (e.g., higher plants, mosses, yeast, or fungal cells).

[0131] Vectors containing the nucleic acids described herein need not contain a promoter or other regulatory sequences, particularly when the vector is used to introduce the nucleic acid into cells for recombination into the genome.

[0132] Preferably, the nucleic acid in the vector is under the control of, and operably linked to, a suitable promoter or other regulatory element for transcription in a host cell such as a microorganism, e.g., yeast or bacteria, or a plant cell. The vector may be a bifunctional expression vector that functions in multiple hosts. In the case of genomic DNA, this can include its own promoter or other regulatory elements (optionally in combination with a heterologous enhancer such as the 35S enhancer discussed in the examples below). The advantage of using a native promoter is that a pleiotropic response can be avoided. In the case of cDNA, this can be under the control of a suitable promoter or other regulatory element for expression in the host cell.

[0133] A promoter is a nucleotide sequence at which transcription of DNA operably linked downstream (i.e., in the 3' direction of the sense strand of double-stranded DNA) is initiated.

[0134] Operably linked means being joined as part of the same nucleic acid molecule and having the appropriate position and orientation for transcription to be initiated from the promoter. DNA operably linked to a promoter is "under the transcriptional control" of the promoter.

[0135] Suitable promoters include inducible promoters. The term "inducible" as applied to a promoter is well understood by those skilled in the art. In essence, expression under the control of an inducible promoter is "switched on" or increased in response to an applied stimulus. The nature of the stimulus varies with the promoter. Some inducible promoters cause little or undetectable levels of expression (or no expression) in the absence of a suitable stimulus. Other inducible promoters cause detectable constitutive expression in the absence of the stimulus. Regardless of the level of expression in the absence of the stimulus, expression from an inducible promoter increases in the presence of a suitable stimulus.

[0136] Accordingly, the nucleic acids described herein can be placed under the control of an externally inducible gene promoter in order to place expression (expression of a heterologous sequence) under the control of the user. The advantage of introducing a heterologous gene into a plant cell is that the expression of the gene can be placed under the control of any promoter so that, particularly when the cell is incorporated into a plant, gene expression, and thus QA or glycosylated QA biosynthesis, can be affected according to preference. Furthermore, mutants and derivatives of wild-type genes, for example those with higher or lower activity than the wild-type, can be used in place of the endogenous gene.

[0137] Also provided are gene constructs, preferably replicable vectors, comprising a promoter (optionally inducible) operably linked to a biosynthetic gene or a mutant thereof described herein.

[0138] Particularly interesting in this context are nucleic acid constructs that function as plant vectors. Certain procedures and vectors that have been widely successful in plants have previously been described in Guerineau and Mullineaux (1993) (Plant transformation and expression vectors. In: Plant Molecular Biology Labfax (Croy RRD ed.) Oxford, BIOS Scientific Publishers, pp 121-148). Suitable vectors include vectors derived from plant viruses (see, for example, EP-A-194809).

[0139] Preferably, the vector for use in plants contains border sequences that allow the introduction and integration of the expression cassette into the plant genome. Preferably, the construct is a plant binary vector. Preferably, the binary transformation vector is based on pPZP (Hajdukiewicz, et al. 1994). Other exemplary constructs include pBin19 (see Frisch, D. A., L. W. Harris-Haller, et al. (1995). “Complete Sequence of the binary vector Bin 19.” Plant Molecular Biology 27: 405-409).

[0140] Suitable promoters that act in plants include the cauliflower mosaic virus 35S (CaMV 35S). Other examples are disclosed on page 120 of Lindsey & Jones (1989) “Plant Biotechnology in Agriculture” Pub. OU Press, Milton Keynes, UK. The promoter can be selected to contain one or more sequence motifs or elements that provide developmental and / or tissue-specific control of expression. Inducible plant promoters include the ethanol-inducible promoter of Caddick et al (1998) Nature Biotechnology 16: 177-180.

[0141] Optionally, the construct can include a selectable genetic marker that confers a selectable phenotype such as resistance to an antibiotic or herbicide (e.g., kanamycin, hygromycin, phosphinothricin, chlorosulfuron, methotrexate, gentamicin, spectinomycin, imidazolinone, and glyphosate). A positive selection system as described in Haldrup et al. 1998 Plant molecular Biology 37, 287-296 can be used to create a construct that does not rely on antibiotics.

[0142] As described above, a preferred vector is a "CPMV-HT" vector as described in WO2009 / 087391. The following examples demonstrate the use of these pEAQ-HT expression plasmids.

[0143] These vectors (typically binary vectors) used in the present invention typically contain an expression cassette comprising: (i) a promoter, optionally linked to: (ii) an enhancer sequence derived from the RNA-2 genomic segment of a bipartite RNA virus, wherein the target start site in the RNA-2 genomic segment is mutated; (iii) the nucleic acid sequence as described above; (iv) a terminator sequence; and optionally (v) a 3' UTR located upstream of said terminator sequence.

[0144] The enhancer sequence (or enhancer element) is a sequence derived from (or sharing homology with) the RNA-2 genomic segment of a bipartite RNA virus such as a comovirus in which the target start site is mutated. Such a sequence can enhance expression downstream of the heterologous ORF to which it is bound. When present in the transcribed RNA, such a sequence can also enhance the translation of the heterologous ORF to which it is bound.

[0145] The target start is the start site (start codon) in the wild-type RNA-2 genomic segment of the bipartite virus (e.g., comovirus) from which the enhancer sequence is derived, and functions as the start site for the production (translation) of the longer of the two carboxy co-terminal proteins encoded by the wild-type RNA-2 genomic segment.

[0146] Typically, the RNA virus is a comovirus as described above.

[0147] The most preferred vector is the pEAQ vector of WO2009 / 087391 that enables direct cloning by use of a polylinker between the 5' leader and 3' UTR of an expression cassette containing the translation enhancer of the present invention, and is arranged on the T-DNA that also contains a suppressor of gene silencing and an NPTII cassette.

[0148] In such a gene expression system, it is preferred but not essential that a gene silencing suppressor be present. Suppressors of gene silencing are known in the art and are described in WO / 2007 / 135480. These include HcPro from potato virus Y, He-Pro from TEV, P19 from TBSV, rgsCam, the B2 protein from FHV, the small coat protein of CPMV, and the coat protein from TCV. A preferred suppressor for producing stable transgenic plants is the P19 suppressor incorporating the R43W mutation.

[0149] As described herein, a host can be converted from a phenotype in which the host cannot perform the effective biosynthesis described herein to a phenotype in which the host can perform said biosynthesis, and the product can be recovered therefrom or utilized in vivo for synthesizing downstream products.

[0150] Biosynthesis includes (i) the conversion from OS to QA, or the conversion to an intermediate such as oleanolic acid or echinocystic acid, (ii) the conversion from QA to QA-Tri, or the conversion to an intermediate such as QA-Mono or QA-Di (iii) the conversion from QA-Tri to QA-TriFRXX, or the conversion to an intermediate such as QA-TriF, QA-TriFR or QA-TriFX (iv) the conversion from QA to SpB, or the conversion to an intermediate such as QA-TriF(Q)RXX.

[0151] Biosynthesis may also include (i) the conversion of OS to β-amyrin, (ii) the conversion of β-amyrin to oleanolic acid, (iii) the conversion of oleanolic acid to echinocystic acid, (iv) the conversion of echinocystic acid to QA, (v) the conversion of QA to 3-O-{[β-D-glucopyranuronosyloxy]}-kirayaic acid (「QA-GlcA」), (vi) the conversion of 3-O-{[β-D-glucopyranuronosyloxy]}-kirayaic acid (「QA-GlcA」) to 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronosyloxy}-kirayaic acid (「QA-GlcA-Gal」), (vii) the conversion of 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronosyloxy}-kirayaic acid (「QA-GlcA-Gal」) to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronosyloxy}-kirayaic acid (「QA-GlcA-Gal-Xyl or 「QA-Tri」), (viii) the conversion of QA-Tri to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronosyloxy}-28-O-{β-D-fucopyranosyl ester}-kirayaic acid (QA-TriF), (ix) the conversion of QA-TriF to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronosyloxy}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFR), and (x) the conversion of QA-TriFR to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronosyloxy}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid (QA-TriFRX);(xi) Conversion of QA-TriFRX to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-glucopyranosyl ester}-oleanolic acid 3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRXX); (xii) Conversion of QA-TriFRXX to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl ester}-oleanolic acid α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX) and / or (xii) 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)4)-α-L-rhamnopyranosyl-(1→2)-2)-[β-D-4-O-acetylquinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (SpB).;

[0152] As described above, the triterpenoid biosynthetic genes described herein can also be engineered in plants. Suitable techniques are available in the art (see, for example, WO2019 / 122259). Since 2,3-oxidosqualene is ubiquitous in higher plants due to its role in sterol biosynthesis, the biosynthesis described herein has broad applicability in plant hosts. Suitable plant hosts include any plant that is amenable to transformation with Agrobacterium. As discussed herein, additional activities can be employed when practicing the methods described herein in microorganisms.

[0153] Examples of suitable hosts include plants such as Nicotiana benthamiana and microorganisms such as yeast. These will be described in detail below.

[0154] The present invention may comprise transforming a host as described above with a heterologous nucleic acid by introducing a biosynthetic nucleic acid into a host cell via a vector, causing or allowing recombination between the vector and the host cell genome to introduce the nucleic acid according to the present invention into the genome.

[0155] In another aspect of the present invention, there is provided a host cell transformed with a heterologous nucleic acid comprising each of a plurality of triterpenoid biosynthetic nucleotide sequences encoding a polypeptide having the biosynthetic activities described herein, wherein the expression of said nucleic acid confers on the transformed host the ability to perform biosynthesis or improves said ability in the host.

[0156] The present invention further encompasses host cells transformed with a triterpenoid biosynthetic nucleic acid or a vector as described above (e.g., comprising a biosynthetically modified nucleotide sequence), particularly plant or microbial cells. In a transgenic host cell (i.e., transgenic for the nucleic acid in question), the transgene is on an extrachromosomal vector or is integrated into the genome, preferably stably. There may be one or more heterologous nucleotide sequences per haploid genome.

[0157] The methods and materials described herein can be used, inter alia, to generate stable crop-plants that accumulate biosynthetic triterpenoid saponins or other products. Examples of plants include row crops such as sunflower, potato, canola, kidney bean, field bean, flax, safflower, buckwheat, cotton, corn, soybean, sugar beet, etc. Major crop plants such as corn, wheat, rapeseed and rice can also be preferred hosts.

[0158] There is also provided a plant comprising a plant cell according to the present invention.

[0159] Also provided are methods that include introducing such constructs into plant cells or microbial (e.g., bacterial, yeast, or fungal) cells and / or inducing expression of the construct in plant cells by applying a suitable stimulus, such as an effective exogenous inducer.

[0160] As an alternative to the microorganism, cell suspensions of artificial glycoprotein QA-producing plant species including the moss Physcomitrella patens may be cultured in a fermentation tank (see, for example, Grotewold et al. (Engineering Secondary Metabolites in Maize Cells by Ectopic Expression of Transcription Factors, Plant Cell, 10, 721-740, 1998)).

[0161] Also provided are host cells, particularly plant or microbial cells, containing the above-described heterologous constructs.

[0162] The above discussion of the host cells related to the reconstitution of QA or glycosylated QA biosynthesis in heterologous organisms is also incorporated herein by reference.

[0163] Also provided is a method for transforming a plant cell that includes introducing the above-described construct into the plant cell and causing or allowing recombination between the vector and the plant cell genome to introduce the nucleic acid described herein into the genome.

[0164] The present invention further encompasses host cells, particularly plant or microbial cells, transformed with the nucleic acids or vectors described herein (e.g., containing triterpenoid biosynthetic nucleotide sequences). In transgenic plant cells (i.e., transgenic with respect to the nucleic acid in question), the transgene is on an extrachromosomal vector or is preferably stably integrated into the genome. There may be one or more heterologous nucleotide sequences per haploid genome.

[0165] Yeasts are widely used as triterpene-producing hosts and may thus be well adapted to the biosynthesis of QA and, subsequently, glycosylated QA as described herein, for example, the biosynthesis of triterpenoid saponins.

[0166] In some preferred embodiments, the host is a yeast. In the case of such a host, it may be desirable to introduce additional genes to improve the flux of QA and thus the production of QA or glycosylated QA as described above. Examples include one or more plant cytochrome P450 reductases (CPRs) that function as the redox partner of the introduced P450, similar to HMGR. Similarly, it may be desirable to introduce additional genes that contribute to other elements of QA or improve the QA glycosylation pathway. These include enzymes that provide UDP-sugar donors (see, for example, Ohashi T, Hasegawa Y, Misaki R, Fujiyama K (2016) “Substrate preference of citrus naringenin rhamnosyl transferases and their application to flavonoid glycoside production in fission yeast” (2016). Applied Microbiology and Biotechnology. 100(2): 687-696.); Oka T, Jigami Y. (2006). “Reconstruction of de novo pathway for synthesis of UDP-glucuronic acid and UDP-xylose from intrinsic UDP-glucose in Saccharomyces cerevisiae”. FEBS J.273(12):2645-57). In light of the present disclosure, those skilled in the art can provide such auxiliary activities as needed.

[0167] Also provided are plants containing the plant cells transformed as described above.

[0168] Optionally, after transformation of the plant cells, plants can be regenerated from, for example, single cells, callus tissue or leaf discs, as is standard in the art. Almost all plants can be completely regenerated from plant cells, tissues and organs. Available techniques are outlined in Vasil et al., Cell Culture and Somatic Cell Genetics of Plants, Vol I, II and III, Laboratory Procedures and Their Applications, Academic Press, 1984, and Weissbach and Weissbach, Methods for Plant Molecular Biology, Academic Press, 1989.

[0169] In addition to the regenerated plants, the following are also provided: clones of such plants, seeds, progeny and descendants of selfed or hybrid plants (e.g., F1 progeny and F2 progeny). Also provided are plant propagules from such plants, i.e., any part that can be used for sexual or asexual reproduction or propagation, including cuttings, seeds, etc. In all cases, these plants or parts contain, for example, the above-described plant cells or heterologous biosynthetically modified nucleic acids introduced into the progenitor plant.

[0170] Also provided are any parts of these plants (e.g., leaves, stems, dried or ground products, edible parts, etc.), which in all cases contain the above-described plant cells or heterologous triterpenoid biosynthetic DNA.

[0171] The invention also encompasses any expression product of the disclosed coded triterpenoid biosynthetic nucleic acid sequences, and thus a method for producing an expression product by expression from the coded nucleic acid under suitable conditions, which may also be in a suitable host cell.

[0172] As described below, such a plant background may be natural or transgenic for one or more other genes related to the biosynthesis of triterpenoids such as QA or glycosylated QA, or alternatively, may affect its phenotype or trait.

[0173] In modifying the host phenotype, the triterpenoid biosynthesis nucleic acids described herein can be used in combination with any other genes such as genes introduced that affect the biosynthesis rate or yield of triterpenoids such as QA or glycosylated QA, or their modification, or any other phenotypic trait or desirable property.

[0174] By gene combination, plants or microorganisms (e.g., bacteria, yeast or fungi) can be engineered to enhance the production of desirable precursors or reduce unwanted metabolism.

[0175] The triterpenoid biosynthesis sequences described herein can be used in vitro or in vivo to catalyze their respective biological activities.

[0176] For example, a method for converting 2,3-oxidosqualene (OS) to β-amyrin may consist of contacting OS with Saponaria officinalis synthase (SobAS) comprising the amino acid sequence of SEQ ID NO: 8 or a variant thereof such that the OS is converted to β-amyrin. Also provided is the use of Saponaria officinalis synthase (SobAS) comprising the amino acid sequence of SEQ ID NO: 8 or a variant thereof for converting OS to β-amyrin.

[0177] The method for oxidizing β-amyrin to carboxylic acid at the C28 position may comprise contacting β-amyrin with a SoC28 oxidase polypeptide comprising the amino acid sequence of SEQ ID NO: 2 or a variant thereof such that the C28 position of the β-amyrin is oxidized to carboxylic acid to produce oleanolic acid. Also provided is the use of a SoC28 oxidase polypeptide comprising the amino acid sequence of SEQ ID NO: 2 or a variant thereof for oxidizing the C28 position of β-amyrin to carboxylic acid.

[0178] The method for oxidizing oleanolic acid to alcohol at the C16 position to produce echinocystic acid may include contacting oleanolic acid with a SoC28C16 oxidase polypeptide comprising the amino acid sequence of SEQ ID NO: 4 or a variant thereof such that the C16 position of the oleanolic acid is oxidized to alcohol, thereby producing echinocystic acid. Also provided is the use of a SoC28C16 oxidase polypeptide comprising the amino acid sequence of SEQ ID NO: 4 or a variant thereof for oxidizing the C16 position of oleanolic acid to alcohol to produce echinocystic acid.

[0179] The method for oxidizing β-amyrin to alcohol at the C16 position and to carboxylic acid at the C28 position to produce echinocystic acid includes contacting β-amyrin with a SoC28C16 oxidase polypeptide comprising the amino acid sequence of SEQ ID NO: 4 or a variant thereof, as a result of which the C16 position of the β-amyrin is oxidized to alcohol and the C28 position is oxidized to carboxylic acid, thereby producing echinocystic acid. Also provided is a method of using a SoC28C16 oxidase polypeptide comprising the amino acid sequence of SEQ ID NO: 4 or a variant thereof to oxidize the C28 and C16 positions of β-amyrin to produce echinocystic acid.

[0180] A method for oxidizing echinocystic acid to an alcohol at the C-23 position to produce quillaic acid (QA) includes contacting echinocystic acid with a SoC23 oxidase polypeptide comprising the amino acid sequence of SEQ ID NO: 6 or a variant thereof to oxidize the C-23 position of the echinocystic acid to an aldehyde, thereby producing quillaic acid (QA). Also provided is the use of a SoC23 oxidase polypeptide comprising the amino acid sequence of SEQ ID NO: 6 or a variant thereof for oxidizing the C23 position of β-amyrin or an oxidized derivative thereof to produce quillaic acid (QA).

[0181] A method for converting quillaic acid (QA) to 3-O-{β-D-glucopyranuronosyluronic acid}-quillaic acid (“QA-GlcA”) may consist of contacting QA with a Saponaria officinalis QA 3-O glucuronosyltransferase (“SoCSL”) polypeptide comprising the amino acid sequence of SEQ ID NO: 10 or a variant thereof such that the QA is converted to QA-GlcA. Also provided is the use of a SoCSL polypeptide comprising the amino acid sequence of SEQ ID NO: 10 or a variant thereof for converting QA to QA-GlcA.

[0182] A method for converting 3-O-{β-D-glucopyranuronosyluronic acid}-quillaic acid (“QA-GlcA”) to 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronosyluronic acid}-quillaic acid (“QA-GlcA-Gal”) may consist of: contacting QA-GlcA with a Saponaria officinalis QA-GlcA galactosyltransferase (“SoC3Gal”) polypeptide comprising the amino acid sequence of SEQ ID NO: 12 or a variant thereof such that the QA-GlcA is converted to QA-GlcA-Gal. Also provided is the use of a Saponaria officinalis QA-GlcA galactosyltransferase (“SoC3Gal”) polypeptide comprising the amino acid sequence of SEQ ID NO: 12 or a variant thereof for converting QA-GlcA to QA-GlcA-Gal.

[0183] A method for converting 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-cichoric acid (「QA-GlcA-Gal」) to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-cichoric acid (「QA-GlcA-[Gal]-Xyl」 or 「QA-Tri」) may include contacting QA-GlcA with a Saponaria officinalis QA-GlcA-Gal xylosyltransferase (「SoC3Xyl」) polypeptide comprising the amino acid sequence of SEQ ID NO: 14 or a variant thereof, and converting the said QA-GlcA-Gal to QA-Tri. Also provided is the use of a Saponaria officinalis QA-GlcA-Gal xylosyltransferase (「SoC3Xyl」) polypeptide comprising the amino acid sequence of SEQ ID NO: 14 or a variant thereof to convert QA-GlcA-Gal to QA-Tri.

[0184] Converting QA-Tri to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosyl ester}-cichoric acid (QA-TriF) may consist of contacting QA-Tri with a Saponaria officinalis QA-Tri fucosyltransferase (「SoC28Fu」) polypeptide comprising the amino acid sequence of SEQ ID NO: 16 or a variant thereof, and the said QA-Tri is converted to QA-TriF. Also provided is the use of a Saponaria officinalis QA-Tri fucosyltransferase (「SoC28Fu」) polypeptide comprising the amino acid sequence of SEQ ID NO: 16 or a variant thereof for converting QA-Tri to QA-TriF.

[0185] The method for converting QA-TriF to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFR) may include contacting QA-TriF with a Saponaria officinalis QA-TriF rhamnosyltransferase (「SoC28Rha」) polypeptide comprising the amino acid sequence of SEQ ID NO: 18 or a variant thereof to convert QA-TriF to QA-TriFR. Also provided is the use of a Saponaria officinalis QA-TriF rhamnosyltransferase (「SoC28Rha」) polypeptide comprising the amino acid sequence of SEQ ID NO: 18 or a variant thereof to convert QA-TriF to QA-TriFR.

[0186] The method for converting QA-TriFR to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRX) may include contacting QA-TriFR with a Saponaria officinalis QA-TriFR xylosyltransferase (「SoC28Xyl1」) polypeptide comprising the amino acid sequence of SEQ ID NO: 20 or a variant thereof to convert QA-TriFR to QA-TriFRX. Also provided is converting QA-TriFR to QA-TriFRX using a Saponaria officinalis QA-TriFR xylosyltransferase (「SoC28Xyl1」) polypeptide comprising the amino acid sequence of SEQ ID NO: 20 or a variant thereof.

[0187] The method of converting QA-TriFRX to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRXX) may include contacting QA-TriFRX with a Saponaria officinalis QA-TriFRX xylosyltransferase ("SoC28Xyl2") polypeptide comprising the amino acid sequence of SEQ ID NO: 22 or a variant thereof, thereby converting QA-TriFRX to QA-TriFRXX. Also provided is the use of a Saponaria officinalis QA-TriFRX xylosyltransferase ("SoC28Xyl2") comprising the amino acid sequence of SEQ ID NO: 22 or a variant thereof to convert QA-TriFRX to QA-TriFRXX.

[0188] The method of converting QA-TriFRXX to 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX) includes contacting QA-TriFRXX with a Saponaria officinalis QA-TriFRXX quinovosyltransferase ("SoGH1") polypeptide comprising the amino acid sequence of SEQ ID NO: 34 or a variant thereof, whereby QA-TriFRXX is converted to QA-TriF(Q)RXX. Also provided is the use of a Saponaria officinalis QA-TriFRXX quinovosyltransferase ("SoC28Xyl2") comprising the amino acid sequence of SEQ ID NO: 34 or a variant thereof to convert QA-TriFRXX to QA-TriF(Q)RXX.

[0189] A method for converting QA-TriF(Q)RXX into 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-2)-[β-D-4-O-acetylquinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q-Ac)RXX) may involve contacting QA-TriF(Q)RXX with a Saponaria officinalis QA-TriF(Q)RXX acetyltransferase ("SoBAHD1") polypeptide comprising the amino acid sequence of SEQ ID NO: 36 or a variant thereof, wherein the QA-TriF(Q)RXX is converted to QA-TriF(Q-Ac)RXX. Also provided is the use of a Saponaria officinalis QA-TriF(Q)RXX acetyltransferase ("SoBAHD1") comprising the amino acid sequence of SEQ ID NO: 36 or a variant thereof to convert QA-TriF(Q)RXX to QA-TriF(Q-Ac)RXX(SpB).

[0190] In some embodiments, one or more of the above nucleic acids or proteins can be used for the heterologous reconstitution of biosynthetic pathways. The biosynthetic pathways are described above and can include one or more of the conversion from OS to QA, from QA to QA-Tri, from QA-Tri to QA-TriFRXX, and from QA-TriFRXX to QA-TriF(Q-Ac)RXX.

[0191] Furthermore, a method of affecting or influencing biosynthesis in a host such as a plant, the method comprising causing or enabling the transcription of a heterologous triterpenoid biosynthetic nucleic acid as discussed above in a cell of the plant is provided. This step may precede the step of introducing the nucleic acid into a cell of the plant or its ancestor. Biosynthesis may include the production of glycosylated QAs such as QA; QA-Tri, QA-TriFRXX or QA-TriF(Q-Ac)RXX; or an intermediate of any of these.

[0192] Such methods are usually part of a method for producing glycosylated QA (e.g., QA-Tri, QA-TriFRXX, or QA-TriF(Q-Ac)RXX) in a host such as a plant, and in some cases form a single step. Preferably, this method uses a triterpenoid biosynthetic polypeptide as described above or a variant thereof, or a nucleic acid encoding any of them.

[0193] The above method may be used to generate QA or glycosylated QA such as QA-Tri, QA-TriFRXX, QA-TriF(Q-Ac)RXX in a heterologous host, or may be used to generate an intermediate. Glycosylated QA generally does not naturally occur in the introduced biological species.

[0194] Triterpenoids containing glycosylated forms of QA such as QA-Tri, QA-TriFRXX, or QA-TriF(Q-Ac)RXX can be isolated from the plants or methods described herein and can be commercially utilized.

[0195] The above method may be part of a method for producing downstream products such as QS-21 in a host, and in some cases may form a single step. This method involves culturing the host (if it is a microorganism) or growing the host (if it is a plant), and then harvesting it, and purifying triterpenoids such as glycosylated QA such as QA-Tri, QA-TriFRXX, or QA-TriF(Q-Ac)RXX, or downstream products or derivatives (e.g., QS-21) therefrom. The products thus produced form a further aspect of the present invention. The usefulness of QS-21 is described above.

[0196] Alternatively, glycosylated QA such as QA-Tri, QA-TriFRXX, or QA-TriF(Q-Ac)RXX may be recovered to enable further chemical synthesis of downstream compounds.

[0197] The methods described herein encompass both in vitro and in vivo production, or manipulation, of triterpenoids such as QA and / or one or more glycosylated QAs. For example, triterpenoid biosynthetic polypeptides can be employed in fermentation via expression in microorganisms such as Escherichia coli, yeast, and filamentous fungi. In some embodiments, one or more newly characterized triterpenoid biosynthetic sequences described herein can be used in these organisms in combination with one or more other biosynthetic genes.

[0198] In vivo methods are described extensively above and generally involve causing transcription, and then translation, of a recombinant nucleic acid molecule encoding a triterpenoid biosynthetic polypeptide.

[0199] In other embodiments, triterpenoid biosynthetic polypeptides (enzymes) can be used in vitro, for example, in isolated, purified, or semi-purified form. Optionally, they may be the product of expression of a recombinant nucleic acid molecule.

[0200] Down-regulation of genes in a host may be desirable to reduce unwanted metabolism and flux that can affect the yield of triterpenoids such as QA and glycosylated QA. Such down-regulation can be achieved by methods known in the art, for example, using antisense technology.

[0201] When using an antisense gene or a partial gene sequence to down-regulate gene expression, the nucleotide sequence is placed under the control of a promoter in the "reverse" direction such that transcription results in an RNA that is complementary to the normal mRNA transcribed from the "sense" strand of the target gene. See, for example, Rothstein et al, 1987; Smith et al, (1988) Nature 334, 724-726; Zhang et al, (1992) The Plant Cell 4, 1575-1588, English et al., (1996) The Plant Cell 8, 179-188. The antisense technique is also reviewed in Bourque, (1995), Plant Science 105, 125-149, and Flavell, (1994) PNAS USA 91, 3490-3496.

[0202] As an alternative to antisense, there is a method of inserting all or part of a copy of the target gene in the sense, i.e., the same orientation as the target gene, to reduce expression of the target gene by co-suppression. See, for example, van der Krol et al., (1990) The Plant Cell 2, 291-299; Napoli et al., (1990) The Plant Cell 2, 279-289; Zhang et al., (1992) The Plant Cell 4, 1575-1588, and US-A-5,231,020. Further improvements in gene silencing or co-suppression techniques can be found in WO95 / 34668 (Biosource); Angell & Baulcombe (1997) The EMBO Journal 16,12:3675-3684; and Voinnet & Baulcombe (1997) Nature 389: pg 553.

[0203] Double-stranded RNA (dsRNA) has been found to be even more effective in gene silencing than either the sense or antisense strand alone (Fire A. et al Nature, Vol 391, (1998)). Silencing via dsRNA is gene-specific and is often referred to as RNA interference (RNAi) (see also Fire (1999) Trends Genet. 15: 358-363, Sharp (2001) Genes Dev. 15: 485-490, Hammond et al. (2001) Nature Rev. Genes 2: 1110-1119 and Tuschl (2001) Chem. Biochem. 2: 239-245).

[0204] RNA interference occurs in a two-step process. First, dsRNA is cleaved intracellularly to produce short interfering RNAs (siRNAs) approximately 21-23 nucleotides in length with 5'-terminal phosphate and 3'-short overhangs (approximately 2 nt). These siRNAs specifically target and destroy the corresponding mRNA sequences (Zamore P.D. Nature Structural Biology, 8, 9, 746-750, (2001)).

[0205] Another methodology known in the art for down-regulation of target sequences is the use of "microRNA" (miRNA) as described, for example, in Schwab et al 2006, Plant Cell 18, 1121-1133. This technique uses artificial miRNAs that can be encoded by stem-loop precursors incorporating appropriate oligonucleotide sequences.

[0206] In some embodiments, a method for affecting or influencing the biosynthesis of QA or glycosylated QA in a host, comprising any of the following steps: (i) Causing or enabling transcription from a nucleic acid comprising a complementary sequence of the host nucleotide sequence described herein such that the activity of each encoded polypeptide is reduced by an antisense mechanism; (ii) Causing or enabling transcription from a nucleic acid encoding a stem-loop precursor, comprising 20-25 nucleotides of the host nucleotide sequence, optionally including one or more mismatches, such that the activity of each encoded polypeptide is reduced by the miRNA mechanism; (iii) Causing or enabling transcription from a nucleic acid encoding a double-stranded RNA corresponding to 20-25 nucleotides of the host nucleotide sequence (optionally including one or more mismatches) such that the activity of each encoded polypeptide is reduced by the siRNA mechanism.

[0207] Those skilled in the art will appreciate that, in light of the present disclosure, additional genes may be utilized in the practice of the present invention to provide additional activities and improve expression or activity. These include those that express cofactors or helper proteins, or other factors.

[0208] It will be understood that when these general terms are used in connection with any aspect or embodiment, their meaning or disclosure applies mutatis mutandis to any of these sequences individually.

[0209] Other aspects and embodiments of the present invention provide the above-described aspects and embodiments with the term "comprising" replaced by the term "consisting of", and the above-described aspects and embodiments with the term "comprising" replaced by the term "consisting essentially of".

[0210] It should be understood that this application discloses all combinations of any of the above-described aspects and embodiments with each other, unless the context requires otherwise. Similarly, this application discloses all combinations of preferred features and / or any features, alone or in combination with any of the other aspects, unless the context requires otherwise.

[0211] Changes to the above embodiments, further embodiments, and modifications thereof will be apparent to those skilled in the art upon reading this disclosure, and as such, these are within the scope of the present invention.

[0212] All documents and array database items referred to in this specification are hereby incorporated by reference in their entirety for all purposes.

[0213] As used herein, "and / or" is considered to specifically disclose each of the two specified features or components, regardless of the presence or absence of the other. For example, "A and / or B" is considered a specific disclosure of (i) A, (ii) B, and (iii) A and B, and is interpreted as if each were individually described herein.

[0214] Abbreviations QA-GlcA-[Gal]-xyl or QA-Tri-3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-kirayaic acid QA-GlcA-Gal or QA-Di-3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-kirayaic acid QA-GlcA or QA-Mono-3-O-{β-D-glucopyranuronate}-kirayaic acid QA-TriF-3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-fucopyranosyl ester}-kirayaic acid QA-TriFR-3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-kirayaic acid QA-TriFRX-3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid QA-TriFRXX-3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid QA-TriF(Q)RXX-3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid QA-TriF(Q-Ac)RXX or SpB or saponarioside B-3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-2)-[β-D-4-O-acetylquoride 3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-2)-[β-D-4-O-acetylquinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid QA-oleanolic acid OS-2,3-oxidosqualene Gal-D-galactopyranose GlcA-D-glucopyranuronic acid (the additional number indicates a specific carbon) Xyl-D-xylopyranose Rha-L-rhamnopyranose Ac-acetyl group Qui-D-Quinovose (or Q) SobAS or S officinalis β -amylin synthase of SobAS1 SoC28 oxidase or SoC28 or CYP716A378 - S officinalis kiraya acid C28 oxidase SoC16 oxidase or SoC28C16 oxidase or SoC28C16 or CYP716A379 - S officinalis kiraya acid C28 and C16 oxidase SoC23 oxidase or SoC23 or CYP72A984 - S officinalis kiraya acid C23 oxidase SoQA - GlcAT or SoCSL or SoCSL1 - S officinalis QA 3 - O glucuronosyltransferase SoQA - GalT or SoC3Gal or UGT73DL1 - S officinalis QA - GlcA galactose transferase SoQA - XylT or SoC3Xyl or UGT3CC6 - S officinalis QA - GlcA - Gal xylosyltransferase SoQA - TriFuT or SoC28Fu or UGT74CD1 - S officinalis QA - Tri fucosyltransferase SoFuSyn or SoSDR - S officinalis short - chain dehydrogenase SoQA - TriFRhaT or SoC28Rha or UGT79T1 - S officinalis QA - TriF rhamnosyltransferase SoQA - TriFRXylT or SoC28Xyl1 or UGT79L3 - S officinalis QA - TriFR xylosyltransferase SoQA - TriFRXXylT or SoC28Xyl2 or UGT73M2 - S officinalis QA - TriFRX xylosyltransferase SoGH1 - S officinalis QA - TriFRXX quinovosyltransferase SoBAHD1 - S. officinalis QA - TriF(Q)RXX Acetyltransferase tHMGR - Avena strigosa (diploid oat) truncated 3 - hydroxy, 3 - methylbutyryl - CoA reductase

[0215] Experiment Materials and Methods RNA Synthesis and RNA-seq Analysis [MacKenzie et al (1997) Plant Disease. 81:222 - 226] protocol was modified, and total RNA was extracted from leaves and roots of representative soapwort plants using the RNeasy Plant Mini kit (Qiagen). Simultaneously with RNA extraction, on - column DNase digestion was performed using RQ1 RNase - Free DNase (Promega). cDNA used for the amplification of putative soapwort bAS was prepared from 0.8 μg of DNase - treated RNA according to the manufacturer's instructions using GoScript™ Reverse Transcriptase (Promega).

[0216] A total of 24 RNA samples were sent to the Earlham Institute (EI) for transcriptome sequencing and RNA - seq analysis. NEBNext Ultra II Directional RNA - Seq libraries were constructed from the 24 samples and sequenced on 2 lanes of a NovaSeq 6000 SP flow cell (150 - paired - end reads). Transcriptome assembly was performed by EI using the Trinity de novo assembler (ver.2.8.5), and ORF prediction and functional annotation were assigned using TransDecoder (ver.5.5.0) and Human Readable Descriptions (ver.AHRD, ver.3.3.3), respectively. Also, quantification of transcripts was performed by EI using salmon (ver.0.14.1).

[0217] Identification of Candidate Genes To identify bAS candidates in soapwort, the transcriptome of S. officinalis was obtained from the 1,000 Plants (1KP) project (www.onekp.com) [Wicket et al (2014) PNAS 45 E4859-4868]. BLASTP searches were performed against the translated S. officinalis protein database using OSCs previously characterized from other plant species listed in Table 1 as queries. The list of soapwort candidates was filtered by removing sequences less than 500 amino acids (aa) in length. This list was further filtered by performing phylogenetic analysis with MEGA-X (http: / / www.megasoftware.net). Amino acid alignments were performed using the MUSCLE algorithm (https: / / www.ebi.ac.uk / Tools / msa / muscle / ) between the putative soapwort genes shown in Table 1 and the OSCs published from other plants. Using this alignment, a phylogenetic tree was created using the neighbor-joining algorithm (Poisson model) with 1,000 bootstrap replicates. Based on the phylogenetic analysis, candidates with low likelihood of being bAS were removed from the list.

[0218] After identifying SobAS, all other pathway candidates were identified using the newly assembled S. officinalis transcriptome by EI. Preliminary lists of soapwort CYP450, CSL, and UGT candidates were created by performing BLASTP searches against the new soapwort transcriptome using the literature gene families as queries. The lists were filtered by excluding candidates less than 500 aa in length. To further narrow down the lists, correlation analysis was performed to find candidates with similar expression patterns to SobAS. All bioinformatics analyses were performed in R. The quantification results of transcripts from Salmon were read using tximport (ver. 1.18.0). rlog library-normalized reads were generated using DESeq2 (ver. 1.30.1) and used for hierarchical clustering. Pearson's method was used for correlation analysis.

[0219] cDNA Synthesis and Gateway (registered trademark) Cloning For the cloning of candidate subtilisin-like genes, cDNA pools were prepared from leaf and root RNAs. First-strand cDNA synthesis was performed using the GoScript Reverse transcription system (Promega) according to the manufacturer's protocol.

[0220] The coding sequences of candidate subtilisin-like genes except SoGH1 were PCR-amplified from the cDNA pool using gene-specific primers with 5’ AttB sites. The coding sequence of SoGH1 with 5’ AttB sites was synthesized by IDT. PCR products were purified using the QIAquick PCR Purification kit according to the manufacturer's protocol. The purified PCR products were introduced into an entry vector using Gateway (registered trademark) technology (Invitrogen) and finally into an expression vector. Briefly, the pDONR207 vector and the purified PCR products were used in equimolar amounts (~150 ng each), and a BP recombination reaction was performed according to the manufacturer's instructions. Subsequently, chemically competent Escherichia coli cells (DH5α, ThermoFisher Scientific) were heat-shock transformed. Plasmids were recovered by plasmid preparation using the QIAprep Spin Miniprep Kit (Qiagen), and the sequences were confirmed. To generate expression clones, an entry vector containing the gene of interest and the pEAQ-HT-DEST1 expression vector were used in equimolar amounts (~150 ng each), and an LR recombination reaction was performed according to the manufacturer's protocol [Sainsbury et al (2009) Plant Biotechnol J 7(7): 682-693]. Plasmids were recovered again using the QIAprep Spin Miniprep Kit (Qiagen) according to the manufacturer's protocol.

[0221] Transient Expression of Candidate Genes in N. benthamiana For transient expression of candidate genes in Nicotiana benthamiana, the Agrobacteria tumefaciens strain LBA4404 (Invitrogen) was used. Agroinfiltration, sample collection and preparation were carried out as previously described in [Reed et al (2017) Metabolic Engineering, 42, 185-193].

[0222] GC-MS Analysis GC-MS analysis was performed using an Agilent 7890B equipped with a Zebron AB5-HT Inferno column (Phenomenex) and a 20-minute method program developed by Dr. James Reed (Osbourn laboratory). Briefly, 1 μL of each sample was injected into the inlet (250 °C) in pulse splitless mode (pulse pressure 30 psi). The oven temperature was held at 170 °C for 2 minutes, then raised to 300 °C at a rate of 20 °C / min and held at 300 °C for 11.5 minutes. Mass spectrometry was performed using an Agilent 5977A Mass Selector Detector in scan mode from 60 - 800 m / z after an 8-minute solvent delay. MassHunter Workstation (Agilent) was used for analysis of the obtained data.

[0223] LC-MS Analysis LC-MS analysis was performed using a Shimadzu Prominence HPLC system equipped with an IT-TOF mass spectrometer (manufactured by Shimadzu Corporation). Aqueous formic acid solution (0.1% v / v) was used as solvent A, and acetonitrile was used as solvent B. The sample was analyzed using a Kinetex XB-C18 100A (50×2.1 mm, 2.6 μm; manufactured by Phenomenex) column at a flow rate of 0.5 mL / min, 40 °C, and an injection volume of 5 μL. The mass spectrometer was equipped with electrospray in negative ionization mode (capillary temperature 250 °C, nebulizing gas 1.3 min / L, heat block temperature 300 °C, spray voltage -3.5 kV). The elution profile was as follows: 0 - 1 min, 5% B in A; 1 - 10 min, 55% B in A; 10 - 12 min, 100% B; 12 - 13 min, 100% B; 13 - 13.1 min, 5% B in A; 13.1 - 15.6 min, 5% B in A. LCMSolution software (Shimadzu Corporation) was used for data acquisition and processing. All authentic saponario-side pathway intermediate standards were provided by members of the Osbourn group.

[0224] Generation and Transformation of Hairy Roots of officinalis Seeds of S. officinalis were collected from plants grown in the glasshouse at JIC. After washing with sterile water, the seeds were held in sterile water for 3 - 4 h, surface - sterilized with sodium hypochlorite (5% w / v) for 30 min, and then washed three times with sterile water. Further, the seeds were washed with 70% ethanol (v / v) for 1 min and washed three times with sterile water. The seeds were germinated on MS medium (Murashige and Skoog 1962) (pH 5.88). Four weeks later, sub - cultures of the plantlets were carried out and maintained on MS medium (pH 5.88), 3% sucrose, 0.8% agar, 16 - h photoperiod, and 25 °C. Induction of hairy roots was carried out with ATCC15834, which was efficient (100% induction) among other test strains (A4, A4RS, LBA1334). Briefly, each bacterial suspension (100 μM acetosyringone, 1% sucrose, OD: 0.6 in MS) was injected into ~5 pieces per leaf excision using a syringe needle. The infected excised plants were cultured in the dark at 25 °C for 4 days in a co - culture medium containing semi - solid (0.8% agar) MS medium supplemented with 3% sucrose and 100 μM acetosyringone. Further, they were transplanted onto semi - solid (0.8% agar) MS medium supplemented with 3% sucrose, 500 mg / l cefotaxime, and 50 mg / l kanamycin and cultured at 25 °C with a 16 - h photoperiod until the bacteria were removed and the desired hairy roots appeared.

[0225] Primers for silencing were designed from unique regions of S. officinalis β - amylase synthase (SoβAS) and cloned into pDONR207 (Gateway - compatible vector). Sub - cloning was carried out with pK7WGIGW - 2R, which enables silencing of the transgene via dsRNA. To over - express SoβAS, the full - length sequence was cloned into pK7WG2R using Gateway technology. Control hairy roots were grown using empty pK7WG2R (Zhao et al.). All constructs were transformed with ATCC15834, co - cultured with wounded leaves, and then the transgeneity of the hairy roots was evaluated by dsRED fluorescence and PCR. Three - week - old dsRED - expressing hairy roots grown on liquid B5 (containing vitamins and sucrose) medium in the dark were evaluated for metabolite analysis.

[0226] Result Identification and Characterization of SobAS Based on Phylogenetic Trees The first step of triterpenoid biosynthesis is predicted to be the production of β-amyrin, catalyzed by oxidosqualene cyclase (OSC), β-amyrin synthase (bAS). To identify candidates for bAS in soapwort, the translated transcriptome of S. officinalis available from the 1,000 Plants (1KP) project (www.onekp.com; [Wickett et al supra]) was searched, and reciprocal BLASTP searches were performed using OSCs previously characterized from other plant species as query sequences (Table 1). As a result of phylogenetic analysis, SobAS was identified as a candidate for bAS in soapwort.

[0227] To examine the activity of SobAS, we transiently expressed SobAS together with truncated HMG-CoA reductase (tHMGR) in Nicotiana benthamiana to increase the flux into the MVA pathway [Reed et al 2017, supra]. The complete open reading frames of SobAS and tHMGR were transformed into Agrobacterium tumefaciens and co-infiltrated into the leaves of N. benthamiana. Leaves were harvested 4 days after infiltration, metabolites were extracted, and analyzed using GC-MS. Transient expression of SobAS in N. benthamiana formed peak 1 at m / z 498, which corresponded to a commercially available β-amyrin standard in both retention time and mass spectrum (Figure 4a). Peak 1 was not present in leaves expressing only tHMGR as a negative control (Figure 4a). From these results, candidate SobAS was identified as an OSC that cyclizes oxidosqualene to β-amyrin (Figure 4b).

[0228] Identification of Saponarioside Pathway Genes by Co-expression Analysis Since the soapwort transcriptome published in the 1KP project does not contain organ-specific transcriptome data, RNA-seq analysis was performed on six different soapwort organs (flowers, flower buds, young leaves, old leaves, stems, roots) with different saponin contents. This new soapwort transcriptome was used for further gene identification instead of the transcriptome available from the 1KP project.

[0229] Following the biosynthesis of β-amyrin, the next step in saponarioside biosynthesis is predicted to be the oxidation of β-amyrin to kirenolic acid by three cytochrome P450 (CYP450). To create a list of candidate soapwort CYP450s, a BLASTP search was performed against the newly assembled soapwort transcriptome using the literature CYP450s from the TriForC database (http: / / bioinformatics.psb.ugent.be / triforc / , [Miettinen et al (2017) Nature Comms 8(1) 1-13]) as queries. This list was refined by removing candidates less than 500 aa in length. To further narrow down the candidate list, Pearson co-expression analysis was performed using the expression patterns of SobAS characterized above. Candidates with a Pearson correlation coefficient (PCC) of less than 0.80 were excluded from the candidate list (Table 3).

[0230] The next step in the saponarioside biosynthetic pathway is predicted to be the decoration of kielactinic acid by a family 1 UDP-dependent glycosyltransferase (UGT). To identify candidate UGTs in soapwort, a list of UGTs previously characterized from other plant species was obtained from [Louveau et al (2019) Cold Spring Harbor Perspectives in Biology, 11(12), a034744] and used as a BLASTP query against the S. officinalis transcriptome. The candidate list was further narrowed down as described above. Pearson co-expression analysis was performed using the expression profile of SobAS, and candidates with a Pearson correlation coefficient (PCC) value of less than 0.90 were excluded from the list (Table 4).

[0231] In addition to UGTs, recent discoveries by Jozwiak et al. and members of the Osbourn group have shown that cellulose synthase-like (CSL) genes have the ability to glucuronidate triterpenoid saponins (Jozwiak et al., 2020; WO / 2020 / 260475). Therefore, CSL candidates in the soapwort transcriptome were also searched for. A list of CSLs from other plant species was obtained from Reed et al. (in preparation) and used as a BLASTP query against the soapwort transcriptome. The list of candidate soapwort CSLs was further refined by performing Pearson co-expression analysis using the expression profile of SobAS. Soapwort CSL candidates with a PCC value of less than 0.85 were excluded from the list (Table 5). All of the identified putative saponarioside biosynthetic genes shared similar expression profiles along different soapwort organs, suggesting that they are involved in the same biosynthetic pathway (Figure 3).

[0232] The candidate list was further selected and refined based on the height of co-expression (PCC > 0.88) with the SobAS1 bait gene ranked using PCC, annotation, and the abundance of transcript absolute numbers in floral organs (Figure 3).

[0233] Characterization Analysis of Candidate Genes by Transient Expression in N. benthamiana The candidate genes for saponarioside biosynthesis identified above were transiently expressed in N. benthamiana, and their activities were examined. The open reading frames (ORFs) of the candidate genes were PCR-amplified using the primers shown in Table 2 or synthesized using the upstream 5’attb site to enable Gateway (registered trademark) cloning. The amplified or synthesized gene fragments were cloned into pDONR207 and introduced into the plant expression vector pEAQ-HT-DEST1 [Sainsbury et al 2009, supra]. The expression constructs were individually transformed into Agrobacterium tumefaciens (LBA4404) and transiently expressed in N. benthamiana. In all experiments, an A. tumefaciens strain carrying tHMGR was co-infiltrated to enhance triterpene production in N. benthamiana. By screening the activities of the top candidates in Tables 3-5 and Figure 3, SoC28, SoC28C16, SoC23, SoCSL, SoC3Gal, SoC3Xyl, SoC28Fu, SoC28Rha, SoC28Xyl1, SoC28Xyl2, SoGH1 and SoBAHD1 were identified.

[0234] The leaves of N. benthamiana were co-infiltrated with Agrobacterium tumefaciens strains carrying the ORF of (i) tHMGR SobAS+SoC28 or (ii) tHMGR+SobAS+SoC28C16, and the activities of SoC28 and SoC28C16 were examined. The leaves were harvested 4 days after infiltration, and metabolites were extracted and analyzed using GC-MS. Co-expression of SobAS with SoC28 in N. benthamiana resulted in the formation of peak 2 at m / z 585 (Figure 5b). Peak 2 was identified as oleanolic acid because the retention time (RT), m / z, and mass spectrum of peak 2 present in the N. benthamiana extract were consistent with those of peak 2 found in the commercially available oleanolic acid standard (Figure 5b). Interestingly, an additional metabolite peak (peak 3) at m / z 570 and oleanolic acid were also produced from the extract of N. benthamiana leaves co-infiltrated with SobAS and SoC28C16. Peak 3 was identified as echinocystic acid because the RT and mass spectrum of peak 3 detected from the N. benthamiana extract were consistent with those of peak 3 detected from the echinocystic acid standard. Neither peak 2 nor peak 3 was detected in the leaves of N. benthamiana expressing only tHMGR used as a negative control. From these results, SoC28 is a CYP450 with C28 oxidation activity and is likely to produce oleanolic acid from β-amyrin, and SoC28C16 is a CYP450 with oxidation activities for both C28 and C16 and is likely to produce both oleanolic acid and echinocystic acid.

[0235] The leaves of N. benthamiana were co-infiltrated with an Agrobacterium tumefaciens strain carrying the OFR of tHMGR+SobAS+SoC28C16+SoC23, and the activity of SoC23 was tested. The extract of the harvested leaves was analyzed using HPLC-MS in negative ionization mode. Expression of SoC23 resulted in [M-H] of kirayaic acid -Peak 4 with an m / z of 485.3 corresponding to it was generated (Figure 5b). Furthermore, the retention time and mass spectrum of Peak 4 were consistent with the peaks observed for the kirenol acid standard. In the negative control expressing only tHMGR, Peak 4 was not detected. From these results, SoC23 was considered to be a CYP450 having C-23 oxidation activity.

[0236] Following the biosynthesis of kirenol acid using the gene of S. officinalis, candidate SoCSL was co-expressed with the genes necessary for the production of kirenol acid (tHMGR + SobAS + SoC28C16 + SoC23). The extract of the harvested leaves was analyzed using HPLC-MS in the negative ionization mode. As a result of the HPLC-MS analysis, the addition of SoCSL generated Peak 5 with an m / z of 661.3, which is the predicted [M-H] of QA-Mono - (Figure 6). Peak 5 was not detected in the negative control expressing only tHMGR, and the RT and mass spectrum of Peak 5 were consistent with the Osbourn group's QA-Mono authentic standard (Figure 6). Also, the MS / MS fragmentation pattern of Peak 5 indicated that the main fragment ion was at m / z 485.33, which corresponds to the predicted [M-H] of kirenol acid - From the above results, Peak 5 was identified as QA-Mono, and SoCSL was identified as the CSL that glucuronidates kirenol acid.

[0237] Next, candidate SoC3Gal was co-expressed with the genes necessary for the production of kirenol acid (tHMGR + SobAS + SoC28C16 + SoC23) and the newly characterized SoCSL. Similar to the above, the harvested leaf extract was analyzed using HPLC-MS. As a negative control, a plant extract expressing only the genes that produce kirenol acid and SoCSL was used. The addition of SoC3Gal resulted in the [M-H] of QA-Di -A new peak at m / z 823.4 corresponding to - was generated (Figure 7). Furthermore, the retention time (RT) and mass spectrum of peak 6 generated by SoC3Gal were consistent with those of the peaks generated by the authentic QA-Di standard. Additionally, from the MS / MS fragmentation pattern, the major fragment ion of peak 6 was m / z 485.32, which

[0238] corresponds to [M-H] of kiyamycin, suggesting the fragmentation of the sugar chain from QA-Di (Figure 7). From these results, peak 6 in Figure 7 was presumed to be QA-Di, and SoC3Gal was presumed to be a galactosyltransferase derived from S. officinalis. - Next, the properties of the SoC3Xyl candidate were investigated. The genes required for QA-Di production (tHMGR + SobAS + SoC28C16 + SoC23 + SoCSL + SoC3Gal) were co-expressed in N. benthamiana together with the addition of SoC3Xyl. When the extract of the harvested leaves was analyzed by HPLC-MS, a new gene product with a predicted mass of m / z 955.4 corresponding to - [M-H] of QA-Tri was detected. In the negative control where only the genes required for production up to QA-Di were co-expressed, no peak occurred at the predicted m / z, but when SoC3Xyl was additionally expressed, a new peak at m / z 955.4 was observed (Figure 8). Peak 7 not only showed the same RT and mass spectrum as the authentic QA-Tri standard, but it was also revealed by MS / MS fragmentation that the major ions were m / z 823.42 [M-H-Xyl] - and m / z 485.33 [M-H-Xyl-Gal] (Figure 8). From this result, peak 7 in Figure 8 was presumed to be QA-Tri, and SoC3Xyl was presumed to be a candidate for xylosyltransferase.

[0239] Our next focus was to characterize the glycosyltransferase with the activity to transfer D-fucose to QA-Tri. In previous studies by the Osbourn group, two genes, QsC28Fu and QsFuSyn, involved in the addition of D-fucose in the QS-21 biosynthetic pathway were identified. QsC28Fu was found to have UDP-4-keto-6-deoxy-glucose transferase activity, and QsFuSyn was revealed to be a 4-keto-reductase (Reed, Orme, El-Demerdash et al., 2023). In the process of this discovery, SoFuSyn, which converts UPD-4-keto-6-deoxy-glucose to UPD-D-fucose, was also identified and characterized. The candidate gene SoC28Fu was identified by co-expression analysis with SobAS, and the activity of the candidate gene was tested by transient expression in N. benthamiana. A combination of genes required to produce QA-Tri (tHMGR + SobAS + SoC28C16 + SoC23 + SoCSL + SoC3Gal + SoC3Xyl), plus the candidate gene SoC28Fu and the previously characterized SoFuSyn, was co-expressed in N. benthamiana. After harvesting, the leaves were extracted and analyzed by HPLC-MS. Peak 8 (m / z 1101.5) was generated by the additional activity of SoC28Fu and matched the peak generated by the authentic QA-TriF standard in both the RT and mass spectra (Figure 9). This peak was not detected in the negative control without SoFuC28 (Figure 9). Furthermore, from the MS / MS fragmentation pattern of Peak 8, the major daughter ions were found to be m / z 955.4, which is the [M-H] of QA-Tri, and m / z 485.3, which is the [M-H] of kiyamycin acid (Figure 9). These results suggest that the candidate SoC28Fu may transfer the fucose moiety to QA-Tri together with SoFuSyn. - was m / z 955.4, and the [M-H] of kiyamycin acid - was m / z 485.3 (Figure 9). These results suggest that the candidate SoC28Fu may transfer the fucose moiety to QA-Tri together with SoFuSyn.

[0240] Next, the activity of candidate SoC28Rha was tested. The combination of genes required for the production of QA-TriF (tHMGR + SobAS + SoC28C16 + SoC23 + SoCSL + SoC3Gal + SoC3Xyl + SoFuSyn + SoC28Fu) with the addition of SoC28Rha was co-expressed in N. benthamiana. The harvested leaf extracts were analyzed in negative ionization mode using HPLC-MS. Leaf extracts expressing only the genes required for the production of QA-TriF were used as a negative control. The predicted [M-H] of QA-TriFR - , peak 9 with m / z 1247.5 was detected only in the leaf extracts overexpressing SoC28Rha (Figure 10). Furthermore, from the MS / MS fragmentation of peak 9, the major fragment ions were m / z 955.4 corresponding to [M-H] of QA-Tri - and m / z 485.3 corresponding to [M-H] of kiranoyl acid - , suggesting that they were fragments of the C28 sugar chain and then the C-3 sugar chain (Figure 10). From these results, peak 9 in Figure 10 was determined to be QA-TriFR, and SoC28Rha was presumed to be a rhamnose transferase.

[0241] The next two enzymes to be characterized were SoC28Xyl1 and SoC28Xyl2. To examine SoC28Xyl1, the genes required for the production of QA-TriFR (tHMGR + SobAS + SoC28C16 + SoC23 + SoCSL + SoC3Gal + SoC3Xyl + SoFuSyn + SoC28Fu + SoC28Rha) were co-expressed with candidate SoC28Xyl1 in N. benthamiana. Extracts from leaves expressing only the genes required for the production of QA-TriFR were used as a negative control. The [M-H] of QA-TriFR -Peak 10 with an expected m / z of 1379.6 was detected only in samples expressing the genes necessary to produce the substrate QA-TriFR and SoC28Rha (Figure 11). From MS / MS fragmentation, the major fragment ions of peak 10 were m / z 955.4 and m / z 485.3, suggesting that the C28 sugar chain was lost to produce QA-Tri, followed by the loss of the C3 sugar chain to produce quillaic acid (Figure 11). The activity of SoC28Xyl2 was determined in the same manner as SoC28Xyl1. The genes necessary for the production of QA-TriFRX (tHMGR + SobAS + SoC28C16 + SoC23 + SoCSL + SoC3Gal + SoC3Xyl + SoFuSyn + SoC28Fu + SoC28Rha + SoC28Xyl1) were co-expressed with candidate SoC28Xyl2 in N. benthamiana. Tobacco leaves expressing the genes necessary for the production of QA-TriFRX without adding the SoC28Xyl2 candidate were used as a negative control. As a result of HPLC-MS analysis, the generation of peak 11 with an expected [M-H] - of m / z 1511.6 was observed only in samples expressing SoC28Xyl2 (Figure 12). Furthermore, in the MS / MS analysis, the major fragment ions of peak 11 were m / z 1379.6 [M-H-X] of quillaic acid, m / z 955.4 [M-H-FRXX] - and m / z 485.3 [M-H] - . These results suggested that SoC28Xyl1 and SoC28Xyl2 are xylose transferases in S. officinalis.

[0242] So far, we have elucidated the genes and enzymes necessary for the biosynthesis of QA-TriFRXX (11). To complete the biosynthetic pathway to saponarioside B, it is necessary to elucidate the step responsible for the transfer of 4-O-acetylquinovose to 13. GTs related to the biosynthesis of plant natural products usually belong to Family 1 of the GT superfamily, but among the UGTs included in our main candidate list, none showed quinovosyltransferase activity towards 11. Therefore, we expanded the candidates by examining genes with high co-expression with SobAS1 and focused on a candidate from glycoside hydrolase family 1 (GH1) that showed high co-expression with SobAS1 (PCC = 0.971) (Figure 3). Transient expression via Agrobacterium was performed in N. benthamiana to examine the activity of SoGH1 towards 11. When SoGH1 was co-expressed with the biosynthesis genes of 11, two new products (12’ and 12’’) with the same mass ([M-H] - = 1657.7 m / z) but different RTs corresponding to the predicted masses of 11 and deoxyhexose were observed (Figure 13). To distinguish between these two products, tandem MS analysis of 12 and 12’ was performed, and the same fragmentation pattern was obtained. The main fragmentation ions were 1525.7 m / z (the [M-H] - ) of QA-TriFRXX and 955.4 m / z (the [M-H] -) and the loss of deoxyhexose was suggested, followed by the loss of the entire C-28 sugar chain and the occurrence of QA-Tri. Next, 12 and 12’ were compared with the standard substance of 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-kijic acid (12, hereinafter abbreviated as QA-TriF(Q)RXX). As a result, the fragmentation of both 12 and 12’ was consistent with QA-TriF(Q)RXX, but only 12 showed the same RT as the QA-TriF(Q)RXX standard. These results suggested the possibility that SoGH1 is involved in the transfer of D-quinovose to QA-TriFRXX. With the successful elucidation of the pathway up to 12, only the acetylation step remained to complete the biosynthetic pathway to SpB(13). The functions of the BAHD ATs in the main candidate list in Figure 3 were screened by transient expression in N. benthamiana leaves. LC-MS analysis of the obtained leaf extracts revealed that when SoBAHD1 was co-expressed in combination with the gene set that produces 12, two new products (13 and 13’) with the predicted mass corresponding to SpB were formed ([M-H] - =1699.7 m / z). Furthermore, tandem MS analysis revealed the same fragmentation pattern for both 13 and 13’. The major fragment ions were 1657.7 m / z ([M-H] - 12) and 955.5 m / z ([M] -7), suggesting that the entire C-28 sugar chain disappeared following the fragmentation of the acetyl group (Figure 14). However, only 13 produced by the heterologous expression of SoBAHD1 matched the authentic SpB standard in both RT and fragmentation pattern. From these results, 13 was identified as SpB(13) produced by SoBAHD1 acetylating D-quinovose in 12, and SoBAHD1 was identified as an acetyltransferase with the ability to introduce an acetyl group into QA-TriF(Q)RXX to produce SpB.

[0243] The sequence similarity between the saponarioside biosynthetic gene identified here and the corresponding gene of Q. saponaria involved in QS-21 biosynthesis was compared using the amino acid sequence (Table 6). The first several genes showed high similarity in the amino acid sequence, while the remaining pathway genes showed overall low sequence similarity. This indicates that the two pathways were likely established independently, suggesting evidence of convergent evolution. The biosynthetic pathway of saponarioside discussed here is shown in Figure 15. However, the actual biosynthesis order can occur in any order in the plant body.

[0244] To investigate the role of characteristic genes in plants, hairy roots were successfully generated from soapwort seedlings. As a proof of concept, the expression of SobAS1 was silenced in soapwort hairy roots, and the metabolic profiles of SobAS1-silenced hairy roots were compared with those of DsRED-expressing control hairy roots. β-Amyrin was not detected in either the control or silenced hairy roots (Figure 16), but cycloartenol accumulated only in the SobAS1-silenced lines (Figure 17). These results suggest that silencing SobAS1 in soapwort hairy roots increases the flux through the sterol biosynthetic pathway. Furthermore, LC / MS analysis revealed that hairy roots with silenced SobAS1 did not accumulate kirayaic acid, whereas a large amount of kirayaic acid was detected in the control hairy roots (Figure 18). Consistent with this result, SpB was not detected in the hairy roots with silenced SobAS1, while SpB was detected in the control hairy roots (Figure 19). These results indicate that SobAS1 is the OSC responsible for β-amyrin biosynthesis in S. officinalis.

[0245] sequence SEQ ID NO:1 SoC28 Oxidase (SoC28) Nucleotide Sequence

[0246] MELFFICGLVLFSTLSLISLFLLHNHSSARGYRLPPGRMGWPFIGESYEFLANGWKGYPEKFIFSRLAKYKPNQVFKTSILGEKVAVMCGATCNKFLFSNEGKLVNAWWPNSVNKIFPSSTQTSSKEEAKKMRKLLPTFFKPEALQRYIPIMDEIAIRHMEDEWEGKSKIEVFPLAKRYTFWLACRLFLSIDDPVHVAKFADPFNDIASGIISIPIDLPGTPFNRGIKASNVVRQELKTIIKQRKLDLSDNKASPTQDILSHMLLTPDEDGRYMNELDIADKILGLLIGGHDTASAACTFVVKFLAELPHIYDGVYKEQMEIAKSKKEGERLNWEDIQKMKYSWNVACEVMRLAPPLQGAFREALSDFMYAGFQIPKGWKLYWSANSTHRNPECFPEPEKFDPARFDGSGPAPYTYVPFGGGPRMCPGKEYARLEILVFMHNIVKRFKWEKLIPDETIVVNPMPTPAKGLPVRLRPHSKPVTVSA* SEQ ID NO:2 SoC28 Oxidase (SoC28) Amino Acid Sequence

[0247] Sequence number 3 C28C16 oxidase (SoC28C16) nucleotide sequence

[0248] MELITLLSALLVLAIVSLSTFFVLYYNTPTKDGKTLPPGRMGWPFIGESYDFFAAGWKGKPESFIFDRLKKFAKGNLNGQFRTSLFGNKSIVVAGAAANKLLFSNEKKLVTMWWPPSIDKAFPSTAQLSANEEALLMRKFFPSFLIRREALQRYIPIMDDCTRRHFATGAWGPSDKIEAFNVTQDYTFWVACRVFMSIDAQEDPETVDSLFRHFNVLKAGIYSMHIDLPWTNFHHAMKASHAIRSAVEQIAKKRRAELAEGKAFPTQDMLSYMLETPITSAEDSKDGKAKYLNDADIGTKILGLLVGGHDTSSTVIAFFFKFMAENPHVYEAIYKEQMEVAATKAPGELLNWDDLQKMKYSWCAICEVMRLTPPVQGAFRQAITDFTHNGYLIPKGWKIYWSTHSTHRNPEIFPQPEKFDPTRFEGNGPPAFSFVPFGGGPRMCPGKEYARLQVLTFVHHIVTKFKWEQILPNEKIIVSPMPYPEKNLPLRMIARSESATLA* Sequence number 4 C28C16 oxidase (SoC28C16) amino acid sequence

[0249] SEQ ID NO:5 SoC23 Oxidase (SoC23) Nucleotide Sequence

[0250] MEYLPYIATSIACIVILRWALNMMQWLWFEPRRLEKLLRKQGLQGNSYKFLFGDMKESSMLRNEALAKPMPMPFDNDYFPRINPFVDQLLNKYGMNCFLWMGPVPAIQIGEPELVREAFNRMHEFQKPKTNPLSALLATGLVSYEGDKWAKHRRLINPSFHVEKLKLMIPAFRESIVEVVNQWEKKVPENGSAEIDVWPSLTSLTGDVISRAAFGSVYGDGRRIFELLAVQKELVLSLLKFSYIPGYTYLPTEGNKKMKAVNNEIQRLLENVIQNRKKAMEAGEAAKDDLLGLLMDSNYKESMLEGGGKNKKLIMSFQDLIDECKLFFLAGHETTAVLLVWTLILLCKHQDWQTKAREEVLATFGMSEPTDYDALNRLKIVTMILNEVLRLYPPVVSTNRKLFKGETKLGNLVIPPGVGISLLTIQANRDPKVWGEDASEFRPDRFAEGLVKATKGNVAFFPFGWGPRICIGQNFALTESKMAVAMILQRFTFDLSPSYTHAPSGLITLNPQYGAPLMFRRR* SEQ ID NO:6 SoC23 Oxidase (SoC23) Amino Acid Sequence

[0251] Array number 7 bAS (SobAS) nucleotide sequence

[0252] MWRLKIAEGGNDPYLYSTNNFVGRQTWEFDSEYGTPEAIKEVEEARQIFYKNRFQVKPCGDLLWRFQFLREKNFKQTIPQVKVGDGEEVTYEAASTTLKRSVNLLTALQADDGHWPAEIAGPQFFLPPLVFCLYITGHLNVVFNVHHREEILRSIYYHQNEDGGWGLHIEGHSTMFCTALNYICLRMLGVGPDEGDDNACPRARKWILDHGSVTHIPSWGKTWLSILGLFDWSGSNPMPPEFWILPTFMPMYPAKMWCYCRMVYMPMSYLYGKRFVGPITPLIKQLREELFSEPFEEIKWKKVRHLCAPEDLYYPHPLIQDLMWDSLYLFTEPLLTRWPFNNLIRQKALQVTMDHIHYEDENSRYITIGCVEKVLCMLACWVEDPNGVCYKKHLARVPDYIWIAEDGLKMQSFGSQQWDCGFAVQALLASNMSLDEIGPALKKGHFFIKESQVKDNPSGDFKSMHRHISKGSWTFSDQDHGWQVSDCTAEGLKCCLILSTMPPEIVGEKMDPERLYDSVNVLLSLQSENGGLSAWEPAGAQAWLELLNPTEFFADIVIEHEYVECTGASIQALVLFKKMYPGHRKKEIENFIAKAAKYLEDTQYPNGSWYGNWGVCFTYGTWFALGGLAAAGKTYANCAAMRKGVEFLLKSQKEDGGWGESYVSCPKKDFVPLEGPSNLTQTAWALMGLIYARQMERDPTPLHQAAKLLINSQLENGDFPQQEITGVFMKNCMLHYPMYRTIYPLWAIAEYRTHVPLRLS* Array number 8 bAS (SobAS) amino acid sequence

[0253] Array No. 9 SoQA-GlcAT (SoCSL) nucleotide sequence

[0254] MSPHNTCTLQITRALLSRLHILFHSALVASVFYYRFSNFSSGPAWALMTFAELTLAFIWALTQAFRWRPVVRAVFGPEEIDPAQLPGLDVFICTADPRKEPVMEVMNSVVSALALDYPAEKLAVYLSDDGGSPLTREVIREAAVFGKYWVGFCGKYNVKTRCPEAYFSSFCDGERVDHNQDYLNDELSVKSKFEAFKKYVQKASEDATKCIVVNDRPSCVEIIHDSKQNGEGEVKMPLLVYVAREKRPGFNHHAKAGAINTLLRVSGLLSNSPFFLVLDCDMYCNDPTSARQAMCFHLDPKLAPSLAFVQYPQIFYNTSKNDIYDGQARAAFKTKYQGMDGLRGPVMSGTGYFLKRKALYGKPHDQDELLREQPTKAFGSSKIFIASLGENTCVALKGLSKDELLQETQKLAACTYESNTLWGSEVGYSYDCLLESTYCGYLLHCKGWISVYLYPKKPCFLGCATVDMNDAMLQIMKWTSGLIGVGISKFSPFTYAMSRISIMQSLCYAYFAFSGLFAVFFLIYGVVLPYSLLQGVPLFPKAGDPWLLAFAGVFISSLLQHLYEVLSSGETVKAWWNEQRIWIIKSITACLFGLLDAMLNKIGVLKASFRLTNKAVDKQKLDKYEKGRFDFQGAQMFMVPLMILVVFNLVSFFGGLRRTVIHKNYEDMFAQLFLSLFILALSYPIMEEIVRKARKGRS* Array No. 10 SoQA-GlcAT (SoCSL) amino acid sequence

[0255] SEQ ID NO:11 QA-GalT(SoC3Gal) Nucleotide Sequence

[0256] MGSNTEATEIPKMPLKIVFLTLPIAGHMLHIVDTASTFAIHGVECTIITTPANVPFIEKSISATNTTIRQFLSIRLVDFPHEAVGLPPGVENFSAVTCPDMRPKISKGLSIIQKPTEDLIKEISPDCIVSDMFYPWTSDFALEIGVPRVVFRGCGMFPMCCWHSIKSHLPHEKVDRDDEMIVLPTLPDHIEMRKSTLPDWVRKPTGYSYLMKMIDAAELKSYGVIVNSFSDLERDYEEYFKNVTGLKVWTVGPISLHVGRNEELEGSDEWVKWLDGKKLDSVIYVSFGGVAKFPPHQLREIAAGLESSGHDFVWVVRASDENGDQAEADEWSLQKFKEKMKKTNHGLVIESWVPQLMFLEHKAIGGMLTHVGWGTMLEGITAGLPLVTWPLYAEQFYNERLVVDVLKIGVGVGVKEFCGLDDIGKKETIGRENIEASVRLVMGDGEEAAAMRLRVKELSEASMKAVREGGSSKANIHDFLNELSTLRSLRQA* SEQ ID NO:12 QA-GalT(SoC3Gal) Amino Acid Sequence

[0257] SEQ ID NO:13 SoQA-R XylT (SoC3Xyl) Nucleotide Sequence

[0258] MKSPLKLYFLPYISPGHMIPLSEMARLFANQGHHVTIITTTSNATLLQKYTTATLSLHLIPLPTKEAGLPDGLENFISVNDLETAGKLYYALSLLQPVIEEFITSNPPDCIVSDMFYPWTADLASQLQVPRMVFHAACIFAMCMKESMRGPDAPHLKVSSDYELFEVKGLPDPVFMTRAQLPDYVRTPNGYTQLMEMWREAEKKSYGVMVNNFYELDPAYTEHYSKIMGHKVWNIGPAAQILHRGSGDKIERVHKAVVGENQCLSWLDTKEPNSVFYVCFGSAIRFPDDQLYEIASALESSGAQFIWAVLGKDSDNSDSNSDSEWLPAGFEEKMKETGRGMIIRGWAPQVLILDHPSVGGFMTHCGWNSTIEGVSAGVGMVTWPLYAEQFYNEKLITQVLKIGVEAGVEEWNLWVDVGRKLVKREKIEAAIRAVMGEAGVEMRRKAKELSVKAKKAVQDGGSSHRNLMALIEDLQRIRDDKMSKVAN* SEQ ID NO:14 SoQA-R XylT (SoC3Xyl) Amino Acid Sequence

[0259] SEQ ID NO:15 QATriFuT (SoC28F) nucleotide sequence

[0260] MSDQNDKKVEIIVFPYHGQGHMNTMLQFAKRIAWKNAKVTIATTLSTTNKMKSKVENAWGTSITLDSIYDDSDESQIKFMDRMARFEAAAASSLSKLLVQKKEEADNKVLLVYDGNLPWALDIAHEHGVRGAAFFPQSCATVATYYSLYQETQGKELETELPAVFPPLELIQRNVPNVFGLKFPEAVVAKNGKEYSPFVLFVLRQCINLEKADLLLFNQFDKLVEPGEVLQWMSKIFNVKTIGPTLPSSYIDKRIKDDVDYGFHAFNLDNNSCINWLNSKPARSVIYIAFGSSVHYSVEQMTEIAEALKSQPNNFLWAVRETEQKKLPEDFVQQTSEKGLMLSWCPQLDVLVHESISCFVTHCGWNSITEALSFGVPMLSVPQFLDQPVDAHFVEQVWGAGITVKRSEDGLVTRDEIVRCLEVLNNGEKAEEIKANVARWKVLAKEALDEGGSSDKHIDEIIEWVSSF* SEQ ID NO:16 Amino acid sequence of QATriFuT (SoC28F)

[0261] SEQ ID NO: 17 QA-TriFR(SoC28Rha) Nucleotide Sequence

[0262] MSAKMLHVVMYPWFAYGHMIPFLHLSNKLAETGHKVTYILPPKALTRLQNLNLNPTQITFRTITVPRVDGLPAGAENVTDIPDITLHTHLATALDRTRPEFETIVELIKPDVIMYDVAYWVPEVAVKYGAKSVAYSVVSAASVSLSKTVVDRMTPLEKPMTEEERKKKFAQYPHLIQLYGPFGEGITMYDRLTGMLSKCDAIACRTCREIEGKYCQYLSTQYEKKVTLTGPVLPEPEVGATLEAPWSEWLSRFKLGSVLFCAFGSQFYLDKDQFQEIILGLEMTNLPFLMAVQPPKGCATIEEAYPEGFAERVKDRGVVTSQWVQQLVILAHPAVGCFVNHCAFGTMWEALLSEKQLVMIPQLGDQILNTKMLADELKVGVEVERGIGGWVSKENLCKAIKSVMDEDSEIGKDVKQSHEKWRATLSSKDLMSTYIDSFIKDLQALVE* SEQ ID NO: 18 QA-TriFR(SoC28Rha) Amino Acid Sequence

[0263] SEQ ID NO:19 SoQA-TriFRXylT (SoC28Xyl1) Nucleotide Sequence

[0264] MGTKELHIVMYPWLAFGHFIPYLHLSNKLAQKGHKITFLLPHRAKLQLDSQNLYPSLITLVPITVPQVDTLPLGAESTADIPLSQHGDLSIAMDRTRPEIESILSKLDPKPDLIFFDMAQWVPVIASKLGIKSVSYNIVCAISLDLVRDWYKKDDGSNVPSWTLKHDKSSHFGENISILERALIALGTPDAIGIRSCREIEGEYCDSIAERFKKPVLLSGTTLPEPSDDPLDPKWVKWLGKFEEGSVIFCCLGSQHVLDKPQLQELALGLEMTGLPFFLAIKPPLGYATLDEVLPEGFSERVRDRGVAHGGWVQQPQMLAHPSVGCFLCHCGSSSMWEALVSDTQLVLFPQIPDQALNAVLMADKLKVGVKVEREDDGGVSKEVWSRAIKSVMDKESEIAAEVKKNHTKWRDMLINEEFVNGYIDSFIKDLQDLVEK* SEQ ID NO:20 SoQA-TriFRXylT (SoC28Xyl1) Amino Acid Sequence

[0265] SEQ ID NO: 21 SoQA-TriFRXXylT (SoC28Xyl2) Nucleotide Sequence

[0266] MEESKEEVHVAFFPFMTPGHSIPMLDLVRLFIARGVKTTVFTTPLNAPNISKYLNIIQDSSSNKNTIYVTPFPSKEAGLPEGVESQDSTTSPEMTLKFFVAMELLQDPLDVFLKETKPHCLVADNFFPYATDIASKYGIPRFVFQFTGFFPMSVMMALNRFHPQNSVSSDDDPFLVPSLPHDIKLTKSQLQREYEGSDGIDTALSRLCNGAGRALFTSYGVIFNSFYQLEPDYVDYYTNTMGKRSRVWHVGPVSLCNRRHVEGKSGRGRSASISEHLCLEWLNAKEPNSVIYVCFGSLTCFSNEQLKEIATALERCEEYFIWVLKGGKDNEQEWLPQGFEERVEGKGLIIRGWAPQVLILDHEAIGGFVTHCGWNSTLESISAGVPMVTWPIYAEQFYNEKLVTDVLKVGVKVGSMKWSETTGATHLKHEEIEKALKQIMVGEEVLEMRKRASKLKEMAYNAVEEGGSSYSHLTSLIDDLMASKAVLQKF* SEQ ID NO: 22 SoQA-TriFRXXylT (SoC28Xyl2) Amino Acid Sequence

[0267] The full-length sequence of HMGR is shown below. Removing the 5' region (underlined) can generate a truncated feedback-insensitive form (tHMGR). The sequence of tHMGR is also shown separately below.

[0268] ATGGCTGTGGAGGTTCACCGCCGGGCTCCCGCGCCCCATGGCCGGGGCACCGGGGAGAAGGGCCGCGTGCAGGCCGGGGACGCGCTGCCGCTGCCGATCCGCCACACCAACCTCATCTTCTCGGCGCTCTTCGCCGCCTCCCTCGCATACCTCATGCGCCGCTGGAGGGAGAAGATCCGCAACTCCACGCCGCTCCACGTCGTGGGGCTCACCGAGATCTTCGCCATCTGCGGCCTCGTCGCCTCCCTCATCTACCTCCTCAGCTTCTTCGGCATCGCCTTCGTGCAGTCCGTCGTATCCAACAGCGACGACGAGGACGAGGACTTCCTCATCGCGGCTGCAGCATCCCAGGCCCCCCCGCCGCCCTCCTCCAAGCCCGCGCCGCAGCAGTGCGCCCTGCTGCAGAGCGCCGGAGTC Array number 23 - AsHMGR (Avena strigosa HMG-CoA reductase) coding sequence (1689 bp):

[0269] MAVEVHRRAPAPHGRGTGEKGRVQAGDALPLPIRHTNLIFSALFAASLAYLMRRWREKIRNSTPLHVVGLTEIFAICGLVASLIYLLSFFGIAFVQSVVSNSDDEDEDFLIAAAASQAPPPPSSKPAPQQCALLQSAGV APEKMPEEDEEIVAGVVAGKIPSYVLETRLGDCRRAAGIRREALRRITGREIDGLPLDGFDYDSILGQCCEMPVGYVQLPVGVAGPLVLDGRRIYVPMATTEGCLIASTNRGCKAIAESGGASSVVYRDGMTRAPVARFPSARRAAELKGFLENPANYDTLSVVFNRSSRFARLQGVKCAMAGRNLYMRFTCSTGDAMGMNMVSKGVQNVLDYLQEDFPDMDVVSISGNFCSDKKSAAVNWIEGRGKSVVCEAVIREEVVHKVLKTNVQSLVELNVIKNLAGSAVAGALGGFNAHASNIVTAIFIATGQDPAQNVESSQCITMLEAVNDGRDLHISVTMPSIEVGTVGGGTQLASQSACLDLLGVKGANRESPGSNARLLATVVAGAVLAGELSLISAQAAGHLVQSHMKYNRSSKDMSKIAC* Array number 24 - AsHMGR (Avena strigosa HMG-CoA reductase) translated nucleotide sequence (562 aa)

[0270] Array No. 25 - AstHMGR (Avena strigosa truncated HMG-CoA reductase) coding sequence (1275bp):

[0271] MAPEKMPEEDEEIVAGVVAGKIPSYVLETRLGDCRRAAGIRREALRRITGREIDGLPLDGFDYDSILGQCCEMPVGYVQLPVGVAGPLVLDGRRIYVPMATTEGCLIASTNRGCKAIAESGGASSVVYRDGMTRAPVARFPSARRAAELKGFLENPANYDTLSVVFNRSSRFARLQGVKCAMAGRNLYMRFTCSTGDAMGMNMVSKGVQNVLDYLQEDFPDMDVVSISGNFCSDKKSAAVNWIEGRGKSVVCEAVIREEVVHKVLKTNVQSLVELNVIKNLAGSAVAGALGGFNAHASNIVTAIFIATGQDPAQNVESSQCITMLEAVNDGRDLHISVTMPSIEVGTVGGGTQLASQSACLDLLGVKGANRESPGSNARLLATVVAGAVLAGELSLISAQAAGHLVQSHMKYNRSSKDMSKIAC* Array No. 26 - AstHMGR (Avena strigosa truncated HMG-CoA reductase) translated nucleotide sequence (424aa):

[0272] Array number 27 - AsSQS (Avena strigosa squalene synthase) coding sequence (1212bp):

[0273] MGALSRPEEVVALVKLRVAAGQIKRQIPAEEHWAFAYDMLQKVSRSFALVIQQLGPELRNAVCIFYLVLRALDTVEDDTSIPNDVKLPILRDFYRHVYNPDWRYSCGTNHYKVLMDKFRLVSTAFLELGEGYQKAIEEITRRMGAGMAKFICQEVETIDDYNEYCHYVAGLVGYGLSRLFHAAGTEDLASDQLSNSMGLFLQKTNIIRDYLEDINEIPKCRMFWPREIWSKYADKLEDLKYEENSEKAVQCLNDMVTNALVHAEDCLQYMSALKDNTNFRFCAIPQIMAIGTCAICYNNVKVFRGVVKMRRGLTARIIDETKSMSDVYSAFYEFSSLLESKIDDNDPSSALTRKRVEAIKRTCKSSGLLKRRGYDLEKSKYRHMLIMLALLLVAIIFGVLYAK* Array number 28 - AsSQS (Avena strigosa squalene synthase) translation nucleotide sequence (403aa):

[0274] Array number 29 - AtATR2 (Arabidopsis thaliana cytochrome P450 reductase 2) coding sequence (2325 bp):

[0275] MKNMMNYKLKLCSVSKNSKGVSLSPTPHLTKPPTIHTERDLLLPSSSFFFLLLSSSSYNIYNAMSSSSSSSTSMIDLMAAIIKGEPVIVSDPANASAYESVAAELSSMLIENRQFAMIVTTSIAVLIGCIVMLVWRRSGSGNSKRVEPLKPLVIKPREEEIDDGRKKVTIFFGTQTGTAEGFAKALGEEAKARYEKTRFKIVDLDDYAADDDEYEEKLKKEDVAFFFLATYGDGEPTDNAARFYKWFTEGNDRGEWLKNLKYGVFGLGNRQYEHFNKVAKVVDDILVEQGAQRLVQVGLGDDDQCIEDDFTAWREALWPELDTILREEGDTAVATPYTAAVLEYRVSIHDSEDAKFNDINMANGNGYTVFDAQHPYKANVAVKRELHTPESDRSCIHLEFDIAGSGLTYETGDHVGVLCDNLSETVDEALRLLDMSPDTYFSLHAEKEDGTPISSSLPPPFPPCNLRTALTRYACLLSSPKKSALVALAAHASDPTEAERLKHLASPAGKDEYSKWVVESQRSLLEVMAEFPSAKPPLGVFFAGVAPRLQPRFYSISSSPKIAETRIHVTCALVYEKMPTGRIHKGVCSTWMKNAVPYEKSENCSSAPIFVRQSNFKLPSDSKVPIIMIGPGTGLAPFRGFLQERLALVESGVELGPSVLFFGCRNRRMDFIYEEELQRFVESGALAELSVAFSREGPTKEYVQHKMMDKASDIWNMISQGAYLYVCGDAKGMARDVHRSLHTIAQEQGSMDSTKAEGFVKNLQTSGRYLRDVW* Array number 30 - AtATR2 (Arabidopsis thaliana cytochrome P450 reductase 2) translated nucleotide sequence (774 aa):

[0276] ATGGCTGAAGCATCCTCATTTCTTGCACAGAAAAGGTATGCGGTCGTGACAGGAGCAAACAAAGGACTAGGACTAGAAATATGCGGACAGCTTGCTTCACAGGGGGTGACGGTACTGCTGACATCCAGAGATGAAAAACGAGGCTTAGAAGCCATTGAGGAGCTTAAGAAATCGGGGATTAATTCGGAAAATCTTGAATATCATCAGCTGGATGTTACTAAGCCAGCTAGTTTCGCTTCTCTGGCCGATTTCATCAAGGCCAAATTTGGCAAGCTTGATATCCTGGTGAACAATGCAGGGATCAGCGGTGTTATTGTAGATTATGCAGCTTTAATGGAAGCCATTCGCCGTCGAGGGGCAGAGATCAATTACGATGGAGTGATGAAACAGACCTACGAGCTAGCAGAGGAATGCTTGCAAACAAATTACTATGGTGTGAAAAGAACCATTAATGCTCTCCTTCCGCTACTTCAGTTTTCCGATTCACCAAGGATCGTCAATGTTTCCTCCGATGTTGGCCTCCTTAAGAAAATACCCGGCGAGAGAATCAGAGAAGCCTTAGGCGACGTGGAAAAACTTACGGAAGAAAGCGTGGACGGGATTTTAGACGAGTTTCTAAGAGATTTCAAGGAAGGCAAGATCGCAGAGAAAGGTTGGCCTACGTTTAAGAGCGCCTATTCAATCTCAAAGGCGGCGCTCAATTCGTACACGAGGGTTTTAGCACGGAAATACCCGTCGATCATCATCAACTGTGTCTGCCCGGGTGTCGTCAAAACCGATATCAATCTTAAAATGGGCCACTTGACGGTTGAAGAAGGCGCGGCCAGTCCCGTGAGGTTAGCACTCATGCCCCTTGGTTCGCCTTCCGGCCTGTTCTATACTCGAAACGAAGTAACTCCATTTGAATGA SEQ ID NO: 31 SoFuSyn coding sequence

[0277] MAEASSFLAQKRYAVVTGANKGLGLEICGQLASQGVTVLLTSRDEKRGLEAIEELKKSGINSENLEYHQLDVTKPASFASLADFIKAKFGKLDILVNNAGISGVIVDYAALMEAIRRRGAEINYDGVMKQTYELAEECLQTNYYGVKRTINALLPLLQFSDSPRIVNVSSDVGLLKKIPGERIREALGDVEKLTEESVDGILDEFLRDFKEGKIAEKGWPTFKSAYSISKAALNSYTRVLARKYPSIIINCVCPGVVKTDINLKMGHLTVEEGAASPVRLALMPLGSPSGLFYTRNEVTPFE* Sequence number 32 SoFuSyn translated nucleotide sequence

[0278] Array number 33 SoGH1 coding sequence

[0279] MVLSRLDFPSDFIFGSGTSASQVEGAALEDGKTSTAFEGFLTRMSGNDLSKGVEGYYKYKEDVQLMVQTGLDAYRFSISWSRLIPGGKGPVNPKGLQYYNNFIDELIKNGIQPHVTLLHFDIPDTLMTAYNGLKGQEFVEDFTAFADVCFKEFGDRVLYWTTVNEANNFASLTLDEGNFMPSTEPYIRGHNIILAHASAVKLYREKYKKTQNGFIGLNLYASWYFPETDDEQDSIAAQRAIDFTIGWIMQPLIYGEYPETLKKQVGERLPTFTKEESTFVKNSFDFIGVNCYVGTAVKDDPDSCNSKNKTIITDMSAKLSPKGELGGAYMKGLLEYFKRDYGNPPIYIQENGYWTPRELGVNDASRIEYHTASLASMHDAMKNGA Array number 34 SoGH1 translated nucleotide sequence

[0280] Array number 35 SoBAHD1 coding sequence

[0281] MEPSKMEVKIISSETIKPSSPTPSHLRKYTLSLLDQKYTPIVVPAILFYERPQGVAPLDMDRLRTCLSQTLTAFYPLAGRAESRDVIICNDEGIPFVEAHVDCELSSVVKSLSSLGSDLRSFYPPRDGLLEGGIQFAIQMNVFSCGGFAFAWYCTHNVTDGTSTANFFRYWTALYAQRSEYAVQDLMDFNSVVTAFPPVPPRVPQEEKPVTTELKPEKQEGQEKEEKKKSSFNFSFQSHIVARSFLIKSKAVAELKAKSVSEEVPYPSRFEAVSAFLWKSIVSSSTTEGKTMINMPVNLRPRVDPPLPLDSVGNIFENALVQSEKKAELHEFVARIRGSISKMKDFATEYQGEKREEAKDAHWKRFIKAVIECKGKDAYVISPWYKSSGFTDIDFGFGTPIRVVPMDDVVNHNQRNTIMLMEFVDSDGDGFEAWMFLEEECIKFLESNPEFLAFASPNF Array number 36 SoBAHD1 translated nucleotide sequence

[0282] Table

[0283]

Table 1

[0284]

Table 2

[0285]

Table 3

[0286]

Table 4

[0287]

Table 5

[0288]

Table 6

[0289] References [1] Jia, Z., Koike, K. and Nikaido, T. (1998). Major triterpenoid saponins from Saponaria officinalis. Journal of Natural Products. 61: 1368-1373. [2] Eastman, J. (2014). Wildflowers of the Eastern United States: An Introduction to Common Species of Woods, Wetlands and Fields. Stackpole Books. [3] Rees, A. (1819). The cyclopaedia; or, universal dictionary of arts, sciences, and literature (Vol. 4). Longman, Hurst, Rees, Orme and Brown. [4] Korkmaz, M. and Ozcelik, H. (2011). Economic importance of Gypsophila L., Ankyropetalum fenzl and Saponaria L.(Caryophyllaceae) taxa of Turkey. African journal of Biotechnology, 10(47), 9533-9541. [5] Bottger, S. and Melzig, M. F. (2011). Triterpenoid saponins of the Caryophyllaceae and Illecebraceae family. Phytochemistry Letters. 4: 59-68. [6] Smulek, W., Zdarta, A., Pacholak, A., Zgo;a-Grzeskowiak, A., Marczak, L., Jarzebski, M., and Kaczorek, E. (2017). Saponaria officinalis L. extract: Surface active properties and impact on environmental bacterial strains. Colloids and Surfaces B: Biointerfaces, 150, 209-215. [7] Gonzalez, P. J. and Sorensen, P. M. (2020). Characterization of saponin foam from Saponaria officinalis for food applications. Food Hydrocolloids, 101, 105541. [8] Sadowska, B., Budzynska, A., Wieckowska-Szakiel, M., Paszkiewicz, M., Stochmal, A., Moniuszko-Szajwaj, B., Kowalczyk, M. and Rozalska, B. (2014). New pharmacological properties of Medicago sativa and Saponaria officinalis saponin-rich fractions addressed to Candida albicans. Journal of medical microbiology, 63(8), 1076-1086. [9] Gilabert-Oriol, R., Thakur, M., Haussmann, K., Niesler, N., Bhargava, C., Gorick, C., Fuchs, H. and Weng, A. (2016). Saponins from Saponaria officinalis L. augment the efficacy of a rituximab-immunotoxin. Planta medica, 82(18), 1525-1531.

[10] Reed, J., Orme, A., El-Demerdash, A., Owen, C., Martin, L. B., Misra, R. C.,... & Osbourn, A. (2023). Elucidation of the pathway for biosynthesis of saponin adjuvants from the soapbark tree. Science, 379(6638), 1252-1264.

Claims

1. (i) a step of contacting an OS with a Saponaria officinalis β-amyrin synthase (SoAS) comprising an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 8 such that the OS is converted into β-amyrin; (ii) a) a step of contacting β-amyrin with a SoC28 oxidase polypeptide comprising an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 2, thereby oxidizing the C28 position of the β-amyrin to a carboxylic acid to produce oleanolic acid, and contacting the oleanolic acid with a SoC28C16 oxidase polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 4, oxidizing the alcohol at the C16 position of the oleanolic acid, thereby producing echinocystic acid, or b) a step of contacting oleanolic acid with a SoC28C16 oxidase polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 4, as a result, oxidizing the C16 position of the oleanolic acid to an alcohol, oxidizing the C28 position of the β-amyrin to a carboxylic acid, thereby producing echinocystic acid; (iii) a step of contacting echinocystic acid with a SoC23 oxidase polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 6, as a result, oxidizing the C-23 position of the echinocystic acid to an aldehyde, thereby producing quillaic acid (QA); (iv) a step of contacting QA with a Saponaria officinalis QA 3-O glucuronosyltransferase ("SoCSL") polypeptide comprising an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 10, as a result, the QA is converted to QA-GlcA; (v) a step of contacting QA-GlcA with a Saponaria officinalis QA-GlcA galactosyltransferase ("SoC3Gal") polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 12, as a result, the QA-GlcA is converted to QA-GlcA-Gal; (vi) contacting QA-GlcA with a Saponaria officinalis QA-GlcA-Gal xylosyltransferase (“SoC3Xyl”) polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 14, whereby said QA-GlcA-Gal is converted to QA-Tri; (vii) contacting QA-Tri with a Saponaria officinalis QA-Tri fucosyltransferase (“SoC28Fu”) polypeptide comprising an amino acid sequence having at least 60% sequence identity with SEQ ID NO: 16, whereby said QA-Tri is converted to QA-TriF; (viii) contacting QA-TriF with a Saponaria officinalis QA-TriF rhamnosyltransferase (“SoC28Rha”) polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 18, whereby said QA-TriF is converted to QA-TriFR; (ix) contacting QA-TriFR with a Saponaria officinalis QA-TriFR xylosyltransferase (“SoC28Xyl1”) polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 20, whereby said QA-TriFR is converted to QA-TriFRX; (x) contacting QA-TriFRX with a Saponaria officinalis QA-TriFRX xylosyltransferase (“SoC28Xyl2”) polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 22, whereby said QA-TriFRX is converted to QA-TriFRXX; (xi) contacting QA-TriFRXX with a Saponaria officinalis QA-TriFRXX quinovosyltransferase (“SoGH1”) polypeptide comprising an amino acid sequence having at least 50% sequence identity with SEQ ID NO: 34, whereby said QA-TriFRXX is converted to QA-TriF(Q)RXX; and / or (xii) Contacting QA-Trif(Q)RXX with a Saponaria officinalis QA-Trif(Q)RXX acetyltransferase ("SoBAHD1") polypeptide having an amino acid sequence with at least 50% sequence identity to SEQ ID NO: 36, whereby said QA-Trif(Q)RXX is converted to saponarioside B (SpB). A method for producing a triterpenoid, comprising: **Claim 2** (i) (a) Contacting β-amyrin with Saponaria officinalis C28 oxidase (SoC28 oxidase) to oxidize the C28 position of β-amyrin to a carboxylic acid to form oleanolic acid (the amino acid sequence of SoC28 oxidase has at least 80% sequence identity to SEQ ID NO: 2), and contacting oleanolic acid with Saponaria officinalis C28C16 oxidase (SoC28C16 oxidase) to oxidize the C16 position of oleanolic acid to an alcohol to form oleanolic acid (the amino acid sequence of C16 oxidase has at least 50% sequence identity to SEQ ID NO: 4); or (b) Contacting β-amyrin with Saponaria officinalis C28C16 oxidase (SoC28C16 oxidase) to oxidize the C28 position of β-amyrin to a carboxylic acid and the C16 position to an alcohol to form echinocystic acid (the amino acid sequence of C28C16 oxidase has at least 50% sequence identity to SEQ ID NO: 4). (iii) Contacting echinocystic acid with Saponaria officinalis C-23 oxidase (SoC23 oxidase) to oxidize the C23 position of echinocystic acid to an aldehyde to form quillaic acid (QA) (the amino acid sequence of SoC23 oxidase has at least 50% sequence identity to SEQ ID NO: 6). The method according to claim 1, comprising: **Claim 3** The method according to claim 2, wherein β-amyrin is produced by contacting 2,3-oxidosqualene (OS) with a β-amyrin synthase (SobSAS) having an amino acid sequence with at least 80% sequence identity to SEQ ID NO: 8, thereby cyclizing OS to produce β-amyrin. **Claim 4** (iv) Contacting QA with Saponaria officinalis QA 3-O-glucuronosyltransferase (“SoCSL”) to covalently bond D-glucuronic acid (“GlcA”) to the 3-O position of kirenol acid to form 3-O-{β-D-glucopyranosiduronic acid}-kirenol acid (“QA-GlcA”) (the amino acid sequence of SoCSL has at least 60% sequence identity with SEQ ID NO: 10); (v) Contacting QA-GlcA with Saponaria officinalis QA-GlcA galactosyltransferase (“SoC3Gal”) to covalently bond D-galactose (“Gal”) to QA-GlcA via a β-1→2 linkage to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-kirenol acid (“QA-GlcA-Gal”) (the amino acid sequence of QA-GlcA-Gal has at least 50% sequence identity with SEQ ID NO: 12); and (vi) Contacting QA-GlcA-Gal with Saponaria officinalis QA-GlcA-Gal xylosyltransferase (“SoC3Xyl”) to covalently bond D-xylose (“Xyl”) to QA-GlcA-Gal via a 1,3 linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-kirenol acid (“QA-GlcA-Gal”) (the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14), the method according to claim 2 or claim 3, further comprising. [

5. ] (vii) Contacting 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-oleanolic acid (QA-Tri) with Saponaria officinalis QA-trifucosyltransferase (SoC28Fu) to bind fucose to the QA-Tri at the 28O position to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF) (the amino acid sequence of QATriFuT has at least 60% sequence identity with SEQ ID NO: 16); (viii) Contacting QA-TriF with Saponaria officinalis QA-TriF rhamnosyltransferase (SoC28Rha) to covalently bind rhamnose to QA-TriF via a 1,2 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester (the amino acid sequence of SoC28Rha has at least 50% sequence identity with SEQ ID NO: 18); (ix) Contacting QA-TriFR with Saponaria officinalis QA-TriFR xylosyltransferase (SoC28Xyl1) to covalently bind xylose to QA-TriFR via a 1,4 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRX) (the amino acid sequence of SoC28Xyl1 has at least less than 50% sequence identity with SEQ ID NO: 20); and Contacting (x) QA-TriFRX with Saponaria officinalis QA-TriFRX-xylosyltransferase (“SoC28Xyl2”) and covalently attaching xylose to QA-TriFRX via a 1,3 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRXX) (the amino acid sequence of SoC28Xyl2 has at least 50% sequence identity with SEQ ID NO: 22). The method according to claim 4, further comprising the above. **Claim 6** Contacting QA-TriFRXX with Saponaria officinalis QA-TriFRXX quinovosyltransferase (“SoGH1”) and covalently attaching quinovose to QA-TriFRXX via a 1,4 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)]-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX) (the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34); and Contacting QA-TriF(Q)RXX with Saponaria officinalis QA-TriF(Q)RXX acetyltransferase (“SoBAHD1”) and covalently attaching an acetyl group to QA-TriF(Q)RXX to form QA-TriF(Q-Ac)RXX (saponarioside B) (the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36). The method according to claim 5, further comprising the above. **Claim 7** A method for converting a host from a phenotype unable to perform triterpenoid biosynthesis from 2,3-oxide squalene (OS) to a phenotype capable of performing said triterpenoid biosynthesis, said method comprising, subsequent to a step prior to introducing nucleic acid into the host or one of its ancestors, expressing a heterologous nucleic acid in the host or one or more of its cells, wherein the heterologous nucleic acid is (i) SoC28 oxidase capable of oxidizing β-amyrin to a carboxylic acid at the C28 position to form oleanolic acid (said SoC28 oxidase having at least 80% sequence identity with SEQ ID NO: 2); (ii) SoC28C16 oxidase (``C16 oxidase'') capable of oxidizing β-amyrin to a carboxylic acid at the C28 position and to an alcohol at the C16 position to form echinocystic acid (said SoC28C16 oxidase having at least 50% sequence identity with SEQ ID NO: 4); (iii) SoC23 oxidase capable of oxidizing echinocystic acid to an aldehyde at the C23 position to form kirenolic acid (QA) (said SoC23 oxidase having at least 50% sequence identity with SEQ ID NO: 6); (iv) Saponaria officinalis QA 3-O glucuronosyltransferase (``SoCSL'') for binding D-glucuronic acid (``GlcA'') to the 3-O position of kirenolic acid to form 3-O-{β-D-glucopyranosiduronic acid}-kirenolic acid (``QA-GlcA'') (said SoCSL having at least 60% sequence identity with SEQ ID NO: 10); (v) Saponaria officinalis QA-GlcA galactosyltransferase (``SoC3Gal'') for binding D-galactose (``Gal'') to QA-GlcA via a β-1→2 bond to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-kirenolic acid (``QA-GlcA-Gal'') (the amino acid sequence of SoC3Gal having at least 50% sequence identity with SEQ ID NO: 12); and (vi) A Saponaria officinalis QA-GlcA-Gal xylosyltransferase ("SoC3Xyl") (the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14) for binding D-xylose ("Xyl") to QA-GlcA-Gal via a 1,3-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-gypsogenin ("QA-GlcA-[Gal]-Xyl", QA-Tri); and (vii) A Saponaria officinalis QA-Tri fucosyltransferase (SoC28Fu) (the SoC28Fu has at least 60% sequence identity with SEQ ID NO: 16) for binding fucose ("Fuc") to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-fucopyranosyl ester}-gypsogenin (QA-TriF); (viii) A Saponaria officinalis QA-TriF rhamnosyltransferase (SoC28Rha) (the SoC28Rha has at least 50% sequence identity with SEQ ID NO: 18) for binding rhamnose ("Rha") to QA-TriF via a 1,2-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-gypsogenin (QA-TriFR); (ix) Binding D-xylose ("Xyl") to QA-TriFR via a 1,4 linkage to form "3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRX)" using Saponaria officinalis QA-TriFR xylosyltransferase ("SoC28Xyl1") (the amino acid sequence of SoC28Xyl1 has at least 50% sequence identity with SEQ ID NO: 20); (x) Binding D-xylose ("Xyl") to QA-TriFRX via a 1,3 linkage to form "3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRXX)" using Saponaria officinalis QA-TriFRX xylosyltransferase ("SoC28Xyl12") (the amino acid sequence of SoC28Xyl2 has at least 80% sequence identity with SEQ ID NO: 22); (xi) Binding quinovose to QA-TrFRXX via a 1,4 linkage to form "3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl(1→3)-β-D-xylopyranosyl-(1→4)-α-l-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX)" using Saponaria officinalis QA-TriFRXX xylosyltransferase ("SoGH1") (the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34); and / or (xii) A method for encoding one or more of Saponaria officinalis QA-TrIF(Q)RXX acetyltransferase ("SoBAHD1") (the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36) for binding an acetyl group to QA-TrIF(Q)RXX to form saponarioside B. A method for encoding one or more of the above. **Claim 8** The heterologous nucleic acid encodes the following polypeptides: (i) SoC28 oxidase capable of oxidizing β-amyrin to carboxylic acid at the C28 position to form oleanolic acid (the SoC28 oxidase has at least 80% sequence identity with SEQ ID NO: 2); (ii) SoC28C16 oxidase ("C16 oxidase") capable of oxidizing β-amyrin to carboxylic acid at the C28 position and to alcohol at the C16 position to form echinocystic acid (the SoC28C16 oxidase has at least 50% sequence identity with SEQ ID NO: 4); and (iii) SoC23 oxidase capable of oxidizing echinocystic acid to aldehyde at the C-23 position to form quillaic acid (QA) (the SoC23 oxidase has at least 50% sequence identity with SEQ ID NO: 6) The method according to claim 7, wherein the heterologous nucleic acid encodes the above. (xii) A method for encoding one or more of Saponaria officinalis QA-TrIF(Q)RXX acetyltransferase ("SoBAHD1") (the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36) for binding an acetyl group to QA-TrIF(Q)RXX to form saponarioside B. The method according to claim 8, wherein the heterologous nucleic acid further encodes Saponaria officinalis β-amyrin synthase (SobAS) for cyclizing OS to triterpene; and the SobAS has at least 80% sequence identity with SEQ ID NO:

8. **Claim 10** The heterologous nucleic acid further encodes the following polypeptides: (iv) Saponaria officinalis QA 3-O glucuronosyltransferase ("SoCSL") for binding D-glucuronic acid ("GlcA") to the 3-O position of quillaic acid to form 3-O-{β-D-glucopyranosiduronic acid}-quillaic acid ("QA-GlcA") (the SoCSL has at least 60% sequence identity with SEQ ID NO: 10); (v) Saponaria officinalis QA-GlcA galactosyltransferase (SoC3Gal) (the amino acid sequence of SoC3Gal has at least 50% sequence identity with SEQ ID NO: 12) for binding D-galactose ("Gal") to QA-GlcA via a β-1→2 linkage to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-oleanolic acid ("QA-GlcA-Gal"); and (vi) Saponaria officinalis QA-GlcA-Gal xylosyltransferase (SoC3Xyl) (the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14) for binding D-xylose ("xylose") to QA-GlcA-Gal via a 1,3-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-oleanolic acid ("QA-GlcA-Gal-Xyl" or "QA-Tri") The method according to claim 8 or 9, encoding [

11. ] The heterologous nucleic acid further comprises the following polypeptide: (vii) Saponaria officinalis QA-Tri fucosyltransferase (SoC28Fu) (the SoC28Fu has at least 60% sequence identity with SEQ ID NO: 16) for binding fucose ("Fuc") to the 28O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF); (viii) Saponaria officinalis QA-TriF rhamnosyltransferase ("SoC28Rha") (said SoC28Rha having at least 50% sequence identity with SEQ ID NO: 18) for attaching rhamnose ("Rha") to QA-TriF via a 1,2-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-gypsogenic acid (QA-TrFR); (ix) Saponaria officinalis QA-TrFR xylosyltransferase ("SoC28Xyl1") (the amino acid sequence of SoC28Xyl1 having at least 50% sequence identity with SEQ ID NO: 20) for attaching D-xylose ("Xyl") to QA-TrFR via a 1,4-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-gypsogenic acid (QA-TrFRX); and (x) Saponaria officinalis QA-TrFRX xylosyltransferase ("SoC28Xyl2") (the amino acid sequence of SoC28Xyl2 having at least 80% sequence identity with SEQ ID NO: 22) for attaching D-xylose ("Xyl") to QA-TrFRX via a 1,3-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-gypsogenic acid (QA-TrFRXX) The method according to claim 10, encoding the same. [

12. ] The heterologous nucleic acid further comprises the following polypeptide: (xi) binding quinovose (Q) to QA-TrFRXX via a 1,4 linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX) using Saponaria officinalis QA-TriFXX xylosyltransferase ("SoGH1") (the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34); and (xii) binding an acetyl group to QA-TriF(Q)RXX using Saponaria officinalis QA-TriF(Q)RXX acetyltransferase ("SoBAHD1") (the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36) to form ponarioside B The method according to claim 11, encoding the same. [

13. ] A host cell comprising a heterologous nucleic acid comprising a plurality of nucleotide sequences encoding polypeptides having triterpenoid biosynthetic activity in combination, or a host cell transformed with the heterologous nucleic acid, wherein the plurality of nucleotide sequences are the following polypeptides: (i) Saponaria officinalis β-amyrin synthase (SobAS) for the cyclization of OS to triterpenes (said SobAS has at least 80% sequence identity with SEQ ID NO: 8); (ii) SoC28 oxidase capable of oxidizing β-amyrin to a carboxylic acid at the C28 position to form oleanolic acid (said SoC28 oxidase has at least 80% sequence identity with SEQ ID NO: 2); (iii) SoC28C16 oxidase ("C16 oxidase") capable of oxidizing β-amyrin to a carboxylic acid at the C28 position and to an alcohol at the C16 position to form echinocystic acid (said SoC28C16 oxidase has at least 50% sequence identity with SEQ ID NO: 4); and (iv) SoC23 oxidase capable of oxidizing echinocystic acid to an aldehyde at the C-23 position to form quillaic acid (QA) (the SoC23 oxidase has at least 50% sequence identity with SEQ ID NO: 6); (v) Saponaria officinalis QA 3-O glucuronosyltransferase (SoCSL) for binding D-glucuronic acid ("GlcA") to the 3-O position of quillaic acid to form 3-O-{β-D-glucopyranosiduronic acid}-quillaic acid ("QA-GlcA") (the SoQA-GlcT has at least 60% sequence identity with SEQ ID NO: 10); (vi) Saponaria officinalis QA-GlcA galactosyltransferase (SoCdGal) for binding D-galactose ("Gal") to QA-GlcA via a β-1→2 bond to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-quillaic acid ("QA-GlcA-Gal") (the amino acid sequence of QA-GlcA-Gal has at least 50% sequence identity with SEQ ID NO: 12); (vii) Saponaria officinalis QA-GlcA-Gal xylosyltransferase (ScC3Xyl) for binding D-xylose ("Xyl") to QA-GlcA-Gal via a 1,3 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-quillaic acid ("QA-GlcA-[Gal]-Xyl" QA-Tri) (the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14); (viii) Saponaria officinalis QA-Tri fucosyltransferase (SoC28Fu) for binding fucose ("Fuc") to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-fucopyranosyl ester}-quillaic acid (QA-TriF) (the SoC28Fu has at least 60% sequence identity with SEQ ID NO: 16); (ix) attaching rhamnose ("Rha") to QA-TriF via a 1,2-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-gypsogenic acid (QA-TrFR) with Saponaria officinalis QA-TrF rhamnosyltransferase ("SoC28Rha") (said SoC28Rha having at least 50% sequence identity with SEQ ID NO: 18); (x) attaching D-xylose ("Xyl") to QA-TrFR via a 1,4-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-gypsogenic acid (QA-TrFRX) with Saponaria officinalis QA-TrFR xylosyltransferase ("SoC28Xyl1") (the amino acid sequence of SoC28Xyl1 having at least 50% sequence identity with SEQ ID NO: 20); and / or (xi) attaching D-xylose ("Xyl") to QA-TrFRX via a 1,3-linkage to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-gypsogenic acid (QA-TrFRXX) with Saponaria officinalis QA-TrFRX xylosyltransferase ("SoC28Xyl2") (the amino acid sequence of SoC28Xyl2 having at least 50% sequence identity with SEQ ID NO: 22); (xii) Saponaria officinalis QA-TriFRXX quinovosyltransferase ("SoGH1") (the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34) for forming 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX) via a 1,4 linkage with quinovose; and / or (xiii) Saponaria officinalis QA-TriF(Q)RXX acetyltransferase ("SoBAHD1") (the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36) for attaching an acetyl group to QA-TriF(Q)RXX to form saponarioside B encoding one or more of the above, and expression of said nucleic acid confers on the transformed host the ability to perform triterpenoid biosynthesis. **Claim 14** A plurality of nucleotide sequences encode the following polypeptides: (i) SoC28 oxidase capable of oxidizing β-amyrin to carboxylic acid at the C28 position to form oleanolic acid (said SoC28 oxidase has at least 80% sequence identity with SEQ ID NO: 2); (ii) SoC28C16 oxidase ("C16 oxidase") capable of oxidizing β-amyrin to carboxylic acid at the C28 position and to alcohol at the C16 position to form echinocystic acid (said SoC28C16 oxidase has at least 50% sequence identity with SEQ ID NO: 4); and (iii) SoC23 oxidase capable of oxidizing echinocystic acid to aldehyde at the C-23 position to form oleanolic acid (QA) (said SoC23 oxidase has at least 50% sequence identity with SEQ ID NO: 6) The host cell according to claim 13, wherein expression of said nucleic acid confers on the transformed host the ability to perform QA biosynthesis. (xiv) Saponaria officinalis QA-TriF(Q)RXX acetyltransferase ("SoBAHD1") (the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36) for attaching an acetyl group to QA-TriF(Q)RXX to form saponarioside B The heterologous nucleic acid further encodes Saponaria officinalis β-amyrin synthase (SoAS) for cyclizing the OS to triterpenes; the host cell according to claim 14, wherein the SoAS has at least 80% sequence identity with SEQ ID NO:

8.

16. The heterologous nucleic acid further encodes the following polypeptides: (iv) Saponaria officinalis QA 3-O glucuronosyltransferase (SoCSL) for binding D-glucuronic acid ("GlcA") to the 3-O position of kiwifruit acid to form 3-O-{β-D-glucopyranosiduronic acid}-kiwifruit acid ("QA-GlcA") (the SoCSL has at least 60% sequence identity with SEQ ID NO: 10); (v) Saponaria officinalis QA-GlcA galactosyltransferase (SoC3Gal) for binding D-galactose ("Gal") to QA-GlcA via a β-1→2 bond to form 3-O-{[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-kiwifruit acid (the amino acid sequence of SoC3Gal has at least 60% sequence identity with SEQ ID NO: 12); and (vi) Saponaria officinalis QA-GlcA-Gal xylosyltransferase (SoC3Xyl) for binding D-xylose ("Xyl") to QA-GlcA-Gal via a 1,3 bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-kiwifruit acid ("QA-GlcA-Gal-Xyl" QA-Tri) (the amino acid sequence of SoC3Xyl has at least 50% sequence identity with SEQ ID NO: 14) The host cell according to claim 14 or 15.

17. The heterologous nucleic acid further encodes the following polypeptides (vii)Saponaria officinalis QA-Tri fucosyltransferase (SoC28Fu) for attaching fucose ("Fuc") to the 28-O position of QA-Tri to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-fucopyranosyl ester}-gypsogenic acid (QA-TriF) (the SoC28Fu has at least 60% sequence identity with SEQ ID NO: 16); (viii)Saponaria officinalis QA-Tri rhamnosyltransferase (SoC28Rha) for attaching rhamnose ("Rha") to QA-TriF via a 1,2-bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-gypsogenic acid (QA-TriFR) (the SoC28Rha has at least 50% sequence identity with SEQ ID NO: 18); (ix)Saponaria officinalis QA-TriFR xylosyltransferase (SoC28Xyl1) for attaching D-xylose ("Xyl") to QA-TriFR via a 1,4-bond to form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranuronate}-28-O-{β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-gypsogenic acid (QA-TriFRX) (the amino acid sequence of SoC28Xyl1 has at least 50% sequence identity with SEQ ID NO: 20); and (x) To form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriFRXX) by binding D-xylose ("Xyl") to QA-TriFRX via a 1,3-linkage, Saponaria officinalis QA-TriFR xylosyltransferase ("SoC28Xyl2") (the amino acid sequence of SoC28Xyl2 has at least 50% sequence identity with SEQ ID NO: 22) The host cell according to claim 16, encoding one, two, three or all four of them. **Claim 18** The heterologous nucleic acid is the following polypeptide: (xi) To form 3-O-{β-D-xylopyranosyl-(1→3)-[β-D-galactopyranosyl-(1→2)]-β-D-glucopyranosiduronic acid}-28-O-{β-D-xylopyranosyl-(1→3)-β-D-xylopyranosyl-(1→4)-α-L-rhamnopyranosyl-(1→2)-[β-D-quinovopyranosyl-(1→4)]-β-D-fucopyranosyl ester}-oleanolic acid (QA-TriF(Q)RXX) by binding quinovose (Q) to QA-TrFRXX via a 1,4-linkage, Saponaria officinalis QA-TriFRXX quinovosyltransferase ("SoGH1") (the amino acid sequence of SoGH1 has at least 50% sequence identity with SEQ ID NO: 34), and (xii) Saponaria officinalis QA-TriF(Q)RXX acetyltransferase ("SoBAHD1") (the amino acid sequence of SoBAHD1 has at least 50% sequence identity with SEQ ID NO: 36) for binding an acetyl group to QA-TriF(Q)RXX to form saponarioside B The host cell according to claim 17, further encoding one or both of them. **Claim 19** (i) The So bAS amino acid sequence having at least 80% sequence identity with SEQ ID NO: 8; (ii) The SoC28 oxidase amino acid sequence having at least 80% sequence identity with SEQ ID NO: 2; (iii) an SoC16C28 oxidase amino acid sequence having at least 50% sequence identity with SEQ ID NO: 4; (iv) an SoC23 oxidase amino acid sequence having at least 50% sequence identity with SEQ ID NO: 6; (v) an SoCSL amino acid sequence having at least 60% sequence identity with SEQ ID NO: 10; (vi) an SoC3Gal amino acid sequence having at least 50% sequence identity with SEQ ID NO: 12; (vii) an SoQA-RXylT amino acid sequence having at least 50% sequence identity with SEQ ID NO: 14; (viii) an SoC28Fu amino acid sequence having at least 60% sequence identity with SEQ ID NO: 16; (ix) an SoC28Rha amino acid sequence having at least 50% sequence identity with SEQ ID NO: 18; (x) an SoC28Xyl1 amino acid sequence having at least 50% sequence identity with SEQ ID NO: 20; (xi) an SoC28Xyl2 amino acid sequence having at least 50% sequence identity with SEQ ID NO: 22; (xii) an SoGH1 amino acid sequence having at least 50% sequence identity with SEQ ID NO: 34; and / or (xiii) an SoBAHD1 amino acid sequence having at least 50% sequence identity with SEQ ID NO: 36 An isolated polypeptide comprising the same.

20. An isolated nucleic acid encoding one or more polypeptides according to Claim 19.

21. A vector comprising the nucleic acid according to Claim 20.

22. A host cell comprising the nucleic acid according to Claim 20 or the vector according to Claim 21.

23. A method for generating a host cell comprising transforming or transfecting a host cell with a heterologous nucleic acid comprising a plurality of nucleotide sequences according to any one of Claims 7 to 18 and 20.

24. The method according to Claim 23, wherein the host cell is a plant cell.

25. A method for generating a transgenic plant, comprising: (a) performing the method according to Claim 24, and (b) regenerating a plant from the transformed plant cell A method comprising the same.

26. A transgenic plant obtainable by the method according to Claim 25, or a clone of said transgenic plant, or a transgenic plant which is a self-propagated or crossbred offspring or other offspring of said transgenic plant, A transgenic plant in which the expression of a heterologous nucleic acid confers an increased ability to perform triterpenoid biosynthesis compared to a wild-type plant corresponding to the transgenic plant.

27. A method for producing a triterpenoid in a heterologous host, the method comprising culturing the host cell according to any one of claims 13 to 18 and 22 and purifying the triterpenoid therefrom.

28. A method for producing a triterpenoid in a heterologous host, the method comprising growing the plant according to claim 26, then harvesting it, and purifying the triterpenoid therefrom.

29. The method according to claim 27 or 28, wherein the triterpenoid is QA or glycosylated QA.

30. The method according to claim 29, wherein the glycosylated QA is QA-Tri, QA-TriFRXX, or QA-TriF(Q-Ac)RXX.