Constructs and methods for the biosynthesis of gastrodin
By employing host cells with transgenes encoding heterologous UGTs with specific amino acid sequences and gastrodin synthases, the inefficiencies in gastrodin biosynthesis are addressed, resulting in enhanced production efficiency.
Patent Information
- Application Number
- JP2025528445
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-17
- Filing Date
- 2023-11-17
- Publication Date
- 2025-12-09
Smart Images

Figure 2025539777000001_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims the benefit of U.S. Provisional Application No. 63 / 384,169, filed November 17, 2022, the entire teachings of which are incorporated herein by reference.
[0002] Including material by reference in XML This application incorporates by reference the sequence listing contained in the following extensible markup language (XML) file filed concurrently herewith: a) File name: 5767_1006-001_SL.xml; created on November 17, 2023, and is 59,909 bytes in size. [Background technology]
[0003] Gastrodin is a natural product with various biological activities, including neuroprotective, analgesic, and anti-inflammatory effects in both humans and model organisms. Gastrodin is produced by the plant Gastrodia elata, also known as Tianma in traditional Chinese medicine. Gastrodin is one of the major bioactive components of Gastrodia plant extracts. Gastrodin has shown efficacy in several pain models and is considered a potential treatment for chronic pain, neuropathic pain, and chemotherapy-induced pain, both as a monotherapy and in combination with other therapeutic agents.
[0004] To achieve the final conversion of 4-hydroxybenzyl alcohol to gastrodin, glycosyltransferase (UGT) enzymes that utilize UDP-glucose as a sugar donor are required. Several enzymes, including UGT73B6 and AsUGT, have been described as gastrodin synthases and engineered into heterologous organisms to produce gastrodin. (CN113755354A; incorporated herein by reference in its entirety).
[0005] However, efficient biosynthesis of gastrodin in such heterologous organisms depends on the successful engineering of multiple glucose-dependent enzymes into the gastrodin biosynthetic pathway. Such processes are cost- and time-inefficient and result in poor yields. Therefore, there is a need for constructs and methods for the efficient biosynthesis of gastrodin. Summary of the Invention
[0006] In one aspect, the disclosure provides a host cell comprising a transgene encoding a heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least 75% amino acid sequence identity to SEQ ID NO:2, wherein the amino acid sequence of the heterologous UGT comprises one or more of M at position 88 of SEQ ID NO:2, F, Y, or W at position 119 of SEQ ID NO:2, F, Y, or W at position 145 of SEQ ID NO:2, L or M at position 149 of SEQ ID NO:2, M at position 198 of SEQ ID NO:2, and F at position 383 of SEQ ID NO:2.
[0007] In another aspect, the disclosure provides a host cell comprising a transgene encoding a heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least 75% amino acid sequence identity to SEQ ID NO:28, and the amino acid sequence of the heterologous UGT comprises one or more of M at position 93 of SEQ ID NO:28, F, Y, or W at position 129 of SEQ ID NO:28, F, Y, or W at position 150 of SEQ ID NO:28, L or M at position 154 of SEQ ID NO:28, M at position 203 of SEQ ID NO:28, and F at position 391 of SEQ ID NO:28.
[0008] In another aspect, the disclosure provides a method for producing gastrodin in a host cell, the method comprising culturing the host cell in a cell culture medium comprising 4-hydroxybenzyl alcohol, wherein the host cell expresses a transgene encoding a heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least 75% amino acid sequence identity to SEQ ID NO:2, wherein the amino acid sequence of the heterologous UGT comprises one or more of M at position 88 of SEQ ID NO:2, F, Y, or W at position 119 of SEQ ID NO:2, F, Y, or W at position 145 of SEQ ID NO:2, L or M at position 149 of SEQ ID NO:2, M at position 198 of SEQ ID NO:2, and F at position 383 of SEQ ID NO:2.
[0009] In another aspect, the disclosure provides a method for producing gastrodin in a host cell, the method comprising culturing the host cell in a cell culture medium comprising 4-hydroxybenzyl alcohol, wherein the host cell expresses a transgene encoding a heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least 75% amino acid sequence identity to SEQ ID NO:28, wherein the amino acid sequence of the heterologous UGT comprises one or more of M at position 93 of SEQ ID NO:28, F, Y, or W at position 129 of SEQ ID NO:28, F, Y, or W at position 150 of SEQ ID NO:28, L or M at position 154 of SEQ ID NO:28, M at position 203 of SEQ ID NO:28, and F at position 391 of SEQ ID NO:28.
[0010] In another aspect, the present disclosure provides a vector comprising a nucleic acid encoding a gastrodin synthase for converting 4-hydroxybenzyl alcohol to gastrodin, wherein the gastrodin synthase can have at least about 75% amino acid sequence identity to SEQ ID NO:2.
[0011] In another aspect, the disclosure provides a vector comprising a nucleic acid encoding a gastrodin synthase for converting 4-hydroxybenzyl alcohol to gastrodin, wherein the nucleic acid can have at least about 75% amino acid sequence identity to SEQ ID NO:28.
[0012] In another aspect, the disclosure provides a method of producing a genetically modified host cell, the method comprising introducing a vector into the host cell, the vector comprising a nucleic acid encoding a heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least 75% amino acid sequence identity to SEQ ID NO:2, wherein the amino acid sequence of the heterologous UGT comprises one or more of M at position 88 of SEQ ID NO:2; F, Y, or W at position 119 of SEQ ID NO:2; F, Y, or W at position 145 of SEQ ID NO:2; L or M at position 149 of SEQ ID NO:2; M at position 198 of SEQ ID NO:2; and F at position 383 of SEQ ID NO:2.
[0013] In another aspect, the disclosure provides a method of producing a genetically modified host cell, the method comprising introducing a vector into the host cell, the vector comprising a nucleic acid encoding a heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least 75% amino acid sequence identity to SEQ ID NO:28, wherein the amino acid sequence of the heterologous UGT comprises one or more of M at position 93 of SEQ ID NO:28; F, Y, or W at position 129 of SEQ ID NO:28; F, Y, or W at position 150 of SEQ ID NO:28; L or M at position 154 of SEQ ID NO:28; M at position 203 of SEQ ID NO:28; and F at position 391 of SEQ ID NO:28.
[0014] In another aspect, the present disclosure provides a pharmaceutical composition comprising a gastrodin, wherein the gastrodin is produced by a genetically modified plant or plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
[0015] The foregoing will become more apparent from the following more detailed description of exemplary embodiments. As illustrated in the accompanying drawings, like reference numerals refer to the same parts in the different views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the embodiments. [Brief explanation of the drawings]
[0016] [Figure 1] This figure shows the chemical conversion of 4-hydroxybenzyl alcohol to gastrodin. The use of the UDP-glucose glycotransferase (UGT) enzyme, GeUGT, to catalyze the transfer of glucose to the 4-hydroxy group of 4-hydroxybenzyl alcohol using UDP-glucose as the glucose donor is described. This reaction produces gastrodin and UDP as by-products. [Figure 2] This study demonstrates that GeUGT is more efficient than the previously described AsUGT in converting 4-HBA to gastrodin. 4-HBA was incubated with E. coli strains containing either GeUGT or AsUGT for 48 hours, and 4-HBA conversion was measured by HPLC analysis. GeUGT demonstrated complete conversion of 4-HBA to gastrodin, unlike AsUGT, which did not completely convert 4-HBA in the medium. Standards were run in the growth medium and used to identify both 4-HBA and gastrodin in the test strains. [Figure 3] A total of 11 UGTs were identified that could potentially convert 4-HBA to gastrodin. In addition to GeUGT, the 11 aforementioned UGTs were found to have gastrodin synthase activity, an activity not previously reported for these enzymes. These enzymes were assayed in the same manner as GeUGT, and their activities were compared after 24 and 48 hours. [Figure 4]A sequence alignment of GeUGT (SEQ ID NO: 2), AsUGT (SEQ ID NO: 26), and 11 additional described gastrodin synthase enzymes (SEQ ID NOs: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, and 24) and the consensus sequence (SEQ ID NO: 28) is shown. Sequence analysis reveals that residues M88, F119, I139, F145, L149, M198, and F383 of GeUGT are almost entirely unique to this enzyme and may potentially account for the highly active nature of this enzyme. [Figure 5A] An image of the GeUGT active site and identified amino acid residues is shown. [Figure 5B] An image of the AsUGT active site and identified amino acid residues is shown. DETAILED DESCRIPTION OF THE INVENTION
[0017] A description of exemplary embodiments follows.
[0018] Some aspects of the present disclosure are described below with reference to examples for illustrative purposes only. It should be understood that certain specific details, relationships, and methods are set forth to provide a thorough understanding of the present disclosure. However, one of ordinary skill in the art will readily recognize that the present disclosure can be practiced without one or more of the specific details, or can be practiced using other methods, protocols, reagents, cell lines, and animals. The present disclosure is not limited to the illustrated order of acts or events, as acts may occur in different orders and / or concurrently with other acts or events. Furthermore, not all illustrated acts, steps, or events are required to implement a methodology in accordance with the present disclosure. Many of the techniques and procedures described or referenced herein are well understood and commonly employed by those of ordinary skill in the art using conventional methodology.
[0019] Unless otherwise defined, all technical terms, notations, and other scientific or technical terms used herein are intended to have the meaning commonly understood by one of ordinary skill in the art to which this disclosure pertains. In various instances, terms having commonly understood meanings are defined herein for clarity and / or ease of reference, but the inclusion of such definitions herein should not necessarily be construed as representing a substantial difference from that commonly understood in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be construed to have a meaning consistent with their meaning in the context of the relevant art and / or as otherwise defined herein.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0021] As used herein, the indefinite articles "a," "an," and "the" are to be understood as including plural references unless the context clearly indicates otherwise.
[0022] Unless the context requires otherwise, throughout this specification and in the claims that follow, the word "comprise," as well as variations such as "comprises" or "comprising," will be understood to imply, for example, the inclusion of a stated integer or step or group of integers or steps, but not the exclusion of any other integer or step or group of integers or steps. As used herein, the term "comprising" can be replaced with the words "containing" or "including."
[0023] As used herein, "consisting of" excludes any element, step, or ingredient not specified in the claim element. As used herein, "consisting essentially of" does not exclude materials or steps that do not materially affect the basic and novel characteristics of the claim. When used herein in the context of aspects or embodiments of the present disclosure, any of the terms "comprising," "containing," "including," and "having" can be substituted with the terms "consisting of" or "consisting essentially of" in various embodiments to vary the scope of the present disclosure.
[0024] As used herein, the conjunction "and / or" between multiple listed elements is understood to encompass both individual and combined options. For example, when two elements are joined by "and / or," the first option refers to the first element being applicable without the second element. The second option refers to the second element being applicable without the first element. The third option refers to the first and second elements being applicable together. Any one of these options is understood to be within the meaning and thus meets the requirements of the term "and / or" as used herein. If two or more of the options are simultaneously applicable, it is also understood to be within the meaning and therefore meets the requirements of the term "and / or."
[0025] Where lists are presented, unless otherwise expressly stated, it is to be understood that each individual element of that list, and every combination of that list, is a separate embodiment. For example, a list of embodiments presented as "A, B, or C" should be interpreted as including the embodiments "A," "B," "C," "A or B," "A or C," "B or C," or "A, B, or C."
[0026] nucleic acid As used herein, the term "nucleic acid" refers to a polymer comprising multiple nucleotide monomers (e.g., ribonucleotide monomers or deoxyribonucleotide monomers). "Nucleic acid" includes, for example, DNA (e.g., genomic DNA and cDNA), RNA, and DNA-RNA hybrid molecules. Nucleic acid molecules can be naturally occurring, recombinant, or synthetic. Furthermore, nucleic acid molecules can be single-stranded, double-stranded, or triple-stranded. In certain embodiments, nucleic acid molecules can be modified. In the case of a double-stranded polymer, "nucleic acid" can refer to either or both strands of the molecule.
[0027] The terms "nucleotide" and "nucleotide monomer" refer to naturally occurring ribonucleotide or deoxyribonucleotide monomers, as well as non-naturally occurring derivatives and analogs thereof. Thus, nucleotides can include, for example, nucleotides containing naturally occurring bases (e.g., adenosine, thymidine, guanosine, cytidine, uridine, inosine, deoxyadenosine, deoxythymidine, deoxyguanosine, or deoxycytidine), as well as nucleotides containing modified bases known in the art.
[0028] As used herein, "wild-type" refers to the standard amino acid sequence found in nature. As one of skill in the art will appreciate, nucleic acid sequences can be modified (e.g., for codon optimization in host cells, such as bacterial, yeast, and plant host cells).
[0029] As used herein, the term "sequence identity" refers to the degree, expressed as a percentage, to which two nucleotide sequences or two amino acid sequences have the same residues at the same positions when the sequences are aligned to achieve the maximum level of identity. For sequence alignment and comparison, typically, one sequence is designated as a reference sequence to which a test sequence is compared. Sequence identity is expressed as the percentage of positions over the entire length of a reference sequence that the reference sequence and the test sequence share the same nucleotide or amino acid when aligned to achieve the maximum level of identity. As an example, if the test sequence has the same nucleotide or amino acid residue at 70% of the same positions over the entire length of the reference sequence when aligned to achieve the maximum level of identity, the two sequences are considered to have 70% sequence identity.
[0030] Alignment of sequences for comparison to achieve the maximum level of identity can be easily performed by those skilled in the art using an appropriate alignment technique or algorithm. In some cases, alignment can include the introduction of gaps to maximize the level of identity. Examples include the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), computer implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), and visual inspection (see generally Ausubel et al., Current Protocols in Molecular Biology).
[0031] When using sequence comparison algorithm, test sequence and reference sequence are input into computer, and if necessary, subsequent coordinates are designated, and sequence algorithm program parameters are designated.The sequence comparison algorithm then calculates the percent sequence identity of test sequence (s) to reference sequence based on designated program parameters.The tool commonly used to determine percent sequence identity is the Basic Local Alignment of Proteins Search Tool (BLASTP), which can be obtained through the National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health (USA) (Altschul et al., 1990).
[0032] In various embodiments, two nucleotide sequences or two amino acid sequences can have at least, e.g., 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity. When determining percent sequence identity to one or more sequences described herein, the sequences described herein are the reference sequences.
[0033] Various embodiments of the present invention relate to nucleic acid coding sequences (e.g., dsDNA, cDNA) encoding one or more of the enzymes described herein, including those nucleic acid sequences provided in SEQ ID NO:1 and SEQ ID NO:27. Many different nucleic acids can encode the UGTs of the present disclosure due to the degeneracy of the genetic code. Nucleic acids can also differ, for example, as a result of one or more mutations (e.g., silent mutations).
[0034] enzyme As used herein, the term 5'-diphospho-glucosyltransferase (UGT) refers to an enzyme that catalyzes the conversion of 4-hydroxybenzyl alcohol to gastrodin. Methods and assays for determining whether an enzyme catalyzes the conversion of 4-hydroxybenzyl alcohol to gastrodin are known in the art and include enzyme activity assays and liquid chromatography to assess the retention time of metabolites. Chemical structure can also be assessed by nuclear magnetic resonance (NMR) or liquid chromatography-mass spectrometry. An example of a UGT is SEQ ID NO: 2, which is the amino acid sequence of a UGT identified in Gastrodia elata (GeUGT).
[0035] Aspects of the present disclosure provide UGTs, or biologically active fragments thereof, having at least about 70% or more sequence identity to SEQ ID NO:2. In other aspects, the present disclosure provides UGTs, or biologically active fragments thereof, having at least about 75% or more sequence identity to SEQ ID NO:2. The present disclosure provides UGTs, or biologically active fragments thereof, having at least about 76% or more sequence identity to SEQ ID NO:2. In yet other aspects, the present disclosure provides UGTs, or biologically active fragments thereof, having at least about 77% or more sequence identity to SEQ ID NO:2. In other aspects, the present disclosure provides UGTs, or biologically active fragments thereof, having at least about 78% or more sequence identity to SEQ ID NO:2. In yet other embodiments, the present disclosure provides UGTs, or biologically active fragments thereof, having at least about 78% or more sequence identity to SEQ ID NO:2. In yet other embodiments, the present disclosure provides UGTs, or biologically active fragments thereof, having at least about 79% or more sequence identity to SEQ ID NO:2. In yet another aspect, the disclosure provides a UGT having at least about 80% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In yet another aspect, the disclosure provides a UGT having at least about 81% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In yet another embodiment, the disclosure provides a UGT having at least about 82% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In another embodiment, the disclosure provides a UGT having at least about 83% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In yet another embodiment, the disclosure provides a UGT having at least about 84% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In a further embodiment, the disclosure provides a UGT having at least about 85% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In another aspect, the disclosure provides a UGT having at least about 86% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In yet another aspect, the disclosure provides a UGT having at least about 87% or greater sequence identity to SEQ ID NO:2, or a biologically active fragment thereof.In another aspect, the disclosure provides a UGT having at least about 88% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In a further embodiment, the disclosure provides a UGT having at least about 89% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. The disclosure also provides a UGT having at least about 90% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In a further aspect, the disclosure provides a UGT having at least about 91% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In yet another embodiment, the disclosure provides a UGT having at least about 92% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In an aspect, the disclosure provides a UGT having at least about 93% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In another embodiment, the disclosure provides a UGT having at least about 94% or more sequence identity to SEQ ID NO:2, or a biologically active fragment thereof. In further embodiments, the present disclosure also provides a UGT, or biologically active fragment thereof, having at least about 95% or more sequence identity to SEQ ID NO:2. In embodiments, the present disclosure provides a UGT, or biologically active fragment thereof, having at least about 96% or more sequence identity to SEQ ID NO:2. Additionally, aspects of the present disclosure provide a UGT, or biologically active fragment thereof, having at least about 97% or more sequence identity to SEQ ID NO:2. In other embodiments, the present disclosure provides a UGT, or biologically active fragment thereof, having at least about 98% or more sequence identity to SEQ ID NO:2. In yet other embodiments, the present disclosure provides a UGT, or biologically active fragment thereof, having at least about 99% or more sequence identity to SEQ ID NO:2. The present disclosure also provides a UGT, or biologically active fragment thereof, that shares sequence identity with SEQ ID NO:2.
[0036] In one aspect, the disclosure provides a heterologous UGT operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least 75% amino acid sequence identity to SEQ ID NO:2, wherein the amino acid sequence of the heterologous UGT comprises one or more of: M at position 88 of SEQ ID NO:2, F, Y, or W at position 119 of SEQ ID NO:2, F, Y, or W at position 145 of SEQ ID NO:2, L or M at position 149 of SEQ ID NO:2, M at position 198 of SEQ ID NO:2, and F at position 383 of SEQ ID NO:2. In an embodiment, the heterologous UGT comprises at least two of: M at position 88 of SEQ ID NO:2, F, Y, or W at position 119 of SEQ ID NO:2, F, Y, or W at position 145 of SEQ ID NO:2, L or M at position 149 of SEQ ID NO:2, M at position 198 of SEQ ID NO:2, or F at position 383 of SEQ ID NO:2. As a non-limiting example, the heterologous UGT can comprise an M at position 88 of SEQ ID NO:2 and an F at position 119 of SEQ ID NO:2. Embodiments of the present disclosure provide heterologous UGTs, wherein the heterologous UGT comprises at least three of: M at position 88 of SEQ ID NO:2, F, Y, or W at position 119 of SEQ ID NO:2, F, Y, or W at position 145 of SEQ ID NO:2, L or M at position 149 of SEQ ID NO:2, M at position 198 of SEQ ID NO:2, or F at position 383 of SEQ ID NO:2. As a non-limiting example, the heterologous UGT can comprise M at position 88 of SEQ ID NO:2, F at position 119 of SEQ ID NO:2, and F at position 145 of SEQ ID NO:2. In embodiments, the heterologous UGT comprises at least four of: M at position 88 of SEQ ID NO:2, F, Y, or W at position 119 of SEQ ID NO:2, F, Y, or W at position 145 of SEQ ID NO:2, L or M at position 149 of SEQ ID NO:2, M at position 198 of SEQ ID NO:2, or F at position 383 of SEQ ID NO:2. As a non-limiting example, a heterologous UGT can include an M at position 88 of SEQ ID NO:2, an F at position 119 of SEQ ID NO:2, an F at position 145 of SEQ ID NO:2, and an F at position 383 of SEQ ID NO:2.In other embodiments, the heterologous UGT comprises at least five of: M at position 88 of SEQ ID NO:2, F, Y, or W at position 119 of SEQ ID NO:2, F, Y, or W at position 145 of SEQ ID NO:2, L or M at position 149 of SEQ ID NO:2, M at position 198 of SEQ ID NO:2, or F at position 383 of SEQ ID NO:2. As a non-limiting example, the heterologous UGT can comprise M at position 88 of SEQ ID NO:2, F at position 119 of SEQ ID NO:2, F at position 145 of SEQ ID NO:2, L at position 149 of SEQ ID NO:2, and F at position 383 of SEQ ID NO:2. In yet other embodiments, the heterologous UGT comprises all of: M at position 88 of SEQ ID NO:2, F, Y, or W at position 119 of SEQ ID NO:2, F, Y, or W at position 145 of SEQ ID NO:2, L or M at position 149 of SEQ ID NO:2, M at position 198 of SEQ ID NO:2, or F at position 383 of SEQ ID NO:2. As a non-limiting example, a heterologous UGT can include an M at position 88 of SEQ ID NO:2, an F at position 119 of SEQ ID NO:2, an F at position 145 of SEQ ID NO:2, an L at position 149 of SEQ ID NO:2, an M at position 198 of SEQ ID NO:2, and an F at position 383 of SEQ ID NO:2.
[0037] In yet other embodiments, the disclosure provides a UGT operably linked to a promoter, wherein the UGT comprises an amino acid sequence that does not have one or more of the following residues: I at position 88 of SEQ ID NO:2; L at position 119 of SEQ ID NO:2; C at position 145 of SEQ ID NO:2; F at position 149 of SEQ ID NO:2; L at position 198 of SEQ ID NO:2; or Y at position 383 of SEQ ID NO:2.
[0038] In any of the foregoing embodiments, the disclosure provides a UGT having at least about 70% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In other aspects, the disclosure provides a UGT having at least about 75% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In some aspects, the disclosure provides a UGT having at least about 76% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In yet other aspects, the disclosure provides a UGT having at least about 77% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In other aspects, the disclosure provides a UGT having at least about 78% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In yet other embodiments, the disclosure provides a UGT having at least about 79% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In yet another aspect, the disclosure provides a UGT having at least about 80% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In yet another aspect, the disclosure provides a UGT having at least about 81% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In yet another embodiment, the disclosure provides a UGT having at least about 82% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In another embodiment, the disclosure provides a UGT having at least about 83% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In yet another embodiment, the disclosure provides a UGT having at least about 84% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In a further embodiment, the disclosure provides a UGT having at least about 85% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In another aspect, the disclosure provides a UGT having at least about 86% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In yet another aspect, the disclosure provides a UGT having at least about 87% or greater sequence identity to SEQ ID NO: 28, or a biologically active fragment thereof.In another aspect, the disclosure provides a UGT having at least about 88% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In a further embodiment, the disclosure provides a UGT having at least about 89% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. The disclosure also provides a UGT having at least about 90% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In a further aspect, the disclosure provides a UGT having at least about 91% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In yet another embodiment, the disclosure provides a UGT having at least about 92% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In an aspect, the disclosure provides a UGT having at least about 93% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In another embodiment, the disclosure provides a UGT having at least about 94% or more sequence identity to SEQ ID NO:28, or a biologically active fragment thereof. In further embodiments, the present disclosure also provides a UGT having at least about 95% or more sequence identity to SEQ ID NO: 28, or a biologically active fragment thereof. In embodiments, the present disclosure provides a UGT having at least about 96% or more sequence identity to SEQ ID NO: 28, or a biologically active fragment thereof. Additionally, aspects of the present disclosure provide a UGT having at least about 97% or more sequence identity to SEQ ID NO: 28, or a biologically active fragment thereof. In other embodiments, the present disclosure provides a UGT having at least about 98% or more sequence identity to SEQ ID NO: 28, or a biologically active fragment thereof. In yet other embodiments, the present disclosure provides a UGT having at least about 99% or more sequence identity to SEQ ID NO: 28, or a biologically active fragment thereof. The present disclosure also provides a UGT sharing sequence identity with SEQ ID NO: 28, or a biologically active fragment thereof.
[0039] In one aspect, the disclosure provides a heterologous UGT operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least about 75% amino acid sequence identity to SEQ ID NO:28, wherein the amino acid sequence of the heterologous UGT comprises one or more of M at position 93 of SEQ ID NO:28, F, Y, or W at position 129 of SEQ ID NO:28, F, Y, or W at position 150 of SEQ ID NO:28, L or M at position 154 of SEQ ID NO:28, M at position 203 of SEQ ID NO:28, and F at position 391 of SEQ ID NO:28.
[0040] vector The terms "vector," "vector construct," and "expression vector" refer to a vehicle by which DNA or RNA sequences (e.g., foreign genes) can be introduced into a host cell to transform the host and promote expression (e.g., transcription and translation) of the introduced sequences. A vector typically contains a transmissible DNA element into which foreign DNA encoding a protein can be inserted by restriction enzyme technology. A common type of vector is the "plasmid," which is generally a self-contained molecule of double-stranded DNA that can readily accept additional (foreign) DNA and can be easily introduced into a suitable host cell. Numerous vectors, including plasmids and fungal vectors, have been described for replication and / or expression in a variety of eukaryotic and prokaryotic hosts.
[0041] The terms "express" and "expression" mean enabling or causing the information in a gene or DNA sequence to be made manifest, e.g., by activating cellular functions involved in the transcription and translation of the corresponding gene or DNA sequence, to produce a protein. A DNA sequence is expressed in or by a cell to form an "expression product," such as a protein. The expression product itself (e.g., the resulting protein) may also be said to be "expressed" by the cell. A polynucleotide or polypeptide is recombinantly expressed when, for example, it is expressed or produced in a foreign host cell under the control of a foreign or native promoter, or in a native host cell under the control of a foreign promoter. A gene delivery vector generally contains a transgene (e.g., a nucleic acid encoding an enzyme) operably linked to a promoter and other nucleic acid elements necessary for expression of the transgene in a host cell into which the vector is introduced. Suitable promoters for gene expression and delivery constructs are known in the art. For bacterial host cells, suitable promoters include, but are not limited to, promoters from the E. coli lac operon, the agarase gene (dagA) of Streptomyces coelicolor, the levansucrase gene (sacB) of Bacillus subtilis, the α-amylase gene (amyL) of Bacillus licheniformis, the maltogenic amylase gene (amyM) of Bacillus stearothermophilus, the α-amylase gene (amyQ) of Bacillus amyloliquefaciens, the penicillinase gene (penP) of Bacillus licheniformis, the xy1A and xy1B genes of Bacillus subtilis, and prokaryotic β-lactamase genes (e.g., Villa-Kamaroff et al., Proc. Natl. Acad. Sci. USA 75: 3727-3731, 1978), as well as promoters derived from the tac promoter (see, e.g., DeBoer et al., Proc. Natl. Acad. Sci. USA 80: 21-25, 1983).Examples of promoters for filamentous fungal host cells include, but are not limited to, promoters obtained from the genes for Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral α-amylase, Aspergillus niger acid-stable α-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triosephosphate isomerase, Aspergillus nidulans acetamidase, Fusarium oxysporum trypsin-like protease (see, e.g., WO 96 / 00787), and the NA2-tpi promoter (Aspergillus niger neutral α-amylase and Aspergillus Examples of yeast cell promoters include those derived from the Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase genes. Other useful promoters for yeast host cells are known in the art (see, e.g., Romanos et al., Yeast 8:423-488, 1992). Selection of a suitable promoter is within the skill of the art. The recombinant plasmid can also contain an inducible or regulatable promoter for expression of the enzyme in the cell.
[0042] Various gene delivery vehicles are known in the art, including both viral and non-viral (e.g., naked DNA, plasmid) vectors. Viral vectors suitable for gene delivery are known to those skilled in the art. Such viral vectors include, but are not limited to, vectors derived from herpes viruses, baculovirus vectors, lentivirus vectors, retrovirus vectors, adenovirus vectors, and adeno-associated virus vectors (AAV). Vectors derived from plant viruses can also be used, including the viral backbones of the RNA viruses tobacco mosaic virus (TMV), potato virus X (PVX), and cowpea mosaic virus (CPMV), and the DNA geminivirus bean mosaic yellows virus. Viral vectors can be replicating or non-replicating.
[0043] Non-viral vectors include, inter alia, naked DNA and plasmids, non-limiting examples of which include pKK plasmids (Clonetech), pUC plasmids, pET plasmids (Novagen, Inc., Madison, Wis.), pRSET or pREP plasmids (Invitrogen, San Diego, Calif.), or pMAL plasmids (New England Biolabs, Beverly, Mass.), which can be introduced into many suitable host cells using methods disclosed or cited herein or otherwise known to those skilled in the art.
[0044] In certain embodiments, the vector comprises a transgene operably linked to a promoter, the transgene encoding a biologically active molecule, such as an enzyme (e.g., a heterologous UGT) described herein.
[0045] To facilitate the introduction of gene delivery vectors into host cells, the vectors can be combined with different chemical means, such as colloidal dispersion systems (e.g., polymer complexes, nanocapsules, microspheres, beads) or lipid-based systems (e.g., oil-in-water emulsions, micelles, liposomes).
[0046] The present disclosure also provides embodiments relating to vectors comprising nucleic acids encoding the enzymes described herein. In certain embodiments, the vector is a plasmid and includes any one or more plasmid sequences (e.g., promoter sequences, selectable marker sequences, and / or locus targeting sequences).
[0047] Although the genetic code is degenerate in that most amino acids are represented by multiple codons (termed "synonymous" or "synonymous" codons), it is understood in the art that codon usage by particular organisms is non-random and biased toward particular codon triplets. Thus, in various embodiments, vectors contain nucleotide sequences that can be optimized (e.g., through codon optimization) for expression in a particular type of host cell. Codon optimization refers to a process in which a polynucleotide encoding a protein of interest is modified to replace particular codons in the polynucleotide with codons that encode the same amino acid but are more commonly used / recognized in the host cell in which the nucleic acid is expressed. In various aspects, the polynucleotides described herein are codon-optimized for expression in bacterial cells (e.g., E. coli) or yeast cells (e.g., S. cerevisiae).
[0048] host cell A wide variety of host cells can be used, including fungal cells, bacterial cells, plant cells, insect cells, and mammalian cells.
[0049] In embodiments of the present disclosure, the host cell is a fungal cell, such as a yeast cell or an Aspergillus cell. A wide variety of yeast cells are suitable, including Pichia cells, including Pichia pastoris and Pichia stipitis, Saccharomyces cells, including Saccharomyces cerevisiae, Schizosaccharomyces cells, including Schizosaccharomyces pombe, and Candida cells, including Candida albicans.
[0050] In other embodiments, the host cell is a bacterial cell. A wide variety of bacterial cells are suitable, including cells of the genus Escherichia, including Escherichia coli, Bacillus, including Bacillus subtilis, Pseudomonas, including Pseudomonas aeruginosa, and Streptomyces, including Streptomyces griseus.
[0051] In another embodiment, the host cell is a plant cell. Cells from a wide variety of plants are suitable, including cells from Nicotiana benthamiana plants. In another embodiment, the plant belongs to a genus selected from the group consisting of Arabidopsis, Beta, Glycine, Helianthus, Solanum, Triticum, Oryza, Brassica, Medicago, Prunus, Malus, Hordeum, Musa, Phaseolus, Citrus, Piper, Sorghum, Daucus, Manihot, Capsicum, and Zea.
[0052] In yet other embodiments, the host cell is an insect cell, including, for example, a Spodoptera frugiperda cell, such as the Spodoptera frugiperda Sf9 cell line, and Spodoptera frugiperda Sf21.
[0053] In a further embodiment, the host cell is a mammalian cell.
[0054] In further embodiments, the host cell is an Escherichia coli cell. In embodiments of the disclosure, the host cell is a Nicotiana benthamiana cell. In other embodiments, the cell is a Saccharomyces cerevisiae cell.
[0055] As used herein, the term "host cell" also encompasses cells in cell culture and cells within an organism (e.g., a plant).
[0056] Various embodiments relate to host cells comprising the vectors described herein. In certain embodiments, the host cell is an Escherichia coli cell, a Nicotiana benthamiana cell, or a Saccharomyces cerevisiae cell.
[0057] In embodiments, the host cells are cultured in a cell culture medium, such as a standard cell culture medium known in the art to be suitable for the particular host cell.
[0058] In another aspect, the disclosure provides a method for producing gastrodin in a host cell, the method comprising culturing a host cell in a cell culture medium comprising 4-hydroxybenzyl alcohol, wherein the host cell expresses a transgene encoding a heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to a promoter.
[0059] In embodiments of the present disclosure, the host cell is a plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell. In yet another aspect, the present disclosure provides heterologous UGTs that are codon-optimized for expression in a host cell.
[0060] In an aspect of the disclosure, the method provides a cell culture medium further comprising glucose. In another aspect, the disclosure provides a method comprising making a host cell, the method comprising introducing into the host cell a vector comprising a nucleic acid encoding a heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to a promoter.
[0061] In yet another aspect, the present disclosure provides a method wherein gastrodin is extracted using maceration, permeation, decoction, reflux extraction, Soxhlet extraction, pressurized liquid extraction, supercritical fluid extraction, ultrasound-assisted extraction, pulsed field extraction, enzyme-assisted extraction, steam distillation, steam distillation, or any combination thereof.
[0062] In further aspects of the disclosure, the methods provided herein include providing a concentration of gastrodin in the cell culture medium after 24 hours of incubation, the concentration being at least about 4 mM, 5 mM, 6 mM, 7 mM, or 8 mM. In other aspects, the disclosure provides a concentration of 4-hydroxybenzyl alcohol in the cell culture medium after 24 hours of incubation, the concentration being no more than 2 mM, 1.5 mM, or 1 mM.
[0063] Methods for producing genetically modified host cells Methods for producing genetically modified host cells are described herein. Genetically modified host cells can be produced, for example, by introducing one or more of the vector embodiments described herein into a host cell.
[0064] In a further embodiment, the present disclosure provides a method of producing a genetically modified host cell, the method comprising introducing into the host cell a vector comprising a nucleic acid encoding a heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to a promoter.
[0065] In other embodiments, one or more of the nucleic acids are integrated into the genome of the host cell. In still other embodiments, the nucleic acids to be integrated into the host genome can be introduced into the host cell using any of a variety of suitable methodologies known in the art, including, for example, CRISPR-based systems (e.g., CRISPR / Cas9, CRISPR / Cpf1), TALEN systems, and Agrobacterium-mediated transformation. However, as those skilled in the art will recognize, transient transformation techniques can be used that do not require integration into the genome of the host cell. In embodiments of the present disclosure, nucleic acids (e.g., plasmids) that are maintained as episomes and do not need to be integrated into the host cell genome can be introduced.
[0066] In certain embodiments, the nucleic acid is introduced into the tissue, cell, or seed of a plant cell. Various methods for introducing nucleic acid into plant tissue, cell, or seed are known to those skilled in the art, including, for example, protoplast transformation. A particular method can be selected based on several considerations, such as the type of plant used. For example, the floral immersion method described herein is a suitable method for introducing genetic material into a plant. In certain embodiments, the nucleic acid can be delivered to the plant by Agrobacterium.
[0067] Methods for Making Gastrodin Methods for making gastrodin are described herein. In various embodiments, the present disclosure provides pharmaceutical compositions comprising gastrodin, wherein the gastrodin is produced by a genetically modified plant or plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
[0068] Values and Ranges Unless otherwise indicated or otherwise apparent from the context and the understanding of one of ordinary skill in the art, values expressed as ranges can, in various embodiments, assume any specific value or subrange within the stated range, unless the context clearly dictates otherwise. "About" with respect to numerical values generally refers to values within a range of ±8%, in some embodiments within ±6%, in some embodiments within ±4%, in some embodiments within ±2%, in some embodiments within ±1%, and in some embodiments within ±0.5%, unless otherwise specified or apparent from the context.
[0069] array Table 1 is a summary of the nucleotide and amino acid sequences disclosed in the sequence listing incorporated herein. [Table 1] [Example]
[0070] Example 1 Materials and Methods Transcriptomics of G. elata: To discover enzymes capable of converting 4-hydroxybenzyl alcohol to gastrodin, the raw transcriptome of G. elata was processed as previously described. 10, 11 The transcriptome constructed from this data was searched using NCBI BLAST with AsUGT as the search query. From this transcriptome, five candidate genes were selected as likely UGT enzymes.
[0071] Cloning into expression vectors: Each putative UGT enzyme was cloned into a pCL1921-derived plasmid containing a constitutive pL promoter to drive expression and a spectinomycin resistance cassette. The UGTs were assembled into the pCL1921 backbone using Gibson assembly, transformed into DH5-a cells, and plated on spectinomycin. Sequenced plasmids containing the UGT were electroporated into wild-type E. coli C and selected on medium containing spectinomycin.
[0072] Test strains for gastrodin production: Three colonies from each plasmid transformation were inoculated into 1 mL of LB medium containing 10 mM 4-hydroxybenzyl alcohol and 2% glucose. These strains were grown in deep-well plates at 37°C with shaking at 1000 rpm for 48 hours. Samples were taken at both 24 and 48 hours for analysis.
[0073] Detection and quantification of gastrodin: Using standards for both gastrodin and 4-hydroxybenzyl alcohol, we developed an HPLC-based method for detecting these two compounds in strain testing experiments. At 24 and 48 hours postinoculation, each strain was sampled, centrifuged to remove cells, and the supernatant was removed for HPLC analysis. Strains with gastrodin synthase activity showed the appearance of a peak corresponding to gastrodin, while the signal for 4-hydroxybenzyl alcohol was reduced. After 48 hours, the conversion of 4-hydroxybenzyl alcohol to gastrodin was assessed (Figure 2).
[0074] Other UGT options: Other similar UGTs predicted to have gastrodin synthase activity based on their sequences were tested to determine whether they could convert 4-hydroxybenzyl alcohol to gastrodin. A combination of enzyme screening and BLAST searches of the NCBI database identified 11 additional UGTs (Figure 3) that were capable of converting 4-hydroxybenzyl alcohol to gastrodin when assayed as described above. No explanation for gastrodin production by any of these 11 UGTs has been identified. These genes were SiUGT (SEQ ID NO: 4), CaUGT (SEQ ID NO: 6), PpUGT (SEQ ID NO: 8), NbUGT1 (SEQ ID NO: 10), NbUGT2 (SEQ ID NO: 12), PcUGT (SEQ ID NO: 14), WsUGT (SEQ ID NO: 16), AtUGT1 (SEQ ID NO: 18), AtUGT2 (SEQ ID NO: 20), PtUGT (SEQ ID NO: 22), and RrUGT (SEQ ID NO: 24).
[0075] Structural analysis of AsUGT and comparison of active site residues: To better understand the residues that contribute to GeUGT substrate binding and higher activity, we investigated the structural model of AsUGT. After comparing the sequences of GeUGT (SEQ ID NO: 2) and AsUGT (SEQ ID NO: 26), we identified GeUGT active site residues M88, F119, I139, F145, L149, M198, and F383, corresponding to I84, L115, Y135, C141, F145, L194, and Y382, respectively, in AsUGT. These residues were also found to be unique to GeUGT compared to other identified gastrodin synthase enzymes described herein (Figure 4). These active site differences directly surround the putative binding site for 4-hydroxybenzyl alcohol and may increase the binding affinity of GeUGT for 4-hydroxybenzyl alcohol, resulting in the observed increased rate of gastrodin synthesis by GeUGT.
[0076] Fermentation conditions: To investigate the scalability of gastrodin production, E. coli C transformed with a plasmid carrying GeUGT under the pL promoter was cultivated in a 1 L bioreactor. The culture was fed with glucose and 4-hydroxybenzyl alcohol for 86 hours. A total of 38 g of 4-hydroxybenzyl alcohol was fed to the culture, resulting in a final gastrodin titer of 48.8 g / L.
[0077] Consensus sequence: A consensus nucleotide sequence was calculated using all nucleotide sequences and using the MUSCLE alignment algorithm.
[0078] Results and Discussion Gastrodin is derived from the plant Gastrodia elata, but the enzyme from this plant has not been described. Therefore, we investigated potential enzymes from this plant by analyzing existing transcriptome data. Transcriptome analysis of G. elata enabled the search for UGTs similar to the known gastrodin synthase AsUGT. Herein, the present disclosure describes the use of a native Gastrodia UGT enzyme (GeUGT, SEQ ID NO: 2) to efficiently convert 4-hydroxybenzyl alcohol to gastrodin using an E. coli host (Figure 1). Furthermore, avoiding the previously used initial biosynthetic step favors a single-step biotransformation process in which 4-hydroxybenzyl alcohol is directly fed to a microorganism expressing a single UGT gene.
[0079] This enzyme has not been previously described in the literature. This disclosure describes the cloning and use of GeUGT to produce gastrodin at high titers using a biotransformation approach by feeding 4-hydroxybenzyl alcohol. Additionally, the following 11 UGT enzymes have been identified and described herein that have not previously been described as enzymes that convert 4-hydroxybenzyl alcohol to gastrodin: SiUGT (SEQ ID NO: 4), CaUGT (SEQ ID NO: 6), PpUGT (SEQ ID NO: 8), NbUGT1 (SEQ ID NO: 10), NbUGT2 (SEQ ID NO: 12), PcUGT (SEQ ID NO: 14), WsUGT (SEQ ID NO: 16), AtUGT1 (SEQ ID NO: 18), AtUGT2 (SEQ ID NO: 20), PtUGT (SEQ ID NO: 22), and RrUGT (SEQ ID NO: 24).
[0080] From this search, five candidate sequences were selected, transformed into E. coli C, and tested for enzymatic activity. One of these sequences was found to have significant 4-hydroxybenzyl alcohol UGT activity and was designated GeUGT (SEQ ID NO: 2). To further explore additional enzymes capable of producing gastrodin, BLAST searches of these sequences were performed against publicly available sequences in the NCBI database. Fifteen additional enzymes were selected and tested. Of these enzymes, 11 were able to produce gastrodin with varying efficiencies (SEQ ID NOs: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, and 24 in Figure 4). None of the 11 enzymes have previously been reported as UGTs capable of converting 4-hydroxybenzyl alcohol to gastrodin.
[0081] After evaluating the gastrodin production efficiency, it was clear that GeUGT was the most efficient at converting 4-hydroxybenzyl alcohol to gastrodin, with no 4-hydroxybenzyl alcohol remaining after 48 h. The enzyme from R. rosea also converted all 4-hydroxybenzyl alcohol to gastrodin, but GeUGT showed a higher conversion rate at 24 h after inoculation.
[0082] Example 2 To understand the kinetics of GeUGT, we investigated the biochemical basis of this activity. Due to the lack of an experimental crystal structure for AsUGT, we generated PyMOL 3D visualizations of the GeUGT active site (Figure 5A) and the AsUGT active site (Figure 5B) using publicly available predicted structures of AsUGT from the Alphafold protein structure database. This allowed us to view the active site surrounding the region where UDP-glucose is bound in greater detail. By overlaying this model with a closely related experimental structure containing UDP-glucose (2VCE), we were able to approximate the location of UDP-glucose in AsUGT. This is because UDP-glucose binding is well conserved within the GT1 superfamily of glycosyltransferases. Finally, by modifying the model in PyMOL, we converted all active site residues in AsUGT to their corresponding GeUGT residues based on sequence alignment.
[0083] Clear differences in the active site structure were observed, providing an answer to the question of why GeUGT is much more efficient than AsUGT in producing gastrodin. To biosynthesize gastrodin, the phenolic oxygen of hydroxybenzyl alcohol must be close to the anomeric carbon of UDP-glucose. In the case of 4-hydroxybenzyl alcohol, this means that the hydroxybenzyl moiety must also be close to UDP-glucose. In the GeUGT model, a concentration of aromatic side chains close to the UDP-glucose molecule was observed, particularly at F119, F145, and F383. Notably, in AsUGT, L115 and C141 (corresponding to F119 and F145 in GeUGT) are residues close to the UDP-glucose molecule, which is not ideal for ensuring the hydroxybenzyl moiety is attached in the correct position. While AsUGTs are known to be broad-spectrum UGTs, GeUGTs are likely specifically tuned to glycosylate 4-hydroxybenzyl alcohol, potentially via these aromatic residues clustered near the UDP-glucose binding site. While phenylalanine (F) was identified at positions 119, 145, and 383, other amino acids with aromatic hydrophobic side chains (e.g., tyrosine (Y) and / or tryptophan (W)) are contemplated by the present disclosure. Similarly, the present disclosure provides for the use of amino acids with hydrophobic side chains, as well as other amino acids with similar chemical properties (e.g., leucine (L) and / or methionine (M)).
[0084] The active site variability was obtained from an alignment of all UGTs identified in Example 1. At each position, differences between AsUGT and GeUGT were observed. All other sequences were then examined to see if the corresponding residues were more closely related to AsUGT or GeUGT. In some cases, other UGTs have residues similar to GeUGT but different from AsUGT, and therefore some of the residues have additional substitutions.
[0085] Incorporation by Reference The teachings of all patents, published applications, and references cited herein are incorporated by reference in their entirety.
[0086] While exemplary embodiments have been particularly shown and described, it will be understood by those skilled in the art that various changes in form and detail can be made therein without departing from the scope of the embodiments encompassed by the appended claims.
Claims
1. 1. A host cell comprising a transgene encoding a heterologous uridine 5′-diphospho-glucosyltransferase (UGT) operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least about 75% amino acid sequence identity to SEQ ID NO:2, wherein the amino acid sequence of the heterologous UGT comprises: a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; A host cell comprising one or more of:
2. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; The host cell of claim 1 , comprising two or more of:
3. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; The host cell of claim 1 , comprising three or more of:
4. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; The host cell of claim 1 , comprising four or more of:
5. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; The host cell of claim 1 , comprising five or more of:
6. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; The host cell of claim 1 , comprising:
7. 2. The host cell of claim 1 , wherein the amino acid sequence of the heterologous UGT comprises an M at position 88 of SEQ ID NO:
2.
8. 2. The host cell of claim 1 , wherein the amino acid sequence of the heterologous UGT comprises an F at position 119 of SEQ ID NO:
2.
9. 2. The host cell of claim 1 , wherein the amino acid sequence of the heterologous UGT comprises a Y at position 119 of SEQ ID NO:
2.
10. 2. The host cell of claim 1 , wherein the amino acid sequence of the heterologous UGT comprises a W at position 119 of SEQ ID NO:
2.
11. 2. The host cell of claim 1, wherein the amino acid sequence of the heterologous UGT comprises an F at position 145 of SEQ ID NO:
2.
12. 2. The host cell of claim 1, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 145 of SEQ ID NO:
2.
13. 2. The host cell of claim 1, wherein the amino acid sequence of the heterologous UGT comprises a W at position 145 of SEQ ID NO:
2.
14. 2. The host cell of claim 1, wherein the amino acid sequence of the heterologous UGT comprises an L at position 149 of SEQ ID NO:
2.
15. 2. The host cell of claim 1, wherein the amino acid sequence of the heterologous UGT comprises an M at position 149 of SEQ ID NO:
2.
16. 2. The host cell of claim 1, wherein the amino acid sequence of the heterologous UGT comprises an M at position 198 of SEQ ID NO:
2.
17. 2. The host cell of claim 1, wherein the amino acid sequence of the heterologous UGT comprises an F at position 383 of SEQ ID NO:
2.
18. 18. The host cell of any one of claims 1 to 17, wherein the heterologous UGT comprises an amino acid sequence having at least about 78% amino acid sequence identity to SEQ ID NO:
2.
19. 18. The host cell of any one of claims 1 to 17, wherein the heterologous UGT comprises an amino acid sequence having at least about 80% amino acid sequence identity to SEQ ID NO:
2.
20. 18. The host cell of any one of claims 1 to 17, wherein the heterologous UGT comprises an amino acid sequence having at least about 85% amino acid sequence identity to SEQ ID NO:
2.
21. 18. The host cell of any one of claims 1 to 17, wherein the heterologous UGT comprises an amino acid sequence having at least about 90% amino acid sequence identity to SEQ ID NO:
2.
22. 18. The host cell of any one of claims 1 to 17, wherein the heterologous UGT comprises an amino acid sequence having at least about 95% amino acid sequence identity to SEQ ID NO:
2.
23. 18. The host cell of any one of claims 1 to 17, wherein the heterologous UGT comprises an amino acid sequence having at least about 97% amino acid sequence identity to SEQ ID NO:
2.
24. 18. The host cell of any one of claims 1 to 17, wherein the heterologous UGT comprises SEQ ID NO:
2.
25. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 25. The host cell of claim 1 , wherein the host cell does not have one or more of the following:
26. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 25. The host cell of claim 1 , wherein the host cell does not have two or more of:
27. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 25. The host cell of claim 1 , wherein the host cell does not have three or more of:
28. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 25. The host cell of claim 1 , wherein the host cell does not have four or more of:
29. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 25. The host cell of claim 1 , wherein the host cell does not have five or more of the following:
30. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 25. The host cell of claim 1 , wherein the host cell does not have:
31. 31. The host cell of any one of claims 1 to 30, wherein the host cell is a plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
32. 32. The host cell of any one of claims 1 to 31, wherein the heterologous UGT is codon-optimized for expression in the host cell.
33. 1. A host cell comprising a transgene encoding a heterologous uridine 5′-diphospho-glucosyltransferase (UGT) operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least about 75% amino acid sequence identity to SEQ ID NO:28, wherein the amino acid sequence of the heterologous UGT comprises: a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; A host cell comprising one or more of:
34. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 34. The host cell of claim 33, comprising two or more of:
35. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 34. The host cell of claim 33, comprising three or more of:
36. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 34. The host cell of claim 33, comprising four or more of:
37. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 34. The host cell of claim 33, comprising five or more of:
38. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 34. The host cell of claim 33, comprising:
39. 34. The host cell of claim 33, wherein the amino acid sequence of the heterologous UGT comprises an M at position 93 of SEQ ID NO:
28.
40. 34. The host cell of claim 33, wherein the amino acid sequence of the heterologous UGT comprises an F at position 129 of SEQ ID NO:
28.
41. 34. The host cell of Claim 33, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 129 of SEQ ID NO:
28.
42. 34. The host cell of Claim 33, wherein the amino acid sequence of the heterologous UGT comprises a W at position 129 of SEQ ID NO:
28.
43. 34. The host cell of claim 33, wherein the amino acid sequence of the heterologous UGT comprises an F at position 150 of SEQ ID NO:
28.
44. 34. The host cell of Claim 33, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 150 of SEQ ID NO:
28.
45. 34. The host cell of Claim 33, wherein the amino acid sequence of the heterologous UGT comprises a W at position 150 of SEQ ID NO:
28.
46. 34. The host cell of claim 33, wherein the amino acid sequence of the heterologous UGT comprises an L at position 154 of SEQ ID NO:
28.
47. 34. The host cell of claim 33, wherein the amino acid sequence of the heterologous UGT comprises an M at position 154 of SEQ ID NO:
28.
48. 34. The host cell of claim 33, wherein the amino acid sequence of the heterologous UGT comprises an M at position 203 of SEQ ID NO:
28.
49. 34. The host cell of Claim 33, wherein the amino acid sequence of the heterologous UGT comprises an F at position 391 of SEQ ID NO:
28.
50. 50. The host cell of any one of claims 33-49, wherein the heterologous UGT comprises an amino acid sequence having at least about 78% amino acid sequence identity to SEQ ID NO:
28.
51. 50. The host cell of any one of claims 33-49, wherein the heterologous UGT comprises an amino acid sequence having at least about 80% amino acid sequence identity to SEQ ID NO:
28.
52. 50. The host cell of any one of claims 33-49, wherein the heterologous UGT comprises an amino acid sequence having at least about 85% amino acid sequence identity to SEQ ID NO:
28.
53. 50. The host cell of any one of claims 33-49, wherein the heterologous UGT comprises an amino acid sequence having at least about 90% amino acid sequence identity to SEQ ID NO:
28.
54. 50. The host cell of any one of claims 33-49, wherein the heterologous UGT comprises an amino acid sequence having at least about 95% amino acid sequence identity to SEQ ID NO:
28.
55. 50. The host cell of any one of claims 33-49, wherein the heterologous UGT comprises an amino acid sequence having at least about 97% amino acid sequence identity to SEQ ID NO:
28.
56. 50. The host cell of any one of claims 33 to 49, wherein the heterologous UGT comprises SEQ ID NO:
28.
57. 57. The host cell of any one of claims 33 to 56, wherein the host cell is a plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
58. 58. The host cell of any one of claims 33-57, wherein the heterologous UGT is codon-optimized for expression in the host cell.
59. 1. A method for producing gastrodin in a host cell, comprising culturing the host cell in a cell culture medium comprising 4-hydroxybenzyl alcohol, wherein the host cell expresses a transgene encoding a heterologous uridine 5′-diphospho-glucosyltransferase (UGT) operably linked to a promoter, the heterologous UGT comprising an amino acid sequence having at least about 75% amino acid sequence identity to SEQ ID NO:2, wherein the amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; The method includes one or more of:
60. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 60. The method of claim 59, comprising two or more of:
61. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 60. The method of claim 59, comprising three or more of:
62. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 60. The method of claim 59, comprising four or more of:
63. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 60. The method of claim 59, comprising five or more of:
64. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 60. The method of claim 59, comprising:
65. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises an M at position 88 of SEQ ID NO:
2.
66. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises an F at position 119 of SEQ ID NO:
2.
67. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 119 of SEQ ID NO:
2.
68. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises a W at position 119 of SEQ ID NO:
2.
69. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises an F at position 145 of SEQ ID NO:
2.
70. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 145 of SEQ ID NO:
2.
71. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises a W at position 145 of SEQ ID NO:
2.
72. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises an L at position 149 of SEQ ID NO:
2.
73. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises an M at position 149 of SEQ ID NO:
2.
74. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises an M at position 198 of SEQ ID NO:
2.
75. 60. The method of Claim 59, wherein the amino acid sequence of the heterologous UGT comprises an F at position 383 of SEQ ID NO:
2.
76. 76. The method of any one of claims 59-75, wherein the heterologous UGT comprises an amino acid sequence having at least about 78% amino acid sequence identity to SEQ ID NO:
2.
77. 76. The method of any one of claims 59-75, wherein the heterologous UGT comprises an amino acid sequence having at least about 80% amino acid sequence identity to SEQ ID NO:
2.
78. 76. The method of any one of claims 59-75, wherein the heterologous UGT comprises an amino acid sequence having at least about 85% amino acid sequence identity to SEQ ID NO:
2.
79. 76. The method of any one of claims 59-75, wherein the heterologous UGT comprises an amino acid sequence having at least about 90% amino acid sequence identity to SEQ ID NO:
2.
80. 76. The method of any one of claims 59-75, wherein the heterologous UGT comprises an amino acid sequence having at least about 95% amino acid sequence identity to SEQ ID NO:
2.
81. 76. The method of any one of claims 59-75, wherein the heterologous UGT comprises an amino acid sequence having at least about 97% amino acid sequence identity to SEQ ID NO:
2.
82. 76. The method of any one of claims 59-75, wherein the heterologous UGT comprises SEQ ID NO:
2.
83. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 83. The method of any one of claims 59 to 82, wherein the method does not have one or more of:
84. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 83. The method of any one of claims 59 to 82, wherein the method does not have two or more of:
85. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 83. The method of any one of claims 59 to 82, wherein the method does not have three or more of:
86. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 83. The method of any one of claims 59 to 82, wherein the method does not have four or more of:
87. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 83. The method of any one of claims 59 to 82, wherein no more than five of:
88. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 83. The method of any one of claims 59 to 82, wherein the method is free of
89. 89. The method of any one of claims 59 to 88, wherein the host cell is a plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
90. 90. The method of any one of claims 59-89, wherein the heterologous UGT is codon-optimized for expression in the host cell.
91. 1. A method for producing gastrodin in a host cell, comprising culturing the host cell in a cell culture medium comprising 4-hydroxybenzyl alcohol, wherein the host cell expresses a transgene encoding a heterologous uridine 5′-diphospho-glucosyltransferase (UGT) operably linked to a promoter, the heterologous UGT comprising an amino acid sequence having at least about 75% amino acid sequence identity to SEQ ID NO:28, wherein the amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; The method includes one or more of:
92. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 92. The method of claim 91, comprising two or more of:
93. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 92. The method of claim 91, comprising three or more of:
94. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 92. The method of claim 91, comprising four or more of:
95. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 92. The method of claim 91, comprising five or more of:
96. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 92. The method of claim 91, comprising:
97. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises an M at position 93 of SEQ ID NO:
28.
98. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises an F at position 129 of SEQ ID NO:
28.
99. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 129 of SEQ ID NO:
28.
100. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises a W at position 129 of SEQ ID NO:
28.
101. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises an F at position 150 of SEQ ID NO:
28.
102. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 150 of SEQ ID NO:
28.
103. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises a W at position 150 of SEQ ID NO:
28.
104. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises an L at position 154 of SEQ ID NO:
28.
105. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises an M at position 154 of SEQ ID NO:
28.
106. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises an M at position 203 of SEQ ID NO:
28.
107. 92. The method of Claim 91, wherein the amino acid sequence of the heterologous UGT comprises an F at position 391 of SEQ ID NO:
28.
108. 108. The method of any one of claims 91-107, wherein the heterologous UGT comprises an amino acid sequence having at least about 78% amino acid sequence identity to SEQ ID NO:
28.
109. 108. The method of any one of claims 91-107, wherein the heterologous UGT comprises an amino acid sequence having at least about 80% amino acid sequence identity to SEQ ID NO:
28.
110. 108. The method of any one of claims 91-107, wherein the heterologous UGT comprises an amino acid sequence having at least about 85% amino acid sequence identity to SEQ ID NO:
28.
111. 108. The method of any one of claims 91-107, wherein the heterologous UGT comprises an amino acid sequence having at least about 90% amino acid sequence identity to SEQ ID NO:
28.
112. 108. The method of any one of claims 91-107, wherein the heterologous UGT comprises an amino acid sequence having at least about 95% amino acid sequence identity to SEQ ID NO:
28.
113. 108. The method of any one of claims 91-107, wherein the heterologous UGT comprises an amino acid sequence having at least about 97% amino acid sequence identity to SEQ ID NO:
28.
114. 108. The method of any one of claims 91-107, wherein the heterologous UGT comprises SEQ ID NO:
28.
115. 115. The method of any one of claims 91 to 114, wherein the host cell is a plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
116. 116. The method of any one of claims 91-115, wherein the heterologous UGT is codon-optimized for expression in the host cell.
117. 117. The method of any one of claims 59 to 116, wherein the cell culture medium further comprises glucose.
118. 118. The method of any one of claims 59 to 117, further comprising generating the host cell, said method comprising introducing a vector into the host cell, said vector comprising a nucleic acid encoding the heterologous uridine 5'-diphospho-glucosyltransferase (UGT) operably linked to the promoter.
119. 119. The method of any one of claims 59 to 118, wherein the gastrodin is extracted using maceration, percolation, decoction, reflux extraction, Soxhlet extraction, pressurized liquid extraction, supercritical fluid extraction, ultrasound-assisted extraction, pulsed field extraction, enzyme-assisted extraction, steam distillation, steam distillation, or any combination thereof.
120. 120. The method of any one of claims 59 to 119, wherein the concentration of gastrodin in the cell culture medium after 24 hours of incubation is at least about 4 mM, 6 mM, or 8 mM.
121. 121. The method of any one of claims 59 to 120, wherein the concentration of 4-hydroxybenzyl alcohol in the cell culture medium after 24 hours of incubation is no more than 2 mM, 1.5 mM, or 1 mM.
122. A vector comprising a nucleic acid encoding a gastrodin synthase for converting 4-hydroxybenzyl alcohol to gastrodin, wherein the gastrodin synthase has at least about 75% amino acid sequence identity with SEQ ID NO:
2.
123. The vector of claim 122, wherein the nucleic acid encoding gastrodin synthase for converting 4-hydroxybenzyl alcohol to gastrodin has at least about 95% nucleotide sequence identity to SEQ ID NO:
1.
124. 124. The vector of claim 122 or 123, wherein the gastrodin synthase is derived from a plant.
125. The vector of claim 124, wherein the plant is an orchid.
126. 126. The vector of claim 125, wherein the plant is Gastrodia elata.
127. The amino acid sequence of the gastrodin synthase is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 127. The vector of any one of claims 122 to 126, comprising one or more of:
128. The amino acid sequence of the gastrodin synthase is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 127. The vector of any one of claims 122 to 126, comprising two or more of:
129. The amino acid sequence of the gastrodin synthase is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 127. The vector of any one of claims 122 to 126, comprising three or more of:
130. The amino acid sequence of the gastrodin synthase is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 127. The vector of any one of claims 122 to 126, comprising four or more of:
131. The amino acid sequence of the gastrodin synthase is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 127. The vector of any one of claims 122 to 126, comprising five or more of:
132. The amino acid sequence of the gastrodin synthase is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 127. The vector of any one of claims 122 to 126, comprising:
133. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises M at position 88 of SEQ ID NO:
2.
134. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises F at position 119 of SEQ ID NO:
2.
135. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises Y at position 119 of SEQ ID NO:
2.
136. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises W at position 119 of SEQ ID NO:
2.
137. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises F at position 145 of SEQ ID NO:
2.
138. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises Y at position 145 of SEQ ID NO:
2.
139. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises W at position 145 of SEQ ID NO:
2.
140. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises an L at position 149 of SEQ ID NO:
2.
141. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises M at position 149 of SEQ ID NO:
2.
142. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises M at position 198 of SEQ ID NO:
2.
143. 127. The vector of any one of claims 122 to 126, wherein the amino acid sequence of the gastrodin synthase comprises F at position 383 of SEQ ID NO:
2.
144. 144. The vector of any one of claims 122 to 143, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 78% amino acid sequence identity to SEQ ID NO:
2.
145. 144. The vector of any one of claims 122 to 143, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 80% amino acid sequence identity to SEQ ID NO:
2.
146. 144. The vector of any one of claims 122 to 143, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 85% amino acid sequence identity to SEQ ID NO:
2.
147. 144. The vector of any one of claims 122 to 143, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 90% amino acid sequence identity to SEQ ID NO:
2.
148. 144. The vector of any one of claims 122 to 143, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 95% amino acid sequence identity to SEQ ID NO:
2.
149. 144. The vector of any one of claims 122 to 143, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 97% amino acid sequence identity to SEQ ID NO:
2.
150. 144. The vector of any one of claims 122 to 143, wherein the amino acid sequence of the gastrodin synthase comprises SEQ ID NO:
2.
151. The amino acid sequence of the gastrodin synthase comprises the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 151. The vector of any one of claims 122 to 150, wherein the vector does not have one or more of:
152. The amino acid sequence of the gastrodin synthase comprises the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 151. The vector of any one of claims 122 to 150, wherein the vector does not have two or more of:
153. The amino acid sequence of the gastrodin synthase comprises the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 151. The vector of any one of claims 122 to 150, wherein the vector does not have three or more of:
154. The amino acid sequence of the gastrodin synthase comprises the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 151. The vector of any one of claims 122 to 150, wherein the vector does not have four or more of:
155. The amino acid sequence of the gastrodin synthase comprises the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 151. The vector of any one of claims 122 to 150, wherein the vector does not have five or more of:
156. The amino acid sequence of the gastrodin synthase is a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 151. The vector of any one of claims 122 to 150, wherein the vector does not have
157. 157. The vector of any one of claims 122 to 156, wherein the host cell is a plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
158. 158. The vector of any one of claims 122 to 157, wherein the amino acid sequence of the gastrodin synthase is codon-optimized for expression in the host cell.
159. A vector comprising a nucleic acid encoding a gastrodin synthase for converting 4-hydroxybenzyl alcohol to gastrodin, wherein the gastrodin synthase has at least about 75% amino acid sequence identity with SEQ ID NO:
28.
160. 160. The vector of claim 159, wherein the nucleic acid encoding a gastrodin synthase for converting 4-hydroxybenzyl alcohol to gastrodin has at least about 95% nucleotide sequence identity to SEQ ID NO:
27.
161. The amino acid sequence of the gastrodin synthase is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 161. The vector of any one of claims 159 to 160, comprising one or more of:
162. The amino acid sequence of the gastrodin synthase is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 161. The vector of any one of claims 159 to 160, comprising two or more of:
163. The amino acid sequence of the gastrodin synthase is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 161. The vector of any one of claims 159 to 160, comprising three or more of:
164. The amino acid sequence of the gastrodin synthase is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 161. The vector of any one of claims 159 to 160, comprising four or more of:
165. The amino acid sequence of the gastrodin synthase is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 161. The vector of any one of claims 159 to 160, comprising five or more of:
166. The amino acid sequence of the gastrodin synthase is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 161. The vector of any one of claims 159 to 160, comprising:
167. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises M at position 93 of SEQ ID NO:
28.
168. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises F at position 129 of SEQ ID NO:
28.
169. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises Y at position 129 of SEQ ID NO:
28.
170. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises W at position 129 of SEQ ID NO:
28.
171. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises F at position 150 of SEQ ID NO:
28.
172. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises Y at position 150 of SEQ ID NO:
28.
173. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises W at position 150 of SEQ ID NO:
28.
174. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises an L at position 154 of SEQ ID NO:
28.
175. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises M at position 154 of SEQ ID NO:
28.
176. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises M at position 203 of SEQ ID NO:
28.
177. 161. The vector of any one of claims 159 to 160, wherein the amino acid sequence of the gastrodin synthase comprises F at position 391 of SEQ ID NO:
28.
178. 178. The vector of any one of claims 159 to 177, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 78% amino acid sequence identity to SEQ ID NO:
28.
179. 178. The vector of any one of claims 159 to 177, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 80% amino acid sequence identity to SEQ ID NO:
28.
180. 178. The vector of any one of claims 159 to 177, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 85% amino acid sequence identity to SEQ ID NO:
28.
181. 178. The vector of any one of claims 159 to 177, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 90% amino acid sequence identity to SEQ ID NO:
28.
182. 178. The vector of any one of claims 159 to 177, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 95% amino acid sequence identity to SEQ ID NO:
28.
183. 178. The vector of any one of claims 159 to 177, wherein the amino acid sequence of the gastrodin synthase comprises an amino acid sequence having at least about 97% amino acid sequence identity to SEQ ID NO:
28.
184. 178. The vector of any one of claims 159 to 177, wherein the amino acid sequence of the gastrodin synthase comprises SEQ ID NO:
28.
185. 185. The method of any one of claims 159 to 184, wherein the host cell is a plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
186. 186. The method of any one of claims 159-185, wherein the heterologous UGT is codon-optimized for expression in the host cell.
187. 1. A method for producing a genetically modified host cell, comprising introducing a vector into a host cell, the vector comprising a nucleic acid encoding a heterologous uridine 5′-diphospho-glucosyltransferase (UGT) operably linked to a promoter, the heterologous UGT comprising an amino acid sequence having at least 75% amino acid sequence identity to SEQ ID NO:2, wherein the amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; The method includes one or more of:
188. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 188. The method of claim 187, comprising two or more of:
189. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 188. The method of claim 187, comprising three or more of:
190. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 188. The method of claim 187, comprising four or more of:
191. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 188. The method of claim 187, comprising five or more of:
192. The amino acid sequence of the heterologous UGT is a) M at position 88 of SEQ ID NO:2; b) F, Y, or W at position 119 of SEQ ID NO:2; c) F, Y, or W at position 145 of SEQ ID NO:2; d) L or M at position 149 of SEQ ID NO: 2; e) M at position 198 of SEQ ID NO:2, and f) F at position 383 of SEQ ID NO: 2; 188. The method of claim 187, comprising:
193. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises an M at position 88 of SEQ ID NO:
2.
194. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises an F at position 119 of SEQ ID NO:
2.
195. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 119 of SEQ ID NO:
2.
196. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises a W at position 119 of SEQ ID NO:
2.
197. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises an F at position 145 of SEQ ID NO:
2.
198. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 145 of SEQ ID NO:
2.
199. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises a W at position 145 of SEQ ID NO:
2.
200. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises an L at position 149 of SEQ ID NO:
2.
201. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises an M at position 149 of SEQ ID NO:
2.
202. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises an M at position 198 of SEQ ID NO:
2.
203. 188. The method of claim 187, wherein the amino acid sequence of the heterologous UGT comprises an F at position 383 of SEQ ID NO:
2.
204. 204. The method of any one of claims 187-203, wherein the heterologous UGT comprises an amino acid sequence having at least about 78% amino acid sequence identity to SEQ ID NO:
2.
205. 204. The method of any one of claims 187-203, wherein the heterologous UGT comprises an amino acid sequence having at least about 80% amino acid sequence identity to SEQ ID NO:
2.
206. 204. The method of any one of claims 187-203, wherein the heterologous UGT comprises an amino acid sequence having at least about 85% amino acid sequence identity to SEQ ID NO:
2.
207. 204. The method of any one of claims 187-203, wherein the heterologous UGT comprises an amino acid sequence having at least about 90% amino acid sequence identity to SEQ ID NO:
2.
208. 204. The method of any one of claims 187-203, wherein the heterologous UGT comprises an amino acid sequence having at least about 95% amino acid sequence identity to SEQ ID NO:
2.
209. 204. The method of any one of claims 187 to 203, wherein the heterologous UGT comprises SEQ ID NO:
2.
210. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 210. The method of any one of claims 187 to 209, wherein the method does not have one or more of:
211. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 210. The method of any one of claims 187 to 209, wherein the method does not have two or more of:
212. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 210. The method of any one of claims 187 to 209, wherein the method does not have three or more of:
213. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 210. The method of any one of claims 187 to 209, wherein the method does not have four or more of:
214. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 210. The method of any one of claims 187 to 209, wherein the method does not have five or more of:
215. The heterologous UGT comprises an amino acid sequence, wherein the amino acid sequence consists of the following residues: a) I at position 88 of SEQ ID NO:2, b) L at position 119 of SEQ ID NO:2; c) C at position 145 of SEQ ID NO:2; d) F at position 149 of SEQ ID NO:2; e) L at position 198 of SEQ ID NO:2, or f) Y at position 383 of SEQ ID NO: 2; 210. The method of any one of claims 187 to 209, wherein the method does not have
216. 216. The method of any one of claims 187 to 215, wherein the host cell is a plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
217. 217. The method of any one of claims 187-216, wherein the heterologous UGT is codon-optimized for expression in the host cell.
218. 1. A method of producing a genetically modified host cell, comprising introducing a vector into a host cell, wherein the vector comprises a nucleic acid encoding a heterologous uridine 5′-diphospho-glucosyltransferase (UGT) operably linked to a promoter, wherein the heterologous UGT comprises an amino acid sequence having at least 75% amino acid sequence identity to SEQ ID NO:28, wherein the amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; The method includes one or more of:
219. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 219. The method of claim 218, comprising two or more of:
220. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 219. The method of claim 218, comprising three or more of:
221. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 219. The method of claim 218, comprising four or more of:
222. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 219. The method of claim 218, comprising five or more of:
223. The amino acid sequence of the heterologous UGT is a) M at position 93 of SEQ ID NO: 28; b) F, Y, or W at position 129 of SEQ ID NO: 28; c) F, Y, or W at position 150 of SEQ ID NO: 28; d) L or M at position 154 of SEQ ID NO: 28; e) M at position 203 of SEQ ID NO: 28, and f) F at position 391 of SEQ ID NO: 28; 219. The method of claim 218, comprising:
224. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises an M at position 93 of SEQ ID NO:
28.
225. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises an F at position 129 of SEQ ID NO:
28.
226. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 129 of SEQ ID NO:
28.
227. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises a W at position 129 of SEQ ID NO:
28.
228. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises an F at position 150 of SEQ ID NO:
28.
229. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises a Y at position 150 of SEQ ID NO:
28.
230. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises a W at position 150 of SEQ ID NO:
28.
231. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises an L at position 154 of SEQ ID NO:
28.
232. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises an M at position 154 of SEQ ID NO:
28.
233. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises an M at position 203 of SEQ ID NO:
28.
234. 219. The method of claim 218, wherein the amino acid sequence of the heterologous UGT comprises an F at position 391 of SEQ ID NO:
28.
235. 235. The method of any one of claims 218-234, wherein the heterologous UGT comprises an amino acid sequence having at least about 78% amino acid sequence identity to SEQ ID NO:
28.
236. 235. The method of any one of claims 218-234, wherein the heterologous UGT comprises an amino acid sequence having at least about 80% amino acid sequence identity to SEQ ID NO:
28.
237. 235. The method of any one of claims 218-234, wherein the heterologous UGT comprises an amino acid sequence having at least about 85% amino acid sequence identity to SEQ ID NO:
28.
238. 235. The method of any one of claims 218-234, wherein the heterologous UGT comprises an amino acid sequence having at least about 90% amino acid sequence identity to SEQ ID NO:
28.
239. 235. The method of any one of claims 218-234, wherein the heterologous UGT comprises an amino acid sequence having at least about 95% amino acid sequence identity to SEQ ID NO:
28.
240. 235. The method of any one of claims 218-234, wherein the heterologous UGT comprises an amino acid sequence having at least about 97% amino acid sequence identity to SEQ ID NO:
28.
241. 235. The method of any one of claims 218-234, wherein the heterologous UGT comprises SEQ ID NO:
28.
242. 242. The method of any one of claims 218 to 241, wherein the host cell is a plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.
243. 243. The method of any one of claims 218-242, wherein the heterologous UGT is codon-optimized for expression in the host cell.
244. A pharmaceutical composition comprising a gastrodin, wherein the gastrodin is produced by a genetically modified plant, or plant cell, a fungal cell, a yeast cell, an insect cell, or a bacterial cell.