The polypeptide beta-1,2-glycosyltransferase was designed, along with a method for converting the sugar component to the steviol glycoside substrate, and related polynucleotides and host microorganisms.
Patent Information
- Authority / Receiving Office
- VN · VN
- Patent Type
- Applications
- Filing Date
- 2024-06-05
- Publication Date
- 2026-06-15
AI Technical Summary
The production of rebaudiosides D and M, which offer improved sweetness and bitterness profiles compared to rebaudioside A, is limited by their low quantities in Stevia rebaudiana leaves and the high cost of native glycosyltransferases that use UDP-glucose as a sugar donor, making large-scale production economically and efficiently challenging.
Engineered beta-1,2-glycosyltransferases (B12GTs) that are thermostable and capable of using ADP-glucose as a sugar donor, allowing for the conversion of lower-order steviol glycosides like stevioside and rebaudioside A to higher-order glycosides such as rebaudioside D and M, with a sucrose synthase recycling system for efficient sugar donor cofactor utilization.
The engineered enzymes reduce the amount of enzyme required and increase flexibility in reaction conditions, enabling cost-effective and efficient production of rebaudioside-based sweeteners by enhancing the conversion of lower-order to higher-order steviol glycosides, thus addressing the economic and efficiency limitations of native enzyme systems.
Smart Images

Figure VN1202509821_0
Abstract
Description
COMPOSITIONS AND METHODS FOR PRODUCING REBAUDIOSIDE D ANDREBAUDIOSIDE MCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 506,566, filed June 6, 2023, the disclosure of which is incorporated herein by reference in its entirety. This application is related to U.S. Patent Application No. 18 / 546,881, filed August 14, 2023, and International Patent Application No. PCT / US2023 / 073344, filed September 1, 2023, the dis- closure of each of which is incorporated herein by reference in its entirety.INCORPORATION OF THE SEQUENCE LISTING
[0002] The contents of the electronic sequence listing (ARZE_040_01WO_Se- qList_ST26.xml; Size: 1,309,826 bytes; and Date of Creation: May 29, 2024) are herein incor- porated by reference in its entirety.FIELD OF THE DISCLOSURE
[0003] The present disclosure relates to enzymes and biocatalytic processes for producing ste- viol glycosides. The present disclosure particularly relates to use of glycosyltransferases that can transfer a glucose moiety from an ADP-glucose sugar donor to steviol glycosides.BACKGROUND
[0004] Excess sugar consumption has been linked to worldwide health epidemics including diabetes and heart disease. Healthcare systems incur exorbitant costs associated with treating these diseases. Replacing added sugar in food with a low calorie, high-intensity sweetener would have significant health and economic impact.
[0005] The species Stevia rebaudiana is commonly grown for its sweet leaves, which have traditionally been used as a sweetener. Stevia extract is 200-300 times sweeter than sugar and is used commercially as a high intensity sweetener. The main glycoside components of stevia leaf are steviosides and rebaudiosides. Over ten different steviol glycosides are present in ap- preciable quantities in the leaf. The principal sweetening compounds are stevioside and rebau- dioside A. Rebaudioside A (Reb A) is considered a higher value compared to stevioside be- cause of its increased sweetness and decreased bitterness.
[0006] The sweetness and bitterness profiles of rebaudioside D (Reb D) and rebaudioside M (Reb M) are improved compared to Reb A. However, Reb D and Reb M are present at verylow quantities in the stevia leaf. Reb D and Reb M can be made by the addition of one or two glucose molecules to Reb A, respectively. Native glycosyltransferases that make Reb D and Reb M use UDP-glucose as the source for transferring glucose to the lower-order steviol gly- cosides.BRIEF SUMMARY
[0007] The present disclosure provides engineered glycosyltransferases designed to exhibit high thermostability while increasing expression and conversion of lower order steviol glyco- sides, such as steviol and rebaudioside A, to higher order steviol glycosides, such as rebaudi- oside M and rebaudioside D. As a result, the engineered glycosyltransferases of the present disclosure help to make rebaudioside-based sweeteners a reality by reducing the amount of enzyme required to produce the product and increasing flexibility in reaction conditions.
[0008] To this end, the present disclosure provides methods to use the engineered glycosyl- transferases to transfer one or more sugar moieties to a substrate steviol glycoside (also referred to herein as a “SG”), thereby performing the conversion from a lower order SG to a higher order SG. Specifically, the disclosed beta-l,2-glycosyltransferases (also referred to herein as “B12GTs”) transfer a glucose to a SG by making a beta- 1,2 glycosidic bond with the first glucose at either the C 13 or C 19 end of the SG. This includes, but is not limited to, the conver- sion of stevioside to rebaudioside E (Reb E) (FIG. 3), rebaudioside A (Reb A) to rebaudioside D (Reb D) (FIG. 2), rebaudioside I (Reb I) to rebaudioside M (Reb M) (FIG. 4) and rebaudi- oside E2 (Reb E2) to rebaudioside AM (Reb AM). The disclosed B12GTs are used to convert stevioside to Reb E and Reb A to Reb D. The B12GTs are used in combination with sucrose synthase (also referred to herein as a “SuSy”) in a one-pot reaction which allows for efficient recycling of the sugar donor cofactor during steviol glycoside conversion. Additionally, the B12GTs may be used in combination with a sucrose synthase and a beta- 1,3 -glycosyltransfer- ase (B13GT) to convert a starting composition of Reb A and stevioside to Reb M.
[0009] Moreover, in contrast to native glycosyltransferases, the disclosure provides glycosyl- transferase polypeptides that can utilize ADP -glucose as the sugar donor for SG conversion. The disclosure provides glycosyltransferase polypeptides that comprise an amino acid se- quence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879. The glycosyltransferase polypeptide may comprise, or consist of, an amino acid sequence selected from the group con- sisting of SEQ ID NOs: 449-879. The polypeptides may comprise one or more peptide tagsused for solubility, expression and / or purification; for example, a polyhistidine tag of between 4 and 10 histidine residues, and preferably 6 histidine residues. Other suitable tags include, but are not limited to, glutathione .S'-transfcrasc (GST), FLAG, maltose binding protein (MBP), calmodulin binding peptide (CBP), and Myc tag. Suitable linkers include, but are not limited to, polypeptides composed of glycine and serine, such as GSGS, polyglycine linkers, EAAAK repeats, and sequences containing cleavage sites for enzymes such as factor Xa, enterokinase, and thrombin.
[0010] Nucleotide sugar donors, including both UDP-glucose and ADP-glucose, are expensive co-substrates and add significant costs to any process that utilizes the compounds. Sucrose synthases (SuSy; EC 2.4.1.13) catalyze the chemical reaction of nucleoside diphosphate (NDP) and sucrose to form NDP-glucose and fructose. Therefore, sucrose synthases can be used to convert an NDP into an NDP-glucose required by B 12GTs (an exemplary glycosyltransferase). The disclosure provides SuSy polypeptides that comprise an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 2-448. Specifically, the disclosed sucrose synthases can convert ADP into the ADP-glucose cofactor required by the disclosed B12GTs.
[0011] The disclosure additionally provides a method to utilize a SuSy ADP-glucose recycling system combined with a B12GT polypeptide in a one-pot reaction to convert Reb A and / or stevioside into Reb D and Reb E, respectively (FIG. 5). In embodiments, the method comprises contacting a stevia leaf extract purified to contain greater than 50% Reb A (RA50), ADP, and sucrose with a Bl, 2 glycosyltransferase and sucrose synthase to produce Reb D and / or Reb E. In embodiments, the method comprises contacting a stevia leaf extract purified to contain greater than 60% Reb A (RA60), ADP, and sucrose with a B12GT and SuSy to produce Reb D and / or Reb E.
[0012] The disclosure additionally provides a method to utilize B12GT, SuSy, and one or more additional glycosyltransferases to make higher-order steviol glycosides. In embodiments, the method comprises contacting a B12GT, a SuSy, and a B13GT with RA50, ADP, and sucrose to make Reb M. In embodiments, the method comprises contacting a B12GT, a SuSy, and a B13GT with RA60, ADP, and sucrose to make Reb M.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are included to provide a further understanding of the dis- closure. The drawings illustrate embodiments of the disclosure and together with the descrip- tion, serve to explain the principles of the embodiments of the disclosure.
[0014] FIG. 1 depicts the steviol glycoside backbone.
[0015] FIG. 2 depicts the conversion (i.e., glycosylation) of rebaudioside A (Reb A) to rebau- dioside D (Reb D).
[0016] FIG. 3 depicts the conversion (i.e., glycosylation) of stevioside to rebaudioside E (Reb E).
[0017] FIG. 4 depicts the conversion (i.e., glycosylation) of rebaudioside I (Reb I) to rebaudi- oside M (Reb M).
[0018] FIG. 5 depicts the reaction scheme to produce Reb D and Reb E by contacting a beta-1.2-glycosyltransferase (B12GT) and sucrose synthase with sucrose, ADP, and a mixture of stevioside and Reb A.
[0019] FIG. 6 depicts an exemplary LCMS chromatogram of the Reb D and Reb E reaction product produced by the disclosed B12GTs.
[0020] FIG. 7 depicts a schematic of the B12GT active site with catalytically important resi- dues shown.
[0021] FIG. 8 depicts an SDS-PAGE gel of clarified E. coli lysate from 1 L fermentations of E. coli strains expressing B12GT designs. Lanes from left to right: Ladder, Ladder, SEQ ID NO: 465, SEQ ID NO: 469, SEQ ID NO: 496, SEQ ID NO: 500. The arrow points to the expressed B12GT enzyme.
[0022] FIG. 9 depicts an annotated version of the amino acid sequence of SEQ ID NO: 1. Bolded residues indicate amino acid positions within the active site and underlined residues indicate amino acid positions outside the active site. The active site is defined by aligning the structural model of SEQ ID NO: 714 from U.S. Patent Application No. 18 / 546,881 to crystal structures of homologous glycosyltransferases from Oryza saliva and Stevia rebaudiana as de- scribed below. Briefly, residues of the structural model within 8 A of a modeled crystal struc- ture substrate or product was defined as an active site residue. Bracketed residues indicate re- gions of the amino acid sequence displaying a secondary structure of a beta strand or an alpha helix.
[0023] FIG. 10 depicts a partial steviol glycoside network showing conversion of stevioside and Reb A to Reb M through additions of glucose monomers catalyzed by a B 12GT and a beta-1.3 -glycosyltransferase (B13GT). For each molecule, the central hexagon represents the steviolglycoside core, circles represent beta- 1,2-linked glucose monomers and squares represent beta- 1,3 -linked glucose monomers.DETAILED DESCRIPTION
[0024] The present disclosure provides enzymes and biocatalytic processes for preparing a composition comprising one or more target steviol glycosides (SG) by contacting a starting composition comprising one or more substrate steviol glycosides, sucrose, and NDP with one or more NDP-glycosyltransferase polypeptides and a sucrose synthase, thereby producing a composition comprising the target steviol glycoside(s) comprising one or more additional glu- cose units than the substrate steviol glycoside(s).
[0025] As used herein, “biocatalysis” or “biocatalytic” refers to the use of natural catalysts, such as protein enzymes, to perform chemical transformations on organic compounds. Bio- catalysis is alternatively known as biotransformation or biosynthesis. Both isolated and whole cell biocatalysis methods are known in the art. Biocatalyst protein enzymes can be naturally occurring or recombinant proteins.
[0026] As used herein, the term “steviol glycoside(s)” refers to a glycoside of steviol, includ- ing, but not limited to, naturally occurring steviol glycosides, e.g. steviol- 13 -O-glucoside, ste- viol- 19-O-glucoside, rubusoside, steviol- 1,2-bioside, steviol- 1,3 -bioside, rubusoside, dulcoside B, dulcoside A, rebaudioside B, rebaudioside G, stevioside, rebaudioside C, rebau- dioside F, rebaudioside A, rebaudioside I, rebaudioside E, rebaudioside E2, rebaudioside H, rebaudioside L, rebaudioside K, rebaudioside J, rebaudioside AM, rebaudioside M, rebaudi- oside D, rebaudioside N, rebaudioside O, rebaudioside Q, rebaudioside X, synthetic steviol glycosides, e.g. enzymatically glucosylated steviol glycosides and combinations thereof.
[0027] As used herein, “starting composition” refers to any composition (generally an aqueous solution) containing one or more steviol glycosides, where the one or more steviol glycosides serve as the substrate for the biotransformation.
[0028] As used herein, the terms “polynucleotide" or “nucleic acid” are used interchangeably, unless indicated by context, and is used to refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides, typically DNA.
[0029] As used herein, "expression" refers to either or both steps, depending on context, of the two-step process by which polynucleotides are transcribed into mRNA and the transcribed mRNA is subsequently translated into polypeptides.
[0030] "Under transcriptional control" means that transcription of a polynucleotide, usually a DNA sequence, depends on its being operatively linked to an element that promotes transcrip- tion.
[0031] "Operatively linked" means that the polynucleotide elements are arranged in a manner that allows them to function in a cell; typically to produce polypeptides in the cell; for example, the disclosure provides promoters operatively linked to the downstream sequences encoding polypeptides.
[0032] The term "encode" refers to the ability of a polynucleotide to produce an mRNA or a polypeptide if it can be transcribed to produce the mRNA and then translated to produce the polypeptide or a fragment thereof. In each case, the polynucleotide is referred to as encoding the mRNA and encoding the polypeptide. The antisense strand is the complement of such a nucleic acid, and the encoding sequence can be deduced therefrom. Similarly, a “coding se- quence” refers to a region of a nucleic acid that encodes an mRNA or a polypeptide.
[0033] The term "promoter" as used herein refers to a control sequence that is a portion of a polynucleotide sequence that controls the initiation and rate of transcription of a coding se- quence. An “enhancer” is a regulatory element that increases the expression of a target se- quence. A "promoter / enhancer" is a polynucleotide with sequences that provide both promoter and enhancer functions.
[0034] The regulatory elements, e.g., enhancers and promoters, may be "homologous" or "het- erologous." A "homologous" regulatory element is one which is naturally linked with a given polynucleotide in the genome; for example, it may be the promoter found natively in the or- ganism upstream of the encoded polypeptide. A "heterologous" regulatory element is one which is placed in juxtaposition to a polynucleotide by means of recombinant molecular bio- logical techniques but is not a combination found in nature. Often, promoters, enhancers and other regulatory elements are heterologous so as to facilitate expression of a polypeptide in a host cell other than one in which a polypeptide naturally occurs. Thus, “heterologous expres- sion”, as used herein, refers to producing an mRNA and / or a polypeptide in a host cell, such as a microorganism, where the polynucleotide is not found naturally, or one or more regulatory elements are not naturally found operably linked to the polynucleotide in the host cell.
[0035] The term "polypeptide" is used here to refer to a molecule of two or more subunits of amino acids linked by peptide bonds. Typically, though not always, the polypeptides contain several hundred amino acids; for example, about 400 to about 900 amino acids.
[0036] A "plasmid" or “vector” is a DNA molecule that is typically separate from and capable of replicating independently of the chromosomal DNA. In many cases, it is circular and double-stranded. It is known in the art that while plasmid vectors often exist as extrachromosomal circular DNA molecules, plasmid vectors may also be designed to be stably integrated into a host chromosome either randomly or in a targeted manner. Many plasmids are commercially available for varied uses. The gene to be replicated is inserted into copies of a plasmid contain- ing genes that make cells resistant to particular antibiotics, and a multiple cloning site (MCS, or polylinker), which is a short region containing several commonly used restriction sites al- lowing the easy insertion of DNA fragments at this location. Typically, the polypeptides dis- closed herein are expressed from plasmids.
[0037] The term “about” or “approximately” when immediately preceding a numerical value means a range (e.g., plus or minus 10% of that value). For example, “about 50” can mean 45 to 55, “about 25,000” can mean 22,500 to 27,500, etc., unless the context of the disclosure indicates otherwise, or is inconsistent with such an interpretation. For example, in a list of numerical values such as “about 49, about 50, about 55, ... ”, “about 50” means a range extend- ing to less than half the interval(s) between the preceding and subsequent values, e.g., more than 49.5 to less than 52.5. Furthermore, the phrases “less than about” a value or “greater than about” a value should be understood in view of the definition of the term “about” provided herein. Similarly, the term “about” when preceding a series of numerical values or a range of values (e.g., “about 10, 20, 30” or “about 10-30”) refers, respectively to all values in the series, or the endpoints of the range.
[0038] As used herein the terms “microorganism” or “microbe” should be taken broadly. These terms are used interchangeably and include, but are not limited to, the two prokaryotic domains, Bacteria and Archaea, as well as certain eukaryotic fungi and protists. In embodiments, the disclosure refers to the “microorganisms” or “microbes” of lists and figures present in the dis- closure. This characterization can refer to not only the identified taxonomic genera but also the identified taxonomic species, as well as the various novel and newly identified or designed strains of any organism in said tables or figures. The same characterization holds true for the recitation of these terms in other parts of the Specification, such as in the Examples.
[0039] Amino acids are the compounds or building blocks that make up peptides and proteins. Each amino acid is structured from an amino group (N-terminus) and a carboxyl group (C- terminus) bound to a tetrahedral carbon. This carbon is designated as the a-carbon (alpha-car- bon). Amino acids differ from each other with respect to their side chains, which are referred to as R groups. Though the R group for each of the amino acids will differ in structure, electrical charge, and polarity, amino acids can be grouped into several groups of amino acids having similar properties. These similar properties permit these amino acids, in certain instances, tobe reasonably interchangeably within an amino acid sequence. These groups include aliphatic amino acids, aromatic amino acids, amino acids with polar neutral side chains, acidic amino acids with electrically charged side chains, basic amino acids with electrically charged side chains, and “unique” amino acids. Aliphatic amino acids are amino acids with hydrophobic side chains and include alanine, isoleucine, leucine, methionine, and valine. Aromatic amino acids are amino acids with hydrophobic side chains and include phenylalanine, tryptophan, and tyrosine. Amino acids with polar neutral side chains include asparagine, cysteine, glutamine, serine, and threonine. Acidic amino acids with electrically charged side chains include aspartic acid and glutamic acid. Basic amino acids with electrically charged side chains include argi- nine, histidine, and lysine. “Unique” amino acids, which are amino acids that cannot be other- wise grouped together, include glycine and proline.
[0040] When referring to a nucleic acid sequence or protein sequence, the term “identity” is used to denote similarity between two sequences. Sequence similarity or identity may be de- termined using standard techniques known in the art, including, but not limited to, the local sequence identity algorithm of Smith & Waterman, Adv. Appl. Math. 2, 482 (1981), by the sequence identity alignment algorithm of Needleman & Wunsch, J Mol. Biol. 48,443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Natl. Acad. Sci. USA 85, 2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Sci- ence Drive, Madison, WI), the Best Fit sequence program described by Devereux et al., Nucl. Acid Res. 12, 387-395 (1984), or by inspection. Another suitable algorithm is the BEAST al- gorithm, described in Altschul et al., J Mol. Biol. 215, 403-410, (1990) and Karlin et al., Proc. Natl. Acad. Sci. USA 90, 5873-5787 (1993). An exemplary BLAST program is the WU- BLAST-2 program which was obtained from Altschul et al., Methods in Enzymology, 266, 460-480 (1996); blast.wustl / edu / blast / README.html. WU-BLAST-2 uses several search pa- rameters, which are optionally set to the default values. The parameters are dynamic values and are established by the program itself depending upon the composition of the sequence and composition of the particular database against which the sequence of interest is being searched; however, the values may be adjusted to increase sensitivity. Another algorithm is gapped BLAST as reported by Altschul et al, (1997) Nucleic Acids Res. 25, 3389-3402. Other algo- rithms may be described herein.
[0041] In bioinformatics, several methods have been developed to find and determine related polypeptide sequences. For example, percent sequence identity, position-specific scoring ma-trices (PSSMs) and hidden Markov models (HMMs) are all commonly employed to find se- quences that are similar to a given query sequence. Percent sequence identity calculates the number of amino acids that are shared between two sequences. Percent sequence identity is calculated in the context of a given alignment between two sequences. Percentage identity may be calculated using the alignment program Clustal Omega (available at / www.ebi.ac.uk / Tools / msa / clustalo / ) with default settings. The default transition matrix is Gonnet, gap opening penalty is 6 bits, and gap extension is 1 bit. Clustal Omega uses the HHalign algorithm and its default settings as its core alignment engine. The algorithm is de- scribed in Soding, J. (2005) 'Protein homology detection by HMM-HMM comparison'. Bioin- formatics 21, 951-960.
[0042] Position-specific scoring matrices (PSSMs) are a concise way to represent many related sequences. PSSMs are often generated using multiple sequence alignments. The sequence search tool PSI-BLAST generates PSSMs and uses them to search for related polypeptide se- quences. A PSSM used to score polypeptide sequences is a matrix (i.e. table) composed of 21 columns by N rows, where N is the length of the related sequences. Each row corresponds to a position within the polypeptide sequence and each column represents a different amino acid (or gap) that the residue position can take on. Each entry in the PSSM represents a score for the specific amino acid at the specific position within the polypeptide sequence. A sequence can be scored with a PSSM by first aligning the sequence to a reference sequence, and then calculating the following sum: SPSSM=aaL). where i is the sequence position and ctcti is the amino acid at position i. Related polypeptide sequences will all have high PSSM scores, while unrelated sequences will yield low scores.
[0043] Steviol glycosides are composed of a diterpenoid steviol core and two variable glycans. The two variable glycans are attached to the C13-hydroxyl (Rl) and C19-carboxylate (R2) of the steviol core, as shown in FIG. 1. Steviol and a selection of its glycosides are shown below in the context of their variable glycans:
[0044] Rebaudioside D (Reb D) can be produced by adding a beta- 1 ,2-linked glucose monomer to the first C19 glucose of Reb A, as shown in FIG. 2. Rebaudioside E (Reb E) can be produced by adding a beta-l,2-linked glucose monomer to the first C19 glucose of stevioside (FIG. 3). Rebaudioside M (Reb M) can be produced by adding a beta-1, 3-linked glucose monomer to the lower-order steviol glycosides Rebaudioside D (Reb D) or Rebaudioside AM (Reb AM) or by adding a beta-l,2-linked glucose monomer to Rebaudioside I (Reb I), as shown in FIG. 4. FIG. 10 illustrates pathways, which proceed generally from left to right, between stevioside as a starting steviol glycoside and Reb M as a target steviol glycoside. Enzyme-catalyzed path- ways between steviol glycosides are indicated. It should be appreciated that more than one enzyme and more than steviol glycoside may be present concurrently. Native glycosyltransfer- ases that perform these conversions use uridine diphosphate (UDP)-glucose as the glucose source for transferring to the lower-order steviol glycosides.
[0045] The present disclosure provides non-natural, engineered beta- 1,2 glycosyltransferases (B12GTs) that can use an ADP-glucose sugar donor to make a beta-l,2-glycosidic bond with the first glucose at either the C13 or C19 end of a steviol glycoside (SG). In aspects, the gly- cosyltransferase polypeptide may be one of SEQ ID NOs: 449-879. For example, the glycosyl- transferase polypeptide may be a polypeptide sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to one of SEQ ID NOs: 449-879.
[0046] The present disclosure also provides non-natural, engineered sucrose synthases (SuSys) that can use a sucrose sugar donor to convert ADP to ADP-glucose. In aspects, the SuSy poly- peptide is one of SEQ ID NOs: 2-448. For example, the SuSy polypeptide may be a polypeptide sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to one of SEQ ID NOs: 2- 448.
[0047] In embodiments, the glycosyltransferase and / or sucrose synthase polypeptides are pre- pared by expression in a host microorganism. Suitable host microorganisms include, but are not limited to, E. coli, Saccharomyces sp., Aspergillus sp., Pichia sp., Bacillus sp. For exam- ple, the glycosyltransferase and sucrose synthase may be expressed in E. coli. For example, the glycosyltransferase and sucrose synthase may be expressed in Pichia pastoris.
[0048] The B12GT and / or SuSy polypeptide can be provided in any suitable form, including free, immobilized, or as a whole cell system. The degree of purity of the glycosyltransferase polypeptide may vary, e.g., it may be provided as a crude, semi-purified, or purified enzyme preparation(s). In embodiments, the glycosyltransferase polypeptide is free. In embodiments, the glycosyltransferase polypeptide is immobilized to a solid support, for example on an inor- ganic or organic support. The solid support may be derivatized cellulose, glass, ceramic, meth- acrylate, styrene, acrylic, a metal oxide, or a membrane. In embodiments, the glycosyltransfer- ase polypeptide is immobilized to the solid support by covalent attachment, adsorption, cross- linking, entrapment, or encapsulation.
[0049] In embodiments, the B 12GT and / or SuSy polypeptide is provided in the form of a whole cell system, for example as a living fermentative microbial cell, or as dead and stabilized mi- crobial cell, or in the form of a cell lysate.
[0050] The present disclosure provides a biocatalytic process for the preparation of a compo- sition comprising one or more target steviol glycosides from a starting composition comprising one or more substrate steviol glycosides, wherein the target steviol glycoside(s) comprises one or more additional glucose units than the substrate steviol glycoside(s). In embodiments, the biocatalytic process comprises contacting a B12GT with a starting composition comprising one or more steviol glycosides and a non-UDP nucleoside diphosphate-sugar. In embodiments, the biocatalytic process comprises contacting an engineered B12GT with a starting composition comprising one or more steviol glycosides and a non-UDP nucleoside diphosphate-sugar. In embodiments, the biocatalytic process comprises contacting an engineered Bl 2GT with a start- ing composition comprising one or more steviol glycosides and ADP-glucose. In embodiments, the biocatalytic process comprises contacting an engineered B12GT with a starting composi- tion comprising stevioside and Reb A and ADP-glucose to produce Reb E and Reb D.
[0051] In embodiments, the biocatalytic process comprises contacting a B12GT and a SuSy with a starting composition comprising one or more steviol glycosides, a non-UDP nucleoside diphosphate, and sucrose. In embodiments, the biocatalytic process comprises contacting an engineered B12GT and a SuSy with a starting composition comprising one or more steviol glycosides, a non-UDP nucleoside diphosphate, and sucrose. In embodiments, the biocatalytic process comprises contacting an engineered B12GT and an engineered SuSy with a starting composition comprising one or more steviol glycosides, a non-UDP nucleoside diphosphate, and sucrose. In embodiments, the method comprises contacting a mixture of stevioside and Reb A, ADP, and sucrose with an engineered B12GT and a SuSy to produce Reb D and RebE. In embodiments, the method comprises contacting RA50, ADP, and sucrose with an engi- neered B12GT and a SuSy to produce Reb D and Reb E. In embodiments, the method com- prises contacting RA60, ADP, and sucrose with an engineered B12GT and a SuSy to produce Reb D and Reb E.
[0052] In embodiments, the biocatalytic process comprises contacting a B12GT, a SuSy, and one or more additional glycosyltransferases with a starting composition comprising one or more steviol glycosides, anon-UDP nucleoside diphosphate, and sucrose. In embodiments, the biocatalytic process further comprises contacting the engineered Bl 2GT, SuSy, and a beta- 1,3- glycosyltransferase (B13GT) with a starting composition comprising one or more steviol gly- cosides, anon-UDP nucleotide diphosphate, and sucrose. In embodiments, the biocatalytic pro- cess comprises contacting an engineered B12GT, a SuSy, and an engineered B13GT with a starting composition comprising one or more steviol glycosides, anon-UDP nucleoside diphos- phate, and sucrose. In embodiments, the biocatalytic process comprises contacting an engi- neered B12GT, an engineered SuSy, and an engineered B13GT with a starting composition comprising one or more steviol glycosides, a non-UDP nucleoside diphosphate, and sucrose. In embodiments, the method comprises contacting a mixture of stevioside and Reb A, ADP, and sucrose with an engineered B12GT, a SuSy and a B13GT to produce Reb M. In embodi- ments, the method comprises contacting RA60, ADP, and sucrose with an engineered B12GT, a SuSy, and a B13GT to produce Reb M. In embodiments, the method comprises contacting RA50, ADP, and sucrose with an engineered B12GT, a SuSy, and a B13GT to produce Reb M.
[0053] In embodiments, the B12GT polypeptide is one of SEQ ID NOs: 449-879. The glyco- syltransferase polypeptide may be a polypeptide sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to one of SEQ ID NOs: 449-879. Preferably, the catalytic domain in the B12GT polypeptide contains residues corresponding to H at position 15, D at position 114, H at position 333, S at position 338, E at position 341, D or E at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
[0054] In embodiments, the sucrose synthase is any polypeptide with sucrose synthase activity. In embodiments, the sucrose synthase is derived from an organism from the Bacterial domain. In embodiments, the sucrose synthase is derived from an organism from the Plantae kingdom. In embodiments, the sucrose synthase is derived from an organism from the proteobacteria, deferribacteres, or cyanobacteria phylum. In embodiments, the sucrose synthase is derived from the species Acidithiobacillus caldus. Nitrosomonas europaea. Denitrovibrio acetiphilus,Thermosynechococcus elongatus, Oryza sativa. Arabidopsis thaliana. or Coffea arabica. In embodiments, the sucrose synthase is one of SEQ ID NOs: 2-448. The sucrose synthase may be an engineered sucrose synthase with a polypeptide sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to one of SEQ ID NOs: 2-448. Preferably, the catalytic domain in the SuSy polypeptide contains residues corresponding to H at position 425, R at position 567, K at position 572, and E at position 663, numbered according to SEQ ID NO: 4 or residues corresponding to H at position 436, R at position 578, K at position 583, and E at position 674, numbered according to SEQ ID NO: 7.
[0055] In embodiments, the B13GT is any polypeptide with beta- 1,3 -glycosyltransferase ac- tivity. The B13GT may be derived from an organism from the Plantae kingdom. For example, the B13GT may be derived from the species Stevia rebaudiana. The B13GT may be an engi- neered B13GT with a polypeptide sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to the B13GT from Stevia rebaudiana. The B13GT polypeptide may be a polypeptide sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to one of SEQ ID NOs: 2-517 from International Patent Application No. PCT / US2023 / 073344, fded September 1, 2023.
[0056] The B 12GT, sucrose synthase, and / or B 13GT polypeptides may be prepared by expres- sion in a host microorganism. Suitable host microorganisms include, but are not limited to, E. coli, Saccharomyces sp., Aspergillus sp., Pichia sp., Bacillus sp. For example, the B 12GT, su- crose synthase, and / or Bl 3GT may be expressed in E. coli. For example, the B12GT, sucrose synthase, and / or B13GT may be expressed in Pichia pastoris. In embodiments, the B12GT, sucrose synthase, and / or Bl 3GT polypeptides are prepared by cell-free expression.
[0057] The B12GT, sucrose synthase, and / or B13GT polypeptides can be provided in any suit- able form, including free, immobilized, or as a whole cell system. The degree of purity of the polypeptides may vary, e.g., they may be provided as a crude, semi-purified, or purified en- zyme preparation(s). In embodiments, the B12GT, sucrose synthase, and / or B13GT polypep- tide is free. In embodiments, the B12GT, sucrose synthase, and / or B13GT polypeptide is im- mobilized to a solid support, for example on an inorganic or organic support. The solid support is derivatized cellulose, glass, ceramic, methacrylate, styrene, acrylic, a metal oxide, or a mem- brane. In some embodiments, the B12GT, SuSy, and / or B13GT polypeptide is immobilized tothe solid support by covalent attachment, adsorption, cross-linking, entrapment, or encapsula- tion.
[0058] The B12GT, sucrose synthase, and / or B13GT polypeptide may be provided in the form of a whole cell system, for example as a living fermentative microbial cell, or as dead and stabilized microbial cell, or in the form of a cell lysate.
[0059] The steviol glycoside component(s) of the starting composition serves as a substrate(s) for the production of the target steviol glycoside(s), as described herein. The target steviol gly- coside (s) differ chemically from their corresponding substrate steviol glycoside(s) by the ad- dition of one or more glucose units.
[0060] The starting steviol glycoside composition can contain at least one substrate steviol glycoside. In embodiments, the substrate steviol glycoside is selected from the group consisting of steviol, steviol-13-O-glucoside, steviol- 19-O-glucoside, rubusoside, steviol- 1,2-bioside, ste- viol- 1,3-bioside, rubusoside, dulcoside B, dulcoside A, rebaudioside B, rebaudioside G, stevi- oside, rebaudioside C, rebaudioside F, rebaudioside A, rebaudioside I, rebaudioside E, rebau- dioside E2, rebaudioside H, rebaudioside L, rebaudioside K, rebaudioside J, rebaudioside AM, rebaudioside M, rebaudioside D, rebaudioside N, rebaudioside O, rebaudioside Q, rebaudioside X, an isomer thereof, a synthetic steviol glycoside or combinations thereof. In embodiments, the starting steviol glycoside composition is composed of stevioside and Reb A. In embodi- ments, the starting steviol glycoside composition is composed of stevioside. In embodiments, the starting steviol glycoside composition is composed of Reb A.
[0061] The starting steviol glycoside composition may be synthetic or purified (partially or entirely), commercially available or prepared. One example of a starting composition useful in the method of the present disclosure is an extract obtained from purification of Stevia rebaudi- ana plant material (e.g., leaves). Another example of a starting composition is a commercially available stevia extract brought into solution with a solvent. Yet another example of a starting composition is a commercially available mixture of steviol glycosides brought into solution with a solvent. Other suitable starting compositions include by-products of processes to isolate and purify steviol glycosides.
[0062] In embodiments, the starting composition comprises a purified substrate steviol glyco- side. For example, the starting composition may comprise greater than about 50%, greater than about 60%, greater than about 70%, greater than about 80%, greater than about 85%, greater than about 90%, greater than about 91%, greater than about 92%, greater than about 93%, greater than about 94%, greater than about 95%, greater than about 96%, greater than about97%, greater than about 98%, greater than about 99%, or greater than about 99.6% of one or more substrate steviol glycosides by weight on an anhydrous basis.
[0063] In embodiments, the starting composition comprises a partially purified substrate ste- viol glycoside composition. For example, the starting composition contains greater than about 0.5%, greater than about 1%, greater than about 2%, greater than about 3%, greater than about 4%, greater than about 5%, greater than about 10%, greater than about 20%, greater than about 30%, greater than about 40%, or greater than about 50%, of one or more substrate steviol gly- cosides by weight on an anhydrous basis.
[0064] In embodiments, the substrate steviol glycoside is purified rebaudioside A, or isomers thereof. The substrate steviol glycoside may contain greater than 99% rebaudioside A, or iso- mers thereof, by weight on an anhydrous basis. In embodiments, the substrate steviol glycoside comprises partially purified rebaudioside A. The substrate steviol glycoside may contain greater than about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% rebaudioside A by weight on an anhydrous basis.
[0065] In embodiments, the substrate steviol glycoside comprises purified stevioside, or iso- mers thereof. For example, the substrate steviol glycoside may contain greater than 99% stevi- oside, or isomers thereof, by weight on an anhydrous basis. In embodiments, the substrate ste- viol glycoside comprises partially purified stevioside. For example, the substrate steviol gly- coside may contain greater than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% stevioside by weight on an anhydrous basis.
[0066] In embodiments, the substrate steviol glycoside is a combination of stevioside and re- baudioside A. The substrate steviol glycoside may contain greater than about 5% stevioside and greater than about 5% Reb A, greater than about 10% stevioside and greater than about 10% Reb A, greater than about 20% stevioside and greater than about 20% Reb A, greater than about 30% stevioside and greater than about 30% Reb A, greater than about 40% stevioside and greater than about 40% Reb A, greater than about 45% stevioside and greater than about 45% Reb A, greater than about 40% stevioside and greater than about 50% Reb A, greater than about 30% stevioside and greater than about 60% Reb A, greater than about 20% stevioside and greater than about 70% Reb A, greater than about 10% stevioside and greater than about 80% Reb A, greater than about 5% stevioside and greater than about 90% Reb A, greater than about 50% stevioside and greater than about 40% Reb A, greater than about 60% stevioside and greater than about 30% Reb A, greater than about 70% stevioside and greater than about 20% Reb A, greater than about 80% stevioside and greater than about 10% Reb A, or greater than about 90% stevioside and greater than about 5% Reb A by weight on an anhydrous basis.
[0067] The substrate steviol glycoside may be derived from stevia leaf extract. RA50, a stevia leaf extract purified to contain greater than 50% Reb A, may be used as the steviol glycoside substrate. In embodiments, RA50 is used at a concentration between about 1 and 800 mg / mL; for example, about 100 mg / mL. RA50 may be used at a concentration of about 60 mg / mL. In aspects, the range of RA50 is about 60 mg / mL to about 100 mg / mL.
[0068] In embodiments, the substrate steviol glycoside is derived from stevia leaf extract. In embodiments, RA60, stevia leaf extract purified to contain greater than 60% Reb A, is used as the steviol glycoside substrate. In embodiments, RA60 is used at a concentration between about 1 and 800 mg / mL. For example, RA60 may be used at a concentration of about 100 mg / mL. For example, RA60 may be used at a concentration of about 60 mg / mL. In aspects, the range of RA60 is about 60 mg / mL to about 100 mg / mL.
[0069] The one pot reaction can be carried out with a nucleoside diphosphate cofactor that can be converted to an NDP-glucose by sucrose synthase. In embodiments, the cofactor can be a non-UDP nucleoside diphosphate (i.e. ADP-glucose, GDP-glucose, CDP-glucose, or TDP-glu- cose). In embodiments, the nucleoside diphosphate may be ADP. For example, the one pot reaction can be carried out with ADP at a concentration between about 0.01 and 10 mM, such as, for example, between 0.01 mM and 0.05 mM, between 0.05 mM and 0.1 mM, between 0. 1 mM and 0.5 mM, between 0.5 mM and 1 mM, between 1 mM and 5 mM, or between 5 mM and 10 mM. For example, ADP is used at a concentration of 0.5 mM.
[0070] The one pot reaction can be carried out with a sucrose concentration between about 10 mM and 2M, such as, for example, greater than 10 mM, greater than 50 mM, greater than 100 mM, greater than 250 mM, greater than 500 mM, greater than 1 M, greater than 1.5 M and greater than 2 M. For example, sucrose is used at a concentration of 250 mM.
[0071] In embodiments, the reaction is run at any temperature. The one-pot reaction may be run at a temperature between about 10 °C and 80 °C such as, for example, between 10° C to 20° C, between 20° C to 30° C, between 30° C to 40° C, between 40° C to 50° C, between 50° C to 60° C, between 60° C to 70° C, between 70° C to 80° C or 80° C. For example, the one- pot reaction may be carried out at 60°C.
[0072] The reaction medium for conversion is generally aqueous, e.g., purified water, buffer, or a combination thereof. For example, the reaction medium may be a buffer. Suitable buffers include, but are not limited to, acetate buffer, citrate buffer, HEPES, and phosphate buffer. For example, the reaction medium may be phosphate buffer. The reaction medium can have a pH between about 4 and 10. For example, the reaction medium has a pH of 6. The reaction medium can also be, alternatively, an organic solvent.
[0073] The step of contacting the starting composition with the glycosyltransferase and sucrose synthase polypeptides can be carried out in a duration of time between about 1 hour and 1 week, such as, for example, between 30 minutes and 1 hours, between 1 hour and 4 hours, between 4 hours and 6 hours, between 6 hours and 12 hours, between 12 hours and 24 hours, between 1 day and 2 days, between 2 days and 3 days, 3 days and 4 days, between 4 days and 5 days, between 6 days and 7 days. For example, the reaction may be carried out for 24 hours. For example, the reaction may be carried out for 4 hours.
[0074] The reaction can be monitored by suitable methods including, but not limited to, HPLC, LCMS, TLC, IR, orNMR.
[0075] The target steviol glycoside can be any steviol glycoside. In embodiments, the target steviol glycoside is steviol- 13 -O-glucoside, steviol- 19-O-glucoside, rubusoside, steviol- 1,2-bi- oside, steviol-l,3-bioside, rubusoside, dulcoside B, dulcoside A, rebaudioside B, rebaudioside G, stevioside, rebaudioside C, rebaudioside F, rebaudioside A, rebaudioside I, rebaudioside E, rebaudioside E2, rebaudioside H, rebaudioside L, rebaudioside K, rebaudioside J, rebaudioside AM, rebaudioside M, rebaudioside D, rebaudioside N, rebaudioside O, rebaudioside Q, rebau- dioside X, a rebaudioside with 7 covalently attached glucose units (e.g. rebaudioside M plus 1 glucose unit), a synthetic steviol glycoside, an isomer thereof, and / or a steviol glycoside com- position. In embodiments, the target steviol glycoside is rebaudioside E, or isomers thereof. In embodiments, the target steviol glycoside is rebaudioside D, or isomers thereof. In embodi- ments, the target steviol glycosides are Reb D and Reb E. In embodiments, the target steviol glycoside is Reb M or isomers thereof.
[0076] In embodiments, the conversion of Reb A to Reb D and / or Reb D isomer(s) is at least about 2% complete, as determined by any of the methods mentioned above. In embodiments, conversion of Reb A to Reb D and / or Reb D isomer(s) may be at least about 10% complete, at least about 20% complete, at least about 30% complete, at least about 40% complete, at least about 50% complete, at least about 60% complete, at least about 70% complete, at least about 80% complete, or at least about 90% complete. For example, the conversion of Reb A to Reb D and / or Reb D isomer(s) may be at least about 95% complete. In embodiments, at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the Reb A in the starting compo- sition is converted to Reb D and / or Reb D isomer(s).
[0077] In embodiments, the conversion of stevioside to Reb E and / or Reb E isomer(s) is at least about 2% complete, as determined by any of the methods mentioned above. In embodi- ments, the conversion of stevioside to Reb E and / or Reb E isomer(s) may be at least about 10%complete, at least about 20% complete, at least about 30% complete, at least about 40% com- plete, at least about 50% complete, at least about 60% complete, at least about 70% complete, at least about 80% complete, or at least about 90% complete. For example, the conversion of stevioside to Reb E and / or Reb E isomer(s) may be at least about 95% complete. In embodi- ments, at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the stevioside in the starting composition is converted to Reb E and / or Reb E isomer(s).
[0078] In embodiments, the conversion of stevioside and / or Reb A to Reb M and / or Reb M isomer(s) is at least about 2% complete, as determined by any of the methods mentioned above. In embodiments, the conversion of stevioside and / or Reb A to Reb M and / or Reb M isomer(s) is at least about 10% complete, at least about 20% complete, at least about 30% complete, at least about 40% complete, at least about 50% complete, at least about 60% complete, at least about 70% complete, at least about 80% complete, or at least about 90% complete. For exam- ple, the conversion of stevioside and / or Reb A to Reb M and / or Reb M isomer(s) may be at least about 95% complete. In embodiments, at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the steviol glycosides in the starting composition is converted to Reb M and / or Reb M isomer(s).
[0079] The target steviol glycoside(s) can be in any polymorphic or amorphous form, including hydrates, solvates, anhydrous or combinations thereof.
[0080] Optionally, the method of the present disclosure further comprises separating the target steviol glycoside from the target composition. The target steviol glycoside(s) can be separated by any suitable method, such as, for example, crystallization, fdtration, separation by mem- branes, centrifugation, extraction, chromatographic separation or a combination of such meth- ods.
[0081] In embodiments, the separation of target steviol glycosides produces a composition comprising greater than about 80% by weight of the target steviol glycoside(s) on an anhydrous basis, i.e., a highly purified steviol glycoside composition. In embodiments, separation pro- duces a composition comprising greater than about 0.5%, greater than about 1%, greater than about 2%, greater than about 3%, greater than about 4%, greater than about 5%, greater than about 10%, greater than about 20%, greater than about 30%, greater than about 40%, greater than about 50%, greater than about 60%, greater than about 70%, greater than about 80%, greater than about 85%, greater than about 90%, greater than about 91%, greater than about 92%, greater than about 93%, greater than about 94%, greater than about 95%, greater than about 96%, greater than about 97%, greater than about 98%, greater than about 99%, or greaterthan about 99.6% by weight of the target steviol glycoside(s). For example, the composition may comprise greater than about 95% by weight of the target steviol glycoside(s).
[0082] Purified target steviol glycoside(s) can be used in consumable products as a sweetener. Suitable consumer products include, but are not limited to, food, beverages, pharmaceutical compositions, tobacco products, nutraceutical compositions, oral hygiene compositions, and cosmetic compositions.
[0083] In addition to the primary structure (i.e., amino acid sequence) of the enzymes described above and elsewhere herein, each of the enzymes (e.g., B12GT, SuSy, and B13GT) of the present disclosure exhibit a secondary structure and a tertiary structure. The secondary struc- ture, which refers to local folded structures that form within a polypeptide due to interactions between atoms of the backbone, of the engineered B12GT polypeptides of the present disclo- sure may include beta strands, beta sheets, and alpha helices. In an alpha helix, the carbonyl (C=O) of one amino acid is hydrogen bonded to the amino H (N-H) of an amino acid that is four down the chain. This pattern of bonding pulls the polypeptide chain into a helical structure that resembles a curled ribbon, with each turn of the helix containing 3.6 amino acids. The R groups of the amino acids stick outward from the alpha helix, where they are free to interact. In a beta sheet, two or more segments of a polypeptide chain line up next to each other, forming a sheet-like structure held together by hydrogen bonds. The hydrogen bonds form between carbonyl and amino groups of backbone, while the R groups extend above and below the plane of the beta sheet. The beta strands of the beta sheet may be parallel, pointing in the same direc- tion, or antiparallel. Certain amino acids are more or less likely to be found in alpha helices or beta strands / sheets. For instance, the amino acid proline is sometimes called a helix breaker because its unusual R group (which bonds to the amino group to form a ring) creates a bend in the chain and is not compatible with helix formation. Proline is typically found in bends, un- structured regions between secondary structures. Similarly, amino acids such as tryptophan, tyrosine, and phenylalanine, which have large ring structures in their R groups, are often found beta sheets. The tertiary structure of the engineered B12GT polypeptides of the present disclo- sure refers to the three-dimensional structure of the polypeptides. The tertiary structure is pri- marily due to interactions between the R groups of the amino acids that make up the protein. R group interactions that contribute to tertiary structure include hydrogen bonding, ionic bond- ing, dipole-dipole interactions, and London dispersion forces. Tertiary structure is also im- pacted by hydrophobic interactions, in which amino acids with nonpolar, hydrophobic R groups cluster together on the inside of the protein, leaving hydrophilic amino acids on the outside to interact with surrounding water molecules. Tertiary structure is also impacted bydisulfide bonds, which are covalent linkages between sulfur-containing side chains of cyste- ines.
[0084] As stated above, each of the enzymes (e.g., B12GT, SuSy, and B13GT) of the present disclosure exhibit a secondary structure and atertiary structure. Moreover, each ofthe enzymes of the present disclosure can be described according to any one of their primary structure, sec- ondary structure, or tertiary structure. To this end, the secondary structure of SEQ ID NO: 1, on whose structure the design ofthe engineered Bl 2GTs is based, is shown in FIG. 9. Each of the residues forming a beta strand is shown as bracketed and marked with pn, where n denotes a particular beta strand. Each of the residues forming an alpha helix is shown as bracketed and marked with an, where n denotes a particular alpha helix. The table below shows the alpha helices and beta strands of SEQ ID NO: 1, where each of the amino acids indicated as forming an alpha helix or a beta strand correspond to the range of amino acid positions shown in the first column.
[0085] The enzymes disclosed herein may replace one or more alpha helix or one or more beta strand with those from other sequences disclosed herein. Positions of alpha helices or beta strands may be determined by alignment to SEQ ID NO: 1. The enzymes disclosed herein may contain alpha helices or beta strands at the same position as those described in the table with respect to SEQ ID NO: 1 but with each strand or helix having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, or 42 substitutions. In embodiments, the active site amino acids are not substituted. Thus, as an example, amino acids corresponding to the amino acids of FIG. 9 that are shown in bold are not substituted. In embodiments, the catalytic amino acids within the active site are not substituted. To identify key active site residues within the active site involved in catalysis (i.e., the catalytic amino acids), visual analysis and multiple sequence alignment of structurally related glycosyltransferases were performed (FIG. 7). Histidine 15 (sequence num- bering relative to SEQ ID NO: 1) initiates the reaction by removing the 02 hydroxyl proton from the first glucose attached to either the C13 or C19 of the steviol glycoside substrate. As- partate 114 activates the catalytic H15, glutamate 341 binds the ADP-glucose ribose, histidine 333 and serine 338 bind the ADP phosphate groups, and glutamine 358 and the carboxylic acid of aspartate 357 bind and orient the glucose moiety from ADP-glucose. Both aspartate and glutamate at residue 357 can bind the glucose moiety and enable catalysis.
[0086] The tertiary structure of the enzymes can be examined to identify an active site of the enzyme. In the present disclosure, the active site of the B 12GT enzyme was defined by aligning the structural model of SEQ ID NO: 714 from U.S. Patent Application no. 18 / 546,881 to crystal structures of homologous glycosyltransferases from Oryza sativa and Stevia rebaudiana. The structural model was generated by AlphaFold software. The crystal structures of each glyco- syltransferase, which included publicly available crystal structures (e.g., PDB-6INI, PDB- 7ES0, and PDB-7ES2), were then aligned to the structural model. For each instance of the aligned structural model and crystal structure, residues of the enzyme within 8 angstroms of a modeled crystal structure substrate or product were defined as active site residues. Residues not within 8 angstroms of the modeled crystal structure substrate or product were defined gen- erally as surface site residues, though it should be appreciated these residues may or may notbe surface accessible residues. A collective active site was identified by combining the active sites of each instance of the alignment. The collective active site, which may be referred to below as the active site, was numbered according to SEQ ID NO: 1, as will be described below in detail. Referring again to FIG. 9, the active site residues and the surface site residues of a B12GT are shown. In particular, bolded residues indicate amino acid positions of the B12GT that are part of the active site and underlined residues indicate amino acid positions of the B 12GT that are not within the active site (i .e . , are part of the surface site) . The tertiary structure of each enzyme can also be examined to determine which beta strands identified as secondary structures may be associated with a beta sheet of the tertiary structure of the enzyme. In the present disclosure, the B12GT enzymes feature two extended beta-sheets. Numbered according to SEQ ID NO: 1, and with reference to the above table of secondary structures of SEQ ID NO: 1, amino acids at positions corresponding to the first 7 beta strands (e.g., 4-8, 32-37, 57-61, 110-114, 131-135, 191-195, and 216-219 form a first beta sheet and the amino acids at posi- tions corresponding to the remaining 6 beta strands (e.g., 249-253, 278-283, 309-312, 329-332, 349-351, and 371-373) form a second beta sheet.
[0087] In embodiments, the primary structure of an amino acid sequence may be modified to include replacements, additions, and / or deletions. In embodiments, the total length of the amino acid sequence may be shortened or lengthened. Highly mobile residues may be residues that can be deleted, reduced, or lengthened without impacting function of the enzyme. By way of example, residues 151-180 of a B12GT described herein, numbered with respect to SEQ ID NO: 1, may be determined to be highly mobile. These highly mobile residues form a mobile motif which may be replaced with any number of amino acids. For instance, the 30 residues of the mobile motif may be deleted or may be replaced with 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, or any other number of amino acids.EXAMPLESExample 1: In vivo Production of a Native Beta-1, 2-Glycosyltransferase (B12GT)
[0088] A polynucleotide encoding the amino acid sequence of the native glycosyltransferase from Solcinum tuberosum (SEQ ID NO: 1) was synthesized (Twist Bioscience) and inserted into the pARZ4 expression vector. The recombinant vectors were used in a heat shock method to transform E. coli HMS174(DE3) (Novagen), thereby preparing recombinant microorgan- isms.
[0089] The transformed recombinant microorganism was inoculated to 1ml LB- kanamycin medium, cultured by shaking at 37°C overnight. The culture was inoculated to 5ml TB-kana- mycin medium and grown for 2 hours at 37°C, followed by 25°C for 1 hour. The culture was induced with 50 uL 50 mM IPTG and grown overnight. Finally, the culture was centrifuged at top-speed for 5 minutes and stored at -80°C.
[0090] The cell pellet was dissolved in a lysis buffer (lysozyme, DNAsel, Bugbuster, 300mL 20 mM KPOr pH 7.5, 500mM NaCl, and 20mM Imidazole). Two to three glass beads were added to the resuspension and the microorganisms were disrupted by shaking at 25°C and 220rpm for 30 minutes. The disrupted liquid was centrifuged at 2200 x g for 6-10 minutes. The obtained supernatant was loaded onto aNi-NTA plate and shaken for 10 minutes at room tem- perature. The plate was centrifuged for 4 minutes at 100 x g followed by two washes of 500 uL binding buffer (300mL 20mM KPOr pH 7.5, 500mM NaCl, 20mM Imidazole) and two-minute centrifugation (500 x g). The proteins were eluted with 150 uL elution buffer (15mL 20mM KPOr pH 7.5, 500mM NaCl, 500mM Imidazole) and shaken for 1 minute at 0.25 maximum shaking speed followed by centrifugation for 2 minutes at 500 x g. The recovered protein was desalted into a buffer solution for enzyme activity evaluation (50mM KPOr pH 6, 250 mM sodium acetate).
[0091] The melting temperature (Tm) of the enzyme was assessed using the GloMelt™ Ther- mal Shift Protein Stability Assay (Biotium). A 9 pL sample of protein (0.5 mg / mL) in desalt buffer (20 mM KPOr pH6, 50 mM NaCl) was mixed with I pL 10X GloMelt™ dye. Then, the reaction mixture was incrementally heated from 25°C to 99°C at a rate of 0.05°C / second in a Quantstudio3 PCR System (Thermofisher). Protein unfolding was monitored through dye binding to the denatured protein measured by fluorescence at the Excitation / Emission 470 ±15 nm / 520 ± 15 nm wavelengths. The protein melting temperature (Tm) was determined by plotting temperature vs fluorescence, calculating the first derivative of the curve, and taking the local minimum of the derivative (-dF / dT) reported in degrees Celsius. The melting temper- ature of the native B12GT was 45°C.
[0092] The B12GT enzyme was assayed for activity in a one-pot reaction with an engineered variant of the sucrose synthase from Acidithiobacillus caldus (SEQ ID NO: 4). The purified B12GT and SuSy were reacted with 50 mM KPO4 at pH 6, 250 mM NaOAc, 71.5 mg / mL RA60, 250 mM sucrose, and 0.5 mM ADP for 4 hours at 60°C with shaking. Control samples without B12GT were also included. Conversion of stevioside and Reb A to Reb E and Reb D, as schematized respectively in FIG. 3 and FIG. 2, was monitored by liquid chromatography- mass spectrometry (LCMS) using an Agilent 6470 QQQ mass spectrometer (column: WatersACQUITY UPLC HSS T3 Column, 100 mm x 2. 1 mm). The QQQ was run with multi reaction monitoring (MS / MS) to accurately quantitate steviol glycosides of interest. The native B12GT showed <1% conversion of stevioside and Reb A to Reb E and Reb D.Example 2: Computational Design of ADP-Glucose Dependent B12GTs
[0093] Structural models of previously engineered B12GTs (SEQ ID NOs: 549 and 714 from U.S. Patent Application No. 18 / 546,881 and a variant of SEQ ID NO: 1) were generated using a deep learning structural folding method. Each structural model was used as a starting point for computational design of novel B 12GTs. Computational designs were conducted to improve the stability, expression, and activity of the B12GTs. A computational structure-based deep learning design method was used to select B 12GT designs for experimental validation. Expres- sion plasmids for the computational designs were built as in Example 1. Each B12GT variant was expressed and purified as in Example 1 using a desalt buffer composed of 20 mM KPOr pH 7.5 and 50 mM NaCl. Each B12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPO4 pH 6, 50 mM NaCl, 71.5 mg / mL RA50, 250 mM sucrose, and 0.5 mM ADP for 24 hours at 60°C with shaking. Control samples without B12GT were also included. Product rebaudiosides were monitored by LCMS similar to Example 1. The designed enzymes expressed and were active for conversion of stevioside to Reb E and Reb A to Reb D. The total measured Reb D and Reb E conversion and concentration of purified protein obtained for the top computationally de- signed B12GTs (SEQ ID NOs: 449-452) is shown in Table 1.Table 1. B12GT Computational Designs
[0094] Additional designs were conducted where the enzyme active site of the parent scaffold was not allowed to mutate during the design process. The active site was held fixed to ensure that the catalytic activity of the parent scaffold was maintained or improved in the B12GT designs. As introduced previously, the active site was defined by aligning the structural model of SEQ ID NO: 714 from U.S. Patent Application No. 18 / 546,881 to crystal structures of ho- mologous glycosyltransferases from Oryza sativa and Stevia rebaudiana. All residues of thestructural model within 8 A of a modeled crystal structure substrate or product was defined as an active site residue. The active site corresponded to the following 118 residues (numbered according to SEQ ID NO: 1): 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 48, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 114, 115, 116, 117, 134, 135, 136, 137, 138, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 152, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173,174, 175, 176, 177, 178, 179, 180, 181, 182, 184, 185, 196, 223, 224, 225, 226, 230, 232, 251,252, 253, 254, 255, 256, 257, 258, 259, 282, 284, 312, 314, 315, 316, 317, 318, 319, 320, 321,322, 330, 332, 333, 334, 335, 336, 337, 338, 339, 340, 341, 342, 345, 353, 354, 355, 356, 357,358, 359, 360, and 361.
[0095] The structure-based deep learning design method with active site constraints was ap- plied to identify B12GT designs for evaluation. Expression plasmids for the computational designs were built as in Example 1. Each B12GT variant was expressed and purified as in Example 1 using a desalt buffer composed of 20 mM KPOi pH 7.5 and 50 mM NaCl. Each B12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPOr pH 6, 50 mM NaCl, 71.5 mg / mL RA50, 250 mM sucrose, and 0.5 mM ADP for 24 hours at 60°C with shaking. Control samples without B12GT were also included. Product rebaudiosides were monitored by LCMS similar to Example 1 and enzyme melting temperatures were measured as described in Example 1.
[0096] Several designed enzymes expressed well, were thermostable, and actively converted stevioside and Reb A to Reb D and Reb E (SEQ ID NOs: 453-511; Table 2).Table 2. B12GT Computational DesignsaMelting temperature is defined as: less than 60°C, ‘+’: between 60°C and 70°C, ‘++’: between 70°C and 80°C, ‘+++’: between 80°C and 90°C, ‘++++’: above 90°C, ‘ND’: not de- termined.
[0097] To bias the general structure-based deep learning design model towards improved B12GT designs, the deep learning model was updated by re-training the model on native gly- cosyltransferase structures. The updated and re-trained design model was then used to identify B12GT designs for evaluation. Expression plasmids for the computational designs were built as in Example 1. Each B12GT variant was expressed and purified as in Example 1 using a desalt buffer composed of 20 mM KPOr pH 7.5 and 50 mM NaCl. Each B12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPOr pH 6, 50 mM NaCl, 71.5 mg / mL RA50, 250 mM su- crose, and 0.5 mM ADP for 24 hours at 60°C with shaking. Control samples without B12GT were also included. Product rebaudiosides were monitored by LCMS similar to Example 1 and melting temperatures were measured as described above. Several designed enzymes expressed well, were thermostable, and actively converted stevioside and Reb A to Reb E and Reb D (SEQ ID NOs: 512-571; Table 3).Table 3. B12GT Computational DesignsExample 3: Computational Design of ADP-Glucose Dependent B12GTs
[0098] Engineered B12GT experimental data from U.S. Patent Application No. 18 / 546,881 was used to fine-tune a computational sequence-based deep learning protein design method to model B12GT activity and expression. The design method was used to select B12GT designs for experimental testing. Expression plasmids for the computational designs were built as in Example 1. Each B12GT variant was expressed and purified as in Example 2. Each B12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPOr pH 6, 50 mM NaCl, 71.5 mg / mL RA50, 250 mM sucrose, and 0.5 mM ADP for 24 hours at 60°C with shaking. Control samples without B12GT were also included. Product rebaudiosides were monitored by LCMS similar to Example 1 and enzyme melting temperatures were measured as described in Example 1. Several designed enzymes expressed well, were thermostable, and actively converted stevioside and Reb A to Reb E and Reb D (SEQ ID NOs: 572-594; Table 4Table).Table 4. B12GT Computational Designs
[0099] A similar computational sequence-based deep learning model was fine-tuned on B 12GT experimental data. The model was used to select B 12GT designs for experimental test- ing. Expression plasmids for the computational designs were built as in Example 1. Each B12GT variant was expressed and purified as in Example 2. Each B12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPOr pH 6, 50 mM NaCl, 71.5 mg / mL RA50, 250 mM sucrose, and 0.5 mM ADP for 24 hours at 60°C with shaking. Control samples without B12GT were also included. Product rebaudiosides were monitored by LCMS similar to Example 1 and enzyme melting temperatures were measured as described in Example 1. Several designed enzymes expressed well, were thermostable, and actively converted stevioside and Reb A to Reb E and Reb D (SEQ ID NOs: 595-604; Table 5).Table 5. B12GT Computational DesignsExample 4: Computational Design of ADP-Glucose Dependent B12GTs
[0100] Structural models of SEQ ID NOs: 614 and 617 from U.S. Patent Application No. 18 / 546,881 were made with a deep learning protein folding method to perform further compu- tational designs. A computational structure-based protein design method was used to incorpo- rate beneficial mutations into the previously engineered B12GTs. New B12GT designs werechosen for experimental testing. Expression plasmids for the computational designs were built as in Example 1. Each B 12GT variant was expressed and purified as in Example 2. Each B 12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPOr pH 6, 50 mM NaCl, 71.5 mg / mL RA50, 250 mM sucrose, and 0.5 mM ADP for 22.5 hours at 60°C with shaking. Control samples without B12GT were also included. Product rebaudiosides were monitored by LCMS similar to Example 1 and enzyme melting temperatures were measured as described in Example 1. Several designed enzymes expressed well, were thermostable, and actively converted stevio- side and Reb A to Reb E and Reb D (SEQ ID NOs: 605-635; Table 6Table).Table 6. B12GT Computational DesignsExample 5: Computational Design of Improved of ADP-Glucose Dependent B12GTs
[0101] Structural models of the top designs from Examples 2-4 (SEQ ID NOs: 465, 469, 496, 500) were generated using a deep learning protein folding method. The structural models were used as input to the computational structure-based deep learning design method from Example 2 to design new B12GTs for experimental validation. Expression plasmids for the computa- tional designs were built as in Example 1. Each B12GT variant was expressed and purified as in Example 1. Each B12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPOr at pH 6, 250 mM sodium acetate, 71.5 mg / mL RA60, 250 mM sucrose, and 0.5 mM ADP for 4 hours at 60°C with shaking. Control samples without B12GT were also included. Product rebaudiosides were monitored by LCMS similar to Example 1 and enzyme melting temperatures were meas- ured as described in Example 1. Several designed enzymes expressed well, were thermostable, and actively converted stevioside and Reb A to Reb E and Reb D (SEQ ID NOs: 636-706; Table 7Table).Table 7. B12GT Computational DesignsExample 6: Computational Design of Improved of ADP-Glucose Dependent B12GTs
[0102] The deep learning sequence-based computational design method from Example 3 was fine-tuned using the experimental results of Examples 2-4. The method was used to select B12GT designs for experimental validation. Expression plasmids for the computational de- signs were built as in Example 1. Each B 12GT variant was expressed and purified as in Example 1. Each B12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B 12GT and SuSy were reacted with 50 mM KPOr at pH 6, 250 mM sodium acetate, 71.5 mg / mL RA60, 250 mM sucrose, and 0.5 mM ADP for 4 hours at 60°C with shak- ing. Control samples without B12GT were also included. Product rebaudiosides were moni- tored by LCMS similar to Example 1 and enzyme melting temperatures were measured as de- scribed in Example 1. Several designed enzymes expressed well, were thermostable, and ac- tively converted stevioside and Reb A to Reb E and Reb D (SEQ ID NOs: 707-788; Table 8Table).Table 8. B12GT Computational DesignsExample 7: Active Site Computational Designs of ADP-Glucose Dependent B12GTs
[0103] In Example 7. computational protein designs focused on mutating the active site to create ADP-glucose dependent B12GTs. A structural model of SEQ ID NO: 500 was built using a deep learning protein folding method. A transition state model of the substrates Reb A and ADP-glucose was made and modeled into the active site of the SEQ ID NO: 500 structural model. Computational structure-based designs were conducted to mutate the active site of SEQ ID NO: 500 using the structural model described above, but catalytic residues were not allowedto mutate during design to maintain B12GT activity. To this end, key active site residues in- volved in catalysis were determined by visual analysis and multiple sequence alignment of structurally related glycosyltransferases (FIG. 7). Histidine 15 (sequence numbering relative to SEQ ID NO: 1) initiates the reaction by removing the 02 hydroxyl proton from the first glucose attached to either the C13 or C19 of the steviol glycoside substrate. Aspartate 114 activates the catalytic H15, glutamate 341 binds the ADP-glucose ribose, histidine 333 and serine 338 bind the ADP phosphate groups, and glutamine 358 and the carboxylic acid of aspartate 357 bind and orient the glucose moiety from ADP-glucose. Both aspartate and glutamate at residue 357 can bind the glucose moiety and enable catalysis.
[0104] B12GT designs were chosen for experimental validation. Expression plasmids for the computational designs were built as in Example 1. Each B12GT variant was expressed and purified as in Example 1. Each B12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPOr at pH 6, 250 mM sodium acetate, 71.5 mg / mL RA60, 250 mM sucrose, and 0.5 mM ADP for 4 hours at 60°C with shaking. Control samples without B12GT were also included. Product re- baudiosides were monitored by LCMS similar to Example 1 and enzyme melting temperatures were measured as described in Example 1. Several designed enzymes expressed well, were thermostable, and actively converted stevioside and Reb A to Reb E and Reb D (SEQ ID NOs: 789-831; Table 9Table).Table 9. B12GT Computational Designs
[0105] Additional B12GTs were designed by predicting beneficial active site mutations with a computational structure -based deep learning method. The method used the structural model of SEQ ID NO: 500 with bound transition state as input. B12GT designs were chosen for ex- perimental validation. Expression plasmids for the computational designs were built as in Example 1. Each B 12GT variant was expressed and purified as in Example 1. Each B 12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPOr at pH 6, 250 mM sodium acetate, 71.5 mg / mL RA60, 250 mM sucrose, and 0.5 mM ADP for 4 hours at 60 °C with shaking. Control samples without B 12GT were also included. Product rebaudiosides were monitored by LCMS similar to Example 1 and enzyme melting temperatures were measured as described in Example 1. Several designedenzymes expressed well, were thermostable, and actively converted stevioside and Reb A to Reb E and Reb D (SEQ ID NOs: 832-873; Table lOTable).Table 10. B12GT Computational DesignsExample 8: Computational Design of Improved ADP-Glucose Dependent B12GTs by Loop Replacement
[0106] Analysis of B12GT structural models and homologous crystal structures revealed that residues 151-180 (sequence numbering with respect to SEQ ID NO: 1) are highly mobile. Computational designs were conducted to replace the mobile protein segment with a protein segment of 2, 3, or 4 residues. A deep learning method was used to build protein backbone segments of length 2, 3, or 4 residues into the structural model of SEQ ID NO: 500. Amino acid sequences for the new backbone segments were designed using the structural deep learn- ing design method used in Example 2. Six B12GT designs were chosen for experimental vali- dation (SEQ ID NOs: 874-879). Expression plasmids for the computational designs were built as in Example 1. Each B 12GT variant was expressed and purified as in Example 1. Each B 12GT variant was assayed in a one-pot reaction with a SuSy variant of SEQ ID NO: 4. The purified B12GT and SuSy were reacted with 50 mM KPOr at pH 6, 250 mM sodium acetate, 71.5 mg / mL RA60, 250 mM sucrose, and 0.5 mM ADP for 4 hours at 60°C with shaking. Control samples without B 12GT were also included. Product rebaudiosides were monitored by LCMS similar to Example 1 and enzyme melting temperatures were measured as described in Example 1. The designed enzymes were thermostable, and actively converted stevioside and Reb A to Reb E and Reb D (Table 1 ITable). Two B 12GT designs (SEQ ID NOs: 876 and 879) expressed very well.Table 11. B12GT Computational DesignsExample 9: Representing Successful B12GTs Designs with a PSSM
[0107] The successful B12GT designs from Examples 2-8 were used to generate position- specific scoring matrices (PSSMs) that represent the B12GT designs. Each PSSM is a concise way to represent the successful designs and related B12GT sequences. To score a sequence with the PSSM, it must first be aligned with the representative sequence SEQ ID NO: 1.
[0108] The top 26% of the B12GT designs from Examples 2-8 were collected and used to generate a representative PSSM (Table 12). Sequences that have a score greater than 96.0 when scored with the PSSM from Table 12 are considered related to the top active computational designs described in Examples 2-8. For example, the following successfully designed Bl 2GTs, SEQ ID NOs: 465, 469, 496, and 500, have the following PSSM scores: 149.7, 129.4, 145.3, 151.3, while the native B12GT (SEQ ID NO: 1), has a PSSM score of only 90.2.Table 12. Position Specific Scoring Matrix (PSSM) of Successful B12GT Designs
[0109] The top 15% of the B12GT designs from Example 2-8 were collected and used to generate a representative PSSM (Table 13). Sequences that have a score greater than 94.1 when scored with the PSSM from Table 13 are considered related to the top active computational designs described in Examples 2-8. For example, the following successfully designed B12GTs, SEQ ID NOs: 465, 469, 496, and 500, have the following PSSM scores: 148.4, 126.8, 144.4, 150.0, while the native B12GT SEQ ID NO: 1, has a PSSM score of only 86.0.Table 13. Position Specific Scoring Matrix (PSSM) of Successful B12GT DesignsExample 10: Scaled Up Expression of Successful B12GTs Designs
[0110] Four top performing B12GTs from Example 2 (SEQ ID NOs: 465, 469, 496, and 500) were chosen for expression scale-up. A polynucleotide encoding the amino acid sequence of each B12GT design was inserted into expression vectors for constitutive expression. The ex- pression vectors were transformed into E. coli and the B12GT E. coli strains were grown in 1 L fermenters for 48 hours. Cell pellets were harvested by centrifugation. All strains grew to a wet cell weight of greater than 130 g / L (Table 14). The resuspended cells were lysed by FrenchPress, clarified by centrifugation, and examined by SDS-PAGE for protein expression (see FIG. 8, wherein the arrow indicates relevant expression). The designed B 12GTs enzymes were purified by immobilized metal affinity chromatography (IMAC), dialyzed into desalt buffer (20mM KPO4 pH6, 50mM NaCl) and assayed for Reb D and Reb E conversion. All four en- zymes successfully expressed and had the desired B12GT activity.Table 14. Biomass Analysis ofB12GT FermentationsExample 11: Production of Reb M with a Designed B12GT
[0111] E. coli strains expressing a designed B12GT (SEQ ID NO: 500), designed SuSy (SEQ ID NO: 251), and designed B13GT (SEQ ID NO: 219 from International Application No. PCT / US2023 / 073344) were grown as described in Example 1. The enzymes were expressed and purified as in Example 1 using a desalt buffer composed of 20 mM KPOr pH 6.0 mM and 250 mM sodium acetate. The enzymes were assayed in one-pot reactions where in each reaction the total protein mass was kept constant, but the ratio of the three enzymes was varied. In each one-pot reaction, the purified B12GT, SuSy, and B13GT were reacted with 50 mM KPO4 pH 6, 250 mM sodium acetate, 60 g / L RA60, 250 g / L sucrose, and 1 mM ADP for 5 hours at 55 °C with shaking. By varying the ratio of B12GT, SuSy, and B13GT in the reaction, the per- centage of RebM conversion ranged from less than 3% to greater than 95%.INCORPORATION BY REFERENCE
[0112] All references, articles, publications, patents, patent publications, and patent applica- tions cited herein are incorporated by reference in their entireties for all purposes. However, mention of any reference, article, publication, patent, patent publication, and patent application cited herein is not, and should not be taken as an acknowledgment or any form of suggestion that they constitute valid prior art or form part of the common general knowledge in any coun- try in the world.NUMBERED EMBODIMENTS OF THE INVENTION
[0113] Notwithstanding the appended claims, the disclosure sets forth the following num- bered embodiments:
[0114] (1) An engineered beta-l,2-glycosyltransferase polypeptide that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879.
[0115] (2) The engineered beta-l,2-glycosyltransferase polypeptide of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99% or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-452.
[0116] (3) The engineered beta-l,2-glycosyltransferase polypeptide of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99% or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 453-511.
[0117] (4) The engineered beta-l,2-glycosyltransferase polypeptide of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 512-571.
[0118] (5) The engineered beta-l,2-glycosyltransferase polypeptide of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 572-594.
[0119] (6) The engineered beta-l,2-glycosyltransferase of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 595-604.
[0120] (7) The engineered beta-l,2-glycosyltransferase of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 605-635.
[0121] (8) The engineered beta-l,2-glycosyltransferase of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 636-706.
[0122] (9) The engineered beta-l,2-glycosyltransferase of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 707-788.
[0123] (10) The engineered beta-l,2-glycosyltransferase of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 789-831.
[0124] (11) The engineered beta-l,2-glycosyltransferase of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 832-873.
[0125] (12) The engineered beta-l,2-glycosyltransferase of (1) that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 874-879.
[0126] (13) The engineered beta-l,2-glycosyltransferase polypeptide of (1) wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S at position 338, E at position 341, D at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
[0127] (14) The engineered beta-l,2-glycosyltransferase polypeptide of (1) wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S at position 338, E at position 341, E at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
[0128] (15) The engineered beta-l,2-glycosyltransferase polypeptide of (1) wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S at position 338, E at position 341, a carboxylic acid-presenting amino acid at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
[0129] (16) An engineered beta-l,2-glycosyltransferase polypeptide that has a score greater than 96.0 when scored by the PSSM shown in Table 12.
[0130] (17) An engineered beta-l,2-glycosyltransferase polypeptide that has a score greater than 94.1 when scored by the PSSM shown in Table 13.
[0131] (18) A polypeptide having the sequence of:XXXXXXXXPXLXXGHXXPXLRXAXXLXXXXXXXXXXXXXXXXXXXXXXXXXXYXXXIXLXEXXLXEXPELPXXXHTTNGLP- PHLXXXXXXXXXXXXXXXXXXXXXXXXXXXXXDXXXXXXXXXXXXXXXPXVX XXXXXAXXXXXXXXXXXXPGXXFPFXXIXXXXXXXXXXXXXXXXXPXXXDXLX XXXXXXXLXXTSXXXEXXXXXXXXXXXXXXXXPVGXXFXXXXXXXXXXXXXXX WLXXXXXXSXXXXSFGXEXXLXXXXXXXXXXXLXXXXXXFIXVXRF- PKGXEXXLXXXLPXXXLEXXXXXXXVXDXXVPXXXILXHXXXGXFXSHCGWXSV XESXXXGX- PIIXMPMXXDQPINXXXXXXXGVXXXXXRXXXGXXXXXXXXXXXXXXXXXXXX XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX, wherein residue 1 is Q or L or P or S or T, wherein residue 2 is R or N or Q or L or K or P or S or T, wherein residue 3 is Q or E or I or L or K or M or F or P or T or V, wherein residue 4 is R or D or Q or H or K or S or T or V, wherein residue 5 is I or V, wherein residue 6 is A or L or M or F or T or V, wherein residue 7 is M or F, wherein residue 8 is F or V, wherein residue 10 is F or W, wherein residue 12 is A or G, wherein residue 13 is G or L or F or Y, wherein residue 16 is I or V, wherein residue 17 is N or D or L or K or F or S, wherein residue 19 is L or F, wherein residue 22 is I or L, wherein residue 24 is E or K or M, wherein residue 25 is A or R or D or Q or E or L or K, wherein residue 27 is A or D, wherein residue 28 is A or D or K or P or S, wherein residue 29 is R or L or K, wherein residue 30 is N or D or G, wherein residue 31 is L or M or F or V, wherein residue 32 is R or N or Q or E or H or I or L or K or F or T or Y or V, wherein residue 33 is I or V, wherein residue 34 is H or Y, wherein residue 35 is I or L or V, wherein residue 36 is C or G or L or F or P or S or Y or V, wherein residue 37 is N or L or S, wherein residue 38 is P or S or T, wherein residue 39 is A or R or N or Q or E or K or P or S, wherein residue 40 is I or P or T or V, wherein residue 41 is N or G or V, wherein residue 42 is L or P, wherein residue 43 is N or D or E or L or S, wherein residue 44 is L or M or S, wherein residue 45 is A or E or I or T or V, wherein residue 46 is A or R or E or K or S, wherein residue 47 is K or F, wherein residue 48 is R or E or L, wherein residue 49 is R or I, wherein residue 50 is P or T, wherein residue 51 is A or N or D or E or L or K or P or Y, wherein residue 52 is R or E or G or H or K, wherein residue 54 is A or R or E or K or S, wherein residue 55 is N or D or H or P or S, wherein residue 56 is R or K or S or V, wherein residue 58 is R or Q or E or H or K or F or T, wherein residue 60 is I or V, wherein residue 62 is L or F or Y, wherein residue 63 is A or R or N or Q or E or H or K or P or T, wherein residue 65 is E or P, wherein residue 67 is L or Y, wherein residue 72 is R or E or P or S, wherein residue 73 is E or H or Y, wherein residue 74 is L or Y, wherein residue 85 is N or I or M, wherein residue 86 is G or K or P, wherein residue 87 is I or L or T or V, wherein residue 88 is L or Y, wherein residue 89 is Dor H or K or F or Y, wherein residue 90 is R or D or E or K, wherein residue 91 is A or L, wherein residue 92 is A or L, wherein residue 93 is A or R or N or E or L or K, wherein residue 94 is A or D or E or L or K or M, wherein residue 95 is A or S, wherein residue 96 is A or R or N or Q or L or K or S, wherein residue 97 is D or P, wherein residue 98 is R or N or D or Q or E or K or S or T, wherein residue 99 is L or F or V, wherein residue 100 is A or R or D or E or K or S or V, wherein residue 101 is A or R or N or D or Q or E or H or K, wherein residue 102 is Q or I or L or M or V, wherein residue 103 is I or L or V, wherein residue 104 is A or R or N or D or Q or E or K, wherein residue 105 is A or R or N or D or E or K or S or T, wherein residue 106 is E or I or L or V, wherein residue 107 is R or N or D or Q or E or K, wherein residue 108 is P or V, wherein residue 109 is N or D or T or V, wherein residue 110 is A or L, wherein residue 111 is L or V, wherein residue 112 is I or L or V, wherein residue 113 is A or T or Y, wherein residue 115 is I or F or V, wherein residue 116 is L or M or V, wherein residue 117 is Q or I or L or M or V, wherein residue 118 is L or M or P, wherein residue 119 is W or V, wherein residue 120 is A or N, wherein residue 121 is A or Q or E or P or S, wherein residue 122 is A or R or D or E or G or H or L or K or S or T, wherein residue 123 is A or E or I or L or S or V, wherein residue 124 is A or S, wherein residue 125 is A or R or N or Q or E or L or S or V, wherein residue 126 is A or E or K or S, wherein residue 127 is R or Q or H or L or K or F or T, wherein residue 128 is N or D or G or K, wherein residue 129 is I or V, wherein residue 131 is A or G or S or V, wherein residue 133 is R or Q or G or K or P, wherein residue 134 is L or F, wherein residue 135 is L or F or V, wherein residue 136 is A or T, wherein residue 137 is A or F or S or T or V, wherein residue 138 is G or S, wherein residue 140 is A or R or E, wherein residue 141 is A or L or F or T or V, wherein residue 142 is L or F or W or Y, wherein residue 143 is A or L or S, wherein residue 144 is R or T or Y, wherein residue 145 is A or I or F or Y or V, wherein residue 146 is R or F or W or Y or V, wherein residue 147 is N or D or Q or H, wherein residue 148 is H or F or V, wherein residue 149 is L or V, wherein residue 150 is Q or K or T, wherein residue 151 is R or K, wherein residue 154 is E or H or I or K or S or V, wherein residue 155 is E or P, wherein residue 159 is D or E or P, wherein residue 160 is A or D or E or G, wherein residue 162 is R or P or Y, wherein residue 163 is L or T, wherein residue 164 is A or P or S, wherein residue 165 is E or K, wherein residue 166 is R or I or S or V, wherein residue 167 is E or L, wherein residue 168 is R or Q or L, wherein residue 169 is A or D or P, wherein residue 170 is N or Q or K, wherein residue 171 is A or R or L or M or T, wherein residue 172 is R or Y, wherein residue 173 is A or E, wherein residue 174 is A or L or M or F, wherein residue 175 is L or M or F, wherein residue 176 is A or E or G, wherein residue 177 is D or K or T, wherein residue 178 is A or E or G or Y or V, whereinresidue 180 is N or L or P or T, wherein residue 181 is D or E, wherein residue 182 is D or E or L or K or F, wherein residue 184 is R or F, wherein residue 186 is A or D or P or V, wherein residue 187 is E or K or P or V, wherein residue 188 is A or Q or G or F, wherein residue 189 is A or R or N or Q or K or P or S or T, wherein residue 190 is A or N or I or K or M or S, wherein residue 191 is Q or G or P, wherein residue 192 is I or L or Y or V, wherein residue 193 is A or I or M or T or V, wherein residue 195 is I or L or M or V, wherein residue 196 is C or F, wherein residue 199 is R or Q or E or P, wherein residue 200 is A or E or I or S or T or V, wherein residue 201 is I or V, wherein residue 203 is A or R or E or G or K or P or S, wherein residue 204 is R or E or K or P, wherein residue 205 is E or Y, wherein residue 206 is I or L or V, wherein residue 207 is A or D or E or K, wherein residue 208 is A or Y, wherein residue 209 is A or C or L, wherein residue 210 is A or R or N or Q or E or G or H or S or T, wherein residue 211 is A or N or D or Q or E or G or L or K or S or T or V, wherein residue 212 is L or P, wherein residue 213 is G or L or M or S or T, wherein residue 214 is R or N or E or G or K, wherein residue 215 is A or R or S or T or W or V, wherein residue 216 is R or E or K, wherein residue 217 is I or L or V, wherein residue 218 is I or V, wherein residue 222 is A or P, wherein residue 223 is P or S or T, wherein residue 225 is R or Q or I or L or K or M or S, wherein residue 226 is D or V, wherein residue 227 is L or P, wherein residue 228 is N or D or L, wherein residue 229 is R or P or T, wherein residue 230 is D or E or M, wherein residue 231 is D or S, wherein residue 232 is N or D or I or T or V, wherein residue 233 is D or M or S or Y, wherein residue 234 is D or K, wherein residue 235 is D or Q or E or K or M or P or S, wherein residue 236 is D or E or G or P, wherein residue 237 is L or V, wherein residue 238 is A or I or L or M, wherein residue 239 is A or R or N or D or E or K or S or T, wherein residue 242 is N or D or G, wherein residue 243 is A or N or E or L or K or S or T, wherein residue 244 is Q or K, wherein residue 245 is A or R or D or E or G or K or P, wherein residue 246 is A or D or E or K or P, wherein residue 247 is A or R or N or D or G or H or L or S or Y, wherein residue 249 is T or V, wherein residue 250 is I or V, wherein residue 251 is F or Y or V, wherein residue 252 is C or V, wherein residue 256 is S or T, wherein residue 258 is A or Y, wherein residue 259 is F or Y, wherein residue 261 is P or S, wherein residue 262 is R or E or K or P, wherein residue 263 is E or Y, wherein residue 264 is N or D or Q or E, wherein residue 265 is R or I or L or M or V, wherein residue 266 is A or R or E or H or K or F or T, wherein residue 267 is A or N or E, wherein residue 268 is I or L or V, wherein residue 269 is A or C, wherein residue 270 is R or N or Q or E or H or L or F or S or W or Y, wherein residue 271 is A or G or S, wherein residue 273 is A or E or L, wherein residue 274 is A or R or E or I or L or K, wherein residue 275 is A or S, wherein residue 276 is N or D or G or K,wherein residue 277 is A or Q or E or L or K or Y or V, wherein residue 278 is N or D or Y, wherein residue 281 is F or W, wherein residue 283 is A or L or T or V, wherein residue 289 is E or V, wherein residue 291 is R or Q or G or K or V, wherein residue 292 is R or N or D or H or K or P or S, wherein residue 294 is E or I, wherein residue 295 is N or D or E or G or K or S or T, wherein residue 296 is A or V, wherein residue 299 is A or R or D or E or K or P or V, wherein residue 300 is E or G, wherein residue 301 is F or Y, wherein residue 304 is A or R or E or K, wherein residue 305 is A or I or T or V, wherein residue 306 is G or K, wherein residue 307 is D or E or G, wherein residue 308 is R or K, wherein residue 309 is A or G, wherein residue 310 is R or L or V, wherein residue 312 is R or L or F, wherein residue 314 is H or K, wherein residue 315 is L or M or F, wherein residue 318 is Q or L, wherein residue 319 is A or D or E or L, wherein residue 320 is E or H or L or K, wherein residue 323 is A or R or N or G or S, wherein residue 325 is R or E or G or K or P or S, wherein residue 326 is A or S, wherein residue 327 is L or T, wherein residue 329 is A or G, wherein residue 331 is I or L or M or V, wherein residue 337 is N or S, wherein residue 340 is C or M, wherein residue 343 is I or L or V, wherein residue 344 is R or N or D or Q or H or K or W or Y, wherein residue 345 is N or H or F or Y, wherein residue 347 is I or V, wherein residue 351 is A or C, wherein residue 355 is Q or H or I, wherein residue 356 is L or W or V, wherein residue 362 is A or S, wherein residue 363 is R or E or K, wherein residue 364 is L or F, wherein residue 365 is I or L or M, wherein residue 366 is E or K or V, wherein residue 367 is A or R or D or E or K or S, wherein residue 368 is I or L or K or M or V, wherein residue 371 is A or G, wherein residue 372 is I or K or V, wherein residue 373 is R or E or I or T, wherein residue 374 is I or V, wherein residue 375 is R or E or K or M or P or V, wherein residue 377 is R or N or D or Q, wherein residue 378 is A or D or E or P, wherein residue 379 is N or D or E, wherein residue 381 is R or N or K or S, wherein residue 382 is I or L or F or V, wherein residue 383 is R or N or D or E or H or S, wherein residue 384 is A or R or P or S, wherein residue 385 is A or R or D or E or G or H, wherein residue 386 is A or N or D or E or T or W or V, wherein residue 387 is I or M, wherein residue 388 is A or C or L or K or M, wherein residue 389 is A or R or Q or E or K, wherein residue 390 is A or C or T or V, wherein residue 391 is I or L or V, wherein residue 392 is R or Q or E or K, wherein residue 393 is A or R or N or D or Q or E or G, wherein residue 394 is M or V, wherein residue 395 is I or L or V, wherein residue 396 is A or R or N or D or E or G or H or L or K or M or F or S or T or Y or V, wherein residue 397 is D or E or G, wherein residue 398 is R or D or E or K or P, wherein residue 399 is R or E or I or L or K or T or V, wherein residue 400 is A or G, wherein residue 401 is D or E or I or K or S, wherein residue 402 is R or N or E or G or I or K or S or T or V,wherein residue 403 is I or L, wherein residue 404 is R or H or W, wherein residue 405 is A or R or Q or E or K or M or S, wherein residue 406 is R or N or Q or E or K, wherein residue 407 is A or T or V, wherein residue 408 is A or R or E or L or K, wherein residue 409 is A or R or D or Q or E or K, wherein residue 410 is I or L or M or V, wherein residue 411 is A or G or S, wherein residue 412 is A or R or N or D or Q or E or G or L or K or T, wherein residue 413 is A or R or N or Q or E or L or K or M, wherein residue 414 is I or L or M, wherein residue 415 is R or N or E or K, wherein residue 416 is A or N or Q or E or L or K or S or T, wherein residue 417 is A or R or N or Q or E or I or K or S or T or W or Y or V, wherein residue 418 is R or D or Q or E or G or K or Y, wherein residue 419 is R or D or Q or E or K, wherein residue 420 is A or R or E or K or V, wherein residue 421 is N or E or L or T or V, wherein residue 422 is I or L or M or F or V, wherein residue 423 is A or N or D or E or G or T, wherein residue 424 is A or Ror D or E or G or i or K or W or V, wherein residue 425 is A or F, wherein residue 426 is A or R or L or S or Y or V, wherein residue 427 is A or R or D or Q or E or G or L, wherein residue 428 is A or R or Q or E or L or K or T or V, wherein residue 429 is L or M, wherein residue 430 is A or R or E or I or L or K or M or T or V, wherein residue 431 is A or R or N or Q or E or K, wherein residue 432 is I or L or V, wherein residue 433 is A or C or G or H, wherein residue 434 is A or R or Q or E or L or K or T, wherein residue 435 is A or R or N or D or Q or E or K, wherein residue 436 is R or L or P, wherein residue 437 is A or R or N or L or S or V, wherein residue 438 is A or K, wherein residue 439 is R or H or L or Y, or wherein residue 440 is A or K or V.
[0132] (19) The engineered beta-l,2-glycosyltransferase polypeptide of any one of (1) to (18) that comprises an amino acid sequence having an active site defined by amino acids at positions numbered according to SEQ ID NO: 1, the positions including 9-19, 48, 76-89, 114-117, 134- 138, 141-150, 152, 163-182, 184, 185, 196, 223-226, 230, 232, 251-259, 282, 284, 312, 314- 322, 330, 332-342, 345, 353-359-361).
[0133] (20) The engineered beta-l,2-glycosyltransferase polypeptide of (19) wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S at position 338, E at position 341, D at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
[0134] (21) The engineered beta-l,2-glycosyltransferase polypeptide of (19) wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S at position 338, E at position 341, E at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
[0135] (22) The engineered beta-l,2-glycosyltransferase polypeptide of (19) wherein the amino acid sequence comprises H at position 17, D at position 114, H at position 333, S at position 338, E at position 341, a carboxylic acid-presenting amino acid at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
[0136] (23) The engineered beta-l,2-glycosyltransferase polypeptide of any one of (1) to (22) wherein amino acids at positions 151-180, numbered according to SEQ ID NO: 1, are replaced with 2, 3, or 4 amino acids.
[0137] (24) The engineered beta-l,2-glycosyltransferase polypeptide of any one of (1) to (22) wherein amino acids at positions 151-180, numbered according to sequence number SEQ ID NO: 1, are replaced with an amino acid motif comprising Xn, wherein X is any amino acid and n is any integer equal to or greater than 1.
[0138] (24.1) The engineered beta-l,2-glycosyltransferase polypeptide of any one of (1) to (22) and (24), wherein n is any integer up to 50.
[0139] (25) A method for transferring a sugar moiety to a substrate steviol glycoside, the method comprising contacting a B12GT with one or more steviol glycosides and a non-UDP nucleoside diphosphate-sugar.
[0140] (26) A method for transferring a sugar moiety to a substrate steviol glycoside, the method comprising contacting a B12GT and a sucrose synthase with one or more steviol gly- cosides, a non-UDP nucleoside diphosphate, and sucrose.
[0141] (27) A method for transferring a sugar moiety to a substrate steviol glycoside, the method comprising contacting a B12GT, a B13GT, and a sucrose synthase with one or more steviol glycosides, a non-UDP nucleoside diphosphate, and sucrose.
[0142] (28) A method for transferring a sugar moiety to a substrate steviol glycoside, the method comprising contacting a B12GT, a sucrose synthase, and one or more additional gly- cosyltransferases with one or more steviol glycosides, a non-UDP nucleoside diphosphate, and sucrose.
[0143] (29) The method of (26), wherein the B12GT polypeptide is an engineered B12GT that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449- 879.
[0144] (30) The method of (27), wherein the B12GT polypeptide is an engineered B12GT that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449- 879.
[0145] (31) The method of (26), wherein the SuSy polypeptide is an engineered sucrose syn- thase that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 9-448.
[0146] (32) The method of (27), wherein the SuSy polypeptide is an engineered sucrose syn- thase that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 9-448.
[0147] (33) The method of (26), wherein (a) the B12GT polypeptide is an engineered B12GT that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449- 879, and (b) the SuSy polypeptide is an engineered sucrose synthase that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 9-448.
[0148] (34) The method of (26), wherein the B12GT polypeptide is an engineered B12GT that has a score greater than 96.0 when scored by the PSSM shown in Table 13.
[0149] (35) The method of (26), wherein the B12GT polypeptide is an engineered B12GT that has a score greater than 94.1 when scored by the PSSM shown in Table 14.
[0150] (36) The method of any one of (26) to (35), wherein the substrate steviol glycoside is steviol, steviol- 13-O-glucoside, steviol- 19-O-glucoside, rubusoside, steviol- 1,2-bioside, ste- viol- 1,3-bioside, rubusoside, dulcoside B, dulcoside A, rebaudioside B, rebaudioside G, stevi- oside, rebaudioside C, rebaudioside F, rebaudioside A, rebaudioside I, rebaudioside E, rebau- dioside E2, rebaudioside H, rebaudioside L, rebaudioside K, rebaudioside J, rebaudioside AM, rebaudioside M, rebaudioside D, rebaudioside N, rebaudioside O, rebaudioside Q, rebaudioside X, an isomer thereof, a synthetic steviol glycoside or combinations thereof.
[0151] (37) The method of any one of (26) to (35), wherein the substrate steviol glycoside is a mixture of stevioside and rebaudioside A.
[0152] (38) The method of any one of (26) to (35), wherein the target steviol glycoside is steviol, steviol- 13-O-glucoside, steviol- 19-O-glucoside, rubusoside, steviol- 1,2-bioside, ste- viol- 1,3-bioside, rubusoside, dulcoside B, dulcoside A, rebaudioside B, rebaudioside G, stevi- oside, rebaudioside C, rebaudioside F, rebaudioside A, rebaudioside I, rebaudioside E, rebau- dioside E2, rebaudioside H, rebaudioside L, rebaudioside K, rebaudioside J, rebaudioside AM, rebaudioside M, rebaudioside D, rebaudioside N, rebaudioside O, rebaudioside Q, rebaudioside X, an isomer thereof, a synthetic steviol glycoside or combinations thereof.
[0153] (39) The method of any one of (26) to (35), wherein the target steviol glycoside is a mixture of rebaudioside E, rebaudioside M, and / or rebaudioside D.
[0154] (40) The method of any one of (26) to (35), wherein the target steviol glycoside is rebaudioside D.
[0155] (41) The method of any one of (26) to (35), wherein the target steviol glycoside is rebaudioside E.
[0156] (42) The method of any one of (26) to (35), wherein the target steviol glycoside is rebaudioside M.
[0157] (43) The method of any one of (26) to (42), wherein the non-UDP nucleoside diphos- phate is ADP, GDP, CDP, or TDP.
[0158] (44) The method of any one of (26) to (42), wherein the non-UDP nucleoside diphos- phate is ADP.
[0159] (45) A polynucleotide encoding a polypeptide of any one of (1) to (23).
[0160] (46) A host microorganism heterologously expressing a polynucleotide of (45).
[0161] (Al) An engineered beta-l,2-glycosyltransferase polypeptide comprising a substituted amino acid sequence of SEQ ID NO: 1, wherein the amino acid at position 9 is P, the amino acid at position 10 is W, the amino acid at position 11 is L, the amino acid at position 12 is A or G, the amino acids at positions 13-19 are at least 90% identical to the amino acid sequence of SEQ ID NO: 883, the amino acid at position 48 is R, the amino acids at positions 76-85 are at least 90% identical to the amino acid sequence of SEQ ID NO: 884, the amino acid at posi- tion 86 is G or K, the amino acid at position 87 is T or L, the amino acid at position 88 is L, the amino acid at position 89 is H or K, the amino acids at positions 114-116 comprise the amino acid sequence of SEQ ID NO: 885, the amino acid at position 117 is I, V, or Q, the amino acids at positions 134-196 are replaced with a first cassette, the amino acids at positions 223-259 are replaced with a second cassette, the amino acid at position 282 is V, the amino acid at position 284 is R, and the amino acids at positions 312-361 are replaced with a thirdcassete, wherein the amino acid positions are numbered with reference to the amino acid num- bering of SEQ ID NO: 1.
[0162] (A2) The polypeptide of (Al), wherein the first cassete comprises LLTX1GAALX2X3YX4X5X6FX7X8X9PGX10X11FPFX12X13IRLX14KX15EQX16X17X18REMX 19GTEPX20X21X22DFLX23X24X25X26X27X28X29MLX30C, and wherein Xi is S, T, or V, X2is F or Y, X3is S or A, X4is F or V, X5is F or V, X6is N or H, X7is L or V, X8is K or T, X9is R or K, X10 is H, E, or V, X11 is E or P, X12 is P or E, X13 is A, E, or G, X14 is S or P, X15 is R or V, Xi6 is D or P, X17 is K or Q, Xis is M or L, X19 is F or M, X20 is T or P, X21 is E or D, X22 is L, E, or D, X23 is V, A, or D, X24 is P or E, X25 is A or G, X26 is N, A, Q, P, K, or S, X27 is A, N, S, K, or M, X28 is G, Q, or P, X29 is I, V, L, and X30 is M or I.
[0163] (A3) The polypeptide of either (Al) or (A2), wherein the second cassete comprises PFQDPX1TX2DX3X4X5X6X7LX8X9WLX10X11X12X13X14X15SX16VYVSFGSEX17F, and wherein Xi is L or D, X2 is E, M, or D, X3 is I, D, or V, X4 is Y or D, Xs is D or K, Xe is P, E, K, or D, X7 is E or D, Xs is I or M, X9 is A, R, D, K, or S, X10 is D, G, or N, X11 is T, K, or S, X12 is K or Q, X13 is P, D, E, or K, X14 is E, P, or D, X15 is G, N, H, R, or S, Xi6 is V or T, or X17 is Y or A.
[0164] (A4) The polypeptide of any one of (Al) to (A3), wherein the third cassete comprises X1DX2LVPQX3X4ILNHX5X6TGX7FX8SHCGWNSVMESIX9FGVPIIAMPMQWDQPINA, and wherein Xi is L or R, X2 is H or K, X3 is A or L, X4 is H or K, Xs is P, K, R, or S, Xe is A or S, X7 is G or A, Xs is I, V, or L, and X9 is D or Y.
[0165] (B 1) An engineered beta-l,2-glycosyltransferase polypeptide comprising a substituted amino acid sequence of SEQ ID NO: 1, wherein the amino acid at position 9 is P, the amino acid at position 10 is W, the amino acid at position 11 is L, the amino acid at position 12 is A or G, the amino acids at positions 13-19 are at least 90% identical to the amino acid sequence of SEQ ID NO: 883, the amino acid at position 48 is R, the amino acids at positions 76-85 are at least 90% identical to the amino acid sequence of SEQ ID NO: 884, the amino acid at posi- tion 86 is G or K, the amino acid at position 87 is T or L, the amino acid at position 88 is L, the amino acid at position 89 is H or K, the amino acids at positions 114-116 comprise the amino acid sequence of SEQ ID NO: 885, the amino acid at position 117 is I, V, or Q, the amino acids at positions 134-196 are replaced with a first cassete, the amino acids at positions 223-259 are replaced with a second cassete, the amino acid at position 282 is V, the amino acid at position 284 is R, and the amino acids at positions 312-361 are replaced with a third cassete, wherein the amino acid positions are numbered with reference to the amino acid num-bering of SEQ ID NO: 1, and wherein remaining positions of the substituted amino acid se- quence are at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 1.
[0166] (Cl) An engineered beta-l,2-glycosyltransferase polypeptide comprising a substituted amino acid sequence of SEQ ID NO: 1, wherein the amino acid at position 9 is P, the amino acid at position 10 is W, the amino acid at position 11 is L, the amino acid at position 12 is A or G, the amino acids at positions 13-19 are at least 90% identical to the amino acid sequence of SEQ ID NO: 883, the amino acid at position 48 is R, the amino acids at positions 76-85 are at least 90% identical to the amino acid sequence of SEQ ID NO: 884, the amino acid at posi- tion 86 is G or K, the amino acid at position 87 is T or L, the amino acid at position 88 is L, the amino acid at position 89 is H or K, the amino acids at positions 114-116 comprise the amino acid sequence of SEQ ID NO: 885, the amino acid at position 117 is I, V, or Q, the amino acids at positions 134-196 are replaced with a first cassette, the amino acids at positions 223-259 are replaced with a second cassette, the amino acid at position 282 is V, the amino acid at position 284 is R, and the amino acids at positions 312-361 are replaced with a third cassette, wherein the amino acid positions are numbered with reference to the amino acid num- bering of SEQ ID NO: 1, and wherein the substituted amino acid sequence is at least at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449- 879.
[0167] (D 1) An engineered beta-l,2-glycosyltransferase polypeptide comprising a substituted amino acid sequence of SEQ ID NO: 1, wherein the amino acids at positions 9-19 are at least 90% identical to the amino acid sequence of SEQ ID NO: 880, the amino acid at position 48 is K, the amino acids at positions 76-89 comprise the amino acid sequence of SEQ ID NO: 881, the amino acids at positions 114-117 comprise the amino acid sequence of SEQ ID NO: 882, the amino acids at positions 134-196 are replaced with a first cassette, the amino acids at posi- tions 223-259 are replaced with a second cassette, the amino acid at position 282 is V, the amino acid at position 284 is R, and the amino acids at positions 312-361 are replaced with a third cassette, wherein the amino acid positions are numbered with reference to the amino acid numbering of SEQ ID NO: 1, and wherein the substituted amino acid sequence is at least at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879.
[0168] (El) An engineered beta-l,2-glycosyltransferase polypeptide that comprises a substi- tuted amino acid sequence of SEQ ID NO: 1 in which amino acids 13-19 are at least 90% identical to the amino acid sequence of SEQ ID NO: 883 and in which amino acids 114-116 are at least 90% identical to the amino acid sequence of SEQ ID NO: 885, wherein the amino acid sequence comprises D at position 357 and Q at position 358, wherein the amino acid po- sitions are numbered with reference to the amino acid numbering of SEQ ID NO: 1.
[0169] (Fl) An engineered beta-l,2-glycosyltransferase polypeptide that comprises a substi- tuted amino acid sequence of SEQ ID NO: 1 in which amino acids 9-19 are at least 90% iden- tical to a first cassette and in which amino acids 114-117 are at least 90% identical to a second cassette, wherein the amino acid sequence comprises D at position 357 and Q at position 358, wherein the amino acid positions are numbered with reference to the amino acid numbering of SEQ ID NO: 1.
[0170] (F2) The polypeptide of (Fl), wherein the first cassette comprises PWLXiLGHVNPF, and wherein Xi is A or G.
[0171] (F3) The polypeptide of either (Fl) or (F2), wherein the second cassette comprises DFLXi, and wherein Xi is I, V, or Q.
[0172] (Gl) An engineered beta-l,2-glycosyltransferase polypeptide that comprises a substi- tuted amino acid sequence of SEQ ID NO: 1 in which amino acids 9-19 are at least 90% iden- tical to the amino acid sequence of SEQ ID NO: 880 and in which amino acids 114-117 are at least 90% identical to the amino acid sequence of SEQ ID NO: 882, wherein the amino acid sequence comprises D at position 357 and Q at position 358, wherein the amino acid positions are numbered with reference to the amino acid numbering of SEQ ID NO: 1, and wherein the substituted amino acid sequence is at least at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879.
[0173] (Hl) An engineered beta-l,2-glycosyltransferase polypeptide that is at least at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449- 879 and comprises an active site sequence and a surface site sequence, wherein the active site sequence comprises amino acid positions positioned within 8 angstroms of a ligand when the polypeptide is bound to the ligand.
[0174] (JI) An engineered beta-l,2-glycosyltransferase polypeptide that is at least at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879 and comprises an active site sequence and a surface site sequence, wherein the active site sequence comprises amino acid positions 9-19, 48, 76-89, 114-117, 134-138, 141-150, 152, 163-182, 184-196, 223-226, 230, 232, 251-259, 282, 284, 312, 314-322, 330, 332-342, 345, and 353-361, wherein the surface site sequence comprises amino acid positions 1-8, 20-47, 49- 75, 90-113, 118-133, 139, 140, 151, 153-162, 183, 186-195, 197-222, 227-229, 231, 233-250, 260-281, 283, 285-311, 313, 323-329, 331, 343, and 346-352, and wherein the amino acid positions are numbered with reference to the amino acid numbering of SEQ ID NO: 1.
[0175] (J2) The polypeptide of (JI), wherein the surface site sequence comprises the follow- ing amino acids at respective positions numbered according to the sequence of SEQ ID NO: 1 : the amino acid at position 3 is P, L, I, Q, or F, the amino acid at position 4 is K, R, Q, or S, the amino acid at position 5 is I or V, the amino acid at position 32 is H, K, E, Y, V, I, L, or T, the amino acid at position 54 is K, S, or A, the amino acid at position 55 is D, N, S, or H, the amino acid at position 56 is K, S, or R, the amino acid at position 58 is H, T, Q, E, or K, the amino acid at position 63 is E, R, H, K, A, or Q, the amino acid at position 98 is R, N, E, K, or T, the amino acid at position 100 is D, S, R, A, or E, the amino acid at position 101 is E, K, N, or R, the amino acid at position 104 is K, Q, E, R, or A, the amino acid at position 105 is E, N, T, D, or A, the amino acid at position 107 is K, Q, or N, the amino acid at position 123 is V, I, or L, the amino acid at position 124 is A, the amino acid at position 126 is K, S, E, or A, the amino acid at position 127 is H, Q, L, K, or R, the amino acid at position 199 is R, E, or P, the amino acid at position 204 is K, P, or E, the amino acid at position 207 is D, E, A, or K, the amino acid at position 216 is E, K, or R, the amino acid at position 239 is A, R, D, K, or S, the amino acid at position 247 is G, N, H, R, or S, the amino acid at position 266 is E, H, or R, the amino acid at position 270 is E, R, F, Y, L, or H, the amino acid at position 273 is L or E, the amino acid at position 274 is E, L, A, K, I, or R, the amino acid at position 292 is P, N, R, or H, the amino acid at position 295 is E or D, the amino acid at position 299 is E, K, or P, the amino acid at position 326 is A or S, the amino acid at position 385 is E, D, G, A, or H, the amino acid at position 386 is E, W, D, V, T, or A, the amino acid at position 391 is I, L, or V, the amino acid at position 392 is R or K, the amino acid at position 393 is E, D, A, or R, the amino acid at position 396 is Y, S, F, V, N, G, or H, the amino acid at position 402 is E, K, T, I, R, N, or V, the amino acid at position 406 is K, Q, N, or R, the amino acid at position 408 is A, R, or K, the amino acid at position 409 is K, E, D, R, or A, the amino acid at position 412 is A, E, K, or L, the amino acid at position 413 is R, Q, K, N, E, or A, the amino acid at position 417 is A, R, E, K, I, or N, the amino acid at position 420 is E, A, R, or K, the amino acid at position423 is D, N, T, or A, the amino acid at position 424 is R, A, E, G, or I, the amino acid at position 428 is R, E, A, T, or V, and the amino acid at position 434 is R, L, E, Q, or T.
[0176] (KI) An engineered beta-l,2-glycosyltransferase polypeptide that comprises 13 beta strands designated Pi, P2, P3, P4, Ps, Pe, P?, Ps, P$>, Pio, P11, P12, and P13, and 18 alpha helices designated ai, 02, OB, a4, as, ae, a?, as, a$>, aio, an, an, ai3, ai4, ais, aie, ai7, and ais,
[0177] (K2) The polypeptide of (KI), wherein Pi comprises amino acid positions 4-8, P2 com- prises amino acid positions 32-37, P3 comprises amino acid positions 57-61, P4 comprises amino acid positions 110-114, Ps comprises amino acid positions 131-135, Pe comprises amino acid positions 191-195, P? comprises amino acid positions 216-219, Ps comprises amino acid positions 249-253, P9 comprises amino acid positions 278-283, Pio comprises amino acid po- sitions 309-312, P11 comprises amino acid positions 329-332, P12 comprises amino acid posi- tions 349-351 and P13 comprises amino acid positions 371-373, wherein the amino acid posi- tions are numbered with reference to the amino acid numbering of SEQ ID NO: 1.
[0178] (K3) The polypeptide of claim either (KI) or (K2), wherein ai comprises amino acid positions 13-29, 02 comprises amino acid positions 39-48, as comprises amino acid positions 51-56, a4 comprises amino acid positions 82-106, as comprises amino acid positions 118-127, ae comprises amino acid positions 139-150, a? comprises amino acid positions 165-177, as comprises amino acid positions 181-184, a9 comprises amino acid positions 199-213, aiocom- prises amino acid positions 232-243, an comprises amino acid positions 262-275, an com- prises amino acid positions 293-296, ai3 comprises amino acid positions 301-305, ai4 com- prises amino acid positions 318-323, ais comprises amino acid positions 336-345, ai6 com- prises amino acid positions 357-368, ai7 comprises amino acid positions 384-396, and ais com- prises amino acid positions 398-439, wherein the amino acid positions are numbered with ref- erence to the amino acid numbering of SEQ ID NO: 1
[0179] (K4) The polypeptide of any one of (KI) to (K3), wherein at least one amino acid of one or more of beta strands designated Pi, P2, P3, P4, Ps, Pe, P7, Ps, P9, Pio, P11, P12, and P13, and alpha helices designated ai, a2, a3, a4, as, ae, a7, as, a9, aio, an, an, ai3, ai4, ais, aie, ai7, and ais of SEQ ID NOS: 449-879 are substituted with a corresponding at least one amino acid from a corresponding beta strand or alpha helix from a different polypeptide selected from the group consisting of SEQ ID NOS: 449-879.
[0180] (K5) The polypeptide of (K4), wherein the entire corresponding beta strand or helix is substituted.
[0181] (K6) The polypeptide of any one of (KI) to (K5), wherein at least one amino acid of one or more of beta strands designated Pi, P2, P3, P4, Ps, Pe, P7, Ps, P9, Pio, P11, P12, and P13, andalpha helices designated ai, a2, as, a4, as, ae, a?, as, a9, aio, an, an, ai3, ai4, ais, ai6, ai7, and ais of SEQ ID NOS: 449-879 are substituted with a corresponding at least one amino acid from a corresponding beta strand or alpha helix from SEQ ID NO: 500.
[0182] (K7) The polypeptide of any one of (KI) to (K5), wherein a wherein at least one amino acid of one or more of the beta strand designated [34 and the alpha helix designated ai of SEQ ID NOS: 449-879 are substituted with a corresponding at least one amino acid from a corre- sponding beta strand or alpha helix from SEQ ID NO: 500.
[0183] (K8) The polypeptide of any one of (KI) to (K5), wherein a wherein at least one amino acid of one or more of the beta strand designated [34 and the alpha helix designated ai of SEQ ID NOS: 449-879 are substituted with a corresponding at least one amino acid from a corre- sponding beta strand or alpha helix from SEQ ID NO: 449-879.
Claims
CLAIMS:
1. An engineered beta-l,2-glycosyltransferase polypeptide that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879.
2. The engineered beta-l,2-glycosyltransferase polypeptide of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99% or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-452.
3. The engineered beta-l,2-glycosyltransferase polypeptide of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99% or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 453-511.
4. The engineered beta-l,2-glycosyltransferase polypeptide of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 512-571.
5. The engineered beta-l,2-glycosyltransferase polypeptide of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 572-594.
6. The engineered beta-l,2-glycosyltransferase of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 595-604.
7. The engineered beta-l,2-glycosyltransferase of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 605-635.
8. The engineered beta-l,2-glycosyltransferase of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 636-706.
9. The engineered beta-l,2-glycosyltransferase of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 707-788.
10. The engineered beta-l,2-glycosyltransferase of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 789-831.
11. The engineered beta-l,2-glycosyltransferase of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 832-873.
12. The engineered beta-l,2-glycosyltransferase of claim 1 that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 874-879.
13. The engineered beta-l,2-glycosyltransferase polypeptide of claim 1 wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S at position 338, E at position 341, D at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
14. The engineered beta-l,2-glycosyltransferase polypeptide of claim 1 wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S atposition 338, E at position 341, E at position 357, and Q at position 358, numbered according to SEQ ID NO:
1.
15. The engineered beta-1,2-glycosyltransferase polypeptide of claim 1 wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S at position 338, E at position 341, a carboxylic acid-presenting amino acid at position 357, and Q at position 358, numbered according to SEQ ID NO:
1.
16. An engineered beta-1,2-glycosyltransferase polypeptide that has a score greater than 96.0 when scored by the PSSM shown in Table 12.
17. An engineered beta-1,2-glycosyltransferase polypeptide that has a score greater than 94.1 when scored by the PSSM shown in Table 13.
18. A polypeptide having the sequence of: XXXXXXXXPXLXXGHXXPXLRXAXXLXXXXXXXXXXXXXXXXXXXXXXXXXXY XXXIXLXEXXLXEXPELPXXXHTTNGLPPHLXXXXXXXXXXXXXXXXXXXXXXXX XXXXXDXXXXXXXXXXXXXXXPXVXXXXXXAXXXXXXXXXXXXPGXXFPFXXI XXXXXXXXXXXXXXXXXPXXXDXLXXXXXXXXLXXTSXXXEXXXXXXXXXXXX XXXXPVGXXFXXXXXXXXXXXXXXXWLXXXXXXSXXXXSFGXEXXLXXXXXXX XXXXLXXXXXXFIXVXRFPKGXEXXLXXXLPXXXLEXXXXXXXVXDXXVPXXXIL XHXXXGXFXSHCGWXSVXESXXXGXPIIXMPMXXDQPINXXXXXXXGVXXXXXR XXXGXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX XXXXXXXXXXXX wherein residue 1 is Q or L or P or S or T, wherein residue 2 is R or N or Q or L or K or P or S or T, wherein residue 3 is Q or E or I or L or K or M or F or P or T or V, wherein residue 4 is R or D or Q or H or K or S or T or V, wherein residue 5 is I or V, wherein residue 6 is A or L or M or F or T or V, wherein residue 7 is M or F, wherein residue 8 is F or V, wherein residue 10 is F or W, wherein residue 12 is A or G,wherein residue 13 is G or L or F or Y, wherein residue 16 is I or V, wherein residue 17 is N or D or L or K or F or S, wherein residue 19 is L or F, wherein residue 22 is I or L, wherein residue 24 is E or K or M, wherein residue 25 is A or R or D or Q or E or L or K, wherein residue 27 is A or D, wherein residue 28 is A or D or K or P or S, wherein residue 29 is R or L or K, wherein residue 30 is N or D or G, wherein residue 31 is L or M or F or V, wherein residue 32 is R or N or Q or E or H or I or L or K or F or T or Y or V, wherein residue 33 is I or V, wherein residue 34 is H or Y, wherein residue 35 is I or L or V, wherein residue 36 is C or G or L or F or P or S or Y or V, wherein residue 37 is N or L or S, wherein residue 38 is P or S or T, wherein residue 39 is A or R or N or Q or E or K or P or S, wherein residue 40 is I or P or T or V, wherein residue 41 is N or G or V, wherein residue 42 is L or P, wherein residue 43 is N or D or E or L or S, wherein residue 44 is L or M or S, wherein residue 45 is A or E or I or T or V, wherein residue 46 is A or R or E or K or S, wherein residue 47 is K or F, wherein residue 48 is R or E or L, wherein residue 49 is R or I, wherein residue 50 is P or T, wherein residue 51 is A or N or D or E or L or K or P or Y, wherein residue 52 is R or E or G or H or K, wherein residue 54 is A or R or E or K or S,wherein residue 55 is N or D or H or P or S, wherein residue 56 is R or K or S or V, wherein residue 58 is R or Q or E or H or K or F or T, wherein residue 60 is I or V, wherein residue 62 is L or F or Y, wherein residue 63 is A or R or N or Q or E or H or K or P or T, wherein residue 65 is E or P, wherein residue 67 is L or Y, wherein residue 72 is R or E or P or S, wherein residue 73 is E or H or Y, wherein residue 74 is L or Y, wherein residue 85 is N or I or M, wherein residue 86 is G or K or P, wherein residue 87 is I or L or T or V, wherein residue 88 is L or Y, wherein residue 89 is D or H or K or F or Y, wherein residue 90 is R or D or E or K, wherein residue 91 is A or L, wherein residue 92 is A or L, wherein residue 93 is A or R or N or E or L or K, wherein residue 94 is A or D or E or L or K or M, wherein residue 95 is A or S, wherein residue 96 is A or R or N or Q or L or K or S, wherein residue 97 is D or P, wherein residue 98 is R or N or D or Q or E or K or S or T, wherein residue 99 is L or F or V, wherein residue 100 is A or R or D or E or K or S or V, wherein residue 101 is A or R or N or D or Q or E or H or K, wherein residue 102 is Q or I or L or M or V, wherein residue 103 is I or L or V, wherein residue 104 is A or R or N or D or Q or E or K, wherein residue 105 is A or R or N or D or E or K or S or T, wherein residue 106 is E or I or L or V, wherein residue 107 is R or N or D or Q or E or K,wherein residue 108 is P or V, wherein residue 109 is N or D or T or V, wherein residue 110 is A or L, wherein residue 111 is L or V, wherein residue 112 is I or L or V, wherein residue 113 is A or T or Y, wherein residue 115 is I or F or V, wherein residue 116 is L or M or V, wherein residue 117 is Q or I or L or M or V, wherein residue 118 is L or M or P, wherein residue 119 is W or V, wherein residue 120 is A orN, wherein residue 121 is A or Q or E or P or S, wherein residue 122 is A or R or D or E or G or H or L or K or S or T, wherein residue 123 is A or E or I or L or S or V, wherein residue 124 is A or S, wherein residue 125 is A or R or N or Q or E or L or S or V, wherein residue 126 is A or E or K or S, wherein residue 127 is R or Q or H or L or K or F or T, wherein residue 128 is N or D or G or K, wherein residue 129 is I or V, wherein residue 131 is A or G or S or V, wherein residue 133 is R or Q or G or K or P, wherein residue 134 is L or F, wherein residue 135 is L or F or V, wherein residue 136 is A or T, wherein residue 137 is A or F or S or T or V, wherein residue 138 is G or S, wherein residue 140 is A or R or E, wherein residue 141 is A or L or F or T or V, wherein residue 142 is L or F or W or Y, wherein residue 143 is A or L or S, wherein residue 144 is R or T or Y, wherein residue 145 is A or I or F or Y or V,wherein residue 146 is R or F or W or Y or V, wherein residue 147 is N or D or Q or H, wherein residue 148 is H or F or V, wherein residue 149 is L or V, wherein residue 150 is Q or K or T, wherein residue 151 is R or K, wherein residue 154 is E or H or I or K or S or V, wherein residue 155 is E or P, wherein residue 159 is D or E or P, wherein residue 160 is A or D or E or G, wherein residue 162 is R or P or Y, wherein residue 163 is L or T, wherein residue 164 is A or P or S, wherein residue 165 is E or K, wherein residue 166 is R or I or S or V, wherein residue 167 is E or L, wherein residue 168 is R or Q or L, wherein residue 169 is A or D or P, wherein residue 170 is N or Q or K, wherein residue 171 is A or R or L or M or T, wherein residue 172 is R or Y, wherein residue 173 is A or E, wherein residue 174 is A or L or M or F, wherein residue 175 is L or M or F, wherein residue 176 is A or E or G, wherein residue 177 is D or K or T, wherein residue 178 is A or E or G or Y or V, wherein residue 180 is N or L or P or T, wherein residue 181 is D or E, wherein residue 182 is D or E or L or K or F, wherein residue 184 is R or F, wherein residue 186 is A or D or P or V, wherein residue 187 is E or K or P or V, wherein residue 188 is A or Q or G or F,wherein residue 189 is A or R or N or Q or K or P or S or T, wherein residue 190 is A or N or I or K or M or S, wherein residue 191 is Q or G or P, wherein residue 192 is I or L or Y or V, wherein residue 193 is A or I or M or T or V, wherein residue 195 is I or L or M or V, wherein residue 196 is C or F, wherein residue 199 is R or Q or E or P, wherein residue 200 is A or E or I or S or T or V, wherein residue 201 is I or V, wherein residue 203 is A or R or E or G or K or P or S, wherein residue 204 is R or E or K or P, wherein residue 205 is E or Y, wherein residue 206 is I or L or V, wherein residue 207 is A or D or E or K, wherein residue 208 is A or Y, wherein residue 209 is A or C or L, wherein residue 210 is A or R or N or Q or E or G or H or S or T, wherein residue 211 is A or N or D or Q or E or G or L or K or S or T or V, wherein residue 212 is L or P, wherein residue 213 is G or L or M or S or T, wherein residue 214 is R or N or E or G or K, wherein residue 215 is A or R or S or T or W or V, wherein residue 216 is R or E or K, wherein residue 217 is I or L or V, wherein residue 218 is I or V, wherein residue 222 is A or P, wherein residue 223 is P or S or T, wherein residue 225 is R or Q or I or L or K or M or S, wherein residue 226 is D or V, wherein residue 227 is L or P, wherein residue 228 is N or D or L, wherein residue 229 is R or P or T, wherein residue 230 is D or E or M,wherein residue 231 is D or S, wherein residue 232 is N or D or I or T or V, wherein residue 233 is D or M or S or Y, wherein residue 234 is D or K, wherein residue 235 is D or Q or E or K or M or P or S, wherein residue 236 is D or E or G or P, wherein residue 237 is L or V, wherein residue 238 is A or I or L or M, wherein residue 239 is A or R or N or D or E or K or S or T, wherein residue 242 is N or D or G, wherein residue 243 is A or N or E or L or K or S or T, wherein residue 244 is Q or K, wherein residue 245 is A or R or D or E or G or K or P, wherein residue 246 is A or D or E or K or P, wherein residue 247 is A or R or N or D or G or H or L or S or Y, wherein residue 249 is T or V, wherein residue 250 is I or V, wherein residue 251 is F or Y or V, wherein residue 252 is C or V, wherein residue 256 is S or T, wherein residue 258 is A or Y, wherein residue 259 is F or Y, wherein residue 261 is P or S, wherein residue 262 is R or E or K or P, wherein residue 263 is E or Y, wherein residue 264 is N or D or Q or E, wherein residue 265 is R or I or L or M or V, wherein residue 266 is A or R or E or H or K or F or T, wherein residue 267 is A or N or E, wherein residue 268 is I or L or V, wherein residue 269 is A or C, wherein residue 270 is R or N or Q or E or H or L or F or S or W or Y, wherein residue 271 is A or G or S, wherein residue 273 is A or E or L,wherein residue 274 is A or R or E or I or L or K, wherein residue 275 is A or S, wherein residue 276 is N or D or G or K, wherein residue 277 is A or Q or E or L or K or Y or V, wherein residue 278 is N or D or Y, wherein residue 281 is F or W, wherein residue 283 is A or L or T or V, wherein residue 289 is E or V, wherein residue 291 is R or Q or G or K or V, wherein residue 292 is R or N or D or H or K or P or S, wherein residue 294 is E or I, wherein residue 295 is N or D or E or G or K or S or T, wherein residue 296 is A or V, wherein residue 299 is A or R or D or E or K or P or V, wherein residue 300 is E or G, wherein residue 301 is F or Y, wherein residue 304 is A or R or E or K, wherein residue 305 is A or I or T or V, wherein residue 306 is G or K, wherein residue 307 is D or E or G, wherein residue 308 is R or K, wherein residue 309 is A or G, wherein residue 310 is R or L or V, wherein residue 312 is R or L or F, wherein residue 314 is H or K, wherein residue 315 is L or M or F, wherein residue 318 is Q or L, wherein residue 319 is A or D or E or L, wherein residue 320 is E or H or L or K, wherein residue 323 is A or R or N or G or S, wherein residue 325 is R or E or G or K or P or S, wherein residue 326 is A or S, wherein residue 327 is L or T, wherein residue 329 is A or G,wherein residue 331 is I or L or M or V, wherein residue 337 is N or S, wherein residue 340 is C or M, wherein residue 343 is I or L or V, wherein residue 344 is R or N or D or Q or H or K or W or Y, wherein residue 345 is N or H or F or Y, wherein residue 347 is I or V, wherein residue 351 is A or C, wherein residue 355 is Q or H or I, wherein residue 356 is L or W or V, wherein residue 362 is A or S, wherein residue 363 is R or E or K, wherein residue 364 is L or F, wherein residue 365 is I or L or M, wherein residue 366 is E or K or V, wherein residue 367 is A or R or D or E or K or S, wherein residue 368 is I or L or K or M or V, wherein residue 371 is A or G, wherein residue 372 is I or K or V, wherein residue 373 is R or E or I or T, wherein residue 374 is I or V, wherein residue 375 is R or E or K or M or P or V, wherein residue 377 is R or N or D or Q, wherein residue 378 is A or D or E or P, wherein residue 379 is N or D or E, wherein residue 381 is R or N or K or S, wherein residue 382 is I or L or F or V, wherein residue 383 is R or N or D or E or H or S, wherein residue 384 is A or R or P or S, wherein residue 385 is A or R or D or E or G or H, wherein residue 386 is A or N or D or E or T or W or V, wherein residue 387 is I or M, wherein residue 388 is A or C or L or K or M, wherein residue 389 is A or R or Q or E or K,wherein residue 390 is A or C or T or V, wherein residue 391 is I or L or V, wherein residue 392 is R or Q or E or K, wherein residue 393 is A or R or N or D or Q or E or G, wherein residue 394 is M or V, wherein residue 395 is I or L or V, wherein residue 396 is A or R or N or D or E or G or H or L or K or M or F or S or T or Y or V, wherein residue 397 is D or E or G, wherein residue 398 is R or D or E or K or P, wherein residue 399 is R or E or I or L or K or T or V, wherein residue 400 is A or G, wherein residue 401 is D or E or I or K or S, wherein residue 402 is R or N or E or G or I or K or S or T or V, wherein residue 403 is I or L, wherein residue 404 is R or H or W, wherein residue 405 is A or R or Q or E or K or M or S, wherein residue 406 is R or N or Q or E or K, wherein residue 407 is A or T or V, wherein residue 408 is A or R or E or L or K, wherein residue 409 is A or R or D or Q or E or K, wherein residue 410 is I or L or M or V, wherein residue 411 is A or G or S, wherein residue 412 is A or R or N or D or Q or E or G or L or K or T, wherein residue 413 is A or R or N or Q or E or L or K or M, wherein residue 414 is I or L or M, wherein residue 415 is R or N or E or K, wherein residue 416 is A or N or Q or E or L or K or S or T, wherein residue 417 is A or R or N or Q or E or I or K or S or T or W or Y or V, wherein residue 418 is R or D or Q or E or G or K or Y, wherein residue 419 is R or D or Q or E or K, wherein residue 420 is A or R or E or K or V, wherein residue 421 is N or E or L or T or V, wherein residue 422 is I or L or M or F or V,wherein residue 423 is A or N or D or E or G or T, wherein residue 424 is A or R or D or E or G or I or K or W or V, wherein residue 425 is A or F, wherein residue 426 is A or R or L or S or Y or V, wherein residue 427 is A or R or D or Q or E or G or L, wherein residue 428 is A or R or Q or E or L or K or T or V, wherein residue 429 is L or M, wherein residue 430 is A or R or E or I or L or K or M or T or V, wherein residue 431 is A or R or N or Q or E or K, wherein residue 432 is I or L or V, wherein residue 433 is A or C or G or H, wherein residue 434 is A or R or Q or E or L or K or T, wherein residue 435 is A or R or N or D or Q or E or K, wherein residue 436 is R or L or P, wherein residue 437 is A or R or N or L or S or V, wherein residue 438 is A or K, wherein residue 439 is R or H or L or Y, or wherein residue 440 is A or K or V.
19. The engineered beta-l,2-glycosyltransferase polypeptide of claim 1 that comprises an amino acid sequence having an active site defined by amino acids at positions numbered according to SEQ ID NO: 1, the positions including 9-19, 48, 76-89, 114-117, 134-138, 141- 150, 152, 163-182, 184, 185, 196, 223-226, 230, 232, 251-259, 282, 284, 312, 314-322, 330, 332-342, 345, 353-359-361, wherein the polypeptide is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to the active site sequence of an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879.
20. The engineered beta-l,2-glycosyltransferase polypeptide of claim 19 wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S at position 338, E at position 341, D at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
21. The engineered beta-l,2-glycosyltransferase polypeptide of claim 19 wherein the amino acid sequence comprises H at position 15, D at position 114, H at position 333, S atposition 338, E at position 341, E at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
22. The engineered beta-l,2-glycosyltransferase polypeptide of claim 19, wherein the amino acid sequence comprises H at position 17, D at position 114, H at position 333, S at position 338, E at position 341, a carboxylic acid-presenting amino acid at position 357, and Q at position 358, numbered according to SEQ ID NO: 1.
23. The engineered beta-l,2-glycosyltransferase polypeptide of claim 1 wherein amino acids at positions 151-180, numbered according to SEQ ID NO: 1, are replaced with 2, 3, or 4 amino acids.
24. The engineered beta-l,2-glycosyltransferase polypeptide of claim 1 wherein amino acids at positions 151-180, numbered according to sequence number SEQ ID NO: 1, are replaced with an amino acid motif comprising Xn, wherein X is any amino acid and n is any integer equal to or greater than 1.
25. A method for transferring a sugar moiety to a substrate steviol glycoside, the method comprising contacting a B 12GT with one or more steviol glycosides and a non-UDP nucleoside diphosphate-sugar.
26. A method for transferring a sugar moiety to a substrate steviol glycoside, the method comprising contacting a B12GT and a sucrose synthase with one or more steviol glycosides, a non-UDP nucleoside diphosphate, and sucrose.
27. A method for transferring a sugar moiety to a substrate steviol glycoside, the method comprising contacting a B12GT, a B13GT, and a sucrose synthase with one or more steviol glycosides, a non-UDP nucleoside diphosphate, and sucrose.
28. A method for transferring a sugar moiety to a substrate steviol glycoside, the method comprising contacting a B12GT, a sucrose synthase, and one or more additional glycosyltransferases with one or more steviol glycosides, a non-UDP nucleoside diphosphate, and sucrose.
29. The method of claim 26, wherein the B12GT polypeptide is an engineered B12GT that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879.
30. The method of claim 27, wherein the B12GT polypeptide is an engineered B12GT that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879.
31. The method of claim 26, wherein the SuSy polypeptide is an engineered sucrose synthase that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 9-448.
32. The method of claim 27, wherein the SuSy polypeptide is an engineered sucrose synthase that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 9-448.
33. The method of claim 26, wherein(a) the B12GT polypeptide is an engineered Bl 2GT that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 449-879, and(b) the SuSy polypeptide is an engineered sucrose synthase that comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 9-448.
34. The method of claim 26, wherein the B12GT polypeptide is an engineered B12GT that has a score greater than 96.0 when scored by the PSSM shown in Table 13.
35. The method of claim 26, wherein the B12GT polypeptide is an engineered B12GT that has a score greater than 94. 1 when scored by the PSSM shown in Table 14.
36. The method of any of claims 26 to 35, wherein the substrate steviol glycoside is steviol, steviol- 13 -O-glucoside, steviol- 19-O-glucoside, rubusoside, steviol- 1,2-bioside, steviol- 1,3 -bioside, rubusoside, dulcoside B, dulcoside A, rebaudioside B, rebaudioside G, stevioside, rebaudioside C, rebaudioside F, rebaudioside A, rebaudioside I, rebaudioside E, rebaudioside E2, rebaudioside H, rebaudioside L, rebaudioside K, rebaudioside J, rebaudioside AM, rebaudioside M, rebaudioside D, rebaudioside N, rebaudioside O, rebaudioside Q, rebaudioside X, an isomer thereof, a synthetic steviol glycoside or combinations thereof.
37. The method of any of claims 26 to 35, wherein the substrate steviol glycoside is a mixture of stevioside and rebaudioside A.
38. The method of any of claims 26 to 35, wherein the target steviol glycoside is steviol, steviol- 13-O-glucoside, steviol- 19-O-glucoside, rubusoside, steviol- 1,2-bioside, steviol-1,3- bioside, rubusoside, dulcoside B, dulcoside A, rebaudioside B, rebaudioside G, stevioside, rebaudioside C, rebaudioside F, rebaudioside A, rebaudioside I, rebaudioside E, rebaudioside E2, rebaudioside H, rebaudioside L, rebaudioside K, rebaudioside J, rebaudioside AM, rebaudioside M, rebaudioside D, rebaudioside N, rebaudioside O, rebaudioside Q, rebaudioside X, an isomer thereof, a synthetic steviol glycoside or combinations thereof.
39. The method of any of claims 26 to 35, wherein the target steviol glycoside is a mixture of rebaudioside E, rebaudioside M, and / or rebaudioside D.
40. The method of any of claims 26 to 35, wherein the target steviol glycoside is rebaudioside D.
41. The method of any of claims 26 to 35, wherein the target steviol glycoside is rebaudioside E.
42. The method of any of claims 26 to 35, wherein the target steviol glycoside is rebaudioside M.
43. The method of any of claims 26 to 42, wherein the non-UDP nucleoside diphosphate is ADP, GDP, CDP, or TDP.
44. The method of any of claims 26 to 42, wherein the non-UDP nucleoside diphosphate is ADP.
45. A polynucleotide encoding a polypeptide of any of claims 1-23.
46. A host microorganism heterologously expressing a polynucleotide of claim 45.