Compositions and methods for producing dihydrofurans from keto sugars

JP2025525311A5Pending Publication Date: 2026-06-19ARZEDA CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ARZEDA CORP
Filing Date
2023-06-14
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Current methods for producing furandicarboxylic acid (FDCA) and dihydrofuran rely on chemical catalytic processes that use high-cost catalysts and lack efficient biocatalytic alternatives.

Method used

A biocatalytic method using glycoside hydrolases to convert 2-keto-3-deoxygluconic acid (KDG) into dihydrofuran under specific pH and temperature conditions, followed by dehydration to produce 5-hydroxymethyl-2-furoic acid (HMFA) and further oxidation to FDCA, utilizing engineered enzymes with specific sequence identities.

Benefits of technology

Achieves yields of at least 40% HMFA and enables the production of FDCA through biocatalytic pathways comparable to chemical catalytic routes, offering a cost-effective and environmentally friendly alternative.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Compositions and methods are provided for producing dihydrofurans by glycosyl hydrolases capable of dehydrating 2-keto-3-deoxy-gluconic acid (KDG) to K4. Compositions and methods are also provided for further processing K4 to produce HMFA (5-hydroxymethyl-2-furoic acid) and / or FDCA (2,5-furandicarboxylic acid).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 352,145, filed June 14, 2022, the entire contents of which are incorporated herein by reference.

[0002] Electronic Sequence Listing Reference The contents of the electronic sequence listing (ARZE_035_03WO_SeqList_ST26.xml, size: 144,098 bytes, created on June 12, 2023) are incorporated herein by reference in their entirety. [Background technology]

[0003] Furandicarboxylic acid (FDCA) is one of the key components of polyethylene 2,5-furandicarboxylate (PEF), a plant-based polymer with 50-70% lower carbon emissions than its petroleum-based competitor, polyethylene terephthalate (PET). Currently, PET dominates the global packaging market due to its strength, light weight, and shatter-resistant properties. The majority of PET is used to make synthetic fibers, with the remaining 30% used to manufacture bottles. However, PEF offers several advantages over PET. When combined with PET, it can be reused up to five times more than PET alone, decomposes much faster than PET, and can replace film packaging that is currently non-recyclable.

[0004] Previous methods for producing FDCA or more basic natural carbohydrate starting materials such as dihydrofuran have focused on chemical catalytic methods in industry. However, these methods involve the use of high-cost catalysts, which also present significant limitations and process drawbacks. Currently, there are no biocatalytic processes employed in industry that can produce FDCA yields comparable to those of chemical catalytic routes.

[0005] Therefore, there is a need to provide new enzymes and methods for producing dihydrofurans from keto sugars using biocatalytic pathways under thermochemically favorable conditions. Summary of the Invention

[0006] Provided herein are methods for producing dihydrofurans from keto sugars.

[0007] In one embodiment, the present disclosure relates to a biocatalytic method for producing dihydrofuran, the method comprising contacting 2-keto-3-deoxygluconic acid (KDG) with a glycoside hydrolase, thereby producing dihydrofuran, the contacting comprising: a. a pH of about 4 to about 7 as measured by a pH meter; b. a temperature of 45°C to 74°C, or both, thereby producing dihydrofuran. In one embodiment, the method comprises a pH of about 4 to 5. In one embodiment, the method comprises b., and a temperature of 70°C to 74°C. In one embodiment, the method comprises a. and b., and a pH of about 4 to 5 and a temperature of about 62°C to 72°C. In one embodiment, the method comprises a. and b., and the pH and temperature are selected from the group consisting of: a. a pH of about 4 and a temperature of about 63°C; b. a pH of about 4.5 and a temperature of about 69°C; and c. a pH of about 5 and a temperature of about 72°C. In one embodiment, the method comprises c. In one embodiment, the KDG is 180 mM to 300 mM. In one embodiment, the KDG is 180 mM to 220 mM. In one embodiment, the glycoside hydrolase comprises a protein having at least 80% sequence identity to a sequence selected from the group consisting of SEQ ID NOs: 1-116. In one embodiment, the sequence identity is at least 85%, 90%, 95%, 98%, 99%, or 100%. In one embodiment, the glycoside hydrolase comprises a first motif that binds to KDG and a second motif comprising a catalytic residue, wherein the catalytic residue comprises an aspartic acid, the first motif comprises at least two residues, the first residue comprises an arginine, and the second residue comprises a tryptophan, phenylalanine, or tyrosine, and the glycoside hydrolase is a homolog of SEQ ID NO: 1 or SEQ ID NO: 19 as determined by SWISS-MODEL modeling. In one embodiment, the glycoside hydrolase comprises a protein having 100% sequence identity to SEQ ID NO: 27. In one embodiment, the contacting is for 0.5 hours to 24 hours. In one embodiment, the contacting is for 0.5 hours to 5 hours. In one embodiment, the contacting is for about 3 hours.In one embodiment, the method further comprises dehydrating the dihydrofuran to produce 5-hydroxymethyl-2-furoic acid (HMFA), and a yield of HMFA of at least 40% is observed after dehydration. In one embodiment, the dehydration comprises contacting the dihydrofuran with an acid selected from the group consisting of formic acid, hydrochloric acid, sulfuric acid, phosphoric acid, nitric acid, hydrobromic acid, and Ci-6 carboxylic acid. In one embodiment, the acid is formic acid. In one embodiment, the method further comprises oxidizing the HMFA to produce 2,5-furandicarboxylic acid (FDCA). In one embodiment, the oxidation comprises a chemical oxidation reaction. In one embodiment, the oxidation comprises an enzymatic oxidation reaction.

[0008] In one embodiment, the disclosure relates to an isolated polypeptide comprising at least 85% identity to any one of SEQ ID NOs: 35 to 116. In embodiments, the disclosure further relates to a biocatalytic method for producing dihydrofuran, the method comprising contacting 2-keto-3-deoxygluconic acid (KDG) with the isolated polypeptide to produce dihydrofuran. In one embodiment, the contacting comprises: a. a pH of about 4 to about 7 as measured by a pH meter; b. a temperature of 45°C to 74°C, or both, thereby producing dihydrofuran. In one embodiment, the method comprises a) wherein the pH is about 4 to 5. In one embodiment, the method comprises b) wherein the temperature is 70°C to 74°C. In one embodiment, the method comprises a) and b) wherein the pH is about 4 to 5 and the temperature is about 62°C to 72°C. In one embodiment, the method includes a. and b., wherein the pH and temperature are selected from the group consisting of: a. pH about 4, temperature about 63°C; b. pH about 4.5, temperature about 69°C; and c. pH about 5, temperature about 72°C. In one embodiment, the method includes c. In one embodiment, the KDG is about 180 mM to 300 mM. In one embodiment, the KDG is about 180 mM to 220 mM. In one embodiment, the contacting is for 0.5 hours to 24 hours. In one embodiment, the contacting is for up to 5 hours. In one embodiment, the contacting is for about 3 hours. In one embodiment, the method further includes dehydrating the dihydrofuran to produce 5-hydroxymethyl-2-furoic acid (HMFA), wherein a yield of HMFA of at least 40% is observed after dehydration. In one embodiment, the dehydration comprises contacting dihydrofuran with an acid selected from the group consisting of formic acid, hydrochloric acid, sulfuric acid, phosphoric acid, nitric acid, hydrobromic acid, and Ci-6 carboxylic acid; in one embodiment, the acid is formic acid. In one embodiment, the method further comprises oxidizing HMFA to produce 2,5-furandicarboxylic acid (FDCA). In one embodiment, the oxidation comprises a chemical oxidation reaction. In one embodiment, the oxidation comprises an enzymatic oxidation reaction. In one embodiment, the present disclosure relates to a composition comprising an isolated polypeptide.

[0009] In one embodiment, the disclosure relates to an engineered microorganism comprising an exogenous glycoside hydrolase, wherein the exogenous glycoside hydrolase comprises a sequence having at least 85% sequence identity to a sequence selected from the group consisting of SEQ ID NOs: 1-116. In one embodiment, the exogenous glycoside hydrolase comprises a sequence having at least 90%, 95%, 97%, 98%, 99%, or 100% sequence identity to a sequence selected from the group consisting of SEQ ID NOs: 1-116. In one embodiment, the exogenous glycoside hydrolase comprises SEQ ID NO: 27. In one embodiment, the engineered microorganism is a bacterium. In one embodiment, the bacterium is selected from the group consisting of E. coli, Saccharomyces spp., Aspergillus spp., Pichia spp., Pseudomonas spp., and Bacillus spp. In one embodiment, the bacterium is E. coli. In one embodiment, the disclosure relates to a composition comprising the engineered microorganism. In one embodiment, the present disclosure relates to a method for producing polyethylene 2,5-furandicarboxylate (PEF), the method comprising a biocatalytic method.

[0010] In one embodiment, the disclosure relates to an isolated non-natural glycoside hydrolase, the isolated non-natural glycoside hydrolase comprising a first motif that binds 2-keto-3-deoxygluconate, the first motif comprising at least two residues, the first residue being arginine and the second residue being tryptophan, phenylalanine, or tyrosine, and a second motif comprising a catalytic residue, the catalytic residue being aspartic acid or glutamic acid, wherein the isolated non-natural glycoside hydrolase has at least 20% identity to SEQ ID NO: 1 or SEQ ID NO: 19. In one embodiment, the isolated non-natural glycoside hydrolase comprises at least 25%, 45%, 65%, 85%, 95%, or 99% identity to SEQ ID NO: 1 or SEQ ID NO: 19. In one embodiment, the second residue of the at least two residues in the first motif is tryptophan and / or the catalytic residue of the second motif is aspartic acid. In one embodiment, the first motif has a sequence of RxQTW, where x is serine or an aliphatic amino acid. In one embodiment, the first motif has a sequence of Rx1QTW(2x2)Yx2Y, where x1 is serine and x2 is an aliphatic amino acid. In one embodiment, the second motif has a sequence of xD, where x is an aliphatic amino acid. In one embodiment, the second motif has a sequence of (2x)KSE(3x)DT(2M)xSxPFx, where x is serine or an aliphatic amino acid. In one embodiment, the arginine of the first motif and the catalytic residue of the second motif are separated by about 70 residues.

[0011] In one embodiment, the disclosure relates to an isolated glycoside hydrolase, the isolated glycoside hydrolase having formula 1: P(19x)LPP(4x)HYHQGVxLxG(4x)(W / Y)(10x)Y(3x)(Y / W)x(D / E)(6x)G(9x)D(2x)Q(P / A)Gx(L / I)L(2x)L( 7x)(R / K)Y(2x)(A / G)(3x)(L / I)(9x)(T / N)xE(G / Q)G(F / Y)(W / F)H(K / N)(3x)Px(Q / E )(M / Q)WLDGLYMxG(5x)Y(A / G)(9x)D(4x)Q(6x)(H / K)(T / M)(R / K)(3x)TGL(2x)H(A / G) (W / F)(D / S)(2x)(R / K)(3x)W(A / S)(D / N)(2x)(T / S)Gx(S / A)PExW(G / A)R(S / A)xGW(9 x)(D / E)x(I / L)P(2x)H(20x)Q(4x)GxWxQ(V / I)x(D / N)(K / R)(G / V)(4x)NW(L / P)ExSx The isolated glycoside hydrolase has at least 50% identity to SEQ ID NO:1 with a sequence of (S / T)xL(6x)K(G / A)(15x)(K / Q)(A / G)(F / Y)xG(18x)(I / V)C(I / V)GT(S / G)xGxY(5x)R(5x)D(L / M)HG(V / A)GA(F / L), where x is any amino acid. In one embodiment, the isolated glycoside hydrolase has at least 55%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, or 100% identity to SEQ ID NO:1. In one embodiment, the isolated glycoside hydrolase further comprises a Y at residue 41, a D at residue 88, an H at residue 132, a W at residue 141, a D at residue 143, an M at residue 147, an H at residue 189, a W at residue 211, an A at residue 212, an R at residue 213, a W at residue 217, an S at residue 278, an L or M at residue 282, a C at residue 330, and an H or K at residue 352, wherein these residues are numbered according to SEQ ID NO: 1. In one embodiment, the isolated glycoside hydrolase further comprises a loop region having at least 21 residues. In one embodiment, the isolated glycoside hydrolase further comprises a loop region having at least one modification selected from the group consisting of residues 331-336 of SEQ ID NO: 1.In one embodiment, the at least one modification selected from the group consisting of positions 331-336 of SEQ ID NO: 1 comprises one or more of: V331K, G332E, G332M, G332V, S334A, S334C, S334D, S334E, S334G, S334I, S334K, S334M, S334N, S334Q, S334R, S334T, S334V, A335V, A335P, A335L, A335C. In one embodiment, the isolated glycoside hydrolase further comprises two amino acids added to the N-terminus of the sequence and five amino acids added to the C-terminus of the sequence.

[0012] In one embodiment, the disclosure relates to a biocatalytic method for producing dihydrofuran, the method comprising contacting 2-keto-3-deoxygluconic acid (KDG) with a glycoside hydrolase, thereby producing dihydrofuran, the contacting comprising: a. a pH of about 4 to about 7 as measured by a pH meter; and b. a temperature of 45°C to 74°C, or both, thereby producing dihydrofuran, wherein the glycoside hydrolase has at least 50% identity to SEQ ID NO: 1, and the glycoside hydrolase comprises a D at residue 143, an R at residue 213, and a W at residue 217 relative to SEQ ID NO: 1. In one embodiment, residue 143 of the glycoside hydrolase is a catalytic residue, and residues 213 and 217 are substrate-binding residues.

[0013] In one embodiment, the disclosure relates to an isolated glycoside hydrolase, the isolated glycoside hydrolase having formula 2: (F / Y)P(8x)(W / Y)(7x)W(T / M)(2x)F(2x)G(2x)(W / Y)(2x)Y(11x)(A / G)(10x)(L / I)(8x)(H / F)D(L / I)GF(4x)(S / T)(4x)(W / Y)(15x)(A / G)(13x)(L / I)(16x)IDx(L / M)(L / M)(N / S) (22x)H(3x)(T / S)(5x)RxDxS(S / T)(6x)(D / N)(10x)TxQG(4x)SxW(A / S / T)RG(Q / L)(A / T)W(2x)YG(28x)PxD(4x)(Y / W)D(F / L)(12x)S(6x)(S / C)(33x)Y(30x)(W / F / Y)(G / A)DY(Y / F)(2x)ExL, wherein x is any amino acid, and has at least 50% homology to SEQ ID NO: 19. In one embodiment, the isolated glycoside hydrolase has at least 55%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, or 100% homology to SEQ ID NO: 19. In one embodiment, the isolated glycoside hydrolase further comprises an I or L at residue 26, an H at residue 41, a W at residue 42, an M at residue 43, an H at residue 87, a D at residue 88, a G at residue 134, a D at residue 149, a T at residue 150, an M at residue 152, a Q at residue 193, a W at residue 219, an R at residue 221, a W at residue 225, an S at residue 280, an I at residue 284, and an F at residue 352, wherein the residues are numbered according to SEQ ID NO: 19. In one embodiment, the isolated glycoside hydrolase comprises a sequence selected from the group consisting of SEQ ID NOs: 24-27. In one embodiment, the isolated glycoside hydrolase further comprises 13 amino acids added to the N-terminus of the sequence and 5 amino acids added to the C-terminus of the sequence.

[0014] These and other embodiments are described below.

[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate some exemplary embodiments and / or features, but are not the only or exclusive of them. It is intended that the embodiments and drawings disclosed herein be considered illustrative and not limiting. [Brief explanation of the drawings]

[0016] [Figure 1A] The conversion (dehydration) of 2-keto-3-deoxygluconic acid (KDG) (1) to 4,5-dihydro-4-hydroxy-5-hydroxymethyl-2-furancarboxylic acid (K4) (2) and the further conversion (dehydration) of K4 to 5-hydroxymethyl-2-furoic acid (HMFA) (3) are shown. [Figure 1B] This is a two-dimensional view of the KDG. [Figure 2A] Illustrated is SEQ ID NO: 1, a KDG-linked glycoside hydrolase. The catalytic residues arginine, tryptophan, and aspartic acid are conserved between both glycosyl hydrolase 88 and 105 family enzymes. [Figure 2B] The active site of the KDG-bound glycoside hydrolase SEQ ID NO: 1 is depicted. The catalytic residues arginine (1), tryptophan (2), and aspartate (3) are conserved among both glycosyl hydrolase 88 and 105 family enzymes. [Figure 3] 1 shows an illustration of the geometric parameters in the active site of glycosyl hydrolase of SEQ ID NO: 1. Eight geometric parameters (d indicates distance, θ indicates bond angle) are specified to describe the spatial position of functional groups relative to KDG. [Figure 4]Figure 1 shows K4 detection as determined by liquid chromatography / mass spectrometry (LC / MS) traces of glycosyl hydrolase of SEQ ID NO:27 converting KDG to K4 under favorable thermochemical conditions (sodium acetate, pH 4-5, 150 mM NaCl, 63-74 °C, 3 h, 1 uM enzyme, and 180 mM KDG substrate), ultimately producing HMFA. Using glycoside hydrolase of SEQ ID NO:27, KDG is dehydrated to K4, which then spontaneously and irreversibly further dehydrates to HMFA. [Figure 5] Figure 1 shows K4 detection as determined by LC / MS trace of glycoside hydrolase of SEQ ID NO: 1 converting KDG to K4 under favorable thermochemical conditions (sodium acetate, pH 5, 25 mM KDG, 45 °C, 3 h), ultimately producing HMFA. Glycoside hydrolase of SEQ ID NO: 1 is used to dehydrate KDG to the detectable intermediate, K4. [Figure 6] An alignment of the sequences disclosed herein (SEQ ID NOS: 24-27) is shown. The catalytic aspartic acid residue (D) is indicated by an *, as are the substrate binding residues arginine (R) and tryptophan (W). [Figure 7] An alignment of residues important for functionality is shown with reference to SEQ ID NO: 19. Corresponding residues from other sequences are provided. Note the high degree of correspondence despite the spatial separation of the amino acids. For SEQ ID NOs: 23-27, additional common residues are shown in bold. DETAILED DESCRIPTION OF THE INVENTION

[0017] The following description includes information that may be useful in understanding the present disclosure. No admission is made that any information provided herein is prior art or relevant to the presently claimed disclosure, or that any publication specifically or implicitly referenced is prior art.

[0018] definition Although the following terms are believed to be well understood by those of ordinary skill in the art, the following definitions are provided to facilitate description of the presently disclosed subject matter.

[0019] Unless otherwise defined below, all technical and scientific terms used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to technology used herein are intended to refer to technologies as commonly understood in the art, and include variations of these technologies and / or equivalent technology alternatives that are apparent to those skilled in the art.

[0020] As used herein, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise.

[0021] When immediately preceding a numerical value, the term "about" or "approximately" refers to a range (e.g., ±10% of that value). For example, unless the context of the disclosure indicates otherwise or contradicts such an interpretation, "about 50" can refer to 45 to 55, "about 25,000" can refer to 22,500 to 27,500, etc. For example, in a list of numerical values, e.g., "about 49, about 50, about 55, ...," "about 50" refers to a range spanning less than half the interval(s) of the preceding or following values, e.g., a range greater than 49.5 and less than 52.5. Furthermore, the terms "about," "less than," or "above," should be understood in light of the definition of the term "about" provided herein. Similarly, when preceding a series of numerical values or ranges of values (e.g., "about 10, 20, 30," or "about 10 to 30"), the term "about" refers to all values in the series or to the endpoints of the range, respectively.

[0022] As used herein, the terms "microorganism" or "microorganism" should be interpreted broadly. These terms are used interchangeably and include, but are not limited to, the two prokaryotic domains, bacteria and archaea, and certain eukaryotic fungi and protists. In some embodiments, the present disclosure refers to the "microorganisms" or "microorganisms" in the lists and figures present in this disclosure. This characterization can refer not only to the identified taxonomic genera, but also to the identified taxonomic species, as well as various novel and newly identified or engineered strains of any organism in the tables or figures. The same characterization applies to the listing of these terms elsewhere in this specification, such as in the Examples.

[0023] When referring to nucleic acid or protein sequences, the term "identity" is used to indicate the similarity between two sequences. Sequence similarity or identity can be determined using standard techniques known in the art, for example, but not limited to, the local identity algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), the sequence identity alignment algorithm of Needleman and Wunsch, J. Mol. Biol. 48:443 (1970), the search for similarity method of Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444 (1988), computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), the BestFit sequence program described by Devereux et al., Nucl. Acid Res. 12, 387-395 (1984), or by inspection. Another suitable algorithm is the BLAST algorithm described in Altschul et al., J. Mol. Biol. 215, 403-410, (1990) and Karlin et al., Proc. Natl. Acad. Sci. USA 90, 5873-5787 (1993). A particularly useful BLAST program is the WU-BLAST-2 program, Altschul et al., Methods in Enzymology, 266, 460-480 (1996), obtained from blast.wustl / edu / blast / README.html. WU-BLAST-2 uses several search parameters, which are preferably set to default values. These parameters are dynamic values established by the program itself depending on the composition of the sequence and the composition of the particular database in which the sequence of interest is searched, but the values can also be adjusted to increase sensitivity.Yet another useful algorithm is gapped BLAST, as reported by Altschul et al. (1997) Nucleic Acids Res. 25, 3389-3402. Unless otherwise specified, percent identity as described herein is determined using the algorithm available at the following internet address: blast.ncbi.nlm.nih.gov / Blast.cgi.

[0024] As used herein, an "isolated" or "purified" polynucleotide or polypeptide, or a biologically active portion thereof, is substantially or essentially free from components that normally accompany or normally interact with the polynucleotide or polypeptide found in its natural environment. Thus, an isolated or purified polynucleotide or polypeptide is substantially free of other cellular material or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. Optionally, an "isolated" polynucleotide is free of sequences (optimally protein-coding sequences) that naturally flank the polynucleotide in the genomic DNA of the organism from which the polynucleotide is derived (i.e., sequences located at the 5' and 3' ends of the polynucleotide). For example, in various embodiments, an isolated polynucleotide can contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleotide sequences that naturally flank the polynucleotide in the genomic DNA of the cell from which the polynucleotide is derived. A polypeptide that is substantially free of cellular material includes preparations of polypeptide having less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of contaminating protein.

[0025] As used herein, the term "nucleic acid" refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides, or analogs thereof. The term refers to the primary structure of the molecule and thus includes double- and single-stranded DNA, as well as double- and single-stranded RNA. It also includes modified nucleic acids, such as methylated and / or capped nucleic acids, nucleic acids containing modified bases, and backbone modifications.

[0026] As used herein, "SWISS-MODEL" refers to a fully automated protein structure homology modeling server for homology modeling of 3D protein structures, accessible via the Expasy web server or through the program DeepView (Swiss Pdb-Viewer). SWISS-MODEL consists of three integrated components: (1) the SWISS-MODEL pipeline—a suite of software tools and databases for automated protein structure modeling; (2) the SWISS-MODEL Workspace—a web-based graphical user workbench; and (3) the SWISS-MODEL Repository—a continuously updated database of homology models for a set of model organism proteomes of biomedical interest. Use of the SWISS-MODEL pipeline involves four major steps involved in building a homology model of a given protein structure: (1) Identification of structural template(s). BLAST and HHblist are used to identify templates. Templates are stored in the SWISS-MODEL Template Library (SMTL), which is derived from the PDB. (2) Alignment of the target sequence and template sequence(s). (3) Model building and energy minimization. SWISS-MODEL implements a rigid fragment assembly approach for modeling. (4) Assessment of model quality using QMEAN, the statistical likelihood of mean force.

[0027] The present disclosure provides enzymatic and biocatalytic processes for producing dihydrofuran and downstream products such as 5-hydroxymethyl-2-furoic acid (HMFA), 2,5-furandicarboxylic acid (FDCA), furandicarboxylic acid methyl ester (FDME), and polyethylene 2,5-furandicarboxylate (PEF). FDME is the methyl ester of FDCA, a derivative that can be polymerized with ethylene glycol to produce PEF. Also provided are methods that include biocatalytic processes for producing dihydrofuran by contacting a substrate keto sugar with a glycoside hydrolase, thereby producing dihydrofuran. The dihydrofuran can be further processed to produce HMFA. In some embodiments, HMFA can also be further processed chemically or biocatalytically to produce FDCA. FDCA can be utilized to produce PEF.

[0028] Also provided are isolated and / or modified glycoside hydrolases that can be used to carry out the dehydration of keto sugars, such as 2-keto-3-deoxygluconic acid (KDG).

[0029] glycoside hydrolase Provided herein are one or more glycoside hydrolases or motifs thereof. Glycoside hydrolases are enzymes that can hydrolyze glycosidic bonds between carbohydrates or between carbohydrates and non-carbohydrate moieties. In some embodiments, glycoside hydrolases can be used in biocatalytic reactions to dehydrate substrates such as KDG.

[0030] Glycosyl hydrolases are classified into families based on sequence similarity. This classification is available on the CAZy (CArbohydrate-Active EnZymes) website. Because protein folds are better conserved than their sequences, some families can be grouped into "clans." In some embodiments, a glycoside hydrolase can be part of any of the 128 families and / or any of the identified clans of glycosyl hydrolases.

[0031] In some embodiments, the glycoside hydrolase is from the GH88 and / or GH105 family. In some embodiments, the glycoside hydrolase has the classification E.C3.2.1.179 and / or E.C3.2.1.172. In some embodiments, the glycoside hydrolase comprises the shape and / or active site provided in Figure 2A, Figure 2B, and / or Figure 3.

[0032] In some embodiments, computational methods can be used to identify and / or design glycoside hydrolase geometries capable of effecting dehydration reactions with substrates such as KDG. A linear representation of KDG is shown in Figure 1B. An example of such a reaction is shown in Figure 1A. The biological dehydration reaction of a ketosugar (e.g., KDG1) to a dihydrofuran (e.g., FDCA2) can be carried out via a glycosyl hydrolase (which then dehydrates to HMFA3). The protonation of the hydroxyl group of KDG1 occurs, in one example, by aspartate to sufficiently activate it for elimination. Referring to Figures 2A and 2B, the overall structural fold of glycosyl hydrolase enzymes is partially shown to be an (alpha / alpha)6 fold. Furthermore, Figure 2B illustrates the key residues of glycosyl hydrolases involved in the reaction. For example, aspartate 3 can act as a general acid / base. In the first step, aspartate 3 acts as an acid to donate a proton to the leaving hydroxyl group at the anomeric carbon C1 of the substrate (e.g., KDG). Aspartate 3 (i.e., the catalytic residue) is located approximately 3.5 Å from the anomeric carbon C1 of the substrate. In the second step, aspartate 3 acts as a base to abstract a proton from the anomeric carbon C2, thereby facilitating the dehydration reaction. Substrate binding occurs through interactions between the ketosugar carboxylic acid group on C1 and the arginine 1 and tryptophan 2 residues in the active site.

[0033] In some embodiments, the glycoside hydrolase comprises an active site geometry set forth in Table 1. Table 1 provides exemplary active site / residue geometries for glycoside hydrolase polypeptides useful for KDG dehydration, with reference to Figure 3. For example, Table 1 lists the distances (e.g., d1, d2, d3, d4) and angles (e.g., θ1, θ2, θ3, θ4) at each key residue, as illustrated in Figure 3, for each of the key residues.

[0034] In some embodiments, residues #1 and #2 in Table 1 coordinate a carboxylate group at the anomeric carbon C1 of KDG in the active site. In some embodiments, residues #1 and #2 are arginine (Arg) and tryptophan (Trp), respectively. In some embodiments, residue #3 is catalytic and functions as a proton donor, adding a proton to the leaving hydroxyl group at the anomeric carbon C1 of KDG. As noted above, catalytic residue #3 is approximately 3.5 Å away from the anomeric carbon C1 bearing the leaving hydroxyl group. In some embodiments, residues #1, #2, and #3 perform glycoside hydrolase activity in the KDG dehydration reaction.

[0035] [Table 1]

[0036] In some embodiments, glycoside hydrolases comprising substantially similar shapes as presented in Table 1 are also contemplated. For example, glycoside hydrolases can also comprise residues within about 1.5, 1, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, or 0.1 angstroms of the values shown in Table 1, e.g., residue #1, residue #2, and residue #3. In some embodiments, glycoside hydrolases comprise residues within at least 0.5 angstroms of the values shown in Table 1, e.g., residue #1, residue #2, and residue #3. Similarly, for example, glycoside hydrolases can comprise residues within about 1.0, 2.5, 5.0, 10, 15, 30, 45, 60, 90, 120, 150, or 180 degrees of the values shown in Table 1, e.g., residue #1, residue #2, and residue #3.

[0037] In embodiments, a glycoside hydrolase is provided that comprises a shape described in Table 1. In embodiments, an isolated polypeptide encoding a glycoside hydrolase that comprises a shape described in Table 1 is also provided.

[0038] Glycoside hydrolase motifs Also provided are motifs for the described glycoside hydrolases. A "motif" is intended to refer to a portion of a polynucleotide or a portion of an amino acid sequence. The motif may retain activity for KDG dehydration. Thus, the polynucleotide sequence motif may range from at least about 2 nucleotides, about 10 nucleotides, about 20 nucleotides, about 50 nucleotides, or about 100 nucleotides, up to the full-length polynucleotide sequence corresponding to the glycoside hydrolase. In some embodiments, the glycoside hydrolase motif is at least about 2, 5, 8, 10, 35, 60, 85, 110, 135, 160, 185, 210, 235, 260, 285, 310, 335, 360, 385, 410, 435, 460, 485, 510, 535, 560, 585, 610, 635, 660, 685, 710, 735, 760, 785, 810, 835, 860, 885, 910, 935, 960, 985, or up to about 1000 amino acid residues, or about the total number of amino acid residues present in a full-length glycoside hydrolase, such as any of SEQ ID NOs: 1-116.

[0039] In some embodiments, the glycoside hydrolase motif comprises a biologically active portion of a glycoside hydrolase that is capable of at least partially dehydrating KDG. In some embodiments, the motif comprises or has a substantially similar residue shape as set forth in any of Table 1.

[0040] In some embodiments, a glycoside hydrolase motif can be prepared by isolating a portion of a polynucleotide encoding a polypeptide capable of KDG dehydration, expressing the encoded portion of the polypeptide capable of KDG dehydration (e.g., by recombinant expression in vitro), and assaying for KDG dehydration activity.

[0041] In some embodiments, the glycoside hydrolase is provided in any form. In some embodiments, the glycoside hydrolase is present in DNA, RNA, protein, or a combination thereof. In some embodiments, the glycoside hydrolase is provided in isolated form. In some embodiments, the glycoside hydrolase is provided as a whole cell system.

[0042] In some embodiments, provided herein is a recombinant glycoside hydrolase polypeptide comprising an amino acid sequence that is at least 10% to at least 99.73% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-116. In some embodiments, the recombinant glycoside hydrolase polypeptide comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-116. In some embodiments, the glycoside hydrolase polypeptide further comprises a tag amino acid sequence. In some embodiments, the tag amino acid sequence is His6.

[0043] In some embodiments, the composition comprising the glycoside hydrolase is at least partially pure. In some embodiments, the composition comprising the glycoside hydrolase is substantially pure. The purity of the glycoside hydrolase can vary, and can be provided, for example, as a crude, semi-purified, or purified enzyme preparation. In some embodiments, the glycoside hydrolase polypeptide is free of impurities.

[0044] Modified glycoside hydrolases The present disclosure also provides modified glycoside hydrolases. In some embodiments, the glycoside hydrolases provided herein comprise one or more modifications. The modifications can be in any region of the glycoside hydrolase. In some embodiments, the modifications are within the active site. Modifications to the active site increase the binding efficiency and / or catalytic efficiency of the glycoside hydrolase for a substrate, such as KDG. In some embodiments, the modified glycoside hydrolases provided herein comprise modifications around catalytic residues to recognize, bind, and / or be more catalytically efficient for KDG. In some embodiments, the glycoside hydrolase is modified to comprise a shape described in Table 1 or a shape substantially similar thereto. In some embodiments, the glycoside hydrolase is modified to improve a shape described in Table 1 to increase recognition, binding, and / or catalytic efficiency for a substrate, such as KDG. Such modifications can be informed by the experimental results shown in Table 3. Various reaction conditions were evaluated for each of SEQ ID NOS: 1-116, and the reaction results were evaluated.

[0045] Modified glycoside hydrolases can be produced using any means. In some embodiments, a nucleotide sequence or amino acid sequence is modified to produce a recombinant glycoside hydrolase. In some embodiments, the amino acid sequence is modified. Modifications include one or more substitutions, deletions, insertions, and combinations thereof. Modifications can include the use of natural amino acid residues, synthetic amino acid residues, or combinations thereof. In some embodiments, modifications include substitutions.

[0046] In some embodiments, polynucleotides encoding glycoside hydrolase polypeptides are modified. Modified polynucleotides can include deletions. In some cases, the deletions are truncations of bases at the 5' and / or 3' ends and / or deletions of one or more nucleotides at one or more internal sites within the native polynucleotide. In some cases, the modification includes the insertion of one or more bases at either the 5', 3', and / or one or more internal sites of the polynucleotide. In some embodiments, the modification includes the substitution of one or more nucleotides at one or more sites in the polynucleotide. In the case of polynucleotides, the modification can include conservative modifications. Conservative modifications can include amino acid substitutions in proteins that change a given amino acid to a different amino acid with similar biochemical properties (e.g., charge, hydrophobicity, and / or size). In some embodiments, conservative modifications include sequences that encode any amino acid sequence of a polypeptide capable of KDG dehydration due to the degeneracy of the genetic code.

[0047] In some embodiments, modified glycoside hydrolase refers to a modified sequence encoding a polypeptide or protein. Protein modifications can include deletions, truncations, additions, substitutions, or a combination thereof. In some embodiments, the glycoside hydrolase protein is modified by truncation at either the 5' and / or 3' end. In some embodiments, the glycoside hydrolase-encoding sequence is modified by addition, deletion, or both of one or more residues in either the 5', 3', and / or internal regions. Modified glycoside hydrolases can retain biological activity. In some embodiments, modified glycoside hydrolases retain equivalent biological activity compared to unmodified glycoside hydrolases. In some cases, biological activity can be reduced. In some cases, biological activity can be increased by modification.

[0048] In some embodiments, the modified glycoside hydrolase comprises a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to a sequence selected from SEQ ID NOs: 1-116, or an active variant, fragment, or modified version thereof. In some embodiments, the modified glycoside hydrolase comprises an amino acid sequence that is at least 10% to at least 99.73% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-116. In some embodiments, the modified glycoside hydrolase comprises an amino acid sequence that is at least 82% identical to SEQ ID NO: 1. In some embodiments, the modified glycoside hydrolase comprises an amino acid sequence that is at least 88% identical to SEQ ID NO: 19. In some embodiments, the modified glycoside hydrolase comprises an amino acid sequence that is at least 85%, 87%, 89%, 91%, 93%, 95%, 97%, 99%, or 100% identical to SEQ ID NO: 27. In some embodiments, the computationally designed glycoside hydrolase comprises SEQ ID NOs: 35-116.

[0049] In some embodiments, a polynucleotide encoding a modified glycoside hydrolase is also provided. In some embodiments, a polynucleotide encoding a modified glycoside hydrolase comprising any one of SEQ ID NOs: 1-116 is provided.

[0050] In some embodiments, a glycoside hydrolase or motif thereof comprises a KDG dehydrating activity and comprises an active site having, or a substantially similar catalytic residue shape, a catalytic residue shape as set forth in Table 1. In some embodiments, a glycoside hydrolase or motif thereof comprises a KDG dehydrating activity and comprises an active site having a catalytic residue shape as set forth in Table 1, further comprises an amino acid sequence having at least 10%, 20%, 30%, 40%, 75%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 1-116.

[0051] In some embodiments, the glycoside hydrolase comprises an active site having a catalytic residue shape set forth in any of Table 1, or a catalytic residue shape substantially similar thereto, and further has (a) at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-34; (i) the amino acid residue in the encoded polypeptide corresponding to amino acid position 40 of SEQ ID NO:1 comprises histidine or glycine or cystine or serine or tyrosine or phenylalanine or isoleucine or asparagine or glutamine or aspartic acid or glutamic acid, (ii) the amino acid residue in the encoded polypeptide corresponding to amino acid position 41 of SEQ ID NO:1 comprises tyrosine or tryptophan, and (iii) the amino acid residue in the encoded polypeptide corresponding to amino acid position 42 of SEQ ID NO:1 comprises histidine or glycine or cystine or serine or tyrosine or phenylalanine or isoleucine or asparagine or glutamine or aspartic acid or glutamic acid. (iv) the amino acid residue in the encoded polypeptide corresponding to amino acid position 88 of SEQ ID NO:1 comprises aspartic acid, leucine, isoleucine, phenylalanine, asparagine, or lysine; (v) the amino acid residue in the encoded polypeptide corresponding to amino acid position 132 of SEQ ID NO:1 comprises histidine, alanine, leucine, valine, serine, cystine, proline, aspartic acid, glutamic acid, asparagine, arginine, glycine, or glutamine; (vi) the amino acid residue in the encoded polypeptide corresponding to amino acid position 146 of SEQ ID NO:1 comprises tyrosine, phenylalanine, or methionine; (vii) the amino acid residue in the encoded polypeptide corresponding to amino acid position 147 of SEQ ID NO:1 comprises methionine, alanine, leucine, or phenylalanine;(viii) the amino acid residue in the encoded polypeptide corresponding to amino acid position 189 of SEQ ID NO:1 comprises histidine or arginine or glutamine or valine or alanine, (ix) the amino acid residue in the encoded polypeptide corresponding to amino acid position 211 of SEQ ID NO:1 comprises tryptophan, (x) the amino acid residue in the encoded polypeptide corresponding to amino acid position 214 of SEQ ID NO:1 comprises serine or alanine or glycine, (xi) the amino acid residue in the encoded polypeptide corresponding to amino acid position 220 of SEQ ID NO:1 comprises methionine or tyrosine or valine or leucine or alanine or glycine or phenylalanine, (xii) the amino acid residue in the encoded polypeptide corresponding to amino acid position 278 of SEQ ID NO:1 comprises serine, and (xiii) the amino acid residue in the encoded polypeptide corresponding to amino acid position 282 of SEQ ID NO:1 comprises leucine or methionine or isoleucine or valine or phenylalanine. (xiv) the amino acid residue in the encoded polypeptide corresponding to amino acid position 352 of SEQ ID NO: 1 is histidine, tyrosine, tryptophan, phenylalanine, lysine, valine, or arginine; (xv) the amino acid residues in the encoded polypeptide corresponding to amino acids 211 to 217 of SEQ ID NO: 1 are W(A / G / S / T)R(G / A / S)(N / Q / I / L / M)(G / A / T)W. (xvi) amino acid residues in the encoded polypeptide corresponding to amino acids 141 to 145 of SEQ ID NO: 1 include a fragment of (W / I)(L / I / C / A / V / S)D(G / D / A / C / T / V / I / N)(L / M / I / V), and / or (xvii) amino acid residues in the encoded protein that correspond to the amino acid positions of SEQ ID NO: 1 listed in Table 1 and also correspond to the specific amino acid substitutions listed in (d)(i) to (d)(xvii) above, or any combination of these residues.

[0052] In some embodiments, the glycoside hydrolase comprises a point mutation selected from the group consisting of H42I, H42T, H42V, H42W, D88C, D88N, H132E, K133M, W141Y, H189A, H189V, V331K, G332E, G332M, G332V, S334A, S334C, S334D, S334E, S334G, S334I, S334K, S334M, S334N, S334Q, S334R, S334T, S334V, A335V, A335P, A335L, A335C, H352R, and H352V. In some embodiments, the glycoside hydrolase comprises a modification at a residue selected from the group consisting of 332, 334, and 335 of SEQ ID NO: 1. In some embodiments, the glycoside hydrolase comprises a modification in the loop region. In some embodiments, the modification in the loop region comprises any one of the residues at amino acid positions 331-336 of SEQ ID NO: 1. In some embodiments, the glycoside hydrolase comprises a point mutation in SEQ ID NO: 19 selected from the group consisting of D41A, D41C, D41E, D41N, D41S, D41T, H87A, H87C, H87E, H87G, H87Q, H87R, H87S, L152M, L152N, W225Y, S337A, Y338V, H339A, H339N, W352F, Y356A, Y356C, Y356F, and Y356H. In embodiments, the point mutation results in increased activity compared to an otherwise equivalent glycoside hydrolase lacking the point mutation.

[0053] In some embodiments, the glycoside hydrolases provided herein may not contain catalytic residue configurations as set forth in any of Table 1, yet retain KDG dehydration activity.

[0054] In some embodiments, an isolated polypeptide is provided comprising a sequence encoding any of the provided glycoside hydrolases. In several embodiments, the isolated polypeptide comprises a first motif that binds to 2-keto-3-deoxygluconate and a second motif comprising a catalytic residue. In several embodiments, the first motif comprises at least two residues, the first residue comprising arginine and the second residue comprising tryptophan, phenylalanine, or tyrosine. In several embodiments, the second residue comprises tryptophan. In several embodiments, the second residue comprises phenylalanine. In several embodiments, the second residue comprises tyrosine. In some embodiments, the catalytic residue comprises aspartic acid. In some embodiments, the isolated polypeptide is a homolog of SEQ ID NO: 1 or SEQ ID NO: 19, as determined by SWISS-MODEL homology modeling. In some embodiments, the isolated polypeptide comprises at least about 25%, 35%, 45%, 55%, 65%, 75%, 85%, 95%, 97%, or 100% identity to SEQ ID NO:1 or SEQ ID NO:19.

[0055] In some embodiments, an isolated polypeptide is provided comprising a sequence encoding any of the provided glycoside hydrolases. In some embodiments, the isolated polypeptide comprises a first motif that binds 2-keto-3-deoxygluconate and a second motif comprising a catalytic residue (or vice versa). In some embodiments, the first motif comprises at least two residues. The at least two residues can be arginine and one of tryptophan, phenylalanine, and tyrosine. In some embodiments, the second residue of the first motif can be tryptophan. In some embodiments, the second residue of the first motif can be phenylalanine. In some embodiments, the second residue of the first motif can be tyrosine. In some embodiments, the catalytic residue of the second motif comprises aspartic acid or glutamic acid. In some embodiments, the catalytic residue of the second motif is aspartic acid. In some embodiments, the isolated polypeptide is a homolog of SEQ ID NO: 1 or SEQ ID NO: 19, as determined by SWISS-MODEL homology modeling. In some embodiments, the isolated polypeptide comprises at least about 25%, 35%, 45%, 55%, 65%, 75%, 85%, 95%, 97%, or 100% identity to SEQ ID NO:1 or SEQ ID NO:19.

[0056] In embodiments, alanine scanning, compatible with other site-directed mutagenesis techniques, can be used to determine the contribution of specific residues to the stability and function of modified glycoside hydrolases, and the results of these analyses are reflected throughout this specification.

[0057] In embodiments, the glycoside hydrolase is selected from the group consisting of D88 (G / I / L / N / Q / A / R / V / W / Y), H132 (M / P / R / T / W / K / Y / D / C / G), H189 (F / I / L / N / Y / T / V / S), H352 (W / G / E / K / Q), H40 (I / K / L / P / R / W), H42 (C / K / Q / S), M147 (A / D / F / K / N / H / S), and / or M147 (M / P / R / T / W / K / Y / D / C / G). W141(A / D / H / K / M / P / R / T / Y / E / I / N / Q / S / V / C / F), W211(E / L / V / A / F / M / R / G / I / K / N / Q), and Y41(A / D / F / M / R / G / C / E / H / K / N / P / Q / V / I / T).

[0058] According to one embodiment, the isolated polypeptide comprises a first motif that binds to KDG and a second motif comprising a catalytic residue. In one embodiment, the first motif and the second motif are separated by about 70 residues. In one embodiment, the arginine of the first motif and the catalytic residue of the second motif are separated by 70 residues. In one embodiment, the arginine of the first motif and the catalytic residue of the second motif are separated by 72 residues.

[0059] In one embodiment, the first motif can be at least 5 residues, at least 10 residues, at least 15 residues, at least 20 residues, at least 30 residues, at least 50 residues, at least 100 residues, etc. In one embodiment, the first motif can be 5 residues and have the sequence RxQTW, where R is arginine, x is serine or an aliphatic amino acid such as glycine, alanine, valine, leucine, isoleucine, and proline, Q is glutamine, T is threonine, W is tryptophan, and R and W, independently and in combination, are substrate binding residues. In one embodiment, the first motif may be 10 residues and have the sequence Rx1QTW(2x2)Yx2Y, where R is arginine, x1 is serine, x2 is an aliphatic amino acid such as glycine, alanine, valine, leucine, isoleucine, and proline, Q is glutamine, T is threonine, W is tryptophan, Y is tyrosine, and R and W, independently and in combination, are substrate binding residues.

[0060] In one embodiment, the first motif can be at least 5 residues, at least 10 residues, at least 15 residues, at least 20 residues, at least 30 residues, at least 50 residues, at least 100 residues, etc. In one embodiment, the second motif can be 2 residues and have the sequence xD, where x is an aliphatic amino acid such as glycine, alanine, valine, leucine, isoleucine, and proline, D is aspartic acid, and D is a catalytic residue. In one embodiment, the second motif can be 18 residues and have the sequence (2x)KSE(3x)DT(2M)xSxPFx, where x is an aliphatic amino acid such as glycine, alanine, valine, leucine, isoleucine, and proline, K is lysine, S is serine, E is glutamic acid, D is aspartic acid, T is threonine, M is methionine, P is proline, F is phenylalanine, and D is a catalytic residue.

[0061] In one embodiment, the modified glycoside hydrolase has formula 1: P(19x)LPP(4x)HYHQGVxLxG(4x)(W / Y)(10x)Y(3x)(Y / W)x(D / E)(6x)G(9x)D(2x)Q(P / A)Gx(L / I)L(2x)L(7x)(R / K)Y(2x)(A / G)(3x) (L / I)(9x)(T / N)xE(G / Q)G(F / Y)(W / F)H(K / N)(3x)Px(Q / E)(M / Q)WLDGLYMxG(5x) Y(A / G)(9x)D(4x)Q(6x)(H / K)(T / M)(R / K)(3x)TGL(2x)H(A / G)(W / F)(D / S)(2x)( and (I / V)C(I / V)GT(S / G)xGxY(5x)R(5x)D(L / M)HG(V / A)GA(F / L), wherein Formula 1 has at least 50% homology to SEQ ID NO: 1, and wherein x is any amino acid. In one embodiment, the modified glycoside hydrolase has at least 55%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, or 100% homology to SEQ ID NO: 1. In one embodiment, the modified glycoside hydrolase comprises a Y at residue 41, a D at residue 8, an H at residue 132, a W at residue 141, a D at residue 143, an M at residue 147, an H at residue 189, a W at residue 211, a G at residue 212, an R at residue 213, a W at residue 217, an S at residue 278, an L or M at residue 282, a C at residue 330, and an H or K at residue 352, wherein these residues are numbered according to SEQ ID NO: 1. In one embodiment, the modified glycoside hydrolase comprises a loop region having at least 21 residues. In one embodiment, the modified glycoside hydrolase comprises a loop region having at least one modification selected from the group consisting of 331-336 of SEQ ID NO:1.In one embodiment, the modified glycoside hydrolase comprising at least one modification selected from the group consisting of 331-336 of SEQ ID NO:1 includes one or more of V331K, G332E, G332M, G332V, S334A, S334C, S334D, S334E, S334G, S334I, S334K, S334M, S334N, S334Q, S334R, S334T, S334V, A335V, A335P, A335L, A335C.

[0062] In one embodiment, the modified glycoside hydrolase is as follows: (F / Y)P(8x)(W / Y)(7x)W(T / M)(2x)F(2x)G(2x)(W / Y)(2x)Y(11x)(A / G)(10x)(L / I)(8x)(H / F)D(L / I)GF(4x)(S / T)(4x)(W / Y)(15x)(A / G)(13x)(L / I)(16x)IDx(L / M)(L / M)(N / S)(22x)H(3x)(T / S )(5x)RxDxS(S / T)(6x)(D / N)(10x)TxQG(4x)SxW(A / S / T)RG(Q / L)(A / T)W(2x)YG(28x)PxD(4x)(Y / W)D(F / L)(12x)S(6x)(S / C)(33x)Y(30x)(W / F / Y)(G / A)DY(Y / F)(2x)ExL, wherein Formula 2 has at least 50% homology to SEQ ID NO: 19, and x is any amino acid. In one embodiment, the modified glycoside hydrolase has at least 55%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, or 100% homology to SEQ ID NO: 19. In one embodiment, the modified glycoside hydrolase comprises an I or L at residue 26, an H at residue 41, a W at residue 42, an M at residue 43, an H at residue 87, a D at residue 88, a G at residue 134, a D at residue 149, a T at residue 150, an M at residue 152, a Q at residue 193, a W at residue 219, an R at residue 221, a W at residue 225, an S at residue 280, an I at residue 284, and an F at residue 352, wherein the residues are numbered according to SEQ ID NO: 19. In one embodiment, the modified glycoside hydrolase comprises a sequence selected from the group consisting of SEQ ID NOs: 24-27, or a sequence homologous thereto. An alignment of SEQ ID NOs: 24-27 is shown in Figure 6. *Residues marked with * correspond to catalytic and substrate binding residues as outlined above, which are also shown in Figure 7. Additionally, the relationship to other sequences with the same residues in similar positions relative to SEQ ID NO: 19 is shown.

[0063] In one embodiment, the modified glycoside hydrolase comprises a sequence of (W / Y)7x(A / G)(92x)WxD(35x)(L / M)(9x)(H / R)(22x)W(A / G / S)R(2x)(G / S)W(8x)(L / I)(27x)Q(3x)(G / K)xW(3x)(I / L)(9x)ExSx(S / T)(9x)(A / G)(52x)Gx(G / A), as in Formula 3, which has at least 50% homology to SEQ ID NO: 2. In one embodiment, the modified glycoside hydrolase has at least 55%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, 99%, or 100% homology to SEQ ID NO: 2.

[0064] Qualification methodology In some embodiments, glycoside hydrolase polypeptides and / or motifs thereof can be modified using one or more methodologies. In some embodiments, modifications are identified through rational design modeling. Methods for such engineering are generally known in the art. For example, amino acid sequence variants and fragments of KDG dehydrated polypeptides can be prepared by modifications in polynucleotide sequences. Methods for mutagenesis and polynucleotide modification are well known in the art. See, e.g., Kinkel (1985) Proc. Natl. Acad. Sci. USA 82:488-492; Kinkel et al. (1987) Methods in Enzymol. 154:367-382; U.S. Pat. No. 4,873,192; Walter and Gastra, eds. (1983) Techniques in Molecular Biology (MacMillan Publishing Company, New York) and references cited therein. Guidance regarding appropriate amino acid substitutions that do not affect the biological activity of the protein of interest can be found in the model Dayhoff et al. (1978) Atlas of Protein Sequence and Structure (Natl. Biomed. Res. Foun., Washington, DC), incorporated herein by reference in its entirety. Conservative substitutions, e.g., exchanging one amino acid for another with similar properties, are also contemplated. In some embodiments, modifications such as mutations may not position the sequence outside the reading frame. In some embodiments, modifications do not create complementary regions that could generate secondary mRNA structure. See EP Patent Application Publication No. 75,444.

[0065] In some embodiments, suitable modifications can be identified using well-known molecular biology techniques, such as polymerase chain reaction (PCR) and hybridization techniques, and sequencing techniques. Mutant polynucleotides also include synthetically derived polynucleotides, such as polynucleotides generated using site-directed mutagenesis or gene synthesis, but that still encode polypeptides capable of KDG dehydration, or polynucleotides generated by computational modeling.

[0066] In some embodiments, homology modeling can be performed on the designed and / or generated sequences using SWISS-MODEL homology modeling. In some embodiments, homology modeling is performed using the target glycoside hydrolase sequence as a template. In some embodiments, homology modeling includes (a) performing a template search for related homologs; (b) ranking the templates identified from (a) according to Global Model Quality Estimate (GMQE) and / or quaternary Structure Quality Estimate (QSQE); (c) determining whether the top templates cover different regions of the target protein (e.g., glycoside hydrolase) and / or whether the top templates represent different conformations; and (d) selecting templates. In some embodiments, the top templates can include 1 to 10, 1 to 30, 1 to 50, or 1 to 100 templates. In embodiments, the selected template may have about 1-5%, 1-10%, 1-20%, 1-30%, 10-30%, 20-40%, 20-60%, 25-65%, 30-70%, 50-80%, 60-95%, or 25-85% identity to the target glycoside hydrolase sequence. In some embodiments, the selected template may have at least about 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, or 100% identity to the target glycoside hydrolase sequence. In some embodiments, the selected template may comprise at least about 27% identity to the target glycoside hydrolase.

[0067] How to Make Dihydrofuran Also provided are biocatalytic processes for producing dihydrofuran from a composition comprising a substrate ketosugar. As used herein, "biocatalyst" or "biocatalytic" refers to the use of a natural catalyst, such as a protein enzyme, to perform a chemical transformation on an organic compound. Biocatalysis is also known as biotransformation or biosynthesis. The biocatalytic protein enzyme can be a naturally occurring protein or a recombinant protein. In some embodiments, also provided herein are methods for making dihydrofuran via dehydration of a ketosugar. In some embodiments, a glycoside hydrolase polypeptide can convert a ketosugar moiety to dihydrofuran. The provided methods can include any of the described glycoside hydrolases, modified glycoside hydrolases, and portions thereof, including, but not limited to, glycoside hydrolases comprising a sequence selected from SEQ ID NOs: 1-116.

[0068] The method can be biocatalytic, i.e., utilizes a biological catalyst. In some embodiments, the biocatalyst is a protein enzyme. In some embodiments, the biocatalyst is a glycoside hydrolase polypeptide. In some embodiments, the substrate is a keto sugar that is dehydrated by the glycoside hydrolase, thereby producing dihydrofuran.

[0069] The glycoside hydrolase may be provided in free or immobilized form. In some embodiments, the glycoside hydrolase preparation may be crude, semi-purified, and / or purified. In some embodiments, the glycoside hydrolase is provided as a whole cell system, e.g., viable or non-viable microbial cells, or whole microbial cells, cell lysates, and / or any other form known in the art.

[0070] In some embodiments, provided herein are methods for producing a dihydrofuran composition, the methods including: (a) providing a substrate keto sugar, e.g., KDG; (b) contacting the keto sugar with a glycoside hydrolase polypeptide; (c) producing a composition comprising dihydrofuran; and (d) dehydrating the dihydrofuran at elevated temperature.

[0071] In some embodiments, provided herein are methods for producing a dihydrofuran composition, the methods including: (a) providing a composition comprising greater than about 0.5%, about 1%, about 2%, about 3%, about 4%, 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 99.6% by weight of a substrate keto sugar on an anhydrous basis; (b) contacting the composition with a glycoside hydrolase polypeptide; (c) producing a composition comprising dihydrofuran; and (d) dehydrating the dihydrofuran at elevated temperature.

[0072] substrate In some embodiments, the methods provided herein include contacting a substrate, such as a keto sugar, with a glycoside hydrolase or motif thereof.

[0073] In some embodiments, the methods provided herein include contacting a substrate with a glycoside hydrolase or motif thereof. The substrate can include a keto sugar. In some embodiments, the substrate is the keto sugar 2-keto-3-deoxy-gluconic acid, of which 2-keto-3-deoxy-gluconic acid serves as a substrate for biotransformation by the glycoside hydrolase. The keto sugar can be synthetic, purified (partially or wholly), commercially available, or prepared. One example of a composition useful in the methods of the present disclosure is chemically synthesized 2-keto-3-deoxy-gluconic acid brought into solution with a solvent. Another example of a substrate is enzymatically synthesized 2-keto-3-deoxy-gluconic acid in water. Another example of a substrate is fermented 2-keto-3-deoxy-gluconic acid in broth.

[0074] In some embodiments, the composition comprises a purified substrate keto sugar. For example, the composition can comprise, on an anhydrous basis, greater than about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, 95%, about 96%, about 97%, about 98%, about 99%, or about 99.6% by weight of the substrate keto sugar.

[0075] In some embodiments, the composition comprises a partially purified substrate keto sugar, for example, the composition contains greater than about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 20%, about 30%, about 40%, or about 50% by weight of the substrate keto sugar on an anhydrous basis.

[0076] In some embodiments, the composition includes a substrate keto sugar (e.g., KDG). For example, the composition may include about 250 μM to about 2 M, about 500 μM to about 1.5 M, about 1 mM to about 1 M, about 25 mM to about 750 mM, about 250 mM to about 500 mM, and about 300 mM to about 400 mM of the substrate keto sugar. In one embodiment, the composition may include about 180 mM of the substrate keto sugar. In one embodiment, the composition may include about 750 mM of the substrate keto sugar. In one embodiment, the composition may include about 1.7 M of the substrate keto sugar.

[0077] In some embodiments, the composition comprises purified KDG. In some embodiments, the composition contains greater than about 99% by weight of KDG on an anhydrous basis. In some embodiments, the composition comprises partially purified KDG. In some embodiments, the composition contains greater than about 50%, about 60%, about 70%, about 80%, or about 90% by weight of KDG on an anhydrous basis.

[0078] In some embodiments, provided herein are methods for producing a dihydrofuran composition, the methods comprising: (a) providing a composition comprising a substrate keto sugar, such as KDG; (b) contacting the keto sugar with a glycoside hydrolase polypeptide; and (c) producing dihydrofuran. In some embodiments, the composition comprises an enzymatically produced keto sugar. In some embodiments, the glycoside hydrolase utilized in the methods herein is expressed in an engineered microorganism.

[0079] In some embodiments, provided herein are methods for producing dihydrofuran compositions, the methods comprising: (a) providing a composition comprising greater than about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 99.6% by weight of a substrate keto sugar on an anhydrous basis; (b) contacting the substrate-containing composition with a glycoside hydrolase polypeptide; and (c) producing dihydrofuran. In some embodiments, the composition comprises an enzymatically produced keto sugar. In some embodiments, the glycoside hydrolase utilized in the methods herein is expressed in an engineered microorganism.

[0080] In some embodiments, a composition comprising KDG is contacted with a glycoside hydrolase, thereby catalyzing the reaction of KDG (2-keto-3-deoxy-gluconic acid) to produce dihydrofuran. In some embodiments, the composition comprises partially purified KDG. In some embodiments, the composition comprises purified KDG. In some embodiments, the composition comprises at least about >95% KDG. In some embodiments, the composition comprises at least about 95%, 96%, 97%, 98%, 99%, or 100% KDG. In some embodiments, the composition comprises more than about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 99.6% KDG.

[0081] The present invention also provides a method for producing a dihydrofuran composition, the method comprising: (a) providing a composition comprising greater than about 0.5%, about 1%, about 2%, about 3%, about 4%, 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 99.6% by weight 2-keto-3-deoxy-gluconic acid (KDG) on an anhydrous basis; (b) contacting the composition with a glycoside hydrolase; and (c) producing a composition comprising dihydrofuran. In some embodiments, the composition comprises greater than about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 99.6% by weight of 4,5-dihydro-4-hydroxy-5-hydroxymethyl-2-furancarboxylic acid (K4) on an anhydrous basis.

[0082] In some embodiments, K4 can be further converted to 5-hydroxymethyl-2-furoic acid (HMFA), as described in WO2021016220. In some embodiments, HMFA can be further oxidized to FDCA using chemical oxidation as described in WO2021016220, or using enzymatic oxidation as described, for example, in Dijkman et al. 2014, "Discovery and Characterization of a 5-Hydroxymethylfurfuran Oxidase from Meticovorus sp. strain MP688," incorporated herein by reference.

[0083] [Table 2]

[0084] In some embodiments, the method produces a composition comprising greater than about 80% HMFA by weight. In some embodiments, purification produces a composition comprising greater than about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 99.6% HMFA by weight. In some embodiments, the composition comprises greater than about 95% HMFA by weight.

[0085] In some embodiments, the step of contacting the composition with the glycoside hydrolase polypeptide and the keto sugar is carried out at a temperature of about 0.5° C. to about 75° C. In some embodiments, the temperature is from about 0.5° C., about 5° C., about 10° C., about 15° C., about 20° C., about 25° C., about 30° C., about 35° C., about 40° C., about 45° C., about 50° C., about 55° C., about 60° C., about 65° C., or about 70° C., up to about 75° C. In some embodiments, the temperature is from about −0.5° C. to 5° C., 10° C. to 15° C., 20° C. to 30° C., 35° C. to 45° C., 35° C. to 50° C., 35° C. to 55° C., 30° C. to 60° C., or 30° C., up to about 75° C. In some embodiments, the step of contacting the composition with the glycoside hydrolase polypeptide and the keto sugar occurs at a temperature of about 45°C.

[0086] In some embodiments, the step of contacting the composition with the glycoside hydrolase polypeptide and the keto sugar is carried out at a pH of about 3 to about 8. In some embodiments, the pH is about 0.5 to about 1, about 1.5, about 2, about 2.5, about 3, about 3.5, about 4, about 4.5, about 5, about 5.5, about 6, about 6.5, about 7, about 7.5, and about 8. In some embodiments, the pH is about 3 to about 6. In some embodiments, the pH is about 3 to about 5. In some embodiments, the pH is about 3 to about 4. In some embodiments, the pH is about 3 to about 3.5. In some embodiments, a decrease in pH results in a reduced need for NaOH in the contacting process. In some embodiments, the step of contacting the composition with the glycoside hydrolase and the keto sugar is carried out at about pH 5.

[0087] In some embodiments, the step of contacting the composition with the glycoside hydrolase polypeptide and the ketosugar is carried out for a period of 1 hour to 14 days. In some embodiments, the contacting is for a period of about 1 hour, about 3 hours, about 5 hours, about 7 hours, about 9 hours, about 11 hours, about 13 hours, about 15 hours, about 17 hours, about 19 hours, about 21 hours, about 23 hours, about 25 hours, about 27 hours, about 29 hours, about 31 hours, about 33 hours, about 35 hours, about 37 hours, about 39 hours, about 41 hours, about 43 hours, about 45 hours, about 48 hours, about 49 hours, about 51 hours, about 53 hours, about 55 hours, about 57 hours, about 59 hours, about 61 hours, about 63 hours, about 65 hours, about 67 hours, about 69 hours, about 71 hours, or about 73 hours, up to about 75 hours. In some embodiments, the step of contacting the composition with the glycoside hydrolase is carried out for a period of about 48 hours. In some embodiments, the contacting is for a period of about 1 day to about 14 days. In some embodiments, the contacting is for a period of about 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, or about 13 days, up to about 15 days. In some embodiments, the contacting is for a period of about 1 day to about 14 days. In some embodiments, the step of contacting the composition with the glycoside hydrolase is carried out for a period of about 6 hours.

[0088] In some embodiments, the resulting composition comprising K4 comprises greater than about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 99.6% by weight of K4 on an anhydrous basis.

[0089] In some embodiments, at least about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the substrate KDG in the composition is converted to K4.

[0090] The reaction medium for the conversion can be aqueous. In some embodiments, the reaction medium can be purified water, a buffer solution, or a combination thereof. In some embodiments, the reaction medium is a buffer solution. Suitable buffer solutions include, but are not limited to, acetate buffer, citrate buffer, phosphate buffer, and Bis-Tris buffer. In some embodiments, the reaction medium is acetate buffer. In some embodiments, the reaction medium is phosphate buffer. In some embodiments, the reaction medium is Bis-Tris buffer. Alternatively, the reaction medium can be an organic solvent. In some embodiments, the reaction medium is supplemented with glycerol, Tween-20, sucrose, or sorbitol as an enzyme stabilizer. In some embodiments, the reaction medium is supplemented with 20% glycerol, 0.1% Tween, 2 M sucrose, or 2 M sorbitol.

[0091] In some embodiments, the conversion of KDG to dihydrofuran is at least about 2% complete, as determined by any of the methods described above. In some embodiments, the conversion of KDG to dihydrofuran is at least about 10% complete, at least about 20% complete, at least about 30% complete, at least about 40% complete, at least about 50% complete, at least about 60% complete, at least about 70% complete, at least about 80% complete, at least about 90% complete, at least about 95% complete, or at least about 100% complete. In some embodiments, the conversion of KDG to dihydrofuran is greater than about 80% complete. In some embodiments, at least about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the KDG in the composition is converted to dihydrofuran.

[0092] In some embodiments, the reaction may be monitored by means including, but not limited to, HPLC, LCMS, TLC, IR, UV, or NMR. In several embodiments, the reaction is monitored using LCMS, UV, or both LCMS and UV.

[0093] In some embodiments, contacting the composition with the glycoside hydrolase and / or glycoside hydrolase polypeptide and KDG can be carried out for a period of 1 hour to 14 days, e.g., about 1 hour, about 6 hours, about 12 hours, about 24 hours, about 48 hours, about 72 hours, about 120 hours, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 8 days, about 9 days, or about 10 days. In some embodiments, the reaction is carried out for about 6 days. In some embodiments, the reaction is carried out for about 7 days.

[0094] Conversion of dihydrofuran to HMFA and FDCA In some embodiments, the dihydrofuran undergoes further processing, such as purification and / or dehydration, to produce HMFA and / or FDCA, e.g., as depicted in Figure 1, where HMFA is shown as 3. In some embodiments, the dihydrofuran is chemically converted to HMFA using acidic conditions. In some embodiments, the HMFA is purified. In some embodiments, the HMFA is purified before proceeding with subsequent reactions.

[0095] In some embodiments, provided herein are methods for producing HMFA, the methods including: (a) providing a composition comprising greater than about 0.5%, about 1%, about 2%, about 3%, about 4%, 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 99.6% by weight of a substrate keto sugar on an anhydrous basis; (b) contacting the composition with a glycoside hydrolase polypeptide; (c) producing a composition comprising dihydrofuran; and (d) dehydrating the dihydrofuran to HMFA under acidic conditions. In embodiments, a method is provided that includes contacting a composition comprising HMFA with a glycoside hydrolase polypeptide, thereby producing dihydrofuran, which in embodiments is dehydrated to HMFA under acidic conditions.

[0096] In some embodiments, dihydrofuran is converted to HMFA at elevated temperatures. In some embodiments, the step of contacting the composition with the glycoside hydrolase polypeptide and KDG can be carried out at a temperature of about 0.5°C to about 110°C, e.g., about 10°C, about 20°C, about 30°C, about 40°C, about 50°C, about 60°C, about 70°C, about 80°C, about 90°C, about 100°C, or about 110°C. In a preferred embodiment, the reaction is carried out at about 74°C. In some embodiments, the reaction is carried out at about 69°C. In another embodiment, the reaction is carried out at about 63°C.

[0097] In some embodiments, dihydrofuran is contacted with an acid to convert it to HMFA. The acid can be selected from inorganic acids such as hydrochloric acid, sulfuric acid, phosphoric acid, nitric acid, and hydrobromic acid. The acid can be selected from organic acids such as C1-6 carboxylic acids. In some embodiments, the contacting comprises dehydrating dihydrofuran with an acid selected from the group consisting of formic acid, hydrochloric acid, sulfuric acid, phosphoric acid, nitric acid, hydrobromic acid, and a C1-6 carboxylic acid. In some embodiments, the acid comprises formic acid.

[0098] In some embodiments, dihydrofuran is converted to HMFA at a low pH. In some embodiments, the pH can be acidic. In some embodiments, the pH is neutral or basic. In some embodiments, the pH is acidic and is between 0 and 6. The pH can be 0, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, or up to about 10. In some embodiments, the pH is 4. In some embodiments, the pH is 4.5. In some embodiments, the pH is 5. In some embodiments, the pH is selected from the group consisting of 4, 4.5, and 5.

[0099] In some embodiments, the pH is about 4 and the temperature is about 63° C. In some embodiments, the pH is about 4.5 and the temperature is about 69° C. In some embodiments, the pH is about 5 and the temperature is about 72° C.

[0100] In some embodiments, conversion of dihydrofuran to HMFA (e.g., dehydration) results in about 5% to 100% HMFA. In some embodiments, about 5-10%, 10-30%, 25-40%, 30-50%, 35-60%, 40-70%, 45-85%, 50-80%, 50-90%, 55-90%, or 60-100% conversion to HMFA. In some embodiments, conversion of dihydrofuran to HMFA results in at least about or up to about 5%, 15%, 25%, 35%, 45%, 55%, 65%, 75%, 85%, 95%, or 100% HMFA.

[0101] In some embodiments, HMFA can be further chemically and / or biocatalytically oxidized to FDCA. In some embodiments, HMFA is purified before oxidation. In some embodiments, the purification comprises increasing the HMFA in the composition. In some embodiments, the purification comprises increasing the HMFA in the composition by at least about 1-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 100-fold, 150-fold, 200-fold, 300-fold, or 500-fold compared to an otherwise equivalent composition lacking purification.

[0102] In some embodiments, the resulting FDCA can be used as a building block for polymers, which can reduce carbon emissions in the plastics industry compared to PET.

[0103] modified microorganisms In some embodiments, any of the described glycoside hydrolases or motifs thereof can be produced in a host, such as a microorganism. A modified microorganism may refer to a host cell that has been genetically altered by the cloning and transformation methods of the present disclosure. Thus, the term includes host cells (such as bacteria, yeast cells, fungal cells, CHO cells, human cells, etc.) that have been genetically altered, modified, or engineered to exhibit an altered, modified, or different genotype and / or phenotype (e.g., when the genetic alteration affects the encoding nucleic acid sequence of the microorganism) compared to the naturally occurring organism from which it was derived.

[0104] In some cases, the microorganism is engineered to produce a glycoside hydrolase. In some embodiments, the engineered microorganism contains and / or expresses any of the glycoside hydrolases provided herein. The engineered microorganism can also contain a polynucleotide encoding any of the glycoside hydrolases provided herein. Accordingly, also provided herein are engineered microorganisms that contain and / or express glycoside hydrolases, as well as methods for making the same.

[0105] In some embodiments, a DNA sequence encoding a glycoside hydrolase or its motif is cloned into an expression vector and inserted into a production host, such as a microorganism, e.g., a bacterium. The protein can be isolated from cell extracts based on its physical and chemical properties using techniques known in the art. In some embodiments, the sequences of the present disclosure can be introduced into host cells using any of a variety of techniques, including transformation, transfection, transduction, viral infection, gene guns, or Ti-mediated gene transfer (see Christie PJ, and Gordon JE, 2014 "The Agrobacterium Ti Plasmids," Microbiol SPectr. 2014;2(6);10.1128). Specific methods include calcium phosphate transfection, DEAE-dextran-mediated transfection, lipofection, or electroporation (Davis, L., Dibner, M., Battey, I., 1986 "Basic Methods in Molecular Biology"). Other transformation methods include, for example, lithium acetate transformation and electroporation. For example, Gietz et al., Nucleic Acids Res. 27:69-74 (1992); I T et al., J. Bacterol. 153:163-168 (1983); and Becker and Guarente, Methods in Enzymology 194:182-187 (1991). Further, exemplary, non-limiting techniques for isolating glycosyl hydrolases from engineered microorganisms include centrifugation, electrophoresis, liquid chromatography, ion exchange chromatography, gel filtration chromatography, and / or affinity chromatography.

[0106] In some embodiments, the microorganism is modified to comprise and / or express an amino acid sequence comprising 80% to 100% sequence identity to any one of SEQ ID NOs: 1-116. In some embodiments, the microorganism is modified to comprise and / or express an amino acid sequence comprising at least 10%, or at least 99.73%, identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-116.

[0107] In some embodiments, the engineered microorganism of the present disclosure comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1. In some embodiments, the engineered microorganism comprises a glycoside hydrolase polypeptide comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-116. In some embodiments, the engineered microorganism comprises a glycoside hydrolase polypeptide of SEQ ID NO: 19.

[0108] In some embodiments, the glycoside hydrolase polypeptide contained in or expressed by the engineered microorganism further comprises a tag amino acid sequence, hi some embodiments, the tag amino acid sequence is His6.

[0109] In some embodiments, the modified microorganism is a bacterium, including at least 11 different groups of bacteria: (1) Gram-positive (Gram+) bacteria, with two major classifications: (1) high G+C groups (Actinomycetes, Mycobacteria, Micrococcus, etc.) (2) low G+C groups (Bacillus, Clostridia, Lactobacillus, Staphylococci, Streptococci, Mycoplasmas), (2) Proteobacteria, e.g., purple photosynthetic + non-photosynthetic Gram-negative bacteria (including the most "common" Gram-negative bacteria), (3) Cyanobacteria, e.g., oxygenic photosynthetic organisms, (4) Spirochetes and related species, (5) Plantomyces, (6) Bacteroides, Flavobacteria, (7) Chlamydia, (8) green sulfur bacteria, (9) green non-sulfur bacteria (also anaerobic photosynthetic organisms), (10) radioresistant microorganisms and related species, (11) Thermotoga and Thermosipho thermophiles. The bacterium can be any one of the genera E. coli, Saccharomyces, Aspergillus, Pichia, Pseudomonas, or Bacillus. In some embodiments, the bacterium is E. coli. [Example]

[0110] The following examples are provided for the purpose of illustrating various embodiments of the present disclosure and are not meant to limit the disclosure in any way. Those skilled in the art will recognize modifications of the examples and other uses that are encompassed within the spirit of the disclosure, as defined by the scope of the claims.

[0111] Example 1: In vivo production of glycoside hydrolases Polynucleotides encoding amino acids SEQ ID NO:1 to SEQ ID NO:116 were synthesized (TMT Bioscience) and inserted into the pARZ4 expression vector. The recombinant vector was used to transform E. coli NEBT7EL (New England Biolabs) using the heat shock method, thereby preparing recombinant microorganisms.

[0112] The transformed microorganisms were inoculated into 1 ml of TB-kanamycin medium and grown overnight at 37°C with shaking. 100 μL of the culture was inoculated into 5 ml of TB-kanamycin medium and grown at 37°C for 2 hours, followed by 25°C for 1 hour at 400 RPM. The culture was induced with 1 mM IPTG and expression was continued at 25°C and 400 RPM for 20-24 hours. Finally, the culture was harvested by centrifugation at 2,200 × g for 10 minutes. The supernatant was discarded, and the pellet was stored at -20°C.

[0113] Example 2: Purification of glycoside hydrolases The microorganisms prepared in Example 1 were thawed from storage at -20°C. Once thawed, the pellet was resuspended in lysis buffer (2 mg / mL lysozyme, 0.1 mg / mL DNAse I, 5% BugBuster® Protein Extraction Reagent, 20 mM PO4 pH 7.5, 500 mM NaCl, and 20 mM imidazole). The resuspended cells were disrupted by incubating at 5°C with shaking at 200 RPM for 30 minutes. The disrupted lysate was centrifuged at 2,200 x g for 7 minutes. The resulting supernatant was loaded onto a Ni-NTA plate equilibrated with binding buffer. The plate was centrifuged at 100 x g for 4 minutes, followed by two washes with 500 uL of binding buffer (20 mM PO4 pH 7.5, 500 mM NaCl, 20 mM imidazole) and centrifugation (500 x g) for 2 minutes. The protein was eluted with 150 μL of elution buffer (20 mM PO4 pH 7.5, 500 mM NaCl, 500 mM imidazole) followed by centrifugation at 500 × g for 2 minutes. The recovered protein was desalted into a buffer for enzyme activity assessment (50 mM acetate pH 5, 150 mM NaCl).

[0114] Example 3: Measurement of glycoside hydrolase activity using keto sugar substrates by UV quantification The purified enzyme from Example 2 was incubated in 50 mM acetate buffer (pH 5, 150 mM NaCl, 180 mM KDG (2-keto-3-deoxy-gluconate)) at 74°C for 3 hours. The reaction mixture was then filtered through a 10K MWCO filter plate to remove protein from the product, substrate, and other reaction components. Product formation was detected using a Biotek™ Synergy HTX Multi-Mode Microplate Reader. Product was measured at a wavelength of 250 nm in a UV-transparent plate. Quantification of product was achieved by calculation based on a KDG standard curve generated under the same reaction conditions.

[0115] Example 4: Measurement of glycoside hydrolase activity using keto sugar substrates via triple quadrupole (QQQ) quantitation The purified enzyme from Example 2 was incubated in 50 mM acetate buffer (pH 5, 150 mM NaCl, 180 mM KDG (2-keto-3-deoxy-gluconate)) at 74°C for 3 hours. The reaction mixture was then filtered through a 10K MWCO filter plate to remove protein from the product, substrate, and other reaction components. Product formation was detected using an Agilent G6470A triple quadrupole mass spectrometer. The LC component consisted of an Agilent 1290 Multisampler, 1290 Infinity II pump, 1290 Flex pump, and a 1290 MTC alternating at 10°C with a Waters 100 x 2.1 mm HSS T3 1.8 μm C-18 column at 40°C. 50 μL wells were selected for calibration for the QQQ by adding 5 μL of 2 mM, 1 mM, 0.5 mM, and 0.25 mM HMFA solutions to the received 384-well assay plate (resulting in final concentrations of 200, 100, 50, and 25 μM). Samples were delivered alternately, with one column performing the separation and analysis while the second column re-equilibrated with the starting mobile phase. This resulted in a runtime of 2.5 minutes per sample. Analysis was performed using multiple reaction monitoring (MRM) on the QQQ using the following transition settings for 13C-labeled HMFA (m / z 147) and HMFA (m / z 141) in Figures 3 and 4:

[0116] Example 5: Use of rational design approaches to gain or improve KDG dehydratase activity Design of natural enzymes with low to high levels of KDG dehydration activity. To improve the activity of the parent scaffolds that showed initial KDG dehydration activity, computational enzyme design techniques were used to improve the substrate interactions in the active site(s) of SEQ ID NO: 1 and SEQ ID NO: 19, as summarized in Table 3. For these sequences, information on the crystal structure of the native protein provided an accurate picture of how the substrate or transition state fit within the active site(s). Loops were identified in SEQ ID NO: 1 and SEQ ID NO: 19 that were amenable to flexible scaffold design. In total, computational design was used to improve enzymatic efficiency in five enzyme scaffolds that introduced 1 to 12 mutations into the parent sequence.

[0117] [Table 3]

[0118] Site-saturation mutagenesis of glycoside hydrolases SEQ ID NOs: 1 and 19. To discover amino acid positions in SEQ ID NOs: 1 and 19 where point mutations increase KDG dehydration activity, saturation mutagenesis was performed around the active site at positions 40, 41, 42, 88, 132, 133, 141, 147, 189, 211, 331, 332, 333, 334, 335, 336, and 352 in SEQ ID NO: 1 and around the active site at positions 41, 42, 87, 88, 91, 134, 149, 152, 193, 211, 219, 221, 225, 337, 338, 339, 352, and 356 in SEQ ID NO: 19. A total of 323 and 342 protein variants (19 point variants per amino acid position) were tested for KDG dehydration activity for SEQ ID NOs: 1 and 19, respectively.

[0119] Among the SEQ ID NO:1 mutants, 11 point mutations at six amino acid positions in SEQ ID NO:1 increased KDG dehydration activity by up to two-fold (SEQ ID NOs:35, 37, 41, 42, 43, 48, 55, 62, 63, 64, and 68). The top 36 point mutations are: H42I, H42T, H42V, H42W, D88C, D88N, H132E, K133M, W141Y, H189A, H189V, V331K, G332E, G332M, G332V, S334A, S334C, S334D, S334E, S334G, S334I, S334K, S334M, S334N, S334Q, S334R, S334T, S334V, A335V, A335P, A335L, A335C, H352R, and H352V. Flexible positions were discovered where multiple neutral or beneficial amino acid changes were found. For example, 20 amino acids were found at positions 332, 334, and 335 of SEQ ID NO: 1. Positions within the loop region, from amino acid positions 331 to 336 of SEQ ID NO: 1, are more variable.

[0120] Among the variants of SEQ ID NO: 19, eight point mutations at three amino acid positions increased KDG dehydration activity by up to 1.4-fold (SEQ ID NOs: 95, 96, 98, 99, 101, 102, 103, and 105). Point mutations that resulted in increased activity included D41A, D41C, D41E, D41N, D41S, D41T, H87A, H87C, H87E, H87G, H87Q, H87R, H87S, L152M, L152N, W225Y, S337A, Y338V, H339A, H339N, W352F, Y356A, Y356C, Y356F, and Y356H.

[0121] Homology modeling Homology modeling was performed for SEQ ID NOs: 4-18 and 22-34 using the online SWISS-MODEL homology modeling server. To perform modeling, the target sequence was uploaded in fasta format. Template searches were performed using BLAST and HHblits. The retrieved templates were ranked according to the Global Model Quality Estimate (GMQE) and Quaternary Structure Quality Estimate (QSQE). The top templates and alignments were compared to verify whether they covered different regions of the target protein and represented alternative conformations. Multiple templates were automatically selected, and different models were constructed accordingly. A template (e.g., the first template) was selected and used to construct the model. The minimum percentage of sequence identity used for templates was 27%.

[0122] All references, articles, publications, patents, patent publications, and patent applications cited herein are incorporated by reference in their entirety for all purposes. However, mention of any reference, article, publication, patent, patent publication, or patent application cited herein is not, and should not be considered as, an admission or any indication that it constitutes valid prior art or forms part of the common general knowledge in any country in the world.

Claims

1. A biocatalytic method for producing dihydrofuran, The process involves contacting 2-keto-3-deoxygluconic acid (KDG) with glycoside hydrolase, thereby generating the dihydrofuran, and the contact is a. The pH measured with a pH meter is approximately 3 to approximately 7, b. A temperature of 45°C to 74°C, or c. The method comprising both a. and b., thereby producing the dihydrofuran.

2. a. The biocatalyst method according to claim 1, comprising and wherein the pH is approximately 4 to 5.

3. b. The biocatalyst method according to claim 1, comprising, wherein the temperature is 70°C to 74°C.

4. The biocatalytic method according to claim 1, comprising a. and b., wherein the pH is approximately 4 to 5 and the temperature is approximately 62°C to 72°C.

5. a. and b. are included, and the pH and temperature are, a. The pH is approximately 4 and the temperature is approximately 63°C. b. The pH is approximately 4.5 and the temperature is approximately 69°C. c. The biocatalytic method according to claim 1, wherein the pH is approximately 5 and the temperature is approximately 72°C, selected from the group.

6. c. The biocatalytic method according to claim 5, comprising

7. The biocatalytic method according to claim 1, wherein the KDG is 100 mM to 2 M.

8. The biocatalytic method according to claim 7, wherein the KDG is 100 mM to 750 mM.

9. The biocatalytic method according to any one of claims 1 to 8, wherein the glycoside hydrolase comprises a protein having at least 80% sequence identity with respect to a sequence selected from the group consisting of SEQ ID NOs: 1 to 116.

10. The biocatalytic method according to claim 9, wherein the sequence identity is at least 85%, 90%, 95%, 98%, 99%, or 100%.

11. The glycoside hydrolase described above A first motif coupled to the aforementioned KDG, A second motif comprising a catalytic residue, wherein the catalytic residue comprises aspartic acid, The first motif comprises at least two residues, the first residue comprising arginine, and the second residue comprising tryptophan, phenylalanine, or tyrosine. The biocatalytic method according to any one of claims 1 to 8, wherein the glycoside hydrolase is determined by SWISS-MODEL modeling to be a homolog of SEQ ID NO: 1 or SEQ ID NO:

19.

12. The biocatalytic method according to any one of claims 1 to 8, wherein the glycoside hydrolase comprises a protein having 100% sequence identity with SEQ ID NO: 27 or SEQ ID NO:

35.

13. The biocatalyst method according to any one of claims 1 to 8, wherein the contact is 72 hours to 14 days.

14. The biocatalyst method according to any one of claims 1 to 8, further comprising dehydrating the dihydrofuran to produce 5-hydroxymethyl-2-furoic acid (HMFA), wherein the yield of the HMFA after dehydration is observed to be at least 40%.

15. The biocatalyst method according to claim 14, wherein the dehydration comprises contacting the dihydrofuran with an acid selected from the group consisting of formic acid, hydrochloric acid, sulfuric acid, phosphoric acid, nitric acid, hydrobromic acid, and Ci-6 carboxylic acid.

16. The biocatalyst method according to claim 14, further comprising oxidizing the HMFA to produce 2,5-franzicarboxylic acid (FDCA).

17. Isolated non-natural glycoside hydrolase, A first motif that binds to 2-keto-3-deoxygluconate, comprising at least two residues, wherein the first residue is arginine and the second residue is tryptophan, phenylalanine, or tyrosine, and A second motif comprising a catalytic residue, wherein the catalytic residue is aspartic acid or glutamic acid, The isolated non-natural glycoside hydrolase having at least 20% identity with SEQ ID NO: 1 or SEQ ID NO:

19.

18. The isolated non-natural glycoside hydrolase according to claim 17, wherein the arginine of the first motif and the catalytic residue of the second motif are separated by approximately 70 residues.

19. A biocatalytic method for producing dihydrofuran, The process involves contacting 2-keto-3-deoxygluconic acid (KDG) with glycoside hydrolase, thereby generating the dihydrofuran, and the contact is a. The pH measured with a pH meter is approximately 3 to approximately 7. b. A temperature of 45°C to 74°C, or c. Includes both a. and b., thereby generating the dihydrofuran, The glycoside hydrolase has at least 50% homology to SEQ ID NO: 1, The method wherein the glycoside hydrolase contains, compared to SEQ ID NO: 1, residue 143 (D), residue 213 (R), and residue 217 (W).

20. The biocatalytic method according to claim 19, wherein residue 143 of the glycoside hydrolase is a catalytic residue, and residues 213 and 217 are substrate-binding residues.