Polypeptides for biosynthesis of 1, 2-propanediol and uses thereof
By expressing and purifying polypeptides with 1,5-syringol-6-phosphate isomerase activity, these enzymes are used to catalyze 1,5-syringol-6-phosphate to produce 1,2-propylene glycol, solving the complex and high cost of production of 1,2-propylene glycol in the prior art, and achieving efficient and environmentally friendly biocatalytic synthesis.
Patent Information
- Application Number
- CN202311617289.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
The existing 1,2-propylene glycol production technology is complex, consumes a lot of raw materials and energy, has high production costs, and has many by-products, which affects environmental protection.
Biocatalytic synthesis is achieved by expressing and purifying polypeptides with 1,5-syringol-6-phosphate isomerase activity, using these enzymes to catalyze 1,5-syringol-6-phosphate to produce 1,2-propylene glycol.
It reduces the energy consumption and raw material demand for 1,2-propylene glycol production, reduces the generation of by-products, and improves production efficiency and environmental protection performance.
Smart Images

Figure CN120060228A_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to the fields of genetic engineering, enzyme engineering, and bioinformatics; specifically, the present application provides new sugar degradation pathways, 1,5-anhydroglucitol-6-phosphate isomerases involved in these pathways, and their uses, particularly in the synthesis of 1,2-propanediol. Background Art
[0002] 1,2-Propanediol is mainly used in the production of coatings and unsaturated polyester resins, and can also be used in the production of plasticizers and hydraulic brake fluids. In addition, 1,2-propanediol can be used as a good solvent for the preparation of inks and epoxy resins; it can be used in non-ionic detergents to act as an enzyme stabilizer and solvent; and it can be used as an antifreeze and as a humectant in the pharmaceutical, cosmetic, animal food, and tobacco industries.
[0003] Common 1,2-propanediol production technologies include direct and indirect hydration of propylene oxide, direct catalytic oxidation of propylene, and co-production of dimethyl carbonate (DMC) / propylene glycol. However, these methods often require complex chemical devices, consume large amounts of raw materials and energy, produce by-products, etc., resulting in high production costs and poor atom economy.
[0004] In view of the fact that 1,2-propanediol is an important industrial raw material with important values and uses, the research on its synthesis pathway has important scientific research and industrial application values. Summary of the Invention
[0006] In a first aspect, the present application provides a polypeptide comprising the amino acid sequence shown in SEQ ID NO: 1 or 97 or a functional variant thereof, wherein the functional variant has 1,5-anhydroglucitol-6-phosphate isomerase activity.
[0007] In some embodiments of the first aspect, the polypeptide comprises the amino acid sequence shown in SEQ ID NO: 1 or a functional variant thereof, wherein the functional variant has 1,5-anhydroglucitol-6-phosphate isomerase activity. In some embodiments, the polypeptide has an active site defined in the following spatial conformation: the active site comprises amino acid residues H165, H282, S283, H334, C441, E443, R453, T455, L562, S662, I664, and G786 that are spatially close to each other with reference to SEQ ID NO: 1.
[0008] In some embodiments of the first aspect, the functional variant is a natural isoenzyme of the amino acid sequence shown in SEQ ID NO: 1.
[0009] In some embodiments of the first aspect, the polypeptide comprises the amino acid sequence shown in SEQ ID NO: 97 or a functional variant thereof, wherein the functional variant has 1,5-anhydro-D-mannitol-6-phosphate isomerase activity. In some embodiments, the polypeptide has an active site defined in terms of its spatial conformation: the active site comprises amino acid residues Q162, H169, S277, S278, R323, F331, P335, C431, E433, D445, Y628, V630, and G752 of SEQ ID NO: 97 that are close to each other in spatial conformation.
[0010] In some embodiments of the first aspect, the functional variant is a natural isoenzyme of the amino acid sequence shown in SEQ ID NO: 97.
[0011] In some embodiments of the first aspect, the substrate of the polypeptide is 1,5-anhydro-D-sorbitol-6-phosphate.
[0012] In some embodiments of the first aspect, the functional variant is generated by insertion, substitution, and / or deletion of one or more amino acids based on the amino acid sequence shown in SEQ ID NO: 1 or 97 or its natural isoenzyme.
[0013] In a second aspect, the present application provides a nucleic acid molecule encoding the polypeptide of the first aspect.
[0014] In a third aspect, the present application provides an expression cassette comprising the nucleic acid molecule of the second aspect.
[0015] In a fourth aspect, the present application provides an expression vector comprising the nucleic acid molecule of the second aspect or the expression cassette of the third aspect.
[0016] In a fifth aspect, the present application provides a host cell comprising the nucleic acid molecule of the second aspect, the expression cassette of the third aspect, or the expression vector of the fourth aspect.
[0017] In some embodiments of the fifth aspect, the host cell is capable of expressing and producing: a polypeptide comprising the amino acid sequence shown in SEQ ID NO: 1 or 97 or a functional variant thereof, wherein the functional variant has 1,5-anhydro-D-sorbitol-6-phosphate isomerase activity.
[0018] In some embodiments of the fifth aspect, the host cell further expresses: an activating enzyme, an aldolase, a hydroxyacetone reductase, and / or a transport complex.
[0019] In some embodiments of the fifth aspect, the transport complex has phosphoenolpyruvate - phosphotransferase system transport activity.
[0020] In some embodiments of the fifth aspect, the host cell is a eukaryotic cell or a prokaryotic cell.
[0021] In some embodiments of the fifth aspect, the eukaryotic cell is a yeast cell.
[0022] In some embodiments of the fifth aspect, the prokaryotic cells are selected from the genera Escherichia, Klebsiella, Streptococcus, Lactobacillus, Bifidobacterium, Bacteroidetes, and Firmicutes, etc.
[0023] In a sixth aspect, the present application provides a composition comprising the polypeptide described in the first aspect.
[0024] In some embodiments of the sixth aspect, the composition is used for catalyzing the formation of 1,2-propanediol from 1,5-anhydroglucitol-6-phosphate.
[0025] In a seventh aspect, the present application provides the use of the polypeptide described in the first aspect, the nucleic acid molecule described in the second aspect, the expression cassette described in the third aspect, the expression vector described in the fourth aspect, and the host cell described in the fifth aspect in the preparation of a composition for catalyzing the formation of 1,2-propanediol from 1,5-anhydroglucitol-6-phosphate.
[0026] In an eighth aspect, the present application provides a method for producing 1,2-propanediol, which comprises culturing the host cell described in the fifth aspect in a medium containing 1,5-anhydroglucitol to obtain 1,2-propanediol.
[0027] In a ninth aspect, the present application provides a method for producing 1,2-propanediol, which comprises contacting the composition described in the sixth aspect with 1,5-anhydroglucitol-6-phosphate to obtain 1,2-propanediol. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 Schematic diagram of radical-dependent anhydroglycolysis. A. Gene cluster in Lactiplantibacillus plantarum (including the gene encoding 1,5-anhydroglucitol-6-phosphate isomerase shown in SEQ ID NO: 2) and YbiW-dependent anhydroglycolysis pathway. B. Gene cluster in Escherichia coli (including the gene encoding 1,5-anhydromannitol-6-phosphate isomerase shown in SEQ ID NO: 98) and PflD-dependent anhydroglycolysis pathway.
[0029] Figure 2Electron paramagnetic resonance (EPR) spectroscopy and LC-MS enzymatic reaction analysis of YbiW. A. EPR spectra of LpYbiW with LpYbiY and titanium(III) citrate in the presence or absence of S-adenosylmethionine (SAM). B. Detection of the formation of 1-deoxyfructose-6-phosphate (1-deoxy-F6P) in the isomerization of 1,5-anhydroglucitol-6-phosphate (1,5-AG-6P) catalyzed by LpYbiW by LC-MS. In the negative ion mode, the extracted ion chromatogram (EIC) at m / z = 243 was used to monitor the formation of 1-deoxy-F6P (t R = 18.40 min). C-D. Mass spectra (negative ion mode) of compound 1 (1-deoxy-F6P) and compound 2 (1,5-AG-6P) corresponding to the EIC peaks in B. E. HPLC elution profiles of the DNPH derivatives of the reaction products and standards in the reaction analysis of LpFsaA-coupled LpYbiW. F-H. ESI(-) m / z mass spectra of DNPH-glyceraldehyde 3-phosphate (peak 3) and DNPH-hydroxyacetone (peaks 4 and 5) in E. Among them, w / o means without.
[0030] Figure 3 Electron paramagnetic resonance (EPR) spectroscopy and LC-MS enzymatic reaction analysis of PflD. A. EPR spectra of EcPflD with EcPflC and titanium(III) citrate in the presence or absence of SAM. B. Detection of the formation of 1-deoxyfructose-6-phosphate (1-deoxy-F6P) in the isomerization of 1,5-anhydromannitol-6-phosphate (1,5-AM-6P) catalyzed by EcPflD by LC-MS. In the negative ion mode, the extracted ion chromatogram (EIC) at m / z = 243 was used to monitor the formation of 1-deoxy-F6P (t R = 18.40 min). C-D. Mass spectra (negative ion mode) of compound 1 (1-deoxy-F6P) and compound 6 (1,5-AM-6P) corresponding to the EIC peaks in B. E. Reaction analysis of EcFsaB-coupled EcPflD. HPLC elution profiles of the DNPH derivatives of the reaction products and standards. F-H. Mass spectra of DNPH-glyceraldehyde 3-phosphate (peak 3) and DNPH-hydroxyacetone (peaks 4 and 5) in E. I. Spectrophotometric determination of EcGldA-EcFsaB-coupled EcPflD, monitoring NADH consumption as hydroxyacetone is reduced. Among them, w / o means without.
[0031] Figure 4X-ray diffraction crystal structures of EcYbiW and SdPflD. A - B. Structures of monomers EcYbiW and SdPflD in each asymmetric unit. C - D. Structural active sites of the complexes formed by EcYbiW with 1,5-AG-6P and SdPflD with 1,5-AM-6P, respectively. The hydrogen atom transfer pathway is indicated by arrows, hydrogen bonds are represented by black dashed lines, and the distances between key atoms are marked. The 2Fo - Fc electron density of the substrate is 1.0σ. E - F. Catalytic mechanisms of EcYbiW and SdPflD.
[0032] Figure 5 In Escherichia coli, the 1,5-anhydroglucitol (1,5-AG)-induced YbiW-dependent glycolytic pathway and the 1,5-anhydromannitol (1,5-AM)-induced PflD-dependent glycolytic pathway. A. Growth of Escherichia coli MG1655 wild-type and ΔybiW strains with 1,5-AG as the sole carbon source. The positive control uses glucose as the sole carbon source, and the negative control (w / o glucose) does not add any carbon source. B. SDS-PAGE analysis of Escherichia coli MG1655 wild-type with glucose (lane 2), 1,5-AG as the sole carbon source (lane 3), or ΔybiW with glucose (lane 4) as the sole carbon source. Arrows indicate two ~95 kDa bands containing PtsA and YbiW, a ~42 kDa band identified as GldA, and a ~27 kDa band identified as containing FsaB and FsaA. C. Growth of Escherichia coli MG1655 wild-type and ΔpflD strains with 1,5-AM as the sole carbon source. The positive control uses glucose as the sole carbon source, and the negative control does not add any carbon source. D. SDS-PAGE analysis of Escherichia coli MG1655 wild-type with glucose (lane 2), 1,5-AM as the sole carbon source (lane 3), or ΔpflD with glucose (lane 4) as the sole carbon source. The ~95 kDa band indicated by the arrow is identified as PtsA, the ~90 kDa band is identified as PflD, the ~42 kDa band is identified as GldA, and the ~27 kDa band is identified as FsaB. E. Escherichia coli neighborhood genomes of YbiW (SEQ ID NO:1) and PflD (SEQ ID NO:98) in E. coli MG1655. F. 1,5-AG anhydroglycolytic pathway and 1,5-AM anhydroglycolytic pathway in Escherichia coli. Among them, w / o means without.
[0033] Figure 6SDS-PAGE analysis of purified proteins for enzymatic assays and biochemical characterization. A. LpYbiW; B. MBP-LpYbiY; C. LpFsaA; D. EcPflD; E. MBP-EcPflC; F. EcFsaB; G. EcGldA. Lanes 1 to 4 in each 4–20% gradient gel (Bis-Tris) contained protein molecular weight markers and 1, 2, and 4 μg of recombinant protein in sequence.
[0034] Figure 7 Characterization of recombinant LpYbiY. A. [Fe–S] cluster quantification in LpYbiY, where the assays were performed in triplicate and represented with standard deviation. B. The UV-Vis absorption spectrum of LpYbiY was separated and reconstructed, corresponding to the [4Fe-4S] 2+ cluster at 410 nm that disappeared upon reduction with the strong reducing agent titanium(III) citrate. C. LC-MS elution curves of the SAM cleavage reaction mixture catalyzed by LpYbiY in the presence and absence of Ti(III), using commercially available 5’-deoxyadenosine (5’-dA) as a standard. D. Positive ionization mass spectrum of the 5’-dA peak eluted in C at 27.22 minutes.
[0035] Figure 8 Characterization of recombinant EcPflC. A. [Fe–S] cluster quantification in EcPflC, where the assays were performed in triplicate and represented with standard deviation. B. The ultraviolet-visible absorption spectrum of EcPflC was separated and reconstructed, corresponding to the [4Fe-4S] 2+ cluster at 410 nm that disappeared upon reduction with the strong reducing agent titanium(III) citrate. C. LC-MS elution curves of the SAM cleavage reaction mixture catalyzed by EcPflC in the presence and absence of Ti(III), using commercially available 5’-deoxyadenosine (5’-dA) as a standard. D. Positive ionization mass spectrum of the 5’-dA peak eluted in C at 27.22 minutes.
[0036] Figure 9 LC-MS enzyme activity assays of LpFsaA and EcFsaB for catalyzing the aldol addition reaction of hydroxyacetone and glyceraldehyde 3-phosphate. A-B. Extracted ion chromatograms (m / z = 243, negative ion mode) of the LpFsaA- and EcFsaB-catalyzed reaction mixtures, monitoring the formation of 1-deoxyfructose-6-phosphate (1-deoxy-F6P). C-D. Negative ionization mass spectra of 1-deoxyfructose-6-phosphate formed by LpFsaA and EcFsaB. Among them, w / o means without.
[0037] Figure 10Enzyme kinetic data for LpYbiW. A. Dose-dependent LpFsaA-EcGldA coupled enzyme activity assay of LpYbiW, showing the amount of LpYbiW used in the assay. B. Michaelis-Menten kinetics of LpYbiW. Error bars represent the standard deviation of three independent experiments.
[0038] Figure 11 Enzyme kinetic data for EcPflD. A. Dose-dependent EcFsaB-EcGldA coupled enzyme activity assay of EcPflD, showing the amount of EcPflD used in the assay. B. Michaelis-Menten kinetics of EcPflD. Error bars represent the standard deviation of three independent experiments.
[0039] Figure 12 SDS-PAGE and SEC analysis of purified EcYbiW (E114A, E115A, and K117A) and SdPflD for protein crystallization. A. 4-20% gradient gel (Bis-Tris): Lane 1 is the protein molecular weight marker; lanes 2-4 are 1, 2, and 4 μg of purified EcYbiW (E114A, E115A, and K117A), respectively. B. 4-20% gradient gel (Bis-Tris): Lane 1 is the protein molecular weight marker; lanes 2-4 are 1, 2, and 4 μg of purified SdPflD, respectively. C. SEC standard curve established using standards bovine thyroglobulin (669 kDa), horse apoferritin (443 kDa), sweet potato β-amylase (200 kDa), yeast alcohol dehydrogenase (150 kDa), BSA (66 kDa), and bovine carbonic anhydrase (29 kDa) (Sigma MWGF 1000-1KT). D-E. Elution curves of EcYbiW (E114A, E115A, and K117A) and SdPflD for determining the molecular weights of EcYbiW and SdPflD using Superdex200 gel filtration chromatography, with estimated molecular weights of 154.9 kDa and 77.9 kDa, respectively.
[0040] Figure 13 Gas chromatography (GC) analysis of the fermentation broth of wild-type E. coli MG1655 grown on glucose, 1,5-AG, and 1,5-AM. Among them, w / o glucose indicates no glucose and is the negative control; glucose is the positive control.
[0041] Figure 14 YbiW cluster and PflD cluster constructed using the sequence similarity network (SSN) tool, with each node representing a sequence with 80% or more similarity. The nodes where EcYbiW and EcPflD are located are marked with arrows in the two clusters, respectively.
[0042] Figure 15 A gene cluster distribution map (the direction of the gene indicates the direction of the coding strand, from 5' to 3') of natural isozymes including 1,5-anhydroglucitol-6-phosphate isomerase shown in SEQ ID NO:1, glycine radical enzyme activating enzyme of the S-adenosylmethionine radical enzyme family, and 1-deoxyfructose-6-phosphate aldolase, which are listed exemplarily.
[0043] Figure 16 A gene cluster distribution map of natural isozymes including 1,5-anhydromannitol-6-phosphate isomerase shown in SEQ ID NO:97, glycine radical enzyme activating enzyme of the S-adenosylmethionine radical enzyme family, 1-deoxyfructose-6-phosphate aldolase, and hydroxyacetone reductase, which are listed exemplarily.
[0044] Figure 17 SDS-PAGE analysis of purified proteins for biochemical characterization. A. EcYbiW; B. MBP-EcYbiY; C. SdPflD; D. MBP-SdPflC. Lanes 1 to 4 in each 4-20% gradient gel (Bis-Tris) contain protein molecular weight markers and 1, 2, and 4 μg of recombinant proteins in sequence.
[0045] Figure 18 LC-MS enzymatic reaction analysis of E. coli species YbiW. A. Detection of the formation of 1-deoxyfructose-6-phosphate (1-deoxy-F6P) in the isomerization of 1,5-anhydroglucitol-6-phosphate (1,5-AG-6P) catalyzed by EcYbiW by LC-MS. In the negative ion mode, the extracted ion chromatogram (EIC) at m / z = 243 monitors the formation of 1-deoxy-F6P (t R = 18.40 min). B-C. Mass spectra (negative ion mode) of compound 1 (1-deoxy-F6P) and compound 2 (1,5-AG-6P) corresponding to the EIC peaks in A, respectively. Among them, w / o means without.
[0046] Figure 19Gene cluster of PflD and LC-MS enzymatic reaction analysis of Streptococcus dysgalactiae subsp. Equisimilis species. A. Gene cluster in Streptococcus dysgalactiae subsp. Equisimilis (including the gene expressing 1,5-anhydro-D-mannitol-6-phosphate isomerase shown in SEQ ID NO: 97). B. Formation of 1-deoxy-D-fructose-6-phosphate (1-deoxy-F6P) in the isomerization of 1,5-anhydro-D-mannitol-6-phosphate (1,5-AM-6P) catalyzed by SdPflD detected by LC-MS. In the negative ion mode, the extracted ion chromatogram (EIC) at m / z = 243 monitors the formation of 1-deoxy-F6P (t R = 18.40 min). C-D. Mass spectra (negative ion mode) of compound 1 (1-deoxy-F6P) and compound 6 (1,5-AM-6P) corresponding to the EIC peaks in B. Among them, w / o means without.
[0047] Sequence description
[0048] SEQ ID NO: 1 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number P75793 expressed in Escherichia coli.
[0049] SEQ ID NO: 2 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A807DR53 expressed in Lactiplantibacillus plantarum.
[0050] SEQ ID NO: 3 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0R2FVZ1 expressed in Lactobacillus selangorensis.
[0051] SEQ ID NO: 4 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0A6S4M4 expressed in Streptococcus uberis.
[0052] SEQ ID NO:5 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0K1F3T2 expressed by Olsenella sp. oral taxon 807.
[0053] SEQ ID NO:6 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A134AK54 expressed by Leptotrichia wadei.
[0054] SEQ ID NO:7 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A134A1S1 expressed by Senella sp. DNF00959.
[0055] SEQ ID NO:8 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0E2UQR7 expressed by Streptococcus parauberis.
[0056] SEQ ID NO:9 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A173RRH3 expressed by Anaerostipes hadrus.
[0057] SEQ ID NO:10 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1E8VV95 expressed by Senella sp. HMSC062G07.
[0058] SEQ ID NO:11 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1I3HGA4 expressed by Selenomonas ruminantium.
[0059] SEQ ID NO:12 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A285HLM7 expressed by Orenia metallireducens.
[0060] SEQ ID NO: 13 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A2M9H9C8 expressed by Bifidobacterium primatium.
[0061] SEQ ID NO: 14 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number C9N0T4 expressed by Leptotrichia hofstadii F0254.
[0062] SEQ ID NO: 15 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0R2AF79 expressed by Ligilactobacillus agilis DSM 20509.
[0063] SEQ ID NO: 16 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number U2TM27 expressed by Senella profusa F0195.
[0064] SEQ ID NO: 17 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0A6PW80 expressed by Clostridium butyricum.
[0065] SEQ ID NO: 18 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1C0A9L3 expressed by Orenia metallireducens.
[0066] SEQ ID NO: 19 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0B8PMP8 expressed by Vibrio ishigakensis.
[0067] SEQ ID NO: 20 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A425WQY7 expressed by bacteria of the family Coriobacteriaceae.
[0068] SEQ ID NO:21 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number F2N714 expressed by Coriobacterium glomerans strain ATCC49209 / DSM 20642 / JCM 10262 / PW2.
[0069] SEQ ID NO:22 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number M2NEK5 expressed by Eggerthia catenaformis OT 569.
[0070] SEQ ID NO:23 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0A7FYZ2 expressed by Clostridium baratiistr. Sullivan.
[0071] SEQ ID NO:24 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A143X8V0 expressed by Clostridiales bacterium CHKCI006.
[0072] SEQ ID NO:25 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A2N2BBI5 expressed by Firmicutes bacterium HGW-Firmicutes-5.
[0073] SEQ ID NO:26 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A2T0BJ14 expressed by Clostridium vincentii.
[0074] SEQ ID NO:27 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A380JDF6 expressed by Streptococcus downei MFe28.
[0075] SEQ ID NO:28 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A847CI73 expressed by Erysipelotrichaceae bacterium.
[0076] SEQ ID NO:29 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A8J6Z417 expressed by Quinella sp. 3Q1.
[0077] SEQ ID NO: 30 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number D1AFQ7 expressed by Sebaldella termitidis (strain ATCC33386 / NCTC 11300).
[0078] SEQ ID NO:31 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A062WZ81 expressed by Ligilactobacillus animalis.
[0079] SEQ ID NO:32 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0R1V613 expressed by Liquorilactobacillus satsumensis DSM 16230.
[0080] SEQ ID NO:33 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0M0A869 expressed by Clostridium botulinum.
[0081] SEQ ID NO: 34 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0V8QAB4 expressed by Acetivibrio ethanolgignens.
[0082] SEQ ID NO:35 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A2N3J357 expressed by Aeromonas sobria.
[0083] SEQ ID NO: 36 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1I2MD83 expressed by Clostridium cadaveris.
[0084] SEQ ID NO:37 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1L8MKJ9 expressed by Streptococcus bovimastitidis.
[0085] SEQ ID NO:38 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A4R1NC63 expressed by Sodalis ligni.
[0086] SEQ ID NO:39 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A5R9CBU2 expressed by Lactococcus raffinolactis.
[0087] SEQ ID NO:40 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number C2ELS4 expressed by Lactobacillus ultunensis DSM 16047.
[0088] SEQ ID NO:41 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A3P1W4E8 expressed by Leptotrichia sp.OH3620.
[0089] SEQ ID NO:42 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A498R759 expressed by Lucifera butyrica.
[0090] SEQ ID NO:43 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A4V6Q2Z2 expressed by Fonticella tunisiensis.
[0091] SEQ ID NO:44 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A7MME9 expressed by Cronobacter sakazakii strain ATCC BAA-894.
[0092] SEQ ID NO:45 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0R1Q920 expressed by Liquorilactobacillus uvarum DSM 19971.
[0093] SEQ ID NO:46 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1Y4F189 expressed by Anaeromassilibacillus sp. An250.
[0094] SEQ ID NO:47 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A498D0T9 expressed by Anaerotruncus sp. 22A2-44.
[0095] SEQ ID NO:48 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A510JEA5 expressed by Leptotrichia hofstadii.
[0096] SEQ ID NO:49 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0A2VZ96 expressed by Beauveria bassiana D1-5.
[0097] SEQ ID NO:50 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A447V187 expressed by Cedecea lapagei.
[0098] SEQ ID NO:51 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number B2ISH7 expressed by Streptococcus pneumoniae strain CGSP14.
[0099] SEQ ID NO:52 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A075SG00 expressed by Streptococcus suis 6407.
[0100] SEQ ID NO:53 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0R2GM47 expressed by Lactiplantibacillus plantarum.
[0101] SEQ ID NO:54 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1Q8F4Y7 expressed by Aeromonas veronii.
[0102] SEQ ID NO:55 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A290PWH6 expressed by Lactococcus raffinolactis.
[0103] SEQ ID NO:56 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A447Y009 expressed by Escherichia coli.
[0104] SEQ ID NO:57 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0B7GIL0 expressed by Streptococcus sanguinis.
[0105] SEQ ID NO:58 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1L7RNK7 expressed by Actinomyces succiniciruminis.
[0106] SEQ ID NO:59 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A376TJI4 expressed by Escherichia coli.
[0107] SEQ ID NO:60 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A3R9YDI0 expressed by Vagococcus humatus.
[0108] SEQ ID NO:61 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A064C019 expressed by Streptococcus pneumoniae.
[0109] SEQ ID NO:62 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0P7FGJ3 expressed by Vibrio alginolyticus.
[0110] SEQ ID NO:63 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0R2KGV0 expressed by Ligilactobacillus acidipiscis.
[0111] SEQ ID NO:64 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1T4P4Y5 expressed by Pilibacter termitis.
[0112] SEQ ID NO:65 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1E3KU28 expressed by Lactiplantibacillus plantarum.
[0113] SEQ ID NO:66 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A349BXL4 expressed by bacteria of the family Lachnospiraceae.
[0114] SEQ ID NO:67 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A3T0ZV83 expressed by Aeromonas hydrophila.
[0115] SEQ ID NO:68 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1L8WPN6 expressed by Enterococcus ratti.
[0116] SEQ ID NO:69 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0R2H5D7 expressed by Kandleria vitulina.
[0117] SEQ ID NO:70 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0H3J4P2 expressed by Clostridium pasteurianum DSM525.
[0118] SEQ ID NO:71 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0T6U3W3 expressed by Aeromonas allosaccharophila.
[0119] SEQ ID NO:72 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number K8CC45 expressed by Enterobacter sakazakii.
[0120] SEQ ID NO:73 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1W6B1J9 expressed by Pantoea alhagi.
[0121] SEQ ID NO:74 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number B6FW14 expressed by the strain Peptacetobacter hiranonis DSM 13275 / JCM 10541 / KCTC 15199 / TO-931.
[0122] SEQ ID NO:75 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A034T1M9 expressed by Edwardsiella piscicida.
[0123] SEQ ID NO:76 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A089PXZ6 expressed by Cedecea neteri.
[0124] SEQ ID NO:77 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0A3AQL5 expressed by Chelonobacter oris.
[0125] SEQ ID NO:78 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0D0QVG1 expressed by Aeromonas sp. L_1B5_3.
[0126] SEQ ID NO:79 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0F4VRP2 expressed by Clostridium sp. IBUN125C.
[0127] SEQ ID NO:80 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1M5EXF5 expressed by Vibrio gazogenes DSM 21264.
[0128] SEQ ID NO:81 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A2Z5Y0I5 expressed by Melissococcus plutonius.
[0129] SEQ ID NO:82 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A5C7QZ11 expressed by Tolumonas sp.
[0130] SEQ ID NO:83 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A3D0GQ55 expressed by bacteria of the family Erysipelotrichaceae.
[0131] SEQ ID NO:84 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0R1F6B2 expressed by Loigolactobacillus coryniformis subsp. coryniformis KCTC 3167.
[0132] SEQ ID NO:85 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A4P9VTG8 expressed by Zooshikella ganghwensis.
[0133] SEQ ID NO:86 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A829T2C0 expressed by Vibrio sp.
[0134] SEQ ID NO:87 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A2T3N2G3 expressed by Photobacterium lipolyticum.
[0135] SEQ ID NO:88 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A8S7XWX0 expressed by Escherichia coli.
[0136] SEQ ID NO:89 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number X0PS46 expressed by Agrilactobacillus composti DSM 18527.
[0137] SEQ ID NO:90 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A061Q2Y5 expressed by Vibrio sp. JCM 19052.
[0138] SEQ ID NO:91 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A174H662 expressed by Clostridium symbiosum.
[0139] SEQ ID NO:92 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number F3Y874 expressed by Melissococcus plutonius strain ATCC35311 / CIP 104052 / LMG 20360 / NCIMB 702443.
[0140] SEQ ID NO:93 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1S1BRH2 expressed by Aerococcus sp. HMSC035B07.
[0141] SEQ ID NO:94 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A1Y4BCL8 expressed by Senella sp. An293.
[0142] SEQ ID NO:95 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A7Z0VH52.
[0143] SEQ ID NO:96 shows the amino acid sequence of 1,5-anhydroglucitol-6-phosphate isomerase with Uniprot accession number A0A0I9WIZ4.
[0144] SEQ ID NO:97 shows the amino acid sequence of 1,5-anhydromannitol-6-phosphate isomerase with NCBI accession number OCX05109.1 expressed by Streptococcus dysgalactiae subsp. Equisimilis.
[0145] SEQ ID NO:98 shows the amino acid sequence of 1,5-anhydromannitol-6-phosphate isomerase with Uniprot accession number P32674 expressed by Escherichia coli strain K12.
[0146] SEQ ID NO:99 shows the amino acid sequence of 1,5-anhydromannitol-6-phosphate isomerase with Uniprot accession number A0A7C9H654 expressed by Firmicutes bacteria.
[0147] SEQ ID NO:100 shows the amino acid sequence of 1,5-anhydromannitol-6-phosphate isomerase with Uniprot accession number A0A510J954 expressed by Pseudoleptotrichia goodfellowii.
[0148] SEQ ID NO:101 shows the amino acid sequence of 1,5-anhydromannitol-6-phosphate isomerase with Uniprot accession number D4F6M7 expressed by Edwardsiella tarda.
[0149] SEQ ID NO:102 shows the amino acid sequence of 1,5-anhydromannitol-6-phosphate isomerase with Uniprot accession number D4E0C4 expressed by Serratia odorifera DSM 4582.
[0150] SEQ ID NO:103 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A2V4E297 expressed by Gilliamella apicola.
[0151] SEQ ID NO:104 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A084T2T3 expressed by Vibrio sp. ER1A.
[0152] SEQ ID NO:105 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A0R2FPV6 expressed by Lactobacillus selangorensis.
[0153] SEQ ID NO:106 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A0N8NM72 expressed by Caloranaerobacter sp.
[0154] SEQ ID NO:107 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A1T4K0K6 expressed by Pilibacter termitis.
[0155] SEQ ID NO:108 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A096BHK0 expressed by Caloranaerobacter azorensis H53214.
[0156] SEQ ID NO:109 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A725B0Z8 expressed by Salmonella enteritidis.
[0157] SEQ ID NO:110 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A0B8Q7B9 expressed by Vibrio sp. JCM 19236.
[0158] SEQ ID NO:111 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A037YM28 expressed by Escherichia coli.
[0159] SEQ ID NO:112 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number C4LFY8 expressed by Tolumonas auensis strain DSM 9187 / TA4.
[0160] SEQ ID NO:113 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A6N7XPT4 expressed by Senella porci.
[0161] SEQ ID NO:114 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A4S2F5T8 expressed by Mycobacteriaceae bacterium.
[0162] SEQ ID NO:115 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A4R7KVA0 expressed by Fonticella tunisiensis.
[0163] SEQ ID NO:116 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A437UWZ5 expressed by Coriobacteriales bacterium OH1046.
[0164] SEQ ID NO:117 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A3R9P0U8 expressed by Bacillus sp. HMF5848.
[0165] SEQ ID NO:118 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A380TM32 expressed by Actinobacillus rossii.
[0166] SEQ ID NO:119 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number C2ELS8 expressed by Lactobacillus ultunensis DSM 16047.
[0167] SEQ ID NO:120 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A2N2B0Z5 expressed by Firmicutes bacterium HGW-Firmicutes-7.
[0168] SEQ ID NO:121 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with the Uniprot accession number A0A1J0GF99 expressed by Clostridium estertheticum subsp. estertheticum.
[0169] SEQ ID NO:122 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with the Uniprot accession number A0A1D8GPW2 expressed by Geosporobacter ferrireducens.
[0170] SEQ ID NO:123 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with the Uniprot accession number A0A0B8PE23 expressed by Vibrio ishigakensis.
[0171] SEQ ID NO:124 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with the Uniprot accession number A0A239SQG5 expressed by Streptococcus merionis.
[0172] SEQ ID NO:125 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with the Uniprot accession number A0A823B0Q1 expressed by Shigella boydii.
[0173] SEQ ID NO:126 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with the Uniprot accession number X0PU04 expressed by Agrilactobacillus composti DSM 18527.
[0174] SEQ ID NO:127 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with the Uniprot accession number A0A0R2FH98 expressed by Lactobacillus selangorensis.
[0175] SEQ ID NO:128 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with the Uniprot accession number A0A430AQ99 expressed by Vagococcus elongatus.
[0176] SEQ ID NO:129 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A1W1WZ75 expressed by Clostridium acidisoli DSM 12555.
[0177] SEQ ID NO:130 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A376L3P4 expressed by Escherichia coli.
[0178] SEQ ID NO:131 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A1T4WMT8 expressed by Caloramator quimbayensis.
[0179] SEQ ID NO:132 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A1I1H9K1 expressed by Clostridium uliginosum.
[0180] SEQ ID NO:133 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A143I1W1 expressed by Enterobacter asburiae.
[0181] SEQ ID NO:134 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number N2BYX9 expressed by Atopobium minutum.
[0182] SEQ ID NO:135 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A4S2EVE2 expressed by bacteria of the family Pyrrhocoraceae.
[0183] SEQ ID NO:136 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A7S7M803 expressed by Thermophilibacter immobilis.
[0184] SEQ ID NO:137 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A2T3N2G1 expressed by Photobacterium lipolyticum.
[0185] SEQ ID NO:138 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A381GG45 expressed by Citrobacter amalonaticus.
[0186] SEQ ID NO:139 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A0A0F8F5 expressed by Escherichia coli G3.
[0187] SEQ ID NO:140 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A0J0GCD9 expressed by Enterobacter asburiae.
[0188] SEQ ID NO:141 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A0H3J2B1 expressed by Clostridium pasteurianum.
[0189] SEQ ID NO:142 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number D1AJA5 expressed by Sebaldella termitidis strain ATCC 33386 / NCTC 11300.
[0190] SEQ ID NO:143 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A8S0W6S0 expressed by Acididesulfobacillus acetoxydans.
[0191] SEQ ID NO:144 shows the amino acid sequence of 1,5-anhydro-D-mannitol-6-phosphate isomerase with Uniprot accession number A0A7W2G1V0 expressed by Clostridium sp.
[0192] SEQ ID NO:145 shows the amino acid sequence of LpYbiY with Uniprot accession number A0A162FGG2 expressed by Lactiplantibacillus plantarum.
[0193] SEQ ID NO:146 shows the amino acid sequence of LpFsaA with Uniprot accession number A0A0M4CJJ6 expressed by Lactiplantibacillus plantarum.
[0194] SEQ ID NO:147 shows the amino acid sequence of EcPflC with Uniprot accession number P32675 expressed by Escherichia coli strain K12.
[0195] SEQ ID NO:148 shows the amino acid sequence of EcFsaB with Uniprot accession number P32669 expressed by Escherichia coli strain K12.
[0196] SEQ ID NO:149 shows the amino acid sequence of EcGldA with Uniprot accession number P0A9S5 expressed by Escherichia coli strain K12.
[0197] SEQ ID NO:150 shows the amino acid sequence of the PTS subunit PtsA with Uniprot accession number P32670 expressed by Escherichia coli strain K12.
[0198] SEQ ID NO:151 shows the amino acid sequence of the PTS subunit FrwB with Uniprot accession number P69816 expressed by Escherichia coli strain K12.
[0199] SEQ ID NO:152 shows the amino acid sequence of the PTS subunit FrwC with Uniprot accession number P32672 expressed by Escherichia coli strain K12.
[0200] SEQ ID NO:153 shows the amino acid sequence of the PTS subunit FrwD with Uniprot accession number P32676 expressed by Escherichia coli strain K12.
[0201] SEQ ID NO:154 shows the nucleotide sequence of primer 1F.
[0202] SEQ ID NO:155 shows the nucleotide sequence of primer 1R.
[0203] SEQ ID NO:156 shows the nucleotide sequence of primer 2F.
[0204] SEQ ID NO:157 shows the nucleotide sequence of primer 2R.
[0205] SEQ ID NO:158 shows the nucleotide sequence of primer 3F.
[0206] SEQ ID NO:159 shows the nucleotide sequence of primer 3R.
[0207] SEQ ID NO:160 shows the nucleotide sequence of primer 4F.
[0208] SEQ ID NO:161 shows the nucleotide sequence of primer 4R.
[0209] SEQ ID NO:162 shows the nucleotide sequence of primer 5F.
[0210] SEQ ID NO:163 shows the nucleotide sequence of primer 5R.
[0211] SEQ ID NO:164 shows the nucleotide sequence of primer 6F.
[0212] SEQ ID NO:165 shows the nucleotide sequence of primer 6R.
[0213] SEQ ID NO:166 shows the nucleotide sequence of primer 7F.
[0214] SEQ ID NO:167 shows the nucleotide sequence of primer 7R.
[0215] SEQ ID NO:168 shows the nucleotide sequence of primer 8F.
[0216] SEQ ID NO:169 shows the nucleotide sequence of primer 8R.
[0217] SEQ ID NO:170 shows the nucleotide sequence of primer 9F.
[0218] SEQ ID NO:171 shows the nucleotide sequence of primer 9R.
[0219] SEQ ID NO:172 shows the nucleotide sequence of primer 10F.
[0220] SEQ ID NO:173 shows the nucleotide sequence of primer 10R.
[0221] SEQ ID NO:174 shows the nucleotide sequence of primer 11F.
[0222] SEQ ID NO:175 shows the nucleotide sequence of primer 11R.
[0223] SEQ ID NO:176 shows the nucleotide sequence of primer 12F.
[0224] SEQ ID NO:177 shows the nucleotide sequence of primer 12R.
[0225] SEQ ID NO:178 shows the nucleotide sequence of primer 13F.
[0226] SEQ ID NO:179 shows the nucleotide sequence of primer 13R.
[0227] SEQ ID NO: 180 shows the nucleotide sequence of primer 14F.
[0228] SEQ ID NO: 181 shows the nucleotide sequence of primer 14R.
[0229] SEQ ID NO: 182 shows the nucleotide sequence of primer 15F.
[0230] SEQ ID NO: 183 shows the nucleotide sequence of primer 16F.
[0231] SEQ ID NO: 184 shows the nucleotide sequence of primer 16R.
[0232] SEQ ID NO: 185 shows the nucleotide sequence of primer 17F.
[0233] SEQ ID NO: 186 shows the nucleotide sequence of primer 17R.
[0234] SEQ ID NO: 187 shows the nucleotide sequence of the guide RNA contained in the pRed_Cas9_recA_ΔybiW plasmid.
[0235] SEQ ID NO: 188 shows the nucleotide sequence of the guide RNA contained in the pRed_Cas9_recA_ΔpflD plasmid.
[0236] SEQ ID NO: 189 shows the nucleotide sequence of primer 18F.
[0237] SEQ ID NO: 190 shows the nucleotide sequence of primer 18R.
[0238] SEQ ID NO: 191 shows the nucleotide sequence of primer 19F.
[0239] SEQ ID NO: 192 shows the nucleotide sequence of primer 19R. DETAILED DESCRIPTION OF THE INVENTION
[0241] Through bioinformatics analysis, the inventors of the present application have discovered two new radical-dependent glycolytic pathways for the first time. Through heterologous expression and purification of the relevant enzymes in these two new pathways, the enzyme activities were verified by LC-MS detection, and the enzyme kinetic parameters were characterized. Combining with protein crystallography analysis, the catalytic mechanism of the key enzymes (such as 1,5-anhydroglucitol-6-phosphate isomerase or 1,5-anhydromannitol-6-phosphate isomerase) was explored. When 1,5-anhydroglucitol and 1,5-anhydromannitol were used as the sole carbon sources respectively, in Escherichia coli cells, the relevant pathway proteins were highly expressed; while the knockout of the key genes (such as the genes encoding 1,5-anhydroglucitol-6-phosphate isomerase or 1,5-anhydromannitol-6-phosphate isomerase) made Escherichia coli unable to grow when 1,5-anhydroglucitol or 1,5-anhydromannitol was used as the sole carbon source, indicating that Escherichia coli can utilize the two new glycolytic pathways discovered by the inventors of the present application to produce energy and grow. These are new glycolytic pathways since the discovery of the EMP, ED, and pentose phosphate pathways in the 1920s to 1950s of the last century, and both of these two new pathways can produce 1,2-propanediol.
[0242] Unless otherwise specified, the terms used in this application have the meanings commonly understood by those skilled in the art.
[0243] Definitions
[0244] Unless otherwise indicated, nucleic acids are written from left to right in the 5' to 3' direction; amino acid sequences are written from left to right in the amino to carboxyl direction. Numerical ranges include the numbers defining the range. Amino acids can be represented herein by their commonly known three-letter symbols or the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Similarly, nucleotides can be represented by the commonly accepted single-letter codes. The above-defined terms are more fully defined with reference to the entire specification. As used herein, the amino acid residue abbreviations are as follows: alanine is Ala or A; arginine is Arg or R; asparagine is Asn or N; aspartic acid is Asp or D; cysteine is Cys or C; glutamic acid is Glu or E; glutamine is Gln or Q; glycine is Gly or G; histidine is His or H; isoleucine is Ile or I; leucine is Leu or L; lysine is Lys or K; methionine is Met or M; phenylalanine is Phe or F; proline is Pro or P; serine is Ser or S; threonine is Thr or T; tryptophan is Trp or W; tyrosine is Tyr or Y; and valine is Val or V.
[0245] The "polypeptides" and "proteins" of the present application are used interchangeably herein and refer to polymers of amino acid residues, their variants, and synthetic and naturally occurring analogs. Thus, these terms apply to naturally occurring amino acid polymers and their naturally occurring chemical derivatives, as well as amino acid polymers in which one or more amino acid residues are synthetic non-naturally occurring amino acids (such as chemical analogs of the corresponding naturally occurring amino acids). Such derivatives include, for example, post-translational modifications and degradation products, including phosphorylated, glycosylated, oxidized, isomerized, carboxylated, and deaminated variants of polypeptide fragments.
[0246] As used herein, the term "enzyme active center" refers to the part of an enzyme molecule that can directly bind to a substrate molecule and catalyze the chemical reaction of the substrate, and this part becomes the active center of the enzyme.
[0247] As used herein, the term "amino acid" refers to a compound in which a hydrogen atom on the carboxylic acid carbon atom is replaced by an amino group. An amino acid molecule contains two functional groups, an amino group and a carboxyl group. It includes naturally occurring and non-naturally occurring amino acids, as well as amino acid analogs and mimetics. Naturally occurring amino acids include the 20 (L)-amino acids used in protein biosynthesis, as well as other amino acids, such as 4-hydroxyproline, hydroxylysine, carboxylated lysine, desmosine, isodesmosine, homocysteine, citrulline, and ornithine. Non-naturally occurring amino acids include, for example, (D)-amino acids, norleucine, norvaline, p-fluorophenylalanine, ethylthreonine, etc., which are known to those skilled in the art. Amino acid analogs include modified forms of naturally occurring and non-naturally occurring amino acids. Such modifications can include, for example, substituting chemical groups and moieties on the amino acid, or derivatization of the amino acid. Amino acid mimetics include, for example, organic structures that exhibit functionally similar properties, such as the charge and charge-space characteristics of amino acids. For example, an organic structure that mimics arginine (Arg or R) has a positive charge moiety located in a similar molecular space and having the same degree of mobility as the e-amino group of the side chain of the naturally occurring Arg amino acid. Mimetics also include constrained structures to maintain the optimal spatial and charge interactions of amino acids or amino acid functional groups. Those skilled in the art can determine what structures constitute functionally equivalent amino acid analogs and amino acid mimetics.
[0248] As used herein, the term "isoenzyme" refers to enzymes in an organism that catalyze the same reaction but have different molecular structures.
[0249] As used herein, the term "nucleic acid" refers to mRNA, RNA, cRNA, cDNA, or DNA, including single-stranded and double-stranded forms of DNA. The term generally refers to polynucleotide forms of nucleotides that are at least 10 bases in length, and the nucleotides are ribonucleotides or deoxyribonucleotides or modified forms of either type of nucleotide.
[0250] As used herein, the term "encoding," when used in the context of a particular nucleic acid, means that the nucleic acid contains the necessary information to direct the translation of that nucleotide sequence into a particular protein. The information for encoding the protein is expressed using codons. A nucleic acid encoding a protein may contain non-translated sequences (e.g., introns) located within the translated region of the nucleic acid or may lack such intervening non-translated sequences (e.g., as in cDNA).
[0251] As used herein, the "full-length sequence" with respect to a particular polynucleotide or the protein encoded thereby refers to the entire nucleic acid sequence or the entire amino acid sequence having the native (non-synthetic) endogenous sequence. The full-length polynucleotide encodes the full-length, catalytically active form of the particular protein.
[0252] As used herein, the term "isolated" refers to a polypeptide or nucleic acid or a biologically active portion thereof that is substantially or essentially free of components that normally accompany or interact with the protein or nucleic acid as found in its natural environment. Thus, an isolated polypeptide or nucleic acid produced by recombinant techniques is substantially free of other cellular material or culture medium, or an isolated polypeptide or nucleic acid chemically synthesized is substantially free of chemical precursors or other chemicals.
[0253] As used herein, the term "expression vector" is a recombinant or synthetically produced nucleic acid construct that has a series of specific nucleic acid elements that permit the transcription of a particular nucleic acid in a host cell.
[0254] As used herein, the term "host cell" refers to a cell that receives an exogenous gene in transformation and transduction (infection). A host cell can be a eukaryotic cell such as a yeast cell or a prokaryotic cell such as Escherichia coli.
[0255] As used herein, the term "phosphoenolpyruvate - phosphotransferase system (PTS)" refers to an enzyme complex widely present in bacteria, fungi, and some archaea, which consists of cytoplasmic enzyme I (EI) or histidine phosphocarrier protein (HPr or NPr) and sugar - specific enzyme II complexes and other phosphotransferases, and has both catalytic transport functions and very broad regulatory functions. The bacterial phosphoenolpyruvate - phosphotransferase system mainly phosphorylates various sugars and their derivatives through a phosphocascade reaction and then transports them into the cell. All PTSs rely on cytoplasmic enzyme I (EI) and histidine phosphocarrier protein (HPr), and the latter phosphorylates sugars through sugar - specific EII complexes, which are composed of two cytoplasmic domains (EIIA and EIIB) and one or two membrane domains (EIIC or EIID). The EIIC or EIID domain binds to the membrane and transfers the sugar into the cytoplasm, where the sugar undergoes a multi - stage phosphorylation process involving the EIIA and EIIB domains. The phosphate group of phosphoenolpyruvate is transferred from EI to HPr, then from HPr to EIIA, then from EIIA to EIIB, and then the phosphate group is transferred to the sugar transported by EIIC or EIID. Most bacteria use the phosphoenolpyruvate - phosphotransferase system (PTS) to transport carbohydrates such as glucose. EI catalyzes the phosphotransferase reaction from the glycolytic intermediate phosphoenolpyruvate (PEP) to HPr. Subsequently, HPr transfers the phosphate group to different EIIA, and then to EIIB. Finally, the sugar is transported across the membrane by EIIC and EIID and phosphorylated by EIIB at the same time. Specific embodiments
[0256] In a first aspect, the present application provides a polypeptide comprising the amino acid sequence shown in SEQ ID NO: 1 or 97, or a functional variant thereof, wherein the functional variant has 1,5 - anhydroglucitol - 6 - phosphate isomerase activity.
[0257] In some embodiments of the first aspect, the present application provides a polypeptide comprising the amino acid sequence shown in SEQ ID NO: 1 or a functional variant thereof, wherein the functional variant has 1,5 - anhydroglucitol - 6 - phosphate isomerase activity. In some embodiments, the polypeptide has an active site defined in terms of its spatial conformation: the active site comprises amino acid residues H165, H282, S283, H334, C441, E443, R453, T455, L562, S662, I664, and G786 that are spatially close to each other with reference to SEQ ID NO: 1.
[0258] In some embodiments of the first aspect, the functional variant is a natural isoenzyme of the amino acid sequence shown in SEQ ID NO: 1.
[0259] In some embodiments of the first aspect, the natural isoenzymes of the amino acid sequence shown in SEQ ID NO: 1 are from: Lactiplantibacillus plantarum, Escherichia coli, Lactobacillus selangorensis, Streptococcus uberis, Olsenella sp., Leptotrichia wadei, Streptococcus parauberis, Anaerostipes hadrus, Senella sp., Selenomonas ruminantium, Orenia metallireducens, Bifidobacterium primatium, Leptotrichia hofstadii, Ligilactobacillus agilis, Senella profusa, Clostridium butyricum, Vibrio ishigakensis, Coriobacteriaceae, Coriobacterium glomerans, Eggerthia catenaformis, Clostridium baratii str. Sullivan, Clostridiales, Firmicutes, Clostridium vincentii, Streptococcus downei, Erysipelotrichaceae, Quinella sp., Sebaldella termitidis, Ligilactobacillus animalis, Liquorilactobacillus satsumensis, Clostridium botulinum, Acetivibrio ethanolgignens, Aeromonas sobria, Clostridium cadaveris, Streptococcus bovimastitidis, Sodalis ligni, Lactococcus raffinolactis, Lactobacillus ultunensis, Leptotrichia sp., Lucifera butyrica, Fonticella tunisiensis, Cronobacter sakazakii, Liquorilactobacillus uvarum, Anaeromassilibacillus sp., Anaerotruncus sp.) Beauveria bassiana, Cedecea lapagei, Streptococcus pneumoniae, Streptococcus suis, Aeromonas veronii, Lactococcus raffinolactis, Streptococcus sanguinis, Actinomyces succiniciruminis, Vagococcus humatus, Vibrio alginolyticus, Ligilactobacillus acidipiscis, Pilibacter termitis, Lachnospiraceae, Aeromonas hydrophila, Enterococcus ratti, Kandleria vitulina, Clostridium pasteurianum, Aeromonas allosaccharophila, Pantoea alhagi, Peptacetobacter hiranonis, Edwardsiella piscicida, Cedecea neteri, Chelonobacter oris, Aeromonas sp., Clostridium sp., Vibrio gazogenes, Melissococcus plutonius, Tolumonas sp., Loigolactobacillus coryniformis, Zooshikella ganghwensis, Vibrio sp., Photobacterium lipolyticum, Agrilactobacillus composti, Clostridium symbiosum, and Aerococcus sp..
[0260] In some specific embodiments of the first aspect, the natural isozyme of the amino acid sequence shown in SEQ ID NO:1 comprises the amino acid sequence shown in any one of SEQ ID NOs:2-96.
[0261] In some embodiments of the first aspect, the present application provides a polypeptide, which comprises the amino acid sequence shown in SEQ ID NO:97 or a functional variant thereof, wherein the functional variant has 1,5-anhydroglucitol-6-phosphate isomerase activity. In some embodiments, the polypeptide has an active site defined as follows in its spatial conformation: the active site comprises amino acid residues Q162, H169, S277, S278, R323, F331, P335, C431, E433, D445, Y628, V630 and G752 that are close to each other in spatial conformation and are referenced to SEQ ID NO:97.
[0262] In some embodiments of the first aspect, the functional variant is a natural isozyme of the amino acid sequence shown in SEQ ID NO:97.
[0263] In some embodiments of the first aspect, the natural isoenzymes of the amino acid sequence shown in SEQ ID NO: 97 are from: Escherichia coli, Streptococcus dysgalactiae subsp. Equisimilis, Firmicutes, Pseudoleptotrichia goodfellowii, Edwardsiella tarda, Serratia odorifera, Gilliamella apicola, Vibrio spp., Lactobacillus selangorensis, Caloranaerobacter sp., Pilibacter termitis, Caloranaerobacter azorensis, Salmonella enteritidis, Tolumonas auensis, Senella porci, Mycobacteriaceae, Fonticella tunisiensis, Coriobacteriales, Bacillus sp., Actinobacillus rossii, Lactobacillus ultunensis, Clostridium estertheticum, Geosporobacter ferrireducens, Vibrio ishigakensis, Streptococcus merionis, Shigella boydii, Agrilactobacillus composti, Lactobacillus selangorensis, Vagococcus elongatus, Clostridium acidisoli, Caloramator quimbayensis, Clostridium uliginosum, Enterobacter asburiae, Atopobium minutum, Thermophilibacter immobilis, Photobacterium lipolyticum, Citrobacter amalonaticus, Clostridium pasteurianum, Sebaldella termitidis, Acididesulfobacillus acetoxydans, and Clostridium spp.
[0264] In some specific embodiments of the first aspect, the natural isozyme of the amino acid sequence shown in SEQ ID NO: 97 comprises the amino acid sequence shown in any one of SEQ ID NOs: 98-144.
[0265] In some specific embodiments of the first aspect, the substrate of the polypeptide is 1,5-anhydroglucitol-6-phosphate.
[0266] In some specific embodiments of the first aspect, the 1,5-anhydroglucitol-6-phosphate is 1,5-anhydroglucose-6-phosphate or 1,5-anhydromannitol-6-phosphate.
[0267] In some specific embodiments of the first aspect, when the polypeptide has 1,5-anhydroglucose-6-phosphate isomerase activity, it isomerizes 1,5-anhydroglucose-6-phosphate, and the product generated is 1-deoxyfructose-6-phosphate.
[0268] In some specific embodiments of the first aspect, when the polypeptide has 1,5-anhydromannitol-6-phosphate isomerase activity, it isomerizes 1,5-anhydromannitol-6-phosphate, and the product generated is 1-deoxyfructose-6-phosphate.
[0269] In some embodiments of the first aspect, the functional variant is generated by one or more amino acid insertions, substitutions, and / or deletions based on the amino acid sequence shown in SEQ ID NO: 1 or 97 or its natural isozyme.
[0270] In some embodiments of the first aspect, the insertions, substitutions, and / or deletions do not occur in the active site.
[0271] In some embodiments of the first aspect, the number of amino acid insertions, substitutions, and / or deletions is 1-30, preferably 1-20, more preferably 1-10, and the obtained functional variant substantially retains the unchanged 1,5-anhydroglucitol-6-phosphate isomerase (such as 1,5-anhydroglucose-6-phosphate isomerase or 1,5-anhydromannitol-6-phosphate isomerase) activity.
[0272] In some embodiments of the first aspect, the functional variant differs from the amino acid sequence shown in SEQ ID NO: 1 or 97 by about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid insertions, substitutions, and / or deletions.
[0273] In some embodiments of the first aspect, the functional variant differs from the amino acid sequence shown in SEQ ID NO: 1 or 97 by about 1, 2, 3, 4, or 5 amino acid insertions, substitutions, and / or deletions.
[0274] In some embodiments of the first aspect, the polypeptide is an isolated polypeptide.
[0275] In some embodiments of the first aspect, the polypeptide belongs to the glycine radical enzyme family. When the polypeptide has 1,5-anhydroglucitol-6-phosphate isomerase or 1,5-anhydromannitol-6-phosphate isomerase activity, its catalysis involves glycine and cysteine radicals.
[0276] In a second aspect, the present application provides a nucleic acid molecule encoding the polypeptide of the first aspect.
[0277] In some embodiments of the second aspect, the nucleic acid molecule of the present application comprises a nucleotide sequence that hybridizes under stringent conditions to a nucleotide sequence encoding a polypeptide shown in any one of SEQ ID NOs: 1-96, or consists of a nucleic acid sequence that specifically hybridizes to a nucleotide sequence encoding a polypeptide in any one of SEQ ID NOs: 1-96 and encodes a polypeptide that is functionally equivalent to the polypeptide shown in any one of SEQ ID NOs: 1-96.
[0278] In some embodiments of the second aspect, the nucleic acid molecule of the present application comprises a nucleotide sequence that hybridizes under stringent conditions to a nucleotide sequence encoding a polypeptide shown in any one of SEQ ID NOs: 97-144, or consists of a nucleic acid sequence that specifically hybridizes to a nucleotide sequence encoding a polypeptide in any one of SEQ ID NOs: 97-144 and encodes a polypeptide that is functionally equivalent to the polypeptide shown in any one of SEQ ID NOs: 97-144.
[0279] Those skilled in the art can routinely select the stringent conditions for DNA hybridization. Generally, longer probes require higher temperatures for proper annealing, while shorter probes require lower temperatures. Hybridization usually depends on the ability of denatured DNA to reanneal when the complementary strands are in an environment below their melting temperature. The higher the degree of homology between the probe and the hybridizable sequence, the higher the relative temperature that can be used. Thus, higher relative temperatures tend to make the reaction conditions more stringent, while at lower temperatures, the stringency is lower. For a detailed description of the stringent conditions for hybridization reactions, reference can be made to Ausubel et al., Current Protocols in Molecular Biology, Wiley Interscience Publishers, (1995).
[0280] In some embodiments of the second aspect, the stringent conditions employed for DNA hybridization include: 1) low ionic strength and high temperature during washing, such as 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate at 50 °C; 2) denaturing agents such as formamide during hybridization, such as 50% (v / v) formamide plus 0.1% bovine serum albumin / 0.1% Ficoll / 0.1% polyvinylpyrrolidone / 50 mM sodium phosphate buffer at pH 6.5 and 750 mM sodium chloride, 75 mM sodium citrate at 42 °C; or (3) overnight hybridization at 42 °C in a hybridization solution containing 50% formamide, 5×SSC (0.75 M sodium chloride, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5×Denhardt's solution, sonicated salmon sperm DNA (50 mg / mL), 0.1% SDS, and 10% dextran sulfate, followed by washing in 0.2×SSC (sodium chloride / sodium citrate) at 42 °C for 10 minutes and then high-stringency washing in 0.1×SSC containing EDTA at 55 °C. Moderate stringent conditions can be determined as described in Sambrook et al., Molecular Cloning: A Laboratory Manual, New York: Cold Spring Harbor Press, 1989. Moderate stringent conditions include using washing solutions and hybridization conditions (such as temperature, ionic strength, and SDS percentage) with a stringency lower than those described above. For example, moderate stringent conditions include hybridization at 42 °C with at least about 16% v / v to at least about 30% v / v formamide and at least about 0.5 M to at least about 0.9 M salt, and washing at 55 °C with at least about 0.1 M to at least about 0.2 M salt. Moderate stringent conditions can also include hybridization at 65 °C with 1% bovine serum albumin (BSA), 1 mM EDTA, 0.5 M NaHPO 4 (pH 7.2), 7% SDS, and washing with (i) 2×SSC, 0.1% SDS; or (ii) 0.5% BSA, 1 mM EDTA, 40 mM NaHPO 4(pH 7.2), wash with 5% SDS at 60 - 65 °C. Professionals will adjust the temperature, ionic strength, etc. according to factors such as the probe length. The stringency during nucleic acid hybridization depends on the nucleic acid molecule length and degree of complementarity, as well as other variables well known in the art. The greater the similarity or homology between two nucleotide sequences, the greater the Tm of the nucleic acid hybrid containing these sequences. The relative stability of nucleic acid hybridization (corresponding to a higher Tm) decreases in the following order: RNA:RNA, DNA:RNA, DNA:DNA. Preferably, the minimum length of the hybridizable nucleic acid is at least about 12 nucleotides, preferably at least about 16, more preferably at least about 24, and most preferably at least about 36 nucleotides.
[0281] The nucleic acid molecules of the present application can be combined with other DNA sequences, such as promoters, polyadenylation signals, other restriction enzyme cleavage sites, multiple cloning sites, other coding segments, etc., such that their total lengths can be significantly different. Therefore, it is contemplated that polynucleotide fragments of almost any length can be utilized; the total length is preferably limited by the convenience of preparation and use in the intended recombinant DNA protocols.
[0282] Any of a variety of well-established techniques known and available in the art can be utilized to prepare, manipulate, and / or express polynucleotides and their fusions. For example, a nucleic acid molecule encoding a polypeptide of the present application or a functional variant thereof can be used in a recombinant DNA molecule to direct the expression of the polypeptide in a suitable host cell. Due to the inherent degeneracy of the genetic code, other DNA sequences encoding substantially the same or functionally equivalent amino acid sequences can also be used in the present application, and these sequences can be used for cloning and expressing a given polypeptide.
[0283] In addition, the nucleic acid molecules of the present application can be modified using methods well known in the art, including but not limited to altering the cloning, processing, expression, and / or activity of the gene product.
[0284] In some embodiments of the second aspect, the nucleic acid molecule is produced by artificial synthesis, such as direct chemical synthesis or enzymatic synthesis.
[0285] In some embodiments of the second aspect, the nucleic acid molecule is produced by recombinant techniques.
[0286] In some embodiments of the second aspect, the nucleic acid molecule is an isolated nucleic acid molecule.
[0287] In a third aspect, the present application provides an expression cassette comprising the nucleic acid molecule described in the second aspect.
[0288] In some specific embodiments of the third aspect, the expression cassette can further comprise a 5' leader sequence that can enhance translation.
[0289] When preparing an expression cassette, various DNA fragments can be manipulated to provide DNA sequences in a suitable orientation and, where appropriate, in a suitable reading frame. For this purpose, linkers or adaptors can be used to join the DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, and so on. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, replacement, such as transition and transversion, can be involved.
[0290] In a fourth aspect, the present application provides an expression vector comprising the nucleic acid molecule described in the second aspect or the expression cassette described in the third aspect.
[0291] In some embodiments of the fourth aspect, any suitable expression vector can be used in the present application. For example, the expression vector can be a vector suitable for the Escherichia coli system of the present application. In some embodiments, the expression vector can be any one of vectors such as HT, pACYC, etc.
[0292] In some embodiments of the fourth aspect, a nucleic acid molecule encoding any polypeptide shown in SEQ ID NO: 1-96, or a nucleic acid molecule encoding any polypeptide shown in SEQ ID NO: 97-144, is cloned into a vector to construct a recombinant vector containing the nucleic acid molecule described in the present application.
[0293] In some embodiments of the fourth aspect, the expression vector for cloning polynucleotides is a plasmid vector.
[0294] In some embodiments of the fourth aspect, the above expression vector further comprises a regulatory sequence for regulating the expression of the nucleic acid molecule, wherein the nucleic acid molecule is operably linked to the regulatory sequence.
[0295] As used herein, the term "regulatory sequence" refers to a polynucleotide sequence required to effect the expression of an encoding sequence to which it is linked. The nature of such regulatory sequences varies with the host organism. In prokaryotes, such regulatory sequences generally include a promoter, a ribosome binding site, and a terminator; in eukaryotes, such regulatory sequences generally include a promoter, a terminator, and, in some cases, an enhancer. Thus, the term "regulatory sequence" includes all sequences that are minimally required for the expression of the target gene to occur, and can also include other sequences that are beneficial for the expression of the target gene, such as a leader sequence.
[0296] As used herein, the term "operably linked" refers to the situation where the sequences involved are in a relationship that allows them to function in the desired manner. Thus, for example, a regulatory sequence "operably linked" to a coding sequence enables the expression of the coding sequence under conditions compatible with the regulatory sequence.
[0297] In some embodiments of the fourth aspect, an expression vector containing a nucleotide sequence encoding a polypeptide shown in any one of SEQ ID NOs: 1-96 or a nucleotide sequence encoding a polypeptide shown in any one of SEQ ID NOs: 97-144 and appropriate transcriptional / translational regulatory elements is constructed using methods well known to those skilled in the art. These methods include in vitro recombinant DNA techniques, DNA synthesis techniques, in vivo recombinant techniques, etc. (Sambroook, et al. Molecular Cloning, a Laboratory Manual, cold Spring Harbor Laboratory. New York, 1989). The nucleotide sequence is operably linked to an appropriate promoter in the expression vector to direct mRNA synthesis. Representative examples of these promoters include: the lac or trp promoter of Escherichia coli; the PL promoter of phage λ; eukaryotic promoters include the CMV immediate early promoter, the HSV thymidine kinase promoter, the early and late SV40 promoters, the LTRs of retroviruses, and some other known promoters that can control gene expression in prokaryotic cells, eukaryotic cells, or their viruses. The expression vector also includes a ribosome binding site for translation initiation and a transcription terminator, etc. Inserting an enhancer sequence into the vector will enhance its transcription in higher eukaryotic cells. Enhancers are cis-acting factors for DNA expression, usually about 10 to 300 base pairs in length, and act on the promoter to enhance gene transcription. Examples include the 100 to 270 base pair SV40 enhancer on the late side of the replication origin, the polyomavirus enhancer on the late side of the replication origin, and the adenovirus enhancer, etc.
[0298] In addition, the expression vector preferably contains one or more selectable marker genes to provide phenotypic traits for selecting transformed host cells, such as those genes encoding resistance to kanamycin sulfate, ampicillin, etc.
[0299] In a fifth aspect, the present application provides a host cell comprising the nucleic acid molecule described in the second aspect, the expression cassette described in the third aspect, or the expression vector described in the fourth aspect.
[0300] In some embodiments of the fifth aspect, the host cell is capable of expressing and producing: a polypeptide comprising the amino acid sequence shown in SEQ ID NO: 1 or 97 or a functional variant thereof, wherein the functional variant has 1,5-sn-glycerol-6-phosphate isomerase activity.
[0301] In some embodiments of the fifth aspect, the host cell is capable of expressing and producing: a polypeptide comprising the amino acid sequence shown in SEQ ID NO: 1 or a functional variant thereof, wherein the functional variant has 1,5-anhydroglucitol-6-phosphate isomerase activity. In some embodiments, the polypeptide has an active site defined in terms of its spatial conformation: the active site comprises amino acid residues H165, H282, S283, H334, C441, E443, R453, T455, L562, S662, I664, and G786 that are spatially close to each other with reference to SEQ ID NO: 1.
[0302] In some embodiments of the fifth aspect, the host cell is capable of expressing and producing: a polypeptide comprising the amino acid sequence shown in SEQ ID NO: 97 or a functional variant thereof, wherein the functional variant has 1,5-anhydromannitol-6-phosphate isomerase activity. In some embodiments, the polypeptide has an active site defined in terms of its spatial conformation: the active site comprises amino acid residues Q162, H169, S277, S278, R323, F331, P335, C431, E433, D445, Y628, V630, and G752 that are spatially close to each other with reference to SEQ ID NO: 97.
[0303] In some embodiments of the fifth aspect, the host cell further expresses: an activating enzyme, an aldolase, a hydroxyacetone reductase, and / or a transport complex.
[0304] In some embodiments of the fifth aspect, the host cell further expresses the S-adenosylmethionine radical enzyme family glycine radical enzyme activating enzyme (enzyme classification number EC: 1.97.1.4) of the polypeptide described in the first aspect, 1-deoxyfructose-6-phosphate aldolase (also known as fructose-6-phosphate aldolase, its enzyme classification number EC: 4.1.2-), hydroxyacetone reductase (also known as glycerol dehydrogenase, its enzyme classification number EC: 1.1.1.6), and / or the transport complex.
[0305] In some embodiments of the fifth aspect, the transport complex has phosphoenolpyruvate-phosphotransferase system transport activity.
[0306] In some embodiments of the fifth aspect, the transport complex is the phosphoenolpyruvate-phosphotransferase system.
[0307] In some embodiments of the fifth aspect, the transport complex comprises at least one of SEQ ID NOs: 150-153 or a functional variant thereof.
[0308] In some embodiments of the fifth aspect, the transport complex comprises at least two of SEQ ID NOs: 150-153 or functional variants thereof.
[0309] In some embodiments of the fifth aspect, the transport complex comprises at least three of SEQ ID NOs: 150-153 or functional variants thereof.
[0310] In some embodiments of the fifth aspect, the transport complex comprises SEQ ID NOs: 150-153 or functional variants thereof.
[0311] In some embodiments of the fifth aspect, SEQ ID NOs: 150-153 respectively represent the amino acid sequences of Multiphosphoryl transfer protein 2 (PtsA), PTS system fructose-like EIIB component 2 (FrwB), Fructose-like permease EIIC component 2 (FrwC), and PTS system fructose-like EIIB component 3 (FrwD).
[0312] In some specific embodiments of the fifth aspect, the four sequences shown in SEQ ID NOs: 150-153 constitute the phosphoenolpyruvate - phosphotransferase system, wherein PtsA contains the EⅠ and EⅡA domains, FrwC is EⅡC, and FrwB and FrwD are EⅡB.
[0313] In some specific embodiments of the fifth aspect, PtsA, FrwB, FrwC, and FrwD interact to form a transport complex and exert the transport activity of the phosphoenolpyruvate - phosphotransferase system. In some embodiments, the transport complex transports 1,5-anhydrohexitols (such as 1,5-anhydroglucitol or 1,5-anhydromannitol) into the host cell for subsequent glycolysis.
[0314] In some embodiments of the fifth aspect, a functional variant of any one of SEQ ID NOs: 150-153 differs from the amino acid sequence shown in any one of the corresponding SEQ ID NOs: 150-153 by about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid insertions, substitutions, and / or deletions.
[0315] In some embodiments of the fifth aspect, the functional variant of any one of SEQ ID NOs: 150-153 differs from the amino acid sequence shown in any one of the corresponding SEQ ID NOs: 150-153 by about 1, 2, 3, 4, or 5 amino acid insertions, substitutions, and / or deletions.
[0316] In some embodiments of the fifth aspect, the available host cells are cells containing the above expression vector, which can be eukaryotic cells. For example, a yeast cell culture system can be used for the expression of the polypeptides of the present application. The host cells can also be prokaryotic cells containing the above expression vector, and can be selected from, for example, the genus Escherichia (such as Escherichia coli), the genus Klebsiella, the genus Streptococcus, the genus Lactobacillus, the genus Bifidobacterium, the phylum Bacteroidetes, and the phylum Firmicutes, etc.
[0317] In some specific embodiments of the fifth aspect, the host cells are yeast cells or Escherichia coli.
[0318] In some specific embodiments of the fifth aspect, the nucleic acid molecule encoding one or more enzymes of the present application can exist in the host cell in the form of a free vector, or can also be integrated into the genome of the host cell.
[0319] In some embodiments of any of the above aspects, the isolated nucleic acid is operably linked to a regulatory sequence that can be recognized by a host cell transformed with the expression vector.
[0320] Any technique known in the art can be used to introduce the expression vector into the host cell, including transformation, transduction, transfection, viral infection, gene gun, or Ti-mediated gene transfer. Specific methods include calcium phosphate transfection, DEAE-dextran-mediated transfection, lipofection, or electroporation, etc. (Davis, L., Dibner, M., Battey, I., Basic Methods in Molecular Biology, (1986)). As an example, when the host is a prokaryote such as Escherichia coli, competent cells can be harvested after the exponential growth phase and transformed by the well-known CaCl 2 method.
[0321] In a sixth aspect, the present application provides a composition comprising the polypeptide described in the first aspect.
[0322] In some embodiments of the sixth aspect, the composition further comprises: an activating enzyme, an aldolase, and / or a hydroxyacetone reductase.
[0323] In some specific embodiments of the sixth aspect, the composition comprises: 1,5-anhydroglucitol-6-phosphate isomerase, S-adenosylmethionine radical enzyme family glycine radical enzyme activating enzyme, 1-deoxyfructose-6-phosphate aldolase, and hydroxyacetone reductase.
[0324] In some embodiments of the sixth aspect, the S-adenosylmethionine radical enzyme family glycine radical enzyme activating enzyme is an activating enzyme of 1,5-anhydroglucitol-6-phosphate isomerase. In some embodiments, the S-adenosylmethionine radical enzyme family glycine radical enzyme activating enzyme comprises an [Fe-S] cluster. In some embodiments, the S-adenosylmethionine radical enzyme family glycine radical enzyme activating enzyme comprises a [4Fe-4S] cluster, i.e., it contains 4 Fe and 4 S. In some embodiments, the S-adenosylmethionine radical enzyme family glycine radical enzyme activating enzyme catalyzes the cleavage of S-adenosylmethionine (SAM).
[0325] In some embodiments of the sixth aspect, the aldolase is 1-deoxyfructose-6-phosphate aldolase.
[0326] In some embodiments of the sixth aspect, the composition is used for catalyzing the formation of 1,2-propanediol from 1,5-anhydroglucitol-6-phosphate.
[0327] In some embodiments of the sixth aspect, the 1,5-anhydroglucitol-6-phosphate is 1,5-anhydroglucitol-6-phosphate or 1,5-anhydromannitol-6-phosphate.
[0328] In a seventh aspect, the present application provides the use of the polypeptide described in the first aspect, the nucleic acid molecule described in the second aspect, the expression cassette described in the third aspect, the expression vector described in the fourth aspect, and the host cell described in the fifth aspect in the preparation of a composition for catalyzing the formation of 1,2-propanediol from 1,5-anhydroglucitol-6-phosphate.
[0329] In some embodiments of the seventh aspect, the 1,5-anhydroglucitol-6-phosphate is 1,5-anhydroglucitol-6-phosphate or 1,5-anhydromannitol-6-phosphate.
[0330] In some specific embodiments, 1,5-anhydroglucitol is transported into a host cell by the transport complex (e.g., a transport complex having phosphoenolpyruvate - phosphotransferase system transport activity) and phosphorylated (e.g., at the C-6 position) to generate 1,5-anhydroglucitol-6-phosphate; 1,5-anhydroglucitol-6-phosphate is converted into 1-deoxyfructose-6-phosphate under the catalysis of 1,5-anhydroglucitol-6-phosphate isomerase; 1-deoxyfructose-6-phosphate is cleaved by aldolase (e.g., 1-deoxyfructose-6-phosphate aldolase) to generate hydroxyacetone and glyceraldehyde 3-phosphate; and hydroxyacetone is reduced by hydroxyacetone reductase to generate 1,2-propanediol.
[0331] In some specific embodiments, 1,5-anhydromannitol is transported into a host cell by the transport complex (e.g., a transport complex having phosphoenolpyruvate - phosphotransferase system transport activity) and phosphorylated (e.g., at the C-6 position) to generate 1,5-anhydromannitol-6-phosphate; 1,5-anhydromannitol-6-phosphate is converted into 1-deoxyfructose-6-phosphate under the catalysis of 1,5-anhydromannitol-6-phosphate isomerase; 1-deoxyfructose-6-phosphate is cleaved by aldolase (e.g., 1-deoxyfructose-6-phosphate aldolase) to generate hydroxyacetone and glyceraldehyde 3-phosphate; and hydroxyacetone is reduced by hydroxyacetone reductase to generate 1,2-propanediol.
[0332] In an eighth aspect, the present application provides a method for producing 1,2-propanediol, which includes culturing the host cell according to the fifth aspect in a culture medium containing 1,5-anhydroglucitol to obtain 1,2-propanediol.
[0333] In some embodiments of the eighth aspect, the 1,5-anhydroglucitol is 1,5-anhydroglucitol or 1,5-anhydromannitol.
[0334] In some embodiments of the eighth aspect, the method further includes the steps of collecting the culture and purifying 1,2-propanediol.
[0335] In a ninth aspect, the present application provides a method for producing 1,2-propanediol, which includes contacting the composition according to the sixth aspect with 1,5-anhydroglucitol-6-phosphate to obtain 1,2-propanediol.
[0336] In some embodiments of the ninth aspect, the 1,5-anhydroglucitol-6-phosphate is 1,5-anhydroglucitol-6-phosphate or 1,5-anhydromannitol-6-phosphate.
[0337] Examples
[0338] The following examples are illustrative only and are not intended to limit the scope of the embodiments of the present application or the scope of the appended claims.
[0339] Example 1. Identification, Characterization, and Testing of the YbiW-Related Pathway
[0340] Materials and Methods
[0341] Experimental Materials
[0342] Tryptone and yeast extract used to prepare LB medium were purchased from Oxoid Limited (Hampshire, UK). Ultra-pure deionized water from Millipore Direct-Q was used. TALON resin was purchased from Clontech Laboratories Inc (California, USA). 1,5-anhydroglucitol (1,5-AG) was purchased from Shanghai Yuanye Bio-Technology Co., Ltd., and 1,5-anhydroglucitol-6-phosphate (1,5-AG-6P) was chemically synthesized by Tianjin Siennes Biochemical Technology Co., Ltd. Direct-Q oligonucleotide primers were synthesized by Beijing Tsingke Biotechnology Co., Ltd. All protein purification chromatography experiments were carried out on a pure FPLC system (GE Healthcare, USA). Anaerobic experiments were carried out in a Lab2000 glove box (Etelux) protected by N 2 (oxygen concentration less than 5 ppm).
[0343] Gene Synthesis and Cloning
[0344] The gene fragments of LpYbiW (Uniprot accession number: A0A807DR53), LpYbiY (Uniprot accession number: A0A162FGG2), and E. coli codon-optimized LpFsaA (Uniprot accession number: A0A0M4CJJ6) were synthesized by Beijing Tsingke Biotechnology Co., Ltd. LpYbiW and LpFsaA were inserted into the SspI site of the HT plasmid (optimized pET28 vector) to express proteins with an N-terminal His 6 tag. While LpYbiY was inserted into the NdeI site of the pACYC-MBP vector to express proteins with an N-terminal His 6 tag and maltose-binding protein (MBP).
[0345] For the purpose of biochemical characterization, EcYbiW (Uniprot accession number: P75793) and EcYbiY (Uniprot accession number: P75794) were amplified from the E. coli MG1655 genome using primers 18F / 18R and 19F / 19R (Table 6). EcYbiW was inserted into the SspI site of the HT plasmid (optimized pET28 vector) to express proteins with an N-terminal His6 protein of the label. And EcYbiY was inserted into the NdeI site of the pACYC-MBP vector to express a protein with an N-terminal His 6 tag and maltose-binding protein (MBP).
[0346] Expression and purification of LpYbiW, LpYbiY and LpFsaA
[0347] The plasmids of HT-LpYbiW, pACYC-MBP-LpYbiY and HT-LpFsaA were transformed into Escherichia coli BL21(DE3) cells to express the corresponding proteins. LpYbiW and LpFsaA were screened on LB agar plates containing 50 μg / mL kanamycin, while LpYbiY positive clones were screened on LB agar plates containing 25 μg / mL chloramphenicol. The cells were cultured overnight in 4 mL of LB medium and then transferred to fresh LB medium (usually 1 L in a 2.6 L flask) and grown in an orbital shaker incubator at 37 °C and 220 rpm. When the OD 600 reached approximately 0.8, the temperature was reduced to 18 °C and isopropyl β-D-1-thiogalactopyranoside (IPTG) was added to a final concentration of 0.3 mM to induce the production of the target protein. After 16 - 20 hours, the cells were collected by centrifugation (8000 g, 10 minutes at 4 °C). The collected cells were resuspended in 40 mL of lysis buffer (50 mM Tris / HCl, pH 8.0, 100 mM KCl, 1 mM phenylmethylsulfonyl fluoride (PMSF), 0.2 mg / mL lysozyme, 0.03% Triton X-100 and 0.02 mg / mL DNase I) and stored frozen in a -80 °C refrigerator.
[0348] The above frozen cells were thawed and incubated at room temperature (RT, 25 °C) for 20 minutes during which cell lysis occurred. 5 mM β-mercaptoethanol (BME) was added and nucleic acids were precipitated with 1% streptomycin sulfate. The cell debris was removed by centrifugation at 10000×g for 10 minutes at 4 °C. For LpYbiW and LpFsaA, the supernatant was filtered through a 0.22 μm filter and loaded onto a 5 mL TALON Co 2+Loaded onto a column (TakaraBio USA, Inc.), and the impurity proteins were washed away with 10 column volumes of buffer A, and then the protein was eluted with 5 column volumes of buffer A containing 150 mM imidazole. The eluted protein (~20 mL) was dialyzed against 2 L of buffer A at 4 °C for 3 hours, concentrated and aliquoted, frozen in liquid nitrogen, and stored at -80 °C. For LpYbiY, the supernatant was filtered and loaded onto a column packed with 10 mL of amylose resin (New England Biolabs, Massachusetts, U.S.A.). The impurity proteins were washed away with 10 column volumes of buffer A, and then the target protein was eluted with buffer A containing 10 mM maltose.
[0349] The purified proteins were detected by SDS-PAGE using a commercial gel (SurePAGE, Bis-Tris, 4-20%). The absorbance of the proteins at 280 nm was measured using an ultra-micro UV-visible spectrophotometer (Hangzhou Mio Instruments Co., Ltd.) to calculate their concentrations. [LpYbiW (ε 280 = 104,630 M -1 cm -1 ), MBP-LpYbiY (ε 280 = 100,730 M -1 cm -1 ), LpFsaA (ε 280 = 20,400 M -1 cm -1 )].
[0350] Expression and purification of EcYbiW and EcYbiY for biochemical characterization
[0351] The methods and procedures for protein expression and purification, SDS-PAGE detection method, and concentration determination method were the same as those for LpYbiW and LpYbiY above. [EcYbiW (ε 280 = 99,700 M -1 cm -1 ), MBP-EcYbiY (ε 280 = 97,750 M -1 cm -1 )].
[0352] Reconstitution and characterization of the cofactor [Fe-S] cluster of LpYbiY
[0353] Sequence alignment showed that LpYbiY contains only one [4Fe–4S] cluster in the radical SAM domain, with a theoretical maximum of 4Fe and 4S per monomer. After degassing and deoxygenating the LpYbiY protein solution with argon, it was transferred into the glove box. 100 mM Tris / HCl, pH 7.5, 10 mM DTT, 4 equivalents of ammonium ferrous sulfate and sodium sulfide were added to the deoxygenated protein solution above, and incubated overnight in a metal bath (Dry Bath H2O3-100C; Coyote Bioscience, Beijing, China) at 4 °C. 4 equivalents of EDTA solution was added, and then it was repeatedly concentrated by using a centrifugal filtration ultrafiltration tube (1.5 mL Ym-30 Amicon; Millipore) and diluted and replaced with a buffer solution (20 mM Tris-HCl, pH 7.5 and 100 mM KCl).
[0354] Phenanthroline [3-(2-pyridyl)-5,6-diphenyl-1,2,4-triazine-p,p’-disulfonic acid monosodium salt] was used to determine the iron content of unconstructed and constructed LpYbiY. Using an AAS iron standard, a standard curve of Fe in the range of 0 - 600 μM was established. The sulfur content in unconstructed and constructed LpYbiY was determined by measuring the absorbance of methylene blue formed by reacting with N,N-dimethyl-p-phenylenediamine dihydrochloride (DPD). Using Na 2 2S, a standard curve of S in the range of 0–600 μM was established.
[0355] The UV-Vis absorption spectrum of LpYbiY in the range of 200 - 800 nm was measured using a nanophotometer NP80 Mobile (Germany). Unconstructed and reconstituted LpYbiY were diluted to 10 μM with a buffer solution containing 20 mM Tris / HCl, pH 7.5 and 100 mM KCl, and transferred to a septum-sealed anaerobic cuvette for measurement. 10 equivalents of titanium(III) citrate were added to the reconstituted LpYbiY solution and incubated for 10 minutes to determine its reduced form.
[0356] LC-MS analysis of the cleavage of S-adenosylmethionine (SAM) catalyzed by LpYbiY
[0357] A reaction system (500 μL) containing 20 mM Tris / HCl, pH 7.5, 100 mM KCl, 200 μM titanium(III) citrate, 20 μM reconstituted LpYbiY, and 1 mM SAM was incubated overnight at room temperature in a glove box. Titanium(III) citrate was not added to the negative control. The reaction was quenched by adding formic acid (final concentration 5% v / v) and incubated in a boiling water bath for 1 minute to denature the protein. The precipitated protein was removed by centrifugation at 14000×g for 15 minutes. The supernatant was filtered through a 0.22 μm PES membrane and subjected to LC-MS analysis. The method was as follows: 20 μL of the sample was loaded onto an Agilent 6420 Triple Quadrupole LC / MS instrument (Agilent Technologies) equipped with a C18 reversed-phase column, and the ultraviolet absorption at 257 nm was detected. The mobile phase system consisted of water (A) and acetonitrile (B), and a linear gradient elution from 0 - 16% B was carried out for 30 minutes at a flow rate of 0.5 mL / min. Commercial 5’-dA was used as a standard, and the formation of the SAM cleavage product catalyzed by LpYbiY was verified by mass spectrometry.
[0358] Electron paramagnetic resonance (EPR) spectroscopy was used to detect the generation of LpYbiW radicals
[0359] The glycine radical of LpYbiW was characterized using continuous-wave X-band electron paramagnetic resonance (EPR) spectroscopy. A reaction system (200 μL) containing 20 mM Tris / HCl, pH 7.5, 100 mM KCl, 100 μM titanium(III) citrate, 1 mM SAM, 80 μM reconstituted LpYbiY, and 40 μM LpYbiW was incubated in a glove box at room temperature for 15 minutes. 10% glycerol was added, and then it was loaded into an EPR tube (Wilmad Lab Glass, 734-LPV-7) with an outer diameter of 4 mm and a length of 8 inches. It was sealed with a rubber stopper, taken out of the glove box, and frozen using liquid nitrogen before EPR analysis. The experimental spectrum of the glycine radical was modeled by Bruker Xepr spin fitting to obtain the g-value, hyperfine coupling constant, and linewidth. The double integral of the simulated spectrum was used to measure the spin concentration. The EPR spectrum was obtained by superimposing 30 scans, and the test conditions were as follows: temperature, 90 K; center field, 3370.00 gauss; range, 200 gauss; microwave power, 10 μW; microwave frequency, 9.43 MHz; modulation amplitude, 0.5 mT; modulation frequency, 100 kHz; time constant, 20.48 ms; conversion time, 25 ms; scan time, 20 seconds; receive gain, 43 dB.
[0360] LC-MS analysis for the activity assay of LpYbiW
[0361] As described in the above EPR experiment, LpYbiW was activated without adding glycerol in this step. A 200 μL reaction mixture containing 10 μM activated LpYbiW, 0.1 mM titanium(III) citrate, and 10 mM 1,5-AG-6P was incubated at room temperature for 1 h in a glove box. Negative controls were prepared without adding 1,5-AG-6P, SAM, or activated LpYbiW, respectively. 200 μL of acetonitrile was added to the reaction system to precipitate proteins, and the precipitate was removed by centrifugation. The supernatant was filtered through a 0.22 μm PES membrane for LC-MS analysis.
[0362] LC-MS analysis was performed using an Agilent 6420 Triple Quadrupole LC / MS instrument (Agilent Technologies). The drying gas temperature was maintained at 300 °C, the flow rate was 9 L / min, and the nebulizer pressure was 15 psi. LC-MS analysis was carried out using a ZIC-HILIC column (5 mm, 150 × 4.6 mm; Merck). The HPLC conditions were as follows: mobile phase A was 90% 20 mM ammonium acetate and 10% acetonitrile, and mobile phase B was acetonitrile; gradient elution was performed from 90% B to 70% B in 10 min and from 70% B to 50% B in 20 min. The flow rate was set at 0.5 mL / min. The mass spectrometer was operated in the ESI negative ion mode.
[0363] LC-MS analysis of EcYbiW activity assay
[0364] The experimental method of LC-MS analysis for EcYbiW activity assay was the same as that for LpYbiW activity assay.
[0365] LC-MS analysis of the activity of LpFsaA-coupled LpYbiW
[0366] A reaction mixture (200 μL) containing 10 μM activated LpYbiW, 0.1 mM titanium(III) citrate, and 10 mM 1,5-AG-6P was incubated at room temperature for 1 h in a glove box, and then the enzyme activity of LpFsaA-coupled LpYbiW was assayed. 10 μM LpFsaA was added to the above reaction system, and after reacting for another 1 h in the glove box, 50 μL of the reaction sample was mixed with 550 μL of 0.73 M sodium acetate (pH 5.0), and then mixed with 400 μL of freshly prepared 2,4-dinitrophenylhydrazine (DNPH) solution (20 mg dissolved in 50 mL of methanol), and subsequently incubated at 50 °C for 1 h. The mixture was centrifuged at 13000×g for 5 min and filtered before LC-MS analysis. The product standard was prepared by treating 3-phosphoglyceraldehyde (5 mM, Sigema-Aldrich) and hydroxyacetone (2 mM, TCI) standards with sodium acetate (pH 5.0) and freshly prepared 2,4-dinitrophenylhydrazine (DNPH) in the same way for 1 h.
[0367] LC-MS analysis was performed using 20 μL of the sample on an Agilent ZORBAX SB-C18 reversed-phase column (4.6×250 mm) in the ESI negative ion mode. The solvent system consisted of solvent C (deionized ultrapure water containing 0.1% formic acid) and D (chromatographic acetonitrile containing 0.1% formic acid). The liquid phase analysis was carried out according to the following procedure: eluting from 15% to 100% of D at a flow rate of 1.0 mL / min within 15 min, and the detection wavelength was set at 360 nm.
[0368] LC-MS analysis of LpFsaA activity
[0369] The activity of LpFsaA catalyzing the aldol condensation reaction of hydroxyacetone and 3-phosphoglyceraldehyde was also assayed: A 200 μL reaction system containing 50 mM Tris / HCl, pH 8.0, 100 mM KCl, 5 mM 3-phosphoglyceraldehyde, 10 mM hydroxyacetone, and 10 μM LpFsaA was incubated at room temperature for 1 h. 3-Phosphoglyceraldehyde, hydroxyacetone, or LpFsaA was not added to the negative control. LC-MS analysis was performed using a ZIC-HILIC column with the same sample preparation method and elution conditions as described in the LpYbiW activity assay.
[0370] Determination of Michaelis-Menten kinetic parameters of LpYbiW
[0371] The Michaelis-Menten kinetic parameters of LpYbiW were determined by coupling LpFsaA and EcGldA. By detecting A 340nmThe consumption of NADH at this point was used to monitor the reaction rate of LpYbiW. Using the cuvette mode (1 cm optical path) of a nano - photometer, an ultra - micro ultraviolet - visible spectrophotometer in the glove box, the absorbance change of the reaction system at 340 nm was detected every 5 s. These reaction systems consisted of 100 nM activated LpYbiW, 20 mM Tris / HCl, pH 7.5, 100 mM KCl, different concentrations of 1,5 - anhydroglucitol - 6 - phosphate, 5 μM LpFsaA, 5 μM EcGldA, and 0.4 mM NADH.
[0372] Construction of the EcYbiW expression plasmid for crystallography research
[0373] By PCR, the Escherichia coli gene encoding YbiW (EcYbiW, Uniprot accession number: P75793) was amplified using the primer pair 5F / 5R and inserted into the NdeI site of the vector pACYC. Then, site - directed mutagenesis of the plasmid pACYC - EcYbiW was performed using the primer pair 6F / 6R to obtain the plasmid pACYC - EcYbiW (E114A, E115A, and K117A, which means that the amino acids at positions 114, 115, and 117 were mutated from glutamate (E) and lysine (K) to alanine (A)) to express mutants with reduced surface entropy of EcYbiW and improve the success rate of protein crystallization screening.
[0374] Expression and purification of EcYbiW for crystallography research
[0375] The plasmid pACYC - EcYbiW (E114A, E115A, and K117A) was transformed into Escherichia coli BL21(DE3) cells, and positive clones were screened using LB agar plates containing 25 μg / mL chloramphenicol. Single colonies were picked and cultured overnight in 4 mL of LB medium, then inoculated into 1 L of fresh LB medium and cultured at 37 °C and 220 rpm until the OD 600 reached approximately 0.8. Then, 0.3 mM IPTG was added to induce expression at 18 °C for 16 h. Cells in 1 L of the culture were collected by centrifugation (8000×g, 10 min) and resuspended in 40 mL of lysis buffer (50 mM Tris / HCl, pH 8.0, 100 mM KCl, 0.03% Triton X - 100). Cells were lysed by sonication and centrifuged at 10000×g for 10 min at 4 °C to remove insoluble cell precipitates. The supernatant was filtered through a 0.45 - μm filter and applied to 10 mL of TALONC Co 2+The column was then washed with 10 column volumes of buffer A (20 mM Tris / HCl, pH 7.5, 200 mM KCl and 5 mM BME) and eluted with 5 column volumes of buffer A containing 150 mM imidazole. The eluate was dialyzed against 2 L of buffer B (20 mM Tris / HCl, pH 7.5, 5 mM BME) for 3 h, then loaded onto a 10 mL DEAE column and eluted linearly with 30 column volumes of buffer B with a salt gradient from 0 to 500 mM KCl. The fractions containing EcYbiW were collected and concentrated to approximately 5 mL by ultrafiltration. The protein solution was then injected into a Superdex 200 gel filtration column (300 mL) and eluted with buffer C (20 mM Tris / HCl, pH 7.5, 100 mM KCl, 1 mM DTT). The sample eluate from the gel filtration column was reconcentrated to 13.3 mg / mL for crystallography studies.
[0376] Crystallization, data collection and structure determination of EcYbiW
[0377] The initial screening of EcYbiW crystals was performed by the sitting-drop method using an automated liquid handling robotic system (Gryphon, Art Robbins) in 96-well format. Screening was carried out at 291 K using various crystallization screening kits from Hampton Research and Molecular Dimensions. After further optimization by the hanging-drop method in 24-well plates, crystals suitable for single-crystal X-ray diffraction studies were obtained. The optimal conditions for growing EcYbiW bulk crystals were 0.2 M NaCl, 0.1 M Tris, pH 8.0, 25% PEG 3350 plus 10 mM 1,5-AG-6P, and a 6.7 mg / mL EcYbiW protein solution. The crystallization solution containing 15% glycerol was used as a cryoprotectant and rapidly cooled in liquid nitrogen. Diffraction data were collected and processed at beamline BL10U2 of the Shanghai Synchrotron Radiation Facility (SSRF) at a resolution of The crystal structure model created with the website PHYRE2 was subjected to molecular replacement using the PHENIX software. The structure was manually built using the Coot software according to the electron density map and further refined in the PHENIX software, and then uploaded to the RCSB Protein Data Bank (accession code: 8ID7). The appendix table contains the crystal structure data for data collection and final refinement (Table 1). All structure diagrams were generated using UCSF Chimera ( Figure 4 A, 4B).
[0378] Identification of the active center of 1,5-anhydroglucitol-6-phosphate isomerase
[0379] For the PDB data of the crystal of the obtained EcYbiW-1,5-AG-6P complex (resolution ), the UCSF Chimera software was used for display. With the substrate 1,5-AG-6P as the center, the amino acid residues within the range of the substrate center were selected for display. Based on the hydrogen bond distances formed between the substrate and the amino acid residues, the key amino acid residues binding to the substrate in the active center of 1,5-anhydroglucitol-6-phosphate isomerase were determined. Combining with the catalytic mechanism of glycine radical 1,2-lyase in the literature, the key amino acid residues in the active center of 1,5-anhydroglucitol-6-phosphate isomerase were determined.
[0380] Identification of Isoenzymes of 1,5-Anhydroglucitol-6-Phosphate Isomerase
[0381] To identify the isozymes of 1,5-anhydroglucitol-6-phosphate isomerase in the glycine radical enzyme (GRE) family, a sequence similarity network (SSN) of 25,347 unique sequences in the InterPro family IPR004184 was constructed using the web-based Enzyme Function Initiative-Enzyme Similarity Tool (EFI-EST). The alignment score threshold was set to 250, the minimum sequence length was set to 650, and the 80% similarity representative node network (RepNode network) was displayed using Cytoscape v3.519. Under these settings, previously characterized GREs with different catalytic activities were divided into different clusters.
[0382] Sequences with 80% or more similarity were represented by one node. The sequences were folded together to reduce the total number of nodes, making the less complex network easier to load in Cytoscape. For the cluster formed by YbiW, a representative sequence was selected from each node according to different nodes. According to the crystal structure of the complex formed by EcYbiW and the substrate, the amino acid residues interacting with the substrate were used as conserved sites for multiple sequence alignment. Sequences containing these conserved sites simultaneously were retained. Further analysis was performed on the neighboring genes of these sequences, and sequences containing aldolase in the gene cluster were retained, and finally representative sequences were selected.
[0383] Results, Analysis and Discussion
[0384] During the bioinformatics analysis of glycine radical enzyme (GRE) sequences in the UniProt database, the inventors of the present application noticed that a GRE with unknown function (YbiW) appears in a gene cluster involved in sugar metabolism. In addition to YbiW and its activating enzyme (YbiY), this gene cluster also contains an aldolase (FsaA). This aldolase can catalyze the formation of corresponding 2-keto-hexose-6-phosphate from hydroxyacetone, 1,3-dihydroxyacetone and glyceraldehyde 3-phosphate. Based on this, the inventors of the present application speculated that this GRE is involved in sugar metabolism and its substrate may be a hexose phosphate. Combining with the molecular docking results, the inventors of the present application speculated that the substrate of YbiW is 1,5-anhydroglucitol-6-phosphate (1,5-AG-6P), and proposed a new radical-dependent glycolysis pathway. In this pathway, 1,5-AG-6P is catalyzed by YbiW to break the C-O bond, generating 1-deoxyfructose-6-phosphate. This product is further cleaved by aldolase FsaA to produce hydroxyacetone and glyceraldehyde 3-phosphate( Figure 1 A).
[0385] To prove the inventors' conjecture, the inventors of the present application selected the YbiW, YbiY and FsaA genes from Lactiplantibacillus plantarum, and then heterologously expressed them in Escherichia coli BL21(DE3) cells and characterized the activities of these proteins.
[0386] The inventors of the present application reconstructed the cofactor [Fe-S] 2+ cluster of the activating enzyme LpYbiY and characterized it. Anaerobic reconstruction of the [4Fe-4S] cluster resulted in 1.54 ± 0.12 Fe and 2.01 ± 0.08 S per monomer (a radical SAM domain, the theoretical maximum of the [4Fe-4S] cluster is 4 Fe and 4 S) and a typical UV-visible spectrum of the [4Fe-4S] protein-containing, molar extinction coefficient extinction coefficient ε 410nm of 8.85 mM -1 cm -1 ( Figure 7 B). The ε 410nm of each [4Fe-4S] cluster is approximately 15 mM -1 cm -1 , so the inventors of the present application estimated that each monomer contains approximately 0.59 [4Fe-4S], which is roughly consistent with the measured Fe and S contents( Figure 7 A). LC-MS detection results showed that, like other radical SAM enzymes, LpYbiY catalyzes the cleavage of SAM in the presence of the reducing agent titanium(III) citrate to form 5'-deoxyadenosine( Figure 7C-7D). The EPR spectrum showed that incubating LpYbiW, LpYbiY, SAM, and Ti(III) formed free radicals, and these generated free radicals were quantified as 0.017 Gly· / dimer ( Figure 2 A).
[0387] To determine the activity of LpYbiW, activated LpYbiW was incubated with 1,5-AG-6P and analyzed by LC-MS ( Figure 2 B-2D). It can be seen that there is a new peak with m / z(-) = 243.0 in the full reaction group at t R = 18.40 min. In the activity experiment of LpFsaA catalyzing the aldol condensation reaction of hydroxyacetone and glyceraldehyde 3-phosphate, a new peak (m / z(-) = 243.0, t R = 18.40 min) was also generated in the full reaction group ( Figure 9 A,9C). The Michaelis-Menten kinetic parameters of LpYbiW were also determined ( Figure 10 , k cat = 43.38 ± 2.96 s -1 / LpYbiW, K M = 36.52 ± 6.66 mM). LpFsaA was coupled with LpYbiW for the reaction, and through LC-MS analysis, the production of the corresponding products hydroxyacetone and glyceraldehyde 3-phosphate was detected ( Figure 2 E-2H). These data indicate that LpYbiW catalyzes the cleavage of the C-O bond of 1,5-AG-6P to produce 1-deoxyfructose-6-phosphate (1-deoxy-F6P).
[0388] According to the LC-MS analysis results, it can be seen that there is a new peak with m / z(-) = 243.0 in the full reaction group at t R = 18.40 min ( Figure 18 A) is generated, and the mass spectrum corresponding to this new peak is Figure 18 B, and the mass spectrum corresponding to the substrate is Figure 18 C. The LC-MS results indicate that EcYbiW has 1,5-anhydroglucitol-6-phosphate isomerase activity.
[0389] To further study the catalytic mechanism of YbiW, the inventors of this application determined the crystal structure of the EcYbiW-1,5-AG-6P complex (resolution ). Each asymmetric unit of YbiW contains one monomer. Each monomer exhibits the typical β / α barrel fold common to other GREs ( Figure 4(A), and a radical-like conformation with Gly· and Cys· rings. From this crystal structure, the inventors of the present application can see that the S atom of the Cys· residue Cys441 is adjacent to the H atom at the C2 position of the substrate, with a distance of This is consistent with the catalytic mechanism involving "C2 hydrogen atom abstraction by Cys·". The 2-OH group of the substrate 1,5-AG-6P forms a hydrogen bond with Glu443; the phosphate group coordinates with His165, His334, and Arg453 of YbiW; the 1-O atom forms a hydrogen bond with His334, and the 3-OH forms a hydrogen bond with S662 ( Figure 4 (C). Glu443 is involved in the protonation of the substrate 2-OH, which is consistent with the role of the base in the catalytic mechanism of GRE 1,2-elimination enzymes.
[0390] Based on the crystal structure of the complex formed by EcYbiW and the substrate 1,5-anhydroglucitol-6-phosphate. By referring to the hydrogen bond distances formed by the substrate and amino acid residues, the key amino acid residues binding to the substrate in the active center of 1,5-anhydroglucitol-6-phosphate isomerase were determined. Combining with the catalytic mechanism of glycine radical 1,2-lyase in the literature, the key amino acid residues in the active center of 1,5-anhydroglucitol-6-phosphate isomerase were determined. In EcYbiW (SEQ ID NO:1), they are the amino acid residues H165, H282, S283, H334, C441, E443, R453, T455, L562, S662, I664, and G786. Figure 4 (E).
[0391] Based on the crystal structure of the complex formed by EcYbiW and 1,5-anhydroglucitol-6-phosphate. By referring to the hydrogen bond distances formed by the substrate and amino acid residues, the key amino acid residues binding to the substrate in the active center of 1,5-anhydroglucitol-6-phosphate isomerase were determined. Combining with the catalytic mechanism of glycine radical 1,2-lyase in the literature, the key amino acid residues in the active center of 1,5-anhydroglucitol-6-phosphate isomerase were determined. In EcYbiW (SEQ ID NO:1), they are the amino acid residues H165, H282, S283, H334, C441, E443, R453, T455, L562, S662, I664, and G786.
[0392] Based on the analysis of SSN (Sequence Similarity Network) data, the YbiW cluster showed a total of 207 nodes and 2333 consensus sequences. The inventors of this application selected a representative sequence from each node. According to the crystal structure of EcYbiW (Uniprot accession number: P75793), the amino acid residues interacting with the substrate (H165, H282, S283, H334, C441, E443, R453, T455, L562, S662, I664, and G786 referring to SEQ ID NO:1) were used as conserved sites for multiple sequence alignment. Sequences containing these conserved sites simultaneously were retained, and a total of 173 isoenzyme sequences were screened out. Further analysis was performed on the neighboring genes of these 173 sequences, and sequences containing aldolase FsaA in the gene cluster were retained. Finally, a total of 95 representative sequences ( Figure 14 and Figure 15 ) were listed. Table 2 shows the accession numbers, strain sources, and amino acid sequence numbers of 95 1,5-anhydroglucitol-6-phosphate isomerase isoenzymes found in the Uniprot database.
[0393] Example 2. Identification, Characterization, and Testing of the PflD-Related Pathway
[0394] Experimental Materials
[0395] Tryptone and yeast extract used to prepare LB medium were purchased from Oxoid Limited (Hampshire, UK). Ultra-pure deionized water from Millipore Direct-Q was used. TALON resin was purchased from Clontech Laboratories Inc (California, USA). 1,5-anhydromannitol (1,5-AM) and 1,5-anhydromannitol-6-phosphate (1,5-AM-6P) were chemically synthesized by Tianjin Sine Biochem Co., Ltd. Direct-Q oligonucleotide primers were synthesized by Beijing Tsingke Biotechnology Co., Ltd. All protein purification chromatography experiments were carried out on a pure FPLC system (GE Healthcare, USA). Anaerobic experiments were carried out in a Lab2000 glove box (Etelux) protected by N 2 (oxygen concentration less than 5 ppm).
[0396] Gene Synthesis and Cloning
[0397] EcPflD (Uniprot accession number: P32674), EcPflC (Uniprot accession number: P32675), EcFsaB (Uniprot accession number: P32669), and EcGldA (Uniprot accession number: P0A9S5) were amplified from the E. coli MG1655 genome using primer pairs 1F / 1R, 2F / 2R, 3F / 3R, and 4F / 4R, respectively (Table 3). EcPflD, EcFsaB, and EcGldA were inserted into the SspI site of the HT plasmid (optimized pET28 vector) to express proteins with an N-terminal His 6 tag. EcPflC was inserted into the SspI site of the HMT vector to express a protein with an N-terminal His 6 tag and maltose-binding protein (MBP).
[0398] For biochemical characterization purposes, gene fragments of SdPflD (NCBI accession number OCX05109.1) and SdPflC (NCBI accession number OCX05103.1) were synthesized by Beijing Tsingke Biotechnology Co., Ltd. SdPflD was inserted into the SspI site of the HT plasmid (optimized pET28 vector) to express a protein with an N-terminal His 6 tag. SdPflC was inserted into the NdeI site of the pACYC-MBP vector to express a protein with an N-terminal His 6 tag and maltose-binding protein (MBP).
[0399] Expression and purification of EcPflD, EcPflC, EcFsaB, and EcGldA
[0400] The plasmids HT-EcPflD, HMT-EcPflC, HT-EcFsaB, and HT-EcGldA were transformed into E. coli BL21(DE3) cells to express the corresponding proteins. Positive clones were screened using LB agar plates containing 50 μg / mL kanamycin. The cells were cultured overnight in 4 mL of LB medium and then transferred to fresh LB medium (usually 1 L in a 2.6 L flask) and cultured in an orbital shaker at 37 °C and 220 rpm. When OD 600When it reaches about 0.8, the temperature is lowered to 18 °C, and isopropyl β-D-1-thiogalactopyranoside (IPTG) with a final concentration of 0.3 mM is added to induce the expression of the target protein. After 16 - 20 hours, the cells are collected by centrifugation (8000×g, 10 minutes at 4 °C). The collected cells are resuspended in 40 mL of lysis buffer (50 mM Tris / HCl, pH 8.0, 100 mM KCl, 1 mM phenylmethylsulfonyl fluoride (PMSF), 0.2 mg / mL lysozyme, 0.03% Triton X-100 and 0.02 mg / mL DNase I), and stored in a -80 °C refrigerator.
[0401] The frozen cells are thawed and incubated at room temperature (RT, 25 °C) for 20 minutes, during which cell lysis occurs. 5 mM β-mercaptoethanol (BME) is added, and nucleic acids are removed by precipitation with 1% streptomycin sulfate. At 4 °C, the cell debris is removed by centrifugation at 10000×g for 10 minutes. For EcPflD, EcFsaB and EcGldA, the supernatant is filtered through a 0.22 μm filter and loaded onto a 5 mL TALON Co 2+ column (Takara Bio USA, Inc.) pre-equilibrated with buffer A (20 mM Tris / HCl, pH 7.5, 200 mM KCl and 5 mM BME), and the column is washed with 10 column volumes of buffer A to remove contaminating proteins, and then the protein is eluted with 5 column volumes of buffer A containing 150 mM imidazole. The eluted protein (~20 mL) is dialyzed against 2 L of buffer A at 4 °C for 3 hours, concentrated and aliquoted, frozen with liquid nitrogen, and stored at -80 °C. For EcPflC, the supernatant is filtered and loaded onto a column packed with 10 mL of amylose resin (New England Biolabs, Massachusetts, U.S.A.), the column is washed with 10 column volumes of buffer A to remove contaminating proteins, and then the target protein is eluted with buffer A containing 10 mM maltose.
[0402] The purified proteins are detected by SDS-PAGE using a commercial gel (SurePAGE, Bis-Tris, 4 - 20%). The absorbance of the proteins at 280 nm is measured using an ultra-micro ultraviolet-visible spectrophotometer (Hangzhou Mio Instruments Co., Ltd.) to calculate their concentrations. [EcPflD (ε 280 = 74,720 M -1 cm -1 ), MBP-EcPflC (ε 280 = 88,810 M -1 cm -1 ), EcFsaB (ε 280 = 18,450 M-1 cm -1 ), EcGldA(ε 280 = 32,890 M -1 cm -1 )].
[0403] Expression and Purification of SdPflD and SdPflC for Biochemical Characterization
[0404] The protein expression and purification methods and procedures, SDS-PAGE detection method, and concentration determination method are the same as those for LpYbiW and LpYbiY. [SdPflD(ε 280 = 86,070 M -1 cm -1 ), MBP-SdPflC(ε 280 = 107,260 M -1 cm -1 )].
[0405] Reconstitution and Characterization of the [Fe-S] Cluster of the EcPflC Prosthetic Group
[0406] Sequence alignment showed that EcPflC contains only one [4Fe–4S] cluster in the radical SAM domain, and the theoretical maximum for each monomer is 4Fe and 4S. After degassing and deoxygenating the EcPflC protein solution with argon, it was transferred into the glove box. To the deoxygenated protein solution described above, 100 mM Tris-HCl, pH 7.5, 10 mM DTT, 4 equivalents of ammonium ferrous sulfate, and sodium sulfide were added, and the mixture was incubated overnight in a metal bath (Dry Bath H2O3-100C; Coyote Bioscience, Beijing, China) at 4 °C. Then 4 equivalents of EDTA solution were added, and the mixture was repeatedly concentrated using a centrifugal filter ultrafiltration tube (1.5 mL Ym-30 Amicon; Millipore) and diluted and exchanged using a buffer (20 mM Tris / HCl, pH 7.5 and 100 mM KCl).
[0407] Phenanthroline [3-(2-pyridyl)-5,6-diphenyl-1,2,4-triazine-p,p'-disulfonic acid monosodium salt] was used to determine the iron content of unconstructed and constructed EcPflC. Using an AAS iron standard, a standard curve for Fe in the range of 0 - 600 μM was established. The sulfur content in unconstructed and constructed EcPflC was determined by measuring the absorbance of methylene blue formed by reacting with N,N-dimethyl-p-phenylenediamine dihydrochloride (DPD). Using Na 2 S, a standard curve for S in the range of 0 - 600 μM was established.
[0408] The UV-Vis absorption spectrum of EcPflC in the range of 200 - 800 nm was measured using a NanoPhotometer NP80 Mobile (Germany). The unconstructed and reconstituted EcPflC was diluted to 10 μM using a buffer containing 20 mM Tris-HCl, pH 7.5, and 100 mM KCl, and transferred to a septum-sealed anaerobic cuvette for measurement. In the reconstituted EcPflC solution, 10 equivalents of titanium(III) citrate were added and incubated for 10 minutes to determine its reduced form.
[0409] LC-MS analysis of the catalysis of SAM cleavage by EcPflC
[0410] A reaction system (500 μL) containing 20 mM Tris / HCl, pH 7.5, 100 mM KCl, 200 μM titanium(III) citrate, 20 μM reconstituted EcPflC, and 1 mM SAM was incubated overnight at room temperature in a glove box. Titanium(III) citrate was not added to the negative control. The reaction was quenched by adding formic acid (final concentration 5% v / v) and incubated in a boiling water bath for 1 minute to denature the protein. The precipitated protein was removed by centrifugation at 14000×g for 15 minutes. The supernatant was filtered through a 0.22 μm PES membrane for LC-MS analysis. The method was as follows: 20 μL of the sample was loaded onto an Agilent 6420 Triple Quadrupole LC / MS instrument (Agilent Technologies) equipped with a C18 reverse-phase column, and the ultraviolet absorption was detected at 257 nm. The mobile phase system consisted of water (A) and acetonitrile (B), and a linear gradient elution from 0 - 16% B was performed for 30 minutes at a flow rate of 0.5 mL / min. Commercial 5’-dA was used as a standard to verify the formation of the EcPflC-catalyzed SAM cleavage product by mass spectrometry.
[0411] Detection of the generation of EcPflD radicals by electron paramagnetic resonance (EPR) spectroscopy
[0412] Characterization of the glycine radical of EcPflD using continuous-wave X-band electron paramagnetic resonance (EPR) spectroscopy. A reaction system (200 μL) containing 20 mM Tris / HCl, pH 7.5, 100 mM KCl, 100 μM citrate peptide Ti(III), 1 mM SAM, 80 μM reconstituted EcPflC, and 40 μM EcPflD was incubated in a glove box at room temperature for 15 minutes. 10% glycerol was added, and then it was loaded into an EPR tube (Wilmad Lab Glass, 734-LPV-7) with an outer diameter of 4 mm and a length of 8 inches. It was sealed with a rubber stopper, taken out of the glove box, and frozen using liquid nitrogen before EPR analysis. The experimental spectrum of the glycine radical was modeled by Bruker Xepr spin fitting to obtain the g-value, hyperfine coupling constant, and line width. The double integral of the simulated spectrum was used to measure the spin concentration. The EPR spectrum was obtained by superimposing 30 scans, and the test conditions were as follows: temperature, 90 K; center field, 3370.00 Gauss; range, 200 Gauss; microwave power, 10 μW; microwave frequency, 9.43 MHz; modulation amplitude, 0.5 mT; modulation frequency, 100 kHz; time constant, 20.48 ms; conversion time, 25 ms; scan time, 20 s; receive gain, 43 dB.
[0413] LC-MS analysis for the activity assay of EcPflD
[0414] A 200 μL reaction mixture containing 10 μM activated EcPflD, 0.1 mM titanium(III) citrate, and 10 mM 1,5-AM-6P was incubated in a glove box at room temperature for 1 h. Negative controls were prepared without adding 1,5-AM-6P, SAM, or activated EcPflD, respectively. 200 μL of acetonitrile was added to the reaction system to precipitate proteins, and the precipitate was removed by centrifugation. The supernatant was filtered through a 0.22 μm PES membrane for LC-MS analysis.
[0415] LC-MS analysis was performed using an Agilent 6420 Triple Quadrupole LC / MS instrument (Agilent Technologies). The drying gas temperature was maintained at 300 °C, the flow rate was 9 L / min, and the nebulizer pressure was 15 psi. LC-MS analysis was carried out using a ZIC-HILIC column (5 mm, 150 × 4.6 mm; Merck). The HPLC conditions were as follows: mobile phase A was 90% 20 mM ammonium acetate and 10% acetonitrile, and mobile phase B was acetonitrile; gradient elution was performed from 90% B to 70% B in 10 minutes and from 70% B to 50% B in 20 minutes. The flow rate was set at 0.5 mL / min. The mass spectrometer was operated in the ESI negative ion mode.
[0416] LC-MS Analysis of SdPflD Activity
[0417] The experimental method of LC-MS analysis for SdPflD of Streptococcus dysgalactiae subsp. Equisimilis species is the same as that of LC-MS analysis for EcPflD activity determination.
[0418] LC-MS Analysis of the Activity of EcFsaB Coupled with EcPflD
[0419] After incubating a reaction mixture (200 μL) containing 10 μM activated EcPflD, 0.1 mM titanium(III) citrate, and 10 mM 1,5-AM-6P at room temperature for 1 h in a glove box, the enzyme activity assay of EcFsaB coupled with EcPflD was carried out. 10 μM EcFsaB was added to the above reaction system, and after reacting for another 1 h in the glove box, 50 μL of the reaction sample was mixed with 550 μL of 0.73 M sodium acetate (pH 5.0), and then mixed with 400 μL of freshly prepared 2,4-dinitrophenylhydrazine (DNPH) solution (20 mg dissolved in 50 mL of methanol), and then incubated at 50 °C for 1 h. The mixture was centrifuged at 13000×g for 5 min and filtered before LC-MS analysis. The positive control was prepared by treating 3-phosphoglyceraldehyde (5 mM, Sigema-Aldrich) and hydroxyacetone (2 mM, TCI) standards in the same way as sodium acetate (pH 5.0) and freshly prepared DNPH solution for 1 h.
[0420] LC-MS analysis was performed with 20 μL of the sample on an Agilent ZORBAX SB-C18 reversed-phase column (4.6×250 mm) using the ESI negative ion mode. The solvent system consisted of solvent C (deionized ultrapure water containing 0.1% formic acid) and D (chromatographic acetonitrile containing 0.1% formic acid). The liquid phase analysis was carried out according to the following procedure: eluting from 15% to 100% of D at a flow rate of 1.0 mL / min within 15 min, and the detection wavelength was set at 360 nm.
[0421] Spectrophotometric Determination of the Coupling of EcPflD, EcFsaB, and EcGldA
[0422] The reaction mixture containing 20 mM Tris / HCl, pH 7.5, 100 mM KCl, 10 mM 1,5-AM-6P, 40 nM activated EcPflD, 5 μM EcFsaB, 5 μM EcGldA, and 0.4 mM NADH was incubated in a 1-cm Eppendorf cuvette in a glove box at room temperature. Using the cuvette mode of an ultra-micro UV-visible spectrophotometer (Hangzhou Mio Instruments Co., Ltd.) in the glove box, the absorbance at 340 nm was monitored every 5 seconds. The negative control was prepared by omitting 1,5-AM-6P, SAM, or EcPflD.
[0423] LC-MS analysis of EcFsaB activity
[0424] The activity of EcFsaB catalyzing the aldol condensation reaction of hydroxyacetone and glyceraldehyde 3-phosphate was also determined: 200 μL of the reaction mixture containing 50 mM Tris / HCl, pH 8.0, 100 mM KCl, 5 mM glyceraldehyde 3-phosphate, 10 mM hydroxyacetone, and 10 μM EcFsaB was incubated at room temperature for 1 hour. Glyceraldehyde 3-phosphate, hydroxyacetone, or EcFsaB was not added to the negative control. LC-MS analysis was performed using a ZIC-HILIC column with the same sample preparation method and elution conditions as described in the EcPflD activity assay.
[0425] Determination of Michaelis-Menten kinetic parameters of EcPflD
[0426] The reaction system containing 20 mM Tris / HCl, pH 7.5, 100 mM KCl, 40 nM activated EcPflD, and different substrate concentrations of 1,5-AM-6P, 2 μM EcFsaB, 2 μM EcGldA, and 0.4 mM NADH was carried out in a 1-cm Eppendorf cuvette using the cuvette mode of an ultra-micro UV-visible spectrophotometer (Hangzhou Mio Instruments Co., Ltd.) in the glove box. The decrease in A 340nm was monitored at 5-s intervals.
[0427] Construction of the SdPflD expression plasmid for crystallography study
[0428] The DNA fragment encoding PflD in Streptococcus dysgalactiae subsp. equisimilis optimized for E. coli codons (SdPflD, NCBI accession number: OCX05109.1) was synthesized by Beijing Tsingke Biotechnology Co., Ltd. and inserted into the SspI site of vector HT.
[0429] Expression and purification of SdPflD for crystallographic study
[0430] For crystallographic study of SdPflD, Escherichia coli BL21(DE3) carrying plasmid HT-SdPflD was cultured in 2 L of LB medium containing 50 μg / mL kanamycin. Cells were collected and lysed using the same method as for EcYbiW, and the supernatant of the cell lysate was purified using a 10 mL TALON Co 2+ column. The eluate was pre-mixed with recombinant His 6 -tagged TEV protease (the molar ratio of TEV protease to SdPflD was approximately 1:10) and dialyzed overnight against 2 L of buffer A (20 mM Tris / HCl, pH 7.5, 200 mM KCl, and 5 mM BME) containing 5 mM BME. The dialyzed sample was loaded onto a 10 mL TALON Co 2+ column to retain TEV protease and His 6 -tag. The flow-through was collected and dialyzed against 2 L of buffer B (20 mM Tris / HCl, pH 7.5, 5 mM BME), and then loaded onto a 10 mL DEAE column using the same method as described above. The main peak containing SdPflD was concentrated and purified using a Superdex200 gel filtration column. The sample from the gel filtration column was concentrated to a final concentration of 9.9 mg / mL.
[0431] Crystallization, data collection, and structure determination of SdPflD
[0432] Preliminary screening of SdPflD crystals was performed by the sitting-drop method using an automated liquid handling robotic system (Gryphon, Art Robbins) in 96-well format. Screening was carried out at 291 K using various crystallization screening kits from Hampton Research and Molecular Dimensions. After further optimization by the hanging-drop method in a 24-well plate, crystals suitable for single-crystal X-ray diffraction study were obtained. The optimal conditions for producing SdPflD block crystals were a 5.0 mg / mL SdPflD protein solution, 1.2 M sodium citrate, 0.1 M sodium HEPES, pH 8.0, and 10 mM 1,5-AM-6P. The block crystals were cryoprotected with 10% glycerol. Diffraction data were collected and processed at BL10U2 of the Shanghai Synchrotron Radiation Facility (SSRF) at a resolution of The crystal structure model created with the website PHYRE2 was subjected to molecular replacement on the PHENIX software. The structure was manually built using the Coot software according to the electron cloud orientation, further optimized in the PHENIX software, and then uploaded to the RCSB Protein Data Bank (accession code: 8ID0). The appendix table contains the crystal structure data for data collection and final optimization (Table 1). All structure diagrams were generated using UCSF Chimera( Figure 4 B and 4D).
[0433] Identification of the active center of 1,5-anhydro-D-mannitol-6-phosphate isomerase
[0434] For the PDB data of the crystal of SdPflD complexed with 1,5-AM-6P obtained (resolution ), the UCSF Chimera software was used for display. With the substrate 1,5-AM-6P as the center, the amino acid residues within the range of the substrate center were selected for display. Based on the hydrogen bond distances formed between the substrate and the amino acid residues, the key amino acid residues binding to the substrate in the active center of 1,5-anhydro-D-mannitol-6-phosphate isomerase were determined. Combining with the catalytic mechanism of glycine radical 1,2-lyase in the literature, the key amino acid residues in the active center of 1,5-anhydro-D-mannitol-6-phosphate isomerase were determined.
[0435] Isozyme identification of 1,5-anhydro-D-mannitol-6-phosphate isomerase
[0436] To identify the isozymes of 1,5-anhydro-D-mannitol-6-phosphate isomerase in the glycine radical enzyme (GRE) family, a sequence similarity network (SSN) of 25,347 unique sequences in the InterPro family IPR004184 was constructed using the web-based Enzyme Function Initiative Enzyme Similarity Tool (EFI-EST). The alignment score threshold was set to 250, the minimum sequence length was set to 650, and the 80% similarity representative node network (RepNode network) was displayed using Cytoscape v3.519. According to these set values, previously characterized GREs with different catalytic activities were divided into different clusters.
[0437] Sequences with 80% or more similarity are represented by a single node. The sequences are folded together to reduce the total number of nodes, making the less complex network easier to load in Cytoscape. For the clusters formed by PflD, a representative sequence is selected from each node according to different nodes. Based on the crystal structure of the complex formed by SdPflD and the substrate, the amino acid residues interacting with the substrate are taken as conserved sites for multiple sequence alignment. Sequences containing these conserved sites simultaneously are retained. Further analysis is carried out on the neighboring genes of these sequences, and sequences containing aldolase in the gene cluster are retained. Finally, representative sequences are selected.
[0438] Results, Analysis and Discussion
[0439] There is a gene cluster involved in sugar metabolism in the genome of Escherichia coli MG1655. Through bioinformatics analysis of this gene cluster, the inventors of the present application noticed that an unknown function GRE (PflD) is involved. In addition to PflD and the S-adenosylmethionine radical enzyme family glycine radical enzyme activator (PflC), this gene cluster also contains a 1-deoxyfructose-6-phosphate aldolase (FsaB), a hydroxyacetone reductase (GldA), and there is also a complete phosphoenolpyruvate-dependent phosphotransferase system (PTS), which includes PtsA, FrwB, FrwC, and FrwD. The function of PTS is to transfer the phosphate group of phosphoenolpyruvate (PEP) to the sugar entering the cell through a series of sequential steps involving different components of PTS. PTS consists of cytoplasmic components and membrane-associated enzymes. The cytoplasmic components lack sugar specificity, and the membrane-associated enzymes are specific for at most a few sugars. This PTS transport complex may be involved in the uptake of hexoses and the phosphorylation of substrates. Aldolase FsaB can catalyze the formation of the corresponding 2-keto-hexose-6-phosphate from hydroxyacetone, 1,3-dihydroxyacetone, and glyceraldehyde 3-phosphate. GldA catalyzes the dehydrogenation of glycerol or propylene glycol to produce 1,3-dihydroxyacetone or hydroxyacetone. Based on this, the inventors of the present application speculated that this GRE is involved in sugar metabolism, and its substrate may be a hexose phosphate. Combining the results of molecular docking, the inventors of the present application speculated that the substrate of PflD is 1,5-anhydro-D-mannitol-6-phosphate (1,5-AM-6P), and proposed a new radical-dependent glycolysis pathway. In this pathway, 1,5-AM-6P is catalyzed by PflD to break the C-O bond, generating 1-deoxy-F6P. This product is further cleaved by aldolase FsaB to produce hydroxyacetone and glyceraldehyde 3-phosphate. Hydroxyacetone is reduced by hydroxyacetone reductase GldA to produce 1,2-propanediol ( Figure 1 B).
[0440] To prove the conjecture of the inventors of the present application, the inventors of the present application selected the PflD, PflC, FsaB and GldA genes from Escherichia coli MG1655, then carried out heterologous expression in Escherichia coli BL21(DE3) cells, and characterized the activities of these proteins.
[0441] The inventors of the present application reconstructed the cofactor [Fe-S] 2+ cluster of the activating enzyme EcPflC and characterized it. The anaerobic reconstruction of the [4Fe-4S] cluster resulted in 2.15 ± 0.12 Fe and 2.30 ± 0.07 S per monomer (a radical SAM domain, the theoretical maximum of the [4Fe-4S] cluster is 4 Fe and 4 S) and a typical UV-visible spectrum of the [4Fe-4S] protein-containing, with a molar extinction coefficient ε 410nm of 9.45 mM -1 cm -1 . The ε 410nm of each [4Fe-4S] cluster is approximately 15 mM -1 cm -1 , so the inventors of the present application estimated that each monomer contains approximately 0.63 [4Fe-4S], which is roughly consistent with the measured Fe and S contents ( Figure 8 A-8B). The LC-MS detection results showed that, like other radical SAM enzymes, EcPflC catalyzed the cleavage of SAM in the presence of the reducing agent titanium(III) citrate to form 5'-deoxyadenosine ( Figure 8 C-8D). EPR spectroscopy showed that incubating EcPflD, EcPflC, SAM and Ti(III) formed radicals, and the radicals generated were quantified as 0.060 Gly· / dimer ( Figure 3 A).
[0442] To determine the activity of EcPflD, activated EcPflD was incubated with 1,5-AM-6P and analyzed by LC-MS. It can be seen that there is a new peak with m / z(-) = 243.0 in the full reaction group at t R = 18.40 min ( Figure 3 B-3D). In the activity experiment of EcFsaB catalyzing the aldol condensation reaction of hydroxyacetone and glyceraldehyde 3-phosphate, a new peak (m / z(-) = 243.0, t R = 18.40 min) was also generated in the full reaction group ( Figure 9 B-9D). The Michaelis-Menten kinetic parameters of EcPflD were also determined ( Figure 11 , k cat = 33.47 ± 0.79 s -1 / EcPflD, K M= 1.33 ± 0.12 mM). The reaction was carried out by coupling EcFsaB with EcPflD, and the production of the corresponding products hydroxyacetone and glyceraldehyde 3-phosphate was detected by LC-MS analysis ( Figure 3 E-3H). The spectrophotometric determination coupled with EcPflD, EcFsaB and EcGldA showed that only in the full reaction group, the decrease of A 340nm indicated that the substrate 1,5-AM-6P was catalyzed by EcPflD to break the C-O bond, generating 1-deoxy-F6P, and the product was cleaved by EcFsaB to produce hydroxyacetone and glyceraldehyde 3-phosphate; hydroxyacetone was reduced by EcGldA to generate 1,2-propanediol Figure 3 I).
[0443] The PflD gene cluster in Streptococcus dysgalactiae subsp. Equisimilis is as Figure 19 shown in A. According to the LC-MS analysis results, it can be seen that there is a new peak with m / z(-) = 243.0 in the full reaction group at t R = 18.40 min ( Figure 19 B), and the mass spectrum corresponding to this new peak is Figure 19 C, and the mass spectrum corresponding to the substrate is Figure 19 D. The LC-MS results indicate that SdPflD has 1,5-anhydro-D-mannitol-6-phosphate isomerase activity.
[0444] To further study the catalytic mechanism of PflD, the inventors of the present application determined the crystal structure of the complex of SdPflD with 1,5-AM-6P (resolution ). Each asymmetric unit of PflD contains a monomer. Each monomer exhibits the typical β / α barrel fold common to other GREs Figure 4 B), as well as a radical-like conformation with Gly· and Cys· loops. From this crystal structure, it can be seen that the S atom of the Cys· residue Cys431 is adjacent to the H atom at the C2 position of the substrate, with a distance of This is consistent with the catalytic mechanism involving "the C2 hydrogen atom is abstracted by Cys·". The 2-OH group of the substrate 1,5-AM-6P forms a hydrogen bond with Glu433; the phosphate group coordinates with Gln162, Ser277 and Arg323 of PflD; the 1-O atom forms a hydrogen bond with His169, Glu445, and the 3-OH forms a hydrogen bond with S278 Figure 4 D). Glu433 is involved in the deprotonation of the 2-OH of the substrate, which is consistent with the role of the base in the catalytic mechanism of GRE 1,2-elimination enzymes.
[0445] Based on the crystal structure of PflD, which is similar to other GRE 1,2-elimination enzymes, the inventors of the present application proposed the following PflD catalytic mechanism. Cys431 abstracts the H atom at the C2 position from the substrate 1,5-AM-6P, generating a substrate radical at the C2 position. Then, Glu433 deprotonates the 2-OH of the substrate, and the substrate radical translocates to the C1 position, followed by cleavage of the C-O bond to form a product radical. The product radical then abstracts the H atom that Cys431 abstracted from the substrate, generating the product 1-deoxy-F6P and regenerating Cys431·( Figure 4 F).
[0446] Based on the crystal structure of the complex formed by SdPflD and the substrate 1,5-anhydro-D-mannitol-6-phosphate, referring to the hydrogen bond distances formed by the substrate and amino acid residues, the key amino acid residues that bind to the substrate in the active center of 1,5-anhydro-D-mannitol-6-phosphate isomerase were determined. Combining with the catalytic mechanism of glycine radical 1,2-lyase in the literature, the key amino acid residues in the active center of 1,5-anhydro-D-mannitol-6-phosphate isomerase were determined. They are the amino acid residues Q162, H169, S277, S278, R323, F331, P335, C431, E433, D445, Y628, V630, and G752 in SdPflD.
[0447] According to the SSN sequence analysis, the PflD cluster showed a total of 89 nodes and 953 consensus sequences. The inventors of the present application selected a representative sequence from each node. Based on the crystal structure of SdPflD (NCBI accession number OCX05109.1), the amino acid residues that interact with the substrate (Q162, H169, S277, S278, R323, F331, P335, C431, E433, D445, Y628, V630, and G752) were used as conserved sites for multiple sequence alignment. Sequences that simultaneously contained these conserved sites were retained, and a total of 77 isoenzyme sequences were screened out. Further analysis of the neighboring genes of these 77 sequences was carried out, and sequences containing aldolase FsaB in the gene cluster were retained. Finally, a total of 47 representative sequences were listed( Figure 14 and Figure 16 ).
[0448] Example 3. Experiments on the utilization of 1,5-anhydroglucitol and 1,5-anhydro-D-mannitol by Escherichia coli
[0449] Gene knockout of Escherichia coli MG1655
[0450] The ybiW and pflD in Escherichia coli MG1655 were knocked out respectively using the CRISPR-assisted homologous recombination strategy. The inventors of this application designed primer pairs 7F / 7R and 8F / 8R to amplify the upstream and downstream sequences of ybiW, 500 bp at the 5' end and 504 bp at the 3' end, to generate fragments 1 and 2. Fragments 1 and 2 were assembled by PCR using primers 7F and 8R to generate fragment 3. Using the pRed_Cas9_recA plasmid as a template, fragments 4 and 5 were obtained by PCR amplification using primer pairs 9F / 9R and 10F / 10R. Fragments 3, 4, and 5 were subjected to Gibson assembly to generate the recombinant pRed_Cas9_recA_ΔybiW plasmid, which contains the guide RNA sequence "GCCTGCCAGAAAGTCTGCGG" (SEQ ID NO:187), aiming to introduce Cas9-catalyzed cleavage near the homologous recombination site to increase the rate of the desired homologous recombination.
[0451] The construction method of the recombinant plasmid pRed_Cas9_recA_ΔpflD was similar to that of pRed_Cas9_recA_ΔybiW. Primer pairs 11F / 11R and 12F / 12R were used to amplify the upstream (500 bp) and downstream (505 bp) sequences of pflD to generate fragments 6 and 7. Similarly, using fragments 6 and 7 as templates and 11F and 12R as primers, fragment 8 was generated. Using primer pairs 9F / 13R and 13F / 10R, fragments 9 and 10 were amplified by PCR using pRed_Cas9_recA as a template. The pRed_Cas9_recA_ΔpflD containing the guide RNA sequence "AAATACCAGAACCCGCGCGG" (SEQ ID NO:188) was constructed by Gibson assembly using fragments 8, 9, and 10. DNA sequencing was performed using primers 14F, 14R, and 15F to confirm the recombinant plasmids pRed_Cas 9_recA_ΔybiW and pRed_Casa9_recA_ΔpflD.
[0452] To obtain Escherichia coli MG1655_ΔybiW and Escherichia coli MG16.55_ΔpflD strains, pRed_Cas9_recA_ΔybiW and pRed_Cas9_recA_ΔpflD were transformed into competent cells of Escherichia coli MG1655 strain by electroporation, respectively. Positive transformants were selected on LB agar plates containing 50 μg / mL kanamycin in a 30 °C incubator. Single colonies were cultured in 5 mL LB culture containing 50 mg / L kanamycin and 2 g / L arabinose at 30 °C for 24 hours. The cells were harvested and plated on LB agar plates containing 50 g / L kanamycin and 2 g / L arabinose at 30 °C. Single colonies were picked and streaked on LB agar plates and cultured at 37 °C to eliminate the recombinant plasmid. Colony PCR verification was performed using primers 16F / 16R for ΔybiW and primers 17F / 17R for ΔpflD, and a DNA fragment of approximately 1500 bp was amplified theoretically. Agarose gel electrophoresis showed successful genomic knockout, and this result was further confirmed by DNA sequencing using the same primer pairs.
[0453] Anaerobic growth of Escherichia coli strains on 1,5-AG / 1,5-AM as the sole carbon source
[0454] Single colonies of Escherichia coli MG1655 WT, ΔybiW, and ΔpflD strains freshly grown on LB agar plates were inoculated into 5 mL LB medium and cultured in a shaker at 37 °C for 4 hours. The cells in 100 μL of the culture were centrifuged and transferred to an anaerobic bottle containing 5 mL anaerobic LB medium and cultured at 37 °C for 6 hours. The cells in 100 μL of the culture were collected by centrifugation and transferred to an anaerobic vial containing 5 mL anaerobic M9 medium and cultured at 37 °C for 3 days. 1 mL of cells was harvested, washed three times with anaerobic M9 medium without glucose, and resuspended in this medium. 100 μL of the cell suspension was transferred to an anaerobic bottle, and each anaerobic bottle contained 5 mL M9 medium without glucose or the same medium supplemented with 20 mg of sugar (glucose, 1,5-AG, or 1,5-AM). Then the anaerobic bottles were placed in a 37 °C incubator, photographed after 7 days, and the cells were collected for SDS-PAGE gel analysis.
[0455] Identification of proteins by SDS / PAGE and mass spectrometry
[0456] Cells were harvested by centrifugation, lysed by boiling in Laemmli loading buffer, and analyzed on a 10% SDS / PAGE gel. Prominent protein bands induced in cells cultured with glucose, 1,5-anhydroglucitol, and 1,5-anhydromannitol were excised manually. After in-gel digestion and extraction, the peptide mixture was analyzed by a Fusion Lumos mass spectrometer coupled with an Easy nLC 1200 system (Thermo Fisher Scientific). The MS / MS spectra of each LC-MS / MS run were searched against the Escherichia coli protein database (released on April 1, 2021), which contains 15,862 sequence entries from UniProt, using an in-house Proteome Discoverer (version 2.2) search algorithm. Protein identification was based on Sequest HT.
[0457] GC analysis of the fermentation broth
[0458] Wild-type Escherichia coli MG1655 was cultured at 37 °C for 7 days in different carbon sources, and the cells were removed by centrifugation. 200 μL of the fermentation broth was taken, 800 μL of chromatographically pure anhydrous ethanol was added, vortexed thoroughly, and the insoluble precipitate was removed by centrifugation. Then it was filtered through a 0.45 μm organic filter membrane, and the resulting filtrate was analyzed by gas chromatography. The (R)-1,2-propanediol commercial standard was dissolved in chromatographically pure anhydrous ethanol.
[0459] Gas chromatography (GC) analysis was performed using an Agilent 6820 Series G1176A gas chromatograph (Agilent Technologies). The chromatographic column used for GC analysis was an AT TM -Aquawax-DA (Alltech) gas chromatographic column (30 m × 0.53 mm, 1.0 μm). The GC conditions were as follows: the carrier gas was high-purity nitrogen, the column flow rate was a constant 1.0 mL / min; the inlet temperature was 230 °C, the detector temperature was 240 °C; the hydrogen flow rate was 20 mL / min; the air flow rate was 200 mL / min. Programmed temperature elevation was used: the initial column temperature was 60 °C, held for 2 minutes, heated at a rate of 20 °C / min to 80 °C, held for 3 minutes, heated at a rate of 20 °C / min to 160 °C, held for 2 minutes; then heated at a rate of 15 °C / min to 220 °C and held for 10 minutes.
[0460] Results, analysis, and discussion
[0461] Escherichia coli MG1655_wild type and Escherichia coli MG1655_ΔybiW strain were anaerobically cultured using M9 medium with 1,5-AG as the sole carbon source; the positive control used glucose as the sole carbon source, and the negative control did not add any carbon source. The Escherichia coli MG1655_wild type strain could grow when glucose and 1,5-AG were used as carbon sources respectively; while the Escherichia coli MG1655_ΔybiW strain could only grow using glucose as the sole carbon source ( Figure 5 A). Further SDS-PAGE analysis results showed that when the Escherichia coli MG1655_wild type strain used 1,5-AG as the sole carbon source, the corresponding proteins in the YbiW-dependent pathway were induced to express. Combining with proteomic analysis, the induced bands were respectively identified. The two bands of ~95 kDa both contained PtsA and YbiW; the band of ~42 kDa was identified as GldA, and the band of ~27 kDa contained FsaB and FasA( Figure 5 B).
[0462] Similarly, Escherichia coli MG1655_wild type and Escherichia coli MG1655_ΔpflD strain were anaerobically cultured using M9 medium with 1,5-AM as the sole carbon source; the positive control used glucose as the sole carbon source, and the negative control did not add any carbon source. The Escherichia coli MG1655_wild type strain could grow when glucose and 1,5-AM were used as carbon sources respectively; while the Escherichia coli MG1655_ΔpflD strain could only grow using glucose as the sole carbon source ( Figure 5 C). Further SDS-PAGE analysis results showed that when the Escherichia coli MG1655_wild type strain used 1,5-AM as the sole carbon source, the corresponding proteins in the PflD-dependent pathway were induced to express. Combining with proteomic analysis, the induced bands were respectively identified. The band of ~95 kDa contained PtsA; the band of ~90 kDa contained PflD; the band of ~42 kDa was identified as GldA, and the band of ~27 kDa contained FsaB( Figure 5 D). GC result analysis showed that when 1,5-AG and 1,5-AM were used as the sole carbon sources respectively, Escherichia coli MG1655 could produce 1,2-propanediol under anaerobic conditions, and its retention time was 8.85 min( Figure 13 ).
[0463] There are YbiW and PflD gene clusters in the Escherichia coli MG1655 genome( Figure 5E). Combining the growth results of the strains on different carbon sources, SDS-PAGE analysis, and proteomic analysis results, the inventors of the present application believe that 1,5-AG can induce the expression of the YbiW and PflD-related gene clusters in Escherichia coli MG1655. 1,5-AG enters the cell using the PTS transport system in the PflD gene cluster, induces the high expression of FsaB and GldA in the PflD gene cluster, and simultaneously induces the high expression of YbiW and FsaA in the YbiW gene cluster. 1,5-AG is transported into the cell by the PTS transport system, and the C-6 position of the substrate is phosphorylated to produce 1,5-AG-6P; under the catalysis of YbiW, the C-O bond is cleaved to produce 1-deoxy-F6P; this product is cleaved by FsaB and FsaA to produce hydroxyacetone and glyceraldehyde 3-phosphate; hydroxyacetone is reduced by EcGldA to generate 1,2-propanediol.
[0464] However, 1,5-AM can only induce the high expression of the proteins related to the PflD pathway. The substrate is transported into the cell by the PTS transport system in the gene cluster, the C-6 position is phosphorylated to produce 1,5-AM-6P, and under the catalysis of PflD, the C-O bond is cleaved to produce 1-deoxy-F6P; this product is cleaved by FsaB to produce hydroxyacetone and glyceraldehyde 3-phosphate; hydroxyacetone is reduced by GldA to generate 1,2-propanediol ( Figure 5 F).
[0465] Table 1. Data collection and refinement statistics for EcYbiW and SdPflD crystals
[0466]
[0467]
[0468] Note: Statistics for the highest resolution shell are shown in parentheses
[0469] Table 2. Accession numbers, strain sources, and amino acid sequence numbers of 1,5-anhydroglucitol-6-phosphate isomerase isoenzymes
[0470]
[0471]
[0472]
[0473]
[0474]
[0475] Table 3. Primers for plasmid construction
[0476]
[0477]
[0478] Note: * represents the genome of Escherichia coli MG1655
[0479] Table 4. Accession numbers, strain sources, and amino acid sequences of 1,5-anhydro-D-mannitol-6-phosphate isomerase isozymes
[0480]
[0481]
[0482] Table 5. Primers used for constructing Escherichia coli ΔybiW and ΔpflD strains
[0483]
[0484]
[0485] Note: represents the genome of Escherichia coli MG1655WT; # represents the genome of ΔybiW; * represents the genome of ΔPflD.
[0486] Table 6. Primers used for plasmid construction
[0487]
[0488] Note: * represents the genome of Escherichia coli MG1655
[0489] The above describes exemplary embodiments of the various inventions of the present application. However, without departing from the essence and scope of the present application, those skilled in the art can modify or improve the exemplary embodiments described in the present application, and the resulting variant or equivalent embodiments also fall within the scope of the present application.
Claims
1. A polypeptide comprising the amino acid sequence shown in SEQ ID NO: 1 or 97, or a functional variant thereof, wherein the functional variant has 1,5-anhydroglucitol-6-phosphate isomerase activity.
2. The polypeptide according to claim 1, wherein the polypeptide comprises the amino acid sequence shown in SEQ ID NO: 1 or a functional variant thereof, wherein the functional variant has 1,5-anhydroglucose-6-phosphate isomerase activity; Preferably, the polypeptide has an active site defined in terms of its spatial conformation: the active site comprises amino acid residues H165, H282, S283, H334, C441, E443, R453, T455, L562, S662, I664 and G786 that are close to each other in spatial conformation with reference to SEQ ID NO:
1.
3. The polypeptide according to claim 2, wherein the functional variant is a natural isoenzyme of the amino acid sequence shown in SEQ ID NO: 1; Preferably, the natural isoenzyme is derived from: Lactiplantibacillus plantarum, Escherichia coli, Lactobacillus selangorensis, Streptococcus uberis, Olsenella sp., Leptotrichia wadei, Streptococcus parauberis, Anaerostipes hadrus, Senella sp., Selenomonas ruminantium, Orenia metallireducens, Bifidobacterium primatium, Leptotrichia hofstadii, Ligilactobacillus agilis, Senella profusa, Clostridium butyricum, Vibrio ishigakensis, Coriobacteriaceae, Coriobacterium glomerans, Eggerthia catenaformis, Clostridium baratii str. Sullivan, Clostridiales, Firmicutes, Clostridium vincentii, Streptococcus downei, Erysipelotrichaceae, Quinella sp., Sebaldella termitidis, Ligilactobacillus animalis, Liquorilactobacillus satumensis, Clostridium botulinum, Acetivibrio ethanolgignens, Aeromonas sobria, Clostridium cadaveris, Streptococcus bovimastitidis, Sodalis ligni, Lactococcus raffinolactis, Lactobacillus ultunensis, Leptotrichia sp., Lucifera butyrica, Fonticella tunisiensis, Cronobacter sakazakii, Liquorilactobacillus uvarum, Anaeromassilibacillus sp., Anaerotruncus sp.) Beauveria bassiana, Cedecea lapagei, Streptococcus pneumoniae, Streptococcus suis, Aeromonas veronii, Lactococcus raffinolactis, Streptococcus sanguinis, Actinomyces succiniciruminis, Vagococcus humatus, Vibrio alginolyticus, Ligilactobacillus acidipiscis, Pilibacter termitis, Lachnospiraceae, Aeromonas hydrophila, Enterococcus ratti, Kandleria vitulina, Clostridium pasteurianum, Aeromonas allosaccharophila, Pantoea alhagi, Peptacetobacter hiranonis, Edwardsiella piscicida, Cedecea neteri, Chelonobacter oris, Aeromonas sp., Clostridium sp., Vibrio gazogenes, Melissococcus plutonius, Tolumonas sp., Loigolactobacillus coryniformis, Zooshikella ganghwensis, Vibrio sp., Photobacterium lipolyticum, Agrilactobacillus composti, Clostridium symbiosum, and Aerococcus sp.;. More preferably, the natural isoenzyme comprises the amino acid sequence shown in any one of SEQ ID NOs: 2-96.
4. The polypeptide according to claim 1, wherein the polypeptide comprises the amino acid sequence shown in SEQ ID NO: 97 or a functional variant thereof, wherein the functional variant has 1,5-anhydromannitol-6-phosphate isomerase activity; Preferably, the polypeptide has an active site defined in terms of its spatial conformation: the active site comprises amino acid residues Q162, H169, S277, S278, R323, F331, P335, C431, E433, D445, Y628, V630 and G752 that are close to each other in spatial conformation with reference to SEQ ID NO:
97.
5. The polypeptide according to claim 4, wherein the functional variant is a natural isoenzyme of the amino acid sequence shown in SEQ ID NO: 97; Preferably, the natural isoenzyme is derived from: Escherichia coli, Streptococcus dysgalactiae subsp. Equisimilis, Firmicutes, Pseudoleptotrichia goodfellowii, Edwardsiella tarda, Serratia odorifera, Gilliamella apicola, Vibrio, Lactobacillus selangorensis, Caloranaerobacter sp., Pilibacter termitis, Caloranaerobacter azorensis, Salmonella enteritidis, Tolumonas auensis, Senella porci, Mycobacteriaceae, Fonticella tunisiensis, Coriobacteriales, Bacillus sp., Actinobacillus rossii, Lactobacillus ultunensis, Clostridium estertheticum, Geosporobacter ferrireducens, Vibrio ishigakensis, Streptococcus merionis, Shigella boydii, Agrilactobacillus composti, Lactobacillus selangorensis, Vagococcus elongatus, Clostridium acidisoli, Caloramator quimbayensis, Clostridium uliginosum, Enterobacter asburiae, Atopobium minutum, Thermophilibacter immobilis, Photobacterium lipolyticum, Citrobacter amalonaticus, Clostridium pasteurianum, Sebaldella termitis, Acididesulfobacillus acetoxydans, and Clostridium; More preferably, the natural isoenzyme comprises the amino acid sequence shown in any one of SEQ ID NO: 98-144.
6. The polypeptide according to any one of claims 1-5, wherein its substrate is 1,5-anhydrohexitol-6-phosphate; preferably, the 1,5-anhydrohexitol-6-phosphate is 1,5-anhydroglucitol-6-phosphate or 1,5-anhydromannitol-6-phosphate.
7. The polypeptide according to any one of claims 1-6, wherein the functional variant is generated by one or more amino acid insertions, substitutions, and / or deletions based on the amino acid sequence shown in SEQ ID NO: 1 or 97 or its natural isozyme. Optionally, the insertions, substitutions, and / or deletions do not occur in the active site.
8. A nucleic acid molecule encoding the polypeptide according to any one of claims 1-7.
9. An expression cassette comprising the nucleic acid molecule according to claim 8.
10. An expression vector comprising the nucleic acid molecule according to claim 8, or the expression cassette according to claim 9.
11. A host cell comprising the nucleic acid molecule according to claim 8, the expression cassette according to claim 9, or the expression vector according to claim 10.
12. The host cell according to claim 11, which further expresses: an activating enzyme, an aldolase, a hydroxyacetone reductase, and / or a transport complex; Preferably, the host cell further expresses an S-adenosylmethionine radical enzyme family glycine radical enzyme activating enzyme, a 1-deoxyfructose-6-phosphate aldolase, a hydroxyacetone reductase, and / or the transport complex of the polypeptide according to any one of claims 1-7.
13. The host cell according to claim 12, wherein the transport complex has phosphoenolpyruvate-phosphotransferase system transport activity; preferably, the transport complex is a phosphoenolpyruvate-phosphotransferase system; more preferably, the transport complex comprises at least one of SEQ ID NOs: 150-153 or its functional variant.
14. The host cell according to any one of claims 11-13, which is a eukaryotic cell or a prokaryotic cell; Preferably, the eukaryotic cell is a yeast cell; and / or Preferably, the prokaryotic cell is selected from: Escherichia, Klebsiella, Streptococcus, Lactobacillus, Bifidobacterium, Bacteroidetes, and Firmicutes, etc.; more preferably, the Escherichia is Escherichia coli.
15. A composition comprising the polypeptide according to any one of claims 1-7; Preferably, the composition further comprises: an activating enzyme, an aldolase, and / or a hydroxyacetone reductase; More preferably, the composition further comprises an S-adenosylmethionine radical enzyme family glycine radical enzyme activating enzyme, a 1-deoxyfructose-6-phosphate aldolase, and / or a hydroxyacetone reductase of the polypeptide according to any one of claims 1-7.
16. The composition according to claim 15, which is used for catalyzing the production of 1,2-propanediol from 1,5-anhydrohexitol-6-phosphate; preferably, the 1,5-anhydrohexitol-6-phosphate is 1,5-anhydroglucitol-6-phosphate or 1,5-anhydromannitol-6-phosphate.
17. Use of the polypeptide according to any one of claims 1-7, the nucleic acid molecule according to claim 8, the expression cassette according to claim 9, the expression vector according to claim 10, and the host cell according to any one of claims 11-14 in the preparation of a composition for catalyzing the formation of 1,2-propanediol from 1,5-anhydroglycitol-6-phosphate; preferably, the 1,5-anhydroglycitol-6-phosphate is 1,5-anhydroglucitol-6-phosphate or 1,5-anhydromannitol-6-phosphate.
18. A method for producing 1,2-propanediol, which comprises culturing the host cell according to any one of claims 11-14 in a medium containing 1,5-anhydroglycitol to obtain 1,2-propanediol; Preferably, the 1,5-anhydroglycitol is 1,5-anhydroglucitol or 1,5-anhydromannitol.
19. A method for producing 1,2-propanediol, which comprises contacting the composition according to claim 15 or 16 with 1,5-anhydroglycitol-6-phosphate to obtain 1,2-propanediol; Preferably, the 1,5-anhydroglycitol-6-phosphate is 1,5-anhydroglucitol-6-phosphate or 1,5-anhydromannitol-6-phosphate.