Metabolic engineering for co-consumption of xylose and glucose to produce chemicals from second-generation sugars
By engineering the E. coli, the problem of low microbial productivity when multiple carbon sources exist is solved, and the simultaneous utilization and efficient conversion of xylose and glucose are achieved, thereby improving productivity.
Patent Information
- Application Number
- CN202080041307.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-04
- Filing Date
- 2020-04-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-04-06
AI Technical Summary
The prior art is difficult to maximize the microbial productivity of producing desired chemicals from renewable feedstocks, especially in the presence of multiple carbon sources, which are susceptible to catabolic repression effects.
By engineering E. coli, deletion or inactivate pentose ATP-binding transporter, expressing C5 sugar homologous transporter, encoding xylose isomerase and xylulose kinase, and deletion or inactivate xylose isomerase and xylulose kinase, to achieve simultaneous utilization of xylose and glucose.
Simultaneous utilization of xylose and glucose is achieved, reducing or eliminating the repressive activity of carbon sources on microorganisms, and improving the productivity of fermentation products from renewable raw materials.
Smart Images

Figure CN114008197B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Application No. 62 / 829,398, filed on April 4, 2019, the content of which is incorporated herein by reference in its entirety.
[0003] Statement Regarding the Sequence Listing
[0004] The sequence listing associated with this application is provided in text format, in lieu of a paper copy, and is hereby incorporated by reference into the specification. The name of the text file containing the sequence listing is BRSK_021_01WO_ST25.txt. The text file is 222 kb, was created on April 4, 2020, and is submitted in electronic form. Background of the Invention
[0005] Producing desired chemicals (such as monoethylene glycol, glycolic acid, C3 compounds (such as acetone, isopropanol, and propylene), amino acids, and polyols) from alternative feedstocks (such as pentoses) is an alternative to obtaining such chemicals from petro - based chemicals.
[0006] The utilization of xylose (a pentose) is a source distinct from most renewable chemical schemes. Considering that using lignocellulosic biomass as a feedstock does not require using plants that would otherwise produce food, lignocellulosic biomass is a promising renewable feedstock. Due to the sustainability and global availability of lignocellulosic biomass, it is even more promising as a renewable feedstock. The separation and isolation of lignocellulosic sugars is an option to increase sugar production without increasing land use. Xylose is the main carbon source in hemicellulose hydrolysates, followed by glucose and arabinose. Xylose typically accounts for 70 - 80% of the sugars present in hemicellulose hydrolysates, while glucose accounts for 10 - 20%.
[0007] In Escherichia coli, even the smallest amount of glucose completely inhibits xylose uptake (even when xylose is the main sugar), thus limiting the overall conversion of sugar to the desired chemical. In industrial processes, productivity (grams of product per liter per hour) is a key factor in ensuring economic viability, and thus microbial strains must be able to continuously convert the main substrate (such as xylose) to product at maximum rate. Considering that glucose is present in lignocellulose hydrolysates, in batch operation, xylose uptake will be delayed until glucose is completely depleted, thus reducing productivity. If we consider fed - batch or continuous operation, a stream containing xylose and glucose is continuously fed into the reactor, enhancing the repressive potential of glucose, and thus the potential productivity of microorganisms capable of producing one or more products from common renewable feedstocks cannot be maximized.
[0008] It is necessary to maximize the productivity of microorganisms capable of producing the desired products from renewable raw materials. In particular, it is necessary to maximize the utilization of multiple carbon sources while minimizing or eliminating the repressive activity of carbon sources on microorganisms.
[0009] As described herein, the present disclosure provides methods and compositions for engineering microorganisms to utilize a mixed sugar stream to produce the desired chemicals without regard to the typical catabolic repression effects caused by the presence of the mixed sugar stream. Summary of the Invention
[0010] In some aspects, the present disclosure generally relates to recombinant microorganisms capable of producing a fermentation product from a feedstock comprising xylose and glucose, wherein the recombinant microorganism simultaneously utilizes xylose and glucose, and wherein the microorganism comprises one or more of the following: (a) deletion or inactivation of a pentose ATP-binding transporter in the genome of the microorganism such that the transporter is not expressed; (b) one or more endogenous or exogenous nucleic acid sequences that encode at least one C5 sugar symporter and are operably linked to one or more constitutive promoters; wherein the C5 sugar symporter comprises: (1) a xylose symporter and / or (2) an arabinose symporter; (c) one or more endogenous or exogenous nucleic acid sequences that (1) encode xylose isomerase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more xylulokinases, and / or (2) encode xylose dehydrogenase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases.
[0011] In some aspects, the fermentation product produced by the microorganism is one or more molecules containing 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbons. In some aspects, two or more molecules are produced simultaneously.
[0012] In some aspects, the present disclosure generally relates to recombinant Escherichia coli capable of producing a fermentation product from a feedstock comprising xylose and glucose, wherein the recombinant microorganism co-utilizes xylose and glucose, and wherein the microorganism comprises one or more of the following: (a) deletion or inactivation of the ATP-binding transporters araFGH and xylFGH in the genome of the microorganism such that the transporters are not expressed; (b) one or more endogenous or exogenous nucleic acid sequences encoding at least one C5 sugar symporter and operably linked to one or more constitutive promoters; wherein the C5 sugar symporter comprises: (1) a xylose symporter and / or (2) an arabinose symporter; (c) one or more endogenous or exogenous nucleic acid sequences that (1) encode xylose isomerase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more xylulokinases, and / or (2) encode xylose dehydrogenase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases.
[0013] In some aspects, the present disclosure generally relates to recombinant microorganisms capable of producing monoethylene glycol (MEG) and / or acetone from a feedstock comprising xylose and glucose, wherein the recombinant microorganisms co-utilize xylose and glucose, and comprise one or more of the following: (a) deletion or inactivation of aldA, araFGH, and xylFGH in the genome of the parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganisms express an MEG and / or acetone production pathway.
[0014] In some aspects, the recombinant microorganism of claim 5, wherein the microorganism further comprises deletion or inactivation of glcDEF. In some aspects, the C5 symporter is controlled by the GAPDH promoter at the araFGH locus. In some aspects, the C5 sugar symporter is the xylose symporter XylE. In some aspects, the XylE comprises the amino acid sequence set forth in SEQ ID NO:49. In some aspects, the XylE is encoded by a nucleic acid sequence comprising SEQ ID NO:48. In some aspects, the xylose symporter is endogenous to the microorganism.
[0015] In some aspects, the C5 sugar symporter is the arabinose symporter AraE. In some aspects, the AraE comprises the amino acid sequence set forth in SEQ ID NO:47. In some aspects, the AraE is encoded by a nucleic acid sequence comprising SEQ ID NO:46. In some aspects, the arabinose symporter is endogenous to the microorganism.
[0016] In some aspects, the uptake of xylose is insensitive to catabolite repression of other monosaccharides. In some aspects, the microorganism comprises a functional phosphotransferase system. In some aspects, the microorganism comprises a native wild-type nucleic acid sequence encoding a cAMP receptor protein (CRP). In some aspects, the CRP comprises an amino acid sequence comprising SEQ ID NO:10. In some aspects, the CRP is encoded by a nucleic acid sequence comprising SEQ ID NO:9.
[0017] In some aspects, constitutive overexpression of the xylose symporter enables continuous input of xylose from the feedstock into the microorganism. In some aspects, constitutive overexpression of the arabinose symporter enables continuous input of xylose from the feedstock into the microorganism. In some aspects, continuous xylose input occurs independently of the presence of other sugars in the feedstock.
[0018] In some aspects, the recombinant microorganism comprises a MEG production pathway having one or more of the following (c) to (e); (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose isomerase and / or ketohexokinase and / or fructose bisphosphate aldolase, the nucleic acid sequences being operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding hydroxyacetaldehyde reductase, the hydroxyacetaldehyde reductase catalyzing the conversion of hydroxyacetaldehyde to MEG; and (e) deletion or inactivation of one or more xylulokinases in the genome of the parental microorganism. In some aspects, (c) and (d) are in an operon controlled by the proD promoter. In some aspects, the proD promoter is encoded by a nucleic acid sequence comprising SEQ ID NO:53.
[0019] In some aspects, the xylose isomerase is XylA. In some aspects, the xylose isomerase is endogenous to the microorganism. In some aspects, the XylA comprises an amino acid sequence comprising SEQ ID NO:6. In some aspects, the XylA is encoded by a nucleic acid sequence comprising SEQ ID NO:5.
[0020] In some aspects, the ketohexokinase is from Homo Sapiens. In some aspects, the ketohexokinase is heterologous to the microorganism. In some aspects, the ketohexokinase is khk-C. In some aspects, the khk-C comprises an amino acid sequence comprising SEQ ID NO:12. In some aspects, the khk-C is encoded by a nucleic acid sequence comprising SEQ ID NO:11.
[0021] In some aspects, the fructose-bisphosphate aldolase is from Homo sapiens. In some aspects, the fructose-bisphosphate aldolase is heterologous to the microorganism. In some aspects, the fructose-bisphosphate aldolase is aldoB. In some aspects, the aldoB comprises the amino acid sequence comprising SEQ ID NO:51. In some aspects, the aldoB is encoded by the nucleic acid sequence comprising SEQ ID NO:50.
[0022] In some aspects, the glyoxaldehyde reductase is endogenous to the microorganism. In some aspects, the glyoxaldehyde reductase is fucO. In some aspects, the fucO comprises the amino acid sequence comprising SEQ ID NO:98. In some aspects, the fucO is encoded by the nucleic acid sequence comprising SEQ ID NO:52.
[0023] In some aspects, the xylulokinase is XylB. In some aspects, the xylB comprises the amino acid sequence comprising SEQ ID NO:14. In some aspects, the xylB is encoded by the nucleic acid sequence comprising SEQ ID NO:13.
[0024] In some aspects, the recombinant microorganism comprises a MEG production pathway having one or more of the following (c) to (e); (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose dehydrogenase and / or xylonic acid lactonase and / or xylose dehydratase, the nucleic acid sequences being operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding a glyoxaldehyde reductase that catalyzes the conversion of glyoxaldehyde to MEG; and (e) deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases in the genome of the parental microorganism.
[0025] In some aspects, the xylose dehydrogenase is from Caulobacter crescentus, Burkholderia xenovorans, Haloferax volcanii. In some aspects, the xylose dehydrogenase is xdh. In some aspects, the xdh comprises the amino acid sequence comprising SEQ ID NO:16, 17 or 19. In some aspects, the xdh is encoded by the nucleic acid sequence comprising SEQ ID NO:15, 18 or 97. In some aspects, the xylose dehydrogenase is heterologous to the microorganism.
[0026] In some aspects, the xylonolactonase is from Caulobacter crescentus, Burkholderia xenovorans, or Halomonas volcanii. In some aspects, the xylonolactonase is xylC. In some aspects, the xylC comprises an amino acid sequence comprising SEQ ID NO:55, 57, or 59. In some aspects, the xylC is encoded by a nucleic acid sequence comprising SEQ ID NO:54, 56, or 58.
[0027] In some aspects, the xylonolactonase is heterologous to the microorganism. In some aspects, the xylonolactonase is endogenous to the microorganism.
[0028] In some aspects, the xylose dehydratase is from Caulobacter crescentus, Burkholderia xenovorans, or Halomonas volcanii. In some aspects, the xylose dehydratase is xylD. In some aspects, the xylD comprises an amino acid sequence comprising SEQ ID NO:61, 63, or 65. In some aspects, the xylD is encoded by a nucleic acid sequence comprising SEQ ID NO:60, 62, or 64. In some aspects, the xylose dehydratase is heterologous to the microorganism. In some aspects, the xylose dehydratase is endogenous to the microorganism.
[0029] In some aspects, the glyoxaldehyde reductase is endogenous to the microorganism. In some aspects, the glyoxaldehyde reductase is fucO. In some aspects, the fucO comprises an amino acid sequence comprising SEQ ID NO:98. In some aspects, the fucO is encoded by a nucleic acid sequence comprising SEQ ID NO:52.
[0030] In some aspects, the glyoxaldehyde reductase is heterologous to the microorganism. In some aspects, the xylose isomerase is XylA. In some aspects, the xylA comprises an amino acid sequence comprising SEQ ID NO:6. In some aspects, the xylA is encoded by a nucleic acid sequence comprising SEQ ID NO:5.
[0031] In some aspects, the xylulokinase is XylB. In some aspects, the xylB comprises an amino acid sequence comprising SEQ ID NO:14. In some aspects, the xylB is encoded by a nucleic acid sequence comprising SEQ ID NO:13.
[0032] In some aspects, the recombinant microorganism further comprises an acetone production pathway having one or more of the following (f) to (h); (f) expression of at least one exogenous nucleic acid molecule encoding acetoacetyl-CoA thiolase; (g) expression of at least one exogenous nucleic acid molecule encoding acetate:acetoacetyl-CoA transferase; and (h) expression of at least one exogenous nucleic acid molecule encoding acetoacetate decarboxylase, which catalyzes the conversion of acetoacetate to acetone.
[0033] In some aspects, (f), (g) and (h) are in an operon controlled by the OXB11 promoter. In some aspects, the OXB11 promoter is encoded by a nucleic acid sequence comprising SEQ ID NO:78. In some aspects, the acetoacetyl-CoA thiolase is from Clostridium acetobutylicum. In some aspects, the acetoacetyl-CoA thiolase comprises an amino acid sequence comprising SEQ ID NO:67 or 69. In some aspects, the acetoacetyl-CoA thiolase is encoded by a nucleic acid sequence comprising SEQ ID NO:66 or 68. In some aspects, the acetate:acetoacetyl-CoA transferase is AtoDA. In some aspects, the AtoDA subunit α comprises an amino acid sequence comprising SEQ ID NO:72. In some aspects, the AtoDA subunit α is encoded by a nucleic acid sequence comprising SEQ ID NO:70. In some aspects, the AtoDA subunit β comprises an amino acid sequence comprising SEQ ID NO:73. In some aspects, the AtoDA subunit β is encoded by a nucleic acid sequence comprising SEQ ID NO:71.
[0034] The recombinant microorganism according to claim 77, wherein the acetoacetate decarboxylase is from Clostridium beijerinckii or Clostridium acetobutylicum. In some aspects, the acetoacetate decarboxylase is Adc. In some aspects, the Adc comprises an amino acid sequence comprising SEQ ID NO:75 or 77. In some aspects, the Adc is encoded by a nucleic acid sequence comprising SEQ ID NO:74 or 76.
[0035] In some aspects, the recombinant microorganism further comprises an isopropanol production pathway having one or more of the following (f) to (i); (f) expression of at least one exogenous nucleic acid molecule encoding acetoacetyl-CoA thiolase; (g) expression of at least one exogenous nucleic acid molecule encoding acetate:acetoacetyl-CoA transferase; and (h) expression of at least one exogenous nucleic acid molecule encoding acetoacetate decarboxylase that catalyzes the conversion of acetoacetate to acetone; (i) expression of at least one exogenous nucleic acid molecule encoding an alcohol dehydrogenase that catalyzes the conversion of acetone to isopropanol.
[0036] In some aspects, the present disclosure generally relates to recombinant microorganisms capable of producing glycolic acid from a feedstock comprising xylose and glucose, wherein the recombinant microorganisms utilize xylose and glucose simultaneously and comprise one or more of the following: (a) deletion or inactivation of fucO, yqhD, araFGH, and xylFGH in the genome of the parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganisms further express one or more glycolic acid production pathways.
[0037] In some aspects, the microorganism further comprises deletion or inactivation of glcDEF. In some aspects, the microorganism further comprises deletion or inactivation of dkgA. In some aspects, the microorganism further comprises deletion or inactivation of yahK. In some aspects, the xylose symporter is controlled by the GAPDH promoter at the araFGH locus.
[0038] In some aspects, the C5 sugar symporter is the xylose symporter XylE. In some aspects, the XylE comprises the amino acid sequence comprising SEQ ID NO:49. In some aspects, the XylE is encoded by the nucleic acid sequence comprising SEQ ID NO:48. In some aspects, the xylose symporter is endogenous to the microorganism.
[0039] In some aspects, the C5 sugar symporter is the arabinose symporter AraE. In some aspects, the arabinose symporter is endogenous to the microorganism. In some aspects, the uptake of xylose is insensitive to catabolite repression by other monosaccharides. In some aspects, the microorganism comprises a functional phosphotransferase system.
[0040] In some aspects, the microorganism comprises a native wild-type nucleic acid sequence encoding a cAMP receptor protein (CRP). In some aspects, the CRP comprises an amino acid sequence comprising SEQ ID NO:10. In some aspects, the CRP is encoded by a nucleic acid sequence comprising SEQ ID NO:9. In some aspects, constitutive overexpression of the xylose symporter enables continuous input of xylose from the feedstock into the microorganism. In some aspects, constitutive overexpression of the arabinose symporter enables continuous input of xylose from the feedstock into the microorganism. In some aspects, continuous xylose input occurs independently of the presence of other sugars in the feedstock.
[0041] In some aspects, the recombinant microorganism comprises a glyoxylate production pathway having one or more of the following (c) and (e); (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose isomerase and / or ketohexokinase and / or fructose bisphosphate aldolase, the nucleic acid sequences being operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding glyoxylate dehydrogenase, the glyoxylate dehydrogenase catalyzing the conversion of glyoxylate to glyoxylate; and (e) deletion or inactivation of one or more xylulokinases in the genome of the parental microorganism. In some aspects, (c) and (d) are in an operon controlled by the proD promoter. In some aspects, the proD promoter is encoded by a nucleic acid sequence comprising SEQ ID NO:53.
[0042] In some aspects, the xylose isomerase is XylA. In some aspects, the XylA comprises an amino acid sequence comprising SEQ ID NO:6. In some aspects, the XylA is encoded by a nucleic acid sequence comprising SEQ ID NO:5. In some aspects, the xylose isomerase is endogenous to the microorganism.
[0043] In some aspects, the ketohexokinase is from Homo sapiens. In some aspects, the ketohexokinase is heterologous to the microorganism. In some aspects, the ketohexokinase is khk-C. In some aspects, the khk-C comprises an amino acid sequence comprising SEQ ID NO:12. In some aspects, the khk-C is encoded by a nucleic acid sequence comprising SEQ ID NO:11.
[0044] In some aspects, the fructose bisphosphate aldolase is from Homo sapiens. In some aspects, the fructose bisphosphate aldolase is aldoB. In some aspects, the aldoB comprises an amino acid sequence comprising SEQ ID NO:51. In some aspects, the aldoB is encoded by a nucleic acid sequence comprising SEQ ID NO:50. In some aspects, the fructose bisphosphate aldolase is heterologous to the microorganism.
[0045] In some aspects, the glycoaldehyde dehydrogenase is aldA. In some aspects, the aldA comprises an amino acid sequence comprising SEQ ID NO:4. In some aspects, the aldA is encoded by a nucleic acid sequence comprising SEQ ID NO:3. In some aspects, the glycoaldehyde dehydrogenase is endogenous to the microorganism.
[0046] In some aspects, the xylulokinase is XylB. In some aspects, the xylB comprises an amino acid sequence comprising SEQ ID NO:14. In some aspects, the xylB is encoded by a nucleic acid sequence comprising SEQ ID NO:13.
[0047] In some aspects, the recombinant microorganism comprises a glycolate production pathway having one or more of the following (c) to (e); (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose dehydrogenase and / or xylonic acid lactonase and / or xylose dehydratase, the nucleic acid sequences being operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding a glycoaldehyde dehydrogenase that catalyzes the conversion of glycoaldehyde to glycolic acid; and (e) deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases in the genome of the parental microorganism. In some aspects, (c) and (d) are controlled by the proD promoter. In some aspects, the proD promoter is encoded by a nucleic acid sequence comprising SEQ ID NO:53.
[0048] In some aspects, the xylose isomerase is XylA. In some aspects, the XylA comprises an amino acid sequence comprising SEQ ID NO:6. In some aspects, the XylA is encoded by a nucleic acid sequence comprising SEQ ID NO:5.
[0049] In some aspects, the xylulokinase is XylB. In some aspects, the xylB comprises an amino acid sequence comprising SEQ ID NO:14. In some aspects, the xylB is encoded by a nucleic acid sequence comprising SEQ ID NO:13.
[0050] In some aspects, the xylose dehydrogenase is from Caulobacter crescentus, Burkholderia xenovorans, Halobacterium volcanii. In some aspects, the xylose dehydrogenase is xdh. In some aspects, the xdh comprises an amino acid sequence comprising SEQ ID NO:16, 17 or 19. In some aspects, the xdh is encoded by a nucleic acid sequence comprising SEQ ID NO:15, 18 or 97. In some aspects, the xylose dehydrogenase is heterologous to the microorganism.
[0051] In some aspects, the xylonic acid lactonase is from Caulobacter crescentus, Burkholderia xenovorans, or Halobacterium volcanii. In some aspects, the xylonic acid lactonase is xylC. In some aspects, the xylC comprises an amino acid sequence comprising SEQ ID NO: 55, 57, or 59. In some aspects, the xylC is encoded by a nucleic acid sequence comprising SEQ ID NO: 54, 56, or 58.
[0052] In some aspects, the xylonic acid lactonase is heterologous to the microorganism. In some aspects, the xylonic acid lactonase is endogenous to the microorganism.
[0053] In some aspects, the glyoxaldehyde dehydrogenase is aldA. In some aspects, the aldA comprises an amino acid sequence comprising SEQ ID NO: 4. In some aspects, the aldA is encoded by a nucleic acid sequence comprising SEQ ID NO: 3. In some aspects, the glyoxaldehyde dehydrogenase is endogenous to the microorganism.
[0054] In some aspects, the microorganism further expresses a glycolic acid production pathway having one or more of the following: (f) expression of at least one endogenous or exogenous nucleic acid molecule encoding isocitrate lyase; and / or (g) expression of at least one endogenous or exogenous nucleic acid molecule encoding glyoxylate reductase. In some aspects, (f) and (g) are in an operon controlled by the OXB20 promoter. In some aspects, the OXB20 promoter is encoded by a nucleic acid sequence comprising SEQ ID NO: 96.
[0055] In some aspects, the isocitrate lyase is AceA. In some aspects, the AceA comprises an amino acid sequence comprising SEQ ID NO: 90. In some aspects, the AceA is encoded by a nucleic acid sequence comprising SEQ ID NO: 89.
[0056] In some aspects, the glyoxylate reductase is YcdW. In some aspects, the YcdW comprises an amino acid sequence comprising SEQ ID NO: 92. In some aspects, the YcdW is encoded by a nucleic acid sequence comprising SEQ ID NO: 91.
[0057] In some aspects, the recombinant microorganism is derived from a parental microorganism selected from the group consisting of: Clostridium sp., Clostridium ljungdahlii, Clostridium autoethanogenum, Clostridium ragsdalei, Eubacterium limosum, Butyribacterium methylotrophicum, Moorella thermoacetica, Clostridium aceticum, Acetobacterium woodii, Alkalibaculum bacchii, Clostridium drakei, Clostridium carboxidivorans, Clostridium formicoaceticum, Clostridium scatologenes, Moorella thermoautotrophica, Acetonema longum, Blautia producta, Clostridium glycolicum, Clostridium magnum, Clostridium mayombei, Clostridium methoxybenzovorans, Clostridium acetobutylicum, Clostridium beijerinckii, Oxobacter pfennigii, Thermoanaerobacter kivui, Sporomusa ovata, Thermoacetogenium phaeum, Acetobacterium carbinolicum, Sporomusa termitida, Moorella glycerini, Eubacterium aggregans, TreponemaAzotonutricium), Escherichia coli, Saccharomyces cerevisiae, Pseudomonas putida, Bacillus sp., Corynebacterium sp., Yarrowia lipolytica, Scheffersomyces stipitis, and Terrisporobacter glycolicus. In some aspects, the parental microorganism is Escherichia coli.
[0058] In some aspects, the present disclosure generally relates to recombinant microorganisms capable of producing a fermentation product from a feedstock comprising xylose and glucose, wherein the recombinant microorganism co-utilizes xylose and glucose, and wherein the microorganism comprises one or more of the following: (a) a deletion or inactivation of a pentose ATP-binding transporter in the genome of the microorganism such that the transporter is not expressed; (b) one or more endogenous or exogenous nucleic acid sequences that encode at least one C5 sugar symporter and are operably linked to one or more constitutive promoters; wherein the C5 sugar symporter comprises: (1) a xylose symporter and / or (2) an arabinose symporter; (c) one or more endogenous or exogenous nucleic acid sequences that (1) encode xylose isomerase and are operably linked to one or more constitutive promoters, and a deletion or inactivation of one or more xylulokinases, and / or (2) encode xylose dehydrogenase and are operably linked to one or more constitutive promoters, and a deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases. In some aspects, the fermentation product produced by the microorganism is one or more molecules comprising 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 carbons. In some aspects, two or more molecules are produced simultaneously.
[0059] In some aspects, the present disclosure generally relates to recombinant Escherichia coli capable of producing a fermentation product from a feedstock comprising xylose and glucose, wherein the recombinant microorganism co-utilizes xylose and glucose, and wherein the microorganism comprises one or more of the following: (a) deletion or inactivation of the ATP-binding transporters araFGH and xylFGH in the genome of the microorganism such that the transporters are not expressed; (b) one or more endogenous or exogenous nucleic acid sequences that encode at least one C5 sugar symporter and are operably linked to one or more constitutive promoters; wherein the C5 sugar symporter comprises: (1) a xylose symporter and / or (2) an arabinose symporter; (c) one or more endogenous or exogenous nucleic acid sequences that (1) encode xylose isomerase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more xylulokinases, and / or (2) encode xylose dehydrogenase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases.
[0060] In some aspects, the present disclosure generally relates to recombinant microorganisms capable of producing monoethylene glycol (MEG) and / or acetone from a feedstock comprising xylose and glucose, wherein the recombinant microorganism co-utilizes xylose and glucose and comprises one or more of the following: (a) deletion or inactivation of aldA, araFGH, and xylFGH in the genome of the parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule that is operably linked to one or more constitutive promoters and encodes a C5 sugar symporter; wherein the recombinant microorganism expresses an MEG and / or acetone production pathway.
[0061] In some aspects, the microorganism further comprises a deletion or inactivation of glcDEF. In some aspects, the C5 symporter is controlled by the GAPDH promoter at the araFGH locus. In some aspects, the C5 sugar symporter is the xylose symporter XylE. In some aspects, the xylose symporter is endogenous to the microorganism. In some aspects, the C5 sugar symporter is the arabinose symporter AraE. In some aspects, the arabinose symporter is endogenous to the microorganism. In some aspects, the uptake of xylose is insensitive to catabolite repression by other monosaccharides. In some aspects, the microorganism comprises a functional phosphotransferase system. In some aspects, the microorganism comprises a native wild-type nucleic acid sequence encoding a cAMP receptor protein (CRP). In some aspects, the one or more nucleic acid molecules encoding aldA comprise the nucleic acid sequence set forth in SEQ ID NO:3. In some aspects, the one or more amino acid sequences encoding aldA comprise the amino acid sequence set forth in SEQ ID NO:4. The recombinant microorganism according to claim 5, wherein constitutive overexpression of the xylose symporter enables continuous input of xylose from the feedstock into the microorganism. The recombinant microorganism according to claim 5, wherein constitutive overexpression of the arabinose symporter enables continuous input of xylose from the feedstock into the microorganism. The recombinant microorganism according to claim 5, wherein continuous xylose input occurs independently of the presence of other sugars in the feedstock.
[0062] In some aspects, the recombinant microorganism comprises a MEG production pathway having one or more of the following (c) to (e): (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose isomerase and / or ketohexokinase and / or fructose bisphosphate aldolase, the nucleic acid sequences being operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding hydroxyacetaldehyde reductase, the hydroxyacetaldehyde reductase catalyzing the conversion of hydroxyacetaldehyde to MEG; and (e) deletion or inactivation of one or more xylulokinases in the genome of the parental microorganism.
[0063] In some aspects, (c) and (d) are in an operon controlled by the proD promoter. In some aspects, the xylose isomerase is XylA. In some aspects, the xylose isomerase is endogenous to the microorganism. In some aspects, the ketohexokinase is from Homo sapiens. In some aspects, the ketohexokinase is heterologous to the microorganism. In some aspects, the fructose-bisphosphate aldolase is from Homo sapiens. In some aspects, the fructose-bisphosphate aldolase is heterologous to the microorganism. In some aspects, the glycoaldehyde reductase is endogenous to the microorganism. In some aspects, the glycoaldehyde reductase is fucO. In some aspects, the xylulokinase is XylB.
[0064] In some aspects, the recombinant microorganism includes a MEG production pathway having one or more of the following (c) to (e): (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose dehydrogenase and / or xylonic acid lactonase and / or xylose dehydratase, the nucleic acid sequences being operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding a glycoaldehyde reductase that catalyzes the conversion of glycoaldehyde to MEG; and (e) deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases in the genome of the parental microorganism.
[0065] In some aspects, the xylose dehydrogenase is from Caulobacter crescentus, Burkholderia xenovorans, Halobacterium volcanii. In some aspects, the xylose dehydrogenase is heterologous to the microorganism. In some aspects, the xylonic acid lactonase is from Caulobacter crescentus, Burkholderia xenovorans, Halobacterium volcanii. In some aspects, the xylonic acid lactonase is heterologous to the microorganism. In some aspects, the xylonic acid lactonase is endogenous to the microorganism. In some aspects, the xylose dehydratase is from Caulobacter crescentus, Burkholderia xenovorans, Halobacterium volcanii.
[0066] In some aspects, the xylose dehydratase is heterologous to the microorganism. In some aspects, the xylose dehydratase is endogenous to the microorganism. In some aspects, the glycoaldehyde reductase is endogenous to the microorganism. In some aspects, the glycoaldehyde reductase is fucO. In some aspects, the glycoaldehyde reductase is heterologous to the microorganism. In some aspects, the xylose isomerase is XylA. In some aspects, the xylulokinase is XylB.
[0067] In some aspects, the recombinant microorganism further comprises an acetone production pathway having one or more of the following (f) to (h): (f) expression of at least one exogenous nucleic acid molecule encoding acetoacetyl-CoA thiolase; (g) expression of at least one exogenous nucleic acid molecule encoding acetate:acetoacetyl-CoA transferase; and (h) expression of at least one exogenous nucleic acid molecule encoding acetoacetate decarboxylase, which catalyzes the conversion of acetoacetate to acetone. In some aspects, (f), (g), and (h) are in an operon controlled by the OXB11 promoter.
[0068] In some aspects, the acetoacetyl-CoA thiolase is from Clostridium acetobutylicum. In some aspects, the acetate:acetoacetyl-CoA transferase is AtoDA. In some aspects, the acetoacetate decarboxylase is from Clostridium beijerinckii.
[0069] In some aspects, the recombinant microorganism further comprises an isopropanol production pathway having one or more of the following (f) to (i): (f) expression of at least one exogenous nucleic acid molecule encoding acetoacetyl-CoA thiolase; (g) expression of at least one exogenous nucleic acid molecule encoding acetate:acetoacetyl-CoA transferase; and (h) expression of at least one exogenous nucleic acid molecule encoding acetoacetate decarboxylase, which catalyzes the conversion of acetoacetate to acetone, (i) expression of at least one exogenous nucleic acid molecule encoding alcohol dehydrogenase, which catalyzes the conversion of acetone to isopropanol.
[0070] In some aspects, the present disclosure generally relates to recombinant microorganisms capable of producing glycolic acid from feedstocks containing xylose and glucose, wherein the recombinant microorganisms co-utilize xylose and glucose and comprise one or more of the following: (a) deletion or inactivation of fucO, yqhD, araFGH, and xylFGH in the genome of a parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism also expresses one or more glycolic acid production pathways. In some aspects, the microorganism also comprises deletion or inactivation of glcDEF. In some aspects, the microorganism also comprises deletion or inactivation of dkgA. In some aspects, the microorganism also comprises deletion or inactivation of yahK. In some aspects, the xylose symporter is controlled by the GAPDH promoter at the araFGH locus. In some aspects, the C5 sugar symporter is the xylose symporter XylE. In some aspects, the xylose symporter is endogenous to the microorganism. In some aspects, the C5 sugar symporter is the arabinose symporter AraE. In some aspects, the arabinose symporter is endogenous to the microorganism. In some aspects, the uptake of xylose is insensitive to catabolite repression by other monosaccharides.
[0071] In some aspects, the microorganism comprises a functional phosphotransferase system. In some aspects, the microorganism comprises a native wild-type nucleic acid sequence encoding a cAMP receptor protein (CRP). In some aspects, the one or more nucleic acid molecules encoding the CRP comprise the nucleic acid sequence set forth in SEQ ID NO:9. In some aspects, the one or more amino acid sequences encoding the CRP comprise the amino acid sequence set forth in SEQ ID NO:10. In some aspects, constitutive overexpression of the xylose symporter enables continuous input of xylose from the feedstock into the microorganism. In some aspects, constitutive overexpression of the arabinose symporter enables continuous input of xylose from the feedstock into the microorganism. In some aspects, continuous xylose input occurs independently of the presence of other sugars in the feedstock.
[0072] In some aspects, the recombinant microorganism comprises a glyoxylate production pathway having one or more of the following (c) and (e): (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose isomerase and / or ketohexokinase and / or fructose bisphosphate aldolase, the nucleic acid sequences being operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding glyoxylate dehydrogenase, the glyoxylate dehydrogenase catalyzing the conversion of glyoxylate to glyoxylate; and (e) deletion or inactivation of one or more xylulokinases in the genome of the parental microorganism. In some aspects, (c) and (d) are in an operon controlled by the proD promoter.
[0073] In some aspects, the xylose isomerase is XylA. In some aspects, the xylose isomerase is endogenous to the microorganism. In some aspects, the ketohexokinase is from Homo sapiens. In some aspects, the ketohexokinase is heterologous to the microorganism. In some aspects, the fructose bisphosphate aldolase is from Homo sapiens. In some aspects, the fructose bisphosphate aldolase is heterologous to the microorganism. In some aspects, the glyoxylate dehydrogenase is aldA. In some aspects, the glyoxylate dehydrogenase is endogenous to the microorganism. In some aspects, the xylulokinase is XylB.
[0074] In some aspects, the recombinant microorganism comprises a glyoxylate production pathway having one or more of the following (c) to (e); (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose dehydrogenase and / or xylonolactonase and / or xylose dehydratase, the nucleic acid sequences being operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding glyoxylate dehydrogenase, the glyoxylate dehydrogenase catalyzing the conversion of glyoxylate to glyoxylate; and (e) deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases in the genome of the parental microorganism. In some aspects, (c) and (d) are controlled by the proD promoter.
[0075] In some aspects, the xylose isomerase is XylA. In some aspects, the xylulokinase is XylB. In some aspects, the xylose dehydrogenase is from Caulobacter crescentus, Burkholderia xenovorans, Halobacterium volcanii. In some aspects, the xylose dehydrogenase is heterologous to the microorganism. In some aspects, the xylonolactonase is from Caulobacter crescentus, Burkholderia xenovorans, Halobacterium volcanii. In some aspects, the xylonolactonase is heterologous to the microorganism. In some aspects, the xylonolactonase is endogenous to the microorganism. In some aspects, the glyoxylate dehydrogenase is aldA. In some aspects, the glyoxylate dehydrogenase is endogenous to the microorganism.
[0076] In some aspects, the microorganism also expresses a glyoxylate production pathway having one or more of the following: (f) expression of at least one endogenous or exogenous nucleic acid molecule encoding isocitrate lyase; and / or (g) expression of at least one endogenous or exogenous nucleic acid molecule encoding glyoxylate reductase. In some aspects, (f) and (g) are in an operon controlled by the OXB20 promoter. In some aspects, the isocitrate lyase is AceA. In some aspects, the glyoxylate reductase is YcdW.
[0077] In some aspects, the recombinant microorganism is derived from a parental microorganism selected from the group consisting of Clostridium, Clostridium ljungdahlii, Clostridium autoethanogenum, Clostridium ragsdalei, Eubacterium limosum, Butyribacterium methylotrophicum, Moorella thermoacetica, Clostridium aceticum, Acetobacter woodii, Alkaliphilus bacchii, Clostridium drakei, Clostridium carboxidivorans, Clostridium formicoaceticum, Clostridium scatologenes, Moorella thermautotrophica, Acetofilamentum elongatum, Blautia producta, Clostridium glycolicum, Clostridium magnum, Clostridium mayombei, Clostridium methoxybenzovorans, Clostridium acetobutylicum, Clostridium beijerinckii, Acetobacterium wieringae, Thermoanaerobacter kivui, Sphaerochaeta ovata, Thermoanaerobacter brockii, Acetobacter methanolicus, Sphaerochaeta termitida, Moorella glycerini, Eubacterium aggregans, Treponema azotonutricium, Escherichia coli, Saccharomyces cerevisiae, Pseudomonas putida, Bacillus sp., Corynebacterium sp., Yarrowia lipolytica, Pichia stipitis, and Terripora glycerophila. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 Depicts multiple pathways for using xylose and glucose to produce the products contemplated herein.
[0079] Figure 2 Is a graph depicting the simultaneous utilization of glucose and xylose detected in co-consuming strains in a 1:1 ratio culture, while xylose in the parental strain only begins to be consumed after glucose is depleted.
[0080] Figure 3 Is a graph depicting that the co-consuming strain consumes 75% of the initial sugar mixture, while the parental strain only consumes 62% (36-hour culture). For a 6:1 ratio culture, both the parental and co-consuming strains completely consume the initial glucose and xylose, and the characteristic curves of xylose consumption and biomass production are similar.
[0081] Figure 4 Is a graph depicting the use of MEG in the strain. The total amount of MEG increases by 12%, while the amount of acetone increases by 197%.
[0082] Figure 5It is a figure depicting the simultaneous utilization of glucose and xylose detected in co - consuming strains in a 1:1 ratio culture, while xylose in the parental strain only starts to decrease 18 hours after glucose depletion.
[0083] Figure 6 It is a figure depicting that the co - consuming strain consumes 61% of the initial sugar mixture, while the parental strain consumes 52% of the sugar (36 - hour culture). For the 6:1 ratio culture, the co - consuming and parental strains completely consume the initial glucose and xylose, and the characteristic curves of xylose consumption and biomass production are similar.
[0084] Figure 7 It is a figure depicting that the total amount of MEG increases by 9%, while the total amount of acetone increases by 119%. Detailed Description
[0085] The present disclosure generally relates to the engineering of microorganisms to maximize the production of desired products from bio - renewable plant feedstocks, which typically cannot achieve any near - maximum yields and productivities due to the repression of multiple carbon sources present in a single type of feedstock. Stemming from the co - consumption of some monosaccharides, the present disclosure describes methods and compositions for reducing or eliminating the repression that causes microorganisms to not operate at maximum productivity.
[0086] The following definitions and abbreviations are used to explain the present disclosure.
[0087] As used herein and in the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" include plural referents. Thus, for example, reference to "an enzyme" includes multiple such enzymes and reference to "the microorganism" includes reference to one or more microorganisms, and so on.
[0088] As used herein, the terms "comprises", "comprising", "includes", "including", "has", "having", "contains", "containing" or any other variation thereof are intended to cover non - exclusive inclusion. A composition, mixture, process, method, article, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not expressly listed or that are non - inherent to such composition, mixture, process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to the inclusive "or" rather than the exclusive "or".
[0089] The terms "polynucleotide", "nucleotide", "nucleotide sequence", "nucleic acid", and "oligonucleotide" are used interchangeably. They refer to polymeric forms of nucleotides (deoxyribonucleotides or ribonucleotides) of any length or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. Following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, multiple loci (a locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides can contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. Modifications to the nucleotide structure can be imparted before or after polymerization of the polymer, if present. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with a labeling component.
[0090] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence through traditional Watson-Crick or other non-traditional types. Percent complementarity represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, or 10 out of 10 being 50%, 60%, 70%, 80%, 90%, and 100% complementary, respectively). "Fully complementary" means that all consecutive residues of a nucleic acid sequence will hydrogen bond with the same number of consecutive residues in a second nucleic acid sequence. As used herein, "substantially complementary" refers to a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions. Sequence identity, such as for the purpose of assessing percent complementarity, can be measured by any suitable alignment algorithm, including but not limited to the Needleman-Wunsch algorithm (see, e.g., the EMBOSS Needle alignment program available at www.ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html, optionally with default settings), the BLAST algorithm (see, e.g., the BLAST alignment tool available at blast.ncbi.nlm.nih.gov / Blast.cgi, optionally with default settings), or the Smith-Waterman algorithm (see, e.g., the Water alignment program available at www.ebi.ac.uk / Tools / psa / emboss_water / nucleotide.html, optionally with default settings). Any suitable parameters of the selected algorithm, including default parameters, can be used to evaluate the optimal alignment.
[0091] As used herein, "expression" refers to the process of transcription of a polynucleotide from a DNA template (such as transcription into mRNA or other RNA transcript) and / or the subsequent translation of the transcribed mRNA into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide can be collectively referred to as a "gene product". If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell.
[0092] The terms "polypeptide", "peptide", and "protein" are used interchangeably herein and refer to polymers of amino acids of any length. The polymers can be linear or branched, can contain modified amino acids, and can be interrupted by non-amino acids. These terms also encompass polymers of amino acids that have been modified; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation (such as conjugation with a label component). As used herein, the term "amino acid" includes natural and / or non-natural or synthetic amino acids, including glycine and D or L optical isomers, as well as amino acid analogs and peptidomimetics.
[0093] As used herein, the term "about" is used synonymously with the term "approximately". Exemplarily, the use of the term "about" with respect to a quantity represents a value that is slightly beyond the recited value, e.g., plus or minus 0.1% to 10%.
[0094] The term "biologically pure culture" or "substantially pure culture" refers to a culture of the bacterial species described herein that is free of other bacterial species in amounts sufficient to interfere with the replication of the culture or detectable by normal bacteriological techniques.
[0095] As used herein, a "control sequence" refers to an operon, promoter, silencer, or terminator.
[0096] As used herein, "introduced" refers to introduction by modern biotechnology, rather than natural occurrence.
[0097] As used herein, a "constitutive promoter" is a promoter that is active under most conditions and / or during most developmental stages. There are several advantages to using a constitutive promoter in an expression vector used in biotechnology, such as: high-level production of a protein for selection of transgenic cells or organisms; high expression levels of a reporter protein or scorable marker, allowing easy detection and quantification; high-level production of a transcription factor as part of a regulatory transcription system; production of a compound that requires ubiquitous activity in an organism; and production of a desired compound during all developmental stages.
[0098] As used herein, a "non-constitutive promoter" is a promoter that is active under certain conditions, in certain types of cells, and / or during certain developmental stages. For example, inducible promoters and promoters under developmental control are non-constitutive promoters.
[0099] As used herein, an "inducible" or "repressible" promoter is a promoter that is under the control of chemical or environmental factors. Examples of environmental conditions that can affect transcription of an inducible promoter include anaerobic conditions, certain chemicals, the presence of light, acidic or alkaline conditions, etc.
[0100] As used herein, the term "operably linked" refers to the association of nucleic acid sequences on a single nucleic acid fragment such that the function of one is regulated by the other. For example, a promoter is operably linked to a coding sequence when it is capable of regulating the expression of the coding sequence (i.e., the coding sequence is under the transcriptional control of the promoter). The coding sequence can be operably linked to the regulatory sequence in the sense or antisense orientation. In another example, the complementary RNA regions of the present disclosure can be directly or indirectly operably linked to the 5', or the 3', or within the target mRNA, or the first complementary region is at the 5' of the target mRNA and its complement is at the 3' of the target mRNA.
[0101] As used herein, the term "signal sequence" refers to an amino acid sequence that targets peptides and polypeptides to a cellular location or the extracellular environment. The signal sequence is typically at the N-terminal portion of the polypeptide and is typically removed enzymatically. A polypeptide with its signal sequence is referred to as full-length and / or unprocessed. A polypeptide with its signal sequence removed is referred to as mature and / or processed.
[0102] As used herein with respect to various molecules, such as polynucleotides, polypeptides, enzymes, etc., the term "exogenous" refers to molecules that are not normally or naturally found in a given yeast, bacterium, organism, microorganism, or cell in nature and / or are not produced by a given yeast, bacterium, organism, microorganism, or cell.
[0103] On the other hand, as used herein with respect to various molecules, such as polynucleotides, polypeptides, enzymes, etc., the term "endogenous" or "native" refers to molecules that are normally or naturally found in a given yeast, bacterium, organism, microorganism, or cell in nature and / or are produced by a given yeast, bacterium, organism, microorganism, or cell.
[0104] As used herein in the context of a modified host cell, the term "heterologous" refers to various molecules, such as polynucleotides, polypeptides, enzymes, etc., where at least one of the following is true: (a) the one or more molecules are foreign to the host cell ("exogenous") (i.e., not naturally found); (b) the one or more molecules are naturally found in a given host microorganism or host cell (e.g., are "endogenous"), but are produced in the cell at a non-natural location or in a non-natural amount; and / or (c) the one or more molecules differ in nucleotide or amino acid sequence from the endogenous nucleotide or amino acid sequence such that a molecule that differs in nucleotide or amino acid sequence from the endogenous nucleotide found endogenously is produced in the cell in a non-natural amount (e.g., greater than the amount naturally found).
[0105] As used herein, the term "homolog" with respect to an original enzyme or gene of a first family or species refers to a different enzyme or gene of a second family or species that is determined by functional, structural, or genomic analysis to correspond to the original enzyme or gene of the first family or species. Homologs most often have functional, structural, or genomic similarity. Techniques are known for readily cloning homologs of enzymes or genes using gene probes and PCR. The identity of a cloned sequence as a homolog can be confirmed using functional assays and / or by genomic mapping of the gene.
[0106] A protein is "homologous" to a second protein or "is homologous with" a second protein if the amino acid sequence encoded by the gene has an amino acid sequence that is similar to the amino acid sequence of the second gene. Alternatively, a protein is homologous with a second protein if the two proteins have "similar" amino acid sequences. Thus, the term "homologous proteins" is intended to mean that the two proteins have similar amino acid sequences. In some cases, homology between two proteins indicates a common lineage due to evolutionary relatedness. The terms "homologous sequences" or "homologs" are considered, believed, or known to be functionally related. The functional relationship can be indicated in any of a variety of ways, including but not limited to: (a) the degree of sequence identity and / or (b) the same or similar biological function. Preferably, both (a) and (b) are indicated. The degree of sequence identity can vary, but in one aspect is at least 50% (when using standard sequence alignment programs known in the art), at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least 98.5%, or at least about 99%, or at least 99.5%, or at least 99.8%, or at least 99.9%. Homology can be determined using software programs readily available in the art, such as those discussed in Current Protocols in Molecular Biology (edited by F.M. Ausubel et al., 1987), Supplement 30, Section 7.718, Table 7.71. Some alignment programs are MacVector (Oxford Molecular Ltd, Oxford, U.K.) and ALIGN Plus (Scientific and Educational Software, Pennsylvania). Other non-limiting alignment programs include Sequencher (Gene Codes, Ann Arbor, Michigan), AlignX, and Vector NTI (Invitrogen, Carlsbad, CA). Similar biological functions can include but are not limited to: catalyzing the same or similar enzymatic reactions; having the same or similar selectivity for substrates or cofactors; having the same or similar stability; having the same or similar tolerance to various fermentation conditions (temperature, pH, etc.); and / or having the same or similar tolerance to various metabolic substrates, products, by-products, intermediates, etc.Based on one or more assays known to those skilled in the art for determining a given biological function, the degree of similarity of biological functions can vary, but in one aspect, it is at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least 98.5% or at least about 99% or at least 99.5% or at least 99.8% or at least 99.9%.
[0107] The term "variant" refers to any polypeptide or enzyme described herein. Variants also encompass one or more components of a multimer, a multimer comprising individual components, a multimer comprising multiple individual components (e.g., a multimer of a reference molecule), chemical degradation products, and biological degradation products. Specifically, in non-limiting aspects, an enzyme can be a "variant" relative to a reference enzyme due to a change in any part of the polypeptide sequence encoding the reference enzyme. In a standard assay for measuring the enzyme activity of a preparation of a reference enzyme, a variant of the reference enzyme can have at least 10%, at least 30%, at least 50%, at least 80%, at least 90%, at least 100%, at least 105%, at least 110%, at least 120%, at least 130% or higher enzyme activity. In some aspects, a variant can also refer to a polypeptide having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the full-length or unprocessed enzyme of the present disclosure. In some aspects, a variant can also refer to a polypeptide having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the mature or processed enzyme of the present disclosure.
[0108] As used herein, the term "microorganism (microorganism or microbe)" should be understood broadly. These interchangeable terms include, but are not limited to, two prokaryotic domains, bacteria and archaea.
[0109] As used herein, terms such as "isolated", "isolated from", "isolated microorganism" are intended to mean that one or more microorganisms have been separated from at least one material with which they were associated in a particular environment (e.g., culture medium, water, reaction chamber, etc.). Thus, an "isolated microorganism" does not exist in its natural environment; rather, through the various techniques described herein, the microorganism is removed from its natural environment and placed in a non-natural state of existence. Thus, an isolated strain or isolated microorganism can exist, for example, as a biopure culture or spore (or other form of strain). In some aspects, the isolated microorganism can be combined with an acceptable carrier, which can be a commercially or industrially acceptable carrier.
[0110] In certain aspects of the present disclosure, the isolated microorganism exists as an "isolated and biopure culture". Those skilled in the art will understand that an isolated and biopure culture of a particular microorganism means that the culture is substantially free of other living organisms and contains only the individual microorganism in question. The culture can contain different concentrations of the microorganism. The present disclosure states that isolated and biopure microorganisms are generally "necessarily different from less pure or impure materials". See, e.g., In re Bergstrom, 427 F.2d 1394, (CCPA 1970) (discussing purified prostaglandins), see also, In re Bergy, 596 F.2d 952 (CCPA 1979) (discussing purified microorganisms), see also, Parke-Davis & Co. v. H.K. Mulford & Co., 189 F. 95 (S.D.N.Y. 1911) (discussing a treatise on purified adrenaline), aff'd in part, rev'd in part, 196 F. 496 (2d Cir. 1912), each of which is incorporated herein by reference. In addition, in some aspects, the present disclosure provides certain quantitative measures of the concentration or purity limits that must be present in an isolated and biopure microorganism culture. In certain aspects, the presence of these purity values is an additional attribute that distinguishes the microorganisms of the present disclosure from those that exist in their natural state. See, e.g., Merck & Co. v. Olin Mathieson Chemical Corp., 253 F.2d 156 (4th Cir. 1958) (discussing purity limits of vitamin B12 produced by microorganisms), which is incorporated herein by reference.
[0111] The microorganisms of the present disclosure may include spores and / or vegetative cells. In some aspects, the microorganisms of the present disclosure include microorganisms in a viable but non-culturable (VBNC) state. As used herein, "spore (or spores)" refers to a structure produced by bacteria and fungi that is adapted for survival and dissemination. Spores are typically characterized as dormant structures; however, spores are capable of differentiating through a process of germination. Germination is the differentiation of a spore into a vegetative cell capable of metabolic activity, growth, and reproduction. The germination of a single spore gives rise to a single fungal or bacterial vegetative cell. Fungal spores are units of asexual reproduction and are, in some cases, essential structures in the fungal life cycle. Bacterial spores are structures for survival conditions that may generally be unfavorable for the survival or growth of vegetative cells.
[0112] As used herein, "microbial composition" refers to a composition comprising one or more microorganisms of the present disclosure.
[0113] As used herein, "carrier", "acceptable carrier", "commercially acceptable carrier", or "industrially acceptable carrier" refers to a diluent, adjuvant, excipient, or vehicle that can be used to administer, store, or transfer microorganisms without having an adverse effect on the microorganisms.
[0114] As used herein, the term "yield potential" refers to the yield of a product from a biosynthetic pathway. In one aspect, the yield potential can be expressed as a percentage based on the final product weight / starting compound weight.
[0115] As used herein, the term "thermodynamic maximum yield" refers to the maximum yield of a product obtained from the fermentation of a given feedstock such as glucose based on the energy value of the product compared to the feedstock. In normal fermentation, no additional energy sources such as light, hydrogen, or methane or electricity are used, e.g., the product cannot contain more energy than the feedstock. The thermodynamic maximum yield represents the product yield when all the energy and mass from the feedstock are converted into the product. This yield can be calculated and is independent of the specific pathway. If the yield of a specific pathway towards the product is lower than the thermodynamic maximum yield, it has lost mass and is most likely to be improved or replaced by a more efficient pathway towards the product.
[0116] The term "redox balance" refers to a set of reactions that together produce as much redox cofactor as they consume. Designing metabolic pathways and engineering organisms to have a balanced or nearly balanced redox cofactor often results in the production of the desired compound in a more efficient and higher-yield manner. Redox reactions always occur together as two simultaneous half-reactions, one a oxidation reaction and the other a reduction reaction. In a redox process, a reducing agent transfers electrons to an oxidizing agent. Thus, in the reaction, the reducing agent or reducing reagent loses electrons and is oxidized, while the oxidizing agent or oxidizing reagent gains electrons and is reduced. In one aspect, redox reactions occur in biological systems. Bioenergy is often stored and released through redox reactions. Photosynthesis involves the reduction of carbon dioxide to sugars and the oxidation of water to molecular oxygen. The reverse reaction, respiration, oxidizes sugars to produce carbon dioxide and water. As an intermediate step, reduced carbon compounds are used to reduce nicotinamide adenine dinucleotide (NAD+), which then helps to generate a proton gradient that drives the synthesis of adenosine triphosphate (ATP) and is maintained by the reduction of oxygen. The term redox state is commonly used to describe the balance of GSH / GSSG, NAD+ / NADH, and NADP+ / NADPH in biological systems such as cells or organs. The redox state is reflected in the balance of several metabolites (e.g., lactate and pyruvate, β-hydroxybutyrate and acetoacetate), whose interconversion depends on these ratios. Abnormal redox states can develop in many harmful situations such as hypoxia, shock, and sepsis.
[0117] As used herein, the term "productivity" refers to the total amount of bioproduct produced per hour - grams of product / (liter per hour).
[0118] As used herein, the terms "substantially free of microorganisms", "substantially free of bacteria", or "substantially free of fungi / yeast" should not be construed to mean the absence of microorganisms / bacteria / fungi / yeast, although in some aspects this may be preferred. Instead, "substantially free of" should be construed to mean, for example, that a composition substantially free of bacteria is a composition in which any bacteria present in the composition are so few that they are below the detection limit. In some aspects, the microorganisms are selected from one or more bacteria, fungi, yeast, viruses, protists, and algae.
[0119] As used herein, the term "free of microorganisms" means the complete absence of microorganisms or the complete absence of live microorganisms capable of undergoing reproductive vegetative growth.
[0120] Consuming xylose and glucose simultaneously
[0121] In an industrial or commercial process, microbial productivity is a key factor that must be considered when contemplating the economic viability of large-scale reactions, which typically have thin profit margins. Microbial productivity, in this sense, is the number of grams of product produced per liter per hour. In the absence of engineered microorganisms, a stream containing xylose and glucose continuously fed into one or more reaction chambers will likely result in repression of the uptake of at least glucose by one or more monosaccharides.
[0122] The underlying mechanism for two-stage growth is carbon catabolite repression (CCR), in which the global transcriptional regulator CRP (cAMP receptor protein) plays a central role in regulating the transcriptional activation of catabolic operons for secondary sugars such as xylose, arabinose, and galactose. The phosphoenolpyruvate:sugar phosphotransferase system (PTS) is also involved in the repression of xylose utilization in glucose-induced Escherichia coli. Xylose can be used by Escherichia coli as the sole carbon and energy source and is metabolized via the pentose phosphate pathway. Xylose can be imported via two uptake systems: a high-affinity ATP-dependent system and a relatively low-affinity D-xylose:H+ symporter. Unlike arabinose transport, which occurs mainly via a more energy-efficient symporter, xylose is mainly transported via a more energy-consuming ATP-dependent transporter, even at high sugar concentrations. All genes responsible for xylose uptake and catabolism are sensitive to CCR.
[0123] To achieve an efficient biological process for converting pentoses into desired chemicals, it is necessary to engineer the host microorganism to utilize mixed sugars efficiently, synchronously, and rapidly, to achieve the yields and productivities required for industrial processes. This disclosure describes a metabolic engineering strategy to effectively promote the simultaneous consumption of xylose and glucose from lignocellulosic biomass, thereby realizing the full potential of the engineered microbial strain and obtaining the desired chemicals from pathways that have D-xylonic acid or D-xylulose-1P or glycolaldehyde as intermediates.
[0124] Common strategies for engineering sugar co-utilization in Escherichia coli rely on the inactivation of PTS components, which may or may not be associated with improvements in galP (galactose:H+ symporter) activity and CRP mutations. However, inactivation of PTS components impairs glucose uptake, and CRP mutants typically have a slow growth phenotype, which may be due to unpredictable changes in the expression of other important genes. These two methods result in decreased productivity, especially for conditions of high sugar concentration and low-cost media.
[0125] The applicant believes that the metabolic engineering strategy for supporting the co-consumption of glucose and xylose to produce the desired chemical substances is developed for the first time, which uses D-xylulose-1P or D-xylonic acid or glycolaldehyde as intermediates, is independent of PTS inactivation and has a deletion of the ATP-binding transporter.
[0126] In some aspects, promoting the simultaneous consumption of xylose and glucose to produce the desired chemical substances is based on: (1) constitutive overexpression of an ATP-independent D-xylose symporter; (2) constitutive expression of genes that convert xylose into D-xylulose-1P or D-xylonic acid; (3) deletion of the native pentose (mainly xylose and arabinose) ABC transporter system; and / or deletion of xylose catabolic genes.
[0127] The subject matter described herein is different from the prior art in that the deletion or inactivation of the ABC transporter and the expression of symporters and pathways that utilize or contain D-xylulose, D-xylonic acid or glycolaldehyde as intermediates not only positively affect the co-utilization of sugars, but also increase the overall yield and productivity of the pathway for producing the desired chemical substances. This improvement is due to the regulation of the total metabolism of the microorganism, the change in the ATP availability curve, and the promotion of the production of the intermediates D-xylulose-1P, D-xylonic acid and / or glycolaldehyde. See Kim et al. (2015. Metabolic Engineering, 30:141-148), Sievert et al. (2017. PNAS, 114(28):7349-7354), Wang et al. (2018. Microbial Cell Factories, 17(12):1-12) and Bai et al. (2016. Metabolic Engineering, 38:285-292).
[0128] The present disclosure includes strategies for overcoming catabolite repression of xylose by glucose, allowing the simultaneous consumption of both sugars. Different from other methods of co-consuming sugars, the design and implementation of this strategy focus on ensuring effective xylose uptake that is insensitive to catabolite repression of sugars, while maintaining effective uptake of glucose by the native PTS system.
[0129] In some aspects, the method includes making the following modifications in the microbial strain of interest: 1(a) overexpressing the native xylose symporter XylE operably linked to a constitutive promoter, and / or 1(b) overexpressing the native arabinose symporter AraE operably linked to a constitutive promoter; 2(a) expressing the native xylose isomerase XylA, the heterologous ketohexokinase khk-C, and the deletion or inactivation of the native xylulokinase XylB under a constitutive promoter, or 2(b) expressing the heterologous xylose dehydrogenase xdh operably linked to a constitutive promoter and the deletion or inactivation of the native xylose isomerase XylA and / or the deletion of the native xylulokinase XylB; and 3 the deletion of the ATP-binding transporters AraFGH, XylFGH, RbsABC, and AlsABC.
[0130] In some aspects, the constitutive expression of the xylose and arabinose symporters renders xylose import independent of CRP regulation and thus independent of other sugars present in the culture medium. In some aspects, the constitutive expression of the xylose isomerase renders xylose utilization independent of CRP regulation and thus also independent of other sugars present in the culture medium. In some aspects, the expression of the ketohexokinase khk-C effectively converts D-xylulose into the intermediate D-xylulose 1-P for the production of the desired chemical. In some aspects, the deletion of the xylulokinase prevents carbon diversion from the chemical production pathway to the native pentose phosphate pathway.
[0131] In some aspects, the constitutive expression of the xylose dehydrogenase renders xylose utilization independent of CRP regulation and thus also independent of other sugars present in the culture medium; and also effectively converts D-xylose into the intermediate D-xylonic acid for the production of the desired chemical. In some aspects, the deletion of the xylulokinase and / or the xylose isomerase prevents carbon diversion from the chemical production pathway to the native pentose phosphate pathway. In some aspects, the deletion of the ATP-binding cassette transporters (such as the arabinose ABC transporter and the xylose ABC transporter) avoids the loss of ATP during the sugar import process. The net content of ATP can alter the activity of the central metabolism of Escherichia coli, thereby potentially increasing the pathway yield.
[0132] Overexpressing the native xylose symporter XylE under a constitutive promoter
[0133] The D-xylose / proton symporter XylE is an ATP-independent low-affinity transporter encoded by the xylE gene and is a member of the major facilitator superfamily (MFS) of transporters. The transcription of xylE is believed to be regulated by XylR (SEQ ID NO:7 or SEQ ID NO:8). XylR is a transcription factor encoded by the xylR gene that positively regulates the transcription of xylose metabolism and transporter genes in response to xylose (xylE, xylFGH, and xylAB genes).
[0134] Constitutive overexpression of the xylose symporter relieves carbon catabolite repression and enables continuous xylose import, which is independent of the sugars present in the culture medium, while glucose uptake will still be carried out by the PTS system components. Thus, both glucose and xylose present in the hydrolysate can be imported by E. coli simultaneously.
[0135] Overexpression of the native arabinose symporter AraE under a constitutive promoter
[0136] The D-arabinose / proton symporter AraE is an ATP-independent low-affinity transporter encoded by the araE gene and is a member of the major facilitator superfamily (MFS) of transporters. The transcription of araE is regulated by AraC (SEQ ID NO:32 or SEQ ID NO:33) and CRP. AraC is a transcription factor encoded by the araC gene that negatively regulates the transcription of xylose metabolism and transporter genes in response to xylose (xylE, xylFGH, and xylAB genes) and positively regulates the transcription of arabinose metabolism and transporter genes in response to arabinose (araE, araFGH, and araBAD genes). In the absence of glucose, araE expression is induced by arabinose. It is well known that the AraE transporter is promiscuous and can transport xylose and other pentoses.
[0137] Constitutive expression of the promiscuous arabinose symporter relieves CCR and enables continuous xylose import, which is independent of the sugars present in the culture medium, while glucose uptake is still carried out by the PTS system components. Thus, both glucose and xylose present in the hydrolysate can be imported by E. coli simultaneously.
[0138] Expression of the native xylose isomerase XylA and the heterologous ketohexokinase khk-C under a constitutive promoter, and deletion of the native xylulokinase XylB
[0139] XylA is an endogenous D-xylose isomerase ( Figure 1, Reaction 5, Pathway B), which catalyzes the conversion of D-xylose to D-xylulose. D-xylose isomerase (E.C. 5.3.1.5) catalyzes the first reaction in the natural catabolism of D-xylose in Escherichia coli. The transcription of xylA is regulated by XylR and CRP; in the absence of glucose, its expression is induced by xylose. Ketohexokinase ( Figure 1 , Reaction 6, Pathway B) catalyzes the phosphorylation of D-xylulose to D-xylulose-1-P. Ketohexokinase (E.C. 2.7.1.3) can be present in a variety of organisms; however, khk-C from human liver is a promising candidate that is active towards xylulose. D-xylulose 1-P is a key intermediate for the production of various chemicals.
[0140] XylB is a xylulokinase (2.7.1.17) encoded by xylB that catalyzes the phosphorylation of D-xylulose ( Figure 1 , Reaction 8, Pathway B). This is the second step of the natural xylose degradation pathway, which produces the pentose phosphate pathway intermediate D-xylulose-5-phosphate. This reaction competes with the phosphorylation of D-xylulose by khk-C, diverting the flux away from the production of D-xylulose-1-P towards the pentose phosphate pathway.
[0141] Constitutive expression of the native xylose isomerase xylA relieves CCR and, when associated with constitutive heterologous expression of the ketohexokinase khk-C, enables continuous xylose utilization independent of the sugars present in the culture medium and produces the chemical-producing intermediate D-xylulose 1-P. Glucose uptake will still be carried out by the PTS system components. Thus, both glucose and xylose present in the hydrolysate can be utilized simultaneously by Escherichia coli. Deletion of the xylulokinase xylB will prevent carbon diversion from the chemical production pathway to the native pentose phosphate pathway.
[0142] Expression of the heterologous xylose dehydrogenase xdh under a constitutive promoter, and deletion and / or inactivation of the native xylose isomerase Xyla and / or the native xylulokinase XylB
[0143] The heterologous xylose dehydrogenase xdh catalyzes the conversion of D-xylose to D-xylonolactone ( Figure 1 , Reaction 1, Pathway A). D-xylose dehydrogenase (E.C. 1.1.1.175) can be present in a variety of organisms; however, xdh from Caulobacter crescentus is a candidate that is active towards D-xylose. D-xylonolactone can spontaneously convert to D-xylonic acid, so expression of xdh towards xylose produces D-xylonic acid, which is a key intermediate for the production of various chemicals.
[0144] The D-xylose isomerase (E.C. 5.3.1.5) XylA encoded by xylA catalyzes the conversion of D-xylose to the pentose phosphate pathway intermediate D-xyluloseFigure 1 , reaction 5, pathway A). The xylulose kinase (2.7.1.17) XylB encoded by xylB catalyzes the phosphorylation of D-xylulose ( Figure 1 , reaction 8, pathway A), which is the second step of the xylose degradation pathway and generates another intermediate of the pentose phosphate pathway, D-xylulose-5-phosphate. Both of these reactions compete with xdh, diverting the flux from D-xylonate production to the pentose phosphate pathway.
[0145] Constitutive heterologous expression of xylose dehydrogenase enables continuous xylose utilization independent of the sugars present in the culture medium and generates the intermediate D-xylonate for the production of the desired chemicals. Glucose uptake will still be carried out by the PTS system components. Thus, both glucose and xylose present in the hydrolysate can be utilized by E. coli simultaneously. The deletion of D-xylose isomerase and / or xylulose kinase will prevent carbon diversion from the chemical production pathway to the native pentose phosphate pathway.
[0146] Deletion of the ATP-binding cassette transporters AraFGH, XylFGH, RbsABC, and AlsABC
[0147] The arabinose ABC transporter AraFGH (E.C. 3.6.3.17, TCDB 3.A.1.2.2) is a high-affinity ATP-driven system encoded by the araFGH genes. AraF is the periplasmic binding protein, AraH is the membrane component, and AraG is the ATP-binding component of this ABC transporter. The transcription of the araFGH operon is regulated by AraC and CRP. In the absence of glucose, araFGH expression is induced by arabinose. It is well known that the AraFGH transporter is promiscuous and can transport xylose and other pentoses.
[0148] The xylose ABC transporter XylFGH (E.C. 3.6.3.17, TCDB 3.A.1.2.4) is a high-affinity ATP-driven system encoded by the xylFGH genes. XylF is the periplasmic binding protein, XylH is the membrane component, and XylG is the ATP-binding component of this ABC transporter. The transcription of the xylFGH operon is regulated by XylR and CRP; in the absence of glucose, its expression is induced by xylose.
[0149] The ribose ABC transporter RbsABC (E.C. 3.6.3.17; TCDB 3.A.1.2.1) is a high-affinity ATP-driven system encoded by the rbsABC genes.
[0150] In some aspects, the one or more nucleic acid molecules encoding the RbsB periplasmic binding protein subunit of RbsABC comprise the nucleic acid sequence set forth in SEQ ID NO: 35. In some aspects, the one or more amino acid sequences encoding the RbsB periplasmic binding protein subunit of RbsABC comprise the amino acid sequence set forth in SEQ ID NO: 38. In some aspects, the one or more nucleic acid molecules encoding the RbsA ATP-binding subunit of RbsABC comprise the nucleic acid sequence set forth in SEQ ID NO: 34. In some aspects, the one or more amino acid sequences encoding the RbsA ATP-binding subunit of RbsABC comprise the amino acid sequence set forth in SEQ ID NO: 37. In some aspects, the one or more nucleic acid molecules encoding the RbsC membrane subunit of RbsABC comprise the nucleic acid sequence set forth in SEQ ID NO: 36. In some aspects, the one or more amino acid sequences encoding the RbsC membrane subunit of RbsABC comprise the amino acid sequence set forth in SEQ ID NO: 39.
[0151] The allose ABC transporter AlsABC (E.C. 3.6.3.17; TCDB 3.A.1.2.6) is an ATP-driven system encoded by the alsABC gene.
[0152] In some aspects, the one or more nucleic acid molecules encoding the alsB periplasmic binding protein subunit of AlsABC comprise the nucleic acid sequence set forth in SEQ ID NO: 41. In some aspects, the one or more amino acid sequences encoding the alsB periplasmic binding protein subunit of AlsABC comprise the amino acid sequence set forth in SEQ ID NO: 44. In some aspects, the one or more nucleic acid molecules encoding the alsA ATP-binding subunit of AlsABC comprise the nucleic acid sequence set forth in SEQ ID NO: 40. In some aspects, the one or more amino acid sequences encoding the alsA ATP-binding subunit of AlsABC comprise the amino acid sequence set forth in SEQ ID NO: 43. In some aspects, the one or more nucleic acid molecules encoding the alsC membrane subunit of AlsABC comprise the nucleic acid sequence set forth in SEQ ID NO: 42. In some aspects, the one or more amino acid sequences encoding the alsC membrane subunit of AlsABC comprise the amino acid sequence set forth in SEQ ID NO: 45.
[0153] Deletion of ATP-binding cassette transporters (such as ribose ABC transporters, allose ABC transporters, arabinose ABC transporters, and xylose ABC transporters, including the preferential xylose transporter in Escherichia coli) in combination with constitutive expression of xylE (see Examples 1 and 2) relieves CCR and enables continuous xylose import. In addition, the deletion will avoid ATP loss during sugar import. Net ATP can alter the activity of central metabolism in Escherichia coli.
[0154] Microorganism
[0155] As described herein, in some aspects, the recombinant microorganism is capable of simultaneously utilizing xylose and glucose.
[0156] As described herein, in some aspects, the recombinant microorganism is a prokaryotic microorganism. In some aspects, the prokaryotic microorganism is a bacterium. "Bacteria" or "eubacteria" refers to the domain Bacteria. Bacteria include at least eleven different classes as follows: (1) Gram-positive (Gram+) bacteria, which have two main branches: (1) high G+C classes (Actinomycete, Mycobacteria, Micrococcus, etc.) (2) low G+C classes (Bacillus, Clostridia, Lactobacillus, Staphylococci, Streptococci, Mycoplasmas); (2) Proteobacteria, such as purple photosynthetic + non-photosynthetic Gram-negative bacteria (including most "common" Gram-negative bacteria); (3) Cyanobacteria, such as oxygenic phototrophs; (4) Spirochetes and related species; (5) Planctomyces; (6) Bacteroides, Flavobacteria; (7) Chlamydia; (8) Green sulfur bacteria; (9) Green non-sulfur bacteria (also anaerobic phototrophs); (10) Radioresistant Micrococcus and related bacteria; (11) Thermotoga and Thermosipho thermophiles.
[0157] "Gram-negative bacteria" include cocci, non-enterobacteria, and enterobacteria. Genera of Gram-negative bacteria include, for example, Neisseria, Spirillum, Pasteurella, Brucella, Yersinia, Francisella, Haemophilus, Bordetella, Escherichia, Salmonella, Shigella, Klebsiella, Proteus, Vibrio, Pseudomonas, Bacteroides, Acetobacter, Aerobacter, Agrobacterium, Azotobacter, Spirilla, Serratia, Vibrio, Rhizobium, Chlamydia, Rickettsia, Treponema, and Fusobacterium.
[0158] "Gram-positive bacteria" include cocci, non-sporulating bacilli, and sporulating bacilli. Genera of Gram-positive bacteria include, for example, Actinomyces, Bacillus, Clostridium, Corynebacterium, Erysipelothrix, Lactobacillus, Listeria, Mycobacterium, Myxococcus, Nocardia, Staphylococcus, Streptococcus, and Streptomyces.
[0159] In some aspects, the microorganisms of the present disclosure are fungi.
[0160] In some aspects, the recombinant microorganism is a eukaryotic microorganism. In some aspects, the eukaryotic microorganism is yeast. In an exemplary aspect, the yeast is a member of a genus selected from the group consisting of: Yarrowia, Candida, Saccharomyces, Pichia, Hansenula, Kluyveromyces, Issatchenkia, Zygosaccharomyces, Debaryomyces, Schizosaccharomyces, Pachysolen, Cryptococcus, Trichosporon, Rhodotorula, and Myxozyma.
[0161] In some aspects, the recombinant microorganism is a prokaryotic microorganism. In an exemplary aspect, the prokaryotic microorganism is a member of a genus selected from the group consisting of: Escherichia, Clostridium, Zymomonas, Salmonella, Rhodococcus, Pseudomonas, Bacillus, Lactobacillus, Enterococcus, Alcaligenes, Klebsiella, Paenibacillus, Arthrobacter, Corynebacterium, and Brevibacterium.
[0162] In some aspects, the microorganism used in the methods of the present disclosure can be selected from the group consisting of: Yarrowia, Candida, Saccharomyces, Pichia, Hansenula, Kluyveromyces, Issatchenkia, Zygosaccharomyces, Debaryomyces, Schizosaccharomyces, Pachysolen, Cryptococcus, Trichosporon, Rhodotorula, Myxozyma, Escherichia, Clostridium, Zymomonas, Salmonella, Rhodococcus, Pseudomonas, Bacillus, Lactobacillus, Enterococcus, Alcaligenes, Klebsiella, Paenibacillus, Arthrobacter, Corynebacterium, and Brevibacterium.
[0163] In some aspects, the microorganisms produced by the methods described herein can be species of any genus selected from the following genera: Neisseria, Spirillum, Pasteurella, Brucella, Yersinia, Francisella, Haemophilus, Bordetella, Escherichia, Salmonella, Shigella, Klebsiella, Proteus, Vibrio, Pseudomonas, Bacteroides, Acetobacter, Aerobacter, Agrobacterium, Azotobacter, Spirochaeta, Serratia, Vibrio, Rhizobium, Chlamydia, Rickettsia, Treponema, Fusobacterium, Actinomyces, Bacillus, Clostridium, Corynebacterium, Erysipelothrix, Lactobacillus, Listeria, Mycobacterium, Myxococcus, Nocardia, Staphylococcus, Streptococcus, Streptomyces, Saccharomyces, Pichia, and Aspergillus.
[0164] In some aspects, the microorganisms used in the methods of the present disclosure include: Clostridium, Clostridium ljungdahlii, Clostridium autoethanogenum, Clostridium ragsdalei, Eubacterium limosum, Butyribacterium methylotrophicum, Moorella thermoacetica, Clostridium aceticum, Acetobacter woodii, Alkaliphilus metalliredigens, Clostridium drakei, Clostridium carboxidivorans, Clostridium formicoaceticum, Clostridium scatologenes, Moorella thermautotrophica, Acetofilamentum elongatum, Blautia producta, Clostridium glycolicum, Clostridium magnum, Clostridium mayombei, Clostridium methoxybenzovorans, Clostridium acetobutylicum, Clostridium beijerinckii, Acetogenium kivui, Thermoanaerobacter oboediens, Sporomusa ovata, Thermoanaerobacter brockii, Acetobacter methanolicus, Sporomusa termitida, Moorella glycerini, Eubacterium aggregans, Treponema azotonutricium, Escherichia coli, Saccharomyces cerevisiae, Pseudomonas putida, Bacillus sp., Corynebacterium sp., Yarrowia lipolytica, Pichia stipitis, and Terrestrispora glycerinilytica.
[0165] The terms "recombinant microorganism" and "recombinant host cell" are used interchangeably herein and refer to a microorganism that has been genetically modified to express or overexpress an endogenous enzyme, express a heterologous enzyme, such as those included in a vector, an integration construct, or has an alteration in the expression of an endogenous gene. "Alteration" means an upregulation or downregulation of the gene expression or level of an RNA molecule or equivalent RNA molecule encoding one or more polypeptides or polypeptide subunits, or the activity of one or more polypeptides or polypeptide subunits, such that the expression, level, or activity is greater than or less than the expression, level, or activity observed in the absence of the alteration. For example, the term "alteration" can mean "inhibition", but the use of the word "alteration" is not limited to this definition. It should be understood that the terms "recombinant microorganism" and "recombinant host cell" refer not only to a particular recombinant microorganism, but also to the progeny or potential progeny of such a microorganism. Because certain modifications may occur in the next generation due to mutation or environmental influences, such progeny may in fact be different from the parental cell, but are still included within the scope of the terms used herein.
[0166] The cultivation of the microorganisms used in the methods of the present disclosure can be carried out using many methods known in the art for culturing and fermenting substrates with the microorganisms of the present disclosure.
[0167] Fermentation can be carried out in any suitable bioreactor, such as a continuous stirred tank bioreactor, a bubble column bioreactor, an airlift bioreactor, a fluidized bed bioreactor, a packed bed bioreactor, a photobioreactor, an immobilized cell reactor, a trickle bed reactor, a moving bed biofilm reactor, a bubble column, an airlift fermenter, a membrane reactor (such as a hollow fiber membrane bioreactor), etc. In some aspects, the bioreactor includes a first growth reactor in which the microorganisms are cultured and a second fermentation reactor to which the fermentation broth from the growth reactor is fed and in which most of the fermentation products are produced. In some aspects, the bioreactor simultaneously achieves the cultivation of the microorganisms and the production of fermentation products from provided carbon sources such as substrates and / or feedstocks.
[0168] Product
[0169] In some aspects, the engineered microorganisms of the present disclosure produce fermentation products from a feedstock comprising xylose and glucose, wherein the recombinant microorganisms utilize both xylose and glucose simultaneously. In some aspects, the fermentation products produced by the microorganisms include one or more molecules comprising at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 carbons.
[0170] In some aspects, the engineered microorganisms of the present disclosure are capable of producing desired chemicals, such as monoethylene glycol, glycolic acid, C3 compounds (such as acetone, isopropanol, and propylene), amino acids, and polyols. See Koch et al. (WO2017156166A1) and McBride et al. (WO2011022651A1).
[0171] In some aspects, due to the absence of catabolite repression of multiple carbon sources present in a single type of feedstock, the engineered microorganisms of the present disclosure are capable of producing the desired chemicals at maximum yield.
[0172] Genetic modification
[0173] Genetic modifications introduced into one or more of the microorganisms of the present disclosure can alter or eliminate the regulatory sequences of target genes. In some aspects, genetic modifications introduced into one or more of the microorganisms of the present disclosure can introduce new traits or phenotypes into said one or more microorganisms. One or more regulatory sequences can also be inserted, including heterologous regulatory sequences and regulatory sequences found within the genomes of animals, plants, fungi, yeast, bacteria, or viruses corresponding to the microorganisms into which genetic variation is introduced. Moreover, regulatory sequences can be selected based on the expression levels of genes in a microbial culture. The genetic variation can be a predetermined genetic variation specifically introduced into a target site. In some aspects, the genetic variation is a nucleic acid sequence introduced into one or more microbial chromosomes. In some aspects, the genetic variation is a nucleic acid sequence introduced into one or more extrachromosomal nucleic acid sequences. The genetic variation can be a random mutation within a target site. The genetic variation can be an insertion or deletion of one or more nucleotides. In some cases, multiple different genetic variations (e.g., 2, 3, 4, 5, 10 or more) are introduced into one or more isolated bacteria. The multiple genetic variations can be of any of the above types, the same or different types, and any combination. In some cases, multiple different genetic variations are introduced sequentially, with the first genetic variation introduced after a first isolation step, the second genetic variation introduced after a second isolation step, and so on, in order to accumulate multiple desired modifications in the microorganism.
[0174] In some aspects, the genetic modification is a deletion or inactivation of a target gene or regulatory sequence. In some aspects, the deletion is the removal of a target gene or most of the target gene. In some aspects, the deletion is the replacement of a target gene or most of the target gene. In other aspects, the deletion results in a complete loss of function of the target gene. In some aspects, the deletion results in a partial loss of function of the target gene. In some aspects, the loss of function or partial loss of function is determined by comparing the activity of the modified target gene sequence with the activity of the unmodified target gene sequence. In some aspects, the inactivation of the target gene is the result of the deletion or disruption of one or more regulatory or control sequences operably linked to the target sequence. In some aspects, the inactivation of the target gene is the result of disrupting the target gene with a heterologous sequence. In some aspects, the inactivation results in a partial loss of function of the target gene. In some aspects, the inactivation results in a complete loss of function of the target gene.
[0175] In some aspects, one or more of the substrates described in the production of the desired chemical are biosynthesized from carbon sources (e.g., xylose and glucose).
[0176] Generally speaking, the term "genetic variation" refers to any change introduced into a polynucleotide sequence relative to a reference polynucleotide, such as a reference genome or a part thereof, or a reference gene or a part thereof. Genetic variations can be referred to as "mutations", and the sequences or organisms containing genetic variations can be called "genetic variants" or "mutants". Genetic variations can have many effects, such as increasing or decreasing some biological activities, including gene expression, metabolism, and cell signaling. Genetic variations can be introduced specifically into target sites or randomly. A variety of molecular tools and methods can be used to introduce genetic variations. For example, genetic variations can be introduced via polymerase chain reaction mutagenesis, oligonucleotide-directed mutagenesis, saturation mutagenesis, fragment shuffling mutagenesis, homologous recombination, recombineering, lambda red-mediated recombination, CRISPR / Cas9 system, chemical mutagenesis, and combinations thereof. Chemical methods for introducing genetic variations include exposing DNA to chemical mutagens, such as ethyl methanesulfonate (EMS), methyl methanesulfonate (MMS), N-ethyl-N-nitrosourea (ENU), N-methyl-N-nitro-N'-nitroguanidine, 4-nitroquinoline N-oxide, diethyl sulfate, benzo[a]pyrene, cyclophosphamide, bleomycin, triethylenemelamine, acrylamide monomer, nitrogen mustard, vincristine, diepoxides (e.g., diepoxybutane), ICR-170, formaldehyde, procarbazine hydrochloride, ethylene oxide, dimethylnitrosamine, 7,12-dimethylbenz[a]anthracene, chlorambucil, hexamethylphosphoramide, busulfan, and so on. Radiation mutagenic agents include ultraviolet radiation, γ-irradiation, X-rays, and fast neutron bombardment. Genetic variations can also be introduced into nucleic acids using, for example, trimethylpsoralen and ultraviolet light. Random or targeted insertion of mobile DNA elements, such as transposable elements, is another suitable method for generating genetic variations. Genetic variations can be introduced into nucleic acids during amplification in a cell-free in vitro system, such as using polymerase chain reaction (PCR) techniques, such as error-prone PCR. Genetic variations can be introduced into nucleic acids in vitro using DNA shuffling techniques (e.g., exon shuffling, domain swapping, etc.).
[0177] Genetic variation can also be introduced into nucleic acids due to a deficiency of DNA repair enzymes in a cell. For example, the presence of a mutant gene encoding a mutant DNA repair enzyme in a cell is expected to generate a high frequency of mutations (i.e., about 1 mutation / 100 genes - 1 mutation / 10,000 genes) in the genome of that cell. Examples of genes encoding DNA repair enzymes include, but are not limited to, MutH, MutS, MutL, and MutU, as well as their homologs in other species (e.g., MSH1-6, PMS1-2, MLH1, GTBP, ERCC-1, etc.). Exemplary descriptions of various methods for introducing genetic variation are provided, for example, in Stemple (2004) Nature 5:1-7; Chiang et al. (1993) PCR Methods Appl 2(3):210-217; Stemmer (1994) Proc. Natl. Acad. Sci. USA 91:10747-10751; and U.S. Patent Nos. 6,033,861 and 6,773,900.
[0178] Genetic variation introduced into a microorganism can be classified as transgenic, cisgenic, intragenomic, intragenic, intergenic, synthetic, evolved, rearranged, or SNP.
[0179] The CRISPR / Cas9 (Clustered Regularly Interspaced Short Palindromic Repeats) / CRISPR-associated (Cas) system can be used to introduce desired mutations. CRISPR / Cas9 provides acquired immunity against viruses and plasmids for bacteria and archaea by using CRISPR RNAs (crRNAs) to guide the silencing of invading nucleic acids. The Cas9 protein (or its functional equivalents and / or its variants, i.e., Cas9-like proteins) naturally contains DNA endonuclease activity, which depends on the association of the protein with two naturally occurring or synthetic RNA molecules called crRNA and tracrRNA (also called guide RNA). In some cases, the two molecules are covalently linked to form a single molecule (also called single-guide RNA (“sgRNA”). Thus, Cas9 or a Cas9-like protein associates with an RNA targeting DNA (the term encompasses both the bimolecular guide RNA configuration and the single-molecule guide RNA configuration), thereby activating the Cas9 or Cas9-like protein and directing the protein to the target nucleic acid sequence. If the Cas9 or Cas9-like protein retains its native enzymatic function, it will cleave the target DNA to create a double-strand break, which can lead to genomic alterations (i.e., editing: deletions, insertions (when a donor polynucleotide is present), substitutions, etc.), thereby altering gene expression. Some variants of Cas9 (which are encompassed by the term Cas9-like) have been altered such that they have reduced DNA cleavage activity (in some cases, they cleave single strands rather than both strands of the target DNA, while in other cases, their DNA cleavage activity is severely reduced or there is no DNA cleavage activity). Further exemplary descriptions of CRISPR systems for introducing genetic variation can be found, for example, in US8795965.
[0180] Oligonucleotide-directed mutagenesis, also known as site-directed mutagenesis, typically utilizes synthetic DNA primers. Such synthetic primers contain the desired mutation and are complementary to the template DNA surrounding the mutation site, enabling them to hybridize to the DNA in the target gene. The mutation can be a single-base change (point mutation), multiple-base changes, deletions or insertions, or a combination of these. Then, a DNA polymerase that copies the rest of the gene is used to extend the single-stranded primer. The gene thus copied contains the mutation site and can then be introduced into a host cell as a vector and cloned. Finally, mutants can be selected by DNA sequencing to check that they contain the desired mutation.
[0181] Genetic variations can be introduced using error-prone PCR. In this technique, the target gene is amplified using a DNA polymerase under conditions of insufficient fidelity of sequence replication. The result is that the amplified product contains at least one error in the sequence. When a gene is amplified and the resulting product of the reaction contains one or more alterations in the sequence compared to the template molecule, the resulting product is mutagenic compared to the template. Another means of introducing random mutations is to expose cells to chemical mutagens such as nitrosoguanidine or ethyl methanesulfonate (Nestmann, Mutat Res Jun 1975;28(3):323-30), and then isolate the vector containing the gene from the host.
[0182] Homologous recombination mutagenesis involves recombination between an exogenous DNA fragment and a targeted polynucleotide sequence. After a double-strand break occurs, the DNA segment around the 5' end of the break is excised in a process called resection. In a subsequent strand invasion step, the protruding 3' end of the broken DNA molecule then "invades" an unbroken, similar or identical DNA molecule. This method can be used to delete genes, remove exons, add genes, and introduce point mutations. Homologous recombination mutagenesis can be permanent or conditional. Typically, a recombination template is also provided. The recombination template can be a component of another vector, contained in a separate vector, or provided as a separate polynucleotide. In some aspects, the recombination template is designed to act as a template in homologous recombination, such as within or near a target sequence that is nicked or cleaved by a site-specific nuclease. The template polynucleotide can have any suitable length, such as a length of about or more than about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1000 or more nucleotides. In some aspects, the template polynucleotide is complementary to a portion of the polynucleotide containing the target sequence. When optimally aligned, the template polynucleotide can overlap with one or more nucleotides of the target sequence (e.g., about or more than about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 or more nucleotides). In some aspects, when the template sequence and the polynucleotide containing the target sequence are optimally aligned, the nearest nucleotide of the template polynucleotide is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 5000, 10000 or more nucleotides of the target sequence. Non-limiting examples of site-specific nucleases that can be used in homologous recombination methods include zinc finger nucleases, CRISPR nucleases, TALE nucleases, and meganucleases. For further descriptions of the use of such nucleases, see, for example, US8795965 and US20140301990.
[0183] The introduction of genetic variation may be an incomplete process, such that some bacteria in the treated bacterial population carry the desired mutation while other bacteria do not. In some cases, it is desirable to impose a selection pressure in order to enrich for bacteria carrying the desired genetic variation. Traditionally, the selection of successful genetic variants has involved selecting for or against some functionality conferred or eliminated by the genetic variation, such as in the case of inserting an antibiotic resistance gene or eliminating a metabolic activity capable of converting a non-lethal compound into a lethal metabolite. Selection pressure can also be imposed based on the polynucleotide sequence itself, such that only the desired genetic variation need be introduced (e.g., a selection marker is not required either). In such cases, the selection pressure can include lysing genomes lacking the genetic variation introduced at the target site, such that selection effectively occurs against the reference sequence into which the genetic variation is sought to be introduced. Typically, lysis occurs within 100 nucleotides of the target site (e.g., within 75, 50, 25, 10 or fewer nucleotides from the target site, including lysis at or within the target site). Lysis can be directed by a site-specific nuclease selected from the group consisting of zinc finger nucleases, CRISPR nucleases, TALE nucleases (TALENs) or meganucleases. Such methods are similar to methods used to enhance homologous recombination at the target site, except that no template for homologous recombination is provided. Thus, bacteria lacking the desired genetic variation are more likely to undergo lysis that is not repaired, resulting in cell death. The bacteria surviving the selection can then be isolated to assess the improved traits conferred.
[0184] CRISPR nucleases can be used as site-specific nucleases to direct lysis to the target site. By using Cas9 to kill non-mutant cells, improved selection of mutant microorganisms can be obtained. The microorganisms can then be re-isolated from the tissue. The CRISPR nuclease system for eliminating non-variants can employ elements similar to those described above with respect to the introduction of genetic variation, except that no template for homologous recombination is provided. Thus, lysis directed to the target site enhances the death of the affected cells.
[0185] Other options are available for specifically inducing cleavage at a target site, such as zinc finger nucleases, TALE nuclease (TALEN) systems, and meganucleases. Zinc finger nucleases (ZFNs) are artificial DNA endonucleases created by fusing zinc finger DNA-binding domains to DNA cleavage domains. ZFNs can be engineered to target desired DNA sequences, and this enables zinc finger nucleases to cleave unique target sequences. When introduced into cells, ZFNs can be used to edit target DNA (e.g., the cell genome) in the cell by inducing double-strand breaks. Transcription activator-like effector nucleases (TALENs) are artificial DNA endonucleases created by fusing TAL (transcription activator-like) effector DNA-binding domains to DNA cleavage domains. TALENs can be rapidly engineered to bind to almost any desired DNA sequence, and when introduced into cells, TALENs can be used to edit target DNA (e.g., the cell genome) in the cell by inducing double-strand breaks. Meganucleases (homing endonucleases) are endogenous deoxyribonucleases characterized by large recognition sites (double-stranded DNA sequences of 12 to 40 base pairs). Meganucleases can be used to replace, eliminate, or modify sequences in a highly targeted manner. By protein engineering to modify their recognition sequences, the targeting sequences can be altered. Meganucleases can be used to modify all genome types, whether bacterial, plant, or animal, and are generally divided into four families: the LAGLIDADG family, the GIY-YIG family, the His-Cyst box family, and the HNH family. Exemplary homing endonucleases include I-SceI, I-CeuI, PI-PspI, PI-Sce, I-SceIV, I-CsmI, I-PanI, I-SceII, I-PpoI, I-SceIII, I-CreI, I-TevI, I-TevII, and I-TevIII.
[0186] In some aspects, the microorganism is a recombinant microorganism. In some aspects, the microorganism has been genetically modified to produce monoethylene glycol. In some aspects, the microorganism has been genetically modified to produce one or more three-carbon compounds, such as acetone, isopropanol, and propylene. In some aspects, the microorganism has been genetically modified to co-produce monoethylene glycol and one or more three-carbon compounds. In some aspects, the microorganism has been genetically modified with a microbial biosynthetic pathway to produce one or more of monoethylene glycol, acetone, isopropanol, and propylene. See Koch et al. (WO2017156166A1), which relates to the prior art for engineering microorganisms to produce one or more of monoethylene glycol, acetone, isopropanol, and propylene from renewable feedstocks.
[0187] In some aspects, the microorganism has been genetically modified by introducing the xylonate pathway. In some aspects, the microorganism has been genetically modified by introducing the xylulose phosphate pathway. In some aspects, the microorganism has been genetically modified by introducing the ribulose phosphate pathway. See Koch et al.
[0188] A recombinant microorganism capable of producing a fermentation product from a feedstock comprising xylose and glucose, wherein the recombinant microorganism simultaneously utilizes xylose and glucose
[0189] In some aspects, the recombinant microorganism comprises one or more of the following: (a) deletion or inactivation of a pentose ATP-binding transporter in the genome of the microorganism such that the transporter is not expressed; (b) one or more endogenous or exogenous nucleic acid sequences that encode at least one C5 sugar symporter and are operably linked to one or more constitutive promoters; wherein the C5 sugar symporter comprises: (1) a xylose symporter and / or (2) an arabinose symporter; (c) one or more endogenous or exogenous nucleic acid sequences that (1) encode xylose isomerase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more xylulokinases, and / or (2) encode xylose dehydrogenase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases.
[0190] In some aspects, the C5 sugar symporter is a symporter capable of transporting a 5-carbon sugar. In some aspects, the 5-carbon sugar can be, but is not limited to, xylose, arabinose, or ribose.
[0191] Normal production of MEG and / or acetone
[0192] In some aspects, the recombinant microorganism comprises (a) deletion or inactivation of aldA, araFGH, and xylFGH in the genome of the parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule that is operably linked to one or more constitutive promoters and encodes a C5 sugar symporter; wherein the recombinant microorganism expresses a MEG and / or acetone production pathway.
[0193] In some aspects, the microorganism further comprises a deletion or inactivation of glycollate dehydrogenase glcDEF. In some aspects, the C5 symporter is controlled by the GAPDH promoter at the araFGH locus. In some aspects, the one or more nucleic acid molecules encoding the GAPDH promoter comprise the nucleic acid sequence set forth in SEQ ID NO:95. In some aspects, the C5 sugar symporter is the xylose symporter XylE. In some aspects, the one or more nucleic acid molecules encoding the XylE comprise the nucleic acid sequence set forth in SEQ ID NO:48. In some aspects, the one or more amino acid sequences encoding the XylE comprise the amino acid sequence set forth in SEQ ID NO:49. In some aspects, the xylose symporter is endogenous to the microorganism. In some aspects, the C5 sugar symporter is the arabinose symporter AraE. In some aspects, the arabinose symporter is endogenous to the microorganism. In some aspects, the one or more nucleic acid molecules encoding the AraE comprise the nucleic acid sequence set forth in SEQ ID NO:46. In some aspects, the one or more amino acid sequences encoding the AraE comprise the amino acid sequence set forth in SEQ ID NO:47. In some aspects, the xylose is insensitive to catabolite repression by other monosaccharides.
[0194] In some aspects, the microorganism comprises a functional phosphotransferase system. In some aspects, the microorganism comprises the native wild-type nucleic acid sequence encoding the cAMP receptor protein (CRP). In some aspects, the one or more nucleic acid molecules encoding the CRP comprise the nucleic acid sequence set forth in SEQ ID NO:9. In some aspects, the one or more amino acid sequences encoding the CRP comprise the amino acid sequence set forth in SEQ ID NO:10. In some aspects, the constitutive overexpression of the xylose symporter enables the continuous input of xylose from the feedstock into the microorganism. In some aspects, the constitutive overexpression of the arabinose symporter enables the continuous input of xylose from the feedstock into the microorganism. In some aspects, the continuous xylose input occurs independently of the presence of other sugars in the feedstock.
[0195] Production of MEG and / or acetone, including the xylulose pathway
[0196] In some aspects, the recombinant microorganism comprises (a) deletion or inactivation of aldA, araFGH, and xylFGH in the genome of a parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism comprises a MEG production pathway having one or more of the following: (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose isomerase and / or ketohexokinase and / or fructose bisphosphate aldolase and operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding glycoaldehyde reductase, which catalyzes the conversion of glycoaldehyde to MEG; and (e) deletion or inactivation of one or more xylulokinases in the genome of the parental microorganism; wherein the recombinant microorganism expresses a MEG and / or acetone production pathway.
[0197] In some aspects, (c) and (d) are in an operon controlled by the proD promoter. In some aspects, the one or more nucleic acid molecules encoding the proD promoter comprise the nucleic acid sequence set forth in SEQ ID NO:53. In some aspects, the xylose isomerase is XylA. In some aspects, the one or more nucleic acid molecules encoding the XylA comprise the nucleic acid sequence set forth in SEQ ID NO:5. In some aspects, the one or more amino acid sequences encoding the XylA comprise the amino acid sequence set forth in SEQ ID NO:6. In some aspects, the xylose isomerase is endogenous to the microorganism. In some aspects, the ketohexokinase is Khk-C. In some aspects, the ketohexokinase is from Homo sapiens. In some aspects, the ketohexokinase is heterologous to the microorganism. In some aspects, the one or more nucleic acid molecules encoding the Khk-C comprise the nucleic acid sequence set forth in SEQ ID NO:11. In some aspects, the one or more amino acid sequences encoding the Khk-C comprise the amino acid sequence set forth in SEQ ID NO:12. In some aspects, the fructose-bisphosphate aldolase is aldoB. In some aspects, the fructose-bisphosphate aldolase is from Homo sapiens. In some aspects, the fructose-bisphosphate aldolase is heterologous to the microorganism. In some aspects, the one or more nucleic acid molecules encoding the aldoB comprise the nucleic acid sequence set forth in SEQ ID NO:50. In some aspects, the one or more amino acid sequences encoding the aldoB comprise the amino acid sequence set forth in SEQ ID NO:51. In some aspects, the glycoaldehyde reductase is endogenous to the microorganism. In some aspects, the glycoaldehyde reductase is fucO. In some aspects, the one or more nucleic acid molecules encoding the fucO comprise the nucleic acid sequence set forth in SEQ ID NO:52. In some aspects, the one or more amino acid sequences encoding the fucO comprise the amino acid sequence set forth in SEQ ID NO:98. In some aspects, the xylulokinase is XylB. In some aspects, the one or more nucleic acid molecules encoding the XylB comprise the nucleic acid sequence set forth in SEQ ID NO:13. In some aspects, the one or more amino acid sequences encoding the XylB comprise the amino acid sequence set forth in SEQ ID NO:14.
[0198] Production of MEG and / or acetone, including the xylonate pathway
[0199] In some aspects, the recombinant microorganism comprises (a) a deletion or inactivation of aldA, araFGH, and xylFGH in the genome of a parental microorganism; and (b) the expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism comprises a MEG production pathway having one or more of the following: (c) the expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose dehydrogenase and / or xylonic acid lactonase and / or xylose dehydratase and operably linked to one or more constitutive promoters; (d) the expression of at least one endogenous or exogenous nucleic acid molecule encoding glycoaldehyde reductase, which catalyzes the conversion of glycoaldehyde to MEG; and (e) a deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases in the genome of the parental microorganism; and wherein the recombinant microorganism expresses a MEG and / or acetone production pathway.
[0200] In some aspects, the xylose dehydrogenase is from Caulobacter crescentus, Burkholderia xenovorans, or Halomonas volcanii. In some aspects, the xylose dehydrogenase is heterologous to the microorganism. In some aspects, the one or more nucleic acid molecules encoding the Caulobacter crescentus xylose dehydrogenase comprise the nucleic acid sequence set forth in SEQ ID NO:15. In some aspects, the one or more amino acid sequences encoding the Caulobacter crescentus xylose dehydrogenase comprise the amino acid sequence set forth in SEQ ID NO:16. In some aspects, the one or more nucleic acid molecules encoding the Burkholderia xenovorans xylose dehydrogenase comprise the nucleic acid sequence set forth in SEQ ID NO:97. In some aspects, the one or more amino acid sequences encoding the Burkholderia xenovorans xylose dehydrogenase comprise the amino acid sequence set forth in SEQ ID NO:17. In some aspects, the one or more nucleic acid molecules encoding the Halomonas volcanii xylose dehydrogenase comprise the nucleic acid sequence set forth in SEQ ID NO:18. In some aspects, the one or more amino acid sequences encoding the Halomonas volcanii xylose dehydrogenase comprise the amino acid sequence set forth in SEQ ID NO:19.
[0201] In some aspects, the xylonic acid lactonase is from Caulobacter crescentus, Burkholderia xenovorans, or Halomonas volcanii. In some aspects, the xylonic acid lactonase is heterologous to the microorganism. In some aspects, the xylonic acid lactonase is endogenous to the microorganism. In some aspects, the one or more nucleic acid molecules encoding the Caulobacter crescentus xylonic acid lactonase comprise the nucleic acid sequence set forth in SEQ ID NO:54. In some aspects, the one or more amino acid sequences encoding the Caulobacter crescentus xylonic acid lactonase comprise the amino acid sequence set forth in SEQ ID NO:55. In some aspects, the one or more nucleic acid molecules encoding the Burkholderia xenovorans xylonic acid lactonase comprise the nucleic acid sequence set forth in SEQ ID NO:56. In some aspects, the one or more amino acid sequences encoding the Burkholderia xenovorans xylonic acid lactonase comprise the amino acid sequence set forth in SEQ ID NO:57. In some aspects, the one or more nucleic acid molecules encoding the Halomonas volcanii xylonic acid lactonase comprise the nucleic acid sequence set forth in SEQ ID NO:58. In some aspects, the one or more amino acid sequences encoding the Halomonas volcanii xylonic acid lactonase comprise the amino acid sequence set forth in SEQ ID NO:59.
[0202] In some aspects, the xylose dehydratase is from Caulobacter crescentus, Burkholderia xenovorans, or Halomonas volcanii. In some aspects, the xylose dehydratase is heterologous to the microorganism. In some aspects, the xylose dehydratase is endogenous to the microorganism. In some aspects, the one or more nucleic acid molecules encoding the Caulobacter crescentus xylose dehydratase comprise the nucleic acid sequence set forth in SEQ ID NO: 60. In some aspects, the one or more amino acid sequences encoding the Caulobacter crescentus xylose dehydratase comprise the amino acid sequence set forth in SEQ ID NO: 61. In some aspects, the one or more nucleic acid molecules encoding the Burkholderia xenovorans xylose dehydratase comprise the nucleic acid sequence set forth in SEQ ID NO: 62. In some aspects, the one or more amino acid sequences encoding the Burkholderia xenovorans xylose dehydratase comprise the amino acid sequence set forth in SEQ ID NO: 63. In some aspects, the one or more nucleic acid molecules encoding the Halomonas volcanii xylose dehydratase comprise the nucleic acid sequence set forth in SEQ ID NO: 64. In some aspects, the one or more amino acid sequences encoding the Halomonas volcanii xylose dehydratase comprise the amino acid sequence set forth in SEQ ID NO: 65. In some aspects, the glyoxaldehyde reductase is endogenous to the microorganism. In some aspects, the glyoxaldehyde reductase is fucO. In some aspects, the glyoxaldehyde reductase is heterologous to the microorganism. In some aspects, the xylose isomerase is XylA. In some aspects, the xylulokinase is XylB.
[0203] Production of MEG and / or acetone, including the xylulose pathway
[0204] In some aspects, the recombinant microorganism comprises (a) a deletion or inactivation of aldA, araFGH, and xylFGH in the genome of a parental microorganism; and (b) the expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism comprises a MEG production pathway having one or more of the following: (c) the expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose isomerase and / or ketohexokinase and / or fructose bisphosphate aldolase and operably linked to one or more constitutive promoters; (d) the expression of at least one endogenous or exogenous nucleic acid molecule encoding hydroxyacetaldehyde reductase, which catalyzes the conversion of hydroxyacetaldehyde to MEG; and (e) a deletion or inactivation of one or more xylulokinases in the genome of the parental microorganism; wherein the recombinant microorganism comprises an acetone production pathway having one or more of the following: (f) the expression of at least one exogenous nucleic acid molecule encoding acetoacetyl-CoA thiolase; (g) the expression of at least one exogenous nucleic acid molecule encoding acetate:acetoacetyl-CoA transferase; and (h) the expression of at least one exogenous nucleic acid molecule encoding acetoacetate decarboxylase, which catalyzes the conversion of acetoacetate to acetone, wherein the recombinant microorganism expresses a MEG and / or acetone production pathway.
[0205] In some aspects, (f), (g), and (h) are in an operon controlled by the OXB11 promoter. In some aspects, the one or more nucleic acid molecules encoding OXB11 comprise the nucleic acid sequence set forth in SEQ ID NO:78. In some aspects, the acetoacetyl-CoA thiolase is Thl. In some aspects, the thiolase is from Clostridium acetobutylicum or Clostridium beijerinckii. In some aspects, the one or more nucleic acid molecules encoding the Clostridium acetobutylicum thl thiolase comprise the nucleic acid sequence set forth in SEQ ID NO:68. In some aspects, the one or more amino acid sequences encoding the Clostridium acetobutylicum thl thiolase comprise the amino acid sequence set forth in SEQ ID NO:69. In some aspects, the one or more nucleic acid molecules encoding the Clostridium beijerinckii thl thiolase comprise the nucleic acid sequence set forth in SEQ ID NO:66. In some aspects, the one or more amino acid sequences encoding the Clostridium beijerinckii thl thiolase comprise the amino acid sequence set forth in SEQ ID NO:67. In some aspects, the acetate:acetoacetyl-CoA transferase is AtoDA. In some aspects, the acetoacetate decarboxylase is Adc. In some aspects, the decarboxylase is from Clostridium acetobutylicum or Clostridium beijerinckii. In some aspects, the one or more nucleic acid molecules encoding the Clostridium acetobutylicum Adc acetoacetate decarboxylase comprise the nucleic acid sequence set forth in SEQ ID NO:74. In some aspects, the one or more amino acid sequences encoding the Clostridium acetobutylicum Adc acetoacetate decarboxylase comprise the amino acid sequence set forth in SEQ ID NO:75. In some aspects, the one or more nucleic acid molecules encoding the Clostridium beijerinckii Adc acetoacetate decarboxylase comprise the nucleic acid sequence set forth in SEQ ID NO:76. In some aspects, the one or more amino acid sequences encoding the Clostridium beijerinckii Adc acetoacetate decarboxylase comprise the amino acid sequence set forth in SEQ ID NO:77.
[0206] Production of MEG and / or acetone specific to the xylonate pathway
[0207] In some aspects, the recombinant microorganism comprises (a) a deletion or inactivation of aldA, araFGH, and xylFGH in the genome of a parental microorganism; and (b) the expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism comprises a MEG production pathway having one or more of the following: (c) the expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose dehydrogenase and / or xylonic acid lactonase and / or xylose dehydratase and operably linked to one or more constitutive promoters; (d) the expression of at least one endogenous or exogenous nucleic acid molecule encoding hydroxyacetaldehyde reductase, which catalyzes the conversion of hydroxyacetaldehyde to MEG; and (e) a deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases in the genome of the parental microorganism; wherein the recombinant microorganism comprises an acetone production pathway having one or more of the following: (f) the expression of at least one exogenous nucleic acid molecule encoding acetoacetyl-CoA thiolase; (g) the expression of at least one exogenous nucleic acid molecule encoding acetate:acetoacetyl-CoA transferase; and (h) the expression of at least one exogenous nucleic acid molecule encoding acetoacetate decarboxylase, which catalyzes the conversion of acetoacetate to acetone, wherein the recombinant microorganism expresses a MEG and / or an acetone production pathway.
[0208] In some aspects, (f), (g), and (h) are in an operon controlled by the OXB11 promoter. In some aspects, the acetoacetyl-CoA thiolase is Thl. In some aspects, the thiolase is from Clostridium acetobutylicum. In some aspects, the acetate:acetoacetyl-CoA transferase is AtoDA. In some aspects, the one or more nucleic acid molecules encoding the AtoD subunit α of the acetate:acetoacetyl-CoA transferase comprise the nucleic acid sequence set forth in SEQ ID NO:70. In some aspects, the one or more amino acid sequences encoding the AtoD subunit α of the acetate:acetoacetyl-CoA transferase comprise the amino acid sequence set forth in SEQ ID NO:72. In some aspects, the one or more nucleic acid molecules encoding the AtoD subunit β of the acetate:acetoacetyl-CoA transferase comprise the nucleic acid sequence set forth in SEQ ID NO:71. In some aspects, the one or more amino acid sequences encoding the AtoD subunit β of the acetate:acetoacetyl-CoA transferase comprise the amino acid sequence set forth in SEQ ID NO:73. In some aspects, the acetoacetate decarboxylase is Adc. In some aspects, the decarboxylase is from Clostridium acetobutylicum or Clostridium beijerinckii. In some aspects, the one or more nucleic acid molecules encoding the Adc acetoacetate decarboxylase of Clostridium acetobutylicum comprise the nucleic acid sequence set forth in SEQ ID NO:74. In some aspects, the one or more amino acid sequences encoding the Adc acetoacetate decarboxylase of Clostridium acetobutylicum comprise the amino acid sequence set forth in SEQ ID NO:75. In some aspects, the one or more nucleic acid molecules encoding the Adc acetoacetate decarboxylase of Clostridium beijerinckii comprise the nucleic acid sequence set forth in SEQ ID NO:76. In some aspects, the one or more amino acid sequences encoding the Adc acetoacetate decarboxylase of Clostridium beijerinckii comprise the amino acid sequence set forth in SEQ ID NO:77.
[0209] Production of isopropanol
[0210] In some aspects, a recombinant microorganism is capable of producing isopropanol from any one or more feedstocks capable of producing acetone. In some aspects, a recombinant microorganism that has been engineered to produce acetone is further engineered to express at least one exogenous nucleic acid molecule encoding an alcohol dehydrogenase that catalyzes the conversion of acetone to isopropanol. In some aspects, the foregoing acetone production (f), (g), and (h) are further modified to (i) – expression of at least one exogenous nucleic acid molecule encoding an alcohol dehydrogenase that catalyzes the conversion of acetone to isopropanol. In some aspects, the one or more nucleic acid molecules encoding the alcohol dehydrogenase comprise the nucleic acid sequence set forth in SEQ ID NO:93. In some aspects, the one or more amino acid sequences encoding the alcohol dehydrogenase comprise the amino acid sequence set forth in SEQ ID NO:94.
[0211] Production of glyoxylic acid
[0212] In some aspects, a recombinant microorganism capable of producing glyoxylic acid from a feedstock comprising xylose and glucose, wherein the recombinant microorganism co-utilizes xylose and glucose, comprises one or more of the following: (a) deletion or inactivation of fucO, yqhD (SEQ ID NO:1 or SEQ ID NO:2), araFGH, and xylFGH in the genome of the parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism also expresses one or more glyoxylic acid production pathways.
[0213] In some aspects, the one or more nucleic acid molecules encoding the AraF periplasmic binding protein subunit of AraFGH comprise the nucleic acid sequence set forth in SEQ ID NO:20. In some aspects, the one or more amino acid sequences encoding the AraF periplasmic binding protein subunit of AraFGH comprise the amino acid sequence set forth in SEQ ID NO:23. In some aspects, the one or more nucleic acid molecules encoding the AraG ATP-binding subunit of AraFGH comprise the nucleic acid sequence set forth in SEQ ID NO:21. In some aspects, the one or more amino acid sequences encoding the AraG ATP-binding subunit of AraFGH comprise the amino acid sequence set forth in SEQ ID NO:24. In some aspects, the one or more nucleic acid molecules encoding the AraH membrane subunit of AraFGH comprise the nucleic acid sequence set forth in SEQ ID NO:22. In some aspects, the one or more amino acid sequences encoding the AraH membrane subunit of AraFGH comprise the amino acid sequence set forth in SEQ ID NO:25.
[0214] In some aspects, the one or more nucleic acid molecules encoding the xylF periplasmic binding protein subunit of xylFGH comprise the nucleic acid sequence set forth in SEQ ID NO:26. In some aspects, the one or more amino acid sequences encoding the xylF periplasmic binding protein subunit of xylFGH comprise the amino acid sequence set forth in SEQ ID NO:29. In some aspects, the one or more nucleic acid molecules encoding the xylG ATP-binding subunit of xylFGH comprise the nucleic acid sequence set forth in SEQ ID NO:27. In some aspects, the one or more amino acid sequences encoding the xylG ATP-binding subunit of xylFGH comprise the amino acid sequence set forth in SEQ ID NO:30. In some aspects, the one or more nucleic acid molecules encoding the xylH membrane subunit of xylFGH comprise the nucleic acid sequence set forth in SEQ ID NO:28. In some aspects, the one or more amino acid sequences encoding the xylH membrane subunit of xylFGH comprise the amino acid sequence set forth in SEQ ID NO:31.
[0215] In some aspects, the microorganism also comprises a deletion or inactivation of glcDEF. In some aspects, the one or more nucleic acid molecules encoding the putative FAD-linked subunit GlcD comprise the nucleic acid sequence set forth in SEQ ID NO:79. In some aspects, the one or more amino acid sequences encoding the putative FAD-linked subunit GlcD comprise the amino acid sequence set forth in SEQ ID NO:82. In some aspects, the one or more nucleic acid molecules encoding the putative FAD-binding subunit GlcE comprise the nucleic acid sequence set forth in SEQ ID NO:80. In some aspects, the one or more amino acid sequences encoding the putative FAD-binding subunit GlcE comprise the amino acid sequence set forth in SEQ ID NO:83. In some aspects, the one or more nucleic acid molecules encoding the putative iron-sulfur subunit GlcF comprise the nucleic acid sequence set forth in SEQ ID NO:81. In some aspects, the one or more amino acid sequences encoding the putative iron-sulfur subunit GlcF comprise the amino acid sequence set forth in SEQ ID NO:84. In some aspects, the microorganism also comprises a deletion or inactivation of aldehyde reductase dkgA. In some aspects, the one or more nucleic acid molecules encoding dkgA comprise the nucleic acid sequence set forth in SEQ ID NO:85. In some aspects, the one or more amino acid sequences encoding dkgA comprise the amino acid sequence set forth in SEQ ID NO:86. In some aspects, the microorganism also comprises a deletion or inactivation of aldehyde reductase yahK. In some aspects, the one or more nucleic acid molecules encoding yahK comprise the nucleic acid sequence set forth in SEQ ID NO:87. In some aspects, the one or more amino acid sequences encoding yahK comprise the amino acid sequence set forth in SEQ ID NO:88. In some aspects, the xylose symporter is controlled by the GAPDH promoter at the araFGH locus. In some aspects, the C5 sugar symporter is the xylose symporter XylE. In some aspects, the one or more nucleic acid molecules encoding xylE comprise the nucleic acid sequence set forth in SEQ ID NO:48. In some aspects, the one or more amino acid sequences encoding xylE comprise the amino acid sequence set forth in SEQ ID NO:49. In some aspects, the xylose symporter is endogenous to the microorganism. In some aspects, the C5 sugar symporter is the arabinose symporter AraE. In some aspects, the one or more nucleic acid molecules encoding araE comprise the nucleic acid sequence set forth in SEQ ID NO:46. In some aspects, the one or more amino acid sequences encoding araE comprise the amino acid sequence set forth in SEQ ID NO:47. In some aspects, the arabinose symporter is endogenous to the microorganism.In some aspects, the uptake of xylose is insensitive to catabolite repression of other monosaccharides. In some aspects, the microorganism comprises a functional phosphotransferase system. In some aspects, the microorganism comprises a native wild-type nucleic acid sequence encoding a cAMP receptor protein (CRP). In some aspects, the one or more nucleic acid molecules encoding the CRP comprise the nucleic acid sequence set forth in SEQ ID NO:9. In some aspects, the one or more amino acid sequences encoding the CRP comprise the amino acid sequence set forth in SEQ ID NO:10. In some aspects, constitutive overexpression of the xylose symporter enables continuous input of xylose from the feedstock into the microorganism. In some aspects, constitutive overexpression of the arabinose symporter enables continuous input of xylose from the feedstock into the microorganism. In some aspects, continuous xylose input occurs independently of the presence of other sugars in the feedstock.
[0216] Glycolate production, including the xylulose pathway
[0217] In some aspects, a recombinant microorganism capable of producing glycolate from a feedstock comprising xylose and glucose, wherein the recombinant microorganism simultaneously utilizes xylose and glucose, comprises one or more of the following: (a) deletion or inactivation of fucO, yqhD, yahK, dkgA, araFGH, and xylFGH in the genome of the parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism further expresses one or more glycolate production pathways having one or more of the following: (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose isomerase and / or ketohexokinase and / or fructose bisphosphate aldolase and operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding hydroxyacetaldehyde dehydrogenase that catalyzes the conversion of hydroxyacetaldehyde to glycolate; and (e) deletion or inactivation of one or more xylulokinases in the genome of the parental microorganism.
[0218] In some aspects, (c) and (d) are in an operon controlled by the proD promoter. In some aspects, the one or more nucleic acid molecules encoding the proD promoter comprise the nucleic acid sequence set forth in SEQ ID NO:53. In some aspects, the xylose isomerase is XylA. In some aspects, the one or more nucleic acid molecules encoding the xylA comprise the nucleic acid sequence set forth in SEQ ID NO:5. In some aspects, the one or more amino acid sequences encoding the xylA comprise the amino acid sequence set forth in SEQ ID NO:6. In some aspects, the xylose isomerase is endogenous to the microorganism. In some aspects, the xylose isomerase is heterologous to the microorganism. In some aspects, the ketohexokinase is Khk-C. In some aspects, the one or more nucleic acid molecules encoding the khk-C comprise the nucleic acid sequence set forth in SEQ ID NO:11. In some aspects, the one or more amino acid sequences encoding the khk-C comprise the amino acid sequence set forth in SEQ ID NO:12. In some aspects, the ketohexokinase is from Homo sapiens. In some aspects, the fructose-bisphosphate aldolase is aldoB. In some aspects, the one or more nucleic acid molecules encoding the alsoB comprise the nucleic acid sequence set forth in SEQ ID NO:50. In some aspects, the one or more amino acid sequences encoding the aldoB comprise the amino acid sequence set forth in SEQ ID NO:51. In some aspects, the fructose-bisphosphate aldolase is from Homo sapiens. In some aspects, the glyoxylate dehydrogenase is endogenous to the microorganism. In some aspects, the glyoxylate dehydrogenase is heterologous to the microorganism. In some aspects, the glyoxylate dehydrogenase is aldA. In some aspects, the one or more nucleic acid molecules encoding the aldA comprise the nucleic acid sequence set forth in SEQ ID NO:3. In some aspects, the one or more amino acid sequences encoding the aldA comprise the amino acid sequence set forth in SEQ ID NO:4. In some aspects, the xylulokinase is XylB. In some aspects, the one or more nucleic acid molecules encoding the xylB comprise the nucleic acid sequence set forth in SEQ ID NO:13. In some aspects, the one or more amino acid sequences encoding the xylB comprise the amino acid sequence set forth in SEQ ID NO:14.
[0219] Production of glycolic acid, including the xylonic acid pathway
[0220] In some aspects, a recombinant microorganism capable of producing glycolic acid from a feedstock comprising xylose and glucose, wherein the recombinant microorganism co-utilizes xylose and glucose, comprises one or more of the following: (a) deletion or inactivation of fucO, yqhD, yahK, dkgA, araFGH, and xylFGH in the genome of a parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism also expresses one or more glycolic acid production pathways having one or more of the following: (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose dehydrogenase and / or xylonic acid lactonase and / or xylose dehydratase and operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding glyoxaldehyde dehydrogenase that catalyzes the conversion of glyoxaldehyde to glycolic acid; and (e) deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases in the genome of the parental microorganism.
[0221] In some aspects, (c) and (d) are controlled by the proD promoter. In some aspects, the xylose isomerase is XylA. In some aspects, the xylulokinase is XylB. In some aspects, the xylose dehydrogenase is from Caulobacter crescentus, Burkholderia xenovorans, or Halomonas volcanii. In some aspects, the xylose dehydrogenase is heterologous to the microorganism. In some aspects, the xylonic acid lactonase is from Caulobacter crescentus, Burkholderia xenovorans, or Halomonas volcanii. In some aspects, the xylonic acid lactonase is heterologous to the microorganism. In some aspects, the xylonic acid lactonase is endogenous to the microorganism. In some aspects, the glyoxaldehyde dehydrogenase is aldA. In some aspects, the glyoxaldehyde dehydrogenase is endogenous to the microorganism.
[0222] Production of glycolic acid (alternative pathway), including the xylulose pathway -
[0223] In some aspects, a recombinant microorganism capable of producing glycolic acid from a feedstock comprising xylose and glucose, wherein the recombinant microorganism co-utilizes xylose and glucose, comprises one or more of the following: (a) deletion or inactivation of fucO, yqhD, yahK, dkgA, araFGH, and xylFGH in the genome of a parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism further expresses one or more glycolic acid-producing pathways having one or more of the following: (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose isomerase and / or ketohexokinase and / or fructose bisphosphate aldolase and operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding hydroxyacetaldehyde dehydrogenase, which catalyzes the conversion of hydroxyacetaldehyde to glycolic acid; and (e) deletion or inactivation of one or more xylulokinases in the genome of the parental microorganism; and wherein the microorganism further expresses a glycolic acid-producing pathway having one or more of the following: (f) expression of at least one endogenous or exogenous nucleic acid molecule encoding isocitrate lyase; and / or (g) expression of at least one endogenous or exogenous nucleic acid molecule encoding glyoxylate reductase. In some aspects, (f) and (g) are in an operon controlled by the OXB20 promoter. In some aspects, the one or more nucleic acid molecules encoding the OXB20 promoter comprise the nucleic acid sequence set forth in SEQ ID NO:96. In some aspects, the isocitrate lyase is AceA. In some aspects, the glyoxylate reductase is YcdW. In some aspects, the one or more nucleic acid molecules encoding ycdW comprise the nucleic acid sequence set forth in SEQ ID NO:91. In some aspects, the one or more amino acid sequences encoding ycdW comprise the amino acid sequence set forth in SEQ ID NO:92.
[0224] Production of glycolic acid (alternative pathway), including the xylonate pathway
[0225] In some aspects, a recombinant microorganism capable of producing glycolic acid from a feedstock comprising xylose and glucose, wherein the recombinant microorganism co-utilizes xylose and glucose, comprises one or more of the following: (a) deletion or inactivation of fucO, yqhD, yahK, dkgA, araFGH, and xylFGH in the genome of a parental microorganism; and (b) expression of at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter; wherein the recombinant microorganism also expresses one or more glycolic acid production pathways having one or more of the following: (c) expression of one or more endogenous or exogenous nucleic acid sequences encoding xylose dehydrogenase and / or xylonic acid lactonase and / or xylose dehydratase and operably linked to one or more constitutive promoters; (d) expression of at least one endogenous or exogenous nucleic acid molecule encoding glyoxaldehyde dehydrogenase that catalyzes the conversion of glyoxaldehyde to glycolic acid; and (e) deletion or inactivation of one or more xylose isomerases and / or one or more xylulokinases in the genome of the parental microorganism; and wherein the microorganism also expresses a glycolic acid production pathway having one or more of the following: (f) expression of at least one endogenous or exogenous nucleic acid molecule encoding isocitrate lyase; and / or (g) expression of at least one endogenous or exogenous nucleic acid molecule encoding glyoxylate reductase.
[0226] In some aspects, (f) and (g) are in an operon controlled by the OXB20 promoter. In some aspects, the isocitrate lyase is AceA. In some aspects, the glyoxylate reductase is YcdW.
[0227] A recombinant microorganism comprising the sequences and modifications described herein
[0228] In some aspects, the present disclosure broadly relates to the recombinant microorganism of any one of the foregoing aspects, wherein the recombinant microorganism is derived from a parental microorganism selected from the group consisting of Clostridium, Clostridium ljungdahlii, Clostridium autoethanogenum, Clostridium ragsdalei, Eubacterium limosum, Butyribacterium methylotrophicum, Moorella thermoacetica, Clostridium aceticum, Acetobacterium woodii, Alkaliphilus bacchii, Clostridium drakei, Clostridium carboxidivorans, Clostridium formicoaceticum, Clostridium scatologenes, Moorella thermoautotrophica, Acetofilamentum longum, Blautia producta, Clostridium glycolicum, Clostridium magnum, Clostridium mayombei, Clostridium methoxybenzovorans, Clostridium acetobutylicum, Clostridium beijerinckii, Acetogenium kivui, Thermoanaerobacter ovoideus, Thermoanaerobacter brockii, Acetobacterium methanolicum, Sporomusa termitida, Moorella glycerini, Eubacterium aggregans, Treponema azotonutricium, Escherichia coli, Saccharomyces cerevisiae, Pseudomonas putida, Bacillus sp., Corynebacterium sp., Yarrowia lipolytica, Pichia stipitis, and Terripora glycerilytica. In some aspects, the parental microorganism is Escherichia coli.
[0229] In some aspects, the enzymes, proteins, promoters, and nucleic acids of the present disclosure are summarized in Table 1.
[0230] Table 1: Proteins and Nucleic Acids of the Present Disclosure
[0231]
[0232]
[0233]
[0234]
[0235]
[0236] Methods for Detecting Genetic Modifications
[0237] The present disclosure teaches primers, probes, and assays that can be used to detect the microorganisms taught herein. In some aspects, the present disclosure provides methods for detecting a WT parental strain. In other aspects, the present disclosure provides methods for detecting engineered or modified microorganisms derived from a parental strain or a WT strain. In some aspects, the present disclosure provides methods for identifying genetic alterations in microorganisms.
[0238] In some aspects, the genome engineering methods of the present disclosure result in the production of non-natural nucleotide "linker" sequences in modified microorganisms. These non-naturally occurring nucleotide linkages can be used as a type of diagnostic that indicates the presence of specific genetic alterations in the microorganisms taught herein.
[0239] The techniques of the present disclosure are capable of detecting these non-naturally occurring nucleotide linkages by using specialized quantitative PCR methods, including uniquely designed primers and probes. In some aspects, the probes of the present disclosure bind to non-naturally occurring nucleotide linker sequences. In some aspects, conventional PCR is employed. In other aspects, real-time PCR is employed. In some aspects, quantitative PCR (qPCR) is employed. In some aspects, PCR methods are used to identify heterologous sequences that have been inserted into the genomic DNA or episomal DNA of a microorganism.
[0240] Accordingly, the present disclosure can cover two common methods for real-time detection of PCR products: (1) non-specific fluorescent dyes that intercalate into any double-stranded DNA, and (2) sequence-specific DNA probes composed of oligonucleotides that are labeled with a fluorescent reporter gene, and the fluorescent reporter gene can only be detected after the probe hybridizes with its complementary sequence. In some aspects, only non-naturally occurring nucleotide linkages are amplified by the taught primers, and thus detection can be carried out by non-specific dyes or by using specific hybridization probes. In other aspects, the primers of the present disclosure are selected such that the primers flank either side of the linkage sequence, such that if an amplification reaction occurs, the linkage sequence is present.
[0241] Aspects of the present disclosure relate to non-naturally occurring nucleotide linkage sequence molecules themselves, and other nucleotide molecules that are capable of binding to the non-naturally occurring nucleotide linkage sequence under mild to stringent hybridization conditions. In some aspects, nucleotide molecules that are capable of binding to the non-naturally occurring nucleotide linkage sequence under mild to stringent hybridization conditions are referred to as "nucleotide probes".
[0242] In some aspects, genomic DNA can be extracted from a sample and used to quantify the presence of the microorganisms of the present disclosure by using qPCR. The primers employed in the qPCR reaction can be primers designed by Primer Blast (https: / / www.ncbi.nlm.nih.gov / tools / primer-blast / ) for amplifying unique regions of the wild-type genome or unique regions of engineered non-intergeneric mutant strains. The qPCR reaction can be carried out using the SYBR GreenER qPCR SuperMix Universal (Thermo Fisher P / N 11762100) kit, using only forward and reverse amplification primers; alternatively, the Kapa Probe Force kit (Kapa Biosystems P / N KK4301) can be used together with amplification primers and a TaqMan probe that contains a FAM dye label at the 5' end, an internal ZEN quencher, a minor groove binder, and a fluorescent quencher at the 3' end (Integrated DNA Technologies).
[0243] Quantitative polymerase chain reaction (qPCR) is a method for real-time quantification of the amplification of one or more nucleic acid sequences. The real-time quantification of PCR assays allows determination of the amount of nucleic acid produced by the PCR amplification step by comparing the amplified nucleic acid of interest and an appropriate control nucleic acid sequence, which can serve as a calibration standard.
[0244] TaqMan probes are commonly used in qPCR assays that require higher specificity to quantify target nucleic acid sequences. A TaqMan probe comprises an oligonucleotide probe with a fluorophore attached to the 5'-end and a quencher attached to the 3'-end of the probe. When the TaqMan probe keeps the 5’ and 3’ ends of the probe in close contact with each other, the quencher prevents the fluorescence signal from the fluorophore from propagating. The TaqMan probe is designed to anneal within a nucleic acid region amplified by a set of specific primers. When Taq polymerase extends the primer and synthesizes the nascent strand, the 5'-to-3' exonuclease activity of Taq polymerase degrades the probe annealed to the template. The degradation of the probe releases the fluorophore, thus disrupting the close proximity to the quencher and allowing the fluorophore to fluoresce. The fluorescence detected in the qPCR assay is proportional to the amount of the released fluorophore and the DNA template present in the reaction.
[0245] The features of qPCR allow practitioners to eliminate the labor-intensive post-amplification step of gel electrophoresis, which is typically required to visualize the amplification products of conventional PCR assays. The advantages of qPCR over conventional PCR are considerable, including increased speed, ease of use, reproducibility, and quantification ability.
[0246] Microbial composition
[0247] In some aspects, the microorganisms of the present disclosure are combined into a microbial composition.
[0248] In some aspects, the microbial composition of the present disclosure is solid. When using a solid composition, it may be necessary to include one or more carrier materials, including but not limited to: mineral soils such as silica, talc, kaolin, limestone, chalk, clay, dolomite, diatomaceous earth; calcium sulfate; magnesium sulfate; magnesium oxide; zeolite, calcium carbonate; magnesium carbonate; trehalose; chitosan; shellac; albumin; starch; skim milk powder; sweet whey powder; maltodextrin; lactose; inulin; dextrose; and products of plant origin such as cereal flour, bark powder, wood powder, and fruit shell powder.
[0249] In some aspects, the microbial composition of the present disclosure is liquid. In other aspects, the liquid contains a solvent, which may include water or alcohol or brine or carbohydrate solution. In some aspects, the microbial composition of the present disclosure includes an adhesive such as a polymer, carboxymethyl cellulose, starch, polyvinyl alcohol, etc.
[0250] In some aspects, the microbial compositions of the present disclosure comprise sugars (e.g., monosaccharides, disaccharides, trisaccharides, polysaccharides, oligosaccharides, etc.), polymeric sugars, lipids, polymeric lipids, lipopolysaccharides, proteins, polymeric proteins, lipoproteins, nucleic acids, nucleic acid polymers, silica, inorganic salts, and combinations thereof. In another aspect, the microbial compositions comprise polymers such as agar, agarose, gellan gum, deacylated gellan gum (gelrite), etc. In some aspects, the microbial compositions comprise plastic capsules, emulsions (e.g., water and oil), membranes, and artificial membranes. In some aspects, an emulsion or a linked polymer solution may comprise the microbial compositions of the present disclosure. See Harel and Bennett (U.S. Patent 8,460,726 B2).
[0251] In some aspects, the microbial compositions of the present disclosure exist in solid form (e.g., dispersed lyophilized spores) or liquid form (microorganisms dispersed in a storage medium). In some aspects, the dried form of the microbial compositions of the present disclosure is added to a liquid prior to use to form a suspension. In some aspects, the microbial compositions comprise ceramized microorganisms.
[0252] In some aspects, the microbial compositions of the present disclosure have a water activity (aw) of less than 0.750, 0.700, 0.650, 0.600, 0.550, 0.500, 0.475, 0.450, 0.425, 0.400, 0.375, 0.350, 0.325, 0.300, 0.275, 0.250, 0.225, 0.200, 0.190, 0.180, 0.170, 0.160, 0.150, 0.140, 0.130, 0.120, 0.110, 0.100, 0.095, 0.090, 0.085, 0.080, 0.075, 0.070, 0.065, 0.060, 0.055, 0.050, 0.045, 0.040, 0.035, 0.030, 0.025, 0.020, 0.015, 0.010, or 0.005.
[0253] In some aspects, the microbial compositions of the present disclosure have a water activity (aw) of less than about 0.750, about 0.700, about 0.650, about 0.600, about 0.550, about 0.500, about 0.475, about 0.450, about 0.425, about 0.400, about 0.375, about 0.350, about 0.325, about 0.300, about 0.275, about 0.250, about 0.225, about 0.200, about 0.190, about 0.180, about 0.170, about 0.160, about 0.150, about 0.140, about 0.130, about 0.120, about 0.110, about 0.100, about 0.095, about 0.090, about 0.085, about 0.080, about 0.075, about 0.070, about 0.065, about 0.060, about 0.055, about 0.050, about 0.045, about 0.040, about 0.035, about 0.030, about 0.025, about 0.020, about 0.015, about 0.010, or about 0.005.
[0254] The water activity value is determined by the saturated aqueous solution method (Multon, “Techniques d’Analyse E DeControle Dans Les Industries Agroalimentaires” APRIA (1981)) or by direct measurement using a powerful Robotronic BT hygrometer or other hygrometers or moisture testers.
[0255] Raw materials
[0256] In some aspects, the present disclosure relates to a method for producing and recovering / separating one or more desired chemical substances. Recovery / collection / separation can be by methods known in the art such as distillation, membrane-based separation, stripping, solvent extraction, and expanded bed adsorption.
[0257] In some aspects, the raw material includes a carbon source. In some aspects, the carbon source can be selected from sugars, glycerol, alcohols, organic acids, alkanes, fatty acids, lignocellulose, proteins, carbon dioxide, and carbon monoxide. In one aspect, the carbon source is a sugar. In one aspect, the sugar is glucose or its glucose oligomers. In one aspect, the glucose oligomers are selected from fructose, sucrose, starch, cellobiose, maltose, lactose, and cellulose. In one aspect, the sugar is a pentose. In one aspect, the sugar is a hexose. In some aspects, the raw material contains one or more pentoses and / or one or more hexoses. In some aspects, the raw material contains one or more of xylose, glucose, arabinose, galactose, maltose, fructose, mannose, sucrose, and / or combinations thereof. In some aspects, the raw material contains one or more of xylose and / or glucose. In some aspects, the raw material contains one or more of arabinose, galactose, maltose, fructose, mannose, sucrose, and / or combinations thereof. In some aspects, the raw material contains xylose and glucose.
[0258] In some aspects, the microorganism utilizes one or more pentoses (pentoses) and / or one or more hexoses (hexoses). In some aspects, the microorganism utilizes one or more of xylose and / or glucose. In some aspects, the microorganism utilizes one or more of arabinose, galactose, maltose, fructose, mannose, sucrose, and / or combinations thereof. In some aspects, the microorganism utilizes one or more of xylose, glucose, arabinose, galactose, maltose, fructose, mannose, sucrose, and / or combinations thereof.
[0259] In some aspects, the hexoses can be selected from D-allose, D-altrose, D-glucose, D-mannose, D-gulose, D-idose, D-galactose, D-talose, D-tagatose, D-sorbose, D-fructose, D-psicose, and other hexoses known in the art. In some aspects, the pentoses can be selected from D-xylose, D-ribose, D-arabinose, D-lyxose, D-xylulose, D-ribulose, and other pentoses known in the art. In some aspects, the hexoses and pentoses can be selected from the left-handed enantiomers or right-handed enantiomers of any hexoses and pentoses disclosed herein.
[0260] In some aspects, the total amount of C5 and / or C6 carbohydrates fed into the bioreactor / growth medium during the growth phase is at least 5 kg carbohydrates / m3, at least 10 kg carbohydrates / m3, at least 20 kg carbohydrates / m3, at least 30 kg carbohydrates / m3, at least 40 kg carbohydrates / m3, at least 50 kg carbohydrates / m3, at least 60 kg carbohydrates / m3, at least 70 kg carbohydrates / m3, at least 80 kg carbohydrates / m3, at least 90 kg carbohydrates / m3, at least 100 kg carbohydrates / m3, at least 150 kg carbohydrates / m3, at least 200 kg carbohydrates / m3, at least 250 kg carbohydrates / m3, at least 300 kg carbohydrates / m3, at least 400 kg carbohydrates / m3, at least 500 kg carbohydrates / m3, at least 600 kg carbohydrates / m3, at least 700 kg carbohydrates / m3, up to 800 kg carbohydrates / m3. In some aspects, the total amount of C5 and / or C6 carbohydrates fed into the bioreactor / growth medium during the growth phase ranges from about 10 kg carbohydrates / m3 to up to 500 kg carbohydrates / m3.
[0261] In some aspects, the time required for the growth phase varies between 1 and 200 hours. In other aspects, the time of the growth phase is between 5 and 50 hours. The time depends on the carbohydrate feed and / or raw materials.
[0262] In some aspects, the total amount of C5 and / or C6 carbohydrates fed into the bioreactor / growth medium during the production phase is at least 50 kg carbohydrates / m3, at least 60 kg carbohydrates / m3, at least 70 kg carbohydrates / m3, at least 80 kg carbohydrates / m3, at least 90 kg carbohydrates / m3, at least 100 kg carbohydrates / m3, at least 150 kg carbohydrates / m3, at least 200 kg carbohydrates / m3, at least 250 kg carbohydrates / m3, at least 300 kg carbohydrates / m3, at least 400 kg carbohydrates / m3, at least 500 kg carbohydrates / m3, at least 600 kg carbohydrates / m3, at least 700 kg carbohydrates / m3, at least 800 kg carbohydrates / m3, at least 900 kg carbohydrates / m3, up to 1000 kg carbohydrates / m3. In some aspects, the total amount of C5 and / or C6 carbohydrates fed into the bioreactor / growth medium during the production phase ranges from about 100 kg carbohydrates / m3 to up to 800 kg carbohydrates / m3.
[0263] In some aspects, the time required for the production phase varies between 5 and 500 hours. In other aspects, for batch and fed-batch operations, the time for the production phase varies between 10 and 300 hours. In other aspects, the time for the production phase is up to 300 hours for continuous fermentation.
[0264] In some aspects, for a single-phase process, the total amount of C5 and / or C6 carbohydrates fed into the bioreactor / growth medium is at least 50 kg carbohydrates / m3, at least 60 kg carbohydrates / m3, at least 70 kg carbohydrates / m3, at least 80 kg carbohydrates / m3, at least 90 kg carbohydrates / m3, at least 100 kg carbohydrates / m3, at least 150 kg carbohydrates / m3, at least 200 kg carbohydrates / m3, at least 250 kg carbohydrates / m3, at least 300 kg carbohydrates / m3, at least 400 kg carbohydrates / m3, at least 500 kg carbohydrates / m3, at least 600 kg carbohydrates / m3, at least 700 kg carbohydrates / m3, at least 800 kg carbohydrates / m3, at least 900 kg carbohydrates / m3, up to 1000 kg carbohydrates / m3. In some aspects, the total amount of C5 and / or C6 carbohydrates fed into the bioreactor / growth medium during the production phase is in the range of about 100 kg carbohydrates / m3 to up to 800 kg carbohydrates / m3.
[0265] In some aspects, in a single-phase process, the time required for the production phase varies between 5 and 500 hours. In other aspects, in a single-phase process, the time required for the production phase varies between 5 and 300 hours.
[0266] In some aspects, a single-phase or multi-phase production process requires about 5, about 10, about 25, about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, or about 500 hours.
[0267] In some aspects, a single-phase or multi-phase production process requires 5, 10, 25, 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475 or 500 hours.
[0268] Improvement of traits
[0269] One or more of a plurality of desired traits can be introduced or improved using the methods of the present disclosure. Examples of traits that can be introduced or improved include: an increase in the rate and / or velocity of MEG, glycolic acid, polyols, acetone, propylene, isopropanol; an increase in the co-consumption of xylose and glucose; and a decrease in the inhibitory effect of one or more sugars on sugar consumption and / or uptake.
[0270] In some aspects, the microorganisms produced by the methods described herein exhibit a trait difference that is at least about 1% greater than a reference value under control conditions, such as at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 9%, at least about 9%, at least about 10%, at least about 11%, at least about 12%, at least about 13%, at least about 14%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 90%, or at least 100%, at least about 200%, at least about 300%, at least about 400% or greater. In additional examples, the microorganisms produced by the methods described herein exhibit a trait difference that is at least about 5% greater than a reference control grown under similar conditions, such as at least about 5%, at least about 8%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 75%, at least about 80%, at least about 80%, at least about 90%, or at least 100%, at least about 200%, at least about 300%, at least about 400% or greater.
[0271] In some aspects, the increase or decrease of any one or more traits of the present disclosure, relative to an unmodified microorganism, is an increase of about 0.1%, about 0.2%, about 0.3%, about 0.4%, about 0.5%, about 0.6%, about 0.7%, about 0.8%, about 0.9%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, about 59%, about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100%.
[0272] In some aspects, the increase or decrease of any one or more traits of the present disclosure, relative to an unmodified microorganism, is an increase of at least 0.1%, at least 0.2%, at least 0.3%, at least 0.4%, at least 0.5%, at least 0.6%, at least 0.7%, at least 0.8%, at least 0.9%, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 11%, at least 12%, at least 13%, at least 14%, at least 15%, at least 16%, at least 17%, at least 18%, at least 19%, at least 20%, at least 21%, at least 22%, at least 23%, at least 24%, at least 25%, at least 26%, at least 27%, at least 28%, at least 29%, at least 30%, at least 31%, at least 32%, at least 33%, at least 34%, at least 35%, at least 36%, at least 37%, at least 38%, at least 39%, at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100%.
[0273] Example
[0274] Example 1: Co-production of monoethylene glycol (MEG) and acetone via the D-xylonate pathway in a strain capable of consuming xylose and glucose simultaneously– Figure 1 Pathway A.
[0275] The Escherichia coli K12 strain MG1655 was used as a host lacking the following three genes: aldA, xylA, and glcDEF, which can divert the carbon flux from the MEG + acetone pathway. The genes were successfully deleted, and the deletion was confirmed by PCR and sequencing. The next step was the integration of the MEG pathway. An operon expressed under the control of the proD promoter containing the xdh gene (xylose dehydrogenase) and the fucO gene (glycolaldehyde reductase), which encode the first and last enzymes of the xylonic acid pathway, respectively, was integrated into the Escherichia coli genome, and an additional copy of the xdh gene was also placed under the control of the proD promoter and integrated into a different locus. The integration of the xdh gene allowed the conversion of xylose into the intermediate D-xylonic acid and glycolaldehyde. The integration of the fucO gene reduced glycolaldehyde to MEG and was specific for MEG production. The second step was the integration of the acetone pathway. An operon expressed under the control of the OXB11 promoter containing the thlA gene (acetoacetyl-CoA thiolase); the AtoDA gene (acetate:acetoacetyl-CoA transferase) and the adc gene (acetoacetate decarboxylase) was integrated into the Escherichia coli genome to generate a basal strain. This basal strain was used as a host for modification to promote the co-consumption of glucose and xylose. The first modification was the integration of another copy of xylE under the control of the GAPDH promoter at the araFGH locus, resulting in the deletion of araFGH. The second modification was the deletion of the xylFGH operon. All integrations and deletions were confirmed by PCR and sequencing.
[0276] Colonies from transformation were inoculated into 5 mL of mineral medium containing 12.85 g / L xylose and 2.15 g / L glucose (6:1 ratio) or 7.5 g / L xylose and 7.5 g / L glucose (1:1 ratio) for pre-culture. After 16 hours of culture, 5% of the pre-culture was transferred to 100 mL of fresh medium. The flasks were incubated at 37 °C and 250 rpm. The initial OD of the culture was 0.1.
[0277] For the 1:1 ratio culture, the simultaneous utilization of glucose and xylose could be detected in the co-consumption strain after 8 hours of culture, while in the parental strain, xylose only started to be consumed after glucose was exhausted after 18 hours ( Figure 2 ). During 36 hours of culture, the co-consumption strain was able to consume 75% of the initial sugar mixture, while the parental strain only consumed 62%.
[0278] For the 6:1 ratio culture, both the parental and co-consumption strains were able to completely consume the initial glucose and xylose, and the xylose consumption and biomass production characteristic curves were similar ( Figure 3 ). In the co-consumption strain, the total amount of MEG increased by 12%, while the amount of acetone increased by 197% ( Figure 4) Modification of xylose uptake provides an improvement in the co-production rate relative to its parental strain.
[0279] Example 2: Co-production of monoethylene glycol (MEG) and acetone via the D-xylulose pathway in a strain capable of co-consuming xylose and glucose Figure 1 Pathway B.
[0280] The Escherichia coli K12 strain MG1655 was used as a host lacking the following three genes: aldA, xylB, and glcDEF, which can divert carbon flux from the MEG + acetone pathway. The genes were successfully deleted, and the deletion was confirmed by PCR and sequencing. The next step was the integration of the MEG pathway. An operon expressed under the control of the proD promoter containing the khk-C gene (ketohexokinase), the aldB gene (fructose-1,6-bisphosphate aldolase), and the fucO gene (hydroxyacetaldehyde reductase) was integrated into the Escherichia coli genome, and an additional copy of the khk-C and aldB genes, also under the control of the proD promoter, was integrated into a different locus. The integration of the khk-C and aldB genes allowed the conversion of xylose to the intermediate hydroxyacetaldehyde. The integration of the fucO gene reduced hydroxyacetaldehyde to MEG and was specific for MEG production. The second step was the integration of the acetone pathway. An operon expressed under the control of the OXB11 promoter containing the thlA gene (acetoacetyl-CoA thiolase); the AtoDA gene (acetate:acetoacetyl-CoA transferase) and the adc gene (acetoacetate decarboxylase) was integrated into the Escherichia coli genome to generate the base strain. This base strain was used as a host for modification to facilitate the co-consumption of glucose and xylose. The first modification was the integration of another copy of xylE under the control of the GAPDH promoter at the araFGH locus, thereby deleting araFGH. The second modification was the deletion of the xylFGH operon and the replacement of the xylA promoter with the OXB15 promoter. Expression of xylA under a constitutive promoter allowed the conversion of xylose to the intermediates D-xylonic acid and hydroxyacetaldehyde. All integrations and deletions were confirmed by PCR and sequencing.
[0281] Colonies from transformation were inoculated into 5 mL of mineral medium containing 12.85 g / L xylose and 2.15 g / L glucose (6:1 ratio) or 7.5 g / L xylose and 7.5 g / L glucose (1:1 ratio) for pre-culture. After 16 hours of culture, 5% of the pre-culture was transferred to 100 mL of fresh medium. The flasks were incubated at 37 °C and 250 rpm. The initial OD of the culture was 0.1.
[0282] For the 1:1 ratio cultures, co - utilization of glucose and xylose could be detected in the co - consuming strain after 12 h of cultivation, while in the parental strain, xylose started to decrease only after glucose was depleted at 18 h( Figure 5 ). During 36 h of cultivation, the co - consuming strain was able to consume 61% of the initial sugar mixture, while the parental strain consumed 52%.
[0283] For the 6:1 ratio cultures, both the parental and co - consuming strains were able to completely consume the initial glucose and xylose, and the xylose consumption and biomass production profiles were similar( Figure 6 ). In the co - consuming strain, the total amount of MEG increased by 9%, while the total amount of acetone increased by 119%( Figure 7 ). Modification of xylose uptake provided an improvement in the co - production rate relative to its parental strain.
[0284] Incorporated herein by reference
[0285] All references, articles, publications, patents, patent publications, and patent applications cited herein are incorporated by reference in their entirety for all purposes. However, the mention of any reference, article, publication, patent, patent publication, and patent application cited herein is not and should not be taken as an admission or any form of implication that they constitute valid prior art or form part of the common general knowledge in any country or region of the world. In addition, the following references are hereby incorporated by reference:
[0286] Yang et al. (2018). One step fermentative production of aromatic polyesters from glucose by metabolically engineered Escherichia coli strains. Nature Communications. 9(1):79.
[0287] Fritzsche et al. (1990). An unusual bacterial polyester with a phenyl pendant group. Die Makromolekulare Chemie: Macromolecular Chemistry and Physics. 191(8):1957 - 1965.
[0288] Garcia et al.(1999).Novel biodegradable aromatic plastics from a bacterial source genetic and biochemical studies on a route of the phenylacetyl-CoA catabolon.Journal of Biological Chemistry.274(41):29228-29241.
[0289] Olivera et al.(2001).Genetically engineered Pseudomonas:a factory of new bioplastics with broad applications.Environmental Microbiology.3(10):612-618. Sequence Listing <110> Blasco, Inc. <120> Metabolic Engineering for the Simultaneous Consumption of Xylose and Glucose to Produce Chemicals from Second Generation Sugars <130> BRSK-021 / 01WO (331051-2072) <150> 62 / 829,398 <151> 2019-04-04 <160> 98 <170> PatentIn version 3.5 <210> 1 <211> 1164 <212> DNA <213> Escherichia coli <400> 1 atgaacaact ttaatctgca caccccaacc cgcattctgt ttggtaaagg cgcaatcgct 60 ggtttacgcg aacaaattcc tcacgatgct cgcgtattga ttacctacgg cggcggcagc 120 gtgaaaaaaa ccggcgttct cgatcaagtt ctggatgccc tgaaaggcat ggacgtgctg 180 gaatttggcg gtattgagcc aaacccggct tatgaaacgc tgatgaacgc cgtgaaactg 240 gttcgcgaac agaaagtgac tttcctgctg gcggttggcg gcggttctgt actggacggc 300 accaaattta tcgccgcagc ggctaactat ccggaaaata tcgatccgtg gcacattctg 360 caaacgggcg gtaaagagat taaaagcgcc atcccgatgg gctgtgtgct gacgctgcca 420 gcaaccggtt cagaatccaa cgcaggcgcg gtgatctccc gtaaaaccac aggcgacaag 480 caggcgttcc attctgccca tgttcagccg gtatttgccg tgctcgatcc ggtttatacc 540 tacaccctgc cgccgcgtca ggtggctaac ggcgtagtgg acgcctttgt acacaccgtg 600 gaacagtatg ttaccaaacc ggttgatgcc aaaattcagg accgtttcgc agaaggcatt 660 ttgctgacgc taatcgaaga tggtccgaaa gccctgaaag agccagaaaa ctacgatgtg 720 cgcgccaacg tcatgtgggc ggcgactcag gcgctgaacg gtttgattgg cgctggcgta 780 ccgcaggact gggcaacgca tatgctgggc cacgaactga ctgcgatgca cggtctggat 840 cacgcgcaaa cactggctat cgtcctgcct gcactgtgga atgaaaaacg cgataccaag 900 cgcgctaagc tgctgcaata tgctgaacgc gtctggaaca tcactgaagg ttccgatgat 960 gagcgtattg acgccgcgat tgccgcaacc cgcaatttct ttgagcaatt aggcgtgccg 1020 acccacctct ccgactacgg tctggacggc agctccatcc cggctttgct gaaaaaactg 1080 gaagagcacg gcatgaccca actgggcgaa aatcatgaca ttacgttgga tgtcagccgc 1140 cgtatatacg aagccgcccg ctaa 1164 <210> 2 <211> 387 <212> PRT <213> Escherichia coli <400> 2 Met Asn Asn Phe Asn Leu His Thr Pro Thr Arg Ile Leu Phe Gly Lys 1 5 10 15 Gly Ala Ile Ala Gly Leu Arg Glu Gln Ile Pro His Asp Ala Arg Val 20 25 30 Leu Ile Thr Tyr Gly Gly Gly Ser Val Lys Lys Thr Gly Val Leu Asp 35 40 45 Gln Val Leu Asp Ala Leu Lys Gly Met Asp Val Leu Glu Phe Gly Gly 50 55 60 Ile Glu Pro Asn Pro Ala Tyr Glu Thr Leu Met Asn Ala Val Lys Leu 65 70 75 80 Val Arg Glu Gln Lys Val Thr Phe Leu Leu Ala Val Gly Gly Gly Ser 85 90 95 Val Leu Asp Gly Thr Lys Phe Ile Ala Ala Ala Ala Asn Tyr Pro Glu 100 105 110 Asn Ile Asp Pro Trp His Ile Leu Gln Thr Gly Gly Lys Glu Ile Lys 115 120 125 Ser Ala Ile Pro Met Gly Cys Val Leu Thr Leu Pro Ala Thr Gly Ser 130 135 140 Glu Ser Asn Ala Gly Ala Val Ile Ser Arg Lys Thr Thr Gly Asp Lys 145 150 155 160 Gln Ala Phe His Ser Ala His Val Gln Pro Val Phe Ala Val Leu Asp 165 170 175 Pro Val Tyr Thr Tyr Thr Leu Pro Pro Arg Gln Val Ala Asn Gly Val 180 185 190 Val Asp Ala Phe Val His Thr Val Glu Gln Tyr Val Thr Lys Pro Val 195 200 205 Asp Ala Lys Ile Gln Asp Arg Phe Ala Glu Gly Ile Leu Leu Thr Leu 210 215 220 Ile Glu Asp Gly Pro Lys Ala Leu Lys Glu Pro Glu Asn Tyr Asp Val 225 230 235 240 Arg Ala Asn Val Met Trp Ala Ala Thr Gln Ala Leu Asn Gly Leu Ile 245 250 255 Gly Ala Gly Val Pro Gln Asp Trp Ala Thr His Met Leu Gly His Glu 260 265 270 Leu Thr Ala Met His Gly Leu Asp His Ala Gln Thr Leu Ala Ile Val 275 280 285 Leu Pro Ala Leu Trp Asn Glu Lys Arg Asp Thr Lys Arg Ala Lys Leu 290 295 300 Leu Gln Tyr Ala Glu Arg Val Trp Asn Ile Thr Glu Gly Ser Asp Asp 305 310 315 320 Glu Arg Ile Asp Ala Ala Ile Ala Ala Thr Arg Asn Phe Phe Glu Gln 325 330 335 Leu Gly Val Pro Thr His Leu Ser Asp Tyr Gly Leu Asp Gly Ser Ser 340 345 350 Ile Pro Ala Leu Leu Lys Lys Leu Glu Glu His Gly Met Thr Gln Leu 355 360 365 Gly Glu Asn His Asp Ile Thr Leu Asp Val Ser Arg Arg Ile Tyr Glu 370 375 380 Ala Ala Arg 385 <210> 3 <211> 1440 <212> DNA <213> Escherichia coli <400> 3 atgtcagtac ccgttcaaca tcctatgtat atcgatggac agtttgttac ctggcgtgga 60 gacgcatgga ttgatgtggt aaaccctgct acagaggctg tcatttcccg catacccgat 120 ggtcaggccg aggatgcccg taaggcaatc gatgcagcag aacgtgcaca accagaatgg 180 gaagcgttgc ctgctattga acgcgccagt tggttgcgca aaatctccgc cgggatccgc 240 gaacgcgcca gtgaaatcag tgcgctgatt gttgaagaag ggggcaagat ccagcagctg 300 gctgaagtcg aagtggcttt tactgccgac tatatcgatt acatggcgga gtgggcacgg 360 cgttacgagg gcgagattat tcaaagcgat cgtccaggag aaaatattct tttgtttaaa 420 cgtgcgcttg gtgtgactac cggcattctg ccgtggaact tcccgttctt cctcattgcc 480 cgcaaaatgg ctcccgctct tttgaccggt aataccatcg tcattaaacc tagtgaattt 540 acgccaaaca atgcgattgc attcgccaaa atcgtcgatg aaataggcct tccgcgcggc 600 gtgtttaacc ttgtactggg gcgtggtgaa accgttgggc aagaactggc gggtaaccca 660 aaggtcgcaa tggtcagtat gacaggcagc gtctctgcag gtgagaagat catggcgact 720 gcggcgaaaa acatcaccaa agtgtgtctg gaattggggg gtaaagcacc agctatcgta 780 atggacgatg ccgatcttga actggcagtc aaagccatcg ttgattcacg cgtcattaat 840 agtgggcaag tgtgtaactg tgcagaacgt gtttatgtac agaaaggcat ttatgatcag 900 ttcgtcaatc ggctgggtga agcgatgcag gcggttcaat ttggtaaccc cgctgaacgc 960 aacgacattg cgatggggcc gttgattaac gccgcggcgc tggaaagggt cgagcaaaaa 1020 gtggcgcgcg cagtagaaga aggggcgaga gtggcgttcg gtggcaaagc ggtagagggg 1080 aaaggatatt attatccgcc gacattgctg ctggatgttc gccaggaaat gtcgattatg 1140 catgaggaaa cctttggccc ggtgctgcca gttgtcgcat ttgacacgct ggaagatgct 1200 atctcaatgg ctaatgacag tgattacggc ctgacctcat caatctatac ccaaaatctg 1260 aacgtcgcga tgaaagccat taaagggctg aagtttggtg aaacttacat caaccgtgaa 1320 aacttcgaag ctatgcaagg cttccacgcc ggatggcgta aatccggtat tggcggcgca 1380 gatggtaaac atggcttgca tgaatatctg cagacccagg tggtttattt acagtcttaa 1440 <210> 4 <211> 479 <212> PRT <213> Escherichia coli <400> 4 Met Ser Val Pro Val Gln His Pro Met Tyr Ile Asp Gly Gln Phe Val 1 5 10 15 Thr Trp Arg Gly Asp Ala Trp Ile Asp Val Val Asn Pro Ala Thr Glu 20 25 30 Ala Val Ile Ser Arg Ile Pro Asp Gly Gln Ala Glu Asp Ala Arg Lys 35 40 45 Ala Ile Asp Ala Ala Glu Arg Ala Gln Pro Glu Trp Glu Ala Leu Pro 50 55 60 Ala Ile Glu Arg Ala Ser Trp Leu Arg Lys Ile Ser Ala Gly Ile Arg 65 70 75 80 Glu Arg Ala Ser Glu Ile Ser Ala Leu Ile Val Glu Glu Gly Gly Lys 85 90 95 Ile Gln Gln Leu Ala Glu Val Glu Val Ala Phe Thr Ala Asp Tyr Ile 100 105 110 Asp Tyr Met Ala Glu Trp Ala Arg Arg Tyr Glu Gly Glu Ile Ile Gln 115 120 125 Ser Asp Arg Pro Gly Glu Asn Ile Leu Leu Phe Lys Arg Ala Leu Gly 130 135 140 Val Thr Thr Gly Ile Leu Pro Trp Asn Phe Pro Phe Phe Leu Ile Ala 145 150 155 160 Arg Lys Met Ala Pro Ala Leu Leu Thr Gly Asn Thr Ile Val Ile Lys 165 170 175 Pro Ser Glu Phe Thr Pro Asn Asn Ala Ile Ala Phe Ala Lys Ile Val 180 185 190 Asp Glu Ile Gly Leu Pro Arg Gly Val Phe Asn Leu Val Leu Gly Arg 195 200 205 Gly Glu Thr Val Gly Gln Glu Leu Ala Gly Asn Pro Lys Val Ala Met 210 215 220 Val Ser Met Thr Gly Ser Val Ser Ala Gly Glu Lys Ile Met Ala Thr 225 230 235 240 Ala Ala Lys Asn Ile Thr Lys Val Cys Leu Glu Leu Gly Gly Lys Ala 245 250 255 Pro Ala Ile Val Met Asp Asp Ala Asp Leu Glu Leu Ala Val Lys Ala 260 265 270 Ile Val Asp Ser Arg Val Ile Asn Ser Gly Gln Val Cys Asn Cys Ala 275 280 285 Glu Arg Val Tyr Val Gln Lys Gly Ile Tyr Asp Gln Phe Val Asn Arg 290 295 300 Leu Gly Glu Ala Met Gln Ala Val Gln Phe Gly Asn Pro Ala Glu Arg 305 310 315 320 Asn Asp Ile Ala Met Gly Pro Leu Ile Asn Ala Ala Ala Leu Glu Arg 325 330 335 Val Glu Gln Lys Val Ala Arg Ala Val Glu Glu Gly Ala Arg Val Ala 340 345 350 Phe Gly Gly Lys Ala Val Glu Gly Lys Gly Tyr Tyr Tyr Pro Pro Thr 355 360 365 Leu Leu Leu Asp Val Arg Gln Glu Met Ser Ile Met His Glu Glu Thr 370 375 380 Phe Gly Pro Val Leu Pro Val Val Ala Phe Asp Thr Leu Glu Asp Ala 385 390 395 400 Ile Ser Met Ala Asn Asp Ser Asp Tyr Gly Leu Thr Ser Ser Ile Tyr 405 410 415 Thr Gln Asn Leu Asn Val Ala Met Lys Ala Ile Lys Gly Leu Lys Phe 420 425 430 Gly Glu Thr Tyr Ile Asn Arg Glu Asn Phe Glu Ala Met Gln Gly Phe 435 440 445 His Ala Gly Trp Arg Lys Ser Gly Ile Gly Gly Ala Asp Gly Lys His 450 455 460 Gly Leu His Glu Tyr Leu Gln Thr Gln Val Val Tyr Leu Gln Ser 465 470 475 <210> 5 <211> 1323 <212> DNA <213> Escherichia coli <400> 5 atgcaagcct attttgacca gctcgatcgc gttcgttatg aaggctcaaa atcctcaaac 60 ccgttagcat tccgtcacta caatcccgac gaactggtgt tgggtaagcg tatggaagag 120 cacttgcgtt ttgccgcctg ctactggcac accttctgct ggaacggggc ggatatgttt 180 ggtgtggggg cgtttaatcg tccgtggcag cagcctggtg aggcactggc gttggcgaag 240 cgtaaagcag atgtcgcatt tgagtttttc cacaagttac atgtgccatt ttattgcttc 300 cacgatgtgg atgtttcccc tgagggcgcg tcgttaaaag agtacatcaa taattttgcg 360 caaatggttg atgtcctggc aggcaagcaa gaagagagcg gcgtgaagct gctgtgggga 420 acggccaact gctttacaaa ccctcgctac ggcgcgggtg cggcgacgaa cccagatcct 480 gaagtcttca gctgggcggc aacgcaagtt gttacagcga tggaagcaac ccataaattg 540 ggcggtgaaa actatgtcct gtggggcggt cgtgaaggtt acgaaacgct gttaaatacc 600 gacttgcgtc aggagcgtga acaactgggc cgctttatgc agatggtggt tgagcataaa 660 cataaaatcg gtttccaggg cacgttgctt atcgaaccga aaccgcaaga accgaccaaa 720 catcaatatg attacgatgc cgcgacggtc tatggcttcc tgaaacagtt tggtctggaa 780 aaagagatta aactgaacat tgaagctaac cacgcgacgc tggcaggtca ctctttccat 840 catgaaatag ccaccgccat tgcgcttggc ctgttcggtt ctgtcgacgc caaccgtggc 900 gatgcgcaac tgggctggga caccgaccag ttcccgaaca gtgtggaaga gaatgcgctg 960 gtgatgtatg aaattctcaa agcaggcggt ttcaccaccg gtggtctgaa cttcgatgcc 1020 aaagtacgtc gtcaaagtac tgataaatat gatctgtttt acggtcatat cggcgcgatg 1080 gatacgatgg cactggcgct gaaaattgca gcgcgcatga ttgaagatgg cgagctggat 1140 aaacgcatcg cgcagcgtta ttccggctgg aatagcgaat tgggccagca aatcctgaaa 1200 ggccaaatgt cactggcaga tttagccaaa tatgctcagg aacatcattt gtctccggtg 1260 catcagagtg gtcgccagga acaactggaa aatctggtaa accattatct gttcgacaaa 1320 taa 1323 <210> 6 <211> 440 <212> PRT <213> Escherichia coli <400> 6 Met Gln Ala Tyr Phe Asp Gln Leu Asp Arg Val Arg Tyr Glu Gly Ser 1 5 10 15 Lys Ser Ser Asn Pro Leu Ala Phe Arg His Tyr Asn Pro Asp Glu Leu 20 25 30 Val Leu Gly Lys Arg Met Glu Glu His Leu Arg Phe Ala Ala Cys Tyr 35 40 45 Trp His Thr Phe Cys Trp Asn Gly Ala Asp Met Phe Gly Val Gly Ala 50 55 60 Phe Asn Arg Pro Trp Gln Gln Pro Gly Glu Ala Leu Ala Leu Ala Lys 65 70 75 80 Arg Lys Ala Asp Val Ala Phe Glu Phe Phe His Lys Leu His Val Pro 85 90 95 Phe Tyr Cys Phe His Asp Val Asp Val Ser Pro Glu Gly Ala Ser Leu 100 105 110 Lys Glu Tyr Ile Asn Asn Phe Ala Gln Met Val Asp Val Leu Ala Gly 115 120 125 Lys Gln Glu Glu Ser Gly Val Lys Leu Leu Trp Gly Thr Ala Asn Cys 130 135 140 Phe Thr Asn Pro Arg Tyr Gly Ala Gly Ala Ala Thr Asn Pro Asp Pro 145 150 155 160 Glu Val Phe Ser Trp Ala Ala Thr Gln Val Val Thr Ala Met Glu Ala 165 170 175 Thr His Lys Leu Gly Gly Glu Asn Tyr Val Leu Trp Gly Gly Arg Glu 180 185 190 Gly Tyr Glu Thr Leu Leu Asn Thr Asp Leu Arg Gln Glu Arg Glu Gln 195 200 205 Leu Gly Arg Phe Met Gln Met Val Val Glu His Lys His Lys Ile Gly 210 215 220 Phe Gln Gly Thr Leu Leu Ile Glu Pro Lys Pro Gln Glu Pro Thr Lys 225 230 235 240 His Gln Tyr Asp Tyr Asp Ala Ala Thr Val Tyr Gly Phe Leu Lys Gln 245 250 255 Phe Gly Leu Glu Lys Glu Ile Lys Leu Asn Ile Glu Ala Asn His Ala 260 265 270 Thr Leu Ala Gly His Ser Phe His His Glu Ile Ala Thr Ala Ile Ala 275 280 285 Leu Gly Leu Phe Gly Ser Val Asp Ala Asn Arg Gly Asp Ala Gln Leu 290 295 300 Gly Trp Asp Thr Asp Gln Phe Pro Asn Ser Val Glu Glu Asn Ala Leu 305 310 315 320 Val Met Tyr Glu Ile Leu Lys Ala Gly Gly Phe Thr Thr Gly Gly Leu 325 330 335 Asn Phe Asp Ala Lys Val Arg Arg Gln Ser Thr Asp Lys Tyr Asp Leu 340 345 350 Phe Tyr Gly His Ile Gly Ala Met Asp Thr Met Ala Leu Ala Leu Lys 355 360 365 Ile Ala Ala Arg Met Ile Glu Asp Gly Glu Leu Asp Lys Arg Ile Ala 370 375 380 Gln Arg Tyr Ser Gly Trp Asn Ser Glu Leu Gly Gln Gln Ile Leu Lys 385 390 395 400 Gly Gln Met Ser Leu Ala Asp Leu Ala Lys Tyr Ala Gln Glu His His 405 410 415 Leu Ser Pro Val His Gln Ser Gly Arg Gln Glu Gln Leu Glu Asn Leu 420 425 430 Val Asn His Tyr Leu Phe Asp Lys 435 440 <210> 7 <211> 1179 <212> DNA <213> Escherichia coli <400> 7 atgtttacta aacgtcaccg catcacatta ctgttcaatg ccaataaagc ctatgaccgg 60 caggtagtag aaggcgtagg ggaatattta caggcgtcac aatcggaatg ggatattttc 120 attgaagaag atttccgcgc ccgcattgat aaaatcaagg actggttagg agatggcgtc 180 attgccgact tcgacgacaa acagatcgag caagcgctgg ctgatgtcga cgtccccatt 240 gttggggttg gcggctcgta tcaccttgca gaaagttacc cacccgttca ttacattgcc 300 accgataact atgcgctggt tgaaagcgca tttttgcatt taaaagagaa aggcgttaac 360 cgctttgctt tttatggtct tccggaatca agcggcaaac gttgggccac tgagcgcgaa 420 tatgcatttc gtcagcttgt cgccgaagaa aagtatcgcg gagtggttta tcaggggtta 480 gaaaccgcgc cagagaactg gcaacacgcg caaaatcggc tggcagactg gctacaaacg 540 ctaccaccgc aaaccgggat tattgccgtt actgacgccc gagcgcggca tattctgcaa 600 gtatgtgaac atctacatat tcccgtaccg gaaaaattat gcgtgattgg catcgataac 660 gaagaactga cccgctatct gtcgcgtgtc gccctttctt cggtcgctca gggcgcgcgg 720 caaatgggct atcaggcggc aaaactgttg catcgattat tagataaaga agaaatgccg 780 ctacagcgaa ttttggtccc accagttcgc gtcattgaac ggcgctcaac agattatcgc 840 tcgctgaccg atcccgccgt tattcaggcc atgcattaca ttcgtaatca cgcctgtaaa 900 gggattaaag tggatcaggt actggatgcg gtcgggatct cgcgctccaa tcttgagaag 960 cgttttaaag aagaggtggg tgaaaccatc catgccatga ttcatgccga gaagctggag 1020 aaagcgcgca gtctgctgat ttcaaccacc ttgtcgatca atgagatatc gcaaatgtgc 1080 ggttatccat cgctgcaata tttctactct gtttttaaaa aagcatatga cacgacgcca 1140 aaagagtatc gcgatgtaaa tagcgaggtc atgttgtag 1179 <210> 8 <211> 392 <212> PRT <213> Escherichia coli <400> 8 Met Phe Thr Lys Arg His Arg Ile Thr Leu Leu Phe Asn Ala Asn Lys 1 5 10 15 Ala Tyr Asp Arg Gln Val Val Glu Gly Val Gly Glu Tyr Leu Gln Ala 20 25 30 Ser Gln Ser Glu Trp Asp Ile Phe Ile Glu Glu Asp Phe Arg Ala Arg 35 40 45 Ile Asp Lys Ile Lys Asp Trp Leu Gly Asp Gly Val Ile Ala Asp Phe 50 55 60 Asp Asp Lys Gln Ile Glu Gln Ala Leu Ala Asp Val Asp Val Pro Ile 65 70 75 80 Val Gly Val Gly Gly Ser Tyr His Leu Ala Glu Ser Tyr Pro Pro Val 85 90 95 His Tyr Ile Ala Thr Asp Asn Tyr Ala Leu Val Glu Ser Ala Phe Leu 100 105 110 His Leu Lys Glu Lys Gly Val Asn Arg Phe Ala Phe Tyr Gly Leu Pro 115 120 125 Glu Ser Ser Gly Lys Arg Trp Ala Thr Glu Arg Glu Tyr Ala Phe Arg 130 135 140 Gln Leu Val Ala Glu Glu Lys Tyr Arg Gly Val Val Tyr Gln Gly Leu 145 150 155 160 Glu Thr Ala Pro Glu Asn Trp Gln His Ala Gln Asn Arg Leu Ala Asp 165 170 175 Trp Leu Gln Thr Leu Pro Pro Gln Thr Gly Ile Ile Ala Val Thr Asp 180 185 190 Ala Arg Ala Arg His Ile Leu Gln Val Cys Glu His Leu His Ile Pro 195 200 205 Val Pro Glu Lys Leu Cys Val Ile Gly Ile Asp Asn Glu Glu Leu Thr 210 215 220 Arg Tyr Leu Ser Arg Val Ala Leu Ser Ser Val Ala Gln Gly Ala Arg 225 230 235 240 Gln Met Gly Tyr Gln Ala Ala Lys Leu Leu His Arg Leu Leu Asp Lys 245 250 255 Glu Glu Met Pro Leu Gln Arg Ile Leu Val Pro Pro Val Arg Val Ile 260 265 270 Glu Arg Arg Ser Thr Asp Tyr Arg Ser Leu Thr Asp Pro Ala Val Ile 275 280 285 Gln Ala Met His Tyr Ile Arg Asn His Ala Cys Lys Gly Ile Lys Val 290 295 300 Asp Gln Val Leu Asp Ala Val Gly Ile Ser Arg Ser Asn Leu Glu Lys 305 310 315 320 Arg Phe Lys Glu Glu Val Gly Glu Thr Ile His Ala Met Ile His Ala 325 330 335 Glu Lys Leu Glu Lys Ala Arg Ser Leu Leu Ile Ser Thr Thr Leu Ser 340 345 350 Ile Asn Glu Ile Ser Gln Met Cys Gly Tyr Pro Ser Leu Gln Tyr Phe 355 360 365 Tyr Ser Val Phe Lys Lys Ala Tyr Asp Thr Thr Pro Lys Glu Tyr Arg 370 375 380 Asp Val Asn Ser Glu Val Met Leu 385 390 <210> 9 <211> 633 <212> DNA <213> Escherichia coli <400> 9 atggtgcttg gcaaaccgca aacagacccg actctcgaat ggttcttgtc tcattgccac 60 attcataagt acccatccaa gagcacgctt attcaccagg gtgaaaaagc ggaaacgctg 120 tactacatcg ttaaaggctc tgtggcagtg ctgatcaaag acgaagaggg taaagaaatg 180 atcctctcct atctgaatca gggtgatttt attggcgaac tgggcctgtt tgaagagggc 240 caggaacgta gcgcatgggt acgtgcgaaa accgcctgtg aagtggctga aatttcgtac 300 aaaaaatttc gccaattgat tcaggtaaac ccggacattc tgatgcgttt gtctgcacag 360 atggcgcgtc gtctgcaagt cacttcagag aaagtgggca acctggcgtt cctcgacgtg 420 acgggccgca ttgcacagac tctgctgaat ctggcaaaac aaccagacgc tatgactcac 480 ccggacggta tgcaaatcaa aattacccgt caggaaattg gtcagattgt cggctgttct 540 cgtgaaaccg tgggacgcat tctgaagatg ctggaagatc agaacctgat ctccgcacac 600 ggtaaaacca tcgtcgttta cggcactcgt taa 633 <210> 10 <211> 210 <212> PRT <213> Escherichia coli <400> 10 Met Val Leu Gly Lys Pro Gln Thr Asp Pro Thr Leu Glu Trp Phe Leu 1 5 10 15 Ser His Cys His Ile His Lys Tyr Pro Ser Lys Ser Thr Leu Ile His 20 25 30 Gln Gly Glu Lys Ala Glu Thr Leu Tyr Tyr Ile Val Lys Gly Ser Val 35 40 45 Ala Val Leu Ile Lys Asp Glu Glu Gly Lys Glu Met Ile Leu Ser Tyr 50 55 60 Leu Asn Gln Gly Asp Phe Ile Gly Glu Leu Gly Leu Phe Glu Glu Gly 65 70 75 80 Gln Glu Arg Ser Ala Trp Val Arg Ala Lys Thr Ala Cys Glu Val Ala 85 90 95 Glu Ile Ser Tyr Lys Lys Phe Arg Gln Leu Ile Gln Val Asn Pro Asp 100 105 110 Ile Leu Met Arg Leu Ser Ala Gln Met Ala Arg Arg Leu Gln Val Thr 115 120 125 Ser Glu Lys Val Gly Asn Leu Ala Phe Leu Asp Val Thr Gly Arg Ile 130 135 140 Ala Gln Thr Leu Leu Asn Leu Ala Lys Gln Pro Asp Ala Met Thr His 145 150 155 160 Pro Asp Gly Met Gln Ile Lys Ile Thr Arg Gln Glu Ile Gly Gln Ile 165 170 175 Val Gly Cys Ser Arg Glu Thr Val Gly Arg Ile Leu Lys Met Leu Glu 180 185 190 Asp Gln Asn Leu Ile Ser Ala His Gly Lys Thr Ile Val Val Tyr Gly 195 200 205 Thr Arg 210 <210> 11 <211> 897 <212> DNA <213> Homo sapiens <400> 11 atggaggaaa agcaaattct gtgcgttggt ctggtggttc tggacgtgat tagcctggtt 60 gataagtacc cgaaagagga tagcgaaatc cgttgcctga gccagcgttg gcaacgtggt 120 ggcaacgcga gcaatagctg caccgttctg agcctgctgg gtgcgccgtg cgcgttcatg 180 ggtagcatgg cgccgggtca tgttgcggac ttcctggtgg cggattttcg tcgtcgtggt 240 gtggacgtta gccaggttgc gtggcaaagc aagggcgata ccccgagctc ctgctgcatc 300 attaacaaca gcaacggtaa ccgtaccatt gtgctgcacg acaccagcct gccggatgtt 360 agcgcgaccg acttcgagaa ggtggatctg acccagttta aatggattca cattgagggc 420 cgtaacgcga gcgaacaggt taaaatgctg caacgtattg atgcgcacaa cacccgtcag 480 ccgccggaac aaaagattcg tgtgagcgtt gaggtggaaa aaccgcgtga ggaactgttc 540 caactgtttg gttacggcga cgtggttttc gttagcaagg atgtggcgaa acacctgggt 600 tttcaaagcg cggaggaagc gctgcgtggt ctgtatggcc gtgtgcgtaa aggcgcggtt 660 ctggtgtgcg cgtgggcgga ggaaggcgcg gatgcgctgg gtccggatgg caaactgctg 720 cacagcgatg cgttcccgcc gccgcgtgtg gttgacaccc tgggtgcggg cgataccttc 780 aacgcgagcg ttatctttag cctgagccag ggccgtagcg tgcaagaggc gctgcgtttc 840 ggctgccaag ttgcgggtaa aaaatgcggt ctgcaaggct ttgacggtat cgtgtaa 897 <210> 12 <211> 298 <212> PRT <213> Homo sapiens <400> 12 Met Glu Glu Lys Gln Ile Leu Cys Val Gly Leu Val Val Leu Asp Val 1 5 10 15 Ile Ser Leu Val Asp Lys Tyr Pro Lys Glu Asp Ser Glu Ile Arg Cys 20 25 30 Leu Ser Gln Arg Trp Gln Arg Gly Gly Asn Ala Ser Asn Ser Cys Thr 35 40 45 Val Leu Ser Leu Leu Gly Ala Pro Cys Ala Phe Met Gly Ser Met Ala 50 55 60 Pro Gly His Val Ala Asp Phe Leu Val Ala Asp Phe Arg Arg Arg Gly 65 70 75 80 Val Asp Val Ser Gln Val Ala Trp Gln Ser Lys Gly Asp Thr Pro Ser 85 90 95 Ser Cys Cys Ile Ile Asn Asn Ser Asn Gly Asn Arg Thr Ile Val Leu 100 105 110 His Asp Thr Ser Leu Pro Asp Val Ser Ala Thr Asp Phe Glu Lys Val 115 120 125 Asp Leu Thr Gln Phe Lys Trp Ile His Ile Glu Gly Arg Asn Ala Ser 130 135 140 Glu Gln Val Lys Met Leu Gln Arg Ile Asp Ala His Asn Thr Arg Gln 145 150 155 160 Pro Pro Glu Gln Lys Ile Arg Val Ser Val Glu Val Glu Lys Pro Arg 165 170 175 Glu Glu Leu Phe Gln Leu Phe Gly Tyr Gly Asp Val Val Phe Val Ser 180 185 190 Lys Asp Val Ala Lys His Leu Gly Phe Gln Ser Ala Glu Glu Ala Leu 195 200 205 Arg Gly Leu Tyr Gly Arg Val Arg Lys Gly Ala Val Leu Val Cys Ala 210 215 220 Trp Ala Glu Glu Gly Ala Asp Ala Leu Gly Pro Asp Gly Lys Leu Leu 225 230 235 240 His Ser Asp Ala Phe Pro Pro Pro Arg Val Val Asp Thr Leu Gly Ala 245 250 255 Gly Asp Thr Phe Asn Ala Ser Val Ile Phe Ser Leu Ser Gln Gly Arg 260 265 270 Ser Val Gln Glu Ala Leu Arg Phe Gly Cys Gln Val Ala Gly Lys Lys 275 280 285 Cys Gly Leu Gln Gly Phe Asp Gly Ile Val 290 295 <210> 13 <211> 1455 <212> DNA <213> Escherichia coli <400> 13 atgtatatcg ggatagatct tggcacctcg ggcgtaaaag ttattttgct caacgagcag 60 ggtgaggtgg ttgctgcgca aacggaaaag ctgaccgttt cgcgcccgca tccactctgg 120 tcggaacaag acccggaaca gtggtggcag gcaactgatc gcgcaatgaa agctctgggc 180 gatcagcatt ctctgcagga cgttaaagca ttgggtattg ccggccagat gcacggagca 240 accttgctgg atgctcagca acgggtgtta cgccctgcca ttttgtggaa cgacgggcgc 300 tgtgcgcaag agtgcacttt gctggaagcg cgagttccgc aatcgcgggt gattaccggc 360 aacctgatga tgcccggatt tactgcgcct aaattgctat gggttcagcg gcatgagccg 420 gagatattcc gtcaaatcga caaagtatta ttaccgaaag attacttgcg tctgcgtatg 480 acgggggagt ttgccagcga tatgtctgac gcagctggca ccatgtggct ggatgtcgca 540 aagcgtgact ggagtgacgt catgctgcag gcttgcgact tatctcgtga ccagatgccc 600 gcattatacg aaggcagcga aattactggt gctttgttac ctgaagttgc gaaagcgtgg 660 ggtatggcga cggtgccagt tgtcgcaggc ggtggcgaca atgcagctgg tgcagttggt 720 gtgggaatgg ttgatgctaa tcaggcaatg ttatcgctgg ggacgtcggg ggtctatttt 780 gctgtcagcg aagggttctt aagcaagcca gaaagcgccg tacatagctt ttgccatgcg 840 ctaccgcaac gttggcattt aatgtctgtg atgctgagtg cagcgtcgtg tctggattgg 900 gccgcgaaat taaccggcct gagcaatgtc ccagctttaa tcgctgcagc tcaacaggct 960 gatgaaagtg ccgagccagt ttggtttctg ccttatcttt ccggcgagcg tacgccacac 1020 aataatcccc aggcgaaggg ggttttcttt ggtttgactc atcaacatgg ccccaatgaa 1080 ctggcgcgag cagtgctgga aggcgtgggt tatgcgctgg cagatggcat ggatgtcgtg 1140 catgcctgcg gtattaaacc gcaaagtgtt acgttgattg ggggcggggc gcgtagtgag 1200 tactggcgtc agatgctggc ggatatcagc ggtcagcagc tcgattaccg tacggggggg 1260 gatgtggggc cagcactggg cgcagcaagg ctggcgcaga tcgcggcgaa tccagagaaa 1320 tcgctcattg aattgttgcc gcaactaccg ttagaacagt cgcatctacc agatgcgcag 1380 cgttatgccg cttatcagcc acgacgagaa acgttccgtc gcctctatca gcaacttctg 1440 ccattaatgg cgtaa 1455 <210> 14 <211> 484 <212> PRT <213> Escherichia coli <400> 14 Met Tyr Ile Gly Ile Asp Leu Gly Thr Ser Gly Val Lys Val Ile Leu 1 5 10 15 Leu Asn Glu Gln Gly Glu Val Val Ala Ala Gln Thr Glu Lys Leu Thr 20 25 30 Val Ser Arg Pro His Pro Leu Trp Ser Glu Gln Asp Pro Glu Gln Trp 35 40 45 Trp Gln Ala Thr Asp Arg Ala Met Lys Ala Leu Gly Asp Gln His Ser 50 55 60 Leu Gln Asp Val Lys Ala Leu Gly Ile Ala Gly Gln Met His Gly Ala 65 70 75 80 Thr Leu Leu Asp Ala Gln Gln Arg Val Leu Arg Pro Ala Ile Leu Trp 85 90 95 Asn Asp Gly Arg Cys Ala Gln Glu Cys Thr Leu Leu Glu Ala Arg Val 100 105 110 Pro Gln Ser Arg Val Ile Thr Gly Asn Leu Met Met Pro Gly Phe Thr 115 120 125 Ala Pro Lys Leu Leu Trp Val Gln Arg His Glu Pro Glu Ile Phe Arg 130 135 140 Gln Ile Asp Lys Val Leu Leu Pro Lys Asp Tyr Leu Arg Leu Arg Met 145 150 155 160 Thr Gly Glu Phe Ala Ser Asp Met Ser Asp Ala Ala Gly Thr Met Trp 165 170 175 Leu Asp Val Ala Lys Arg Asp Trp Ser Asp Val Met Leu Gln Ala Cys 180 185 190 Asp Leu Ser Arg Asp Gln Met Pro Ala Leu Tyr Glu Gly Ser Glu Ile 195 200 205 Thr Gly Ala Leu Leu Pro Glu Val Ala Lys Ala Trp Gly Met Ala Thr 210 215 220 Val Pro Val Val Ala Gly Gly Gly Asp Asn Ala Ala Gly Ala Val Gly 225 230 235 240 Val Gly Met Val Asp Ala Asn Gln Ala Met Leu Ser Leu Gly Thr Ser 245 250 255 Gly Val Tyr Phe Ala Val Ser Glu Gly Phe Leu Ser Lys Pro Glu Ser 260 265 270 Ala Val His Ser Phe Cys His Ala Leu Pro Gln Arg Trp His Leu Met 275 280 285 Ser Val Met Leu Ser Ala Ala Ser Cys Leu Asp Trp Ala Ala Lys Leu 290 295 300 Thr Gly Leu Ser Asn Val Pro Ala Leu Ile Ala Ala Ala Gln Gln Ala 305 310 315 320 Asp Glu Ser Ala Glu Pro Val Trp Phe Leu Pro Tyr Leu Ser Gly Glu 325 330 335 Arg Thr Pro His Asn Asn Pro Gln Ala Lys Gly Val Phe Phe Gly Leu 340 345 350 Thr His Gln His Gly Pro Asn Glu Leu Ala Arg Ala Val Leu Glu Gly 355 360 365 Val Gly Tyr Ala Leu Ala Asp Gly Met Asp Val Val His Ala Cys Gly 370 375 380 Ile Lys Pro Gln Ser Val Thr Leu Ile Gly Gly Gly Ala Arg Ser Glu 385 390 395 400 Tyr Trp Arg Gln Met Leu Ala Asp Ile Ser Gly Gln Gln Leu Asp Tyr 405 410 415 Arg Thr Gly Gly Asp Val Gly Pro Ala Leu Gly Ala Ala Arg Leu Ala 420 425 430 Gln Ile Ala Ala Asn Pro Glu Lys Ser Leu Ile Glu Leu Leu Pro Gln 435 440 445 Leu Pro Leu Glu Gln Ser His Leu Pro Asp Ala Gln Arg Tyr Ala Ala 450 455 460 Tyr Gln Pro Arg Arg Glu Thr Phe Arg Arg Leu Tyr Gln Gln Leu Leu 465 470 475 480 Pro Leu Met Ala <210> 15 <211> 747 <212> DNA <213> Caulobacter crescentus <400> 15 atgagcagcg cgatctaccc gagcctgaaa ggtaaacgtg tggtgattac cggcggcggc 60 agcggcattg gtgcgggcct gaccgcgggc ttcgcgcgtc agggtgcgga agtgatcttt 120 ctggacattg cggacgaaga tagccgtgcg ctggaggcgg aactggcggg cagcccgatc 180 ccgccggtgt acaagcgttg cgatctgatg aacctggagg cgatcaaagc ggttttcgcg 240 gaaattggcg acgtggatgt tctggtgaac aacgcgggta acgacgaccg tcacaagctg 300 gcggatgtga ccggtgcgta ttgggatgag cgtattaacg ttaacctgcg tcacatgctg 360 ttctgcaccc aggcggtggc gccgggtatg aagaaacgtg gtggcggtgc ggttatcaac 420 tttggcagca ttagctggca cctgggtctg gaggacctgg tgctgtacga aaccgcgaaa 480 gcgggcatcg agggtatgac ccgtgcgctg gcgcgtgaac tgggtccgga cgatattcgt 540 gtgacctgcg tggttccggg taacgttaag accaaacgtc aagagaagtg gtataccccg 600 gagggtgaag cgcagattgt tgcggcgcaa tgcctgaaag gtcgtattgt tccggaaaac 660 gtggcggcgc tggttctgtt tctggcgagc gatgatgcga gcctgtgcac cggccatgag 720 tattggattg atgcgggctg gcgttaa 747 <210> 16 <211> 248 <212> PRT <213> Caulobacter crescentus <400> 16 Met Ser Ser Ala Ile Tyr Pro Ser Leu Lys Gly Lys Arg Val Val Ile 1 5 10 15 Thr Gly Gly Gly Ser Gly Ile Gly Ala Gly Leu Thr Ala Gly Phe Ala 20 25 30 Arg Gln Gly Ala Glu Val Ile Phe Leu Asp Ile Ala Asp Glu Asp Ser 35 40 45 Arg Ala Leu Glu Ala Glu Leu Ala Gly Ser Pro Ile Pro Pro Val Tyr 50 55 60 Lys Arg Cys Asp Leu Met Asn Leu Glu Ala Ile Lys Ala Val Phe Ala 65 70 75 80 Glu Ile Gly Asp Val Asp Val Leu Val Asn Asn Ala Gly Asn Asp Asp 85 90 95 Arg His Lys Leu Ala Asp Val Thr Gly Ala Tyr Trp Asp Glu Arg Ile 100 105 110 Asn Val Asn Leu Arg His Met Leu Phe Cys Thr Gln Ala Val Ala Pro 115 120 125 Gly Met Lys Lys Arg Gly Gly Gly Ala Val Ile Asn Phe Gly Ser Ile 130 135 140 Ser Trp His Leu Gly Leu Glu Asp Leu Val Leu Tyr Glu Thr Ala Lys 145 150 155 160 Ala Gly Ile Glu Gly Met Thr Arg Ala Leu Ala Arg Glu Leu Gly Pro 165 170 175 Asp Asp Ile Arg Val Thr Cys Val Val Pro Gly Asn Val Lys Thr Lys 180 185 190 Arg Gln Glu Lys Trp Tyr Thr Pro Glu Gly Glu Ala Gln Ile Val Ala 195 200 205 Ala Gln Cys Leu Lys Gly Arg Ile Val Pro Glu Asn Val Ala Ala Leu 210 215 220 Val Leu Phe Leu Ala Ser Asp Asp Ala Ser Leu Cys Thr Gly His Glu 225 230 235 240 Tyr Trp Ile Asp Ala Gly Trp Arg 245 <210> 17 <211> 255 <212> PRT <213> Burkholderia xenovorans <400> 17 Methionine, Serine, Tyrosine, Alanine, Isoleucine, Tyrosine, Proline, Serine, Leucine, Serine, Glycine, Lysine, Threonine, Valine, Valine, Isoleucine 1 5 10 15 Threonine, Glycine, Glycine, Glycine, Serine, Glycine, Isoleucine, Glycine, Alanine, Alanine, Methionine, Valine, Glutamic acid, Alanine, Phenylalanine, Alanine 20 25 30 Arginine, Glutamine, Glycine, Alanine, Arginine, Valine, Phenylalanine, Phenylalanine, Leucine, Aspartic acid, Valine, Alanine, Glutamic acid, Aspartic acid, Aspartic acid, Serine 35 40 45 Leucine, Alanine, Leucine, Glutamine, Glutamine, Serine, Leucine, Serine, Aspartic acid, Alanine, Proline, Histidine, Proline, Proline, Leucine, Phenylalanine 50 55 60 Arginine, Arginine, Cysteine, Aspartic acid, Leucine, Arginine, Serine, Valine, Aspartic acid, Alanine, Isoleucine, Histidine, Serine, Alanine, Phenylalanine, Alanine 65 70 75 80 Glycine, Isoleucine, Valine, Glutamic acid, Isoleucine, Alanine, Glycine, Proline, Isoleucine, Glutamic acid, Valine, Leucine, Valine, Asparagine, Asparagine, Alanine 85 90 95 Glycine, Asparagine, Aspartic acid, Aspartic acid, Arginine, Histidine, Glutamic acid, Valine, Aspartic acid, Alanine, Isoleucine, Threonine, Proline, Alanine, Tyrosine, Tryptophan 100 105 110 Aspartic acid, Glutamic acid, Arginine, Methionine, Alanine, Valine, Asparagine, Leucine, Arginine, Histidine, Glutamine, Phenylalanine, Phenylalanine, Cysteine, Alanine, Glutamine 115 120 125 Alanine, Alanine, Alanine, Alanine, Glycine, Methionine, Arginine, Lysine, Isoleucine, Glycine, Arginine, Glycine, Valine, Isoleucine, Leucine, Asparagine 130 135 140 Leucine, Glycine, Serine, Valine, Serine, Tryptophan, Histidine, Leucine, Alanine, Leucine, Proline, Asparagine, Leucine, Alanine, Isoleucine, Tyrosine 145 150 155 160 Met Ser Ala Lys Ala Gly Ile Glu Gly Leu Thr Arg Gly Leu Ala Arg 165 170 175 Asp Leu Gly Ala Ala Gly Ile Arg Val Asn Cys Ile Ile Pro Gly Ala 180 185 190 Val Arg Thr Pro Arg Gln Met Gln Leu Trp Gln Ser Pro Glu Ser Glu 195 200 205 Ala Lys Leu Val Ala Ser Gln Cys Leu Arg Leu Arg Ile Glu Pro Glu 210 215 220 His Val Ala Arg Met Ala Leu Phe Leu Ala Ser Asp Asp Ala Ser Arg 225 230 235 240 Cys Ser Gly Arg Asp Tyr Phe Val Asp Ala Gly Trp Tyr Gly Glu 245 250 255 <210> 18 <211> 1173 <212> DNA <213> Haloferax volcanii <400> 18 atgagccccg cccccaccga catcgtcgag gagttcacgc gccgcgactg gcagggagac 60 gacgtgacgg gcaccgtgcg ggtcgccatg atcggcctcg gctggtggac ccgcgacgag 120 gcgattcccg cggtcgaggc gtccgagttc tgcgagacga cggtcgtcgt cagcagttcg 180 aaggagaaag ccgagggcgc gacggcgttg accgagtcga taacccacgg cctcacctac 240 gacgagttcc acgagggggt cgccgccgac gcctacgacg cggtgtacgt cgtcacgccg 300 aacggtctgc atctcccgta cgtcgagacc gccgccgagt tggggaaggc ggtcctctgc 360 gagaaaccgc tggaagcgtc ggtcgagcgg gccgaaaagc tcgtcgccgc ctgcgaccgc 420 gccgacgtgc ccctgatggt cgcctatcgg atgcagaccg agccggccgt ccggcgcgcc 480 cgcgaactcg tcgaggccgg cgtcatcggc gagccggtgt tcgtccacgg ccacatgtcc 540 cagcgcctgc tcgacgaggt cgtccccgac cccgaccagt ggcggctcga ccccgaactc 600 tccggcggcg cgaccgtcat ggacatcggg ctctacccgc tgaacaccgc ccggttcgtc 660 ctcgacgccg accccgtccg cgtcagggcg accgcccgcg tcgacgacga ggcgttcgag 720 gccgtcggcg acgagcacgt cagtttcggc gtcgacttcg acgacggcac gctcgcggtc 780 tgcaccgcca gccagtcggc ttaccagttg agccacctcc gggtgaccgg caccgagggc 840 gaactcgaaa tcgagcccgc gttctacaac cgccaaaagc ggggattccg actgtcgtgg 900 ggggaccagt ccgccgacta cgacttcgag caggtaaacc agatgacgga ggagttcgac 960 tacttcgcgt cccggctcct gtcggattcc gaccccgcgc ccgacggcga ccacgcgctc 1020 gtggacatgc gcgcgatgga cgcgatttac gccgcggcgg agcgcgggac cgatgtcgcc 1080 gtcgacgccg ccgactccga ttccgccgac tccgattccg ccgacgctgc cgccgccaac 1140 cacgacgccg accccgattc cgacgggacg tag 1173 <210> 19 <211> 390 <212> PRT <213> Haloferax volcanii <400> 19 Met Ser Pro Ala Pro Thr Asp Ile Val Glu Glu Phe Thr Arg Arg Asp 1 5 10 15 Trp Gln Gly Asp Asp Val Thr Gly Thr Val Arg Val Ala Met Ile Gly 20 25 30 Leu Gly Trp Trp Thr Arg Asp Glu Ala Ile Pro Ala Val Glu Ala Ser 35 40 45 Glu Phe Cys Glu Thr Thr Val Val Val Ser Ser Ser Lys Glu Lys Ala 50 55 60 Glu Gly Ala Thr Ala Leu Thr Glu Ser Ile Thr His Gly Leu Thr Tyr 65 70 75 80 Asp Glu Phe His Glu Gly Val Ala Ala Asp Ala Tyr Asp Ala Val Tyr 85 90 95 Val Val Thr Pro Asn Gly Leu His Leu Pro Tyr Val Glu Thr Ala Ala 100 105 110 Glu Leu Gly Lys Ala Val Leu Cys Glu Lys Pro Leu Glu Ala Ser Val 115 120 125 Glu Arg Ala Glu Lys Leu Val Ala Ala Cys Asp Arg Ala Asp Val Pro 130 135 140 Leu Met Val Ala Tyr Arg Met Gln Thr Glu Pro Ala Val Arg Arg Ala 145 150 155 160 Arg Glu Leu Val Glu Ala Gly Val Ile Gly Glu Pro Val Phe Val His 165 170 175 Gly His Met Ser Gln Arg Leu Leu Asp Glu Val Val Pro Asp Pro Asp 180 185 190 Gln Trp Arg Leu Asp Pro Glu Leu Ser Gly Gly Ala Thr Val Met Asp 195 200 205 Ile Gly Leu Tyr Pro Leu Asn Thr Ala Arg Phe Val Leu Asp Ala Asp 210 215 220 Pro Val Arg Val Arg Ala Thr Ala Arg Val Asp Asp Glu Ala Phe Glu 225 230 235 240 Ala Val Gly Asp Glu His Val Ser Phe Gly Val Asp Phe Asp Asp Gly 245 250 255 Thr Leu Ala Val Cys Thr Ala Ser Gln Ser Ala Tyr Gln Leu Ser His 260 265 270 Leu Arg Val Thr Gly Thr Glu Gly Glu Leu Glu Ile Glu Pro Ala Phe 275 280 285 Tyr Asn Arg Gln Lys Arg Gly Phe Arg Leu Ser Trp Gly Asp Gln Ser 290 295 300 Ala Asp Tyr Asp Phe Glu Gln Val Asn Gln Met Thr Glu Glu Phe Asp 305 310 315 320 Tyr Phe Ala Ser Arg Leu Leu Ser Asp Ser Asp Pro Ala Pro Asp Gly 325 330 335 Asp His Ala Leu Val Asp Met Arg Ala Met Asp Ala Ile Tyr Ala Ala 340 345 350 Ala Glu Arg Gly Thr Asp Val Ala Val Asp Ala Ala Asp Ser Asp Ser 355 360 365 Ala Asp Ser Asp Ser Ala Asp Ala Ala Ala Ala Asn His Asp Ala Asp 370 375 380 Pro Asp Ser Asp Gly Thr 385 390 <210> 20 <211> 990 <212> DNA <213> Escherichia coli <400> 20 atgcacaaat ttactaaagc cctggcagcc attggtctgg cagccgttat gtcacaatcc 60 gctatggcgg agaacctgaa gctcggtttt ctggtgaagc aaccggaaga gccgtggttc 120 cagaccgaat ggaagtttgc cgataaagcc gggaaggatt tagggtttga ggttattaag 180 attgccgtgc cggatggcga aaaaacattg aacgcgatcg acagcctggc tgccagtggc 240 gcaaaaggtt tcgttatttg tactccggac cccaaactcg gctctgccat cgtcgcgaaa 300 gcgcgtggct acgatatgaa agtcattgcc gtggatgacc agtttgttaa cgccaaaggt 360 aagccaatgg ataccgttcc gctggtgatg atggcggcga ctaaaattgg cgaacgtcag 420 ggccaggaac tgtataaaga gatgcagaaa cgtggctggg atgtcaaaga aagcgcggtg 480 atggcgatta ccgccaacga actggatacc gcccgccgcc gtactacggg atctatggat 540 gcgctgaaag cggccggatt cccggaaaaa caaatttatc aggtacctac caaatctaac 600 gacatcccgg gggcatttga cgctgccaac tcaatgctgg ttcaacatcc ggaagttaaa 660 cattggctga tcgtcggtat gaacgacagc accgtgctgg gcggcgtacg cgcgacggaa 720 ggtcagggct ttaaagcggc cgatatcatc ggcattggca ttaacggtgt ggatgcggtg 780 agcgaactgt ctaaagcaca ggcaaccggc ttctacggtt ccctgctgcc aagcccggac 840 gtacatggct ataaatccag cgaaatgctt tacaactggg tagcaaaaga cgttgaaccg 900 ccaaaattta ccgaagttac cgacgtggta ctgatcacgc gtgacaactt taaagaagaa 960 ctggagaaaa aaggtttagg cggtaagtaa 990 <210> 21 <211> 1515 <212> DNA <213> Escherichia coli <400> 21 atgcaacagt ctaccccgta tctctcattt cgcggcatcg gtaaaacgtt tcccggcgtt 60 aaggcgctga cggatattag ttttgactgc tatgccggtc aggttcatgc gttgatgggt 120 gaaaatggcg caggaaaatc aactctctta aaaatcctca gcggcaacta tgcgccaacc 180 acgggttctg tagtgattaa tgggcaggaa atgtcctttt ccgacacgac cgcagcactt 240 aacgcgggcg tggcgattat ttaccaggaa ctgcatctcg tgccggaaat gaccgtcgcg 300 gaaaacatct atctcggcca gctgccgcat aaaggcggca ttgtgaatcg ctcattgctg 360 aattatgagg cgggtttaca acttaaacat cttggtatgg atattgaccc ggacacgccg 420 ctgaaatatc tctccattgg tcagtggcag atggttgaaa tcgccaaagc gctggcgcgt 480 aacgccaaaa ttatcgcctt tgatgagcca accagctccc tctctgcccg tgaaatcgac 540 aatcttttcc gcgttattcg tgaactgcga aaagaggggc gggtaatctt atacgtttct 600 caccgtatgg aagaaatatt tgccctcagc gatgccatta ctgtctttaa agatggacgt 660 tatgtcaaaa cctttaccga tatgcagcag gttgaccacg acgcgctggt gcaggcgatg 720 gtcgggcgcg acattggcga tatctacggc tggcaaccgc gtagttatgg cgaggagcgc 780 ctacgtcttg atgctgtgaa agcaccaggc gtgcgtacgc caataagtct ggcggttcgc 840 agtggtgaaa ttgttgggct gtttggtctg gtaggggcgg ggcgtagcga attaatgaaa 900 ggcatgtttg gcgggacgca aatcaccgcc ggtcaggttt atatcgacca acagccgatc 960 gatattcgta aaccgagcca cgccattgcc gcaggcatga tgctctgccc ggaagatcgc 1020 aaagcggaag gcattattcc cgtgcactcc gttcgcgaca atatcaacat cagtgccaga 1080 cgtaaacatg tgctcggcgg ttgtgtaatc aacaacggtt gggaagaaaa caatgccgat 1140 caccacattc gttcgctcaa catcaaaacg ccgggcgcgg agcaactgat catgaatctc 1200 tcaggcggaa atcagcaaaa agccattctg ggccgctggt tatcggaaga gatgaaggtc 1260 attttgctgg atgaacctac gcgcggcatt gatgttggcg ctaagcacga aatatataac 1320 gtaatttatg cgctggcggc gcagggcgtg gcggtgctgt ttgcctccag cgacttacct 1380 gaagtcctcg gcgttgccga ccggattgtg gtgatgcggg aaggtgaaat cgccggtgaa 1440 ttgttacacg agcaggcaga tgagcgtcag gcactgagcc ttgcgatgcc taaagtcagc 1500 caggctgttg cctga 1515 <210> 22 <211> 987 <212> DNA <213> Escherichia coli <400> 22 atgtcttctg tttctacatc ggggtctggc gcacctaagt cgtcattcag cttcgggcgt 60 atctgggatc agtacggcat gctggtggtg tttgcggtgc tctttatcgc ctgtgccatt 120 tttgtcccaa attttgccac cttcattaat atgaaagggt tgggcctggc aatttccatg 180 tcggggatgg tggcttgtgg catgttgttc tgcctcgctt ccggtgactt tgacctttct 240 gtcgcctccg taattgcctg tgcgggtgtc accacggcgg tggttattaa cctgactgaa 300 agcctgtgga ttggcgtggc agcggggttg ttgctgggcg ttctctgtgg cctggtcaat 360 ggctttgtta tcgccaaact gaaaataaat gctctgatca cgacattggc aacgatgcag 420 attgttcgag gtctggcgta catcatttca gacggtaaag cggtcggtat cgaagatgaa 480 agcttctttg cccttggtta cgccaactgg ttcggtctgc ctgcgccaat ctggctcacc 540 gtcgcgtgtc tgattatctt tggtttgctg ctgaataaaa ccacctttgg tcgtaacacc 600 ctggcgattg gcgggaacga agaggccgcg cgtctggcgg gtgtaccggt tgttcgcacc 660 aaaattatta tctttgttct ctcaggcctg gtatcagcga tagccggaat tattctggct 720 tcacgtatga ccagtgggca gccaatgacg tcgattggtt atgagctgat tgttatctcc 780 gcctgcgttt taggtggcgt ttctctgaaa ggtggcatcg gaaaaatctc atatgtggtg 840 gcgggtatct taattttagg caccgtggaa aacgccatga acctgcttaa tatttctcct 900 ttcgcgcagt acgtggttcg cggcttaatc ctgctggcag cggtgatctt cgaccgttac 960 aagcaaaaag cgaaacgcac tgtctga 987 <210> 23 <211> 329 <212> PRT <213> Escherichia coli <400> 23 Met His Lys Phe Thr Lys Ala Leu Ala Ala Ile Gly Leu Ala Ala Val 1 5 10 15 Met Ser Gln Ser Ala Met Ala Glu Asn Leu Lys Leu Gly Phe Leu Val 20 25 30 Lys Gln Pro Glu Glu Pro Trp Phe Gln Thr Glu Trp Lys Phe Ala Asp 35 40 45 Lys Ala Gly Lys Asp Leu Gly Phe Glu Val Ile Lys Ile Ala Val Pro 50 55 60 Asp Gly Glu Lys Thr Leu Asn Ala Ile Asp Ser Leu Ala Ala Ser Gly 65 70 75 80 Ala Lys Gly Phe Val Ile Cys Thr Pro Asp Pro Lys Leu Gly Ser Ala 85 90 95 Ile Val Ala Lys Ala Arg Gly Tyr Asp Met Lys Val Ile Ala Val Asp 100 105 110 Asp Gln Phe Val Asn Ala Lys Gly Lys Pro Met Asp Thr Val Pro Leu 115 120 125 Val Met Met Ala Ala Thr Lys Ile Gly Glu Arg Gln Gly Gln Glu Leu 130 135 140 Tyr Lys Glu Met Gln Lys Arg Gly Trp Asp Val Lys Glu Ser Ala Val 145 150 155 160 Met Ala Ile Thr Ala Asn Glu Leu Asp Thr Ala Arg Arg Arg Thr Thr 165 170 175 Gly Ser Met Asp Ala Leu Lys Ala Ala Gly Phe Pro Glu Lys Gln Ile 180 185 190 Tyr Gln Val Pro Thr Lys Ser Asn Asp Ile Pro Gly Ala Phe Asp Ala 195 200 205 Ala Asn Ser Met Leu Val Gln His Pro Glu Val Lys His Trp Leu Ile 210 215 220 Val Gly Met Asn Asp Ser Thr Val Leu Gly Gly Val Arg Ala Thr Glu 225 230 235 240 Gly Gln Gly Phe Lys Ala Ala Asp Ile Ile Gly Ile Gly Ile Asn Gly 245 250 255 Val Asp Ala Val Ser Glu Leu Ser Lys Ala Gln Ala Thr Gly Phe Tyr 260 265 270 Gly Ser Leu Leu Pro Ser Pro Asp Val His Gly Tyr Lys Ser Ser Glu 275 280 285 Met Leu Tyr Asn Trp Val Ala Lys Asp Val Glu Pro Pro Lys Phe Thr 290 295 300 Glu Val Thr Asp Val Val Leu Ile Thr Arg Asp Asn Phe Lys Glu Glu 305 310 315 320 Leu Glu Lys Lys Gly Leu Gly Gly Lys 325 <210> 24 <211> 504 <212> PRT <213> Escherichia coli <400> 24 Met Gln Gln Ser Thr Pro Tyr Leu Ser Phe Arg Gly Ile Gly Lys Thr 1 5 10 15 Phe Pro Gly Val Lys Ala Leu Thr Asp Ile Ser Phe Asp Cys Tyr Ala 20 25 30 Gly Gln Val His Ala Leu Met Gly Glu Asn Gly Ala Gly Lys Ser Thr 35 40 45 Leu Leu Lys Ile Leu Ser Gly Asn Tyr Ala Pro Thr Thr Gly Ser Val 50 55 60 Val Ile Asn Gly Gln Glu Met Ser Phe Ser Asp Thr Thr Ala Ala Leu 65 70 75 80 Asn Ala Gly Val Ala Ile Ile Tyr Gln Glu Leu His Leu Val Pro Glu 85 90 95 Met Thr Val Ala Glu Asn Ile Tyr Leu Gly Gln Leu Pro His Lys Gly 100 105 110 Gly Ile Val Asn Arg Ser Leu Leu Asn Tyr Glu Ala Gly Leu Gln Leu 115 120 125 Lys His Leu Gly Met Asp Ile Asp Pro Asp Thr Pro Leu Lys Tyr Leu 130 135 140 Ser Ile Gly Gln Trp Gln Met Val Glu Ile Ala Lys Ala Leu Ala Arg 145 150 155 160 Asn Ala Lys Ile Ile Ala Phe Asp Glu Pro Thr Ser Ser Leu Ser Ala 165 170 175 Arg Glu Ile Asp Asn Leu Phe Arg Val Ile Arg Glu Leu Arg Lys Glu 180 185 190 Gly Arg Val Ile Leu Tyr Val Ser His Arg Met Glu Glu Ile Phe Ala 195 200 205 Leu Ser Asp Ala Ile Thr Val Phe Lys Asp Gly Arg Tyr Val Lys Thr 210 215 220 Phe Thr Asp Met Gln Gln Val Asp His Asp Ala Leu Val Gln Ala Met 225 230 235 240 Val Gly Arg Asp Ile Gly Asp Ile Tyr Gly Trp Gln Pro Arg Ser Tyr 245 250 255 Gly Glu Glu Arg Leu Arg Leu Asp Ala Val Lys Ala Pro Gly Val Arg 260 265 270 Thr Pro Ile Ser Leu Ala Val Arg Ser Gly Glu Ile Val Gly Leu Phe 275 280 285 Gly Leu Val Gly Ala Gly Arg Ser Glu Leu Met Lys Gly Met Phe Gly 290 295 300 Gly Thr Gln Ile Thr Ala Gly Gln Val Tyr Ile Asp Gln Gln Pro Ile 305 310 315 320 Asp Ile Arg Lys Pro Ser His Ala Ile Ala Ala Gly Met Met Leu Cys 325 330 335 Pro Glu Asp Arg Lys Ala Glu Gly Ile Ile Pro Val His Ser Val Arg 340 345 350 Asp Asn Ile Asn Ile Ser Ala Arg Arg Lys His Val Leu Gly Gly Cys 355 360 365 Val Ile Asn Asn Gly Trp Glu Glu Asn Asn Ala Asp His His Ile Arg 370 375 380 Ser Leu Asn Ile Lys Thr Pro Gly Ala Glu Gln Leu Ile Met Asn Leu 385 390 395 400 Ser Gly Gly Asn Gln Gln Lys Ala Ile Leu Gly Arg Trp Leu Ser Glu 405 410 415 Glu Met Lys Val Ile Leu Leu Asp Glu Pro Thr Arg Gly Ile Asp Val 420 425 430 Gly Ala Lys His Glu Ile Tyr Asn Val Ile Tyr Ala Leu Ala Ala Gln 435 440 445 Gly Val Ala Val Leu Phe Ala Ser Ser Asp Leu Pro Glu Val Leu Gly 450 455 460 Val Ala Asp Arg Ile Val Val Met Arg Glu Gly Glu Ile Ala Gly Glu 465 470 475 480 Leu Leu His Glu Gln Ala Asp Glu Arg Gln Ala Leu Ser Leu Ala Met 485 490 495 Pro Lys Val Ser Gln Ala Val Ala 500 <210> 25 <211> 328 <212> PRT <213> Escherichia coli <400> 25 Met Ser Ser Val Ser Thr Ser Gly Ser Gly Ala Pro Lys Ser Ser Phe 1 5 10 15 Ser Phe Gly Arg Ile Trp Asp Gln Tyr Gly Met Leu Val Val Phe Ala 20 25 30 Val Leu Phe Ile Ala Cys Ala Ile Phe Val Pro Asn Phe Ala Thr Phe 35 40 45 Ile Asn Met Lys Gly Leu Gly Leu Ala Ile Ser Met Ser Gly Met Val 50 55 60 Ala Cys Gly Met Leu Phe Cys Leu Ala Ser Gly Asp Phe Asp Leu Ser 65 70 75 80 Val Ala Ser Val Ile Ala Cys Ala Gly Val Thr Thr Ala Val Val Ile 85 90 95 Asn Leu Thr Glu Ser Leu Trp Ile Gly Val Ala Ala Gly Leu Leu Leu 100 105 110 Gly Val Leu Cys Gly Leu Val Asn Gly Phe Val Ile Ala Lys Leu Lys 115 120 125 Ile Asn Ala Leu Ile Thr Thr Leu Ala Thr Met Gln Ile Val Arg Gly 130 135 140 Leu Ala Tyr Ile Ile Ser Asp Gly Lys Ala Val Gly Ile Glu Asp Glu 145 150 155 160 Ser Phe Phe Ala Leu Gly Tyr Ala Asn Trp Phe Gly Leu Pro Ala Pro 165 170 175 Ile Trp Leu Thr Val Ala Cys Leu Ile Ile Phe Gly Leu Leu Leu Asn 180 185 190 Lys Thr Thr Phe Gly Arg Asn Thr Leu Ala Ile Gly Gly Asn Glu Glu 195 200 205 Ala Ala Arg Leu Ala Gly Val Pro Val Val Arg Thr Lys Ile Ile Ile 210 215 220 Phe Val Leu Ser Gly Leu Val Ser Ala Ile Ala Gly Ile Ile Leu Ala 225 230 235 240 Ser Arg Met Thr Ser Gly Gln Pro Met Thr Ser Ile Gly Tyr Glu Leu 245 250 255 Ile Val Ile Ser Ala Cys Val Leu Gly Gly Val Ser Leu Lys Gly Gly 260 265 270 Ile Gly Lys Ile Ser Tyr Val Val Ala Gly Ile Leu Ile Leu Gly Thr 275 280 285 Val Glu Asn Ala Met Asn Leu Leu Asn Ile Ser Pro Phe Ala Gln Tyr 290 295 300 Val Val Arg Gly Leu Ile Leu Leu Ala Ala Val Ile Phe Asp Arg Tyr 305 310 315 320 Lys Gln Lys Ala Lys Arg Thr Val 325 <210> 26 <211> 993 <212> DNA <213> Escherichia coli <400> 26 atgaaaataa agaacattct actcaccctt tgcacctcac tcctgcttac caacgttgct 60 gcacacgcca aagaagtcaa aataggtatg gcgattgatg atctccgtct tgaacgctgg 120 caaaaagatc gagatatctt tgtgaaaaag gcagaatctc tcggcgcgaa agtatttgta 180 cagtctgcaa atggcaatga agaaacacaa atgtcgcaga ttgaaaacat gataaaccgg 240 ggtgtcgatg ttcttgtcat tattccgtat aacggtcagg tattaagtaa cgttgtaaaa 300 gaagccaaac aagaaggcat taaagtatta gcttacgacc gtatgattaa cgatgcggat 360 atcgattttt atatttcttt cgataacgaa aaagtcggtg aactgcaggc aaaagccctg 420 gtcgatattg ttccgcaagg taattacttc ctgatgggcg gctcgccggt agataacaac 480 gccaagctgt tccgcgccgg acaaatgaaa gtgttaaaac cttacgttga ttccggaaaa 540 attaaagtcg ttggtgacca atgggttgat ggctggttac cggaaaacgc attgaaaatt 600 atggaaaacg cgctaaccgc caataataac aaaattgatg ctgtagttgc ctcaaacgat 660 gccaccgcag gtggggcaat tcaggcatta agcgcgcaag gtttatcagg gaaagtagca 720 atctccggcc aggatgcgga tctcgcaggt attaaacgta ttgctgccgg tacgcaaact 780 atgacggtgt ataaacctat tacgttgttg gcaaatactg ccgcagaaat tgccgttgag 840 ttgggcaatg gtcaggaacc aaaagcagat accacactga ataatggcct gaaagatgtc 900 ccctcccgcc tcctgacacc gatcgatgtg aataaaaaca acatcaaaga tacggtaatt 960 aaagacggat tccacaaaga gagcgagctg taa 993 <210> 27 <211> 1542 <212> DNA <213> Escherichia coli <400> 27 atgccttatc tacttgaaat gaagaacatt accaaaacct tcggcagtgt gaaggcgatt 60 gataacgtct gcttgcggtt gaatgctggc gaaatcgtct cactttgtgg ggaaaatggg 120 tctggtaaat caacgctgat gaaagtgctg tgtggtattt atccccatgg ctcctacgaa 180 ggcgaaatta tttttgcggg agaagagatt caggcgagtc acatccgcga taccgaacgc 240 aaaggtatcg ccatcattca tcaggaattg gccctggtga aagaattgac cgtgctggaa 300 aatatcttcc tgggtaacga aataacccac aatggcatta tggattatga cctgatgacg 360 ctacgctgtc agaagctgct cgcacaggtc agtttatcca tttcacctga tacccgcgtt 420 ggcgatttag ggcttgggca acaacaactg gttgaaattg ccaaggcact taataaacag 480 gtgcgcttgt taattctcga tgaaccgaca gcctcattaa ctgagcagga aacgtcgatt 540 ttactggata ttattcgcga tctacaacag cacggtatcg cctgtattta tatttcgcac 600 aaactcaacg aagtcaaagc gatttccgat acgatttgcg ttattcgcga cggacagcac 660 attggtacgc gtgatgctgc cggaatgagt gaagacgata ttatcaccat gatggtcggg 720 cgagagttaa ccgcgcttta ccctaatgaa ccacatacca ccggagatga aatattacgt 780 attgaacatc tgacggcatg gcatccggtt aatcgtcata ttaaacgagt taatgatgtc 840 tcgttttccc tgaaacgtgg cgaaatattg ggtattgccg gactcgttgg tgccggacgt 900 accgagacca ttcagtgcct gtttggtgtg tggcccggac aatgggaagg aaaaatttat 960 attgatggca aacaggtaga tattcgtaac tgtcagcaag ccatcgccca ggggattgcg 1020 atggtccccg aagacagaaa gcgcgacggc atcgttccgg taatggcggt tggtaaaaat 1080 attaccctcg ccgcactcaa taaatttacc ggtggcatta gccagcttga tgacgcggca 1140 gagcaaaaat gtattctgga atcaatccag caactcaaag ttaaaacgtc gtcccccgac 1200 cttgctattg gacgtttgag cggcggcaat cagcaaaaag cgatcctcgc tcgctgtctg 1260 ttacttaacc cgcgcattct cattcttgat gaacccacca ggggtatcga tattggcgcg 1320 aaatacgaga tctacaaatt aattaaccaa ctcgtccagc agggtattgc cgttattgtc 1380 atctcttccg aattacctga agtgctcggc cttagcgatc gtgtactggt gatgcatgaa 1440 gggaaactaa aagccaacct gataaatcat aacctgactc aggagcaggt gatggaagcc 1500 gcattgagga gcgaacatca tgtcgaaaag caatccgtct ga 1542 <210> 28 <211> 1182 <212> DNA <213> Escherichia coli <400> 28 atgtcgaaaa gcaatccgtc tgaagtgaaa ttggccgtac cgacatccgg tggcttctcc 60 gggctgaaat cactgaattt gcaggtcttc gtgatgattg cagctatcat cgcaatcatg 120 ctgttcttta cctggaccac cgatggtgcc tacttaagcg cccgtaacgt ctccaacctg 180 ttacgccaga ccgcgattac cggcatcctc gcggtaggaa tggtgttcgt cataatttct 240 gctgaaatcg acctttccgt cggctcaatg atggggctgt taggtggcgt cgcggcgatt 300 tgtgacgtct ggttaggctg gcctttgcca cttaccatca ttgtgacgct ggttctggga 360 ctgcttctcg gtgcctggaa cggatggtgg gtcgcgtacc gtaaagtccc ttcatttatt 420 gtcaccctcg cgggcatgtt ggcatttcgc ggcatactca ttggcatcac caacggcacg 480 actgtatccc ccaccagcgc cgcgatgtca caaattgggc aaagctatct ccccgccagt 540 accggcttca tcattggcgc gcttggctta atggcttttg ttggttggca atggcgcgga 600 agaatgcgcc gtcaggcttt gggtttacag tctccggcct ctaccgcagt agtcggtcgc 660 caggctttaa ccgctatcat cgtattaggc gcaatctggc tgttgaatga ttaccgtggc 720 gttcccactc ctgttctgct gctgacgttg ctgttactcg gcggaatgtt tatggcaacg 780 cggacggcat ttggacgacg catttatgcc atcggcggca atctggaagc agcacgtctc 840 tccgggatta acgttgaacg caccaaactt gccgtgttcg cgattaacgg attaatggta 900 gccatcgccg gattaatcct tagttctcga cttggcgctg gttcaccttc tgcgggaaat 960 atcgccgaac tggacgcaat tgcagcatgc gtgattggcg gcaccagcct ggctggcggt 1020 gtgggaagcg ttgccggagc agtaatgggg gcatttatca tggcttcact ggataacggc 1080 atgagtatga tggatgtacc gaccttctgg cagtatatcg ttaaaggtgc gattctgttg 1140 ctggcagtat ggatggactc cgcaaccaaa cgccgttctt ga 1182 <210> 29 <211> 330 <212> PRT <213> Escherichia coli <400> 29 Met Lys Ile Lys Asn Ile Leu Leu Thr Leu Cys Thr Ser Leu Leu Leu 1 5 10 15 Thr Asn Val Ala Ala His Ala Lys Glu Val Lys Ile Gly Met Ala Ile 20 25 30 Asp Asp Leu Arg Leu Glu Arg Trp Gln Lys Asp Arg Asp Ile Phe Val 35 40 45 Lys Lys Ala Glu Ser Leu Gly Ala Lys Val Phe Val Gln Ser Ala Asn 50 55 60 Gly Asn Glu Glu Thr Gln Met Ser Gln Ile Glu Asn Met Ile Asn Arg 65 70 75 80 Gly Val Asp Val Leu Val Ile Ile Pro Tyr Asn Gly Gln Val Leu Ser 85 90 95 Asn Val Val Lys Glu Ala Lys Gln Glu Gly Ile Lys Val Leu Ala Tyr 100 105 110 Asp Arg Met Ile Asn Asp Ala Asp Ile Asp Phe Tyr Ile Ser Phe Asp 115 120 125 Asn Glu Lys Val Gly Glu Leu Gln Ala Lys Ala Leu Val Asp Ile Val 130 135 140 Pro Gln Gly Asn Tyr Phe Leu Met Gly Gly Ser Pro Val Asp Asn Asn 145 150 155 160 Ala Lys Leu Phe Arg Ala Gly Gln Met Lys Val Leu Lys Pro Tyr Val 165 170 175 Asp Ser Gly Lys Ile Lys Val Val Gly Asp Gln Trp Val Asp Gly Trp 180 185 190 Leu Pro Glu Asn Ala Leu Lys Ile Met Glu Asn Ala Leu Thr Ala Asn 195 200 205 Asn Asn Lys Ile Asp Ala Val Val Ala Ser Asn Asp Ala Thr Ala Gly 210 215 220 Gly Ala Ile Gln Ala Leu Ser Ala Gln Gly Leu Ser Gly Lys Val Ala 225 230 235 240 Ile Ser Gly Gln Asp Ala Asp Leu Ala Gly Ile Lys Arg Ile Ala Ala 245 250 255 Gly Thr Gln Thr Met Thr Val Tyr Lys Pro Ile Thr Leu Leu Ala Asn 260 265 270 Thr Ala Ala Glu Ile Ala Val Glu Leu Gly Asn Gly Gln Glu Pro Lys 275 280 285 Ala Asp Thr Thr Leu Asn Asn Gly Leu Lys Asp Val Pro Ser Arg Leu 290 295 300 Leu Thr Pro Ile Asp Val Asn Lys Asn Asn Ile Lys Asp Thr Val Ile 305 310 315 320 Lys Asp Gly Phe His Lys Glu Ser Glu Leu 325 330 <210> 30 <211> 513 <212> PRT <213> Escherichia coli <400> 30 Met Pro Tyr Leu Leu Glu Met Lys Asn Ile Thr Lys Thr Phe Gly Ser 1 5 10 15 Val Lys Ala Ile Asp Asn Val Cys Leu Arg Leu Asn Ala Gly Glu Ile 20 25 30 Val Ser Leu Cys Gly Glu Asn Gly Ser Gly Lys Ser Thr Leu Met Lys 35 40 45 Val Leu Cys Gly Ile Tyr Pro His Gly Ser Tyr Glu Gly Glu Ile Ile 50 55 60 Phe Ala Gly Glu Glu Ile Gln Ala Ser His Ile Arg Asp Thr Glu Arg 65 70 75 80 Lys Gly Ile Ala Ile Ile His Gln Glu Leu Ala Leu Val Lys Glu Leu 85 90 95 Thr Val Leu Glu Asn Ile Phe Leu Gly Asn Glu Ile Thr His Asn Gly 100 105 110 Ile Met Asp Tyr Asp Leu Met Thr Leu Arg Cys Gln Lys Leu Leu Ala 115 120 125 Gln Val Ser Leu Ser Ile Ser Pro Asp Thr Arg Val Gly Asp Leu Gly 130 135 140 Leu Gly Gln Gln Gln Leu Val Glu Ile Ala Lys Ala Leu Asn Lys Gln 145 150 155 160 Val Arg Leu Leu Ile Leu Asp Glu Pro Thr Ala Ser Leu Thr Glu Gln 165 170 175 Glu Thr Ser Ile Leu Leu Asp Ile Ile Arg Asp Leu Gln Gln His Gly 180 185 190 Ile Ala Cys Ile Tyr Ile Ser His Lys Leu Asn Glu Val Lys Ala Ile 195 200 205 Ser Asp Thr Ile Cys Val Ile Arg Asp Gly Gln His Ile Gly Thr Arg 210 215 220 Asp Ala Ala Gly Met Ser Glu Asp Asp Ile Ile Thr Met Met Val Gly 225 230 235 240 Arg Glu Leu Thr Ala Leu Tyr Pro Asn Glu Pro His Thr Thr Gly Asp 245 250 255 Glu Ile Leu Arg Ile Glu His Leu Thr Ala Trp His Pro Val Asn Arg 260 265 270 His Ile Lys Arg Val Asn Asp Val Ser Phe Ser Leu Lys Arg Gly Glu 275 280 285 Ile Leu Gly Ile Ala Gly Leu Val Gly Ala Gly Arg Thr Glu Thr Ile 290 295 300 Gln Cys Leu Phe Gly Val Trp Pro Gly Gln Trp Glu Gly Lys Ile Tyr 305 310 315 320 Ile Asp Gly Lys Gln Val Asp Ile Arg Asn Cys Gln Gln Ala Ile Ala 325 330 335 Gln Gly Ile Ala Met Val Pro Glu Asp Arg Lys Arg Asp Gly Ile Val 340 345 350 Pro Val Met Ala Val Gly Lys Asn Ile Thr Leu Ala Ala Leu Asn Lys 355 360 365 Phe Thr Gly Gly Ile Ser Gln Leu Asp Asp Ala Ala Glu Gln Lys Cys 370 375 380 Ile Leu Glu Ser Ile Gln Gln Leu Lys Val Lys Thr Ser Ser Pro Asp 385 390 395 400 Leu Ala Ile Gly Arg Leu Ser Gly Gly Asn Gln Gln Lys Ala Ile Leu 405 410 415 Ala Arg Cys Leu Leu Leu Asn Pro Arg Ile Leu Ile Leu Asp Glu Pro 420 425 430 Thr Arg Gly Ile Asp Ile Gly Ala Lys Tyr Glu Ile Tyr Lys Leu Ile 435 440 445 Asn Gln Leu Val Gln Gln Gly Ile Ala Val Ile Val Ile Ser Ser Glu 450 455 460 Leu Pro Glu Val Leu Gly Leu Ser Asp Arg Val Leu Val Met His Glu 465 470 475 480 Gly Lys Leu Lys Ala Asn Leu Ile Asn His Asn Leu Thr Gln Glu Gln 485 490 495 Val Met Glu Ala Ala Leu Arg Ser Glu His His Val Glu Lys Gln Ser 500 505 510 Val <210> 31 <211> 330 <212> PRT <213> Escherichia coli <400> 31 Met Lys Ile Lys Asn Ile Leu Leu Thr Leu Cys Thr Ser Leu Leu Leu 1 5 10 15 Thr Asn Val Ala Ala His Ala Lys Glu Val Lys Ile Gly Met Ala Ile 20 25 30 Asp Asp Leu Arg Leu Glu Arg Trp Gln Lys Asp Arg Asp Ile Phe Val 35 40 45 Lys Lys Ala Glu Ser Leu Gly Ala Lys Val Phe Val Gln Ser Ala Asn 50 55 60 Gly Asn Glu Glu Thr Gln Met Ser Gln Ile Glu Asn Met Ile Asn Arg 65 70 75 80 Gly Val Asp Val Leu Val Ile Ile Pro Tyr Asn Gly Gln Val Leu Ser 85 90 95 Asn Val Val Lys Glu Ala Lys Gln Glu Gly Ile Lys Val Leu Ala Tyr 100 105 110 Asp Arg Met Ile Asn Asp Ala Asp Ile Asp Phe Tyr Ile Ser Phe Asp 115 120 125 Asn Glu Lys Val Gly Glu Leu Gln Ala Lys Ala Leu Val Asp Ile Val 130 135 140 Pro Gln Gly Asn Tyr Phe Leu Met Gly Gly Ser Pro Val Asp Asn Asn 145 150 155 160 Ala Lys Leu Phe Arg Ala Gly Gln Met Lys Val Leu Lys Pro Tyr Val 165 170 175 Asp Ser Gly Lys Ile Lys Val Val Gly Asp Gln Trp Val Asp Gly Trp 180 185 190 Leu Pro Glu Asn Ala Leu Lys Ile Met Glu Asn Ala Leu Thr Ala Asn 195 200 205 Asn Asn Lys Ile Asp Ala Val Val Ala Ser Asn Asp Ala Thr Ala Gly 210 215 220 Gly Ala Ile Gln Ala Leu Ser Ala Gln Gly Leu Ser Gly Lys Val Ala 225 230 235 240 Ile Ser Gly Gln Asp Ala Asp Leu Ala Gly Ile Lys Arg Ile Ala Ala 245 250 255 Gly Thr Gln Thr Met Thr Val Tyr Lys Pro Ile Thr Leu Leu Ala Asn 260 265 270 Thr Ala Ala Glu Ile Ala Val Glu Leu Gly Asn Gly Gln Glu Pro Lys 275 280 285 Ala Asp Thr Thr Leu Asn Asn Gly Leu Lys Asp Val Pro Ser Arg Leu 290 295 300 Leu Thr Pro Ile Asp Val Asn Lys Asn Asn Ile Lys Asp Thr Val Ile 305 310 315 320 Lys Asp Gly Phe His Lys Glu Ser Glu Leu 325 330 <210> 32 <211> 879 <212> DNA <213> Escherichia coli <400> 32 atggctgaag cgcaaaatga tcccctgctg ccgggatact cgtttaacgc ccatctggtg 60 gcgggtttaa cgccgattga ggccaacggt tatctcgatt tttttatcga ccgaccgctg 120 ggaatgaaag gttatattct caatctcacc attcgcggtc agggggtggt gaaaaatcag 180 ggacgagaat ttgtctgccg accgggtgat attttgctgt tcccgccagg agagattcat 240 cactacggtc gtcatccgga ggctcgcgaa tggtatcacc agtgggttta ctttcgtccg 300 cgcgcctact ggcatgaatg gcttaactgg ccgtcaatat ttgccaatac gggtttcttt 360 cgcccggatg aagcgcacca gccgcatttc agcgacctgt ttgggcaaat cattaacgcc 420 gggcaagggg aagggcgcta ttcggagctg ctggcgataa atctgcttga gcaattgtta 480 ctgcggcgca tggaagcgat taacgagtcg ctccatccac cgatggataa tcgggtacgc 540 gaggcttgtc agtacatcag cgatcacctg gcagacagca attttgatat cgccagcgtc 600 gcacagcatg tttgcttgtc gccgtcgcgt ctgtcacatc ttttccgcca gcagttaggg 660 attagcgtct taagctggcg cgaggaccaa cgcattagtc aggcgaagct gcttttgagc 720 actacccgga tgcctatcgc caccgtcggt cgcaatgttg gttttgacga tcaacctat 780 ttctcgcgag tatttaaaaa atgcaccggg gccagcccga gcgagtttcg tgccggttgt 840 gaagaaaaag tgaatgatgt agccgtcaag ttgtcataa 879 <210> 33 <211> 292 <212> PRT <213> Escherichia coli <400> 33 Put Ala Glu Ala Gln Donkey Asp Pro Leu Leu Pro Gly Tyr Ser Phe Donkey 1 5 10 15 Ala His Leu Val Ala Gly Leu Thr Pro Ile Glu Ala Asn Gly Tyr Leu 20 25 30 Asp Phe Phe Ile Asp Arg Pro Leu Gly Met Lys Gly Tyr Ile Leu Asn 35 40 45 Leu Thr Ile Arg Gly Gln Gly Val Val Lys Asn Gln Gly Arg Glu Phe 50 55 60 Val Cys Arg Pro Gly Asp Ile Leu Leu Phe Pro Pro Gly Glu Ile His 65 70 75 80 His Tyr Gly Arg His Pro Glu Ala Arg Glu Trp Tyr His Gln Trp Val 85 90 95 Tyr Phe Arg Pro Arg Ala Tyr Trp His Glu Trp Leu Asn Trp Pro Ser 100 105 110 Ile Phe Ala Asn Thr Gly Phe Phe Arg Pro Asp Glu Ala His Gln Pro 115 120 125 His Phe Ser Asp Leu Phe Gly Gln Ile Ile Asn Ala Gly Gln Gly Glu 130 135 140 Gly Arg Tyr Ser Glu Leu Leu Ala Ile Asn Leu Leu Glu Gln Leu Leu 145 150 155 160 Leu Arg Arg Met Glu Ala Ile Asn Glu Ser Leu His Pro Pro Met Asp 165 170 175 Asn Arg Val Arg Glu Ala Cys Gln Tyr Ile Ser Asp His Leu Ala Asp 180 185 190 Ser Asn Phe Asp Ile Ala Ser Val Ala Gln His Val Cys Leu Ser Pro 195 200 205 Ser Arg Leu Ser His Leu Phe Arg Gln Gln Leu Gly Ile Ser Val Leu 210 215 220 Ser Trp Arg Glu Asp Gln Arg Ile Ser Gln Ala Lys Leu Leu Leu Ser 225 230 235 240 Thr Thr Arg Met Pro Ile Ala Thr Val Gly Arg Asn Val Gly Phe Asp 245 250 255 Asp Gln Leu Tyr Phe Ser Arg Val Phe Lys Lys Cys Thr Gly Ala Ser 260 265 270 Pro Ser Glu Phe Arg Ala Gly Cys Glu Glu Lys Val Asn Asp Val Ala 275 280 285 Val Lys Leu Ser 290 <210> 34 <211> 1506 <212> DNA <213> Escherichia coli <400> 34 atggaagcat tacttcagct taaaggcatc gataaagcct tcccgggcgt aaaagccctc 60 tcgggcgcag cgttaaatgt ctatccgggc cgcgtgatgg cgctggtggg cgaaaacggc 120 gcgggtaaat ccaccatgat gaaagtgctt actggcatct atactcgcga tgccggtacg 180 cttttatggc tggggaaaga aacgacattt accgggccaa aatcttccca ggaagccggg 240 attgggatta tccatcagga actgaacctg atcccgcagt tgaccattgc cgaaaacatt 300 ttcctcggtc gtgagtttgt taatcgcttt ggcaaaattg actggaaaac catgtatgcc 360 gaagcggata aattgctggc taaacttaac ctgcgcttta aaagcgacaa gctggtgggc 420 gatctttcca tcggtgacca gcaaatggtt gaaatcgcca aagtgctgag ctttgagtcg 480 aaagtcatca ttatggatga accgaccgat gcgctgaccg ataccgaaac cgaatccctg 540 ttccgcgtca tccgcgagct gaaatcgcaa ggccgcggta ttgtctatat ctcccaccgc 600 atgaaagaaa tcttcgagat ttgcgatgac gttaccgttt ttcgtgatgg gcaatttatt 660 gctgagcgcg aagtggcatc actgaccgaa gattcgctga ttgagatgat ggtgggtcgc 720 aagctggaag atcaatatcc gcacctggac aaagcgccgg gagatatccg cctgaaagtc 780 gataatctct gcggacctgg cgttaacgat gtctctttta ctttacgcaa aggcgaaatt 840 cttggcgtct ctggtttgat gggcgcaggt cgtaccgaac tgatgaaagt gctctacggc 900 gcactaccgc gcaccagcgg ttacgtcacc ctggatgggc atgaagtcgt tacccgttca 960 ccgcaggatg gcctggcaaa cggcattgtg tatatctccg aagaccgtaa acgtgacggt 1020 ttagtgttgg gcatgtcagt aaaagagaac atgtcgctga cagcgctgcg ctacttcagc 1080 cgcgctggcg gcagtttgaa gcatgccgat gaacagcagg ctgtgagtga tttcattcgt 1140 ctgtttaatg tgaaaactcc gtcgatggaa caggcaattg gtctgctttc cggtggcaat 1200 cagcaaaaag tggcgattgc ccgcggtctg atgacacgcc ccaaagtgtt gatccttgat 1260 gaacctaccc gtggcgtaga tgtcggcgcg aaaaaagaga tctatcaact gattaaccag 1320 ttcaaagccg atggcttgag catcattctg gtgtcatcgg agatgccaga agtattaggc 1380 atgagcgatc gcatcatcgt catgcatgaa gggcatctca gcggggaatt tactcgtgag 1440 caggccaccc aggaagtgtt aatggctgcc gctgtgggca agcttaatcg cgtgaatcag 1500 gagtaa 1506 <210> 35 <211> 891 <212> DNA <213> Escherichia coli <400> 35 atgaacatga aaaaactggc taccctggtt tccgctgttg cgctaagcgc caccgtcagt 60 gcgaatgcga tggcaaaaga caccatcgcg ctggtggtct ccacgcttaa caacccgttc 120 tttgtatcgc tgaaagatgg cgcgcagaaa gaggcggata aacttggcta taacctggtg 180 gtgctggact cccagaacaa cccggcgaaa gagctggcga acgtgcagga cttaaccgtt 240 cgcggcacaa aaattctgct gattaacccg accgactccg acgcagtggg taatgctgtg 300 aagatggcta accaggcgaa catcccggtt atcactcttg accgccaggc aacgaaaggt 360 gaagtggtga gccacattgc ttctgataac gtactgggcg gcaaaatcgc tggtgattac 420 atcgcgaaga aagcgggtga aggtgccaaa gttatcgagc tgcaaggcat tgctggtaca 480 tccgcagccc gtgaacgtgg cgaaggcttc cagcaggccg ttgctgctca caagtttaat 540 gttcttgcca gccagccagc agattttgat cgcattaaag gtttgaacgt aatgcagaac 600 ctgttgaccg ctcatccgga tgttcaggct gtattcgcgc agaatgatga aatggcgctg 660 ggcgcgctgc gcgcactgca aactgccggt aaatcggatg tgatggtcgt cggatttgac 720 ggtacaccgg atggcgaaaa agcggtgaat gatggcaaac tagcagcgac tatcgctcag 780 ctacccgatc agattggcgc gaaaggcgtc gaaaccgcag ataaagtgct gaaaggcgag 840 aaagttcagg ctaagtatcc ggttgatctg aaactggttg ttaagcagta g 891 <210> 36 <211> 966 <212> DNA <213> Escherichia coli <400> 36 atgacaaccc agactgtctc tggtcgccgt tatttcacga aagcgtggct gatggagcag 60 aaatcgctta tcgctctgct ggtgctgatc gcgattgtct cgacgttaag cccgaacttt 120 ttcaccatca ataacttatt caatattctc cagcaaacct cagtgaacgc cattatggcg 180 gtcgggatga cgctggtgat cctgacgtcg ggcatcgact tatcggtagg ttctctgttg 240 gcgctgaccg gcgcagttgc tgcatctatc gtcggcattg aagtcaatgc gctggtggct 300 gtcgctgctg ctctcgcgtt aggtgccgca attggtgcgg taaccggggt gattgtagcg 360 aaaggtcgcg tccaggcgtt tatcgctacg ctggttatga tgcttttact gcgcggcgtg 420 accatggttt ataccaacgg tagcccagtg aataccggct ttactgagaa cgccgatctg 480 tttggctggt ttggtattgg tcgtccgctg ggcgtaccga cgccagtctg gatcatgggg 540 attgtcttcc tcgcggcctg gtacatgctg catcacacgc gtctggggcg ttacatctac 600 gcgctgggcg gcaacgaagc ggcaacgcgt ctttctggta tcaacgtcaa taaaatcaaa 660 atcatcgtct attctctttg tggtctgctg gcatcgctgg ccgggatcat tgaagtggcg 720 cgtctctcct ccgcacaacc cacggcgggg actggctatg agctggatgc tattgctgcg 780 gtggttctgg gcggtacgag tctggcgggc ggaaaaggtc gcattgttgg gacgttgatc 840 ggcgcattaa ttcttggctt ccttaataat ggattgaatt tgttaggtgt ttcctcctat 900 taccagatga tcgtcaaagc ggtggtgatt ttgctggcgg tgctggtaga caacaaaaag 960 cagtaa 966 <210> 37 <211> 501 <212> PRT <213> Escherichia coli <400> 37 Met Glu Ala Leu Leu Gln Leu Lys Gly Ile Asp Lys Ala Phe Pro Gly 1 5 10 15 Val Lys Ala Leu Ser Gly Ala Ala Leu Asn Val Tyr Pro Gly Arg Val 20 25 30 Met Ala Leu Val Gly Glu Asn Gly Ala Gly Lys Ser Thr Met Met Lys 35 40 45 Val Leu Thr Gly Ile Tyr Thr Arg Asp Ala Gly Thr Leu Leu Trp Leu 50 55 60 Gly Lys Glu Thr Thr Phe Thr Gly Pro Lys Ser Ser Gln Glu Ala Gly 65 70 75 80 Ile Gly Ile Ile His Gln Glu Leu Asn Leu Ile Pro Gln Leu Thr Ile 85 90 95 Ala Glu Asn Ile Phe Leu Gly Arg Glu Phe Val Asn Arg Phe Gly Lys 100 105 110 Ile Asp Trp Lys Thr Met Tyr Ala Glu Ala Asp Lys Leu Leu Ala Lys 115 120 125 Leu Asn Leu Arg Phe Lys Ser Asp Lys Leu Val Gly Asp Leu Ser Ile 130 135 140 Gly Asp Gln Gln Met Val Glu Ile Ala Lys Val Leu Ser Phe Glu Ser 145 150 155 160 Lys Val Ile Ile Met Asp Glu Pro Thr Asp Ala Leu Thr Asp Thr Glu 165 170 175 Thr Glu Ser Leu Phe Arg Val Ile Arg Glu Leu Lys Ser Gln Gly Arg 180 185 190 Gly Ile Val Tyr Ile Ser His Arg Met Lys Glu Ile Phe Glu Ile Cys 195 200 205 Asp Asp Val Thr Val Phe Arg Asp Gly Gln Phe Ile Ala Glu Arg Glu 210 215 220 Val Ala Ser Leu Thr Glu Asp Ser Leu Ile Glu Met Met Val Gly Arg 225 230 235 240 Lys Leu Glu Asp Gln Tyr Pro His Leu Asp Lys Ala Pro Gly Asp Ile 245 250 255 Arg Leu Lys Val Asp Asn Leu Cys Gly Pro Gly Val Asn Asp Val Ser 260 265 270 Phe Thr Leu Arg Lys Gly Glu Ile Leu Gly Val Ser Gly Leu Met Gly 275 280 285 Ala Gly Arg Thr Glu Leu Met Lys Val Leu Tyr Gly Ala Leu Pro Arg 290 295 300 Thr Ser Gly Tyr Val Thr Leu Asp Gly His Glu Val Val Thr Arg Ser 305 310 315 320 Pro Gln Asp Gly Leu Ala Asn Gly Ile Val Tyr Ile Ser Glu Asp Arg 325 330 335 Lys Arg Asp Gly Leu Val Leu Gly Met Ser Val Lys Glu Asn Met Ser 340 345 350 Leu Thr Ala Leu Arg Tyr Phe Ser Arg Ala Gly Gly Ser Leu Lys His 355 360 365 Ala Asp Glu Gln Gln Ala Val Ser Asp Phe Ile Arg Leu Phe Asn Val 370 375 380 Lys Thr Pro Ser Met Glu Gln Ala Ile Gly Leu Leu Ser Gly Gly Asn 385 390 395 400 Gln Gln Lys Val Ala Ile Ala Arg Gly Leu Met Thr Arg Pro Lys Val 405 410 415 Leu Ile Leu Asp Glu Pro Thr Arg Gly Val Asp Val Gly Ala Lys Lys 420 425 430 Glu Ile Tyr Gln Leu Ile Asn Gln Phe Lys Ala Asp Gly Leu Ser Ile 435 440 445 Ile Leu Val Ser Ser Glu Met Pro Glu Val Leu Gly Met Ser Asp Arg 450 455 460 Ile Ile Val Met His Glu Gly His Leu Ser Gly Glu Phe Thr Arg Glu 465 470 475 480 Gln Ala Thr Gln Glu Val Leu Met Ala Ala Ala Val Gly Lys Leu Asn 485 490 495 Arg Val Asn Gln Glu 500 <210> 38 <211> 296 <212> PRT <213> Escherichia coli <400> 38 Met Asn Met Lys Lys Leu Ala Thr Leu Val Ser Ala Val Ala Leu Ser 1 5 10 15 Ala Thr Val Ser Ala Asn Ala Met Ala Lys Asp Thr Ile Ala Leu Val 20 25 30 Val Ser Thr Leu Asn Asn Pro Phe Phe Val Ser Leu Lys Asp Gly Ala 35 40 45 Gln Lys Glu Ala Asp Lys Leu Gly Tyr Asn Leu Val Val Leu Asp Ser 50 55 60 Gln Asn Asn Pro Ala Lys Glu Leu Ala Asn Val Gln Asp Leu Thr Val 65 70 75 80 Arg Gly Thr Lys Ile Leu Leu Ile Asn Pro Thr Asp Ser Asp Ala Val 85 90 95 Gly Asn Ala Val Lys Met Ala Asn Gln Ala Asn Ile Pro Val Ile Thr 100 105 110 Leu Asp Arg Gln Ala Thr Lys Gly Glu Val Val Ser His Ile Ala Ser 115 120 125 Asp Asn Val Leu Gly Gly Lys Ile Ala Gly Asp Tyr Ile Ala Lys Lys 130 135 140 Ala Gly Glu Gly Ala Lys Val Ile Glu Leu Gln Gly Ile Ala Gly Thr 145 150 155 160 Ser Ala Ala Arg Glu Arg Gly Glu Gly Phe Gln Gln Ala Val Ala Ala 165 170 175 His Lys Phe Asn Val Leu Ala Ser Gln Pro Ala Asp Phe Asp Arg Ile 180 185 190 Lys Gly Leu Asn Val Met Gln Asn Leu Leu Thr Ala His Pro Asp Val 195 200 205 Gln Ala Val Phe Ala Gln Asn Asp Glu Met Ala Leu Gly Ala Leu Arg 210 215 220 Ala Leu Gln Thr Ala Gly Lys Ser Asp Val Met Val Val Gly Phe Asp 225 230 235 240 Gly Thr Pro Asp Gly Glu Lys Ala Val Asn Asp Gly Lys Leu Ala Ala 245 250 255 Thr Ile Ala Gln Leu Pro Asp Gln Ile Gly Ala Lys Gly Val Glu Thr 260 265 270 Ala Asp Lys Val Leu Lys Gly Glu Lys Val Gln Ala Lys Tyr Pro Val 275 280 285 Asp Leu Lys Leu Val Val Lys Gln 290 295 <210> 39 <211> 321 <212> PRT <213> Escherichia coli <400> 39 Met Thr Thr Gln Thr Val Ser Gly Arg Arg Tyr Phe Thr Lys Ala Trp 1 5 10 15 Leu Met Glu Gln Lys Ser Leu Ile Ala Leu Leu Val Leu Ile Ala Ile 20 25 30 Val Ser Thr Leu Ser Pro Asn Phe Phe Thr Ile Asn Asn Leu Phe Asn 35 40 45 Ile Leu Gln Gln Thr Ser Val Asn Ala Ile Met Ala Val Gly Met Thr 50 55 60 Leu Val Ile Leu Thr Ser Gly Ile Asp Leu Ser Val Gly Ser Leu Leu 65 70 75 80 Ala Leu Thr Gly Ala Val Ala Ala Ser Ile Val Gly Ile Glu Val Asn 85 90 95 Ala Leu Val Ala Val Ala Ala Ala Leu Ala Leu Gly Ala Ala Ile Gly 100 105 110 Ala Val Thr Gly Val Ile Val Ala Lys Gly Arg Val Gln Ala Phe Ile 115 120 125 Ala Thr Leu Val Met Met Leu Leu Leu Arg Gly Val Thr Met Val Tyr 130 135 140 Thr Asn Gly Ser Pro Val Asn Thr Gly Phe Thr Glu Asn Ala Asp Leu 145 150 155 160 Phe Gly Trp Phe Gly Ile Gly Arg Pro Leu Gly Val Pro Thr Pro Val 165 170 175 Trp Ile Met Gly Ile Val Phe Leu Ala Ala Trp Tyr Met Leu His His 180 185 190 Thr Arg Leu Gly Arg Tyr Ile Tyr Ala Leu Gly Gly Asn Glu Ala Ala 195 200 205 Thr Arg Leu Ser Gly Ile Asn Val Asn Lys Ile Lys Ile Ile Val Tyr 210 215 220 Ser Leu Cys Gly Leu Leu Ala Ser Leu Ala Gly Ile Ile Glu Val Ala 225 230 235 240 Arg Leu Ser Ser Ala Gln Pro Thr Ala Gly Thr Gly Tyr Glu Leu Asp 245 250 255 Ala Ile Ala Ala Val Val Leu Gly Gly Thr Ser Leu Ala Gly Gly Lys 260 265 270 Gly Arg Ile Val Gly Thr Leu Ile Gly Ala Leu Ile Leu Gly Phe Leu 275 280 285 Asn Asn Gly Leu Asn Leu Leu Gly Val Ser Ser Tyr Tyr Gln Met Ile 290 295 300 Val Lys Ala Val Val Ile Leu Leu Ala Val Leu Val Asp Asn Lys Lys 305 310 315 320 Gln <210> 40 <211> 1533 <212> DNA <213> Escherichia coli <400> 40 atggccacgc catatatatc gatggcgggg atcggcaagt cctttggtcc ggttcacgca 60 ttaaagtcgg ttaatttaac ggtttatcct ggtgaaatac atgcattact aggagaaaat 120 ggcgcgggta aatccacgct aatgaaagtt ttatccggaa tacatgagcc gaccaaaggc 180 accattacca ttaataacat tagctataac aagctggatc ataaattagc ggcacaactc 240 ggtatcggga ttatttatca ggaactcagc gttattgatg aattaaccgt actggaaaat 300 ttatatattg gtcgtcatct gacgaaaaaa atctgtggcg tcaatattat cgactggcga 360 gaaatgcgtg tccgcgccgc catgatgtta ttacgcgtgg gcttgaaagt tgatctagat 420 gagaaagtgg cgaatttatc tatcagccac aagcagatgc tagaaattgc caaaacgctg 480 atgctcgatg ccaaagtcat catcatggat gaacccacct cctcactcac caataaagag 540 gtggactatc tgtttctgat catgaatcag ttgcgtaaag agggtacggc catcgtctat 600 atctcgcata agttggcgga aattcgccgt atttgcgacc gctatacggt gatgaaagac 660 ggcagcagcg tttgcagcgg catagtaagc gatgtgtcaa atgacgatat cgtccgtctg 720 atggtaggcc gcgaactgca aaaccgtttt aacgcgatga aggagaatgt cagcaacctt 780 gcgcacgaaa cggtttttga ggtgcggaac gtcaccagtc gtgacagaaa aaaggtccgg 840 gatatctcat ttagcgtctg ccggggagaa atattaggct ttgccggact ggtcggttcc 900 ggacgtactg aactgatgaa ttgtctgttt ggcgtggata aacgcgctgg cggagaaatc 960 cgtcttaatg gcaaagatat ctctccacgt tcacccctgg atgccgtgaa aaaagggatg 1020 gcttacatca ctgaaagccg ccgggataac ggttttttcc ccaacttttc catcgctcag 1080 aacatggcga tcagccgcag tctgaaagac ggcggctata aaggcgcgat gggcttgttt 1140 catgaagttg acgagcaacg taccgctgaa aatcaacgcg aactgctggc gctgaaatgt 1200 cattcggtaa accagaatat caccgaactc tccgggggaa atcagcagaa agtcctgatc 1260 tccaaatggc tgtgctgttg cccggaagtg attattttcg atgaacctac ccgcggcatc 1320 gacgttggcg cgaaagccga aatttacaaa gtgatgcgcc aactggcgga cgacggaaaa 1380 gtcatcctga tggtgtcatc tgaactacct gaaattatca ccgtctgcga ccgcatcgcc 1440 gtgttctgcg aaggacgact gacgcaaatc ctgacgaatc gcgatgacat gagcgaagag 1500 gagattatgg catgggcttt accacaagag taa 1533 <210> 41 <211> 936 <212> DNA <213> Escherichia coli <400> 41 atgaataaat atctgaaata tttcagcggc acactcgtgg gcttaatgtt gtcaaccagc 60 gcttttgctg ccgccgaata tgctgtcgta ttgaaaaccc tctccaaccc attttgggta 120 gatatgaaaa aaggcattga agatgaagca aaaacactgg gcgtcagcgt tgatattttt 180 gcctctcctt cagaaggcga ttttcaatct caattgcagt tatttgaaga tctcagtaat 240 aaaaattaca aaggtatcgc cttcgctcca ttatcctcag tgaatctggt catgcctgtc 300 gcccgcgcat ggaaaaaagg catttatctg gttaatctcg atgaaaaaat cgacatggat 360 aatctgaaaa aagctggcgg caatgtggaa gcttttgtca ccaccgataa cgttgctgtc 420 ggggcgaaag gcgcgtcgtt cattattgac aaattgggcg ctgaaggtgg tgaagtcgca 480 atcattgagg gtaaagccgg taacgcctcc ggtgaagcgc gtcgtaatgg tgccaccgaa 540 gccttcaaaa aagcaagcca gatcaagctt gtcgccagcc agcctgccga ctgggaccgc 600 attaaagcac tggatgtcgc cactaacgtg ttgcaacgta atccgaatat taaagcgatc 660 tattgcgcga atgacacgat ggcaatgggt gttgctcagg cagtcgcaaa cgccggaaaa 720 acgggaaaag tgctggtcgt cggtacagat ggcattccgg aagcccgcaa aatggtggaa 780 gccggacaaa tgaccgcgac ggttgcccag aacccggcgg atatcggcgc aacgggtctg 840 aagctgatgg ttgacgctga gaaatccggc aaggttatcc cgctggataa agcaccggaa 900 tttaaactgg tcgattcaat cctggtcact caataa 936 <210> 42 <211> 981 <212> DNA <213> Escherichia coli <400> 42 atgggcttta ccacaagagt aaaaagcgaa gcgagcgaga agaaaccgtt caactttgcg 60 ctgttctggg ataaatacgg cacctttttt atcctggcga tcatcgtcgc catctttggt 120 tcgctgtcac cagaatattt tctgaccacc aataatatta cccagatttt tgttcaaagc 180 tccgtgacgg tattgatcgg catgggcgag tttttcgcta tcctggtcgc tggtatcgac 240 ctctcggttg gcgcgattct ggcgctttcc ggtatggtga ccgccaaact gatgttggca 300 ggtgttgacc cgtttctcgc agcgatgatt ggcggtgtac tggttggcgg cgcactgggg 360 gcgatcaacg gctgcctggt caactggacg gggctacacc cgttcatcat cacccttggc 420 accaacgcga ttttccgtgg gatcacgctg gtgatctccg atgccaactc ggtatacggc 480 ttctcatttg acttcgtgaa cttctttgcc gccagcgtaa ttgggatacc tgtccccgtt 540 atcttctcac taattgtcgc gctcatcctt tggtttctga caacgcgtat gcggctcggg 600 cgcaacatct acgcactggg cggcaacaaa aattcggcgt tctattccgg gattgacgtg 660 aaattccaca tcctggtggt gtttatcatc tccggtgttt gtgcaggtct ggcaggcgtc 720 gtctcaactg cacgactcgg tgccgcagaa ccgcttgccg gtatgggttt tgaaacctat 780 gccattgcca gcgccatcat tggcggcacc agtttcttcg gcggcaaggg gcgcattttc 840 tctgtggtga ttggcgggtt gatcatcggc accatcaaca acggtctgaa tattttgcag 900 gtacaaacct attaccaact ggtggtgatg ggcggattaa ttatcgcggc tgtcgccctt 960 gaccgtctta tcagtaagta a 981 <210> 43 <211> 510 <212> PRT <213> Escherichia coli <400> 43 Met Ala Thr Pro Tyr Ile Ser Met Ala Gly Ile Gly Lys Ser Phe Gly 1 5 10 15 Pro Val His Ala Leu Lys Ser Val Asn Leu Thr Val Tyr Pro Gly Glu 20 25 30 Ile His Ala Leu Leu Gly Glu Asn Gly Ala Gly Lys Ser Thr Leu Met 35 40 45 Lys Val Leu Ser Gly Ile His Glu Pro Thr Lys Gly Thr Ile Thr Ile 50 55 60 Asn Asn Ile Ser Tyr Asn Lys Leu Asp His Lys Leu Ala Ala Gln Leu 65 70 75 80 Gly Ile Gly Ile Ile Tyr Gln Glu Leu Ser Val Ile Asp Glu Leu Thr 85 90 95 Val Leu Glu Asn Leu Tyr Ile Gly Arg His Leu Thr Lys Lys Ile Cys 100 105 110 Gly Val Asn Ile Ile Asp Trp Arg Glu Met Arg Val Arg Ala Ala Met 115 120 125 Met Leu Leu Arg Val Gly Leu Lys Val Asp Leu Asp Glu Lys Val Ala 130 135 140 Asn Leu Ser Ile Ser His Lys Gln Met Leu Glu Ile Ala Lys Thr Leu 145 150 155 160 Met Leu Asp Ala Lys Val Ile Ile Met Asp Glu Pro Thr Ser Ser Leu 165 170 175 Thr Asn Lys Glu Val Asp Tyr Leu Phe Leu Ile Met Asn Gln Leu Arg 180 185 190 Lys Glu Gly Thr Ala Ile Val Tyr Ile Ser His Lys Leu Ala Glu Ile 195 200 205 Arg Arg Ile Cys Asp Arg Tyr Thr Val Met Lys Asp Gly Ser Ser Val 210 215 220 Cys Ser Gly Ile Val Ser Asp Val Ser Asn Asp Asp Ile Val Arg Leu 225 230 235 240 Met Val Gly Arg Glu Leu Gln Asn Arg Phe Asn Ala Met Lys Glu Asn 245 250 255 Val Ser Asn Leu Ala His Glu Thr Val Phe Glu Val Arg Asn Val Thr 260 265 270 Ser Arg Asp Arg Lys Lys Val Arg Asp Ile Ser Phe Ser Val Cys Arg 275 280 285 Gly Glu Ile Leu Gly Phe Ala Gly Leu Val Gly Ser Gly Arg Thr Glu 290 295 300 Leu Met Asn Cys Leu Phe Gly Val Asp Lys Arg Ala Gly Gly Glu Ile 305 310 315 320 Arg Leu Asn Gly Lys Asp Ile Ser Pro Arg Ser Pro Leu Asp Ala Val 325 330 335 Lys Lys Gly Met Ala Tyr Ile Thr Glu Ser Arg Arg Asp Asn Gly Phe 340 345 350 Phe Pro Asn Phe Ser Ile Ala Gln Asn Met Ala Ile Ser Arg Ser Leu 355 360 365 Lys Asp Gly Gly Tyr Lys Gly Ala Met Gly Leu Phe His Glu Val Asp 370 375 380 Glu Gln Arg Thr Ala Glu Asn Gln Arg Glu Leu Leu Ala Leu Lys Cys 385 390 395 400 His Ser Val Asn Gln Asn Ile Thr Glu Leu Ser Gly Gly Asn Gln Gln 405 410 415 Lys Val Leu Ile Ser Lys Trp Leu Cys Cys Cys Pro Glu Val Ile Ile 420 425 430 Phe Asp Glu Pro Thr Arg Gly Ile Asp Val Gly Ala Lys Ala Glu Ile 435 440 445 Tyr Lys Val Met Arg Gln Leu Ala Asp Asp Gly Lys Val Ile Leu Met 450 455 460 Val Ser Ser Glu Leu Pro Glu Ile Ile Thr Val Cys Asp Arg Ile Ala 465 470 475 480 Val Phe Cys Glu Gly Arg Leu Thr Gln Ile Leu Thr Asn Arg Asp Asp 485 490 495 Met Ser Glu Glu Glu Ile Met Ala Trp Ala Leu Pro Gln Glu 500 505 510 <210> 44 <211> 311 <212> PRT <213> Escherichia coli <400> 44 Met Asn Lys Tyr Leu Lys Tyr Phe Ser Gly Thr Leu Val Gly Leu Met 1 5 10 15 Leu Ser Thr Ser Ala Phe Ala Ala Ala Glu Tyr Ala Val Val Leu Lys 20 25 30 Thr Leu Ser Asn Pro Phe Trp Val Asp Met Lys Lys Gly Ile Glu Asp 35 40 45 Glu Ala Lys Thr Leu Gly Val Ser Val Asp Ile Phe Ala Ser Pro Ser 50 55 60 Glu Gly Asp Phe Gln Ser Gln Leu Gln Leu Phe Glu Asp Leu Ser Asn 65 70 75 80 Lys Asn Tyr Lys Gly Ile Ala Phe Ala Pro Leu Ser Ser Val Asn Leu 85 90 95 Val Met Pro Val Ala Arg Ala Trp Lys Lys Gly Ile Tyr Leu Val Asn 100 105 110 Leu Asp Glu Lys Ile Asp Met Asp Asn Leu Lys Lys Ala Gly Gly Asn 115 120 125 Val Glu Ala Phe Val Thr Thr Asp Asn Val Ala Val Gly Ala Lys Gly 130 135 140 Ala Ser Phe Ile Ile Asp Lys Leu Gly Ala Glu Gly Gly Glu Val Ala 145 150 155 160 Ile Ile Glu Gly Lys Ala Gly Asn Ala Ser Gly Glu Ala Arg Arg Asn 165 170 175 Gly Ala Thr Glu Ala Phe Lys Lys Ala Ser Gln Ile Lys Leu Val Ala 180 185 190 Ser Gln Pro Ala Asp Trp Asp Arg Ile Lys Ala Leu Asp Val Ala Thr 195 200 205 Asn Val Leu Gln Arg Asn Pro Asn Ile Lys Ala Ile Tyr Cys Ala Asn 210 215 220 Asp Thr Met Ala Met Gly Val Ala Gln Ala Val Ala Asn Ala Gly Lys 225 230 235 240 Thr Gly Lys Val Leu Val Val Gly Thr Asp Gly Ile Pro Glu Ala Arg 245 250 255 Lys Met Val Glu Ala Gly Gln Met Thr Ala Thr Val Ala Gln Asn Pro 260 265 270 Ala Asp Ile Gly Ala Thr Gly Leu Lys Leu Met Val Asp Ala Glu Lys 275 280 285 Ser Gly Lys Val Ile Pro Leu Asp Lys Ala Pro Glu Phe Lys Leu Val 290 295 300 Asp Ser Ile Leu Val Thr Gln 305 310 <210> 45 <211> 326 <212> PRT <213> Escherichia coli <400> 45 Met Gly Phe Thr Thr Arg Val Lys Ser Glu Ala Ser Glu Lys Lys Pro 1 5 10 15 Phe Asn Phe Ala Leu Phe Trp Asp Lys Tyr Gly Thr Phe Phe Ile Leu 20 25 30 Ala Ile Ile Val Ala Ile Phe Gly Ser Leu Ser Pro Glu Tyr Phe Leu 35 40 45 Thr Thr Asn Asn Ile Thr Gln Ile Phe Val Gln Ser Ser Val Thr Val 50 55 60 Leu Ile Gly Met Gly Glu Phe Phe Ala Ile Leu Val Ala Gly Ile Asp 65 70 75 80 Leu Ser Val Gly Ala Ile Leu Ala Leu Ser Gly Met Val Thr Ala Lys 85 90 95 Leu Met Leu Ala Gly Val Asp Pro Phe Leu Ala Ala Met Ile Gly Gly 100 105 110 Val Leu Val Gly Gly Ala Leu Gly Ala Ile Asn Gly Cys Leu Val Asn 115 120 125 Trp Thr Gly Leu His Pro Phe Ile Ile Thr Leu Gly Thr Asn Ala Ile 130 135 140 Phe Arg Gly Ile Thr Leu Val Ile Ser Asp Ala Asn Ser Val Tyr Gly 145 150 155 160 Phe Ser Phe Asp Phe Val Asn Phe Phe Ala Ala Ser Val Ile Gly Ile 165 170 175 Pro Val Pro Val Ile Phe Ser Leu Ile Val Ala Leu Ile Leu Trp Phe 180 185 190 Leu Thr Thr Arg Met Arg Leu Gly Arg Asn Ile Tyr Ala Leu Gly Gly 195 200 205 Asn Lys Asn Ser Ala Phe Tyr Ser Gly Ile Asp Val Lys Phe His Ile 210 215 220 Leu Val Val Phe Ile Ile Ser Gly Val Cys Ala Gly Leu Ala Gly Val 225 230 235 240 Val Ser Thr Ala Arg Leu Gly Ala Ala Glu Pro Leu Ala Gly Met Gly 245 250 255 Phe Glu Thr Tyr Ala Ile Ala Ser Ala Ile Ile Gly Gly Thr Ser Phe 260 265 270 Phe Gly Gly Lys Gly Arg Ile Phe Ser Val Val Ile Gly Gly Leu Ile 275 280 285 Ile Gly Thr Ile Asn Asn Gly Leu Asn Ile Leu Gln Val Gln Thr Tyr 290 295 300 Tyr Gln Leu Val Val Met Gly Gly Leu Ile Ile Ala Ala Val Ala Leu 305 310 315 320 Asp Arg Leu Ile Ser Lys 325 <210> 46 <211> 1419 <212> DNA <213> Escherichia coli <400> 46 atggttacta tcaatacgga atctgcttta acgccacgtt ctttgcggga tacgcggcgt 60 atgaatatgt ttgtttcggt agctgctgcg gtcgcaggat tgttatttgg tcttgatatc 120 ggcgtaatcg ccggagcgtt gccgttcatt accgatcact ttgtgctgac cagtcgtttg 180 caggaatggg tggttagtag catgatgctc ggtgcagcaa ttggtgcgct gtttaatggt 240 tggctgtcgt tccgcctggg gcgtaaatac agcctgatgg cgggggccat cctgtttgta 300 ctcggttcta tagggtccgc ttttgcgacc agcgtagaga tgttaatcgc cgctcgtgtg 360 gtgctgggca ttgctgtcgg gatcgcgtct tacaccgctc ctctgtatct ttctgaaatg 420 gcaagtgaaa acgttcgcgg taagatgatc agtatgtacc agttgatggt cacactcggc 480 atcgtgctgg cgtttttatc cgatacagcg ttcagttata gcggtaactg gcgcgcaatg 540 ttgggggttc ttgctttacc agcagttctg ctgattattc tggtagtctt cctgccaaat 600 agcccgcgct ggctggcgga aaaggggcgt catattgagg cggaagaagt attgcgtatg 660 ctgcgcgata cgtcggaaaa agcgcgagaa gaactcaacg aaattcgtga aagcctgaag 720 ttaaaacagg gcggttgggc actgtttaag atcaaccgta acgtccgtcg tgctgtgttt 780 ctcggtatgt tgttgcaggc gatgcagcag tttaccggta tgaacatcat catgtactac 840 gcgccgcgta tcttcaaaat ggcgggcttt acgaccacag aacaacagat gattgcgact 900 ctggtcgtag ggctgacctt tatgttcgcc acctttattg cggtgtttac ggtagataaa 960 gcagggcgta aaccggctct gaaaattggt ttcagcgtga tggcgttagg cactctggtg 1020 ctgggctatt gcctgatgca gtttgataac ggtacggctt ccagtggctt gtcctggctc 1080 tctgttggca tgacgatgat gtgtattgcc ggttatgcga tgagcgccgc gccagtggtg 1140 tggatcctgt gctctgaaat tcagccgctg aaatgccgcg atttcggtat tacctgttcg 1200 accaccacga actgggtgtc gaatatgatt atcggcgcga ccttcctgac actgcttgat 1260 agcattggcg ctgccggtac gttctggctc tacactgcgc tgaacattgc gtttgtgggc 1320 attactttct ggctcattcc ggaaaccaaa aatgtcacgc tggaacatat cgaacgcaaa 1380 ctgatggcag gcgagaagtt gagaaatatc ggcgtctga 1419 <210> 47 <211> 472 <212> PRT <213> Escherichia coli <400> 47 Met Val Thr Ile Asn Thr Glu Ser Ala Leu Thr Pro Arg Ser Leu Arg 1 5 10 15 Asp Thr Arg Arg Met Asn Met Phe Val Ser Val Ala Ala Ala Val Ala 20 25 30 Gly Leu Leu Phe Gly Leu Asp Ile Gly Val Ile Ala Gly Ala Leu Pro 35 40 45 Phe Ile Thr Asp His Phe Val Leu Thr Ser Arg Leu Gln Glu Trp Val 50 55 60 Val Ser Ser Met Met Leu Gly Ala Ala Ile Gly Ala Leu Phe Asn Gly 65 70 75 80 Trp Leu Ser Phe Arg Leu Gly Arg Lys Tyr Ser Leu Met Ala Gly Ala 85 90 95 Ile Leu Phe Val Leu Gly Ser Ile Gly Ser Ala Phe Ala Thr Ser Val 100 105 110 Glu Met Leu Ile Ala Ala Arg Val Val Leu Gly Ile Ala Val Gly Ile 115 120 125 Ala Ser Tyr Thr Ala Pro Leu Tyr Leu Ser Glu Met Ala Ser Glu Asn 130 135 140 Val Arg Gly Lys Met Ile Ser Met Tyr Gln Leu Met Val Thr Leu Gly 145 150 155 160 Ile Val Leu Ala Phe Leu Ser Asp Thr Ala Phe Ser Tyr Ser Gly Asn 165 170 175 Trp Arg Ala Met Leu Gly Val Leu Ala Leu Pro Ala Val Leu Leu Ile 180 185 190 Ile Leu Val Val Phe Leu Pro Asn Ser Pro Arg Trp Leu Ala Glu Lys 195 200 205 Gly Arg His Ile Glu Ala Glu Glu Val Leu Arg Met Leu Arg Asp Thr 210 215 220 Ser Glu Lys Ala Arg Glu Glu Leu Asn Glu Ile Arg Glu Ser Leu Lys 225 230 235 240 Leu Lys Gln Gly Gly Trp Ala Leu Phe Lys Ile Asn Arg Asn Val Arg 245 250 255 Arg Ala Val Phe Leu Gly Met Leu Leu Gln Ala Met Gln Gln Phe Thr 260 265 270 Gly Met Asn Ile Ile Met Tyr Tyr Ala Pro Arg Ile Phe Lys Met Ala 275 280 285 Gly Phe Thr Thr Thr Glu Gln Gln Met Ile Ala Thr Leu Val Val Gly 290 295 300 Leu Thr Phe Met Phe Ala Thr Phe Ile Ala Val Phe Thr Val Asp Lys 305 310 315 320 Ala Gly Arg Lys Pro Ala Leu Lys Ile Gly Phe Ser Val Met Ala Leu 325 330 335 Gly Thr Leu Val Leu Gly Tyr Cys Leu Met Gln Phe Asp Asn Gly Thr 340 345 350 Ala Ser Ser Gly Leu Ser Trp Leu Ser Val Gly Met Thr Met Met Cys 355 360 365 Ile Ala Gly Tyr Ala Met Ser Ala Ala Pro Val Val Trp Ile Leu Cys 370 375 380 Ser Glu Ile Gln Pro Leu Lys Cys Arg Asp Phe Gly Ile Thr Cys Ser 385 390 395 400 Thr Thr Thr Asn Trp Val Ser Asn Met Ile Ile Gly Ala Thr Phe Leu 405 410 415 Thr Leu Leu Asp Ser Ile Gly Ala Ala Gly Thr Phe Trp Leu Tyr Thr 420 425 430 Ala Leu Asn Ile Ala Phe Val Gly Ile Thr Phe Trp Leu Ile Pro Glu 435 440 445 Thr Lys Asn Val Thr Leu Glu His Ile Glu Arg Lys Leu Met Ala Gly 450 455 460 Glu Lys Leu Arg Asn Ile Gly Val 465 470 <210> 48 <211> 1476 <212> DNA <213> Escherichia coli <400> 48 atgaataccc agtataattc cagttatata ttttcgatta ccttagtcgc tacattaggt 60 ggtttattat ttggctacga caccgccgtt atttccggta ctgttgagtc actcaatacc 120 gtctttgttg ctccacaaaa cttaagtgaa tccgctgcca actccctgtt agggttttgc 180 gtggccagcg ctctgattgg ttgcatcatc ggcggtgccc tcggtggtta ttgcagtaac 240 cgcttcggtc gtcgtgattc acttaagatt gctgctgtcc tgttttttat ttctggtgta 300 ggttctgcct ggccagaact tggttttacc tctataaacc cggacaacac tgtgcctgtt 360 tatctggcag gttatgtccc ggaatttgtt atttatcgca ttattggcgg tattggcgtt 420 ggtttagcct caatgctctc gccaatgtat attgcggaac tggctccagc tcatattcgc 480 gggaaactgg tctcttttaa ccagtttgcg attattttcg ggcaactttt agtttactgc 540 gtaaactatt ttattgcccg ttccggtgat gccagctggc tgaatactga cggctggcgt 600 tatatgtttg cctcggaatg tatccctgca ctgctgttct taatgctgct gtataccgtg 660 ccagaaagtc ctcgctggct gatgtcgcgc ggcaagcaag aacaggcgga aggtatcctg 720 cgcaaaatta tgggcaacac gcttgcaact caggcagtac aggaaattaa acactccctg 780 gatcatggcc gcaaaaccgg tggtcgtctg ctgatgtttg gcgtgggcgt gattgtaatc 840 ggcgtaatgc tctccatctt ccagcaattt gtcggcatca atgtggtgct gtactacgcg 900 ccggaagtgt tcaaaacgct gggggccagc acggatatcg cgctgttgca gaccattatt 960 gtcggagtta tcaacctcac cttcaccgtt ctggcaatta tgacggtgga taaatttggt 1020 cgtaagccac tgcaaattat cggcgcactc ggaatggcaa tcggtatgtt tagcctcggt 1080 accgcgtttt acactcaggc accgggtatt gtggcgctac tgtcgatgct gttctatgtt 1140 gccgcctttg ccatgtcctg gggtccggta tgctgggtac tgctgtcgga aatcttcccg 1200 aatgctattc gtggtaaagc gctggcaatc gcggtggcgg cccagtggct ggcgaactac 1260 ttcgtctcct ggaccttccc gatgatggac aaaaactcct ggctggtggc ccatttccac 1320 aacggtttct cctactggat ttacggttgt atgggcgttc tggcagcact gtttatgtgg 1380 aaatttgtcc cggaaaccaa aggtaaaacc cttgaggagc tggaagcgct ctgggaaccg 1440 gaaacgaaga aaacacaaca aactgctacg ctgtaa 1476 <210> 49 <211> 491 <212> PRT <213> Escherichia coli <400> 49 Met Asn Thr Gln Tyr Asn Ser Ser Tyr Ile Phe Ser Ile Thr Leu Val 1 5 10 15 Ala Thr Leu Gly Gly Leu Leu Phe Gly Tyr Asp Thr Ala Val Ile Ser 20 25 30 Gly Thr Val Glu Ser Leu Asn Thr Val Phe Val Ala Pro Gln Asn Leu 35 40 45 Ser Glu Ser Ala Ala Asn Ser Leu Leu Gly Phe Cys Val Ala Ser Ala 50 55 60 Leu Ile Gly Cys Ile Ile Gly Gly Ala Leu Gly Gly Tyr Cys Ser Asn 65 70 75 80 Arg Phe Gly Arg Arg Asp Ser Leu Lys Ile Ala Ala Val Leu Phe Phe 85 90 95 Ile Ser Gly Val Gly Ser Ala Trp Pro Glu Leu Gly Phe Thr Ser Ile 100 105 110 Asn Pro Asp Asn Thr Val Pro Val Tyr Leu Ala Gly Tyr Val Pro Glu 115 120 125 Phe Val Ile Tyr Arg Ile Ile Gly Gly Ile Gly Val Gly Leu Ala Ser 130 135 140 Met Leu Ser Pro Met Tyr Ile Ala Glu Leu Ala Pro Ala His Ile Arg 145 150 155 160 Gly Lys Leu Val Ser Phe Asn Gln Phe Ala Ile Ile Phe Gly Gln Leu 165 170 175 Leu Val Tyr Cys Val Asn Tyr Phe Ile Ala Arg Ser Gly Asp Ala Ser 180 185 190 Trp Leu Asn Thr Asp Gly Trp Arg Tyr Met Phe Ala Ser Glu Cys Ile 195 200 205 Pro Ala Leu Leu Phe Leu Met Leu Leu Tyr Thr Val Pro Glu Ser Pro 210 215 220 Arg Trp Leu Met Ser Arg Gly Lys Gln Glu Gln Ala Glu Gly Ile Leu 225 230 235 240 Arg Lys Ile Met Gly Asn Thr Leu Ala Thr Gln Ala Val Gln Glu Ile 245 250 255 Lys His Ser Leu Asp His Gly Arg Lys Thr Gly Gly Arg Leu Leu Met 260 265 270 Phe Gly Val Gly Val Ile Val Ile Gly Val Met Leu Ser Ile Phe Gln 275 280 285 Gln Phe Val Gly Ile Asn Val Val Leu Tyr Tyr Ala Pro Glu Val Phe 290 295 300 Lys Thr Leu Gly Ala Ser Thr Asp Ile Ala Leu Leu Gln Thr Ile Ile 305 310 315 320 Val Gly Val Ile Asn Leu Thr Phe Thr Val Leu Ala Ile Met Thr Val 325 330 335 Asp Lys Phe Gly Arg Lys Pro Leu Gln Ile Ile Gly Ala Leu Gly Met 340 345 350 Ala Ile Gly Met Phe Ser Leu Gly Thr Ala Phe Tyr Thr Gln Ala Pro 355 360 365 Gly Ile Val Ala Leu Leu Ser Met Leu Phe Tyr Val Ala Ala Phe Ala 370 375 380 Met Ser Trp Gly Pro Val Cys Trp Val Leu Leu Ser Glu Ile Phe Pro 385 390 395 400 Asn Ala Ile Arg Gly Lys Ala Leu Ala Ile Ala Val Ala Ala Gln Trp 405 410 415 Leu Ala Asn Tyr Phe Val Ser Trp Thr Phe Pro Met Met Asp Lys Asn 420 425 430 Ser Trp Leu Val Ala His Phe His Asn Gly Phe Ser Tyr Trp Ile Tyr 435 440 445 Gly Cys Met Gly Val Leu Ala Ala Leu Phe Met Trp Lys Phe Val Pro 450 455 460 Glu Thr Lys Gly Lys Thr Leu Glu Glu Leu Glu Ala Leu Trp Glu Pro 465 470 475 480 Glu Thr Lys Lys Thr Gln Gln Thr Ala Thr Leu 485 490 <210> 50 <211> 1095 <212> DNA <213> Homo sapiens <400> 50 atggcgcacc gttttccggc gctgacccaa gagcagaaga aggagctgag cgagattgcg 60 cagagcatcg tggcgaatgg taaaggtatt ctggcggcgg atgagagcgt tggtaccatg 120 ggcaaccgtc tgcagcgtat taaggtggag aacaccgagg aaaaccgtcg tcaattccgt 180 gaaatcctgt ttagcgttga tagcagcatc aaccagagca ttggtggcgt gatcctgttc 240 cacgaaaccc tgtaccagaa ggacagccaa ggtaaactgt ttcgtaacat tctgaaggaa 300 aaaggtattg tggttggcat caagctggat caaggtggcg cgccgctggc gggcaccaac 360 aaggaaacca ccatccaggg tctggacggc ctgagcgaac gttgcgcgca atataagaaa 420 gatggtgttg acttcggcaa gtggcgtgcg gtgctgcgta ttgcggacca gtgcccgagc 480 agcctggcga tccaagaaaa cgcgaacgcg ctggcgcgtt acgcgagcat ctgccagcaa 540 aacggtctgg tgccgattgt tgagccggaa gttatcccgg acggcgatca cgacctggag 600 cactgccagt atgtgaccga aaaggttctg gcggcggtgt acaaagcgct gaacgatcac 660 cacgtttatc tggagggtac cctgctgaaa ccgaacatgg tgaccgcggg ccatgcgtgc 720 accaagaaat acaccccgga acaggtggcg atggcgaccg tgaccgcgct gcaccgtacc 780 gttccggcgg cggtgccggg tatttgcttt ctgagcggtg gcatgagcga agaggacgcg 840 accctgaacc tgaacgcgat caacctgtgc ccgctgccga agccgtggaa actgagcttc 900 agctacggcc gtgcgctgca ggcgagcgcg ctggcggcgt ggggtggcaa ggcggcgaac 960 aaagaggcga cccaagaagc gtttatgaag cgtgcgatgg cgaactgcca ggcggcgaaa 1020 ggtcaatatg tgcataccgg cagcagcggt gcggcgagca cccagagcct gtttaccgcg 1080 tgctatacct attaa 1095 <210> 51 <211> 364 <212> PRT <213> Homo sapiens <400> 51 Met Ala His Arg Phe Pro Ala Leu Thr Gln Glu Gln Lys Lys Glu Leu 1 5 10 15 Ser Glu Ile Ala Gln Ser Ile Val Ala Asn Gly Lys Gly Ile Leu Ala 20 25 30 Ala Asp Glu Ser Val Gly Thr Met Gly Asn Arg Leu Gln Arg Ile Lys 35 40 45 Val Glu Asn Thr Glu Glu Asn Arg Arg Gln Phe Arg Glu Ile Leu Phe 50 55 60 Ser Val Asp Ser Ser Ile Asn Gln Ser Ile Gly Gly Val Ile Leu Phe 65 70 75 80 His Glu Thr Leu Tyr Gln Lys Asp Ser Gln Gly Lys Leu Phe Arg Asn 85 90 95 Ile Leu Lys Glu Lys Gly Ile Val Val Gly Ile Lys Leu Asp Gln Gly 100 105 110 Gly Ala Pro Leu Ala Gly Thr Asn Lys Glu Thr Thr Ile Gln Gly Leu 115 120 125 Asp Gly Leu Ser Glu Arg Cys Ala Gln Tyr Lys Lys Asp Gly Val Asp 130 135 140 Phe Gly Lys Trp Arg Ala Val Leu Arg Ile Ala Asp Gln Cys Pro Ser 145 150 155 160 Ser Leu Ala Ile Gln Glu Asn Ala Asn Ala Leu Ala Arg Tyr Ala Ser 165 170 175 Ile Cys Gln Gln Asn Gly Leu Val Pro Ile Val Glu Pro Glu Val Ile 180 185 190 Pro Asp Gly Asp His Asp Leu Glu His Cys Gln Tyr Val Thr Glu Lys 195 200 205 Val Leu Ala Ala Val Tyr Lys Ala Leu Asn Asp His His Val Tyr Leu 210 215 220 Glu Gly Thr Leu Leu Lys Pro Asn Met Val Thr Ala Gly His Ala Cys 225 230 235 240 Thr Lys Lys Tyr Thr Pro Glu Gln Val Ala Met Ala Thr Val Thr Ala 245 250 255 Leu His Arg Thr Val Pro Ala Ala Val Pro Gly Ile Cys Phe Leu Ser 260 265 270 Gly Gly Met Ser Glu Glu Asp Ala Thr Leu Asn Leu Asn Ala Ile Asn 275 280 285 Leu Cys Pro Leu Pro Lys Pro Trp Lys Leu Ser Phe Ser Tyr Gly Arg 290 295 300 Ala Leu Gln Ala Ser Ala Leu Ala Ala Trp Gly Gly Lys Ala Ala Asn 305 310 315 320 Lys Glu Ala Thr Gln Glu Ala Phe Met Lys Arg Ala Met Ala Asn Cys 325 330 335 Gln Ala Ala Lys Gly Gln Tyr Val His Thr Gly Ser Ser Gly Ala Ala 340 345 350 Ser Thr Gln Ser Leu Phe Thr Ala Cys Tyr Thr Tyr 355 360 <210> 52 <211> 1149 <212> DNA <213> Escherichia coli <400> 52 atggctaaca gaatgattct gaacgaaacg gcatggtttg gtcggggtgc tgttggggct 60 ttaaccgatg aggtgaaacg ccgtggttat cagaaggcgc tgatcgtcac cgataaaacg 120 ctggtgcaat gcggcgtggt ggcgaaagtg accgataaga tggatgctgc agggctggca 180 tgggcgattt acgacggcgt agtgcccaac ccaacaatta ctgtcgtcaa agaagggctc 240 ggtgtattcc agaatagcgg cgcggattac ctgatcgcta ttggtggtgg ttctccacag 300 gatacttgta aagcgattgg cattatcagc aacaacccgg agtttgccga tgtgcgtagc 360 ctggaagggc tttccccgac caataaaccc agtgtaccga ttctggcaat tcctaccaca 420 gcaggtactg cggcagaagt gaccattaac tacgtgatca ctgacgaaga gaaacggcgc 480 aagtttgttt gcgttgatcc gcatgatatc ccgcaggtgg cgtttattga cgctgacatg 540 atggatggta tgcctccagc gctgaaagct gcgacgggtg tcgatgcgct cactcatgct 600 attgaggggt atattacccg tggcgcgtgg gcgctaaccg atgcactgca cattaaagcg 660 attgaaatca ttgctggggc gctgcgagga tcggttgctg gtgataagga tgccggagaa 720 gaaatggcgc tcgggcagta tgttgcgggt atgggcttct cgaatgttgg gttagggttg 780 gtgcatggta tggcgcatcc actgggcgcg ttttataaca ctccacacgg tgttgcgaac 840 gccatcctgt taccgcatgt catgcgttat aacgctgact ttaccggtga gaagtaccgc 900 gatatcgcgc gcgttatggg cgtgaaagtg gaaggtatga gcctggaaga ggcgcgtaat 960 gccgctgttg aagcggtgtt tgctctcaac cgtgatgtcg gtattccgcc acatttgcgt 1020 gatgttggtg tacgcaagga agacattccg gcactggcgc aggcggcact ggatgatgtt 1080 tgtaccggtg gcaacccgcg tgaagcaacg cttgaggata ttgtagagct ttaccatacc 1140 gcctggtaa 1149 <210> 53 <211> 144 <212> DNA <213> Artificial Sequence <220> <223> proD promoter <400> 53 cacagctaac accacgtcgt ccctatctgc tgccctaggt ctatgagtgg ttgctggata 60 actttacggg catgcataag gctcgtataa tatattcagg gagtccacaa cggtttccct 120 ctacaaataa ttttgtttaa cttt 144 <210> 54 <211> 870 <212> DNA <213> Caulobacter crescentus <400> 54 atgacggcgc aagttacttg tgtgtgggac ctgaaagcca cattaggaga aggcccgatc 60 tggcatggtg atacattatg gtttgtcgat attaaacagc gtaaaatcca taactaccat 120 ccggccactg gggagcgctt cagcttcgac gcaccggacc aggtgacgtt ccttgcgccc 180 atcgtcggtg cgactggatt tgtcgtagga ttaaagacgg gcatccaccg tttccatcca 240 gcaacaggtt tttccctgtt gcttgaggta gaggacgccg ctttgaataa tcgtccgaac 300 gatgcgaccg tggatgcaca aggacgtttg tggtttggta cgatgcacga tggtgaagag 360 aataacagcg gctctttgta tcgtatggat ttgacgggag ttgcacgcat ggatcgtgat 420 atttgcatca caaacggacc gtgcgtttcc cctgacggta aaaccttcta tcatacggac 480 acgctggaaa agaccattta cgcgttcgat cttgcggaag acggtcttct gagcaataaa 540 cgcgtgtttg ttcaattcgc tctgggtgac gatgtctacc ctgacggctc tgtggtagat 600 agcgaaggtt atctgtggac cgcactgtgg ggagggtttg gcgcggtgcg tttcagcccg 660 cagggtgacg ctgtaacccg cattgaatta ccggctccga acgtgacgaa gccgtgcttt 720 gggggtccgg atctgaaaac gttgtatttc accaccgcgc gtaaaggatt atcagatgag 780 acccttgccc aatatccatt ggctggcggg gttttcgccg taccggtaga tgtcgctggt 840 caacctcaac acgaagtgcg tttagtttaa 870 <210> 55 <211> 289 <212> PRT <213> Caulobacter crescentus <400> 55 Met Thr Ala Gln Val Thr Cys Val Trp Asp Leu Lys Ala Thr Leu Gly 1 5 10 15 Glu Gly Pro Ile Trp His Gly Asp Thr Leu Trp Phe Val Asp Ile Lys 20 25 30 Gln Arg Lys Ile His Asn Tyr His Pro Ala Thr Gly Glu Arg Phe Ser 35 40 45 Phe Asp Ala Pro Asp Gln Val Thr Phe Leu Ala Pro Ile Val Gly Ala 50 55 60 Thr Gly Phe Val Val Gly Leu Lys Thr Gly Ile His Arg Phe His Pro 65 70 75 80 Ala Thr Gly Phe Ser Leu Leu Leu Glu Val Glu Asp Ala Ala Leu Asn 85 90 95 Asn Arg Pro Asn Asp Ala Thr Val Asp Ala Gln Gly Arg Leu Trp Phe 100 105 110 Gly Thr Met His Asp Gly Glu Glu Asn Asn Ser Gly Ser Leu Tyr Arg 115 120 125 Met Asp Leu Thr Gly Val Ala Arg Met Asp Arg Asp Ile Cys Ile Thr 130 135 140 Asn Gly Pro Cys Val Ser Pro Asp Gly Lys Thr Phe Tyr His Thr Asp 145 150 155 160 Thr Leu Glu Lys Thr Ile Tyr Ala Phe Asp Leu Ala Glu Asp Gly Leu 165 170 175 Leu Ser Asn Lys Arg Val Phe Val Gln Phe Ala Leu Gly Asp Asp Val 180 185 190 Tyr Pro Asp Gly Ser Val Val Asp Ser Glu Gly Tyr Leu Trp Thr Ala 195 200 205 Leu Trp Gly Gly Phe Gly Ala Val Arg Phe Ser Pro Gln Gly Asp Ala 210 215 220 Val Thr Arg Ile Glu Leu Pro Ala Pro Asn Val Thr Lys Pro Cys Phe 225 230 235 240 Gly Gly Pro Asp Leu Lys Thr Leu Tyr Phe Thr Thr Ala Arg Lys Gly 245 250 255 Leu Ser Asp Glu Thr Leu Ala Gln Tyr Pro Leu Ala Gly Gly Val Phe 260 265 270 Ala Val Pro Val Asp Val Ala Gly Gln Pro Gln His Glu Val Arg Leu 275 280 285 Val <210> 56 <211> 885 <212> DNA <213> Burkholderia xenovorans <400> 56 atgaaaatcc accctccggt ctgtgtctgg ccagttcacg cggagttggg ggaaggccct 60 ttgtggcagg cgggggaaaa cgcggtttat ttcgtggata ttaagggtcg tcaaattcac 120 cgtttaactg tgaccactgg ccaaacgcaa acctggaaag cgcctggaca gccagggttt 180 attgccccat tggccggtca tgggtttgtt tgtggtctgc cgggaggtct ttaccgcttt 240 gacgccgggt caggccaatt cagcaagctg aaagatgtag aagtacatct tccggggaat 300 cgcttaaacg acggcttcgt cgatgccagc gggcatttat ggtttggatc tatggatgat 360 ggagaagaac agcctagtgg aaccctgtac cgtgtcaatc atgcgggaga ggctgttgcc 420 caggatgacg gatatgtcat cacgaatggt cctgcaagtt caccggatgc ccgcacttta 480 tatcacggcg acactatgcg tcgcgttgtc tacgctttcg acctgggcga agatgggaca 540 cttagccgca agcgcatctt tgccgcgatc tcgggagatg gctatccaga tggcatggcc 600 gtggacgcgg acggattcct gtgggttgcg ctgttcggtg gatggcgtat tgaacgtttc 660 tccccggacg gaaaacgtgt cgagcaagtt ccgtttcctt gcgctaacgt gacgaaattg 720 gcattcggcg gcgacgattt acaaacggtc tatgcgtcaa ccgcctggaa gggtttgtct 780 cccgttgcgc gtcagcaaca gcccgacgca ggggggttat tcagctttcg ctccccggta 840 ccaggccagc cccaggcccg ttgcaaatgg ggcttcgctc agtaa 885 <210> 57 <211> 294 <212> PRT <213> Burkholderia xenovorans <400> 57 Met Lys Ile His Pro Pro Val Cys Val Trp Pro Val His Ala Glu Leu 1 5 10 15 Gly Glu Gly Pro Leu Trp Gln Ala Gly Glu Asn Ala Val Tyr Phe Val 20 25 30 Asp Ile Lys Gly Arg Gln Ile His Arg Leu Thr Val Thr Thr Gly Gln 35 40 45 Thr Gln Thr Trp Lys Ala Pro Gly Gln Pro Gly Phe Ile Ala Pro Leu 50 55 60 Ala Gly His Gly Phe Val Cys Gly Leu Pro Gly Gly Leu Tyr Arg Phe 65 70 75 80 Asp Ala Gly Ser Gly Gln Phe Ser Lys Leu Lys Asp Val Glu Val His 85 90 95 Leu Pro Gly Asn Arg Leu Asn Asp Gly Phe Val Asp Ala Ser Gly His 100 105 110 Leu Trp Phe Gly Ser Met Asp Asp Gly Glu Glu Gln Pro Ser Gly Thr 115 120 125 Leu Tyr Arg Val Asn His Ala Gly Glu Ala Val Ala Gln Asp Asp Gly 130 135 140 Tyr Val Ile Thr Asn Gly Pro Ala Ser Ser Pro Asp Ala Arg Thr Leu 145 150 155 160 Tyr His Gly Asp Thr Met Arg Arg Val Val Tyr Ala Phe Asp Leu Gly 165 170 175 Glu Asp Gly Thr Leu Ser Arg Lys Arg Ile Phe Ala Ala Ile Ser Gly 180 185 190 Asp Gly Tyr Pro Asp Gly Met Ala Val Asp Ala Asp Gly Phe Leu Trp 195 200 205 Val Ala Leu Phe Gly Gly Trp Arg Ile Glu Arg Phe Ser Pro Asp Gly 210 215 220 Lys Arg Val Glu Gln Val Pro Phe Pro Cys Ala Asn Val Thr Lys Leu 225 230 235 240 Ala Phe Gly Gly Asp Asp Leu Gln Thr Val Tyr Ala Ser Thr Ala Trp 245 250 255 Lys Gly Leu Ser Pro Val Ala Arg Gln Gln Gln Pro Asp Ala Gly Gly 260 265 270 Leu Phe Ser Phe Arg Ser Pro Val Pro Gly Gln Pro Gln Ala Arg Cys 275 280 285 Lys Trp Gly Phe Ala Gln 290 <210> 58 <211> 876 <212> DNA <213> Haloferax volcanii <400> 58 atgaccgtca cgagagtggt cgacacgtcg tgccgactcg gcgagggacc ggtctggcac 60 cccgacgaga agcggctgta ctgggtagac atcgaatccg ggcgactcca ccgctacgac 120 cccgagaccg gagcgcacga ctgtccggtc gaaacgtcgg tcatcgccgg cgtgacgata 180 cagcgcgacg ggtcgctcct agcgttcatg gaccgcggtc gcgtcgggcg cgtcgtcgac 240 ggcgaccgcc gggaaagcgc gcgaatcgtc gactcaccga cccggttcaa cgacgtcatc 300 gccgaccccg ccgggcgggt gttctgcggg acgatgccgt cagatacggc gggcgggcgg 360 ctgttccgcc tcgacaccga cggaacggtg acgacggtcg aaaccggcgt cggaatcccc 420 aacgggatgg gattcacccg cgaccgcgag cgattctact tcacggaaac cgaggcgcgc 480 accgtctatc ggtacgccta cgacgaagag acgggagccg tctcggcccg cgagcggttc 540 gtcgaatcgc cggagacgcc ggggctaccg gacgggatga cggtcgattc ggcggggcac 600 atctggtcgg cccgctggga gggcggctgt gtggtcgagt acgacgccga cggcaccgaa 660 ctcggccggt tcgacgtgcc gacggagaag gtgacgagcg tcgccttcgg cgggccggac 720 ctcgattcgc tctacgtgac gaccgccggc ggcgacgggg acgggagcgc tggcgagggc 780 gacgagagca ccggcgacgc tgcgggcgcg ctgttccgcc tcgacgtggc ggcgaccggg 840 cggcccgagt tccgctccga cgttcggttg gggtga 876 <210> 59 <211> 291 <212> PRT <213> Haloferax volcanii <400> 59 Met Thr Val Thr Arg Val Val Asp Thr Ser Cys Arg Leu Gly Glu Gly 1 5 10 15 Pro Val Trp His Pro Asp Glu Lys Arg Leu Tyr Trp Val Asp Ile Glu 20 25 30 Ser Gly Arg Leu His Arg Tyr Asp Pro Glu Thr Gly Ala His Asp Cys 35 40 45 Pro Val Glu Thr Ser Val Ile Ala Gly Val Thr Ile Gln Arg Asp Gly 50 55 60 Ser Leu Leu Ala Phe Met Asp Arg Gly Arg Val Gly Arg Val Val Asp 65 70 75 80 Gly Asp Arg Arg Glu Ser Ala Arg Ile Val Asp Ser Pro Thr Arg Phe 85 90 95 Asn Asp Val Ile Ala Asp Pro Ala Gly Arg Val Phe Cys Gly Thr Met 100 105 110 Pro Ser Asp Thr Ala Gly Gly Arg Leu Phe Arg Leu Asp Thr Asp Gly 115 120 125 Thr Val Thr Thr Val Glu Thr Gly Val Gly Ile Pro Asn Gly Met Gly 130 135 140 Phe Thr Arg Asp Arg Glu Arg Phe Tyr Phe Thr Glu Thr Glu Ala Arg 145 150 155 160 Thr Val Tyr Arg Tyr Ala Tyr Asp Glu Glu Thr Gly Ala Val Ser Ala 165 170 175 Arg Glu Arg Phe Val Glu Ser Pro Glu Thr Pro Gly Leu Pro Asp Gly 180 185 190 Met Thr Val Asp Ser Ala Gly His Ile Trp Ser Ala Arg Trp Glu Gly 195 200 205 Gly Cys Val Val Glu Tyr Asp Ala Asp Gly Thr Glu Leu Gly Arg Phe 210 215 220 Asp Val Pro Thr Glu Lys Val Thr Ser Val Ala Phe Gly Gly Pro Asp 225 230 235 240 Leu Asp Ser Leu Tyr Val Thr Thr Ala Gly Gly Asp Gly Asp Gly Ser 245 250 255 Ala Gly Glu Gly Asp Glu Ser Thr Gly Asp Ala Ala Gly Ala Leu Phe 260 265 270 Arg Leu Asp Val Ala Ala Thr Gly Arg Pro Glu Phe Arg Ser Asp Val 275 280 285 Arg Leu Gly 290 <210> 60 <211> 1776 <212> DNA <213> Caulobacter crescentus <400> 60 atgtcgaatc gcactccacg tcgttttcgc tcacgtgact ggttcgataa tccagaccat 60 atcgacatga cggccttgta tcttgaacgt tttatgaatt atgggattac ccccgaggag 120 ttacgctcgg ggaagccgat tattgggatc gcacagactg gctctgatat ttcaccctgt 180 aaccgtatcc acctggacct ggtgcagcgt gtgcgtgacg ggatccgtga cgcaggaggc 240 attcctatgg aattcccagt tcacccaatc ttcgagaatt gccgtcgccc tactgctgcg 300 cttgatcgta acctttctta cttgggtctt gtcgagactc ttcatgggta ccctatcgat 360 gcggtcgttt taacaacggg gtgcgataaa acgaccccag caggcattat ggctgccacc 420 acagtcaaca tcccggctat tgtgttgtcc gggggcccaa tgcttgatgg gtggcatgaa 480 acagtcaaca tcccggctat tgtgttgtcc gggggcccaa tgcttgatgg gtggcatgaa 480 aatgaattgg tgggttcagg cacagtcatt tggcgtagcc gtcgcaaatt agccgccggc 540 aatgaattgg tgggttcagg cacagtcatt tggcgtagcc gtcgcaaatt agccgccggc 540 gaaatcacag aagaagaatt tattgaccgt gccgctagta gcgcgccaag cgcgggacac 600 gaaatcacag aagaagaatt tattgaccgt gccgctagta gcgcgccaag cgcgggacac 600 tgtaacacga tgggaacagc gtccacgatg aatgccgtgg cggaggcttt gggtctttcc 660 tgtaacacga tgggaacagc gtccacgatg aatgccgtgg cggaggcttt gggtctttcc 660 ctgacaggtt gtgcagccat tcccgcaccc tatcgcgagc gtggtcagat ggcgtacaaa 720 ctgacaggtt gtgcagccat tcccgcaccc tatcgcgagc gtggtcagat ggcgtacaaa 720 acgggacaac gcattgtgga cttggcgtat gacgacgtta agccacttga catccttacc 780 acgggacaac gcattgtgga cttggcgtat gacgacgtta agccacttga catccttacc 780 aagcaggcgt ttgaaaacgc cattgctttg gtcgccgcag ccggaggttc gaccaatgca 840 aagcaggcgt ttgaaaacgc cattgctttg gtcgccgcag ccggaggttc gaccaatgca 840 cagccacaca ttgtagcaat ggcccgccat gccggagtgg aaatcaccgc cgatgattgg 900 cagccacaca ttgtagcaat ggcccgccat gccggagtgg aaatcaccgc cgatgattgg 900 cgtgctgcat atgacattcc gcttatcgtg aatatgcaac ctgctggtaa atacttggga 960 cgtgctgcat atgacattcc gcttatcgtg aatatgcaac ctgctggtaa atacttggga 960 gaacgtttcc atcgcgctgg tggcgcacca gcggtactgt gggagttact tcagcaaggt 1020 gaacgtttcc atcgcgctgg tggcgcacca gcggtactgt gggagttact tcagcaaggt 1020 cgcctgcacg gcgacgtact tactgtcacg ggaaagacaa tgtcagagaa cttacagggg 1080 cgcctgcacg gcgacgtact tactgtcacg ggaaagacaa tgtcagagaa cttacagggg 1080 cgcgaaacct cagatcgtga agtaatcttt ccttatcatg aaccgcttgc cgagaaagcc 1140 cgcgaaacct cagatcgtga agtaatcttt ccttatcatg aaccgcttgc cgagaaagcc 1140 ggcttcttgg ttcttaaggg caatttgttt gatttcgcga tcatgaaaag ctcggtgatt 1200 ggtgaggaat ttcgtaagcg ctacttatca cagccagggc aagagggtgt tttcgaagca 1260 cgtgctattg tttttgacgg ctcggacgac taccataaac gcattaacga tcctgcatta 1320 gaaatcgatg aacgttgtat cttagtgatc cgtggtgcag gaccaatcgg gtggccgggc 1380 tcagcggagg tcgtcaacat gcaaccacca gaccacttac tgaaaaaagg gatcatgtcg 1440 ttaccaacac tgggggacgg tcgtcaaagt ggtactgccg attccccctc tattttgaat 1500 gcgtccccag agtccgctat tggaggcgga cttagctggc ttcgtacagg tgacacaatc 1560 cgtatcgatc tgaacacggg acgttgcgac gcgttagtgg acgaggccac aattgcggct 1620 cgtaagcagg acggaattcc cgccgtcccg gcaaccatga caccgtggca agaaatttat 1680 cgcgcgcatg cgtcacagct tgacaccggt ggggtattgg agtttgctgt gaaatatcag 1740 gatctggcgg cgaagctgcc acgccataat cactaa 1776 <210> 61 <211> 591 <212> PRT <213> Caulobacter crescentus <400> 61 Met Ser Asn Arg Thr Pro Arg Arg Phe Arg Ser Arg Asp Trp Phe Asp 1 5 10 15 Asn Pro Asp His Ile Asp Met Thr Ala Leu Tyr Leu Glu Arg Phe Met 20 25 30 Asn Tyr Gly Ile Thr Pro Glu Glu Leu Arg Ser Gly Lys Pro Ile Ile 35 40 45 Gly Ile Ala Gln Thr Gly Ser Asp Ile Ser Pro Cys Asn Arg Ile His 50 55 60 Leu Asp Leu Val Gln Arg Val Arg Asp Gly Ile Arg Asp Ala Gly Gly 65 70 75 80 Ile Pro Met Glu Phe Pro Val His Pro Ile Phe Glu Asn Cys Arg Arg 85 90 95 Pro Thr Ala Ala Leu Asp Arg Asn Leu Ser Tyr Leu Gly Leu Val Glu 100 105 110 Thr Leu His Gly Tyr Pro Ile Asp Ala Val Val Leu Thr Thr Gly Cys 115 120 125 Asp Lys Thr Thr Pro Ala Gly Ile Met Ala Ala Thr Thr Val Asn Ile 130 135 140 Pro Ala Ile Val Leu Ser Gly Gly Pro Met Leu Asp Gly Trp His Glu 145 150 155 160 Asn Glu Leu Val Gly Ser Gly Thr Val Ile Trp Arg Ser Arg Arg Lys 165 170 175 Leu Ala Ala Gly Glu Ile Thr Glu Glu Glu Phe Ile Asp Arg Ala Ala 180 185 190 Ser Ser Ala Pro Ser Ala Gly His Cys Asn Thr Met Gly Thr Ala Ser 195 200 205 Thr Met Asn Ala Val Ala Glu Ala Leu Gly Leu Ser Leu Thr Gly Cys 210 215 220 Ala Ala Ile Pro Ala Pro Tyr Arg Glu Arg Gly Gln Met Ala Tyr Lys 225 230 235 240 Thr Gly Gln Arg Ile Val Asp Leu Ala Tyr Asp Asp Val Lys Pro Leu 245 250 255 Asp Ile Leu Thr Lys Gln Ala Phe Glu Asn Ala Ile Ala Leu Val Ala 260 265 270 Ala Ala Gly Gly Ser Thr Asn Ala Gln Pro His Ile Val Ala Met Ala 275 280 285 Arg His Ala Gly Val Glu Ile Thr Ala Asp Asp Trp Arg Ala Ala Tyr 290 295 300 Asp Ile Pro Leu Ile Val Asn Met Gln Pro Ala Gly Lys Tyr Leu Gly 305 310 315 320 Glu Arg Phe His Arg Ala Gly Gly Ala Pro Ala Val Leu Trp Glu Leu 325 330 335 Leu Gln Gln Gly Arg Leu His Gly Asp Val Leu Thr Val Thr Gly Lys 340 345 350 Thr Met Ser Glu Asn Leu Gln Gly Arg Glu Thr Ser Asp Arg Glu Val 355 360 365 Ile Phe Pro Tyr His Glu Pro Leu Ala Glu Lys Ala Gly Phe Leu Val 370 375 380 Leu Lys Gly Asn Leu Phe Asp Phe Ala Ile Met Lys Ser Ser Val Ile 385 390 395 400 Gly Glu Glu Phe Arg Lys Arg Tyr Leu Ser Gln Pro Gly Gln Glu Gly 405 410 415 Val Phe Glu Ala Arg Ala Ile Val Phe Asp Gly Ser Asp Asp Tyr His 420 425 430 Lys Arg Ile Asn Asp Pro Ala Leu Glu Ile Asp Glu Arg Cys Ile Leu 435 440 445 Val Ile Arg Gly Ala Gly Pro Ile Gly Trp Pro Gly Ser Ala Glu Val 450 455 460 Val Asn Met Gln Pro Pro Asp His Leu Leu Lys Lys Gly Ile Met Ser 465 470 475 480 Leu Pro Thr Leu Gly Asp Gly Arg Gln Ser Gly Thr Ala Asp Ser Pro 485 490 495 Ser Ile Leu Asn Ala Ser Pro Glu Ser Ala Ile Gly Gly Gly Leu Ser 500 505 510 Trp Leu Arg Thr Gly Asp Thr Ile Arg Ile Asp Leu Asn Thr Gly Arg 515 520 525 Cys Asp Ala Leu Val Asp Glu Ala Thr Ile Ala Ala Arg Lys Gln Asp 530 535 540 Gly Ile Pro Ala Val Pro Ala Thr Met Thr Pro Trp Gln Glu Ile Tyr 545 550 555 560 Arg Ala His Ala Ser Gln Leu Asp Thr Gly Gly Val Leu Glu Phe Ala 565 570 575 Val Lys Tyr Gln Asp Leu Ala Ala Lys Leu Pro Arg His Asn His 580 585 590 <210> 62 <211> 1785 <212> DNA <213> Burkholderia xenovorans <400> 62 atgtcagcca gtacccctcg tcgtcttcgc tctcaggagt ggttcgacga tccgagtcac 60 gctgatatga ccgcattata tgtcgagcgt tttatgaact atggattaac tcgcgaggaa 120 ttacaatctg gtcgtccgat tatcggcatt gcacagaccg gaagcgacct ggcgccctgt 180 aatcgccacc atatcgagtt agccgcgcgt acaaaggccg gaattcgcga cgcgggtgga 240 attccaatgg agttccccgt ccatccgctt gccgaacaat cccgccgccc gacggccgcg 300 cttgaccgta atctggctta cttgggtctg gtggaaatct tacacggttt ccccttagat 360 ggcgtagttt taaccactgg ttgcgataaa accacgccag cgtgtctgat ggcagctgcg 420 acagtggata tgccagcgat cgtgttaagt ggcggcccca tgcttgatgg atggcacgag 480 ggtaaacgcg tgggctctgg tacagtcatt tggcatgcgc gcaacctgtt ggcgacaggg 540 gagattgatt acgaagggtt catgcagttg actacggcgt cgagtccatc tatcggccac 600 tgtaatacta tggggacagc gttatccatg aactcattgg cagaggcgct gggtatgagc 660 ttacccggtt gtgcttcaat tcccgctgca tatcgtgaac gtggccagat ggcgtacgcg 720 acaggcaagc gtattgtgga cttggtccgc gaagatgtcc gtccaagcca tatcatgact 780 aaggcggcat ttgaaaatgc catcgtcgtt gcaagcgccc ttggcgcgag cacaaattgt 840 ccaccacatt taattgctat tgcgcgccat atgggggtgg aattgacttt agatgattgg 900 caacgcgtgg gtgagcaggt gccccttatt gtaaactgca tgcccgctgg cgaatatttg 960 ggtgaatcgt ttcatcgtgc tgggggtgtt ccagctgtct tgcgcgagtt ggtccacgct 1020 caactggttc atcgcgagtg tctgactgtc agtggccgca ctattggaga gatcgcggcg 1080 gatgccccat cgccagatcg tgacgtgatt cgcactactg aagacccgct taaacatggg 1140 gcgggattta tcgtattgtc aggcaacttt ttcgatagtg caattatgaa gatgagtgtg 1200 gtaggcgaag cctttcgtca aacgtacctg tccgagcctg gcaacgagaa cacatttgaa 1260 gctcgtgcca tcgtttttga cggcccagaa gattaccacg cgcgtatcaa tgaccctagc 1320 cttgatattg accaacattg tattttagtc attcgtggag ccggtaccgt aggatatccg 1380 ggcagtgctg aggtagtgaa catggccccg ccggcggaat tagtgcgtca gggtatcacc 1440 tcattgccaa cacttggaga tggccgtcaa agcggcacat cagcttcgcc gtctattttg 1500 aacatgagcc ctgaggccgc cgttggcgga ggattaagtc ttttgcgcac caatgaccgc 1560 attcgtgttg atctgaactc acgctcagta aatgtactgg ttgatgacga agaattagca 1620 cgccgtcgtg aaaccgcaac gttcgctatt cctcctaccc agacgccgtg gcaagaatta 1680 taccgccaga cggtggggca gctgagtact ggcgggtgtc tggaacctgc taccttatat 1740 ttgaaagtta tcgcagaacg tggcaatccg cgtcacagcc actaa 1785 <210> 63 <211> 594 <212> PRT <213> Burkholderia xenovorans <400> 63 Met Ser Ala Ser Thr Pro Arg Arg Leu Arg Ser Gln Glu Trp Phe Asp 1 5 10 15 Asp Pro Ser His Ala Asp Met Thr Ala Leu Tyr Val Glu Arg Phe Met 20 25 30 Asn Tyr Gly Leu Thr Arg Glu Glu Leu Gln Ser Gly Arg Pro Ile Ile 35 40 45 Gly Ile Ala Gln Thr Gly Ser Asp Leu Ala Pro Cys Asn Arg His His 50 55 60 Ile Glu Leu Ala Ala Arg Thr Lys Ala Gly Ile Arg Asp Ala Gly Gly 65 70 75 80 Ile Pro Met Glu Phe Pro Val His Pro Leu Ala Glu Gln Ser Arg Arg 85 90 95 Pro Thr Ala Ala Leu Asp Arg Asn Leu Ala Tyr Leu Gly Leu Val Glu 100 105 110 Ile Leu His Gly Phe Pro Leu Asp Gly Val Val Leu Thr Thr Gly Cys 115 120 125 Asp Lys Thr Thr Pro Ala Cys Leu Met Ala Ala Ala Thr Val Asp Met 130 135 140 Pro Ala Ile Val Leu Ser Gly Gly Pro Met Leu Asp Gly Trp His Glu 145 150 155 160 Gly Lys Arg Val Gly Ser Gly Thr Val Ile Trp His Ala Arg Asn Leu 165 170 175 Leu Ala Thr Gly Glu Ile Asp Tyr Glu Gly Phe Met Gln Leu Thr Thr 180 185 190 Ala Ser Ser Pro Ser Ile Gly His Cys Asn Thr Met Gly Thr Ala Leu 195 200 205 Ser Met Asn Ser Leu Ala Glu Ala Leu Gly Met Ser Leu Pro Gly Cys 210 215 220 Ala Ser Ile Pro Ala Ala Tyr Arg Glu Arg Gly Gln Met Ala Tyr Ala 225 230 235 240 Thr Gly Lys Arg Ile Val Asp Leu Val Arg Glu Asp Val Arg Pro Ser 245 250 255 His Ile Met Thr Lys Ala Ala Phe Glu Asn Ala Ile Val Val Ala Ser 260 265 270 Ala Leu Gly Ala Ser Thr Asn Cys Pro Pro His Leu Ile Ala Ile Ala 275 280 285 Arg His Met Gly Val Glu Leu Thr Leu Asp Asp Trp Gln Arg Val Gly 290 295 300 Glu Gln Val Pro Leu Ile Val Asn Cys Met Pro Ala Gly Glu Tyr Leu 305 310 315 320 Gly Glu Ser Phe His Arg Ala Gly Gly Val Pro Ala Val Leu Arg Glu 325 330 335 Leu Val His Ala Gln Leu Val His Arg Glu Cys Leu Thr Val Ser Gly 340 345 350 Arg Thr Ile Gly Glu Ile Ala Ala Asp Ala Pro Ser Pro Asp Arg Asp 355 360 365 Val Ile Arg Thr Thr Glu Asp Pro Leu Lys His Gly Ala Gly Phe Ile 370 375 380 Val Leu Ser Gly Asn Phe Phe Asp Ser Ala Ile Met Lys Met Ser Val 385 390 395 400 Val Gly Glu Ala Phe Arg Gln Thr Tyr Leu Ser Glu Pro Gly Asn Glu 405 410 415 Asn Thr Phe Glu Ala Arg Ala Ile Val Phe Asp Gly Pro Glu Asp Tyr 420 425 430 His Ala Arg Ile Asn Asp Pro Ser Leu Asp Ile Asp Gln His Cys Ile 435 440 445 Leu Val Ile Arg Gly Ala Gly Thr Val Gly Tyr Pro Gly Ser Ala Glu 450 455 460 Val Val Asn Met Ala Pro Pro Ala Glu Leu Val Arg Gln Gly Ile Thr 465 470 475 480 Ser Leu Pro Thr Leu Gly Asp Gly Arg Gln Ser Gly Thr Ser Ala Ser 485 490 495 Pro Ser Ile Leu Asn Met Ser Pro Glu Ala Ala Val Gly Gly Gly Leu 500 505 510 Ser Leu Leu Arg Thr Asn Asp Arg Ile Arg Val Asp Leu Asn Ser Arg 515 520 525 Ser Val Asn Val Leu Val Asp Asp Glu Glu Leu Ala Arg Arg Arg Glu 530 535 540 Thr Ala Thr Phe Ala Ile Pro Pro Thr Gln Thr Pro Trp Gln Glu Leu 545 550 555 560 Tyr Arg Gln Thr Val Gly Gln Leu Ser Thr Gly Gly Cys Leu Glu Pro 565 570 575 Ala Thr Leu Tyr Leu Lys Val Ile Ala Glu Arg Gly Asn Pro Arg His 580 585 590 Ser His <210> 64 <211> 1239 <212> DNA <213> Haloferax volcanii <400> 64 atggttgagc aagcgaagct tagcgacccg aacgcggagt acacgatgcg cgacctgtcc 60 gcggagacga tagacatcac gaatccgcga ggtggcgtcc gcgacgccga aatcacggac 120 gtacagacga cgatggtcga cgggaactac ccgtggattc tcgtccgcgt ctacaccgac 180 gcgggcgtcg tcggcaccgg cgaggcctac tggggcggcg gcgacaccgc catcatcgag 240 cggatgaagc cgttcctcgt cggcgagaac cccctcgaca tcgaccgcct gtacgagcat 300 ctcgtccaga agatgtccgg cgagggctcc gtctcgggca aggtcatctc cgccatctcg 360 ggcatcgaaa tcgcgctcca cgacgtcgcc ggaaagctcc tcgacgtgcc cgcctatcaa 420 ctcgtcggcg ggaagtaccg cgacgaggtg cgcgtctact gcgacctcca caccgaagac 480 gaggccaacc cgcaggcctg cgccgaggag ggcgtccgcg tggtcgagga actcggctac 540 gacgccatca agttcgacct cgacgtgccc tcgggccacg agaaggaccg cgcgaaccgc 600 cacctccgaa accccgaaat cgaccacaag gtcgaaatcg tcgaggccgt caccgaggcc 660 gtcggcgacc gcgcggacgt ggcgttcgac tgccactggt cctttaccgg cgggagcgcc 720 aagcgcctcg cgtccgagct ggaagactac gacgtgtggt ggctcgaaga ccccgtgccg 780 ccggagaacc acgacgtgca gaagctcgtg acgcagtcca cgacgacgcc catcgcggtc 840 ggtgagaacg tctaccggaa gttcggccag cggacgctgc tcgaaccgca ggcggtggat 900 atcatcgcgc ccgacctgcc ccgcgtcggc ggcatgcgcg agacgcggaa gattgccgac 960 ctcgcggaca tgtactacat ccccgtggcg atgcacaacg tctcgtcgcc catcgggacg 1020 atggcctccg cgcaggtcgc cgcggccatc ccgaactcgc tggccctcga ataccactcc 1080 taccagctcg gctggtggga ggacctcgtc gaagaggacg acctgattca gaacggtcac 1140 atggagattc ccgaaaagcc cggcctcggg ctgacgctcg acctcgacgc cgtcgaagca 1200 cacatggtcg aaggggagac gctcttcgac gaggagtaa 1239 <210> 65 <211> 412 <212> PRT <213> Haloferax volcanii <400> 65 Met Val Glu Gln Ala Lys Leu Ser Asp Pro Asn Ala Glu Tyr Thr Met 1 5 10 15 Arg Asp Leu Ser Ala Glu Thr Ile Asp Ile Thr Asn Pro Arg Gly Gly 20 25 30 Val Arg Asp Ala Glu Ile Thr Asp Val Gln Thr Thr Met Val Asp Gly 35 40 45 Asn Tyr Pro Trp Ile Leu Val Arg Val Tyr Thr Asp Ala Gly Val Val 50 55 60 Gly Thr Gly Glu Ala Tyr Trp Gly Gly Gly Asp Thr Ala Ile Ile Glu 65 70 75 80 Arg Met Lys Pro Phe Leu Val Gly Glu Asn Pro Leu Asp Ile Asp Arg 85 90 95 Leu Tyr Glu His Leu Val Gln Lys Met Ser Gly Glu Gly Ser Val Ser 100 105 110 Gly Lys Val Ile Ser Ala Ile Ser Gly Ile Glu Ile Ala Leu His Asp 115 120 125 Val Ala Gly Lys Leu Leu Asp Val Pro Ala Tyr Gln Leu Val Gly Gly 130 135 140 Lys Tyr Arg Asp Glu Val Arg Val Tyr Cys Asp Leu His Thr Glu Asp 145 150 155 160 Glu Ala Asn Pro Gln Ala Cys Ala Glu Glu Gly Val Arg Val Val Glu 165 170 175 Glu Leu Gly Tyr Asp Ala Ile Lys Phe Asp Leu Asp Val Pro Ser Gly 180 185 190 His Glu Lys Asp Arg Ala Asn Arg His Leu Arg Asn Pro Glu Ile Asp 195 200 205 His Lys Val Glu Ile Val Glu Ala Val Thr Glu Ala Val Gly Asp Arg 210 215 220 Ala Asp Val Ala Phe Asp Cys His Trp Ser Phe Thr Gly Gly Ser Ala 225 230 235 240 Lys Arg Leu Ala Ser Glu Leu Glu Asp Tyr Asp Val Trp Trp Leu Glu 245 250 255 Asp Pro Val Pro Pro Glu Asn His Asp Val Gln Lys Leu Val Thr Gln 260 265 270 Ser Thr Thr Thr Pro Ile Ala Val Gly Glu Asn Val Tyr Arg Lys Phe 275 280 285 Gly Gln Arg Thr Leu Leu Glu Pro Gln Ala Val Asp Ile Ile Ala Pro 290 295 300 Asp Leu Pro Arg Val Gly Gly Met Arg Glu Thr Arg Lys Ile Ala Asp 305 310 315 320 Leu Ala Asp Met Tyr Tyr Ile Pro Val Ala Met His Asn Val Ser Ser 325 330 335 Pro Ile Gly Thr Met Ala Ser Ala Gln Val Ala Ala Ala Ile Pro Asn 340 345 350 Ser Leu Ala Leu Glu Tyr His Ser Tyr Gln Leu Gly Trp Trp Glu Asp 355 360 365 Leu Val Glu Glu Asp Asp Leu Ile Gln Asn Gly His Met Glu Ile Pro 370 375 380 Glu Lys Pro Gly Leu Gly Leu Thr Leu Asp Leu Asp Ala Val Glu Ala 385 390 395 400 His Met Val Glu Gly Glu Thr Leu Phe Asp Glu Glu 405 410 <210> 66 <211> 1179 <212> DNA <213> Clostridium beijerinckii <400> 66 atgaaagaag ttgttattgc gagcgcggtt cgtaccgcga ttggcagcta tggcaagagc 60 ctgaaggatg ttccggcggt ggacctgggt gcgaccgcga tcaaagaggc ggttaagaaa 120 gcgggcatta aaccggagga tgtgaacgaa gttatcctgg gtaacgtgct gcaagcgggt 180 ctgggccaaa acccggcgcg tcaggcgagc ttcaaggcgg gcctgccggt tgaaatcccg 240 gcgatgacca ttaacaaagt ttgcggtagc ggcctgcgta ccgtgagcct ggcggcgcaa 300 atcattaagg cgggtgacgc ggatgttatc attgcgggtg gcatggagaa catgagccgt 360 gcgccgtacc tggcgaacaa cgcgcgttgg ggttatcgta tgggcaacgc gaaattcgtg 420 gacgaaatga ttaccgacgg tctgtgggat gcgtttaacg actaccacat gggcatcacc 480 gcggagaaca ttgcggaacg ttggaacatt agccgtgagg aacaagatga gttcgcgctg 540 gcgagccaga agaaagcgga ggaagcgatc aagagcggcc agtttaaaga cgaaatcgtt 600 ccggtggtta ttaagggtcg taagggtgaa accgtggtgg acaccgatga acacccgcgt 660 ttcggtagca ccattgaggg cctggcgaag ctgaaaccgg cgtttaagaa agatggcacc 720 gtgaccgcgg gtaacgcgag cggcctgaac gactgcgcgg cggtgctggt tatcatgagc 780 gcggagaagg cgaaagaact gggtgtgaag ccgctggcga aaattgttag ctacggtagc 840 gcgggtgtgg acccggcgat catgggttac ggcccgtttt atgcgaccaa ggcggcgatt 900 gagaaagcgg gttggaccgt ggacgaactg gatctgatcg agagcaacga agcgttcgcg 960 gcgcaaagcc tggcggtggc gaaggatctg aaatttgaca tgaacaaggt gaacgtgaac 1020 ggtggtgcga ttgcgctggg tcacccgatt ggtgcgagcg gcgcgcgtat cctggtgacc 1080 ctggttcacg cgatgcagaa acgtgacgcg aagaaaggtc tggcgaccct gtgcattggt 1140 ggtggtcaag gcaccgcgat tctgctggaa aagtgctaa 1179 <210> 67 <211> 393 <212> PRT <213> Clostridium beijerinckii <400> 67 Met Arg Asp Val Val Ile Val Ser Ala Val Arg Thr Ala Leu Gly Ser 1 5 10 15 Phe Gly Gly Ala Leu Lys Asp Val Ser Ala Val Asp Leu Gly Ala Leu 20 25 30 Val Ile Lys Glu Ala Val Asn Arg Ala Gly Val Lys Pro Glu Leu Ile 35 40 45 Glu Glu Val Ile Met Gly Asn Val Ile Gln Ala Gly Leu Gly Gln Asn 50 55 60 Thr Ala Arg Gln Ser Thr Ile Lys Ala Gly Leu Pro Gln Glu Val Ser 65 70 75 80 Ala Met Thr Ile Asn Lys Val Cys Gly Ser Gly Leu Arg Ala Val Ser 85 90 95 Leu Ala Ala Gln Met Ile Lys Ala Gly Asp Ala Asp Val Val Val Ala 100 105 110 Gly Gly Met Glu Asn Met Ser Ala Ala Pro Tyr Ala Leu Asp Lys Ala 115 120 125 Arg Trp Gly Gln Arg Met Gly Asp Gly Lys Leu Val Asp Thr Met Ile 130 135 140 Lys Asp Ala Leu Trp Asp Ala Phe Asn Asn Tyr His Met Gly Val Thr 145 150 155 160 Ala Glu Asn Ile Ala Lys Gln Trp Gly Leu Thr Arg Glu Glu Gln Asp 165 170 175 Ala Phe Ser Ala Ser Ser Gln Gln Lys Ala Glu Ala Ala Ile Lys Ser 180 185 190 Gly Arg Phe Lys Asp Glu Ile Val Pro Val Val Ile Pro Gln Arg Lys 195 200 205 Gly Glu Pro Lys Val Phe Asp Thr Asp Glu Phe Pro Arg Phe Gly Thr 210 215 220 Thr Ala Glu Thr Leu Ala Lys Leu Lys Pro Ala Phe Ile Lys Asp Gly 225 230 235 240 Thr Val Thr Ala Gly Asn Ala Ser Gly Ile Asn Asp Gly Ala Ala Ala 245 250 255 Phe Val Val Met Ser Ala Glu Lys Ala Glu Glu Leu Gly Leu Lys Pro 260 265 270 Met Ala Lys Ile Leu Ser Tyr Gly Ser Lys Gly Leu Asp Pro Ala Ile 275 280 285 Met Gly Tyr Gly Pro Phe His Ala Thr Lys Lys Ala Leu Glu Lys Ala 290 295 300 Asn Leu Thr Val Glu Asp Leu Asp Leu Ile Glu Ala Asn Glu Ala Phe 305 310 315 320 Ala Ala Gln Ser Leu Ala Val Ala Lys Asp Leu Lys Phe Asp Met Ser 325 330 335 Lys Val Asn Val Asn Gly Gly Ala Ile Ala Leu Gly His Pro Val Gly 340 345 350 Ala Ser Gly Ala Arg Ile Leu Val Thr Leu Leu His Glu Met Glu Lys 355 360 365 Arg Asp Ala Lys Lys Gly Leu Ala Thr Leu Cys Ile Gly Gly Gly Met 370 375 380 Gly Thr Ala Leu Ile Val Glu Arg Val 385 390 <210> 68 <211> 1179 <212> DNA <213> Clostridium acetobutylicum <400> 68 atgaaagaag ttgtaatagc tagtgcagta agaacagcga ttggatctta tggaaagtct 60 cttaaggatg taccagcagt agatttagga gctacagcta taaaggaagc agttaaaaaa 120 gcaggaataa aaccagagga tgttaatgaa gtcattttag gaaatgttct tcaagcaggt 180 ttaggacaga atccagcaag acaggcatct tttaaagcag gattaccagt tgaaattcca 240 gctatgacta ttaataaggt ttgtggttca ggacttagaa cagttagctt agcagcacaa 300 attataaaag caggagatgc tgacgtaata atagcaggtg gtatggaaaa tatgtctaga 360 gctccttact tagcgaataa cgctagatgg ggatatagaa tgggaaacgc taaatttgtt 420 gatgaaatga tcactgacgg attgtgggat gcatttaatg attaccacat gggaataaca 480 gcagaaaaca tagctgagag atggaacatt tcaagagaag aacaagatga gtttgctctt 540 gcatcacaaa aaaaagctga agaagctata aaatcaggtc aatttaaaga tgaaatagtt 600 cctgtagtaa ttaaaggcag aaagggagaa actgtagttg atacagatga gcaccctaga 660 tttggatcaa ctatagaagg acttgcaaaa ttaaaacctg ccttcaaaaa agatggaaca 720 gttacagctg gtaatgcatc aggattaaat gactgtgcag cagtacttgt aatcatgagt 780 gcagaaaaag ctaaagagct tggagtaaaa ccacttgcta agatagtttc ttatggttca 840 gcaggagttg acccagcaat aatgggatat ggacctttct atgcaacaaa agcagctatt 900 gaaaaagcag gttggacagt tgatgaatta gatttaatag aatcaaatga agcttttgca 960 gctcaaagtt tagcagtagc aaaagattta aaatttgata tgaataaagt aaatgtaaat 1020 ggaggagcta ttgcccttgg tcatccaatt ggagcatcag gtgcaagaat actcgttact 1080 cttgtacacg caatgcaaaa aagagatgca aaaaaaggct tagcaacttt atgtataggt 1140 ggcggacaag gaacagcaat attgctagaa aagtgctag 1179 <210> 69 <211> 392 <212> PRT <213> Clostridium acetobutylicum <400> 69 Met Lys Glu Val Val Ile Ala Ser Ala Val Arg Thr Ala Ile Gly Ser 1 5 10 15 Tyr Gly Lys Ser Leu Lys Asp Val Pro Ala Val Asp Leu Gly Ala Thr 20 25 30 Ala Ile Lys Glu Ala Val Lys Lys Ala Gly Ile Lys Pro Glu Asp Val 35 40 45 Asn Glu Val Ile Leu Gly Asn Val Leu Gln Ala Gly Leu Gly Gln Asn 50 55 60 Pro Ala Arg Gln Ala Ser Phe Lys Ala Gly Leu Pro Val Glu Ile Pro 65 70 75 80 Ala Met Thr Ile Asn Lys Val Cys Gly Ser Gly Leu Arg Thr Val Ser 85 90 95 Leu Ala Ala Gln Ile Ile Lys Ala Gly Asp Ala Asp Val Ile Ile Ala 100 105 110 Gly Gly Met Glu Asn Met Ser Arg Ala Pro Tyr Leu Ala Asn Asn Ala 115 120 125 Arg Trp Gly Tyr Arg Met Gly Asn Ala Lys Phe Val Asp Glu Met Ile 130 135 140 Thr Asp Gly Leu Trp Asp Ala Phe Asn Asp Tyr His Met Gly Ile Thr 145 150 155 160 Ala Glu Asn Ile Ala Glu Arg Trp Asn Ile Ser Arg Glu Glu Gln Asp 165 170 175 Glu Phe Ala Leu Ala Ser Gln Lys Lys Ala Glu Glu Ala Ile Lys Ser 180 185 190 Gly Gln Phe Lys Asp Glu Ile Val Pro Val Val Ile Lys Gly Arg Lys 195 200 205 Gly Glu Thr Val Val Asp Thr Asp Glu His Pro Arg Phe Gly Ser Thr 210 215 220 Ile Glu Gly Leu Ala Lys Leu Lys Pro Ala Phe Lys Lys Asp Gly Thr 225 230 235 240 Val Thr Ala Gly Asn Ala Ser Gly Leu Asn Asp Cys Ala Ala Val Leu 245 250 255 Val Ile Met Ser Ala Glu Lys Ala Lys Glu Leu Gly Val Lys Pro Leu 260 265 270 Ala Lys Ile Val Ser Tyr Gly Ser Ala Gly Val Asp Pro Ala Ile Met 275 280 285 Gly Tyr Gly Pro Phe Tyr Ala Thr Lys Ala Ala Ile Glu Lys Ala Gly 290 295 300 Trp Thr Val Asp Glu Leu Asp Leu Ile Glu Ser Asn Glu Ala Phe Ala 305 310 315 320 Ala Gln Ser Leu Ala Val Ala Lys Asp Leu Lys Phe Asp Met Asn Lys 325 330 335 Val Asn Val Asn Gly Gly Ala Ile Ala Leu Gly His Pro Ile Gly Ala 340 345 350 Se...
Claims
1. A recombinant microorganism capable of producing a fermentation product from a feedstock comprising xylose and glucose, wherein the recombinant microorganism simultaneously utilizes xylose and glucose, and wherein the recombinant microorganism comprises: (a) deletion or inactivation of a nucleic acid sequence encoding a xylose ABC transporter and an arabinose ABC transporter; (b) one or more endogenous or exogenous nucleic acid sequences encoding at least one C5 sugar symporter and operably linked to one or more constitutive promoters ; wherein the C5 sugar symporter comprises: (1) a xylose symporter and / or (2) Escherichia coli AraE; (c) one or more endogenous or exogenous nucleic acid sequences which (1) encode xylose isomerase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more nucleic acid sequences encoding xylulokinase, or (2) encode xylose dehydrogenase and are operably linked to one or more constitutive promoters, and deletion or inactivation of one or more nucleic acid sequences encoding xylose isomerase and / or xylulokinase; (d) deletion or inactivation of one or more nucleic acid sequences encoding glycollate dehydrogenase; and (e) a native functional phosphotransferase system (PTS) and a native cAMP receptor protein (CRP); wherein the recombinant microorganism is Escherichia coli.
2. The recombinant microorganism according to claim 1, wherein the fermentation product comprises monoethylene glycol (MEG), acetone, isopropanol, glycollate, and / or propylene.
3. The recombinant microorganism according to claim 2, wherein two or more fermentation products are produced simultaneously.
4. A recombinant microorganism capable of producing monoethylene glycol (MEG), isopropanol, and / or acetone from a feedstock comprising xylose and glucose, wherein the recombinant microorganism simultaneously utilizes xylose and glucose, wherein the recombinant microorganism comprises: (a) deletion or inactivation of nucleic acid sequences encoding glyoxaldehyde dehydrogenase, a xylose ABC transporter, and an arabinose ABC transporter; (b) at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter ; wherein the C5 sugar symporter comprises: (1) a xylose symporter and / or (2) Escherichia coli AraE; (c) deletion or inactivation of glycollate dehydrogenase and deletion or inactivation of xylose isomerase or xylulokinase; and (d) a native functional phosphotransferase system (PTS) and a native cAMP receptor protein (CRP), wherein the recombinant microorganism comprises an enzymatic pathway for MEG, isopropanol, and / or acetone production, and wherein the recombinant microorganism is Escherichia coli.
5. The recombinant microorganism according to claim 4, wherein the glycollate dehydrogenase is glcDEF.
6. The recombinant microorganism according to claim 4, wherein the arabinose ABC transporter is araFGH.
7. The recombinant microorganism according to claim 4, wherein the xylose ABC transporter is xylFGH.
8. The recombinant microorganism according to claim 4, wherein the C5 sugar symporter is the xylose symporter XylE.
9. The recombinant microorganism according to claim 8, wherein the XylE comprises the amino acid sequence set forth in SEQ ID NO:
49.
10. The recombinant microorganism according to claim 8, wherein the XylE is encoded by the nucleic acid sequence set forth in SEQ ID NO:
48.
11. The recombinant microorganism according to claim 8, wherein the xylose symporter is endogenous to the microorganism.
12. The recombinant microorganism according to claim 4, wherein the C5 sugar symporter is Escherichia coli AraE.
13. The recombinant microorganism according to claim 12, wherein the AraE comprises the amino acid sequence set forth in SEQ ID NO:
47.
14. The recombinant microorganism according to claim 12, wherein the AraE is encoded by the nucleic acid sequence set forth in SEQ ID NO:
46.
15. The recombinant microorganism according to claim 4, wherein the CRP comprises the amino acid sequence set forth in SEQ ID NO:
10.
16. The recombinant microorganism according to claim 4, wherein the CRP is encoded by the nucleic acid sequence set forth in SEQ ID NO:
9.
17. The recombinant microorganism according to claim 4, wherein the recombinant microorganism further comprises one or more endogenous or exogenous nucleic acid sequences encoding a constitutive promoter, at least one ketohexokinase, fructose-bisphosphate aldolase, and glycoaldehyde reductase.
18. The recombinant microorganism according to claim 17, wherein the constitutive promoter has the nucleic acid sequence set forth in SEQ ID NO:
53.
19. The recombinant microorganism according to claim 4, wherein the xylose isomerase is xylA.
20. The recombinant microorganism according to claim 4, wherein the xylose isomerase is endogenous to the microorganism.
21. The recombinant microorganism according to claim 19, wherein the xylA comprises the amino acid sequence set forth in SEQ ID NO:
6.
22. The recombinant microorganism according to claim 19, wherein the xylA is encoded by the nucleic acid sequence set forth in SEQ ID NO:
5.
23. The recombinant microorganism according to claim 17, wherein the ketohexokinase is from Homo sapiens.
24. The recombinant microorganism according to claim 17, wherein the ketohexokinase is heterologous to the microorganism.
25. The recombinant microorganism according to claim 24, wherein the ketohexokinase is khk-C.
26. The recombinant microorganism according to claim 25, wherein the khk-C comprises the amino acid sequence set forth in SEQ ID NO:
12.
27. The recombinant microorganism according to claim 25, wherein the khk-C is encoded by the nucleic acid sequence set forth in SEQ ID NO:
11.
28. The recombinant microorganism according to claim 17, wherein the fructose-bisphosphate aldolase is from Homo sapiens.
29. The recombinant microorganism according to claim 17, wherein the fructose bisphosphate aldolase is heterologous to the microorganism.
30. The recombinant microorganism according to claim 29, wherein the fructose bisphosphate aldolase is aldoB.
31. The recombinant microorganism according to claim 30, wherein the aldoB comprises the amino acid sequence set forth in SEQ ID NO:
51.
32. The recombinant microorganism according to claim 30, wherein the aldoB is encoded by the nucleic acid sequence set forth in SEQ ID NO:
50.
33. The recombinant microorganism according to claim 17, wherein the glycoaldehyde reductase is endogenous to the microorganism.
34. The recombinant microorganism according to claim 17, wherein the glycoaldehyde reductase is fucO.
35. The recombinant microorganism according to claim 34, wherein the fucO comprises the amino acid sequence set forth in SEQ ID NO:
98.
36. The recombinant microorganism according to claim 34, wherein the fucO is encoded by the nucleic acid sequence set forth in SEQ ID NO:
52.
37. The recombinant microorganism according to claim 4, wherein the xylulokinase is xylB.
38. The recombinant microorganism according to claim 37, wherein the xylB comprises the amino acid sequence set forth in SEQ ID NO:
14.
39. The recombinant microorganism according to claim 37, wherein the xylB is encoded by the nucleic acid sequence set forth in SEQ ID NO:
13.
40. The recombinant microorganism according to claim 4, wherein the recombinant microorganism further comprises one or more endogenous or exogenous nucleic acid sequences encoding a constitutive promoter, at least one xylose dehydrogenase, and a glycoaldehyde reductase.
41. The recombinant microorganism according to claim 40, wherein the xylose dehydrogenase is from Caulobacter crescentus, Burkholderia xenovorans, or Haloferax volcanii.
42. The recombinant microorganism according to claim 41, wherein the xylose dehydrogenase is xdh.
43. The recombinant microorganism according to claim 42, wherein the xdh comprises the amino acid sequence set forth in SEQ ID NO: 16, 17, or 19.
44. The recombinant microorganism according to claim 42, wherein the xdh is encoded by the nucleic acid sequence set forth in SEQ ID NO: 15, 18, or 97.
45. The recombinant microorganism according to claim 40, wherein the xylose dehydrogenase is heterologous to the microorganism.
46. The recombinant microorganism according to claim 40, wherein the glycoaldehyde reductase is endogenous to the microorganism.
47. The recombinant microorganism according to claim 40, wherein the glycoaldehyde reductase is fucO.
48. The recombinant microorganism according to claim 47, wherein the fucO comprises the amino acid sequence set forth in SEQ ID NO:
98.
49. The recombinant microorganism according to claim 47, wherein the fucO is encoded by a nucleic acid sequence comprising SEQ ID NO:
52.
50. The recombinant microorganism according to claim 40, wherein the hydroxymethylglyoxal reductase is heterologous to the microorganism.
51. The recombinant microorganism according to claim 40, wherein the constitutive promoter has a nucleic acid sequence comprising SEQ ID NO:
53.
52. The recombinant microorganism according to claim 17 or 40, wherein the recombinant microorganism further comprises one or more endogenous or exogenous nucleic acid sequences encoding a constitutive promoter, acetoacetyl-CoA thiolase, acetate:acetoacetyl-CoA transferase, and / or acetoacetate decarboxylase.
53. The recombinant microorganism according to claim 52, wherein the constitutive promoter has a nucleic acid sequence comprising SEQ ID NO:
78.
54. The recombinant microorganism according to claim 52, wherein the acetoacetyl-CoA thiolase is from Clostridium acetobutylicum.
55. The recombinant microorganism according to claim 54, wherein the acetoacetyl-CoA thiolase comprises an amino acid sequence comprising SEQ ID NO:67 or 69.
56. The recombinant microorganism according to claim 54, wherein the acetoacetyl-CoA thiolase is encoded by a nucleic acid sequence comprising SEQ ID NO:66 or 68.
57. The recombinant microorganism according to claim 52, wherein the acetate:acetoacetyl-CoA transferase is AtoDA having subunit α and subunit β.
58. The recombinant microorganism according to claim 57, wherein AtoDA subunit α comprises an amino acid sequence comprising SEQ ID NO:
72.
59. The recombinant microorganism according to claim 57, wherein AtoDA subunit α is encoded by a nucleic acid sequence comprising SEQ ID NO:
70.
60. The recombinant microorganism according to claim 57, wherein AtoDA subunit β comprises an amino acid sequence comprising SEQ ID NO:
73.
61. The recombinant microorganism according to claim 57, wherein the AtoDA subunit β is encoded by a nucleic acid sequence comprising SEQ ID NO:
71.
62. The recombinant microorganism according to claim 52, wherein the acetoacetate decarboxylase is from Clostridium beijerinckii or Clostridium acetobutylicum.
63. The recombinant microorganism according to claim 62, wherein the acetoacetate decarboxylase is Adc.
64. The recombinant microorganism according to claim 63, wherein the Adc comprises an amino acid sequence comprising SEQ ID NO:75 or 77.
65. The recombinant microorganism according to claim 63, wherein the Adc is encoded by a nucleic acid sequence comprising SEQ ID NO:74 or 76.
66. The recombinant microorganism according to claim 17 or 40, wherein the recombinant microorganism further comprises one or more endogenous or exogenous nucleic acid sequences encoding a constitutive promoter, acetoacetyl-CoA thiolase, acetate:acetoacetyl-CoA transferase, acetoacetate decarboxylase, and alcohol dehydrogenase.
67. A recombinant microorganism capable of producing glycolic acid from a feedstock comprising xylose and glucose, wherein the recombinant microorganism simultaneously utilizes xylose and glucose, and wherein the recombinant microorganism comprises: (a) deletion or inactivation of a nucleic acid sequence encoding one or more glyoxaldehyde reductases, xylose ABC transporter, and arabinose ABC transporter; (b) at least one endogenous or exogenous nucleic acid molecule operably linked to one or more constitutive promoters and encoding a C5 sugar symporter ; wherein the C5 sugar symporter comprises: (1) a xylose symporter and / or (2) Escherichia coli AraE; (c) deletion or inactivation of xylose isomerase or xylulokinase; and (d) a native functional phosphotransferase system (PTS) and a native cAMP receptor protein (CRP), wherein the recombinant microorganism comprises an enzymatic pathway for glycolic acid production, and wherein the recombinant microorganism is Escherichia coli.
68. The recombinant microorganism according to claim 67, wherein the arabinose ABC transporter is araFGH.
69. The recombinant microorganism according to claim 67, wherein the xylose ABC transporter is xylFGH.
70. The recombinant microorganism according to claim 67, wherein the C5 sugar symporter is the xylose symporter XylE.
71. The recombinant microorganism according to claim 70, wherein the XylE comprises an amino acid sequence comprising SEQ ID NO:
49.
72. The recombinant microorganism according to claim 70, wherein the XylE is encoded by a nucleic acid sequence comprising SEQ ID NO:
48.
73. The recombinant microorganism according to claim 70, wherein the xylose symporter is endogenous to the microorganism.
74. The recombinant microorganism according to claim 67, wherein the C5 sugar symporter is Escherichia coli AraE.
75. The recombinant microorganism according to claim 67, wherein the CRP comprises an amino acid sequence comprising SEQ ID NO:
10.
76. The recombinant microorganism according to claim 67, wherein the CRP is encoded by a nucleic acid sequence comprising SEQ ID NO:
9.
77. The recombinant microorganism according to claim 67, wherein the xylose isomerase is xylA.
78. The recombinant microorganism according to claim 77, wherein the xylA comprises an amino acid sequence comprising SEQ ID NO:
6.
79. The recombinant microorganism according to claim 77, wherein the xylA is encoded by a nucleic acid sequence comprising SEQ ID NO:
5.
80. The recombinant microorganism according to claim 67, wherein the xylose isomerase is endogenous to the microorganism.
81. The recombinant microorganism according to claim 4, wherein the glycolaldehyde dehydrogenase is aldA.
82. The recombinant microorganism according to claim 81, wherein the aldA comprises the amino acid sequence comprising SEQ ID NO:
4.
83. The recombinant microorganism according to claim 81, wherein the aldA is encoded by the nucleic acid sequence comprising SEQ ID NO:
3.
84. The recombinant microorganism according to claim 4, wherein the glycolaldehyde dehydrogenase is endogenous to the microorganism.
85. The recombinant microorganism according to claim 67, wherein the xylulokinase is xylB.
86. The recombinant microorganism according to claim 85, wherein the xylB comprises the amino acid sequence comprising SEQ ID NO:
14.
87. The recombinant microorganism according to claim 85, wherein the xylB is encoded by the nucleic acid sequence comprising SEQ ID NO:
13.
88. The recombinant microorganism according to claim 67, wherein the glycolaldehyde reductase is fucO and yqhD.
89. The recombinant microorganism according to claim 1, wherein the recombinant microorganism comprises the deletion or inactivation of one or more endogenous or exogenous nucleic acid sequences encoding xylose dehydrogenase and glycolaldehyde reductase and operably linked to one or more constitutive promoters, and one or more nucleic acid sequences encoding xylose isomerase.
90. The recombinant microorganism according to claim 1, wherein the fermentation product comprises monoethylene glycol (MEG) and / or acetone.
91. The recombinant microorganism according to claim 4, wherein the recombinant microorganism comprises the deletion or inactivation of xylulokinase.
92. The recombinant microorganism according to claim 4, wherein the recombinant microorganism comprises the deletion or inactivation of xylose isomerase.
93. The recombinant microorganism according to claim 4, wherein the recombinant microorganism is capable of producing monoethylene glycol (MEG) and / or acetone and comprises an enzymatic pathway for MEG and / or acetone production.
Citation Information
Patent Citations
Targeted disruption of t cell receptor genes using engineered zinc finger protein nucleases
US20140301990A1
Methods for obtaining nucleic acid containing a mutation
US6033861A
End selection in directed evolution
US6773900B2
Dry food product containing live probiotic
US8460726B2
CRISPR-Cas component systems, methods and compositions for sequence manipulation
US8795965B2