Aldehyde dehydrogenase variants and methods of use

Aldehyde dehydrogenase variants in engineered cells facilitate the production of 3-hydroxybutyraldehyde and 1,3-butanediol, and 4-hydroxybutyraldehyde and 1,4-butanediol, addressing the need for renewable chemical production with improved efficiency and reduced by-products.

JP7763026B2Active Publication Date: 2025-10-31GENOMATICA INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2019553009
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-03-31
Filing Date
2018-03-29
Publication Date
2025-10-31
Estimated Expiration
2038-03-29

AI Technical Summary

Technical Problem

There is a need for more efficient and renewable methods to produce chemicals like 1,4-butanediol and 1,3-butanediol, which are typically derived from petroleum, to reduce energy and capital intensity in production processes.

Method used

Development of aldehyde dehydrogenase variants and cells expressing these variants to convert substrates like 3-hydroxybutyryl-CoA and 4-hydroxybutyryl-CoA into 3-hydroxybutyraldehyde and 4-hydroxybutyraldehyde, which can then be further processed to produce 1,3-butanediol and 1,4-butanediol, using engineered microbial organisms and in vitro biocatalysts.

Benefits of technology

Enhances the production of 3-hydroxybutyraldehyde and 1,3-butanediol or their esters/amides, and 4-hydroxybutyraldehyde and 1,4-butanediol or their esters/amides, with improved activity and specificity, reducing by-products and increasing efficiency in fermentation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007763026000045
    Figure 0007763026000045
  • Figure 0007763026000046
    Figure 0007763026000046
  • Figure 0007763026000047
    Figure 0007763026000047
Patent Text Reader

Abstract

The present invention provides polypeptides and encoding nucleic acids of aldehyde dehydrogenase variants. The present invention also provides cells expressing the aldehyde dehydrogenase variants. The present invention further provides methods for producing 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or esters or amides thereof, comprising culturing cells expressing the aldehyde dehydrogenase variants or using a lysate of such cells. The present invention further provides methods for producing 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) or esters or amides thereof, comprising culturing cells expressing the aldehyde dehydrogenase variants or using a lysate of such cells.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 480,194, filed March 31, 2017, the entire contents of which are incorporated herein by reference.

[0002] Reference is made to the following provisional international applications, which are incorporated herein by reference in their entireties: (1) U.S. Provisional Application No. 62 / 480,208, entitled "3-HYDROXYBUTYRYL-COA DEHYDROGENASE VARIANTS AND METHODS OF USE," filed March 31, 2017 (Attorney Docket No. 12956-409-888); (2) U.S. Provisional Application No. 62 / 480,270, entitled "PROCESS AND SYSTEMS FOR OBTAINING 1,3-BUTANEDIOL FROM FERMENTATION BROTHS," filed March 31, 2017 (Attorney Docket No. 12956-407-888); (3) U.S. Provisional Application No. 62 / 480,270, entitled "3-HYDROXYBUTYRYL-COA DEHYDROGENASE VARIANTS AND METHODS OF USE," filed March 31, 2017 (Attorney Docket No. 12956-407-888); (4) International Patent Application No. ____ (Attorney Docket No. 12956-409-228), entitled "PROCESS AND SYSTEMS FOR OBTAINING 1,3-BUTANEDIOL FROM FERMENTATION BROTHS," filed on even date herewith (Attorney Docket No. 12956-407-228).

[0003] This application incorporates by reference herein the Sequence Listing as an ASCII text file entitled "12956-408-228_Sequence_Listing.txt", created on March 27, 2018, and having a size of 498,061 bytes.

[0004] The present invention relates generally to organisms engineered to produce desired products, to engineered enzymes that facilitate the production of desired products, and more specifically to enzymes and cells that produce desired products such as 3-hydroxybutyraldehyde, 1,3-butanediol, 4-hydroxybutyraldehyde, 1,4-butanediol, and related and derived products. [Background technology]

[0005] A variety of mass-produced chemicals are used to produce desired products for commercial applications. Many of these mass-produced chemicals are derived from petroleum. These mass-produced chemicals have a variety of uses, including as solvents, resins, polymer precursors, and specialty chemicals. Desired mass-produced chemicals include four-carbon molecules, such as 1,4-butanediol and 1,3-butanediol, which are upstream precursors and downstream products, respectively. It is desirable to develop methods for the production of mass-produced chemicals to provide renewable sources for petroleum-based products and to provide less energy- and capital-intensive processes.

[0006] Thus, there exists a need for methods for enhancing the production of desired products. The present invention fulfills this need and provides related advantages as well. Summary of the Invention [Means for solving the problem]

[0007] The present invention provides polypeptides and encoding nucleic acids of aldehyde dehydrogenase variants. The present invention also provides cells expressing the aldehyde dehydrogenase variants. The present invention further provides methods for producing 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or esters or amides thereof, comprising culturing cells expressing the aldehyde dehydrogenase variants or using a lysate of such cells. The present invention further provides methods for producing 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) or esters or amides thereof, comprising culturing cells expressing the aldehyde dehydrogenase variants or using a lysate of such cells. In an embodiment of the present invention, for example, the following items are provided: (Item 1) (a) a nucleic acid molecule encoding an amino acid sequence referenced as SEQ ID NO: 1, 2 or 3 or in Table 4, wherein said amino acid sequence contains one or more of the amino acid substitutions set forth in Tables 1, 2 and / or 3; (b) a nucleic acid molecule that hybridizes to the nucleic acid of (a) under highly stringent hybridization conditions and comprises a nucleic acid sequence that encodes one or more of the amino acid substitutions set forth in Tables 1, 2 and / or 3; (c) a nucleic acid molecule encoding an amino acid sequence comprising the consensus sequence of Loop A (SEQ ID NO: 5) and / or Loop B (SEQ ID NO: 6), wherein the amino acid sequence comprises one or more of the amino acid substitutions set forth in Tables 1, 2, and / or 3; and (d) a nucleic acid molecule complementary to (a) or (b); An isolated nucleic acid molecule selected from: (Item 2) 2. The isolated nucleic acid molecule of claim 1, wherein the amino acid sequence other than the one or more amino acid substitutions has at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity to or is identical to an amino acid sequence referenced in SEQ ID NO: 1, 2, or 3 or in Table 4. (Item 3) 3. The nucleic acid of item 1 or 2, wherein the amino acid sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or 16 of the amino acid substitutions set out in Tables 1, 2 and / or 3. (Item 4) A vector containing the nucleic acid molecule of any one of items 1 to 3. (Item 5) 5. The vector according to item 4, which is an expression vector. (Item 6) 6. The vector according to item 4 or 5, comprising double-stranded DNA. (Item 7) 1. An isolated polypeptide comprising an amino acid sequence referenced as SEQ ID NO: 1, 2 or 3 or in Table 4, wherein said amino acid sequence comprises one or more of the amino acid substitutions set forth in Tables 1, 2 and / or 3. (Item 8) An isolated polypeptide comprising the consensus amino acid sequence of Loop A (SEQ ID NO: 5) and / or Loop B (SEQ ID NO: 6). (Item 9) 1. An isolated polypeptide comprising an amino acid sequence referenced as SEQ ID NO: 1, 2 or 3 or in Table 4, wherein said amino acid sequence comprises one or more of the amino acid substitutions set forth in Tables 1, 2 and / or 3, and wherein said amino acid sequence other than said one or more amino acid substitutions has at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98% or 99% sequence identity to or is identical to the amino acid sequence referenced as SEQ ID NO: 1, 2 or 3 or in Table 4. (Item 10) 10. The isolated polypeptide of item 9, wherein the amino acid sequence further comprises conservative amino acid substitutions at 1 to 100 amino acid positions, the positions being other than the one or more amino acid substitutions shown in Tables 1, 2 and / or 3. (Item 11) 11. The isolated polypeptide of any one of items 7 to 10, wherein the amino acid sequence does not contain any alterations at 2 to 300 amino acid positions compared to the parent sequence other than the one or more amino acid substitutions shown in Tables 1, 2 and / or 3, and the positions are selected from positions that are identical to 2, 3, 4 or 5 of the amino acid sequences referenced as SEQ ID NO: 1, 2 or 3 or in Table 4. (Item 12) 12. The isolated polypeptide of any one of items 7 to 11, wherein the amino acid sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or 16 of the amino acid substitutions set out in Tables 1, 2 and / or 3. (Item 13) 13. The isolated polypeptide of any one of items 7 to 12, encoding an aldehyde dehydrogenase. (Item 14) 14. The isolated polypeptide of any one of items 7 to 13, which is capable of converting 3-hydroxybutyryl-CoA into 3-hydroxybutyraldehyde. (Item 15) 14. The isolated polypeptide according to any one of items 7 to 13, which is capable of converting 4-hydroxybutyryl-CoA into 4-hydroxybutyraldehyde. (Item 16) 16. The isolated polypeptide of any one of items 7 to 15, which has higher activity compared to the parent polypeptide. (Item 17) 15. The isolated polypeptide of any one of items 7 to 14, having higher activity for 3-hydroxy-(R)-butyryl-CoA than for 3-hydroxy-(S)-butyryl-CoA. (Item 18) 15. The isolated polypeptide of any one of items 7 to 14, having a higher specificity for 3-hydroxybutyryl-CoA than acetyl-CoA. (Item 19) 16. The isolated polypeptide of any one of items 7 to 13 or 15, having a higher specificity for 4-hydroxybutyryl-CoA than acetyl-CoA. (Item 20) 16. The isolated polypeptide of any one of items 7 to 15, which reduces the production of by-products in a cell or cell extract. (Item 21) 21. The isolated polypeptide of claim 20, wherein the by-product is ethanol or 4-hydroxy-2-butanone. (Item 22) 16. The isolated polypeptide of any one of items 7 to 15, having a higher kcat than the parent polypeptide. (Item 23) A cell comprising the vector of any one of items 4 to 6. (Item 24) A cell comprising the nucleic acid of any one of items 1 to 3. (Item 25) 25. The cell of item 24, wherein the nucleic acid molecule is integrated into a chromosome of the cell. (Item 26) 26. The cell of item 25, wherein the integration is site-specific. (Item 27) 27. The cell of any one of items 23 to 26, wherein the nucleic acid molecule is expressed. (Item 28) 23. A cell comprising the polypeptide of any one of items 7 to 22. (Item 29) 29. The cell of any one of items 23 to 28, which is a microbial organism. (Item 30) 30. The cell of item 29, wherein the microbial organism is a bacterium, yeast, or fungus. (Item 31) 29. The cell of any one of items 23 to 28, which is an isolated eukaryotic cell. (Item 32) 32. The cell of any one of items 23 to 31, comprising a pathway to produce 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or esters or amides thereof. (Item 33) 32. The cell of any one of items 23 to 31, comprising a pathway to produce 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) or esters or amides thereof. (Item 34) 34. The cell of any one of items 23 to 33, which is capable of fermentation. (Item 35) 35. The cell of any one of items 23 to 34, further comprising at least one substrate for the polypeptide. (Item 36) 36. The cell of item 35, wherein the substrate is 3-hydroxybutyryl-CoA. (Item 37) 37. The cell according to item 36, wherein the substrate is 3-hydroxy-(R)-butyryl-CoA. (Item 38) 38. The cell of item 36 or 37, having higher activity for 3-hydroxy-(R)-butyryl-CoA than for 3-hydroxy-(S)-butyryl-CoA. (Item 39) 36. The cell of item 35, wherein the substrate is 4-hydroxybutyryl-CoA. (Item 40) 23. Use of a polypeptide according to any one of items 7 to 22 as a biocatalyst. (Item 41) 23. A composition comprising a polypeptide according to any one of items 7 to 22 and at least one substrate for said polypeptide. (Item 42) 42. The composition of claim 41, wherein the polypeptide is capable of reacting with the substrate under in vitro conditions. (Item 43) 43. The composition according to item 41 or 42, wherein the substrate is 3-hydroxybutyryl-CoA. (Item 44) 44. The composition according to item 43, wherein the substrate is 3-hydroxy-(R)-butyryl-CoA. (Item 45) 43. The composition according to item 41 or 42, wherein the substrate is 4-hydroxybutyryl-CoA. (Item 46) 40. A medium comprising the cells of any one of items 23 to 39. (Item 47) 7. A method for constructing a host strain, comprising the step of introducing the vector according to any one of items 4 to 6 into a cell capable of fermentation. (Item 48) 40. A method for producing 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or an ester or amide thereof, comprising culturing the cell of any one of items 23 to 39 to produce 3-HBal and / or 1,3-BDO or an ester or amide thereof. (Item 49) 40. A method for producing 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) or an ester or amide thereof, comprising culturing the cell of any one of items 23 to 39 to produce 4-HBal and / or 1,4-BDO or an ester or amide thereof. (Item 50) 50. The method of claim 48 or 49, wherein the cells are in a substantially anaerobic medium. (Item 51) 51. The method of any one of items 48 to 50, further comprising isolating or purifying the 3-HBal and / or 1,3-BDO or the 4-HBal and / or 1,4-BDO, or esters or amides thereof. (Item 52) 52. The method of claim 51, wherein the isolating or purifying step comprises distillation. (Item 53) 53. A culture medium comprising biogenic 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO, wherein the biogenic 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO has a carbon-12, carbon-13, and carbon-14 isotope ratio that reflects an atmospheric carbon dioxide uptake source, and wherein the biogenic 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO is produced by the cell of any one of items 23 to 39 or the method of any one of items 48 to 52. (Item 54) 54. The medium of item 53, separated from the cells. (Item 55) 53. 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) having carbon-12, carbon-13, and carbon-14 isotope ratios reflective of an atmospheric carbon dioxide uptake source, wherein 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO is produced by the cell of any one of items 23 to 39 or the method of any one of items 48 to 52. (Item 56) 56. 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO according to item 55, having an Fm value of at least 80%, at least 85%, at least 90%, at least 95%, or at least 98%. (Item 57) 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) produced by the cell of any one of items 23 to 39 or the method of any one of items 48 to 52. (Item 58) 53. 3-Hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) having carbon-12, carbon-13, and carbon-14 isotope ratios that reflect an atmospheric carbon dioxide uptake source, wherein the 3-HBal and / or 1,3-BDO is produced by the cell of any one of items 23 to 39 or the method of any one of items 48 to 52, and is enriched in the R-enantiomer. (Item 59) 59. 3-HBal and / or 1,3-BDO according to item 58, having an Fm value of at least 80%, at least 85%, at least 90%, at least 95%, or at least 98%. (Item 60) 53. 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) produced by the cell of any one of items 23 to 39 or the method of any one of items 48 to 52, wherein the 3-HBal and / or 1,3-BDO is enriched in the R-enantiomer. (Item 61) 61. 3-HBal and / or 1,3-BDO according to item 60, wherein the R-configuration is greater than 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9% of the 3-HBal and / or 1,3-BDO. (Item 62) 62. A composition comprising 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO according to any one of items 58 to 61, and a compound other than said 3-HBal and / or 1,3-BDO or 4-HBal or 1,4-BDO, respectively. (Item 63) 63. The composition of claim 62, wherein the compound other than 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO is part of a cell that produces the 3-HBal and / or 1,3-BDO or the 4-HBal and / or 1,4-BDO, respectively, or expresses the polypeptide of any one of claims 7 to 22. (Item 64) 62. A composition comprising 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO according to any one of items 55 to 61, or a cell lysate or culture supernatant of cells that produce said 3-HBal and / or 1,3-BDO or said 4-HBal and / or 1,4-BDO. (Item 65) 62. A product comprising 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO according to any one of items 55 to 61, wherein the product is a plastic, elastic fiber, polyurethane, polyester, polyhydroxyalkanoate, poly-4-hydroxybutyrate (P4HB) or a copolymer thereof, poly(tetramethylene ether) glycol (PTMEG), polybutylene terephthalate (PBT), polyurethane-polyurea copolymer, nylon, organic solvent, polyurethane resin, polyester resin, hypoglycemic agent, butadiene, or a butadiene-based product. (Item 66) 66. The product according to item 65, which is a cosmetic or a food additive. (Item 67) 67. The product of item 65 or 66, comprising at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% bio-sourced 3-HBal and / or 1,3-BDO or bio-sourced 4-HBal and / or 1,4-BDO. (Item 68) 68. The product of any one of items 65 to 67, comprising as a repeating unit a portion of produced 3-HBal and / or 1,3-BDO or produced 4-HBal and / or 1,4-BDO. (Item 69) 69. A molded product obtained by molding a product according to any one of items 65, 67 or 68. (Item 70) 69. A process for producing the product of any one of items 65 to 68, comprising chemically reacting the 3-HBal and / or 1,3-BDO or the 4-HBal and / or 1,4-BDO with itself or another compound in a reaction to produce the product. (Item 71) 70. A process for producing the product of claim 69, comprising chemically reacting the 3-HBal and / or 1,3-BDO or the 4-HBal and / or 1,4-BDO with itself or another compound in a reaction to produce the product. (Item 72) 23. A method for producing 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or an ester or amide thereof, the method comprising the steps of providing a substrate to the polypeptide of any one of items 7 to 22 and converting the substrate to 3-HBal and / or 1,3-BDO, wherein the substrate is a racemic mixture of 1,3-hydroxybutyryl-CoA. (Item 73) 73. The method of claim 72, wherein the 3-HBal and / or 1,3-BDO is enantiomerically enriched in the R form. (Item 74) 23. A method for producing 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) or an ester or amide thereof, the method comprising the steps of providing a substrate to the polypeptide of any one of items 7 to 22 and converting the substrate to 4-HBal and / or 1,4-BDO, wherein the substrate is 1,4-hydroxybutyryl-CoA. (Item 75) 75. The method of any one of items 72 to 74, wherein the polypeptide is present in a cell, in a cell lysate, or isolated from a cell or a cell lysate. (Item 76) 40. A method for producing 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO, comprising incubating a lysate of the cells of any one of paragraphs 23 to 39 to produce 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO. (Item 77) 77. The method of claim 76, wherein the cell lysate is mixed with a second cell lysate, the second cell lysate containing an enzyme activity that produces a substrate for the polypeptide of any one of claims 7 to 22, or a downstream product of 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO. (Item 78) 23. A method for producing a polypeptide according to any one of items 7 to 22, comprising expressing the polypeptide in a cell. (Item 79) 23. A method for producing a polypeptide according to any one of items 7 to 22, comprising in vitro transcription and translation of a nucleic acid according to any one of items 1 to 3 or a vector according to items 4 to 6 to produce said polypeptide. [Brief explanation of the drawings]

[0008] [Figure 1]Figure 1 shows an exemplary 1,3-butanediol (1,3-BDO) pathway comprising an aldehyde dehydrogenase. Figure 1 shows a pathway from acetoacetyl-CoA to 1,3-butanediol. The enzymes are: (A) acetoacetyl-CoA reductase (CoA-dependent, aldehyde forming); (B) 3-oxobutyraldehyde reductase (ketone reducing); (C) 3-hydroxybutyraldehyde reductase, also referred to herein as 1,3-butanediol dehydrogenase; (D) acetoacetyl-CoA reductase (CoA-dependent, alcohol forming); (E) 3-oxobutyraldehyde reductase (aldehyde reducing); (F) 4-hydroxy,2-butanone reductase; (G) acetoacetyl-CoA reductase (ketone reducing); (H) 3-hydroxybutyryl-CoA reductase, also referred to herein as 3-hydroxybutyraldehyde dehydrogenase (aldehyde forming); and (I) 3-hydroxybutyryl-CoA reductase (alcohol forming).

[0009] [Figure 2]2 shows an exemplary 1,4-butanediol (1,4-BDO) pathway that includes aldehyde dehydrogenases. The enzymes catalyzing the biosynthetic reactions are: (1) succinyl-CoA synthetase; (2) CoA-independent succinic semialdehyde dehydrogenase; (3) α-ketoglutarate dehydrogenase; (4) glutamate:succinic semialdehyde transaminase; (5) glutamate decarboxylase; (6) CoA-dependent succinic semialdehyde dehydrogenase; (7) 4-hydroxybutanoate dehydrogenase (also called 4-hydroxybutyrate dehydrogenase); and (8) α-ketoglutarate decarboxylase. (9) 4-hydroxybutyryl-CoA:acetyl-CoA transferase; (10) butyrate kinase (also called 4-hydroxybutyrate kinase); (11) phosphotransbutyrylase (also called phospho-trans-4-hydroxybutyrylase); (12) aldehyde dehydrogenase (also called 4-hydroxybutyryl-CoA reductase); and (13) alcohol dehydrogenase (also called 4-hydroxybutanal reductase or 4-hydroxybutyraldehyde reductase).

[0010] [Figure 3]Figure 3 shows a sequence alignment of ALD-1, ALD-2, and ALD-3. The sequences correspond to SEQ ID NOs: 1, 2, and 3, respectively. Underlined in the figure are two loop regions, the first designated A and the second designated B, both of which are responsible for substrate specificity and enantiomer specificity as determined herein. Loop A of ALD-1 has the sequence LQKNNETQEYSINKKWVGKD (SEQ ID NO: 124), loop A of ALD-2 has the sequence IGPKGAPDRKFVGKD (SEQ ID NO: 125), and loop A of ALD-3 has the sequence ITPKGLNRNCVGKD (SEQ ID NO: 126). Loop B of ALD-1 has the sequence SFAGVGYEAEGFTTFTIA (SEQ ID NO: 127), loop B of ALD-2 has the sequence TYCGTGVATNGAHSGASALTIA (SEQ ID NO: 128), and loop B of ALD-3 has the sequence SYAAIGFGGEGFCTFTIA (SEQ ID NO: 129). The sequences and lengths of the substrate specificity loops A and B from ALD-2 differ from those of ALD-1 and ALD-3; nevertheless, alignment shows sufficient conservation to facilitate identification of corresponding positions for substitutions as described herein, particularly when combined with 3D modeling as shown in FIG. 6. ALD-3 was used as a template for modeling the crystal structure; see particularly FIG. 6, which shows the two loop regions that interact to affect substrate and enantiomer specificity when modified with exemplary substitutions as described herein. ALD-1 and ALD-3 are 51.9% identical. ALD-1 and ALD-2 are 35.9% identical. ALD-3 and ALD-2 are 40% identical. The match for loop A based on the alignment of ALD-1, ALD-2, and ALD-3 is IXPKG-----XXNRKXVGKD (SEQ ID NO: 5). The match for loop B based on the alignment of ALD-1, ALD-2, and ALD-3 is SYAGXGXXXE----GFXTFTIA (SEQ ID NO: 6).It is understood that specifically identified amino acids in the consensus sequence are conserved residues, while positions marked with "X" are variable and can correspond to any amino acid, as desired and disclosed herein. It is further understood that "-----" can correspond to the presence or absence of a variable number of amino acid residues. Examples of such variable numbers of amino acid residues are shown in Figures 3 and 4A-4C. It is further understood that conserved residues in the consensus sequence can be substituted with conservative amino acids, e.g., as described herein (see, e.g., Figures 4A-4C).

[0011] [Figure 4A] Figures 4A-4C show alignments of exemplary aldehyde dehydrogenases (ALDs), identifying positions in ALDs corresponding to positions in a representative template ALD sequence where substitutions of the present invention can be made. As in Figure 3, underlined are two loop regions, the first designated A and the second designated B, both of which are involved in substrate specificity and enantiomer specificity as determined herein. Figure 4A shows an alignment of exemplary ALD sequences at a 40-55% cutoff compared to ALD-1. The sequences correspond to SEQ ID NOS: 1 (ALD-1), 13, 20, and 24 as shown in Figure 4A. Figure 4B shows an alignment of exemplary ALD sequences at a 75-90% cutoff compared to ALD-1. The sequences correspond to SEQ ID NOS: 1 (ALD-1), 30, 33, and 37 as shown in Figure 4B. Loops A and B are underlined. Figure 4C shows an alignment of exemplary ALD sequences at a 90% cutoff compared to ALD-1. The sequences correspond to SEQ ID NOs: 1 (ALD-1), 38, 40, and 44 as shown in Figure 4C. ALD-1 is 99%, 97%, and 95% identical to SEQ ID NOs: 38, 40, and 44, respectively. Figures 4A-4C show that corresponding positions for substitutions taught herein can be identified in ALDs with at least 40% identity to ALD-1, particularly in the loop A and B regions, especially the highly conserved loop B region. [Figure 4B] Same as above [Figure 4C] Same as above

[0012] [Figure 5A] Figures 5A and 5B show the enzymatic activity of various exemplary aldehyde dehydrogenases. Figure 5A shows the specific activity of ALD-2, ALD-1, and ALD-1 variants toward 3-hydroxy-(R)-butyraldehyde (left bar in the set of bars) and 3-hydroxy-(S)-butyraldehyde (right bar in the set of bars). Figure 5B shows the ratio of activity of the R form of 3-hydroxybutyraldehyde to the S form. [Figure 5B] Figures 5A and 5B show the enzymatic activity of various exemplary aldehyde dehydrogenases. Figure 5A shows the specific activity of ALD-2, ALD-1, and ALD-1 variants toward 3-hydroxy-(R)-butyraldehyde (left bar in the set of bars) and 3-hydroxy-(S)-butyraldehyde (right bar in the set of bars). Figure 5B shows the ratio of activity of the R form of 3-hydroxybutyraldehyde to the S form.

[0013] [Figure 6] Figures 6A-6C show ribbon diagrams of the structure of aldehyde dehydrogenase 959. The diagrams show the docking of 3-hydroxy-(R)-butyraldehyde (Figure 6A) or 3-hydroxy-(S)-butyraldehyde (Figure 6B) onto the structure of 959. Figure 6C shows the same orientation as 3-hydroxy-(R)-butyraldehyde (R3HB). DETAILED DESCRIPTION OF THE INVENTION

[0014] The present invention relates to enzyme variants that have desirable properties and are useful for producing desired products. In certain embodiments, the present invention relates to aldehyde dehydrogenase variants, which are enzyme variants that have significantly different structural and / or functional characteristics compared to naturally occurring wild-type enzymes. Thus, the aldehyde dehydrogenases of the present invention are not naturally occurring enzymes. Such aldehyde dehydrogenase variants of the present invention are useful in engineered cells, e.g., microbial organisms, engineered to produce desired products. For example, cells, e.g., microbial organisms, having metabolic pathways as disclosed herein can produce desired products. Aldehyde dehydrogenases of the present invention with desirable characteristics can be introduced into cells, e.g., microbial organisms, having metabolic pathways that use aldehyde dehydrogenase enzyme activity to produce desired products. Such aldehyde dehydrogenase variants are further useful as biocatalysts for performing desired reactions in vitro. Thus, the aldehyde dehydrogenase variants of the present invention can be utilized in engineered cells, such as microbial organisms, to produce desired products or as in vitro biocatalysts to produce desired products.

[0015] As used herein, the term "non-naturally occurring" when used with respect to a cell, microbial organism, or microorganism of the present invention means that the cell has at least one genetic change that is not normally found in natural strains of the reference species, including wild-type strains of the reference species. Genetic changes include, for example, modifications that introduce expressible nucleic acids encoding metabolic polypeptides, additions of other nucleic acids, nucleic acid deletions, and / or other functional disruptions of cellular genetic material. Such modifications include, for example, coding regions and functional fragments of heterologous, homologous, or heterologous and homologous polypeptides of the reference species. Additional modifications include, for example, non-coding regulatory regions that alter the expression of genes or operons. Exemplary metabolic polypeptides include enzymes or proteins in a biosynthetic pathway to produce a desired product.

[0016] A metabolic modification refers to a biochemical reaction that is altered from its native state. Thus, a non-naturally occurring cell can have a genetic modification to a nucleic acid encoding a metabolic polypeptide or a functional fragment thereof. Exemplary metabolic modifications are disclosed herein.

[0017] As used herein, the term "isolated," when used with respect to a cell or microbial organism, is intended to mean a cell that is substantially free of at least one component that the referenced cell would be found in if such cell were found in nature. The term includes cells from which some or all components found in their natural environment have been removed. The term also includes cells from which some or all components found in a non-naturally occurring environment have been removed. Thus, an isolated cell is partially or completely separated from other materials in which it is found in nature or in which it is grown, stored, or subsisted in a non-naturally occurring environment. Examples of isolated cells include partially pure cells, substantially pure cells, and cells cultured in a non-naturally occurring medium.

[0018] As used herein, the terms "microbial," "microbial organism," or "microorganism" are intended to mean any organism existing as a microscopic cell, such as those in the domains Archaea, Bacteria, or Eukarya. The term therefore encompasses prokaryotic or eukaryotic cells or organisms of microscopic size, including all species of bacteria, archaea, and eubacteria, as well as eukaryotic microorganisms such as yeast and fungi. The term also includes cell cultures of any species that can be cultured for the production of biochemicals.

[0019] As used herein, the term "CoA" or "coenzyme A" is intended to mean an organic cofactor or prosthetic group (the non-protein portion of an enzyme) whose presence is required for the activity of many enzymes (apoenzymes) to form an active enzyme system. Coenzyme A functions in certain condensing enzymes and plays a role in acetyl or other acyl group transfer, as well as fatty acid synthesis and oxidation, pyruvate oxidation, and other acetylations.

[0020] As used herein, the term "substantially anaerobic" when used in reference to culture or growth conditions is intended to mean that the amount of oxygen is less than about 10% of the saturation level of dissolved oxygen in the liquid medium. The term is also intended to include sealed chambers of liquid or solid medium maintained in an atmosphere of less than about 1% oxygen.

[0021] As used herein, "exogenous" means that a reference molecule or activity is introduced into a host cell. The molecule can be introduced, for example, by introduction of the encoding nucleic acid into the host genetic material, such as by integration into a host chromosome, or as non-chromosomal genetic material, such as a plasmid. Thus, when used with respect to expression of an encoding nucleic acid, the term refers to the introduction of the encoding nucleic acid into a cell in an expressible form. When used with respect to a biosynthetic activity, the term refers to an activity that is introduced into a host reference organism. The source can be, for example, a homologous or heterologous encoding nucleic acid that expresses the reference activity after introduction into a host microbial organism. Thus, the term "endogenous" refers to a reference molecule or activity that is present in a host. Similarly, when used with respect to expression of an encoding nucleic acid, the term refers to expression of an encoding nucleic acid contained in a cell. The term "heterologous" refers to a molecule or activity derived from a source other than the reference species, and "homologous" refers to a molecule or activity derived from a host cell. Thus, exogenous expression of an encoding nucleic acid of the present invention can utilize either or both heterologous or homologous encoding nucleic acids.

[0022] When two or more exogenous nucleic acids are contained in a microbial organism, it is understood that these two or more exogenous nucleic acids refer to the coding nucleic acid or biosynthetic activity referenced above. As disclosed herein, it is further understood that such two or more exogenous nucleic acids can be introduced into a host cell on separate nucleic acid molecules, on a polycistronic nucleic acid molecule, or in combination, and still be considered two or more exogenous nucleic acids. For example, as disclosed herein, a cell can be engineered to express two or more exogenous nucleic acids encoding desired enzymes or proteins, such as pathway enzymes or proteins. When two exogenous nucleic acids encoding desired activities are introduced into a host cell, it is understood that these two exogenous nucleic acids can be introduced as a single nucleic acid, for example, on a single plasmid or on separate plasmids, and can also be integrated into a single site or multiple sites in the host chromosome, and still be considered two exogenous nucleic acids. Similarly, it is understood that three or more exogenous nucleic acids can be introduced into a host organism in any desired combination, for example, on a single plasmid or on separate plasmids, and can be integrated into a single site or multiple sites on a host chromosome, and still be considered two or more exogenous nucleic acids, for example, three exogenous nucleic acids. Thus, the number of exogenous nucleic acids or biosynthetic activities referenced refers to the number of encoding nucleic acids or biosynthetic activities, rather than the number of separate nucleic acids introduced into the host organism.

[0023] As used herein, the term "gene disruption" or its grammatical equivalents refers to a genetic change that inactivates or attenuates the encoded gene product. Genetic changes can be, for example, the deletion of the entire gene, the deletion of regulatory sequences required for transcription or translation, the deletion of a portion of a gene resulting in a truncated gene product, or various mutation strategies that inactivate or attenuate the encoded gene product. Complete gene disruption is a particularly useful method of gene deletion, as it reduces or eliminates the occurrence of gene reversion in the non-naturally occurring cells of the present invention. Gene disruption also includes null mutations, which refer to mutations within a gene or a region containing a gene that result in the gene not being transcribed into RNA and / or translated into a functional gene product. Such null mutations can result from many types of mutations, including, for example, inactivating point mutations, deletion of a portion of a gene, deletion of the entire gene, or deletion of a chromosome segment.

[0024] As used herein, the term "growth-coupled," when used in reference to the production of a biochemical product, is intended to mean that the biosynthesis of the referenced biochemical product is produced during the growth phase of the microorganism. In certain embodiments, growth-coupled production may be essential, meaning that the biosynthesis of the referenced biochemical is an essential product produced during the growth phase of the microorganism.

[0025] As used herein, the term "attenuate" or its grammatical equivalents means weakening, lowering, or reducing the activity or amount of an enzyme or protein. Attenuation of an enzyme or protein's activity or amount can mimic complete destruction if the attenuation causes the activity or amount to fall below a critical level required for a given function. However, complete destruction, e.g., attenuation of an enzyme or protein's activity or amount that mimics complete destruction of one pathway, may still be sufficient for another pathway to continue functioning. For example, attenuation of an endogenous enzyme or protein may be sufficient to mimic the complete destruction of the same enzyme or protein for the production of a desired product of the invention, but the remaining activity or amount of the enzyme or protein may still be sufficient to maintain other pathways, such as pathways important for the survival, reproduction, or growth of a host cell. Attenuation of an enzyme or protein may also be attenuating, lowering, or reducing the activity or amount of the enzyme or protein enough to increase the yield of a desired product of the invention, but does not necessarily mimic complete destruction of the enzyme or protein.

[0026] The non-naturally occurring cells of the present invention can contain stable genetic alterations, which refers to cells that can be cultured for more than five generations without losing the alteration. Generally, stable genetic alterations include alterations that persist for more than 10 generations, particularly stable alterations that persist for more than about 25 generations, and more particularly stable genetic alterations that persist for more than 50 generations, including indefinitely.

[0027] In the case of gene disruption, gene deletion is a particularly useful stable genetic change.The use of gene deletion to introduce stable genetic changes is particularly useful for reducing the possibility of reversion to the phenotype before genetic change.For example, the stable production of growth-linked biochemicals can be achieved by, for example, deleting the gene encoding the enzyme that catalyzes one or more reactions in a set of metabolic modifications.The stability of the production of growth-linked biochemicals can be significantly reduced and even enhanced through multiple deletions, which significantly reduces the possibility of multiple compensatory reversions occurring for each disrupted activity.

[0028] Those skilled in the art will understand that the genetic changes, including metabolic modifications, exemplified herein are described in terms of suitable host cells or organisms, such as E. coli, and their corresponding metabolic reactions, or suitable source cells or organisms for the desired genetic material, e.g., genes of a desired metabolic pathway. However, given the complete genome sequencing of a wide variety of organisms and the high level of technology in the field of genomics, those skilled in the art can easily apply the teachings and guidance provided herein to virtually any other organism. For example, the metabolic modifications exemplified herein for E. coli can be easily applied to other species by incorporating the same or similar encoding nucleic acid from a species other than the reference species. Such genetic changes include, for example, genetic changes of species homologs in general, and more specifically, orthologous, paralogous, or non-orthologous gene replacements.

[0029] Orthologs are one or more genes that have been vertically transferred and that fulfill substantially the same or identical functions in different organisms. For example, mouse epoxide hydrolase and human epoxide hydrolase can be considered orthologs for the biological function of epoxide hydrolysis. Genes are vertically transferred if they share sufficient sequence similarity to indicate, for example, that they are homologous or related by evolution from a common ancestor. Genes can also be considered orthologs if they share sufficient three-dimensional structure, but not necessarily sequence similarity, to indicate that they evolved from a common ancestor to the extent that primary sequence similarity is not discernible. Orthologous genes can encode proteins with sequence similarity of approximately 25% to 100% amino acid sequence identity. Genes encoding proteins that share less than 25% amino acid similarity can also be considered to have arisen by vertical transfer if their three-dimensional structures also show similarity. Members of the serine protease family of enzymes, including tissue plasminogen activator and elastase, are considered to have arisen by vertical transfer from a common ancestor.

[0030] Orthologs include genes or their encoded gene products that differ in structure or overall activity, for example, through evolution. For example, if one species encodes a gene product that exhibits two functions and those functions are segregated into different genes in a second species, the three genes and their corresponding products are considered orthologous. For the production of biochemical products, those skilled in the art will understand that orthologous genes harboring the metabolic activity to be introduced or disrupted should be selected for construction of a non-naturally occurring cell. An example of an ortholog exhibiting segregable activities is when different activities are segregated into different gene products between two or more species or within a single species. A specific example is the separation of two types of serine protease activity, elastase proteolysis and plasminogen proteolysis, into different molecules such as plasminogen activator and elastase. A second example is the separation of mycoplasma 5'-3' exonuclease and Drosophila DNA polymerase III activity. A DNA polymerase from a first species can be considered to be orthologous to either or both of the exonuclease or polymerase from a second species, or vice versa.

[0031] In contrast, paralogs are homologs related, for example, by duplication and subsequent evolutionary divergence, and have similar or shared, but non-identical, functions. Paralogs can originate from or be derived from the same or different species. For example, microsomal epoxide hydrolase (epoxide hydrolase I) and soluble epoxide hydrolase (epoxide hydrolase II) can be considered paralogs because they represent two distinct enzymes that catalyze different reactions and have different functions in the same species, coevolving from a common ancestor. Paralogs are proteins from the same species that share significant sequence similarity with each other, suggesting they are homologous or related through coevolution from a common ancestor. Examples of paralogous protein families include HipA homologs, luciferase genes, and peptidases.

[0032] Nonorthologous gene substitution is a nonorthologous gene from one species that can replace the function of a reference gene in a different species. Substitutions include, for example, those that can perform substantially the same or similar functions in the source species compared to the reference function in a different species. Generally, nonorthologous gene substitution can be confirmed to be structurally related to a known gene encoding the reference function, but less structurally related but functionally similar genes and their corresponding gene products also fall within the meaning of the term as used herein. For example, functional similarity requires at least some structural similarity in the active site or binding region of the nonorthologous gene product compared to the gene encoding the function to be replaced. Thus, nonorthologous genes include, for example, paralogs or unrelated genes.

[0033] Thus, by applying the teachings and guidance provided herein to a particular species in identifying and constructing a non-naturally occurring cell of the invention capable of biosynthesizing a desired product, one skilled in the art will understand that identifying metabolic modifications can include identifying and incorporating or inactivating orthologs. To the extent that paralogs and / or non-orthologous gene replacements exist in the reference cell that encode enzymes that catalyze similar or substantially similar metabolic reactions, one skilled in the art can also utilize these evolutionarily related genes. As with gene disruption, evolutionarily related genes can also be disrupted or deleted in a host cell to reduce or eliminate functional redundancy of the enzyme activity targeted for disruption.

[0034] Orthologs, paralogs, and nonorthologous gene substitutions can be determined by methods well known to those skilled in the art. For example, examination of the nucleic acid or amino acid sequences of two polypeptides reveals sequence identity and similarity between the compared sequences. Based on such similarity, those skilled in the art can determine whether the similarity is high enough to indicate that the proteins are related through evolution from a common ancestor. Algorithms well known to those skilled in the art, such as Align, BLAST, Clustal W, etc., compare and determine the similarity or identity of raw sequences, and further determine the presence or significance of gaps in the sequences, which can be assigned weights or scores. Such algorithms are also known in the art and can be similarly applied to determining the similarity or identity of nucleotide sequences. The parameter of similarity sufficient to determine relatedness is calculated based on statistical similarity, or the probability of finding a similar match in a random polypeptide, and well-known methods for calculating the significance of the determined match. If desired, computer-based comparison of two or more sequences can also be visually optimized by those skilled in the art. Related gene products or proteins can be expected to have high similarity, e.g., 25% to 100% sequence identity. Unrelated proteins can have identities that are virtually identical to what would be expected by chance when scanning a database of sufficient size (approximately 5%). Sequences between 5% and 24% may or may not exhibit sufficient homology to conclude that the compared sequences are related. To determine the relevance of these sequences, additional statistical analyses can be performed to determine the significance of such matches given the size of the dataset.

[0035] For example, exemplary parameters for determining the relatedness of two or more sequences using the BLAST algorithm can be as follows: Briefly, amino acid sequence alignment can be performed using BLASTP version 2.0.8 (January 5, 1999) and the following parameters: matrix: 0 BLOSUM62; gap open: 11; gap extension: 1; x_dropoff: 50; expectation: 10.0; word size: 3; filter: on. Nucleic acid sequence alignment can be performed using BLASTN version 2.0.6 (September 16, 1998) and the following parameters: match: 1; mismatch: -2; gap open: 5; gap extension: 2; x_dropoff: 50; expectation: 10.0; word size: 11; filter: off. Those skilled in the art will know what modifications can be made to the above parameters, for example, to increase or decrease the stringency of the comparison and determine the relatedness of two or more sequences.

[0036] In one embodiment, the present invention provides aldehyde dehydrogenases that are variants of wild-type or parent aldehyde dehydrogenases. The aldehyde dehydrogenases of the present invention convert acyl-CoA to its corresponding aldehyde. Such enzymes may also be referred to as oxidoreductases that convert acyl-CoA to its corresponding aldehyde. Such aldehyde dehydrogenases of the present invention can be classified as oxidoreductases of reaction 1.2.1.b (acyl-CoA to aldehyde), where the first three digits correspond to the first three digits of the enzyme number, which indicates a general type of conversion independent of substrate specificity. Exemplary enzymatic conversions of the aldehyde dehydrogenases of the present invention include, but are not limited to, the conversion of 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde (also known as 3-HBal) (see Figure 1) and the conversion of 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde (see Figure 2). The aldehyde dehydrogenases of the invention can be used in cells, e.g., microbial organisms, containing suitable metabolic pathways or in vitro to produce downstream products, including desired products such as 3-hydroxybutyraldehyde (3-HBal), 1,3-butanediol (1,3-BDO), 4-hydroxybutyraldehyde (4-HBal), 1,4-butanediol (1,4-BDO), or other desired products, e.g., esters or amides thereof. For example, 1,3-BDO can be converted to an ester by reacting it with an acid, e.g., using a lipase, either in vivo or in vitro.Such esters may have nutritional supplement, pharmaceutical, and food applications, and are advantageous when the R-form of 1,3-butanediol is used because it is the form most commonly utilized by both animals and humans as an energy source (compared to the S-form or racemic mixtures produced from petroleum or ethanol via the acetaldehyde chemical synthesis route) (e.g., ketone esters, such as (R)-3-hydroxybutyl-R-1,3-butanediol monoester and (R)-3-hydroxybutyric acid glycerol monoester or diester (which have Generally Recognized As Safe (GRAS) approval in the United States). Ketone esters can be delivered orally, and the esters release R-1,3-butanediol for use by the body (see, e.g., WO2013150153). Thus, the present invention is particularly useful for providing improved enzymatic pathways and microorganisms to provide improved compositions of 1,3-butanediol, i.e., R-1,3-butanediol, that are highly enriched or essentially enantiomerically pure, and that further have improved purity qualities for by-products.

[0037] 1,3-butanediol, also known as butylene glycol, has additional food-related uses, including as a direct food source, food ingredient, flavoring, or solubilizer for flavorings, stabilizer, emulsifier, and antimicrobial and preservative. 1,3-butanediol is used in the pharmaceutical industry as a parenteral drug solvent. 1,3-butanediol finds use in cosmetics as an emollient, a humectant, a component that prevents crystallization of insoluble ingredients, a solubilizer for less water-soluble ingredients such as fragrances, and an antimicrobial and preservative. For example, it can be used as a humectant, especially in hairsprays and setting lotions; it reduces fragrance loss from essential oils, preserves against microbial spoilage, and is used as a solvent for benzoates. 1,3-butanediol can be used in concentrations from 0.1 percent or less to 50 percent or more. It is used in hair and bath products, eye and facial cosmetics, fragrances, body cleansing products, and shaving and skin care preparations (see, e.g., the Cosmetic Ingredient Review board report: "Final Report on the Safety Assessment of Butylene Glycol, Hexylene Glycol, Ethoxydiglycol, and Dipropylene Glycol," Journal of the American College of Toxicology, Vol. 4, No. 5, 1985, incorporated herein by reference). This report provides specific uses and concentrations of 1,3-butanediol (butylene glycol) in cosmetics; see, e.g., Table 2 of the report therein entitled "Product Formulation Data."

[0038] In one embodiment, the invention provides an isolated nucleic acid molecule selected from: (a) a nucleic acid molecule encoding an amino acid sequence referred to as SEQ ID NO: 1, 2 or 3 or in Table 4, wherein the amino acid sequence comprises one or more of the amino acid substitutions set forth in Tables 1, 2 and / or 3; (b) a nucleic acid molecule that hybridizes to the nucleic acid of (a) under highly stringent hybridization conditions and comprises a nucleic acid sequence encoding one or more of the amino acid substitutions set forth in Tables 1, 2 and / or 3; (c) a nucleic acid molecule encoding an amino acid sequence comprising the consensus sequence of Loop A (SEQ ID NO: 5) and / or Loop B (SEQ ID NO: 6), wherein the amino acid sequence comprises one or more of the amino acid substitutions set forth in Tables 1, 2 and / or 3; and (d) a nucleic acid molecule complementary to (a) or (b). In one embodiment, the amino acid sequence encoded by the nucleic acid molecule, other than the one or more amino acid substitutions, has at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98% or 99% sequence identity to or is identical to an amino acid sequence referenced in SEQ ID NO: 1, 2 or 3 or in Table 4. The amino acid sequence may include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or 16 or more, e.g., 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42 or 43, of the amino acid substitutions set forth in Tables 1, 2 and / or 3, i.e., up to all of the amino acid positions having substitutions.

[0039] The present invention also provides a vector containing the nucleic acid molecule of the present invention. In one embodiment, the vector is an expression vector. In one embodiment, the vector comprises double-stranded DNA.

[0040] The present invention also provides nucleic acids encoding the aldehyde dehydrogenase polypeptides of the present invention. Nucleic acid molecules encoding the aldehyde dehydrogenases of the present invention can also include nucleic acid molecules that hybridize to the nucleic acids disclosed herein by SEQ ID NOs, GenBank, and / or GI numbers, or nucleic acid molecules that hybridize to nucleic acid molecules that encode the amino acid sequences disclosed herein by SEQ ID NOs, GenBank, and / or GI numbers. Hybridization conditions can include highly stringent, moderately stringent, or low stringency hybridization conditions well known to those skilled in the art, such as those described herein. Similarly, nucleic acid molecules that can be used in the present invention can be described as having a certain percent sequence identity to the nucleic acids disclosed herein by SEQ ID NOs, GenBank, and / or GI numbers, or nucleic acid molecules that hybridize to the nucleic acids that encode the amino acid sequences disclosed herein by SEQ ID NOs, GenBank, and / or GI numbers. For example, a nucleic acid molecule can have at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to or be identical to a nucleic acid described herein.

[0041] Stringent hybridization refers to conditions under which the hybridized polynucleotides are stable. As known to those skilled in the art, the stability of the hybridized polynucleotides is determined by the melting temperature (T m) Generally, the stability of hybridized polynucleotides is a function of salt concentration, e.g., sodium ion concentration, and temperature. Hybridization reactions can be performed under conditions of lower stringency, followed by washing at various, but higher, stringency levels. References to hybridization stringency refer to such washing conditions. As contemplated herein, highly stringent hybridization includes conditions that allow hybridization of only nucleic acid sequences that form stable hybridized polynucleotides at 65°C in 0.018M NaCl; for example, if a hybrid is not stable at 65°C in 0.018M NaCl, it is not stable under high stringency conditions. High stringency conditions can be provided, for example, by hybridization in 50% formamide, 5x Denhart's solution, 5x SSPE, 0.2% SDS at 42°C, followed by washing in 0.1x SSPE and 0.1% SDS at 65°C. Hybridization conditions other than highly stringent hybridization conditions can also be used to describe the nucleic acid sequences disclosed herein. For example, the phrase "moderately stringent hybridization" refers to conditions equivalent to hybridization in 50% formamide, 5x Denhart's solution, 5x SSPE, 0.2% SDS at 42°C, followed by washing in 0.2x SSPE, 0.2% SDS at 42°C. The phrase "low stringency hybridization" refers to conditions equivalent to hybridization in 10% formamide, 5x Denhardt's solution, 6x SSPE, 0.2% SDS at 22°C, followed by washing in 1x SSPE, 0.2% SDS at 37°C. Denhardt's solution contains 1% Ficoll, 1% polyvinylpyrrolidone, and 1% bovine serum albumin (BSA). 20x SSPE (sodium chloride, sodium phosphate, ethylenediaminetetraacetic acid (EDTA)) contains 3M sodium chloride, 0.2M sodium phosphate, and 0.025M EDTA.Other suitable low, moderate, and high stringency hybridization buffers and conditions are well known to those of skill in the art and are described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory, New York (2001); and Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, MD (1999).

[0042] Nucleic acid molecules encoding the aldehyde dehydrogenases of the present invention may have at least a certain sequence identity with the nucleotide sequences disclosed herein. Thus, in some embodiments of the present invention, nucleic acid molecules encoding the aldehyde dehydrogenases of the present invention have or are identical to a nucleotide sequence that is at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to a nucleic acid disclosed herein by SEQ ID NO:, GenBank, and / or GI number, or a nucleic acid molecule that hybridizes to a nucleic acid molecule encoding an amino acid sequence disclosed herein by SEQ ID NO:, GenBank, and / or GI number.

[0043] Sequence identity (also known as homology or similarity) refers to the sequence similarity between two nucleic acid molecules or two polypeptides. Identity can be determined by comparing positions in each sequence that can be aligned for comparison purposes. If a position in the compared sequences is occupied by the same base or amino acid, the molecules are identical at that position. The degree of identity between sequences is a function of the number of matching or homologous positions shared by the sequences. Alignment of two sequences to determine their percent sequence identity can be performed using software programs known in the art, such as the software program described in Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, MD (1999). Preferably, default parameters are used for alignment. One alignment program known in the art that can be used is BLAST, set to default parameters. In particular, the programs are BLASTN and BLASTP, using the following default parameters: Genetic coding = standard; Filter = none; Strand = both; Cutoff = 60; Expected = 10; Matrix = BLOSUM62; Description = 50 sequences; Classification = HIGH SCORE; Databases = non-redundant, GenBank + EMBL + DDBJ + PDB + GenBank CDS translations + SwissProtein + SPupdate + PIR. Details of these programs can be found at the National Center for Biotechnology Information (see also Altschul et al., J. Mol. Biol. 215:403-410 (1990)).

[0044] In some embodiments, the nucleic acid molecule is an isolated nucleic acid molecule. In some embodiments, the isolated nucleic acid molecule is a nucleic acid molecule encoding a variant of a reference polypeptide, wherein (i) the reference polypeptide has the amino acid sequence of SEQ ID NO: 1, 2, or 3 or a sequence in Table 4 (SEQ ID NOS: 7-123), (ii) the variant comprises one or more amino acid substitutions compared to the sequence in SEQ ID NO: 1, 2, or 3 or Table 4, and (iii) the one or more amino acid substitutions are selected from the amino acid substitutions set forth in Tables 1-3. Tables 1-3 provide a non-limiting list of exemplary variants of the sequences in SEQ ID NO: 1, 2, or 3 or Table 4. In one embodiment, for each variant in Tables 1-3, all positions except those indicated are identical to the sequence in SEQ ID NO: 1, 2, or 3 or Table 4. Amino acid substitutions are indicated by a letter indicating the identity of the original amino acid, followed by a number indicating the position of the substituted amino acid in the sequence in SEQ ID NO: 1, 2, or 3 or Table 4, followed by a letter indicating the identity of the substituted amino acid. For example, "D12A" indicates that the aspartic acid at position 12 of SEQ ID NO: 1 or 2 has been replaced with an alanine. The one-letter codes used to identify amino acids are standard codes known to those of skill in the art. Some variants in Tables 1-3 contain two or more substitutions, as indicated by the list of substitutions. One or more amino acid substitutions can be selected from any one of the variants listed in Tables 1-3, or from any combination of two or more variants listed in Tables 1-3. When selecting from a single variant in Tables 1-3, the resulting variant can contain one or more of the substitutions of the selected variant in any combination, including all or fewer than all of the indicated substitutions. When substitutions are selected from the substitutions of two or more variants in Tables 1-3, the resulting variant can contain one or more of the substitutions of the selected variant, including all or fewer than all of the indicated substitutions, from each of the two or more selected variants in any combination.For example, the resulting variant may include 1, 2, 3, or 4 substitutions from a single variant of Tables 1-3. As a further example, the resulting variant may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 20, 25, or more substitutions selected from 1, 2, 3, 4, 5, or more selected variants of Tables 1-3. In some embodiments, the resulting variant includes all of the indicated substitutions of a selected variant of Tables 1-3. In some embodiments, the resulting variant differs from the sequence of SEQ ID NO: 1, 2, or 3 or Table 4 by at least one amino acid substitution but by fewer than 25, 20, 10, 5, 4, or 3 amino acid substitutions. In some embodiments, the resulting variant comprises, consists essentially of, or consists of a sequence indicated by a variant selected from Tables 1-3 and differs from the sequence of SEQ ID NO: 1, 2, or 3 or Table 4 only at the indicated amino acid substitutions.

[0045] In some embodiments, the nucleic acid molecule is an isolated nucleic acid molecule encoding a variant of a reference polypeptide (a reference polypeptide having an amino acid sequence of SEQ ID NO: 1, 2, or 3 or a sequence in Table 4), wherein the variant (i) comprises one or more amino acid substitutions of a corresponding variant selected from Tables 1-3 and (ii) has at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to the corresponding variant. When a second variant has 100% sequence identity with the corresponding variant, the second variant comprises the sequence set forth by a variant selected from Tables 1-3, with or without one or more additional amino acids at either or both the amino and carboxy termini. In some embodiments, the resulting variant has at least 80%, 85%, 90%, or 95% sequence identity to the corresponding variant selected from Tables 1-3; in some cases, the identity is at least 90% or higher. If the resulting variant is less than 100% identical to a corresponding variant selected from Tables 1-3, one or more positions of the amino acid substitutions indicated for the corresponding variant may be shifted (e.g., in the case of insertion or deletion of one or more amino acids) but still be contained within the resulting variant. For example, an aspartic acid to alanine substitution corresponding to "D12A" (at position 12 relative to SEQ ID NO: 1 or 2) may be present in the resulting variant, but at a different position. Whether an amino acid corresponds to a indicated substitution despite a different position can be determined by sequence alignment, as is well known in the art. Generally, an alignment showing the identity or similarity of amino acids adjacent to the amino acid to be substituted, such as when a contiguous sequence is aligned with a homologous sequence of another polypeptide, allows the substituted amino acid to be locally located with respect to the corresponding variant in Tables 1-3 and the corresponding position at which the substitution should be made, despite the shifted numerical position in a given polypeptide chain.In one embodiment, a region comprising at least 3-15 amino acids that includes the substituted position locally aligns with the corresponding variant sequence at a relatively high percent identity (e.g., 90%, 95%, or 100% identity), including the position of the substituted amino acid along the corresponding variant sequence. In some embodiments, one or more amino acid substitutions (e.g., all or fewer than all amino acid substitutions) exhibited by a corresponding variant selected from Tables 1-3 are considered to be present in a given variant, even if they occur at different physical positions along the polypeptide chain, if the sequence of the compared polypeptide is aligned with a corresponding variant that has an identical match or a similar amino acid at the indicated position along the corresponding variant sequence using the BLASTP alignment algorithm with default parameters, where a similar amino acid is considered to have sufficient chemical properties to align with the variant position of interest using the default parameters of the alignment algorithm.

[0046] In some embodiments, the nucleic acid molecules of the invention are complementary to the nucleic acids described in connection with any of the various embodiments herein.

[0047] It is understood that the nucleic acids of the present invention or polypeptides of the present invention may exclude a parent sequence, such as a wild-type parent sequence, e.g., a parent sequence such as SEQ ID NO: 1, 2, or 3 or a sequence disclosed in Table 4. Those of skill in the art will readily understand the meaning of a parent wild-type sequence based on what is known in the art. It is further understood that such nucleic acids of the present invention may exclude a nucleic acid sequence that encodes a naturally occurring amino acid sequence as found in nature. Similarly, a polypeptide of the present invention may exclude an amino acid sequence as found in nature. Thus, in certain embodiments, a nucleic acid or polypeptide of the present invention is as set forth herein, provided that the encoded amino acid sequence is not a wild-type parent sequence or a naturally occurring amino acid sequence and / or the nucleic acid sequence is not a wild-type or naturally occurring nucleic acid sequence. It is understood by those of skill in the art that a naturally occurring amino acid or nucleic acid sequence refers to a sequence found in a naturally occurring organism as found in nature. Thus, nucleic acids or amino acid sequences that are not found in the same state as in a naturally occurring organism or that do not have the same nucleotide or encoded amino acid sequence as in a naturally occurring organism are included within the meaning of the nucleic acids and / or amino acid sequences of the present invention. For example, nucleic acid or amino acid sequences that are altered at one or more nucleotide or amino acid positions from a parent sequence, including variants described herein, are included within the meaning of non-naturally occurring nucleic acid or amino acid sequences of the invention. The isolated nucleic acid molecules of the invention exclude naturally occurring chromosomes that contain the nucleic acid sequence, and may further exclude other molecules found in naturally occurring cells, such as DNA-binding proteins, e.g., proteins such as histones that bind to chromosomes in eukaryotic cells.

[0048] Thus, the isolated nucleic acid sequences of the present invention have physical and chemical differences compared to naturally occurring nucleic acid sequences. The isolated or non-naturally occurring nucleic acids of the present invention do not contain, or do not necessarily have, some or all of the chemical bonds, either covalent or non-covalent, of naturally occurring nucleic acid sequences as found in nature. The isolated nucleic acids of the present invention therefore differ from naturally occurring nucleic acids, for example, by having a different chemical structure from naturally occurring nucleic acid sequences found in chromosomes. The different chemical structure can occur, for example, by cleavage of phosphodiester bonds, releasing the isolated nucleic acid sequence from the naturally occurring chromosome. The isolated nucleic acids of the present invention can also differ from naturally occurring nucleic acids by isolating or separating the nucleic acid from proteins that bind to chromosomal DNA in either prokaryotic or eukaryotic cells, thereby differing from naturally occurring nucleic acids by different non-covalent bonds. For nucleic acids of prokaryotic origin, the non-naturally occurring nucleic acids of the present invention do not necessarily have some or all of the naturally occurring chemical bonds of chromosomes, for example, binding to DNA-binding proteins such as polymerases or chromosomal structural proteins, or are not present in a higher-order structure such as supercoiled. For nucleic acids of eukaryotic origin, the non-naturally occurring nucleic acids of the present invention also do not contain the same internal nucleic acid chemical bonds found in chromatin, nor do they contain chemical bonds with structural proteins. For example, the non-naturally occurring nucleic acids of the present invention are not chemically bound to histones or scaffolding proteins, and are not contained in centromeres or telomeres. Thus, the non-naturally occurring nucleic acids of the present invention are chemically distinct from naturally occurring nucleic acids because they lack or contain van der Waals interactions, hydrogen bonds, ionic or electrostatic bonds, and / or covalent bonds that differ from those found in nature. Such bond differences may occur internally within separate regions of the nucleic acid (i.e., in cis), or they may occur in trans interactions, for example, with chromosomal proteins. In the case of nucleic acids of eukaryotic origin, the chemical bonds within the cDNA differ from the covalent bonds, i.e., the sequence, of genes on chromosomal DNA, and therefore the cDNA is considered an isolated or non-naturally occurring nucleic acid.Thus, it will be understood by those skilled in the art that isolated or non-naturally occurring nucleic acids are distinct from naturally occurring nucleic acids.

[0049] In one embodiment, the invention provides an isolated polypeptide comprising an amino acid sequence referred to as SEQ ID NO: 1, 2 or 3 or in Table 4, wherein the amino acid sequence comprises one or more of the amino acid substitutions set forth in Tables 1, 2 and / or 3. In one embodiment, the invention provides an isolated polypeptide comprising the consensus amino acid sequence of Loop A (SEQ ID NO: 5) and / or Loop B (SEQ ID NO: 6).

[0050] In another embodiment, the invention provides an isolated polypeptide comprising an amino acid sequence referenced as SEQ ID NO: 1, 2 or 3 or in Table 4, wherein the amino acid sequence comprises one or more of the amino acid substitutions set forth in Table 1, 2 and / or 3, and the amino acid sequence other than the one or more amino acid substitutions has at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98% or 99% sequence identity to or is identical to the amino acid sequence referenced as SEQ ID NO: 1, 2 or 3 or in Table 4. In one embodiment, the amino acid sequence further comprises conservative amino acid substitutions at 1 to 100 amino acid positions, the positions being other than one or more amino acid substitutions set forth in Table 1, 2 and / or 3. In another embodiment, the amino acid sequence does not comprise an alteration at 2 to 300 amino acid positions compared to the parent sequence other than one or more amino acid substitutions set forth in Table 1, 2 and / or 3, the positions being selected from positions identical to 2, 3, 4 or 5 of the amino acid sequence referenced as SEQ ID NO: 1, 2 or 3 or in Table 4. In one embodiment, the amino acid sequence includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 or more, for example, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, of the amino acid substitutions set forth in Tables 1, 2, and / or 3, i.e., up to all of the amino acid positions having a substitution.

[0051] In one embodiment, the polypeptide of the present invention encodes an aldehyde dehydrogenase. In one embodiment, the polypeptide can convert 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde. In one embodiment, the polypeptide can convert 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde. In one embodiment, the polypeptide has higher activity than the parent polypeptide. In one embodiment, the polypeptide has higher activity for 3-hydroxy-(R)-butyryl-CoA than for 3-hydroxy-(S)-butyryl-CoA. In one embodiment, the polypeptide has higher specificity for 3-hydroxybutyryl-CoA than for acetyl-CoA. In one embodiment, the polypeptide has higher specificity for 4-hydroxybutyryl-CoA than for acetyl-CoA. In one embodiment, the polypeptide reduces the production of by-products in a cell or cell extract. In certain embodiments, the by-product is ethanol or 4-hydroxy-2-butanone. In one embodiment, the polypeptide has a higher kcat than the parent polypeptide.

[0052] In some embodiments, the invention provides isolated polypeptides having an amino acid sequence disclosed herein, e.g., a sequence referenced in SEQ ID NO: 1, 2, or 3, or Table 4, wherein the amino acid sequence includes one or more variant amino acid positions set forth in Tables 1, 2, and / or 3. In particular, such polypeptides encode aldehyde dehydrogenases, which can convert acyl-CoA to the corresponding aldehyde, e.g., 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde, or 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde. In some embodiments, the isolated polypeptides of the present invention comprise an amino acid sequence other than one or more variant amino acid positions set forth in Tables 1, 2, and / or 3 that has at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to or is identical to an amino acid sequence referenced as SEQ ID NO: 1, 2, or 3 or in Table 4. It is understood that a variant amino acid position may comprise any one of the 20 naturally occurring amino acids, a conservative substitution of the wild-type or parent sequence at the corresponding position of the variant amino acid position, or a specific amino acid at the variant amino acid position, such as an amino acid disclosed herein in Tables 1, 2, and / or 3. It is further understood that any of the variant amino acid positions can be combined to generate additional variants. Variants with a combination of two or more variant amino acid positions exhibited greater activity than the wild-type. Thus, as exemplified herein, generating enzyme variants by combining active variant amino acid positions resulted in enzyme variants with improved properties.One of skill in the art can readily generate polypeptides having single variant positions or combinations of variant positions using methods well known to those of skill in the art to produce polypeptides with desired properties, including increased activity, increased specificity for the R-form of 3-hydroxybutyryl-CoA or 3-hydroxybutyraldehyde over the S-form, increased specificity for 3-hydroxybutyryl-CoA and / or 4-hydroxybutyryl-CoA over acetyl-CoA, reduced formation of by-products such as ethanol or 4-hydroxy-2-butanone, increased kcat, increased stability in vivo and / or in vitro, etc., as described herein.

[0053] "Homology" or "identity" or "similarity" refers to the sequence similarity between two polypeptides or two nucleic acid molecules. Homology can be determined by comparing a position in each sequence that can be aligned for comparison purposes. If a position in the compared sequences is occupied by the same base or amino acid, the molecules are identical at that position. The degree of homology between sequences is a function of the number of matching or homologous positions shared by the sequences. A polypeptide or polypeptide region (or polynucleotide or polynucleotide region) having a certain percentage (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99%) of "sequence identity" with another sequence means that, when aligned, the percentage of amino acids (or nucleotide bases) are the same when the two sequences are compared.

[0054] In certain embodiments, the invention provides isolated polypeptides having an amino acid sequence comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more variants of any combination disclosed herein. The variants may include any combination of variants set forth in Tables 1, 2 and / or 3. In some embodiments, the isolated polypeptide is a variant of a reference polypeptide, wherein the reference polypeptide has the amino acid sequence of SEQ ID NO: 1, 2 or 3 or a sequence in Table 4, and the polypeptide variant is selected from Tables 1-3 and has one or more amino acid substitutions compared to the sequence of SEQ ID NO: 1, 2 or 3 or Table 4.

[0055] In some embodiments, the isolated polypeptide is a variant of a reference polypeptide, wherein the reference polypeptide has an amino acid sequence of SEQ ID NO: 1, 2, or 3 or a sequence in Table 4, and the polypeptide variant comprises one or more amino acid substitutions compared to the sequence of SEQ ID NO: 1, 2, or 3 or Table 4, wherein the one or more amino acid substitutions are selected from Tables 1-3, and the polypeptide variant has at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity to the corresponding variant selected from Tables 1-3. The one or more amino acid substitutions may be selected from any one of the variants listed in Tables 1-3, or from any combination of two or more variants listed in Tables 1-3. When selecting from a single variant in Tables 1-3, the resulting variant may comprise one or more of the substitutions of the selected variant in any combination, including all of the indicated substitutions or fewer than all of the indicated substitutions. When substitutions are selected from substitutions of two or more variants in Tables 1-3, the resulting variant can include one or more of the substitutions of the selected variants, including all or fewer than all of the indicated substitutions from each of the two or more selected variants in any combination. For example, the resulting variant can include one, two, three, or four substitutions from a single variant in Tables 1-3. As a further example, the resulting variant can include one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, twenty, twenty-five, or more substitutions selected from one, two, three, four, five, or more selected variants in Tables 1-3, including up to all positions substituted as disclosed herein. In some embodiments, the resulting variant includes all of the indicated substitutions of the selected variants in Tables 1-3. In some embodiments, the resulting variant differs from SEQ ID NO: 1, 2 or 3 or a sequence of Table 4 by at least one amino acid substitution but by fewer than 25, 20, 10, 5, 4 or 3 amino acid substitutions.In some embodiments, the resulting variant comprises, consists essentially of, or consists of a sequence set forth by a variant selected from Tables 1-3, and differs from the sequence of SEQ ID NO: 1, 2 or 3 or Table 4 only in the amino acid substitutions indicated.

[0056] In some embodiments, the resulting variant has at least 80%, 85%, 90%, or 95% sequence identity with the corresponding variant selected from Tables 1-3; in some cases, the identity is at least 90% or higher. If the resulting variant is less than 100% identical to the corresponding variant selected from Tables 1-3, one or more positions of the amino acid substitutions indicated for the corresponding variant may be shifted (e.g., in the case of insertion or deletion of one or more amino acids) but still be contained within the resulting variant. For example, a glycine-to-glutamic acid substitution corresponding to "D12A" (at position 12 relative to SEQ ID NO: 1 or 2) may be present in the resulting variant, but at a different position. Whether an amino acid corresponds to a indicated substitution, despite a different position, can be determined by sequence alignment, as described above and as is well known in the art. In some embodiments, one or more amino acid substitutions (e.g., all or fewer than all amino acid substitutions) exhibited by a corresponding variant selected from Tables 1-3, even if they occur at different physical positions along the polypeptide chain, are considered to be present in a given variant if the sequence of the compared polypeptide aligns with a corresponding variant that has an identical match or a similar amino acid at the indicated position along the corresponding variant sequence when using the BLASTP alignment algorithm with default parameters, where a similar amino acid is considered to have sufficient chemical properties to align with the variant position of interest using the default parameters of the alignment algorithm.

[0057] Variants, alone or in combination, can produce enzymes that retain or improve activity compared to a reference polypeptide, e.g., a wild-type (native) enzyme. In some embodiments, polypeptides of the invention can have any combination of variants shown in Tables 1, 2, and / or 3. In some embodiments, polypeptides of the invention having any combination of variants shown in Tables 1, 2, and / or 3 can convert acyl-CoA to the corresponding aldehyde, e.g., 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde, or 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde. Methods for producing and assaying such polypeptides are well known to those of skill in the art.

[0058] In some embodiments, the isolated polypeptides of the invention may further comprise conservative amino acid substitutions at 1 to 100 amino acid positions, or alternatively 2 to 100 amino acid positions, or alternatively 3 to 100 amino acid positions, or alternatively 4 to 100 amino acid positions, or alternatively 5 to 100 amino acid positions, or alternatively 6 to 100 amino acid positions, or alternatively 7 to 100 amino acid positions, or alternatively 8 to 100 amino acid positions, or alternatively 9 to 100 amino acid positions, or alternatively 10 to 100 amino acid positions, or alternatively 15 to 100 amino acid positions, or alternatively 20 to 100 amino acid positions, or alternatively 30 to 100 amino acid positions, or alternatively 40 to 100 amino acid positions, or alternatively 50 to 100 amino acid positions, or any integer number of amino acid positions therein, wherein the positions are other than the variant amino acid positions set forth in Tables 1, 2 and / or 3. In some embodiments, conservative amino acid sequences are chemically or evolutionarily conservative amino acid substitutions. Methods for identifying conservative amino acids are well known to those of skill in the art, any one of which can be used to generate the isolated polypeptides of the present invention.

[0059] In some embodiments, the isolated polypeptides of the invention have between 2 and 300 amino acid positions, or alternatively between 3 and 300 amino acid positions, or alternatively between 4 and 300 amino acid positions, or alternatively between 5 and 300 amino acid positions, or alternatively between 10 and 300 amino acid positions, or alternatively between 20 and 300 amino acid positions, or alternatively between 30 and 300 amino acid positions, or alternatively between 40 and 300 amino acid positions, or alternatively between 50 and 300 amino acid positions, or alternatively between 50 and 300 amino acid positions, or alternatively between 60 and 650 amino acid positions, or alternatively between 70 and 750 amino acid positions, or alternatively between 80 and 850 amino acid positions, or alternatively between 90 and 950 amino acid positions, or alternatively between 100 and 1200 amino acid positions, or alternatively between 120 and 1250 amino acid positions, or alternatively between 130 and 1400 amino acid positions, or alternatively between 140 and 1500 amino acid positions, or alternatively between 150 and 1600 amino acid positions, or alternatively between 160 and 1700 amino acid positions, or alternatively between 170 and 1800 amino acid positions, or alternatively between 180 and 1900 amino acid positions, or alternatively between 200 and 2000 amino acid positions, or alternatively between 200 and 2500 amino acid positions, or alternatively between 250 and 2600 amino acid positions, or alternatively between 260 and 2700 amino acid positions, or alternatively between 270 and 2800 amino acid positions, or alternatively between 300 and 300 amino acid positions, or alternatively between 400 and 300 amino acid positions, or alternatively between 500 and 300 amino or alternatively, may not comprise alterations at 60 to 300 amino acid positions, or alternatively, 80 to 300 amino acid positions, or alternatively, 100 to 300 amino acid positions, or alternatively, 150 to 300 amino acid positions, or alternatively, 200 to 300 amino acid positions, or alternatively, 250 to 300 amino acid positions, or any integer number of amino acid positions therein, wherein the positions are selected from positions identical to 2, 3, 4 or 5 of the amino acid sequences referenced as SEQ ID NO: 1, 2 or 3 or in Table 4.

[0060] As disclosed herein, it is understood that variant polypeptides, such as polypeptide variants of aldehyde dehydrogenase, can perform the same enzymatic reaction as the parent polypeptide, e.g., converting acyl-CoA to its corresponding aldehyde, e.g., converting 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde or 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde. It is further understood that polypeptide variants of aldehyde dehydrogenase enzymes can include variants that provide distinct benefits to the polypeptide, including, but not limited to, increased activity, increased specificity for the R-form of 3-hydroxybutyryl-CoA or 3-hydroxybutyraldehyde over the S-form, increased specificity for 3-hydroxybutyryl-CoA and / or 4-hydroxybutyryl-CoA over acetyl-CoA, reduced formation of by-products such as ethanol or 4-hydroxy-2-butanone, increased kcat, increased stability in vivo and / or in vitro, etc. (See Examples). In certain embodiments, the aldehyde dehydrogenase variants can exhibit activity at least equal to or higher than that of the wild-type or parent polypeptide, i.e., higher than that of the parent polypeptide not having the variant amino acid position. For example, the aldehyde dehydrogenase variants of the present invention can have variant polypeptides with 1.2, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10-fold or even higher activity than the wild-type or parent polypeptide (see Examples). Activity is understood to refer to the ability of the aldehyde dehydrogenase of the present invention to convert a substrate into a product compared to the wild-type or parent polypeptide under the same assay conditions.

[0061] In another specific embodiment, the aldehyde dehydrogenase variant can exhibit increased specificity for the R-form of 3-hydroxybutyryl-CoA or 3-hydroxybutyraldehyde over the S-form, e.g., about 2-40 fold higher, e.g., 2-35, 2-30, 2-25, 2-20, 2-15, 2-10, or 2-5, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40 fold or even higher activity. Such increased specificity can be measured, for example, by the ratio of activity for the R-form of 3-hydroxybutyryl-CoA or 3-hydroxybutyraldehyde to the S-form.

[0062] In another particular embodiment, the aldehyde dehydrogenase variant has increased specificity for 3-hydroxybutyryl-CoA and / or 4-hydroxybutyryl-CoA over acetyl-CoA, e.g., 1.5-100, 1.5-95, 1.5-90, 1.5-85, 1.5-80, 1.5-75, 1.5-70, 1.5-65, 1.5-60, 1.5-55, 1. The specificity may be increased by 5 to 50, 1.5 to 45, 1.5 to 40, 1.5 to 35, 1.5 to 30, 1.5 to 25, 1.5 to 20, 1.5 to 15, 1.5 to 10, or 1.5 to 5, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 times. Such an increase in specificity can be measured, for example, by the ratio of activity for 3-hydroxybutyryl-CoA or 4-hydroxybutyryl-CoA to acetyl-CoA. Specificity is expressed by the activity for 3HB-CoA or 4HB-CoA divided by the activity for acetyl-CoA.

[0063] In another specific embodiment, the aldehyde dehydrogenase variant may exhibit reduced formation of by-products such as ethanol and / or 4-hydroxy-2-butanone, e.g., 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% reduction in by-product formation. Such aldehyde dehydrogenase variants may exhibit activity with reduced by-product formation, as described above, compared to the wild-type or parent polypeptide, i.e., the parent polypeptide not having the variant amino acid position.

[0064] In another specific embodiment, the aldehyde dehydrogenase variant may exhibit an increase in kcat of, for example, 1.25, 1.5, 1.75, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10-fold or more compared to the wild-type or parent polypeptide, i.e., the parent polypeptide not having the variant amino acid position. kcat is understood to refer to its well-known meaning in enzymology of turnover number, where kcat=Vmax / [E T ], Vmax is the enzyme reaction rate at saturated substrate, [E T

[0047] where k is the total enzyme concentration (see Segel, Enzyme Kinetics: Behavior and Analysis of Rapid Equilibrium and Steady-State Enzyme Kinetics, Wiley-Interscience, New York (1975)). Such aldehyde dehydrogenase variants can exhibit activity with an increased k compared to the wild-type or parent polypeptide, i.e., the parent polypeptide not having the variant amino acid position.

[0065] In another specific embodiment, the aldehyde dehydrogenase variant may exhibit increased stability in vitro or in vivo, or both, compared to the wild-type or parent polypeptide, i.e., the parent polypeptide not having the variant amino acid position. For example, the aldehyde dehydrogenase variant may exhibit increased stability in vitro in cell lysates.

[0066] It is understood that in certain embodiments, an aldehyde dehydrogenase variant can exhibit two or more of the above characteristics, such as, for example, any combination of: (1) increased activity; (2) increased specificity for the R-form of 3-hydroxybutyryl-CoA or 3-hydroxybutyraldehyde over the S-form; (3) increased specificity for 3-hydroxybutyryl-CoA and / or 4-hydroxybutyryl-CoA over acetyl-CoA; (4) reduced formation of by-products such as ethanol and / or 4-hydroxy-2-butanone; (5) increased kcat; and (6) increased stability in vivo and / or in vitro. Such combinations include, for example, features 1 and 2; 1 and 3; 1 and 4; 1 and 5; 1 and 6; 2 and 3; 2 and 4; 2 and 5; 2 and 6; 3 and 4; 3 and 5; 3 and 6; 4 and 5; 4 and 6; 5 and 6; 1, 2 and 3; 1, 2 and 4; 1, 2 and 5; 1, 2 and 6; 1, 3 and 4; 1, 3 and 5; 1, 3 and 6; 1, 4 and 5; 1, 4 and 6; 1, 5 and 6; 2, 3 and 4; 2, 3 and 5; 2, 3 and 6; 2, 4 and 5; 2, 4 and 6; 2, 5 and 6; 3, 4 and 5; 3, 4 and and 6; 3, 5 and 6; 4, 5 and 6; 1, 2, 3 and 4; 1, 2, 3 and 5; 1, 2, 3 and 6; 1, 2, 4 and 5; 1, 2, 4 and 6; 1, 2, 5 and 6; 1, 3, 4 and 5; 1, 3, 4 and 6; 1, 3, 5 and 6; 1, 4, 5 and 6; 2, 3, 4 and 5; 2, 3, 4 and 6; 2, 3, 5 and 6; 3, 4, 5 and 6; 1, 2, 3, 4 and 5; 1, 3, 4, 5 and 6; 1, 2, 4, 5 and 6; 1, 2, 3, 5 and 6; 1, 2, 3, 4 and 6; 2, 3, 4, 5 and 6; 1, 2, 3, 4, 5 and 6.

[0067] The polypeptides of the present invention can be isolated by various methods well known in the art, such as recombinant expression systems, precipitation, gel filtration, ion exchange, reverse-phase and affinity chromatography, etc. Other well-known methods are described in Deutscher et al., Guide to Protein Purification: Methods in Enzymology, Vol. 182, (Academic Press, (1990)). Alternatively, the isolated polypeptides of the present invention can be obtained using well-known recombinant methods (see, e.g., Sambrook et al., supra, 1989; Ausubel et al., supra, 1999). Methods and conditions for biochemical purification of the polypeptides of the present invention can be selected by those skilled in the art, and purification can be monitored, for example, by functional assays.

[0068] One non-limiting example of a method for preparing a polypeptide of the present invention is to express a nucleic acid encoding the polypeptide in a suitable host cell, such as a bacterial cell, yeast cell, or other suitable cell, using methods well known in the art, as described herein, and recover the expressed polypeptide using, again, well-known purification methods. Polypeptides of the present invention can be isolated directly from cells transformed with an expression vector as described herein. Alternatively, recombinantly expressed polypeptides of the present invention can be expressed as fusion proteins with a suitable affinity tag, such as glutathione S-transferase (GST), polyHis, streptavidin, or the like, and affinity purified, if desired. Polypeptides of the present invention may optionally retain the affinity tag, or, if desired, the affinity tag may be removed from the polypeptide using well-known methods for removing the affinity tag, such as by suitable enzymatic or chemical cleavage. Thus, the present invention provides polypeptides of the present invention without or optionally with an affinity tag. In some embodiments, the present invention provides host cells expressing the polypeptides of the present invention disclosed herein. The polypeptides of the present invention can also be produced by chemical synthesis using methods of polypeptide synthesis well known to those skilled in the art (Merrifield, J. Am. Chem. Soc. 85:2149 (1964); Bodansky, M., Principles of Peptide Synthesis (Springer-Verlag, 1984); Houghten, Proc. Natl. Acad. Sci., USA 82:5131 (1985); Grant Synthetic Peptides: A User Guide. W.H. Freeman and Co., NY (1992); Bodansky M and Trost B., eds., Principles of Peptide Synthesis. Springer-Verlag Inc., NY (1993)).

[0069] In some embodiments, the present invention provides for the use of the polypeptides disclosed herein as biocatalysts. "Biocatalyst," as used herein, refers to a biological substance that initiates or modifies the rate of a chemical reaction. The biocatalyst can be an enzyme. The polypeptides of the present invention can be used to increase the rate of conversion of a substrate to a product, as disclosed herein. In the context of industrial reactions, the polypeptides of the present invention can be used, for example, to improve reactions that produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, using in vitro methods without host cells expressing the polypeptide. In one embodiment, the present invention provides for the use of the polypeptides of the present invention as biocatalysts.

[0070] In some embodiments of the present invention, a polypeptide encoding an aldehyde dehydrogenase of the present invention is provided as a cell lysate of a cell expressing the aldehyde dehydrogenase. In such cases, the cell lysate serves as a source of aldehyde dehydrogenase for converting 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde or 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde in an in vitro reaction, or the reverse reaction. In another embodiment, the aldehyde dehydrogenase can be provided in a partially purified form, for example, partially purified from a cell lysate. In another embodiment, the aldehyde dehydrogenase can be provided in a substantially purified form, in which the aldehyde dehydrogenase is substantially purified from other components, for example, components of a cell extract. Methods for partially or substantially purifying a polypeptide encoding an aldehyde dehydrogenase are well known in the art, as described herein. In some embodiments, the aldehyde dehydrogenase is immobilized on a solid support, for example, a bead, a plate, or a membrane. In certain embodiments, the aldehyde dehydrogenase comprises an affinity tag, which is used to immobilize the aldehyde dehydrogenase to a solid support. Such affinity tags may include, but are not limited to, glutathione S-transferase (GST), polyHis, streptavidin, etc., as described herein.

[0071] In some embodiments, the present invention provides a composition comprising a polypeptide disclosed herein and at least one substrate for the polypeptide. The substrate for each of the polypeptides disclosed herein is described herein and illustrated in the figures. The polypeptide in the composition of the present invention can react with the substrate under in vitro or in vivo conditions. In this context, in vitro conditions refer to reactions in the absence of cells or outside cells, including cells of the present invention.

[0072] In one embodiment, the present invention provides a composition comprising a polypeptide of the present invention and at least one substrate for the polypeptide. In one embodiment, the polypeptide can react with the substrate under in vitro conditions. In one embodiment, the substrate is 3-hydroxybutyryl-CoA. In one embodiment, the substrate is 3-hydroxy-(R)-butyryl-CoA. In one embodiment, the substrate is 4-hydroxybutyryl-CoA.

[0073] In some embodiments, the present invention provides methods for constructing host strains, which may include, among other steps, introducing a vector disclosed herein into a host cell capable of, for example, expressing and / or fermenting the amino acid sequence encoded by the vector. The vectors of the present invention can be stably or transiently introduced into host cells using techniques well known in the art, including, but not limited to, conjugation, electroporation, chemical conversion, transduction, transfection, and ultrasonic transformation. Additional methods are disclosed herein, any one of which can be used in the methods of the present invention.

[0074] In further embodiments, the present invention provides cells comprising a polypeptide of the present invention, i.e., an aldehyde dehydrogenase of the present invention. Accordingly, the present invention provides non-naturally occurring cells comprising a polypeptide encoding an aldehyde dehydrogenase of the present invention. Optionally, the cells can comprise a 3-HBal or 1,3-BDO pathway, or a 4-HBal or 1,4-BDO pathway, and optionally, a pathway for producing an associated downstream product, such as an ester or amide thereof. In some embodiments, the non-naturally occurring cells comprise at least one exogenous nucleic acid encoding an aldehyde dehydrogenase that converts acyl-CoA to its corresponding aldehyde. Those skilled in the art will understand that these are merely examples, and that any of the substrate-product pairs disclosed herein that are suitable for producing a desired product and in which appropriate activity is available for conversion of the substrate to the product can be readily determined by those skilled in the art based on the teachings herein. Thus, in certain embodiments, the invention provides cells, particularly non-naturally occurring cells, that contain at least one exogenous nucleic acid encoding an aldehyde dehydrogenase, where the aldehyde dehydrogenase functions in a 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway, such as those depicted in Figures 1 and 2.

[0075] In one embodiment, the present invention provides a cell comprising a vector of the present invention comprising a nucleic acid of the present invention. The present invention also provides a cell comprising a nucleic acid of the present invention. In one embodiment, the nucleic acid molecule is integrated into a chromosome of the cell. In certain embodiments, the integration is site-specific. In one embodiment of the present invention, the nucleic acid molecule is expressed. In one embodiment, the present invention provides a cell comprising a polypeptide of the present invention.

[0076] In one embodiment, the cell comprising the vector, nucleic acid, or polypeptide is a microbial organism. In certain embodiments, the microbial organism is a bacterium, yeast, or fungus. In certain embodiments, the cell is an isolated eukaryotic cell.

[0077] In one embodiment, the cells comprise a pathway for producing 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or an ester or amide thereof. In another embodiment, the cells comprise a pathway for producing 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) or an ester or amide thereof. In one embodiment, the cells are capable of fermentation. In one embodiment, the cells further comprise at least one substrate for a polypeptide of the invention expressed in the cells. In a specific embodiment, the substrate is 3-hydroxybutyryl-CoA. In a specific embodiment, the substrate is 3-hydroxy-(R)-butyryl-CoA. In one embodiment, the cells have a higher activity for 3-hydroxy-(R)-butyryl-CoA than for 3-hydroxy-(S)-butyryl-CoA. In another specific embodiment, the substrate is 4-hydroxybutyryl-CoA. The present invention also provides culture media comprising the cells of the present invention.

[0078] The aldehyde dehydrogenases of the invention can be utilized in pathways that convert acyl-CoA to its corresponding aldehyde. Exemplary pathways for 3-HBal and / or 1,3-BDO that include aldehyde dehydrogenases are described, for example, in WO 2010 / 127319, WO 2013 / 036764, U.S. Patent No. 9,017,983, and US 2013 / 0066035, each of which is incorporated herein by reference.

[0079] An exemplary 3-HBal and / or 1,3-BDO pathway is shown in Figure 1 and described in WO 2010 / 127319, WO 2013 / 036764, U.S. Patent No. 9,017,983, and US 2013 / 0066035. Such 3-HBal and / or 1,3-BDO pathways containing aldehyde dehydrogenases include, for example, (G) acetoacetyl-CoA reductase (ketone reduction); (H) 3-hydroxybutyryl-CoA reductase (aldehyde formation), also referred to herein as 3-hydroxybutyraldehyde dehydrogenase, aldehyde dehydrogenase (ALD); and (C) 3-hydroxybutyraldehyde reductase, also referred to herein as 1,3-BDO dehydrogenase (see Figure 1). Acetoacetyl-CoA can be formed by utilizing a thiolase to convert two molecules of acetyl-CoA into one molecule of acetoacetyl-CoA. Acetoacetyl-CoA thiolase converts two molecules of acetyl-CoA into one molecule each of acetoacetyl-CoA and CoA (see WO2013 / 036764 and US2013 / 0066035).

[0080] An exemplary 1,3-BDO pathway is shown in Figure 2 of WO 2010 / 127319. Briefly, acetoacetyl-CoA can be converted to 3-hydroxybutyryl-CoA by acetoacetyl-CoA reductase (ketone reduction) (EC 1.1.1.a) (Step G of Figure 1). 3-Hydroxybutyryl-CoA can be converted to 3-hydroxybutyraldehyde by 3-hydroxybutyryl-CoA reductase (aldehyde formation) (EC 1.2.1.b), also referred to herein as 3-hydroxybutyraldehyde dehydrogenase, including aldehyde dehydrogenases of the invention (Step H of Figure 1). 3-Hydroxybutyraldehyde can be converted to 1,3-butanediol by 3-hydroxybutyraldehyde reductase (EC 1.1.1.a), also referred to herein as 1,3-BDO dehydrogenase (Step C of Figure 1).

[0081] As disclosed herein, the aldehyde dehydrogenases of the present invention can function in pathways that convert 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde. In the pathway described above, which includes an aldehyde dehydrogenase that converts 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde, the pathway converts acetoacetyl-CoA to 3-hydroxybutyryl-CoA (see FIG. 1). The aldehyde dehydrogenases of the present invention can also be used in other 3-HBal and / or 1,3-BDO pathways that include 3-hydroxybutyryl-CoA as a pathway substrate / product. One skilled in the art can readily utilize the aldehyde dehydrogenases of the present invention to convert 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde in any desired pathway that includes such reactions.

[0082] Exemplary 4-HBal and / or 1,4-BDO pathways are shown in FIG. 2 and described in WO2008 / 115840, WO2010 / 030711, WO2010 / 141920, WO2011 / 047101, WO2013 / 184602, WO2014 / 176514, U.S. Pat. No. 8,067,214, U.S. Pat. No. 7,858,350, U.S. Pat. No. 8,129,169, U.S. Pat. No. 8,377,666, US2013 / 0029381, US2014 / 0030779, US2015 / 0148513, and US2014 / 0371417. Such 4-HBal and / or 1,4-BDO pathways containing aldehyde dehydrogenases include, for example, (1) succinyl-CoA synthetase; (2) CoA-independent succinic semialdehyde dehydrogenase; (3) α-ketoglutarate dehydrogenase; (4) glutamate:succinic semialdehyde transaminase; (5) glutamate decarboxylase; (6) CoA-dependent succinic semialdehyde dehydrogenase; (7) 4-hydroxybutanoate dehydrogenase; (8) α-ketoglutarate decarboxylase; (9) 4-hydroxybutanoate dehydrogenase; (10) butyrate kinase (also called 4-hydroxybutyrate kinase); (11) phosphotransbutyrylase (also called phospho-trans-4-hydroxybutyrylase); (12) aldehyde dehydrogenase (also called 4-hydroxybutyryl-CoA reductase); and (13) alcohol dehydrogenase, for example, 1,4-butanediol dehydrogenase (also called 4-hydroxybutanal reductase or 4-hydroxybutyraldehyde reductase) (see Figure 2).

[0083] Similar to Figure 2, an exemplary 1,4-BDO pathway is shown in Figure 8A of WO 2010 / 141920. Briefly, succinyl-CoA can be converted to succinic semialdehyde by succinyl-CoA reductase (or succinic semialdehyde dehydrogenase) (EC 1.2.1.b). Succinic semialdehyde can be converted to 4-hydroxybutyrate by 4-hydroxybutyrate dehydrogenase (EC 1.1.1.a). Alternatively, succinyl-CoA can be converted to 4-hydroxybutyrate by succinyl-CoA reductase (alcohol formation) (EC 1.1.1.c). 4-Hydroxybutyrate can be converted to 4-hydroxybutyryl-CoA by 4-hydroxybutyryl-CoA transferase (EC 2.8.3.a), 4-hydroxybutyryl-CoA hydrolase (EC 3.1.2.a), or 4-hydroxybutyryl-CoA ligase (or 4-hydroxybutyryl-CoA synthetase) (EC 6.2.1.a). Alternatively, 4-hydroxybutyrate can be converted to 4-hydroxybutyryl-phosphate by 4-hydroxybutyrate kinase (EC 2.7.2.a). 4-Hydroxybutyryl-phosphate can be converted to 4-hydroxybutyryl-CoA by phosphotrans-4-hydroxybutyrylase (EC 2.3.1.a). Alternatively, 4-hydroxybutyryl-phosphate can be converted to 4-hydroxybutanal by 4-hydroxybutanal dehydrogenase (phosphorylating) (EC 1.2.1.d). 4-Hydroxybutyryl-CoA can be converted to 4-hydroxybutanal by 4-hydroxybutyryl-CoA reductase (or 4-hydroxybutanal dehydrogenase) (EC 1.2.1.b), including by the aldehyde dehydrogenase variants of the present invention. Alternatively, 4-hydroxybutyryl-CoA can be converted to 1,4-butanediol by 4-hydroxybutyryl-CoA reductase (alcohol formation) (EC 1.1.1.c). 4-Hydroxybutanal can be converted to 1,4-butanediol by 1,4-butanediol dehydrogenase (EC 1.1.1.a).

[0084] An exemplary 1,4-BDO pathway is also shown in Figure 8B of WO 2010 / 141920. Briefly, alpha-ketoglutarate can be converted to succinic semialdehyde by alpha-ketoglutarate decarboxylase (EC 4.1.1.a). Alternatively, alpha-ketoglutarate can be converted to glutamate by glutamate dehydrogenase (EC 1.4.1.a). 4-Aminobutyrate can be converted to succinic semialdehyde by 4-aminobutyrate oxidoreductase (deaminating) (EC 1.4.1.a) or 4-aminobutyrate transaminase (EC 2.6.1.a). Glutamate can be converted to 4-aminobutyrate by glutamate decarboxylase (EC 4.1.1.a). Succinic semialdehyde can be converted to 4-hydroxybutyrate by 4-hydroxybutyrate dehydrogenase (EC 1.1.1.a). 4-Hydroxybutyrate can be converted to 4-hydroxybutyryl-CoA by 4-hydroxybutyryl-CoA transferase (EC 2.8.3.a), 4-hydroxybutyryl-CoA hydrolase (EC 3.1.2.a), or 4-hydroxybutyryl-CoA ligase (or 4-hydroxybutyryl-CoA synthetase) (EC 6.2.1.a). 4-Hydroxybutyrate can be converted to 4-hydroxybutyryl-phosphate by 4-hydroxybutyrate kinase (EC 2.7.2.a). 4-Hydroxybutyryl-phosphate can be converted to 4-hydroxybutyryl-CoA by phosphotrans-4-hydroxybutyrylase (EC 2.3.1.a). Alternatively, 4-hydroxybutyryl phosphate can be converted to 4-hydroxybutanal by 4-hydroxybutanal dehydrogenase (phosphorylating) (EC 1.2.1.d). 4-Hydroxybutyryl-CoA can be converted to 4-hydroxybutanal by 4-hydroxybutyryl-CoA reductase (or 4-hydroxybutanal dehydrogenase) (EC 1.2.1.b), including by an aldehyde dehydrogenase of the invention.4-Hydroxybutyryl-CoA can be converted to 1,4-butanediol by 4-hydroxybutyryl-CoA reductase (alcohol formation) (EC 1.1.1.c), and 4-hydroxybutanal can be converted to 1,4-butanediol by 1,4-butanediol dehydrogenase (EC 1.1.1.a).

[0085] As disclosed herein, the aldehyde dehydrogenases of the present invention can function in pathways that convert 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde. In the above pathways that include an aldehyde dehydrogenase that converts 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde, the pathway converts 4-hydroxybutyrate to 4-hydroxybutyryl-CoA or 4-hydroxybutyryl-phosphate to 4-hydroxybutyryl-CoA (see Figure 2). The aldehyde dehydrogenases of the present invention can also be used in other 4-HBal and / or 1,4-BDO pathways that include 4-hydroxybutyryl-CoA as a pathway substrate / product. One skilled in the art can readily utilize the aldehyde dehydrogenases of the present invention to convert 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde in any desired pathway that includes such reactions. For example, 4-oxobutyryl-CoA can be converted to 4-hydroxybutyryl-CoA as described and shown in WO2010 / 141290, Figure 9A. In addition, 5-hydroxy-2-oxopentanoic acid can be converted to 4-hydroxybutyryl-CoA as described and shown in WO2010 / 141290, Figures 10 and 11. Also, acetoacetyl-CoA, 3-hydroxybutyryl-CoA, crotonyl-CoA, and / or vinylacetyl-CoA can be converted to 4-hydroxybutyryl-CoA as described and shown in WO2010 / 141290, Figure 12. In addition, 4-hydroxybut-2-enoyl-CoA can be converted to 4-hydroxybutyryl-CoA as described and shown in WO2010 / 141290, Figure 13. Thus, one of skill in the art will readily understand how to use the aldehyde dehydrogenases of the invention, as desired, in 4-HBal and / or 1,4-BDO pathways involving the conversion of 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde.

[0086] The enzyme types required to convert common central metabolic intermediates to 1,3-BDO or 1,4-BDO are listed above with representative enzyme (EC) numbers (WO2010 / 127319, WO2013 / 036764, WO2008 / 115840, WO2010 / 030711, WO2010 / 141920, WO2011 / 047101, WO2013 / 184602, WO (See also US2014 / 176514, US9,017,983, US8,067,214, US7,858,350, US8,129,169, US8,377,666, US2013 / 0066035, US2013 / 0029381, US2014 / 0030779, US2015 / 0148513, and US2014 / 0371417.) The first three digits of each label correspond to the first three digits of the enzyme number, which indicates the general type of conversion independent of substrate specificity. Exemplary enzymes include: 1.1.1.a, oxidoreductase (ketone to hydroxyl, or aldehyde to alcohol); 1.1.1.c, oxidoreductase (two-step, acyl-CoA to alcohol); 1.2.1.b, oxidoreductase (acyl-CoA to aldehyde); 1.2.1.c, oxidoreductase (2-oxoacid to acyl-CoA, decarboxylation); 1.2.1.d, oxidoreductase (phosphorylation / dephosphorylation); 1.3.1.a, oxidoreductase acting on CH-CH donors; 1.4.1.a, oxidoreductase acting on amino acids. ductases (deaminating); 2.3.1.a, acyltransferases (transfer of phosphate groups); 2.6.1.a, aminotransferases; 2.7.2.a, phosphotransferases, carboxyl group acceptors; 2.8.3.a, coenzyme A transferases; 3.1.2.a, thiol ester hydrolases (CoA specific); 4.1.1.a, carboxy-lyases; 4.2.1.a, hydro-lyases; 4.3.1.a, ammonia-lyases; 5.3.3.a, isomerases; 5.4.3.a, aminomutases; and 6.2.1.a, acid-thiol ligases.

[0087] The aldehyde dehydrogenases of the present invention can be utilized in cells or in vitro to convert acyl-CoA to its corresponding aldehyde. As disclosed herein, the aldehyde dehydrogenases of the present invention have beneficial and useful properties, including, but not limited to, increased specificity for the R enantiomer of 3-hydroxybutyryl-CoA over the S enantiomer, increased specificity for 3-hydroxybutyryl-CoA and / or 4-hydroxybutyryl-CoA over acetyl-CoA, increased activity, reduced by-product production, increased kcat, etc. The aldehyde dehydrogenases of the present invention can be used to produce the R form of 1,3-butanediol (also referred to as (R)-1,3-butanediol) by enzymatically converting the product of the aldehyde dehydrogenases of the present invention, 3-hydroxy-(R)-butyraldehyde, to (R)-1,3-butanediol using 1,3-butanediol dehydrogenase.

[0088] The bio-derived R-form of 1,3-butanediol can be utilized for the production of downstream products where the R-form is preferred. In some embodiments, the R-form can be utilized as a pharmaceutical and / or dietary supplement (see WO 2014 / 190251). For example, (R)-1,3-butanediol can be used to produce (3R)-hydroxybutyl (3R)-hydroxybutyrate, which can have beneficial effects such as increasing blood ketone body levels. Increasing ketone body levels can result in various clinical benefits, including enhanced physical and cognitive performance, cardiovascular conditions, treatment of diabetes, and mitochondrial dysfunction disorders, as well as treatment of muscle fatigue and muscle disorders (see WO 2014 / 190251). The bio-derived R-form of 1,3-butanediol can be utilized for the production of downstream products where a non-petroleum-based product is desired, for example, by replacing petroleum-derived racemic 1,3-butanediol, its S-form, or its R-form with the bio-derived R-form.

[0089] In one embodiment, the present invention provides 3-HBal or 1,3-BDO or related downstream products, such as esters or amides thereof, that are enriched in the R-enantiomer of the compound. In some embodiments, the 3-HBal or 1,3-BDO is an R-enantiomer-enriched racemate, i.e., contains more R-enantiomers than S-enantiomers. For example, a 3-HBal or 1,3-BDO racemate may contain 55% or more R-enantiomers and 45% or less S-enantiomers. For example, a 3-HBal or 1,3-BDO racemate may contain 60% or more R-enantiomers and 40% or less S-enantiomers. For example, a 3-HBal or 1,3-BDO racemate may contain 65% or more R-enantiomers and 35% or less S-enantiomers. For example, a 3-HBal or 1,3-BDO racemate may contain 70% or more R-enantiomers and 30% or less S-enantiomers. For example, a 3-HBal or 1,3-BDO racemate may contain 75% or more R-enantiomers and 25% or less S-enantiomers. For example, a 3-HBal or 1,3-BDO racemate may contain 80% or more R-enantiomers and 20% or less S-enantiomers. For example, a 3-HBal or 1,3-BDO racemate may contain 85% or more R-enantiomers and 15% or less S-enantiomers. For example, a 3-HBal or 1,3-BDO racemate may contain 90% or more R-enantiomers and 10% or less S-enantiomers. For example, a 3-HBal or 1,3-BDO racemate may contain 95% or more of the R-enantiomer and 5% or less of the S-enantiomer.In some embodiments, 3-HBal or 1,3-BDO or a downstream product related thereto, e.g., an ester or amide thereof, is greater than 90% R-form, e.g., greater than 95%, 96%, 97%, 98%, 99% or 99.9% R-form. In one embodiment, 3-HBal and / or 1,3-BDO or downstream products related thereto, e.g., esters or amides thereof, are 55% or more R-enantiomer, 60% or more R-enantiomer, 65% or more R-enantiomer, 70% or more R-enantiomer, 75% or more R-enantiomer, 80% or more R-enantiomer, 85% or more R-enantiomer, 90% or more R-enantiomer, or 95% or more R-enantiomer, and can be highly chemically pure, e.g., 99% or more, e.g., 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, 99.1% or more, 99.2% or more, 99.3% or more, 99.4% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, or 99.9% or more R-enantiomer.

[0090] In one embodiment, a petroleum-derived racemic mixture of precursors to 3-HBal and / or 1,3-BDO, particularly a racemic mixture of 3-hydroxybutyryl-CoA, is used as a substrate for an aldehyde dehydrogenase of the invention, which exhibits increased specificity for the R-form over the S-form, to produce 3-HBal or 1,3-BDO or related downstream products, such as esters or amides thereof, that are enriched in the R-enantiomer. Such reactions can be carried out by feeding the petroleum-derived precursor to a cell expressing an aldehyde dehydrogenase of the invention, particularly a cell capable of converting the precursor to 3-hydroxybutyryl-CoA, or can be carried out in vitro or a combination of in vivo and in vitro reactions using one or more enzymes that convert the petroleum-derived precursor to 3-hydroxybutyryl-CoA. The reaction to produce 4-hydroxybutyryl-CoA with an aldehyde dehydrogenase of the invention can also be carried out by supplying petroleum-derived precursors to cells that express the aldehyde dehydrogenase of the invention, particularly cells that can convert the precursors to 4-hydroxybutyryl-CoA, or can be carried out in vitro or in a combination of in vivo and in vitro reactions using one or more enzymes that convert the petroleum-derived precursors to 4-hydroxybutyryl-CoA.

[0091] While cells containing a 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway including an aldehyde dehydrogenase of the invention are generally described herein, it is understood that the invention also provides cells containing at least one exogenous nucleic acid encoding an aldehyde dehydrogenase of the invention. The aldehyde dehydrogenase can be expressed in sufficient amounts to produce a desired product, e.g., a product of the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway, or a related downstream product, e.g., an ester or amide thereof. Exemplary 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathways are shown in Figures 1 and 2 and described herein.

[0092] As described in the Examples and illustrated in the Figures, it is understood that any of the pathways disclosed herein, including the pathways of Figures 1 and 2, can be utilized, if desired, to generate cells that produce intermediates or products of any pathway, particularly pathways that utilize the aldehyde dehydrogenases of the invention. As disclosed herein, such cells that produce intermediates can be used in combination with other cells that express one or more upstream or downstream pathway enzymes that produce the desired product. However, it is understood that cells that produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates can be utilized to produce the intermediate as a desired product.

[0093] The present invention will be described herein generally with reference to metabolic reactions, their reactants or products, or in particular with reference to one or more nucleic acids or genes that code the enzymes that are related to or catalyze the metabolic reactions, reactants or products that are mentioned, or the proteins that are related to the metabolic reactions, reactants or products that are mentioned.Unless otherwise expressly stated herein, those skilled in the art will understand that a reference to a reaction also refers to the reactants and products of the reaction.Similarly, unless otherwise expressly stated herein, a reference to a reactant or product also refers to the reaction, and a reference to any of these metabolic components also refers to the enzymes that catalyze the reaction, reactant or product that are mentioned, or one or more genes that code the proteins that are involved in the reaction, reactant or product that are mentioned.Similarly, in light of the well-known fields of metabolic biochemistry, enzymology and genomics, a reference to a gene or encoding nucleic acid herein also refers to the corresponding encoded enzyme and the reaction that it catalyzes or the proteins that are related to the reaction, and the reactants and products of that reaction.

[0094] As disclosed herein, carboxylic acid products or pathway intermediates can exist in various ionized forms, including fully protonated, partially protonated, and fully deprotonated forms. Thus, the suffix "-ate" or acid form can be used interchangeably to describe both the free acid form and any deprotonated form, particularly since the ionized form is known to depend on the pH in which the compound is found. It is understood that carboxylic acid products or intermediates include ester forms of carboxylic acid products or pathway intermediates, such as O-carboxylate esters and S-carboxylate esters. O- and S-carboxylates can include C1-C6 lower alkyl, branched or straight-chain carboxylates. Some such O- or S-carboxylates include, but are not limited to, methyl, ethyl, n-propyl, n-butyl, i-propyl, sec-butyl and tert-butyl, pentyl, hexyl O- or S-carboxylates, any of which can further possess unsaturation, such as, for example, propenyl, butenyl, pentyl, and hexenyl O- or S-carboxylates. O-carboxylates may be products of biosynthetic pathways. Other biosynthetically accessible O-carboxylates include medium to long chain groups, i.e., C7 to C22 O-carboxylate esters, derived from fatty alcohols such as heptyl alcohol, octyl alcohol, nonyl alcohol, decyl alcohol, undecyl alcohol, lauryl alcohol, tridecyl alcohol, myristyl alcohol, pentadecyl alcohol, cetyl alcohol, palmitryl alcohol, heptadecyl alcohol, stearyl alcohol, nonadecyl alcohol, arachidyl alcohol, heneicosyl alcohol, and behenyl alcohol, any of which may be branched and / or contain unsaturation, as appropriate.O-carboxylate esters can also be obtained through biochemical or chemical processes, such as esterification of free carboxylic acid products or transesterification of O- or S-carboxylates. S-carboxylates are, for example, CoA S-esters, cysteinyl S-esters, alkyl thioesters, and various aryl and heteroaryl thioesters.

[0095] The cells of the invention can be produced by introducing an expressible nucleic acid encoding an aldehyde dehydrogenase of the invention, and optionally, an expressible nucleic acid encoding one or more of the enzymes or proteins involved in one or more 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO biosynthetic pathways, and, optionally, an enzyme encoding a downstream product related to 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO, such as an ester or amide thereof. Depending on the host cell selected, nucleic acids for some or all of a particular 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO biosynthetic pathway or downstream product can be expressed. For example, if one or more enzymes or proteins for a desired biosynthetic pathway are defective in a selected host, expressible nucleic acids for the missing enzyme(s) or protein(s) are introduced into the host, followed by exogenous expression. Alternatively, if a selected host exhibits endogenous expression of some pathway genes but is deficient in others, to achieve biosynthesis of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO, nucleic acids encoding the missing enzyme(s) or protein(s) or exogenously expressed genes can be included, or exogenous expression can be provided to increase expression of pathway enzymes, if desired. Thus, cells of the invention can be produced by introducing an aldehyde dehydrogenase of the invention, and, if desired, exogenous enzyme or protein activity, to obtain a desired biosynthetic pathway, or by introducing one or more exogenous enzyme or protein activity that, together with one or more endogenous enzymes or proteins, includes an aldehyde dehydrogenase of the invention to produce a desired product, e.g., 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO, or a related downstream product, e.g., an ester or amide thereof.

[0096] For example, host cells can be selected from bacteria, yeast, fungi, or any of a variety of microorganisms applicable or suitable for fermentation processes, in which non-naturally occurring cells expressing the aldehyde dehydrogenase of the invention can be generated. Exemplary bacteria include bacteria from the order Enterobacteriales, family Enterobacteriaceae, including the genera Escherichia and Klebsiella; the order Aeromonadales, family Succinivibrionaceae, including the genus Anaerobiospirillum; the order Pasteurellales, family Pasteurellaceae, including the genera Actinobacillus and Mannheimia; the order Rhizobiales, family Bradyrhizobiaceae, including the genus Rhizobium; the order Bacillales, family Bacillaceae, including the genus Bacillus; and the order Actinomycetales, including the genera Corynebacterium and Streptomyces, respectively. The family includes any species selected from the families Bacteriaceae and Streptomycetaceae; the order Rhodospirillales, family Acetobacteraceae, including the genus Gluconobacter; the order Sphingomonadales, family Sphingomonadaceae, including the genus Zymomonas; the order Lactobacillales, families Lactobacillaceae and Streptococcaceae, including the genera Lactobacillus and Lactococcus, respectively; the order Clostridiales, family Clostridiaceae, genus Clostridium; and the order Pseudomonadales, family Pseudomonadaceae, including the genus Pseudomonas.Non-limiting species of host bacteria include Escherichia coli, Klebsiella oxytoca, Anaerobiospirillum succiniciproducens, Actinobacillus succinogenes, Mannheimia succiniciproducens, Rhizobium etli, Bacillus subtilis, Corynebacterium glutamicum, Gluconobacter oxydans, Zymomonas mobilis, Lactococcus lactis, Lactobacillus plantarum, Streptomyces coelicolor, Clostridium acetobutylicum, Pseudomonas fluorescens, and Pseudomonas putida. E. coli is a particularly useful host organism because it is a well-characterized microbial organism suitable for genetic engineering.

[0097] Similarly, exemplary species of yeast or fungal species include any species selected from the order Saccharomycetales, family Saccaromycetaceae, including the genera Saccharomyces, Kluyveromyces, and Pichia; the order Saccharomycetales, family Dipodascaceae, including the genus Yarrowia; the order Schizosaccharomycetales, family Schizosaccharomycetaceae, including the genus Schizosaccharomyces; the order Eurotiales, family Trichocomaceae, including the genus Aspergillus; and the order Mucorales, family Mucoraceae, including the genus Rhizopus. Non-limiting species of host yeast or fungi include Saccharomyces cerevisiae, Schizosaccharomyces pombe, Kluyveromyces lactis, Kluyveromyces marxianus, Aspergillus terreus, Aspergillus niger, Pichia pastoris, Rhizopus arrhizus, Rhizobus oryzae, Yarrowia lipolytica, etc. Particularly useful host organisms that are yeast include Saccharomyces cerevisiae.

[0098] While the present specification generally describes the use of microbial cells as host cells, particularly for producing 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, it is understood that the host cells may be higher eukaryotic cell lines, e.g., mammalian or insect cell lines. Thus, references herein to microbial host cells indicate that higher eukaryotic cell lines may alternatively be utilized to produce the desired product. Exemplary higher eukaryotic cell lines include, but are not limited to, Chinese hamster ovary (CHO), human (Hela, human embryonic kidney (HEK) 293, Jurkat), mouse (3T3), primate (Vero), insect (Sf9), and the like. Such cell lines are commercially available (see, e.g., American Type Culture Collection (ATCC; Manassas, VA); Life Technologies, Carlsbad, CA). It is understood that any suitable host cell can be used to introduce the aldehyde dehydrogenase of the invention and, if necessary, metabolic and / or genetic modifications to produce the desired product.

[0099] Depending on the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO biosynthetic pathway components of the selected host cell, non-naturally occurring cells of the invention can include at least one exogenously expressed 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway-encoding nucleic acid, and up to all of the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO biosynthetic pathway or its associated downstream products, e.g., esters or amides, including an aldehyde dehydrogenase of the invention. For example, biosynthesis of 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO can be established in a host deficient in a pathway enzyme or protein through exogenous expression of the corresponding encoding nucleic acid, including an aldehyde dehydrogenase of the invention. In a host deficient in all enzymes or proteins of the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway or an associated downstream product thereof, e.g., an ester or amide thereof, exogenous expression of all enzymes or proteins in the pathway can be included, although it is understood that all enzymes or proteins in the pathway can be expressed even when the host contains at least one of the enzymes or proteins of the pathway. For example, exogenous expression of all enzymes or proteins in the pathway to produce the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway or an associated downstream product thereof, e.g., an ester or amide thereof, can be included, including the aldehyde dehydrogenase of the invention.

[0100] Given the teachings and guidance provided herein, it will be understood by those skilled in the art that the number of encoding nucleic acids introduced in expressible form will at least parallel the deficiencies of the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway in the selected host cell, if such a pathway is to be contained in the cell. Thus, non-naturally occurring cells of the invention can contain one, two, three, four, five, six, seven, eight, etc., nucleic acids encoding the enzymes or proteins that make up the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO biosynthetic pathway disclosed herein, depending on the particular pathway. In some embodiments, the non-naturally occurring cells can also contain other genetic modifications that enhance or optimize 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO biosynthesis or confer other useful functions on the host microbial organism. One such other functionality can include, for example, increased synthesis of one or more precursors of the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathways, such as acetyl-CoA or acetoacetyl-CoA.

[0101] Generally, host cells are selected so that they are capable of expressing an aldehyde dehydrogenase of the invention and, as needed, producing a precursor of a 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway, either as a naturally occurring molecule or as an engineered product in cells containing such pathways, resulting in either de novo production of a desired precursor or increased production of a precursor naturally produced by the host cell. As disclosed herein, host organisms can be engineered to increase precursor production. Furthermore, cells engineered to produce a desired precursor can also be used as host organisms and further engineered to express enzymes or proteins of the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway or related downstream products, such as esters or amides thereof (if desired).

[0102] In some embodiments, non-naturally occurring cells of the invention are produced from hosts containing enzymatic capabilities to synthesize 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof. In this particular embodiment, it may be useful to increase the synthesis or accumulation of a product of the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway, for example, to drive 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway reactions toward the production of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof. Increased synthesis or accumulation can be achieved, for example, by overexpression of nucleic acids encoding one or more of the above-mentioned 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway enzymes or proteins, including the aldehyde dehydrogenases of the invention. Overexpression of one or more enzymes and / or one or more proteins of the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway can occur, for example, through exogenous expression of an endogenous gene or genes or through exogenous expression of a heterologous gene or genes, including exogenous expression of an aldehyde dehydrogenase of the invention. Thus, for example, naturally occurring organisms can be readily converted into non-naturally occurring cells of the invention that produce 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, e.g., esters or amides thereof, through overexpression of nucleic acids encoding enzymes that produce one, two, three, four, five, six, seven, eight or more, i.e., up to all, of the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO biosynthetic pathway enzymes or proteins, or related downstream products, e.g., esters or amides thereof, depending on the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway.Additionally, non-naturally occurring organisms can be generated by mutagenesis of endogenous genes that result in increased activity of enzymes in the biosynthetic pathways of 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, e.g., esters or amides thereof.

[0103] In particularly useful embodiments, exogenous expression of the encoding nucleic acid is used. Exogenous expression provides the ability to tailor expression and / or regulatory elements to the host and application, achieving a desired expression level controlled by the user. However, other embodiments can utilize endogenous expression, such as by removing a negative regulatory effector or inducing the gene's promoter when linked to an inducible promoter or other regulatory element. Thus, endogenous genes with naturally occurring inducible promoters can be upregulated by providing an appropriate inducer, or the regulatory region of the endogenous gene can be engineered to incorporate an inducible regulatory element, thereby allowing for regulated increased expression of the endogenous gene at a desired time. Similarly, an inducible promoter can be included as a regulatory element for an exogenous gene introduced into a non-naturally occurring cell.

[0104] It is understood that one or more exogenous nucleic acids can be introduced into a cell to generate a non-naturally occurring cell of the invention. The introduction of a nucleic acid can provide the cell with a biosynthetic pathway for 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or a related downstream product, such as an ester or amide thereof, including, for example, the introduction of a nucleic acid encoding an aldehyde dehydrogenase of the invention. Alternatively, the introduction of an encoding nucleic acid can generate a cell with biosynthetic capabilities to catalyze some of the reactions necessary to obtain 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO biosynthetic capabilities and generate intermediates. For example, a non-naturally occurring cell with a biosynthetic pathway for 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO can contain at least two exogenous nucleic acids encoding desired enzymes or proteins, including an aldehyde dehydrogenase of the invention. Thus, it is understood that any combination of two or more enzymes or proteins of a biosynthetic pathway, including an aldehyde dehydrogenase of the invention, can be included in a non-naturally occurring cell of the invention. Similarly, it is understood that any combination of three or more enzymes or proteins of a biosynthetic pathway can be included in a non-naturally occurring cell of the invention, if desired, so long as the combination of enzymes and / or proteins of the desired biosynthetic pathway results in the production of the corresponding desired product. Similarly, any combination of four, or more, enzymes or proteins of the biosynthetic pathways disclosed herein can also be included in a non-naturally occurring cell of the invention, if desired, so long as the combination of enzymes and / or proteins of the desired biosynthetic pathway results in the production of the corresponding desired product.

[0105] In addition to the biosynthesis of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, described herein, the non-naturally occurring cells and methods of the invention can also be utilized in various combinations, i.e., with each other and / or with other cells and methods known in the art, to achieve the biosynthesis of products by other routes. For example, one alternative to using a 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO producer to produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO is by the addition of additional cells capable of converting a 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediate to 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO. One such procedure involves, for example, fermentation of cells that produce a 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediate, which can then be used as a substrate for a second cell that converts the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediate to 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO. The 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediate can be added directly to a separate culture of a second organism, or the original culture of the producer of the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediate can be depleted of these cells, for example, by cell separation, and then the second organism can be added to the fermentation broth and used to produce the final product without an intermediate purification step. Cells that produce downstream products related to 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO, such as esters or amides thereof, can be included, if desired, to produce such downstream products.

[0106] Alternatively, such enzymatic conversions can be carried out in vitro by a combination of enzymes or sequential exposure of the substrate to enzymes that results in the conversion of the substrate to the desired product. As another alternative, a combination of cell-based conversion and in vitro enzymatic conversion can be used, if desired.

[0107] In other embodiments, the non-naturally occurring cells and methods of the invention can be assembled into a variety of subpathways to achieve, for example, the biosynthesis of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof. In these embodiments, the biosynthetic pathways of a desired product of the invention can be separated into different cells, and these different cells can be co-cultured to produce the final product. In such biosynthetic schemes, the product of one cell serves as a substrate for a second cell until the final product is synthesized. For example, the biosynthesis of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, can be achieved by constructing cells containing biosynthetic pathways to convert intermediates of one pathway into intermediates or products of another pathway. Alternatively, 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO can also be biosynthetically produced from cells through co-culture or co-fermentation using two different cells in the same vessel, where a first cell produces a 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO intermediate and a second cell converts the intermediate to 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or a related downstream product, such as an ester or amide thereof.

[0108] Given the teachings and guidance provided herein, it will be understood by those of skill in the art that there are a wide variety of combinations and permutations of the non-naturally occurring cells and methods of the invention that utilize other cells, co-cultures of other non-naturally occurring cells with sub-pathways, and combinations of other chemical and / or biochemical procedures known in the art to produce 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, e.g., esters or amides thereof.

[0109] Sources of encoding nucleic acids for 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway enzymes or proteins or related downstream products, e.g., esters or amides thereof, can include, for example, any species in which the encoded gene product is capable of catalyzing the referenced reaction. Such species include both prokaryotes and eukaryotes, including, but not limited to, bacteria, including archaea and eubacteria, and eukaryotes, including yeast, plants, insects, animals, and mammals, including humans. Exemplary species of such sources include, for example, Escherichia coli, Saccharomyces cerevisiae, Saccharomyces kluyveri, Clostridium kluyveri, Clostridium acetobutylicum, Clostridium beijerinckii, Clostridium saccharoperbutylacetonicum, Clostridium perfringens, Clostridium difficile, Clostridium botulinum, Clostridium tyrobutyricum, Clostridium tetanomorphum, Clostridium tetani, Clostridium propionicum, Clostridium aminobutyricum, Clostridium subterminale, Clostridium sticklandii, Ralstonia eutropha, Mycobacterium bovis, Mycobacterium tuberculosis, Porphyromonas gingivalis, Arabidopsis thaliana, Thermus thermophilus, Pseudomonas species (e.g., Pseudomonas aeruginosa, Pseudomonas putida, Pseudomonas stutzeri, Pseudomonas fluorescens), Homo sapiens, Oryctolagus cuniculus, Rhodobacter spaeroides, Thermoanaerobacterbrockii, Metallosphaera sedula, Leuconostoc mesenteroides, Chloroflexus aurantiacus, Roseiflexus castenholzii, Erythrobacter, Simmondsia chinensis, Acinetobacter species (e.g., Acinetobacter calcoaceticus and Acinetobacter baylyi), Porphyromonas gingivalis, Sulfolobus tokodaii, Sulfolobus solfataricus, Sulfolobus acidocaldarius, Bacillus subtilis, Bacillus cereus, Bacillus megaterium, Bacillus brevis, Bacillus pumilus, Rattus norvegicus, Klebsiella pneumonia, Klebsiella oxytoca, Euglena gracilis, Treponema denticola, Moorella thermoacetica, Thermotoga maritima, Halobacterium salinarum, Geobacillus stearothermophilus, Aeropyrum pernix, Sus scrofa, Caenorhabditis elegans, Corynebacterium glutamicum, Acidaminococcus fermentans, Lactococcus lactis, Lactobacillus plantarum, Streptococcus thermophilus, Enterobacter aerogenes、Candida、Aspergillus terreus、Pedicoccus pentosaceus、Zymomonas mobilus、Acetobacter pasteurians、Kluyveromyces lactis、Eubacterium barkeri、Bacteroides capillosus、Anaerotruncus colihominis、Natranaerobiusthermophilusm, Campylobacter jejuni, Haemophilus influenzae, Serratia marcescens, Citrobacter amalonaticus, Myxococcus xanthus, Fusobacterium nuleatum, Penicillium chrysogenum, Marine gammaproteobacteria, Butyrate-producing bacteria, Nocardia iowensis, Nocardia farcinica, Streptomyces griseus, Schizosaccharomyces pombe, Geobacillus thermoglucosidasius, Salmonella typhimurium, Vibrio cholera, Heliobacter pylori, Nicotiana tabacum, Oryza sativa, Haloferax mediterranei, Agrobacterium tumefaciens, Achromobacter denitrificans, Fusobacterium nucleatum, Streptomyces clavuligenus, Acinetobacter baumanii, Mus musculus, Lachancea kluyveri, Trichomonas vaginalis, Trypanosoma brucei, Pseudomonas stutzeri, Bradyrhizobium japonicum, Mesorhizobium loti, Bos taurus, Nicotiana glutinosa, Vibrio vulnificus, Selenomonas ruminantium, Vibrio parahaemolyticus, Archaeoglobus fulgidus, Haloarcula marismortui, Pyrobaculum aerophilum, Mycobacterium smegmatis MC2 155, Mycobacterium avium subspecies paratuberculosis K-10, Mycobacterium marinum M, Tsukamurella paurometabola DSM 20162, CyanobiumPCC7001, Dictyostelium discoideum AX4, Acidaminococcus fermentans, Acinetobacter baylyi, Acinetobacter calcoaceticus, Aquifex aeolicus, Arabidopsis thaliana, Archaeoglobus fulgidus, Aspergillus niger, Aspergillus terreus, Bacillus subtilis, Bos Taurus, Candida albicans, Candida tropicalis, Chlamydomonas reinhardtii, Chlorobium tepidum, Citrobacter koseri, Citrus junos, Clostridium acetobutylicum, Clostridium kluyveri, Clostridium saccharoperbutylacetonicum, Cyanobium PCC7001, Desulfatibacillum alkenivorans, Dictyostelium discoideum, Fusobacterium nucleatum, Haloarcula marismortui, Homo sapiens, Hydrogenobacter thermophilus, Klebsiella pneumoniae, Kluyveromyces lactis, Lactobacillus brevis, Leuconostoc mesenteroides, Metallosphaera sedula, Methanothermobacter thermautotrophicus, Mus musculus, Mycobacterium avium, Mycobacterium bovis, Mycobacterium marinum, Mycobacterium smegmatis, Nicotiana tabacum, Nocardia iowensis, Oryctolagus cuniculus, Penicillium chrysogenum, Pichia pastoris, Porphyromonas gingivalis, Porphyromonasgingivalis, Pseudomonas aeruginos, Pseudomonas putida, Pyrobaculum aerophilum, Ralstonia eutropha, Rattus norvegicus, Rhodobacter sphaeroides, Saccharomyces cerevisiae, Salmonella enteric, Salmonella typhimurium, Schizosaccharomyces pombe, Sulfolobus acidocaldarius, Sulfolobus solfataricus, Sulfolobus tokodaii, Thermoanaerobacter tengcongensis, Thermus thermophilus, Trypanosoma brucei, Tsukamurella paurometabola, Yarrowia lipolytica, Zoogloea ramigera and Zymomonas mobilis, Clostridium species (e.g., Clostridium saccharoperbutylacetonicum, Clostridium beijerinckii, Clostridium saccharobutylicum, Clostridium botulinum, Clostridium methylpentosum, Clostridium sticklandii, Clostridium phytofermentans, Clostridium saccharolyticum, Clostridium asparagiforme, Clostridium celatum, Clostridium carboxidivorans, Clostridium clostridioforme, Clostridium bolteae、Caldalkalibacillus thermarum、Clostridium botulinum (but not limited to these)、Pelosinus fermentans、Thermoanaerobacteriumthermosaccharolyticum, Desulfosporosinus species, Thermoanaerobacterium species (including, but not limited to, Thermoanaerobacterium saccharolyticum and Thermoanaerobacterium xylanolyticum), Acetonema longum, Geobacillus species (including, but not limited to, Geobacillus thermogluco sidans), Bacillus azotoformans, Thermincola potens, Fusobacterium species (including, but not limited to, Fusobacterium nucleatum and Fusobacterium ulcerans), Fusobacterium varium, Ruminococcus species (e.g., Ruminococcus gnavus and Ruminococcus obeum), Lachnospiraceae bacterium, Flavonifractor plautii, Roseburia inulinivorans, Acetobacterium woodii, Eubacterium species (including, but not limited to, Eubacterium plexicaudatum, Eubacterium hallii, Eubacterium limosum, and Eubacterium yurii), Eubacteriaceae bacterium, Thermosediminibacter oceani, Ilyobacter polytropus, Shuttleworthia satelles, Halanaerobium saccharolyticum, Thermoanaerobacter ethanolicus, Rhodospirillum rubrum, Vibrio, Propionibacterium propionicum, and other exemplary species that can be used as source organisms for the corresponding genes, including source organisms encoding the aldehyde dehydrogenases disclosed herein or listed in Table 4.However, because complete genome sequences are currently available for over 550 species, including the genomes of 395 microbial species and a variety of yeast, fungi, plants, and mammalian genomes (more than half of which are available in public databases, e.g., NCBI), it is routine and well-known in the art to identify genes encoding 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO biosynthetic activity and to exchange genetic alterations between organisms for one or more genes in related or distant species, including, for example, homologs, orthologs, paralogs, and nonorthologous gene replacements of known genes. Thus, metabolic alterations that enable the biosynthesis of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, including expression of the aldehyde dehydrogenase of the invention, as described herein with reference to a particular organism, e.g., E. coli, can be readily applied to other microorganisms, including prokaryotes and eukaryotes alike. Given the teachings and guidance provided herein, one of skill in the art will know that metabolic changes exemplified in one organism can be equally applied to other cells, e.g., organisms.

[0110] In some instances, such as when an alternative 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO biosynthetic pathway exists in an unrelated species, 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO biosynthesis can be conferred to the host species by, for example, exogenous expression of one or more paralogs from the unrelated species that catalyze a similar, but not identical, metabolic reaction in place of the reference reaction. It will be understood by those skilled in the art that actual gene usage may vary between different organisms because certain differences exist between metabolic networks in different organisms. However, given the teachings and guidance provided herein, those skilled in the art will also understand that the teachings and methods of the present invention can be applied to all cells, including the introduction of an aldehyde dehydrogenase of the present invention, by using metabolic alterations cognate to those exemplified herein to construct cells in a subject species that synthesize 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, e.g., esters or amides thereof.

[0111] Methods for constructing and testing expression levels of non-naturally occurring hosts that produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, including the aldehyde dehydrogenases of the invention, can be performed, for example, by recombinant and detection methods well known in the art. Such methods can be found, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory, New York (2001); and Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, MD (1999).

[0112] Exogenous nucleic acids encoding the aldehyde dehydrogenase of the present invention, and, if desired, exogenous nucleic acid sequences involved in the pathway for producing 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, can be stably or transiently introduced into host cells using techniques well known in the art, including, but not limited to, conjugation, electroporation, chemical conversion, transformation, transfection, and ultrasonic transformation. For exogenous expression in E. coli or other prokaryotic cells, some nucleic acid sequences of genes or cDNAs of eukaryotic nucleic acids can encode targeting signals, such as N-terminal mitochondrial or other targeting signals, which can be removed, if desired, before transformation into prokaryotic host cells. For example, removal of the mitochondrial leader sequence increased expression in E. coli (Hoffmeister et al., J. Biol. Chem. 280:4329-4338 (2005)). For exogenous expression in yeast or other eukaryotic cells, genes can be expressed in the cytosol without the addition of a leader sequence, or can be targeted to mitochondria or other organelles, or can be targeted for secretion by adding an appropriate targeting sequence, such as a mitochondrial targeting or secretion signal appropriate to the host cell. It is therefore understood that appropriate modifications can be incorporated into exogenous nucleic acid sequences to remove or include targeting sequences to confer desired properties. Additionally, genes can be codon-optimized using techniques well known in the art to achieve optimized expression of the protein.

[0113] One or more expression vectors can be constructed to contain a nucleic acid encoding an aldehyde dehydrogenase of the invention, as exemplified herein, and / or optionally one or more nucleic acids encoding one or more 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO biosynthetic pathways, or nucleic acids encoding enzymes that produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, operably linked to expression control sequences functional in the host organism. Expression vectors applicable for use in the host cells of the invention include, for example, plasmids, phage vectors, viral vectors, episomes, and artificial chromosomes, including vectors and selection sequences or markers operable for stable integration into the host chromosome. Furthermore, the expression vector can contain one or more selectable marker genes and appropriate expression control sequences. For example, selectable marker genes can be included that provide resistance to antibiotics or toxins, complement auxotrophic deficiencies, or supply critical nutrients absent from the culture medium. Expression control sequences can include constitutive and inducible promoters, transcription enhancers, transcription terminators, and the like, as are well known in the art. When two or more exogenous encoding nucleic acids are coexpressed, both nucleic acids can be inserted, for example, into a single expression vector or into separate expression vectors. For single-vector expression, the encoding nucleic acids can be operably linked to a common expression control sequence or to different expression control sequences, for example, one inducible promoter and one constitutive promoter. Transformation of an exogenous nucleic acid sequence encoding an aldehyde dehydrogenase of the present invention or encoding a polypeptide involved in a metabolic or synthetic pathway can be confirmed using methods well known in the art. Such methods include, for example, nucleic acid analysis, such as Northern blot or polymerase chain reaction (PCR) amplification of mRNA, or immunoblotting for expression of gene products, or other appropriate analytical methods for testing the expression of the introduced nucleic acid sequence or its corresponding gene product.It will be understood by those skilled in the art that the exogenous nucleic acid will be expressed in an amount sufficient to produce the desired product, and it will be further understood that expression levels can be optimized to obtain sufficient expression using methods well known in the art and disclosed herein.

[0114] A vector or expression vector can also be used to express the encoded nucleic acid by in vitro transcription and translation to produce the encoded polypeptide. Such vectors or expression vectors contain at least one promoter and include the vectors described hereinabove. Vectors for such in vitro transcription and translation are generally double-stranded DNA. Methods for in vitro transcription and translation are well known to those skilled in the art (see Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory, New York (2001); and Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, MD (1999)). Kits for in vitro transcription and translation are also commercially available (see, for example, Promega, Madison, WI; New England Biolabs, Ipswich, MA; Thermo Fisher Scientific, Carlsbad, CA).

[0115] In one embodiment, the present invention provides a method for producing 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO), or an ester or amide thereof, comprising culturing cells of the present invention to produce 3-HBal and / or 1,3-BDO, or an ester or amide thereof. Such cells express a polypeptide of the present invention. In one embodiment, the present invention provides a method for producing 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO), or an ester or amide thereof, comprising culturing cells of the present invention to produce 4-HBal and / or 1,4-BDO, or an ester or amide thereof. In one embodiment, the cells are in a substantially anaerobic medium. In one embodiment, the method may further comprise isolating or purifying 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO, or an ester or amide thereof. In certain embodiments, the isolating or purifying step comprises distillation.

[0116] In one embodiment, the invention provides a process for producing a product of the invention, the process comprising chemically reacting 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO with itself or another compound in a reaction to produce the product.

[0117] In one embodiment, the present invention provides a method for producing 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO), or an ester or amide thereof, comprising providing a substrate to a polypeptide of the present invention and converting the substrate to 3-HBal and / or 1,3-BDO, wherein the substrate is a racemic mixture of 1,3-hydroxybutyryl-CoA. In one embodiment, the 3-HBal and / or 1,3-BDO are enriched in the R-enantiomer. In one embodiment, the present invention provides a method for producing 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO), or an ester or amide thereof, comprising providing a substrate to a polypeptide of the present invention and converting the substrate to 4-HBal and / or 1,4-BDO, wherein the substrate is 1,4-hydroxybutyryl-CoA. In one embodiment, the polypeptide is present in a cell, in a cell lysate, or isolated from a cell or cell lysate.

[0118] In one embodiment, the invention provides a method for producing 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO, comprising incubating a lysate of a cell of the invention to produce 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO. In one embodiment, the cell lysate is mixed with a second cell lysate, the second cell lysate containing an enzymatic activity that produces a substrate for a polypeptide of the invention or a downstream product of 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO.

[0119] The present invention also provides a method for producing a polypeptide of the present invention, the method comprising the step of expressing the polypeptide in a cell. The present invention further provides a method for producing a polypeptide of the present invention, the method comprising the step of in vitro transcribing and translating a nucleic acid of the present invention or a vector of the present invention to produce the polypeptide.

[0120] As described herein, cells can be used to express the aldehyde dehydrogenases of the invention, and optionally, the cells can include a metabolic pathway that utilizes the aldehyde dehydrogenases of the invention to produce a desired product, e.g., 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO. Such methods for expressing a desired product are described herein. Alternatively, as described herein, the aldehyde dehydrogenases of the invention can be expressed and / or the desired product can be produced in a cell lysate, e.g., a cell lysate of a cell expressing the aldehyde dehydrogenases of the invention, or a cell that expresses the aldehyde dehydrogenases of the invention and a metabolic pathway to produce the desired product. In another embodiment, the aldehyde dehydrogenases of the invention can be expressed by in vitro transcription and translation, and the aldehyde dehydrogenases are produced in a cell-free system. The aldehyde dehydrogenases expressed by in vitro transcription and translation can be used to perform reactions in vitro. If desired, other enzymes or cell lysates containing such enzymes can be used to convert the products of the aldehyde dehydrogenase enzymatic reaction in vitro to desired downstream products.

[0121] Suitable purification and / or assays for testing aldehyde dehydrogenase expression or the production of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, including assays for aldehyde dehydrogenase activity, can be performed using well-known methods (see also the Examples). For each engineered strain to be tested, appropriate replicates, such as triplicate cultures, can be grown. For example, the formation of products and by-products in the engineered production host can be monitored. End products and intermediates, as well as other organic compounds, can be analyzed by methods such as HPLC (high-performance liquid chromatography), GC-MS (gas chromatography-mass spectrometry), and LC-MS (liquid chromatography-mass spectrometry), or other suitable analytical methods using routine procedures known in the art. Product release in the fermentation broth can also be tested in the culture supernatant. By-products and residual glucose can be quantified, for example, by HPLC using a refractive index detector for glucose and alcohols and an ultraviolet detector for organic acids (Lin et al., Biotechnol. Bioeng. 90:775-779 (2005)), or other suitable assays and detection methods known in the art. The activity of individual enzymes or proteins derived from the exogenous DNA sequence can also be assayed using methods known in the art.

[0122] 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or other desired products, such as related downstream products, such as esters or amides thereof, can be separated from other components in the culture using a variety of methods known in the art. Such separation methods include, for example, extraction methods, as well as methods involving continuous liquid-liquid extraction, pervaporation, membrane filtration, membrane separation, reverse osmosis, electrodialysis, distillation, crystallization, centrifugation, extractive filtration, ion exchange chromatography, size exclusion chromatography, adsorption chromatography, and ultrafiltration. All of the above methods are well known in the art.

[0123] Cells expressing any of the non-naturally occurring aldehyde dehydrogenases of the invention described herein can be cultured to produce and / or secrete biosynthetic products of the invention. For example, cells producing 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, can be cultured for the biosynthesis of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof. Thus, in some embodiments, the invention provides culture media containing 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates described herein. In some aspects, culture medium can be separated from non-naturally occurring cells of the invention that produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates. Methods for separating cells from medium are well known in the art. Exemplary methods include filtration, flocculation, sedimentation, centrifugation, sedimentation, and the like.

[0124] To produce the aldehyde dehydrogenase of the invention, or 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO, or related downstream products, such as esters or amides thereof, in cells expressing the aldehyde dehydrogenase of the invention, the recombinant strain is cultured in a medium having a carbon source and other essential nutrients. To reduce the cost of the overall process, it is sometimes desirable, and in some cases highly desirable, to maintain anaerobic conditions in the fermentor. Such conditions can be achieved, for example, by first sparging the medium with nitrogen and then sealing the flask with a septum and crimp cap. For strains for which growth is not observed under anaerobic conditions, microaerobic or substantially anaerobic conditions can be applied by drilling a small hole in the septum for limited aeration. Exemplary anaerobic conditions have been described previously and are well known in the art. Exemplary aerobic and anaerobic conditions are described, for example, in U.S. Publication No. 2009 / 0047719, filed August 10, 2007. As disclosed herein, fermentation can be carried out in batch, fed-batch, or continuous mode. Fermentation can also be carried out in two phases, if desired. The first phase can be aerobic, allowing for high growth and therefore high productivity, followed by an anaerobic phase for high yields of the desired product, e.g., 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO, or related downstream products, e.g., esters or amides thereof.

[0125] When necessary to maintain the medium at a desired pH, the pH of the medium can be optionally maintained at a desired pH, particularly a neutral pH, for example a pH of about 7, by the addition of a base, such as NaOH or other base, or an acid. The growth rate can be determined by measuring the optical density using a spectrophotometer (600 nm), and the glucose uptake rate can be determined by monitoring carbon source depletion over time.

[0126] The growth medium can include, for example, any carbohydrate source capable of providing a carbon source to non-naturally occurring cells. Such sources include, for example: sugars, such as glucose, xylose, arabinose, galactose, mannose, fructose, sucrose, and starch; or glycerol. It is understood that the carbon source can be used alone as the sole carbon source or in combination with other carbon sources described herein or known in the art. Other sources of carbohydrate include, for example, renewable feedstocks and biomass. Exemplary types of biomass that can be used as feedstocks in the methods of the present invention include cellulosic biomass, hemicellulosic biomass, and lignin feedstocks or portions of feedstocks. Such biomass feedstocks contain, for example, carbohydrate substrates useful as carbon sources, such as glucose, xylose, arabinose, galactose, mannose, fructose, and starch. Given the teachings and guidance provided herein, those skilled in the art will understand that renewable feedstocks and biomass other than those exemplified above can also be used to culture the cells of the invention for expression of the aldehyde dehydrogenases of the invention and, optionally, production of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof.

[0127] In addition to renewable feedstocks such as those exemplified above, cells of the invention that produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, can also be engineered to grow on syngas as a carbon source. In this particular embodiment, one or more proteins or enzymes are expressed in the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO-producing organism to provide a metabolic pathway for utilizing syngas or other gaseous carbon sources.

[0128] Synthesis gas, also known as syngas or producer gas, is the primary product of the gasification of coal and carbonaceous materials, such as biomass materials, including agricultural crops and residues. Syngas is primarily a mixture of H and CO and can be obtained from the gasification of any organic feedstock, including, but not limited to, coal, petroleum, natural gas, biomass, and waste organics. Gasification is typically performed under a high fuel-to-oxygen ratio. While mostly H and CO, synthesis gas can also contain smaller amounts of CO and other gases. Thus, synthesis gas provides a cost-effective source of gaseous carbon, such as CO, as well as CO.

[0129] The Wood-Ljungdahl pathway catalyzes the conversion of CO and H to other products such as acetyl-CoA and acetate. Organisms that can utilize CO and syngas generally also have the ability to utilize CO and CO / H mixtures through the same basic set of enzymes and transformations encompassed by the Wood-Ljungdahl pathway. The H-dependent conversion of CO to acetate by microorganisms was recognized long before it was determined that the same organisms could also use CO and involve the same pathway. Many acetogenic bacteria have been shown to grow in the presence of CO and produce compounds such as acetate, as long as hydrogen is present to supply the necessary reducing equivalents (see, e.g., Drake, Acetogenesis, pp. 3-60, Chapman and Hall, New York, (1994)). This can be summarized in the following equation: 2CO2+4H2+nADP+nPi→CH3COOH+2H2O+nATP Thus, non-naturally occurring microorganisms possessing the Wood-Ljungdahl pathway can also utilize CO and H mixtures for the production of acetyl-CoA and other desired products.

[0130] The Wood-Ljungdahl pathway is well known in the art and consists of 12 reactions that can be divided into two branches: (1) the methyl branch and (2) the carbonyl branch. The methyl branch converts syngas to methyltetrahydrofolate (methyl-THF), and the carbonyl branch converts methyl-THF to acetyl-CoA. The reactions in the methyl branch are catalyzed in sequence by the following enzymes or proteins: ferredoxin oxidoreductase, formate dehydrogenase, formyltetrahydrofolate synthetase, methenyltetrahydrofolate cyclodehydratase, methylenetetrahydrofolate dehydrogenase, and methylenetetrahydrofolate reductase. The reaction in the carbonyl branch is catalyzed in turn by the following enzymes or proteins: methyltetrahydrofolate:corrinoid protein methyltransferase (e.g., AcsE), corrinoid iron-sulfur protein, nickel-protein assembly protein (e.g., AcsF), ferredoxin, acetyl-CoA synthase, carbon monoxide dehydrogenase, and nickel-protein assembly protein (e.g., CooC) (see WO 2009 / 094485). Following the teachings and guidance provided herein for introducing a sufficient number of encoding nucleic acids to produce the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway or related downstream products, e.g., esters or amides thereof, including nucleic acids encoding the aldehyde dehydrogenases of the invention, one skilled in the art will understand that the same engineering can also be performed, at least with respect to introducing nucleic acids encoding Wood-Ljungdahl enzymes or proteins not present in the host organism. Thus, introduction of one or more encoding nucleic acids into a cell of the invention such that the modified organism contains an entire Wood-Ljungdahl pathway confers syngas utilization ability.

[0131] Additionally, the reductive (reverse) tricarboxylic acid cycle in conjunction with carbon monoxide dehydrogenase and / or hydrogenase activity can be used to convert CO, CO, and / or H to acetyl-CoA and other products, such as acetate. Organisms capable of fixing carbon via the reductive TCA pathway can utilize one or more of the following enzymes: ATP citrate lyase, citrate lyase, aconitase, isocitrate dehydrogenase, alpha-ketoglutarate:ferredoxin oxidoreductase, succinyl-CoA synthetase, succinyl-CoA transferase, fumarate reductase, fumarase, malate dehydrogenase, NAD(P)H:ferredoxin oxidoreductase, carbon monoxide dehydrogenase, and hydrogenase. Specifically, reducing equivalents extracted from CO and / or H by carbon monoxide dehydrogenase and hydrogenase are utilized to fix CO to acetyl-CoA or acetate via the reductive TCA cycle. Acetate can be converted to acetyl-CoA by enzymes such as acetyl-CoA transferase, acetate kinase / phosphotransacetylase, and acetyl-CoA synthetase. Acetyl-CoA can be converted to glyceraldehyde-3-phosphate, phosphoenolpyruvate, and pyruvate by pyruvate:ferredoxin oxidoreductase and enzymes of gluconeogenesis. Acetyl-CoA can also be converted to acetoacetyl-CoA by, for example, acetoacetyl-CoA thiolase, as disclosed herein (see FIG. 1), and fed into the 1,3-BDO pathway. Following the teachings and guidance provided herein for introducing a sufficient number of encoding nucleic acids to generate a pathway that produces a 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway or a related downstream product, e.g., an ester or amide thereof, it will be understood by those skilled in the art that the same engineering can also be performed with respect to introducing nucleic acids encoding at least one reductive TCA pathway enzyme or protein that is not present in the host organism.Thus, introduction of one or more encoding nucleic acids into a cell of the invention can be performed such that the modified organism contains a reductive TCA pathway.

[0132] Thus, given the teachings and guidance provided herein, it will be understood by those skilled in the art that non-naturally occurring cells can be generated that, when grown on a carbon source, e.g., a carbohydrate, produce and / or secrete biosynthesized compounds of the invention, including, for example, 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, as well as any of the intermediate metabolites in the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway. All that is required is to engineer one or more of the required enzyme or protein activities to achieve biosynthesis of the desired compound or intermediate, including, for example, engineering part or all of the biosynthetic pathway for 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, including, for example, the aldehyde dehydrogenases of the invention. Thus, the invention provides non-naturally occurring cells that, when grown on carbohydrates or other carbon sources, produce and / or secrete 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, as well as cells that, when grown on carbohydrates or other carbon sources, produce and / or secrete any of the intermediate metabolites shown in the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway. Cells that produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO of the present invention or related downstream products, e.g., esters or amides thereof, can initiate synthesis from an intermediate, e.g., an intermediate in the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway.

[0133] Non-naturally occurring cells of the invention are constructed using methods well known in the art, as exemplified herein, to exogenously express at least one nucleic acid encoding an aldehyde dehydrogenase of the invention and, optionally, an enzyme or protein of the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway or a related downstream product thereof, e.g., an ester or amide thereof. The enzyme or protein can be expressed in an amount sufficient to produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or a related downstream product thereof, e.g., an ester or amide thereof. It is understood that the cells of the invention are cultured under conditions sufficient to express the aldehyde dehydrogenase of the invention or to produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or a related downstream product thereof, e.g., an ester or amide thereof. Following the teachings and guidance provided herein, non-naturally occurring cells of the invention can achieve biosynthesis of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, resulting in intracellular concentrations of between about 0.1 and 300 mM or higher, e.g., 0.1 to 1.3 M or higher. Generally, intracellular concentrations of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, are between about 3 and 150 mM, particularly between about 5 and 125 mM, and more particularly between about 8 and 100 mM, including concentrations of about 10 mM, 20 mM, 50 mM, 80 mM, or higher. Intracellular concentrations between and above each of these exemplary ranges can also be achieved from non-naturally occurring cells of the invention. For example, the intracellular concentration of 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or a related downstream product thereof, e.g., an ester or amide thereof, can be between about 100 mM and 1.3 M, including about 100 mM, 200 mM, 500 mM, 800 mM, 1 M, 1.1 M, 1.2 M, 1.3 M or higher.

[0134] The cells of the present invention are cultured using well-known methods. Culture conditions can include, for example, liquid culture methods and fermentation and other large-scale culture methods. As described herein, particularly useful yields of the biosynthetic products of the present invention can be obtained under anaerobic or substantially anaerobic culture conditions.

[0135] In some embodiments, culture conditions include anaerobic or substantially anaerobic growth or maintenance conditions. Exemplary anaerobic conditions have been previously described and are well known in the art. Exemplary anaerobic conditions for fermentation processes are described herein, for example, in U.S. Publication No. 2009 / 0047719, filed August 10, 2007. Any of these conditions can be utilized with non-naturally occurring cells and other anaerobic conditions known in the art. Under such anaerobic or substantially anaerobic conditions, 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO processes can synthesize 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, at intracellular concentrations of 5-10 mM or higher, as well as all other concentrations exemplified herein. Although the above description refers to intracellular concentrations, it is understood that a 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO producing cell may produce 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or a related downstream product thereof, e.g., an ester or amide thereof, intracellularly and / or secrete the product into the medium.

[0136] As described herein, one exemplary growth condition for achieving biosynthesis of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, includes anaerobic culture or fermentation conditions. In certain embodiments, non-naturally occurring cells of the invention can be maintained, cultured, or fermented under anaerobic or substantially anaerobic conditions. Briefly, anaerobic conditions refer to an oxygen-deficient environment. Substantially anaerobic conditions include, for example, cultures, batch fermentations, or continuous fermentations in which the dissolved oxygen concentration in the medium remains between 0 and 10% of saturation. Substantially anaerobic conditions also include growing or resting cells in liquid medium or solid agar in a sealed chamber maintained in an atmosphere of less than 1% oxygen. The oxygen percentage can be maintained, for example, by sparging the culture with a N2 / CO2 mixture or other suitable non-oxygen gas(es).

[0137] The culture conditions described herein can be scaled up and continuously grown to produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, using the cells of the invention. Exemplary growth methods include, for example, fed-batch fermentation and batch separation; fed-batch fermentation and continuous separation; or continuous fermentation and continuous separation. All of these processes are well known in the art. Fermentation methods are particularly useful for the biosynthetic production of commercial quantities of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof. Generally, and similar to discontinuous culture methods, continuous and / or near-continuous production of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, involves culturing non-naturally occurring cells of the invention that produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, in nutrients and medium sufficient to maintain and / or nearly maintain exponential-phase growth. Continuous culture under such conditions can include growth or culture for, e.g., 1, 2, 3, 4, 5, 6, or 7 days or longer. Furthermore, continuous culture can also include longer periods of 1, 2, 3, 4, or 5 weeks or more, and up to several months. Alternatively, organisms of the invention can be cultured for several hours, as appropriate for a particular application. It should be understood that continuous and / or near-continuous culture conditions can also include all periods between these exemplary periods. It is further understood that the time for culturing the cells of the present invention is for a period of time sufficient to produce a sufficient amount of product for the desired purpose.

[0138] Exemplary fermentation processes include, but are not limited to, fed-batch fermentation and batch separation; fed-batch fermentation and continuous separation; and continuous fermentation and continuous separation. In an exemplary batch fermentation protocol, the production organism is grown in a suitably sized bioreactor sparged with an appropriate gas. Under anaerobic conditions, the culture is sparged with an inert gas or combination of gases, such as nitrogen, N2 / CO2 mixtures, argon, helium, etc. As the cells grow and utilize the carbon source, additional carbon source and / or other nutrients are fed to the bioreactor at a rate that approximately balances the consumption of the carbon source and / or nutrients. The temperature of the bioreactor is maintained at a desired temperature, generally within the range of 22-37°C, although the temperature may be maintained at a higher or lower temperature depending on the growth characteristics of the production organism and / or the desired conditions for the fermentation process. Growth continues for a desired period of time to achieve the desired characteristics of the culture in the fermenter, such as cell density, product concentration, etc. In a batch fermentation process, the duration of the fermentation generally ranges from a few hours to several days, e.g., 8 to 24 hours, or 1, 2, 3, 4, or 5 days, or up to a week, depending on the desired culture conditions. The pH may or may not be controlled, as desired; if the pH is not controlled, the culture typically decreases to a pH of 3 to 6 by the end of the run. Upon completion of the culture period, the fermenter contents may be passed through a cell separation unit, e.g., a centrifuge, filtration unit, etc., to remove cells and cell debris. If a desired product is expressed within the cells, the cells may be lysed or enzymatically or chemically disrupted, if desired, before or after separation from the fermentation broth to release additional product. The fermentation broth may be transferred to a product separation unit. Product isolation is performed by standard separation procedures utilized in the art for separating the desired product from a dilute aqueous solution.Such methods include, but are not limited to, liquid-liquid extraction using a water-immiscible organic solvent (e.g., toluene or other suitable solvents including, but not limited to, diethyl ether, ethyl acetate, tetrahydrofuran (THF), methylene chloride, chloroform, benzene, pentane, hexane, heptane, petroleum ether, methyl tertiary butyl ether (MTBE), dioxane, dimethylformamide (DMF), dimethyl sulfoxide (DMSO), etc.) to provide an organic solution of the product, depending on the chemical characteristics of the product of the fermentation process, and, where appropriate, standard distillation methods.

[0139] In an exemplary fully continuous fermentation protocol, the production organism is generally initially grown in a batch mode to achieve a desired cell density. Once the carbon source and / or other nutrients are exhausted, a feed medium of the same composition is continuously fed at a desired rate, and fermentation broth is removed at the same rate. Under such conditions, the product concentration and cell density within the bioreactor are generally kept constant. The temperature of the fermenter is maintained at a desired temperature, as described above. During the continuous fermentation phase, it is generally desirable to maintain a suitable pH range to optimize production. pH may be monitored and maintained using routine methods, including the addition of a suitable acid or base to maintain the desired pH range. The bioreactor is operated continuously for an extended period of time, generally at least one week to several weeks and up to one month or longer, as appropriate and desired. The fermentation broth and / or culture are optionally monitored periodically, including up to daily sampling, to ensure consistency of product concentration and / or cell density. In a continuous mode, fermenter contents are constantly removed as new feed medium is added. The effluent stream containing cells, medium, and product is generally subjected to a continuous product separation procedure, with or without removal of cells and cell debris, as desired. Continuous separation methods utilized in the art, including but not limited to continuous liquid-liquid extraction with a water-immiscible organic solvent (e.g., toluene or other suitable solvents, including but not limited to, diethyl ether, ethyl acetate, tetrahydrofuran (THF), methylene chloride, chloroform, benzene, pentane, hexane, heptane, petroleum ether, methyl tertiary butyl ether (MTBE), dioxane, dimethylformamide (DMF), dimethyl sulfoxide (DMSO), etc.), standard continuous distillation methods, etc., or other methods known in the art, can be used to separate the product from the dilute aqueous solution.

[0140] Fermentation procedures are well known in the art. Briefly, fermentation for the biosynthetic production of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, can be utilized, for example, in fed-batch fermentation and batch separation; fed-batch fermentation and continuous separation; or continuous fermentation and continuous separation. Examples of batch and continuous fermentation procedures are well known in the art and are described herein.

[0141] In addition to the fermentation procedures described herein that utilize the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, e.g., esters or amides thereof, of the present invention for the continuous production of substantial quantities of 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, e.g., esters or amides thereof, the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, e.g., esters or amides thereof, can also, if desired, be simultaneously subjected to, for example, chemical synthesis and / or enzymatic procedures to convert the products to other compounds, or the products can be separated from the fermentation culture and sequentially subjected to chemical and / or enzymatic conversions to convert the products to other compounds.

[0142] In addition to the culture and fermentation conditions disclosed herein, growth conditions for expression of the aldehyde dehydrogenase of the invention or biosynthesis of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, can also include the addition of an osmoprotectant to the culture conditions. In certain embodiments, non-naturally occurring cells of the invention can be maintained, cultured, or fermented as described herein in the presence of an osmoprotectant. Briefly, an osmoprotectant refers to a compound that acts as an osmolyte and helps the cells described herein survive osmotic stress. Osmoprotectants include, but are not limited to, betaine, amino acids, and the sugar trehalose. Non-limiting examples of such osmoprotectants include glycine betaine, praline betaine, dimethylthetin, dimethylsulfonioproprionate, 3-dimethylsulfonio-2-methylproprionate, pipecolic acid, dimethylsulfonioacetate, choline, L-carnitine, and ectoine. In one embodiment, the osmoprotectant is glycine betaine. Those skilled in the art will appreciate that the amount and type of osmoprotectant suitable for protecting the cells described herein from osmotic stress will depend on the cells used. The amount of osmoprotectant in the culture conditions can be, for example, about 0.1 mM or less, about 0.5 mM or less, about 1.0 mM or less, about 1.5 mM or less, about 2.0 mM or less, about 2.5 mM or less, about 3.0 mM or less, about 5.0 mM or less, about 7.0 mM or less, about 10 mM or less, about 50 mM or less, about 100 mM or less, or about 500 mM or less.

[0143] In some embodiments, carbon feedstocks and other cellular uptake sources, such as phosphate, ammonia, sulfate, chloride, and other halogens, can be chosen to vary the isotopic distribution of atoms present in 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or its related downstream products, such as esters or amides thereof, or any 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates. The various carbon feedstocks and other uptake sources listed above are collectively referred to herein as "uptake sources." The incorporation source can provide isotopic enrichment of any atom present in the product 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, or in a 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediate, or in a by-product produced in a reaction that deviates from the 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway. Isotopic enrichment can be achieved for any target atom, including, for example, carbon, hydrogen, oxygen, nitrogen, sulfur, phosphorus, chloride, or other halogens.

[0144] In some embodiments, an incorporation source can be selected to vary the ratio of carbon-12, carbon-13, and carbon-14. In some embodiments, an incorporation source can be selected to vary the ratio of oxygen-16, oxygen-17, and oxygen-18. In some embodiments, an incorporation source can be selected to vary the ratio of hydrogen, deuterium, and tritium. In some embodiments, an incorporation source can be selected to vary the ratio of nitrogen-14 and nitrogen-15. In some embodiments, an incorporation source can be selected to vary the ratio of sulfur-32, sulfur-33, sulfur-34, and sulfur-35. In some embodiments, an incorporation source can be selected to vary the ratio of phosphorus-31, phosphorus-32, and phosphorus-33. In some embodiments, an incorporation source can be selected to vary the ratio of chlorine-35, chlorine-36, and chlorine-37.

[0145] In some embodiments, the isotope ratio of a target atom can be varied to a desired ratio by selecting one or more incorporation sources. Incorporation sources can be obtained from natural sources, such as those found in nature, or artificial sources, and those skilled in the art can select natural sources, artificial sources, or a combination thereof to achieve a desired ratio of isotopes of a target atom. Examples of artificial incorporation sources include, for example, incorporation sources obtained at least in part from chemical synthesis reactions. Such isotope-enriched incorporation sources can be purchased commercially or prepared in a laboratory, and / or can be mixed with natural sources of incorporation sources as needed to achieve a desired ratio of isotopes. In some embodiments, the isotope ratio of a target atom of an incorporation source can be achieved by selecting an incorporation source of a desired origin, such as one found in nature. For example, as discussed herein, natural sources can be biobased sources, derived from or synthesized by biological organisms, or can be sources such as petroleum-based products or the atmosphere. In some such embodiments, the carbon source can be selected from, for example, fossil fuel-derived carbon sources that may be relatively depleted in carbon-14, or environmental or atmospheric carbon sources, e.g., CO, that may possess greater amounts of carbon-14 than their petroleum-derived counterparts.

[0146] Carbon-14, or radiocarbon, an unstable isotope of carbon, is present in 10% of Earth's atmosphere. 12 It makes up roughly one in every 10 carbon atoms and has a half-life of about 5700 years. Carbon stocks are reduced in the upper atmosphere by cosmic rays and ordinary nitrogen ( 14 Carbon-14 is replenished by nuclear reactions involving carbon-14 (N). Fossil fuels do not contain carbon-14 because it decayed long ago. Burning fossil fuels reduces the proportion of carbon-14 in the atmosphere, which is the so-called "Suess effect."

[0147] The method for determining the ratio of the isotopes of atoms in a compound is well known to those skilled in the art.Isotopic enrichment can be easily evaluated by mass spectrometry using techniques known in the art, for example, accelerator mass spectrometry (AMS), stable isotope ratio mass spectrometry (SIRMS) and site-specific natural isotope distribution by nuclear magnetic resonance (SNIF-NMR).Such mass spectrometry techniques can be combined with separation techniques, for example, liquid chromatography (LC), high performance liquid chromatography (HPLC) and / or gas chromatography.

[0148] In the case of carbon, ASTM D6866 was developed in the United States by the American Society for Testing and Materials (ASTM) International as a standardized analytical method for determining the biobased content of solid, liquid, and gas samples using radiocarbon dating. This standard method relies on the use of radiocarbon dating to determine the biobased content of a product. ASTM D6866 was first published in 2004, and the currently valid version of this standard method is ASTM D6866-11 (valid since April 1, 2011). Radiocarbon dating techniques, including those described herein, are well known to those skilled in the art.

[0149] The biobased content of a compound is determined by carbon-14( 14 C) vs. carbon-12 ( 12 Specifically, the fraction modern (Fm) is calculated from the formula: Fm = (SB) / (MB), where B, S, and M are the values ​​of the blank, sample, and modern reference, respectively. 14 C / 12 The current ratio indicates the ratio of the sample from the "present" 14 C / 12 It is a measure of the deviation of the C ratio. 13 C VPDBIt is defined as 95% of the radiocarbon concentration (in 1950 AD) of the National Bureau of Standards (NBS) Oxalic Acid I (i.e., Standard Reference Material (SRM) 4990b), normalized to δ = -19 / mil (Olsson, The use of Oxalic acid as a Standard, Radiocarbon Variations and Absolute Chronology, Nobel Symposium, 12th Proc., John Wiley & Sons, New York (1970)). For example, mass spectrometry results measured by ASM are calculated using the internationally agreed definition, i.e., δ 13 C VPDB = -19 / mil, calculated using 0.95 times the specific activity of NBS oxalate I (SRM 4990b), normalized to 1.176 ± 0.010 × 10 -12 of (1950 AD) 14 C / 12 The ratio is equivalent to the absolute ratio of C (Karlen et al., Arkiv Geofysik, 4:465-471 (1968)). Standard calculations involve differential incorporation of one isotope with respect to another, e.g., C 12 >C 13 >C 14 These corrections take into account the preferential uptake of δ 13 is reflected as the corrected Fm.

[0150] An oxalic acid standard (SRM 4990b or HOx1) was prepared from the 1955 sugar beet crop. 1,000 pounds were prepared, but this oxalic acid standard is no longer commercially available. An oxalic acid II standard (HOx2; NIST designation SRM 4990C) was prepared from molasses from the 1977 French beet crop. In the early 1980s, a group of 12 laboratories measured the ratio of the two standards. The activity ratio of oxalic acid II to 1 is 1.2933 ± 0.001 (weighted average). The isotopic ratio of HOxII is -17.8 / mil. ASTM D6866-11 suggests the use of the available oxalic acid II standard SRM 4990C (Hox2) for the current standard (see discussion of the original oxalic acid standard versus currently available oxalic acid standards in Mann, Radiocarbon, Vol. 25(No. 2):519-527 (1983)). Fm=0% indicates a complete lack of carbon-14 atoms in the material and thus indicates a fossil (e.g., petroleum-based) carbon source. Fm=100%, after correcting for post-1950 release of carbon-14 into the atmosphere from nuclear bomb testing, indicates an entirely current carbon source. As described herein, such "current" sources include bio-based sources.

[0151] The percent modern carbon (pMC), as described in ASTM D6866, can be greater than 100%. This is due to the continuing, albeit diminishing, effects of the 1950s nuclear testing program, which resulted in significant atmospheric carbon-14 enrichment, as described in ASTM D6866-11. Because all sample carbon-14 activity is referenced to "pre-bomb" standards and because almost all new biobased products are produced in post-bomb environments, all pMC values ​​(after correction for isotopic ratios) must be multiplied by 0.95 (as of 2010) to better reflect the sample's true biobased content. A biobased content greater than 103% suggests that either analytical error has occurred or the biobased carbon source is more than several years old.

[0152] ASTM D6866 quantifies the biobased content relative to the total organic content of a material, disregarding the inorganic carbon and other non-carbon-containing substances present. For example, based on ASTM D6866, a product that is 50% starch-based material and 50% water is considered to have a biobased content = 100% (100% biobased, 50% organic content). In another example, a product that is 50% starch-based material, 25% petroleum-based, and 25% water has a biobased content = 66.7% (75% organic content, but only 50% of the product is biobased). In another example, a product that is 50% organic carbon and a petroleum-based product is considered to have a biobased content = 0% (50% organic carbon, but derived from fossil sources). Thus, based on well-known methods and known standards for determining the biobased content of a compound or material, one of skill in the art can readily determine the biobased content of a compound or material and / or prepare downstream products utilizing the compounds or materials of the present invention having a desired biobased content.

[0153] The application of carbon-14 dating techniques to quantify the biobased content of materials is known in the art (Currie et al., Nuclear Instruments and Methods in Physics Research B, 172:281-287 (2000)). For example, carbon-14 dating has been used to quantify the biobased content in terephthalate-containing materials (Colonna et al., Green Chemistry, 13:2543-2548 (2011)). In particular, polypropylene terephthalate (PPT) polymers derived from renewable 1,3-propanediol and petroleum-derived terephthalic acid yielded Fm values ​​of nearly 30% (i.e., 3 / 11 of the polymer carbon comes from renewable 1,3-propanediol and 8 / 11 comes from the fossil end-member terephthalic acid) (Currie et al., supra, 2000). In contrast, polybutylene terephthalate polymers obtained from both renewable 1,4-butanediol and renewable terephthalic acid have yielded biobased contents of over 90% (Colonna et al., supra, 2011).

[0154] Thus, in some embodiments, the invention provides 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, such as esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediates produced by the cells of the invention, having a ratio of carbon-12, carbon-13 and carbon-14 that reflects an uptake source of atmospheric carbon, also referred to as ambient carbon. For example, in some embodiments, 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or a downstream product related thereto, such as an ester or amide thereof, or a 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediate, can have an Fm value of at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100%. In some such embodiments, the uptake source is CO2. In some embodiments, the invention provides 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates, having carbon-12, carbon-13, and carbon-14 ratios that reflect a petroleum-based carbon uptake source. In this embodiment, 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or a downstream product related thereto, such as an ester or amide thereof, or a 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediate, can have an Fm value of less than 95%, less than 90%, less than 85%, less than 80%, less than 75%, less than 70%, less than 65%, less than 60%, less than 55%, less than 50%, less than 45%, less than 40%, less than 35%, less than 30%, less than 25%, less than 20%, less than 15%, less than 10%, less than 5%, less than 2%, or less than 1%.In some embodiments, the invention provides 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates, having ratios of carbon-12, carbon-13, and carbon-14 that are obtained by combining atmospheric carbon uptake sources with petroleum-based uptake sources. The use of such combinations of uptake sources provides one way in which the ratios of carbon-12, carbon-13, and carbon-14 can be varied, with the respective ratios reflecting the ratios of the uptake sources.

[0155] Additionally, the present invention relates to biologically produced 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products thereof, such as esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediates disclosed herein, and products derived therefrom, wherein the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products thereof, such as esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediates have carbon-12, carbon-13, and carbon-14 isotope ratios that are approximately the same as those found in environmental CO. For example, in some embodiments, the present invention provides bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, or bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO intermediates, that have carbon-12 to carbon-13 to carbon-14 isotope ratios that are approximately the same as those found in environmental CO, or any of the other ratios disclosed herein. As disclosed herein, the product may have a carbon-12 to carbon-13 to carbon-14 isotope ratio approximately equal to that of environmental CO, or any of the ratios disclosed herein, and the product may be produced from bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or a related downstream product, e.g., an ester or amide thereof, or a bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediate, as disclosed herein, with the understanding that the bio-derived product is chemically modified to produce the final product. Methods for chemically modifying a bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or a related downstream product, e.g., an ester or amide thereof, or a 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO intermediate to produce a desired product are well known to those skilled in the art, as described herein.The present invention relates to 3-HBal and / or 1,3-BDO or downstream products related thereto, such as plastics that may be based on esters or amides thereof, elastic fibers, polyurethanes, polyesters containing polyhydroxyalkanoates, nylons, organic solvents, polyurethane resins, polyester resins, blood glucose lowering agents, butadiene and / or butadiene-based products, and 4-HBal and / or 1,4-BDO or downstream products related thereto, such as plastics that may be based on esters or amides thereof, elastic fibers, polyurethanes, polyesters containing polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or copolymers thereof, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT), and spandex, elastane, or Lycra ( Further provided herein are plastics, elastic fibers, polyurethanes, polyesters comprising polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or copolymers thereof, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, polyurethane-polyurea copolymers referred to as elastane or Lycra™, nylons, organic solvents, polyurethane resins, polyester resins, blood glucose lowering agents, butadiene and / or butadiene-based products produced directly from or in combination with bio-sourced 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, such as esters or amides thereof, or bio-sourced 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediates, as disclosed herein.Methods for producing butadiene and / or butadiene-based products have been previously described (see, e.g., WO 2010 / 127319, WO 2013 / 036764, U.S. Pat. No. 9,017,983, US 2013 / 0066035, WO / 2012 / 018624, US 2012 / 0021478, each of which is incorporated herein by reference). 1,3-BDO can be converted to an ester either in vivo or in vitro by reaction with an acid, for example, using a lipase. Such esters may have nutritional supplement, pharmaceutical, and food applications, and are advantageous when the R-form of 1,3-BDO is used because it is the form most commonly utilized by both animals and humans as an energy source (compared to the S-form or racemic mixture) (e.g., ketone esters, such as (R)-3-hydroxybutyl-R-1,3-butanediol monoester and (R)-3-hydroxybutyric acid glycerol monoester or diester (which have Generally Recognized As Safe (GRAS) approval in the United States). Ketone esters can be delivered orally, and the esters release R-1,3-butanediol for use by the body (see, e.g., WO2013150153). Methods for producing amides are well known in the art (see, e.g., Goswami and Van Lanen, Mol. Biosyst. 11(2):338-353 (2015)).

[0156] Thus, the present invention is particularly useful for providing microorganisms that provide R-1,3-butanediol with improved enzymatic pathways and improved 1,3-BDO compositions, i.e., highly enriched or essentially enantiomerically pure, with improved purity qualities for by-products. 1,3-BDO has additional food-related uses, including as a direct food source, food ingredient, flavor, solvent or solubilizer for flavorings, stabilizer, emulsifier, and antimicrobial and preservative. 1,3-BDO is used in the pharmaceutical industry as a parenteral drug solvent. 1,3-BDO finds use in cosmetics as an emollient, a humectant, a component that prevents crystallization of insoluble components, a solubilizer for less water-soluble components such as fragrances, and an antimicrobial and preservative. For example, it can be used as a humectant in hairsprays and setting lotions, among others; it reduces fragrance loss from essential oils, preserves against microbial spoilage, and is used as a solvent for benzoates. 1,3-BDO can be used at concentrations of 0.1% to 50%, and even below 0.1% and above 50%. It is used in hair and bath products, eye and facial cosmetics, fragrances, body cleansing products, and shaving and skin care preparations (see, for example, the Cosmetic Ingredient Review Board report: "Final Report on the Safety Assessment of Butylene Glycol, Hexylene Glycol, Ethoxydiglycol, and Dipropylene Glycol," Journal of the American College of Toxicology, Vol. 4, No. 5, 1985, incorporated herein by reference). This report provides specific uses and concentrations of 1,3-BDO in cosmetics; see, for example, Table 2 of the report entitled "Product Formulation Data."

[0157] In one embodiment, the invention provides a medium comprising biogenic 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO, wherein the biogenic 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO has a carbon-12, carbon-13, and carbon-14 isotope ratio that reflects a source of atmospheric carbon dioxide uptake, and wherein the biogenic 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO is produced by or in a cell lysate of the invention or method of the invention. In one embodiment, the medium is separated from the cells.

[0158] In one embodiment, the invention provides 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) produced by a cell or in a cell lysate of the invention or method of the invention, having a carbon-12, carbon-13, and carbon-14 isotope ratio that reflects the source of atmospheric carbon dioxide uptake. In one embodiment, the 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO has an Fm value of at least 80%, at least 85%, at least 90%, at least 95%, or at least 98%.

[0159] In one embodiment, the invention provides 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) produced by cells or in cell lysates of the invention or methods of the invention. In one embodiment, the invention provides 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) having a carbon-12, carbon-13, and carbon-14 isotope ratio that reflects a source of atmospheric carbon dioxide uptake, wherein the 3-HBal and / or 1,3-BDO produced by cells or in cell lysates of the invention or methods of the invention is enriched in the R-enantiomer. In one embodiment, the 3-HBal and / or 1,3-BDO have an Fm value of at least 80%, at least 85%, at least 90%, at least 95%, or at least 98%.

[0160] In one embodiment, the invention provides 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) produced by a cell or in a cell lysate of the invention or a method of the invention, wherein the 3-HBal and / or 1,3-BDO is enriched in the R-enantiomer. In one embodiment, the R-form of the 3-HBal and / or 1,3-BDO is greater than 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%. In one embodiment, the 3-HBal and / or 1,3-BDO is 55% or more R-enantiomer, 60% or more R-enantiomer, 65% or more R-enantiomer, 70% or more R-enantiomer, 75% or more R-enantiomer, 80% or more R-enantiomer, 85% or more R-enantiomer, 90% or more R-enantiomer, or 95% or more R-enantiomer, and can be highly chemically pure, e.g., 99% or more, e.g., 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, 99.1% or more, 99.2% or more, 99.3% or more, 99.4% or more, 99.5% or more, 99.6% or more, 99.7% or more, 99.8% or more, or 99.9% or more R-enantiomer.

[0161] In one embodiment, the invention provides a composition comprising 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO produced by or in a cell lysate of the invention or method of the invention, and a compound other than 3-HBal and / or 1,3-BDO or 4-HBal or 1,4-BDO, respectively. In one embodiment, the compound other than 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO is part of a cell producing 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO, respectively, or expressing a polypeptide of the invention.

[0162] In one embodiment, the invention provides a composition comprising 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO produced by or in a cell lysate, or a cell lysate or culture supernatant of cells producing 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO, of the invention or a method of the invention.

[0163] In one embodiment, the present invention provides a product comprising 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO produced by cells or in cell lysates of the present invention or methods, the product being a plastic, elastic fiber, polyurethane, polyester, polyhydroxyalkanoate, poly-4-hydroxybutyrate (P4HB) or copolymers thereof, poly(tetramethylene ether) glycol (PTMEG), polybutylene terephthalate (PBT), polyurethane-polyurea copolymer, nylon, organic solvent, polyurethane resin, polyester resin, hypoglycemic agent, butadiene, or butadiene-based product. In one embodiment, the product is a cosmetic or a food additive. In one embodiment, the product contains at least 0.1%, at least 0.5%, at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% bio-sourced 3-HBal and / or 1,3-BDO or bio-sourced 4-HBal and / or 1,4-BDO. In one embodiment, the product comprises a portion of the produced 3-HBal and / or 1,3-BDO or produced 4-HBal and / or 1,4-BDO as a repeating unit. In one embodiment, the invention provides a shaped product obtained by shaping a product made with or derived from 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO produced by cells or in cell lysates of the invention or methods of the invention.

[0164] The present invention further provides compositions comprising biologically-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, and compounds other than biologically-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof. The non-biologically-derived compounds may be cellular fractions, e.g., trace cellular fractions, of non-naturally-occurring cells of the invention having a pathway to produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, or may be fermentation broths or media, or purified or partially purified fractions thereof, produced in the presence of non-naturally-occurring cells of the invention having a pathway to produce 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof. The compositions may contain reduced levels of by-products, for example, if produced by an organism with reduced by-product formation, as disclosed herein. The compositions can include, for example, biologically derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products thereof, such as esters or amides thereof, or cell lysates or culture supernatants of cells of the invention.

[0165] 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as their esters or amides, are chemicals used in commercial and industrial applications, including, but not limited to, the production of plastics, elastic fibers, polyurethanes, polyesters containing polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or its copolymers, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT), and spandex, polyurethane-polyurea copolymers known as elastane or Lycra™, nylons, organic solvents, polyurethane resins, polyester resins, hypoglycemic agents, and butadiene and / or butadiene-based products. Additionally, 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO are also used as raw materials in the production of a wide range of products, including plastics, elastic fibers, polyurethanes, polyesters containing polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or its copolymers, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT), and spandex, polyurethane-polyurea copolymers known as elastane or Lycra™, nylons, organic solvents, polyurethane resins, polyester resins, blood sugar lowering agents, butadiene, and / or butadiene-based products.Thus, in some embodiments, the invention provides bio-based plastics, elastic fibers, polyurethanes, polyesters comprising polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or copolymers thereof, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, elastane or the polyurethane-polyurea copolymer known as Lycra™, nylon, organic solvents, polyurethane resins, polyester resins, blood glucose lowering agents, butadiene and / or butadiene-based products, e.g., comprising one or more bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products thereof, e.g., esters or amides thereof, or bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates, produced by a non-naturally occurring cell of the invention expressing an aldehyde dehydrogenase of the invention or produced using the methods disclosed herein.

[0166] As used herein, the term "biogenic" means obtained from or synthesized by a biological organism and can be produced by a biological organism and therefore can be considered a renewable resource. Such biological organisms, particularly the cells of the invention disclosed herein, can utilize feedstock or biomass, e.g., sugars or carbohydrates, obtained from agricultural, plant, bacterial, or animal sources. Alternatively, the biological organism can utilize atmospheric carbon. As used herein, the term "biobased" refers to a product, as described above, that is composed in whole or in part of the biologically derived compounds of the invention. Biobased or bioderived products are in contrast to petroleum-derived products, which are obtained from or synthesized from petroleum or petrochemical feedstocks.

[0167] In some embodiments, the present invention provides plastics, elastic fibers, polyurethanes, polyesters including polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or copolymers thereof, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, elastane or the polyurethane-polyurea copolymer known as Lycra™, nylon, organic solvents, polyurethane resins, polyester resins, blood glucose lowering agents, butadiene and / or butadiene-based products, including bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, e.g., esters or amides thereof, or bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates. Bio-derived 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediates include all or a portion of 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, such as esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediates used in the production of plastics, elastic fibers, polyurethanes, polyhydroxyalkanoate-containing polyesters such as poly-4-hydroxybutyrate (P4HB) or copolymers thereof, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, elastane or the polyurethane-polyurea copolymer known as Lycra™, nylon, organic solvents, polyurethane resins, polyester resins, hypoglycemic agents, butadiene and / or butadiene-based products.For example, the final plastics, elastic fibers, polyurethanes, polyesters containing polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or its copolymers, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, polyurethane-polyurea copolymers known as elastane or Lycra™, nylon, organic solvents, polyurethane resins, polyester resins, hypoglycemic agents, butadiene and / or butadiene-based products are generally made from plastics, elastic fibers, polyurethanes, polyesters containing polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or its copolymers, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, elastane or polyurethane-polyurea copolymers known as Lycra™, nylon, organic solvents, polyurethane resins, polyester resins, hypoglycemic agents, butadiene and / or butadiene-based products are generally made from plastics, elastic fibers, polyurethanes, polyesters containing polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or its copolymers, The following may contain bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products thereof, such as esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates, or portions thereof, resulting from the manufacture of butadiene and / or butadiene-based products: polyesters containing 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO; poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide); polybutylene terephthalate (PBT) and spandex; polyurethane-polyurea copolymers called elastane or Lycra™; nylon; organic solvents; polyurethane resins; polyester resins; blood glucose lowering agents; butadiene and / or butadiene-based products; or related downstream products, such as esters or amides thereof, or 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates, or portions thereof.Such production can include chemically reacting (e.g., chemical conversion, chemical functionalization, chemical coupling, oxidation, reduction, polymerization, copolymerization, etc.) bio-based 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof, or bio-based 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates into a final plastic, elastic fiber, polyurethane, polyester containing polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or its copolymers, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, elastane, or a polyurethane-polyurea copolymer called Lycra™, nylon, organic solvents, polyurethane resins, polyester resins, blood glucose lowering agents, butadiene, and / or butadiene-based products.Thus, in some embodiments, the present invention provides a method for producing at least 2%, at least 3%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or 100% bio-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products thereof, such as esters or amides thereof, or bio-derived 3-HBal, 1,3-BDO, 4-BDO, as disclosed herein. The Company offers bio-based plastics, including HBal or 1,4-BDO pathway intermediates; elastic fibers; polyurethanes; polyesters containing polyhydroxyalkanoates, such as poly-4-hydroxybutyrate (P4HB) or its copolymers; poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide); polybutylene terephthalate (PBT) and spandex; polyurethane-polyurea copolymers known as elastane or Lycra™; nylons; polyurethane resins; polyester resins; blood glucose lowering agents; and butadiene and / or butadiene-based products.

[0168] Additionally, in some embodiments, the invention provides compositions comprising a compound other than the biologically-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO disclosed herein or a related downstream product thereof, such as an ester or amide thereof, or a 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediate, and a biologically-derived 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or a related downstream product thereof, such as an ester or amide thereof, or a 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediate. For example, in some embodiments, the present invention provides bio-based plastics, elastic fibers, polyurethanes, polyesters including polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or copolymers thereof, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, polyurethane-polyurea copolymers known as elastane or Lycra™, nylon, organic solvents, polyurethane resins, polyester resins, hypoglycemic agents, butadiene and / or butadiene-based products, including butadiene-based polymers, such as 3-HBal, 1,3-BDO, 4-HBal, or 5-HBal-based polymers, and the like, used in the production thereof. The 1,4-BDO or related downstream product, e.g., an ester or amide thereof, or a 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediate is a combination of a biologically derived 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream product, e.g., an ester or amide thereof, or a 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediate, and a petroleum-derived 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream product, e.g., an ester or amide thereof, or a 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway intermediate.For example, bio-based plastics, elastic fibers, polyurethanes, polyesters including polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or copolymers thereof, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, polyurethane-polyurea copolymers known as elastane or Lycra™, nylon, organic solvents, polyurethane resins, polyester resins, blood glucose lowering agents, butadiene and / or butadiene-based products, at least a portion of which may be produced by the cells disclosed herein. The bio-based precursors can be produced using 50% bio-based 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, e.g., esters or amides thereof, and 50% petroleum-derived 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO or related downstream products, e.g., esters or amides thereof, or any other desired ratio, e.g., 60% / 40%, 70% / 30%, 80% / 20%, 90% / 10%, 95% / 5%, 100% / 0%, 40% / 60%, 30% / 70%, 20% / 80%, 10% / 90% bio-based precursors / petroleum-derived precursors, so long as the bio-based precursors include bio-based precursors.It is understood that methods for using the bio-based 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO of the present invention or related downstream products, such as esters or amides thereof, or bio-based 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO pathway intermediates to produce plastics, elastic fibers, polyurethanes, polyesters containing polyhydroxyalkanoates such as poly-4-hydroxybutyrate (P4HB) or copolymers thereof, poly(tetramethylene ether) glycol (PTMEG) (PTMO, also known as polytetramethylene oxide), polybutylene terephthalate (PBT) and spandex, the polyurethane-polyurea copolymer known as elastane or Lycra™, nylon, organic solvents, polyurethane resins, polyester resins, blood glucose lowering agents, butadiene, and / or butadiene-based products are well known in the art.

[0169] To obtain better producers, metabolic modeling can be used to optimize growth conditions. Modeling can be used to design gene knockouts that further optimize pathway utilization (see, for example, U.S. Patent Publications US2002 / 0012939, US2003 / 0224363, US2004 / 0029149, US2004 / 0072723, US2003 / 0059792, US2002 / 0168654 and US2004 / 0009466, and U.S. Patent No. 7,127,379). Modeling analysis allows reliable prediction of the effects on cell growth that shift metabolism toward more efficient production of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO or related downstream products, such as esters or amides thereof.

[0170] One computational method for identifying and designing metabolic modifications that favor the biosynthesis of a desired product is the OptKnock computational framework (Burgard et al., Biotechnol. Bioeng. 84:647-657 (2003)). OptKnock is a metabolic modeling and simulation program that suggests gene deletion or disruption strategies that result in genetically stable microorganisms that overproduce a target product. Specifically, the framework examines the complete metabolic and / or biochemical network of a microorganism to suggest genetic manipulations that force the desired biochemical to become an essential by-product of cell growth. By coupling biochemical production with cell growth through strategically placed gene deletions or other functional gene disruptions, growth selection pressure applied to engineered strains after extended periods in bioreactors results in improved performance as a result of forced, growth-coupled biochemical production. Finally, once gene deletions are constructed, the genes selected by OptKnock are completely removed from the genome, making it highly unlikely that the engineered strains will revert to their wild-type state. Thus, this computational method can be used to identify alternative pathways that lead to the biosynthesis of a desired product, or can be used with non-naturally occurring cells for further optimization of the biosynthesis of a desired product.

[0171] Briefly, OptKnock is a term used herein to refer to computational methods and systems for modeling cellular metabolism. The OptKnock program relates to a model and method framework that incorporates specific constraints into flux balance analysis (FBA) models. These constraints include, for example, qualitative kinetic information, qualitative regulatory information, and / or DNA microarray experimental data. OptKnock also computes solutions to various metabolic problems, for example, by tightening flux boundaries induced through flux balance models and then investigating the performance limits of metabolic networks in the presence of gene additions or deletions. The OptKnock computational framework allows for the construction of model formulations that allow for efficient querying of the performance limits of metabolic networks and provides methods for solving the resulting mixed integer linear programming problems. The metabolic modeling and simulation method referred to herein as OptKnock is described, for example, in U.S. Patent Publication No. 2002 / 0168654, filed January 10, 2002, International Patent No. PCT / US02 / 00660, filed January 10, 2002, and U.S. Publication No. 2009 / 0047719, filed August 10, 2007.

[0172] Another computational method for identifying and designing metabolic changes that favor the biosynthetic production of a product is a metabolic modeling and simulation system called SimPheny®. This computational method and system are described, for example, in U.S. Patent Publication No. 2003 / 0233218, filed June 14, 2002, and International Patent Application No. PCT / US03 / 18838, filed June 13, 2003. SimPheny® is a computational system that can be used to generate a network model in silico and simulate the flux of mass, energy, or charge through the chemical reactions of a biological system to define a solution space that includes all possible functions of the chemical reactions in the system, thereby determining the range of activities that are permissible for the biological system. This approach is referred to as constraint-based modeling, because the solution space is defined by constraints such as the known stoichiometries of the reactions involved, as well as thermodynamic and capacity constraints of the reactions related to the maximum flux through the reactions. The space defined by these constraints can be explored to determine the phenotypic capabilities and behavior of the biological system or its biochemical components.

[0173] These computational methods are consistent with biological reality because biological systems are flexible and can reach the same result in many different ways. Biological systems are designed through evolutionary mechanisms that are limited by fundamental constraints that all biological systems must face. Constraint-based modeling strategies therefore embrace these general realities. Furthermore, the ability to continuously impose further constraints on a network model through constraint tightening results in a reduction in the size of the solution space, thereby increasing the accuracy with which physiological performance or phenotype can be predicted.

[0174] Given the teachings and guidance provided herein, those skilled in the art can apply various computational frameworks for metabolic modeling and simulation to design and implement the biosynthesis of desired compounds in host cells. Such metabolic modeling and simulation methods include, for example, the above-exemplified computer systems, such as SimPheny® and OptKnock. For purposes of illustrating the present invention, some methods are described herein in relation to the OptKnock computational framework for modeling and simulation. Those skilled in the art will know how to apply the identification, design, and implementation of metabolic changes using OptKnock to any of such other metabolic modeling and simulation computational frameworks and methods known in the art.

[0175] The above method provides a set of metabolic reactions to disrupt. Elimination or metabolic modification of each reaction in the set can result in a desired product as an essential product during the growth phase of the organism. Since the reactions are known, solving the bilevel OptKnock problem also provides one or more associated genes encoding one or more enzymes that catalyze each reaction in the reaction set. Identification of the reaction set and their corresponding genes encoding the enzymes involved in each reaction is generally an automated process, achieved through correlation of the reactions with a reaction database containing relationships between enzymes and their encoding genes.

[0176] Once identified, in target cell or organism, the reaction set to be disrupted to achieve the production of desired product is carried out by functional disruption of at least one gene encoding each metabolic reaction in the set.One particularly useful way to achieve functional disruption of reaction set is by deleting each coding gene.However, in some cases, it may be beneficial to disrupt reaction by other genetic abnormalities, including, for example, mutation, deletion of regulatory regions such as promoters or cis-binding sites for regulatory factors, or cutting of coding sequences at any of several locations.For example, when it is desired to quickly evaluate the linkage of product or when gene reversion is less likely to occur, these latter abnormalities that result in the deletion of less than the entire gene set may be useful.

[0177] To identify further productive solutions to the above two-tiered OptKnock problem that result in additional reaction sets or metabolic modifications to be disrupted that can lead to biosynthesis, including growth-coupled biosynthesis of the desired product, an optimization method called integer cuts can be implemented. This method proceeds by iteratively solving the OptKnock problem exemplified above, incorporating additional constraints called integer cuts at each iteration. The integer cut constraints effectively prevent the solution from selecting the exact same reaction set identified in any previous iteration that unavoidably couples product biosynthesis to growth. For example, if a previously identified growth-coupled metabolic modification identifies reactions 1, 2, and 3 for disruption, the following constraint prevents the same reactions from being simultaneously considered in subsequent solutions. Integer cut methods are well known in the art and can be found, for example, in Burgard et al., Biotechnol. Prog., Vol. 17, pp. 791-797 (2001). As with all methods described herein for their use in conjunction with the OptKnock computational framework for metabolic modeling and simulation, the integer cut methods for reducing redundancy in iterative computational analyses can also be applied with other computational frameworks known in the art, including, for example, SimPheny®.

[0178] The methods exemplified herein allow for the construction of cells and organisms that biosynthetically produce desired products, including the essential linking of the production of target biochemical products to the growth of cells or organisms engineered to carry specified genetic modifications.Thus, the computer methods described herein allow for the identification and implementation of metabolic modifications identified by a computer method selected from OptKnock or SimPheny®.A set of metabolic modifications can include, for example, the addition of enzymes of one or more biosynthetic pathways and / or the functional disruption of one or more metabolic reactions, including, for example, disruption by gene deletion.

[0179] As mentioned above, the OptKnock method was developed on the premise that mutant microbial networks can evolve toward their computationally predicted maximum growth phenotypes when subjected to long-term growth selection. In other words, the method introduces the ability of organisms to self-optimize under selective pressure. The OptKnock framework allows for the exhaustive enumeration of gene deletion combinations that enforce a coupling between biochemical production and cellular growth based on network stoichiometry. Identifying optimal gene / reaction knockouts requires the solution of a two-level optimization problem: selecting a set of active reactions such that the optimal growth solution for the resulting network overproduces the biochemical of interest (Burgard et al., Biotechnol. Bioeng. 84:647-657 (2003)).

[0180] As previously exemplified, and described in, for example, U.S. Patent Publications US2002 / 0012939, US2003 / 0224363, US2004 / 0029149, US2004 / 0072723, US2003 / 0059792, US2002 / 0168654, and US2004 / 0009466, and U.S. Patent No. 7,127,379, computer stoichiometric models of E. coli metabolism can be used to identify essential genes in metabolic pathways. As disclosed herein, the OptKnock mathematical framework can be applied to precise gene deletions that result in growth-coupled production of a desired product. Furthermore, solving the two-tiered OptKnock problem only provides one set of deletions. To enumerate all meaningful solutions, i.e., all sets of knockouts that result in the formation of growth-coupled products, an optimization technique called integer cuts can be performed. This requires iteratively solving the OptKnock problem, as described above, with the incorporation of additional constraints called integer cuts at each iteration.

[0181] As disclosed herein, the present invention relates to aldehyde dehydrogenase variants (see Examples). The generation of such variants is described in the Examples. Any of a variety of methods can be used to generate aldehyde dehydrogenase variants, for example, the aldehyde dehydrogenase variants disclosed herein. Such methods include, but are not limited to, site-directed mutagenesis, random mutagenesis, combinatorial libraries, and other mutagenesis methods described below (see Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory, New York (2001); Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, MD (1999); Gillman et al., Directed Evolution Library Creation: Methods and Protocols (Methods in Molecular Biology), Springer, 2nd ed. (2014)).

[0182] As disclosed herein, nucleic acids encoding desired activities in a pathway for 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO, or related downstream products, such as esters or amides thereof, can be introduced into a host organism. In some cases, it may be desirable to modify the activity of an enzyme or protein in a pathway for 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO, or related downstream products, such as esters or amides thereof, to increase production of 3-HBal, 1,3-BDO, 4-HBal, or 1,4-BDO, or related downstream products, such as esters or amides thereof. For example, known mutations that increase protein or enzyme activity can be introduced into the encoding nucleic acid molecule. Furthermore, optimization methods can be applied to increase enzyme or protein activity and / or decrease inhibitory activity, e.g., to decrease the activity of negative regulators.

[0183] One such optimization method is directed evolution. Directed evolution is a powerful approach that involves the targeted introduction of mutations into specific genes to improve and / or alter the properties of an enzyme. 4Improved and / or altered enzymes can be identified through the development and implementation of sensitive, high-throughput screening assays that allow for automated screening of enzyme variants (number of variants). Typically, mutagenesis and screening are performed iteratively to obtain enzymes with optimized properties. Computational algorithms have also been developed that can assist in identifying regions of genes for mutagenesis, thereby significantly reducing the number of enzyme variants that need to be generated and screened. Numerous directed evolution techniques have been developed to efficiently generate diverse variant libraries (for reviews, see Hibbert et al., Biomol. Eng. 22:11-19 (2005); Huisman and Lalonde, Biocatalysis in the pharmaceutical and biotechnology industries, pp. 717-742 (2007), Patel (ed.), CRC Press; Otten and Quax, Biomol. Eng. 22:1-9 (2005); and Sen et al., Appl. Biochem. Biotechnol. 143:212-223 (2007)), and these methods have been successfully applied to improve a wide range of properties across many enzyme classes. Enzyme characteristics that have been improved and / or altered through directed evolution techniques include, for example, selectivity / specificity for the conversion of unnatural substrates; temperature stability for robust high temperature processing; pH stability for bioprocessing under lower or higher pH conditions; substrate or product tolerance so that high product titers can be achieved; binding affinity (K), including broadening substrate binding affinity to include unnatural substrates. m ); inhibition (K i ); activity (kcat) to increase enzyme reaction rate to achieve desired flux; expression level to increase protein yield and overall pathway flux; oxygen stability to operate air-sensitive enzymes in aerobic conditions; and anaerobic activity to operate aerobic enzymes in the absence of oxygen.

[0184] Several exemplary methods have been developed for gene mutagenesis and diversification to target the desired properties of specific enzymes. Such methods are well known to those skilled in the art. Any of these methods can be used to change and / or optimize the activity of enzymes or proteins in the 3-HBal, 1,3-BDO, 4-HBal or 1,4-BDO pathway or their related downstream products, such as esters or amides, or the activity of the aldehyde dehydrogenase of the present invention. Such methods include EpPCR, which introduces random point mutations by reducing the fidelity of DNA polymerase in the PCR reaction (Pritchard et al., J Theor. Biol. 234:497-509 (2005)); Error-prone Rolling Circle Amplification (epRCA), which is similar to epPCR except that the entire circular plasmid is used as a template and a random six-base stretch bearing exonuclease-resistant thiophosphate linkages on the last two nucleotides is used to amplify the plasmid, followed by transformation into cells in which the plasmid recircularizes at the tandem repeat (Fujii et al., Nucleic Acids Res. 32:e145 (2004); and Fujii et al., Nat. Protoc.1:2493-2497 (2006); DNA Shuffling or Family Shuffling, which typically involves digesting two or more variant genes with a nuclease, e.g., Dnase I or EndoV, to generate a pool of random fragments, which are then reassembled by cycles of annealing and extension in the presence of DNA polymerase, resulting in a library of chimeric genes (Stemmer, Proc Natl Acad Sci USA 91:10747-10751 (1994); and Stemmer, Nature 370:389-391 (1994)); Staggered Extension, which involves template priming followed by repeated cycles of two-step PCR with denaturation and very short periods of annealing / extension (as short as 5 seconds). Recombination techniques include, but are not limited to, Streptomyces Recombination (StEP) (Zhao et al., Nat. Biotechnol. 16:258-261 (1998)); Random Priming Recombination (RPR) (Shao et al., Nucleic Acids Res 26:681-683 (1998)), which uses primers of random sequence to generate many short DNA fragments complementary to different segments of a template.

[0185] Additional methods include Heteroduplex Recombination, which uses linearized plasmid DNA to form heteroduplexes that are repaired by mismatch repair (Volkov et al., Nucleic Acids Res. 27:e18 (1999); and Volkov et al., Methods Enzymol. 328:456-463 (2000)); Random Chimeragenesis on Transient Templates (RACHITT), which uses Dnase I fragmentation and size fractionation of single-stranded DNA (ssDNA) (Coco et al., Nat. Biotechnol. 19:354-359 (2001)); and Recombinant Extension on Truncated Templates, which causes template switching of unidirectionally growing strands from primers in the presence of unidirectional ssDNA fragments used as a pool of templates. Degenerate Oligonucleotide Gene Shuffling (DOGS), which uses degenerate primers to control intermolecular recombination (Bergquist and Gibbs, Methods Mol. Biol 352:191-204 (2007); Bergquist et al., Biomol. Eng 22:63-72 (2005); Gibbs et al., Gene 271:13-20 (2001)); Incremental Truncation for the Creation of Hybrid Enzymes (ITCHY), which generates combinatorial libraries with single base pair deletions of genes or gene fragments of interest (Ostermeier et al., Proc. Natl. Acad. Sci. USA). 96:3562-3567 (1999); and Ostermeier et al., Nat. Biotechnol.17:1205-1209 (1999)); Thio-Incremental Truncation for the Creation of Hybrid Enzymes (THIO-ITCHY), which is similar to ITCHY except that phosphothioate dNTPs are used to generate the cleavage sites (Lutz et al., Nucleic Acids Res 29:E16 (2001)); SCRATCHY, which combines ITCHY and DNA shuffling, two methods for recombining genes (Lutz et al., Proc. Natl. Acad. Sci. USA 98:11248-11253 (2001)); Random Drift Mutagenesis (RNDM), in which mutations are made via epPCR followed by screening / selection for mutations that maintain usable activity (Bergquist et al., Biomol. Eng. 22:63-72 (2005); Sequence Saturation Mutagenesis (SeSaM), a method of random mutagenesis in which a pool of fragments of random length is generated using random incorporation and cleavage of phosphothioate nucleotides, and this pool is used as a template to extend in the presence of a "universal" base, e.g., inosine, with replication of the inosine-containing complement resulting in incorporation of random bases and, consequently, mutagenesis (Wong et al., Biotechnol. J. 3:74-82 (2008); Wong et al., Nucleic Acids Res. 32:e26 (2004); and Wong et al., Anal. Biochem. 341:187-189 (2005); Synthetic Shuffling, which uses overlapping oligonucleotides designed to encode "all the genetic diversity in the target" and allow for very high diversity in the shuffled progeny (Ness et al., Nat. Biotechnol.20:1251-1255 (2002); Nucleotide Exchange and Excision Technology (NexT), which utilizes a combination of dUTP incorporation followed by treatment with uracil DNA glycosylase and then piperidine to achieve end-point DNA fragmentation (Muller et al., Nucleic Acids Res. 33:e117 (2005)).

[0186] Additional methods include Sequence Homology-Independent Protein Recombination (SHIPREC) (Sieber et al., Nat. Biotechnol. 19:456-460 (2001)), which uses a linker to facilitate fusion between two loosely related or unrelated genes, generating a variety of chimeras between the two genes and resulting in a library of single-crossover hybrids; Gene Site Saturation Mutagenesis™ (GSSM™) (Kretz et al., Methods Enzymol. 388:3-11 (2004)), in which the starting material comprises a supercoiled double-stranded DNA (dsDNA) plasmid containing an insert and two primers that are degenerate at the desired site of mutation; and Combinatorial Cassette Mutagenesis (Cassette Mutagenesis) (Cassette Mutagenesis) (Kretz et al., Methods Enzymol. 388:3-11 (2004)), in which short oligonucleotide cassettes are used to replace defined regions with a large number of possible amino acid sequence changes. Combinatorial Multiple Cassette Mutagenesis (CCM) (Reidhaar-Olson et al., Methods Enzymol. 208:564-586 (1991); and Reidhaar-Olson et al., Science 241:53-57 (1988)); essentially, similar to CCM, epPCR is used at a high mutation rate to identify hot spots and hot regions, which are then extended by CMCM to cover defined regions of protein sequence space (Reetz et al., Angew. Chem. Int. Ed Engl.40:3589-3591 (2001); the Mutator Strains technique, which utilizes the mutD5 gene, encoding a mutant subunit of DNA polymerase III, to allow conditional ts mutator plasmids to increase the frequency of random and natural mutations 20-4000-fold during selection, blocking the accumulation of deleterious mutations when selection is not required (Selifonova et al., Appl. Environ. Microbiol. 67:3645-3649 (2001); Low et al., J. Mol. Biol. 260:359-3680 (1996)).

[0187] Additional exemplary methods include Look-Through Mutagenesis (LTM), a multidimensional mutagenesis method that evaluates and optimizes combinatorial mutations of selected amino acids (Rajpal et al., Proc. Natl. Acad. Sci. USA 102:8466-8471 (2005)); Gene Reassembly (Tunable GeneReassembly™ (TGR™) technology supplied by Verenium Corporation), a DNA shuffling method that can be applied to multiple genes at once or to generate large libraries of chimeras (multiple mutations) of a single gene; and in silico Protein Design Automation (PDA), an optimization algorithm that searches sequence space for amino acid substitutions that can anchor a structurally defined protein scaffold with a specific fold and stabilize the fold and overall protein energetics, and is generally most effective for proteins with known three-dimensional structure (Hayes et al., Proc. Natl. Acad. Sci. USA 99:15926-15931 (2002)); and Iterative Saturation Mutagenesis (ISM) (Reetz et al., Nat. Protoc. 2:891-903 (2007); and Reetz et al., Angew. Chem. Int. Ed Engl. 45:7745-7751 (2006)), which involves using structure / function knowledge to select promising sites for enzyme improvement, performing saturation mutagenesis at the selected sites using mutagenesis methods such as Stratagene QuikChange (Stratagene; San Diego CA), screening / selecting for the desired properties, and using the improved clone(s) to begin again at another site, repeating until the desired activity is achieved.

[0188] Any of the above methods for mutagenesis can be used alone or in any combination. Additionally, as described herein, directed evolution methods may be used in conjunction with adaptive evolution techniques, either alone or in combination.

[0189] It is understood that modifications that do not substantially affect the activity of the various embodiments of this invention are also provided within the definition of the invention provided herein. Accordingly, the following examples are intended to illustrate, but not limit, the present invention. [Example]

[0190] Aldehyde dehydrogenase variants This example describes the generation of aldehyde dehydrogenase variants with desirable properties.

[0191] Mutagenesis techniques were used to generate variant aldehyde dehydrogenases based on template ALD-1. Variants were generated using error-prone PCR, site-directed mutagenesis, and by spontaneous mutations during gene selection. Template ALD-1 corresponds to the aldehyde dehydrogenase provided below: [ka]

[0192] Additional ALD sequences for ALD-2 and ALD-3 are provided below: [ka] [ka]

[0193] ALD-1 is slightly more specific for the R enantiomer of 3-hydroxybutyryl-CoA compared to the S enantiomer. A sequence alignment of ALD-1 to ALD-2 and ALD-3 is shown in Figure 3. The sequences correspond to SEQ ID NOs: 1, 2, and 3, respectively. A crystal structure also exists for ALD-3 (PDBID 4C3S), and ALD-2 is more closely related to ALD-3 than to ALD-1. Therefore, ALD-3 was used as a template. Underlined in Figure 3 are two loop regions, the first designated A and the second designated B, both of which are responsible for substrate specificity and enantiomer specificity as determined herein. Loop A of ALD-1 has the sequence LQKNNETQEYSINKKWVGKD (SEQ ID NO: 124), loop A of ALD-2 has the sequence IGPKGAPDRKFVGKD (SEQ ID NO: 125), and loop A of ALD-3 has the sequence ITPKGLNRNCVGKD (SEQ ID NO: 126). Loop B of ALD-1 has the sequence SFAGVGYEAEGFTTFTIA (SEQ ID NO: 127), loop B of ALD-2 has the sequence TYCGTGVATNGAHSGASALTIA (SEQ ID NO: 128), and loop B of ALD-3 has the sequence SYAAIGFGGEGFCTFTIA (SEQ ID NO: 129). The sequences and lengths of the substrate specificity loops A and B from ALD-2 are different from those of ALD-1 and ALD-3; nevertheless, the alignment shows sufficient conservation to facilitate identification of corresponding positions for substitutions as described herein, particularly when combined with 3D modeling as shown in FIG. 6 , which indicates that the two loop regions interact to affect substrate and enantiomer specificity, particularly when modified with exemplary substitutions as described herein. ALD-1 and ALD-3 are 51.9% identical. ALD-1 and ALD-2 are 35.9% identical. ALD-3 and ALD-2 are 40% identical. A consensus ALD sequence was generated based on the alignment in FIG. 3 .The consensus for loop A based on the alignment of ALD-1, ALD-2 and ALD-3 is IXPKG-----XXNRKXVGKD (SEQ ID NO: 5). The consensus for loop B based on the alignment of ALD-1, ALD-2 and ALD-3 is SYAGXGXXXE----GFXTFTIA (SEQ ID NO: 6).

[0194] Additional alignments were performed (Figure 4). Figure 4A shows an alignment at a 40-55% cutoff compared to ALD-1. Figure 4B shows an alignment at a 75-90% cutoff compared to ALD-1. Figure 4C shows an alignment at a 90% cutoff compared to ALD-1. The alignments of exemplary aldehyde dehydrogenases (ALDs) shown in Figures 4A-4C indicate the identification of positions in ALDs corresponding to positions in a representative template ALD sequence where substitutions of the present invention can be made. Underlined are two important loop regions, the first designated A and the second designated B, both of which are involved in substrate specificity and enantiomer specificity as determined herein. Figures 4A-4C show that corresponding positions for substitutions taught herein can be identified in ALDs that are at least 40% identical to ALD-1, particularly in the loop A and B regions, and particularly in the highly conserved loop B region.

[0195] Mutations to increase the specificity of variant 45 for 3HB-CoA relative to acetyl-CoA resulted in several variants with increased 1,3 BDO production and reduced ethanol. Because acetaldehyde produced from acetyl-CoA can be converted to ethanol by the enzyme natively in the host cell or by pathway enzymes that convert 3-hydroxybutyraldehyde to 1,3-butanediol, mutagenesis that increases the specificity for 3-hydroxybutyryl-CoA over acetyl-CoA results in reduced ethanol. Variants that increase the enzymatic activity of aldehyde dehydrogenase or increase its specificity for 3-hydroxybutyryl-CoA reduce 4-hydroxy-2-butanone by increasing the flux to 1,3-butanediol through the enzymatic pathway, which pulls acetoacetyl-CoA into 1,3-butanediol formation, which reduces its availability for the two-step conversion to 4-hydroxy-2-butanone by the native enzyme or by less specific pathway enzymes. The sequence of variant 45 is provided below: [ka]

[0196] The assay performed was an in vitro assay that examines activity toward 3HB-CoA by monitoring the decrease in absorbance as NADH is converted to NAD. An assay using acetyl-CoA (AcCoA) as a substrate was also performed to identify improved enzymes for improved 3HB-CoA to AcCoA activity ratios. Because acetaldehyde produced from acetyl-CoA can be converted to ethanol by enzymes native to the host cell or by pathway enzymes that convert 3-hydroxybutyraldehyde to 1,3-butanediol, mutations that increase the specificity for 3-hydroxybutyryl-CoA over acetyl-CoA result in reduced ethanol production.

[0197] Further investigation of a subset of these variants with (R) and (S) 3-hydroxybutyraldehyde showed that five of the variants tested (952, 955, 957, 959, 961) had improved selectivity for the R enantiomer compared to the parent enzyme (variant 45) and wild-type ALD-1 (Figure 5). Figure 5A shows the specific activity of ALD-2, ALD-1, and ALD-1 variants toward 3-hydroxy-(R)-butyraldehyde (left bar in the set of bars) and 3-hydroxy-(S)-butyraldehyde (right bar in the set of bars). Purified streptavidin-tagged proteins were incubated at 35°C in IVI buffer pH 7.5, 0.5 mM NADPH. + The assay was performed in the presence of 10 mM either R or S 3-hydroxybutyraldehyde in 2 mM CoA, and activity was monitored by the change in NADH absorbance at 340 nm. The IVI buffer contained 5 mM monobasic potassium phosphate, 20 mM dibasic potassium phosphate, 10 mM sodium glutamate monohydrate, and 150 mM potassium chloride, pH 7.5. Therefore, the enzymatic reaction in the assay was performed in the opposite direction to that shown in Figure 1, i.e., the reaction measured the conversion of 3-hydroxybutyraldehyde to 3-hydroxybutyryl-CoA. As shown in Figure 5B, certain aldehyde dehydrogenase variants exhibited selectivity for R-3-hydroxybutyraldehyde (R-3HB-aldehyde) over S-3-hydroxybutyraldehyde (S-3HB-aldehyde).

[0198] Computer modeling of mutant 959 using the ALD-1 crystal structure suggests that the amino acid substitution F442N allows a hydrogen bond network to form with the hydroxyl on carbon 3 in the R isomer but not in the (S) isomer (Figure 6). Figures 6A-6C show ribbon diagrams of the structure of aldehyde dehydrogenase 959. The diagrams show the docking of 3-hydroxy-(R)-butyraldehyde (Figure 6A) or 3-hydroxy-(S)-butyraldehyde (Figure 6B) into the 959 structure. Figure 6C shows that when 3-hydroxy-(S)-butyraldehyde is docked in the same energetically most favorable orientation for docking of 3-hydroxy-(R)-butyraldehyde shown in Figure 6A, an unfavorable interaction (circled) is created with the isoleucine located in the active site. The model shows that the mutation F442N creates a hydrogen bond between the protein and the hydroxyl of 3-hydroxy-(R)-butyraldehyde, which is not possible in the S enantiomer.

[0199] Exemplary aldehyde dehydrogenase variants are shown in Tables 1A-1D. [Table 1A-1] [Table 1A-2] [Table 1A-3] [Table 1A-4] [Table 1A-5] [Table 1B-1] [Table 1B-2] [Table 1B-3] [Table 1B-4] [Table 1C-1] [Table 1C-2] [Table 1C-3] [Table 1C-4] [Table 1C-5] [Table 1D-1] [Table 1D-2] [Table 1D-3] [Table 1D-4]

[0200] The activities of various ALD variants were determined and are shown in Table 2. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8]

[0201] Additional activities of exemplary ALD variants are shown in Table 3. Levels of 1,3-BDO production at 48 hours as high as greater than 50 g / liter, greater than 60 g / liter, greater than 70 g / liter, greater than 80 g / liter, and greater than 90 g / liter were obtained with the ALD variants. [Table 3-1] [Table 3-2]

[0202] Such aldehyde dehydrogenase variants can be used to produce stereoisomers of R-3-hydroxybutyraldehyde, or a mixture of R and S forms with a higher proportion of the R form. Such stereoisomers can be used to produce stereoisomers of downstream products, such as R-1,3-butanediol. Such stereoisomers are useful as pharmaceuticals or dietary supplements.

[0203] These results demonstrate the production of aldehyde dehydrogenase variants with desirable properties that are useful for the commercial production of 3-hydroxybutyraldehyde, 1,3-butanediol, 4-hydroxybutyraldehyde, or 1,4-butanediol, or other desired products produced by metabolic pathways involving aldehyde dehydrogenase.

[0204] The above variants are based on the ALD-1 parent sequence. It is understood that the variant amino acid positions shown in Tables 1, 2, or 3 can be applied to homologous aldehyde dehydrogenase sequences. Table 4 shows exemplary ALD sequences based on homology. Those skilled in the art will readily understand that such sequences can be analyzed using routine and well-known methods for aligning sequences (e.g., BLAST, blast.ncbi.nlm.nih.gov; Altschul et al., J. Mol. Biol. 215:403-410 (1990)). Furthermore, additional homologous ALD sequences can be identified by searching publicly available sequence databases, such as those found in the National Center for Biotechnology Information (NCBI) GenBank database, the European Molecular Biology Laboratory (EBM), ExPasy Prosite, or other publicly available sequence databases using BLAST. Such alignments can provide information about conserved residues that can be utilized to identify consensus sequences for preserving enzyme activity and positions for generating additional enzyme variants. [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] [Table 4-7] [Table 4-8] Table 4-9 Table 4-10 Table 4-11 Table 4-12

[0205] It is understood that individual ALD variants, such as those described above, may be used alone or combined with any other variant amino acid positions, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16, i.e., up to all of the variant amino acid positions disclosed herein (see Tables 1-3), to generate additional variants with desired activity. Exemplary ALD variants include, but are not limited to, a single substitution or one or more combinations of substitutions at the amino acid positions disclosed in any of Tables 1-3, for example, at amino acid positions 12, 19, 33, 44, 65, 72, 73, 107, 122, 129, 139, 143, 145, 155, 163, 167, 174, 189, 204, 220, 227, 229, 230, 243, 244, 254, 267, 315, 353, 356, 396, 429, 432, 437, 440, 441, 442, 444, 447, 450, 460, 464, or 467 corresponding to the amino acid sequence of ALD-1 (SEQ ID NO: 1) (see Tables 1-3). For example, ALD variants include those at amino acid positions D12, V19, C33, I44, K65, K72, A73, Y107, D122, E129, I139, T143, P145, G155, V163, G167, C174, C189, M204, C220, M227, K229, T230, A243, G250, and G304, which correspond to the amino acid sequence of ALD-1 (SEQ ID NO: 1). These include, but are not limited to, amino acid substitutions, single substitutions, or one or more combinations of substitutions at 244, A254, C267, V315, C353, C356, R396, F429, V432, E437, T440, T441, F442, I444, S447, E450, R460, C464, or A467 (see Tables 1-3). It is understood that substitutions of any of the other 19 amino acids may be made at one or more desired amino acid positions.

[0206] In one embodiment, the variant ALD comprises an amino acid substitution at position 12 that is D12A. In one embodiment, the variant ALD comprises an amino acid substitution at position 19 that is V19I. In one embodiment, the variant ALD comprises an amino acid substitution at position 33 that is C33R. In one embodiment, the variant ALD comprises an amino acid substitution at position 44 that is I44L. In one embodiment, the variant ALD comprises an amino acid substitution at position 65 that is K65A. In one embodiment, the variant ALD comprises an amino acid substitution at position 72 that is K72N. In one embodiment, the variant ALD comprises an amino acid substitution at position 73 selected from A73S, A73D, A73G, A73L, A73Q, A73F, A73E, A73W, A73R, A73C, and A73M. In one embodiment, the variant ALD comprises an amino acid substitution at position 107 that is Y107K. In one embodiment, the variant ALD comprises an amino acid substitution at position 122 that is D122N. In one embodiment, the variant ALD comprises an amino acid substitution at position 129 that is E129I. In one embodiment, the variant ALD comprises an amino acid substitution at position 139 that is selected from I139S, I139V, and I139L. In one embodiment, the variant ALD comprises an amino acid substitution at position 143 that is T143N or T143S. In one embodiment, the variant ALD comprises an amino acid substitution at position 163 that is selected from V163C, V163G, and V163T. In one embodiment, the variant ALD comprises an amino acid substitution at position 167 that is G167S. In one embodiment, the variant ALD comprises an amino acid substitution at position 174 that is C174S. In one embodiment, the variant ALD comprises an amino acid substitution at position 189 that is C189A. In one embodiment, the variant ALD comprises an amino acid substitution at position 204 that is M204R. In one embodiment, the variant ALD comprises an amino acid substitution at position 220 that is C220V. In one embodiment, the variant ALD comprises an amino acid substitution at position 227 selected from M227K, M227Q, M227I, M227V, M227C, M227L and M227A.In one embodiment, the variant ALD comprises an amino acid substitution at position 229 that is K229S. In one embodiment, the variant ALD comprises an amino acid substitution at position 230 that is selected from T230R, T230K, T230H, T230A, T230M, T230C, T230L, T230S, T230Y, T230G, T230T, T230I, T230W, T230N, T230V, and T230Q. In one embodiment, the variant ALD comprises an amino acid substitution at position 243 that is selected from A243P, A243Q, A243E, A243S, A243N, A243K, A243L, A243C, A243M, and A243I. In one embodiment, the variant ALD comprises an amino acid substitution at position 254 that is A254T. In one embodiment, the variant ALD comprises an amino acid substitution at position 267 that is C267A. In one embodiment, the variant ALD comprises an amino acid substitution at position 315 that is V315A. In one embodiment, the variant ALD comprises an amino acid substitution at position 353 that is C353A. In one embodiment, the variant ALD comprises an amino acid substitution at position 356 that is C356T or C356L. In one embodiment, the variant ALD comprises an amino acid substitution at position 396 that is R396H. In one embodiment, the variant ALD comprises an amino acid substitution at position 429 selected from F429Y, F429Q, F429H, F429M, F429D, and F429L. In one embodiment, the variant ALD comprises an amino acid substitution at position 432 that is V432V or V432N. In one embodiment, the variant ALD comprises an amino acid substitution at position 437 that is E437P. In one embodiment, the variant ALD comprises an amino acid substitution at position 440 that is T440H. In one embodiment, the variant ALD comprises an amino acid substitution at position 441 that is T441G. In one embodiment, the variant ALD comprises an amino acid substitution at position 442 that is selected from F442T, F442Y, F442H, F442N, F442Q, F442M, and F442F. In one embodiment, the variant ALD comprises an amino acid substitution at position 444 that is I444V.In one embodiment, the variant ALD comprises an amino acid substitution at position 447 selected from S447M, S447P, S447H, S447K, S447R, S447T, S447E, and S447S. In one embodiment, the variant ALD comprises an amino acid substitution at position 460 that is R460K. In one embodiment, the variant ALD comprises an amino acid substitution at position 464 that is C464V or C464I. In one embodiment, the variant ALD comprises an amino acid substitution at position 467 that is A467V. Any of the above amino acid positions can be used for a single amino acid substitution or one or more combinations of substitutions to generate the ALD variants of the present invention.

[0207] Based on the teachings herein, one of skill in the art can readily identify amino acid positions in a homologous ALD sequence that correspond to any of amino acid positions 12, 19, 33, 44, 65, 72, 73, 107, 122, 129, 139, 143, 145, 155, 163, 167, 174, 189, 204, 220, 227, 229, 230, 243, 244, 254, 267, 315, 353, 356, 396, 429, 432, 437, 440, 441, 442, 444, 447, 450, 460, 464, or 467 of the amino acid sequence of ALD-1 (SEQ ID NO: 1). For example, as shown in the alignment in Figure 4A, amino acid I139 of ALD-1 corresponds to amino acid I133 of SEQ ID NOs: 13 and 20. For SEQ ID NO: 24, the corresponding position is V199. Using well-known methods for aligning amino acid sequences, generally using default parameters as disclosed herein, one of skill in the art can readily determine the amino acid position in another ALD sequence that corresponds to any of amino acid positions 12, 19, 33, 44, 65, 72, 73, 107, 122, 129, 139, 143, 145, 155, 163, 167, 174, 189, 204, 220, 227, 229, 230, 243, 244, 254, 267, 315, 353, 356, 396, 429, 432, 437, 440, 441, 442, 444, 447, 450, 460, 464, or 467 in the amino acid sequence of ALD-1 (SEQ ID NO: 1).

[0208] It is further understood that the ALD variants can contain, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 variant amino acid positions in Tables 1-3, i.e., up to all variant amino acid positions disclosed herein. One skilled in the art can readily generate ALD variants based on any single or combination of amino acid substitutions, e.g., amino acid variant positions described above and in Tables 1-3, as disclosed herein. In certain embodiments, the ALD variants are variants disclosed in Tables 1-3.

[0209] Throughout this application, various publications are referenced. The disclosures of these publications in their entireties, including GenBank accession update numbers and / or GI number publications, are hereby incorporated by reference into this application in order to more fully describe the state of the art to which this invention pertains. While the invention has been described with reference to the examples provided above, it should be understood that various modifications can be made without departing from the spirit of the invention.

Claims

1. (a) An isolated nucleic acid molecule encoding an aldehyde dehydrogenase variant comprising an amino acid sequence comprising the following amino acid substitutions in the amino acid sequence referenced as SEQ ID NO:1: C174S, M204R, C220V, C267A, C356T, R396H, E437P, C464I, and A467V, wherein the amino acid sequence other than the amino acid substitutions has at least 90%, 95%, 98%, or 99% sequence identity to, or is identical to, the amino acid sequence referenced as SEQ ID NO:

1.

2. A vector containing the nucleic acid molecule of claim 1.

3. An aldehyde dehydrogenase variant comprising an amino acid sequence comprising the following amino acid substitutions in the amino acid sequence referenced as SEQ ID NO:1: C174S, M204R, C220V, C267A, C356T, R396H, E437P, C464I, and A467V, wherein the amino acid sequence other than the amino acid substitutions has at least 90%, 95%, 98%, or 99% sequence identity to, or is identical to, the amino acid sequence referenced as SEQ ID NO:

1.

4. 4. The isolated nucleic acid molecule of claim 1, the vector of claim 2, or the aldehyde dehydrogenase variant of claim 3, wherein the aldehyde dehydrogenase variant further comprises one or more additional amino acid substitutions selected from D12A and I139S.

5. The aldehyde dehydrogenase variant is (a) capable of converting 3-hydroxybutyryl-CoA to 3-hydroxybutyraldehyde or 4-hydroxybutyryl-CoA to 4-hydroxybutyraldehyde; (b) has higher activity compared to an aldehyde dehydrogenase comprising SEQ ID NO: 1; (c) has higher activity for 3-hydroxy-(R)-butyryl-CoA than for 3-hydroxy-(S)-butyryl-CoA; (d) has a higher specificity for 3-hydroxybutyryl-CoA than acetyl-CoA; (e) has a higher specificity for 4-hydroxybutyryl-CoA than for acetyl-CoA; (f) reducing the production of a by-product in a cell or cell extract, wherein the by-product is ethanol or 4-hydroxy-2-butanone; or (g) has a higher kcat compared to an aldehyde dehydrogenase comprising SEQ ID NO: 1; An isolated nucleic acid molecule according to claim 1 or 4, a vector according to claim 2 or 4, or an aldehyde dehydrogenase variant according to claim 3 or 4.

6. A cell, a) comprising the vector of claim 2; b) comprising the isolated nucleic acid of claim 1, or c) comprising the aldehyde dehydrogenase variant of claim 3; cell.

7. The cell of claim 6, (a) a microbial organism, wherein the microbial organism is a bacterium, yeast, or fungus; (b) is an isolated eukaryotic cell; or (c) capable of fermentation; cell.

8. The cell of claim 6, (a) comprising a pathway to produce 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or an ester or amide thereof; or (b) a pathway that produces 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) or an ester or amide thereof; cell.

9. 7. The cell of claim 6, comprising at least one substrate for the aldehyde dehydrogenase variant, wherein the substrate is 3-hydroxybutyryl-CoA, 3-hydroxy-(R)-butyryl-CoA, or 4-hydroxybutyryl-CoA.

10. 7. The cell of claim 6, wherein the cell has higher activity for 3-hydroxy-(R)-butyryl-CoA than for 3-hydroxy-(S)-butyryl-CoA.

11. Use of an aldehyde dehydrogenase variant according to any one of claims 3, 4 and 5 as a biocatalyst.

12. 6. A composition comprising the aldehyde dehydrogenase variant of claim 3 and at least one substrate for the aldehyde dehydrogenase variant, wherein the aldehyde dehydrogenase variant is capable of reacting with the substrate under in vitro conditions, and the substrate is 3-hydroxybutyryl-CoA, 3-hydroxy-(R)-butyryl-CoA, or 4-hydroxybutyryl-CoA.

13. A method for constructing a host strain, comprising the step of introducing the vector of claim 2 into a cell capable of fermentation.

14. 1. A method comprising: a) a method for producing 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or an ester or amide thereof, the method comprising culturing the cell according to any one of claims 6 to 10 to produce 3-HBal and / or 1,3-BDO or an ester or amide thereof, or 4-HBal and / or 1,4-BDO or an ester or amide thereof, i) the cells are in a substantially anaerobic medium; ii) the method further comprises a step of isolating or purifying the 3-HBal and / or 1,3-BDO or the 4-HBal and / or 1,4-BDO, or esters or amides thereof, wherein the isolating or purifying step comprises distillation; b) a method for producing 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) or an ester or amide thereof, the method comprising culturing the cell according to any one of claims 6 to 10 to produce 3-HBal and / or 1,3-BDO or an ester or amide thereof or 4-HBal and / or 1,4-BDO or an ester or amide thereof; i) the cells are in a substantially anaerobic medium; ii) the method further comprises a step of isolating or purifying the 3-HBal and / or 1,3-BDO or the 4-HBal and / or 1,4-BDO, or esters or amides thereof, wherein the isolating or purifying step comprises distillation; c) A method for producing 3-hydroxybutyraldehyde (3-HBal) and / or 1,3-butanediol (1,3-BDO) or esters or amides thereof, comprising the steps of providing a substrate to the aldehyde dehydrogenase variant of claim 3 and converting the substrate to 3-HBal and / or 1,3-BDO, wherein the substrate is a racemic mixture of 1,3-hydroxybutyryl-CoA, the 3-HBal and / or 1,3-BDO are enantiomerically enriched in the R form, and the aldehyde dehydrogenase variant is present in a cell, a cell lysate, or isolated from a cell or a cell lysate. d) a method for producing 4-hydroxybutyraldehyde (4-HBal) and / or 1,4-butanediol (1,4-BDO) or an ester or amide thereof, comprising the steps of providing a substrate to an aldehyde dehydrogenase variant according to claim 3 and converting the substrate to 4-HBal and / or 1,4-BDO, wherein the substrate is 1,4-hydroxybutyryl-CoA and the aldehyde dehydrogenase variant is present in a cell, in a cell lysate, or isolated from a cell or a cell lysate; or e) A method for producing 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO, comprising the step of incubating a lysate of cells according to any one of claims 6 to 10 to produce 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO, wherein the cell lysate is mixed with a second cell lysate, and the second cell lysate contains an enzyme activity that produces a substrate for the aldehyde dehydrogenase variant according to claim 3 or a downstream product of 3-HBal and / or 1,3-BDO or 4-HBal and / or 1,4-BDO. method.

15. 4. A method for producing the aldehyde dehydrogenase variant of claim 3, comprising: (i) expressing the aldehyde dehydrogenase variant in a cell; or (ii) A method comprising the step of in vitro transcribing and translating the nucleic acid of claim 1 or the vector of claim 2, thereby producing the aldehyde dehydrogenase variant.

Citation Information

Patent Citations

  • Microorganisms and methods for production of 4-hydroxybutyrate, 1,4-butanediol and related compounds

    WO2014176514A2