Thioesterase variants with improved activity for the production of medium-chain fatty acid derivatives
Engineered thioesterase variants and recombinant host cells address the instability of medium-chain fatty acid production by enhancing catalytic activity and selectivity, ensuring a reliable and sustainable supply.
Patent Information
- Application Number
- JP2020502555
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-04-03
- Filing Date
- 2018-04-03
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2038-04-03
AI Technical Summary
The supply of medium-chain fatty acids is unstable and unreliable due to the limitations of existing production methods, which often rely on suboptimal thioesterases with broad specificity and microbial toxicity issues, hindering efficient and sustainable production.
Development of engineered thioesterase variants with improved activity and selectivity for producing medium-chain fatty acid derivatives, along with recombinant host cells equipped with specific biochemical pathways to enhance tolerance and conversion efficiency.
The engineered thioesterase variants and host cells enable stable and sustainable production of medium-chain fatty acid derivatives, overcoming toxicity issues and improving catalytic activity and selectivity, thereby ensuring a reliable supply.
Smart Images

Figure 0007705243000031 
Figure 0007705243000032 
Figure 0007705243000033
Abstract
Description
Technical Field
[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 481,078, filed Apr. 3, 2017, which is hereby incorporated by reference in its entirety.
[0002] Field The present disclosure relates to molecular tools useful for the production of medium-chain length fatty acids and fatty acid derivatives. Accordingly, the present disclosure relates to genes that confer resistance to medium-chain length fatty acids and fatty acid derivatives in microorganisms. The present disclosure further relates to novel engineered thioesterase variants having improved activity and / or selectivity for the production of medium-chain fatty acid derivatives, including, for example, fatty acids and fatty acid derivatives having 8 and 10 carbon atoms, and polynucleotides encoding the same. Accordingly, the present disclosure also relates to host cells comprising engineered thioesterase variants and polynucleotides encoding the same, and related cell cultures. Further included is a method for producing medium-chain fatty acid derivatives by employing host cells expressing engineered thioesterase variants and compositions of biologically produced medium-chain fatty acid derivatives.
Background Art
[0003] Background There is great interest in producing products derived from medium-chain fatty acids (MCFAs). Medium-chain fatty acids and medium-chain fatty acid derivatives are widely applied in industry as, for example, feedstocks for biofuels, lubricants and greases, metalworking fluids, coating agents and adhesives, cosmetics and personal care products, fragrances, food nutrition, pharmaceuticals, plastics and rubbers, and other applications in the chemical industry.
[0004] In addition to their value in industry, medium-chain fatty acids are being put to valuable uses as nutritional supplements and foods for special dietary uses (see, for example, Stig Bengmark (2013) Nutrients 5(1): 162-207 (Non-Patent Document 1)). In fact, medium-chain fatty acids and their derivatives exhibit antibacterial properties (see, for example, Nobmann et al. International Journal of Food Microbiology. 2009;128(3):440-445 (Non-Patent Document 2); B W Petschow, et al., (1996) Antimicrob. Agents Chemother. 40(2):302-306 (Non-Patent Document 3)), suppress the accumulation of body fat, and prevent metabolic syndrome (see, for example, Takeuchi H., et al. (2008) Asia Pac J Clin Nutr. 17 Suppl 1:320-3 (Non-Patent Document 4); Koji Nagao (2010) Pharmacological Research 61:208-212 (Non-Patent Document 5); Omura Y., et al. (2011) Acupunct Electrother Res. 36(1-2):19-64 (Non-Patent Document 6)) and have an anti-convulsant effect at clinically meaningful concentrations (see, for example, Chang et al., (2013) Neuropharmacology 2013; 69: 105-14 (Non-Patent Document 7); Wlaz et al., (2015) Prog Neuropsychopharmacol Biol Psychiatry 2015; 57: 110-16 (Non-Patent Document 8)).
[0005] Considering the many useful applications, it is not surprising that the demand for medium-chain fatty acids has been on the rise over the past few years. Unfortunately, the supply of medium-chain fatty acids has always been associated with the production of longer-chain free fatty acid (FFA) products from plants (palm oil) or chemical synthesis, and the medium-chain length chains are produced as shoulders corresponding to less than 20% of the total fatty acid acyl species (see, for example, Kostik, V. et al. (2013) J. Hyg. Eng. Des. 4:112-116 (Non-Patent Document 9)). This makes the supply of medium-chain fatty acids very variable and unstable. Therefore, there is a need in the art for a method that can provide a reliable, stable, and sustainable supply of these compounds.
[0006] An alternative to the current sources of medium-chain fatty acids is their production using biological systems such as microbial fermentation. However, the production of free fatty acids by biological systems presents two major challenges. First, this often relies on thioesterases that act on alkylthioester molecules produced by the host organism. Available thioesterases active towards medium-chain alkylthioesters either have suboptimal catalytic activity or their specificity is too broad to act on a range of alkylthioester chain lengths. Second, medium-chain aliphatic compounds are often highly toxic to microbial cells, hindering their production at high levels. Additionally, the toxicity of medium-chain acyl compounds can disadvantage the selection and manipulation of highly active thioesterases. Therefore, for biological systems to provide an alternative supply of medium-chain fatty acids, a biological system is needed that has an improved thioesterase with higher activity and selectivity towards medium-chain alkylthioesterases and that shows improved tolerance to medium-chain aliphatic compounds.
[0007] Fortunately, as will be apparent from the following disclosure, the present invention meets these and other needs.
Prior Art Documents
Non-Patent Documents
[0008]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Non-Patent Document 6
Non-Patent Document 7
Non-Patent Document 8
Non-Patent Document 9
Summary of the Invention
[0009] Summary One aspect of the present disclosure provides engineered thioesterase variants having improved activity for the production of medium-chain fatty acid derivatives. Thus, in one embodiment, the present disclosure provides engineered thioesterase variants having improved activity for the production of medium-chain fatty acid derivatives. In one embodiment, the engineered thioesterase variant has improved activity for the production of C8 fatty acid derivatives, the engineered thioesterase variant of claim 1. In one embodiment, the engineered thioesterase variant has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 and at least one substitution mutation at an amino acid position selected from the group consisting of 3, 4, 6, 14, 15, 17, 22, 37, 44, 45, 50, 54, 56, 64, 67, 73, 76, 91, 99, 102, 110, 111, 114, 129, 132, 137, 158, 162, 165, 176, 178, 185, 186, 196, 197, 198, 203, 213, 217, 225, 227, 236, 244, 254, 256, 258, 278, 282, 292, 297, 298, 299, 300, 301, 302, 316, 321, and 322. In one embodiment of the engineered thioesterase, the at least one substitution mutation is (a) lysine at amino acid position 3; (b) methionine at amino acid position 4; (c) arginine at amino acid position 6; (d) glycine or arginine at amino acid position 14; (e) leucine or tryptophan at amino acid position 15; (f) alanine or cysteine at amino acid position 17; (g) arginine at amino acid position 22; (h) proline at amino acid position 37; (i) glycine or isoleucine at amino acid position 44; (j) serine at position 45; (k) tryptophan at amino acid position 50; (l) arginine at amino acid position 54; (m) lysine or cysteine at amino acid position 56; (n) arginine or proline at amino acid position 64; (o) leucine at amino acid position 67; (p) valine at position 73; (q) phenylalanine or leucine or tyrosine at amino acid position 76; (r) methionine at amino acid position 91; (s) lysine or proline at amino acid position 99; (t) isoleucine at amino acid position 102; (u) leucine at amino acid position 110; (v) threonine at position 111; (w) lysine at position 114; (x) valine at amino acid position 129;(y) Tryptophan at amino acid position 132; (z) Cysteine at amino acid position 137; (aa) Glutamine at amino acid position 158; (bb) Glutamic acid at amino acid position 162; (cc) Valine at amino acid position 176; (dd) Proline at amino acid position 178; (ee) Alanine at amino acid position 185; (ff) Glycine at amino acid position 186; (gg) Valine at amino acid position 196; (hh) Asparagine at amino acid position 197; (ii) Tryptophan at amino acid position 198; (jj) Arginine at amino acid position 203; (kk) Histidine or Arginine at amino acid position 213; (ll) Arginine at amino acid position 217; (mm) Leucine at amino acid position 225; (nn) Glycine at amino acid position 227; (oo) Threonine at amino acid position 236; (pp) Methionine or Arginine at amino acid position 244; (qq) Glycine at amino acid position 254; (rr) Cysteine or Arginine at amino acid position 256; (ss) Threonine or Valine at amino acid position 258; (tt) Lysine or Valine at amino acid position 278; (uu) Serine or Valine at amino acid position 282; (vv) Phenylalanine at amino acid position 292; (ww) Threonine or Aspartic acid or Valine at amino acid position 297; (xx) Valine or Cysteine at amino acid position 298; (yy) Leucine at amino acid position 299; (zz) Lysine or Tryptophan or Leucine at amino acid position 300; (aaa) Cysteine at amino acid position 301; (bbb) Threonine at amino acid position 302; (ccc) Arginine at amino acid position 316; (ddd) Arginine at amino acid position 321; and (eee) Lysine at amino acid position 322, which is a member selected from the group consisting of;
[0010] In one aspect, the engineered thioesterase variant is a member selected from SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59.
[0011] In one aspect, the engineered thioesterase variant has an overall increased effective positive charge compared to the thioesterase having SEQ ID NO:1. In one aspect, the engineered thioesterase variant has an overall increased effective positive charge compared to the thioesterase having variant SEQ ID NO:4.
[0012] In one aspect, the engineered thioesterase variant is selected from SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, and SEQ ID NO:46.
[0013] In one aspect, the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:15. In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50 and SEQ ID NO:51.
[0014] In one aspect, the engineered thioesterase variant has improved solubility. In one aspect, the engineered thioesterase variant has improved solubility as compared to SEQ ID NO:49.
[0015] In one aspect, the engineered thioesterase variant has a truncated mutation at amino acids 2 - 40 of SEQ ID NO:49. In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59.
[0016] In one aspect, the variant thioesterase has improved activity for the production of C10 fatty acid derivatives. In one aspect, the variant thioesterase has improved activity for the production of C8 fatty acid derivatives.
[0017] In one aspect, the present disclosure provides a recombinant host cell comprising one or more heterologous genes encoding a biochemical pathway for converting a first fatty acid derivative to a second fatty acid derivative, wherein the second fatty acid derivative has a higher minimum inhibitory concentration (MIC) than the first fatty acid derivative, and the presence of the second fatty acid derivative increases the MIC of the first fatty acid derivative.
[0018] In one aspect, the biochemical pathway comprises carboxylic acid reductase, carboxylic acid reductase and alcohol dehydrogenase, carboxylic acid reductase and alcohol-O-acetyltransferase, carboxylic acid reductase, and alcohol dehydrogenase, and alcohol O-acetyltransferase, ester synthase, ester synthase and fatty acid acyl-CoA synthetase, acyl-CoA reductase, acyl-CoA reductase and acyl-CoA synthetase, acyl-CoA reductase and alcohol O-acetyltransferase, acyl-CoA reductase, alcohol O-acetyltransferase, and acyl-CoA synthetase, O-methyltransferase, acyl-ACP reductase, acyl-ACP reductase and aldehyde decarbonylase, acyl-ACP reductase and aldehyde oxidative deformylase, acyl-ACP reductase and alcohol O-acetyltransferase, acyl-ACP reductase, alcohol-O-acetyltransferase, and alcohol dehydrogenase, OleA protein, OleA, OleC, and OleD proteins, OleA protein and fatty acid acyl-CoA synthetase, or OleA, OleC, and OleD proteins and one of fatty acid acyl-CoA synthetase.
[0019] In one aspect, the first fatty acid derivative is a fatty acid, the second fatty acid derivative is a fatty acid alkyl ester, and the biochemical pathway includes ester synthase and fatty acid acyl-CoA synthetase.
[0020] In one aspect, the fatty acid alkyl ester is a fatty acid methyl ester or a fatty acid ethyl ester.
[0021] In one aspect, the first fatty acid derivative is a fatty alcohol, the second fatty acid derivative is a fatty alcohol acetate ester, and the biochemical pathway includes carboxylic acid reductase and alcohol-O-acetyltransferase.
[0022] In one aspect, the first fatty acid derivative and the second fatty acid derivative are medium-chain fatty acid derivatives.
[0023] In one aspect, the recombinant host cell further includes an engineered thioesterase variant.
[0024] In one aspect, the engineered thioesterase variant is a member selected from the group consisting of variant thioesterases having an amino acid sequence with at least 90% sequence identity to SEQ ID NO:1 and having at least one substitution mutation at an amino acid position selected from the group consisting of 3, 4, 6, 14, 15, 17, 22, 37, 44, 45, 50, 54, 56, 64, 67, 73, 76, 91, 99, 102, 110, 111, 114, 129, 132, 137, 158, 162, 165, 176, 178, 185, 186, 196, 197, 198, 203, 213, 217, 225, 227, 236, 244, 254, 256, 258, 278, 282, 292, 297, 298, 299, 300, 301, 302, 316, 321, and 322.
[0025] In one aspect, at least one substitution mutation is: (a) lysine at amino acid position 3; (b) methionine at amino acid position 4; (c) arginine at amino acid position 6; (d) glycine or arginine at amino acid position 14; (e) leucine or tryptophan at amino acid position 15; (f) alanine or cysteine at amino acid position 17; (g) arginine at amino acid position 22; (h) proline at amino acid position 37; (i) glycine or isoleucine at amino acid position 44; (j) serine at position 45; (k) tryptophan at amino acid position 50; (l) arginine at amino acid position 54; (m) lysine or cysteine at amino acid position 56; (n) arginine or proline at amino acid position 64; (o) leucine at amino acid position 67; (p) valine at position 73; (q) phenylalanine or leucine or tyrosine at amino acid position 76; (r) methionine at amino acid position 91; (s) lysine or proline at amino acid position 99; (t) isoleucine at amino acid position 102; (u) leucine at amino acid position 110; (v) threonine at position 111; (w) lysine at position 114; (x) valine at amino acid position 129; (y) tryptophan at amino acid position 132; (z) cysteine at amino acid position 137; (aa) glutamine at amino acid position 158; (bb) glutamate at amino acid position 162; (cc) valine at amino acid position 176; (dd) proline at amino acid position 178; (ee) alanine at amino acid position 185; (ff) glycine at amino acid position 186; (gg) valine at amino acid position 196; (hh) asparagine at amino acid position 197; (ii) tryptophan at amino acid position 198; (jj) arginine at amino acid position 203; (kk) histidine or arginine at amino acid position 213; (ll) arginine at amino acid position 217; (mm) leucine at amino acid position 225; (nn) glycine at amino acid position 227; (oo) threonine at amino acid position 236; (pp) methionine or arginine at amino acid position 244; (qq) glycine at amino acid position 254; (rr) cysteine or arginine at amino acid position 256; (ss) threonine or valine at amino acid position 258; (tt) lysine or valine at amino acid position 278; (uu) serine or valine at amino acid position 282;(vv) phenylalanine at amino acid position 292; (ww) threonine or aspartic acid or valine at amino acid position 297; (xx) valine or cysteine at amino acid position 298; (yy) leucine at amino acid position 299; (zz) lysine or tryptophan or leucine at amino acid position 300; (aaa) cysteine at amino acid position 301; (bbb) threonine at amino acid position 302; (ccc) arginine at amino acid position 316; (ddd) arginine at amino acid position 321; and (eee) lysine at amino acid position 322, which is a member selected from the group consisting of;
[0026] In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59.
[0027] In one aspect, the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:1. In one aspect, the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:4. In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, and SEQ ID NO:46.
[0028] In one aspect, the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:15. In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50 and SEQ ID NO:51.
[0029] In one aspect, the engineered thioesterase variant has improved solubility. In one aspect, the engineered thioesterase variant has improved solubility compared to SEQ ID NO:49. In one aspect, the engineered thioesterase variant has a truncated mutation in amino acids 2-40 of SEQ ID NO:49. In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59.
[0030] In another aspect, the present disclosure provides a method for producing a medium-chain fatty acid derivative at a commercial titer, comprising culturing a recombinant host cell comprising an engineered thioesterase variant under conditions suitable for the production of the medium-chain fatty acid derivative in the presence of a carbon source, wherein the recombinant host cell comprises one or more heterologous genes encoding a biochemical pathway for converting a first fatty acid derivative to a second fatty acid derivative, the second fatty acid derivative has a higher minimum inhibitory concentration (MIC) than the first fatty acid derivative, and the presence of the second fatty acid derivative increases the MIC of the first fatty acid derivative.
[0031] In one aspect, the first fatty acid derivative is a medium-chain fatty acid, the second fatty acid derivative is a medium-chain fatty acid alkyl ester, and the biochemical pathway comprises an ester synthase and a fatty acid acyl-CoA synthetase.
[0032] In one aspect, the fatty acid alkyl ester is a medium-chain fatty acid methyl ester or a medium-chain fatty acid ethyl ester.
[0033] In one aspect, the first fatty acid derivative is a medium-chain fatty alcohol, the second fatty acid derivative is a medium-chain fatty alcohol acetate ester, and the biochemical pathway comprises a carboxylic acid reductase and an alcohol-O-acetyltransferase.
[0034] In one aspect, the engineered thioesterase variant has an amino acid sequence having at least 90% sequence identity with SEQ ID NO:1 and at least one substitution mutation at an amino acid position selected from the group consisting of 3, 4, 6, 14, 15, 17, 22, 37, 44, 45, 50, 54, 56, 64, 67, 73, 76, 91, 99, 102, 110, 111, 114, 129, 132, 137, 158, 162, 165, 176, 178, 185, 186, 196, 197, 198, 203, 213, 217, 225, 227, 236, 244, 254, 256, 258, 278, 282, 292, 297, 298, 299, 300, 301, 302, 316, 321, and 322. In one aspect, the at least one substitution mutation is (a) lysine at amino acid position 3; (b) methionine at amino acid position 4; (c) arginine at amino acid position 6; (d) glycine or arginine at amino acid position 14; (e) leucine or tryptophan at amino acid position 15; (f) alanine or cysteine at amino acid position 17; (g) arginine at amino acid position 22; (h) proline at amino acid position 37; (i) glycine or isoleucine at amino acid position 44; (j) serine at position 45; (k) tryptophan at amino acid position 50; (l) arginine at amino acid position 54; (m) lysine or cysteine at amino acid position 56; (n) arginine or proline at amino acid position 64; (o) leucine at amino acid position 67; (p) valine at position 73; (q) phenylalanine or leucine or tyrosine at amino acid position 76; (r) methionine at amino acid position 91; (s) lysine or proline at amino acid position 99; (t) isoleucine at amino acid position 102; (u) leucine at amino acid position 110; (v) threonine at position 111; (w) lysine at position 114; (x) valine at amino acid position 129; (y) tryptophan at amino acid position 132; (z) cysteine at amino acid position 137; (aa) glutamine at amino acid position 158; (bb) glutamate at amino acid position 162; (cc) valine at amino acid position 176; (dd) proline at amino acid position 178; (ee) alanine at amino acid position 185; (ff) glycine at amino acid position 186; (gg) valine at amino acid position 196; (hh) asparagine at amino acid position 197;(ii) Tryptophan at amino acid position 198; (jj) Arginine at amino acid position 203; (kk) Histidine or Arginine at amino acid position 213; (ll) Arginine at amino acid position 217; (mm) Leucine at amino acid position 225; (nn) Glycine at amino acid position 227; (oo) Threonine at amino acid position 236; (pp) Methionine or Arginine at amino acid position 244; (qq) Glycine at amino acid position 254; (rr) Cysteine or Arginine at amino acid position 256; (ss) Threonine or Valine at amino acid position 258; (tt) Lysine or Valine at amino acid position 278; (uu) Serine or Valine at amino acid position 282; (vv) Phenylalanine at amino acid position 292; (ww) Threonine or Aspartic acid or Valine at amino acid position 297; (xx) Valine or Cysteine at amino acid position 298; (yy) Leucine at amino acid position 299; (zz) Lysine or Tryptophan or Leucine at amino acid position 300; (aaa) Cysteine at amino acid position 301; (bbb) Threonine at amino acid position 302; (ccc) Arginine at amino acid position 316; (ddd) Arginine at amino acid position 321; and (eee) Lysine at amino acid position 322, which is a member selected from the group consisting of;
[0035] In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59.
[0036] In one aspect, the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:1.
[0037] In one aspect, the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:4. In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, and SEQ ID NO:46.
[0038] In one aspect, the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:15. In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50 and SEQ ID NO:51.
[0039] In one aspect, the engineered thioesterase variant has improved solubility. In one aspect, the engineered thioesterase variant has improved solubility as compared to SEQ ID NO:49. In one aspect, the engineered thioesterase variant has a truncated mutation at amino acids 2-40 of SEQ ID NO:49. In one aspect, the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59.
[0040] In another aspect, the present disclosure provides a composition of medium-chain fatty acid derivatives having a ratio of C8 fatty acid derivatives to C10 fatty acid derivatives (C8 / C10) of at least 3.6. In one aspect, the ratio of C8 fatty acid derivatives to C10 fatty acid derivatives is 7.7. [The present invention 1001] Engineered thioesterase variants having improved activity for the production of medium-chain fatty acid derivatives. [The present invention 1002] Engineered thioesterase variants of the present invention 1001 having improved activity for the production of C8 fatty acid derivatives. [The present invention 1003] Engineered thioesterase variants of the present invention 1001 having an amino acid sequence with at least 90% sequence identity to SEQ ID NO:1 and at least one substitution mutation at an amino acid position selected from the group consisting of 3, 4, 6, 14, 15, 17, 22, 37, 44, 45, 50, 54, 56, 64, 67, 73, 76, 91, 99, 102, 110, 111, 114, 129, 132, 137, 158, 162, 165, 176, 178, 185, 186, 196, 197, 198, 203, 213, 217, 225, 227, 236, 244, 254, 256, 258, 278, 282, 292, 297, 298, 299, 300, 301, 302, 316, 321, and 322. [The present invention 1004] At least one substitution mutation is present at (a) lysine at amino acid position 3; (b) methionine at amino acid position 4; (c) arginine at amino acid position 6; (d) glycine or arginine at amino acid position 14; (e) leucine or tryptophan at amino acid position 15; (f) alanine or cysteine at amino acid position 17; (g) arginine at amino acid position 22; (h) proline at amino acid position 37; (i) glycine or isoleucine at amino acid position 44; (j) serine at position 45; (k) tryptophan at amino acid position 50; (l) arginine at amino acid position 54; (m) lysine or cysteine at amino acid position 56; (n) arginine or proline at amino acid position 64; (o) leucine at amino acid position 67; (p) valine at position 73; (q) phenylalanine or leucine or tyrosine at amino acid position 76; (r) methionine at amino acid position 91; (s) lysine or proline at amino acid position 99; (t) isoleucine at amino acid position 102; (u) leucine at amino acid position 110; (v) threonine at position 111; (w) lysine at position 114; (x) valine at amino acid position 129; (y) tryptophan at amino acid position 132; (z) cysteine at amino acid position 137; (aa) glutamine at amino acid position 158; (bb) glutamate at amino acid position 162; (cc) valine at amino acid position 176; (dd) proline at amino acid position 178; (ee) alanine at amino acid position 185; (ff) glycine at amino acid position 186; (gg) valine at amino acid position 196; (hh) asparagine at amino acid position 197; (ii) tryptophan at amino acid position 198; (jj) arginine at amino acid position 203; (kk) histidine or arginine at amino acid position 213; (ll) arginine at amino acid position 217; (mm) leucine at amino acid position 225; (nn) glycine at amino acid position 227; (oo) threonine at amino acid position 236; (pp) methionine or arginine at amino acid position 244; (qq) glycine at amino acid position 254; (rr) cysteine or arginine at amino acid position 256; (ss) threonine or valine at amino acid position 258; (tt) lysine or valine at amino acid position 278; (uu) serine or valine at amino acid position 282;(vv) Phenylalanine at amino acid position 292; (ww) Threonine or aspartic acid or valine at amino acid position 297; (xx) Valine or cysteine at amino acid position 298; (yy) Leucine at amino acid position 299; (zz) Lysine or tryptophan or leucine at amino acid position 300; (aaa) Cysteine at amino acid position 301; (bbb) Threonine at amino acid position 302; (ccc) Arginine at amino acid position 316; (ddd) Arginine at amino acid position 321; and (eee) Lysine at amino acid position 322, which is a member selected from the group consisting of, the engineered thioesterase variant of the present invention 1003.; [The present invention 1005] The engineered thioesterase variant of the present invention 1004, which is a member selected from the group consisting of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59. [The present invention 1006] The engineered thioesterase variant of the present invention 1005, which has an overall increased effective positive charge compared to the thioesterase having SEQ ID NO:1. [The present invention 1007] The engineered thioesterase variant of the present invention 1006, which has an overall increased effective positive charge compared to the variant thioesterase having SEQ ID NO:4. [The present invention 1008] An engineered thioesterase variant of the present invention 1007, which is a member selected from the group consisting of SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, and SEQ ID NO:46. [The present invention 1009] An engineered thioesterase variant of the present invention 1006, which has an increased surface positive charge as compared with SEQ ID NO:15. [The present invention 1010] An engineered thioesterase variant of the present invention 1009, which is a member selected from the group consisting of SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50 and SEQ ID NO:51. [The present invention 1011] An engineered thioesterase variant of the present invention 1005, which has improved solubility. [The present invention 1012] An engineered thioesterase variant of the present invention 1011, which has improved solubility as compared with SEQ ID NO:49. [The present invention 1013] An engineered thioesterase variant of the present invention 1012, which has a truncated mutation at amino acids 2-40 of SEQ ID NO:49. [The present invention 1014] An engineered thioesterase variant of the present invention 1016, which is a member selected from the group consisting of SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59. [The present invention 1015] Variant thioesterase enzyme of the present invention 1001, wherein the variant thioesterase has improved activity for the production of C10 fatty acid derivatives. [The present invention 1016] A recombinant host cell comprising one or more heterologous genes encoding a biochemical pathway for converting a first fatty acid derivative to a second fatty acid derivative, wherein the second fatty acid derivative has a higher minimum inhibitory concentration (MIC) than the first fatty acid derivative, and the presence of the second fatty acid derivative increases the MIC of the first fatty acid derivative. [The present invention 1017] The biochemical pathway is a. Carboxylic acid reductase, b. Carboxylic acid reductase and alcohol dehydrogenase, c. Carboxylic acid reductase and alcohol-O-acetyltransferase, d. Carboxylic acid reductase, alcohol dehydrogenase, and alcohol O-acetyltransferase, e. Ester synthase, f. Ester synthase and fatty acid acyl-CoA synthetase, g. Acyl-CoA reductase, h. Acyl-CoA reductase and acyl-CoA synthetase, i. Acyl-CoA reductase and alcohol O-acetyltransferase, j. Acyl-CoA reductase, alcohol O-acetyltransferase, and acyl-CoA synthetase, k. O-methyltransferase, l. Acyl-ACP reductase, m. Acyl-ACP reductase and aldehyde decarbonylase, n. Acyl-ACP reductase and aldehyde oxidative deformylase, o. Acyl-ACP reductase and alcohol O-acetyltransferase, p. Acyl-ACP reductase, alcohol-O-acetyltransferase, and alcohol dehydrogenase, q. OleA protein, r. OleA, OleC, and OleD proteins, s. OleA protein and fatty acid acyl-CoA synthetase, or t. OleA, OleC, and OleD proteins and fatty acid acyl-CoA synthetase The recombinant host cell of the present invention 1016 comprising one of the above. [The present invention 1018] wherein the first fatty acid derivative is a fatty acid and the second fatty acid derivative is a fatty acid alkyl ester, and the biochemical pathway comprises ester synthase and fatty acid acyl-CoA synthetase. The recombinant host cell of the present invention 1017. [The present invention 1019] The fatty acid alkyl ester is a fatty acid methyl ester or a fatty acid ethyl ester, The recombinant host cell of the present invention 1018. [The present invention 1020] The first fatty acid derivative is a fatty alcohol, and the second fatty acid derivative is a fatty alcohol acetate ester, The biochemical pathway includes carboxylic acid reductase and alcohol - O - acetyltransferase, The recombinant host cell of the present invention 1017. [The present invention 1021] The first fatty acid derivative and the second fatty acid derivative are medium - chain fatty acid derivatives, The recombinant host cell of the present invention 1016. [The present invention 1021] Further comprising an engineered thioesterase variant, The recombinant host cell of the present invention 1016. [The present invention 1022] The engineered thioesterase variant is a member selected from the group consisting of variants having at least 90% sequence identity with SEQ ID NO:1 and having at least one substitution mutation at amino acid positions selected from the group consisting of 3, 4, 6, 14, 15, 17, 22, 37, 44, 45, 50, 54, 56, 64, 67, 73, 76, 91, 99, 102, 110, 111, 114, 129, 132, 137, 158, 162, 165, 176, 178, 185, 186, 196, 197, 198, 203, 213, 217, 225, 227, 236, 244, 254, 256, 258, 278, 282, 292, 297, 298, 299, 300, 301, 302, 316, 321, and 322, The recombinant host cell of the present invention 1023. [The present invention 1023] At least one substitution mutation is present at (a) lysine at amino acid position 3; (b) methionine at amino acid position 4; (c) arginine at amino acid position 6; (d) glycine or arginine at amino acid position 14; (e) leucine or tryptophan at amino acid position 15; (f) alanine or cysteine at amino acid position 17; (g) arginine at amino acid position 22; (h) proline at amino acid position 37; (i) glycine or isoleucine at amino acid position 44; (j) serine at position 45; (k) tryptophan at amino acid position 50; (l) arginine at amino acid position 54; (m) lysine or cysteine at amino acid position 56; (n) arginine or proline at amino acid position 64; (o) leucine at amino acid position 67; (p) valine at position 73; (q) phenylalanine or leucine or tyrosine at amino acid position 76; (r) methionine at amino acid position 91; (s) lysine or proline at amino acid position 99; (t) isoleucine at amino acid position 102; (u) leucine at amino acid position 110; (v) threonine at position 111; (w) lysine at position 114; (x) valine at amino acid position 129; (y) tryptophan at amino acid position 132; (z) cysteine at amino acid position 137; (aa) glutamine at amino acid position 158; (bb) glutamate at amino acid position 162; (cc) valine at amino acid position 176; (dd) proline at amino acid position 178; (ee) alanine at amino acid position 185; (ff) glycine at amino acid position 186; (gg) valine at amino acid position 196; (hh) asparagine at amino acid position 197; (ii) tryptophan at amino acid position 198; (jj) arginine at amino acid position 203; (kk) histidine or arginine at amino acid position 213; (ll) arginine at amino acid position 217; (mm) leucine at amino acid position 225; (nn) glycine at amino acid position 227; (oo) threonine at amino acid position 236; (pp) methionine or arginine at amino acid position 244; (qq) glycine at amino acid position 254; (rr) cysteine or arginine at amino acid position 256; (ss) threonine or valine at amino acid position 258; (tt) lysine or valine at amino acid position 278; (uu) serine or valine at amino acid position 282;(vv) phenylalanine at amino acid position 292; (ww) threonine or aspartic acid or valine at amino acid position 297; (xx) valine or cysteine at amino acid position 298; (yy) leucine at amino acid position 299; (zz) lysine or tryptophan or leucine at amino acid position 300; (aaa) cysteine at amino acid position 301; (bbb) threonine at amino acid position 302; (ccc) arginine at amino acid position 316; (ddd) arginine at amino acid position 321; and (eee) lysine at amino acid position 322, which is a member selected from the group consisting of, the recombinant cell of the present invention 1022; [The present invention 1024] The recombinant host cell of the present invention 1023, wherein the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59. [The present invention 1025] The recombinant host cell of the present invention 1024, wherein the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:1. [The present invention 1026] The recombinant host cell of the present invention 1025, wherein the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:4. [The present invention 1027] The recombinant host cell of the present invention 1026, wherein the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, and SEQ ID NO:46. [The present invention 1028] The recombinant host cell of the present invention 1025, wherein the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:15. [The present invention 1029] The recombinant host cell of the present invention 1028, wherein the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50 and SEQ ID NO:51. [The present invention 1030] The recombinant host cell of the present invention 1024, wherein the engineered thioesterase variant has improved solubility. [The present invention 1031] The recombinant host cell of the present invention 1030, wherein the engineered thioesterase variant has improved solubility as compared to SEQ ID NO:49. [The present invention 1032] The recombinant host cell of the present invention 1031, wherein the engineered thioesterase variant has a truncated mutation at amino acids 2-40 of SEQ ID NO:49. [The present invention 1033] The recombinant host cell of the present invention 1032, wherein the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59. [The present invention 1034] A method for producing a medium-chain fatty acid derivative at a commercial titer, comprising culturing a recombinant host cell containing an engineered thioesterase variant under conditions suitable for the production of the medium-chain fatty acid derivative in the presence of a carbon source, wherein the recombinant host cell contains one or more heterologous genes encoding a biochemical pathway for converting a first fatty acid derivative to a second fatty acid derivative, the second fatty acid derivative has a higher minimum inhibitory concentration (MIC) than the first fatty acid derivative, and the presence of the second fatty acid derivative increases the MIC of the first fatty acid derivative. A method. [The present invention 1035] wherein the first fatty acid derivative is a medium-chain fatty acid and the second fatty acid derivative is a medium-chain fatty acid alkyl ester, and the biochemical pathway includes an ester synthase and a fatty acid acyl-CoA synthetase. The method of the present invention 1034. [The present invention 1036] The method of the present invention 1035, wherein the fatty acid alkyl ester is a medium-chain fatty acid methyl ester or a medium-chain fatty acid ethyl ester. The method of the present invention 1035. [The present invention 1037] wherein the first fatty acid derivative is a medium-chain fatty alcohol and the second fatty acid derivative is a medium-chain fatty alcohol acetate ester, and the biochemical pathway includes a carboxylic acid reductase and an alcohol-O-acetyltransferase. The method of the present invention 1035. [The present invention 1038] The method of the present invention 1034, wherein the engineered thioesterase variant has an amino acid sequence with at least 90% sequence identity to SEQ ID NO:1 and has at least one substitution mutation at an amino acid position selected from the group consisting of 3, 4, 6, 14, 15, 17, 22, 37, 44, 45, 50, 54, 56, 64, 67, 73, 76, 91, 99, 102, 110, 111, 114, 129, 132, 137, 158, 162, 165, 176, 178, 185, 186, 196, 197, 198, 203, 213, 217, 225, 227, 236, 244, 254, 256, 258, 278, 282, 292, 297, 298, 299, 300, 301, 302, 316, 321, and 322. [The present invention 1039] At least one substitution mutation is at (a) lysine at amino acid position 3; (b) methionine at amino acid position 4; (c) arginine at amino acid position 6; (d) glycine or arginine at amino acid position 14; (e) leucine or tryptophan at amino acid position 15; (f) alanine or cysteine at amino acid position 17; (g) arginine at amino acid position 22; (h) proline at amino acid position 37; (i) glycine or isoleucine at amino acid position 44; (j) serine at position 45; (k) tryptophan at amino acid position 50; (l) arginine at amino acid position 54; (m) lysine or cysteine at amino acid position 56; (n) arginine or proline at amino acid position 64; (o) leucine at amino acid position 67; (p) valine at position 73; (q) phenylalanine or leucine or tyrosine at amino acid position 76; (r) methionine at amino acid position 91; (s) lysine or proline at amino acid position 99; (t) isoleucine at amino acid position 102; (u) leucine at amino acid position 110; (v) threonine at position 111; (w) lysine at position 114; (x) valine at amino acid position 129; (y) tryptophan at amino acid position 132; (z) cysteine at amino acid position 137; (aa) glutamine at amino acid position 158; (bb) glutamate at amino acid position 162; (cc) valine at amino acid position 176; (dd) proline at amino acid position 178; (ee) alanine at amino acid position 185; (ff) glycine at amino acid position 186; (gg) valine at amino acid position 196; (hh) asparagine at amino acid position 197; (ii) tryptophan at amino acid position 198; (jj) arginine at amino acid position 203; (kk) histidine or arginine at amino acid position 213; (ll) arginine at amino acid position 217; (mm) leucine at amino acid position 225; (nn) glycine at amino acid position 227; (oo) threonine at amino acid position 236; (pp) methionine or arginine at amino acid position 244; (qq) glycine at amino acid position 254; (rr) cysteine or arginine at amino acid position 256; (ss) threonine or valine at amino acid position 258; (tt) lysine or valine at amino acid position 278; (uu) serine or valine at amino acid position 282;(vv) phenylalanine at amino acid position 292; (ww) threonine or aspartic acid or valine at amino acid position 297; (xx) valine or cysteine at amino acid position 298; (yy) leucine at amino acid position 299; (zz) lysine or tryptophan or leucine at amino acid position 300; (aaa) cysteine at amino acid position 301; (bbb) threonine at amino acid position 302; (ccc) arginine at amino acid position 316; (ddd) arginine at amino acid position 321; and (eee) lysine at amino acid position 322, a member selected from the group consisting of: the method of the present invention 1036; [The present invention 1038] The engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59, The method of the present invention 1037. [The present invention 1039] The engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:1, The method of the present invention 1038. [The present invention 1040] The engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:4, The method of the present invention 1039. [The present invention 1041] The method of the present invention 1040, wherein the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, and SEQ ID NO:46. [The present invention 1042] The method of the present invention 1039, wherein the engineered thioesterase variant has an increased surface positive charge as compared to SEQ ID NO:15. [The present invention 1043] The method of the present invention 1042, wherein the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50 and SEQ ID NO:51. [The present invention 1044] The method of the present invention 1038, wherein the engineered thioesterase variant has improved solubility. [The present invention 1045] The method of the present invention 1044, wherein the engineered thioesterase variant has improved solubility as compared to SEQ ID NO:49. [The present invention 1046] The method of the present invention 1045, wherein the engineered thioesterase variant has a truncated mutation in amino acids 2-40 of SEQ ID NO:49. [The present invention 1047] The method of the present invention 1046, wherein the engineered thioesterase variant is a member selected from the group consisting of SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 and SEQ ID NO:59. [The present invention 1048] A composition of medium-chain fatty acid derivatives having a ratio (C8 / C10) of C8 fatty acid derivatives to C10 fatty acid derivatives of at least 3.6. [Invention 1049] The composition of Invention 1048, wherein the ratio of C8 fatty acid derivatives to C10 fatty acid derivatives is 7.7.
Brief Description of the Drawings
[0041] [Figure 1] Illustrate the minimum inhibitory concentration (MIC) curves for different C8 aliphatic compounds. [Figure 2] Illustrate the partition coefficients (logPwo) of different medium-chain aliphatic compounds. [Figure 3] Illustrate the protection from the toxicity of 1-octanol in the presence of octyl acetate. When exposed to 1-octanol, the viability of Escherichia coli (E. coli) cells was completely lost after 5 hours of exposure. However, when 50 g / L (non-toxic concentration for E. coli cells) of octyl acetate was also added, the decrease in cell viability was less than 20% in the presence of up to 10 g / L of 1-octanol. [Figure 4] Illustrate the pathways for the production of medium-chain fatty alcohols and their acetylation to their fatty acid acetates. R: CH3(CH2)n [where n = 1, 2, 3, 4 or 5]; FFA: free fatty acid; FALD: fatty aldehyde; FALC: fatty alcohol; FACE: fatty alcohol acetate; ACP: acyl carrier protein; AAR: acyl-ACP reductase; ADH: aldehyde / alcohol dehydrogenase; TE: thioesterase; ACR: acyl-CoA reductase; CAR: carboxylic acid reductase; AAT: o-alcohol acetyltransferase. [Figure 5A] Illustrate different metrics showing the improved tolerance and production of medium-chain fatty alcohol (FALC) compounds by the expression of alcohol acetyltransferase. Figure 5A illustrates that the FALC-producing strain (sRG.674) could not grow on minimal salt medium with glucose as the carbon source. In contrast, there was no growth inhibition on the same medium with the expression of o-alcohol acetyltransferase (AAT) in the sJN.209 strain. [Figure 5B]Illustrates different metrics showing improved tolerance and production of medium-chain fatty alcohol (FALC) compounds due to the expression of alcohol acetyltransferase. Figure 5B illustrates the levels of total fatty species (FAS) produced by the FALC-producing strain (sRG.674) and the AAT-expressing strain sJN.209. [Figure 5C] Illustrates different metrics showing improved tolerance and production of medium-chain fatty alcohol (FALC) compounds due to the expression of alcohol acetyltransferase. Figure 5C illustrates a comparison of the levels and composition of fatty species produced by the FALC-producing strain (sRG.674) with the AAT-expressing strain sJN.209. [Figure 6] Illustrates the pathway for the esterification of free fatty acids. R: CH3(CH2)n [where n = 1, 2, 3, 4 or 5]; FFA: free fatty acid; FAEE: fatty acid ethyl ester; TE: thioesterase; ES: ester synthase. [Figure 7A] Illustrates different metrics showing improved survival rate and production of medium-chain fatty acid derivatives by strains expressing the medium-chain alkyl ester biosynthesis pathway, compared to strains expressing only the medium-chain length fatty acid biosynthesis pathway. The sRS.786 strain has been engineered to express the medium-chain length thioesterase (chFatB2) and produces only free fatty acids (FFA). The Stpay.179 strain is isogenic to sRS.786 and also expresses fatty acid acyl-CoA synthetase and ester synthase, and produces medium-chain length fatty alkyl esters when short-chain alcohols (e.g., methanol, ethanol, etc.) are provided in the medium. The sRS.786 and Stpay.179 strains were grown in minimal salt medium with glucose as the carbon source. Additionally, ethanol was supplied during the course of the fermentation run to maintain the alcohol concentration at approximately 2 g / L. In Figure 7A, the strain producing FFA alone (sRS.786) stopped growing and consuming glucose approximately 10 hours after the addition of IPTG to induce the expression of the medium-chain length acyl-ACP thioesterase. In contrast, the Stpay.179 strain expressing the esterification pathway was able to continue growing after IPTG induction. [Figure 7B]Figure 7B. The sRS.786 strain stopped producing medium-chain fatty acid species (FAS) approximately 10 hours after the addition of IPTG, induced the expression of medium-chain-length acyl-ACP thioesterase, and ultimately produced only approximately 5 g of C8 + C10. In contrast, the Stpay.179 strain continued to grow and produce FAS throughout the fermentation run, ultimately producing over 84 g / kg of total fatty acid species. [Figure 7C] Figure 7C. The Stpay.179 strain expressing the esterification pathway was able to grow and produce total fatty acid species at a titer over 84 g / kg, and 93% of the total fatty acid species were C8-C10 FFA. [Figure 8] Plasmid pIR.108 is illustrated. [Figure 9] The structure-based sequence alignment used to construct the model of SEQ ID NO:1 disclosed in Example 6 is illustrated. [Figure 10] The final full-length model for the 3D structure of SEQ ID NO:1 is illustrated. Surface residues are shown as balls and sticks. [Figure 11] Western blot (1 = whole cell fraction, 2 = soluble fraction) to evaluate the solubility of various FatB2 truncations. [Figure 12] The characteristic final product composition of the production of medium-chain-length fatty alcohol acetates by the sRG.825 and sDH.377 strains when cultured under the conditions of Example 8 is illustrated. [Figure 13] The characteristic final product composition of the production of medium-chain-length fatty acid ethyl esters by the sAZ918 strain when cultured under the conditions of Example 11 is illustrated.
Mode for Carrying Out the Invention
[0042] Detailed Description Definitions As used in this specification and the appended claims in the context of describing an element, the articles “a,” “an,” and “the” and similar designations referring to the singular, such as those used herein, are to be construed to include the singular and the plural unless specifically indicated otherwise herein or clearly contradicted by the context. Thus, for example, reference to “a host cell” includes two or more such host cells, reference to “a nucleic acid sequence” includes one or more nucleic acid sequences, reference to “an enzyme” includes one or more enzymes, and so forth for the others.
[0043] As used herein, “about” is understood by those of ordinary skill in the art and may vary to some extent depending on the context in which it is used. When the term “about” is not clear to those of ordinary skill in the art and is used in a context where it is not clear, “about” means up to plus or minus 10% of a particular term.
[0044] As will be understood by those of ordinary skill in the art, for any and all purposes, all ranges disclosed herein also include any and all possible subranges and combinations of subranges thereof. Further, as will be understood by those of ordinary skill in the art, ranges include each individual member. Thus, for example, a group having 1 to 3 atoms represents a group having 1, 2, or 3 atoms. Similarly, a group having 1 to 5 atoms represents a group having 1, 2, 3, 4, or 5 atoms, and so forth.
[0045] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. In particular, the present disclosure utilizes conventional techniques in the fields of recombinant genetics, organic chemistry, fermentation, and biochemistry. Basic texts that disclose general terms in molecular biology and genetics include, for example, Lackie, Dictionary of Cell and Molecular Biology, Elsevier (5th ed. 2013). Basic texts that disclose general methods and terms in biochemistry include, for example, Lehninger Principles of Biochemistry Sixth edition, David L. Nelson and Michael M. Cox eds. W.H. Freeman (2012). Basic texts that disclose general methods and terms in fermentation include, for example, Principles of Fermentation Technology, 3rd Edition by Peter F Stanbury, Allan Whitaker and Stephen J Hall. Butterworth-Heinemann (2016). Basic texts that disclose general methods and terms in organic chemistry include, for example, Favre, Henri A. and Powell, Warren H. Nomenclature of Organic Chemistry. IUPAC Recommendations and Preferred Name 2013. Cambridge, UK: The Royal Society of Chemistry, 2013; Practical Synthetic Organic Chemistry: Reactions, Principles, and Techniques, Stephane Caron ed., John Wiley and Sons Inc. (2011); Organic Chemistry, 9th Edition - Francis Carey and Robert Giuliano, McGraw Hill (2013).
[0046] Throughout this specification, sequence accession numbers are from databases provided by the NCBI (National Center for Biotechnology Information), which is maintained by the National Institutes of Health (referred to herein as "NCBI accession numbers" or alternatively "GenBank accession numbers" or alternatively simply "accession numbers"), as well as from the UniProt Knowledgebase (UniProtKB) and Swiss-Prot databases provided by the Swiss Institute of Bioinformatics (referred to herein as "UniProtKB accession numbers").
[0047] Enzyme Commission (EC) numbers are established by the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (IUBMB), and their descriptions are available from the IUBMB Enzyme Nomenclature website on the World Wide Web. EC numbers classify an enzyme according to the reaction it catalyzes. For example, the enzyme activity of thioesterases is classified into E.C. 3.1.2.1 - 3.1.2.27 and 3.1.2.-. Specific classifications are based on the activities of different thioesterases towards different substrates.
[0048] For example, in some exemplary embodiments, thioesterases that catalyze the hydrolysis of thioester bonds of C6 - C18 alkyl thioesters, such as acyl - acyl carrier protein thioesters (acyl - ACP) and acyl - coenzyme A thioesters (acyl - CoA), are classified from E.C. 3.1.2.- to 3.1.2.14. Thioesterases are present in most prokaryotes and in the chloroplasts of most plants and algae. The functionality of thioesterases is conserved on a species - by - species basis in most prokaryotes. Thus, different microbial species can exhibit the same thioesterase enzyme activity as those classified into E.C. 3.1.2.1 - 3.1.2.27 and 3.1.2.-.
[0049] As used herein, the term "fatty acid" refers to an aliphatic carboxylic acid having the formula RCOOH, where R is an aliphatic group having at least 4 carbons, typically from about 4 to about 28 carbon atoms. The aliphatic R group can be saturated or unsaturated, branched or unbranched. Unsaturated "fatty acids" can be mono-unsaturated or poly-unsaturated.
[0050] One or more "fatty acids" as used herein can be produced intracellularly or supplied to cells via the fatty acid biosynthetic process or via the reverse of fatty acid beta-oxidation. As is well known in the art, fatty acid biosynthesis is generally the malonyl-CoA-dependent synthesis of acyl-ACP, whereas the reverse of beta-oxidation results in acyl-CoA. Fatty acids supplied to cells are converted to acyl-CoA.
[0051] The biosynthesis and breakdown of fatty acids occur in all biological forms, including prokaryotes, single-celled eukaryotes, higher eukaryotes, and archaea. The tools and methods disclosed herein are useful for the production of medium-chain fatty acid derivatives derivatized by any one or more of fatty acid synthesis, breakdown, or supply in any organism that naturally produces alkyl thioesters.
[0052] As used herein, the term "medium-chain fatty acid" or "medium-chain length fatty acid" in a similar sense refers to a fatty acid having a carbon chain length of 6 to 10. Thus, in some exemplary embodiments, a "medium-chain fatty acid" is a fatty acid having a carbon chain length of 6 carbons, 7 carbons, 8 carbons, 9 carbons, or 10 carbons.
[0053] As used herein, the term "fatty acid derivative" refers to a product produced by derivatization from a fatty acid. Thus, "fatty acid derivatives" include the "fatty acids" and "medium-chain fatty acids" defined above. Generally, "fatty acid derivatives" include malonyl-CoA-derived compounds including acyl-ACP or acyl-ACP derivatives. "Fatty acid derivatives" also include malonyl-CoA-derived compounds such as acyl-CoA or acyl-CoA derivatives. Thus, "fatty acid derivatives" include molecules / compounds obtained from metabolic pathways including the thioesterase reaction. Exemplary fatty acid derivatives include fatty acids, fatty acid esters (e.g., waxes, fatty acid esters, fatty acid methyl esters (FAME), fatty acid ethyl esters (FAEE)), fatty alcohol acetic esters (FACE), fatty amines, fatty aldehydes, fatty alcohols, hydrocarbons such as alkanes, alkenes, etc., ketones, terminal olefins, internal olefins, 3-hydroxy fatty acid derivatives, bifunctional fatty acid derivatives (e.g., ω-hydroxy fatty acids, 1,3 fatty diols, α,ω-diols, α,ω-3-hydroxy triols, ω-hydroxy FAME, ω-OH FAEE, etc.), and unsaturated fatty acid derivatives including unsaturated compounds of each of the above fatty acid derivatives.
[0054] As used herein, the expression "fatty acid derivative composition" refers to a composition of fatty acid derivatives, for example, a fatty acid composition produced by an organism. The "fatty acid derivative composition" may contain a single fatty acid derivative species or a mixture of fatty acid derivative species. In some exemplary embodiments, the mixture of fatty acid derivatives contains more than one of fatty acid derivative products (e.g., fatty acids, fatty acid esters, fatty alcohols, fatty alcohol acetic acid esters, fatty aldehydes, fatty amines, bifunctional fatty acid derivatives, etc.). In other exemplary embodiments, the mixture of fatty acid derivatives contains a mixture of fatty acid esters (or another fatty acid derivative) having different chain lengths, degrees of saturation and / or branching characteristics. In other exemplary embodiments, the mixture of fatty acid derivatives mainly contains one kind of fatty acid derivative, for example, a medium-chain fatty acid derivative composition. In still other exemplary embodiments, the mixture of fatty acid derivatives contains more than one fatty acid derivative product, for example, a mixture of fatty acid derivatives having different chain lengths, degrees of saturation and / or branching characteristics. In still other exemplary embodiments, the mixture of fatty acid derivatives contains a mixture of fatty esters and beta-hydroxy esters. In still other exemplary embodiments, the fatty acid derivative composition contains a mixture of fatty alcohols and fatty aldehydes. In still other exemplary embodiments, the fatty acid derivative composition contains a mixture of FAME and / or FAEE, particularly a mixture of medium-chain FAME and / or FAEE. In still other exemplary embodiments, the fatty acid derivative composition contains a mixture of fatty alcohol acetic acid esters (FACE), particularly a mixture of medium-chain fatty alcohol acetic acid esters (FACE).
[0055] As used herein, the term "nucleotide" takes on its conventional meaning known in the art. In addition to representing natural ribonucleotide or deoxyribonucleotide monomers, the term "nucleotide" encompasses nucleotide analogs and modified nucleotides such as amino-modified nucleotides. Additionally, "nucleotide" includes non-natural analog structures. Thus, for example, the individual units of peptide nucleic acids each containing a base may be referred to as nucleotides herein.
[0056] The term "polynucleotide" typically refers to a polymer of phosphodiester-linked ribonucleotides (RNA) or deoxyribonucleotides (DNA), which can be single-stranded or double-stranded and may contain natural and / or non-natural and / or modified nucleotides. The terms "polynucleotide", "nucleic acid sequence", and "nucleotide sequence" are used interchangeably herein to refer to a polymeric form of nucleotides of either RNA or DNA of any length. These terms represent the primary structure of the molecule and thus include polynucleotides such as single-stranded, double-stranded, triple-stranded, quadruple-stranded, partially double-stranded, branched, hairpin-shaped, circular, padlocked conformation, etc. These terms include, by way of non-limiting equivalents, any analogs of RNA or DNA made from nucleotide analogs and modified polynucleotides such as methylated polynucleotides and / or capped polynucleotides. A polynucleotide can be in any form including, without limitation, plasmids, viruses, chromosomes, ESTs, cDNAs, mRNAs, and rRNAs, and may be prepared by any known method including synthesis, recombination, ex vivo generation, or a combination thereof, and utilizing any purification method known in the art.
[0057] As used herein, the terms "polypeptide" and "protein" are used interchangeably and typically refer to polymers of amino acid residues that are at least 12 amino acids in length. Polypeptides less than 12 amino acids in length are referred to herein as "peptides". The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical mimics of the corresponding natural amino acids, as well as to natural and non-natural amino acid polymers. The term "recombinant polypeptide" generally refers to a polypeptide produced by recombinant techniques where DNA or RNA encoding the expressed protein is inserted into an appropriate expression vector, which is then used to transform a host cell to produce the polypeptide. In some exemplary embodiments, the DNA or RNA encoding the expressed peptide, polypeptide or protein is inserted into the host chromosome by homologous recombination or other means well known in the art, and then used to transform the host cell to produce the peptide or polypeptide. Similarly, the terms "recombinant polynucleotide" or "recombinant nucleic acid" or "recombinant DNA" refer to those produced by recombinant techniques known to those of skill in the art (e.g., see the methods described in Sambrook et al., Molecular Cloning--A Laboratory Manual, Cold Spring Harbor Press 4 th Edition (Cold Spring Harbor, N.Y. 2012) or Current Protocols in Molecular Biology Volumes 1-3, John Wiley & Sons, Inc. (1994-1998) and Supplements 1-115 (1987-2016)).
[0058] The term "amino acid" refers to natural and synthetic amino acids, as well as amino acid analogs and mimetics that function in a manner similar to natural amino acids. Natural amino acids are those encoded by the genetic code, as well as amino acids that are later modified, such as hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. Amino acid analogs are compounds that have the same basic chemical structure as natural amino acids, i.e., a compound having a carbon bonded to hydrogen, a carboxyl group, an amino group, and an R group, such as homoserine, norleucine, methionine sulfoxide, and methionine methyl sulfonium. Such analogs have a modified R group (e.g., norleucine) or a modified peptide backbone, but retain the same basic chemical structure as natural amino acids. Naturally encoded amino acids are the 20 common amino acids (alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine) as well as pyrroline and selenocysteine. In some exemplary embodiments, the single-letter codes shown in the table below are used to represent specific members of the 20 common natural amino acids. Single-letter amino acid codes are well known in the art (see, e.g., Lehninger, supra). TIFF0007705243000001.tif56128
[0059] When referring to two nucleotide sequences or polypeptide sequences, the "sequence identity percentage" between the two sequences is determined by comparing the two optimally aligned sequences over a comparison window, where for optimal alignment of the two sequences, the polynucleotide sequence portion in the comparison window may include additions or deletions (i.e., gaps) as compared to the reference sequence (excluding additions and deletions). The "sequence identity percentage" is calculated by determining the number of positions at which the identical nucleic acid bases or amino acid residues occur in both sequences, yielding the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to yield the sequence identity percentage.
[0060] Accordingly, the term "percent identity" or "sequence identity" in a similar sense, as related to two or more nucleic acid sequences or peptides or polypeptides, refers to sequences or subsequences that are the same or have a specified percentage of the same nucleotides or amino acids when measured, for example, using the default parameters of the BLAST or BLAST 2.0 sequence comparison algorithms (e.g., Altschul et al. (1990) J. Mol. Biol. 215(3):403-410) and / or referring to the NCBI website at ncbi.nlm.nih.gov / BLAST / or by manual alignment and visual inspection (e.g., about 50% identity, preferably 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a particular region when compared over a comparison window or specified region and aligned for maximum correspondence). The percent sequence identity between two nucleic acid sequences or amino acid sequences can also be determined, for example, using the Needleman and Wunsch algorithm incorporated into the GAP program in the GCG software package, using either a Blossum 62 matrix or a PAM250 matrix and gap weights of 16, 14, 12, 10, 8, 6, or 4 and length weights of 1, 2, 3, 4, 5, or 6 (Needleman and Wunsch (1970) J. Mol. Biol. 48:444-453). The percent sequence identity between two nucleotide sequences can also be determined using the GAP program in the GCG software package, using the NWSgapdna.CMP matrix and gap weights of 40, 50, 60, 70, or 80 and length weights of 1, 2, 3, 4, 5, or 6. One skilled in the art can perform the initial calculation of sequence identity and appropriately adjust the parameters of the algorithm.A set of parameters that can be used when the practitioner does not know which parameters to apply to determine whether a molecule is within the limits of claim equivalence is the Blossum 62 scoring matrix with a gap penalty of 12, a gap extension penalty of 4, and a frameshift gap penalty of 5. Additional methods of sequence alignment are known in the art of biotechnology (see, for example, Rosenberg (2005) BMC Bioinformatics 6:278; Altschul et al. (2005) FEBS J. 272(20):5101-5109).
[0061] When two or more nucleic acid or amino acid sequences are aligned and analyzed as described above and are found to share at least about 50% identity, preferably 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a particular region, they are said to be "substantially identical." Two nucleic acid or polypeptide sequences are said to be "identical" if, when aligned for maximum correspondence as described above, the sequences of nucleotides or amino acid residues in the two sequences are the same. This definition can also apply to, or represent, the complement of a test sequence. Identity is typically calculated over a region that is at least about 25 amino acids or nucleotides in length, more preferably over a region that is 50 to 100 amino acids or nucleotides in length, or over the full length of a given sequence.
[0062] The phrase "hybridizes under low stringency, medium stringency, high stringency, or ultra-high stringency conditions" describes the conditions for hybridization and washing. Guidance for performing the hybridization reactions can be found, for example, in Current Protocols in Molecular Biology, John Wiley & Sons, N.Y. (1989), 6.3.1 - 6.3.6. Both aqueous and non-aqueous methods are described in the references cited and either method can be used. The specific hybridization conditions referred to herein are as follows: (1) Low stringency hybridization conditions -- 6× sodium chloride / sodium citrate (SSC) at about 45° C. followed by washing twice in 0.2× SSC, 0.1% SDS at at least 50° C. (the wash temperature can be raised to 55° C. for low stringency conditions); (2) Medium stringency hybridization conditions -- 6× SSC at about 45° C. followed by washing one or more times in 0.2× SSC, 0.1% SDS at 60° C.; (3) High stringency hybridization conditions -- 6× SSC at about 45° C. followed by washing one or more times in 0.2× SSC, 0.1% SDS at 65° C.; and (4) Ultra-high stringency hybridization conditions -- 0.5 M sodium phosphate, 7% SDS at 65° C. followed by washing one or more times in 0.2× SSC, 1% SDS at 65° C. Unless otherwise stated, ultra-high stringency conditions (4) are the preferred conditions.
[0063] As used herein, the term "endogenous" refers to substances produced within a cell, such as nucleic acids, proteins, etc. Thus, an "endogenous" polynucleotide or polypeptide refers to a polynucleotide or polypeptide produced by a cell. In some exemplary embodiments, an "endogenous" polypeptide or polynucleotide is encoded by the genome of a parental cell (or host cell). In other exemplary embodiments, an "endogenous" polypeptide or polynucleotide is encoded by an autonomously replicating plasmid carried by a parental cell (or host cell). In some exemplary embodiments, an "endogenous" gene is a gene that was present in a cell when the cell was originally isolated from nature, i.e., the gene is "native to the cell". In other exemplary embodiments, an "endogenous" gene has been modified by recombinant techniques, for example, by altering the relationship between regulatory and coding sequences. Thus, a "heterologous" gene can, in some exemplary embodiments, be "endogenous" to a host cell.
[0064] In contrast, an "exogenous" polynucleotide or polypeptide, or other substance (e.g., fatty acid derivatives, small molecule compounds, etc.) refers to a polynucleotide or polypeptide or other substance that is not produced by a parental cell and is thus added to a cell, cell culture, or assay from outside the cell.
[0065] As used herein, the term "native" refers to a nucleic acid, protein, polypeptide, or fragment thereof in the form isolated from nature, or a nucleic acid, protein, polypeptide, or fragment thereof that has no intentionally introduced mutations.
[0066] As used herein, the term "fragment" of a polypeptide refers to a shorter portion of a full-length polypeptide or protein in the size range from 2 amino acid residues to the entire amino acid sequence minus 1 amino acid residue. In certain embodiments of the present disclosure, a fragment represents the entire amino acid sequence of a domain (e.g., a substrate-binding domain or a catalytic domain) of a polypeptide or protein.
[0067] The term "mutation induction" refers to the process by which the genetic information of an organism is stably altered to produce a "mutant" or "variant". By inducing mutations in a protein-coding nucleic acid sequence to produce a mutant nucleic acid sequence, a mutant protein is produced. Mutation induction also refers to the alteration of non-coding nucleic acid sequences. In some exemplary embodiments, mutations in non-coding nucleic acid sequences result in altered protein activity.
[0068] Accordingly, as used herein, "mutation" refers to a permanent change at a nucleic acid position of a gene or an amino acid position (residue) of a polypeptide or protein. In fact, the term "mutation" in the context of a polynucleotide refers to a modification to a polynucleotide sequence that results in a change in the polynucleotide sequence relative to a control or reference polynucleotide sequence. In some exemplary embodiments, a mutant polynucleotide sequence represents a modification related to codon optimization for expression purposes that, for example, does not alter the encoded amino acid sequence. In other exemplary embodiments, a mutation in a polynucleotide sequence modifies the codons to result in a modification of the encoded amino acid sequence. Accordingly, a polynucleotide encoding an engineered thioesterase variant having an improved ability to produce medium-chain fatty acid derivatives has at least one mutation compared to a polynucleotide encoding a control thioesterase.
[0069] Similarly, with respect to proteins, the terms "variant" or "varied" refer to a modification to an amino acid sequence that results in a change in a protein sequence relative to a control or reference protein sequence. A variant can represent a substitution of one amino acid by another amino acid, or an insertion or deletion of one or more amino acid residues. In some exemplary embodiments, a "variant" is a substitution of an amino acid with a non-natural amino acid or a chemically modified amino acid residue. In other exemplary embodiments, a "variant" is a shortening (e.g., deletion or interruption) of a sequence or subsequence compared to a precursor sequence, or a sequence shortening by deletion from one or the other terminus. In other exemplary embodiments, a variant is an addition of an amino acid or a subsequence either within a protein or at either terminus of the protein (e.g., two or more amino acids in a stretch inserted between two adjacent amino acids in a precursor protein sequence), whereby the length of the protein is increased (or extended). Variants can be introduced into polynucleotides by a number of methods known to those of skill in the art, including, for example, random mutagenesis, site-directed mutagenesis, oligonucleotide-directed mutagenesis, gene shuffling, directed evolution techniques, combinatorial mutagenesis, chemical synthesis, site-saturation mutagenesis, and the like.
[0070] As used herein, the term "variant" or "variety" in a similar sense refers to a polynucleotide sequence or a polypeptide sequence that contains at least one variant. Thus, an engineered thioesterase variant having an improved ability to produce medium-chain fatty acid derivatives has at least one variant in its polypeptide sequence compared to a control thioesterase.
[0071] As used herein, the term "engineered thioesterase variant" refers to a variant or variant thioesterase having at least one variant compared to SEQ ID NO:1, wherein the thioesterase variant has an improved activity for the production of medium-chain fatty acid derivatives.
[0072] As used herein, the term "gene" refers to a nucleic acid sequence, such as a DNA sequence encoding either an RNA product or a protein product, and a functionally linked nucleic acid sequence (e.g., an expression control sequence, such as a promoter, enhancer, ribosome binding site, translational control sequence, etc.) that affects the expression of the RNA product or protein product. The term "gene product" refers to either an RNA, such as tRNA, mRNA, and / or a protein expressed from a particular gene.
[0073] As used herein, the terms "expression" or "expressed" with respect to a gene refer to the production of one or more transcriptional and / or translational products of the gene. In an exemplary embodiment, the expression level of a DNA molecule in a cell is determined based on either the amount of the corresponding mRNA present in the cell or the amount of the protein encoded by the DNA produced by the cell. The term "expressed gene" refers to a gene that is transcribed into messenger RNA (mRNA) and then translated into a protein, as well as other types of RNA, such as transfer RNA (tRNA), ribosomal RNA (rRNA), and regulatory RNA that are transcribed but not translated into a protein.
[0074] The expression level of a nucleic acid molecule in a cell line or cell-free system is affected by an "expression control sequence" or a "regulatory sequence" in a similar sense. "Expression control sequences" or "regulatory sequences" are known in the art and include, for example, promoters, enhancers, polyadenylation signals, transcription termination factors, nucleotide sequences that affect RNA stability, internal ribosome entry sites (IRES) within the sequence, etc., which provide for the expression of a polynucleotide sequence in a host cell. In an exemplary embodiment, an "expression control sequence" specifically interacts with a cellular protein involved in transcription (see, for example, Maniatis et al., Science, 236: 1237-1245 (1987); Goeddel, Gene Expression Technology: Methods in Enzymology, Vol. 185, Academic Press, San Diego, Calif. (1990)). In an exemplary method, an expression control sequence is operably linked to a polynucleotide sequence. By "operably linked", it is meant that the polynucleotide sequence and the expression control sequence are functionally connected such that the expression of the polynucleotide sequence is enabled when an appropriate molecule (e.g., a transcriptional activator protein) contacts the expression control sequence. In an exemplary embodiment, a promoter that is operably linked is located upstream of the selected polynucleotide sequence with respect to the direction of transcription and translation. In some exemplary embodiments, an enhancer that is operably linked can be located upstream, within, or downstream of the selected polynucleotide.
[0075] Generally, the "minimum inhibitory concentration" (MIC) is the lowest concentration of an antimicrobial agent that inhibits visible growth of a microorganism after an overnight incubation. The MIC can be determined on plates of solid growth medium or by the broth dilution method. For example, to identify the MIC by liquid dilution, the same amount of bacteria is cultured in wells of a liquid medium containing serially decreasing concentrations of the drug. The minimum inhibitory concentration of the antibiotic lies between the concentration of the last well in which the bacteria did not grow and the next lower amount in which the bacteria grew. As used herein, the expression "minimum inhibitory concentration" or "MIC" refers to the concentration of a compound that results in a 50% decrease in the growth of a microbial culture during a 24-hour incubation period as compared to a control. In one embodiment, the "minimum inhibitory concentration" of a potentially toxic compound, such as octanol, is measured by growing a culture of cells, e.g., E. coli cells, in various concentrations of the potentially toxic compound and then determining how much growth of the culture occurred over a 24-hour period in the presence of the potentially toxic compound. In an exemplary embodiment, growth of the culture is measured by measuring total protein from a lysed culture after 24-hour growth as a measure of the total cell number in the culture.
[0076] As used herein, "modified activity" or "altered level of activity" of a protein / polypeptide, such as an engineered thioesterase variant, refers to a difference in one or more characteristics of the activity of the protein / polypeptide as compared to the characteristics of a suitable control protein, such as the corresponding parental protein or the corresponding wild-type protein. Thus, in an exemplary embodiment, the difference in activity of a protein having "modified activity" as compared to the corresponding control protein is determined by measuring the activity of the modified protein in a recombinant host cell and comparing it to the same measure of the activity of the corresponding control protein in an otherwise isogenic host cell. Modified activity can be the result of, for example, a change in the structure of the protein (e.g., a change in the primary structure such as a change in the nucleotide coding sequence of the protein that results in, for example, a change in substrate specificity, a change in observed rate parameters, a change in solubility, etc.); a change in the stability of the protein (e.g., an increase or decrease in proteolysis). In some exemplary embodiments, a polypeptide having "modified activity" is a variant or engineered TE variant disclosed herein.
[0077] In an exemplary embodiment, a polypeptide disclosed herein has "modified activity," such as, for example, an "improved level of activity." As used herein, the expression "improved level of activity" refers to a polypeptide having a higher level of biochemical or biological function (e.g., DNA binding activity or enzymatic activity) as compared to the level of the biochemical and / or biological function of the corresponding control polypeptide under the same conditions. The degree of improved activity can be about 10% or more, about 20% or more, about 50% or more, about 75% or more, about 100% or more, about 200% or more, about 500% or more, about 1000% or more, or any range therein.
[0078] Accordingly, "improved activity" may refer to improved catalytic activity or improved catalytic efficiency of a polypeptide, where catalytic efficiency represents, for example, an increase in the reaction rate of a reaction catalyzed by such an enzyme of the polypeptide. Catalytic activity / catalytic efficiency can be improved by improving one or more rate parameters (measures or calculated values) of the reaction, such as Vmax (the maximum rate at which the reaction can proceed), Km (the Michaelis constant), kcat (the number of substrate molecules turnover per second per molecule of enzyme), or any ratio between these parameters such as kcat / Km (a measure of enzyme efficiency). Accordingly, the "improved catalytic activity" or "improved catalytic efficiency" of a polypeptide can be measured in several ways. For example, "improved activity" may be measured as an increase in titer (concentration: g / L, or mg / L, or g / Kg), a change in composition (the amount of a specific fatty acid species / total fatty acid derivatives (FAS) produced), an improved ratio of molecular components (e.g., C8 / C10 content or C10 / C12 content, etc.), or an increase in the FOC (fold increase relative to the control, see below) of the product produced by a recombinant cell expressing the enzyme with improved activity.
[0079] Accordingly, as used herein, the expressions "having improved activity for the production of medium-chain fatty acid derivatives" or "having improved activity for the production of medium-chain length fatty acid derivative compounds" or "having improved activity for the production of medium-chain aliphatic compounds" or "having improved ability to produce medium-chain length fatty acid derivatives" or "having improved ability to produce medium-chain fatty acid derivatives" refer to "improved activity" of a polypeptide / protein, such as "improved catalytic activity", that leads to an increase in the production of medium-chain fatty acid derivative species (fatty acids and fatty acid derivatives having an alkyl chain of 6 to 10 carbon lengths) when compared under the same conditions with an appropriate control polypeptide / protein.
[0080] In some exemplary embodiments, a polypeptide / protein having "improved activity for the production of medium-chain fatty acid derivatives" or having a similar meaning "having an improved ability to produce medium-chain fatty acid derivatives" has improved activity for the production of medium-chain fatty acid derivatives of a specific chain length. Thus, for example, as used herein, the expression "having improved activity for the production of C8 fatty acid derivatives" refers to a polypeptide / protein having "improved catalytic activity" or "improved activity" that leads to an increase in the production of 8-carbon fatty acid derivatives (measured, for example, as an increase in %C8 FAS, C8 / C10 ratio, etc.).
[0081] Similarly, in some exemplary embodiments, a polypeptide / protein having "improved activity for the production of medium-chain fatty acid derivatives" or having a similar meaning "having an improved ability to produce medium-chain fatty acid derivatives" has "improved activity for the production of C10 fatty acid derivatives". Thus, such a polypeptide / protein has "improved activity" that leads to an increase in the production of 10-carbon fatty acid derivatives (measured, for example, as an increase in %C10 FAS, C10 / C12 ratio, etc.).
[0082] As used herein, the term "fold over control" or the expression "FOC" in a similar sense refers to the ratio of a measured specific metric of cells containing an engineered thioesterase variant to the same metric measured in a suitable control cell, e.g., an isogenic host cell containing a control thioesterase that does not have the engineered mutation. Thus, generally, FOC is equal to variant metric A / control metric A (in some exemplary embodiments, the FOC of %C8 means the %C8 produced by cells containing an engineered thioesterase variant compared to the %C8 of an isogenic control containing a suitable control, e.g., a thioesterase that was not engineered to contain a particular mutation. Thus, in an exemplary embodiment, a recombinant cell containing an engineered thioesterase variant with an FOC of %C8 of 1.1 shows a 10% improvement (increase) in the percent of 8-carbon fatty acid derivatives produced by the cells containing the engineered thioesterase variant compared to the %C8 of the isogenic control containing the control thioesterase.
[0083] A "control" sample, e.g., a "control" nucleotide sequence, a "control" polypeptide sequence, a "control" cell, etc., or value represents a reference for comparison to a test sample, usually a sample that serves as a known reference. For example, in an exemplary embodiment, while the test sample contains a fatty acid derivative composition made by an engineered thioesterase variant, the control sample contains a fatty acid derivative composition made by the corresponding or designated unmodified / non-variant thioesterase (e.g., SEQ ID NO:1). One of skill in the art recognizes that controls can be designed for the evaluation of any number of parameters. Further, one of ordinary skill in the art understands which controls are valuable in a given situation and can analyze data based on comparison to control values.
[0084] As used herein, the term "recombinant" refers to a genetically modified polynucleotide, polypeptide, cell, tissue, or organism. The term "recombinant" equally applies to first-generation genetically modified polynucleotides, polypeptides, cells, tissues, or organisms, and the progeny of genetically modified polynucleotides, polypeptides, cells, tissues, or organisms that have the genetic modification.
[0085] The term "recombinant" when used with respect to a cell indicates that the cell has been modified by the introduction of a heterologous nucleic acid or protein, or by a change in a native nucleic acid or protein, or that the cell is derived from a cell so modified and the derived cell contains the modification. Thus, for example, a "recombinant cell" or "recombinant host cell" in a similar sense may be modified to express a gene not found in the native (non-recombinant) form of the cell, or to modify the abnormal expression of a native gene. For example, a native gene may be overexpressed, underexpressed or not expressed at all. In an exemplary embodiment, a "recombinant cell" or "recombinant host cell" is engineered to express a heterologous thioesterase, such as an engineered thioesterase variant having improved activity for the production of medium-chain fatty acid derivatives. Recombinant cells can be derived from microorganisms such as bacteria, viruses or fungi. In addition, recombinant cells can be derived from plant or animal cells. In an exemplary embodiment, a "recombinant host cell" or "recombinant cell" is used to produce one or more fatty acid derivatives including, but not limited to, fatty acids, fatty esters (e.g., waxes, fatty acid esters, fatty esters, fatty acid methyl esters (FAME), fatty acid ethyl esters (FAEE)), fatty alcohol acetate esters (FAce), fatty alcohols, fatty aldehydes, hydrocarbons, fatty amines, terminal olefins, internal olefins, ketones, bifunctional fatty acid derivatives (e.g., omega-hydroxy fatty acids, omega-hydroxy diols, omega-hydroxy FAME, omega-hydroxy FAEE). Thus, in some exemplary embodiments, a "recombinant host cell" is a "production host" or "production host cell" in a similar sense. In some exemplary embodiments, a recombinant cell contains one or more polynucleotides, each polynucleotide encoding a polypeptide having fatty acid biosynthetic enzyme activity, wherein the recombinant cell produces a fatty acid derivative composition when cultured under conditions effective to express the polynucleotide in the presence of a carbon source.
[0086] When used in connection with a polynucleotide, the term "recombinant" or a term of similar import such as "heterologous" indicates that the polynucleotide has been modified compared to the native or naturally occurring form of the polynucleotide, or compared to a naturally occurring variant of the polynucleotide. In an exemplary embodiment, a recombinant polynucleotide (or a copy or complement of the recombinant polynucleotide) is a polynucleotide that has been engineered by human hand so as to be different from its native form. Thus, in an exemplary embodiment, a recombinant polynucleotide is a variant of a native gene or a variant of a naturally occurring variant of a native gene, where the variant has been created by intentional human manipulation, for example, by saturation mutagenesis using mutagenic oligonucleotides, by use of UV radiation or mutagenic chemicals, etc. Such a recombinant polynucleotide may contain one or more point mutations, deletions and / or insertions compared to the native or naturally occurring form of the gene. Similarly, a polynucleotide containing a promoter operably linked to a second polynucleotide (e.g., a coding sequence) is a "recombinant" polynucleotide. Thus, a recombinant polynucleotide includes combinations of polynucleotides not found in nature. A recombinant protein (as described above) is typically a recombinant protein expressed from a recombinant polynucleotide, and recombinant cells, tissues, and organisms are those that contain recombinant sequences (polynucleotides and / or polypeptides).
[0087] As used herein, the term "microorganism" generally refers to organisms at the microscopic level. Microorganisms can be prokaryotic or eukaryotic. Exemplary prokaryotic microorganisms include, for example, bacteria, archaea, cyanobacteria, etc. An exemplary bacterium is Escherichia coli. Exemplary eukaryotic microorganisms include, for example, yeast, protozoa, algae, etc. In an exemplary embodiment, a "recombinant microorganism" is a microorganism that has been genetically modified such that it expresses or contains a heterologous nucleic acid sequence and / or a heterologous protein.
[0088] A "production host" or "production host cell" in a similar sense is a cell used to produce a product. The "production host" disclosed herein is typically modified to express or overexpress a selected gene or to have attenuated expression of a selected gene. Thus, a "production host" or "production host cell" is a "recombinant host" or "recombinant host cell" in a similar sense. Non-limiting examples of production hosts include plants, animals, humans, bacteria, yeast, cyanobacteria, algae, and / or filamentous fungal cells. An exemplary "production host" is a recombinant Escherichia coli cell.
[0089] As used herein, "acyl-ACP" represents an acyl thioester formed between the carbonyl carbon of an acyl chain and the sulfhydryl group of the phosphopantetheinyl moiety of acyl carrier protein (ACP). In some exemplary embodiments, acyl-ACP is a synthetic intermediate of a fully saturated acyl-ACP. In other exemplary embodiments, acyl-ACP is a synthetic intermediate of an unsaturated acyl-ACP. In some exemplary embodiments, the carbon chain of the acyl group of acyl-ACP has 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28 carbons. In other exemplary embodiments, the carbon chain of the acyl group of acyl-ACP is a medium-chain length chain and has 6, 7, 8, 9, 10, 11, or 12 carbons. In other exemplary embodiments, the carbon chain of the acyl group of acyl-ACP is 8 carbons in length. In still other exemplary embodiments, the carbon chain of the acyl group of acyl-ACP is 10 carbons in length. Each of these acyl-ACPs is a substrate for an enzyme, such as a thioesterase, such as an engineered thioesterase variant that converts acyl-ACP to a fatty acid derivative.
[0090] As used herein, the expression "fatty acid derivative biosynthetic pathway" refers to the biochemical pathway that produces fatty acid derivatives. Enzymes that include the "fatty acid derivative biosynthetic pathway" are thus referred to herein as "fatty acid derivative biosynthetic polypeptides" or, in a similar sense, "fatty acid derivative enzymes". As noted above, the term "fatty acid derivative" includes molecules / compounds derived from a biochemical pathway that includes a thioesterase reaction. Thus, a thioesterase enzyme (e.g., an enzyme having thioesterase activity EC 3.1.1.14) is a "fatty acid derivative biosynthetic polypeptide" or, in a similar sense, a "fatty acid derivative enzyme". In addition to thioesterase, the fatty acid derivative biosynthetic pathway may include additional enzymes for producing fatty acid derivatives having desired characteristics. Thus, the term "fatty acid derivative enzyme" or, in a similar sense, "fatty acid derivative biosynthetic polypeptide" refers, collectively and individually, to enzymes that can be expressed or overexpressed to produce fatty acid derivatives. Non-limiting examples of "fatty acid derivative enzymes" or, in a similar sense, "fatty acid derivative biosynthetic polypeptides" include, for example, fatty acid synthase, thioesterase, acyl-CoA synthase, acyl-CoA reductase, acyl-ACP reductase, alcohol dehydrogenase, alcohol O-acyltransferase, fatty alcohol-forming acyl-CoA reductase, fatty acid decarboxylase, fatty aldehyde decarbonylase and / or oxidative deformylase, carboxylic acid reductase, fatty alcohol O-acetyltransferase, ester synthase, and the like. A "fatty acid derivative enzyme" or, in a similar sense, a "fatty acid derivative biosynthetic polypeptide" converts a substrate into a fatty acid derivative. In an exemplary embodiment, a suitable substrate for a fatty acid derivative enzyme can be a first fatty acid derivative that is converted by the fatty acid derivative enzyme into a different second fatty acid derivative.
[0091] As used herein, the term "culture" refers to a liquid medium containing viable cells. In one aspect, the culture includes cells growing under controlled conditions in a pre-determined medium, e.g., a recombinant host cell culture grown in a liquid medium containing a selected carbon source and nitrogen. "Culturing" or "culture" refers to growing a population of host cells (e.g., recombinant host cells) in a liquid or solid medium under suitable conditions. In certain aspects, culturing refers to the bioconversion of a substrate to a final product. Culture media are well known, and the individual components of such culture media are available from commercial sources, e.g., Difco™ media and BBL™ media. By way of non-limiting example, a liquid nutrient medium is a "rich medium" containing a complex nitrogen source, a salt source, and a carbon source, such as YP medium containing 10 g / L peptone and 10 g / L yeast extract.
[0092] As used herein, the term "titer" refers to the amount of fatty acid derivative, e.g., medium-chain fatty acid derivative, produced per unit volume of host cell culture. The titer can represent the amount of a specific fatty acid derivative, e.g., medium-chain fatty acid derivative, or a combination of fatty acid derivatives of different chain lengths or different functionalities, such as a mixture or fatty acid derivative composition of saturated and unsaturated medium-chain fatty acid derivatives produced by a given recombinant host cell culture.
[0093] As used herein, the expression "commercial titer" for one or more refers to the amount of fatty acid derivative, e.g., medium-chain fatty acid derivative, produced per unit volume of host cell culture that makes a commercially available product economically viable. Typically, the commercial titer is in the range of about 10 g / L (or 10 g / Kg in a similar sense) to about 200 g / L or more. Thus, the commercial titer is 10 g / L or more, 20 g / L or more, 30 g / L or more, 40 g / L or more, 50 g / L or more, 60 g / L or more, 70 g / L or more, 80 g / L or more, 90 g / L or more, 100 g / L or more, 110 g / L or more, 120 g / L or more, 130 g / L or more, 140 g / L or more, 150 g / L or more, 160 g / L or more, 170 g / L or more, 180 g / L or more, 190 g / L or more, 200 g / L or more.
[0094] As used herein, the "yield of fatty acid derivative" produced by a "host cell", for example, the yield of a medium-chain fatty acid derivative or other compound, represents the efficiency with which the input carbon source is converted to the product (i.e., the medium-chain fatty acid derivative) in the host cell. Thus, the expression "yield of fatty acid derivative" represents the amount of product produced from a given amount of carbon substrate. The percent yield is the percent relative to the theoretical yield (the product synthesized under ideal conditions without loss of carbon or energy). Thus, percent yield = (mass of product / mass of theoretical yield) × 100. The yield may represent a specific medium-chain fatty acid derivative or a combination of fatty acid derivatives.
[0095] As used herein, the term "productivity" refers to the amount of medium-chain fatty acid derivatives produced per unit time per unit volume of host cell culture, such as 6-carbon fatty acid derivatives, 8-carbon fatty acid derivatives, 10-carbon fatty acid derivatives, etc. Productivity can represent a specific 8- and / or 10-carbon fatty acid derivative, or a combination of fatty acid derivatives or other compounds produced by a given host cell culture. Thus, in an exemplary embodiment, for example, expression of an engineered thioesterase variant in a recombinant host cell such as Escherichia coli results in an increase in productivity of 8- and / or 10-carbon fatty acid derivatives and / or other compounds compared to a recombinant host cell expressing the corresponding control thioesterase or other appropriate control. As used herein, the terms "total lipid species" and "total fatty acid product" and "total fatty acid derivative" may be used interchangeably herein in relation to the amount (titer) of fatty acid derivatives produced by a host cell, e.g., a host cell expressing an engineered thioesterase variant. Total lipid species, etc., can be evaluated by gas chromatography equipped with a flame ionization detector (GC-FID). When referring to total fatty acid derivative analysis, the same terms may be used, for example, to mean total fatty esters, total fatty alcohols, total fatty aldehydes, total fatty amines, and total free fatty acids. In particular, the same terms may be used to mean total fatty acid methyl esters, fatty acid ethyl esters, or fatty alcohol acetates.
[0096] As used herein, the term "carbon source" refers to a substrate or compound suitable for use as a carbon source for the growth of prokaryotic or simple eukaryotic cells. The carbon source can be in various forms, including, but not limited to, polymers, carbohydrates, acids, alcohols, aldehydes, ketones, amino acids, peptides, and gases (e.g., CO and CO2). Exemplary carbon sources include monosaccharides such as glucose, fructose, mannose, galactose, xylose, and arabinose; oligosaccharides such as fructooligosaccharides and galactooligosaccharides; polysaccharides such as starch, cellulose, pectin, and xylan; disaccharides such as sucrose, maltose, cellobiose, and turanose; cellulose materials and variants such as hemicellulose, methylcellulose, and sodium carboxymethylcellulose; saturated or unsaturated fatty acids, succinic esters, lactic esters, and acetic esters; alcohols such as ethanol, methanol, and glycerol, or mixtures thereof. The carbon source can also be a photosynthetic product such as glucose. In certain embodiments, the carbon source is biomass. In other embodiments, the carbon source is glucose. In other embodiments, the carbon source is sucrose. In other embodiments, the carbon source is glycerol. In other embodiments, the carbon source is a simple carbon source. In other embodiments, the carbon source is a renewable carbon source. In other examples, the carbon source is natural gas or a natural gas component such as methane, ethane, or propane.
[0097] As used herein, the term "biomass" refers to any biological material from which a carbon source is derived. In some embodiments, the biomass is processed into a carbon source suitable for bioconversion. In other embodiments, the biomass does not require further processing into a carbon source. The carbon source can be converted into a composition containing medium-chain fatty acid derivatives.
[0098] Exemplary sources of biomass include plant material or vegetation, such as plant material or vegetation derived from corn, sugarcane, switchgrass, rice, wheat, hardwoods, conifers, palm, hemp, etc. Another exemplary source of biomass is metabolic waste, such as animal material (e.g., cow manure). Further exemplary sources of biomass include algae and marine plants such as macroalgae and kelp. Biomass includes, without limitation, waste from industry, agriculture, forestry, and households, including glycerol, fermentation waste, silage, straw, wood, pulp, sewage, garbage, cellulose general waste, municipal solid waste, oleochemical waste, and food waste (e.g., soaps, oils, and fatty acids). The term "biomass" can also represent a carbon source such as a saccharide (e.g., a monosaccharide, disaccharide, or polysaccharide).
[0099] As used herein with respect to a product (such as a medium-chain fatty acid derivative), the term "isolated" refers to a product separated from cell components, cell culture medium, or a chemical or synthetic precursor. Medium-chain fatty acid derivatives produced by the methods disclosed herein may be relatively immiscible in the fermentation broth and cytoplasm. Thus, in an exemplary embodiment, the medium-chain fatty acid derivative accumulates extracellularly in the organic phase, thereby being "isolated".
[0100] As used herein, the terms "purify", "purified", or "purification" mean, for example, the removal or isolation of a molecule from its environment by isolation or separation. A "substantially purified" molecule has less than at least about 60% of the other components with which it is associated (e.g., less than at least about 65%, less than at least about 70%, less than at least about 75%, less than at least about 80%, less than at least about 85%, less than at least about 90%, less than at least about 95%, less than at least about 96%, less than at least about 97%, less than at least about 98%, less than at least about 99%). As used herein, these terms also represent the removal of contaminants from a sample. For example, the removal of contaminants can result in an increase in the percentage of medium-chain fatty acid derivatives or other compounds in the sample. For example, when medium-chain fatty acid derivatives or other compounds are produced in recombinant host cells, the medium-chain fatty acid derivatives or other compounds can be purified by the removal of host cell biomass or its components such as proteins, nucleic acids, and other cell components if the host cells are lysed. After purification, the percentage of malonyl-CoA-derived compounds containing medium-chain fatty acid derivatives or other compounds in the sample increases. The terms "purify", "be purified", and "purification" are relative terms that do not require absolute purity. Thus, for example, when medium-chain fatty acid derivatives are produced in recombinant host cells, the medium-chain fatty acid derivatives are substantially separated from other cell components (e.g., nucleic acids, polypeptides, lipids, carbohydrates, or other hydrocarbons).
[0101] As used herein, the term "attenuate" means to weaken, reduce, or decrease. For example, the activity of a polypeptide can be attenuated, for example, by modifying the polypeptide structure to reduce its activity (e.g., by modifying the nucleotide sequence encoding the polypeptide).
[0102] I. Introduction As described above, there has been a great deal of interest in medium-chain fatty acid (MCFA) derivatives and products derived from medium-chain fatty acids (MCFAs). The numerous favorable properties of MCFAs have been evaluated. In fact, MCFAs are used as renewable and biodegradable components, such as, for example, surfactants, adhesives, emulsifiers, edible oils, flavorings, fragrances, monomers, polymers, natural product pesticides, and antibacterial agents.
[0103] Due to their numerous uses, the demand for medium-chain fatty acid derivative compounds in industrial applications and nutraceutical applications has been on the rise and continues to increase over the past several years. Unfortunately, however, the supply of medium-chain fatty acid derivatives is highly associated with the production of other, longer-chain free fatty acid products from plants or chemical synthesis, and thus the supply is extremely variable and unstable.
[0104] Accordingly, what is needed in the art are materials and methods that can provide a robust and stable supply chain for MCFAs and their derivatives. Fortunately, the present disclosure provides the tools and methods necessary to support a robust, selective, and stable supply chain for medium-chain fatty acid derivatives, and thus addresses this and other needs.
[0105] II. Engineered Thioesterase Variants with Improved Activity for the Production of Medium-Chain Fatty Acid Derivatives A. General Methods The present disclosure utilizes conventional techniques in the field of recombinant genetics. Basic texts that disclose general methods and terms in molecular biology and genetics include, for example, Sambrook et al., Molecular Cloning, a Laboratory Manual, Cold Spring Harbor Press 4th edition (Cold Spring Harbor, N.Y. 2012); Current Protocols in Molecular Biology Volumes 1-3, John Wiley & Sons, Inc. (1994-1998) and Supplements 1-115 (1987-2016). The present disclosure also utilizes conventional techniques in the field of biochemistry. Basic texts that disclose general methods and terms in biochemistry include, for example, Lehninger Principles of Biochemistry sixth edition, David L. Nelson and Michael M. Cox eds. W.H. Freeman (2012). The present disclosure also utilizes conventional techniques in industrial fermentation. Basic texts that disclose general methods and terms in fermentation include, for example, Principles of Fermentation Technology, 3rd Edition by Peter F. Stanbury, Allan Whitaker and Stephen J. Hall. Butterworth-Heinemann (2016); Fermentation Microbiology and Biotechnology, 2nd Edition, E. M. T. El-Mansi, C. F. A. Bryce, Arnold L. Demain and A.R. Allman eds. CRC Press (2007). The present disclosure also utilizes conventional techniques in the field of organic chemistry.Basic texts that disclose general methods and terms in organic chemistry include, for example, Practical Synthetic Organic Chemistry: Reactions, Principles, and Techniques, Stephane Caron ed., John Wiley and Sons Inc. (2011); The Synthetic Organic Chemist's Companion, Michael C. Pirrung, John Wiley and Sons Inc. (2007); Organic Chemistry, 9th Edition - Francis Carey and Robert Giuliano, McGraw Hill (2013).
[0106] For nucleic acids, size is indicated in either kilobases (kb) or base pairs (bp). Estimates are typically obtained from agarose or acrylamide gel electrophoresis, sequenced nucleic acids, or published DNA sequences. For proteins, size is indicated in kilodaltons (kDa) or the number of amino acid residues. The size of a protein is estimated from gel electrophoresis, a sequenced protein, the obtained amino acid sequence, or a published protein sequence.
[0107] Oligonucleotides that are not commercially available can be chemically synthesized, for example, by the solid-phase phosphoramidite triester method first described by Beaucage & Caruthers, Tetrahedron Letts. 22:1859-1862 (1981), using an automated synthesizer described by Van Devanter et al., Nucleic Acids Res. 12:6159-6168 (1984). Purification of oligonucleotides is by either native acrylamide gel electrophoresis or anion-exchange HPLC as described by Pearson & Reanier, J. Chrom. 255:137-149 (1983).
[0108] The sequences of cloned genes and synthetic oligonucleotides can be verified, for example, after cloning using the chain termination method for sequencing double-stranded templates such as Wallace et al., Gene 16:21-26 (1981).
[0109] B. Thioesterase 1. General Thioesterase or thiolester hydrolase catalyzes the hydrolysis of thioesters to acids and thiols. Thioesterase (TE) is classified from EC 3.1.2.1 to EC 3.1.2.27 based on its activity towards different substrates, and many remain unclassified (EC 3.1.2.-) (see, for example, Cantu, D.C., et al. (2010) Protein Science 19:1281-1295). TE can be obtained from a variety of sources. Exemplary TEs include plant TEs (see, for example, Voelker and Davies, J. Bact., Vol., 176, No. 23, pp. 7320-27, 1994, U.S. Patent Nos. 5,667,997 and 5,455,167), bacterial TEs (see, for example, U.S. Patent No. 9,175,234); cyanobacterial TEs, as well as those of algal, mammalian, insect, and fungal origin.
[0110] In particular, acyl-acyl carrier protein (ACP) thioesterase (TE) classified as EC number 3.1.2.14 selectively hydrolyzes the thioester bond of acyl-ACP, releasing free fatty acid (FFA) and ACP. Thus, acyl-ACP thioesterase plays an important role in determining the carbon chain length of the fatty acid derivatives resulting from the hydrolysis products of alkyl thioesters.
[0111] The FatB2 thioesterase (ChFatB2) derived from Cuphea hookeriana is an exemplary acyl-ACP thioesterase. ChFatB2 inherently has high selectivity for medium-chain length fatty acid derivatives. However, this plant enzyme has low activity when expressed in microorganisms such as Escherichia coli, an industrial microorganism. As disclosed in detail herein, the low ability of microorganisms to produce medium-chain length fatty acids is the result of its low activity, insufficient selectivity for C8 and C10, and poor solubility.
[0112] The polypeptide / protein sequence of wild-type ChFatB2 derived from Cuphea hookeriana has the GenBank accession number AAC49269 (see, for example, Dehesh, K., et al. (1996) The Plant Journal 9(2):167-72). The amino acid sequence of the ChFatB2 polypeptide disclosed herein contains the wild-type sequence in which the first 88 amino acids containing the plant translocation leader sequence at the N-terminus of the wild-type protein are removed and replaced with methionine (M) to facilitate the production of the active enzyme in bacterial cytoplasm. Thus, the amino acid sequence of the wild-type ChFatB2 thioesterase disclosed herein as wild-type (wt) ChFatB2 is shown below as SEQ ID NO:1. TIFF0007705243000002.tif38149
[0113] The activity of SEQ ID NO:1 is known to be specific for saturated 8-carbon (8:0) and saturated 10-carbon (10:0) ACP substrates (see, for example, Dehesh, K., et al. (1996), supra). Unfortunately, however, the medium-chain length fatty acid derivative production capacity possessed by SEQ ID NO:1 is insufficient for the large-scale production of medium-chain fatty acid derivatives. Thus, in an exemplary embodiment, SEQ ID NO:1 is modified to produce engineered TE variants with improved activity for the production of medium-chain fatty acid derivatives to address the need for a stable and reliable supply of medium-chain fatty acid derivatives.
[0114] Thus, in an exemplary aspect, the present disclosure provides an engineered TE variant polypeptide having improved activity for the production of medium-chain fatty acid derivatives, such as medium-chain fatty acid methyl esters (FAME) and medium-chain fatty acid ethyl esters (FAEE), medium-chain fatty acid ethyl esters (FACE), medium-chain fatty amines, medium-chain fatty aldehydes, medium-chain fatty alcohols, medium-chain hydrocarbons, medium-chain fatty ketones, medium-chain alkanes, medium-chain terminal olefins, medium-chain internal olefins, medium-chain hydroxy fatty acid derivatives, medium-chain difunctional fatty acid derivatives, such as medium-chain fatty diacids, medium-chain fatty diols, unsaturated medium-chain fatty acid derivatives, etc., compared to the enzyme having SEQ ID NO:1.
[0115] In some exemplary aspects, engineered TE variants of SEQ ID NO:1 having improved activity for the production of medium-chain fatty acid derivatives (e.g., SEQ ID NOs: 16 - 46) have an increased effective surface positive charge compared to a non-variant / non-engineered control thioesterase, e.g., SEQ ID NO:1.
[0116] 2. Assays for Engineered Thioesterase Variants Having Improved Activity for the Production of Medium-Chain Fatty Acid Derivatives In an exemplary aspect, an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives is identified by measuring medium-chain fatty acid derivatives (e.g., free fatty acids (FFA), fatty acid ethyl esters, FAEE, fatty alcohols (FALC), fatty alcohol acetate esters (FACe), etc.) produced by a bacterial strain containing the engineered TE variant (i.e., the test strain), and comparing these medium-chain fatty acid derivatives to the measured values of medium-chain fatty acid derivatives (e.g., FFA, FAEE, FALC, FACE, etc.) produced by a suitable control test strain that is isogenic to the test strain except for containing the control TE.
[0117] In some exemplary embodiments, the total titer of the medium-chain fatty acid derivatives is measured and compared between the test strain and the control strain. In some exemplary embodiments, the percentage of the total titer of the medium-chain fatty acid derivatives comprising a specific medium-chain fatty acid derivative (e.g., C8 fatty acid derivative) produced by the test strain is measured and compared to the percentage of the total titer of the medium-chain fatty acid derivatives comprising a specific medium-chain fatty acid derivative produced by a suitable control strain that is isogenic to the test strain except for containing the control TE (e.g., SEQ ID NO:1).
[0118] In an exemplary embodiment, gas chromatography with a flame ionization detector (GC-FID) is used to assay the medium-chain fatty acid derivatives. GC-FID is known in the art (see, for example, Adlard, E. R.; Handley, Alan J. (2001). Gas chromatographic techniques and applications. London: Sheffield Academic). However, any method suitable for quantification and analysis, such as mass spectrometry (MS), gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), thin layer chromatography (TLC), etc., can be used.
[0119] C. Method for Producing an Engineered Thioesterase Variant The engineered TE variants can be prepared by any method known in the art (see, e.g., Current Protocols in Molecular Biology, supra). Thus, in an exemplary embodiment, mutagenesis is used to prepare a polynucleotide sequence encoding an engineered TE variant, which can then be screened for improved activity for the production of medium-chain fatty acid derivatives. In another exemplary embodiment, a polynucleotide sequence encoding an engineered TE variant that can then be screened for improved activity for the production of medium-chain fatty acid derivatives is prepared by chemical synthesis of the polynucleotide sequence (see, e.g., M.H. Caruthers et al. (1987) Methods in Enzymology Volume 154, Pages 287-313; Beaucage, S.L. and Iyer, R.P. (1992) Tetrahedron 48(12):2223-2311).
[0120] Mutagenesis methods are well known in the art. Exemplary mutagenesis techniques for preparing engineered TE variants having improved activity for the production of medium-chain fatty acid derivatives include, for example, site-saturation mutagenesis (see, e.g., Chronopoulou EG1, Labrou NE. Curr. Protoc. Protein Sci. 2011 Feb; Chapter 26:Unit 26.6, John Wiley and Sons, Inc; Steffens, D.L. and Williams., J.G.K (2007) J Biomol Tech. 18(3): 147-149; Siloto, R.M.P and Weselake, R.J. (2012) Biocatalysis and Agricultural Biotechnology 1(3):181-189).
[0121] Another exemplary mutagenesis technique for preparing engineered TE variants with improved activity for the production of medium-chain fatty acid derivatives involves transfer PCR (tPCR). See, for example, Erijman A., et al. (2011) J. Struct. Biol. 175(2):171-7.
[0122] Other exemplary mutagenesis techniques include, for example, error-prone polymerase chain reaction (PCR) (see, for example, Leung et al. (1989) Technique 1:11-15; and Caldwell et al. (1992) PCR Methods Applic. 2:28-33).
[0123] Another exemplary mutagenesis technique for preparing engineered TE variants with improved activity for the production of medium-chain fatty acid derivatives involves using oligonucleotide-directed mutagenesis to create site-specific mutations in any cloned DNA of interest (see, for example, Reidhaar-Olson et al. (1988) Science 241:53-57).
[0124] The mutagenized polynucleotide resulting from any method of synthesis or mutagenesis, such as those described above, is then cloned into an appropriate vector, and the activity of the polypeptide of interest encoded by the mutagenized polynucleotide is evaluated as disclosed above.
[0125] One of ordinary skill in the art will recognize that the protocols and procedures disclosed herein can be modified and that such modifications are within the scope of the present disclosure. For example, where method steps are recited in a particular order, the order of these steps can be modified and / or performed in parallel or sequentially.
[0126] III. Host Cells and Host Cell Cultures In view of the present disclosure, one of ordinary skill in the art would recognize that any of the aspects contemplated herein can be practiced using any host cell or microorganism that can be genetically modified via the introduction of one or more nucleic acid sequences encoding the disclosed engineered TE variants. Thus, the recombinant microorganisms disclosed herein comprise one or more polynucleotide sequences that function as a host cell and contain an open reading frame encoding an engineered TE variant polypeptide having improved activity for the production of medium-chain fatty acid derivatives, together with regulatory sequences operably linked to facilitate the expression of the engineered TE variant polypeptide in the host cell.
[0127] Exemplary microorganisms that provide suitable host cells include, without limitation, cells of the genus Escherichia, Bacillus, Lactobacillus, Zymomonas, Rhodococcus, Pseudomonas, Aspergillus, Trichoderma, Neurospora, Fusarium, Humicola, Rhizomucor, Kluyveromyces, Pichia, Mucor, Myceliophtora, Marinobacter, Penicillium, Phanerochaete, Pleurotus, Trametes, Chrysosporium, Saccharomyces, Stenotrophomonas, Schizosaccharomyces, Yarrowia, or Streptomyces. In some exemplary embodiments, the host cell is a Gram-positive bacterial cell. In other exemplary embodiments, the host cell is a Gram-negative bacterial cell. In some embodiments, the host cell is an E. coli cell.In other exemplary embodiments, the host cell is a Bacillus lentus cell, a Bacillus brevis cell, a Bacillus stearothermophilus cell, a Bacillus lichenoformis cell, a Bacillus alkalophilus cell, a Bacillus coagulans cell, a Bacillus circulans cell, a Bacillus pumilis cell, a Bacillus thuringiensis cell, a Bacillus clausii cell, a Bacillus megaterium cell, a Bacillus subtilis cell, or a Bacillus amyloliquefaciens cell.
[0128] In still other exemplary embodiments, the host cell is a Trichoderma koningii cell, a Trichoderma viride cell, a Trichoderma reesei cell, a Trichoderma longibrachiatum cell, an Aspergillus awamori cell, an Aspergillus fumigates cell, an Aspergillus foetidus cell, an Aspergillus nidulans cell, an Aspergillus niger cell, an Aspergillus oryzae cell, a Humicola insolens cell, a Humicola lanuginose cell, a Rhodococcus opacus cell, a Rhizomucor miehei cell, or a Mucor michei cell. In still other exemplary embodiments, the host cell is a Streptomyces lividans cell or a Streptomyces murinus cell. In yet other embodiments, the host cell is an Actinomycetes cell. In some exemplary embodiments, the host cell is a Saccharomyces cerevisiae cell.
[0129] In still other exemplary embodiments, the host cell is a cell of a eukaryotic plant, alga, cyanobacterium, green sulfur bacterium, green non-sulfur bacterium, red sulfur bacterium, red non-sulfur bacterium, extremophilic microorganism, yeast, fungus, engineered organism, or synthetic organism. In some exemplary embodiments, the host cell is from Arabidopsis thaliana, Panicum virgatums, Miscanthus giganteus, Zea mays, botryococcuse braunii, Chalamydomonas reinhardtii, Dunaliela salina, Thermosynechococcus elongatus, Synechococcus elongatus, Synechococcus sp., Synechocystis sp., Chlorobium tepidum, Chloroflexus auranticus, Chromatiumm vinosum, Rhodospirillum rubrum, Rhodobacter capsulatus, Rhodopseudomonas palusris, Clostridium ljungdahlii, Clostridiuthermocellum, or Pencillium chrysogenum.In some other exemplary embodiments, the host cell is a cell derived from Pichia pastories, Saccharomyces cerevisiae, Yarrowia lipolytica, Schizosaccharomyces pombe, Pseudomonas fluorescens, Pseudomonas putida or Zymomonas mobilis. In yet further exemplary embodiments, the host cell is a cell of Synechococcus species PCC 7002, Synechococcus species PCC 7942, or Synechocystis species PCC6803. In some exemplary embodiments, the host cell is a CHO cell, a COS cell, a VERO cell, a BHK cell, a HeLa cell, a Cv1 cell, an MDCK cell, a 293 cell, a 3T3 cell, or a PC12 cell. In some exemplary embodiments, the host cell is an E. coli cell. In some exemplary embodiments, the E. coli cell is an E. coli cell of strain B, strain C, strain K, or strain W.
[0130] In some exemplary embodiments, the host cell includes any genetic manipulation and genetic modification that can be used interchangeably for each host cell depending on what other heterologous enzymes and what native enzyme pathways are present in the host cell. In one exemplary embodiment, the host cell optionally includes a fadE and / or fhuA deletion. In other exemplary embodiments, the host cell is optionally engineered to have the ability to produce more than 200 mg / L of a fatty acid derivative, more than 1000 mg / L of a fatty acid derivative, more than 1200 mg / L of a fatty acid derivative, more than 1700 mg / L of a fatty acid derivative, more than 2000 mg / L of a fatty acid derivative, or more than 3000 mg / L of a fatty acid derivative. The optionally engineered strains described above are useful for the identification and characterization of useful engineered TE variants having an improved ability to produce medium-chain fatty acid derivatives, and for the selective production of medium-chain fatty acid derivatives when expressing an engineered TE variant having an improved ability to produce medium-chain fatty acid derivatives.
[0131] As discussed in detail below in this specification, in some exemplary embodiments, the host cell or host microorganism used to express the engineered TE variant polypeptide further expresses a gene having an enzyme activity that can increase the production of one or more specific fatty acid derivatives such as, for example, fatty acid esters, fatty alcohols, fatty alcohol acetate esters, fatty acid methyl esters, fatty acid ethyl esters, fatty amines, fatty aldehydes, bifunctional fatty acid derivatives, diacids, alkanes, alkenes or olefins, ketones, etc.
[0132] For example, the entD gene encodes phosphopantetheinyl transferase. Overexpression of native E. coli entD, phosphopantetheinyl transferase, is any genetic modification to cells expressing a carboxylic acid reductase such as CarB. It enables improved activation of CarB from apo-CarB to holo-CarB, thereby allowing improved conversion of free fatty acids to fatty aldehydes by holo-CarB, which can then be converted to fatty alcohols by fatty aldehyde reductase. See, for example, U.S. Patent No. 9,340,801.
[0133] In an exemplary embodiment, the host cell or host microorganism used to express the engineered TE variant polypeptide further expresses ester synthase activity (E.C. 2.3.1.75) for the production of fatty esters. In another exemplary embodiment, the host cell has acyl-ACP reductase (AAR) (E.C. 1.2.1.80) activity and / or alcohol dehydrogenase activity (E.C. 1.1.1.1.) and / or fatty alcohol acyl-CoA reductase (FAR) (E.C. 1.1.1.*) activity and / or carboxylic acid reductase (CAR) (EC 1.2.99.6) activity for the production of fatty alcohols. In another exemplary embodiment, the host cell has acyl-ACP reductase (AAR) (E.C. 1.2.1.80) activity for the production of fatty aldehydes. In another exemplary embodiment, the host cell has acyl-ACP reductase (AAR) (E.C. 1.2.1.80) activity and decarbonylase or fatty aldehyde oxidative deformylation activity for the production of alkanes and alkenes. In another exemplary embodiment, the host cell has acyl-CoA reductase (E.C. 1.2.1.50) activity and acyl-CoA synthetase (FadD) (E.C. 2.3.1.86) activity for the production of fatty alcohols. In another exemplary embodiment, the host cell has ester synthase activity (E.C. 2.3.1.75) and acyl-CoA synthetase (FadD) (E.C. 2.3.1.86) activity for the production of fatty esters. In another exemplary embodiment, the host cell has OleA activity for the production of ketones. In another exemplary embodiment, the host cell has OleBCD activity for the production of internal olefins. In another exemplary embodiment, the host cell has acyl-ACP reductase (AAR) (E.C. 1.2.1.80) activity and alcohol dehydrogenase activity (E.C. 1.1.1.1.) for the production of fatty alcohols. In another exemplary embodiment, the host cell has decarboxylase activity for the production of terminal olefins.The expression of enzyme activity in microorganisms and microbial cells is taught, for example, by the following U.S. Pat. Nos. 9,133,406; 9,340,801; 9,200,299; 9,068,201; 8,999,686; 8,658,404; 8,597,922; 8,535,916; 8,530,221; 8,372,610; 8,323,924; 8,313,934; 8,283,143; 8,268,599; 8,183,028; 8,110,670; 8,110,093; and 8,097,439.
[0134] In some exemplary embodiments, the host cell or microorganism used to express the engineered TE variant polypeptide contains certain native enzyme activities that are upregulated or overexpressed, for example, to produce one or more specific fatty acid derivatives such as fatty esters, fatty acid methyl esters, fatty acid ethyl esters, fatty alcohols, fatty alcohol acetate esters, fatty amines, fatty amides, fatty aldehydes, bifunctional fatty acid derivatives, diacids, and the like.
[0135] In some exemplary embodiments, the recombinant host cell produces medium-chain fatty esters such as medium-chain fatty acid methyl esters (FAME) or medium-chain fatty acid ethyl esters (FAEE), medium-chain fatty alcohol acetate esters (FACE), medium-chain fatty alcohols (FALC), medium-chain fatty amines, medium-chain fatty aldehydes, medium-chain bifunctional fatty acid derivatives, medium-chain diacids, medium-chain alkanes, medium-chain olefins, and the like.
[0136] Medium-chain fatty acid derivatives are typically recovered from the culture medium and / or isolated from the host cell. In one exemplary embodiment, the fatty acid derivative is recovered from the culture medium (extracellular). In another exemplary embodiment, the fatty acid derivative is isolated from the host cell (intracellular). In another exemplary embodiment, the fatty acid derivative or non-fatty acid compound is recovered from the culture medium and isolated from the host cell.
[0137] To determine the distribution of specific fatty acid derivatives and the chain length and degree of saturation of the components of the fatty acid derivative composition, the fatty acid derivative composition produced by a host cell can be analyzed using methods known in the art, such as gas chromatography with a flame ionization detector (GC FID). Similarly, other compounds can be analyzed by methods well known in the art.
[0138] IV. Methods for Producing Recombinant Host Cells and Cultures Any method known in the art can be used to manipulate host cells to produce fatty acid derivatives and / or fatty acid derivative compositions or other compounds. Exemplary methods include the use of vectors, such as expression vectors, that contain polynucleotide sequences encoding mutant or engineered TE variants and / or other fatty acid derivative biosynthetic pathway polypeptides disclosed herein. Those skilled in the art will recognize that a variety of viral and non-viral vectors can be used in the methods disclosed herein.
[0139] In some exemplary embodiments, the polynucleotide (or gene) sequence encoding a mutant or engineered TE variant is provided to a host cell by a recombinant vector that contains a promoter operably linked to the polynucleotide sequence encoding the mutant or engineered TE variant. In some exemplary embodiments, the promoter is a developmental regulatory, organelle-specific, tissue-specific, inducible, constitutive, or cell-specific promoter. In some exemplary embodiments, the promoter is inducible by the addition of lactose or isopropylthiogalactoside (IPTG).
[0140] After a polynucleotide sequence encoding a variant or engineered TE variant and / or a polynucleotide sequence encoding another fatty acid derivative biosynthetic pathway polypeptide has been prepared and isolated, various methods can be used to construct expression cassettes, vectors, and other DNA constructs. Expression cassettes containing a polynucleotide sequence encoding a variant or engineered TE variant and / or a polynucleotide sequence encoding another fatty acid biosynthetic pathway polypeptide can be constructed in a variety of ways. Experts are well aware of the genetic elements that must be present on an expression construct / vector in order to successfully transform, select, and propagate in a host cell. Techniques for the manipulation of polynucleotide sequences, such as those encoding a variant or engineered TE variant, e.g., subcloning a nucleic acid sequence into an expression vector, labeled probes, DNA hybridization, etc., are generally described, for example, in Sambrook, et al., supra; Current Protocols in Molecular Biology, supra.
[0141] DNA constructs comprising a polynucleotide sequence encoding a variant or engineered TE variant (e.g., SEQ ID NO:3, SEQ ID NO;16 - SEQ NO:46, etc.) and / or a polynucleotide sequence encoding another fatty acid biosynthetic pathway polypeptide, linked to a heterologous DNA sequence, such as a promoter sequence, can be inserted into a variety of vectors. In some exemplary embodiments, the selected vector is an expression vector useful for the transformation of bacteria, such as E. coli. Expression vectors can be plasmids, viruses, cosmids, artificial chromosomes, nucleic acid fragments, etc. Such vectors are readily constructed by use of recombinant DNA techniques well known to those of skill in the art (see, e.g., Sambrook et al., supra). Expression vectors containing a polynucleotide sequence encoding a variant or engineered TE variant may then be transfected / transformed into a target host cell. Successfully transformed cells are then selected based on the presence of an appropriate marker gene by methods well known in the art.
[0142] Several recombinant vectors are available to those skilled in the art for use in the stable transformation / transfection of bacteria and other microorganisms (see, e.g., Sambrook, et al., supra). Suitable vectors are readily selected by those skilled in the art. In an exemplary embodiment, known vectors are used to generate an expression construct that contains a polynucleotide sequence encoding a variant or engineered TE variant.
[0143] Typically, a transformation vector contains one or more polynucleotide sequences encoding one or more variants or engineered TE variants and / or polynucleotide sequences encoding other fatty acid derivative biosynthetic pathway polypeptides that are operably linked to, for example, a promoter sequence and a selectable marker. Such transformation vectors also typically contain a transcription initiation site, a ribosome binding site, an RNA processing signal, a transcription termination site, and / or a polyadenylation signal as needed.
[0144] Thus, in addition to the polynucleotide sequence encoding a variant or engineered TE variant and / or the polynucleotide sequence encoding other fatty acid derivative biosynthetic pathway polypeptides, the expression constructs prepared as disclosed herein may contain additional elements. In an exemplary embodiment, an expression construct containing a polynucleotide sequence encoding a variant or engineered TE variant and / or a polynucleotide sequence encoding other fatty acid derivative biosynthetic pathway polypeptides also contains enhancer sequences so as to enhance the expression of heterologous proteins. As is known in the art, enhancers are typically found 5' to the transcription initiation site and they can often be inserted either 5' or 3' to the coding sequence, either in the forward or reverse orientation.
[0145] As described above, the transformation / expression vector typically contains a selectable marker gene and / or a screenable marker gene to enable easy identification of transformants. Exemplary selectable marker genes include, but are not limited to, genes encoding antibiotic resistance (e.g., resistance to kanamycin, ampicillin, etc.). Exemplary screenable markers include a six - amino - acid histidine tag introduced at the C - terminus of the recombinant protein.
[0146] In an exemplary embodiment, a selectable marker gene or a screenable marker gene is employed as or in addition to the specific gene of interest to provide or enhance the ability to identify transformants. Numerous selectable marker genes are known in the art (e.g., see Sambrook et al, supra).
[0147] In some exemplary embodiments, the expression vector further contains sequences linked to the coding sequence of the heterologous nucleic acid to be expressed, which are removed post - translationally from the initial translation product. In an exemplary embodiment, the sequences removed post - translationally facilitate the transport of the protein into or through the intracellular or extracellular membrane, thereby promoting the transport of the protein into intracellular and / or extracellular compartments. In an exemplary embodiment, the sequences removed post - translationally protect the nascent protein from intracellular proteolysis. In an exemplary embodiment, a nucleic acid segment encoding a leader peptide sequence is used for recombinant expression of the coding sequence in a host cell, together with the selected coding sequence, upstream and in its reading frame.
[0148] In another exemplary embodiment, the expression construct contains a bacterial origin of replication, such as the ColE1 origin. In yet another exemplary embodiment, the expression construct / vector contains a bacterial selectable marker, such as an ampicillin, tetracycline, hygromycin, neomycin, or chloramphenicol resistance gene.
[0149] As is well known in the art, expression constructs typically contain restriction endonuclease sites to facilitate vector construction. Exemplary restriction endonuclease recognition sites include, but are not limited to, recognition sites for restriction endonucleases NotI, AatII, SacII, PmeI, HindIII, PstI, EcoRI, and BamHI.
[0150] A DNA construct, a polynucleotide sequence that functionally encodes a variant or engineered TE variant and / or a heterologous DNA sequence, e.g., a polynucleotide sequence encoding another fatty acid derivative biosynthetic pathway polypeptide linked to a promoter sequence, a marker sequence; a purified moiety; a secretion sequence functionally coupled to the polynucleotide sequence; a targeting sequence, etc. are used to transform cells to produce recombinant host cells having improved activity for the production of medium-chain fatty acid derivatives. Exemplary host cells for transformation using an expression construct containing a polynucleotide sequence encoding a variant or engineered TE variant are discussed in detail in Section III above.
[0151] Appropriate transformation techniques are readily selected by the expert. Exemplary transformation / transfection methods available to those skilled in the art include, for example, electroporation, calcium chloride transformation, etc., and such methods are well known to the expert (see, e.g., Sambrook, supra). Thus, a polynucleotide sequence comprising an open reading frame encoding a protein and functionally linked regulatory sequences can be integrated into the chromosome of a recombinant host cell, incorporated into one or more plasmid expression systems endogenous to the recombinant host cell, or both.
[0152] The expression vectors disclosed herein typically contain a polynucleotide sequence encoding a variant or engineered TE variant and / or a polynucleotide sequence encoding a fatty acid derivative biosynthetic pathway polypeptide in a form suitable for the expression of the polynucleotide sequence in a host cell. As will be appreciated by those skilled in the art, the design of the expression vector can depend on factors such as, for example, the choice of host cell to be transformed, the level of expression of the desired polypeptide, and the like.
[0153] V. Evaluation of Recombinant Host Cells In an exemplary embodiment, the activity of the engineered TE variant polypeptide is determined by culturing the recombinant host cell and measuring the characteristics of, for example, a fatty acid derivative composition (e.g., medium-chain fatty esters, medium-chain fatty alcohols, medium-chain fatty aldehydes, etc.) or other compounds produced by the recombinant host cell. In an exemplary embodiment, the composition, titer, yield, and / or productivity of the fatty acid derivative or other compound are analyzed.
[0154] The engineered TE variant polypeptide and fragments thereof can be assayed using conventional methods for having improved activity for the production of medium-chain fatty acid derivatives (see, for example, Example 4 below).
[0155] IV. Products Obtained from Recombinant Host Cells Strategies for increasing the production of medium-chain fatty acid derivatives by recombinant host cells include, for example, increasing the flux through the fatty acid biosynthetic pathway by overexpression of native fatty acid biosynthetic genes in the production host and / or expression of heterologous fatty acid biosynthetic genes from the same or different organisms.
[0156] Accordingly, in an exemplary embodiment, a recombinant host cell having improved activity for the production of medium-chain fatty acid derivatives is engineered to contain, in addition to the engineered TE variant, one or more polynucleotide sequences encoding one or more "fatty acid derivative biosynthetic polypeptides" or "fatty acid derivative enzymes" in a similar sense. Metabolic engineering of fatty acid derivative biosynthetic pathways to produce fatty acid derivative compounds (e.g., fatty acid esters, alkanes, olefins, fatty ketones, fatty alcohols, fatty alcohol acetate esters, etc.) using microorganisms to convert biomass-derived sugars into the desired products is known in the art. See, for example, U.S. Patent Nos. 9,133,406; 9,340,801; 9,200,299; 9,068,201; 8,999,686; 8,658,404; 8,597,922; 8,535,916; 8,530,221; 8,372,610; 8,323,924; 8,313,934; 8,283,143; 8,268,599; 8,183,028; 8,110,670; 8,110,093; and 8,097,439. The metabolically engineered strains can be cultured in industrial-scale bioreactors, and the resulting products can be purified using conventional chemical and biochemical engineering techniques.
[0157] As is well known in the art, thioesterase catalyzes the hydrolysis of alkyl thioesters to free fatty acids (FFAs). Thus, thioesterase plays a role in determining the acyl chain length distribution of fatty acids and fatty acid derivatives (see, for example, Dehesh (1996) supra, PNAS (1995) 92(23): 10639-10643). Thus, recombinant host cells having improved activity for the production of medium-chain fatty acid derivatives typically contain engineered TE variants having improved activity for the production of medium-chain fatty acid derivatives (e.g., SEQ ID NO:3, and the variants shown in Table 7). In some exemplary embodiments, such recombinant host cells provide increased amounts of medium-chain fatty acid derivatives, such as medium-chain fatty alcohols, medium-chain fatty acids, FAEEs, FAMEs, FACEs, etc., as compared to appropriate control host cells that do not contain the engineered TE variant, e.g., isogenic control host cells having a control thioesterase (e.g., SEQ ID NO:1) instead of the engineered TE variant.
[0158] Thus, in some embodiments, a fatty acid derivative composition comprising a fatty acid is produced by culturing a recombinant host cell comprising an engineered TE variant under conditions effective to express thioesterase in the presence of a carbon source.
[0159] In some embodiments, substantially all of the fatty acid derivatives produced by culturing a recombinant host cell comprising an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives under conditions effective to express TE are produced extracellularly. Thus, in some exemplary embodiments, the fatty acid derivatives produced are recovered from the culture medium. In some exemplary embodiments, the recovered fatty acid derivative composition is analyzed using any suitable method known in the art, such as GC FID, to determine and quantify the distribution of specific fatty acid derivatives and the chain length and degree of saturation of the components of the fatty acid derivative composition.
[0160] In other embodiments, the recombinant host cell comprises a polynucleotide sequence encoding a variant or engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, and one or more additional polynucleotides encoding polypeptides having other fatty acid derivative biosynthetic enzyme activities. Thus, in some embodiments, a first medium-chain fatty acid derivative (e.g., medium-chain fatty acid, medium-chain fatty alcohol, etc.) produced by the action of an engineered TE variant is converted by one or more fatty acid derivative biosynthetic enzymes into a second fatty acid derivative, e.g., medium-chain fatty acid ester, medium-chain fatty aldehyde, medium-chain fatty alcohol acetate ester, hydrocarbon, e.g., linear alkane, linear alkene, etc.
[0161] Table 1 provides a list of exemplary fatty acid derivative biosynthetic polypeptides that can be expressed in a recombinant host cell and facilitate the production of medium-chain fatty acid derivatives, in addition to an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives.
[0162] (Table 1) Gene names of fatty acid derivative enzymes TIFF0007705243000003.tif127152TIFF0007705243000004.tif251152TIFF0007705243000005.tif201152TIFF0007705243000006.tif88152TIFF0007705243000007.tif219152
[0163] Production of medium-chain fatty acid derivatives As described above, a recombinant host cell comprising an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives produces an increased amount of medium-chain fatty acids as compared to a suitable control host cell that does not contain the engineered TE variant, e.g., an isogenic control host cell having a control TE (such as SEQ ID NO:1).
[0164] In other exemplary embodiments described in detail below, in addition to engineered TE variants having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises additional fatty acid derivative biosynthetic polypeptides that facilitate the production of certain types of fatty acid derivatives.
[0165] Production of fatty aldehydes In some exemplary embodiments, in addition to engineered TE variants having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises carboxylic acid reductase (''CAR'') activity, whereby the recombinant host cell synthesizes fatty aldehydes and fatty alcohols. See, e.g., 9,340,801.
[0166] Thus, in some exemplary embodiments, fatty aldehydes are produced by expressing or overexpressing in a recombinant host cell a polynucleotide encoding a polypeptide having fatty aldehyde biosynthetic activity, e.g., carboxylic acid reductase (CAR) activity. Exemplary carboxylic acid reductase (CAR) polypeptides and polynucleotides encoding the same include, e.g., FadD9 (EC 6.2.1.-, UniProtKB Q50631, GenBank NP_217106), CarA (GenBank ABK75684), CarB (GenBank YP889972) and related polypeptides disclosed, e.g., in U.S. Patent Nos. 8,097,439 and 9,340,801.
[0167] In some exemplary embodiments, the fatty aldehydes produced by the recombinant host cell are then converted to fatty alcohols or hydrocarbons. Thus, in some exemplary embodiments, in addition to engineered TE variants having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises acyl-CoA reductase (''FAR'' or ''ACR'') activity, whereby the recombinant host cell synthesizes fatty aldehydes and fatty alcohols (see, e.g., U.S. Patent Nos. 8,658,404, 8,268,599, U.S. Patent Application Publication No. 2015 / 0361454).
[0168] In some embodiments, the fatty aldehyde produced by the recombinant host cell is converted to a fatty alcohol via the activity of a native or heterologous fatty alcohol biosynthetic polypeptide, such as an aldehyde reductase or alcohol dehydrogenase (see, for example, U.S. Patent Application Publication No. 2011 / 0250663). Thus, in some exemplary embodiments, in addition to the engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises aldehyde reductase activity or, equivalently, alcohol dehydrogenase activity (EC 1.1.1.1), such that the recombinant host cell synthesizes a fatty alcohol. Exemplary fatty alcohol biosynthetic genes include, but are not limited to, alcohol dehydrogenases such as AlrA or an AlrA homolog of Acenitobacter sp. M-1; and endogenous Escherichia coli alcohol dehydrogenases such as DkgA (NP_417485), DkgB (NP_414743), YjgB, (AAC77226), YdjL (AAC74846), YdjJ (NP_416288), AdhP (NP_415995), YhdH (NP_417719), YahK (NP_414859), YphC (AAC75598), and YqhD (Q46856).
[0169] Production of fatty amines In some exemplary embodiments, a recombinant host cell (e.g., as disclosed hereinabove) that comprises an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives and produces a fatty aldehyde is further modified to include a heterologous biosynthetic enzyme having aminotransferase or amine dehydrogenase activity that converts the fatty aldehyde to a fatty amine (see, for example, PCT Publication No. WO2015 / 085271).
[0170] Production of fatty alcohols In some exemplary embodiments, in addition to the engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises a polynucleotide encoding a polypeptide having fatty alcohol biosynthetic activity, whereby a fatty alcohol is produced by the recombinant host cell. Thus, in an exemplary embodiment, a composition comprising a medium-chain fatty alcohol, e.g., octanol, is produced by culturing the recombinant host cell under conditions effective to express the engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives and a fatty alcohol biosynthetic enzyme in the presence of a carbon source.
[0171] Thus, in some exemplary embodiments, in addition to the engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises carboxylic acid reductase (CAR) activity and alcohol dehydrogenase activity, whereby the recombinant host cell synthesizes a medium-chain fatty alcohol, e.g., octanol (see, e.g., U.S. Patent No. 9,340,801).
[0172] In some exemplary embodiments, native fatty aldehyde biosynthetic polypeptides such as aldehyde reductase / alcohol dehydrogenase present in the host cell convert medium-chain fatty aldehydes to medium-chain fatty alcohols. In other exemplary embodiments, the native fatty aldehyde reductase / alcohol dehydrogenase is overexpressed to convert medium-chain fatty aldehydes to medium-chain fatty alcohols. In other exemplary embodiments, a heterologous aldehyde reductase / alcohol dehydrogenase is introduced into the recombinant host cell and expressed or overexpressed to convert medium-chain fatty aldehydes to medium-chain fatty alcohols. Exemplary aldehyde reductase / alcohol dehydrogenase polypeptides useful for converting medium-chain fatty aldehydes to medium-chain fatty alcohols are disclosed above in this specification and in International Publication Nos. 2007 / 136762; 2010 / 062480; U.S. Patent Nos. 8,110,670; 9,068,201.
[0173] In some exemplary embodiments, in addition to the engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises a heterologous polynucleotide encoding a polypeptide having carboxylic acid reductase (EC 6.2.1.3 or EC 1.2.1.42) activity, whereby the recombinant host cell produces 1,3-fatty diols when grown in a fermentation broth having a simple carbon source. In other exemplary embodiments, in addition to the engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises a heterologous polynucleotide encoding a polypeptide having carboxylic acid reductase (EC 6.2.1.3 or EC 1.2.1.42) activity and a heterologous polynucleotide encoding a polypeptide having alcohol dehydrogenase (EC 1.1.1.) activity, wherein the recombinant host cell produces 1,3-fatty diols, such as medium-chain 1,3-fatty diols, when grown in a fermentation broth having a simple carbon source (see, for example, International Publication No. WO 2016 / 011430).
[0174] Production of fatty alcohol acetate esters In some embodiments, the fatty alcohol produced in the cell or supplied to the cell in some embodiments is further processed by the recombinant cell to provide a fatty alcohol acetate ester (FACE). In an exemplary embodiment, an alcohol O-acetyltransferase (EC 2.8.1.14) enzyme processes the fatty alcohol into a fatty alcohol acetate ester (FACE). See, for example, Gabriel M Rodriguez, et al. (2014) Nature Chemical Biology 10, 259-265; Jyun-Liang Lin and Ian Wheeldon (2014) PLoS One. 2014; 9(8): PMCID: PMC4122449.
[0175] Exemplary alcohol O-acetyltransferases include yeast Aft1, e.g., GenBank accession number AY242062; GenBank accession number AY242063. See, e.g., Kevin J. Verstrepen K.J., et al (2003) Appl Environ Microbiol. 2003 Sep; 69(9): 5228-5237.
[0176] In an exemplary embodiment, a recombinant host cell comprising an engineered TE variant having an improved ability to produce a medium-chain fatty acid derivative further comprises carboxylic acid reductase activity (EC 1.2.99.6) sufficient to produce a fatty aldehyde and a fatty alcohol, and further comprises fatty alcohol O-acetyltransferase activity to convert the fatty alcohol to a fatty alcohol acetate ester.
[0177] In a further exemplary embodiment, a recombinant host cell comprising an engineered TE variant having an improved ability to produce a medium-chain fatty acid derivative further comprises carboxylic acid reductase activity (EC 1.2.99.6) that results in the production of a first fatty acid derivative, and further comprises fatty alcohol O-acetyltransferase activity to convert the first fatty acid derivative to a second fatty acid derivative, wherein the second fatty acid derivative has a higher MIC than the first fatty acid derivative.
[0178] In a further exemplary embodiment, a recombinant host cell comprising an engineered TE variant having an improved ability to produce a medium-chain fatty acid derivative further comprises carboxylic acid reductase activity (EC 1.2.99.6) that results in the production of a first fatty acid derivative, and further comprises fatty alcohol O-acetyltransferase activity to convert the first fatty acid derivative to a second fatty acid derivative, wherein the second fatty acid derivative has a higher LogP than the first fatty acid derivative.
[0179] In a further exemplary aspect, a recombinant host cell comprising an engineered TE variant having an improved ability to produce a medium-chain fatty acid derivative further comprises a carboxylic acid reductase activity (EC 1.2.99.6) that results in the production of a first fatty acid derivative, and further comprises a fatty alcohol O-acetyltransferase activity that converts the first fatty acid derivative to a second fatty acid derivative, wherein the presence of the second fatty acid derivative results in an increase in the MIC of the first fatty acid derivative.
[0180] In a further exemplary aspect, a recombinant host cell comprising an engineered TE variant having an improved ability to produce a medium-chain fatty acid derivative further comprises a carboxylic acid reductase activity (EC 1.2.99.6) that results in the production of a first fatty acid derivative, and further comprises a fatty alcohol O-acetyltransferase activity that converts the first fatty acid derivative to a second fatty acid derivative, wherein the second fatty acid derivative is less toxic than the first fatty acid derivative.
[0181] Production of fatty esters In some aspects, in addition to an engineered TE variant having improved activity for the production of a medium-chain fatty acid derivative, a recombinant host cell further comprises a polynucleotide encoding a polypeptide having fatty ester biosynthetic activity, whereby a medium-chain fatty ester is produced by the recombinant host cell.
[0182] As used herein, the term "fatty ester" or "fatty acid ester" in a similar sense refers to any ester made from a fatty acid. In an exemplary aspect, a fatty ester contains an "A side" and a "B side". As used herein, the "A side" of an ester refers to the carbon chain bonded to the carboxylate oxygen of the ester. The "B side" of an ester as used herein refers to the carbon chain containing the parent carboxylate of the ester. In an aspect where the fatty ester is obtained from a fatty acid derivative biosynthetic pathway, the A side is imparted by an alcohol and the B side is imparted by a fatty acid or an alkyl thioester.
[0183] Any alcohol can be used to form the A side of the fatty ester. In an exemplary embodiment, the alcohol is obtained from the fatty acid derivative biosynthetic pathway. In other exemplary embodiments, the alcohol is produced via a non-fatty acid derivative biosynthetic pathway. For example, the alcohol is provided exogenously. For example, the alcohol is supplied in a fermentation broth.
[0184] The carbon chain containing the A side or the B side can be of any length. However, in an exemplary embodiment where a fatty acid derivative biosynthetic pathway containing an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives provides either the A side and / or the B side of the fatty ester, the A side and / or the B side are medium-chain fatty acid derivatives and thus have a carbon chain length of 6, 7, 8, 9, or 10 carbons. Thus, in an exemplary embodiment, a fatty acid derivative biosynthetic pathway containing an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives provides the A side of the ester, and thus the A side of the fatty ester is 6, 7, 8, 9, or 10 carbons in length. In other exemplary embodiments, a fatty acid biosynthetic pathway containing an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives provides the B side of the ester, and thus the B side of the fatty ester is 6, 7, 8, 9, or 10 carbons in length.
[0185] In one exemplary embodiment, the fatty ester is a fatty acid methyl ester, for example, methyl octanoate, where the B side is provided by a fatty acid biosynthetic pathway containing an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, and the A side of the ester is 1 carbon in length. Thus, in an exemplary embodiment, the fatty acid ester is methyl octanoate. In one exemplary embodiment, the A side is provided by the action of a fatty acid O-methyltransferase (FAMT) (EC 2.1.1.15) enzyme (see, for example, Applied and Environmental Microbiology 77(22): 8052-8061).
[0186] In another exemplary embodiment, the fatty ester is a fatty acid ethyl ester, wherein the B side is provided by a fatty acid biosynthetic pathway comprising an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, and the A side of the ester has a length of 2 carbons.
[0187] In one exemplary embodiment, the A side is straight-chain. In another exemplary embodiment, the A side is branched-chain. In one exemplary embodiment, the B side is straight-chain. In another exemplary embodiment, the B side is branched-chain. The branched chain can have one or more branch points. In one exemplary embodiment, the A side is saturated. In another exemplary embodiment, the A side is unsaturated. In one exemplary embodiment, the B side is saturated. In another exemplary embodiment, the B side is unsaturated.
[0188] In an exemplary embodiment, in addition to an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell comprises a polynucleotide encoding a polypeptide having ester synthase activity (EC 3.1.1.67). Ester synthases are known in the art. See, for example, WO 2011 / 038134.
[0189] In some exemplary embodiments, the fatty acid ester is produced by a recombinant host cell comprising an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, as well as an acyl-CoA synthetase (fadD) enzyme and an ester synthase enzyme (see, for example, WO 2011 / 038134; WO 2007 / 136762; US Patent No. 8,110,670).
[0190] In an exemplary embodiment, a recombinant host cell comprising an engineered TE variant having an improved ability to produce medium-chain fatty acid derivatives further comprises sufficient ester synthase activity (EC 3.1.1.67) to produce a fatty ester (such as FAME or FAEE).
[0191] In a further aspect, a recombinant host cell comprising an engineered TE variant having improved activity that results in the production of a first fatty acid derivative further comprises ester synthase activity that converts the first fatty acid derivative to a second fatty acid derivative.
[0192] In a further aspect, a recombinant host cell comprising an engineered TE variant having improved activity that results in the production of a first fatty acid derivative further comprises ester synthase activity that converts the first fatty acid derivative to a second fatty acid derivative, wherein the second fatty acid derivative has a higher MIC than the first fatty acid derivative.
[0193] In a further aspect, a recombinant host cell comprising an engineered TE variant having improved activity that results in the production of a first fatty acid derivative further comprises ester synthase activity that converts the first fatty acid derivative to a second fatty acid derivative, wherein the second fatty acid derivative has a higher partition coefficient (LogP) than the first fatty acid derivative.
[0194] In a further aspect, a recombinant host cell comprising an engineered TE variant having improved activity that results in the production of a first fatty acid derivative further comprises ester synthase activity that converts the first fatty acid derivative to a second fatty acid derivative, wherein the presence of the second fatty acid derivative results in an increase in the MIC of the first fatty acid derivative.
[0195] In a further aspect, a recombinant host cell comprising an engineered TE variant having improved activity that results in the production of a first fatty acid derivative further comprises ester synthase activity that converts the first fatty acid derivative to a second fatty acid derivative, wherein the second fatty acid derivative is less toxic than the first fatty acid derivative.
[0196] Production of hydrocarbons In some embodiments, in addition to the engineered TE variants having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises a polynucleotide encoding a polypeptide having fatty aldehyde biosynthetic activity, such as an acyl-ACP reductase polypeptide (EC 6.4.1.2), and a polynucleotide encoding a polypeptide having hydrocarbon biosynthetic activity, such as a decarbonylase (EC 4.1.99.5), an oxidative deformylase, or a fatty acid decarboxylase, and thus, the recombinant host cell exhibits enhanced production of hydrocarbons (see, e.g., U.S. Patent Application Publication No. 2011 / 0124071). Thus, in an exemplary embodiment, a recombinant host cell comprising an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives produces hydrocarbons, such as alkanes or alkenes (e.g., terminal olefins or internal olefins) or ketones.
[0197] In some exemplary embodiments, the fatty aldehyde produced by a recombinant host cell comprising an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives is converted by decarbonylation, which removes one carbon atom, to form a hydrocarbon (see, e.g., U.S. Patent No. 8,110,670 and International Publication No. 2009 / 140695).
[0198] In other exemplary embodiments, the fatty acid produced by the recombinant host cell is converted by decarboxylation, which removes one carbon atom, to form a terminal olefin. Thus, in some exemplary embodiments, in addition to expressing an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant cell further expresses or overexpresses a polynucleotide encoding a hydrocarbon biosynthetic polypeptide, such as a polypeptide having decarboxylase activity as disclosed in U.S. Patent No. 8,597,922.
[0199] In other exemplary embodiments, the alkyl thioester intermediate is converted by enzymatic decarboxylative condensation to form an internal olefin or a ketone. Thus, in some exemplary embodiments, in addition to expressing an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant cell further expresses or overexpresses a polynucleotide encoding a hydrocarbon biosynthetic polypeptide, such as a polypeptide having OleA activity, thereby producing a ketone (see, e.g., U.S. Patent No. 9,200,299). In other exemplary embodiments, in addition to expressing an engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant cell further expresses or overexpresses a polynucleotide encoding a hydrocarbon biosynthetic polypeptide, such as OleCD or OleBCD, together with a polypeptide having OleA activity, thereby producing an internal olefin (see, e.g., U.S. Patent No. 9,200,299).
[0200] Some exemplary hydrocarbon biosynthetic polypeptides are shown in Table 2 below.
[0201] (Table 2) Exemplary hydrocarbon biosynthetic polynucleotides and polypeptides TIFF0007705243000008.tif52128
[0202] Production of omega (ω)-hydroxylated fatty acid derivatives In some embodiments, in addition to the engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises a polynucleotide encoding a polypeptide having ω-hydroxylase activity (EC 1.14.15.3). In an exemplary embodiment, the modified ω-hydroxylase has modified cytochrome P450 monooxygenase (P450) enzyme activity and efficiently catalyzes the hydroxylation of the ω-position of the hydrocarbon chain in vivo. Thus, the recombinant microorganism produces medium-chain omega-hydroxylated (ω-hydroxylated) fatty acid derivatives in vivo when grown in the presence of a carbon source derived from a renewable feedstock in the fermentation broth (see, e.g., PCT Application Publication WO2014 / 201474).
[0203] In other exemplary embodiments, in addition to the engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, the recombinant host cell further comprises a polynucleotide encoding an alkane hydroxylase such as alkA, a CYP153A-reductase or a CYP153A-reductase hybrid fusion polypeptide variant (see, e.g., International Publication No. 2015 / 195697), such that the recombinant host cell, when cultured in a medium containing a carbon source under conditions effective to express an alkane hydroxylase such as AlkA, CYP153 or a CYP153A-reductase hybrid fusion polypeptide variant and the engineered TE variant having improved activity for the production of medium-chain fatty acid derivatives, produces omega-hydroxylated (ω-hydroxylated) fatty acid derivatives and bifunctional fatty acid derivatives and compositions thereof, including ω-hydroxylated fatty acids, ω-hydroxylated fatty esters, α,ω-diacids, α,ω-diesters, α,ω-diols and chemicals derived therefrom, such as macrolactones and macrocyclic ketones.
[0204] V. Culturing and Fermenting the Recombinant Host Cell As used herein, fermentation broadly refers to the conversion of an organic matter into a target substance by a recombinant host cell. For example, this includes the conversion of a carbon source by a recombinant host cell, by increasing a culture of the recombinant host cell in a medium containing the carbon source, into a fatty acid derivative such as a medium-chain fatty acid, a medium-chain fatty acid ester, a medium-chain fatty alcohol, a medium-chain fatty alcohol acetate ester, etc. For example, the permissive conditions for the production of a target substance such as a fatty acid, a fatty ester, a fatty alcohol, a fatty alcohol acetate ester, etc. are any conditions that allow the host cell to produce a desired product such as a fatty acid derivative composition. Suitable conditions include, for example, typical fermentation conditions. See, for example, Principles of Fermentation Technology, 3rd Edition (2016), supra; Fermentation Microbiology and Biotechnology, 2nd Edition, (2007), supra.
[0205] Fermentation conditions can include a number of parameters well known in the art, including, without limitation, temperature range, pH level, aeration level, feed rate, and medium composition. Each of these conditions, individually and in combination, grows the host cell. Fermentation can be aerobic, anaerobic, or a variation thereof (such as microaerobic). Exemplary media include broth (liquid) or gel (solid). Generally, the medium includes a carbon source (e.g., a simple carbon source derived from a renewable feedstock) that can be directly metabolized by the host cell. In addition, enzymes can be used in the medium to facilitate the metabolism of the carbon source following mobilization (e.g., depolymerization of starch or cellulose into fermentable sugars) for the production of medium-chain fatty acid derivatives.
[0206] For small-scale production, host cells engineered to produce a medium-chain fatty acid derivative composition are grown; fermented; and induced to express a desired polynucleotide sequence, such as a polynucleotide encoding a polypeptide having specific enzyme activity (e.g., thioesterase (TE), carboxylic acid reductase (CAR), alcohol dehydrogenase (ADH), fatty acid acyl-CoA / ACP reductase (FAR), acyl-CoA reductase (ACR), acetyl-CoA carboxylase (ACC) and / or acyl-ACP / CoA reductase (AAR) enzyme activity) in a batch of, for example, about 100 μL, 200 μL, 300 μL, 400 μL, 500 μL, 1 mL, 5 mL, 10 mL, 15 mL, 25 mL, 50 mL, 75 mL, 100 mL, 500 mL, 1 L, 2 L, 5 L, or 10 L. For large-scale production, host cells are grown; fermented; and induced to express any desired polynucleotide sequence in a culture having a batch volume of about 10 L, 100 L, 1000 L, 10,000 L, 100,000 L, 1,000,000 L or more.
[0207] The fatty acid derivative compositions disclosed herein can often be found in the extracellular environment of recombinant host cell cultures and can be readily isolated from the culture medium. Medium-chain fatty acid derivatives, such as medium-chain fatty acids, medium-chain fatty acid esters, medium-chain fatty aldehydes, medium-chain fatty ketones, medium-chain fatty alcohols, and medium-chain fatty alcohol acetate esters, may be secreted by recombinant host cells, transported into the extracellular environment of recombinant host cell cultures, or passively migrate into the extracellular environment. Medium-chain fatty acid derivative compositions may be isolated from recombinant host cell cultures using conventional methods known in the art, including, but not limited to, centrifugation.
[0208] Exemplary microorganisms suitable for use as production host cells include, for example, bacteria, cyanobacteria, yeast, algae, filamentous fungi, and the like. To produce a fatty acid derivative composition, the production host cell (or host cell in a similar sense) is modified compared to an unmanipulated or native host cell, for example, engineered as described above and as disclosed in, for example, U.S. Patent Application Publication No. 2015 / 0064782, to include a fatty acid biosynthetic pathway. A production host engineered to include a modified fatty acid biosynthetic pathway can efficiently convert glucose or other renewable feedstocks into fatty acid derivatives. Protocols and procedures for high-density fermentation for the production of various compounds have been established (see, for example, U.S. Patent Nos. 8,372,610; 8,323,924; 8,313,934; 8,283,143; 8,268,599; 8,183,028; 8,110,670; 8,110,093; and 8,097,439).
[0209] In some exemplary embodiments, the production host cell is cultured in a culture medium (e.g., a fermentation medium) containing an initial concentration of a carbon source (e.g., a simple carbon source) of from about 20 g / L to about 900 g / L. In other embodiments, the culture medium contains an initial carbon source concentration of from about 2 g / L to about 10 g / L; from about 10 g / L to about 20 g / L; from about 20 g / L to about 30 g / L; from about 30 g / L to about 40 g / L; or from about 40 g / L to about 50 g / L. In some embodiments, the level of carbon source available in the culture medium can be monitored during the course of the fermentation. In some embodiments, the method further comprises adding an additional carbon source to the culture medium when the initial level of carbon source in the medium is less than about 0.5 g / L.
[0210] In some exemplary embodiments, an additional carbon source is added to the culture medium when the level of the carbon source in the medium is less than about 0.4 g / L, less than about 0.3 g / L, less than about 0.2 g / L, or less than about 0.1 g / L. In some embodiments, the additional carbon source is added to maintain a carbon source level of from about 1 g / L to about 25 g / L. In some embodiments, the additional carbon source is added to maintain a carbon source level of at least about 2 g / L (e.g., at least about 2 g / L, at least about 3 g / L, at least about 4 g / L). In certain embodiments, the additional carbon source is added to maintain a carbon source level of at most about 5 g / L (e.g., at most about 5 g / L, at most about 4 g / L, at most about 3 g / L). In some embodiments, the additional carbon source is added to maintain a carbon source level of from about 2 g / L to about 5 g / L, from about 5 g / L to about 10 g / L, or from about 10 g / L to about 25 g / L.
[0211] In an exemplary embodiment, the carbon source for fermentation is derived from renewable feedstocks. In some embodiments, the carbon source is glucose. In other embodiments, the carbon source is glycerol. Other possible carbon sources include, but are not limited to, fructose, mannose, galactose, xylose, arabinose, starch, cellulose, hemicellulose, pectin, xylan, sucrose, maltose, cellobiose, turanose, acetic acid, ethane, ethanol, methane, methanol, formic acid, and carbon monoxide; cellulosic materials and variants such as hemicellulose, methyl cellulose and sodium carboxymethyl cellulose; saturated or unsaturated fatty acids, succinic esters, lactic esters and acetic esters; alcohols such as ethanol, methanol, and glycerol, or mixtures thereof. In one embodiment, the carbon source is derived from corn, sugarcane, sorghum, sugar beet, switchgrass, silage, straw, wood, pulp, sewage, garbage, general cellulosic waste, flue-gas, syngas, or carbon dioxide. The simple carbon source can also be a photosynthetic product such as glucose or sucrose. In one embodiment, the carbon source is a waste such as glycerol, flue-gas, or syngas; or a rearrangement of organic matter such as biomass; or a rearrangement of natural gas or methane, or these substances to syngas; or is derived from photosynthetically fixed carbon dioxide, for example, medium-chain fatty acid derivatives can be produced using CO2 as a carbon source by photosynthetically growing recombinant cyanobacteria or algae. In some exemplary embodiments, the carbon source is derived from biomass. Exemplary sources of biomass are plant materials or vegetation such as corn, sugarcane, or switchgrass. Another exemplary source of biomass is metabolic waste such as animal matter (e.g., cow manure). Further exemplary sources of biomass include algae and other marine plants. Biomass also includes industrial, agricultural, forestry, and household waste including, but not limited to, fermentation waste, silage, straw, wood, sewage, garbage, general cellulosic waste, municipal solid waste, and food scraps.
[0212] In some exemplary embodiments, fatty acid derivatives, such as medium-chain fatty acids, medium-chain fatty esters, medium-chain fatty alcohols, etc., are produced at a concentration of about 0.5 g / L to about 40 g / L. In some embodiments, the fatty acid derivative is produced at a concentration of about 1 g / L or more (e.g., about 1 g / L or more, about 10 g / L or more, about 20 g / L or more, about 50 g / L or more, about 100 g / L or more). In some embodiments, the fatty acid derivative is produced at a concentration of about 1 g / L to about 170 g / L, about 1 g / L to about 10 g / L, about 40 g / L to about 170 g / L, about 100 g / L to about 170 g / L, about 10 g / L to about 100 g / L, about 1 g / L to about 40 g / L, about 40 g / L to about 100 g / L, or about 1 g / L to about 100 g / L.
[0213] In other exemplary embodiments, fatty acid derivatives, such as medium-chain fatty acid derivatives, are produced at titers of about 25 mg / L, about 50 mg / L, about 75 mg / L, about 100 mg / L, about 125 mg / L, about 150 mg / L, about 175 mg / L, about 200 mg / L, about 225 mg / L, about 250 mg / L, about 275 mg / L, about 300 mg / L, about 325 mg / L, about 350 mg / L, about 375 mg / L, about 400 mg / L, about 425 mg / L, about 450 mg / L, about 475 mg / L, about 500 mg / L, about 525 mg / L, about 550 mg / L, about 575 mg / L, about 600 mg / L, about 625 mg / L, about 650 mg / L, about 675 mg / L, about 700 mg / L, about 725 mg / L, about 750 mg / L, about 775 mg / L, about 800 mg / L, about 825 mg / L, about 850 mg / L, about 875 mg / L, about 900 mg / L, about 925 mg / L, about 950 mg / L, about 975 mg / L, about 1000 mg / L, about 1050 mg / L, about 1075 mg / L, about 1100 mg / L, about 1125 mg / L, about 1150 mg / L, about 1175 mg / L, about 1200 mg / L, about 1225 mg / L, about 1250 mg / L, about 1275 mg / L, about 1300 mg / L, about 1325 mg / L, about 1350 mg / L, about 1375 mg / L, about 1400 mg / L, about 1425 mg / L, about 1450 mg / L, about 1475 mg / L, about 1500 mg / L, about 1525 mg / L, about 1550 mg / L, about 1575 mg / L, about 1600 mg / L, about 1625 mg / L, about 1650 mg / L, about 1675 mg / L, about 1700 mg / L, about 1725 mg / L, about 1750 mg / L, about 1775 mg / L, about 1800 mg / L, about 1825 mg / L, about 1850 mg / L, about 1875 mg / L, about 1900 mg / L, about 1925 mg / L, about 1950 mg / L, about 1975 mg / L, about 2000 mg / L (2 g / L), 3 g / L, 5 g / L, 10 g / L, 20 g / L, 30 g / L, 40 g / L, 50 g / L, 60 g / L, 70 g / L, 80 g / L, 90 g / L, 100 g / L or titers in the range bounded by any two of said values. In other embodiments, fatty acid derivatives or other compounds are produced at titers greater than 100 g / L, greater than 200 g / L, or greater than 300 g / L.In an exemplary embodiment, the titer of a fatty acid derivative or other compound produced by a recombinant host cell by the method disclosed herein is 5 g / L to 200 g / L, 10 g / L to 150 g / L, 20 g / L to 120 g / L, and 30 g / L to 100 g / L. The titer can represent a particular fatty acid derivative or combination of fatty acid derivatives or another compound or combination of other compounds produced by a given recombinant host cell culture. In an exemplary embodiment, expression of an engineered TE variant in a recombinant host cell such as E. coli results in production of a higher titer compared to a recombinant host cell expressing the corresponding wild-type polypeptide. In one embodiment, the higher titer is in the range of at least about 5 g / L to about 200 g / L.
[0214] In other exemplary embodiments, a host cell engineered to produce a fatty acid derivative, such as a medium-chain fatty acid derivative, by the methods of the present disclosure has a yield of at least 1%, at least 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, at least about 11%, at least about 12%, at least about 13%, at least about 14%, at least about 15%, at least about 16%, at least about 17%, at least about 18%, at least about 19%, at least about 20%, at least about 21%, at least about 22%, at least about 23%, at least about 24%, at least about 25%, at least about 26%, at least about 27%, at least about 28%, at least about 29%, or at least about 30%, or a range bounded by any two of the foregoing values. In other embodiments, one or more fatty acid derivatives or other compounds are produced at a yield greater than about 30%, greater than about 35%, greater than about 40%, greater than about 45%, greater than about 50%, greater than about 55%, greater than about 60%, greater than about 65%, greater than about 70%, greater than about 75%, greater than about 80%, greater than about 85%, greater than about 90%. Alternatively or additionally, the yield is about 30% or less, about 27% or less, about 25% or less, or about 22% or less. In another embodiment, the yield is about 50% or less, about 45% or less, or about 35% or less. In another embodiment, the yield is about 95% or less, or 90% or less, or 85% or less, or 80% or less, or 75% or less, or 70% or less, or 65% or less, or 60% or less, or 55% or less, or 50% or less. Thus, the yield can be bounded by any two of the foregoing endpoints.For example, the yield of medium-chain fatty acid derivatives produced by recombinant host cells by the method disclosed herein, such as 8- and / or 10-carbon fatty acid derivatives, can be about 5% to about 15%, about 10% to about 25%, about 10% to about 22%, about 15% to about 27%, about 18% to about 22%, about 20% to about 28%, about 20% to about 30%, about 30% to about 40%, about 40% to about 50%, about 50% to about 60%, about 60% to about 70%, about 70% to about 80%, about 80% to about 90%, about 90% to about 100%, about 100% to about 200%, about 200% to about 300%, about 300% to about 400%, about 400% to about 500%, about 500% to about 600%, about 600% to about 700%, or about 700% to about 800%. The yield can represent a particular medium-chain fatty acid derivative or combination of fatty acid derivatives. In one aspect, higher yields are in the range of about 10% to about 800% of the theoretical yield. Additionally, the yield also depends on the feedstock used.
[0215] In some exemplary embodiments, the productivity of a host cell engineered to produce a fatty acid derivative, such as a medium-chain fatty acid derivative, by the methods of the present disclosure is at least 100 mg / L / hour, at least 200 mg / L / hour, at least 300 mg / L / hour, at least 400 mg / L / hour, at least 500 mg / L / hour, at least 600 mg / L / hour, at least 700 mg / L / hour, at least 800 mg / L / hour, at least 900 mg / L / hour, at least 1000 mg / L / hour, at least 1100 mg / L / hour, at least 1200 mg / L / hour, at least 1300 mg / L / hour, at least 1400 mg / L / hour, at least 1500 mg / L / hour, at least 1600 mg / L / hour, at least 1700 mg / L / hour, at least 1800 mg / L / hour, at least 1900 mg / L / hour, at least 2000 mg / L / hour, at least 2100 mg / L / hour, at least 2200 mg / L / hour, at least 2300 mg / L / hour, at least 2400 mg / L / hour, 2500 mg / L / hour, or as high as 10 g / L / hour (depending on cell mass). For example, the productivity of a malonyl-CoA-derived compound comprising one or more fatty acid derivatives or other compounds produced by a recombinant host cell by the methods of the present disclosure can be from 500 mg / L / hour to 2500 mg / L / hour, or from 700 mg / L / hour to 2000 mg / L / hour. Productivity can represent a particular C8 and / or C10 fatty acid derivative or combination of fatty acid derivatives or other compounds produced by a given host cell culture. For example, expression of an engineered TE variant in a recombinant host cell, such as E. coli, results in an increase in the productivity of C8 and / or C10 fatty acid derivatives or other compounds as compared to a recombinant host cell expressing the corresponding wild-type polypeptide. In exemplary embodiments, higher productivity ranges from about 0.3 g / L / h to about 3 g / L / h, about 10 g / L / h, about 100 g / L / h, about 1000 g / L / h.
[0216] VI. Isolation Biological products, such as compositions, containing the medium-chain fatty acid derivatives disclosed herein, produced using the recombinant host cells as described above, are typically isolated from the fermentation broth by methods known in the art. In an exemplary embodiment, a composition containing the medium-chain fatty acid derivatives disclosed herein, produced using recombinant host cells, is described above and is isolated from the fermentation broth by gravity sedimentation, centrifugation, or decantation.
[0217] VII. Compositions and Formulations of Medium-Chain Fatty Acid Derivatives Biological products, such as compositions containing medium-chain fatty acids and medium-chain fatty acid derivatives, produced using the recombinant host cells described in detail above, are produced from renewable sources (e.g., simple carbon sources obtained from renewable feedstocks), and thus are new substance compositions. These new biological products can be distinguished from petrochemical carbon-derived organic compounds based on dual carbon-isotope fingerprinting or 14 14C dating. Additionally, the origin-specific biogenic carbon (e.g., glucose vs. glycerol) can be determined by dual carbon-isotope fingerprinting known in the art (see, for example, U.S. Patent No. 7,169,588, WO2016 / 011430A1, etc.).
[0218] Furthermore, as shown below, the composition of the biological product defines the unique composition of the natural fatty acid derivatives produced from the organism. These unique compositions, which are unusually high in medium-chain fatty acid derivatives, provide a novel and unique source of these valuable medium-chain length products.
[0219] The following examples are provided to illustrate the invention and not to limit it.
Examples
[0220] The following specific examples are intended to illustrate the present disclosure and should not be construed as limiting the scope of the claims.
[0221] Example 1 The following examples illustrate that chemical modification of medium-chain length fatty acid derivative compounds reduces the toxicity experienced by microorganisms to medium-chain fatty acid derivative compounds compared to the toxicity experienced by microorganisms when grown in the presence of unmodified medium-chain length fatty acid derivative compounds.
[0222] As discussed above herein, the production of medium-chain length fatty acid derivative compounds using biological systems (e.g., fermentation of microbial cells) is a desirable route for the selective production of medium-chain length fatty acids / aliphatic compounds. Unfortunately, medium-chain fatty acid derivative compounds can be highly toxic to microbial cells, and this toxicity is a barrier to the production of medium-chain length fatty acid derivative compounds on a commercial scale via fermentation.
[0223] In this example, the toxicity was evaluated by determining the minimum inhibitory concentration (MIC) (the concentration of the compound sufficient to kill 50% of the culture) for related compounds that differed only in terms of being modified or unmodified. Compounds with lower toxicity (i.e., relatively high MIC) are compounds that are easier to produce by fermentation.
[0224] Escherichia coli cell cultures were grown at various concentrations of these compounds, and their growth was determined as a measure of the total cell number in the culture by measuring the total protein from the lysed culture 24 hours after growth.
[0225] Specifically, Escherichia coli cell cultures were grown in the presence of octanol, octanoic acid, methyl octanoate, and octyl acetate. The results are shown in Figure 1.
[0226] As can be seen in Figure 1, octanoic acid, which is a medium-chain fatty acid, and octanol, which is a medium-chain alcohol, have a MIC of 1 - 5 g / L. On the other hand, octyl acetate, methyl octanoate, and ethyl octanoate (not shown), which are esters of these medium-chain alcohols and acids, have a MIC that is 10 - 100 times higher than that of the corresponding unmodified alcohol and acid. Therefore, Escherichia coli can exhibit tolerance to chemically modified compounds at concentrations 10 - 100 times higher compared to unmodified compounds.
[0227] Therefore, the above examples demonstrate that esters of medium-chain fatty alcohols and acids can be produced and tolerated at high concentrations by industrial fermentation processes. Further, the above examples demonstrate that the toxicity of aliphatic compounds of a given chain length can be significantly reduced by modifying the functional groups associated with the toxic molecule or by slightly increasing its molecular weight.
[0228] Example 2 The following examples illustrate that the toxicity of medium-chain fatty acid derivative compounds correlates with the partition coefficient (LogP).
[0229] As shown in Example 1, esters of medium-chain fatty alcohols and esters of medium-chain fatty acids are less toxic (have a higher MIC) than the corresponding medium-chain fatty alcohols and medium-chain fatty acids.
[0230] Many water-soluble compounds have a low partition coefficient (LogP). LogP is a measure of the partitioning of a compound between water and octanol (see, for example, that compounds with a low LogP such as acetic acid, lactic acid, pyruvic acid, 1,3-propanediol, amino acids, etc. can be produced and tolerated at high concentrations by microorganisms such as Escherichia coli). Therefore, it can be concluded that compounds that are less hydrophobic (or equivalently more hydrophilic) and thus have a lower LogP will be less toxic. To evaluate whether this is true, the inventors measured the logP of the compounds disclosed in Figure 1 (i.e., octanol, octanoic acid, octyl acetate, and methyl octanoate).
[0231] Surprisingly, as shown in FIG. 2, for medium-chain aliphatic compounds, the toxicity as a function of LogP is contrary to expectations. That is, octanol and octanoic acid, which are compounds with low logP, have high toxicity (i.e., low MIC). Compounds with high logP such as octyl acetate, methyl octanoate, and ethyl ocanoate have lower toxicity (higher MIC).
[0232] Therefore, this example demonstrates that the modification from a toxic medium-chain aliphatic compound with low LogP to a compound with higher logP wo is a useful method for reducing the toxicity of medium-chain aliphatic compounds that are toxic to industrial microorganisms such as Escherichia coli.
[0233] Example 3 The following examples illustrate that the expression of a novel biochemical pathway that catalyzes the conversion of toxic medium-chain aliphatic compounds to their less toxic derivatives can render microorganisms tolerant to the pathway leading to the toxic compounds and that the microorganisms can produce the derivatives at high levels.
[0234] As discussed in Examples 1 and 2 above, medium-chain fatty acid derivative compounds such as fatty alcohols and fatty acids are toxic to host cells, but their slightly higher molecular weight derivatives and high logP derivatives are not. Therefore, the inventors determined that it is possible to produce toxic compounds in microorganisms without killing the cells by biochemically converting more toxic compounds to less toxic compounds in vivo. The less toxic compounds can then be produced and tolerated at high levels. These less toxic compounds can, once produced, be isolated and used as such, or isolated and chemically converted back to more toxic compounds.
[0235] As shown below, manipulating cells to modify the functional groups of toxic medium-chain fatty acid derivatives, for example, by esterifying them with short-chain acids or alcohols, eliminates the cytotoxic response of cells to the unesterified compounds and enables the engineered cells to survive from the expression of highly productive biochemical pathways that lead to toxic compounds. This enables a novel and selective process for producing these medium-chain length fatty acid derivatives at concentrations far exceeding their inhibitory levels.
[0236] Moreover, providing esterified medium-chain fatty acids and / or esterified medium-chain fatty alcohols by modifying medium-chain fatty acids and / or medium-chain fatty alcohols by esterification further reduces the toxicity of medium-chain intermediates by acting as an extractant in the biosynthetic pathway.
[0237] Esterified medium-chain fatty acid derivatives as extractants Figure 3 shows the results of an experiment designed to test whether the presence of octyl acetate can protect cells from 1-octanol toxicity. As is clear from Figure 3, after exposure to 1-octanol at a concentration of 0.5 grams per liter (g / L) for 5 hours, the viability of E. coli cells was completely lost. However, interestingly, when 50 g / L (a concentration non-toxic to E. coli cells) of octyl acetate was also added, the cell viability was maintained at 100% of the control level even when the cells were exposed to 1-octanol at a concentration of 0.5 g / L and even when exposed to 1-octanol at a concentration of 1 g / L. When the cells were exposed to 1-octanol at a concentration of 10 g / L (far exceeding the MIC observed for 1-octanol), the viability decreased by less than 20%.
[0238] Modification of tolerance to medium-chain fatty alcohols by expressing fatty alcohol acetyltransferase As discussed and shown above, the production of medium-chain (C6-C10) fatty alcohols by microorganisms is limited by their toxicity. Considerable efforts have been devoted to identifying genetic and biochemical mechanisms to increase tolerance to these medium-chain fatty acid derivative compounds (see, for example, Lennen and Pflefer, 2013; Royce et al., 2015; Tan, et al., 2016; Tan, et al., 2017). However, to date, no solution has been found that enables production at commercial titers (e.g., concentrations between about 10 g / l and 200 g / l or higher).
[0239] In Example 1, the inventors demonstrated that medium-chain fatty alcohol acetates are less toxic than the corresponding medium-chain length fatty alcohols when added to the culture medium. In the following experiments, the inventors demonstrate that the expression of the pathway for producing medium-chain length fatty alcohols intracellularly is cytotoxic, resulting in insufficient cell growth and limited medium-chain alcohol production by these cells. The inventors further show that when the same strain was further engineered to express the biochemical pathway for converting medium-chain alcohols to alcohol acetates, the cells grew well and produced significant amounts of fatty alcohol acetates. Thus, the biochemical conversion of medium-chain length fatty alcohols synthesized in the cells to their acetate alcohols eliminates the toxicity of the intermediate medium-chain fatty alcohols and enables the production of high levels of fatty alcohol acetates. This further demonstrates that the gene encoding medium-chain alcohol-O-acetyltransferase confers tolerance to medium-chain fatty alcohols produced intracellularly (Figure 4).
[0240] Cells can be engineered to produce fatty alcohols through various biochemical pathways (see, e.g., FIG. 4). Such biochemical pathways include thioesterases (TEs) that hydrolyze fatty acid thioesters intracellularly to produce fatty acids (see, e.g., PCT / US1998 / 011697, U.S. Pat. No. 9,765,368, PCT / US2010 / 04049), carboxylic acid reductases that catalyze the ATP- and NAD(P)H-dependent reduction of fatty acids to fatty aldehydes, and alcohol dehydrogenases that catalyze the NAD(P)H-dependent reduction of fatty aldehydes to fatty alcohols (note that most cells have sufficient alcohol dehydrogenase activity to catalyze this reaction intracellularly, but overexpression of these or similar enzymes can ensure that fatty aldehydes do not accumulate; see, e.g., WO 2010 / 062480), but are not limited thereto. Other pathways that can be engineered to produce fatty alcohols include fatty acyl reductases that catalyze the reduction of fatty acid acyl thioesters (see, e.g., Kim et al, 2015) to fatty aldehydes.
[0241] Acetylation of fatty alcohols can be achieved, for example, by expression of an alcohol-O-acetyltransferase (EC 2.3.1.84) that catalyzes the acetyl-CoA-dependent acetylation of alcohols (FIG. 4). Alcohol acetyltransferases (AATs) are diverse, and an appropriate AAT can be selected from families such as plant AATs (e.g., strawberry SAAT or FaAAT2, petunia PhcFATB2), yeast ATF (Saccharomyces cerevisiae ATF1) (see, e.g., PCT / US2014 / 053587), etc. As a non-limiting example, here the inventors demonstrate the effect of expressing S. cerevisiae ATF1 in E. coli cells engineered to produce fatty alcohols.
[0242] To determine whether it was beneficial to express the acetylation pathway, the viability of a strain expressing a pathway for the biosynthesis of medium-chain fatty alcohols was compared with the viability of an isogenic strain expressing ATF, which would convert (toxic) medium-chain fatty alcohols to fatty alcohol acetates (less toxic). In this evaluation, the sRG.674 strain produces fatty alcohol species in which 85-90% of the total fatty species (FAS) produced have a carbon chain length of 8 or 10 (C8+C10 fatty alcohol (C8+C10 FALC)). The sJN.209 strain is isogenic to sRG.674 except that the S. cerevisiae atf1 gene was added to the plasmid expressing the fatty alcohol (FALC) pathway.
[0243] (Table 3) Strains producing medium-chain fatty alcohols (FALC) or fatty alcohol acetates (FACE) TIFF0007705243000009.tif25128
[0244] The sRG.674 and sJN.209 strains were grown in a 5 L bioreactor using minimal salt medium containing glucose as the carbon source supplied at the maximum consumption rate as described in Examples 9 and 10 (Figure 5). Even before the addition of IPTG to induce the expression of the FALC pathway and the production of medium-chain compounds, the FALC-producing strain (sRG.674) could not grow (Figure 5A). Without being bound by theory, the growth failure of the FALC-producing strain may be due to the low-level constitutive expression of enzymes involved in medium-chain FALC synthesis and the early production of inhibitory concentrations of C8 and C10 FALC. On the other hand, in the sJN.209 strain, this growth inhibition was not observed due to the expression of AAT. Instead, complete growth and the production of fatty alcohol acetate (FACE) were observed throughout the 72-hour fermentation.
[0245] Comparison of the levels and compositions of the fatty species produced (Figures 5B and 5C) further demonstrates the strong ability of the acetyltransferase gene to transiently enable high levels of medium-chain length fatty alcohol production in cells by converting the medium-chain length fatty alcohol to the less toxic fatty alcohol acetate ester.
[0246] Modifications for tolerance to medium-chain free fatty acids Similar to medium-chain fatty alcohols, medium-chain free fatty acids are toxic to microbial cells (see, e.g., Figure 1). Here, the inventors show that production of such compounds can be greatly improved by providing cells with the ability to convert free fatty acids to less toxic alkyl esters such as fatty acid methyl esters (FAMEs) or fatty acid ethyl esters (FAEEs).
[0247] Esterification of medium-chain FFA can be achieved through the expression of a fatty acid acyl-CoA synthetase (such as FadD from E. coli) that catalyzes coenzyme A (CoA)- and adenosine triphosphate (ATP)-dependent acyl-CoA (acyl-CoA) synthesis, and the expression of an ester synthetase that catalyzes the alcoholysis of thioesters such as acyl-CoA (the product of fatty acid acyl-CoA synthetase or an intermediate in the β-oxidation or reverse β-oxidation pathway) (Figure 6).
[0248] Esterification of medium-chain FFA can also be achieved through the expression of a medium-chain length selective ester synthetase (such as, e.g., Example?) that catalyzes the direct alcoholysis of medium-chain length acyl-ACP (which is also an alkyl thioester).
[0249] The benefit of expressing the ester synthesis pathway was demonstrated by comparing a strain engineered to express a thioesterase with improved activity for the production of medium-chain length fatty acid derivatives to an isogenic strain that also expressed acyl-CoA synthetase and ester synthetase for viability and the medium-chain fatty acid derivatives produced.
[0250] The sRS.786 strain was engineered to express medium-chain-length thioesterase (SEQ ID NO:49), which produces free fatty acids (FFAs) that are mostly C8 and C10 FFAs (Figure 7C). The Stpay.179 strain is isogenic to sRS.786 but also expresses fatty acid acyl-CoA synthetase and ester synthase. Stpay.179 produces medium-chain-length fatty alkyl esters when short-chain alcohols such as methanol or ethanol are provided in the medium (Figure 7C).
[0251] (Table 4) Strains producing medium-chain fatty acids (FFAs) or fatty alkyl esters TIFF0007705243000010.tif29128
[0252] The sRS.786 and Stpay.179 strains were grown in a 5 L bioreactor fed-batch culture using minimal salt medium containing glucose as the carbon source supplied at a rate of 14 / g / h as described in Examples 11 and 12 below. In addition, either ethanol (Figure 7) or methanol (not shown) was supplied during the fermentation process to maintain the alcohol concentration at around 2 g / L.
[0253] The strain that produces only FFA (sRS.786) stopped growing and consuming glucose approximately 10 hours after the addition of IPTG, which induces the expression of medium-chain-length acyl-ACP thioesterase (SEQ ID NO:49), and produced approximately 5 g of C8+C10 FFA. On the other hand, the Stpay.179 strain, which expresses the esterification pathway, was able to grow and produce total fatty acid species with a titer exceeding 84 g / kg, 93% of which were C8-C10 FFAs (Figure 7B and Figure 7C). Similar results were observed when ethanol or methanol was used as the alcohol supplemented for ester synthesis.
[0254] These data demonstrate that the expression of an ester synthesis pathway that catalyzes the conversion of toxic intracellular medium-chain free fatty acids to less toxic alkyl esters (e.g., fatty acid methyl esters or fatty acid ethyl esters) enables the production of high levels of medium-chain length fatty acid derivatives. These data further show that the expression of the ester synthesis pathway enables high levels of medium-chain length selective thioesterase expression by eliminating its toxicity.
[0255] Example 4 The following examples illustrate engineered thioesterase variants that contain a single amino acid substitution and have improved activity and / or selectivity for the production of medium-chain length fatty acid derivatives.
[0256] The production of medium-chain (C6-C10) length fatty acid derivatives using biotechnology is currently limited in part by the activity and selectivity of available thioesterases (TEs). One of the most active and selective available TEs is the thioesterase chFatB2 from Cuphea hookeriana, an enzyme having the amino acid sequence set forth by SEQ ID NO:1.
[0257] However, unfortunately, this enzyme found in nature has many limitations. This enzyme is not sufficiently expressed as a soluble protein in microorganisms, its specific activity is low, and it is more selective for the hydrolysis of C10 thioesters than C8 thioesters. To generate new TEs with improved activity, selectivity, and solubility, the inventors have made extensive engineering efforts to identify amino acid substitutions in SEQ ID NO:1 that can result in novel engineered TE variants having improved activity for the production of medium-chain fatty acid derivatives. Such TE variants having improved activity for the production of medium-chain fatty acid derivatives can achieve the improvement in their activity for the production of medium-chain fatty acid derivatives through any one or more of improved catalytic activity, improved selectivity, and / or improved solubility.
[0258] SEQ ID NO:1 has 328 amino acids (6560 possible single amino acid variants), but no three-dimensional crystal structure has been reported that could assist in reasonable enzyme engineering efforts. The inventors first made efforts to identify and engineer single mutations to SEQ ID NO:1 that would result in engineered TE variants that show a significant increase in enzyme activity and / or medium-chain length selectivity compared to those of the parental sequence SEQ ID NO:1.
[0259] To evaluate such mutations, genes encoding novel engineered TEs with selected single amino acid substitutions were expressed in E. coli and grown under conditions that support TE-dependent fatty acid derivative production. The amount and composition of the medium-chain fatty acid derivatives produced by the strain were then quantified and compared to the amount and composition of the medium-chain fatty acid derivatives produced by a control strain that was identical except for expressing the enzyme with SEQ ID NO:1.
[0260] E. coli, which do not naturally produce free fatty acids, can produce free fatty acids when engineered to express a heterologous TE, and the amount and composition of these fatty acids are directly correlated with the activity and selectivity of the expressed TE (see, for example, Yuan et al., (1995) supra; International Patent Application Publication WO2007136762; International Patent Application Publication WO2008119082).
[0261] As discussed in Examples 1-3 above, the production of medium-chain length fatty acids and medium-chain length fatty alcohols is toxic to microorganisms such as E. coli. To enable the host E. coli used to evaluate engineered TEs to be tolerant to engineered TEs that produce potentially toxic levels of medium-chain length fatty acids, the E. coli used was also engineered to express a gene that increases the cells' tolerance to medium-chain length fatty acids by affecting the conversion from medium-chain length fatty acids to fatty alcohol acetates such that the levels and composition of the fatty alcohol acetates produced by the engineered cells are directly correlated with the activity and selectivity of the expressed TE.
[0262] Generation of control evaluation strains A gene encoding the polypeptide of SEQ ID NO:1 was synthesized for optimal translation in Escherichia coli, which is shown as SEQ ID NO:60. This gene was cloned into a pACYC-based plasmid (Genbank Accession X06403) conferring resistance to kanamycin such that the gene was under the transcriptional control of thePtrc promoter (see, for example, Camsund et al. Journal of Biological Engineering 2014, 8:4) induced in the presence of isopropylthiogalactoside (IPTG). The resulting plasmid pIR.108 (Figure 8) was transformed into Escherichia coli derived from MG1655, which has been engineered to overexpress the gene EntD from the chromosome (see, for example, International Patent Application Publication WO2010062480) and contains aPtrc-controlled operon expressing the genes carB, alrA and aftA1 that together affect the biochemical conversion from free fatty acid (FFA) to fatty alcohol acetate ester (FACE) as described in Example 3 above.
[0263] To enable the effective testing of engineered TE with high activity and specificity, several control assessment strains were used (Table 5), each engineered to have a different fatty acid derivative production capacity, i.e., to support different levels of carbon flux through the fatty acid pathway. Further, in some cases, novel engineered TE variants with improved activity for the production of medium-chain fatty acid derivatives were used as control TE instead of SEQ ID NO: 1 to identify highly active and improved engineered TE variants. For example, TE with a single amino acid substitution was compared to SEQ ID NO: 1 expressed in a strain engineered to have a moderate fatty acid flux. Once a novel highly active TE variant was developed, SEQ ID NO: 1 expressed in the moderate fatty acid flux strain no longer served well as a control. Instead, a novel highly active engineered TE variant expressed in a strain with high fatty acid capacity and having multiple amino acid substitutions was used as a control. In summary, the inventors used five different TEs and strains to support the evaluation of novel engineered TE variants, enabling the best quantification and identification of the performance improvement of each TE variant. A list of control base strains with various fatty acid derivative production capacities is shown in Table 5. Table 6 shows the performance of control TE compared to the wild-type sequence (SEQ ID NO: 1).
[0264] (Table 5) Description of control base strains with different fatty acid derivative production capacities TIFF0007705243000011.tif27131The range of FAS titers (mg / L) in the HTP screening of "Quantification of relative performance of engineered TE variants" below depends on the level of flux to the alkyl thioester engineered in each strain.
[0265] (Table 6) Performance of control engineered thioesterase variants compared to the wild-type sequence (SEQ ID NO: 1) TIFF0007705243000012.tif67147
[0266] Identification of engineered TE with improved activity compared to SEQ ID NO: 1 Strains expressing the manipulation TE described in Table 7 were each grown under conditions that resulted in the expression of the gene encoding its unique manipulation TE that affects the production of medium-chain fatty acids and the genes encoding CarB, AlrA, and Aft1 that affect the conversion of medium-chain fatty acids to medium-chain fatty alcohol acetates. The resulting fatty acid-derived products were extracted, quantified, and then compared to the fatty acid derivative products produced by control assessment strain 1 expressing SEQ ID NO:1 (Table 6) grown under the same conditions. Detailed methods regarding growth and analysis of the resulting fatty acid derivatives are described below.
[0267] Table 7 describes manipulation TEs with improved performance with respect to selectivity (%C8 FAS / %C10 FAS) of C8 products compared to C10 products, reported as (1) activity, i.e., total fatty acid-derived products produced by the culture, (2) C8 selectivity, i.e., %C8 FAS relative to total FAS produced by the culture, and (3) performance as fold over control (FOC). The single mutants shown in Table 7 are relative to SEQ ID NO:1. Thus, for example, P3K shows a substitution mutation (from proline to lysine) at amino acid position 3 of SEQ ID NO:1.
[0268] (Table 7) Manipulation thioesterase variants with improved ability to produce total FAS, %C8 FAS relative to total FAS, and / or %C8 FAS / %C10 FAS FOC: Fold over control TIFF0007705243000013.tif218159TIFF0007705243000014.tif232159TIFF0007705243000015.tif255157
[0269] Quantification of the relative performance of manipulation TE variants To quantify the performance of each TE variant, cultures of cells expressing the variant were grown under conditions that support the expression of TE, CarB, AlrA, and Atf1 as follows, the resulting fatty acid derivatives were extracted, and quantified by gas chromatography with a flame ionization detector (GC / FID).
[0270] The composition and amount of the resulting fatty acid derivatives (fatty acids, fatty alcohols, and fatty alcohol acetates) were determined and then compared to the fatty acid derivatives produced under the same conditions by a control evaluation strain expressing the control TE. Briefly, single colonies of each strain were inoculated into wells of a 96-well plate containing 200 μL of Luria Bertani medium containing the appropriate antibiotic. 40 μL of this culture was used to inoculate 360 μL of the same medium in a 96-deep well plate, which was shaken at 32 °C for 4 hours. 40 μL of this culture was used to inoculate 360 μL of the production medium (Table 8) in the final 96-deep well plate. 60 μL of hexadecane was overlaid on these cultures, shaken at 32 °C for 2 hours, IPTG was added (up to 1 mM) to induce the expression of TE, CarB, AlrA, and Atf1, and shaking was continued for an additional 20 hours, after which the cultures were evaluated as follows.
[0271] (Table 8) Production medium TIFF0007705243000016.tif123128
[0272] Sample preparation and quantification of fatty acid derivatives (FAS) 400 μL of butyl acetate (containing 500 mg / L undecanol as an internal analysis standard) was added to each well, the plate was heat-sealed, shaken at 2000 rpm for 15 minutes, centrifuged at 4500 rpm for 10 minutes at room temperature, and then 100 μL of the upper organic layer was transferred to a 96-well plate containing 100 μL of N,O-bis(trimethylsilyl)trifluoroacetamide (BSTFA) (see, for example, Stalling DL, et al. Biochemical and Biophysical Research Communications. 1968 May 23;31(4):616-22). The plate was sealed and evaluated by gas chromatography with a flame ionization detector (GC-FID).
[0273] As an "internal plate control" for strains expressing engineered TE variants, control evaluation strains were included on each plate. To determine the relative performance of the engineered TE variants, the total amount of fatty acid derivatives (products resulting from the action of the expressed TE and the downstream conversion enzymes CarB, AlrA, and Atf1: fatty acids, fatty aldehydes, fatty alcohols, and fatty alcohol acetate esters) or a specific fatty acid derivative (a specific chain length, i.e., C8 or C10, etc.) was quantified and then compared to the same parameter for the control evaluation strains reported as the fold over control (FOC) relative to the control from the same plate. For example, the total FOC of the total FAS of variant A was determined by adding up the total titers of all fatty acid species identified in the extract of variant A and dividing it by the total FAS titer of the internal control evaluation strain. Engineered TE variants with improved activity relative to the control are expected to show an FOC greater than 1.0 for the reported parameters. The total FOC of the total C8 FAS of variant A was determined by adding up the total concentrations of all fatty acid species with a C8 chain length identified in the extract of variant A and dividing it by the total concentrations of all fatty acid species with a C8 chain length identified for the internal control evaluation strain. For engineered thioesterase variants containing a single amino acid substitution, the primary metrics used to identify hits were as follows: (a) improved FOC of total FAS, (b) improved FOC of %C8 FAS relative to total FAS, and / or (c) improved %C8 / %C10.
[0274] The mutations shown in Table 7 (above) were surprisingly identified as having the ability to significantly (a) improve the FOC of total FAS; and (b) improve the FOC of %C8 FAS relative to total FAS. Thus, engineered thioesterase (TE) variants containing the mutations listed in Table 7 represent novel engineered TE variants with improved activity for the production of fatty acid derivatives. In particular, the engineered TE variants shown in Table 7 represent novel TE variants with improved activity for the production of C8 and / or C10 fatty acid derivatives.
[0275] Example 5 TE variants were engineered to contain multiple amino acid substitutions that create new TEs with improved activity for the production of medium-chain fatty acid derivatives. The variants had improved activity and selectivity against the native thioesterase enzyme (SEQ ID NO:1) and engineered TE variants with single amino acid substitutions (Example 4, Table 7).
[0276] As in Example 4, genes encoding engineered TE variants with multiple amino acid substitutions were synthesized and cloned into expression vectors that affect their expression when grown in the presence of IPTG. These were transformed into an E. coli strain derived from MG1655 that had been engineered to overexpress the gene EntD from the chromosome (see, for example, WO2010062480) and that harbored the Ptrc control operon that expresses the genes carB, alrA, and aftA1 that together affect the biochemical conversion of free fatty acids (FFA) to fatty alcohol acetates (FACE) as described in Example 3 above. Engineered TEs with multiple amino acid substitutions were compared to specific control assessment strains that were identical except that the TE was expressed.
[0277] Table 7 lists novel engineered TEs that have multiple amino acid substitutions and show improved activity for the production of medium-chain fatty acid derivatives (FAS) and improved selectivity for the production of medium-chain length fatty acid derivatives (SEQ ID NOs: 2 - SEQ ID NO: 15) compared to the listed control assessment strains and TEs. Thus, since the novel engineered TEs are thioesterases with improved activity for the production of medium-chain fatty acid derivatives, the novel engineered TEs, their individual mutations, and their unique combinations of mutations are each useful tools for the production of medium-chain length fatty acid derivatives.
[0278] Example 6 The following examples illustrate engineered TE variants with increased surface charge and improved activity for the production of medium-chain length fatty acid derivatives. Thioesterase variants / mutants with improved activity for the production of medium-chain fatty acid derivatives were engineered using three-dimensional modeling.
[0279] In some embodiments, SEQ ID NO:1 appears to be toxic when overexpressed in E. coli. Without being bound by theory, SEQ ID NO:1 is thought to be unstable at high concentrations within cells or to be able to aggregate easily. Therefore, in order to reduce the potential toxicity of the protein, a three-dimensional model was computationally constructed and used to modify the charge on the surface of SEQ ID NO:1.
[0280] In this example, a three-dimensional molecular model of SEQ ID NO:1 was computationally constructed by templating the x-ray crystal structure of other acyl-ACP thioesterases. Based on this model, specific residues of the SEQ ID NO:1 enzyme were identified that were predicted to change the net surface charge. In particular, negatively charged residues (Asp or Glu) on the enzyme surface were mutated to positively charged residues (Arg or His), thereby modifying the net surface charge from +15 to +25. As shown in Table 7, the resulting engineered TE variants with increased surface charge produce a higher percentage of C8 fatty acid derivatives. Therefore, the engineered TE variants have improved activity for the production of medium-chain fatty acid derivatives.
[0281] 3-D modeling of SEQ ID NO:1 thioesterase. Since the experimental 3D structure of SEQ ID NO:1 was not available, a 3D model based on the homology of this enzyme was computationally constructed as disclosed in the following steps 1-5.
[0282] (1) Identification of homologous thioesterases of known structure The Protein Data Bank (PDB) is the single worldwide archive of structural data of biological macromolecules (see Berman, H.M. et al, Nucl. Acids Res. (2000) 28 (1): 235-242). The PDB Protein Data Bank is available on the World Wide Web at rcsb.org / pdb / home / home.do. The PDB data bank was used to identify three resolved x-ray crystal structures of thioesterase. In particular, the three resolved structures of the identified thioesterase were (1) an acyl-ACP thioesterase from Bacteroides thetaiotaomicron having a Protein Data Bank identification number (PDB ID: 2ESS), (2) an oleoyl thioesterase from Lactobacillus plantarum (PDB ID: 2OWN), and (3) an acyl-ACP thioesterase from Spirosoma linguale (PDB ID: 4GAK). These structures, which show overall sequence identity of about 25% to SEQ ID NO:1, were used as templates.
[0283] (2) Alignment of the query sequence to the template structure Three resolved 3D structures of thioesterases (2ESS, 2OWN, and 4GAK) identified in the PDB and their sequences were aligned using the PROMALS3D multiple sequence and structure alignment server available at prodata.swmed.edu / promals3d / promals3d.php on the World Wide Web (see, for example, J. Pei and N.V.Grishin (2007) Bioinformatics. 23(7): 802-808; J. Pei et al., (2008) Nucl. Acids Res. 36 (7): 2295-2300). After aligning the sequences and structures of 2ESS, 2OWN, and 4GAK, the query sequence of SEQ ID NO:1 was aligned to the sequence alignment based on the existing structure using MMFFT version 7 (see, for example, Katoh, K., et al. (2013) Mol. Biol. Evol. Apr; 30(4): 772-780). The software is available at mafft.cbrc.jp / alignment / software on the World Wide Web. The alignment of SEQ ID NO:1 and acyl-ACP thioesterases (2ESS, 2OWN, and 4GAK) identified in the PDB is shown in FIG. 9.
[0284] (3) Construction of a 3D structural model of the homology of SEQ ID NO:1 thioesterase Using MODELLER software (see, e.g., B. Webb, A. Sali. Comparative Protein Structure Modeling Using Modeller. Current Protocols in Bioinformatics, John Wiley & Sons, Inc., 5.6.1-5.6.32, 2014), a homology model of amino acids 37 - 310 was constructed using all three templates 2ESS, 2OWN, and 4GAK as well as the alignment based on the structure described in step 2 above. Further structure refinement was carried out by the built-in refinement mode in MODELLER. Refinement was performed with all default parameters using the VTFM optimization and MD refinement modules. Information on MODELLER software and downloads is available at salilab.org / modeller on the World Wide Web.
[0285] (4) Construction of ab initio models of the N-terminal and C-terminal domains As shown in Figure 9, the SEQ ID NO:1 thioesterase used in these experiments has N-terminal and C-terminal residues (36 amino acids at the N-terminus and 18 amino acids at the C-terminus) that are not included in the template x-ray crystal structure. Thus, there is no suitable template for constructing a homology-based model for these portions. Therefore, ab initio models (see, for example, J. Lee et al., (2009) Ab Initio Protein Structure Prediction pgs. 3-25 In: From Protein Structure to Function with Bioinformatics, D.J. Rigden (ed.) Springer) for both the N-terminus and C-terminus were constructed by the ROBETTA server (see, for example, Kim, D.E., et al. (2004) Nucleic Acids Res. Jul 1; 32(Web Server issue):W526-W531; available at robetta.org on the World Wide Web).
[0286] (5) Construction of full-length models for SEQ ID NO:1 thioesterase and engineered TE variants Full-length models were created using MODELLER software with three templates: a major portion homology model, an N-terminal ab initio model, and a C-terminal ab initio model.
[0287] Engineered TE variants with amino acid substitutions P3K, L176V, D196V, K203R, V282S (SEQ ID NO:4) relative to the wild-type control demonstrated improved medium-chain length fatty acid derivative production ability (Example 5, Table 7). Therefore, the model of SEQ ID NO:1 was remodeled to SEQ ID NO:4 by substantially substituting five variant residues and re-performing structure refinement in the MODELLER built-in refinement mode. Subsequently, surface residues were defined based on the final model (Figure 10).
[0288] Preparation of Engineered TE Variants with Increased Modeled Surface Charge Based on the 3D structure model of SEQ ID NO:4 above, 12 aspartic acid (D) and glutamic acid (E) residues were modeled to impart negative charge to the surface of SEQ ID NO:4. Then, as described in Examples 4 and 5, genes encoding engineered TE variants with negative substitutions from various positive substitutions of these 12 residues were synthesized and evaluated for the production of improved medium-chain length fatty acid derivatives compared to control TE (SEQ ID NO:4).
[0289] Table 7 describes a series of engineered TE variants (SEQ ID NOs: 16 - SEQ ID NO:46) that have amino acid substitutions resulting in an increase in modeled surface charge compared to SEQ ID NO:4 and have improved activity for the production of medium-chain length fatty acid derivatives. Therefore, the TEs listed as SEQ ID NOs: 16 - SEQ ID NO:46 in Table 7 are novel engineered TE variants useful for the production of medium-chain length fatty acid derivatives. Furthermore, engineered variant TEs with amino acid substitutions that increase the modeled surface charge are useful for the production of improved medium-chain length fatty acid derivatives compared to TEs that do not have an increase in the modified modeled surface charge.
[0290] Preparation of Novel Thioesterases Containing Multiple Engineered Mutations with Increased Modeled Surface Charge and Multiple Engineered Mutations that Increase Activity and / or Selectivity for the Production of Medium-Chain Length Fatty Acid Derivatives Amino acid substitutions (Table 7) predicted to increase the thioesterase surface charge identified by 3-D modeling as described above and that resulted in improved production of medium-chain length fatty acid derivatives were combined with an engineered TE variant having SEQ ID NO:15, the activity of which for the production of medium-chain fatty acid derivatives was improved compared to its corresponding control (Example 6, Table 7).
[0291] Similar to Example 5, genes encoding engineered TE variants with multiple amino acid substitutions were synthesized and cloned into expression vectors that affect their expression when grown in the presence of IPTG. These were transformed into an Escherichia coli strain derived from MG1655 that has been engineered to overexpress the gene EntD from the chromosome (see, for example, WO2010062480) and that harbors the Ptrc control operon expressing the genes carB, alrA, and aftA1 that together affect the biochemical conversion of free fatty acids (FFA) to fatty alcohol acetate esters (FACE) as described in Example 3 above. The engineered TE with multiple amino acid substitutions was compared to a control evaluation strain where the only difference was that the expressed TE was the control TE SEQ ID NO:15. Table 7 describes a series of engineered TE variants (SEQ ID NO:47 - SEQ ID NO:51) derived from this example that have improved production activity of fatty acid derivatives relative to the control (SEQ ID NO:15). The TEs listed in Table 7 are novel engineered TE variants useful for the production of medium-chain length fatty acid derivatives as they are thioesterases with improved activity for the production of medium-chain fatty acid derivatives.
[0292] Example 7 The following examples illustrate engineered TE variants that have a shortened N-terminus, increased solubility, and improved activity for the production of medium-chain length fatty acid derivatives.
[0293] Plant FatB-like thioesterases have a signal peptide that mediates their translocation from the endoplasmic reticulum to the plastid. These enzymes are known to contain an N-terminal hydrophobic region that remains after processing of the signal peptide. This region is thought to be involved in the association of the thioesterase with the thylakoid membrane. When expressed in microorganisms such as Escherichia coli, the wild-type (SEQ ID NO:1) and novel engineered TE variants described above are insoluble and associate with the membrane pellet during cell lysis and centrifugation. The low enzyme solubility suggested that much of the enzyme could associate with the membrane or misfold and become inactive.
[0294] To create novel engineered TE variants with improved solubility and activity for the production of medium-chain fatty acid derivatives, the polypeptide having SEQ ID NO:49 was engineered to have a truncation between amino acid 2 and amino acid 40, which is a region modeled to carry important hydrophobic residues that are suspected to be responsible for the poor solubility of this enzyme. Then, the solubility and activity of these engineered TE variants were evaluated by comparison with control TE of the same amino acid sequence without the truncation between amino acid 2 and amino acid 40.
[0295] Evaluation of the solubility of engineered TE variants having a truncation between amino acid 2 and amino acid 40 Genes encoding engineered TE variants having deletions between amino acid 2 and amino acid 40 of SEQ ID NO:49 were synthesized and cloned under the control of the Ptrc promoter in a medium-copy-number pACYC-based expression plasmid (Gen Bank Accession X06403). These plasmids were then transformed into E. coli, and their ability to direct the expression of TE variants with increased solubility was evaluated by comparison with the same strain carrying a plasmid that directs the expression of control TE (SEQ ID NO:49). The only difference between the strains expressing engineered TE variants having a truncation between amino acid 2 and amino acid 40 and the strain expressing control TE (SEQ ID NO:49) was the sequence of the TE being expressed. The resulting strains expressing control and truncated TE were each grown in 96-well plates under conditions that result in the expression of the gene encoding TE as described in Example 4. Cells were collected by centrifugation and resuspended in 50 μL of 50 mM Tris-HCl (pH 7.8) containing 25 mM NaCl, 5 mM EDTA, and 1 mg / mL of lysozyme. Samples were incubated at 25°C and 1500 rpm. After 20 minutes, 10 μL of 1 mg / mL DNase I solution and 10 μL of 1 M MgSO4 were added to each sample. Samples were then shaken at 1500 rpm for an additional 20 minutes. The resulting whole cell lysates (WCL) were centrifuged at 4500 rpm for 10 minutes to separate the insoluble fraction (pellet) from the soluble fraction (supernatant).
[0296] Western blot using an antibody directed against the C-terminus of TE was used to follow the presence of control and engineered TE truncation variants in the soluble fraction of WCL. As shown in Figure 11, the control polypeptide (thioesterase having SEQ ID NO:49) is visible in the WCL (this indicates that this polypeptide is expressed in the host cell), but is not present in the soluble fraction, indicating low solubility. On the other hand, the engineered truncated TE variants were found in both the WCL and the soluble fraction, and all TE variants showed a significant increase in the presence of TE in the soluble fraction.
[0297] This demonstrates that the solubility of plant FatB-like thioesterases such as SEQ ID NO:1, which is insoluble when expressed in microorganisms, can be improved by expressing engineered variants in which the N-terminus of the enzyme is truncated. It further shows that the engineered truncated variants have increased solubility when expressed in E. coli compared to TE without this truncation, such as control TE (SEQ ID NO:49).
[0298] Growth and production of fatty acid derivatives Increased solubility of the poorly soluble medium-chain length TE should result in a more active medium-chain length TE present in the cell, and thus an increase in the production of medium-chain fatty acids.
[0299] To evaluate the relative activity of the engineered TE shortened variants, each enzyme was cloned into a vector such that the gene was under the transcriptional control of the Ptrc promoter (Camsund et al., supra) which is induced in the presence of isopropylthiogalactoside (IPTG) as described in Example 4. These were transformed into an E. coli parental strain derived from MG1655 which has been engineered to overexpress the gene EntD from the chromosome and which harbors a Ptrc-controlled operon expressing the genes carB, alrA and aftA1 (described in Example 3) for the biochemical conversion of free fatty acids (FFA) to fatty alcohol acetate esters (FACE). The engineered E. coli parent was designed to adapt for high production of fatty acid derivatives to the expected high activity of these more soluble engineered TE variants.
[0300] The performance of each of the engineered shortened thioesterases was compared to a control evaluation strain expressing the control TE (SEQ ID NO:49). The only difference between the evaluation strains expressing the engineered shortened TE variants and the control evaluation strain expressing the control TE (SEQ ID NO:49) was the sequence of the gene encoding the expressed TE. Each of those strains was grown for the production of medium-chain fatty alcohol acetate esters and the resulting medium-chain fatty acid-derived products were extracted and quantified as described in Example 4.
[0301] The activity of each of the engineered shortened TE variants was evaluated by comparing the resulting fatty acid derivative products to the products produced by the control evaluation strain (expressing SEQ ID NO:49) as described in Example 4.
[0302] Table 7 lists the engineered shortened TE variants (SEQ ID NO:52 - SEQ ID NO:59) having improved performance in terms of (1) solubility (Figure 11) and (2) activity for the production of medium-chain fatty acid derivatives, with the performance reported as fold over control (FOC). Thus, the shortened variants are thioesterase variants having improved activity for the production of medium-chain fatty acid derivatives.
[0303] To further demonstrate the improved activity of engineered truncated TE variants with increased solubility, strains expressing SEQ ID NO:55 and SEQ ID NO:56 were grown in a 5 L bioreactor as described in Example 9 and compared to a control evaluation strain expressing TE SEQ ID NO:49 grown under the same conditions. Table 9 describes the performance of these strains reported as fold over control (FOC) at the 72 hour time point. Both SEQ ID NO:55 and SEQ ID NO:56 show improved activity (FAS FOC) and selectivity (%C8 FAS, and %C8 / %C10 FOC) at this larger scale.
[0304] (Table 9) Engineered truncated TE variants showing improved in vivo activity for the production of medium chain length fatty acid derivatives when grown in a 5 L bioreactor TIFF0007705243000017.tif39141
[0305] Example 8 The following examples illustrate processes that can be used to produce fatty acid derivatives using genetically modified microorganisms having improved activity for the production of medium chain fatty acid derivatives. The composition of fatty acid derivatives produced by this process includes, but is not limited to, medium chain fatty acids, medium chain fatty alcohols, medium chain fatty alcohol acetates (FACE), medium chain fatty acid methyl esters (FAME), medium chain fatty acid ethyl esters (FAEE), and further other medium chain fatty acid esters.
[0306] Generation of seed culture enlargement Using a frozen cell bank vial of the selected engineered E. coli strain, inoculate 20 mL of LB medium in a 125 mL baffled shake flask containing the appropriate antibiotic. Incubate this shake flask at 32 °C in an orbital shaker for approximately 6 hours, then transfer 1.25 mL of the medium (1% v / v) to 125 mL of minimal overnight seed medium (2 g / L NH4Cl, 0.5 g / L NaCl, 0.3 g / L KH2PO4, 1 mM MgSO4, 0.1 mM CaCl2, 20 g / L glucose, 1 mL / L of trace mineral solution (2 g / L of ZnCl2·4H2O, 2 g / L of CaCl2·6H2O, 2 g / L of Na2MoO4·2H2O, 1.9 g / L of CuSO4·5H2O, 0.5 g / L of H3BO3, and 10 mL / L of concentrated HCl), 10 mg / L of ferric citrate, 100 mM of Bis-Tris buffer (pH 7.0), and the appropriate antibiotic) in a 500 mL baffled Erlenmeyer shake flask and incubate overnight at 32 °C on a shaker.
[0307] Bioreactor culture protocol Using 75 mL (5% v / v) of the overnight seed culture described above, inoculation was first carried out into a 5 L Biostat Aplus bioreactor (Sartorius BBI) containing 1.5 L of sterile bioreactor fermentation medium. This medium consisted of 2 g / L KH2PO4, 0.5 g / L (NH4)2SO4, 2.2 g / L MgSO4 heptahydrate, 10 g / L of sterile filtered glucose, 80 mg / L ferric citrate, 1 mL / L of the trace mineral solution described above, 0.25 mL / L of vitamin solution (0.42 g / L riboflavin, 5.4 g / L pantothenic acid, 6 g / L niacin, 1.4 g / L pyridoxine, 0.06 g / L biotin, and 0.04 g / L folic acid), 1 g / L NaCl, 1 g / L citric acid, 140 mg / L CaCl2 dihydrate, 10 mg / L ZnCl2, and appropriate antibiotics. The pH of the culture was adjusted to between 6.9 and 7.2 using 28% w / v aqueous ammonia, the culture temperature was set to 33 - 35 °C depending on the specific product, the aeration rate was set to 0.75 lpm (0.5 v / v / m), and the dissolved oxygen partial pressure was maintained at 30% saturation using a DO controller and a stirrer loop connected to oxygen supply. The formation of foam was controlled by the automatic addition of a silicone emulsion-based antifoaming agent (Dow Corning 1430).
[0308] A nutrient supplement consisting of approximately 50% w / w glucose (600 g / L) was started when the glucose in the initial medium was completely depleted (approximately 7 hours after inoculation) and supplied at a rate of 10 g / l / h using a DO-stat controller strategy as needed (each supply shot had a duration of 1 hour). The gene involved in the production of medium-chain fatty acid derivatives was induced by adding isopropylthiogalactoside (IPTG) to a final concentration of 1 mM. The operation of the bioreactor was stopped after approximately 72 hours of fermentation time. Samples of the fermentation broth were taken throughout the fermentation process into the tank.
[0309] Analysis of broth composition Fatty acid derivatives present in samples of fermentation broth were extracted and separated in a single run using conventional GC-FID. For this purpose, 0.5 mL of each of the homogeneous fermentation broth samples was aliquoted into 15 mL Falcon tubes. The mass of the samples was recorded and 5.0 mL of butyl acetate with an internal standard (C11 FAME or C9 / C11 / C15 FALC) at 500 ppm was added to the broth to achieve 10-fold extraction. The samples were mechanically shaken at 2500 rpm for 30 minutes and centrifuged at 4500 rpm at 25 °C for 10 minutes. 50 μl of the extract (upper layer) was transferred to a GC vial, derivatized with 50 uL of BSTFA w / 10% TCMS, and subsequently vortexed for approximately 15 seconds. The samples were then run on a conventional GC-FID system using an Agilent DB1 column 10 m × 180 μm × 0.2 μm to separate all fatty acid derivatives present in the extracted samples. The concentration of each fatty acid derivative was reported in g / Kg.
[0310] Example 9 The following example illustrates a process that can be used to produce medium-chain fatty alcohol acetates using a genetically modified microorganism having improved activity for the production of medium-chain fatty acid derivatives. The composition of the fatty acid derivatives produced by this process can include medium-chain fatty acids, medium-chain fatty alcohols, and medium-chain fatty alcohol acetates (FACE) having acyl chains of 6 to 12 carbons. The production of medium-chain fatty alcohol acetates in a 5 L bioreactor was carried out as described in Example 8.
[0311] In this example, an Escherichia coli strain derived from MG1655, which has been engineered to overexpress the gene EntD from the chromosome and has a high production capacity for medium-chain fatty acid derivatives, was used. These strains contain an engineered thioesterase with improved activity for the production of medium-chain fatty acids, as well as an operon that expresses the genes carB, alrA, and aftA1 (described in Example 3) for the biochemical conversion of free fatty acids (FFA) to fatty alcohol acetate esters (FACE). The engineered thioesterase having SEQ ID NO:9 was expressed in the sRG.825 strain, while the engineered thioesterase having SEQ ID NO:49 was expressed in the sDH.377 strain. The genes encoding the engineered thioesterase, carB, alrA, and aftA1 were all under the transcriptional control of an inducible (Ptrc) promoter that was activated by adding isopropylthiogalactoside (IPTG) to the bioreactor at a fermentation time of about 24 hours. The operation of the bioreactor was stopped at a fermentation time of about 72 hours, the fermentation broth was collected and analyzed as in Example 8 above. The results are shown in Table 10 and Figure 12.
[0312] (Table 10) Total fatty acid species (FAS) concentration produced by representative strains engineered for the production of medium-chain fatty alcohol acetate esters or fatty acid alkyl esters TIFF0007705243000018.tif22132
[0313] Example 10 The following examples illustrate a process that can be used to produce medium-chain fatty alcohols using a genetically modified microorganism with improved ability to produce fatty acid derivatives. The composition of the fatty acid derivatives produced by this process can include fatty acids, fatty aldehydes, and fatty alcohols having an acyl chain of 6 to 12 carbons. The production of medium-chain fatty alcohols in a 5 L bioreactor was carried out using an Escherichia coli strain containing an operon that was engineered to overexpress the gene EntD from the chromosome and express the genes carB and alrA for the biochemical conversion of medium-chain thioesterase and free fatty acid (FFA) to fatty alcohol (FALC), as described in Example 8. The genes encoding thioesterase, carB, and alrA were all under the transcriptional control of an inducible (Ptrc) promoter that was activated by adding isopropylthiogalactoside (IPTG) to the bioreactor at an elapsed fermentation time of about 7 hours. Medium-chain fatty alcohols are very toxic to E. coli, and therefore, the accumulation of these compounds during production in a 5 L bioreactor stopped growth and production immediately after reaching the inhibitory concentration (less than 1 g / L, see Example 3).
[0314] Example 11 The following examples illustrate a process for producing medium-chain fatty acid alkyl esters using a genetically modified microorganism containing a thioesterase variant with improved activity for the production of medium-chain fatty acid derivatives. The composition of the fatty acid derivatives produced by this process can include medium-chain fatty acids, medium-chain fatty acid methyl esters (FAME) and / or medium-chain fatty acid ethyl esters (FAEE) having an acyl chain of 6 to 12 carbons.
[0315] This example illustrates the production of fatty acid ethyl esters (FAEE) using an E. coli strain (sAZ918) derived from MG1655 that was engineered to have a high production capacity for medium-chain fatty acid derivatives. This strain contains an operon that expresses an engineered thioesterase with improved activity for the production of medium-chain fatty acids (SEQ ID NO:49), as well as acyl-CoA synthetase and esterase for the biochemical conversion of free fatty acids (FFA) to fatty acid alkyl esters (FAME or FAEE) (described in Example 3). The engineered thioesterase, acyl-CoA synthetase, and esterase were all under the transcriptional control of an inducible (Ptrc) promoter that was activated by the addition of isopropylthiogalactoside (IPTG).
[0316] The production of medium-chain fatty acid ethyl esters in a 5 L bioreactor was carried out as described in Example 8, but with the addition of ethanol for nutrient supplementation. After inoculating the 5 L bioreactor with a seed culture, nutrient supplementation consisting of 47.5% w / w glucose and 50 mL / L ethanol was started when the glucose in the initial medium was completely depleted (approximately 7 hours after inoculation) and supplied at a rate of 10 g / L / h using a pH-stat controller strategy as required (each supply shot had a duration of 1 hour). The minimum agitation speed was fixed at 1200 rpm once this parameter value was achieved so that biofilm would not coat the dissolved oxygen probe and give a falsely low signal reading. Additional ethanol was added to the culture when the residual concentration dropped below 10 g / L. The ethyl octanoate production pathway of the strain was induced by adding IPTG to a final concentration of 1 mM after approximately 24 hours of fermentation time. The operation of the bioreactor was stopped after approximately 72 hours of fermentation time. The fermentation broth was collected and analyzed as described in Example 8 above.
[0317] The results are shown in Figure 13.
[0318] Example 12 The following examples illustrate a process for the production of medium-chain fatty acids using a genetically modified microorganism containing a thioesterase having improved activity for the production of medium-chain fatty acid derivatives. The composition of the fatty acids produced by this process includes medium-chain fatty acids having an acyl chain of 6 to 12 carbons. In this example, the production of medium-chain fatty acids in a 5 L bioreactor was carried out under the transcriptional control of an inducible (Ptrc) promoter activated by adding isopropylthiogalactoside (IPTG) to the bioreactor at an elapsed fermentation time of about 13 hours, as described in Example 8, using an Escherichia coli strain engineered to overexpress a medium-chain thioesterase. Medium-chain fatty acids are highly toxic to E. coli, and therefore, the accumulation of these compounds during production in a 5 L bioreactor stopped growth and production immediately after reaching the inhibitory concentration (less than 5 g / L, see Example 3).
[0319] Appendix A: Sequences TIFF0007705243000019.tif216157TIFF0007705243000020.tif216157TIFF0007705243000021.tif206157TIFF0007705243000022.tif215157TIFF0007705243000023.tif215157TIFF0007705243000024.tif218157TIFF0007705243000025.tif204157TIFF0007705243000026.tif215157TIFF0007705243000027.tif215157TIFF0007705243000028.tif218157TIFF0007705243000029.tif216157TIFF0007705243000030.tif206157
[0320] As will be apparent to those skilled in the art, various modifications and variations of the above aspects and embodiments can be made without departing from the spirit and scope of the present disclosure.
Claims
**Claim 1** An engineered thioesterase variant having an amino acid sequence with at least 90% sequence identity to SEQ ID NO:1, wherein The engineered thioesterase variant has a substitution mutation selected from the group consisting of P3K, D4M, S6R, T14G, T14R, V15L, V15W, V17A, V17C, P22R, D37P, T44G, V45S, V50W, S54R, S56C, S56K, T64P, T64R, T67L, L73V, L91M, C102I, V110L, I129V, G137C, R158Q, L176V, Y178P, P186G, D196V, D198W, K203R, Q213R, T217R, V225L, Q227G, G236T, T244M, T244R, S254G, A256C, E258T, E258V, S278K, V282S, L292F, A297T, A297V, I298C, I298V, V299L, N300L, N300W, A302T, I316R, T321R, and S322K, or the engineered thioesterase variant is SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41,having the amino acid sequence of SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58 or SEQ ID NO:59, the engineered thioesterase variant has improved catalytic activity, selectivity, and / or solubility for the production of medium-chain fatty acid derivatives compared to the enzyme having SEQ ID NO:
1. **Claim 2** The engineered thioesterase variant according to claim 1, having improved activity for the production of a C8 fatty acid derivative or a C10 fatty acid derivative. **Claim 3** (i) having an overall increased effective positive charge compared to the thioesterase having SEQ ID NO:1, or (ii) having an overall increased effective positive charge compared to the variant thioesterase having SEQ ID NO:4, or (iii) having the amino acid sequence of SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, or SEQ ID NO:46, or (iv) having an increased surface positive charge compared to SEQ ID NO:15, or (v) having the amino acid sequence of SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, or SEQ ID NO:51, or (vi) having improved solubility, or (vii) having improved solubility compared to SEQ ID NO:49, or (viii) having a truncated mutation in amino acids 2 - 40 of SEQ ID NO:49, or (ix) having the amino acid sequence of SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, or SEQ ID NO:59, or (x) the engineered thioesterase variant has improved activity for the production of C10 fatty acid derivatives, The engineered thioesterase variant according to claim 1.
4. A recombinant microbial cell comprising the engineered thioesterase variant according to any one of claims 1 to 3.
5. The recombinant microbial cell according to claim 4, further comprising one or more heterologous genes encoding a biochemical pathway for converting a first fatty acid derivative to a second fatty acid derivative, wherein the second fatty acid derivative has a higher minimum inhibitory concentration (MIC) than the first fatty acid derivative, the presence of the second fatty acid derivative increases the MIC of the first fatty acid derivative, and (i) the first fatty acid derivative is a fatty acid, the second fatty acid derivative is a fatty acid alkyl ester, a fatty acid methyl ester, or a fatty acid ethyl ester, and the biochemical pathway comprises an ester synthase or comprises an ester synthase and a fatty acid acyl-CoA synthase, or (ii) the first fatty acid derivative is a fatty alcohol, the second fatty acid derivative is a fatty alcohol acetate ester, and the biochemical pathway comprises a carboxylic acid reductase and an alcohol-O-acetyltransferase or comprises a carboxylic acid reductase, an alcohol dehydrogenase, and an alcohol-O-acetyltransferase, The recombinant microbial cell according to claim 4.
6. a. Carboxylic acid reductase, b. Carboxylic acid reductase and alcohol dehydrogenase, c. Carboxylic acid reductase and alcohol-O-acetyltransferase, d. Carboxylic acid reductase, alcohol dehydrogenase, and alcohol O-acetyltransferase, e. Ester synthase, f. Ester synthase and fatty acid acyl-CoA synthase, g. Acyl-CoA reductase, h. Acyl-CoA reductase and acyl-CoA synthase, i. Acyl-CoA reductase and alcohol O-acetyltransferase, j. Acyl-CoA reductase, alcohol O-acetyltransferase, and acyl-CoA synthetase, k. O-methyltransferase, l. acyl-ACP reductase, m. acyl-ACP reductase and aldehyde decarbonylase, n. acyl-ACP reductase and aldehyde oxidative deformylase, o. acyl-ACP reductase and alcohol O-acetyltransferase, p. acyl-ACP reductase, alcohol-O-acetyltransferase, and alcohol dehydrogenase, q. OleA protein, r. OleA, OleC, and OleD proteins, s. OleA protein and fatty acid acyl-CoA synthetase, or t. OleA, OleC, and OleD proteins and fatty acid acyl-CoA synthetase The recombinant microbial cell according to claim 4, further comprising.
7. A method for producing a medium-chain fatty acid derivative at a commercial titer, the method comprising culturing the recombinant microbial cell according to claim 5 or 6 under conditions suitable for the production of the medium-chain fatty acid derivative in the presence of a carbon source.
8. The method according to claim 7, wherein the recombinant microbial cell is derived from a genus including Escherichia, Bacillus, Lactobacillus, Zymomonas, Rhodococcus, Pseudomonas, Aspergillus, Trichoderma, Neurospora, Fusarium, Humicola, Rhizomucor, Kluyveromyces, Pichia, Mucor, Myceliophtora, Marinobacter, Penicillium, Phanerochaete, Pleurotus, Trametes, Chrysosporium, Saccharomyces, Stenotrophomonas, Schizosaccharomyces, Yarrowia, or Streptomyces.
9. The method according to claim 7, wherein the recombinant microbial cell is derived from a genus including Synechococcus or Synechocystis.
10. The method according to claim 7, further comprising the step of recovering the medium-chain fatty acid derivative from the culture medium or isolating the medium-chain fatty acid derivative from the recombinant microbial cell.
11. The method according to claim 7, wherein the biochemical pathway includes ester synthase and fatty acid acyl-CoA synthetase, and the recombinant microbial cell produces medium-chain fatty acid esters.
12. The method according to claim 11, wherein the medium-chain fatty acid ester is medium-chain fatty acid methyl ester and / or medium-chain fatty acid ethyl ester. **Claim 13**: An engineered thioesterase variant, wherein the engineered thioesterase variant has substitution mutations selected from the group consisting of T44I, I111T, Q114K, R132W, A162E, M165T, V185A, S197N, Q213H, S278T, A297D, N300K, and G301C, the engineered thioesterase variant has the amino acid sequence of SEQ ID NO:1 except for the said substitution mutations, and the engineered thioesterase variant has improved catalytic activity, selectivity and / or solubility for the production of medium-chain fatty acid derivatives as compared to the enzyme having SEQ ID NO:1.
Citation Information
Patent Citations
Oil and fat composition
JP2015091228A
Preservative for food product and preservation method of food product
JP2015123062A
Mutant thioesterase and method of use
JP2016508717A
Thioesterases and cells for producing modified oils
JP2016518112A
Acyl-ACP thioesterase genes and uses therefor
US20110020883A1