Operated nitroaldolase
Engineered nitroaldolase polypeptides from Baliospermum montanum catalyze the production of chiral β-nitroalcohols and tertiary aminoalcohols, addressing inefficiencies in current synthesis methods by enhancing yields and selectivities while minimizing toxic by-products.
Patent Information
- Application Number
- JP2024571179
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-08
- Filing Date
- 2023-06-06
- Publication Date
- 2025-06-26
AI Technical Summary
Current methods for producing chiral tertiary aminoalcohols are inefficient, requiring long synthetic routes, expensive substrates, and resulting in poor yields and selectivities, along with the generation of toxic by-products.
Engineered nitroaldolase polypeptides derived from Baliospermum montanum are used to catalyze the conversion of trifluoroacetone and nitromethane to β-nitroalcohol, a precursor for tertiary aminoalcohol production.
This approach enables the cost-effective and environmentally considerate synthesis of chiral β-nitroalcohols and tertiary aminoalcohols with improved yields and selectivities, reducing the formation of toxic by-products.
Smart Images

Figure 2025519396000001 
Figure 2025519396000002 
Figure 2025519396000003
Abstract
Description
Technical Field
[0001] This application claims priority based on U.S. Provisional Patent Application No. 63 / 350,128, filed on January 8, 2022, the entire disclosure of which is hereby incorporated by reference herein for all purposes.
[0002] Technical Field The present invention provides engineered nitroaldolase polypeptides useful for the production of β-nitroalcohols and tertiary aminoalcohols, as well as compositions and methods that utilize these engineered polypeptides.
[0003] Reference to a Sequence Listing, Table, or Computer Program A formal copy of the Sequence Listing is submitted herewith as an XML file named "CX2-230WO1 ST26.xml", created on June 1, 2023, and having a size of 2,633,728 bytes. This Sequence Listing, which is part of this application, is hereby incorporated by reference herein in its entirety.
Background Art
[0004] Background Chiral tertiary aminoalcohols are important components in the pharmaceutical industry. In particular, chiral tertiary aminoalcohols are intermediates of compounds used in the treatment of cystic fibrosis and chronic obstructive pulmonary disease (COPD). However, traditional chemical synthesis methods of chiral tertiary aminoalcohols may have drawbacks such as requiring long synthetic routes, using expensive substrates, resulting in poor yields and selectivities, and generating toxic by-products.
[0005] Therefore, more cost-effective and environmentally considerate biocatalytic synthesis routes are of concern for efficiently producing these compounds. However, only a few biocatalytic routes for producing chiral tertiary aminoalcohol compounds are known.
[0006] One approach may involve the production of chiral β-amino alcohols via β-hydroxy-α-amino acid intermediates using a two-enzyme cascade (Duckers et al. Appl Microbiol Biotechnol, 2010, 88:409-424). However, in all reported literature examples, this route is limited to the production of chiral secondary alcohols and not the desired chiral tertiary alcohols.
[0007] Another possible route is the enantioselective addition of nitromethane to trifluoroacetone, catalyzed by an enzyme to form a chiral tertiary nitroaldol product, followed by chemical reduction to the chiral amino alcohol. This is not a known natural reaction. However, there are quite a few literature precedents for hydroxynitrile lyases and other enzymes that catalyze this reaction.
[0008] These include the (R)-selective nitroaldolase reaction catalyzed by hydroxynitrile lyases from Acidobacterium capsulatum and Granulicella tundricula (M. Bekerle-Bogner, M. Gruber-Khadjawi, H. Wiltsche, R. Wiedner, H. Schwab, K. Steiner, ChemCatChem 2016, 8, 2214). The Henry (nitroaldol) reaction using hydroxynitrile lyase from Hevea brasiliensis (Gruber-Khadjawi, M., Purkarthofer, T., Skranc, W. and Griengl, H. (2007) Adv. Synth. Catal., 349: 1445-1450) and transglutaminase from Streptorerticillium griseoverticillatum (Tang R., Guan Z., He Y., Zhu W., J of Molec Catal B: Enzymatic, 63: 1-2, 2010, 62-67) has also been reported. Other examples of reactions catalyzed by similar enzymes are known in the literature. However, there is currently no report demonstrating the activity of hydroxynitrile lyase / nitroaldolase towards a ketone substrate, nor is there a report demonstrating the highly (S)-selective activity of hydroxynitrile lyase as a nitroaldolase. A nitroaldolase having (S)-selective activity towards a ketone would enable the synthesis of various useful chiral amino alcohol compounds.
Prior Art Documents
Non-Patent Documents
[0009]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Summary of the Invention
Means for Solving the Problems
[0010] Abstract The present invention provides novel biocatalysts and related methods useful for the synthesis of β-nitroalcohols and tertiary aminoalcohols. The nitroaldolase biocatalysts of the present disclosure are engineered variants of a polypeptide (SEQ ID NO: 2) encoded by a homologous gene (SEQ ID NO: 1) from Baliospermum montanum. These engineered polypeptides can catalyze the conversion of trifluoroacetone and nitromethane to β-nitroalcohol, which is a precursor of the tertiary aminoalcohol product.
[0011] The present invention includes amino acid sequences having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 2, and 7 / 9, 7 / 9 / 33 / 68 / 72 / 158, 7 / 9 / 33 / 72 / 175 / 197 / 230, 7 / 9 / 33 / 72 / 230, 7 / 9 / 68 / 72 / 197, 7 / 9 / 158 / 175, 7 / 9 / 175, 7 / 9 / 197, 7 / 33, 7 / 68 / 72 / 175, 7 / 72, 7 / 72 / 97 / 175 / 230, 9, 9 / 11 / 33 / 72 / 197 / 230, 9 / 33 / 68 / 158 / 197, 9 / 33 / 72 / 230, 9 / 33 / 97 / 230, 9 / 68 / 72, 9 / 97 / 175 / 202 / 218, 9 / 175, 9 / 202, 11, 13, 14 / 33 / 68 / 72 / 97, 18, 18 / 33 / 68 / 97 / 158 / 175 / 239, 18 / 33 / 72 / 97 / 158 / 175, 18 / 33 / 72 / 97 / 158 / 239, 18 / 33 / 72 / 97 / 175 / 202, 18 / 33 / 72 / 175 / 202 / 237 / 239, 18 / 33 / 72 / 197 / 237, 18 / 68 / 72, 18 / 68 / 72 / 158 / 175 / 237 / 239, 18 / 68 / 72 / 175 / 202 / 239, 18 / 68 / 72 / 239, 18 / 68 / 175 / 223 / 237 / 239, 18 / 68 / 239, 18 / 72 / 97 / 175 / 237 / 239, 18 / 72 / 175 / 197, 18 / 72 / 237, 18 / 72 / 239, 18 / 97 / 175 / 237 / 239, 18 / 97 / 223 / 237 / 239, 18 / 175, 18 / 197 / 239, 18 / 239, 20 / 124, 33 / 68 / 72, 33 / 68 / 72 / 175 / 230, 33 / 68 / 223 / 239, 33 / 72, 33 / 72 / 97 / 197, 33 / 72 / 175, 33 / 175, 39, 44 / 97, 52, 68 / 72, 68 / 72 / 97, 68 / 72 / 97 / 197, 68 / 72 / 197, 68 / 72 / 197 / 239, 68 / 97 / 175, 68 / 97 / 175 / 197 / 223, 72 / 97 / 158 / 223 / 237, 72 / 175 / 197, 97 / 158 / 175 / 197, 97 / 175, 103, 121, 122, 124, 125, 128, 133, 148, 175, 197,Provided is an engineered polypeptide comprising at least one substitution or a set of substitutions at one or more positions selected from 197 / 202, wherein the positions are numbered with reference to SEQ ID NO:2. In some additional embodiments, the engineered polypeptide is SEQ ID NO:7V / 9V, 7V / 9V / 33V / 68L / 72E / 158Y, 7V / 9V / 33V / 72E / 175P / 197V / 230I, 7V / 9V / 33V / 72E / 230I, 7V / 9V / 68L / 72E / 197V, 7V / 9V / 158Y / 175P, 7V / 9V / 175P, 7V / 9V / 197V, 7V / 33V, 7V / 68L / 72E / 175P, 7V / 72E, 7V / 72E / 97I / 175P / 230I, 9V, 9V / 11G / 33V / 72E / 197V / 230I, 9V / 33V / 68L / 158Y / 197V, 9V / 33V / 72E / 230I, 9V / 33V / 97I / 230I, 9V / 68L / 72E, 9V / 97I / 175P / 202I / 218M, 9V / 175P, 9V / 202I, 11S, 13G, 14Y / 33V / 68L / 72E / 97I, 18C / 33V / 68L / 97I / 158Y / 175P / 239F, 18C / 33V / 72E / 97I / 158Y / 175P, 18C / 33V / 72E / 97I / 158Y / 239F, 18C / 33V / 72E / 97I / 175P / 202I, 18C / 33V / 72E / 175P / 202I / 237P / 239F, 18C / 33V / 72E / 197V / 237P, 18C / 68L / 72E, 18C / 68L / 72E / 158Y / 175P / 237P / 239F, 18C / 68L / 72E / 175P / 202I / 239F, 18C / 68L / 72E / 239F, 18C / 68L / 175P / 223P / 237P / 239F, 18C / 68L / 239F, 18C / 72E / 97I / 175P / 237P / 239F, 18C / 72E / 175P / 197V, 18C / 72E / 237P, 18C / 72E / 239F, 18C / 97I / 175P / 237P / 239F, 18C / 97I / 223P / 237P / 239F, 18C / 175P, 18C / 197V / 239F, 18C / 239F, 18I, 20H / 124C, 33V / 68L / 72E, 33V / 68L / 72E / 175P / 230I, 33V / 68L / 223P / 239F, 33V / 72E, 33V / 72E / 97I / 197V,Comprising at least one substitution or set of substitutions selected from 33V / 72E / 175P, 33V / 175P, 39A, 39G, 44N / 97I, 52S, 68L / 72E, 68L / 72E / 97I, 68L / 72E / 97I / 197V, 68L / 72E / 197V, 68L / 72E / 197V / 239F, 68L / 97I / 175P, 68L / 97I / 175P / 197V / 223P, 72E / 97I / 158Y / 223P / 237P, 72E / 175P / 197V, 97I / 158Y / 175P / 197V, 97I / 175P, 103M, 103S, 121Y, 122C, 122Q, 124C, 124E, 124T, 124Y, 125R, 128H, 128Y, 133A, 133C, 148M, 175P, 197V, and 197V / 202I, wherein the positions are numbered with reference to SEQ ID NO: 2. In some embodiments, the engineered polypeptide is SEQ ID NO: I7V / I9V, I7V / I9V / A33V / I68L / K72E / F158Y, I7V / I9V / A33V / K72E / K175P / I197V / V230I, I7V / I9V / A33V / K72E / V230I, I7V / I9V / I68L / K72E / I197V, I7V / I9V / F158Y / K175P, I7V / I9V / K175P, I7V / I9V / I197V, I7V / A33V, I7V / I68L / K72E / K175P, I7V / K72E, I7V / K72E / V97I / K175P / V230I, I9V, I9V / T11G / A33V / K72E / I197V / V230I, I9V / A33V / I68L / F158Y / I197V, I9V / A33V / K72E / V230I, I9V / A33V / V97I / V230I, I9V / I68L / K72E, I9V / V97I / K175P / V202I / Q218M, I9V / K175P, I9V / V202I, T11S, C13G, H14Y / A33V / I68L / K72E / V97I, L18C / A33V / I68L / V97I / F158Y / K175P / I239F, L18C / A33V / K72E / V97I / F158Y / K175P, L18C / A33V / K72E / V97I / F158Y / I239F, L18C / A33V / K72E / V97I / K175P / V202I, L18C / A33V / K72E / K175P / V202I / I237P / I239F,At least one substitution or set of substitutions selected from L18C / A33V / K72E / I197V / I237P, L18C / I68L / K72E, L18C / I68L / K72E / F158Y / K175P / I237P / I239F, L18C / I68L / K72E / K175P / V202I / I239F, L18C / I68L / K72E / I239F, L18C / I68L / K175P / K223P / I237P / I239F, L18C / I68L / I239F, L18C / K72E / V97I / K175P / I237P / I239F, L18C / K72E / K175P / I197V, L18C / K72E / I237P, L18C / K72E / I239F, L18C / V97I / K175P / I237P / I239F, L18C / V97I / K223P / I237P / I239F, L18C / K175P, L18C / I197V / I239F, L18C / I239F, L18I, Y20H / V124C, A33V / I68L / K72E, A33V / I68L / K72E / K175P / V230I, A33V / I68L / K223P / I239F, A33V / K72E, A33V / K72E / V97I / I197V, A33V / K72E / K175P, A33V / K175P, V39A, V39G, D44N / V97I, G52S, I68L / K72E, I68L / K72E / V97I, I68L / K72E / V97I / I197V, I68L / K72E / I197V, I68L / K72E / I197V / I239F, I68L / V97I / K175P, I68L / V97I / K175P / I197V / K223P, K72E / V97I / F158Y / K223P / I237P, K72E / K175P / I197V, V97I / F158Y / K175P / I197V, V97I / K175P, H103M, H103S, F121Y, S122C, S122Q, V124C, V124E, V124T, V124Y, F125R, W128H, W128Y, F133A, F133C, L148M, K175P, I197V, and I197V / V202I, where the positions are numbered with reference to SEQ ID NO:2. In some further embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs:4 to 176.,
[0012] The present invention includes an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 2, and also provides an engineered polypeptide comprising at least at least one substitution or a set of substitutions selected from 7 / 11 / 33 / 72 / 97, 7 / 11 / 33 / 97, 7 / 11 / 72, 7 / 11 / 158, 7 / 11 / 237, 7 / 33 / 72 / 97 / 158 / 175 / 197 / 237 / 239, 9 / 11 / 33, 9 / 11 / 33 / 72, 9 / 11 / 72, 9 / 33 / 68 / 97 / 237 / 239, 9 / 72 / 97 / 158 / 237 / 239, 9 / 218 / 239, 10, 11, 12, 13, 14, 18, 33 / 68 / 72 / 97 / 237 / 239, 33 / 72 / 158 / 239, 39, 40, 41, 48, 49, 51, 54, 54 / 126, 57, 68 / 175 / 239, 82, 84, 85, 86, 103, 104, 105, 107, 115, 117, 119, 119 / 120, 120, 122, 123, 124, 127, 130, 131, 133, 134, 137 / 145, 144, 145, 146, 147, 148, 149, 150, 152, 157, 158, 159, 169, 173, 174, 175, 180, 181, 183, 203, 204, 208, 209, 210, 215, 218, 236, and 237, where the positions are numbered with reference to SEQ ID NO: 2. In some embodiments, the engineered polypeptide is SEQ ID NO: 7V / 11G / 33V / 72E / 97I, 7V / 11G / 33V / 97I, 7V / 11G / 72E, 7V / 11G / 158Y, 7V / 11G / 237P, 7V / 33V / 72E / 97I / 158Y / 175P / 197V / 237P / 239F, 9V / 11G / 33V, 9V / 11G / 33V / 72E, 9V / 11G / 72E, 9V / 33V / 68L / 97I / 237P / 239F, 9V / 72E / 97I / 158Y / 237P / 239F, 9V / 218M / 239F, 10G, 10P, 10R, 11R, 11S, 12K, 13L, 14R, 18S, 18T, 33V / 68L / 72E / 97I / 237P / 239F, 33V / 72E / 158Y / 239F, 39R, 39W, 40L, 40S, 40T, 40V, 41C, 48W, 49Q, 49R, 51A, 51Q, 51S,Comprising at least one substitution or a set of substitutions selected from 51T, 54N / 126I, 54Y, 57N, 68L / 175P / 239F, 82T, 84L, 85A, 85L, 85M, 85Q, 86W, 103P, 104A, 104G, 104S, 105D, 107G, 107W, 115C, 115K, 115L, 115V, 117C, 117L, 119L, 119N / 120F, 120C, 120L, 120Y, 122G, 122T, 123A, 123P, 123T, 124P, 127A, 127L, 127M, 130M, 130V, 131L, 133W, 134D, 134F, 134L, 134R, 134V, 134W, 137I / 145V, 144L, 145F, 145L, 145P, 146D, 146G, 146N, 147R, 148P, 148V, 149W, 150I, 150L, 150V, 152P, 157V, 158I, 158L, 158M, 158R, 158W, 158Y, 159G, 159H, 159L, 169G, 169I, 169S, 173E, 173Y, 174C, 174G, 175D, 180G, 180V, 181C, 183K, 203Q, 204E, 204R, 204V, 208A, 208C, 208E, 208G, 208L, 208T, 208V, 209C, 210V, 215L, 218T, 236G, 236L, 237L, 237S, and 237W, wherein the positions are numbered with reference to SEQ ID NO: 2. In some further embodiments, the engineered polypeptide is SEQ ID NO: I7V / T11G / A33V / K72E / V97I, I7V / T11G / A33V / V97I, I7V / T11G / K72E, I7V / T11G / F158Y, I7V / T11G / I237P, I7V / A33V / K72E / V97I / F158Y / K175P / I197V / I237P / I239F, I9V / T11G / A33V, I9V / T11G / A33V / K72E, I9V / T11G / K72E, I9V / A33V / I68L / V97I / I237P / I239F, I9V / K72E / V97I / F158Y / I237P / I239F, I9V / Q218M / I239F, H10G, H10P, H10R, T11R, T11S, I12K, C13L, H14R, L18S, L18T, A33V / I68L / K72E / V97I / I237P / I239F,At least one substitution or set of substitutions selected from A33V / K72E / F158Y / I239F, V39R, V39W, A40L, A40S, A40T, A40V, S41C, L48W, E49Q, E49R, I51A, I51Q, I51S, I51T, W54N / T126I, W54Y, Y57N, I68L / K175P / I239F, G82T, I84L, N85A, N85L, N85M, N85Q, I86W, H103P, N104A, N104G, N104S, A105D, M107G, M107W, A115C, A115K, A115L, A115V, V117C, V117L, K119L, K119N / K120F, K120C, K120L, K120Y, S122G, S122T, E123A, E123P, E123T, V124P, D127A, D127L, D127M, D130M, D130V, S131L, F133W, S134D, S134F, S134L, S134R, S134V, S134W, T137I / A145V, T144L, A145F, A145L, A145P, V146D, V146G, V146N, E147R, L148P, L148V, G149W, D150I, D150L, D150V, T152P, I157V, F158I, F158L, F158M, F158R, F158W, F158Y, S159G, S159H, S159L, A169G, A169I, A169S, V173E, V173Y, R174C, R174G, K175D, E180G, E180V, Q181C, L183K, Y203Q, G204E, G204R, G204V, Q208A, Q208C, Q208E, Q208G, Q208L, Q208T, Q208V, I209C, F210V, Q215L, Q218T, K236G, K236L, I237L, I237S, and I237W, where the positions are numbered with reference to SEQ ID NO:2. In some further embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 178 to 462.,
[0013] The present invention includes an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 24, and further provides an engineered polypeptide comprising at least at least one substitution or a set of substitutions selected from 9 / 39 / 122 / 157, 18, 18 / 33 / 39 / 230, 18 / 33 / 175, 18 / 39 / 44 / 118, 18 / 175, 20, 33 / 39, 33 / 39 / 44 / 118 / 121, 33 / 39 / 44 / 121, 33 / 39 / 44 / 152, 33 / 39 / 44 / 175, 33 / 39 / 44 / 197, 33 / 39 / 44 / 210 / 230, 33 / 39 / 118 / 121 / 175, 33 / 39 / 121 / 175, 33 / 39 / 197 / 210, 33 / 39 / 210, 33 / 39 / 210 / 230, 33 / 118 / 121, 33 / 118 / 121 / 175, 39, 39 / 44 / 103, 39 / 44 / 118, 39 / 44 / 152 / 230, 39 / 44 / 175 / 197 / 230, 39 / 103, 39 / 103 / 125, 39 / 103 / 125 / 127 / 146, 39 / 103 / 125 / 127 / 146 / 150, 39 / 103 / 125 / 146, 39 / 103 / 127 / 150, 39 / 103 / 150, 39 / 118 / 121, 39 / 197, 39 / 197 / 210, 43, 44 / 103, 81, 103, 103 / 125 / 127, 103 / 125 / 146, 103 / 127, 103 / 150, 118 / 121, 118 / 121 / 175, 118 / 121 / 197, 121, 152, 152 / 197, 154, 159, 163, 175, 210, 232, and 238, wherein the positions are numbered with reference to SEQ ID NO: 24. In some additional embodiments, the engineered polypeptide is SEQ ID NO: 9V / 39W / 122T / 157V, 18I, 18I / 33V / 39W / 230I, 18I / 33V / 175P, 18I / 39W / 44N / 118L, 18I / 175P, 20H, 33V / 39W, 33V / 39W / 44N / 118L / 121Y,At least one substitution or set of substitutions selected from 33V / 39W / 121Y / 175P, 33V / 39W / 197V / 210M, 33V / 39W / 210M, 33V / 39W / 210M / 230I, 33V / 118L / 121W / 175P, 33V / 118L / 121Y, 39W, 39W / 44N / 118L, 39W / 44N / 152F / 230I, 39W / 44N / 175P / 197V / 230I, 39W / 44Y / 103S, 39W / 103M, 39W / 103M / 125R / 146L, 39W / 103M / 150P, 39W / 103Q, 39W / 103S, 39W / 103S / 125P / 127G / 146L / 150P, 39W / 103S / 125R, 39W / 103S / 125R / 127G / 146L, 39W / 103S / 127G / 150P, 39W / 118L / 121Y, 39W / 197V, 39W / 197V / 210M, 43M, 43Q, 43R, 43S, 43T, 44N / 103S, 81F, 103L, 103M, 103M / 125P / 146L, 103M / 127G, 103Q, 103S, 103S / 125P / 127G, 103S / 150P, 103V, 118L / 121Y, 118L / 121Y / 175P, 118L / 121Y / 197V, 121Y, 152A / 197V, 152F, 154I, 159Q, 163P, 175P, 210M, 232G, and 238M, where the positions are numbered with reference to SEQ ID NO: 24. In some further embodiments, the engineered polypeptide is SEQ ID NO: I9V / V39W / S122T / I157V, L18I, L18I / A33V / V39W / V230I, L18I / A33V / K175P, L18I / V39W / D44N / Y118L, L18I / K175P, Y20H, A33V / V39W, A33V / V39W / D44N / Y118L / F121Y, A33V / V39W / D44N / F121Y, A33V / V39W / D44N / T152F, A33V / V39W / D44N / K175P, A33V / V39W / D44N / I197V, A33V / V39W / D44N / F210M / V230I, A33V / V39W / Y118L / F121Y / K175P, A33V / V39W / F121Y / K175P, A33V / V39W / I197V / F210M, A33V / V39W / F210M,Comprising at least at least one substitution or set of substitutions selected from A33V / V39W / F210M / V230I, A33V / Y118L / F121W / K175P, A33V / Y118L / F121Y, V39W, V39W / D44N / Y118L, V39W / D44N / T152F / V230I, V39W / D44N / K175P / I197V / V230I, V39W / D44Y / H103S, V39W / H103M, V39W / H103M / F125R / V146L, V39W / H103M / D150P, V39W / H103Q, V39W / H103S, V39W / H103S / F125P / D127G / V146L / D150P, V39W / H103S / F125R, V39W / H103S / F125R / D127G / V146L, V39W / H103S / D127G / D150P, V39W / Y118L / F121Y, V39W / I197V, V39W / I197V / F210M, I43M, I43Q, I43R, I43S, I43T, D44N / H103S, G81F, H103L, H103M, H103M / F125P / V146L, H103M / D127G, H103Q, H103S, H103S / F125P / D127G, H103S / D150P, H103V, Y118L / F121Y, Y118L / F121Y / K175P, Y118L / F121Y / I197V, F121Y, T152A / I197V, T152F, A154I, S159Q, I163P, K175P, F210M, S232G, and Q238M, wherein the positions are numbered with reference to SEQ ID NO: 24. In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 464 to 596.,
[0014] The present invention includes an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 24, and 7, 9 / 13, 9 / 13 / 39, 9 / 13 / 39 / 68 / 72 / 239, 9 / 13 / 39 / 68 / 239, 9 / 13 / 39 / 72 / 122, 9 / 13 / 39 / 97 / 239, 9 / 13 / 68 / 72 / 97 / 239, 9 / 13 / 68 / 97 / 134 / 239, 9 / 13 / 68 / 97 / 239, 9 / 13 / 72, 9 / 13 / 134, 11, 12 / 42 / 125 / 127, 12 / 103, 12 / 103 / 150, 13, 13 / 39, 13 / 39 / 68, 13 / 39 / 68 / 72 / 97, 13 / 39 / 72 / 97 / 122 / 134 / 223, 13 / 39 / 97 / 239, 13 / 39 / 239, 13 / 68, 13 / 68 / 72 / 97 / 122 / 239, 16, 18 / 33 / 39 / 44 / 118 / 121 / 175 / 197 / 210, 18 / 33 / 39 / 44 / 121, 18 / 33 / 39 / 112 / 118 / 121 / 197, 18 / 33 / 39 / 121, 18 / 33 / 39 / 121 / 175, 18 / 33 / 39 / 121 / 175 / 197, 18 / 33 / 39 / 121 / 175 / 210, 18 / 33 / 39 / 121 / 210, 18 / 33 / 118 / 121, 18 / 33 / 121, 18 / 33 / 121 / 197, 18 / 33 / 121 / 210 / 230, 18 / 39 / 44 / 118 / 121 / 230, 18 / 39 / 44 / 121 / 175 / 210, 18 / 39 / 118 / 121, 18 / 39 / 118 / 121 / 210, 18 / 39 / 121 / 175, 18 / 39 / 121 / 175 / 210, 18 / 39 / 121 / 197 / 210, 18 / 118 / 121, 18 / 118 / 121 / 175, 18 / 118 / 121 / 210, 18 / 118 / 121 / 230, 18 / 121, 18 / 121 / 175, 18 / 121 / 210, 18 / 121 / 230, 20, 21, 33 / 39 / 44 / 118 / 121, 33 / 39 / 44 / 118 / 121 / 210 / 230, 33 / 39 / 44 / 121, 33 / 39 / 44 / 121 / 152, 33 / 39 / 44 / 121 / 175 / 210, 33 / 39 / 44 / 121 / 230, 33 / 39 / 73 / 121 / 197, 33 / 39 / 118 / 121 / 175, 33 / 39 / 118 / 121 / 175 / 230,At least one substitution or set of substitutions selected from 33 / 39 / 121, 33 / 39 / 121 / 175, 33 / 39 / 121 / 210 / 230, 33 / 118 / 121, 33 / 118 / 121 / 175, 33 / 118 / 121 / 197, 33 / 121, 33 / 121 / 230, 37 / 91, 39 / 44 / 118 / 121 / 175 / 210, 39 / 68 / 72 / 239, 39 / 118 / 121, 39 / 118 / 121 / 210, 39 / 121, 39 / 121 / 210 / 230, 43, 44, 76, 79, 101, 118 / 121, 118 / 121 / 175, 118 / 121 / 197, 121, 121 / 175 / 230, 121 / 210, 148, 153, 159, 161, 172, 200, 207, 233, 245, and 248 are further provided, where the positions are numbered with reference to SEQ ID NO: 24. In some additional embodiments, the engineered polypeptide is SEQ ID NO: 7E, 9V / 13G, 9V / 13G / 39W, 9V / 13G / 39W / 68L / 72E / 239F, 9V / 13G / 39W / 68L / 239F, 9V / 13G / 39W / 72E / 122T, 9V / 13G / 39W / 97I / 239F, 9V / 13G / 68L / 72E / 97I / 239F, 9V / 13G / 68L / 97I / 134D / 239F, 9V / 13G / 68L / 97I / 239F, 9V / 13G / 72E, 9V / 13G / 134D, 11N, 12A / 42W / 125R / 127G, 12A / 103S, 12A / 103S / 150P, 13F, 13G / 39W, 13G / 39W / 68L, 13G / 39W / 68L / 72E / 97I, 13G / 39W / 72E / 97I / 122T / 134D / 223P, 13G / 39W / 97I / 239F, 13G / 39W / 239F, 13G / 68L, 13G / 68L / 72E / 97I / 122T / 239F, 16M, 16V, 18C / 33V / 39W / 112R / 118L / 121W / 197V, 18C / 33V / 39W / 121W / 175P / 210M, 18C / 33V / 39W / 121W / 210M, 18C / 33V / 39W / 121Y / 175P, 18C / 33V / 118L / 121W, 18C / 33V / 118L / 121Y, 18C / 33V / 121Y, 18C / 33V / 121Y / 197V, 18C / 33V / 121Y / 210M / 230I,18C / 39W / 44N / 121Y / 175P / 210M, 18C / 39W / 118L / 121W, 18C / 39W / 121W / 175P / 210M, 18C / 39W / 121W / 197V / 210M, 18C / 118L / 121W, 18C / 118L / 121Y, 18C / 121V, 18C / 121V / 175P, 18C / 121W, 18C / 121W / 175P, 18C / 121W / 210M, 18C / 121Y, 18C / 121Y / 230I, 18I / 33V / 39W / 44N / 118L / 121W / 175P / 197V / 210M, 18I / 33V / 39W / 44N / 121T, 18I / 33V / 39W / 121T / 175P, 18I / 33V / 39W / 121V / 175P / 197V, 18I / 33V / 39W / 121W, 18I / 33V / 118L / 121W, 18I / 39W / 44N / 118L / 121W / 230I, 18I / 39W / 44N / 121V / 175P / 210M, 18I / 39W / 118L / 121W / 210M, 18I / 39W / 121V / 175P, 18I / 118L / 121T, 18I / 118L / 121W, 18I / 118L / 121W / 175P, 18I / 118L / 121W / 210M, 18I / 118L / 121Y, 18I / 118L / 121Y / 230I, 18I / 121T / 175P, 18I / 121W, 18I / 121Y / 210M, 20G, 20K, 20R, 20T, 21V, 33V / 39G / 44N / 121W / 152A, 33V / 39W / 44N / 118L / 121W / 210M / 230I, 33V / 39W / 44N / 118L / 121Y, 33V / 39W / 44N / 121V, 33V / 39W / 44N / 121Y, 33V / 39W / 44N / 121Y / 175P / 210M, 33V / 39W / 44N / 121Y / 230I, 33V / 39W / 73N / 121Y / 197V, 33V / 39W / 118L / 121W / 175P, 33V / 39W / 118L / 121Y / 175P, 33V / 39W / 118L / 121Y / 175P / 230I, 33V / 39W / 121V, 33V / 39W / 121W, 33V / 39W / 121W / 210M / 230I, 33V / 39W / 121Y / 175P, 33V / 118L / 121W / 175P, 33V / 118L / 121W / 197V, 33V / 118L / 121Y, 33V / 121W, 33V / 121Y / 230IComprising at least one substitution or a set of substitutions selected from 37L / 91G, 39G / 68L / 72E / 239F, 39W / 44N / 118L / 121W / 175P / 210M, 39W / 118L / 121W / 210M, 39W / 118L / 121Y, 39W / 118L / 121Y / 210M, 39W / 121W, 39W / 121W / 210M / 230I, 43E, 43Q, 43S, 43W, 44V, 76P, 79F, 101L, 118L / 121W, 118L / 121Y, 118L / 121Y / 175P, 118L / 121Y / 197V, 121V, 121W, 121Y, 121Y / 175P / 230I, 121Y / 210M, 148F, 153W, 159Q, 161M, 172A, 172E, 172G, 172Q, 172R, 172V, 200G, 207N, 233R, 245D, 245G, 248G, and 248S, wherein the positions are numbered with reference to SEQ ID NO: 24. In some further embodiments, the engineered polypeptide is SEQ ID NO: I7E, I9V / C13G, I9V / C13G / V39W, I9V / C13G / V39W / I68L / K72E / I239F, I9V / C13G / V39W / I68L / I239F, I9V / C13G / V39W / K72E / S122T, I9V / C13G / V39W / V97I / I239F, I9V / C13G / I68L / K72E / V97I / I239F, I9V / C13G / I68L / V97I / S134D / I239F, I9V / C13G / I68L / V97I / I239F, I9V / C13G / K72E, I9V / C13G / S134D, S11N, I12A / G42W / F125R / D127G, I12A / H103S, I12A / H103S / D150P, C13F, C13G / V39W, C13G / V39W / I68L, C13G / V39W / I68L / K72E / V97I, C13G / V39W / K72E / V97I / S122T / S134D / K223P, C13G / V39W / V97I / I239F, C13G / V39W / I239F, C13G / I68L, C13G / I68L / K72E / V97I / S122T / I239F, A16M, A16V, L18C / A33V / V39W / H112R / Y118L / F121W / I197V, L18C / A33V / V39W / F121W / K175P / F210M,L18C / A33V / V39W / F121W / F210M, L18C / A33V / V39W / F121Y / K175P, L18C / A33V / Y118L / F121W, L18C / A33V / Y118L / F121Y, L18C / A33V / F121Y, L18C / A33V / F121Y / I197V, L18C / A33V / F121Y / F210M / V230I, L18C / V39W / D44N / F121Y / K175P / F210M, L18C / V39W / Y118L / F121W, L18C / V39W / F121W / K175P / F210M, L18C / V39W / F121W / I197V / F210M, L18C / Y118L / F121W, L18C / Y118L / F121Y, L18C / F121V, L18C / F121V / K175P, L18C / F121W, L18C / F121W / K175P, L18C / F121W / F210M, L18C / F121Y, L18C / F121Y / V230I, L18I / A33V / V39W / D44N / Y118L / F121W / K175P / I197V / F210M, L18I / A33V / V39W / D44N / F121T, L18I / A33V / V39W / F121T / K175P, L18I / A33V / V39W / F121V / K175P / I197V, L18I / A33V / V39W / F121W, L18I / A33V / Y118L / F121W, L18I / V39W / D44N / Y118L / F121W / V230I, L18I / V39W / D44N / F121V / K175P / F210M, L18I / V39W / Y118L / F121W / F210M, L18I / V39W / F121V / K175P, L18I / Y118L / F121T, L18I / Y118L / F121W, L18I / Y118L / F121W / K175P, L18I / Y118L / F121W / F210M, L18I / Y118L / F121Y, L18I / Y118L / F121Y / V230I, L18I / F121T / K175P, L18I / F121W, L18I / F121Y / F210M, Y20G, Y20K, Y20R, Y20T, K21V, A33V / V39G / D44N / F121W / T152A, A33V / V39W / D44N / Y118L / F121W / F210M / V230I, A33V / V39W / D44N / Y118L / F121YA33V / V39W / D44N / F121V, A33V / V39W / D44N / F121Y, A33V / V39W / D44N / F121Y / K175P / F210M, A33V / V39W / D44N / F121Y / V230I, A33V / V39W / K73N / F121Y / I197V, A33V / V39W / Y118L / F121W / , K175P, A33V / V39W / Y118L / F121Y / K175P, A33V / V39W / Y118L / F121Y / K175P / V230I, A33V / V39W / F121V, A33V / V39W / F121W, A33V / V39W / F121W / F210M / V230I, A33V / V39W / F121Y / K175P, A33V / Y118L / F121W / K175P, A33V / Y118L / F121W / I197V, A33V / Y118L / F121Y, A33V / F121W, A33V / F121Y / V230I, D37L / E91G, V39G / I68L / K72E / I239F, V39W / D44N / Y118L / F121W / K175P / F210M, V39W / Y118L / F121W / F210M, V39W / Y118L / F121Y, V39W / Y118L / F121Y / F210M, V39W / F121W, V39W / F121W / F210M / V230I, I43E, I43Q, I43S, I43W, D44V, L76P, E79F, V101L, Y118L / F121W, Y118L / F121Y, Y118L / F121Y / K175P, Y118L / F121Y / I197V, F121V, F121W, F121Y, F121Y / K175P / V230I, F121Y / F210M, L148F, L153W, S159Q, S161M, L172A, L172E, L172G, L172Q, L172R, L172V, V200G, D207N, A233R, L245D, L245G, I248G, and I248S, and includes at least at least one substitution or a set of substitutions selected from, where the positions are numbered with reference to SEQ ID NO: 24. In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 598 to 842.
[0015] The present invention further provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 496 and comprising at least at least one substitution or a set of substitutions selected from 43 / 103, 43 / 103 / 172 / 238, 43 / 103 / 238 / 241, 43 / 238, 81 / 232, 103, 103 / 159 / 238, 103 / 194 / 238, 103 / 238, 122, 232, 238, and 238 / 241, wherein the positions are numbered with reference to SEQ ID NO: 496. In some additional embodiments, the engineered polypeptide comprises at least one substitution or a set of substitutions selected from SEQ ID NO: 43E / 103S / 238M / 241R, 43M / 103L, 43Q / 103S / 238M / 241R, 43Q / 238M, 43S / 103S / 172R / 238M, 43S / 238M, 81F / 232G, 103L, 103S, 103S / 159Q / 238M, 103S / 194L / 238M, 103S / 238M, 103V, 122C, 122I, 122K, 122L, 122M, 122Q, 122R, 122T, 122V, 232G, 238M, and 238M / 241R, wherein the positions are numbered with reference to SEQ ID NO: 496. In some further embodiments, the engineered polypeptide comprises at least at least one substitution or a set of substitutions selected from SEQ ID NO: I43E / H103S / Q238M / K241R, I43M / H103L, I43Q / H103S / Q238M / K241R, I43Q / Q238M, I43S / H103S / L172R / Q238M, I43S / Q238M, G81F / S232G, H103L, H103S, H103S / S159Q / Q238M, H103S / Y194L / Q238M, H103S / Q238M, H103V, S122C, S122I, S122K, S122L, S122M, S122Q, S122R, S122T, S122V, S232G, Q238M, and Q238M / K241R, wherein the positions are numbered with reference to SEQ ID NO: 496.In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 844 to 892.
[0016] The present invention further provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 496 and comprising at least at least one substitution or set of substitutions selected from 12, 43, 43 / 47 / 103, 43 / 47 / 103 / 232, 43 / 103, 43 / 103 / 121 / 172 / 241, 43 / 103 / 172, 43 / 103 / 172 / 238, 43 / 103 / 238, 43 / 103 / 238 / 241, 43 / 121, 43 / 194 / 238, 43 / 238, 44 / 103 / 121 / 238, 44 / 103 / 238, 44 / 103 / 238 / 241, 44 / 121 / 238 / 241, 47 / 103, 63, 103, 103 / 121, 103 / 121 / 172 / 238, 103 / 159 / 238, 103 / 172 / 238, 103 / 194 / 238, 103 / 238, 122, 131, 146, 178, 180, and 221, where the positions are numbered with reference to SEQ ID NO: 496.In some additional embodiments, the engineered polypeptide comprises at least one substitution or set of substitutions selected from 12V, 43E / 103S / 121W / 172R / 241R, 43E / 103S / 238M, 43E / 103S / 238M / 241R, 43M / 47G / 103L, 43M / 47G / 103V, 43M / 103L, 43M / 103S, 43M / 103V, 43Q / 47G / 103S / 232G, 43Q / 103L / 172R, 43Q / 103S, 43Q / 103S / 238M / 241R, 43Q / 121W, 43Q / 238M, 43S, 43S / 103L, 43S / 103S, 43S / 103S / 172R, 43S / 103S / 172R / 238M, 43S / 103S / 238M / 241R, 43S / 194L / 238M, 43S / 238M, 44V / 103S / 121W / 238M, 44V / 103S / 238M, 44V / 103S / 238M / 241R, 44V / 121W / 238M / 241R, 47G / 103L, 63P, 103L, 103S, 103S / 121W, 103S / 121W / 172R / 238M, 103S / 159Q / 238M, 103S / 172R / 238M, 103S / 194L / 238M, 103S / 238M, 103V, 122H, 122I, 122K, 122L, 122M, 122N, 122R, 122T, 122V, 131Y, 146I, 178A, 180L, 180V, and 221G, where positions are numbered with reference to SEQ ID NO: 496.In some further embodiments, the engineered polypeptide comprises at least at least one substitution or set of substitutions selected from I12V, I43E / H103S / Y121W / L172R / K241R, I43E / H103S / Q238M, I43E / H103S / Q238M / K241R, I43M / Q47G / H103L, I43M / Q47G / H103V, I43M / H103L, I43M / H103S, I43M / H103V, I43Q / Q47G / H103S / S232G, I43Q / H103L / L172R, I43Q / H103S, I43Q / H103S / Q238M / K241R, I43Q / Y121W, I43Q / Q238M, I43S, I43S / H103L, I43S / H103S, I43S / H103S / L172R, I43S / H103S / L172R / Q238M, I43S / H103S / Q238M / K241R, I43S / Y194L / Q238M, I43S / Q238M, N44V / H103S / Y121W / Q238M, N44V / H103S / Q238M, N44V / H103S / Q238M / K241R, N44V / Y121W / Q238M / K241R, Q47G / H103L, T63P, H103L, H103S, H103S / Y121W, H103S / Y121W / L172R / Q238M, H103S / S159Q / Q238M, H103S / L172R / Q238M, H103S / Y194L / Q238M, H103S / Q238M, H103V, S122H, S122I, S122K, S122L, S122M, S122N, S122R, S122T, S122V, S131Y, V146I, F178A, E180L, E180V, and N221G, where the positions are numbered with reference to SEQ ID NO: 496. In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 894 to 980.
[0017] The present invention includes an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 860, and further provides an engineered polypeptide comprising at least at least one substitution or a set of substitutions selected from 12, 12 / 43 / 63 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 63 / 103 / 146 / 241, 12 / 43 / 81 / 103 / 241, 12 / 43 / 81 / 146 / 232 / 241, 12 / 43 / 81 / 180 / 241, 12 / 43 / 81 / 232 / 241, 12 / 43 / 81 / 241, 12 / 43 / 103, 12 / 43 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 103 / 146 / 180 / 241, 12 / 43 / 103 / 180 / 241, 12 / 43 / 103 / 241, 12 / 43 / 146 / 180 / 232 / 241, 12 / 43 / 180 / 241, 12 / 63 / 81 / 103 / 241, 12 / 81, 12 / 81 / 180 / 241, 12 / 103 / 146 / 241, 12 / 103 / 180 / 241, 12 / 146, 12 / 146 / 180 / 241, 12 / 180 / 232 / 241, 12 / 241, 18, 43 / 81 / 232 / 241, 43 / 103 / 146 / 180, 43 / 103 / 180 / 241, 43 / 103 / 241, 43 / 180 / 232, 43 / 232 / 241, 63 / 103 / 180 / 232, 78, 81, 81 / 103 / 146, 81 / 146 / 180 / 241, 81 / 146 / 241, 81 / 180 / 241, 84, 103, 103 / 146, 103 / 146 / 180, 103 / 146 / 180 / 241, 103 / 180, 103 / 180 / 241, 103 / 261, 144, 145, 146 / 180 / 232, 146 / 180 / 241, 146 / 241, 150, 166, 172, 180 / 232, 232 / 241, and 241, where the positions are numbered with reference to SEQ ID NO: 860. In some additional embodiments, the engineered polypeptide is SEQ ID NO: 12V, 12V / 43M / 63P / 103L / 146I / 180L / 232G / 241R, 12V / 43M / 81F / 103L / 241R, 12V / 43M / 81F / 146I / 232G / 241R, 12V / 43M / 81F / 180V / 241R,At least one substitution or set of substitutions selected from 12V / 43M / 81F / 232G / 241R, 12V / 43M / 81F / 241R, 12V / 43M / 103L / 180L / 241R, 12V / 43M / 103L / 241R, 12V / 43M / 103V / 146I / 180V / 241R, 12V / 43M / 103V / 241R, 12V / 43M / 146I / 180L / 232G / 241R, 12V / 43M / 180L / 241R, 12V / 43M / 180V / 241R, 12V / 43Q / 63P / 103L / 146I / 241R, 12V / 43Q / 103L / 146I / 180V / 232G / 241R, 12V / 43Q / 103L / 146I / 180V / 241R, 12V / 43Q / 103L / 180V / 241R, 12V / 43Q / 103V, 12V / 63P / 81F / 103L / 241R, 12V / 81F, 12V / 81F / 180V / 241R, 12V / 103L / 146I / 241R, 12V / 103L / 180L / 241R, 12V / 103V / 180V / 241R, 12V / 146I, 12V / 146I / 180L / 241R, 12V / 180L / 232G / 241R, 12V / 180V / 232G / 241R, 12V / 241R, 18V, 43M / 81F / 232G / 241R, 43M / 103L / 146I / 180V, 43M / 103L / 180V / 241R, 43M / 103L / 241R, 43M / 103V / 146I / 180V, 43M / 180L / 232G, 43M / 232G / 241R, 43Q / 103V / 146I / 180L, 63P / 103L / 180L / 232G, 78A, 81F / 103L / 146I, 81F / 146I / 180L / 241R, 81F / 146I / 241R, 81F / 180L / 241R, 81Q, 84Y, 103A, 103L, 103L / 146I, 103L / 146I / 180V, 103L / 180L, 103L / 180V / 241R, 103N, 103T, 103T / 261V, 103V / 146I / 180L / 241R, 144V, 145V, 146I / 180L / 241R, 146I / 180V / 232G, 146I / 241R, 150H, 166L, 172L, 172V, 180L / 232G, 232G / 241R, and 241R, where the positions are numbered with reference to SEQ ID NO: 860. In some further embodiments,The engineered polypeptides are SEQ ID NO: I12V, I12V / S43M / T63P / S103L / V146I / E180L / S232G / K241R, I12V / S43M / G81F / S103L / K241R, I12V / S43M / G81F / V146I / S232G / K241R, I12V / S43M / G81F / E180V / K241R, I12V / S43M / G81F / S232G / K241R, I12V / S43M / G81F / K241R, I12V / S43M / S103L / E180L / K241R, I12V / S43M / S103L / K241R, I12V / S43M / S103V / V146I / E180V / K241R, I12V / S43M / S103V / K241R, I12V / S43M / V146I / E180L / S232G / K241R, I12V / S43M / E180L / K241R, I12V / S43M / E180V / K241R, I12V / S43Q / T63P / S103L / V146I / K241R, I12V / S43Q / S103L / V146I / E180V / S232G / K241R, I12V / S43Q / S103L / V146I / E180V / K241R, I12V / S43Q / S103L / E180V / K241R, I12V / S43Q / S103V, I12V / T63P / G81F / S103L / K241R, I12V / G81F, I12V / G81F / E180V / K241R, I12V / S103L / V146I / K241R, I12V / S103L / E180L / K241R, I12V / S103V / E180V / K241R, I12V / V146I, I12V / V146I / E180L / K241R, I12V / E180L / S232G / K241R, I12V / E180V / S232G / K241R, I12V / K241R, L18V, S43M / G81F / S232G / K241R, S43M / S103L / V146I / E180V, S43M / S103L / E180V / K241R, S43M / S103L / K241R, S43M / S103V / V146I / E180V, S43M / E180L / S232G, S43M / S232G / K241R, S43Q / S103V / V146I / E180L, T63P / S103L / E180L / S232G, G78A, G81F / S103L / V146I, G81F / V146I / E180L / K241R,Comprising at least at least one substitution or set of substitutions selected from G81F / V146I / K241R, G81F / E180L / K241R, G81Q, I84Y, S103A, S103L, S103L / V146I, S103L / V146I / E180V, S103L / E180L, S103L / E180V / K241R, S103N, S103T, S103T / A261V, S103V / V146I / E180L / K241R, T144V, A145V, V146I / E180L / K241R, V146I / E180V / S232G, V146I / K241R, D150H, V166L, R172L, R172V, E180L / S232G, S232G / K241R, and K241R, where the positions are numbered with reference to SEQ ID NO: 860. In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 982 to 1118.,
[0018] The present invention further provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 860, and comprising at least at least one substitution or a set of substitutions selected from 9, 12, 12 / 43 / 63 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 63 / 103 / 146 / 241, 12 / 43 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 103 / 146 / 180 / 241, 12 / 43 / 103 / 146 / 232 / 241, 12 / 43 / 103 / 146 / 241, 12 / 43 / 103 / 180 / 241, 12 / 43 / 146 / 180 / 232 / 241, 12 / 43 / 146 / 241, 12 / 43 / 180 / 232, 12 / 43 / 180 / 241, 12 / 63 / 103 / 146 / 232 / 241, 12 / 103 / 146 / 241, 12 / 103 / 180 / 241, 12 / 146, 12 / 146 / 180 / 241, 12 / 146 / 232 / 241, 12 / 146 / 241, 12 / 180, 12 / 180 / 232 / 241, 12 / 180 / 241, 12 / 241, 18, 43 / 63 / 103 / 180 / 232 / 241, 43 / 103 / 146 / 180, 43 / 103 / 180, 43 / 103 / 180 / 232 / 241, 43 / 103 / 180 / 241, 43 / 180 / 232, 63 / 103 / 180 / 232, 103, 103 / 146 / 180, 103 / 146 / 180 / 241, 103 / 180, 103 / 180 / 241, 103 / 241, 118, 119, 122, 144, 145, 146 / 180 / 241, 154, 160, 172, 180 / 232, 180 / 232 / 241, 180 / 241, 208, 211, and 212, wherein the positions are numbered with reference to SEQ ID NO: 860. In some additional embodiments, the engineered polypeptide is SEQ ID NO: 9V, 12V, 12V / 43M / 63P / 103L / 146I / 180L / 232G / 241R, 12V / 43M / 103L / 146I / 241R, 12V / 43M / 103L / 180L / 241R, 12V / 43M / 103V / 146I / 180V / 241R, 12V / 43M / 103V / 146I / 232G / 241R,12V / 43M / 146I / 180L / 232G / 241R, 12V / 43M / 180L / 232G, 12V / 43M / 180L / 241R, 12V / 43Q / 63P / 103V / 146I / 241R, 12V / 43Q / 103L / 146I / 180V / 232G / 241R, 12V / 43Q / 103L / 146I / 180V / 241R, 12V / 43Q / 103L / 146I / 241R, 12V / 43Q / 103L / 180V / 241R, 12V / 43Q / 146I / 241R, 12V / 63P / 103L / 146I / 232G / 241R, 12V / 103L / 180L / 241R, 12V / 103L / 180V / 241R, 12V / 103V / 146I / 241R, 12V / 103V / 180L / 241R, 12V / 103V / 180V / 241R, 12V / 146I, 12V / 146I / 180L / 241R, 12V / 146I / 232G / 241R, 12V / 146I / 241R, 12V / 180L, 12V / 180L / 232G / 241R, 12V / 180L / 241R, 12V / 180V, 12V / 180V / 232G / 241R, 12V / 180V / 241R, 12V / 241R, 18V, 43M / 63P / 103L / 180V / 232G / 241R, 43M / 103L / 146I / 180V, 43M / 103L / 180V, 43M / 103L / 180V / 241R, 43M / 103V / 146I / 180V, 43M / 180L / 232G, 43Q / 103L / 180V / 232G / 241R, 43Q / 103V / 146I / 180L, 43Q / 103V / 180L / 241R, 43Q / 103V / 180V / 232G / 241R, 63P / 103L / 180L / 232G, 103A, 103L, 103L / 146I / 180V, 103L / 146I / 180V / 241R, 103L / 180L, 103L / 180V / 241R, 103L / 241R, 103N, 103T, 103V / 146I / 180L / 241R, 103V / 180V / 241R, 118L, 118V, 119H, 122H, 144H, 144V, 145V, 146I / 180L / 241R, 154R, 160L, 172C, 172L, 172V, 180L / 232G, 180L / 241R, 180V / 232G, 180V / 232G / 241R, 208F, 211Fcomprising at least one substitution or a set of one substitution selected from 212P, wherein the positions are numbered with reference to SEQ ID NO: 860. In some further embodiments, the engineered polypeptide is SEQ ID NO: I9V, I12V, I12V / S43M / T63P / S103L / V146I / E180L / S232G / K241R, I12V / S43M / S103L / V146I / K241R, I12V / S43M / S103L / E180L / K241R, I12V / S43M / S103V / V146I / E180V / K241R, I12V / S43M / S103V / V146I / S232G / K241R, I12V / S43M / V146I / E180L / S232G / K241R, I12V / S43M / E180L / S232G, I12V / S43M / E180L / K241R, I12V / S43Q / T63P / S103V / V146I / K241R, I12V / S43Q / S103L / V146I / E180V / S232G / K241R, I12V / S43Q / S103L / V146I / E180V / K241R, I12V / S43Q / S103L / V146I / K241R, I12V / S43Q / S103L / E180V / K241R, I12V / S43Q / V146I / K241R, I12V / T63P / S103L / V146I / S232G / K241R, I12V / S103L / E180L / K241R, I12V / S103L / E180V / K241R, I12V / S103V / V146I / K241R, I12V / S103V / E180L / K241R, I12V / S103V / E180V / K241R, I12V / V146I, I12V / V146I / E180L / K241R, I12V / V146I / S232G / K241R, I12V / V146I / K241R, I12V / E180L, I12V / E180L / S232G / K241R, I12V / E180L / K241R, I12V / E180V, I12V / E180V / S232G / K241R, I12V / E180V / K241R, I12V / K241R, L18V, S43M / T63P / S103L / E180V / S232G / K241R, S43M / S103L / V146I / E180V, S43M / S103L / E180V, S43M / S103L / E180V / K241R,Comprising at least at least one substitution or set of substitutions selected from S43M / S103V / V146I / E180V, S43M / E180L / S232G, S43Q / S103L / E180V / S232G / K241R, S43Q / S103V / V146I / E180L, S43Q / S103V / E180L / K241R, S43Q / S103V / E180V / S232G / K241R, T63P / S103L / E180L / S232G, S103A, S103L, S103L / V146I / E180V, S103L / V146I / E180V / K241R, S103L / E180L, S103L / E180V / K241R, S103L / K241R, S103N, S103T, S103V / V146I / E180L / K241R, S103V / E180V / K241R, Y118L, Y118V, K119H, S122H, T144H, T144V, A145V, V146I / E180L / K241R, A154R, N160L, R172C, R172L, R172V, E180L / S232G, E180L / K241R, E180V / S232G, E180V / S232G / K241R, Q208F, S211F, and R212P, wherein the positions are numbered with reference to SEQ ID NO: 860. In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 1120 to 1212.,
[0019] The present invention includes an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 1120, and further provides an engineered polypeptide comprising at least at least one substitution or a set of substitutions selected from 3, 7, 7 / 158, 7 / 158 / 175, 7 / 158 / 197 / 202, 27 / 43 / 115 / 172, 27 / 43 / 172, 29, 32, 43 / 110 / 115 / 172, 44, 46, 48 / 172, 55, 59, 67, 71, 72, 73, 94 / 103 / 118 / 146, 103, 103 / 144 / 146 / 172 / 180, 103 / 146 / 154 / 172 / 180 / 232, 103 / 154, 118, 118 / 144 / 215, 118 / 172 / 180, 137, 142, 146, 146 / 154, 154, 158, 172 / 190, 184, 198, 205, 216, 223, 226, 229, 247, 261, and 263, wherein the positions are numbered with reference to SEQ ID NO: 1120.In some additional embodiments, the engineered polypeptide comprises at least one substitution or set of substitutions selected from the group consisting of SEQ ID NO: 3D, 7V, 7V / 158Y, 7V / 158Y / 175P, 7V / 158Y / 197V / 202I, 27E / 43I / 115S / 172L, 27E / 43I / 172L, 29G, 29R, 32G, 32R, 43I / 110T / 115S / 172L, 44S, 46H, 46V, 48I / 172L, 55A, 55G, 55R, 55S, 59G, 59T, 67N, 71S, 72G, 73R, 94Q / 103L / 118V / 146I, 103L, 103L / 144V / 146I / 172L / 180V, 103L / 146I / 154R / 172L / 180V / 232G, 103L / 154R, 118V, 118V / 144V / 215H, 118V / 172L / 180V, 137M, 137R, 142V, 146I, 146I / 154R, 154R, 158Y, 172L / 190S, 184S, 198L, 205G, 205L, 216M, 216V, 223A, 223C, 223E, 223L, 223M, 223S, 226L, 226V, 229L, 229V, 229W, 247V, 261G, 261H, 263E, 263G, 263L, and 263S, where the positions are numbered with reference to SEQ ID NO: 1120.In some further embodiments, the engineered polypeptide comprises at least at least one substitution or a set of substitutions selected from SEQ ID NO: S3D, I7V, I7V / F158Y, I7V / F158Y / K175P, I7V / F158Y / I197V / V202I, Q27E / S43I / A115S / R172L, Q27E / S43I / R172L, A29G, A29R, N32G, N32R, S43I / I110T / A115S / R172L, N44S, R46H, R46V, L48I / R172L, E55A, E55G, E55R, E55S, E59G, E59T, S67N, G71S, K72G, K73R, P94Q / V103L / Y118V / V146I, V103L, V103L / T144V / V146I / R172L / L180V, V103L / V146I / A154R / R172L / L180V / S232G, V103L / A154R, Y118V, Y118V / T144V / Q215H, Y118V / R172L / L180V, T137M, T137R, T142V, V146I, V146I / A154R, A154R, F158Y, R172L / T190S, D184S, R198L, E205G, E205L, L216M, L216V, K223A, K223C, K223E, K223L, K223M, K223S, K226L, K226V, C229L, C229V, C229W, Q247V, A261G, A261H, A263E, A263G, A263L, and A263S, where the positions are numbered with reference to SEQ ID NO: 1120. In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 1214 to 1344.
[0020] The present invention further provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 1120, and comprising at least at least one substitution or a set of substitutions selected from 73, 80, 103 / 146 / 209, 103 / 172 / 209, 118 / 119 / 172 / 209, 135, 144 / 209 / 232, 172 / 180 / 209, 181, and 223, where the positions are numbered with reference to SEQ ID NO: 1120. In some additional embodiments, the engineered polypeptide comprises at least one substitution or a set of substitutions selected from SEQ ID NO: 73R, 80A, 80G, 103L / 146I / 209C, 103L / 172C / 209C, 118V / 119H / 172C / 209C, 135V, 144V / 209C / 232G, 172C / 180V / 209C, 181S, and 223E, where the positions are numbered with reference to SEQ ID NO: 1120. In some further embodiments, the engineered polypeptide comprises at least at least one substitution or a set of substitutions selected from SEQ ID NO: K73R, S80A, S80G, V103L / V146I / I209C, V103L / R172C / I209C, Y118V / K119H / R172C / I209C, N135V, T144V / I209C / S232G, R172C / L180V / I209C, Q181S, and K223E, where the positions are numbered with reference to SEQ ID NO: 1120. In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 1346 to 1366.
[0021] The present invention further provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 1348, and comprising at least at least one substitution or a set of substitutions selected from 2, 7 / 46, 39, 46, 46 / 55, 46 / 55 / 72, 46 / 55 / 92, 46 / 55 / 118, 46 / 55 / 226 / 257 / 263, 46 / 55 / 263, 46 / 70, 46 / 70 / 92 / 118 / 261 / 263, 46 / 70 / 175 / 263, 46 / 72, 46 / 118, 55 / 257 / 261, 55 / 261 / 263, 55 / 263, 105, 118, 118 / 226 / 261 / 263, and 263, wherein the positions are numbered with reference to SEQ ID NO: 1348. In some additional embodiments, the engineered polypeptide comprises at least one substitution or a set of substitutions selected from 2G, 7V / 46H, 39F, 39G, 39M, 39Q, 39S, 39T, 39V, 46H, 46H / 55R, 46H / 55R / 72C, 46H / 55R / 92A, 46H / 55R / 118V, 46H / 55R / 226V / 257Q / 263L, 46H / 55R / 263L, 46H / 70N, 46H / 70N / 92A / 118V / 261G / 263L, 46H / 70N / 175P / 263G, 46H / 72C, 46H / 118V, 55R / 257Q / 261G, 55R / 261G / 263I, 55R / 263G, 105G, 118V, 118V / 226V / 261G / 263I, and 263G, wherein the positions are numbered with reference to SEQ ID NO: 1348.In some further embodiments, the engineered polypeptide comprises at least at least one substitution or set of substitutions selected from SEQ ID NO: V2G, I7V / R46H, W39F, W39G, W39M, W39Q, W39S, W39T, W39V, R46H, R46H / E55R, R46H / E55R / K72C, R46H / E55R / K92A, R46H / E55R / Y118V, R46H / E55R / K226V / S257Q / A263L, R46H / E55R / A263L, R46H / E70N, R46H / E70N / K92A / Y118V / A261G / A263L, R46H / E70N / K175P / A263G, R46H / K72C, R46H / Y118V, E55R / S257Q / A261G, E55R / A261G / A263I, E55R / A263G, A105G, Y118V, Y118V / K226V / A261G / A263I, and A263G, where positions are numbered with reference to SEQ ID NO: 1348. In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 1368 to 1422.
[0022] The present invention includes an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 1348, and further provides an engineered polypeptide comprising at least at least one substitution or a set of substitutions selected from 2, 7 / 46, 7 / 46 / 175, 22 / 46 / 55 / 118, 27 / 46 / 72, 27 / 46 / 146 / 158 / 202, 27 / 46 / 175, 27 / 118, 29 / 70 / 118, 39, 46, 46 / 55, 46 / 55 / 59 / 261 / 263, 46 / 55 / 72, 46 / 55 / 92, 46 / 55 / 118, 46 / 55 / 118 / 263, 46 / 55 / 146, 46 / 55 / 146 / 202, 46 / 55 / 226 / 257 / 263, 46 / 55 / 263, 46 / 70, 46 / 70 / 92 / 118 / 261 / 263, 46 / 70 / 175 / 263, 46 / 72, 46 / 92, 46 / 92 / 175, 46 / 118, 46 / 146, 46 / 146 / 257 / 263, 55 / 70 / 118, 55 / 257 / 261, 55 / 257 / 263, 55 / 261 / 263, 55 / 263, 70, 92 / 118, 105, 118, 118 / 226 / 261 / 263, 146 / 257 / 263, 146 / 261 / 263, 192 / 257 / 261, 202, 257 / 261 / 263, 261 / 263, and 263, where the positions are numbered with reference to SEQ ID NO: 1348.In some additional embodiments, the engineered polypeptide comprises at least one substitution or set of substitutions selected from SEQ ID NO: 2G, 7V / 46H, 7V / 46H / 175P, 22V / 46H / 55R / 118V, 27G / 46H / 72C, 27G / 46H / 146I / 158Y / 202I, 27G / 46H / 175P, 27G / 118V, 29V / 70N / 118V, 39M, 39Q, 39S, 39T, 39V, 46H, 46H / 55R, 46H / 55R / 59T / 261G / 263L, 46H / 55R / 72C, 46H / 55R / 92A, 46H / 55R / 118V, 46H / 55R / 118V / 263S, 46H / 55R / 146I, 46H / 55R / 146I / 202I, 46H / 55R / 226V / 257Q / 263L, 46H / 55R / 263L, 46H / 70N, 46H / 70N / 92A / 118V / 261G / 263L, 46H / 70N / 175P / 263G, 46H / 72C, 46H / 92A, 46H / 92A / 175P, 46H / 118V, 46H / 146I, 46H / 146I / 257Q / 263I, 55R / 70N / 118V, 55R / 257Q / 261G, 55R / 257Q / 263L, 55R / 261G / 263I, 55R / 263G, 70N, 92A / 118V, 105G, 118V, 118V / 226V / 261G / 263I, 146I / 257Q / 263S, 146I / 261G / 263L, 192V / 257Q / 261G, 202I, 257Q / 261G / 263G, 261G / 263I, and 263G, where the positions are numbered with reference to SEQ ID NO: 1348.In some further embodiments, the engineered polypeptide comprises at least at least one substitution or set of substitutions selected from SEQ ID NO: V2G, I7V / R46H, I7V / R46H / K175P, L22V / R46H / E55R / Y118V, Q27G / R46H / K72C, Q27G / R46H / V146I / F158Y / V202I, Q27G / R46H / K175P, Q27G / Y118V, A29V / E70N / Y118V, W39M, W39Q, W39S, W39T, W39V, R46H, R46H / E55R, R46H / E55R / E59T / A261G / A263L, R46H / E55R / K72C, R46H / E55R / K92A, R46H / E55R / Y118V, R46H / E55R / Y118V / A263S, R46H / E55R / V146I, R46H / E55R / V146I / V202I, R46H / E55R / K226V / S257Q / A263L, R46H / E55R / A263L, R46H / E70N, R46H / E70N / K92A / Y118V / A261G / A263L, R46H / E70N / K175P / A263G, R46H / K72C, R46H / K92A, R46H / K92A / K175P, R46H / Y118V, R46H / V146I, R46H / V146I / S257Q / A263I, E55R / E70N / Y118V, E55R / S257Q / A261G, E55R / S257Q / A263L, E55R / A261G / A263I, E55R / A263G, E70N, K92A / Y118V, A105G, Y118V, Y118V / K226V / A261G / A263I, V146I / S257Q / A263S, V146I / A261G / A263L, E192V / S257Q / A261G, V202I, S257Q / A261G / A263G, A261G / A263I, and A263G, where the positions are numbered with reference to SEQ ID NO: 1348. In some additional embodiments, the engineered polypeptide comprises an amino acid sequence having at least 80% sequence identity to any even-numbered sequence shown in SEQ ID NOs: 1424 to 1484.
[0023] The present invention also provides an engineered polynucleotide encoding at least one engineered polypeptide described in the above paragraph. In some embodiments, the engineered polynucleotide comprises the odd-numbered sequences represented by SEQ ID NO: 3 to SEQ ID NO: 1483.
[0024] The present invention further provides a vector comprising at least one of the above-described engineered polynucleotides. In some embodiments, the vector further comprises at least one control sequence.
[0025] The present invention also provides a host cell comprising the vector provided herein. In some embodiments, the host cell produces at least one engineered polypeptide provided herein.
[0026] The present invention further provides a method for producing an engineered nitroaldolase polypeptide, the method comprising culturing a host cell provided herein under conditions such that the engineered polynucleotide is expressed and the engineered polypeptide is produced. In some embodiments, the method further comprises the step of recovering the engineered polypeptide.
BEST MODE FOR CARRYING OUT THE INVENTION
[0027] Description of the Invention Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In general, the nomenclature used herein, as well as the experimental procedures in cell culture, molecular genetics, microbiology, organic chemistry, analytical chemistry, and nucleic acid chemistry described below, are those well known and commonly employed in the art. Such techniques are well known and described in numerous textbooks and references well known to those skilled in the art. Standard techniques or their modified forms are used in chemical synthesis and chemical analysis. All patents, patent applications, papers, and publications referred to both above and below in this specification are hereby expressly incorporated herein by reference.
[0028] Any suitable methods and materials similar to or equivalent to those described herein may be used in the practice of the present invention, but some methods and materials are described herein. It should be understood that the present invention is not limited to the specific methodologies, protocols, and reagents described, as these may vary depending on the circumstances in which they are used by those skilled in the art. Accordingly, the terms defined immediately below are further described more fully by reference to the present invention as a whole.
[0029] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. The section headings used herein are for purposes of organization only and should not be construed as limiting the subject matter described. Numerical ranges include the numbers defining the range. Accordingly, any numerical range disclosed herein is intended to include any narrower numerical range that falls within such broader numerical range, as if such narrower numerical ranges were all expressly recited herein. Any maximum (or minimum) numerical limitation disclosed herein is also intended to include any numerical lower (or upper) limit, as if such numerical lower (or upper) limits were expressly recited herein.
[0030] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, "a polypeptide" includes more than one polypeptide. Similarly, "comprise", "comprises", "comprising", "include", "includes", and "including" are interchangeable and not intended to be limiting.
[0031] When the description of various embodiments uses the term "comprising", it will be understood by those skilled in the art that in some specific cases, the embodiments may instead be described using the words "consisting essentially of" or "consisting of". When the description of various embodiments uses the terms "optionally" and "optionally", it should be further understood that the event or situation described thereafter may or may not occur, and that the description includes cases where the event or situation occurs and cases where it does not occur. It should be understood that both the above summary and the following detailed description are merely exemplary and explanatory and do not limit the present disclosure. The section headings used in this specification are for purposes of organization only and should not be construed as limiting the subject matter described.
[0032] Abbreviations: The abbreviations used for amino acids encoded by genes are conventional and are as follows:
Table A
[0033] When three-letter abbreviations are used, unless in particular cases where "L" or "D" precedes or it is clear from the context in which the abbreviation is used, the amino acid may have an L-configuration or a D-configuration about the α-carbon (Cα). For example, "Ala" represents alanine without specification of the configuration about the α-carbon, while "D-Ala" and "L-Ala" represent D-alanine and L-alanine, respectively.
[0034] When one-letter abbreviations are used, capital letters represent amino acids with an L-configuration about the α-carbon, and lowercase letters represent amino acids with a D-configuration about the α-carbon. For example, "A" represents L-alanine and "a" represents D-alanine. When a polypeptide sequence is presented as a series of one-letter or three-letter abbreviations (or a mixture thereof), the sequence is presented in the direction from amino (N) to carboxy (C) according to common convention.
[0035] The abbreviations used for the nucleosides encoded by the genetic code are as conventional and are as follows: adenosine (A), guanosine (G), cytidine (C), thymidine (T), and uridine (U). Unless otherwise specifically explained, the abbreviated nucleosides may be ribonucleosides or 2'-deoxyribonucleosides. The nucleosides may be designated, individually or collectively, as either ribonucleosides or 2'-deoxyribonucleosides. When a nucleic acid sequence is presented as a series of one-letter abbreviations, the sequence is presented in the 5' to 3' direction according to common convention, and the phosphates are not shown.
[0036] Definitions For the present invention, the technical and scientific terms used in this specification shall have the meanings commonly understood by those skilled in the art, unless otherwise specifically defined. Accordingly, the following terms are intended to have the following meanings.
[0037] The "EC" number refers to the enzyme nomenclature of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (NC-IUBMB). The IUBMB biochemical classification is a numerical classification system for enzymes based on the chemical reactions they catalyze.
[0038] "ATCC" refers to the American Type Culture Collection, and its biorepository collection includes genes and strains.
[0039] "NCBI" refers to the National Center for Biological Information in the United States and the sequence databases provided therein.
[0040] "Protein", "polypeptide", and "peptide" are used interchangeably herein to denote a polymer of at least two amino acids covalently linked by amide bonds, regardless of length or post-translational modifications (e.g., glycosylation, phosphorylation, lipidation, myristoylation, ubiquitination, etc.). This definition includes D-amino acids, L-amino acids, mixtures of D- and L-amino acids, polymers containing D-amino acids, polymers containing L-amino acids, and polymers containing mixtures of D- and L-amino acids.
[0041] "Amino acid" is referred to herein by either its generally known three-letter symbol or the one-letter symbol recommended by the IUPAC-IUB Commission on Biochemical Nomenclature. Nucleotides may likewise be referred to by their generally accepted one-letter code.
[0042] As used herein, "polynucleotide" and "nucleic acid" refer to two or more nucleotides covalently linked to each other. A polynucleotide may consist entirely of ribonucleotides (i.e., RNA), entirely of 2'-deoxyribonucleotides (i.e., DNA), or may consist of a mixture of ribonucleotides and 2'-deoxyribonucleotides. Nucleosides are typically linked to each other by standard phosphodiester linkages, but a polynucleotide may contain one or more non-standard linkages. A polynucleotide may be single-stranded or double-stranded, or may contain both single-stranded and double-stranded regions. Further, a polynucleotide typically consists of naturally occurring coding nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), but may contain one or more modified and / or synthetic nucleobases, such as inosine, xanthine, hypoxanthine, etc. In some embodiments, such modified or synthetic nucleobases are nucleobases that encode an amino acid sequence.
[0043] "Nitroaldolase", "hydroxynitrile lyase" or "oxynitrilase" refers to an enzyme categorized as EC 4.1.2.47 that has native activity in the interconversion of cyanohydrins to cyanide and the corresponding aldehydes and ketones. These enzymes are also known to have activity as nitroaldolase in the conversion of nitroalkanes and aldehydes to β-nitroalcohols. The nitroaldolase enzyme of the present invention is derived from Baliospermum montanum, but the present disclosure is not so limited, and the nitroaldolase enzyme may be derived from any suitable organism or may be produced synthetically. Nitroaldolase, as used herein, includes not only naturally occurring (wild-type) nitroaldolase but also non-naturally occurring engineered polypeptides produced by human manipulation.
[0044] "Coding sequence" refers to that portion of a nucleic acid (e.g., a gene) that encodes the amino acid sequence of a protein.
[0045] "Naturally occurring" or "wild-type" refers to the form found in nature. For example, a naturally occurring or wild-type polypeptide or polynucleotide sequence is a sequence that exists in an organism, can be isolated from a natural source, and has not been intentionally modified by human manipulation.
[0046] As used herein, "recombinant," "engineered," and "not naturally occurring" when used in reference to a cell, nucleic acid, or polypeptide refer to a substance that has been modified in a manner that does not occur in nature or a substance corresponding to the natural or native form of such a substance. In some embodiments, a cell, nucleic acid, or polypeptide is identical to a naturally occurring cell, nucleic acid, or polypeptide but is produced or derived from synthetic substances and / or by manipulation using recombinant techniques. Non-limiting examples include, among others, recombinant cells that express a gene not found in the native (non-recombinant) form of the cell or that express a native gene at a different level than would otherwise occur.
[0047] The terms "percent sequence identity" and "percent homology" are used interchangeably herein to refer to a comparison between polynucleotides or polypeptides, and are determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may include additions or deletions (i.e., gaps) as compared to the reference sequence for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percent sequence identity. Alternatively, the percentage can be calculated by determining the number of positions at which either the identical nucleic acid base or amino acid residue occurs in both sequences or the number of positions at which the nucleic acid base or amino acid residue is aligned using gaps to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percent sequence identity. It will be understood by those of skill in the art that there are numerous established algorithms available for aligning two sequences. Optimal alignment of sequences for comparison can be conducted, for example, by the local homology algorithm of Smith and Waterman (Smith and Waterman, Adv. Appl. Math., 2:482
[1981] ), the homology alignment algorithm of Needleman and Wunsch (Needleman and Wunsch, J. Mol. Biol., 48:443
[1970] ), by the similarity search method of Pearson and Lipman (Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444
[1988] ), by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin Software Package), or by visual inspection.Examples of algorithms suitable for determining percent sequence identity and percent sequence similarity include, but are not limited to, the BLAST and BLAST 2.0 algorithms described by Altschul et al. (see Altschul et al., J. Mol. Biol., 215: 403-410
[1990] ; and Altschul et al., Nucl. Acids Res., 3389-3402
[1977] , respectively). Software for performing BLAST analyses is publicly available through the website of the National Center for Biotechnology Information. This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that match or satisfy some positive-valued threshold score T when aligned with words of the same length in the database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits serve as seeds for initiating a search to find longer HSPs that contain them. The word hits are then extended in both directions along each sequence as far as possible where the cumulative alignment score can be increased. The cumulative score is calculated using the parameters M (reward score for a pair of matching residues, always >0) and N (penalty score for mismatching residues, always <0) for nucleotide sequences. For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction stops when the cumulative alignment score drops from its maximum achieved value by an amount X; stops when the cumulative score goes to zero or below due to the cumulative of one or more negative-scoring residue alignments; or stops when the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses, by default, a word length (W) of 11, an expectation value (E) of 10, M = 5, N = -4, and a comparison of both strands.For the BLASTP program with respect to amino acid sequences, by default, a word length (W) of 3, an expectation value (E) of 10, and the BLOSUM62 score matrix are used (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915
[1989] ). Exemplary determination of sequence alignment and percent sequence identity can utilize the BESTFIT or GAP programs of the GCG Wisconsin Software package (Accelrys, Madison WI) using the provided default parameters.
[0048] A "reference sequence" refers to a defined sequence used as a basis for sequence comparison. A reference sequence may be a subset of a larger sequence, e.g., a segment of a full-length gene or polypeptide sequence. Generally, a reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 residues in length, at least 50 residues in length, or full-length for nucleic acids or polypeptides. Since two polynucleotides or polypeptides may each (1) contain sequences that are similar between the two sequences (i.e., a portion of the full sequence) and (2) further contain sequences that differ between the two sequences, sequence comparison between two (or more) polynucleotides or polypeptides typically is performed by comparing the sequences of the two polynucleotides or polypeptides over a "comparison window" to identify and compare local regions of sequence similarity. In some embodiments, a "reference sequence" may be based on a primary amino acid sequence, in which case the reference sequence is a sequence that may have one or more changes from the primary sequence. For example, a "reference sequence based on SEQ ID NO: 4 having valine at the residue corresponding to X14", or X14V, refers to a reference sequence in which the corresponding residue at X14 of SEQ ID NO: 4, which is tyrosine, is changed to valine.
[0049] A "comparison window" refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acid residues that can be compared to a reference sequence of at least 20 contiguous nucleotides or amino acids, and the portion of the sequence within the comparison window can contain 20 percent or less addition or deletion (i.e., gap) compared to the reference sequence (excluding addition or deletion) for the optimal alignment of the two sequences. The comparison window can be longer than 20 contiguous residues and can include windows of 30, 40, 50, 100 or longer as needed.
[0050] As used herein, "substantial identity" refers to a polynucleotide or polypeptide sequence having at least 80 percent sequence identity, at least 85 percent identity, at least 89 - 95 percent sequence identity between, or more usually at least 99 percent sequence identity compared to a reference sequence over a comparison window of at least 20 residue positions, often over a window of at least 30 - 50 residues. This percent sequence identity is calculated by comparing the sequences over the comparison window to a reference sequence that includes deletions or additions that total 20 percent or less of the reference sequence. In some specific embodiments applied to polypeptides, the term "substantial identity" means that two polypeptide sequences, when optimally aligned, for example, using default gap weights and aligned optimally by the programs GAP or BESTFIT, share at least 80 percent sequence identity, preferably at least 89 percent sequence identity, at least 95 percent sequence identity or higher sequence identity (e.g., 99 percent sequence identity). In some embodiments, the non - identical residue positions in the sequences being compared are conservative amino acid substitutions.
[0051] As used in the context of numbering of a given amino acid or polynucleotide sequence, "corresponding to", "relative to" and "compared to" refer to the numbering of residues of a designated reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue numbers or residue positions of a given polymer are indicated relative to the reference sequence, rather than by the actual numbered positions of the residues within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence, e.g., the amino acid sequence of engineered nitroaldolase, can be aligned with a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, gaps are present, but the numbering of residues in the given amino acid or polynucleotide sequence is done relative to the reference sequence to which it is aligned.
[0052] "Amino acid difference" or "residue difference" refers to a change in an amino acid residue at a position in a polypeptide sequence compared to the amino acid residue at the corresponding position in a reference sequence. In some cases, the reference sequence has a histidine tag, but the numbering is maintained relative to an equivalent reference sequence without the histidine tag. The position of an amino acid difference is generally referred to herein as "Xn", where n refers to the corresponding position in the reference sequence that serves as the basis for the residue difference. For example, "a residue difference at position X25 compared to SEQ ID NO:2" refers to a change in the amino acid residue at the polypeptide position corresponding to position 25 of SEQ ID NO:2. Thus, if the reference polypeptide of SEQ ID NO:2 has valine at position 25, "a residue difference at position X25 compared to SEQ ID NO:2" is an amino acid substitution of any residue other than valine at the polypeptide position corresponding to position 25 of SEQ ID NO:2. In most cases herein, a particular amino acid residue difference at a position is denoted as "XnY", where "Xn" designates the corresponding position as described above, and "Y" is the one-letter identifier of the amino acid found in the engineered polypeptide (i.e., a residue different from that in the reference polypeptide). In some embodiments, more than one amino acid may occur at the designated residue position (i.e., alternative amino acids may be listed in the form XnY / Z, where Y and Z represent alternative amino acid residues). In some cases (e.g., in Tables 2-1, 2-2, 3-1, 3-2, 4-1, 4-2, 5-1, 5-2, 6-1, 6-2, 7-1 and 7-2), the invention also provides certain amino acid differences denoted by the conventional notation "AnB", where A is the one-letter identifier of the residue in the reference sequence, "n" is the number of the residue position in the reference sequence, and B is the one-letter identifier of the residue substitution in the sequence of the engineered polypeptide. Further, in some cases, the polypeptides of the invention include one or more amino acid residue differences compared to a reference sequence, which are indicated by a list of specific positions that have been changed compared to the reference sequence. In some additional embodiments, the invention provides engineered polypeptide sequences that include both conservative and non-conservative amino acid substitutions. As used herein, "conservative amino acid substitution" refers to the substitution of a residue with a different residue having a similar side chain, and thus typically includes the substitution of an amino acid in a polypeptide with an amino acid within the same or a similar defined class of amino acids. By way of example, and not limitation, an amino acid having an aliphatic side chain is substituted with another aliphatic amino acid (e.g., alanine, valine, leucine, and isoleucine); an amino acid having a hydroxyl side chain is substituted with another amino acid having a hydroxyl side chain (e.g., serine and threonine); an amino acid having an aromatic side chain is substituted with another amino acid having an aromatic side chain (e.g., phenylalanine, tyrosine, tryptophan, and histidine); an amino acid having a basic side chain is substituted with another amino acid having a basic side chain (e.g., lysine and arginine); an amino acid having an acidic side chain is substituted with another amino acid having an acidic side chain (e.g., aspartic acid or glutamic acid); and / or a hydrophobic or hydrophilic amino acid is replaced with another hydrophobic or hydrophilic amino acid, respectively. Exemplary conservative substitutions are provided in Table 1 below. [Table 1]
[0053] "Non-conservative substitution" refers to the substitution of an amino acid in a polypeptide with an amino acid having significantly different side chain characteristics. Non-conservative substitutions may use amino acids from defined groups rather than within defined groups, and affect (a) the structure of the peptide backbone in the region of the substitution (e.g., substitution of glycine with proline), (b) the charge or hydrophobicity, or (c) the bulk of the side chain. By way of example, and not limitation, exemplary non-conservative substitutions can be an acidic amino acid substituted with a basic or aliphatic amino acid; an aromatic amino acid substituted with a small amino acid; and a hydrophilic amino acid substituted with a hydrophobic amino acid.
[0054] "Deletion" refers to the modification of a polypeptide by the removal of one or more amino acids from a reference polypeptide. Deletions can include the removal of one or more amino acids, two or more amino acids, five or more amino acids, ten or more amino acids, fifteen or more amino acids, or twenty or more amino acids, up to a maximum of 10% of the total number of amino acids, or up to a maximum of 20% of the total number of amino acids, from the reference enzyme while retaining the enzymatic activity of the engineered nitroaldolase enzyme and / or retaining improved properties. Deletions can relate to the internal portion and / or the terminal portions of the polypeptide. In various embodiments, deletions can include contiguous segments or can be discontinuous.
[0055] "Insertion" refers to the modification of a polypeptide by the addition of one or more amino acids to a reference polypeptide. In some embodiments, the improved engineered nitroaldolase enzyme includes the insertion of one or more amino acids into a naturally occurring nitroaldolase polypeptide as well as the insertion of one or more amino acids into other improved nitroaldolase polypeptides. Insertions can be in the internal portion of the polypeptide or can be at the carboxy or amino terminus. As used herein, insertions include fusion proteins as are known in the art. Insertions can be a contiguous segment of amino acids in a naturally occurring polypeptide or can be separated by one or more amino acids.
[0056] "Fragment" is used herein to refer to a polypeptide that has an amino-terminal and / or carboxy-terminal deletion, but the remaining amino acid sequence is identical to the corresponding position in the sequence. Fragments can be at least 14 amino acids in length, at least 20 amino acids in length, at least 50 amino acids in length or longer, and can be up to 70%, 80%, 90%, 95%, 98% and 99% of the full-length nitroaldolase polypeptide, e.g., the polypeptide of SEQ ID NO: 4, or the engineered nitroaldolase polypeptide provided by the even-numbered sequences among SEQ ID NOs: 4 to 1484.
[0057] "Isolated polypeptide" refers to a polypeptide that is substantially separated from other contaminants (e.g., proteins, lipids and polynucleotides) that naturally accompany it. This term encompasses polypeptides that have been removed or purified from their natural environment or expression system (e.g., host cells or in vitro synthesis). The engineered nitroaldolase enzyme may be present intracellularly, may be present in the cell culture medium, or may be prepared in various forms such as lysates or isolated preparations. Thus, in some embodiments, the engineered nitroaldolase enzyme may be an isolated polypeptide.
[0058] "Substantially pure polypeptide" refers to a composition in which the polypeptide species is the dominant species present (i.e., it is present in a greater amount than any other individual macromolecular species in the composition, on a molar or weight basis), and generally, when the species of interest constitutes at least about 50 percent in molar or weight percent of the macromolecular species present, it is a substantially purified composition. Generally, a substantially pure nitroaldolase composition is about 60% or higher, about 70% or higher, about 80% or higher, about 90% or higher, about 95% or higher, and about 98% or higher in molar or weight percent of all macromolecular species present in the composition. In some embodiments, the species of interest is purified to essential homogeneity where the composition consists essentially of a single macromolecular species (i.e., contaminating species cannot be detected in the composition by conventional detection methods). Solvent species, small molecules (<500 daltons), and elemental ion species are not considered macromolecular species. In some embodiments, the isolated and engineered nitroaldolase enzyme is a substantially pure polypeptide composition.
[0059] "Stereoselective" refers to a preference for the formation of one stereoisomer over another in a chemical or enzymatic reaction. Stereoselectivity can be partial if one stereoisomer is preferred over another or complete if only one stereoisomer is formed. When the stereoisomers are enantiomers, the stereoselectivity is called enantioselectivity, which is the fraction of one enantiomer in the sum of both enantiomers (typically reported as a percentage). This is generally, in the art, instead reported as the enantiomeric excess (e.e.), calculated therefrom according to the formula [major enantiomer - minor enantiomer] / [major enantiomer + minor enantiomer] (typically as a percentage). When the stereoisomers are diastereoisomers, the stereoselectivity is called diastereoselectivity, which is the fraction of one diastereomer in a mixture of two diastereomers (typically reported as a percentage) and is generally, instead, reported as the diastereomeric excess (d.e.). Enantiomeric excess and diastereomeric excess are types of stereoisomeric excess.
[0060] "Highly stereoselective" refers to a chemical or enzymatic reaction that can convert a substrate(s) (e.g., substrate compounds (2) and (3)) to the corresponding amine product(s) (e.g., compound (1)) with at least about 85% stereoisomeric excess.
[0061] As used herein, "improved enzyme properties" refers to at least one improved property of an enzyme. In some embodiments, the present invention provides engineered nitroaldolase polypeptides that exhibit an improvement in any enzyme property as compared to a reference nitroaldolase polypeptide and / or a nitroaldolase polypeptide and / or another engineered nitroaldolase polypeptide. For the engineered nitroaldolase polypeptides described herein, the comparison is generally made with the parental enzyme from which the nitroaldolase is derived, although in some embodiments, the reference enzyme can be another improved engineered nitroaldolase. Thus, the level of "improvement" can be determined and compared among various nitroaldolase polypeptides, including wild-type as well as engineered nitroaldolases. Properties that can be improved include, but are not limited to, enzyme activity (which can be represented by the percentage of substrate conversion), thermal stability, solvent stability, pH activity profile, cofactor requirements, insensitivity to inhibitors (e.g., substrate or product inhibition), stereospecificity, and / or stereoselectivity (including enantioselectivity).
[0062] "Increase in enzyme activity" refers to an improved property of an engineered nitroaldolase polypeptide that can be represented by an increase in specific activity (e.g., product produced / time / protein weight) or an increase in the percentage of substrate conversion to product (e.g., the percentage of starting amount of substrate converted to product within a specified period using a specified amount of nitroaldolase) as compared to a reference nitroaldolase enzyme. Exemplary methods for determining enzyme activity are provided in the Examples. K m , V max or k catAny property related to enzyme activity, including classical enzyme properties that can lead to an increase in enzyme activity, can be affected. The improvement in enzyme activity can range from about 1.2 times the enzyme activity of the corresponding parental enzyme to 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold or higher enzyme activity of a naturally occurring or otherwise engineered nitroaldolase in which the nitroaldolase polypeptide is induced. Nitroaldolase activity can be measured by any one of standard assays, for example, by monitoring changes in the properties of a substrate, cofactor or product. In some embodiments, the amount of product produced can be measured by gas chromatography (GC). Comparison of enzyme activities is performed using a defined preparation of the enzyme, a defined assay under defined conditions, and one or more defined substrates, as further described in detail herein. Generally, when lysates are compared, the number of cells and the amount of protein assayed are determined, and the use of the same expression system and the same host cell is also determined to minimize variation in the amount of enzyme present in the lysate produced by the host cell.
[0063] "Conversion" refers to the enzymatic conversion of a substrate to the corresponding product. "Percent conversion" refers to the percentage of substrate that is converted to product within a period under specified conditions. Thus, the "enzyme activity" or "activity" of a nitroaldolase polypeptide can be expressed as the "percent conversion" of substrate to product.
[0064] "Thermostable" refers to a nitroaldolase polypeptide that maintains similar activity (e.g., higher than 60% - 80%) compared to the parental enzyme exposed to the same high temperature after exposure to a high temperature (e.g., 40 - 80°C) for a period (e.g., 0.5 - 24 hours).
[0065] "Solvent-stable" refers to a nitroaldolase polypeptide that maintains similar activity (e.g., higher than 60% - 80%) compared to the parental enzyme exposed to the same concentration of the same solvent after exposure to various concentrations (e.g., 5% - 99%) of solvents (such as ethanol, isopropyl alcohol, dimethyl sulfoxide (DMSO), tetrahydrofuran, 2-methyltetrahydrofuran, acetone, toluene, butyl acetate, methyl tert-butyl ether, etc.) for a certain period (e.g., 0.5 - 24 hours).
[0066] "Thermo- and solvent-stable" refers to a nitroaldolase polypeptide that is both thermostable and solvent-stable.
[0067] The term "stringent hybridization conditions" is used herein to refer to conditions under which a nucleic acid hybrid is stable. As is known to those of skill in the art, the stability of a hybrid is reflected in the melting temperature (T m ) of the hybrid. Generally, the stability of a hybrid is a function of ionic strength, temperature, G / C content, and the presence of chaotropic agents. The T mThe value can be calculated using known methods for predicting melting temperature (see, e.g., Baldino et al., Meth. Enzymol., 168:761-777
[1989] ; Bolton et al., Proc. Natl. Acad. Sci. USA 48:1390
[1962] ; Bresslauer et al., Proc. Natl. Acad. Sci. USA 83:8893-8897
[1986] ; Freier et al., 1990, Proc. Natl. Acad. Sci. USA 83:9373-9377
[1986] ; Kierzek et al., Biochem., 25:7840-7846
[1986] ; Rychlik et al., Nucl. Acids Res., 18:6409-6412
[1990] (erratum, Nucl. Acids Res., 19:698
[1991] ); Sambrook et al., supra); Suggs et al., 1981, in Developmental Biology Using Purified Genes, Brown et al. [eds.], pp. 683-693, Academic Press, Cambridge, MA
[1981] ; and Wetmur, Crit. Rev. Biochem. Mol. Biol. 26:227-259
[1991] ). In some embodiments, the polynucleotide encodes a polypeptide disclosed herein and hybridizes under defined conditions, such as moderately stringent or highly stringent conditions, to a complementary sequence of a sequence encoding the engineered nitroaldolase enzyme of the invention.
[0068] "Hybridization stringency" refers to hybridization conditions such as washing conditions during nucleic acid hybridization. Generally, the hybridization reaction is carried out under conditions of lower stringency, followed by various, but more highly stringent, washes. The term "moderately stringent hybridization" refers to conditions that allow the target DNA to bind to a complementary nucleic acid having at least about 60% identity to the target DNA, preferably about 75% identity, about 85% identity, and more than about 90% identity to the target polynucleotide. Exemplary moderately stringent conditions are hybridization in 50% formamide, 5× Denhardt's solution, 5× SSPE, 0.2% SDS at 42 °C, followed by washing in 0.2× SSPE, 0.2% SDS at 42 °C. "High stringency hybridization" generally refers to conditions that are about 10 °C or less below the thermal melting temperature T m determined under solution conditions for a defined polynucleotide sequence. In some embodiments, high stringency conditions refer to conditions that allow hybridization of only nucleic acid sequences that form stable hybrids at 65 °C in 0.018 M NaCl (i.e., if the hybrid is not stable at 65 °C in 0.018 M NaCl, it will not be stable under high stringency conditions as contemplated herein). High stringency conditions can be provided, for example, by hybridization at 42 °C in 50% formamide, 5× Denhardt's solution, 5× SSPE, 0.2% SDS, followed by washing at 65 °C in 0.1× SSPE and 0.1% SDS. Another high stringency condition is hybridization at 65 °C in 5× SSC containing 0.1% (w / v) SDS and washing at 65 °C in 0.1× SSC containing 0.1% SDS. Other high stringency hybridization conditions, as well as moderately stringent conditions, are described in the references above.
[0069] A "heterologous" polynucleotide is any polynucleotide that is introduced into a host cell by experimental techniques, removed from the host cell, subjected to experimental manipulation, and then reintroduced into the host cell.
[0070] "Codon-optimized" refers to a change in the codons of a polynucleotide encoding a protein to the codons preferentially used in a particular organism such that the encoded protein is efficiently expressed in that organism. The genetic code is degenerate in that most amino acids are represented by several codons called "synonymous" or "equivalent" codons, but it is well known that the codon usage frequency by a particular organism is not random and is biased towards certain codon triplets. This codon usage bias can be more pronounced with respect to a given gene, genes of common function or ancestral origin, proteins highly expressed relative to low-copy number proteins, and the aggregated protein-coding regions of an organism's genome. In some embodiments, the polynucleotide encoding the nitroaldolase enzyme can be codon-optimized for optimal production from the host organism selected for expression.
[0071] As used herein, "preferred, optimal, codons with high codon usage bias" refers, synonymously, to codons that are used more frequently in the protein-coding region than other codons that encode the same amino acid. Preferred codons may be determined with respect to the codon usage frequency in a single gene, a set of genes of common function or origin, highly expressed genes, the codon frequency in the aggregated protein-coding region of an entire organism, the codon frequency in the aggregated protein-coding region of related organisms, or a combination thereof. Codons whose frequency increases with the gene expression level are typically codons that are optimal for expression. For example, various methods for determining codon frequencies (e.g., codon usage frequency, relative synonymous codon usage frequency) and codon preference in a particular organism, including multivariate analysis using cluster analysis or correspondence analysis, as well as various methods for determining the effective number of codons used in a gene, are known (see, e.g., GCG CodonPreference, Genetics Computer Group Wisconsin Package; CodonW, Peden, University of Nottingham; McInerney, Bioinform., 14:372-73
[1998] ; Stenico et al., Nucl. Acids Res., 22:2437-46
[1994] ; Wright, Gene 87:23-29
[1990] ). Codon usage frequency tables for many different organisms are available (see, e.g., Wada et al., Nucl. Acids Res., 20:2111-2118
[1992] ; Nakamura et al., Nucl. Acids Res., 28:292
[2000] ; Duret, et al., supra; Henaut and Danchin, in Escherichia coli and Salmonella, Neidhardt, et al. (eds.), ASM Press, Washington D.C., p. 2047-2066
[1996] ).The data source for obtaining codon usage frequencies can depend on any available nucleotide sequence that can encode a protein. These datasets include nucleic acid sequences that are actually known to encode an expressed protein (e.g., a complete protein coding sequence - CDS), encode an expressed sequence tag (ESTs), or encode a predicted coding region of a genomic sequence (see, for example, Mount, Bioinformatics: Sequence and Genome Analysis, Chapter 8, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.
[2001] ; Uberbacher, Meth. Enzymol., 266:259-281
[1996] ; and Tiwari et al., Comput. Appl. Biosci., 13:263-270
[1997] ).
[0072] "Control sequence" is defined herein to include all components that are necessary or advantageous for the expression of the polynucleotides and / or polypeptides of the present invention. Each control sequence may be native or foreign to the nucleic acid sequence encoding the polypeptide. Such control sequences include, but are not limited to, leader, polyadenylation sequence, propeptide sequence, promoter, signal peptide sequence and transcription terminator. At a minimum, the control sequences include a promoter, as well as transcription and translation stop signals. The control sequences may also be provided with linkers to introduce specific restriction sites facilitating ligation of the control sequences to the coding region of the nucleic acid sequence encoding the polypeptide.
[0073] "Operably linked" is defined herein as a configuration in which a control sequence is appropriately (i.e., in a functional relationship) positioned relative to a polynucleotide of interest such that the control sequence directs or regulates the expression of the polynucleotide of interest and / or the polypeptide.
[0074] A "promoter sequence" refers to a nucleic acid sequence recognized by a host cell for the expression of a polynucleotide of interest, such as a coding sequence. The promoter sequence contains transcriptional control sequences that mediate the expression of the polynucleotide of interest. A promoter can be any nucleic acid sequence that exhibits transcriptional activity in a selected host cell, including variants, truncated forms, and hybrid promoters, and can be obtained from a gene encoding an extracellular or intracellular polypeptide that is either homologous or heterologous to the host cell.
[0075] "Suitable reaction conditions" refer to the conditions (e.g., ranges of enzyme loading, substrate loading, cofactor loading, temperature, pH, buffer, co-solvent, etc.) in a biocatalytic reaction solution under which the nitroaldolase polypeptide of the present invention can convert one or more substrate compounds into a product compound (e.g., the conversion of compound (1) and compound (2) to compound (3) as shown in Scheme 1). Exemplary "suitable reaction conditions" are provided in the present invention and illustrated by the examples.
[0076] A "composition" refers to a mixture or combination of one or more substances in which each substance or component of the composition retains its individual properties. As used herein, a biocatalytic composition refers to a combination of one or more substances useful for biocatalysis.
[0077] "Loading" in "compound loading" or "enzyme loading" or "cofactor loading", etc. refers to the concentration or amount of a component in the reaction mixture at the start of the reaction.
[0078] In the context of a biocatalytic process, a "substrate" refers to a compound or molecule that is acted upon by a biocatalyst. For example, compound (1) and compound (2) are substrates of nitroaldolase.
[0079] In the context of a biocatalytic process, a "product" refers to a compound or molecule that results from the action of a biocatalyst. For example, compound (3) is the product of the nitroaldolase-mediated conversion of compound (1) and compound (2).
[0080] "Alkyl" refers to a saturated hydrocarbon group that is either linear or branched and has 1 to 18 carbon atoms (including both end values), more preferably 1 to 8 carbon atoms (including both end values), and most preferably 1 to 6 carbon atoms (including both end values). Alkyl having a specified number of carbon atoms is indicated in parentheses (for example, (C1-C6)alkyl refers to alkyl having 1 to 6 carbon atoms).
[0081] "Alkenyl" refers to a hydrocarbon group that is either linear or branched, contains at least one double bond, and optionally contains more than one double bond, and has 2 to 12 carbon atoms (including both end values).
[0082] "Alkynyl" refers to a hydrocarbon group that is either linear or branched, contains at least one triple bond, optionally contains two or more triple bonds, and additionally optionally contains one or more double bond moieties, and has 2 to 12 carbon atoms including both ends.
[0083] "Alkylene" refers to a linear or branched divalent hydrocarbon radical having 1 to 18 carbon atoms (including both end values), more preferably 1 to 8 carbon atoms (including both end values), and most preferably 1 to 6 carbon atoms (including both end values), which is optionally substituted with one or more suitable substituents. Exemplary "alkylene" includes, but is not limited to, methylene, ethylene, propylene, butylene, etc.
[0084] "Alkenylene" refers to a linear or branched divalent hydrocarbon radical having 2 to 12 carbon atoms (including both end values) and one or more carbon-carbon double bonds, more preferably 2 to 8 carbon atoms (including both end values), and most preferably 2 to 6 carbon atoms (including both end values), which is optionally substituted with one or more suitable substituents.
[0085] "Heteroalkyl", "heteroalkenyl", and "heteroalkynyl" each refer to an alkyl, alkenyl, and alkynyl as defined herein, where one or more of the carbon atoms are each independently replaced by the same or different heteroatoms or heteroatom groups. Heteroatoms and / or heteroatom groups that can replace carbon atoms include, but are not limited to, -O-, -S-, -S-O-, -NR γ -, -PH-, -S(O)-, -S(O)2-, -S(O)NR γ -, -S(O)2NR γ -, etc., including combinations thereof, where each R γ is independently selected from hydrogen, alkyl, cycloalkyl, heterocycloalkyl, aryl, and heteroaryl.
[0086] "Aryl" refers to an unsaturated aromatic carbocyclic group having 6 to 12 carbon atoms (including both end values), having a single ring (e.g., phenyl) or multiple fused rings (e.g., naphthyl or anthryl). Exemplary aryls include phenyl, pyridyl, naphthyl, etc.
[0087] "Arylalkyl" preferably refers to an alkyl substituted with an aryl (i.e., an aryl-alkyl group) having 1 to 6 carbon atoms (including both end values) in the alkyl portion and 6 to 12 carbon atoms (including both end values) in the aryl portion. Such arylalkyl groups are exemplified by benzyl, phenethyl, etc.
[0088] "Aryloxy" refers to an -OR λ group (where R λ in the formula is an aryl group that may be optionally substituted).
[0089] "Cycloalkyl" refers to a cyclic alkyl group having 3 to 12 carbon atoms (including both end values) with a single cyclic ring or multiple fused rings, which may be optionally substituted with 1 to 3 alkyl groups. Exemplary cycloalkyl groups include, but are not limited to, monocyclic structures such as cyclopropyl, cyclobutyl, cyclopentyl, cyclooctyl, 1-methylcyclopropyl, 2-methylcyclopentyl, 2-methylcyclooctyl, etc., or polycyclic structures, and the polycyclic structures include bridged ring systems such as adamantyl, etc.
[0090] "Cycloalkylalkyl" preferably refers to an alkyl substituted with cycloalkyl (i.e., a cycloalkyl-alkyl group) having 1 to 6 carbon atoms (including both end values) in the alkyl part and 3 to 12 carbon atoms (including both end values) in the cycloalkyl part. Such cycloalkylalkyl groups are exemplified by cyclopropylmethyl, cyclohexylethyl, etc.
[0091] "Amino" refers to the group -NH2. Substituted amino refers to the groups -NHR η , NR η R η and NR η R η R η (where each R η is independently selected from substituted or unsubstituted alkyl, cycloalkyl, cycloheteroalkyl, alkoxy, aryl, heteroaryl, heteroarylalkyl, acyl, alkoxycarbonyl, sulfanyl, sulfinyl, sulfonyl, etc.). Typical amino groups include, but are not limited to, dimethylamino, diethylamino, trimethylammonium, triethylammonium, methylsulfonylamino, furanyl-oxy-sulfamino, etc.
[0092] "Aminoalkyl" refers to an alkyl group in which one or more of the hydrogen atoms are replaced by one or more amino groups including substituted amino groups.
[0093] "Aminocarbonyl" refers to -C(O)NH2. Substituted aminocarbonyl is -C(O)NR η R η (wherein the amino group NR η R η is as defined herein).
[0094] "Oxy" refers to the divalent group -O-, which may have various substituents to form different oxy groups including ethers and esters.
[0095] "Alkoxy" or "alkyloxy" refers to the group -OR ζ (wherein R ζ is an alkyl group, which may be optionally substituted alkyl group) and is used synonymously herein.
[0096] "Carboxy" refers to -COOH.
[0097] "Carbonyl" refers to -C(O)-, which may have various substituents to form different carbonyl groups including acids, acid halides, aldehydes, amides, esters and ketones.
[0098] "Carboxyalkyl" refers to an alkyl in which one or more of the hydrogen atoms are replaced by one or more carboxy groups.
[0099] "Aminocarbonylalkyl" refers to an alkyl substituted with an aminocarbonyl group as defined herein.
[0100] "Halogen" or "halo" refers to fluoro, chloro, bromo and iodo.
[0101] "Haloalkyl" refers to an alkyl group in which one or more of the hydrogen atoms are replaced by halogen. Thus, the term "haloalkyl" is intended to include monohaloalkyl, dihaloalkyl, trihaloalkyl, etc. up to perhaloalkyl. For example, the expression "(C1-C2) haloalkyl" includes 1-fluoromethyl, difluoromethyl, trifluoromethyl, 1-fluoroethyl, 1,1-difluoroethyl, 1,2-difluoroethyl, 1,1,1-trifluoroethyl, perfluoroethyl, etc.
[0102] "Hydroxy" refers to -OH.
[0103] "Hydroxyalkyl" refers to an alkyl group in which one or more of the hydrogen atoms are replaced by one or more hydroxy groups.
[0104] "Thiol" or "sulfanyl" refers to -SH. Substituted thiol or sulfanyl is -S-R η (wherein R η is alkyl, aryl or other suitable substituent)
[0105] "Alkylthio" refers to -SR ζ (wherein R ζ is alkyl which may be optionally substituted). Typical alkylthio groups include, but are not limited to, methylthio, ethylthio, n-propylthio, etc.
[0106] "Alkylthioalkyl" refers to an alkyl substituted with an alkylthio group -SR ζ (wherein R ζ is alkyl which may be optionally substituted).
[0107] "Sulfonyl" refers to -SO2-. Substituted sulfonyl is -SO2-R η (wherein R η is alkyl, aryl or other suitable substituent)
[0108] "Alkylsulfonyl" refers to -SO2-R ζ (wherein R ζ is alkyl which may be optionally substituted). Typical alkylsulfonyl groups include, but are not limited to, methylsulfonyl, ethylsulfonyl, n-propylsulfonyl, and the like.
[0109] "Alkylsulfonylalkyl" refers to alkyl substituted with an alkylsulfonyl group -SO2-R ζ (wherein R ζ is alkyl which may be optionally substituted).
[0110] "Heteroaryl" refers to an aromatic heterocyclic group having 1 to 10 carbon atoms (including both ends) and 1 to 4 heteroatoms (including both ends) selected from oxygen, nitrogen, and sulfur in the ring. Such heteroaryl groups may have a single ring (e.g., pyridyl or furyl) or multiple fused rings (e.g., indolizinyl or benzothienyl).
[0111] "Heteroarylalkyl" preferably refers to alkyl substituted with heteroaryl (i.e., heteroaryl-alkyl group) having 1 to 6 carbon atoms (including both ends) in the alkyl part and 5 to 12 ring atoms (including both ends) in the heteroaryl part. Such heteroarylalkyl groups are exemplified by pyridylmethyl and the like.
[0112] "Heterocyclic ring", "heterocyclic ring formula", and synonymously "heterocycloalkyl" refer to a saturated or unsaturated group having a single ring or multiple condensed rings, where the number of carbocyclic ring atoms is from 2 to 10 (including both end values), and the number of hetero ring atoms selected from nitrogen, sulfur, or oxygen in the ring is from 1 to 4 (including both end values). Such heterocyclic ring groups may have a single ring (e.g., piperidinyl or tetrahydrofuryl) or multiple condensed rings (e.g., indolinyl, dihydrobenzofuran, or quinuclidinyl). Examples of heterocycles include, but are not limited to, furan, thiophene, thiazole, oxazole, pyrrole, imidazole, pyrazole, pyridine, pyrazine, pyrimidine, pyridazine, indolizine, isoindole, indole, indazole, purine, quinolidine, isoquinoline, quinoline, phthalazine, naphthylpyridine, quinoxaline, quinazoline, cinnoline, pteridine, carbazole, carboline, phenanthridine, acridine, phenanthroline, isothiazole, phenazine, isoxazole, phenoxazine, phenothiazine, imidazolidine, imidazoline, piperidine, piperazine, pyrrolidine, indoline, etc.
[0113] "Heterocycloalkylalkyl" preferably refers to an alkyl substituted with a heterocycloalkyl (i.e., a heterocycloalkyl-alkyl group) having 1 to 6 (including both end values) carbon atoms in the alkyl portion and 3 to 12 (including both end values) ring atoms in the heterocycloalkyl portion.
[0114] "Membered ring" is intended to encompass any cyclic structure. The number before the term "membered" indicates the number of skeletal atoms constituting the ring. Thus, for example, cyclohexyl, pyridine, pyran, and thiopyran are 6-membered rings, and cyclopentyl, pyrrole, furan, and thiophene are 5-membered rings.
[0115] "Fused bicyclic ring", as used herein, refers to both unsubstituted and substituted carbocyclic and / or heterocyclic ring moieties having from 5 to 8 atoms in each ring, with these rings having two common atoms.
[0116] As used herein with respect to the above chemical groups, "optionally substituted" means that the position of the chemical group occupied by hydrogen is replaced by another atom exemplified by but not limited to carbon, oxygen, nitrogen or sulfur (unless otherwise specified), or by a chemical group exemplified by but not limited to hydroxy, oxo, nitro, methoxy, ethoxy, alkoxy, substituted alkoxy, trifluoromethoxy, haloalkoxy, fluoro, chloro, bromo, iodo, halo, methyl, ethyl, propyl, butyl, alkyl, alkenyl, alkynyl, substituted alkyl, trifluoromethyl, haloalkyl, hydroxyalkyl, alkoxyalkyl, thio, alkylthio, acyl, carboxy, alkoxycarbonyl, carboxamide, substituted carboxamide, alkylsulfonyl, alkylsulfinyl, alkylsulfonylamino, sulfonamide, substituted sulfonamide, cyano, amino, substituted amino, alkylamino, dialkylamino, aminoalkyl, acylamino, amidino, amidoximo, hydroxamoyl, phenyl, aryl, substituted aryl, aryloxy, arylalkyl, arylalkenyl, arylalkynyl, pyridyl, imidazolyl, heteroaryl, substituted heteroaryl, heteroaryloxy, heteroarylalkyl, heteroarylalkenyl, heteroarylalkynyl, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloalkyl, cycloalkenyl, cycloalkylalkyl, substituted cycloalkyl, cycloalkyloxy, pyrrolidinyl, piperidinyl, morpholino, heterocycle, (heterocycle)oxy, and (heterocycle)alkyl, and the preferred heteroatoms are oxygen, nitrogen and sulfur.Furthermore, when free valences are present on the chemical groups of these substituents, these substituents may be further substituted with alkyl, cycloalkyl, aryl, heteroaryl and / or heterocyclic groups; when these free valences are present on carbon, these carbons may be further substituted with halogen and by oxygen-bonded substituents, nitrogen-bonded substituents or sulfur-bonded substituents; and when multiple such free valences are present, these groups may be connected such that they form a ring either by direct formation of a bond or by formation of a bond to a new heteroatom, preferably oxygen, nitrogen or sulfur. It is further contemplated that the above substitutions can be made provided that replacement of a hydrogen with a substituent does not introduce unacceptable instability to the molecules of the present invention and is otherwise chemically reasonable. It will be understood by those skilled in the art that for any of the plurality of optionally substituted chemical groups described, only sterically realistic and / or synthetically feasible chemical groups are intended to be included. As used herein, "optionally substituted" refers to the term chemical group or all subsequent modifiers in a series of chemical groups. For example, in the term "optionally substituted arylalkyl", the "alkyl" and "aryl" portions of the molecule may or may not be substituted; for a series of "optionally substituted alkyl, cycloalkyl, aryl and heteroaryl", the alkyl, cycloalkyl, aryl and heteroaryl groups may or may not be substituted independently of each other.
[0117] "Reaction", as used herein, refers to the process(es) by which one or more substances or compounds or substrates are converted into one or more different substances, compounds.
[0118] Conversion of nitromethane and trifluoroacetone to β-nitroalcohol The compound (4), which is (S)-3-amino-1,1,1-trifluoro-2-methylpropan-2-ol, is an intermediate in the synthesis of novel compounds for the treatment of cystic fibrosis. [Chemical formula]
[0119] The enzymatic synthesis of compound (4), a β-hydroxy-α-amino alcohol, is an attractive alternative to traditional chemical synthesis. The use of enzyme biocatalysts can reduce chemical waste and enable more efficient synthesis methods.
[0120] The production of chiral β-amino alcohols via β-hydroxy-α-amino acid intermediates using a two-enzyme cascade has been reported by others (Duckers et al. Appl Microbiol Biotechnol, 2010, 88:409-424). In the first step, threonine aldolase catalyzes the conversion of glycine (or another nucleophile) and acetaldehyde (or another electrophile) to threonine via a PLP-dependent mechanism. In the second step, amino acid decarboxylase catalyzes the removal of the carboxyl group to produce an amino alcohol (also a PLP-dependent mechanism). The inventors previously reported a two-enzyme cascade for the production of compound (4) utilizing engineered threonine aldolase and engineered amino acid decarboxylase (U.S. Patent Application No. 63 / 285,377).
[0121] However, other biocatalytic routes have been proposed and may offer additional utility and efficiency under industrial synthesis conditions. These include the use of oxynitrilases (also known as hydroxynitrile lyases) to produce cyanohydrin or β-nitroalcohol intermediates, followed by reduction of the resulting compound (4) to β-hydroxy-α-aminoalcohol. This disclosure details the use of oxynitrilases for the nitroaldol conversion of nitromethane and trifluoroacetone to their β-nitroalcohol intermediates. In this disclosure, the term nitroaldolase is used to describe engineered oxynitrilases / hydroxynitrile lyases in order to emphasize the improvement of enzymatic functionality towards the nitroaldol reaction.
[0122] Oxynitrilases are found primarily in plants where the conversion of naturally occurring cyanogenic glycosides or cyanolipid cyanohydrins to cyanide is used as a defense mechanism against herbivores and microorganisms, often in stone fruit plants (e.g., almonds, apricots). Both (S)-selective oxynitrilases and (R)-selective oxynitrilases are known. Since oxynitrilases are members of the α / β hydrolase fold family, they have a catalytic triad similar to esterases. The native substrates are thought to be acetone and benzaldehyde. Some oxynitrilases contain non-catalytic FAD that acts as a structural cofactor.
[0123] The proposed mechanism of oxynitrilases shows an activated serine in the catalytic triad that acts as a general acid / base and not as a nucleophile (Gruber et al. J. Biol. Chem., 2004, 279:20501-20510). This is unusual for members of the α / β hydrolase fold family where serine normally acts as a nucleophile to form an acyl-enzyme intermediate.
[0124] The present disclosure provides novel biocatalysts and related methods useful for chiral β - amino alcohol synthesis via β - nitro alcohol intermediates, followed by chemical reduction. The nitroaldolase biocatalysts of the present disclosure are engineered variants of the polypeptide (SEQ ID NO: 2) encoded by the homologous gene (SEQ ID NO: 1) from Baliospermum montanum. These engineered polypeptides can catalyze the conversion of trifluoroacetone (compound (1)) and nitromethane (compound (2)) to compound (4) via intermediate compound (3), as depicted in Scheme 1 below.
Chemical formula
[0125] In some embodiments, the present disclosure provides a nitroaldolase enzyme having improved activity in the production of compound (3) as compared to a reference polypeptide. In some embodiments, the engineered nitroaldolase enzyme has the activity of Scheme 1.
[0126] As described herein, the engineered polypeptide exhibits stereoselectivity and, therefore, one or more chiral centers of the product can be established using the nitroaldolase of Scheme 1. The nitroaldolase of the present disclosure produces the (R) or (S) enantiomer.
[0127] In some embodiments, the present disclosure provides an engineered nitroaldolase enzyme having improved stereoselectivity for the (S)-β - nitro alcohol product (compound (3)) as compared to a reference polypeptide. In some embodiments, the present disclosure provides an amino acid sequence having at least 80% sequence identity to the amino acid reference sequences of SEQ ID NO: 2, 24, 496, 860, 1120, or 1348, and further comprising one or more amino acid residue differences as compared to the reference amino acid sequence, wherein the engineered nitroaldolase polypeptide has increased enantioselectivity for compound (3).
[0128] In some embodiments, the present disclosure provides an engineered polypeptide comprising an amino acid sequence having at least 80% sequence identity to the amino acid reference sequence of SEQ ID NO: 2, 24, 496, 860, 1120, or 1348, and further comprising one or more amino acid residue differences as compared to the reference amino acid sequence, wherein the engineered nitroaldolase polypeptide has increased activity in the conversion of compounds (1) and (2) to compound (3).
[0129] Engineered nitroaldolase polypeptide The present invention provides a nitroaldolase polypeptide, a polynucleotide encoding the polypeptide, a method for preparing the polypeptide, and a method for using the polypeptide. It should be understood that when the description relates to a polypeptide, the description may also relate to the polynucleotide encoding the polypeptide.
[0130] Suitable reaction conditions for performing the desired reaction by virtue of the above-improved properties of the engineered polypeptide include conditions including the concentration or amount of the polypeptide, substrate, co-substrate, buffer, solvent, pH, temperature, and reaction time, and / or conditions associated with the polypeptide immobilized on a solid support, and can be determined as further described below and in the examples.
[0131] In some embodiments, particularly with respect to the stereoselective conversion of compounds (1) and (2) to compound (3), exemplary engineered nitroaldolase polypeptides having improved properties comprise an amino acid sequence having one or more residue differences at the residue positions shown in Tables 2-1, 2-2, 3-1, 3-2, 4-1, 4-2, 5-1, 5-2, 6-1, 6-2, 7-1, and 7-2 as compared to SEQ ID NO: 2, 24, 496, 860, 1120, or 1348.
[0132] Exemplary information on the structure and function of the non-naturally occurring (or engineered) polypeptides of the present invention is based on the conversion of compound (1) and compound (2) to compound (3), as further described in the Examples, the results of which are shown in Tables 2-1, 2-2, 3-1, 3-2, 4-1, 4-2, 5-1, 5-2, 6-1, 6-2, 7-1 and 7-2 below. The odd-numbered sequence identifiers (i.e., sequence numbers) in these tables refer to nucleotide sequences that encode the amino acid sequences provided by the even-numbered sequence numbers in these tables. Exemplary sequences are provided in the accompanying electronic format sequence listing file of the present invention, which is hereby incorporated by reference into this specification. Amino acid residue differences are based on comparison to the reference sequences of SEQ ID NO: 2, 24, 496, 860, 1120 and / or 1348, as described above.
[0133] A nitroaldolase (SEQ ID NO: 2) homolog from Baliospermum montanum was selected based on the conversion of compound (1) and compound (2) to compound (3). SEQ ID NO: 1 is a codon-optimized polynucleotide for expression in Escherichia coli, synthesized based on the polypeptide sequence of SEQ ID NO: 2.
[0134] The activity of each engineered nitroaldolase polypeptide, compared to the reference polypeptides of SEQ ID NO: 2, 24, 496, 860, 1120 and / or 1348, was determined as the conversion of substrate, as described in the Examples herein. In some embodiments, shake flask powder (SFP) was used as a secondary screen to evaluate the properties of the engineered nitroaldolase polypeptides, and the results are provided in the Examples. In some embodiments, the SFP form provides a more purified powder preparation of the engineered polypeptide and may contain up to 30% of the engineered polypeptide of the total protein.
[0135] In some embodiments, the specific enzyme characteristics are related to residue differences at the residue positions shown herein when compared to SEQ ID NOs: 2, 24, 496, 860, 1120 and / or 1348. In some embodiments, residue differences that affect polypeptide expression can be used to increase the expression of the engineered nitroaldolase polypeptide.
[0136] Based on the guidance provided herein, any of the exemplary engineered polypeptides comprising the even-numbered sequences of SEQ ID NOs: 4 - 1484 are used as starting amino acid sequences for synthesizing other engineered nitroaldolase polypeptides by subsequent rounds of evolution that incorporate new combinations of various amino acid differences with other polypeptides at, for example, those in Tables 2-1, 2-2, 3-1, 3-2, 4-1, 4-2, 5-1, 5-2, 6-1, 6-2, 7-1 and 7-2 and at other residue positions described herein. Further improvements can be achieved by including amino acid differences at residue positions that were maintained as invariant throughout previous rounds of evolution.
[0137] In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 2, which is the reference sequence, and SEQ ID NO: 7 / 9, 7 / 9 / 33 / 68 / 72 / 158, 7 / 9 / 33 / 72 / 175 / 197 / 230, 7 / 9 / 33 / 72 / 230, 7 / 9 / 68 / 72 / 197, 7 / 9 / 158 / 175, 7 / 9 / 175, 7 / 9 / 197, 7 / 33, 7 / 68 / 72 / 175, 7 / 72, 7 / 72 / 97 / 175 / 230, 9, 9 / 11 / 33 / 72 / 197 / 230, 9 / 33 / 68 / 158 / 197, 9 / 33 / 72 / 230, 9 / 33 / 97 / 230, 9 / 68 / 72, 9 / 97 / 175 / 202 / 218, 9 / 175, 9 / 202, 11, 13, 14 / 33 / 68 / 72 / 97, 18, 18 / 33 / 68 / 97 / 158 / 175 / 239, 18 / 33 / 72 / 97 / 158 / 175, 18 / 33 / 72 / 97 / 158 / 239, 18 / 33 / 72 / 97 / 175 / 202, 18 / 33 / 72 / 175 / 202 / 237 / 239, 18 / 33 / 72 / 197 / 237, 18 / 68 / 72, 18 / 68 / 72 / 158 / 175 / 237 / 239, 18 / 68 / 72 / 175 / 202 / 239, 18 / 68 / 72 / 239, 18 / 68 / 175 / 223 / 237 / 239, 18 / 68 / 239, 18 / 72 / 97 / 175 / 237 / 239, 18 / 72 / 175 / 197, 18 / 72 / 237, 18 / 72 / 239, 18 / 97 / 175 / 237 / 239, 18 / 97 / 223 / 237 / 239, 18 / 175, 18 / 197 / 239, 18 / 239, 20 / 124, 33 / 68 / 72, 33 / 68 / 72 / 175 / 230, 33 / 68 / 223 / 239, 33 / 72, 33 / 72 / 97 / 197, 33 / 72 / 175, 33 / 175, 39, 44 / 97, 52, 68 / 72, 68 / 72 / 97, 68 / 72 / 97 / 197, 68 / 72 / 197, 68 / 72 / 197 / 239, 68 / 97 / 175, 68 / 97 / 175 / 197 / 223, 72 / 97 / 158 / 223 / 237, 72 / 175 / 197, 97 / 158 / 175 / 197, 97 / 175, 103, 121, 122, 124, 125, 128, 133, 148, 175,It includes an amino acid sequence selected from 197 and 197 / 202 and having one or more residue differences as compared to SEQ ID NO: 2. In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 2, and SEQ ID NO: 7V / 9V, 7V / 9V / 33V / 68L / 72E / 158Y, 7V / 9V / 33V / 72E / 175P / 197V / 230I, 7V / 9V / 33V / 72E / 230I, 7V / 9V / 68L / 72E / 197V, 7V / 9V / 158Y / 175P, 7V / 9V / 175P, 7V / 9V / 197V, 7V / 33V, 7V / 68L / 72E / 175P, 7V / 72E, 7V / 72E / 97I / 175P / 230I, 9V, 9V / 11G / 33V / 72E / 197V / 230I, 9V / 33V / 68L / 158Y / 197V, 9V / 33V / 72E / 230I, 9V / 33V / 97I / 230I, 9V / 68L / 72E, 9V / 97I / 175P / 202I / 218M, 9V / 175P, 9V / 202I, 11S, 13G, 14Y / 33V / 68L / 72E / 97I, 18C / 33V / 68L / 97I / 158Y / 175P / 239F, 18C / 33V / 72E / 97I / 158Y / 175P, 18C / 33V / 72E / 97I / 158Y / 239F, 18C / 33V / 72E / 97I / 175P / 202I, 18C / 33V / 72E / 175P / 202I / 237P / 239F, 18C / 33V / 72E / 197V / 237P, 18C / 68L / 72E, 18C / 68L / 72E / 158Y / 175P / 237P / 239F, 18C / 68L / 72E / 175P / 202I / 239F, 18C / 68L / 72E / 239F, 18C / 68L / 175P / 223P / 237P / 239F, 18C / 68L / 239F, 18C / 72E / 97I / 175P / 237P / 239F, 18C / 72E / 175P / 197V, 18C / 72E / 237P, 18C / 72E / 239F, 18C / 97I / 175P / 237P / 239F, 18C / 97I / 223P / 237P / 239F, 18C / 175P, 18C / 197V / 239F, 18C / 239F, 18I, 20H / 124C,Comprising an amino acid sequence having one or more residue differences compared to SEQ ID NO: 2, selected from 33V / 68L / 72E, 33V / 68L / 72E / 175P / 230I, 33V / 68L / 223P / 239F, 33V / 72E, 33V / 72E / 97I / 197V, 33V / 72E / 175P, 33V / 175P, 39A, 39G, 44N / 97I, 52S, 68L / 72E, 68L / 72E / 97I, 68L / 72E / 97I / 197V, 68L / 72E / 197V, 68L / 72E / 197V / 239F, 68L / 97I / 175P, 68L / 97I / 175P / 197V / 223P, 72E / 97I / 158Y / 223P / 237P, 72E / 175P / 197V, 97I / 158Y / 175P / 197V, 97I / 175P, 103M, 103S, 121Y, 122C, 122Q, 124C, 124E, 124T, 124Y, 125R, 128H, 128Y, 133A, 133C, 148M, 175P, 197V, and 197V / 202I. In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 2, and SEQ ID NO: I7V / I9V, I7V / I9V / A33V / I68L / K72E / F158Y, I7V / I9V / A33V / K72E / K175P / I197V / V230I, I7V / I9V / A33V / K72E / V230I, I7V / I9V / I68L / K72E / I197V, I7V / I9V / F158Y / K175P, I7V / I9V / K175P, I7V / I9V / I197V, I7V / A33V, I7V / I68L / K72E / K175P, I7V / K72E, I7V / K72E / V97I / K175P / V230I, I9V, I9V / T11G / A33V / K72E / I197V / V230I, I9V / A33V / I68L / F158Y / I197V, I9V / A33V / K72E / V230I, I9V / A33V / V97I / V230I, I9V / I68L / K72E, I9V / V97I / K175P / V202I / Q218M, I9V / K175P, I9V / V202I, T11S, C13G, H14Y / A33V / I68L / K72E / V97I,L18C / A33V / I68L / V97I / F158Y / K175P / I239F, L18C / A33V / K72E / V97I / F158Y / K175P, L18C / A33V / K72E / V97I / F158Y / I239F, L18C / A33V / K72E / V97I / K175P / V202I, L18C / A33V / K72E / K175P / V202I / I237P / I239F, L18C / A33V / K72E / I197V / I237P, L18C / I68L / K72E, L18C / I68L / K72E / F158Y / K175P / I237P / I239F, L18C / I68L / K72E / K175P / V202I / I239F, L18C / I68L / K72E / I239F, L18C / I68L / K175P / K223P / I237P / I239F, L18C / I68L / I239F, L18C / K72E / V97I / K175P / I237P / I239F, L18C / K72E / K175P / I197V, L18C / K72E / I237P, L18C / K72E / I239F, L18C / V97I / K175P / I237P / I239F, L18C / V97I / K223P / I237P / I239F, L18C / K175P, L18C / I197V / I239F, L18C / I239F, L18I, Y20H / V124C, A33V / I68L / K72E, A33V / I68L / K72E / K175P / V230I, A33V / I68L / K223P / I239F, A33V / K72E, A33V / K72E / V97I / I197V, A33V / K72E / K175P, A33V / K175P, V39A, V39G, D44N / V97I, G52S, I68L / K72E, I68L / K72E / V97I, I68L / K72E / V97I / I197V, I68L / K72E / I197V, I68L / K72E / I197V / I239F, I68L / V97I / K175P, I68L / V97I / K175P / I197V / K223P, K72E / V97I / F158Y / K223P / I237P, K72E / K175P / I197V, V97I / F158Y / K175P / I197V, V97I / K175P, H103M, H103S, F121Y, S122C, S122Q, V124C, V124E, V124T, V124Y, F125R, W128H, W128Y, F133A, F133C, L148MIt includes an amino acid sequence having one or more residue differences compared to SEQ ID NO: 2, selected from K175P, I197V, and I197V / V202I.
[0138] In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 2, which is a reference sequence, and one or more residue differences compared to SEQ ID NO: 2, selected from SEQ ID NO: 7 / 11 / 33 / 72 / 97, 7 / 11 / 33 / 97, 7 / 11 / 72, 7 / 11 / 158, 7 / 11 / 237, 7 / 33 / 72 / 97 / 158 / 175 / 197 / 237 / 239, 9 / 11 / 33, 9 / 11 / 33 / 72, 9 / 11 / 72, 9 / 33 / 68 / 97 / 237 / 239, 9 / 72 / 97 / 158 / 237 / 239, 9 / 218 / 239, 10, 11, 12, 13, 14, 18, 33 / 68 / 72 / 97 / 237 / 239, 33 / 72 / 158 / 239, 39, 40, 41, 48, 49, 51, 54, 54 / 126, 57, 68 / 175 / 239, 82, 84, 85, 86, 103, 104, 105, 107, 115, 117, 119, 119 / 120, 120, 122, 123, 124, 127, 130, 131, 133, 134, 137 / 145, 144, 145, 146, 147, 148, 149, 150, 152, 157, 158, 159, 169, 173, 174, 175, 180, 181, 183, 203, 204, 208, 209, 210, 215, 218, 236, and 237. In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 2, which is a reference sequence, and one or more residue differences compared to SEQ ID NO: 2, selected from SEQ ID NO: 7V / 11G / 33V / 72E / 97I, 7V / 11G / 33V / 97I, 7V / 11G / 72E, 7V / 11G / 158Y, 7V / 11G / 237P, 7V / 33V / 72E / 97I / 158Y / 175P / 197V / 237P / 239F, 9V / 11G / 33V, 9V / 11G / 33V / 72E, 9V / 11G / 72E, 9V / 33V / 68L / 97I / 237P / 239F, 9V / 72E / 97I / 158Y / 237P / 239F, 9V / 218M / 239F, 10G, 10P, 10R, 11R, 11S, 12K,Comprising an amino acid sequence having one or more residue differences compared to SEQ ID NO: 2, selected from 13L, 14R, 18S, 18T, 33V / 68L / 72E / 97I / 237P / 239F, 33V / 72E / 158Y / 239F, 39R, 39W, 40L, 40S, 40T, 40V, 41C, 48W, 49Q, 49R, 51A, 51Q, 51S, 51T, 54N / 126I, 54Y, 57N, 68L / 175P / 239F, 82T, 84L, 85A, 85L, 85M, 85Q, 86W, 103P, 104A, 104G, 104S, 105D, 107G, 107W, 115C, 115K, 115L, 115V, 117C, 117L, 119L, 119N / 120F, 120C, 120L, 120Y, 122G, 122T, 123A, 123P, 123T, 124P, 127A, 127L, 127M, 130M, 130V, 131L, 133W, 134D, 134F, 134L, 134R, 134V, 134W, 137I / 145V, 144L, 145F, 145L, 145P, 146D, 146G, 146N, 147R, 148P, 148V, 149W, 150I, 150L, 150V, 152P, 157V, 158I, 158L, 158M, 158R, 158W, 158Y, 159G, 159H, 159L, 169G, 169I, 169S, 173E, 173Y, 174C, 174G, 175D, 180G, 180V, 181C, 183K, 203Q, 204E, 204R, 204V, 208A, 208C, 208E, 208G, 208L, 208T, 208V, 209C, 210V, 215L, 218T, 236G, 236L, 237L, 237S, and 237W. In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 2, which is the reference sequence, and SEQ ID NO: I7V / T11G / A33V / K72E / V97I, I7V / T11G / A33V / V97I, I7V / T11G / K72E, I7V / T11G / F158Y, I7V / T11G / I237P, I7V / A33V / K72E / V97I / F158Y / K175P / I197V / I237P / I239F, I9V / T11G / A33V,An amino acid sequence having one or more residue differences compared to SEQ ID NO: 2, selected from I9V / T11G / A33V / K72E, I9V / T11G / K72E, I9V / A33V / I68L / V97I / I237P / I239F, I9V / K72E / V97I / F158Y / I237P / I239F, I9V / Q218M / I239F, H10G, H10P, H10R, T11R, T11S, I12K, C13L, H14R, L18S, L18T, A33V / I68L / K72E / V97I / I237P / I239F, A33V / K72E / F158Y / I239F, V39R, V39W, A40L, A40S, A40T, A40V, S41C, L48W, E49Q, E49R, I51A, I51Q, I51S, I51T, W54N / T126I, W54Y, Y57N, I68L / K175P / I239F, G82T, I84L, N85A, N85L, N85M, N85Q, I86W, H103P, N104A, N104G, N104S, A105D, M107G, M107W, A115C, A115K, A115L, A115V, V117C, V117L, K119L, K119N / K120F, K120C, K120L, K120Y, S122G, S122T, E123A, E123P, E123T, V124P, D127A, D127L, D127M, D130M, D130V, S131L, F133W, S134D, S134F, S134L, S134R, S134V, S134W, T137I / A145V, T144L, A145F, A145L, A145P, V146D, V146G, V146N, E147R, L148P, L148V, G149W, D150I, D150L, D150V, T152P, I157V, F158I, F158L, F158M, F158R, F158W, F158Y, S159G, S159H, S159L, A169G, A169I, A169S, V173E, V173Y, R174C, R174G, K175D, E180G, E180V, Q181C, L183K, Y203Q, G204E, G204R, G204V, Q208A, Q208C, Q208E, Q208G, Q208L, Q208T, Q208V, I209C, F210V, Q215L, Q218T, K236G, K236L, I237L, I237S, and I237W.
[0139] In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence of SEQ ID NO: 24 and one or more residue differences compared to SEQ ID NO: 24 selected from SEQ ID NOs: 9 / 39 / 122 / 157, 18, 18 / 33 / 39 / 230, 18 / 33 / 175, 18 / 39 / 44 / 118, 18 / 175, 20, 33 / 39, 33 / 39 / 44 / 118 / 121, 33 / 39 / 44 / 121, 33 / 39 / 44 / 152, 33 / 39 / 44 / 175, 33 / 39 / 44 / 197, 33 / 39 / 44 / 210 / 230, 33 / 39 / 118 / 121 / 175, 33 / 39 / 121 / 175, 33 / 39 / 197 / 210, 33 / 39 / 210, 33 / 39 / 210 / 230, 33 / 118 / 121, 33 / 118 / 121 / 175, 39, 39 / 44 / 103, 39 / 44 / 118, 39 / 44 / 152 / 230, 39 / 44 / 175 / 197 / 230, 39 / 103, 39 / 103 / 125, 39 / 103 / 125 / 127 / 146, 39 / 103 / 125 / 127 / 146 / 150, 39 / 103 / 125 / 146, 39 / 103 / 127 / 150, 39 / 103 / 150, 39 / 118 / 121, 39 / 197, 39 / 197 / 210, 43, 44 / 103, 81, 103, 103 / 125 / 127, 103 / 125 / 146, 103 / 127, 103 / 150, 118 / 121, 118 / 121 / 175, 118 / 121 / 197, 121, 152, 152 / 197, 154, 159, 163, 175, 210, 232, and 238. In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence of SEQ ID NO: 24 and one or more residue differences compared to SEQ ID NO: 24 selected from SEQ ID NOs: 9V / 39W / 122T / 157V, 18I, 18I / 33V / 39W / 230I, 18I / 33V / 175P, 18I / 39W / 44N / 118L, 18I / 175P, 20H, 33V / 39W, 33V / 39W / 44N / 118L / 121Y,Comprising an amino acid sequence having one or more residue differences compared to SEQ ID NO: 24, selected from 33V / 39W / 44N / 121Y, 33V / 39W / 44N / 152F, 33V / 39W / 44N / 175P, 33V / 39W / 44N / 197V, 33V / 39W / 44N / 210M / 230I, 33V / 39W / 118L / 121Y / 175P, 33V / 39W / 121Y / 175P, 33V / 39W / 197V / 210M, 33V / 39W / 210M, 33V / 39W / 210M / 230I, 33V / 118L / 121W / 175P, 33V / 118L / 121Y, 39W, 39W / 44N / 118L, 39W / 44N / 152F / 230I, 39W / 44N / 175P / 197V / 230I, 39W / 44Y / 103S, 39W / 103M, 39W / 103M / 125R / 146L, 39W / 103M / 150P, 39W / 103Q, 39W / 103S, 39W / 103S / 125P / 127G / 146L / 150P, 39W / 103S / 125R, 39W / 103S / 125R / 127G / 146L, 39W / 103S / 127G / 150P, 39W / 118L / 121Y, 39W / 197V, 39W / 197V / 210M, 43M, 43Q, 43R, 43S, 43T, 44N / 103S, 81F, 103L, 103M, 103M / 125P / 146L, 103M / 127G, 103Q, 103S, 103S / 125P / 127G, 103S / 150P, 103V, 118L / 121Y, 118L / 121Y / 175P, 118L / 121Y / 197V, 121Y, 152A / 197V, 152F, 154I, 159Q, 163P, 175P, 210M, 232G, and 238M. In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 24, and SEQ ID NO: I9V / V39W / S122T / I157V, L18I, L18I / A33V / V39W / V230I, L18I / A33V / K175P, L18I / V39W / D44N / Y118L, L18I / K175P, Y20H, A33V / V39W, A33V / V39W / D44N / Y118L / F121Y,Comprising an amino acid sequence having one or more residue differences compared to SEQ ID NO: 24, selected from A33V / V39W / D44N / F121Y, A33V / V39W / D44N / T152F, A33V / V39W / D44N / K175P, A33V / V39W / D44N / I197V, A33V / V39W / D44N / F210M / V230I, A33V / V39W / Y118L / F121Y / K175P, A33V / V39W / F121Y / K175P, A33V / V39W / I197V / F210M, A33V / V39W / F210M, A33V / V39W / F210M / V230I, A33V / Y118L / F121W / K175P, A33V / Y118L / F121Y, V39W, V39W / D44N / Y118L, V39W / D44N / T152F / V230I, V39W / D44N / K175P / I197V / V230I, V39W / D44Y / H103S, V39W / H103M, V39W / H103M / F125R / V146L, V39W / H103M / D150P, V39W / H103Q, V39W / H103S, V39W / H103S / F125P / D127G / V146L / D150P, V39W / H103S / F125R, V39W / H103S / F125R / D127G / V146L, V39W / H103S / D127G / D150P, V39W / Y118L / F121Y, V39W / I197V, V39W / I197V / F210M, I43M, I43Q, I43R, I43S, I43T, D44N / H103S, G81F, H103L, H103M, H103M / F125P / V146L, H103M / D127G, H103Q, H103S, H103S / F125P / D127G, H103S / D150P, H103V, Y118L / F121Y, Y118L / F121Y / K175P, Y118L / F121Y / I197V, F121Y, T152A / I197V, T152F, A154I, S159Q, I163P, K175P, F210M, S232G, and Q238M.,
[0140] In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 24 and SEQ ID NO: 7, 9 / 13, 9 / 13 / 39, 9 / 13 / 39 / 68 / 72 / 239, 9 / 13 / 39 / 68 / 239, 9 / 13 / 39 / 72 / 122, 9 / 13 / 39 / 97 / 239, 9 / 13 / 68 / 72 / 97 / 239, 9 / 13 / 68 / 97 / 134 / 239, 9 / 13 / 68 / 97 / 239, 9 / 13 / 72, 9 / 13 / 134, 11, 12 / 42 / 125 / 127, 12 / 103, 12 / 103 / 150, 13, 13 / 39, 13 / 39 / 68, 13 / 39 / 68 / 72 / 97, 13 / 39 / 72 / 97 / 122 / 134 / 223, 13 / 39 / 97 / 239, 13 / 39 / 239, 13 / 68, 13 / 68 / 72 / 97 / 122 / 239, 16, 18 / 33 / 39 / 44 / 118 / 121 / 175 / 197 / 210, 18 / 33 / 39 / 44 / 121, 18 / 33 / 39 / 112 / 118 / 121 / 197, 18 / 33 / 39 / 121, 18 / 33 / 39 / 121 / 175, 18 / 33 / 39 / 121 / 175 / 197, 18 / 33 / 39 / 121 / 175 / 210, 18 / 33 / 39 / 121 / 210, 18 / 33 / 118 / 121, 18 / 33 / 121, 18 / 33 / 121 / 197, 18 / 33 / 121 / 210 / 230, 18 / 39 / 44 / 118 / 121 / 230, 18 / 39 / 44 / 121 / 175 / 210, 18 / 39 / 118 / 121, 18 / 39 / 118 / 121 / 210, 18 / 39 / 121 / 175, 18 / 39 / 121 / 175 / 210, 18 / 39 / 121 / 197 / 210, 18 / 118 / 121, 18 / 118 / 121 / 175, 18 / 118 / 121 / 210, 18 / 118 / 121 / 230, 18 / 121, 18 / 121 / 175, 18 / 121 / 210, 18 / 121 / 230, 20, 21, 33 / 39 / 44 / 118 / 121, 33 / 39 / 44 / 118 / 121 / 210 / 230, 33 / 39 / 44 / 121, 33 / 39 / 44 / 121 / 152, 33 / 39 / 44 / 121 / 175 / 210, 33 / 39 / 44 / 121 / 230, 33 / 39 / 73 / 121 / 197,Comprising an amino acid sequence having one or more residue differences compared to SEQ ID NO: 24, selected from 33 / 39 / 118 / 121 / 175, 33 / 39 / 118 / 121 / 175 / 230, 33 / 39 / 121, 33 / 39 / 121 / 175, 33 / 39 / 121 / 210 / 230, 33 / 118 / 121, 33 / 118 / 121 / 175, 33 / 118 / 121 / 197, 33 / 121, 33 / 121 / 230, 37 / 91, 39 / 44 / 118 / 121 / 175 / 210, 39 / 68 / 72 / 239, 39 / 118 / 121, 39 / 118 / 121 / 210, 39 / 121, 39 / 121 / 210 / 230, 43, 44, 76, 79, 101, 118 / 121, 118 / 121 / 175, 118 / 121 / 197, 121, 121 / 175 / 230, 121 / 210, 148, 153, 159, 161, 172, 200, 207, 233, 245, and 248. In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 24, and SEQ ID NO: 7E, 9V / 13G, 9V / 13G / 39W, 9V / 13G / 39W / 68L / 72E / 239F, 9V / 13G / 39W / 68L / 239F, 9V / 13G / 39W / 72E / 122T, 9V / 13G / 39W / 97I / 239F, 9V / 13G / 68L / 72E / 97I / 239F, 9V / 13G / 68L / 97I / 134D / 239F, 9V / 13G / 68L / 97I / 239F, 9V / 13G / 72E, 9V / 13G / 134D, 11N, 12A / 42W / 125R / 127G, 12A / 103S, 12A / 103S / 150P, 13F, 13G / 39W, 13G / 39W / 68L, 13G / 39W / 68L / 72E / 97I, 13G / 39W / 72E / 97I / 122T / 134D / 223P, 13G / 39W / 97I / 239F, 13G / 39W / 239F, 13G / 68L, 13G / 68L / 72E / 97I / 122T / 239F, 16M, 16V, 18C / 33V / 39W / 112R / 118L / 121W / 197V, 18C / 33V / 39W / 121W / 175P / 210M, 18C / 33V / 39W / 121W / 210M,18C / 33V / 39W / 121Y / 175P, 18C / 33V / 118L / 121W, 18C / 33V / 118L / 121Y, 18C / 33V / 121Y, 18C / 33V / 121Y / 197V, 18C / 33V / 121Y / 210M / 230I, 18C / 39W / 44N / 121Y / 175P / 210M, 18C / 39W / 118L / 121W, 18C / 39W / 121W / 175P / 210M, 18C / 39W / 121W / 197V / 210M, 18C / 118L / 121W, 18C / 118L / 121Y, 18C / 121V, 18C / 121V / 175P, 18C / 121W, 18C / 121W / 175P, 18C / 121W / 210M, 18C / 121Y, 18C / 121Y / 230I, 18I / 33V / 39W / 44N / 118L / 121W / 175P / 197V / 210M, 18I / 33V / 39W / 44N / 121T, 18I / 33V / 39W / 121T / 175P, 18I / 33V / 39W / 121V / 175P / 197V, 18I / 33V / 39W / 121W, 18I / 33V / 118L / 121W, 18I / 39W / 44N / 118L / 121W / 230I, 18I / 39W / 44N / 121V / 175P / 210M, 18I / 39W / 118L / 121W / 210M, 18I / 39W / 121V / 175P, 18I / 118L / 121T, 18I / 118L / 121W, 18I / 118L / 121W / 175P, 18I / 118L / 121W / 210M, 18I / 118L / 121Y, 18I / 118L / 121Y / 230I, 18I / 121T / 175P, 18I / 121W, 18I / 121Y / 210M, 20G, 20K, 20R, 20T, 21V, 33V / 39G / 44N / 121W / 152A, 33V / 39W / 44N / 118L / 121W / 210M / 230I, 33V / 39W / 44N / 118L / 121Y, 33V / 39W / 44N / 121V, 33V / 39W / 44N / 121Y, 33V / 39W / 44N / 121Y / 175P / 210M, 33V / 39W / 44N / 121Y / 230I, 33V / 39W / 73N / 121Y / 197V, 33V / 39W / 118L / 121W / 175P, 33V / 39W / 118L / 121Y / 175P, 33V / 39W / 118L / 121Y / 175P / 230I, 33V / 39W / 121V, 33V / 39W / 121WComprising an amino acid sequence having one or more residue differences compared to SEQ ID NO: 24, selected from 33V / 39W / 121W / 210M / 230I, 33V / 39W / 121Y / 175P, 33V / 118L / 121W / 175P, 33V / 118L / 121W / 197V, 33V / 118L / 121Y, 33V / 121W, 33V / 121Y / 230I, 37L / 91G, 39G / 68L / 72E / 239F, 39W / 44N / 118L / 121W / 175P / 210M, 39W / 118L / 121W / 210M, 39W / 118L / 121Y, 39W / 118L / 121Y / 210M, 39W / 121W, 39W / 121W / 210M / 230I, 43E, 43Q, 43S, 43W, 44V, 76P, 79F, 101L, 118L / 121W, 118L / 121Y, 118L / 121Y / 175P, 118L / 121Y / 197V, 121V, 121W, 121Y, 121Y / 175P / 230I, 121Y / 210M, 148F, 153W, 159Q, 161M, 172A, 172E, 172G, 172Q, 172R, 172V, 200G, 207N, 233R, 245D, 245G, 248G, and 248S. In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 24, which is the reference sequence, and SEQ ID NO: I7E, I9V / C13G, I9V / C13G / V39W, I9V / C13G / V39W / I68L / K72E / I239F, I9V / C13G / V39W / I68L / I239F, I9V / C13G / V39W / K72E / S122T, I9V / C13G / V39W / V97I / I239F, I9V / C13G / I68L / K72E / V97I / I239F, I9V / C13G / I68L / V97I / S134D / I239F, I9V / C13G / I68L / V97I / I239F, I9V / C13G / K72E, I9V / C13G / S134D, S11N, I12A / G42W / F125R / D127G, I12A / H103S, I12A / H103S / D150P, C13F, C13G / V39W, C13G / V39W / I68L, C13G / V39W / I68L / K72E / V97I,C13G / V39W / K72E / V97I / S122T / S134D / K223P, C13G / V39W / V97I / I239F, C13G / V39W / I239F, C13G / I68L, C13G / I68L / K72E / V97I / S122T / I239F, A16M, A16V, L18C / A33V / V39W / H112R / Y118L / F121W / I197V, L18C / A33V / V39W / F121W / K175P / F210M, L18C / A33V / V39W / F121W / F210M, L18C / A33V / V39W / F121Y / K175P, L18C / A33V / Y118L / F121W, L18C / A33V / Y118L / F121Y, L18C / A33V / F121Y, L18C / A33V / F121Y / I197V, L18C / A33V / F121Y / F210M / V230I, L18C / V39W / D44N / F121Y / K175P / F210M, L18C / V39W / Y118L / F121W, L18C / V39W / F121W / K175P / F210M, L18C / V39W / F121W / I197V / F210M, L18C / Y118L / F121W, L18C / Y118L / F121Y, L18C / F121V, L18C / F121V / K175P, L18C / F121W, L18C / F121W / K175P, L18C / F121W / F210M, L18C / F121Y, L18C / F121Y / V230I, L18I / A33V / V39W / D44N / Y118L / F121W / K175P / I197V / F210M, L18I / A33V / V39W / D44N / F121T, L18I / A33V / V39W / F121T / K175P, L18I / A33V / V39W / F121V / K175P / I197V, L18I / A33V / V39W / F121W, L18I / A33V / Y118L / F121W, L18I / V39W / D44N / Y118L / F121W / V230I, L18I / V39W / D44N / F121V / K175P / F210M, L18I / V39W / Y118L / F121W / F210M, L18I / V39W / F121V / K175P, L18I / Y118L / F121T, L18I / Y118L / F121W, L18I / Y118L / F121W / K175P, L18I / Y118L / F121W / F210M, L18I / Y118L / F121YL18I / Y118L / F121Y / V230I, L18I / F121T / K175P, L18I / F121W, L18I / F121Y / F210M, Y20G, Y20K, Y20R, Y20T, K21V, A33V / V39G / D44N / F121W / T152A, A33V / V39W / D44N / Y118L / , Comprising an amino acid sequence having one or more residue differences compared to SEQ ID NO: 24, selected from F121W / F210M / V230I, A33V / V39W / D44N / Y118L / F121Y, A33V / V39W / D44N / F121V, A33V / V39W / D44N / F121Y, A33V / V39W / D44N / F121Y / K175P / F210M, A33V / V39W / D44N / F121Y / V230I, A33V / V39W / K73N / F121Y / I197V, A33V / V39W / Y118L / F121W / K175P, A33V / V39W / Y118L / F121Y / K175P, A33V / V39W / Y118L / F121Y / K175P / V230I, A33V / V39W / F121V, A33V / V39W / F121W, A33V / V39W / F121W / F210M / V230I, A33V / V39W / F121Y / K175P, A33V / Y118L / F121W / K175P, A33V / Y118L / F121W / I197V, A33V / Y118L / F121Y, A33V / F121W, A33V / F121Y / V230I, D37L / E91G, V39G / I68L / K72E / I239F, V39W / D44N / Y118L / F121W / K175P / F210M, V39W / Y118L / F121W / F210M, V39W / Y118L / F121Y, V39W / Y118L / F121Y / F210M, V39W / F121W, V39W / F121W / F210M / V230I, I43E, I43Q, I43S, I43W, D44V, L76P, E79F, V101L, Y118L / F121W, Y118L / F121Y, Y118L / F121Y / K175P, Y118L / F121Y / I197V, F121V, F121W, F121Y, F121Y / K175P / V230I, F121Y / F210M, L148F, L153W, S159Q, S161M, L172A, L172E, L172G, L172Q, L172R, L172V, V200G, D207N, A233R, L245D, L245G, I248G, and I248S.
[0141] In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 496, which is a reference sequence, and one or more residue differences compared to SEQ ID NO: 496, selected from SEQ ID NOs: 43 / 103, 43 / 103 / 172 / 238, 43 / 103 / 238 / 241, 43 / 238, 81 / 232, 103, 103 / 159 / 238, 103 / 194 / 238, 103 / 238, 122, 232, 238, and 238 / 241. In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 496, which is a reference sequence, and one or more residue differences compared to SEQ ID NO: 496, selected from SEQ ID NOs: 43E / 103S / 238M / 241R, 43M / 103L, 43Q / 103S / 238M / 241R, 43Q / 238M, 43S / 103S / 172R / 238M, 43S / 238M, 81F / 232G, 103L, 103S, 103S / 159Q / 238M, 103S / 194L / 238M, 103S / 238M, 103V, 122C, 122I, 122K, 122L, 122M, 122Q, 122R, 122T, 122V, 232G, 238M, and 238M / 241R.In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 496, which is a reference sequence, and one or more residue differences compared to SEQ ID NO: 496 selected from I43E / H103S / Q238M / K241R, I43M / H103L, I43Q / H103S / Q238M / K241R, I43Q / Q238M, I43S / H103S / L172R / Q238M, I43S / Q238M, G81F / S232G, H103L, H103S, H103S / S159Q / Q238M, H103S / Y194L / Q238M, H103S / Q238M, H103V, S122C, S122I, S122K, S122L, S122M, S122Q, S122R, S122T, S122V, S232G, Q238M, and Q238M / K241R.
[0142] In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 496, the reference sequence, and one or more residue differences compared to SEQ ID NO: 496, selected from SEQ ID NOs: 12, 43, 43 / 47 / 103, 43 / 47 / 103 / 232, 43 / 103, 43 / 103 / 121 / 172 / 241, 43 / 103 / 172, 43 / 103 / 172 / 238, 43 / 103 / 238, 43 / 103 / 238 / 241, 43 / 121, 43 / 194 / 238, 43 / 238, 44 / 103 / 121 / 238, 44 / 103 / 238, 44 / 103 / 238 / 241, 44 / 121 / 238 / 241, 47 / 103, 63, 103, 103 / 121, 103 / 121 / 172 / 238, 103 / 159 / 238, 103 / 172 / 238, 103 / 194 / 238, 103 / 238, 122, 131, 146, 178, 180, and 221.In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 496 and one or more residue differences compared to SEQ ID NO: 496, selected from SEQ ID NO: 12V, 43E / 103S / 121W / 172R / 241R, 43E / 103S / 238M, 43E / 103S / 238M / 241R, 43M / 47G / 103L, 43M / 47G / 103V, 43M / 103L, 43M / 103S, 43M / 103V, 43Q / 47G / 103S / 232G, 43Q / 103L / 172R, 43Q / 103S, 43Q / 103S / 238M / 241R, 43Q / 121W, 43Q / 238M, 43S, 43S / 103L, 43S / 103S, 43S / 103S / 172R, 43S / 103S / 172R / 238M, 43S / 103S / 238M / 241R, 43S / 194L / 238M, 43S / 238M, 44V / 103S / 121W / 238M, 44V / 103S / 238M, 44V / 103S / 238M / 241R, 44V / 121W / 238M / 241R, 47G / 103L, 63P, 103L, 103S, 103S / 121W, 103S / 121W / 172R / 238M, 103S / 159Q / 238M, 103S / 172R / 238M, 103S / 194L / 238M, 103S / 238M, 103V, 122H, 122I, 122K, 122L, 122M, 122N, 122R, 122T, 122V, 131Y, 146I, 178A, 180L, 180V, and 221G.In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 496 and one or more residue differences compared to SEQ ID NO: 496 selected from I12V, I43E / H103S / Y121W / L172R / K241R, I43E / H103S / Q238M, I43E / H103S / Q238M / K241R, I43M / Q47G / H103L, I43M / Q47G / H103V, I43M / H103L, I43M / H103S, I43M / H103V, I43Q / Q47G / H103S / S232G, I43Q / H103L / L172R, I43Q / H103S, I43Q / H103S / Q238M / K241R, I43Q / Y121W, I43Q / Q238M, I43S, I43S / H103L, I43S / H103S, I43S / H103S / L172R, I43S / H103S / L172R / Q238M, I43S / H103S / Q238M / K241R, I43S / Y194L / Q238M, I43S / Q238M, N44V / H103S / Y121W / Q238M, N44V / H103S / Q238M, N44V / H103S / Q238M / K241R, N44V / Y121W / Q238M / K241R, Q47G / H103L, T63P, H103L, H103S, H103S / Y121W, H103S / Y121W / L172R / Q238M, H103S / S159Q / Q238M, H103S / L172R / Q238M, H103S / Y194L / Q238M, H103S / Q238M, H103V, S122H, S122I, S122K, S122L, S122M, S122N, S122R, S122T, S122V, S131Y, V146I, F178A, E180L, E180V, and N221G.
[0143] In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence of SEQ ID NO: 860 and one or more residue differences compared to SEQ ID NO: 860 selected from SEQ ID NOs: 12, 12 / 43 / 63 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 63 / 103 / 146 / 241, 12 / 43 / 81 / 103 / 241, 12 / 43 / 81 / 146 / 232 / 241, 12 / 43 / 81 / 180 / 241, 12 / 43 / 81 / 232 / 241, 12 / 43 / 81 / 241, 12 / 43 / 103, 12 / 43 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 103 / 146 / 180 / 241, 12 / 43 / 103 / 180 / 241, 12 / 43 / 103 / 241, 12 / 43 / 146 / 180 / 232 / 241, 12 / 43 / 180 / 241, 12 / 63 / 81 / 103 / 241, 12 / 81, 12 / 81 / 180 / 241, 12 / 103 / 146 / 241, 12 / 103 / 180 / 241, 12 / 146, 12 / 146 / 180 / 241, 12 / 180 / 232 / 241, 12 / 241, 18, 43 / 81 / 232 / 241, 43 / 103 / 146 / 180, 43 / 103 / 180 / 241, 43 / 103 / 241, 43 / 180 / 232, 43 / 232 / 241, 63 / 103 / 180 / 232, 78, 81, 81 / 103 / 146, 81 / 146 / 180 / 241, 81 / 146 / 241, 81 / 180 / 241, 84, 103, 103 / 146, 103 / 146 / 180, 103 / 146 / 180 / 241, 103 / 180, 103 / 180 / 241, 103 / 261, 144, 145, 146 / 180 / 232, 146 / 180 / 241, 146 / 241, 150, 166, 172, 180 / 232, 232 / 241, and 241. In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence of SEQ ID NO: 860 and SEQ ID NO: 12V,12V / 43M / 63P / 103L / 146I / 180L / 232G / 241R, 12V / 43M / 81F / 103L / 241R, 12V / 43M / 81F / 146I / 232G / 241R, 12V / 43M / 81F / 180V / 241R, 12V / 43M / 81F / 232G / 241R, 12V / 43M / 81F / 241R, 12V / 43M / 103L / 180L / 241R, 12V / 43M / 103L / 241R, 12V / 43M / 103V / 146I / 180V / 241R, 12V / 43M / 103V / 241R, 12V / 43M / 146I / 180L / 232G / 241R, 12V / 43M / 180L / 241R, 12V / 43M / 180V / 241R, 12V / 43Q / 63P / 103L / 146I / 241R, 12V / 43Q / 103L / 146I / 180V / 232G / 241R, 12V / 43Q / 103L / 146I / 180V / 241R, 12V / 43Q / 103L / 180V / 241R, 12V / 43Q / 103V, 12V / 63P / 81F / 103L / 241R, 12V / 81F, 12V / 81F / 180V / 241R, 12V / 103L / 146I / 241R, 12V / 103L / 180L / 241R, 12V / 103V / 180V / 241R, 12V / 146I, 12V / 146I / 180L / 241R, 12V / 180L / 232G / 241R, 12V / 180V / 232G / 241R, 12V / 241R, 18V, 43M / 81F / 232G / 241R, 43M / 103L / 146I / 180V, 43M / 103L / 180V / 241R, 43M / 103L / 241R, 43M / 103V / 146I / 180V, 43M / 180L / 232G, 43M / 232G / 241R, 43Q / 103V / 146I / 180L, 63P / 103L / 180L / 232G, 78A, 81F / 103L / 146I, 81F / 146I / 180L / 241R, 81F / 146I / 241R, 81F / 180L / 241R, 81Q, 84Y, 103A, 103L, 103L / 146I, 103L / 146I / 180V, 103L / 180L, 103L / 180V / 241R, 103N, 103T, 103T / 261V, 103V / 146I / 180L / 241R, 144V, 145V, 146I / 180L / 241R, 146I / 180V / 232G, 146I / 241R, 150H, 166L, 172LIt comprises an amino acid sequence selected from 172V, 180L / 232G, 232G / 241R, and 241R and having one or more residue differences as compared to SEQ ID NO: 860. In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 860, and SEQ ID NOs: I12V, I12V / S43M / T63P / S103L / V146I / E180L / S232G / K241R, I12V / S43M / G81F / S103L / K241R, I12V / S43M / G81F / V146I / S232G / K241R, I12V / S43M / G81F / E180V / K241R, I12V / S43M / G81F / S232G / K241R, I12V / S43M / G81F / K241R, I12V / S43M / S103L / E180L / K241R, I12V / S43M / S103L / K241R, I12V / S43M / S103V / V146I / E180V / K241R, I12V / S43M / S103V / K241R, I12V / S43M / V146I / E180L / S232G / K241R, I12V / S43M / E180L / K241R, I12V / S43M / E180V / K241R, I12V / S43Q / T63P / S103L / V146I / K241R, I12V / S43Q / S103L / V146I / E180V / S232G / K241R, I12V / S43Q / S103L / V146I / E180V / K241R, I12V / S43Q / S103L / E180V / K241R, I12V / S43Q / S103V, I12V / T63P / G81F / S103L / K241R, I12V / G81F, I12V / G81F / E180V / K241R, I12V / S103L / V146I / K241R, I12V / S103L / E180L / K241R, I12V / S103V / E180V / K241R, I12V / V146I, I12V / V146I / E180L / K241R, I12V / E180L / S232G / K241R, I12V / E180V / S232G / K241R, I12V / K241R, L18V, S43M / G81F / S232G / K241R,Comprising an amino acid sequence having one or more residue differences compared to SEQ ID NO: 860, selected from S43M / S103L / V146I / E180V, S43M / S103L / E180V / K241R, S43M / S103L / K241R, S43M / S103V / V146I / E180V, S43M / E180L / S232G, S43M / S232G / K241R, S43Q / S103V / V146I / E180L, T63P / S103L / E180L / S232G, G78A, G81F / S103L / V146I, G81F / V146I / E180L / K241R, G81F / V146I / K241R, G81F / E180L / K241R, G81Q, I84Y, S103A, S103L, S103L / V146I, S103L / V146I / E180V, S103L / E180L, S103L / E180V / K241R, S103N, S103T, S103T / A261V, S103V / V146I / E180L / K241R, T144V, A145V, V146I / E180L / K241R, V146I / E180V / S232G, V146I / K241R, D150H, V166L, R172L, R172V, E180L / S232G, S232G / K241R, and K241R.,
[0144] In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 860 and one or more residue differences compared to SEQ ID NO: 860, selected from SEQ ID NOs: 9, 12, 12 / 43 / 63 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 63 / 103 / 146 / 241, 12 / 43 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 103 / 146 / 180 / 241, 12 / 43 / 103 / 146 / 232 / 241, 12 / 43 / 103 / 146 / 241, 12 / 43 / 103 / 180 / 241, 12 / 43 / 146 / 180 / 232 / 241, 12 / 43 / 146 / 241, 12 / 43 / 180 / 232, 12 / 43 / 180 / 241, 12 / 63 / 103 / 146 / 232 / 241, 12 / 103 / 146 / 241, 12 / 103 / 180 / 241, 12 / 146, 12 / 146 / 180 / 241, 12 / 146 / 232 / 241, 12 / 146 / 241, 12 / 180, 12 / 180 / 232 / 241, 12 / 180 / 241, 12 / 241, 18, 43 / 63 / 103 / 180 / 232 / 241, 43 / 103 / 146 / 180, 43 / 103 / 180, 43 / 103 / 180 / 232 / 241, 43 / 103 / 180 / 241, 43 / 180 / 232, 63 / 103 / 180 / 232, 103, 103 / 146 / 180, 103 / 146 / 180 / 241, 103 / 180, 103 / 180 / 241, 103 / 241, 118, 119, 122, 144, 145, 146 / 180 / 241, 154, 160, 172, 180 / 232, 180 / 232 / 241, 180 / 241, 208, 211, and 212. In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 860 and one or more residue differences compared to SEQ ID NO: 860, selected from SEQ ID NOs: 9V, 12V, 12V / 43M / 63P / 103L / 146I / 180L / 232G / 241R,12V / 43M / 103L / 146I / 241R, 12V / 43M / 103L / 180L / 241R, 12V / 43M / 103V / 146I / 180V / 241R, 12V / 43M / 103V / 146I / 232G / 241R, 12V / 43M / 146I / 180L / 232G / 241R, 12V / 43M / 180L / 232G, 12V / 43M / 180L / 241R, 12V / 43Q / 63P / 103V / 146I / 241R, 12V / 43Q / 103L / 146I / 180V / 232G / 241R, 12V / 43Q / 103L / 146I / 180V / 241R, 12V / 43Q / 103L / 146I / 241R, 12V / 43Q / 103L / 180V / 241R, 12V / 43Q / 146I / 241R, 12V / 63P / 103L / 146I / 232G / 241R, 12V / 103L / 180L / 241R, 12V / 103L / 180V / 241R, 12V / 103V / 146I / 241R, 12V / 103V / 180L / 241R, 12V / 103V / 180V / 241R, 12V / 146I, 12V / 146I / 180L / 241R, 12V / 146I / 232G / 241R, 12V / 146I / 241R, 12V / 180L, 12V / 180L / 232G / 241R, 12V / 180L / 241R, 12V / 180V, 12V / 180V / 232G / 241R, 12V / 180V / 241R, 12V / 241R, 18V, 43M / 63P / 103L / 180V / 232G / 241R, 43M / 103L / 146I / 180V, 43M / 103L / 180V, 43M / 103L / 180V / 241R, 43M / 103V / 146I / 180V, 43M / 180L / 232G, 43Q / 103L / 180V / 232G / 241R, 43Q / 103V / 146I / 180L, 43Q / 103V / 180L / 241R, 43Q / 103V / 180V / 232G / 241R, 63P / 103L / 180L / 232G, 103A, 103L, 103L / 146I / 180V, 103L / 146I / 180V / 241R, 103L / 180L, 103L / 180V / 241R, 103L / 241R, 103N, 103T, 103V / 146I / 180L / 241R, 103V / 180V / 241R, 118L, 118V, 119H, 122H, 144H, 144V, 145V, 146I / 180L / 241R, 154RIt comprises an amino acid sequence having one or more residue differences compared to SEQ ID NO: 860, selected from 160L, 172C, 172L, 172V, 180L / 232G, 180L / 241R, 180V / 232G, 180V / 232G / 241R, 208F, 211F, and 212P. In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 860, and SEQ ID NOs: I9V, I12V, I12V / S43M / T63P / S103L / V146I / E180L / S232G / K241R, I12V / S43M / S103L / V146I / K241R, I12V / S43M / S103L / E180L / K241R, I12V / S43M / S103V / V146I / E180V / K241R, I12V / S43M / S103V / V146I / S232G / K241R, I12V / S43M / V146I / E180L / S232G / K241R, I12V / S43M / E180L / S232G, I12V / S43M / E180L / K241R, I12V / S43Q / T63P / S103V / V146I / K241R, I12V / S43Q / S103L / V146I / E180V / S232G / K241R, I12V / S43Q / S103L / V146I / E180V / K241R, I12V / S43Q / S103L / V146I / K241R, I12V / S43Q / S103L / E180V / K241R, I12V / S43Q / V146I / K241R, I12V / T63P / S103L / V146I / S232G / K241R, I12V / S103L / E180L / K241R, I12V / S103L / E180V / K241R, I12V / S103V / V146I / K241R, I12V / S103V / E180L / K241R, I12V / S103V / E180V / K241R, I12V / V146I, I12V / V146I / E180L / K241R, I12V / V146I / S232G / K241R, I12V / V146I / K241R, I12V / E180L, I12V / E180L / S232G / K241R, I12V / E180L / K241R, I12V / E180V,Comprising an amino acid sequence having one or more residue differences compared to SEQ ID NO: 860, selected from I12V / E180V / S232G / K241R, I12V / E180V / K241R, I12V / K241R, L18V, S43M / T63P / S103L / E180V / S232G / K241R, S43M / S103L / V146I / E180V, S43M / S103L / E180V, S43M / S103L / E180V / K241R, S43M / S103V / V146I / E180V, S43M / E180L / S232G, S43Q / S103L / E180V / S232G / K241R, S43Q / S103V / V146I / E180L, S43Q / S103V / E180L / K241R, S43Q / S103V / E180V / S232G / K241R, T63P / S103L / E180L / S232G, S103A, S103L, S103L / V146I / E180V, S103L / V146I / E180V / K241R, S103L / E180L, S103L / E180V / K241R, S103L / K241R, S103N, S103T, S103V / V146I / E180L / K241R, S103V / E180V / K241R, Y118L, Y118V, K119H, S122H, T144H, T144V, A145V, V146I / E180L / K241R, A154R, N160L, R172C, R172L, R172V, E180L / S232G, E180L / K241R, E180V / S232G, E180V / S232G / K241R, Q208F, S211F, and R212P.,
[0145] In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence of SEQ ID NO: 1120 and one or more residue differences compared to SEQ ID NO: 1120, selected from SEQ ID NOs: 3, 7, 7 / 158, 7 / 158 / 175, 7 / 158 / 197 / 202, 27 / 43 / 115 / 172, 27 / 43 / 172, 29, 32, 43 / 110 / 115 / 172, 44, 46, 48 / 172, 55, 59, 67, 71, 72, 73, 94 / 103 / 118 / 146, 103, 103 / 144 / 146 / 172 / 180, 103 / 146 / 154 / 172 / 180 / 232, 103 / 154, 118, 118 / 144 / 215, 118 / 172 / 180, 137, 142, 146, 146 / 154, 154, 158, 172 / 190, 184, 198, 205, 216, 223, 226, 229, 247, 261, and 263.In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 1120 and one or more residue differences compared to SEQ ID NO: 1120, selected from SEQ ID NOs: 3D, 7V, 7V / 158Y, 7V / 158Y / 175P, 7V / 158Y / 197V / 202I, 27E / 43I / 115S / 172L, 27E / 43I / 172L, 29G, 29R, 32G, 32R, 43I / 110T / 115S / 172L, 44S, 46H, 46V, 48I / 172L, 55A, 55G, 55R, 55S, 59G, 59T, 67N, 71S, 72G, 73R, 94Q / 103L / 118V / 146I, 103L, 103L / 144V / 146I / 172L / 180V, 103L / 146I / 154R / 172L / 180V / 232G, 103L / 154R, 118V, 118V / 144V / 215H, 118V / 172L / 180V, 137M, 137R, 142V, 146I, 146I / 154R, 154R, 158Y, 172L / 190S, 184S, 198L, 205G, 205L, 216M, 216V, 223A, 223C, 223E, 223L, 223M, 223S, 226L, 226V, 229L, 229V, 229W, 247V, 261G, 261H, 263E, 263G, 263L, and 263S.In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 1120, and one or more residue differences compared to SEQ ID NO: 1120 selected from SEQ ID NO: S3D, I7V, I7V / F158Y, I7V / F158Y / K175P, I7V / F158Y / I197V / V202I, Q27E / S43I / A115S / R172L, Q27E / S43I / R172L, A29G, A29R, N32G, N32R, S43I / I110T / A115S / R172L, N44S, R46H, R46V, L48I / R172L, E55A, E55G, E55R, E55S, E59G, E59T, S67N, G71S, K72G, K73R, P94Q / V103L / Y118V / V146I, V103L, V103L / T144V / V146I / R172L / L180V, V103L / V146I / A154R / R172L / L180V / S232G, V103L / A154R, Y118V, Y118V / T144V / Q215H, Y118V / R172L / L180V, T137M, T137R, T142V, V146I, V146I / A154R, A154R, F158Y, R172L / T190S, D184S, R198L, E205G, E205L, L216M, L216V, K223A, K223C, K223E, K223L, K223M, K223S, K226L, K226V, C229L, C229V, C229W, Q247V, A261G, A261H, A263E, A263G, A263L, and A263S.
[0146] In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 1120, which is a reference sequence, and one or more residue differences compared to SEQ ID NO: 1120, selected from SEQ ID NO: 73, 80, 103 / 146 / 209, 103 / 172 / 209, 118 / 119 / 172 / 209, 135, 144 / 209 / 232, 172 / 180 / 209, 181, and 223. In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 1120, which is a reference sequence, and one or more residue differences compared to SEQ ID NO: 1120, selected from SEQ ID NO: 73R, 80A, 80G, 103L / 146I / 209C, 103L / 172C / 209C, 118V / 119H / 172C / 209C, 135V, 144V / 209C / 232G, 172C / 180V / 209C, 181S, and 223E. In some embodiments, the engineered nitroaldolase polypeptide comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to SEQ ID NO: 1120, which is a reference sequence, and one or more residue differences compared to SEQ ID NO: 1120, selected from SEQ ID NO: K73R, S80A, S80G, V103L / V146I / I209C, V103L / R172C / I209C, Y118V / K119H / R172C / I209C, N135V, T144V / I209C / S232G, R172C / L180V / I209C, Q181S, and K223E.
[0147] In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 1348 and one or more residue differences compared to SEQ ID NO: 1348, selected from SEQ ID NOs: 2, 7 / 46, 39, 46, 46 / 55, 46 / 55 / 72, 46 / 55 / 92, 46 / 55 / 118, 46 / 55 / 226 / 257 / 263, 46 / 55 / 263, 46 / 70, 46 / 70 / 92 / 118 / 261 / 263, 46 / 70 / 175 / 263, 46 / 72, 46 / 118, 55 / 257 / 261, 55 / 261 / 263, 55 / 263, 105, 118, 118 / 226 / 261 / 263, and 263. In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 1348 and one or more residue differences compared to SEQ ID NO: 1348, selected from SEQ ID NOs: 2G, 7V / 46H, 39F, 39G, 39M, 39Q, 39S, 39T, 39V, 46H, 46H / 55R, 46H / 55R / 72C, 46H / 55R / 92A, 46H / 55R / 118V, 46H / 55R / 226V / 257Q / 263L, 46H / 55R / 263L, 46H / 70N, 46H / 70N / 92A / 118V / 261G / 263L, 46H / 70N / 175P / 263G, 46H / 72C, 46H / 118V, 55R / 257Q / 261G, 55R / 261G / 263I, 55R / 263G, 105G, 118V, 118V / 226V / 261G / 263I, and 263G.In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 1348, and one or more residue differences compared to SEQ ID NO: 1348, selected from SEQ ID NO: V2G, I7V / R46H, W39F, W39G, W39M, W39Q, W39S, W39T, W39V, R46H, R46H / E55R, R46H / E55R / K72C, R46H / E55R / K92A, R46H / E55R / Y118V, R46H / E55R / K226V / S257Q / A263L, R46H / E55R / A263L, R46H / E70N, R46H / E70N / K92A / Y118V / A261G / A263L, R46H / E70N / K175P / A263G, R46H / K72C, R46H / Y118V, E55R / S257Q / A261G, E55R / A261G / A263I, E55R / A263G, A105G, Y118V, Y118V / K226V / A261G / A263I, and A263G.
[0148] In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 1348 and one or more residue differences compared to SEQ ID NO: 1348, selected from SEQ ID NO: 2, 7 / 46, 7 / 46 / 175, 22 / 46 / 55 / 118, 27 / 46 / 72, 27 / 46 / 146 / 158 / 202, 27 / 46 / 175, 27 / 118, 29 / 70 / 118, 39, 46, 46 / 55, 46 / 55 / 59 / 261 / 263, 46 / 55 / 72, 46 / 55 / 92, 46 / 55 / 118, 46 / 55 / 118 / 263, 46 / 55 / 146, 46 / 55 / 146 / 202, 46 / 55 / 226 / 257 / 263, 46 / 55 / 263, 46 / 70, 46 / 70 / 92 / 118 / 261 / 263, 46 / 70 / 175 / 263, 46 / 72, 46 / 92, 46 / 92 / 175, 46 / 118, 46 / 146, 46 / 146 / 257 / 263, 55 / 70 / 118, 55 / 257 / 261, 55 / 257 / 263, 55 / 261 / 263, 55 / 263, 70, 92 / 118, 105, 118, 118 / 226 / 261 / 263, 146 / 257 / 263, 146 / 261 / 263, 192 / 257 / 261, 202, 257 / 261 / 263, 261 / 263, and 263.In some embodiments, the engineered nitroaldolase comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence of SEQ ID NO: 1348, and one or more residue differences compared to SEQ ID NO: 1348, selected from SEQ ID NOs: 2G, 7V / 46H, 7V / 46H / 175P, 22V / 46H / 55R / 118V, 27G / 46H / 72C, 27G / 46H / 146I / 158Y / 202I, 27G / 46H / 175P, 27G / 118V, 29V / 70N / 118V, 39M, 39Q, 39S, 39T, 39V, 46H, 46H / 55R, 46H / 55R / 59T / 261G / 263L, 46H / 55R / 72C, 46H / 55R / 92A, 46H / 55R / 118V, 46H / 55R / 118V / 263S, 46H / 55R / 146I, 46H / 55R / 146I / 202I, 46H / 55R / 226V / 257Q / 263L, 46H / 55R / 263L, 46H / 70N, 46H / 70N / 92A / 118V / 261G / 263L, 46H / 70N / 175P / 263G, 46H / 72C, 46H / 92A, 46H / 92A / 175P, 46H / 118V, 46H / 146I, 46H / 146I / 257Q / 263I, 55R / 70N / 118V, 55R / 257Q / 261G, 55R / 257Q / 263L, 55R / 261G / 263I, 55R / 263G, 70N, 92A / 118V, 105G, 118V, 118V / 226V / 261G / 263I, 146I / 257Q / 263S, 146I / 261G / 263L, 192V / 257Q / 261G, 202I, 257Q / 261G / 263G, 261G / 263I, and 263G.In some embodiments, the engineered nitroaldolase polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity to the reference sequence SEQ ID NO: 1348 and has one or more residue differences compared to SEQ ID NO: 1348, selected from SEQ ID NO: V2G, I7V / R46H, I7V / R46H / K175P, L22V / R46H / E55R / Y118V, Q27G / R46H / K72C, Q27G / R46H / V146I / F158Y / V202I, Q27G / R46H / K175P, Q27G / Y118V, A29V / E70N / Y118V, W39M, W39Q, W39S, W39T, W39V, R46H, R46H / E55R, R46H / E55R / E59T / A261G / A263L, R46H / E55R / K72C, R46H / E55R / K92A, R46H / E55R / Y118V, R46H / E55R / Y118V / A263S, R46H / E55R / V146I, R46H / E55R / V146I / V202I, R46H / E55R / K226V / S257Q / A263L, R46H / E55R / A263L, R46H / E70N, R46H / E70N / K92A / Y118V / A261G / A263L, R46H / E70N / K175P / A263G, R46H / K72C, R46H / K92A, R46H / K92A / K175P, R46H / Y118V, R46H / V146I, R46H / V146I / S257Q / A263I, E55R / E70N / Y118V, E55R / S257Q / A261G, E55R / S257Q / A263L, E55R / A261G / A263I, E55R / A263G, E70N, K92A / Y118V, A105G, Y118V, Y118V / K226V / A261G / A263I, V146I / S257Q / A263S, V146I / A261G / A263L, E192V / S257Q / A261G, V202I, S257Q / A261G / A263G, A261G / A263I, and A263G.
[0149] As will be understood by those skilled in the art, in some embodiments, one or a combination of the above-described residue differences selected may be kept (i.e., maintained) as an essential feature in the engineered nitroaldolase, and may be additional residue differences at other residue positions incorporated into the sequence to generate additional engineered nitroaldolase polypeptides having improved properties. Thus, for any engineered nitroaldolase containing one or a subset of the above-described residue differences, it should be understood that the present invention contemplates other engineered nitroaldolases that include one or a subset of the residue differences and additionally include one or more residue differences at other residue positions disclosed herein.
[0150] As described above, the engineered nitroaldolase polypeptide can also convert a substrate (e.g., compound (1) and compound (2)) to a product (e.g., compound (3)). In some embodiments, the engineered nitroaldolase polypeptide can convert a substrate compound to a product compound with an activity that is at least 1.2-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, or higher compared to the activity of the reference polypeptide of SEQ ID NO: 2, 24, 496, 860, 1120, or 1348.
[0151] In some embodiments, the engineered nitroaldolase polypeptide capable of converting a substrate compound to a product compound with an activity that is at least 2-fold compared to SEQ ID NO: 2, 24, 496, 860, 1120, or 1348 includes an amino acid sequence selected from the even-numbered sequences among SEQ ID NOs: 4 to 1484.
[0152] In some embodiments, the engineered nitroaldolase polypeptide can convert a substrate compound to a product compound with stereoselectivity for the (S)-product, as compared to the reference polypeptides of SEQ ID NO: 2, 24, 496, 860, 1120, or 1348. In some embodiments, the engineered nitroaldolase polypeptide can convert a substrate compound to a product compound with at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or higher stereoselectivity for the (S)-product (measured by the amount of (S)-product relative to the total amount of product).
[0153] In some embodiments, the engineered nitroaldolase polypeptide capable of converting a substrate compound to a product compound with at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or higher stereoselectivity for the (S)-product comprises an amino acid sequence selected from the even-numbered sequences in SEQ ID NOs: 4 - 1484.
[0154] In some embodiments, the engineered nitroaldolase has an amino acid sequence that contains one or more residue differences as compared to SEQ ID NO: 2, 24, 496, 860, 1120, or 1348 and increases the expression of engineered nitroaldolase activity in a bacterial host cell, particularly E. coli.
[0155] In some embodiments, the engineered nitroaldolase polypeptide having improved properties has an amino acid sequence that comprises a sequence selected from the even-numbered sequences within the range of SEQ ID NOs: 4 - 1484.
[0156] In some embodiments, the engineered nitroaldolase has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to one of the even-numbered sequences within the range of SEQ ID NOs: 4-1484, as provided in the Examples, and contains amino acid residue differences present in any one of the even-numbered sequences within the range of SEQ ID NOs: 4-1484, as compared to SEQ ID NOs: 2, 24, 496, 860, 1120 or 1348.
[0157] In addition to the residue positions specified above, any of the engineered nitroaldolase polypeptides disclosed herein may further include other residue differences at other residue positions (i.e., residue positions other than those included herein) as compared to SEQ ID NOs: 2, 24, 496, 860, 1120, and / or 1348. Residue differences at these other residue positions can result in additional variations of the amino acid sequence without adversely affecting the ability of the polypeptide to convert substrate to product. Thus, in some embodiments, in addition to the amino acid residue differences present in any one of the engineered nitroaldolase polypeptides selected from the even-numbered sequences within the range of SEQ ID NOs: 4 - 1484, the sequence may further include 1 - 2, 1 - 3, 1 - 4, 1 - 5, 1 - 6, 1 - 7, 1 - 8, 1 - 9, 1 - 10, 1 - 11, 1 - 12, 1 - 14, 1 - 15, 1 - 16, 1 - 18, 1 - 20, 1 - 22, 1 - 24, 1 - 26, 1 - 30, 1 - 35, 1 - 40, 1 - 45, or 1 - 50 residue differences at other amino acid residue positions as compared to SEQ ID NOs: 2, 24, 496, 860, 1120, or 1348. In some embodiments, the number of amino acid residue differences as compared to the reference sequence can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 residue positions. In some embodiments, the number of amino acid residue differences as compared to the reference sequence can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 residue positions. The residue differences at these other positions can be conservative changes or non-conservative conversions. In some embodiments, the residue differences can include conservative substitutions and non-conservative substitutions as compared to the nitroaldolase polypeptides of SEQ ID NOs: 2, 24, 496, 860, 1120, or 1348.
[0158] In some embodiments, the invention also provides an engineered polypeptide comprising a fragment of any of the engineered nitroaldolase polypeptides described herein that retains the functional activity and / or improved properties of the engineered nitroaldolase. Thus, in some embodiments, the invention provides a polypeptide fragment capable of converting a substrate to a product under suitable reaction conditions, the fragment comprising at least about 90%, 95%, 96%, 97%, 98% or 99% of the full-length amino acid sequence of an engineered nitroaldolase of the invention, such as an exemplary engineered nitroaldolase polypeptide selected from the even-numbered sequences within the range of SEQ ID NOs: 4-1484. In some embodiments, the engineered nitroaldolase may have an amino acid sequence comprising a deletion in any one of the nitroaldolase polypeptide sequences described herein, such as an exemplary engineered polypeptide of an even-numbered sequence within the range of SEQ ID NOs: 4-1484.
[0159] Accordingly, for any and all embodiments of the engineered nitroaldolase polypeptide of the present invention, the amino acid sequence may contain deletions of one or more amino acids, two or more amino acids, three or more amino acids, four or more amino acids, five or more amino acids, six or more amino acids, eight or more amino acids, ten or more amino acids, fifteen or more amino acids, or twenty or more amino acids, up to 10% of the total number of amino acids, up to 20% of the total number of amino acids, or up to 30% of the total number of amino acids, provided that the related functional activity and / or improved properties of the engineered nitroaldolase described herein are maintained. In some embodiments, the deletion may comprise 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-30, 1-35, 1-40, 1-45, or 1-50 amino acid residues. In some embodiments, the number of deletions may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45 or 50 amino acid residues. In some embodiments, the deletion may comprise a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24 or 25 amino acid residues.
[0160] In some embodiments, the engineered nitroaldolase polypeptide described herein can have an amino acid sequence that includes an insertion as compared to any one of the engineered nitroaldolase polypeptides described herein, such as an exemplary engineered polypeptide of an even-numbered sequence within the range of SEQ ID NOs: 4 to 1484. Thus, for any and all embodiments of the nitroaldolase polypeptides of the present disclosure, the insertion can include one or more amino acids, two or more amino acids, three or more amino acids, four or more amino acids, five or more amino acids, six or more amino acids, eight or more amino acids, ten or more amino acids, fifteen or more amino acids, twenty or more amino acids, thirty or more amino acids, forty or more amino acids, or fifty or more amino acids, provided that the related functional activity and / or improved properties of the engineered nitroaldolase described herein are maintained. The insertion can be an insertion at the amino or carboxy terminus of the nitroaldolase polypeptide, or it can be an internal insertion.
[0161] In some embodiments, the engineered nitroaldolase herein can have an amino acid sequence selected from even-numbered sequences within the range of SEQ ID NOs: 4 to 1484 and optionally including one or several (e.g., up to 3, 4, 5, or up to 10) amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence can optionally have 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-30, 1-35, 1-40, 1-45, or 1-50 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the number of amino acid sequences can optionally have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence can optionally have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the substitutions can be conservative substitutions or non-conservative substitutions.
[0162] In the above embodiments, the reaction conditions suitable for the engineered polypeptide are provided as described in the examples herein.
[0163] In some embodiments, the polypeptide of the present invention is a fusion polypeptide in which the engineered polypeptide is fused to other polypeptides, such as, by way of example and not limitation, antibody tags (e.g., myc epitope), purification sequences (e.g., His tag for binding to metal), and cell localization signals (e.g., secretion signal). Thus, the engineered polypeptides described herein can be used with or without fusion to other polypeptides.
[0164] It should be understood that the polypeptides described herein are not limited to the amino acids encoded by the genetic code. In addition to the amino acids encoded by the genetic code, the polypeptides described herein may be composed in whole or in part of naturally occurring amino acids and / or synthetic non-coded amino acids. Certain commonly encountered non-coded amino acids that may constitute the polypeptides described herein include, but are not limited to, the D stereoisomers of the amino acids encoded by the genetic code; 2,3-diaminopropionic acid (Dpr); α-aminoisobutyric acid (Aib); ε-aminohexanoic acid (Aha); δ-aminovaleric acid (Ava); N-methylglycine or sarcosine (MeGly or Sar); ornithine (Orn); citrulline (Cit); t-butylalanine (Bua); t-butylglycine (Bug); N-methylisoleucine (MeIle); phenylglycine (Phg); cyclohexylalanine (Cha); norleucine (Nle); naphthylalanine (Nal); 2-chlorophenylalanine (Ocf); 3-chlorophenylalanine (Mcf); 4-chlorophenylalanine (Pcf); 2-fluorophenylalanine (Off); 3-fluorophenylalanine (Mff); 4-fluorophenylalanine (Pff); 2-bromophenylalanine (Obf); 3-bromophenylalanine (Mbf); 4-bromophenylalanine (Pbf); 2-methylphenylalanine (Omf); 3-methylphenylalanine (Mmf); 4-methylphenylalanine (Pmf); 2-nitrophenylalanine (Onf); 3-nitrophenylalanine (Mnf); 4-nitrophenylalanine (Pnf); 2-cyanophenylalanine (Ocf); 3-cyanophenylalanine (Mcf); 4-cyanophenylalanine (Pcf); 2-trifluoromethylphenylalanine (Otf); 3-trifluoromethylphenylalanine (Mtf); 4-trifluoromethylphenylalanine (Ptf); 4-aminophenylalanine (Paf); 4-iodophenylalanine (Pif); 4-aminomethylphenylalanine (Pamf); 2,4-dichlorophenylalanine (Opef); 3,4-dichlorophenylalanine (Mpcf); 2,4-difluorophenylalanine (Opff);3,4-difluorophenylalanine (Mpff); pyrid-2-ylalanine (2pAla); pyrid-3-ylalanine (3pAla); pyrid-4-ylalanine (4pAla); naphth-1-ylalanine (1nAla); naphth-2-ylalanine (2nAla); thiazolylalanine (taAla); benzothienylalanine (bAla); thienylalanine (tAla); furylalanine (fAla); homophenylalanine (hPhe); homotyrosine (hTyr); homotryptophan (hTrp); pentafluorophenylalanine (5ff); styrylkalanine (sAla); authrylalanine (aAla); 3,3-diphenylalanine (Dfa); 3-amino-5-phenylpentanoic acid (Afp); penicillamine (Pen); 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid (Tic); β-2-thienylalanine (Thi); methionine sulfoxide (Mso); N(w)-nitroarginine (nArg); homolysine (hLys); phosphonomethylphenylalanine (pmPhe); phosphoserine (pSer); phosphothreonine (pThr); homoaspartic acid (hAsp); homoglutanic acid (hGlu); 1-aminocyclopent-(2 or 3)-ene-4-carboxylic acid; pipecolic acid (PA), azetidine-3-carboxylic acid (ACA); 1-aminocyclopentane-3-carboxylic acid; allylglycine (aGly); propargylglycine (pgGly); homoalanine (hAla); norvaline (nVal); homoleucine (hLeu), homovaline (hVal); homoisoleucine (hIle); homoarginine (hArg); N-acetyllysine (AcLys); 2,4-diaminobutyric acid (Dbu); 2,3-diaminobutyric acid (Dab); N-methylvaline (MeVal); homocysteine (hCys); homoserine (hSer);Examples include hydroxyproline (Hyp) and homoproline (hPro). Additional non-coded amino acids that can constitute the polypeptides described herein will be apparent to those skilled in the art (see, for example, Fasman, CRC Practical Handbook of Biochemistry and Molecular Biology, CRC Press, Boca Raton, FL, pp. 3-70
[1989] , and the various amino acids provided in the references cited therein, all of which are incorporated herein by reference). These amino acids can be in either the L-configuration or the D-configuration.;
[0165] It will also be appreciated by those skilled in the art that amino acids or residues with side-chain protecting groups can also constitute the polypeptides described herein. Non-limiting examples of such protected amino acids belonging to the aromatic category in this case include, but are not limited to, Arg(tos), Cys(methylbenzyl), Cys(nitropyridinesulfenyl), Glu(δ-benzyl ester), Gln(xanthyl), Asn(N-δ-xanthyl), His(bom), His(benzyl), His(tos), Lys(fmoc), Lys(tos), Ser(O-benzyl), Thr(O-benzyl), and Tyr(O-benzyl) (the protecting groups are listed within parentheses).
[0166] Sterically constrained non-coded amino acids that can constitute the polypeptides described herein include, but are not limited to, N-methyl amino acid (L-configuration); 1-aminocyclopenta-(2 or 3)-ene-4-carboxylic acid; pipecolic acid; azetidine-3-carboxylic acid; homoproline (hPro); and 1-aminocyclopentane-3-carboxylic acid.
[0167] In some embodiments, the engineered polypeptide can be in various forms such as, for example, a substantially purified enzyme, whole cells transformed with a gene encoding the enzyme, and / or isolated preparations such as cell extracts and / or lysates of such cells. The enzyme can be lyophilized, spray-dried, precipitated, or in the form of a crude paste, as further discussed below.
[0168] In some embodiments, the engineered polypeptide can be in the form of a biocatalytic composition. In some embodiments, the biocatalytic composition comprises (a) means for converting a ketone compound and an alkane compound to a β-nitroalcohol intermediate upon contact with a nitroaldolase polypeptide and (b) a suitable buffer. In some embodiments, the biocatalytic composition comprises a nitroaldolase having activity against a ketone substrate. In some further embodiments, the β-nitroalcohol intermediate is further reduced to a β-hydroxy-α-aminoalcohol. In some embodiments, the biocatalytic composition comprises a sodium citrate buffer.
[0169] In some embodiments, the engineered polypeptide can be provided on a solid support, such as a membrane, resin, solid support, or other solid phase material. The solid support can be composed of organic polymers such as polystyrene, polyethylene, polypropylene, polyfluoroethylene, polyethyleneoxy, and polyacrylamide, as well as copolymers and grafts thereof. The solid support can be inorganic, such as glass, silica, controlled pore glass (CPG), reversed-phase silica, or a metal such as gold or platinum. The solid support can be in the form of beads, spheres, particles, granules, gels, membranes, or surfaces. The surface can be planar, substantially planar, or non-planar. The solid support can be porous or non-porous and can have swelling or non-swelling characteristics. The solid support can also be configured in the form of wells, depressions, or other receptacles, containers, features, or locations.
[0170] In some embodiments, the engineered nitroaldolase polypeptides of the invention can be immobilized on a solid support such that they retain their improved activity and / or other improved properties as compared to the reference polypeptides of SEQ ID NO: 2, 24, 496, 860, 1120 or 1348. In such embodiments, the immobilized polypeptides can facilitate the biocatalytic conversion of a substrate compound or other suitable substrate to a product and, after the reaction is complete, can be easily retained (e.g., by retaining the beads to which the polypeptide is immobilized) and thus reused or recycled in subsequent reactions. Such immobilized enzyme processes allow for further efficiency and cost savings. Accordingly, it is further contemplated that any of the methods of using the nitroaldolase polypeptides of the invention can be carried out using a nitroaldolase polypeptide bound or immobilized on a solid support.
[0171] Methods of enzyme immobilization are well known in the art. The engineered polypeptide can be non-covalently or covalently attached. A variety of methods for conjugation and immobilization of enzymes to solid supports (e.g., resins, membranes, beads, glass, etc.) are well known in the art (e.g., Yi et al., Proc. Biochem., 42(5): 895-898
[2007] ; Martin et al., Appl. Microbiol. Biotechnol., 76(4): 843-851
[2007] ; Koszelewski et al., J. Mol. Cat. B: Enzymatic, 63: 39-44
[2010] ; Truppo et al., Org. Proc. Res. Dev., published online: dx.doi.org / 10.1021 / op200157c; Hermanson, Bioconjugate Techniques, 2 ndSee, e.g., “Protein Engineering,” 2nd ed., Academic Press, Cambridge, MA
[2008] ; Mateo et al., Biotechnol. Prog., 18(3):629-34
[2002] ; and “Bioconjugation Protocols: Strategies and Methods,” In Methods in Molecular Biology, Niemeyer (ed.), Humana Press, New York, NY
[2004] , the disclosures of each of which are incorporated herein by reference). Solid supports useful for immobilizing the engineered nitroaldolase of the present invention include, but are not limited to, polymethacrylates having epoxide functional groups, polymethacrylates having aminoepoxide functional groups, styrene / DVB copolymers having octadecyl functional groups, or beads or resins comprising polymethacrylates. Exemplary solid supports useful for immobilizing the engineered nitroaldolase polypeptide of the present invention include, but are not limited to, chitosan beads, Eupergit C, and SEPABEAD (Mitsubishi Chemical Corporation), and SEPABEAD includes the following different types of SEPABEAD: EC-EP, EC-HFA / S, EXA252, EXE119, and EXE120.
[0172] In some embodiments, the polypeptides described herein are provided in the form of a kit. The enzymes in the kit may be present individually or as a plurality of enzymes. The kit may further include reagents for performing the enzyme reaction, substrates for assessing the activity of the enzyme, and reagents for detecting the product. The kit may also include a reagent dispenser and instructions for using the kit.
[0173] In some embodiments, the kit of the present invention comprises an array containing a plurality of different nitroaldolase polypeptides at different addressable positions, wherein the different polypeptides are different variants of a reference sequence, each having at least one different improved enzymatic property. In some embodiments, the plurality of polypeptides immobilized on a solid support are configured at various positions on the array that are addressable by robotic delivery of reagents or by detection methods and / or instrumentation. The array can be used to test various substrate compounds for conversion by the polypeptides. Such arrays containing a plurality of engineered polypeptides and methods of their use are known in the art (see, for example, WO2009 / 008908A2).
[0174] Polynucleotides, expression vectors and host cells encoding engineered nitroaldolases In another aspect, the present invention provides a polynucleotide encoding an engineered nitroaldolase polypeptide as described herein. The polynucleotide may be operably linked to one or more heterologous regulatory sequences that control gene expression to create a recombinant polynucleotide capable of expressing the polypeptide. An expression construct containing at least one heterologous polynucleotide encoding an engineered nitroaldolase is introduced into a suitable host cell for expressing the corresponding nitroaldolase polypeptide.
[0175] As will be apparent to those skilled in the art, knowledge of the protein sequence and the codons corresponding to the various amino acids provides an explanation for all polynucleotides capable of encoding the polypeptide of interest. The degeneracy of the genetic code, where the same amino acid is encoded by alternative or synonymous codons, allows for the creation of a very large number of nucleic acids, all of which encode an improved nitroaldolase enzyme. Thus, given knowledge of a particular amino acid sequence, one skilled in the art will be able to create any number of different nucleic acids by simply modifying the sequence of one or more codons in a way that does not change the amino acid sequence of the protein. In this regard, the present invention specifically contemplates every possible variation of polynucleotides that can be made to encode the polypeptides described herein by selecting combinations based on the possible codon options, and every such variation is specifically disclosed as being any polypeptide disclosed in the sequence listing incorporated herein by reference and in the range of SEQ ID NOs: 4 to 1484 as even-numbered sequences that include the amino acid sequences presented in Tables 2-1, 2-2, 3-1, 3-2, 4-1, 4-2, 5-1, 5-2, 6-1, 6-2, 7-1 and 7-2.
[0176] In various embodiments, the codons are preferably selected to be compatible with the host cell in which the protein is to be produced. For example, the preferred codons used in bacteria are those used to express genes in bacteria, the preferred codons used in yeast are those used for expression in yeast, and the preferred codons used in mammals are those used for expression in mammalian cells. In some embodiments, since the native sequence is likely to contain preferred codons and the use of preferred codons may not be required for all amino acid residues, it is not necessary to replace all codons to optimize the codon usage frequency of nitroaldolase. Therefore, the codon-optimized polynucleotide encoding the nitroaldolase enzyme may contain preferred codons at about 40%, 50%, 60%, 70%, 80%, or more than 90% of the codon positions in the full-length coding region.
[0177] In some embodiments, the polynucleotide comprises a codon-optimized nucleotide sequence encoding the nitroaldolase polypeptide amino acid sequence represented by SEQ ID NO: 2, 24, 496, 860, 1120, or 1348. In some embodiments, the polynucleotide has a nucleic acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity to the codon-optimized nucleic acid sequence encoding an even-numbered sequence within the range of SEQ ID NOs: 4 - 1484. In some embodiments, the polynucleotide has a nucleic acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity to the codon-optimized nucleic acid sequence in the odd-numbered sequences within the range of SEQ ID NOs: 3 - 1483. In some embodiments, the codon-optimized sequence of the odd-numbered sequences within the range of SEQ ID NOs: 3 - 1483 results in a preparation of an enzyme that can enhance the expression of the encoded nitroaldolase and thus convert the substrate to the product.
[0178] In some embodiments, the polynucleotide can hybridize under highly stringent conditions with a reference sequence selected from the odd-numbered sequences in SEQ ID NOs: 3 to 1483 or its complementary sequence, and encodes a nitroaldolase polypeptide.
[0179] In some embodiments, as described above, the polynucleotide encodes an engineered nitroaldolase polypeptide having improved properties as compared to SEQ ID NOs: 2, 24, 496, 860, 1120 or 1348, and the polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher identity with an amino acid sequence to a reference sequence selected from SEQ ID NOs: 2, 24, 496, 860, 1120 or 1348, and includes one or more residue differences as compared to SEQ ID NOs: 2, 24, 496, 860, 1120 or 1348, and the sequence is selected from the even-numbered sequences within the range of SEQ ID NOs: 4 to 1484. In some embodiments, the reference amino acid sequence is selected from the even-numbered sequences within the range of SEQ ID NOs: 4 to 1484. In some embodiments, the reference amino acid sequence is SEQ ID NO: 2, while in some other embodiments, the reference sequence is SEQ ID NO: 24, and in some other embodiments, the reference sequence is SEQ ID NO: 496. In some embodiments, the reference amino acid sequence is SEQ ID NO: 860, while in some other embodiments, the reference sequence is SEQ ID NO: 1120, and in some other embodiments, the reference sequence is SEQ ID NO: 1348.
[0180] In some embodiments, the polynucleotide encodes a nitroaldolase polypeptide that can convert one or more substrates into products and has improved properties as compared to SEQ ID NOs: 2, 24, 496, 860, 1120 or 1348, and the polypeptide includes an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity to the reference sequence SEQ ID NOs: 2, 24, 496, 860, 1120 or 1348.
[0181] In some embodiments, the polynucleotide encoding the engineered nitroaldolase comprises a polynucleotide sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 1, 23, 495, 860, 1119, and / or 1347. In some embodiments, the polynucleotide encoding the engineered nitroaldolase comprises a polynucleotide sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to a sequence selected from the odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483. In some embodiments, the polynucleotide encoding the engineered nitroaldolase comprises a polynucleotide sequence selected from the odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483.
[0182] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 2, with one or more residue differences compared to SEQ ID NO: 2 at positions 7 / 9, 7 / 9 / 33 / 68 / 72 / 158, 7 / 9 / 33 / 72 / 175 / 197 / 230, 7 / 9 / 33 / 72 / 230, 7 / 9 / 68 / 72 / 197, 7 / 9 / 158 / 175, 7 / 9 / 175, 7 / 9 / 197, 7 / 33, 7 / 68 / 72 / 175, 7 / 72, 7 / 72 / 97 / 175 / 230, 9, 9 / 11 / 33 / 72 / 197 / 230, 9 / 33 / 68 / 158 / 197, 9 / 33 / 72 / 230, 9 / 33 / 97 / 230, 9 / 68 / 72, 9 / 97 / 175 / 202 / 218, 9 / 175, 9 / 202, 11, 13, 14 / 33 / 68 / 72 / 97, 18, 18 / 33 / 68 / 97 / 158 / 175 / 239, 18 / 33 / 72 / 97 / 158 / 175, 18 / 33 / 72 / 97 / 158 / 239, 18 / 33 / 72 / 97 / 175 / 202, 18 / 33 / 72 / 175 / 202 / 237 / 239, 18 / 33 / 72 / 197 / 237, 18 / 68 / 72, 18 / 68 / 72 / 158 / 175 / 237 / 239, 18 / 68 / 72 / 175 / 202 / 239, 18 / 68 / 72 / 239, 18 / 68 / 175 / 223 / 237 / 239, 18 / 68 / 239, 18 / 72 / 97 / 175 / 237 / 239, 18 / 72 / 175 / 197, 18 / 72 / 237, 18 / 72 / 239, 18 / 97 / 175 / 237 / 239, 18 / 97 / 223 / 237 / 239, 18 / 175, 18 / 197 / 239, 18 / 239, 20 / 124, 33 / 68 / 72, 33 / 68 / 72 / 175 / 230, 33 / 68 / 223 / 239, 33 / 72, 33 / 72 / 97 / 197, 33 / 72 / 175,It encodes a nitroaldolase comprising an amino acid sequence having amino acid residues at positions selected from 33 / 175, 39, 44 / 97, 52, 68 / 72, 68 / 72 / 97, 68 / 72 / 97 / 197, 68 / 72 / 197, 68 / 72 / 197 / 239, 68 / 97 / 175, 68 / 97 / 175 / 197 / 223, 72 / 97 / 158 / 223 / 237, 72 / 175 / 197, 97 / 158 / 175 / 197, 97 / 175, 103, 121, 122, 124, 125, 128, 133, 148, 175, 197, and 197 / 202.
[0183] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 2, and has one or more residue differences compared to SEQ ID NO: 2 at residue positions selected from SEQ ID NOs: 7 / 11 / 33 / 72 / 97, 7 / 11 / 33 / 97, 7 / 11 / 72, 7 / 11 / 158, 7 / 11 / 237, 7 / 33 / 72 / 97 / 158 / 175 / 197 / 237 / 239, 9 / 11 / 33, 9 / 11 / 33 / 72, 9 / 11 / 72, 9 / 33 / 68 / 97 / 237 / 239, 9 / 72 / 97 / 158 / 237 / 239, 9 / 218 / 239, 10, 11, 12, 13, 14, 18, 33 / 68 / 72 / 97 / 237 / 239, 33 / 72 / 158 / 239, 39, 40, 41, 48, 49, 51, 54, 54 / 126, 57, 68 / 175 / 239, 82, 84, 85, 86, 103, 104, 105, 107, 115, 117, 119, 119 / 120, 120, 122, 123, 124, 127, 130, 131, 133, 134, 137 / 145, 144, 145, 146, 147, 148, 149, 150, 152, 157, 158, 159, 169, 173, 174, 175, 180, 181, 183, 203, 204, 208, 209, 210, 215, 218, 236, and 237, and encodes a nitroaldolase comprising an amino acid sequence having an amino acid sequence containing residue positions selected therefrom.
[0184] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 24, and has one or more residue differences compared to SEQ ID NO: 24 at residue positions including an amino acid sequence having an amino acid sequence selected from SEQ ID NOs: 9 / 39 / 122 / 157, 18, 18 / 33 / 39 / 230, 18 / 33 / 175, 18 / 39 / 44 / 118, 18 / 175, 20, 33 / 39, 33 / 39 / 44 / 118 / 121, 33 / 39 / 44 / 121, 33 / 39 / 44 / 152, 33 / 39 / 44 / 175, 33 / 39 / 44 / 197, 33 / 39 / 44 / 210 / 230, 33 / 39 / 118 / 121 / 175, 33 / 39 / 121 / 175, 33 / 39 / 197 / 210, 33 / 39 / 210, 33 / 39 / 210 / 230, 33 / 118 / 121, 33 / 118 / 121 / 175, 39, 39 / 44 / 103, 39 / 44 / 118, 39 / 44 / 152 / 230, 39 / 44 / 175 / 197 / 230, 39 / 103, 39 / 103 / 125, 39 / 103 / 125 / 127 / 146, 39 / 103 / 125 / 127 / 146 / 150, 39 / 103 / 125 / 146, 39 / 103 / 127 / 150, 39 / 103 / 150, 39 / 118 / 121, 39 / 197, 39 / 197 / 210, 43, 44 / 103, 81, 103, 103 / 125 / 127, 103 / 125 / 146, 103 / 127, 103 / 150, 118 / 121, 118 / 121 / 175, 118 / 121 / 197, 121, 152, 152 / 197, 154, 159, 163, 175, 210, 232, and 238.
[0185] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 24, and has one or more residue differences compared to SEQ ID NO: 24 at positions 7, 9 / 13, 9 / 13 / 39, 9 / 13 / 39 / 68 / 72 / 239, 9 / 13 / 39 / 68 / 239, 9 / 13 / 39 / 72 / 122, 9 / 13 / 39 / 97 / 239, 9 / 13 / 68 / 72 / 97 / 239, 9 / 13 / 68 / 97 / 134 / 239, 9 / 13 / 68 / 97 / 239, 9 / 13 / 72, 9 / 13 / 134, 11, 12 / 42 / 125 / 127, 12 / 103, 12 / 103 / 150, 13, 13 / 39, 13 / 39 / 68, 13 / 39 / 68 / 72 / 97, 13 / 39 / 72 / 97 / 122 / 134 / 223, 13 / 39 / 97 / 239, 13 / 39 / 239, 13 / 68, 13 / 68 / 72 / 97 / 122 / 239, 16, 18 / 33 / 39 / 44 / 118 / 121 / 175 / 197 / 210, 18 / 33 / 39 / 44 / 121, 18 / 33 / 39 / 112 / 118 / 121 / 197, 18 / 33 / 39 / 121, 18 / 33 / 39 / 121 / 175, 18 / 33 / 39 / 121 / 175 / 197, 18 / 33 / 39 / 121 / 175 / 210, 18 / 33 / 39 / 121 / 210, 18 / 33 / 118 / 121, 18 / 33 / 121, 18 / 33 / 121 / 197, 18 / 33 / 121 / 210 / 230, 18 / 39 / 44 / 118 / 121 / 230, 18 / 39 / 44 / 121 / 175 / 210, 18 / 39 / 118 / 121, 18 / 39 / 118 / 121 / 210, 18 / 39 / 121 / 175, 18 / 39 / 121 / 175 / 210, 18 / 39 / 121 / 197 / 210, 18 / 118 / 121, 18 / 118 / 121 / 175, 18 / 118 / 121 / 210,It encodes a nitroaldolase comprising an amino acid sequence having amino acids at residue positions selected from 18 / 118 / 121 / 230, 18 / 121, 18 / 121 / 175, 18 / 121 / 210, 18 / 121 / 230, 20, 21, 33 / 39 / 44 / 118 / 121, 33 / 39 / 44 / 118 / 121 / 210 / 230, 33 / 39 / 44 / 121, 33 / 39 / 44 / 121 / 152, 33 / 39 / 44 / 121 / 175 / 210, 33 / 39 / 44 / 121 / 230, 33 / 39 / 73 / 121 / 197, 33 / 39 / 118 / 121 / 175, 33 / 39 / 118 / 121 / 175 / 230, 33 / 39 / 121, 33 / 39 / 121 / 175, 33 / 39 / 121 / 210 / 230, 33 / 118 / 121, 33 / 118 / 121 / 175, 33 / 118 / 121 / 197, 33 / 121, 33 / 121 / 230, 37 / 91, 39 / 44 / 118 / 121 / 175 / 210, 39 / 68 / 72 / 239, 39 / 118 / 121, 39 / 118 / 121 / 210, 39 / 121, 39 / 121 / 210 / 230, 43, 44, 76, 79, 101, 118 / 121, 118 / 121 / 175, 118 / 121 / 197, 121, 121 / 175 / 230, 121 / 210, 148, 153, 159, 161, 172, 200, 207, 233, 245, and 248.,
[0186] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 496, and has one or more residue differences compared to SEQ ID NO: 496 at residue positions including an amino acid sequence having an amino acid sequence selected from SEQ ID NOs: 43 / 103, 43 / 103 / 172 / 238, 43 / 103 / 238 / 241, 43 / 238, 81 / 232, 103, 103 / 159 / 238, 103 / 194 / 238, 103 / 238, 122, 232, 238, and 238 / 241.
[0187] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 496, and has one or more residue differences compared to SEQ ID NO: 496 at residue positions selected from SEQ ID NOs: 12, 43, 43 / 47 / 103, 43 / 47 / 103 / 232, 43 / 103, 43 / 103 / 121 / 172 / 241, 43 / 103 / 172, 43 / 103 / 172 / 238, 43 / 103 / 238, 43 / 103 / 238 / 241, 43 / 121, 43 / 194 / 238, 43 / 238, 44 / 103 / 121 / 238, 44 / 103 / 238, 44 / 103 / 238 / 241, 44 / 121 / 238 / 241, 47 / 103, 63, 103, 103 / 121, 103 / 121 / 172 / 238, 103 / 159 / 238, 103 / 172 / 238, 103 / 194 / 238, 103 / 238, 122, 131, 146, 178, 180, and 221, and encodes a nitroaldolase comprising an amino acid sequence having an amino acid sequence containing the residue positions.
[0188] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein.In some embodiments, a polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 860, and has one or more residue differences compared to SEQ ID NO: 860 at residue positions including an amino acid sequence having an amino acid sequence selected from SEQ ID NO: 12, 12 / 43 / 63 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 63 / 103 / 146 / 241, 12 / 43 / 81 / 103 / 241, 12 / 43 / 81 / 146 / 232 / 241, 12 / 43 / 81 / 180 / 241, 12 / 43 / 81 / 232 / 241, 12 / 43 / 81 / 241, 12 / 43 / 103, 12 / 43 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 103 / 146 / 180 / 241, 12 / 43 / 103 / 180 / 241, 12 / 43 / 103 / 241, 12 / 43 / 146 / 180 / 232 / 241, 12 / 43 / 180 / 241, 12 / 63 / 81 / 103 / 241, 12 / 81, 12 / 81 / 180 / 241, 12 / 103 / 146 / 241, 12 / 103 / 180 / 241, 12 / 146, 12 / 146 / 180 / 241, 12 / 180 / 232 / 241, 12 / 241, 18, 43 / 81 / 232 / 241, 43 / 103 / 146 / 180, 43 / 103 / 180 / 241, 43 / 103 / 241, 43 / 180 / 232, 43 / 232 / 241, 63 / 103 / 180 / 232, 78, 81, 81 / 103 / 146, 81 / 146 / 180 / 241, 81 / 146 / 241, 81 / 180 / 241, 84, 103, 103 / 146, 103 / 146 / 180, 103 / 146 / 180 / 241, 103 / 180, 103 / 180 / 241, 103 / 261, 144, 145, 146 / 180 / 232, 146 / 180 / 241, 146 / 241, 150, 166, 172, 180 / 232, 232 / 241, and 241. Such polynucleotides encode a nitroaldolase.
[0189] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved characteristics described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 860, and has one or more residue differences compared to SEQ ID NO: 860 at residue positions including an amino acid sequence selected from SEQ ID NOs: 9, 12, 12 / 43 / 63 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 63 / 103 / 146 / 241, 12 / 43 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 103 / 146 / 180 / 241, 12 / 43 / 103 / 146 / 232 / 241, 12 / 43 / 103 / 146 / 241, 12 / 43 / 103 / 180 / 241, 12 / 43 / 146 / 180 / 232 / 241, 12 / 43 / 146 / 241, 12 / 43 / 180 / 232, 12 / 43 / 180 / 241, 12 / 63 / 103 / 146 / 232 / 241, 12 / 103 / 146 / 241, 12 / 103 / 180 / 241, 12 / 146, 12 / 146 / 180 / 241, 12 / 146 / 232 / 241, 12 / 146 / 241, 12 / 180, 12 / 180 / 232 / 241, 12 / 180 / 241, 12 / 241, 18, 43 / 63 / 103 / 180 / 232 / 241, 43 / 103 / 146 / 180, 43 / 103 / 180, 43 / 103 / 180 / 232 / 241, 43 / 103 / 180 / 241, 43 / 180 / 232, 63 / 103 / 180 / 232, 103, 103 / 146 / 180, 103 / 146 / 180 / 241, 103 / 180, 103 / 180 / 241, 103 / 241, 118, 119, 122, 144, 145, 146 / 180 / 241, 154, 160, 172, 180 / 232, 180 / 232 / 241, 180 / 241, 208, 211, and 212.
[0190] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 1120, and has one or more residue differences compared to SEQ ID NO: 1120 at residue positions including amino acid sequences selected from SEQ ID NOs: 3, 7, 7 / 158, 7 / 158 / 175, 7 / 158 / 197 / 202, 27 / 43 / 115 / 172, 27 / 43 / 172, 29, 32, 43 / 110 / 115 / 172, 44, 46, 48 / 172, 55, 59, 67, 71, 72, 73, 94 / 103 / 118 / 146, 103, 103 / 144 / 146 / 172 / 180, 103 / 146 / 154 / 172 / 180 / 232, 103 / 154, 118, 118 / 144 / 215, 118 / 172 / 180, 137, 142, 146, 146 / 154, 154, 158, 172 / 190, 184, 198, 205, 216, 223, 226, 229, 247, 261, and 263, and encodes a nitroaldolase comprising an amino acid sequence having an amino acid sequence.
[0191] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 1120, and has one or more residue differences compared to SEQ ID NO: 1120 at residue positions selected from SEQ ID NOs: 73, 80, 103 / 146 / 209, 103 / 172 / 209, 118 / 119 / 172 / 209, 135, 144 / 209 / 232, 172 / 180 / 209, 181, and 223, and encodes a nitroaldolase comprising an amino acid sequence having an amino acid sequence containing the residue differences at the selected residue positions.
[0192] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 1348, and has one or more residue differences compared to SEQ ID NO: 1348 at residue positions selected from SEQ ID NOs: 2, 7 / 46, 39, 46, 46 / 55, 46 / 55 / 72, 46 / 55 / 92, 46 / 55 / 118, 46 / 55 / 226 / 257 / 263, 46 / 55 / 263, 46 / 70, 46 / 70 / 92 / 118 / 261 / 263, 46 / 70 / 175 / 263, 46 / 72, 46 / 118, 55 / 257 / 261, 55 / 261 / 263, 55 / 263, 105, 118, 118 / 226 / 261 / 263, and 263, and encodes a nitroaldolase comprising an amino acid sequence having an amino acid sequence containing the residue differences at the selected residue positions.
[0193] In some embodiments, the polynucleotide can hybridize under highly stringent conditions to a reference polynucleotide sequence selected from odd-numbered sequences within the range of SEQ ID NOs: 3 to 1483 or its complement, and encodes a nitroaldolase polypeptide having one or more of the improved properties described herein. In some embodiments, the polynucleotide that can hybridize under highly stringent conditions has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 1348, and has one or more residue differences compared to SEQ ID NO: 1348 at residue positions selected from SEQ ID NOs: 2, 7 / 46, 7 / 46 / 175, 22 / 46 / 55 / 118, 27 / 46 / 72, 27 / 46 / 146 / 158 / 202, 27 / 46 / 175, 27 / 118, 29 / 70 / 118, 39, 46, 46 / 55, 46 / 55 / 59 / 261 / 263, 46 / 55 / 72, 46 / 55 / 92, 46 / 55 / 118, 46 / 55 / 118 / 263, 46 / 55 / 146, 46 / 55 / 146 / 202, 46 / 55 / 226 / 257 / 263, 46 / 55 / 263, 46 / 70, 46 / 70 / 92 / 118 / 261 / 263, 46 / 70 / 175 / 263, 46 / 72, 46 / 92, 46 / 92 / 175, 46 / 118, 46 / 146, 46 / 146 / 257 / 263, 55 / 70 / 118, 55 / 257 / 261, 55 / 257 / 263, 55 / 261 / 263, 55 / 263, 70, 92 / 118, 105, 118, 118 / 226 / 261 / 263, 146 / 257 / 263, 146 / 261 / 263, 192 / 257 / 261, 202, 257 / 261 / 263, 261 / 263, and 263, and encodes a nitroaldolase comprising an amino acid sequence having an amino acid sequence containing the residue positions.
[0194] In some embodiments, a polynucleotide capable of hybridizing under highly stringent conditions encodes an engineered nitroaldolase polypeptide having improved properties, comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 2, 24, 496, 860, 1120, or 1348. In some embodiments, the polynucleotide encodes a polypeptide described herein, but has at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or greater sequence identity at the nucleotide level to a reference polynucleotide encoding an engineered nitroaldolase. In some embodiments, the reference polynucleotide sequence is selected from SEQ ID NOs: 3 - 1483.
[0195] In some embodiments, a polynucleotide capable of hybridizing under highly stringent conditions encodes an engineered nitroaldolase polypeptide having improved properties, comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 2. In some embodiments, the polynucleotide encodes a polypeptide described herein, but has at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or greater sequence identity at the nucleotide level to a reference polynucleotide encoding an engineered nitroaldolase. In some embodiments, the reference polynucleotide sequence is selected from SEQ ID NOs: 3 - 461.
[0196] In some embodiments, a polynucleotide that can hybridize under highly stringent conditions encodes an engineered nitroaldolase polypeptide having improved properties, comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 24. In some embodiments, the polynucleotide encodes a polypeptide described herein, but has at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or greater sequence identity at the nucleotide level to a reference polynucleotide encoding an engineered nitroaldolase. In some embodiments, the reference polynucleotide sequence is selected from SEQ ID NOs: 459 - 841.
[0197] In some embodiments, a polynucleotide that can hybridize under highly stringent conditions encodes an engineered nitroaldolase polypeptide having improved properties, comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 496. In some embodiments, the polynucleotide encodes a polypeptide described herein, but has at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or greater sequence identity at the nucleotide level to a reference polynucleotide encoding an engineered nitroaldolase. In some embodiments, the reference polynucleotide sequence is selected from SEQ ID NOs: 843 - 979.
[0198] In some embodiments, a polynucleotide that can hybridize under highly stringent conditions encodes an engineered nitroaldolase polypeptide having improved properties, comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 860. In some embodiments, the polynucleotide encodes a polypeptide as described herein, but has at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or greater sequence identity at the nucleotide level to a reference polynucleotide encoding an engineered nitroaldolase. In some embodiments, the reference polynucleotide sequence is selected from SEQ ID NOs: 981 - 1211.
[0199] In some embodiments, a polynucleotide that can hybridize under highly stringent conditions encodes an engineered nitroaldolase polypeptide having improved properties, comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 1120. In some embodiments, the polynucleotide encodes a polypeptide as described herein, but has at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or greater sequence identity at the nucleotide level to a reference polynucleotide encoding an engineered nitroaldolase. In some embodiments, the reference polynucleotide sequence is selected from SEQ ID NOs: 1213 - 1365.
[0200] In some embodiments, a polynucleotide capable of hybridizing under highly stringent conditions encodes an engineered nitroaldolase polypeptide having improved properties, comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity to SEQ ID NO: 1348. In some embodiments, the polynucleotide encodes a polypeptide described herein, but has at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% or greater sequence identity at the nucleotide level to a reference polynucleotide encoding an engineered nitroaldolase. In some embodiments, the reference polynucleotide sequence is selected from SEQ ID NOs: 1367 - 1483.
[0201] In some embodiments, an isolated polynucleotide encoding any of the engineered nitroaldolase polypeptides provided herein is engineered in various ways to facilitate expression of the enzyme polypeptide. In some embodiments, the polynucleotide encoding the polypeptide is provided as an expression vector in which one or more control sequences are present to regulate the expression of the polynucleotide and / or polypeptide. Manipulation of the isolated polynucleotide prior to its insertion into the vector may be desirable or necessary depending on the expression vector. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art.
[0202] In some embodiments, the control array, among a number of arrays, in particular includes a promoter, a leader sequence, a polyadenylation sequence, a propeptide sequence, a signal peptide sequence, and a transcription terminator. As is known in the art, a suitable promoter can be selected based on the host cell used. For bacterial host cells, suitable promoters for directing the transcription of the nucleic acid constructs of the present application include, but are not limited to, the E. coli lac operon, the Streptomyces coelicolor agarase gene (dagA), the Bacillus subtilis levansucrase gene (sacB), the Bacillus licheniformis alpha-amylase gene (amyL), the Bacillus stearothermophilus maltogenic amylase gene (amyM), the Bacillus amyloliquefaciens alpha-amylase gene (amyQ), the Bacillus licheniformis penicillinase gene (penP), the Bacillus subtilis xylA and xylB genes, and promoters obtained from prokaryotic beta-lactamase genes (see, for example, Villa-Kamaroff et al., Proc. Natl Acad. Sci. USA 75: 3727-3731
[1978] ), as well as the tac promoter (see, for example, DeBoer et al., Proc. Natl Acad. Sci. USA 80: 21-25
[1983] ).Exemplary promoters for filamentous fungal host cells include, but are not limited to, promoters obtained from the genes of Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral alpha-amylase, Aspergillus niger acid stable alpha-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (see, e.g., WO96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoter from the Aspergillus niger neutral alpha-amylase gene and the promoter from the Aspergillus oryzae triose phosphate isomerase gene), and variants, truncated forms, and hybrid promoters thereof. Exemplary yeast cell promoters can be derived from genes encoding Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase. Other useful promoters for yeast host cells are known in the art (see, e.g., Romanos et al., Yeast 8:423-488
[1992] ).
[0203] In some embodiments, the control sequence is a suitable transcription terminator sequence, a sequence recognized by the host cell to terminate transcription. The terminator sequence is operably linked to the 3' end of the nucleic acid sequence encoding the polypeptide. Any terminator that is functional in the selected host cell is used in the present invention. For example, exemplary transcription terminators for filamentous fungal host cells can be obtained from the genes of Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger alpha-glucosidase, and Fusarium oxysporum trypsin-like protease. Exemplary terminators for yeast host cells can be obtained from the genes of Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are known in the art (see, e.g., Romanos et al., supra).
[0204] In some embodiments, the control sequences are suitable leader sequences, which are untranslated regions of mRNA that are important for translation by the host cell. The leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in the selected host cell can be used. Exemplary leaders for filamentous fungal host cells are obtained from the genes of Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase. Leaders suitable for yeast host cells include, but are not limited to, those obtained from the genes of Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae alpha factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP). The control sequence can also be a polyadenylation sequence, a sequence operably linked to the 3' end of the nucleic acid sequence, which, when transcribed, is recognized by the host cell as a signal for adding polyadenosine residues to the transcribed mRNA. Any polyadenylation sequence that is functional in the selected host cell can be used in the present invention. Exemplary polyadenylation sequences for filamentous fungal host cells include, but are not limited to, those derived from the genes of Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger alpha-glucosidase. Useful polyadenylation sequences for yeast host cells are also known in the art (see, for example, Guo and Sherman, Mol. Cell. Bio., 15:5983-5990
[1995] ).
[0205] In some embodiments, the control sequence is a coding region that encodes a signal peptide, an amino acid sequence linked to the amino terminus of the polypeptide, and directs the encoded polypeptide into the secretory pathway of the cell. The 5' end of the coding sequence of the nucleic acid sequence may essentially contain a signal peptide coding region that is naturally linked in-frame with a segment of the coding region encoding the secreted polypeptide. Alternatively, the 5' end of the coding sequence may contain a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the secretory pathway of the selected host cell is used for the expression of the engineered nitroaldolase polypeptide provided herein. Effective signal peptide coding regions for bacterial host cells include, but are not limited to, signal peptide coding regions from the genes of Bacillus NCIB 11837 maltogenic amylase, Bacillus stearothermophilus alpha-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus stearothermophilus neutral protease (nprT, nprS, nprM), and Bacillus subtilis prsA. Further signal peptides are known in the art (see, for example, Simonen and Palva, Microbiol. Rev., 57:109-137
[1993] ). Effective signal peptide coding regions for filamentous fungal host cells include, but are not limited to, signal peptide coding regions obtained from the genes of Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase.Useful signal peptides for yeast host cells include, but are not limited to, those from the Saccharomyces cerevisiae alpha factor and the Saccharomyces cerevisiae invertase genes.
[0206] In some embodiments, the control sequence is a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is sometimes referred to as a "proenzyme", "propolypeptide" or "zymogen". The propolypeptide can be converted to a mature active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide. Propeptide coding regions include, but are not limited to, the genes for Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Saccharomyces cerevisiae alpha factor, Rhizomucor miehei aspartic proteinase, and Myceliophthora thermophila lactase (see, for example, WO95 / 33836). When both a signal peptide region and a propeptide region are present at the amino terminus of the polypeptide, the propeptide region is located adjacent to the amino terminus of the polypeptide and the signal peptide region is located adjacent to the amino terminus of the propeptide region.
[0207] In some embodiments, regulatory sequences are also utilized. These sequences facilitate the regulation of polypeptide expression in relation to the growth of the host cell. Examples of regulatory systems are those that turn gene expression on or off in response to chemical or physical stimuli, including the presence of regulatory compounds. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory sequences include, but are not limited to, the ADH2 system or the GAL1 system. In filamentous fungi, suitable regulatory sequences include, but are not limited to, the TAKA alpha-amylase promoter, the Aspergillus niger glucoamylase promoter, and the Aspergillus oryzae glucoamylase promoter.
[0208] The present invention also includes recombinant expression vectors comprising a polynucleotide encoding an engineered nitroaldolase polypeptide and one or more expression regulatory regions such as a promoter, a terminator, an origin of replication, etc., depending on the type of host to be introduced. In some embodiments, the various nucleic acids and control sequences described above are combined together to generate a recombinant expression vector that includes one or more convenient restriction sites to allow insertion or substitution of a nucleic acid sequence encoding a variant nitroaldolase polypeptide at such sites. Alternatively, the polynucleotide sequence of the present invention is expressed by inserting the polynucleotide sequence or a nucleic acid construct comprising the polynucleotide sequence into an appropriate vector for expression. In constructing the expression vector, the coding sequence is positioned within the vector so as to be operably linked to appropriate control sequences for expression.
[0209] A recombinant expression vector can be any vector (e.g., a plasmid or virus) that can be conveniently subjected to recombinant DNA procedures and can result in the expression of a variant nitroaldolase polynucleotide sequence. The choice of vector typically depends on the compatibility of the vector with the host cell into which it is to be introduced. The vector can be a linear plasmid or a closed circular plasmid.
[0210] In some embodiments, the expression vector is a self-replicating vector (i.e., an entity that exists as an extrachromosomal entity and whose replication does not depend on chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome). The vector can contain any means for ensuring self-replication. In some alternative embodiments, the vector can be one that integrates into the genome when introduced into the host cell and is replicated along with the chromosome into which it has integrated. Further, a single vector or plasmid, or two or more vectors or plasmids that together contain all of the DNA to be introduced into the genome of the host cell or into a transposon, can be used.
[0211] In some embodiments, the expression vector preferably contains one or more selectable markers that enable easy selection of transformed cells. A "selectable marker" is a gene whose product confers biocide or virus resistance, resistance to heavy metals, prototrophy to auxotrophic strains, and the like. Examples of selectable markers for bacteria include, but are not limited to, the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers that confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, or tetracycline resistance. Suitable markers for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for use in filamentous fungal host cells include, but are not limited to, amdS (acetamidase), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase), sC (sulfate adenylyltransferase), and trpC (anthranilate synthase), and equivalents thereof. In another aspect, the invention provides a host cell comprising a polynucleotide encoding at least one engineered nitroaldolase polypeptide of the invention, wherein the polynucleotide is operably linked to one or more control sequences for expression of the engineered nitroaldolase enzyme(s) in the host cell.Host cells for use in expressing the polypeptide encoded by the expression vector of the present invention are well known in the art and include, but are not limited to, bacterial cells such as E. coli, Vibrio fluvialis, Streptomyces and Salmonella typhimurium cells; fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae and Pichia pastoris [ATCC accession number 201178]); insect cells such as Drosophila S2 and Spodoptera Sf9 cells; animal cells such as CHO, COS, BHK, 293 and Bowes melanoma cells; and plant cells. Exemplary host cells are Escherichia coli strains (e.g., W3110 (ΔfhuA) and BL21).
[0212] Accordingly, in another aspect, the present invention provides a method for producing an engineered nitroaldolase polypeptide, the method comprising culturing a host cell capable of expressing a polynucleotide encoding the engineered nitroaldolase under conditions suitable for the expression of the polypeptide. In some embodiments, the method further comprises isolating and / or purifying the nitroaldolase polypeptide as described herein.
[0213] Suitable culture media and growth conditions for the above host cells are well known in the art. The polynucleotide for the expression of the nitroaldolase polypeptide can be introduced into cells by various methods known in the art. The techniques include, among others, electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion.
[0214] An engineered nitroaldolase having the characteristics disclosed herein can be obtained by subjecting a polynucleotide encoding a naturally occurring or engineered nitroaldolase polypeptide to mutagenesis and / or directed evolution as known in the art and as described herein. Exemplary directed evolution techniques include mutagenesis and / or DNA shuffling (see, e.g., Stemmer, Proc. Natl. Acad. Sci. USA 91:10747-10751
[1994] ; WO95 / 22625, WO97 / 0078, WO97 / 35966, WO98 / 27230, WO00 / 42651, WO01 / 75767, and U.S. Patent No. 6,537,746). Other directed evolution procedures that can be used include, among others, the staggered extension process (StEP), in vitro recombination (see, e.g., Zhao et al., Nat. Biotechnol., 16:258-261
[1998] ), mutagenic PCR (see, e.g., Caldwell et al., PCR Methods Appl., 3:S136-S140
[1994] ), and cassette mutagenesis (see, e.g., Black et al., Proc. Natl. Acad. Sci. USA 93:3525-3529
[1996] ).
[0215] For example, mutagenesis and directed evolution methods can be easily applied to polynucleotides to generate variant libraries that can be expressed, screened, and assayed. Mutagenesis and directed evolution methods are well known in the art (e.g., U.S. Patent Nos. 5,605,793, 5,811,238, 5,830,721, 5,834,252, 5,837,458, 5,928,905, 6,096,548, 6,117,679, 6,132,970, 6,165,793, 6,180,406, 6,251,674, 6,265,201, 6,277,638, 6,287,861, 6,287,862, 6,291,242, 6,297,053, 6,303,344, 6,309,883, 6,319,713, 6,319,714, 6,323,030, 6,326,204, 6,335,160, 6,335,198, 6,344,356, 6,352,859, 6,355,484, 6,358,740, 6,358,742, 6,365,377, 6,365,408, 6,368,861, 6,372,497, 6,337,186, 6,376,246, 6,379,964, 6,387,702, 6,391,552, 6,391,640, 6,395,547, 6,406,855, 6,406,910, 6,413,745, 6,413,774, 6,420,175, 6,423,542, 6,426,224, 6,436,675, 6,444,468, 6,455,253, 6,479,652, 6,482,647, 6,483,011, 6,484,105, 6,489,146, 6,500,617, 6,500,639, 6,506,602, 6,506,603, 6,518,065, 6,519,065, 6,521,453, 6,528,311, 6,537,746, 6,573,098, 6,576,467, 6,579,678, 6,586,182, 6,602,986, 6,605,430, 6,613,514, 6,653,072, 6,686,515, 6,703,240, 6,716,631, 6,825,001, 6,902,922, 6,917,882, 6,946,296, 6,961,664, 6,995,017, 7,024,312, 7,058,515, 7,105,297, 7,148,054, 7,220,566, 7,288,375, 7,384,387, 7,421,347, 7,430,477, 7,462,469, 7,534,564, 7,620,500, 7,620,502, 7,629,170, 7,702,464, 7,747,391, 7,747,393, 7,751,986, 7,776,598, 7,783,428, 7,795,030, 7,853,410, 7,868,138, 7,783,428, 7,873,477, 7,873,499, 7,904,249, 7,957,912, 7,981,614, 8,014,961, 8,029,988, 8,048,674, 8,058,001, 8,076,138, 8,108,150, 8,170,806, 8,224,580, 8,377,681, 8,383,346, 8,457,903, 8,504,498, 8,589,085, 8,762,066, 8,768,871, 9,593,326, and all related U.S. patents, as well as PCT and non-U.S. counterpart patents; Ling et al., Anal. Biochem., 254(2):157-78
[1997] ; Dale et al., Meth. Mol. Biol., 57:369-74
[1996] ; Smith, Ann. Rev. Genet., 19:423-462
[1985] ; Botstein et al., Science, 229:1193-1201
[1985] ; Carter, Biochem. J., 237:1-7
[1986] ; Kramer et al., Cell, 38:879-887
[1984] ; Wells et al., Gene, 34:315-323
[1985] ; Minshull et al., Curr. Op. Chem. Biol., 3:284-290
[1999] ; Christians et al., Nat. Biotechnol., 17:259-264
[1999] ; Crameri et al., Nature, 391:288-291
[1998] ; Crameri, et al., Nat. Biotechnol., 15:436-438
[1997] ; Zhang et al., Proc. Nat. Acad. Sci. U.S.A.,See 94:4504-4509
[1997] ; Crameri et al., Nat. Biotechnol., 14:315-319
[1996] ; Stemmer, Nature, 370:389-391
[1994] ; Stemmer, Proc. Nat. Acad. Sci. USA, 91:10747-10751
[1994] ; WO95 / 22625; WO97 / 0078; WO97 / 35966; WO98 / 27230; WO00 / 42651; WO01 / 75767; and WO2009 / 152336, all of which are hereby incorporated by reference).
[0216] In some embodiments, enzyme clones obtained after mutagenesis treatment are screened by subjecting the enzyme to a predetermined temperature (or other assay conditions, e.g., testing the activity of the enzyme against a wide range of substrates) and measuring the amount of enzyme activity remaining after heat treatment or other assay conditions. Clones containing polynucleotides encoding the nitroaldolase polypeptide are then sequenced to identify nucleotide sequence changes (if any) and used to express the enzyme in host cells. Measurement of enzyme activity from an expression library can be performed using any suitable method known in the art (e.g., standard biochemical techniques such as HPLC analysis).
[0217] In some embodiments, clones obtained after mutagenesis treatment can be screened for engineered nitroaldolases having one or more desired improved enzyme properties (e.g., improved regioselectivity). Measurement of enzyme activity from an expression library can be performed using standard biochemical techniques, e.g., GC analysis, HPLC analysis, and / or derivatization of the product (before or after separation) using, for example, dansyl chloride or OPA (see, e.g., Yaegaki et al., J Chromatogr. 356(1):163-70
[1986] ).
[0218] If the sequence of the engineered polypeptide is known, the polynucleotide encoding the enzyme can be prepared by standard solid-phase methods according to known synthetic methods. In some embodiments, fragments of up to about 100 bases can be synthesized individually and then joined (e.g., by enzymatic or chemical ligation methods, or polymerase-mediated methods) to form any desired contiguous sequence. For example, polynucleotides and oligonucleotides encoding portions of nitroaldolase can be prepared by chemical synthesis known in the art, typically as performed by automated synthesis methods (e.g., the classical phosphoramidite method of Beaucage et al., Tet. Lett. 22:1859-69
[1981] , or the method described by Matthes et al., EMBO J. 3:801-05
[1984] ). According to the phosphoramidite method, oligonucleotides are synthesized (e.g., on an automated DNA synthesizer), purified, annealed, ligated, and cloned into an appropriate vector. In addition, essentially any nucleic acid can be obtained from any of a variety of commercial sources. In some embodiments, additional variations can be created by synthesizing oligonucleotides containing deletions, insertions, and / or substitutions, and by combining oligonucleotides in various permutations to create engineered nitroaldolases with improved properties.
[0219] Thus, in some embodiments, a method of preparing an engineered nitroaldolase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide having an amino acid sequence having at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity to an amino acid sequence selected from the even-numbered sequences of SEQ ID NOs: 4-1484, the polypeptide having one or more residue differences compared to SEQ ID NOs: 2, 24, 496, 860, 1120, or 1348; and (b) expressing the nitroaldolase polypeptide encoded by the polynucleotide.
[0220] In some embodiments of the method, the polynucleotide encodes an engineered nitroaldolase having one or several (e.g., up to 3, 4, 5, or up to 10) amino acid residue deletions, insertions, and / or substitutions, as needed. In some embodiments, the amino acid sequence has 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-30, 1-35, 1-40, 1-45, or 1-50 amino acid residue deletions, insertions, and / or substitutions, as needed. In some embodiments, the amino acid sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 amino acid residue deletions, insertions, and / or substitutions, as needed. In some embodiments, the amino acid sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 amino acid residue deletions, insertions, and / or substitutions, as needed. In some embodiments, the substitutions may be conservative substitutions or non-conservative substitutions.
[0221] In some embodiments, any of the engineered nitroaldolase enzymes expressed in a host cell can be recovered from the cells and / or culture medium using any one or more of well-known techniques for protein purification, including but not limited to lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography. Solutions suitable for the lysis of bacteria such as E. coli and the efficient extraction of proteins from bacteria are commercially available (e.g., CelLytic B™, Sigma-Aldrich, St. Louis MO).
[0222] Chromatography techniques for the isolation of nitroaldolase polypeptides include, but are not limited to, reverse-phase chromatography, high-performance liquid chromatography, ion-exchange chromatography, gel electrophoresis, and affinity chromatography. The conditions for purifying a particular enzyme will in part depend on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and will be apparent to those skilled in the art.
[0223] In some embodiments, affinity techniques can be used to isolate the improved nitroaldolase enzyme. For affinity chromatography purification, any antibody that specifically binds to the nitroaldolase polypeptide can be used. For antibody production, various host animals including but not limited to rabbits, mice, rats, etc. can be immunized by injection with the nitroaldolase polypeptide or a fragment thereof. The nitroaldolase polypeptide or fragment can be conjugated to a suitable carrier such as BSA by side-chain functional groups or by linkers attached to the side-chain functional groups. In some embodiments, affinity purification can use a specific ligand to which the nitroaldolase binds, or a dye affinity column (see, e.g., EP0641862; Stellwagen, “Dye Affinity Chromatography,” In Current Protocols in Protein Science, Unit 9.2-9.2.16
[2001] ).
[0224] Method of using an engineered nitroaldolase enzyme In some embodiments, the nitroaldolase enzymes described herein are used in a process for the conversion of one or more suitable substrates to a product.
[0225] In another aspect, the engineered nitroaldolase polypeptides disclosed herein can be used in a process for the conversion of a substrate compound (1) or a structural analog thereof, and a substrate compound (2) or a structural analog thereof, to a product of compound (3) or a corresponding structural analog.
[0226] Structural analogs of compound (1) include other ketones having halogen modifications and / or ketones having various modifications to the alpha carbon. Structural analogs of compound (2) include other alkanes having a nitro group or other nitro-containing compounds having an available α-hydrogen. Structural analogs of compound (3) include various β-nitroalcohols.
[0227] In some embodiments, the present disclosure provides a process for preparing compound (3)
Chemical formula
Chemical formula
Chemical formula
[0228] In the embodiments provided herein and illustrated in the examples, suitable reaction conditions for various ranges that can be used in the process include, but are not limited to, substrate loading, cosubstrate loading, pH, temperature, buffer, solvent system, polypeptide loading, and reaction time. Further reaction conditions suitable for performing a biocatalytic conversion process from a substrate compound to a product compound using the engineered nitroaldolase described herein can be readily optimized by routine experimentation taking into account the guidance provided herein, such experimentation including, but not limited to, contacting the engineered nitroaldolase polypeptide with one or more substrate compounds under experimental reaction conditions for concentration, pH, temperature, and solvent conditions, and detecting the product compound.
[0229] The substrate compound in the reaction mixture can be varied, for example, considering the desired amount of the product compound, the effect of each substrate concentration on enzyme activity, the stability of the enzyme under the reaction conditions, and the percent conversion of the substrate to the product. In some embodiments, suitable reaction conditions include a substrate compound loading for each of one or more substrates of at least about 0.5 to about 60 g / L, 1 to about 40 g / L, 20 to about 35 g / L, about 10 to about 25 g / L, or 25 to about 35 g / L. In some embodiments, suitable reaction conditions include a substrate compound loading for each of one or more substrates of at least about 0.5 g / L, at least about 1 g / L, at least about 5 g / L, at least about 10 g / L, at least about 15 g / L, at least about 20 g / L, at least about 30 g / L, at least about 35 g / L, at least about 40 g / L, or higher.
[0230] When performing the nitroaldolase-mediated process described herein, the engineered polypeptide can be added to the reaction mixture in the form of a purified enzyme, a partially purified enzyme, whole cells transformed with the gene encoding the enzyme, a cell extract and / or lysate of such cells, and / or an enzyme immobilized on a solid support. Whole cells or cell extracts, lysates thereof, and isolated enzymes transformed with the gene encoding the engineered nitroaldolase enzyme can be utilized in a variety of different forms, including solids (e.g., lyophilized, spray-dried, etc.) or semi-solids (e.g., crude paste). The cell extract or cell lysate can be partially purified by precipitation (e.g., ammonium sulfate, polyethyleneimine, heat treatment, etc.), followed by a desalting procedure (e.g., ultrafiltration, dialysis, etc.) prior to lyophilization. Any of the enzyme preparations (including whole cell preparations) can be stabilized by cross-linking using a known cross-linking agent such as glutaraldehyde, or immobilization on a solid phase (e.g., Eupergit C, etc.).
[0231] The gene encoding the engineered nitroaldolase polypeptide can be transformed into host cells separately or together into the same host cells. For example, in some embodiments, one set of host cells can be transformed with the gene encoding one engineered nitroaldolase polypeptide and another set with the gene encoding another engineered nitroaldolase polypeptide. Both sets of transformed cells can be utilized together in the reaction mixture in the form of whole cells or in the form of lysates or extracts obtained therefrom. In other embodiments, host cells can be transformed with genes encoding multiple engineered nitroaldolase polypeptides. In some embodiments, the engineered polypeptide can be expressed in the form of a secreted polypeptide, and the culture medium containing the secreted polypeptide can be used in the nitroaldolase reaction.
[0232] In some embodiments, the improved activity and / or stereoselectivity of the engineered nitroaldolase polypeptide disclosed herein results in a process that can achieve a higher percentage of conversion at a lower concentration of the engineered polypeptide. In some embodiments of the process, suitable reaction conditions include an amount of engineered polypeptide of about 1% (w / w), 2% (w / w), 5% (w / w), 10% (w / w), 20% (w / w), 30% (w / w), 40% (w / w), 50% (w / w), 75% (w / w), 100% (w / w), 150% (w / w), 175% (w / w), 200% (w / w) or more of the substrate compound load.
[0233] In some embodiments, the engineered polypeptide is present at about 0.01 g / L to about 130 g / L; about 0.05 g / L to about 100 g / L; about 0.1 g / L to about 40 g / L; about 1 g / L to about 40 g / L; about 2 g / L to about 40 g / L; about 5 g / L to about 40 g / L; about 5 g / L to about 30 g / L; about 0.1 g / L to about 10 g / L; about 0.5 g / L to about 10 g / L; about 1 g / L to about 10 g / L; about 0.1 g / L to about 5 g / L; about 0.5 g / L to about 5 g / L; or about 0.1 g / L to about 2 g / L. In some embodiments, the nitroaldolase polypeptide is present at about 0.01 g / L, 0.05 g / L, 0.1 g / L, 0.2 g / L, 0.5 g / L, 1, 2 g / L, 5 g / L, 10 g / L, 15 g / L, 20 g / L, 25 g / L, 30 g / L, 35 g / L, 40 g / L, 50 g / L, 80 g / L, or 100 g / L.
[0234] During the reaction process, the pH of the reaction mixture can change. The pH of the reaction mixture can be maintained at a desired pH or within a desired pH range. This can be done by adding an acid or a base before and / or during the reaction. Alternatively, the pH can be controlled by using a buffer. Thus, in some embodiments, the reaction conditions include a buffer. Suitable buffers for maintaining a desired pH range are known in the art and include, by way of example and not limitation, sodium citrate, borate, phosphate, 2-(N-morpholino)ethanesulfonic acid (MES), 3-(N-morpholino)propanesulfonic acid (MOPS), acetate, triethanolamine, and 2-amino-2-hydroxymethyl-propane-1,3-diol (Tris). In some embodiments, the reaction conditions include water as a suitable solvent in the absence of a buffer.
[0235] In process embodiments, the reaction conditions include a suitable pH. The desired pH or desired pH range can be maintained by using an acid or a base, a suitable buffer, or a combination of buffering and acid or base addition. The pH of the reaction mixture can be controlled before and / or during the reaction. In some embodiments, suitable reaction conditions include a solution pH of from about 4 to about 10, a pH of from about 4 to about 6, a pH of from about 5 to about 6, a pH of from about 6 to about 9, a pH of from about 5 to about 7. In some embodiments, the reaction conditions include a solution pH of about 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10.
[0236] In the process embodiments described herein, for example, a temperature suitable for the reaction conditions is used, taking into account the increase in reaction rate at higher temperatures and the activity of the enzyme during the reaction period. Thus, in some embodiments, suitable reaction conditions include temperatures of about 10°C to about 60°C, about 10°C to about 55°C, about 15°C to about 60°C, about 20°C to about 60°C, about 20°C to about 55°C, about 25°C to about 55°C, or about 30°C to about 50°C. In some embodiments, suitable reaction conditions include temperatures of about 10°C, 15°C, 20°C, 22°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, or 60°C. In some embodiments, the temperature during the enzyme reaction can be maintained at a specific temperature throughout the reaction process. In some embodiments, the temperature during the enzyme reaction can be adjusted according to the temperature profile during the reaction process.
[0237] In some embodiments, the process of the present invention is carried out in a solvent. Suitable solvents include water, aqueous buffer solutions, organic solvents, polymeric solvents, and / or cosolvent systems, and cosolvent systems generally include aqueous solvents, organic solvents, and / or polymeric solvents. The aqueous solvent (water or aqueous cosolvent system) may or may not be pH buffered. In some embodiments, the process using the engineered nitroaldolase polypeptide is carried out in an aqueous cosolvent system that includes an organic solvent (e.g., ethanol, isopropanol (IPA), dimethyl sulfoxide (DMSO), dimethylformamide (DMF), ethyl acetate, butyl acetate, 1-octanol, heptane, octane, methyl tert-butyl ether (MTBE), toluene, etc.), an ionic or polar solvent (e.g., 1-ethyl-4-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium hexafluorophosphate, glycerol, polyethylene glycol, etc.). In some embodiments, the cosolvent can be a polar solvent, such as a polyol, dimethyl sulfoxide (DMSO), or a lower alcohol. The non-aqueous cosolvent component of the aqueous cosolvent system may be miscible with the aqueous component and thus may form a single liquid phase, or may be partially miscible or immiscible with the aqueous component and thus may form two liquid phases. Exemplary aqueous cosolvent systems can include water and one or more cosolvents selected from organic solvents, polar solvents, and polyol solvents. Generally, the cosolvent component of the aqueous cosolvent system is selected such that it does not detrimentally inactivate the nitroaldolase enzyme under the reaction conditions. Suitable cosolvent systems can be readily identified by measuring the enzyme activity of the specified engineered nitroaldolase enzyme using a defined substrate of interest in a candidate solvent system, utilizing an enzyme activity assay such as those described herein.
[0238] In some embodiments of the process, suitable reaction conditions include an aqueous co-solvent that contains from about 1% to about 50% (v / v), from about 1% to about 40% (v / v), from about 2% to about 40% (v / v), from about 5% to about 30% (v / v), from about 10% to about 30% (v / v), or from about 10% to about 20% (v / v) DMSO. In some embodiments of the process, suitable reaction conditions may include an aqueous co-solvent that contains about 1% (v / v), about 5% (v / v), about 10% (v / v), about 15% (v / v), about 20% (v / v), about 25% (v / v), about 30% (v / v), about 35% (v / v), about 40% (v / v), about 45% (v / v), or about 50% (v / v) ethanol.
[0239] In some embodiments, the reaction conditions include a surfactant to stabilize or enhance the reaction. The surfactant can include nonionic, cationic, anionic, and / or amphiphilic surfactants. Exemplary surfactants include, by way of example and not limitation, nonylphenoxypolyethoxyethanol (NP40), TRITON® X-100 polyethylene glycol tert-octylphenyl ether, polyoxyethylene-stearylamine, cetyltrimethylammonium bromide, sodium oleylamide sulfate, polyoxyethylene-sorbitan monostearate, hexadecyldimethylamine, and the like. Any surfactant that can stabilize or enhance the reaction can be utilized. The concentration of the surfactant utilized in the reaction can generally be from 0.1 to 50 mg / ml, particularly from 1 to 20 mg / ml.
[0240] In some embodiments, the reaction conditions may include an antifoaming agent that is useful for reducing or preventing the formation of bubbles in the reaction solution, for example when the reaction solution is mixed or sparged. Antifoaming agents include non-polar oils (e.g., mineral, silicone, etc.), polar oils (e.g., fatty acids, alkyl amines, alkyl amides, alkyl sulfates, etc.), and hydrophobic materials (e.g., treated silica, polypropylene, etc.), some of which also function as surfactants. Exemplary antifoaming agents include Y-30 (registered trademark) (Dow Corning), polyglycol copolymers, oxy / ethoxylated alcohols, and polydimethylsiloxane. In some embodiments, the antifoaming agent may be present at about 0.001% (v / v) to about 5% (v / v), about 0.01% (v / v) to about 5% (v / v), about 0.1% (v / v) to about 5% (v / v), or about 0.1% (v / v) to about 2% (v / v). In some embodiments, the antifoaming agent may be present at about 0.001% (v / v), about 0.01% (v / v), about 0.1% (v / v), about 0.5% (v / v), about 1% (v / v), about 2% (v / v), about 3% (v / v), about 4% (v / v), or about 5% (v / v), or higher, as desired to promote the reaction.
[0241] The amount of reactants used in the nitroaldolase reaction will generally vary depending on the amount of desired product and, concomitantly, on the amount of nitroaldolase substrate utilized. The methods for varying these amounts to suit the desired productivity level and production scale will be readily understood by those skilled in the art.
[0242] In some embodiments, the order of addition of the reactants is not critical. The reactants may be added together, simultaneously, to a solvent (e.g., a single-phase solvent, a two-phase aqueous cosolvent system, etc.), or alternatively, some of the reactants may be added separately and some together at different times. For example, the cofactor, cosubstrate, and substrate may first be added to the solvent.
[0243] Solid reactants (e.g., enzymes, salts, etc.) can be supplied to the reaction in various different forms, such as powders (e.g., lyophilized, spray-dried, etc.), solutions, oil emulsions, suspensions, etc. The reactants can be easily lyophilized or spray-dried using methods and equipment known to those skilled in the art. For example, a protein solution can be frozen in small portions at -80°C and then added to a pre-cooled lyophilization chamber, followed by applying a vacuum.
[0244] When using an aqueous co-solvent system, for improving the mixing efficiency, nitroaldolase and the co-substrate can first be added to the aqueous phase and mixed. The organic layer can then be added and mixed following the nitroaldolase substrate, or the substrate can be dissolved in the organic layer and mixed. Alternatively, the nitroaldolase substrate can be premixed in the organic phase prior to addition to the aqueous phase.
[0245] The process of the present invention is generally allowed to proceed until further conversion of the substrate to the product no longer significantly changes with the reaction time (e.g., less than 10% of the substrate is converted, or less than 5% of the substrate is converted). In some embodiments, the reaction is allowed to proceed until the substrate is completely or nearly completely converted to the product. Known methods for detecting the substrate and / or product, with or without derivatization, can be used to monitor the conversion of the substrate to the product. Suitable analytical methods include gas chromatography, HPLC, MS, etc.
[0246] In some embodiments of the process, suitable reaction conditions include a substrate loading for each of one or more substrates of at least about 5 g / L, 10 g / L, 20 g / L, 30 g / L, 40 g / L, or more, and this method results in at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more conversion of the substrate compound to the product compound in about 24 hours or less, about 12 hours or less, about 6 hours or less, or about 4 hours or less.
[0247] When the engineered nitroaldolase polypeptide of the present invention is used in a process under suitable reaction conditions, it provides an excess of the desired product with an enantiomeric excess of at least 30%, 40%, 50%, 60%, or higher over the undesired product(s).
[0248] In further embodiments of a process for converting one or more substrate compounds to a product compound using the engineered nitroaldolase polypeptide, suitable reaction conditions include an initial substrate loading for each of the one or more substrates into the reaction solution, such that the polypeptide contacts it. Subsequently, additional substrate compounds are then added to this reaction solution over time as a continuous or batch addition, at a rate of at least about 1 g / L / h, at least about 2 g / L / h, at least about 4 g / L / h, at least about 6 g / L / h, or higher, for each of the one or more substrate compounds. Thus, according to these suitable reaction conditions, the polypeptide is added to a solution having an initial substrate loading of at least about 1 g / L, 5 g / L, or 10 g / L for each of the one or more substrate compounds. Subsequently, following this addition of the polypeptide, additional substrate is continuously added to the solution at a rate of about 2 g / L / h, about 4 g / L / h, or 6 g / L / h for each of the one or more substrate compounds until a final substrate loading of at least about 30 g / L or much higher is reached. Thus, in some embodiments of the process, suitable reaction conditions include the addition of the polypeptide to a solution having an initial substrate loading of at least about 1 g / L, 5 g / L, or 10 g / L, followed by the addition of additional substrate to the solution at a rate of about 2 g / L / h, 4 g / L / h, or 6 g / L / h until a final substrate loading of at least about 30 g / L or higher is reached. These substrate replenishment reaction conditions make it possible to achieve higher substrate loadings while maintaining a high conversion rate of the substrate to the product, such as at least about 5%, 25%, 50%, 75%, 90%, or higher conversion of the substrate for any or both of the one or more substrate compounds.
[0249] Any of the processes disclosed herein that use the polypeptide engineered for the preparation of compound (3) can be carried out under a range of suitable reaction conditions including, but not limited to, the range of ketone substrates, nitroalkane substrates, temperature, pH, solvent system, substrate loading, polypeptide loading, and reaction time. In one example, in some embodiments, compound (3) can be prepared, and suitable reaction conditions in this case are: (a) a ketone substrate loading of about 0.01 M to 1 M of the substrate compound; (b) an amino acid substrate loading of about 0.01 M to 1 M of the substrate compound; (c) that of the engineered polypeptide of 0.5 g / L to 100 g / L; (d) a sodium citrate buffer of 0.01 M to 1 M; (e) a pH of 5 to 8; and (f) a temperature of about 15 °C to 60 °C. In some embodiments, suitable reaction conditions include: (a) about 0.5 M of compound (1) substrate compound); (b) about 0.3 M of compound (2) substrate compound); (c) about 50 g / L of each engineered polypeptide; (d) a sodium citrate buffer of 0.1 M; (e) a static pH of 6, and (g) about 22 °C.
[0250] In some embodiments, additional reaction components or additional techniques are carried out to supplement the reaction conditions. These can include taking measures to stabilize the enzyme or prevent enzyme inactivation, measures to reduce product inhibition, and measures to shift the reaction equilibrium towards the formation of the desired product.
[0251] In further embodiments, any of the above processes for the conversion of one or more substrate compounds to a product compound may further include one or more steps selected from extraction, isolation, purification, and crystallization of the product compound. Methods, techniques, and protocols for extracting, isolating, purifying, and / or crystallizing the product from the biocatalytic reaction mixture produced by the disclosed processes above are known to those skilled in the art and / or can be obtained by routine experimentation. In addition, exemplary methods are provided in the examples below.
[0252] The various features and embodiments of the present invention are illustrated in the following representative examples, which are intended to be helpful in the description and not intended to be limiting.
Example
[0253] Experiment The following examples, including the experiments conducted and the results, are provided solely for the purpose of being helpful in the description and should not be construed as limiting the present invention.
[0254] In the following examples, the following abbreviations are applied: ppm (parts per million); M (molar concentration); mM (millimolar concentration), uM and μΜ (micromolar concentration); nM (nanomolar concentration); mol (mole); gm and g (gram); mg (milligram); ug and μg (microgram); L and l (liter); ml and mL (milliliter); cm (centimeter); mm (millimeter); um and μm (micrometer); sec. (second); min (minute); h and hr (hour); U (unit); MW (molecular weight); rpm (revolutions per minute); psi and PSI (pounds per square inch); °C (degrees Celsius); RT and rt (room temperature); CAM and cam (chloramphenicol); DMSO (dimethyl sulfoxide); PMBS (polymyxin B sulfate); IPTG (isopropyl β-D-1-thiogalactopyranoside); LB (Luria Bertani broth); TB (Terrific broth; 12 g / L of bacto-tryptone, 24 g / L of yeast extract, 4 mL / L of glycerol, 65 mM potassium phosphate, pH 7.0, 1 mM MgSO4); PLP (pyridoxal 5'-phosphate), TEoA (triethanolamine buffer), HEPES (HEPES zwitterionic buffer; 4-(2-hydroxyethyl)-piperazineethanesulfonic acid); SFP (shaken flask powder); CDS (coding sequence); DNA (deoxyribonucleic acid); RNA (ribonucleic acid); E. coli W3110 (a commonly used laboratory E. coli strain available from the Coli Genetic Stock Center [CGSC], New Haven, CT); HTP (high throughput); HPLC (high performance liquid chromatography); GC (gas chromatography); FIOPC (improvement factor relative to positive control); Microfluidics (Microfluidics, Corp., Westwood, MA); Sigma-Aldrich (Sigma-Aldrich, St. Louis, MO; Difco (Difco Laboratories, BD Diagnostic Systems, Detroit, MI); Agilent (Agilent Technologies, Inc., Santa Clara, CA); Corning (Corning, Inc., Palo Alto, CA); Dow Corning (Dow Corning, Corp., Midland, MI); and Gene Oracle (Gene Oracle, Inc., Mountain View, CA). (Example 1) Generation of Engineered Polypeptides in pCK110900
[0255] A polynucleotide (SEQ ID NO: 1) encoding a polypeptide (SEQ ID NO: 2) having nitroaldolase activity was cloned into the pCK110900 vector system (see, e.g., U.S. Patent No. 9,714,437, which is hereby incorporated by reference in its entirety), and then expressed in E. coli W3110fhuA under the control of the lac promoter. This polynucleotide, and related polypeptides, were derived from oxynitrilase found in Baliospermum montanum.
[0256] In a 96-well format, single colonies were picked and grown in 190 μL of LB medium containing 1% glucose and 30 μg / mL CAM at 30 °C, 200 rpm, and 85% humidity. After overnight growth, 20 μL of the grown culture was transferred to a deep well plate containing 380 μL of TB medium with 30 μg / mL CAM. The culture was grown at 250 rpm, 85% humidity, and 30 °C for approximately 2.5 hours. When the optical density (OD600) of the culture reached 0.4 - 0.6, expression of the nitroaldolase gene was induced by the addition of IPTG to a final concentration of 1 mM. After induction, growth was continued at 85% humidity, 30 °C, and 250 rpm for 18 - 20 hours. Cells were harvested by centrifugation at 4,000 rpm for 10 minutes at 4 °C, and then the supernatant was discarded. The cell pellet was stored at -80 °C until ready for use.
[0257] Before performing the assay, the cell pellet was thawed and resuspended in 200 μL of lysis buffer containing 1 g / L lysozyme, 0.5 g / L PMBS and 0.025 μL / mL of commercially available DNAse (New England BioLabs, M0303L) in 0.1 M sodium citrate buffer at pH 6. The plate was shaken at medium speed on a microtiter plate shaker for 2 hours at room temperature. The plate was then centrifuged at 4,000 rpm for 15 minutes at 4 °C and the clear supernatant was used in the HTP assay reaction described in the following examples.
[0258] Using the shake flask procedure, engineered nitroaldolase or polypeptide shake flask powder useful for secondary screening assays and / or use in the biocatalytic processes described herein can be created. Preparation of the enzyme shake flask powder (SFP) provides a more purified preparation of the engineered enzyme (e.g., up to 30% of total protein) compared to cell lysates used in HTP assays and also allows for the use of more concentrated enzyme solutions. To initiate the culture, a single colony of E. coli containing the plasmid encoding the engineered polypeptide of interest was inoculated into 5 mL of LB cell culture medium with 30 μg / mL of CAM and 1% glucose. The culture was grown overnight (at least 16 hours) at 37 °C in an incubator with shaking at 250 rpm. The grown culture was then added to 250 mL of TB medium with 30 μg / mL of CAM in a 1 L shake flask. The 250 mL culture was grown at 30 °C and 250 rpm for 3.5 hours until the OD600 reached 0.6 - 0.8. Expression of the nitroaldolase gene was induced by addition of IPTG to a final concentration of 1 mM and growth was continued for an additional 18 - 20 hours. The cells were harvested by transferring the culture to a pre-weighed centrifuge bottle and then centrifuged at 4,000 rpm for 20 minutes at 4 °C. The supernatant was discarded and the remaining cell pellet was weighed. In some embodiments, the cell pellet was stored at -80 °C until ready for use. For lysis, the cell pellet was resuspended in 6 mL / wet cell weight g of 25 mM triethanolamine-HCl buffer at pH 7 or 25 mM sodium citrate buffer at pH 6 and lysed using a 110L MICROFLUIDIZER® processor system (Microfluidics). Cell debris was removed by centrifugation at 10,000 rpm for 60 minutes at 4 °C. The clear lysate was collected, frozen at -80 °C and then lyophilized using standard methods known in the art. Lyophilization of the frozen clear lysate provides a dry shake flask powder containing the crude engineered polypeptide. (Example 2) Evolution and Screening of Engineered Polypeptides Derived from SEQ ID NO: 2 for Improved Production of Compound (3)
[0259] Using the engineered polynucleotide (SEQ ID NO: 1) encoding a polypeptide having nitroaldolase activity of SEQ ID NO: 2, the engineered polypeptides in Table 2-1 were generated. These polypeptides showed improved nitroaldolase activity under desired conditions compared to the starting polypeptide, for example, improved formation of the β-nitro alcohol, compound (3), from the substrates trifluoroacetone and nitro methane, compounds (1) and (2) respectively. Similarly, the engineered polypeptides in Table 2-2 showed improved nitroaldolase selectivity for the desired (S)-alcohol. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the "skeleton" amino acid sequence of SEQ ID NO: 2 as described below.
[0260] Directed evolution was initiated using the polynucleotide shown in SEQ ID NO: 1. A library of engineered polypeptides was generated using various well-known techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods to measure the ability of the polypeptides to produce compound (3).
[0261] The enzyme assay was performed in 96-well deep-well (total volume 1.1 mL) plates with a total reaction volume of 100 μL per well. The reaction contained 67 v / v% undiluted nitroaldolase lysate prepared as described in Example 1, dissolved in 100 mM sodium citrate at pH 6, 40 mM trifluoroacetone (1), and 500 mM nitro methane (2). The reaction plates were heat-sealed and shaken at 250 rpm at 22 °C for approximately 22 hours.
[0262] After overnight incubation (about 22 hours), 300 μL / well of methyl tert-butyl ether (MTBE) was added to the reaction plate and mixed well. The plate was sealed and centrifuged at 4,000 rpm for 10 minutes. An aliquot of the upper organic phase was removed and added to a shallow 96-well plate for analysis by gas chromatography (GC) as described in Example 8.
[0263] The hit variants were grown in 250 mL shake flasks as described in Example 1 to produce lyophilized enzyme powder. The activity of the enzyme powder was evaluated by measuring the enzyme dosage response with 0 - 25 g / L of nitroaldolase shake flask powder under the assay conditions described in this example. These hit variants are indicated by * in Tables 2-1 and 2-2.
Table 2-1-1
Table 2-1-2
Table 2-1-3
Table 2-1-4
Table 2-2-1
Table 2-2-2
Table 2-2-3
Table 2-2-4
Table 2-2-5
Table 2-2-6
[0264] Using an engineered polynucleotide (SEQ ID NO: 23) encoding a polypeptide having nitroaldolase activity of SEQ ID NO: 24, the engineered polypeptides in Table 3-1 were generated. These polypeptides showed improved nitroaldolase activity under desired conditions compared to the starting polypeptide, for example, improved formation of compound (3), which is a β-nitroalcohol, from the substrate trifluoroacetone and compounds (1) and (2) which are nitromethane. Similarly, the engineered polypeptides in Table 3-2 showed improved nitroaldolase selectivity for the desired (S)-alcohol. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the "skeleton" amino acid sequence of SEQ ID NO: 24 as described below.
[0265] Directed evolution was initiated using the polynucleotide shown in SEQ ID NO: 23. A library of engineered polypeptides was generated using various well-known techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods to measure the ability of the polypeptides to produce compound (3).
[0266] Enzyme assays and analyses were performed using the same methods as described in Example 2. High-throughput cell pellets were lysed in 400 μL / well of lysis buffer and the reaction was carried out using 93 v / v% clarified lysate, 0.3 M trifluoroacetone (1), and 0.5 M nitromethane (2).
[0267] The hit variants were grown in 250 mL shake flasks and enzyme powder was produced as described in Example 1. Using the same assay as described in Example 2, the activity of the enzyme powder was evaluated at 0 - 100 g / L nitroaldolase SFP, 0.3 M trifluoroacetone (1), 0.5 M nitromethane (2), 0.1 M sodium citrate, pH 6, 22 °C for 22 hours. These hit variants are indicated by * in Tables 3-1 and 3-2.
Table 3-1-1
Table 3-1-2
Table 3-1-3
Table 3-2-1
Table 3-2-2
Table 3-2-3
Table 3-2-4
Table 3-2-5
[0268] Using the engineered polynucleotide (SEQ ID NO: 495) encoding a polypeptide having nitroaldolase activity of SEQ ID NO: 496, the engineered polypeptides in Table 4-1 were generated. These polypeptides showed improved nitroaldolase activity under desired conditions compared to the starting polypeptide, for example, improved formation of the β-nitro alcohol compound (3) from the substrate trifluoroacetone and the nitro compounds (1) and (2) which are nitromethane. Similarly, the engineered polypeptides in Table 4-2 showed improved nitroaldolase selectivity for the desired (S)-alcohol. The engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the "backbone" amino acid sequence of SEQ ID NO: 496 as described below.
[0269] Directed evolution was initiated using the polynucleotide shown in SEQ ID NO: 495. A library of engineered polypeptides was generated using various well-known techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences), and screened using HTP assays and analytical methods to measure the ability of the polypeptides to produce compound (3).
[0270] Enzyme assays and analysis were performed using the same method as described in Example 3.
[0271] The hit variants were grown in 250 mL shake flasks and enzyme powder was generated as described in Example 1. The activity of the enzyme powder was evaluated as described in Example 3. These hit variants are indicated by * in Tables 4-1 and 4-2.
Table 4-1-1
Table 4-1-2
Table 4-2-1
Table 4-2-2
Table 4-2-3
[0272] Using an engineered polynucleotide (SEQ ID NO: 859) encoding a polypeptide having nitroaldolase activity of SEQ ID NO: 860, the engineered polypeptides of Table 5-1 were generated. These polypeptides showed improved nitroaldolase activity under desired conditions compared to the starting polypeptide, for example, improved formation of compound (3), which is a β-nitro alcohol, from each of compound (1) and (2), which are substrate trifluoroacetone and nitromethane, respectively. Similarly, the engineered polypeptides of Table 5-2 showed improved nitroaldolase selectivity for the desired (S)-alcohol. Engineered polypeptides having amino acid sequences of even-numbered sequence identifiers were generated from the "skeleton" amino acid sequence of SEQ ID NO: 860 as described below.
[0273] Directed evolution was initiated using the polynucleotide shown in SEQ ID NO: 859. A library of engineered polypeptides was generated using various well-known techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods to measure the ability of the polypeptides to produce compound (3).
[0274] Enzyme assays and analyses were performed using the same method as described in Example 3.
[0275] Hit variants were grown in 250 mL shake flasks and enzyme powder was generated as described in Example 1. The activity of the enzyme powder was evaluated as described in Example 3. These hit variants are indicated with * in Tables 5-1 and 5-2.
Table 5-1-1
Table 5-1-2
Table 5-1-3
Table 5-1-4
Table 5-2-1
Table 5-2-2
Table 5-2-3
[0276] Engineered polypeptides of Table 6-1 were generated using an engineered polynucleotide (SEQ ID NO: 1119) encoding a polypeptide having nitroaldolase activity of SEQ ID NO: 1120. These polypeptides showed improved nitroaldolase activity under desired conditions compared to the starting polypeptide, e.g., improved formation of the β-nitro alcohol compound (3) from the substrates trifluoroacetone and the nitro compounds (1) and (2) which are nitromethane. Similarly, the engineered polypeptides of Table 6-2 showed improved nitroaldolase selectivity for the desired (S)-alcohol. Engineered polypeptides having amino acid sequences of even-numbered sequence identifiers were generated from the "scaffold" amino acid sequence of SEQ ID NO: 1120 as described below.
[0277] Directed evolution was initiated using the polynucleotide shown in SEQ ID NO: 1119. A library of engineered polypeptides was generated using various well-known techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences), and screened using HTP assays and analytical methods to measure the ability of the polypeptides to produce compound (3).
[0278] Enzyme assays and analyses were performed using the method as described in Example 3, except that the concentration of nitromethane (2) was increased to 0.75 M and the pH was decreased to pH 5.5.
[0279] Hit variants were grown in 250 mL shake flasks and enzyme powder was produced as described in Example 1. The activity of the enzyme powder was evaluated as described in Example 3, except that the concentration of trifluoroacetone (1) was increased to 0.525 M, the concentration of nitromethane was increased to 0.75 M, and the pH was decreased to pH 5.5. Hit variants are indicated by * in Tables 6-1 and 6-2. [Table 6-1-1] [Table 6-1-2] [Table 6-2] (Example 7) Evolution and screening of engineered polypeptides derived from SEQ ID NO: 1348 for improved production of compound (3)
[0280] Using the engineered polynucleotide (SEQ ID NO: 1347) encoding a polypeptide having nitroaldolase activity of SEQ ID NO: 1348, the engineered polypeptides in Table 6-1 were generated. These polypeptides showed improved nitroaldolase activity under desired conditions, for example, an improvement in the formation of the compound (3), which is a β-nitroalcohol, from each of the compounds (1) and (2), which are the substrates trifluoroacetone and nitromethane, compared to the starting polypeptide. Similarly, the engineered polypeptides in Table 6-2 showed improved nitroaldolase selectivity for the desired (S)-alcohol. The engineered polypeptides having the amino acid sequences of even-numbered sequence identifiers were generated from the "skeleton" amino acid sequence of SEQ ID NO: 1348 as described below.
[0281] Directed evolution was initiated using the polynucleotide shown in SEQ ID NO: 1347. A library of engineered polypeptides was generated using various well-known techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences), and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce the compound (3).
[0282] Enzyme assays and analyses were performed using the method as described in Example 6, except that the concentration of trifluoroacetone (1) was increased to 0.525 M and the reaction time was decreased to 6 hours.
[0283] Hit variants were grown in 250 mL shake flasks and enzyme powder was generated as described in Example 1. The activity of the enzyme powder was evaluated as described in Example 6, except that the reaction time was decreased to 6 hours. The hit variants are indicated by * in Tables 7-1 and 7-2.
Table 7-1
Table 7-2-1
Table 7-2-2
[0284] (Example 8) Analytical Detection of the Conversion of Compounds (1) and (2) to Compound (3) The data described in Examples 2 to 4 were collected using the analytical methods provided in Tables 8-1 and 8-2. The methods provided herein are used for the analysis of variants produced using the present invention. However, since there are other suitable methods known in the art that can be applied to the analysis of the variants provided herein and / or variants produced using the methods provided herein, the present invention is not intended to be limited to the methods described herein. [Table 8-1] [Table 8-2]
[0285] The data described in Example 5 were collected using the analytical methods provided in Tables 8-3 and 8-4. The data described in Examples 6 and 7 were collected using the analytical methods provided in Table 8-4. The methods provided herein are used for the analysis of variants produced using the present invention. However, since there are other suitable methods known in the art that can be applied to the analysis of the variants provided herein and / or variants produced using the methods provided herein, the present invention is not intended to be limited to the methods described herein. [Table 8-3] [Table 8-4]
[0286] All publications, patents, patent applications, and other documents cited in this application are hereby incorporated by reference in their entirety for all purposes to the same extent as if each individual publication, patent, patent application, or other document were individually indicated to be incorporated by reference for all purposes.
[0287] Although various specific embodiments have been illustrated and described, it will be understood that various changes can be made without departing from the spirit and scope of the invention.
Claims
**Claim 1** An engineered nitroaldolase polypeptide comprising a polypeptide sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 2, 24, 496, 860, 1120 and / or 1348, or a functional fragment thereof, wherein the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 2, 24, 496, 860, 1120 and / or 1348. **Claim 2** The polypeptide sequence contains at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 2, and the engineered nitroaldolase polypeptide is 11, 7 / 9, 7 / 9 / 33 / 68 / 72 / 158, 7 / 9 / 33 / 72 / 175 / 197 / 230, 7 / 9 / 33 / 72 / 230, 7 / 9 / 68 / 72 / 197, 7 / 9 / 158 / 175, 7 / 9 / 175, 7 / 9 / 197, 7 / 33, 7 / 68 / 72 / 175, 7 / 72, 7 / 72 / 97 / 175 / 230, 9, 9 / 11 / 33 / 72 / 197 / 230, 9 / 33 / 68 / 158 / 197, 9 / 33 / 72 / 230, 9 / 33 / 97 / 230, 9 / 68 / 72, 9 / 97 / 175 / 202 / 218, 9 / 175, 9 / 202, 11, 13, 14 / 33 / 68 / 72 / 97, 18, 18 / 33 / 68 / 97 / 158 / 175 / 239, 18 / 33 / 72 / 97 / 158 / 175, 18 / 33 / 72 / 97 / 158 / 239, 18 / 33 / 72 / 97 / 175 / 202, 18 / 33 / 72 / 175 / 202 / 237 / 239, 18 / 33 / 72 / 197 / 237, 18 / 68 / 72, 18 / 68 / 72 / 158 / 175 / 237 / 239, 18 / 68 / 72 / 175 / 202 / 239, 18 / 68 / 72 / 239, 18 / 68 / 175 / 223 / 237 / 239, 18 / 68 / 239, 18 / 72 / 97 / 175 / 237 / 239, 18 / 72 / 175 / 197, 18 / 72 / 237, 18 / 72 / 239, 18 / 97 / 175 / 237 / 239, 18 / 97 / 223 / 237 / 239, 18 / 175, 18 / 197 / 239, 18 / 239, 20 / 124, 33 / 68 / 72, 33 / 68 / 72 / 175 / 230, 33 / 68 / 223 / 239, 33 / 72, 33 / 72 / 97 / 197, 33 / 72 / 175, 33 / 175, 39, 44 / 97, 52, 68 / 72, 68 / 72 / 97, 68 / 72 / 97 / 197, 68 / 72 / 197, 68 / 72 / 197 / 239, 68 / 97 / 175, 68 / 97 / 175 / 197 / 223, 72 / 97 / 158 / 223 / 237, 72 / 175 / 197, 97 / 158 / 175 / 197, 97 / 175, 103, 121, 122, 124, 125, 128, 133, 148, 175, 197,and at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 197 / 202, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 2, the engineered nitroaldolase polypeptide according to claim 1. **Claim 3** The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 2, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 7 / 11 / 33 / 72 / 97, 7 / 11 / 33 / 97, 7 / 11 / 72, 7 / 11 / 158, 7 / 11 / 237, 7 / 33 / 72 / 97 / 158 / 175 / 197 / 237 / 239, 9 / 11 / 33, 9 / 11 / 33 / 72, 9 / 11 / 72, 9 / 33 / 68 / 97 / 237 / 239, 9 / 72 / 97 / 158 / 237 / 239, 9 / 218 / 239, 10, 11, 12, 13, 14, 18, 33 / 68 / 72 / 97 / 237 / 239, 33 / 72 / 158 / 239, 39, 40, 41, 48, 49, 51, 54, 54 / 126, 57, 68 / 175 / 239, 82, 84, 85, 86, 103, 104, 105, 107, 115, 117, 119, 119 / 120, 120, 122, 123, 124, 127, 130, 131, 133, 134, 137 / 145, 144, 145, 146, 147, 148, 149, 150, 152, 157, 158, 159, 169, 173, 174, 175, 180, 181, 183, 203, 204, 208, 209, 210, 215, 218, 236, and 237, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:
2. The engineered nitroaldolase polypeptide according to claim 1.
4. The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 24, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 9 / 39 / 122 / 157, 18, 18 / 33 / 39 / 230, 18 / 33 / 175, 18 / 39 / 44 / 118, 18 / 175, 20, 33 / 39, 33 / 39 / 44 / 118 / 121, 33 / 39 / 44 / 121, 33 / 39 / 44 / 152, 33 / 39 / 44 / 175, 33 / 39 / 44 / 197, 33 / 39 / 44 / 210 / 230, 33 / 39 / 118 / 121 / 175, 33 / 39 / 121 / 175, 33 / 39 / 197 / 210, 33 / 39 / 210, 33 / 39 / 210 / 230, 33 / 118 / 121, 33 / 118 / 121 / 175, 39, 39 / 44 / 103, 39 / 44 / 118, 39 / 44 / 152 / 230, 39 / 44 / 175 / 197 / 230, 39 / 103, 39 / 103 / 125, 39 / 103 / 125 / 127 / 146, 39 / 103 / 125 / 127 / 146 / 150, 39 / 103 / 125 / 146, 39 / 103 / 127 / 150, 39 / 103 / 150, 39 / 118 / 121, 39 / 197, 39 / 197 / 210, 43, 44 / 103, 81, 103, 103 / 125 / 127, 103 / 125 / 146, 103 / 127, 103 / 150, 118 / 121, 118 / 121 / 175, 118 / 121 / 197, 121, 152, 152 / 197, 154, 159, 163, 175, 210, 232, and 238, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:
24. The engineered nitroaldolase polypeptide according to claim 1.
5. The polypeptide sequence contains at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 24, and the engineered nitroaldolase polypeptide is 7, 9 / 13, 9 / 13 / 39, 9 / 13 / 39 / 68 / 72 / 239, 9 / 13 / 39 / 68 / 239, 9 / 13 / 39 / 72 / 122, 9 / 13 / 39 / 97 / 239, 9 / 13 / 68 / 72 / 97 / 239, 9 / 13 / 68 / 97 / 134 / 239, 9 / 13 / 68 / 97 / 239, 9 / 13 / 72, 9 / 13 / 134, 11, 12 / 42 / 125 / 127, 12 / 103, 12 / 103 / 150, 13, 13 / 39, 13 / 39 / 68, 13 / 39 / 68 / 72 / 97, 13 / 39 / 72 / 97 / 122 / 134 / 223, 13 / 39 / 97 / 239, 13 / 39 / 239, 13 / 68, 13 / 68 / 72 / 97 / 122 / 239, 16, 18 / 33 / 39 / 44 / 118 / 121 / 175 / 197 / 210, 18 / 33 / 39 / 44 / 121, 18 / 33 / 39 / 112 / 118 / 121 / 197, 18 / 33 / 39 / 121, 18 / 33 / 39 / 121 / 175, 18 / 33 / 39 / 121 / 175 / 197, 18 / 33 / 39 / 121 / 175 / 210, 18 / 33 / 39 / 121 / 210, 18 / 33 / 118 / 121, 18 / 33 / 121, 18 / 33 / 121 / 197, 18 / 33 / 121 / 210 / 230, 18 / 39 / 44 / 118 / 121 / 230, 18 / 39 / 44 / 121 / 175 / 210, 18 / 39 / 118 / 121, 18 / 39 / 118 / 121 / 210, 18 / 39 / 121 / 175, 18 / 39 / 121 / 175 / 210, 18 / 39 / 121 / 197 / 210, 18 / 118 / 121, 18 / 118 / 121 / 175, 18 / 118 / 121 / 210, 18 / 118 / 121 / 230, 18 / 121, 18 / 121 / 175, 18 / 121 / 210, 18 / 121 / 230, 20, 21, 33 / 39 / 44 / 118 / 121, 33 / 39 / 44 / 118 / 121 / 210 / 230, 33 / 39 / 44 / 121, 33 / 39 / 44 / 121 / 152, 33 / 39 / 44 / 121 / 175 / 210, 33 / 39 / 44 / 121 / 230, 33 / 39 / 73 / 121 / 197,at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 33 / 39 / 118 / 121 / 175, 33 / 39 / 118 / 121 / 175 / 230, 33 / 39 / 121, 33 / 39 / 121 / 175, 33 / 39 / 121 / 210 / 230, 33 / 118 / 121, 33 / 118 / 121 / 175, 33 / 118 / 121 / 197, 33 / 121, 33 / 121 / 230, 37 / 91, 39 / 44 / 118 / 121 / 175 / 210, 39 / 68 / 72 / 239, 39 / 118 / 121, 39 / 118 / 121 / 210, 39 / 121, 39 / 121 / 210 / 230, 43, 44, 76, 79, 101, 118 / 121, 118 / 121 / 175, 118 / 121 / 197, 121, 121 / 175 / 230, 121 / 210, 148, 153, 159, 161, 172, 200, 207, 233, 245, and 248, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 24, the engineered nitroaldolase polypeptide of claim 1.,
6. The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 496, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 43 / 103, 43 / 103 / 172 / 238, 43 / 103 / 238 / 241, 43 / 238, 81 / 232, 103, 103 / 159 / 238, 103 / 194 / 238, 103 / 238, 122, 232, 238, and 238 / 241, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 496, the engineered nitroaldolase polypeptide according to claim 1.
7. The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 496, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 12, 43, 43 / 47 / 103, 43 / 47 / 103 / 232, 43 / 103, 43 / 103 / 121 / 172 / 241, 43 / 103 / 172, 43 / 103 / 172 / 238, 43 / 103 / 238, 43 / 103 / 238 / 241, 43 / 121, 43 / 194 / 238, 43 / 238, 44 / 103 / 121 / 238, 44 / 103 / 238, 44 / 103 / 238 / 241, 44 / 121 / 238 / 241, 47 / 103, 63, 103, 103 / 121, 103 / 121 / 172 / 238, 103 / 159 / 238, 103 / 172 / 238, 103 / 194 / 238, 103 / 238, 122, 131, 146, 178, 180, and 221, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 496, the engineered nitroaldolase polypeptide according to claim 1.
8. The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 860, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 12, 12 / 43 / 63 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 63 / 103 / 146 / 241, 12 / 43 / 81 / 103 / 241, 12 / 43 / 81 / 146 / 232 / 241, 12 / 43 / 81 / 180 / 241, 12 / 43 / 81 / 232 / 241, 12 / 43 / 81 / 241, 12 / 43 / 103, 12 / 43 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 103 / 146 / 180 / 241, 12 / 43 / 103 / 180 / 241, 12 / 43 / 103 / 241, 12 / 43 / 146 / 180 / 232 / 241, 12 / 43 / 180 / 241, 12 / 63 / 81 / 103 / 241, 12 / 81, 12 / 81 / 180 / 241, 12 / 103 / 146 / 241, 12 / 103 / 180 / 241, 12 / 146, 12 / 146 / 180 / 241, 12 / 180 / 232 / 241, 12 / 241, 18, 43 / 81 / 232 / 241, 43 / 103 / 146 / 180, 43 / 103 / 180 / 241, 43 / 103 / 241, 43 / 180 / 232, 43 / 232 / 241, 63 / 103 / 180 / 232, 78, 81, 81 / 103 / 146, 81 / 146 / 180 / 241, 81 / 146 / 241, 81 / 180 / 241, 84, 103, 103 / 146, 103 / 146 / 180, 103 / 146 / 180 / 241, 103 / 180, 103 / 180 / 241, 103 / 261, 144, 145, 146 / 180 / 232, 146 / 180 / 241, 146 / 241, 150, 166, 172, 180 / 232, 232 / 241, and 241, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:
860. The engineered nitroaldolase polypeptide according to claim 1. **Claim 9** The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 860, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 9, 12, 12 / 43 / 63 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 63 / 103 / 146 / 241, 12 / 43 / 103 / 146 / 180 / 232 / 241, 12 / 43 / 103 / 146 / 180 / 241, 12 / 43 / 103 / 146 / 232 / 241, 12 / 43 / 103 / 146 / 241, 12 / 43 / 103 / 180 / 241, 12 / 43 / 146 / 180 / 232 / 241, 12 / 43 / 146 / 241, 12 / 43 / 180 / 232, 12 / 43 / 180 / 241, 12 / 63 / 103 / 146 / 232 / 241, 12 / 103 / 146 / 241, 12 / 103 / 180 / 241, 12 / 146, 12 / 146 / 180 / 241, 12 / 146 / 232 / 241, 12 / 146 / 241, 12 / 180, 12 / 180 / 232 / 241, 12 / 180 / 241, 12 / 241, 18, 43 / 63 / 103 / 180 / 232 / 241, 43 / 103 / 146 / 180, 43 / 103 / 180, 43 / 103 / 180 / 232 / 241, 43 / 103 / 180 / 241, 43 / 180 / 232, 63 / 103 / 180 / 232, 103, 103 / 146 / 180, 103 / 146 / 180 / 241, 103 / 180, 103 / 180 / 241, 103 / 241, 118, 119, 122, 144, 145, 146 / 180 / 241, 154, 160, 172, 180 / 232, 180 / 232 / 241, 180 / 241, 208, 211, and 212, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 860, the engineered nitroaldolase polypeptide according to claim 1.
10. The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 1120, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 3, 7, 7 / 158, 7 / 158 / 175, 7 / 158 / 197 / 202, 27 / 43 / 115 / 172, 27 / 43 / 172, 29, 32, 43 / 110 / 115 / 172, 44, 46, 48 / 172, 55, 59, 67, 71, 72, 73, 94 / 103 / 118 / 146, 103, 103 / 144 / 146 / 172 / 180, 103 / 146 / 154 / 172 / 180 / 232, 103 / 154, 118, 118 / 144 / 215, 118 / 172 / 180, 137, 142, 146, 146 / 154, 154, 158, 172 / 190, 184, 198, 205, 216, 223, 226, 229, 247, 261, and 263, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 1120, the engineered nitroaldolase polypeptide according to claim 1.
11. The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 1120, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 73, 80, 103 / 146 / 209, 103 / 172 / 209, 118 / 119 / 172 / 209, 135, 144 / 209 / 232, 172 / 180 / 209, 181, and 223, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 1120, the engineered nitroaldolase polypeptide according to claim 1.
12. The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 1348, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 2, 7 / 46, 39, 46, 46 / 55, 46 / 55 / 72, 46 / 55 / 92, 46 / 55 / 118, 46 / 55 / 226 / 257 / 263, 46 / 55 / 263, 46 / 70, 46 / 70 / 92 / 118 / 261 / 263, 46 / 70 / 175 / 263, 46 / 72, 46 / 118, 55 / 257 / 261, 55 / 261 / 263, 55 / 263, 105, 118, 118 / 226 / 261 / 263, and 263, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 1348, the engineered nitroaldolase polypeptide according to claim 1.
13. The polypeptide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 1348, and the engineered nitroaldolase polypeptide comprises at least one substitution or set of substitutions in the polypeptide sequence at one or more positions selected from 2, 7 / 46, 7 / 46 / 175, 22 / 46 / 55 / 118, 27 / 46 / 72, 27 / 46 / 146 / 158 / 202, 27 / 46 / 175, 27 / 118, 29 / 70 / 118, 39, 46, 46 / 55, 46 / 55 / 59 / 261 / 263, 46 / 55 / 72, 46 / 55 / 92, 46 / 55 / 118, 46 / 55 / 118 / 263, 46 / 55 / 146, 46 / 55 / 146 / 202, 46 / 55 / 226 / 257 / 263, 46 / 55 / 263, 46 / 70, 46 / 70 / 92 / 118 / 261 / 263, 46 / 70 / 175 / 263, 46 / 72, 46 / 92, 46 / 92 / 175, 46 / 118, 46 / 146, 46 / 146 / 257 / 263, 55 / 70 / 118, 55 / 257 / 261, 55 / 257 / 263, 55 / 261 / 263, 55 / 263, 70, 92 / 118, 105, 118, 118 / 226 / 261 / 263, 146 / 257 / 263, 146 / 261 / 263, 192 / 257 / 261, 202, 257 / 261 / 263, 261 / 263, and 263, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 1348, the engineered nitroaldolase polypeptide according to claim 1.
14. The engineered nitroaldolase polypeptide according to any one of claims 1 to 13, comprising an amino acid sequence having at least 80% sequence identity to any even-numbered sequence represented by SEQ ID NOs: 4 to 1484.
15. The engineered nitroaldolase polypeptide according to any one of claims 1 to 14, comprising a polypeptide sequence represented by an even-numbered sequence among SEQ ID NOs: 4 to 1484.
16. The engineered nitroaldolase polypeptide according to any one of claims 1 to 15, comprising a polypeptide sequence exhibiting at least one improved property as compared to the engineered nitroaldolase polypeptide of SEQ ID NO:
2. Claim 17 The engineered nitroaldolase polypeptide according to claim 16, wherein the improved property comprises improved production of compound (3). 【Chemical Formula 6】 Claim 18 The engineered nitroaldolase polypeptide according to claim 16, wherein the improved property comprises improved stereoselectivity for at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or higher (S)-product. Claim 19 The engineered nitroaldolase polypeptide according to claim 16, wherein the improved property comprises improved activity against the substrates of compound (1) which is trifluoroacetone and compound (2) which is nitromethane. Claim 20 The engineered nitroaldolase polypeptide according to any one of claims 1 to 19, which is purified. Claim 21 A composition comprising at least one engineered nitroaldolase polypeptide according to any one of claims 1 to 20. Claim 22 An engineered polynucleotide encoding at least one engineered nitroaldolase polypeptide according to any one of claims 1 to 20. Claim 23 An engineered polynucleotide sequence encoding at least one engineered polypeptide, wherein the sequence of the polypeptide comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 2, 24, 496, 860, 1120 and / or 1348, and the sequence of the polypeptide comprises at least one substitution at one or more positions. Claim 24 The engineered polynucleotide sequence according to claim 22 or 23, wherein the polynucleotide sequence comprises at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 1, 23, 495, 860, 1119 and / or 1347. Claim 25 The engineered polynucleotide sequence according to any one of claims 22 to 24, wherein the polynucleotide sequence comprises SEQ ID NO: 3, 23, 495, 859, 1119 or 1347. Claim 26 An engineered polynucleotide comprising an odd-numbered array represented by SEQ ID NOs: 3 to 1483.
27. A vector comprising the engineered polynucleotide according to any one of Claims 22 to 26.
28. The vector according to Claim 27, further comprising at least one control sequence.
29. A host cell comprising the vector according to Claim 27 and / or 28.
30. A host cell that produces at least one engineered nitroaldolase polypeptide according to any one of Claims 1 to 20.
31. A method for producing an engineered nitroaldolase polypeptide in a host cell, the method comprising culturing the host cell according to Claim 29 and / or 30 in a culture medium under suitable conditions such that at least one engineered nitroaldolase polypeptide is produced.
32. The method according to Claim 31, further comprising the step of recovering the engineered nitroaldolase polypeptide.
33. The method according to Claim 31 and / or 32, further comprising the step of purifying the at least one engineered nitroaldolase polypeptide.