Biocatalysts and methods for hydroxylation of chemical compounds

By improving the biocatalyst through engineered proline hydroxylase, the problem of preparing hydroxylated proline has been solved, and the efficient and large-scale preparation of high-purity trans-3-hydroxyproline has been achieved, overcoming the limitations of existing technologies.

CN115175997BActive Publication Date: 2025-12-09CODEXIS INC
View PDF 147 Cites 0 Cited by

Patent Information

Application Number
CN202080082591.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-26
Filing Date
2020-11-19
Publication Date
2025-12-09
Estimated Expiration
2040-11-19

AI Technical Summary

Technical Problem

Existing technologies for preparing hydroxylated proline suffer from limitations in the availability of raw materials from natural sources, difficulties in purifying background contaminants, complex and difficult-to-scale chemical synthesis methods, limited whole-cell reaction conditions, and the use of non-optimized growth media, resulting in difficulties in product separation and low isomer purity.

Method used

By engineering proline hydroxylase, its activity, substrate tolerance, stereoselectivity, and thermal stability were improved. L-proline was hydroxylated to trans-3-hydroxyproline in the presence of oxygen and iron using α-ketoglutarate as a co-substrate, and the preparation was carried out using an improved biocatalyst.

Benefits of technology

This method enables efficient and large-scale preparation of high-purity trans-3-hydroxyproline, solving the preparation problems existing in the prior art and improving the isomer purity and reaction efficiency of the product.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003664691860000121
    Figure BDA0003664691860000121
  • Figure BDA0003664691860000131
    Figure BDA0003664691860000131
  • Figure BDA0003664691860000151
    Figure BDA0003664691860000151
Patent Text Reader

Abstract

The present invention provides engineered proline hydroxylase polypeptides, polynucleotides encoding the engineered proline hydroxylases, host cells capable of expressing the engineered proline hydroxylases, and methods of using the engineered proline hydroxylases to make compounds useful for producing active pharmaceutical agents.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 940,647, filed November 26, 2019, which is incorporated by reference in its entirety for all purposes. TECHNICAL FIELD

[0002] The present invention relates to biocatalysts for the hydroxylation of chemical compounds.

[0003] Reference to Sequence Listing, Tables, or Computer Program

[0004] The official copy of the sequence listing, as submitted via EFS-Web, is filed concurrently with the specification and appended hereto pursuant to 37 C.F.R. § 1.52(e) and is hereby incorporated by reference in its entirety. Pursuant to 37 C.F.R. § 1.52(e)(l)(ii), the sequence listing is submitted as an ASCII formatted text file, created on November 17, 2020, named “CX2-193WO1_ST25.txt” and having a size of 1.39 megabytes. The sequence listing submitted via EFS-Web is part of the specification and is herein incorporated by reference in its entirety.

[0005] BACKGROUND

[0006] Because of the constrained conformation of proline, proline derivatives with functional groups on the carbocyclic ring are useful building blocks for the synthesis of pharmaceutical compounds. One such derivative, hydroxylated proline, is a starting material for the synthesis of a variety of therapeutic compounds including carbapenem antibiotics (see, e.g., Altamura et al., J. Med., Chem. 38(21): 4244-56

[1995] ), angiotensin converting enzyme inhibitors, protease inhibitors (see, e.g., Chen et al., J. Org. Chem., 67(8): 2730-3

[2002] ; Chen et al., 2006, J Med Chem. 49(3): 995-1005), nucleic acid analogs (see, e.g., Efimov et al., Nucleic Acids Res., 34(8): 2247-2257

[2006] ), isoprenyl transferase inhibitors (O’Connell et al., Chem. Pharm. Bull., 48(5): 740-742

[2000] ), and drug library construction (Vergnon et al., J. Comb. Chem., 6(1): 91-8

[2004] ; and Remuzon, Tetrahedron 52: 13803-13835

[1996] ).

[0007] Hydroxyproline can be obtained from natural sources, such as plant material and collagen hydrolysates. Hydroxyproline can also be chemically synthesized, such as from the starting materials allyl bromide and diethyl acetylamino malonic acid (Kyun Lee et al., Bull. Chem. Soc. Japan, 46:2924

[1973] ), D-glutamic acid (Eguchi et al., Bull. Chem. Soc. Japan, 47: 1704-08

[1974] ), glyoxal and oxalacetic acid (Ramaswamy et al., J. Org. Chem., 42(21):3440-3443

[1977] ), and a-alanine (Sinha et al., Proc. ECSOC-4, The Fourth International Electronic Conference on Synthetic Organic Chemistry, ISBN 3-906980-05-7

[2000] ).

[0008] Isolation from natural sources is limited by the availability of raw materials, requires purification from large amounts of background contaminants, and lacks certain desired diastereomers. Chemical synthesis methods can require complex steps, are difficult to scale up to an industrial scale level, and require additional purification steps due to the formation of multiple hydroxylated products.

[0009] Another method for preparing hydroxylated proline uses proline hydroxylases, which are 2-oxoglutarate-dependent dioxygenases that utilize 2-oxoglutarate (a-ketoglutarate) and O2 as co-substrates and ferrous ion as a cofactor (see, e.g., Klein et al., Adv. Synth. Catal., 353: 1375-1383

[2011] ; U.S. Patent No. 5,364,775; and Shibasaki et al., Appl. Environ. Microbiol., 65(9):4028-4031

[1999] ). Unlike prolyl hydroxylases, which specifically recognize peptidyl prolines in procollagen and related peptides, proline hydroxylases are capable of converting free proline to hydroxyproline. Several microbial enzymes that produce cis-3-, cis-4- or trans-4-hydroxyproline are known in the art (see, e.g., U.S. Patent Nos. 5,962,292, 5,963,254, and 5,854,040; WO2009139365; and EP2290065) and an enzyme that produces trans-3-hydroxyproline has been identified in extracts of fungi. Many proline hydroxylases are found in bacteria and fungi, where the proline hydroxylases are associated with the biosynthesis of peptide antibiotics.

[0010] A naturally occurring proline hydroxylase selective for trans-3-hydroxyproline is not known in the art. The fungal proline hydroxylase GloF from Glarea lozoyensis produces trans-3-hydroxyproline as the minor isomer and trans-4-hydroxyproline as the major isomer (Petersen et al., Appl. Microbiol. Biotechnol. 2003, 62, 263; Houwaart et al., ChemBioChem 2014, 15, 2365). Another fungal proline hydroxylase, HtyE from Emericella rugulosa NRRL 11440, sharing about 64% sequence identity with GloF, was reported as part of the echinocandin B biosynthetic gene cluster (Cacho et al., J. Am. Chem. Soc. 2012, 134, 16781). HtyE was also found to produce trans-3-hydroxyproline as the minor isomer and trans-4-hydroxyproline as the major isomer. Recently, a gene cluster comprising three hydroxylase genes was identified in the fungus sp. 11243 (Matsui et al., J. Biosci. Bioeng. 2017, Feb; 123(2): 147-153), and one of the genes was subsequently identified as having homology to HtyE.

[0011] While recombinant whole cells expressing cloned proline hydroxylases are more amenable to large scale industrial processes, the use of whole cells limits reaction conditions such as changes in high substrate concentration; constrains the types of substrates that can be used to those that are permeable to the cell; and results in undesirable byproducts that must be separated from the final product. In addition, in vivo systems can require non-optimal or cost effective defined growth media because the use of rich growth media prepared from protein hydrolysates contains free proline which can be a competitive inhibitor when targeting substrates other than proline. Alternative methods for synthesizing hydroxylated forms of proline and proline analogs and other chemical compounds that can be readily scaled up and result in substantially pure isomer products are necessary. SUMMARY

[0013] The present invention provides engineered proline hydroxylase biocatalysts, polynucleotides encoding the biocatalysts, methods of making the same, and methods of using the engineered biocatalysts to make hydroxylated compounds. The proline hydroxylases of the present invention have been engineered to have one or more improved properties relative to the naturally occurring proline hydroxylase from Aspergillus sp. No. 11243 (SEQ ID NO: 2, to which an N-terminal his-tag has been added). The improved biocatalyst properties of the engineered proline hydroxylases include, among others, activity, substrate tolerance, stereoselectivity, regioselectivity, and thermal stability. The engineered proline hydroxylases have also been found to hydroxylate a variety of substrate compounds, including hydroxylating L-proline to trans-3-hydroxyproline using alpha-ketoglutarate as a co-substrate. In some embodiments, the method is performed in the presence of oxygen (i.e., air) and iron (i.e., Fe(II)).

[0014] The engineered enzymes having one or more improved properties have one or more residue differences compared to the naturally occurring proline hydroxylase, wherein the residue differences occur at residue positions that affect one or more of the above-mentioned enzyme properties.

[0015] Accordingly, in one aspect, the present invention provides an engineered polypeptide having proline hydroxylase activity, wherein the polypeptide comprises an amino acid sequence that is at least about 80% identical to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630. In some embodiments, the present invention provides an engineered polypeptide having proline hydroxylase activity, wherein the polypeptide comprises the amino acid sequence set forth in an even-numbered sequence in the range of SEQ ID NO: 6-658. The following detailed description provides guidance on the selection of residue differences that can be used to make engineered proline hydroxylases having the desired improved biocatalyst properties.

[0016] The present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 4. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 4 and one or more residue differences as compared to SEQ ID NO: 4 at a residue position selected from: 21, 28, 58 / 247, 65, 80, 85, 95, 98, 117, 120, 159, 185, 194, 199, 200, 233, 237, 243, 250, 268, 281, 282, 287, 289, 307, 324, 326, 327, 330, 338, 343, 346, and 348. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 4 and one or more residue differences as compared to SEQ ID NO: 4 at a residue position selected from: 21, 28, 45, 65, 95, 112, 117, 139, 177, 185, 199, 233, 243, 250, 281, 282, 287, 289, 307, 324, 326, 327, 335, 338, 343, and 346. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 4 and one or more residue differences as compared to SEQ ID NO: 4 at a residue position selected from: 48 / 66 / 189 / 194, 48 / 66 / 194, and 66 / 82 / 85 / 135 / 189 / 194 / 267.In some embodiments, the present application provides an engineered polypeptide having proline hydroxylase activity, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 4 and one or more residue differences as compared to SEQ ID NO: 4 at a residue position selected from 20 / 56 / 76 / 168 / 169 / 296, 20 / 56 / 232 / 294, 20 / 119 / 294 / 296, 56 / 76 / 119 / 124 / 147 / 232, 56 / 76 / 294, 76 / 168 / 232 / 294, 76 / 294 / 296, 76 / 296, 147, and 232. In some embodiments, the engineered polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to at least one of the even-numbered sequences in SEQ ID NOs: 4-658.

[0017] The present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 116. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 116 and one or more residue differences as compared to SEQ ID NO: 116 at a residue position selected from 123, 189, 195, 233, and 296. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 116 and one or more residue differences as compared to SEQ ID NO: 116 at a residue position selected from 20 / 21 / 56, 20 / 21 / 56 / 76 / 95 / 232 / 294 / 307 / 335, 20 / 21 / 56 / 76 / 147 / 225 / 232 / 233 / 281 / 294 / 296 / 307 / 335, 20 / 21 / 56 / 95 / 147 / 281 / 294 / 307, 20 / 21 / 56 / 281 / 307, 20 / 21 / 76 / 232 / 243, 20 / 21 / 95 / 232 / 307, 20 / 21 / 95 / 281 / 294 / 296, 20 / 21 / 147 / 189 / 233 / 243 / 281 / 307, 20 / 56, 20 / 56 / 76 / 95 / 281 / 307, 20 / 56 / 76 / 147 / 294 / 296 / 307, 20 / 56 / 95 / 147 / 294, 20 / 56 / 281, 20 / 76, 20 / 76 / 95 / 281 / 294 / 296, 20 / 76 / 95 / 281 / 296 / 307, 20 / 76 / 233 / 294 / 307, 20 / 76 / 243 / 281 / 294, 21 / 76 / 147 / 233 / 294 / 307, 21 / 76 / 147 / 243 / 296 / 307 / 335, 21 / 95 / 185 / 189 / 232 / 281 / 296, 21 / 95 / 233 / 243 / 281 / 296, 21 / 95 / 294 / 296 / 307 / 335,21 / 95 / 307, 21 / 281 / 307, 29 / 76 / 281, 56 / 76 / 95 / 232 / 243 / 281, 56 / 76 / 147 / 281 / 307, 56 / 76 / 243 / 294, 56 / 76 / 281 / 294, 56 / 76 / 296, 56 / 76 / 307, 56 / 95 / 147 / 307 / 335 / 348, 56 / 95 / 232 / 233 / 281 / 294 / 307, 56 / 95 / 243 / 281, 56 / 147 / 281, 56 / 232 / 243 / 281, 56 / 232 / 281, 56 / 232 / 281 / 294 / 296, 56 / 233 / 281 / 294 / 296, 56 / 281 / 307, 76 / 95 / 232 / 243 / 281 / 307, 76 / 95 / 243 / 281 / 307 / 335, 76 / 95 / 294 / 307, 76 / 147, 76 / 147 / 233 / 243 / 294, 76 / 147 / 233 / 281 / 294 / 307, 76 / 147 / 243 / 294 / 296 / 307 / 335, 76 / 147 / 281 / 307, 76 / 189 / 296, 76 / 232 / 233 / 243 / 294 / 296 / 307, 76 / 281, 76 / 281 / 294, 76 / 294 / 296, 95 / 120, 95 / 147 / 335, 95 / 232 / 243 / 281 / 294 / 307, 95 / 232 / 281 / 294 / 296, 95 / 281 / 294 / 296, 95 / 335, 147, 147 / 225 / 232 / 243 / 281 / 296 / 307 / 335, 147 / 233 / 243 / 281 / 307, 147 / 233 / 281 / 307 / 335, 147 / 243 / 281, 147 / 307, 232 / 233 / 281 / 294 / 296 / 307, 232 / 281, 232 / 284 / 307, 233 / 243 / 281 / 296 / 307 / 335, 233 / 281 / 296 / 307, 243 / 281 / 294 / 296, 281, 281 / 294, 281 / 307, 307, and 335. In some embodiments, the application provides an engineered polypeptide having proline hydroxylase activity, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 116 and one or more residue differences as compared to SEQ ID NO: 116 at a residue position selected from: 21 / 76 / 147 / 243 / 296 / 307 / 335,56 / 76 / 147 / 281 / 307 and 95 / 147 / 335. In some embodiments, the engineered polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to at least one of the even numbered sequences in SEQ ID NOs: 4-658.

[0018] The present application provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 162. In some embodiments, the present application provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 162 and one or more residue differences as compared to SEQ ID NO: 162 at a residue position selected from 2 / 85 / 123 / 237, 28 / 115 / 117 / 120 / 123 / 268 / 270 / 343 / 346 / 348, 45 / 123 / 326, 65 / 117 / 120 / 123 / 343 / 346, 85 / 123 / 281 / 282, 114 / 115 / 117 / 120 / 123 / 268 / 271 / 313 / 326 / 343 / 346, 123 / 139 / 233 / 237 / 281 / 282 / 289 / 324 / 326, and 123 / 199 / 200 / 247 / 250 / 338. In some embodiments, the engineered polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to at least one of the even numbered sequences in SEQ ID NOs: 4-658.

[0019] The present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 322. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 322 and one or more residue differences as compared to SEQ ID NO: 322 at a residue position selected from: 26, 54, 61, 129, 132, 149, 156, 175, 189, 201, 209, 228, 236, 248, 262, 272, 277, 291, and 345. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 322 and one or more residue differences as compared to SEQ ID NO: 322 at a residue position selected from: 25, 43, 54, 58, 61, 79, 129, 132, 143, 156, 163, 175, 179, 201, 209, 236, 248, 278, 291, 345, and 347. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 322 and one or more residue differences as compared to SEQ ID NO: 322 at a residue position selected from: 85 / 117 / 120 / 135 / 208 / 270 / 324 / 343 / 346, 85 / 117 / 120 / 135 / 208 / 281 / 282 / 289, 85 / 117 / 120 / 270 / 281 / 289, 85 / 117 / 135 / 139 / 208, and 117 / 120 / 208 / 270 / 324 / 343 / 346.In some embodiments, the engineered polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to at least one of the even numbered sequences in SEQ ID NOs: 6-658.

[0020] The present application provides engineered polypeptides having proline hydroxylase activity, said engineered polypeptides comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 412. In some embodiments, the present application provides engineered polypeptides having proline hydroxylase activity, said engineered polypeptides comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 412 and one or more residue differences as compared to SEQ ID NO: 412 at a residue position selected from: 47, 48, 56 / 118, 85, 95, 95 / 289, 113, 118, 118 / 247, 154, 162, 162 / 204, 164, 164 / 198 / 271, 168, 169, 187, 195, 243, 271, 275, 281, 314, 330, and 342. In some embodiments, the present application provides engineered polypeptides having proline hydroxylase activity, said engineered polypeptides comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 412 and one or more residue differences as compared to SEQ ID NO: 412 at a residue position selected from: 25 / 129 / 163 / 236 / 262 / 345 / 347, 120 / 156 / 175 / 179 / 201, 129 / 189 / 236 / 262 / 277 / 278, 129 / 236 / 262, 156 / 175 / 179 / 228, and 162. In some embodiments, the engineered polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to at least one of the even numbered sequences in SEQ ID NOs: 6-658.

[0021] The present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 492. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 492 and one or more residue differences as compared to SEQ ID NO: 492 at a residue position selected from: 15, 17, 28, 29, 65, 135, 167, 177, 199, 208, 228, 235, 287, 294, 307, and 343. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 492 and one or more residue differences as compared to SEQ ID NO: 492 at a residue position selected from: 85 / 187 / 281 / 347, 85 / 187 / 347, 118 / 120 / 162 / 175 / 179 / 330, 118 / 120 / 162 / 175 / 330, 162 / 175 / 179 / 330, 175 / 228 / 330, 195 / 347, and 278 / 314 / 347. In some embodiments, the engineered polypeptides have at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to at least one of the even-numbered sequences in SEQ ID NOs: 6-658.

[0022] The present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 562. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 562 and one or more residue differences as compared to SEQ ID NO: 562 at a residue position selected from: 15, 40, 43, 44, 59, 79, 82, 149, 164, 179, 345, and 347. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 562 and one or more residue differences as compared to SEQ ID NO: 562 at a residue position selected from: 29 / 85 / 177 / 208 / 228 / 347, 29 / 85 / 208 / 228 / 343 / 347, 29 / 177 / 195 / 228 / 343, 29 / 208 / 228 / 278 / 294 / 347, 56 / 195 / 278, 85 / 187 / 205 / 208 / 278, 113 / 177 / 187 / 195 / 208 / 278 / 294 / 343 / 347, and 177 / 205 / 208 / 228. In some embodiments, the engineered polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to at least one of the even-numbered sequences in SEQ ID NOs: 6-658.

[0023] The present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 598. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 598 and one or more residue differences as compared to SEQ ID NO: 598 at a residue position selected from: 47, 162, 209, 219, 227, and 342. In some embodiments, the present invention provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 598 and one or more residue differences as compared to SEQ ID NO: 598 at a residue position selected from: 17 / 44 / 179 / 195 / 250 / 313 / 345, 17 / 44 / 199 / 313, 43 / 44 / 195 / 199, 44 / 149 / 164 / 171 / 187, 44 / 179 / 195 / 199, 44 / 179 / 195 / 199 / 345, 79 / 163 / 164 / 171 / 187 / 201 / 286 / 288, 82 / 163 / 164, 82 / 163 / 164 / 171 / 187 / 201 / 203 / 208 / 286 / 288 / 320, 149 / 164 / 171 / 288, and 187 / 286. In some embodiments, the engineered polypeptide has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to at least one of the even-numbered sequences in SEQ ID NOs: 6-658.

[0024] The present application provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 630. In some embodiments, the present application provides engineered polypeptides having proline hydroxylase activity comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 630 and one or more residue differences as compared to SEQ ID NO: 630 at a residue position selected from: 82 / 164 / 171 / 203 / 208, 135 / 163 / 164 / 201 / 203 / 208, 162, 162 / 219 / 236, 162 / 219 / 313 / 338, 162 / 236 / 342, 162 / 313 / 342, and 164 / 171 / 201 / 203 / 282. In some embodiments, the engineered polypeptides have at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to at least one of the even-numbered sequences in SEQ ID NOs: 6-658.

[0025] The present application also provides engineered polypeptides having proline hydroxylase activity capable of converting L-proline to trans-3-hydroxyproline. In some embodiments, the engineered polypeptides are capable of converting L-proline to trans-3-hydroxyproline with at least 1.2-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold or more activity of a naturally occurring enzyme. In some further embodiments, the engineered polypeptides are capable of converting L-proline to trans-3-hydroxyproline with greater than 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more isomeric excess of trans-3-hydroxyproline.

[0026] The present application also provides polynucleotides encoding engineered polypeptides having proline hydroxylase activity. In some embodiments, the polynucleotides comprise nucleic acid sequences optimized for expression in E. coli.

[0027] The present application also provides an expression vector comprising a polynucleotide encoding an engineered polypeptide having proline hydroxylase activity. In some embodiments, the expression vector comprises at least one control sequence.

[0028] The present application also provides a host cell comprising a polynucleotide encoding an engineered polypeptide having proline hydroxylase activity. In some embodiments, the host cell is E. coli.

[0029] The present application also provides a method of making an engineered polypeptide having proline hydroxylase activity, comprising culturing a host cell comprising an expression vector comprising at least one polynucleotide encoding an engineered polypeptide having proline hydroxylase activity under conditions suitable for expression of the polypeptide. In some embodiments, the method further comprises the step of isolating the engineered polypeptide.

[0030] Invention Description

[0031] Unless defined otherwise, all technical and scientific terms and any acronyms used herein have the same meanings as commonly understood by one of ordinary skill in the art in the field of the application. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, microbiology, organic chemistry, analytical chemistry, and nucleic acid chemistry described below are those well-known and commonly used in the art. Such techniques are performed according to conventional methods formany references that are standard in the art. Such techniques are performed according to conventional methods well-known and commonly used in the art. Such techniques are performed according to conventional methods well-known and commonly used in the art. Many textbooks and reference works describe these techniques in sufficient detail. All patents, patent applications, articles, and publications mentioned herein, both supra and infra, are hereby expressly incorporated by reference.

[0032] Although any suitable methods and materials known to those skilled in the art can be used, some methods and materials are described herein. The present application should not be construed as limited to the particular methods, reagents, and conditions described, as such can vary. Hence, the terminology used herein is for the purpose of describing only the particular embodiments and is not intended to be limiting, as the present application is intended to cover all methods, reagents, and conditions that are functionally equivalent.

[0033] It should be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the application.

[0034] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0035] Numerical ranges include the numbers defining the range. Thus, every numerical range explicitly recited herein includes each narrower numerical range falling within such broader numerical range as if such narrower numerical range were explicitly recited herein. Also, every maximum numerical limitation recited herein includes every lower numerical limitation, as if such lower numerical limitations were expressly recited herein.

[0036] Abbreviations

[0037] Abbreviations for genetically encoded amino acids are conventional and are as follows:

[0038]

[0039]

[0040] When using the three-letter abbreviations, unless preceded by a specific “L” or “D,” or it is clear from the context of the use of the abbreviation, the amino acid can be either the L- or D- configuration about the alpha-carbon (C α ). For example, “Ala” denotes alanine without specifying the configuration about the alpha-carbon, while “D-Ala” and “L-Ala” denote D-alanine and L-alanine, respectively. When using the one-letter abbreviations, capital letters denote amino acids in the L-configuration about the alpha-carbon, and lower case letters denote amino acids in the D-configuration about the alpha-carbon. For example, “A” denotes L-alanine and “a” denotes D-alanine. When a polypeptide sequence is presented in a string of one-letter or three-letter abbreviations (or a mixture thereof), the sequence is presented in the amino (N) to carboxy (C) direction according to conventional practice.

[0041] Abbreviations for genetically encoded nucleosides are conventional and are as follows: adenosine (A); guanosine (G); cytidine (C); thymidine (T); and uridine (U). Unless specifically described, a nucleoside in an abbreviation can be either a ribonucleoside or a 2’-deoxyribonucleoside. The nucleoside can be specified as a ribonucleoside or a 2’-deoxyribonucleoside, either individually or on a total basis. When a nucleic acid sequence is presented in a string of one-letter abbreviations, the sequence is presented in the 5’ to 3’ direction and without showing the phosphates according to conventional practice.

[0042] Definitions

[0043] With reference to the present application, the technical and scientific terms used in the description herein will have the meanings commonly understood by one of ordinary skill in the art, unless specifically defined otherwise. Thus, the following terms are intended to have the following meanings.

[0044] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polypeptide" includes more than one polypeptide.

[0045] Similarly, "comprise," "comprises," "comprising," "include," "includes," and "including" are interchangeable, and not intended to be limiting. As used herein, the term "comprising" and its cognate forms are used in their inclusive sense (i.e., equivalent to the term "including" and its corresponding cognate forms).

[0046] It is also to be understood that where the description of various embodiments uses terms like "comprising", those skilled in the art will understand that in some specific instances, "consisting essentially of" or "consisting of" language can be used instead.

[0047] The term "about" means an acceptable error in a specified value. In some examples, "about" means within 0.05%, 0.5%, 1.0%, or 2.0% of a given value. In some examples, "about" means within 1, 2, 3, or 4 standard deviations of a given value.

[0048] "EC" number refers to the enzyme nomenclature of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (NC-IUBMB). The IUBMB biochemical classification is a numerical classification system based on the chemical reactions catalyzed by enzymes.

[0049] "ATCC" refers to the American Type Culture Collection, a collection of biological deposits including genes and strains.

[0050] "NCBI" refers to the National Center for Biological Information and the sequence databases provided therein.

[0051] "Protein," "polypeptide," and "peptide" are used interchangeably herein to refer to a polymer of at least two amino acids covalently linked by amide bonds, regardless of length or post-translational modification (e.g., glycosylation or phosphorylation). Included in this definition are D-amino acids and L-amino acids, as well as mixtures of D-amino acids and L-amino acids, and polymers comprising D-amino acids and L-amino acids and mixtures of D-amino acids and L-amino acids.

[0052] "Amino acids" are referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, can be referred to by their commonly accepted single-letter codes.

[0053] As used herein, "polynucleotide" and "nucleic acid" refer to two or more nucleosides covalently linked together. A polynucleotide can comprise entirely ribonucleotides (i.e., RNA), entirely 2' deoxyribonucleotides (i.e., DNA), or a mixture of ribonucleotides and 2' deoxyribonucleotides. While nucleosides will typically be linked together via standard phosphodiester linkages, a polynucleotide can comprise one or more non-standard linkages. A polynucleotide can be single-stranded or double-stranded, or can comprise both single-stranded and double-stranded regions. Furthermore, while a polynucleotide will typically comprise naturally occurring coding nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), it can comprise one or more modified nucleobases and / or synthetic nucleobases, such as, for example, inosine, xanthine, hypoxanthine, etc. In some embodiments, such modified nucleobases or synthetic nucleobases are nucleobases that encode an amino acid sequence.

[0054] "Coding sequence" refers to a portion of nucleic acid (e.g., a gene) that encodes an amino acid sequence of a protein.

[0055] "Proline hydroxylase" refers to a polypeptide having the enzymatic ability to convert free proline to hydroxyproline in the presence of the co-substrate alpha-ketoglutarate and molecular oxygen (dioxygen), as exemplified below:

[0056]

[0057] It is understood that proline hydroxylases are not limited to the above-described reaction with proline, but can hydroxylate other substrates or produce various isomers of hydroxyproline, such as trans-3-hydroxyproline. Proline hydroxylases as used herein include naturally occurring (wild-type) proline hydroxylases as well as non-naturally occurring engineered polypeptides generated by human manipulation. In some embodiments, the proline hydroxylase variants of the present application are capable of converting L-proline to trans-3-hydroxyproline, as exemplified below in Scheme 1:

[0058]

[0059] "Co-substrate" for proline hydroxylases refers to alpha-ketoglutarate and co-substrate analogs that can substitute for alpha-ketoglutarate in the hydroxylation reaction of proline and proline substrate analogs. Co-substrate analogs include, for example, but are not limited to, 2-oxoadipate (see, e.g., Majamaa et al., Biochem. J., 229: 127-133

[1985] ).

[0060] As used herein, "wild type" and "naturally occurring" refer to the form found in nature. For example, a wild type polypeptide or polynucleotide sequence is one that occurs in an organism, which can be isolated from a natural source and has not been intentionally modified by human manipulation.

[0061] "Recombinant" or "engineered" or "non-naturally occurring" when used with reference to a cell, nucleic acid, or polypeptide means a material that has been altered in a manner that would not otherwise exist in nature or that corresponds to a material in its natural or native form. In some embodiments, the cell, nucleic acid, or polypeptide is identical to a naturally occurring cell, nucleic acid, or polypeptide, but is produced or derived from synthetic materials and / or manipulated through the use of recombinant techniques. Non-limiting examples include, among others, recombinant cells that express genes not found in the natural (non-recombinant) form of the cell or that express natural genes at levels different than found in the natural form.

[0062] The term "percent (%) sequence identity" is used herein to refer to a comparison between polynucleotides or polypeptides, and is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window can comprise additions or deletions (i.e., gaps) as compared to the reference sequence for optimal alignment of the two sequences. The percent identity can be calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to yield the percent sequence identity. Alternatively, the percent identity can be calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue or a nucleic acid base or amino acid residue aligned with a gap occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to yield the percent sequence identity. Those of skill in the art understand that there are a number of established algorithms for aligning two sequences. Optimal alignment of sequences for comparison can be conducted by any suitable method, including but not limited to the local homology algorithm of Smith and Waterman (Smith and Waterman, Adv. Appl. Math., 2:482

[1981] ), by the homology alignment algorithm of Needleman and Wunsch (Needleman and Wunsch, J. Mol. Biol., 48:443

[1970] ), by the search for similarity method of Pearson and Lipman (Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444

[1988] ), by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin software package), or by visual inspection, as is known in the art. Examples of algorithms that are suitable for determining percent sequence identity and sequence similarity include, but are not limited to, the BLAST and BLAST 2.0 algorithms, which are described by Altschul et al. (see, respectively, Altschul et al., J. Mol. Biol., 215:403-410

[1990] ; and Altschul et al., Nucl. Acids Res., 3389-3402

[1977] ). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information website. This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that either match or satisfy some positive-valued score threshold T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (see, Altschul et al., supra).These initial neighborhood word hits act as seeds to initiate a search to find longer HSPs containing them. Then the word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always > 0) and N (penalty score for mismatching residues; always < 0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieve value by the inclusion of one or more negative-scoring residue alignments; the total length of the alignment between the two sequences reaches the maximum length defined by the lengths of the two sequences; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as default values a word length (W) of 11, an expectation value (E) of 10, M=5, N=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation value (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89:10915

[1989] ). An exemplary determination of sequence alignment and % sequence identity can use the BESTFIT or GAP programs in the GCG Wisconsin Software Package (Accelrys, Madison WI), using default parameters.

[0063] A "reference sequence" is a defined sequence used as a basis for sequence and / or activity comparisons. A reference sequence can be a subset of a larger sequence, for example, a segment of a full-length gene or polypeptide sequence. Generally, a reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 residues in length, at least 50 residues in length, at least 100 residues in length, or is the full-length of a nucleic acid or polypeptide. Because two polynucleotides or polypeptides can each (1) comprise a sequence that is similar between the two sequences (i.e., a portion of the complete sequence), and (2) can also contain divergent sequences between the two sequences, sequence comparisons between two (or more) polynucleotides or polypeptides are typically performed by comparing sequences of the two polynucleotides or polypeptides over a "comparison window" to identify and compare local regions of sequence similarity. In some embodiments, a "reference sequence" can be based on a primary amino acid sequence, wherein the reference sequence is a sequence that can have one or more variations in the primary sequence.

[0064] As used herein, a "comparison window" refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acid residues wherein sequences can be compared to a reference sequence of at least 20 contiguous nucleotides or amino acids and wherein the portion of the sequences in the comparison window can include up to 20% or less additions or deletions (i.e., gaps) for optimal alignment of the two sequences. A comparison window can be longer than 20 contiguous residues, and optionally includes windows of 30, 40, 50, 100 or longer.

[0065] "Corresponding to," "reference to," or "relative to," when used in the context of numbering of a given amino acid or polynucleotide sequence, refers to the numbering of the residues of the specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue numbering or residue position of a given polymer is specified with respect to the reference sequence, rather than by the actual numerical position of the residue within the given amino acid or polynucleotide sequence. For example, given an amino acid sequence, such as the amino acid sequence of an engineered proline hydroxylase, can be optimized for residue matches between the two sequences by introducing gaps to align the two sequences. In these cases, the residues in the given amino acid or polynucleotide sequence are numbered with respect to the reference sequence to which it is aligned, despite the presence of gaps.

[0066] "Substantial identity" means, in the context of a polynucleotide or polypeptide sequence, a sequence having at least 80% sequence identity, at least 85% identity, at least between about 89-95% sequence identity, or more usually at least 99% sequence identity as compared to a reference sequence over a comparison window of at least 20 residues, usually 30-50 residues, wherein the percentage of sequence identity is calculated by comparing the reference sequence to the sequence including up to 20% of the total number of residues in the reference sequence that are deleted or added to the sequence. In some embodiments applying to polypeptides, the term "substantial identity" means that the two polypeptide sequences have at least 80% sequence identity, preferably at least 89% sequence identity or at least 95% sequence identity or more (e.g., 99% sequence identity) when optimally aligned, such as by the program GAP or BESTFIT using default gap weights. In some embodiments, residue positions that are not identical differ by conservative amino acid substitutions.

[0067] As used herein, “amino acid difference” and “residue difference” refer to a difference in an amino acid residue at a position of a polypeptide sequence relative to the amino acid residue at the corresponding position in a reference sequence. The position of an amino acid difference is often referred to herein as “Xn,” where n refers to the corresponding position in the reference sequence upon which the residue difference is based. For example, “a residue difference at position X93 compared to SEQ ID NO: 4” refers to a difference in the amino acid residue at the polypeptide position corresponding to position 93 of SEQ ID NO: 4. Thus, if the reference polypeptide SEQ ID NO: 4 has a serine at position 93, “a residue difference at position X93 compared to SEQ ID NO: 4” refers to an amino acid substitution of any residue other than serine at the polypeptide position corresponding to position 93 of SEQ ID NO: 4. In most cases herein, a particular amino acid residue difference at a position is indicated as “XnY,” where “Xn” designates the corresponding position as described above, and “Y” is the one-letter identifier for the amino acid found in the engineered polypeptide (i.e., the residue that is different from that in the reference polypeptide). In some cases (e.g., in Tables 4.1, 4.2, 4.3, 4.4, 5.1, 5.2, 5.3, 6.1, 7.1, 7.2, 7.3, 8.1, 8.2, 9.1, 9.2, 10.1, 10.2, 11.1, 11.2, and / or 12.1), the application also provides for particular amino acid differences represented by the conventional notation “AnB,” where A is the one-letter identifier for the residue in the reference sequence, “n” is the number of the residue position in the reference sequence, and B is the one-letter identifier for the residue substitution in the sequence of the engineered polypeptide. In some cases, the polypeptides of the application comprise one or more amino acid residue differences relative to a reference sequence, which are indicated by a list of the specified positions at which a residue difference exists relative to the reference sequence. In some embodiments, where more than one amino acid can be used in a particular residue position of a polypeptide, the various amino acid residues that can be used are separated by a “ / ” (e.g., X307H / X307P or X307H / P). A slash can also be used to indicate more than one substitution within a given variant (i.e., more than one substitution exists in a given sequence such as in a combinatorial variant). In some embodiments, the application includes engineered polypeptide sequences that contain one or more amino acid differences, which include conservative amino acid substitutions or non-conservative amino acid substitutions. In some further embodiments, the application provides engineered polypeptide sequences that contain both conservative amino acid substitutions and non-conservative amino acid substitutions.

[0068] As used herein, "conservative amino acid substitution" refers to the substitution of one residue for another residue with similar side chains, and thus generally involves the substitution of an amino acid in a polypeptide with an amino acid in the same or similar amino acid definition category. For example, but not by way of limitation, in some embodiments, an amino acid with an aliphatic side chain is substituted with another aliphatic amino acid (e.g., alanine, valine, leucine, and isoleucine); an amino acid with a hydroxyl side chain is substituted with another amino acid with a hydroxyl side chain (e.g., serine and threonine); an amino acid with an aromatic side chain is substituted with another amino acid with an aromatic side chain (e.g., phenylalanine, tyrosine, tryptophan, and histidine); an amino acid with a basic side chain is substituted with another amino acid with a basic side chain (e.g., lysine and arginine); an amino acid with an acidic side chain is substituted with another amino acid with an acidic side chain (e.g., aspartic acid or glutamic acid); and / or a hydrophobic or a hydrophilic amino acid is substituted with another hydrophobic or a hydrophilic amino acid, respectively.

[0069] As used herein, "non-conservative substitution" refers to the substitution of an amino acid in a polypeptide with an amino acid having significantly different side chain properties. Non-conservative substitutions can use amino acids between, but not within, defined groups and affect: (a) the structure of the peptide backbone in the region of the substitution (e.g., a proline in place of glycine); (b) the charge or hydrophobicity of the region; or (c) the bulk of the side chain.

[0070] As used herein, "deletion" refers to a modification of a polypeptide by removing one or more amino acids from a reference polypeptide. A deletion can include removal of 1 or more amino acids, 2 or more amino acids, 5 or more amino acids, 10 or more amino acids, 15 or more amino acids, or 20 or more amino acids, up to 10% of the total number of amino acids comprising the reference enzyme or up to 20% of the total number of amino acids, while retaining enzymatic activity and / or retaining improved properties of the engineered proline hydroxylase. A deletion can involve an internal portion and / or a terminal portion of a polypeptide. In various embodiments, a deletion can include contiguous segments or can be non-contiguous.

[0071] As used herein, "insertion" refers to a modification of a polypeptide by adding one or more amino acids to a reference polypeptide. An insertion can be at an internal portion of a polypeptide or can be an insertion to the carboxyl or amino terminus. Insertions as used herein include fusion proteins as known in the art. An insertion can be a contiguous segment of amino acids or separated by one or more amino acids in a naturally occurring polypeptide.

[0072] "Functional fragment" or "biologically active fragment," as used interchangeably herein, refers to a polypeptide that has an amino-terminal deletion and / or a carboxyl-terminal deletion and / or internal deletion, but in which the remaining amino acid sequence is identical to the corresponding positions in the sequence with which it is compared (e.g., a full-length engineered proline hydroxylase of the application) and retains substantially all of the activity of the full-length polypeptide.

[0073] As used herein, "isolated polypeptide" refers to a polypeptide that is substantially free of other contaminant materials (e.g., proteins, lipids, and polynucleotides) with which it can be found in nature. The term includes polypeptides that have been removed from their naturally occurring environment or expression system (e.g., within a host cell or via in vitro synthesis) or purified. A recombinant proline hydroxylase polypeptide can be present within a cell, present in cell culture media, or prepared in a variety of forms, such as a lysate or an isolated preparation. Thus, in some embodiments, a recombinant proline hydroxylase polypeptide can be an isolated polypeptide.

[0074] As used herein, "substantially pure polypeptide" refers to a composition in which the polypeptide material is the predominant species (i.e., present in a molar or mass excess as compared to any other individual macromolecular species in the composition) and the composition is typically a substantially purified composition when the polypeptide of interest constitutes at least about 50% by mole or % weight of the macromolecular species present. However, in some embodiments, a composition comprising a proline hydroxylase comprises less than 50% pure (e.g., about 10%, about 20%, about 30%, about 40%, or about 50%) proline hydroxylase. Typically, a substantially pure proline hydroxylase composition comprises about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, and about 98% or more by mole or % weight of all macromolecular species present in the composition. In some embodiments, the polypeptide of interest is purified to essentially homogeneity (i.e., no detectable contaminating material by conventional detection methods) in which the composition consists essentially of a single macromolecular species. Solvent species, small molecules (< 500 daltons), and elemental ion species are not considered macromolecular species. In some embodiments, an isolated recombinant proline hydroxylase polypeptide is a substantially pure polypeptide composition.

[0075] As used herein, "improved enzyme properties" refers to at least one improved property of an enzyme. In some embodiments, the present application provides an engineered proline hydroxylase polypeptide that exhibits an improvement in any enzyme property as compared to a reference proline hydroxylase polypeptide and / or a wild-type proline hydroxylase polypeptide and / or another engineered proline hydroxylase polypeptide. Thus, a variety of proline hydroxylases, including wild-type and engineered proline hydroxylases, can be determined and compared for the level of "improvement." Improved properties include, but are not limited to, properties such as: increased protein expression, increased thermoactivity, increased thermal stability, increased pH activity, increased stability, increased enzymatic activity, increased substrate specificity or affinity, increased specific activity, increased resistance to substrate or end product inhibition, increased chemical stability, improved chemical selectivity, improved solvent stability, increased tolerance to acidic pH, increased tolerance to basic pH, increased tolerance to proteolytic activity (i.e., decreased sensitivity to proteolysis), decreased aggregation, increased solubility, and altered temperature profile.

[0076] As used herein, "increased enzyme activity" and "enhanced catalytic activity" refer to an improved property of an engineered proline hydroxylase polypeptide that can be expressed as an increase in specific activity (e.g., product produced / time / weight protein) or an increase in the percentage conversion of substrate to product (e.g., the percentage conversion of an initial amount of substrate to product using a specified amount of proline hydroxylase enzyme over a specified period of time) as compared to a reference proline hydroxylase. Exemplary methods of determining enzyme activity are provided in the Examples. Any property related to enzyme activity can be affected, including the classical enzyme properties K m , V max or k cat , the change of which can result in an increase in enzymatic activity. The improvement in enzyme activity can be from about 1.1-fold to as much as 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold or more of the enzymatic activity of the corresponding wild-type enzyme, a naturally occurring proline hydroxylase, or another engineered proline hydroxylase from which the proline hydroxylase polypeptide is derived.

[0077] As used herein, "conversion" refers to the enzymatic conversion (or biotransformation) of one or more substrates to one or more corresponding products. "Percentage conversion" refers to the percentage of substrate that is converted to product over a certain period of time under specified conditions. Thus, the "enzymatic activity" or "activity" of a proline hydroxylase polypeptide can be expressed as the "percentage conversion" of substrate to product over a specified period of time.

[0078] Enzymes having "generalist properties" (or "generalist enzymes") refer to enzymes that exhibit improved activity on a broad range of substrates compared to the parent sequence. Generalist enzymes do not necessarily exhibit improved activity on every possible substrate. In some embodiments, the present application provides proline hydroxylase variants having generalist properties in that they exhibit similar or improved activity on a broad range of spatially and electronically diverse substrates relative to the parent gene. Further, the generalist enzymes provided herein are engineered to be improved across a broad range of diverse API-like molecules to increase production of metabolites / products.

[0079] The term "stringent hybridization conditions" is used herein to refer to conditions under which a nucleic acid hybrid is stable. As known to those skilled in the art, stability of a hybrid is reflected in the melting temperature (T m ) of the hybrid. Generally, stability of a hybrid is a function of ionic strength, temperature, G / C content, and presence of a chaotropic agent. The T m value of a polynucleotide can be calculated using known methods for predicting melting temperature (see, e.g., Baldino et al., Meth. Enzymol., 168:761-777

[1989] ; Bolton et al., Proc. Natl. Acad. Sci. USA 48: 1390

[1962] ; Bresslauer et al., Proc. Natl. Acad. Sci. USA 83:8893-8897

[1986] ; Freier et al., Proc. Natl. Acad. Sci. USA 83:9373-9377

[1986] ; Kierzek et al., Biochem., 25:7840-7846

[1986] ; Rychlik et al., Nucl. Acids Res., 18:6409-6412

[1990] (erratum, Nucl. Acids Res., 19:698

[1991] ); Sambrook et al., supra); Suggs et al., 1981, in Developmental Biology Using Purified Genes , Brown et al. [eds.], pp. 683-693, Academic Press, Cambridge, MA

[1981] ; and Wetmur, Crit. Rev. Biochem. Mol. Biol. 26:227-259

[1991] ). In some embodiments, a polynucleotide encodes a polypeptide disclosed herein and hybridizes under defined conditions, such as at moderate stringency or at high stringency, to a complement of a sequence encoding an engineered proline hydroxylase of the application.

[0080] "Hybridization stringency" refers to the conditions in nucleic acid hybridization, such as wash conditions. Typically, hybridization reactions are performed under conditions of lower stringency, followed by different but higher stringency washes. The term "moderate stringency hybridization" refers to conditions which allow a target DNA to bind to a complementary nucleic acid having about 60% identity, preferably about 75% identity, about 85% identity, greater than about 90% identity to the target polynucleotide. Exemplary moderate stringency conditions are conditions equivalent to those in 50% formamide, 5 x Denhart's solution, 5 x SSPE, 0.2% SDS, at 42°C, followed by a wash in 0.2 x SSPE, 0.2% SDS at 42°C. "High stringency hybridization" typically refers to conditions which differ from solution conditions by about 10°C or less from the thermal melting point Tm of the defined polynucleotide sequence in solution. In some embodiments, high stringency conditions refer to conditions which allow hybridization of only those nucleic acid sequences which form stable hybrids in 0.018 M NaCl at 65°C (i.e., if a hybrid is not stable in 0.018 M NaCl at 65°C, it will not be stable under high stringency conditions as contemplated herein). For example, high stringency conditions can be provided by hybridization in conditions equivalent to 50% formamide, 5 x Denhart's solution, 5 x SSPE, 0.2% SDS at 42°C, followed by a wash in 0.1 x SSPE and 0.1% SDS at 65°C. Another high stringency condition is hybridization in conditions equivalent to hybridization in 5 x SSC containing 0.1% (w:v) SDS at 65°C and a wash in 0.1 x SSC containing 0.1% SDS at 65°C. Other high stringency hybridization conditions, as well as moderate stringency conditions, are described in the references cited above. m about 10°C or less from the thermal melting point Tm of the defined polynucleotide sequence in solution. In some embodiments, high stringency conditions refer to conditions which allow hybridization of only those nucleic acid sequences which form stable hybrids in 0.018 M NaCl at 65°C (i.e., if a hybrid is not stable in 0.018 M NaCl at 65°C, it will not be stable under high stringency conditions as contemplated herein). For example, high stringency conditions can be provided by hybridization in conditions equivalent to 50% formamide, 5 x Denhart's solution, 5 x SSPE, 0.2% SDS at 42°C, followed by a wash in 0.1 x SSPE and 0.1% SDS at 65°C. Another high stringency condition is hybridization in conditions equivalent to hybridization in 5 x SSC containing 0.1% (w:v) SDS at 65°C and a wash in 0.1 x SSC containing 0.1% SDS at 65°C. Other high stringency hybridization conditions, as well as moderate stringency conditions, are described in the references cited above.

[0081] " Codon-optimized" refers to the alteration of the codons of a polynucleotide encoding a protein to those codons preferentially used in a particular organism, such that the encoded protein is efficiently expressed in the organism of interest. Although the genetic code is degenerate, i.e., most amino acids are represented by several codons known as "synonyms" or "synonymous" codons, it is well known that the codon usage of a particular organism is non-random and biased for particular codon triplets. This codon usage bias can be higher for given genes, genes of common function or ancestral origin, highly expressed proteins versus low copy number proteins, and aggregated protein coding regions of a genome. In some embodiments, the polynucleotide encoding proline hydroxylase can be codon-optimized for optimal production in a host organism selected for expression.

[0082] "Preferred, optimal, high-codon-use biased codons" can be interchangeably referred to as codons used more frequently in protein-coding regions than other codons encoding the same amino acid. Preferred codons can be determined based on codon usage in a single gene, a group of genes with a common function or origin, codon usage in highly expressed genes, codon frequency in aggregated protein-coding regions throughout an organism, codon frequency in aggregated protein-coding regions of related organisms, or combinations thereof. Codons whose frequency increases with the level of gene expression are generally the optimal codons for expression. A variety of methods are known for determining codon frequencies (e.g., codon usage, relative synonymous codon usage) and codon preferences in a particular organism, including multivariate analysis, such as using cluster analysis or correlation analysis, and the effective number of codons used in genes (see, for example, GCG CodonPreference, Genetics Computer Group Wisconsin Package; Codon W, Peden, University of Nottingham; McInerney, Bioinform., 14:372-73

[1998] ; Stenico et al., Nucl. Acids Res., 222437-46

[1994] ; Wright, Gene 87:23-29

[1990] ). For many different organisms, codon usage tables are available (see, for example, Wada et al., Nucl. Acids Res., 20:2111-2118

[1992] ; Nakamura et al., Nucl. Acids Res., 28:292

[2000] ; Duret et al., above; Henaut and Danchin, at...). Escherichia coli and Salmonella Neidhardt, et al. (edited), ASMPress, Washington DC, pp. 2047-2066

[1996] ). Data sources used to obtain codon usage can rely on any available nucleotide sequence capable of encoding a protein. These datasets include nucleic acid sequences that actually encode expressed proteins (e.g., complete protein-coding sequences - CDS), expressed sequence tags (ESTS), or predicted coding regions of genomic sequences (see, for example, Mount, Neidhardt, et al. (edited), ASMPress, Washington DC, pp. 2047-2066

[1996] ). Bioinformatics:Sequence and Genome Analysis Chapter 8, ColdSpring Harbor Laboratory Press, Cold Spring Harbor, NY

[2001] ; Uberbacher, Meth. Enzymol., 266:259-281

[1996] ; and Tiwari et al., Comput. Appl. Biosci., 13:263-270

[1997] ).

[0083] "Control sequences" herein refer to all components which are necessary or advantageous for the expression of a polynucleotide and / or polypeptide of the application. Each of the control sequences can be native or foreign to the nucleic acid sequence encoding the polypeptide. Such control sequences include, but are not limited to, a leader, polyadenylation sequence, propeptide sequence, promoter, signal peptide sequence, initiation signal and transcription terminator. At a minimum, the control sequences include a promoter and transcriptional and translational stop signals. The control sequences can be provided with linkers for the purpose of introducing specific restriction sites facilitating ligation of the control sequences with the coding region of the nucleic acid sequence encoding the polypeptide.

[0084] "Operably linked" is defined herein as a configuration in which control sequences are appropriately placed (i.e., in functional relationship) with respect to a polynucleotide of interest such that the control sequences direct or modulate expression of the polynucleotide and / or polypeptide of interest.

[0085] "Promoter sequence" refers to a nucleic acid sequence recognized by a host cell for the initiation of expression of a polynucleotide of interest, such as a coding sequence. The promoter sequence contains transcriptional control sequences that mediate the expression of a polynucleotide of interest. The promoter can be any nucleic acid sequence which shows transcriptional activity in the host cell of choice including mutant, truncated, and hybrid promoters, and can be derived from genes encoding proteins either homologous or heterologous to the host cell.

[0086] "Suitable reaction conditions" refer to those conditions (e.g., ranges of enzyme loading, substrate loading, temperature, pH, buffers, cosolvents, etc.) in the enzymatic conversion reaction solution under which a proline hydroxylase polypeptide of the application is capable of converting a substrate to a desired product compound. Some exemplary "suitable reaction conditions" are provided herein.

[0087] As used herein, "loading", such as in "compound loading" or "enzyme loading", refers to the concentration or amount of a component in a reaction mixture at the initiation of a reaction.

[0088] As used herein, in the context of an enzymatic conversion reaction process, "substrate" refers to a compound or molecule acted upon by a proline hydroxylase polypeptide.

[0089] As used herein, in the context of an enzymatic conversion process, "product" refers to a compound or molecule produced from the action of a proline hydroxylase polypeptide on a substrate.

[0090] The term "culturing" as used herein refers to the growth of a population of microbial cells under any suitable conditions (e.g., using liquid, gel or solid culture media).

[0091] Recombinant polypeptides can be produced using any suitable method known in the art. Genes encoding wild-type polypeptides of interest can be cloned into vectors, such as plasmids, and expressed in a desired host, such as E. coli, and the like. Variants of recombinant polypeptides can be produced by various methods known in the art. In fact, there are a wide variety of different mutagenesis techniques well known to those of skill in the art. Moreover, mutagenesis kits are also available from many commercial molecular biology suppliers. Methods can be used to make specific substitutions at defined amino acids (site-directed), specific (region-specific) or random mutations in a local region of the gene, or random mutagenesis (e.g., saturation mutagenesis) throughout the gene. Numerous suitable methods for producing enzyme variants are known to those of skill in the art, including but not limited to, site-directed mutagenesis using PCR on single- or double-stranded DNA, cassette mutagenesis, gene synthesis, error-prone PCR, shuffling, and chemical saturation mutagenesis, or any other suitable method known in the art. Non-limiting examples of methods for DNA and protein engineering are provided in U.S. Patent No. 6,117,679; U.S. Patent No. 6,420,175; U.S. Patent No. 6,376,246; U.S. Patent No. 6,586,182; U.S. Patent No. 7,747,391; U.S. Patent No. 7,747,393; U.S. Patent No. 7,783,428; and U.S. Patent No. 8,383,346. After variants are produced, they can be screened for any desired property (e.g., high or increased activity, or low or decreased activity, increased thermal activity, increased thermal stability, and / or acid pH stability, etc.). In some embodiments, a “recombinant proline hydroxylase polypeptide” (also referred to herein as an “engineered proline hydroxylase polypeptide,” an “engineered proline hydroxylase polypeptide,” a “variant proline hydroxylase,” and a “proline hydroxylase variant”) can be used.

[0092] As used herein, a “vector” is a DNA construct used to introduce a DNA sequence into a cell. In some embodiments, a vector is an expression vector that is operably linked to suitable control sequences capable of effecting expression of a polypeptide encoded in the DNA sequence in a suitable host. In some embodiments, an “expression vector” has a promoter sequence operably linked to a DNA sequence (e.g., a transgene) to drive expression in a host cell, and in some embodiments, also includes a transcription terminator sequence.

[0093] As used herein, the term “expression” includes any step involved in the process of polypeptide production, including but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of a polypeptide from a cell.

[0094] As used herein, the term "producing" refers to the production of a protein and / or other compound by a cell. It is intended that the term encompass any steps involved in the production of a polypeptide, including but not limited to transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses the secretion of a polypeptide from a cell.

[0095] As used herein, an amino acid or nucleotide sequence (e.g., a promoter sequence, a signal peptide, a terminator sequence, etc.) is "heterologous" to another sequence with which it is operably linked if the two sequences are not associated in nature. For example, a "heterologous polynucleotide" is any polynucleotide that is introduced into a host cell by laboratory techniques, and includes polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into a host cell.

[0096] As used herein, the terms "host cell" and "host strain" refer to a suitable host for an expression vector comprising the DNA provided herein (e.g., a polynucleotide encoding a proline hydroxylase variant). In some embodiments, a host cell is a prokaryotic or eukaryotic cell that has been transformed or transfected with a vector constructed using recombinant DNA techniques as known in the art.

[0097] The term "analog" means a polypeptide having more than 70% sequence identity, but less than 100% sequence identity (e.g., more than 75%, 78%, 80%, 83%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity) to a reference polypeptide. In some embodiments, an analog means a polypeptide that comprises one or more non-naturally occurring amino acid residues (including but not limited to homoarginine, ornithine, and norvaline) as well as naturally occurring amino acids. In some embodiments, an analog also comprises one or more D-amino acid residues and a non-peptide linkage between two or more amino acid residues.

[0098] The term "effective amount" means an amount that is sufficient to produce a desired result. One of ordinary skill in the art can determine what is an effective amount by using routine experimentation.

[0099] The terms "isolated" and "purified" are used to refer to a molecule (e.g., an isolated nucleic acid, polypeptide, etc.) or other component that has been removed from at least one other component with which it is naturally associated. The term "purified" does not require absolute purity; rather, it is intended as a relative definition.

[0100] "stereoselectivity" refers to the preferential formation of one stereoisomer over another in a chemical or enzymatic reaction. Stereoselectivity can be partial, in which the formation of one stereoisomer is favored over the other, or stereoselectivity can be complete, in which only one stereoisomer is formed. When the stereoisomers are enantiomers, the stereoselectivity is referred to as enantioselectivity, i.e., the fraction of one enantiomer (typically reported as a percentage) out of the sum of the two. It is often alternatively reported (typically as a percentage) as the enantiomeric excess (e.e.) calculated therefrom according to the formula: [major enantiomer - minor enantiomer] / [major enantiomer + minor enantiomer]. When the stereoisomers are diastereomers, the stereoselectivity is referred to as diastereoselectivity, i.e., the fraction of one diastereomer (typically reported as a percentage) out of the mixture of the two diastereomers, often alternatively reported as the diastereomeric excess (d.e.). Enantiomeric excess and diastereomeric excess are types of stereomeric excess.

[0101] "highly stereoselective" refers to a chemical or enzymatic reaction that is capable of converting a substrate (e.g., L-proline) to its corresponding hydroxylated product (e.g., trans-3-hydroxyproline) with a stereomeric excess of at least about 85%.

[0102] "regioselectivity" or "regioselective reaction" refers to a reaction in which one direction of bond formation or cleavage is favored over all other possible directions. If the distinction is absolute, the reaction can be completely (100%) regioselective, if the reaction product at one site is favored over the reaction product at other sites, e.g., preferential formation of the product compound is (i.e., trans-3-hydroxyproline over the undesired product trans-4-hydroxyproline), the reaction can be substantially regioselective (at least 75%), or partially regioselective (x%, where the percentage is set depending on the reaction of interest).

[0103] "selective" or "selectivity" can refer to either stereoselectivity or regioselectivity, as defined above, or can refer to both stereoselectivity and regioselectivity.

[0104] "stereomeric excess" refers to the percentage calculated according to the formula: [major isomer - minor isomer] / [major isomer + minor isomer]. This percentage represents the preferential formation of one isomer over another in a chemical or enzymatic reaction. Enantiomeric excess is a form of stereomeric excess.

[0105] As used herein, "heat stable" refers to a proline hydroxylase polypeptide that maintains similar activity (e.g., more than 60% to 80%) after exposure to an elevated temperature (e.g., 40°C to 80°C) for a period of time (e.g., 0.5h-24h) as compared to a wild-type enzyme exposed to the same elevated temperature for the same period of time.

[0106] As used herein, "solvent stable" refers to a proline hydroxylase polypeptide that maintains similar activity (e.g., more than 60% to 80%) after exposure to a solvent (e.g., ethanol, isopropanol, dimethyl sulfoxide [DMSO], tetrahydrofuran, 2-methyltetrahydrofuran, acetone, toluene, butyl acetate, methyl tert-butyl ether, etc.) at varying concentrations (e.g., 5%-99%) for a period of time (e.g., 0.5h-24h) as compared to a wild-type enzyme exposed to the same solvent at the same concentration for the same period of time.

[0107] As used herein, "heat stable and solvent stable" refers to a proline hydroxylase polypeptide that is both heat stable and solvent stable.

[0108] As used herein, "reducing agent" refers to a compound or agent that is capable of converting Fe +3 to Fe +2 . An exemplary reducing agent is ascorbic acid, which is typically in the form of L-ascorbic acid.

[0109] "Alkyl" refers to a saturated hydrocarbon group of from 1 to 18 carbon atoms (inclusive), straight-chained or branched, more preferably 1 to 8 carbon atoms (inclusive), and most preferably 1 to 6 carbon atoms (inclusive). Alkyl groups having a specified number of carbon atoms are indicated in parentheses (e.g., (C1-C6) alkyl refers to an alkyl group of from 1 to 6 carbon atoms).

[0110] "Alkenyl" refers to a hydrocarbon group of from 2 to 12 carbon atoms (inclusive), straight-chained or branched, containing at least one double bond, but optionally containing more than one double bond.

[0111] "Alkynyl" refers to a hydrocarbon group of from 2 to 12 carbon atoms (inclusive), straight-chained or branched, containing at least one triple bond, but optionally containing more than one triple bond, and additionally optionally containing one or more double bond linkage moieties.

[0112] "Alkylene" refers to a straight-chained or branched divalent hydrocarbon group of from 1 to 18 carbon atoms (inclusive), more preferably 1 to 8 carbon atoms (inclusive), and most preferably 1 to 6 carbon atoms (inclusive), optionally substituted with one or more suitable substituents. Exemplary "alkylene" groups include, but are not limited to, methylene, ethylene, propylene, butylene, and the like.

[0113] "Alkenyl" refers to a straight or branched hydrocarbon chain radical having from 2 to 12 carbon atoms (inclusive), and preferably 2 to 8 carbon atoms (inclusive), and most preferably 2 to 6 carbon atoms (inclusive), which contains at least one carbon-carbon double bond, and is optionally substituted with one or more suitable substituents.

[0114] "Heteroalkyl," "heteroalkenyl," and "heteroalkynyl" refer to alkyl, alkenyl, and alkynyl radicals, respectively, as defined herein, in which one or more of the carbon atoms are each independently replaced by the same or different heteroatoms or heteroatom groups. Heteroatoms and / or heteroatom groups that can replace a carbon atom include, but are not limited to, -0-, -S-, -S-O-, -NR γ -, -PH-, -S(O)-, -S(O)2-, -S(O)NR γ -, -S(O)2NR γ -, and the like, including combinations thereof, wherein each R γ is independently selected from hydrogen, alkyl, cycloalkyl, heterocycloalkyl, aryl, and heteroaryl.

[0115] "Aryl" refers to an unsaturated aromatic carbocyclic radical of 6 to 12 carbon atoms (inclusive), having a single ring (e.g., phenyl) or multiple condensed rings (e.g., naphthyl or anthracyl). Exemplary aryl groups include phenyl, pyridyl, naphthyl, and the like.

[0116] "Arylalkyl" refers to an alkyl group (i.e., an "aryl-alkyl-" group) that is substituted with an aryl group, preferably having from 1 to 6 carbon atoms (inclusive) in the alkyl portion and from 6 to 12 carbon atoms (inclusive) in the aryl portion. Such arylalkyl groups are exemplified by phenyl, naphthyl, and the like.

[0117] "Aryloxy" refers to a -OR λ group, wherein R λ is an aryl group, which can be optionally substituted.

[0118] "Cycloalkyl" refers to a cyclic alkyl radical of 3 to 12 carbon atoms (inclusive), having a single ring or multiple condensed rings, which single ring or multiple condensed rings can be optionally substituted with 1 to 3 alkyl groups. Exemplary cycloalkyl groups include, but are not limited to, single ring structures such as cyclopropyl, cyclobutyl, cyclopentyl, cyclooctyl, 1-methylcyclopropyl, 2-methylcyclopentyl, 2-methylcyclooctyl, and the like, or multiple ring structures, including bridged ring systems, such as adamantyl and the like.

[0119] "Cycloalkylalkyl" refers to an alkyl group (i.e., a "cycloalkyl-alkyl-" group) substituted with a cycloalkyl group, preferably having from 1 to 6 carbon atoms, inclusive, in the alkyl portion and from 3 to 12 carbon atoms, inclusive, in the cycloalkyl portion. Such cycloalkylalkyl groups are exemplified by cyclopropylmethyl, cyclohexylethyl, and the like.

[0120] "Amino" refers to the group -NH2. Substituted amino refers to the group -NHR η , NR η R η , and NR η R η R η , where each R η is independently selected from substituted or unsubstituted alkyl, cycloalkyl, cycloheteroalkyl, alkoxy, aryl, heteroaryl, heteroarylalkyl, acyl, alkoxycarbonyl, sulfanyl, sulfinyl, sulfonyl, and the like. Typical amino groups include, but are not limited to, dimethylamino, diethylamino, trimethylammonium, triethylammonium, methylsulfonylamino, furanyl-oxy-sulfonamino, and the like.

[0121] "Aminoalkyl" refers to an alkyl group in which one or more hydrogen atoms are replaced by one or more amino groups, including substituted amino groups.

[0122] "Aminocarbonyl" refers to -C(O)NH2. Substituted aminocarbonyl refers to -C(O)NR η R η , where the amino group NR η R η is as defined herein.

[0123] "Oxy" refers to the divalent group -O- which can have various substituents to form different oxy groups, including ethers and esters.

[0124] "Alkoxy" or "alkyloxy" are used interchangeably herein to refer to the group -OR ξ , where R ξ is an alkyl group, including optionally substituted alkyl groups.

[0125] "Carboxy" refers to -COOH.

[0126] "Carbonyl" refers to -C(O)-, which can have a variety of substituents to form different carbonyl groups, including acids, acyl halides, aldehydes, amides, esters, and ketones.

[0127] "Carboxyalkyl" refers to an alkyl group in which one or more hydrogen atoms are replaced by one or more carboxy groups.

[0128] "Aminocarbonylalkyl" refers to an alkyl group that has been substituted with an aminocarbonyl group as defined herein.

[0129] "Halogen" or "halo" refers to fluorine, chlorine, bromine, and iodine.

[0130] "Halogenated alkyl" refers to an alkyl group in which one or more hydrogen atoms are replaced by a halogen. Therefore, the term "halogenated alkyl" is intended to include monohalogenated alkyl, dihalogenated alkyl, trihalogenated alkyl, and so on, up to perhalogenated alkyl. For example, the expression "(C1-C2)halogenated alkyl" includes 1-fluoromethyl, difluoromethyl, trifluoromethyl, 1-fluoroethyl, 1,1-difluoroethyl, 1,2-difluoroethyl, 1,1,1-trifluoroethyl, perfluoroethyl, etc.

[0131] "Hydroxy group" refers to -OH.

[0132] "Hydroxyalkyl" refers to an alkyl group in which one or more hydrogen atoms are replaced by one or more hydroxyl groups.

[0133] "Thioyl" or "thioalkyl" refers to -SH. Substituted thioyl or thioalkyl refers to –SR. η , where R η It can be alkyl, aryl, or other suitable substituents.

[0134] "alkyl thio" refers to –SR ξ , where R ξ It is an alkyl group, which may be optionally substituted. Typical alkyl thio groups include, but are not limited to, methyl thio, ethyl thio, n-propyl thio, etc.

[0135] "alkylthioalkyl" refers to an alkylthioyl group –SR ξ Substituted alkyl groups, wherein R ξ It is an alkyl group, which may be optionally substituted.

[0136] "Sulfonyl group" refers to -SO2-. A substituted sulfonyl group refers to –SO2-R. η , where R η It can be alkyl, aryl, or other suitable substituents.

[0137] "alkylsulfonyl" refers to –SO2-R ξ , where R ξ It is an alkyl group, which may be optionally substituted. Typical alkyl sulfonyl groups include, but are not limited to, methanesulfonyl, ethylsulfonyl, n-propylsulfonyl, etc.

[0138] "alkylsulfonylalkyl" refers to an alkylsulfonyl group –SO2-R ξ Substituted alkyl groups, wherein R ξ"Alkyl" means a straight or branched carbon chain radical. The alkyl group can be optionally substituted.

[0139] "Heteroaryl" means an aromatic heterocyclic radical having from 1 to 10 carbon atoms (inclusive) in the ring and from 1 to 4 heteroatoms (inclusive) selected from oxygen, nitrogen and sulfur. Such heteroaryl groups can have a single ring (e.g., pyridinyl or furanyl) or more than one condensed ring (e.g., indolizinyl or benzothienyl).

[0140] "Heteroarylalkyl" means an alkyl group substituted with a heteroaryl group (i.e., a "heteroaryl-alkyl-" group), preferably having from 1 to 6 carbon atoms (inclusive) in the alkyl portion and from 5 to 12 ring atoms (inclusive) in the heteroaryl portion. Such heteroarylalkyl groups are exemplified by pyridinylmethyl and the like.

[0141] "Heterocycle," "heterocyclic" and the interchangeable "heterocycloalkyl" mean a saturated or unsaturated radical having a single ring or more than one condensed ring, having from 2 to 10 carbon ring atoms (inclusive) in the ring and from 1 to 4 hetero ring atoms (inclusive) selected from nitrogen, sulfur or oxygen. Such heterocyclic groups can have a single ring (e.g., piperidinyl or tetrahydrofuranyl) or more than one condensed ring (e.g., dihydroindolyl, dihydrobenzofuranyl or quinuclidinyl). Examples of heterocycles include, but are not limited to, furan, thiophene, thiazole, oxazole, pyrrole, imidazole, pyrazole, pyridine, pyrazine, pyrimidine, pyridazine, indolizine, isoindole, indole, indazole, purine, quinolizine, isoquinoline, quinoline, phthalazine, naphthylpyridine, quinoxaline, quinazoline, cinnoline, pteridine, carbazole, carboline, phenanthridine, acridine, phenanthroline, isothiazole, phenazine, isoxazole, phenoxazole, phenothiazine, imidazolidine, imidazoline, piperidine, piperazine, pyrrolidine, indoline and the like.

[0142] "Heterocycloalkylalkyl" means an alkyl group substituted with a heterocycloalkyl group (i.e., a "heterocycloalkyl-alkyl-" group), preferably having from 1 to 6 carbon atoms (inclusive) in the alkyl portion and from 3 to 12 ring atoms (inclusive) in the heterocycloalkyl portion.

[0143] "Membered ring" means to include any cyclic structure. The number preceding the term "member" indicates the number of skeletal atoms comprising the ring. Thus, for example, cyclohexyl, pyridine, pyran and thiopyran are 6-membered rings, and cyclopentyl, pyrrole, furan and thiophene are 5-membered rings.

[0144] As used herein, "fused bicyclic" refers to both unsubstituted and substituted carbocyclic and / or heterocyclic ring moieties having from 5 to 8 atoms in each ring, the rings having 2 common atoms.

[0145] Unless otherwise indicated, positions occupied by hydrogen in the foregoing groups can be further substituted by, for example, but not limited to, hydroxyl, oxo, nitro, methoxy, ethoxy, alkoxy, substituted alkoxy, trifluoromethoxy, haloalkoxy, fluoro, chloro, bromo, iodo, halo, methyl, ethyl, propyl, butyl, alkyl, alkenyl, alkynyl, substituted alkyl, trifluoromethyl, haloalkyl, hydroxyalkyl, alkoxyalkyl, thio, alkylthio, acyl, carboxyl, alkoxycarbonyl, carboxamido, substituted carboxamido, alkylsulfonyl, alkylsulfinyl, alkylsulfonylamino, sulfonamido, substituted sulfonamido, cyano, amino, substituted amino, alkylamino, dialkylamino, aminoalkyl, acylamino, amidino, amidoximo, hydroxamoyl, phenyl, aryl, substituted aryl, aryloxy, arylalkyl, arylalkenyl, arylalkynyl, pyridyl, imidazolyl, heteroaryl, substituted heteroaryl, heteroaryloxy, heteroarylalkyl, heteroarylalkenyl, heteroarylalkynyl, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloalkyl, cycloalkenyl, cycloalkylalkyl, substituted cycloalkyl, cycloalkyloxy, pyrrolidinyl, piperidinyl, morpholino, heterocycle, (heterocycle)oxy, and (heterocycle)alkyl; and preferred heteroatoms are oxygen, nitrogen, and sulfur. It will be understood that where there are open valences on these substituents, they can be further substituted with alkyl, cycloalkyl, aryl, heteroaryl, and / or heterocyclyl groups, where there are these open valences on carbon, they can be further substituted with halogen and oxygen-, nitrogen-, or sulfur-bonded substituents, and where there is more than one such open valence, these groups can be linked to form a ring, either by direct formation of a bond or by formation of a bond with a new heteroatom (preferably oxygen, nitrogen, or sulfur). It will also be understood that the above substitutions can be made provided that replacing a hydrogen with a substituent does not introduce unacceptable instability into the molecules of the invention, and is otherwise chemically reasonable.

[0146] "Optional" or "optionally" means that the subsequently described event or circumstance can or can not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not. One of ordinary skill in the art will understand that, for any molecule described as containing one or more optional substituents, only compounds that are synthetically feasible and / or achievable in space are intended to be included. "Optionally substituted" refers to all subsequent modifiers in the term or series of chemical groups. For example, in the term "optionally substituted arylalkyl," the "alkyl" portion and the "aryl" portion of the molecule can or can not be substituted, and for the series "optionally substituted alkyl, cycloalkyl, aryl, and heteroaryl," the alkyl group, the cycloalkyl group, the aryl group, and the heteroaryl group can each independently be substituted or unsubstituted.

[0147] Engineered proline hydroxylase polypeptides

[0148] The present invention provides polypeptides having proline hydroxylase activity, polynucleotides encoding the polypeptides, methods of making the polypeptides, and methods for using the polypeptides. Where the description refers to a polypeptide, it is understood that it can describe a polynucleotide encoding the polypeptide.

[0149] Proline hydroxylases belong to the class of dioxygenases, which catalyze the hydroxylation of proline in the presence of a-ketoglutarate and oxygen (O2). a-Ketoglutarate is decarboxylated stoichiometrically during hydroxylation, and one atom of the O2 molecule is incorporated into succinate, and the other into the hydroxyl group formed on the proline residue. As mentioned above, proline hydroxylases are distinguished from prolyl hydroxylases by their ability to hydroxylate free proline.

[0150] Based on the major diastereomeric products formed in the enzymatic reaction, several types of proline hydroxylases have been identified: cis-3-proline hydroxylases (cis-P3H), cis-4-proline hydroxylases (cis-P4H), trans-3-proline hydroxylases (trans-P3H), and trans-4-proline hydroxylases (trans-P4H). Cis-P3H enzymes have been identified in Streptomyces sp. TH1, Streptomyces canus, and Bacillus sp. TH2 and TH3 (Mori et al., Appl. Environ. Microbiol., 62(6): 1903-1907

[1996] ). Cis-P4H enzymes have been identified in Lotus corniculatus rhizobia, Mesorhizobium loti, Sinorhizobium meliloti, and Medicago sativa rhizobia (Hara and Kino, Biochem. Biophys. Res. Commun., 379(4):882-6

[2009] ; U.S. Patent Application Publication No. 2011 / 0091942). Trans-P4H has been identified in Dactylosporangium sp., Amycolatopsis sp., Streptomyces griseoviridus, Streptomyces sp., Glarea lozoyensis, and Emericella rugulosa NRRL 11440 (Shibasaki et al., Appl. Environ. Microbiol., 65(9):4028-31

[1999] ; Petersen et al., Appl. Microbiol. Biotechnol., 62(2-3):263-7

[2003] ; Mori et al., Appl. Environ. Microbiol., 62:1903-1907

[1996] ; Lawrence et al., Biochem. J., 313:185-191

[1996] ; and EP 0641862; Cacho et al., J. Am. Chem. Soc. 2012, 134, 16781).

[0151] Recently, a gene cluster comprising three hydroxylase genes was identified in the fungus sp. No. 11243 (Matsui et al., J. Biosci. Bioeng. 2017, February; 123(2): 147-153). Subsequently, one of these genes was identified as a proline hydroxylase and characterized as a trans-selective proline hydroxylase. The proline hydroxylase from the fungus sp. No. 11243, referred to as ANO11243 or ANO, converts free proline to both trans-4-hydroxyproline and trans-3-hydroxyproline, and slightly enriches the trans-3-hydroxyproline isomer. However, the naturally occurring ANO11243 proline hydroxylase lacks properties that make it useful in large-scale industrial processes, including low specific activity, low thermal stability, and low selectivity for the desired trans-3-hydroxyproline isomer.

[0152] Described herein are engineered proline hydroxylases that overcome the deficiencies of the wild-type proline hydroxylase from the fungus sp. No. 11243. Engineered proline hydroxylase polypeptides derived from the wild-type enzyme ANO of the fungus sp. No. 11243 are capable of efficiently converting L-proline to trans-3-hydroxyproline. The present invention identifies amino acid residue positions in the proline hydroxylase polypeptide sequence and corresponding mutations that improve enzyme properties compared to the naturally occurring enzyme, including, among others, activity, stability, expression, regioselectivity, and stereoselectivity. In particular, the present invention provides engineered polypeptides that are capable of efficiently converting L-proline to trans-3-hydroxyproline in the presence of a co-substrate (e.g., a-ketoglutarate) under appropriate reaction conditions (e.g., in the presence of oxygen and Fe(ll)), as shown in Scheme 1 above.

[0153] In some embodiments, the engineered proline hydroxylase polypeptide exhibits increased activity in the hydroxylation of L-proline to trans-3-hydroxyproline in a defined time with the same amount of enzyme compared to the polypeptide of SEQ ID NO: 4. In some embodiments, the engineered proline hydroxylase polypeptide has at least about 1.2-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold or more activity under suitable reaction conditions compared to the polypeptide represented by SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630.

[0154] In some embodiments, the engineered proline hydroxylase polypeptides have increased regioselectivity compared to wild-type proline hydroxylases. Specifically, naturally occurring enzymes primarily convert proline to, if not exclusively, trans-3-hydroxyproline. In some embodiments, the engineered proline hydroxylase polypeptides herein are capable of selectively forming trans-3-hydroxyproline over trans-4-hydroxyproline. In some embodiments, the engineered polypeptides are capable of selectively forming trans-3-hydroxyproline over trans-4-hydroxyproline, wherein the ratio of trans-3-hydroxyproline to trans-4-hydroxyproline formed under suitable reaction conditions is at least 1.5, 2, 3, 4, 5, 10, 15, 20, 25, 30, or more.

[0155] In some embodiments, the engineered proline hydroxylase polypeptides are capable of converting L-proline to trans-3-hydroxyproline at a percent conversion of at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% at a substrate loading concentration of at least about 10 g / L, about 20 g / L, about 30 g / L, about 40 g / L, about 50 g / L, about 70 g / L, about 100 g / L, about 125 g / L, about 150 g / L, about 175 g / L, or about 200 g / L or more, in about 120 h or less, 72 h or less, about 48 h or less, about 36 h or less, or about 24 h or less reaction time under suitable reaction conditions.

[0156] Suitable reaction conditions for the above improved properties of the engineered polypeptides to perform hydroxylation reactions can be determined depending on the polypeptide, substrate, co-substrate, concentration or amount of transition metal co-factor, reducing agent, buffer, cosolvent, pH, conditions including temperature and reaction time, and / or conditions under which the polypeptide is immobilized on a solid support, as further described below and in the Examples.

[0157] In some embodiments, exemplary engineered polypeptides having improved properties, particularly proline hydroxylase activity to convert L-proline to trans-3-hydroxyproline, comprise an amino acid sequence having one or more residue differences compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 at the residue positions indicated in Table 4.1, Table 4.2, Table 4.3, Table 4.4, Table 5.1, Table 5.2, Table 5.3, Table 6.1, Table 7.1, Table 7.2, Table 7.3, Table 8.1, Table 8.2, Table 9.1, Table 9.2, Table 10.1, Table 10.2, Table 11.1, Table 11.2, and / or Table 12.1.

[0158] Structural and functional information for exemplary non-naturally occurring (or engineered) proline hydroxylase polypeptides of the present invention is based on the conversion of L-proline to trans-3-hydroxyproline, the results of which are shown in Tables 4.1, 4.2, 4.3, 4.4, 5.1, 5.2, 5.3, 6.1, 7.1, 7.2, 7.3, 8.1, 8.2, 9.1, 9.2, 10.1, 10.2, 11.1, 11.2, and / or 12.1 below. Odd-numbered sequence identifiers (i.e., SEQ ID NO) refer to nucleotide sequences encoding the amino acid sequences provided by even-numbered SEQ ID NOs. Exemplary sequences are provided in the electronic sequence listing file accompanying the present invention, which is hereby incorporated by reference herein. Amino acid residue differences are based on comparison to reference sequences SEQ ID NO:4, 116, 162, 322, 412, 492, 562, 598, and / or 630. The naturally occurring amino acid sequence of proline hydroxylase ANO from the fungus sp. No. 11243 is provided herein as SEQ ID NO:2 (the corresponding polynucleotide sequence is SEQ ID NO:1, as provided herein). Activity of each engineered polypeptide relative to the reference polypeptide SEQ ID NO:4, 116, 162, 322, 412, 492, 562, 598, and / or 630 was determined by conversion of substrate as described in the Examples herein. In some embodiments, shake flask powder (SFP) or downstream processing (DSP) powder assays were used as secondary screens to assess properties of engineered proline hydroxylases, the results of which are provided in Tables 4.1, 4.2, 4.3, 4.4, 5.1, 5.2, 5.3, 6.1, 7.1, 7.2, 7.3, 8.1, 8.2, 9.1, 9.2, 10.1, 10.2, 11.1, 11.2, and / or 12.1. SFP format provides a more purified powder preparation of the engineered polypeptide, and can contain up to about 30% of the engineered polypeptide of total protein. Because DSP preparations can contain up to about 80% of the engineered proline hydroxylase of total protein, the preparation can provide an even more purified form of the engineered polypeptide.

[0159] In some embodiments, the specific enzyme properties associated with residue differences at the residue positions shown herein compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 include, among others, enzyme activity, substrate tolerance, thermostability, regioselectivity, and stereoselectivity. Improvements in enzyme activity are associated with residue differences at the residue positions shown in the Examples herein. Improvements in selectivity are associated with residue differences at the residue positions shown in the Examples herein. Improvements in thermostability are associated with residue differences at the residue positions shown in the Examples herein. Thus, residue differences at these residue positions can be used, alone or in various combinations, to produce engineered proline hydroxylase polypeptides with desired improved properties, including, among others, enzyme activity, substrate tolerance, regioselectivity, stereoselectivity, and thermostability. Other residue differences that affect polypeptide expression can be used to increase expression of engineered proline hydroxylases.

[0160] According to the guidance provided herein, it is also contemplated that any of the exemplary engineered polypeptides comprising the even-numbered sequences of SEQ ID NO: 4-658 can be used as a starting amino acid sequence for the synthesis of other engineered proline hydroxylase polypeptides, for example, through subsequent rounds of evolution incorporating new combinations of various amino acids from other polypeptides in Table 4.1, Table 4.2, Table 4.3, Table 4.4, Table 5.1, Table 5.2, Table 5.3, Table 6.1, Table 7.1, Table 7.2, Table 7.3, Table 8.1, Table 8.2, Table 9.1, Table 9.2, Table 10.1, Table 10.2, Table 11.1, Table 11.2, and / or Table 12.1, and other residue positions described herein. Additional improvements can be generated by including amino acid differences at residue positions that remain unchanged throughout earlier rounds of evolution.

[0161] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4 and one or more residue differences compared to SEQ ID NO: 4 at a residue position selected from 21, 28, 58 / 247, 65, 80, 85, 95, 98, 117, 120, 159, 185, 194, 199, 200, 233, 237, 243, 250, 268, 281, 282, 287, 289, 307, 324, 326, 327, 330, 338, 343, 346, and 348. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4 and one or more residue differences selected from 21Q, 28A, 58V / 247V, 65A, 80H, 85L, 95P, 95R, 98L, 117E, 117L, 117R, 117S, 117T, 120F, 159G, 185D, 194L, 194T, 199A, 200V, 233A, 233R, 237E, 243A, 243V, 250Q, 268H, 281S, 282E, 282S, 287E, 289D, 307I, 324D, 326G, 326H, 326K, 327Q, 330G, 338I, 343N, 343P, 346S, and 348S (relative to SEQ ID NO: 4).In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4 and one or more residue differences selected from R21Q, P28A, E58V / P247V, S65A, K80H, E85L, G95P, G95R, Q98L, A117E, A117L, A117R, A117S, A117T, L120F, Q159G, A185D, N194L, N194T, T199A, P200V, V233A, V233R, Q237E, L243A, L243V, V250Q, R268H, R281S, L282E, L282S, D287E, M289D, V307I, A324D, R326G, R326H, R326K, W327Q, L330G, M338I, V343N, V343P, A346S, and Q348S (relative to SEQ ID NO: 4).

[0162] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4, and one or more residue differences at residue positions selected from: 21, 28, 45, 65, 95, 112, 117, 139, 177, 185, 199, 233, 243, 250, 281, 282, 287, 289, 307, 324, 326, 327, 335, 338, 343, and 346. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4, and one or more residue differences selected from: 21Q, 28A, 45S, 65A, 95R, 112L, 117S, 139F, 177P, 185D, 199A, 233A, 243V, 250Q, 250T, 281S, 281T, 282E, 282S, 287E, 289D, 307I, 324D, 326G, 326H, 326K, 327Q, 335A, 335M, 338I, 343N, 343P, and 346S (relative to SEQ ID NO: 4).In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4 and one or more residue differences selected from R21Q, P28A, Y45S, S65A, G95R, R112L, A117S, M139F, S177P, A185D, T199A, V233A, L243V, V250Q, V250T, R281S, R281T, L282E, L282S, D287E, M289D, V307I, A324D, R326G, R326H, R326K, W327Q, S335A, S335M, M338I, V343N, V343P, and A346S (relative to SEQ ID NO: 4).

[0163] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4, and one or more residue differences at residue positions selected from: 48 / 66 / 189 / 194, 48 / 66 / 194, and 66 / 82 / 85 / 135 / 189 / 194 / 267. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4, and one or more residue differences selected from: 48V / 66W / 189N / 194L, 48V / 66W / 194L, and 66W / 82P / 85P / 135P / 189N / 194L / 267D (relative to SEQ ID NO: 4). In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4, and one or more residue differences selected from: A48V / Y66W / A189N / N194L, A48V / Y66W / N194L, and Y66W / K82P / E85P / A135P / A189N / N194L / G267D (relative to SEQ ID NO: 4).

[0164] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4 and one or more residue differences at residue positions selected from 20 / 56 / 76 / 168 / 169 / 296, 20 / 56 / 232 / 294, 20 / 119 / 294 / 296, 56 / 76 / 119 / 124 / 147 / 232, 56 / 76 / 294, 76 / 168 / 232 / 294, 76 / 294 / 296, 76 / 296, 147, and 232. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4 and one or more residue differences selected from 20F / 56P / 76E / 168A / 169L / 296I, 20F / 56P / 232E / 294Y, 20F / 119D / 294Y / 296I, 56P / 76E / 119D / 124F / 147F / 232E, 56P / 76E / 294Y, 76E / 168A / 232E / 294Y, 76E / 294Y / 296I, 76E / 296I, 147F, and 232E (relative to SEQ ID NO: 4).In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 4 and one or more residue differences selected from Y20F / S56P / H76E / C168A / I169L / L296I, Y20F / S56P / Q232E / H294Y, Y20F / E119D / H294Y / L296I, S56P / H76E / E119D / W124F / Y147F / Q232E, S56P / H76E / H294Y, H76E / C168A / Q232E / H294Y, H76E / H294Y / L296I, H76E / L296I, Y147F, and Q232E (relative to SEQ ID NO: 4).

[0165] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 116 and one or more residue differences compared to SEQ ID NO: 116 at a residue position selected from 123, 189, 195, 233, and 296. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 116 and one or more residue differences selected from 123T, 189A, 189S, 195Y, 233A, 233M, and 296V (relative to SEQ ID NO: 116). In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 116 and one or more residue differences selected from S123T, N189A, N189S, H195Y, V233A, V233M, and L296V (relative to SEQ ID NO: 116).

[0166] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 116 and one or more residue differences at residue positions selected from: 20 / 21 / 56, 20 / 21 / 56 / 76 / 95 / 232 / 294 / 307 / 335, 20 / 21 / 56 / 76 / 147 / 225 / 232 / 233 / 281 / 294 / 296 / 307 / 335, 20 / 21 / 56 / 95 / 147 / 281 / 294 / 307, 20 / 21 / 56 / 281 / 307, 20 / 21 / 76 / 232 / 243, 20 / 21 / 95 / 232 / 307, 20 / 21 / 95 / 281 / 294 / 296, 20 / 21 / 147 / 189 / 233 / 243 / 281 / 307, 20 / 56, 20 / 56 / 76 / 95 / 281 / 307, 20 / 56 / 76 / 147 / 294 / 296 / 307, 20 / 56 / 95 / 147 / 294, 20 / 56 / 281, 20 / 76, 20 / 76 / 95 / 281 / 294 / 296, 20 / 76 / 95 / 281 / 296 / 307, 20 / 76 / 233 / 294 / 307, 20 / 76 / 243 / 281 / 294, 21 / 76 / 147 / 233 / 294 / 307, 21 / 76 / 147 / 243 / 296 / 307 / 335, 21 / 95 / 185 / 189 / 232 / 281 / 296, 21 / 95 / 233 / 243 / 281 / 296, 21 / 95 / 294 / 296 / 307 / 335, 21 / 95 / 307, 21 / 281 / 307, 29 / 76 / 281, 56 / 76 / 95 / 232 / 243 / 281, 56 / 76 / 147 / 281 / 307, 56 / 76 / 243 / 294, 56 / 76 / 281 / 294, 56 / 76 / 296, 56 / 76 / 307, 56 / 95 / 147 / 307 / 335 / 348, 56 / 95 / 232 / 233 / 281 / 294 / 307, 56 / 95 / 243 / 281, 56 / 147 / 281, 56 / 232 / 243 / 281, 56 / 232 / 281, 56 / 232 / 281 / 294 / 296, 56 / 233 / 281 / 294 / 296, 56 / 281 / 307, 76 / 95 / 232 / 243 / 281 / 307,76 / 95 / 243 / 281 / 307 / 335, 76 / 95 / 294 / 307, 76 / 147, 76 / 147 / 233 / 243 / 294, 76 / 147 / 233 / 281 / 294 / 307, 76 / 147 / 243 / 294 / 296 / 307 / 335, 76 / 147 / 281 / 307, 76 / 189 / 296, 76 / 232 / 233 / 243 / 294 / 296 / 307, 76 / 281, 76 / 281 / 294, 76 / 294 / 296, 95 / 120, 95 / 147 / 335, 95 / 232 / 243 / 281 / 294 / 307, 95 / 232 / 281 / 294 / 296, 95 / 281 / 294 / 296, 95 / 335, 147, 147 / 225 / 232 / 243 / 281 / 296 / 307 / 335, 147 / 233 / 243 / 281 / 307, 147 / 233 / 281 / 307 / 335, 147 / 243 / 281, 147 / 307, 232 / 233 / 281 / 294 / 296 / 307, 232 / 281, 232 / 284 / 307, 233 / 243 / 281 / 296 / 307 / 335, 233 / 281 / 296 / 307, 243 / 281 / 294 / 296, 281, 281 / 294, 281 / 307, 307, and 335, 117, 139, 177, 185, 199, 233, 243, 250, 281, 282, 287, 289, 307, 324, 326, 327, 335, 338, 343, and 346. In some embodiments, the engineered polypeptide having proline hydroxylase activity having one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 116 and one or more residue differences selected from: 20F / 21Q / 56P, 20F / 21Q / 56P / 76E / 95R / 232E / 294Y / 307I / 335M, 20F / 21Q / 56P / 76E / 147F / 225R / 232E / 233A / 281S / 294Y / 296I / 307L / 335M, 20F / 21Q / 56P / 95R / 147F / 281T / 294Y / 307I, 20F / 21Q / 56P / 281T / 307L, 20F / 21Q / 76E / 232E / 243V, 20F / 21Q / 95R / 232E / 307I,20F / 21Q / 95R / 281T / 294Y / 296I, 20F / 21Q / 147F / 189A / 233R / 243V / 281T / 307I, 20F / 56P, 20F / 56P / 76E / 95R / 281S / 307I, 20F / 56P / 76E / 147F / 294Y / 296I / 307L, 20F / 56P / 95P / 147F / 294Y, 20F / 56P / 281S, 20F / 76E, 20F / 76E / 95R / 281S / 294Y / 296I, 20F / 76E / 95R / 281T / 296I / 307I, 20F / 76E / 233A / 294Y / 307I, 20F / 76E / 243V / 281T / 294Y, 21Q / 76E / 147F / 233R / 294Y / 307I, 21Q / 76E / 147F / 243V / 296I / 307I / 335M, 21Q / 95R / 185L / 189A / 232E / 281T / 296I, 21Q / 95R / 233A / 243V / 281T / 296I, 21Q / 95R / 294Y / 296I / 307I / 335M, 21Q / 95R / 307I, 21Q / 281T / 307L, 29T / 76E / 281T, 56P / 76E / 95R / 232E / 243V / 281T, 56P / 76E / 147F / 281T / 307I, 56P / 76E / 243V / 294Y, 56P / 76E / 281T / 294Y, 56P / 76E / 296I, 56P / 76E / 307I, 56P / 95P / 147F / 307I / 335M / 348K, 56P / 95R / 232E / 233R / 281S / 294Y / 307L, 56P / 95R / 243V / 281T, 56P / 147F / 281T, 56P / 232E / 243V / 281S, 56P / 232E / 281S, 56P / 232E / 281S / 294Y / 296I, 56P / 233R / 281S / 294Y / 296I, 56P / 281T / 307I, 76E / 95P / 232E / 243V / 281S / 307L, 76E / 95R / 243V / 281S / 307I / 335M, 76E / 95R / 294Y / 307L, 76E / 147F, 76E / 147F / 233A / 243V / 294Y, 76E / 147F / 233R / 281T / 294Y / 307L, 76E / 147F / 243V / 294Y / 296I / 307L / 335M, 76E / 147F / 281S / 307L, 76E / 189A / 296I, 76E / 232E / 233R / 243V / 294Y / 296I / 307I,76E / 281S, 76E / 281T / 294Y, 76E / 294Y / 296I, 95P / 232E / 281T / 294Y / 296I, 95P / 335M, 95R / 120P, 95R / 147F / 335M, 95R / 232E / 243V / 281T / 294Y / 307I, 95R / 281T / 294Y / 296I, 95R / 335M, 147F, 147F / 225R / 232E / 243V / 281S / 296I / 307L / 335M, 147F / 233A / 243V / 281S / 307L, 147F / 233R / 281T / 307L / 335M, 147F / 243V / 281S, 147F / 307I, 232E / 233A / 281T / 294Y / 296I / 307I, 232E / 281T, 232E / 284R / 307I, 233A / 243V / 281S / 296I / 307I / 335M, 233A / 281T / 296I / 307I, 243V / 281S / 294Y / 296I, 281T, 281T / 294Y, 281T / 307I, 281T / 307L, 307I, and 335M (relative to SEQ ID NO: 116). In some embodiments, the engineered polypeptide having proline hydroxylase activity having one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 116 and one or more residue differences selected from Y20F / R21Q / S56P, Y20F / R21Q / S56P / H76E / G95R / Q232E / H294Y / V307I / S335M, Y20F / R21Q / S56P / H76E / Y147F / Q225R / Q232E / V233A / R281S / H294Y / L296I / V307L / S335M, Y20F / R21Q / S56P / G95R / Y147F / R281T / H294Y / V307I, Y20F / R21Q / S56P / R281T / V307L, Y20F / R21Q / H76E / Q232E / L243V, Y20F / R21Q / G95R / Q232E / V307I, Y20F / R21Q / G95R / R281T / H294Y / L296I, Y20F / R21Q / Y147F / N189A / V233R / L243V / R281T / V307I,Y20F / S56P, Y20F / S56P / H76E / G95R / R281S / V307I, Y20F / S56P / H76E / Y147F / H294Y / L296I / V307L, Y20F / S56P / G95P / Y147F / H294Y, Y20F / S56P / R281S, Y20F / H76E, Y20F / H76E / G95R / R281S / H294Y / L296I, Y20F / H76E / G95R / R281T / L296I / V307I, Y20F / H76E / V233A / H294Y / V307I, Y20F / H76E / L243V / R281T / H294Y, R21Q / H76E / Y147F / V233R / H294Y / V307I, R21Q / H76E / Y147F / L243V / L296I / V307I / S335M, R21Q / G95R / A185L / N189A / Q232E / R281T / L296I, R21Q / G95R / V233A / L243V / R281T / L296I, R21Q / G95R / H294Y / L296I / V307I / S335M, R21Q / G95R / V307I, R21Q / R281T / V307L, A29T / H76E / R281T, S56P / H76E / G95R / Q232E / L243V / R281T, S56P / H76E / Y147F / R281T / V307I, S56P / H76E / L243V / H294Y, S56P / H76E / R281T / H294Y, S56P / H76E / L296I, S56P / H76E / V307I, S56P / G95P / Y147F / V307I / S335M / Q348K, S56P / G95R / Q232E / V233R / R281S / H294Y / V307L, S56P / G95R / L243V / R281T, S56P / Y147F / R281T, S56P / Q232E / L243V / R281S, S56P / Q232E / R281S, S56P / Q232E / R281S / H294Y / L296I, S56P / V233R / R281S / H294Y / L296I, S56P / R281T / V307I, H76E / G95P / Q232E / L243V / R281S / V307L, H76E / G95R / L243V / R281S / V307I / S335M, H76E / G95R / H294Y / V307L, H76E / Y147F, H76E / Y147F / V233A / L243V / H294Y,H76E / Y147F / V233R / R281T / H294Y / V307L, H76E / Y147F / L243V / H294Y / L296I / V307L / S335M, H76E / Y147F / R281S / V307L, H76E / N189A / L296I, H76E / Q232E / V233R / L243V / H294Y / L296I / V307I, H76E / R281S, H76E / R281T / H294Y, H76E / H294Y / L296I, G95P / Q232E / R281T / H294Y / L296I, G95P / S335M, G95R / L120P, G95R / Y147F / S335M, G95R / Q232E / L243V / R281T / H294Y / V307I, G95R / R281T / H294Y / L296I, G95R / S335M, Y147F, Y147F / Q225R / Q232E / L243V / R281S / L296I / V307L / S335M, Y147F / V233A / L243V / R281S / V307L, Y147F / V233R / R281T / V307L / S335M, Y147F / L243V / R281S, Y147F / V307I, Q232E / V233A / R281T / H294Y / L296I / V307I, Q232E / R281T, Q232E / G284R / V307I, V233A / L243V / R281S / L296I / V307I / S335M, V233A / R281T / L296I / V307I, L243V / R281S / H294Y / L296I, R281T, R281T / H294Y, R281T / V307I, R281T / V307L, V307I, and S335M (relative to SEQ ID NO: 116).

[0167] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 116 and one or more residue differences at residue positions selected from: 21 / 76 / 147 / 243 / 296 / 307 / 335, 56 / 76 / 147 / 281 / 307, and 95 / 147 / 335. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 116 and one or more residue differences selected from: 21Q / 76E / 147F / 243V / 296I / 307I / 335M, 56P / 76E / 147F / 281T / 307I, and 95R / 147F / 335M (relative to SEQ ID NO: 116). In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 116 and one or more residue differences selected from: R21Q / H76E / Y147F / L243V / L296I / V307I / S335M, S56P / H76E / Y147F / R281T / V307I, and G95R / Y147F / S335M (relative to SEQ ID NO: 116).

[0168] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 162 and one or more residue differences as compared to SEQ ID NO: 162 at a residue position selected from 2 / 85 / 123 / 237, 28 / 115 / 117 / 120 / 123 / 268 / 270 / 343 / 346 / 348, 45 / 123 / 326, 65 / 117 / 120 / 123 / 343 / 346, 85 / 123 / 281 / 282, 114 / 115 / 117 / 120 / 123 / 268 / 271 / 313 / 326 / 343 / 346, 123 / 139 / 233 / 237 / 281 / 282 / 289 / 324 / 326, and 123 / 199 / 200 / 247 / 250 / 338. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 162 and one or more residue differences selected from 2L / 85L / 123T / 237E, 28A / 115T / 117V / 120I / 123T / 268T / 270L / 343N / 346S / 348S, 45S / 123T / 326G, 65R / 117V / 120I / 123T / 343N / 346G, 85L / 123T / 281T / 282S, 114G / 115T / 117T / 120P / 123T / 268T / 271A / 313F / 326G / 343N / 346S, 123T / 139F / 233A / 237E / 281M / 282S / 289D / 324Q / 326G, and 123T / 199A / 200V / 247L / 250Q / 338I (relative to SEQ ID NO: 162).In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 162 and one or more residue differences selected from G2L / E85L / S123T / Q237E, P28A / V115T / A117V / L120I / S123T / R268T / R270L / V343N / A346S / Q348S, Y45S / S123T / R326G, S65R / A117V / L120I / S123T / V343N / A346G, E85L / S123T / R281T / L282S, E114G / V115T / A117T / L120P / S123T / R268T / S271A / L313F / R326G / V343N / A346S, S123T / M139F / V233A / Q237E / R281M / L282S / M289D / A324Q / R326G, and S123T / T199A / P200V / P247L / V250Q / M338I (relative to SEQ ID NO: 162).

[0169] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 322 and one or more residue differences compared to SEQ ID NO: 322 at a residue position selected from 26, 54, 61, 129, 132, 149, 156, 175, 189, 201, 209, 228, 236, 248, 262, 272, 277, 291, and 345. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 322 and one or more residue differences selected from 26N, 54P, 61H, 129I, 132P, 149G, 156S, 175S, 175V, 189S, 201C, 201G, 201T, 209S, 228T, 236T, 248R, 262V, 272S, 277A, 291G, and 345R (relative to SEQ ID NO: 322). In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 322 and one or more residue differences selected from G26N, G54P, D61H, A129I, E132P, S149G, V156S, L175S, L175V, N189S, A201C, A201G, A201T, C209S, V228T, Q236T, D248R, S262V, V272S, V277A, P291G, and T345R (relative to SEQ ID NO: 322).

[0170] In some embodiments, the engineered polypeptide having one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 322 and one or more residue differences at a residue position selected from 25, 43, 54, 58, 61, 79, 129, 132, 143, 156, 163, 175, 179, 201, 209, 236, 248, 278, 291, 345, and 347. In some embodiments, the engineered polypeptide having proline hydroxylase activity having one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 322 and one or more residue differences selected from: 25K, 43T, 54P, 54S, 58T, 61H, 79T, 129I, 132N, 143L, 156D, 156S, 163L, 175V, 179L, 201C, 209S, 236T, 248R, 278N, 291G, 345R, and 347E (relative to SEQ ID NO: 322). In some embodiments, the engineered polypeptide having proline hydroxylase activity having one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 322 and one or more residue differences selected from: H25K, A43T, G54P, G54S, E58T, D61H, Q79T, A129I, E132N, D143L, V156D, V156S, Q163L, L175V, E179L, A201C, C209S, Q236T, D248R, S278N, P291G, T345R, and A347E (relative to SEQ ID NO: 322).

[0171] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 322 and one or more residue differences at residue positions selected from 85 / 117 / 120 / 135 / 208 / 270 / 324 / 343 / 346, 85 / 117 / 120 / 135 / 208 / 281 / 282 / 289, 85 / 117 / 120 / 270 / 281 / 289, 85 / 117 / 135 / 139 / 208, and 117 / 120 / 208 / 270 / 324 / 343 / 346. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 322 and one or more residue differences selected from 85L / 117T / 120P / 135S / 208E / 281R / 282L / 289M, 85L / 117T / 135S / 139M / 208E, 85L / 117V / 120I / 135S / 208E / 270L / 324A / 343N / 346G, 85L / 117V / 120P / 270L / 281R / 289M, and 117T / 120I / 208E / 270L / 324A / 343N / 346G (relative to SEQ ID NO: 322).In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 322 and one or more residue differences selected from: E85L / A117T / L120P / A135S / A208E / M281R / S282L / D289M, E85L / A117T / A135S / F139M / A208E, E85L / A117V / L120I / A135S / A208E / R270L / Q324A / V343N / A346G, E85L / A117V / L120P / R270L / M281R / D289M, and A117T / L120I / A208E / R270L / Q324A / V343N / A346G (relative to SEQ ID NO: 322).

[0172] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 412 and one or more residue differences at residue positions selected from 47, 48, 56 / 118, 85, 95, 95 / 289, 113, 118, 118 / 247, 154, 162, 162 / 204, 164, 164 / 198 / 271, 168, 169, 187, 195, 243, 271, 275, 281, 314, 330, and 342. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 412 and one or more residue differences selected from 47M, 48G, 56P / 118W, 85P, 95A / 289V, 95W, 113H, 113N, 113P, 113R, 118D, 118P / 247A, 118V, 118W, 154L, 162A, 162L, 162M, 162V, 162V / 204S, 164D / 198V / 271V, 164T, 168V, 169C, 169T, 169V, 187P, 195Y, 243Y, 271V, 275K, 281L, 314A, 314S, 314T, 330G, 330H, and 342R (relative to SEQ ID NO: 412).In some embodiments, the engineered polypeptide having proline hydroxylase activity having one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 412 and one or more residue differences selected from F47M, V48G, S56P / A118W, L85P, G95A / M289V, G95W, S113H, S113N, S113P, S113R, A118D, A118P / P247A, A118V, A118W, F154L, H162A, H162L, H162M, H162V, H162V / L204S, S164D / A198V / S271V, S164T, C168V, I169C, I169T, I169V, C187P, H195Y, V243Y, S271V, R275K, R281L, F314A, F314S, F314T, L330G, L330H, and N342R (relative to SEQ ID NO: 412).

[0173] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 412 and one or more residue differences compared to SEQ ID NO: 412 at a residue position selected from 25 / 129 / 163 / 236 / 262 / 345 / 347, 120 / 156 / 175 / 179 / 201, 129 / 189 / 236 / 262 / 277 / 278, 129 / 236 / 262, 156 / 175 / 179 / 228, and 162. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 412 and one or more residue differences selected from 25K / 129I / 163L / 236T / 262V / 345R / 347E, 120V / 156S / 175V / 179L / 201G, 129I / 189S / 236T / 262V / 277A / 278N, 129I / 236T / 262V, 156S / 175V / 179L / 228A, 162L, and 162V (relative to SEQ ID NO: 412).In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 412 and one or more residue differences selected from: H25K / A129I / Q163L / Q236T / S262V / T345R / A347E, P120V / V156S / L175V / E179L / A201G, A129I / N189S / Q236T / S262V / V277A / S278N, A129I / Q236T / S262V, V156S / L175V / E179L / V228A, H162L, and H162V (relative to SEQ ID NO: 412).

[0174] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 492 and one or more residue differences at residue positions selected from: 15, 17, 28, 29, 65, 135, 167, 177, 199, 208, 228, 235, 287, 294, 307, and 343. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 492 and one or more residue differences selected from: 15V, 17C, 28I, 29S, 65V, 135G, 135N, 135T, 167G, 177A, 177L, 177P, 199C, 208L, 208M, 208S, 228T, 235E, 287E, 294T, 307L, 343S, and 343T (relative to SEQ ID NO: 492). In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 492 and one or more residue differences selected from: I15V, S17C, P28I, A29S, S65V, S135G, S135N, S135T, Q167G, S177A, S177L, S177P, T199C, E208L, E208M, E208S, V228T, D235E, D287E, H294T, I307L, V343S, and V343T (relative to SEQ ID NO: 492).

[0175] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 492 and one or more residue differences compared to SEQ ID NO: 492 at a residue position selected from 85 / 187 / 281 / 347, 85 / 187 / 347, 118 / 120 / 162 / 175 / 179 / 330, 118 / 120 / 162 / 175 / 330, 162 / 175 / 179 / 330, 175 / 228 / 330, 195 / 347, and 278 / 314 / 347. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 492 and one or more residue differences selected from 85P / 187P / 281L / 347E, 85P / 187P / 347E, 118V / 120V / 162V / 175V / 179L / 330H, 118V / 120V / 162V / 175V / 330H, 162V / 175V / 179L / 330H, 175V / 228A / 330H, 195Y / 347E, and 278S / 314A / 347E (relative to SEQ ID NO: 492).In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 492 and one or more residue differences selected from: L85P / C187P / R281L / A347E, L85P / C187P / A347E, A118V / P120V / H162V / L175V / E179L / L330H, A118V / P120V / H162V / L175V / L330H, H162V / L175V / E179L / L330H, L175V / V228A / L330H, H195Y / A347E, and N278S / F314A / A347E (relative to SEQ ID NO: 492).

[0176] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 562 and one or more residue differences at residue positions selected from: 15, 40, 43, 44, 59, 79, 82, 149, 164, 179, 345, and 347. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 562 and one or more residue differences selected from: 15F, 40A, 43S, 44R, 44V, 59L, 79E, 82A, 149N, 164Q, 179T, 345D, and 347K (relative to SEQ ID NO: 562). In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 562 and one or more residue differences selected from: I15F, K40A, A43S, G44R, G44V, R59L, Q79E, K82A, S149N, S164Q, L179T, T345D, and A347K (relative to SEQ ID NO: 562).

[0177] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 562 and one or more residue differences compared to SEQ ID NO: 562 at a residue position selected from 29 / 85 / 177 / 208 / 228 / 347, 29 / 85 / 208 / 228 / 343 / 347, 29 / 177 / 195 / 228 / 343, 29 / 208 / 228 / 278 / 294 / 347, 56 / 195 / 278, 85 / 187 / 205 / 208 / 278, 113 / 177 / 187 / 195 / 208 / 278 / 294 / 343 / 347, and 177 / 205 / 208 / 228. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 562 and one or more residue differences selected from 29S / 85P / 177A / 208S / 228T / 347E, 29S / 85P / 208L / 228T / 343T / 347E, 29S / 177P / 195Y / 228T / 343T, 29S / 208S / 228T / 278S / 294T / 347E, 56P / 195Y / 278S, 85P / 187P / 205S / 208L / 278S, 113N / 177P / 187P / 195Y / 208S / 278S / 294Y / 343T / 347E, and 177A / 205S / 208L / 228T (relative to SEQ ID NO: 562).In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 562 and one or more residue differences selected from A29S / L85P / S177A / E208S / V228T / A347E, A29S / L85P / E208L / V228T / V343T / A347E, A29S / S177P / H195Y / V228T / V343T, A29S / E208S / V228T / N278S / H294T / A347E, S56P / H195Y / N278S, L85P / C187P / A205S / E208L / N278S, S113N / S177P / C187P / H195Y / E208S / N278S / H294Y / V343T / A347E, and S177A / A205S / E208L / V228T (relative to SEQ ID NO: 562).

[0178] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 598 and one or more residue differences at residue positions selected from: 47, 162, 209, 219, 227, and 342. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 598 and one or more residue differences selected from: 47Q, 162S, 209H, 219V, 227R, 342L, and 342M (relative to SEQ ID NO: 598). In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 598 and one or more residue differences selected from: F47Q, V162S, C209H, T219V, S227R, N342L, and N342M (relative to SEQ ID NO: 598).

[0179] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 598 and one or more residue differences as compared to SEQ ID NO: 598 at a residue position selected from 17 / 44 / 179 / 195 / 250 / 313 / 345, 17 / 44 / 199 / 313, 43 / 44 / 195 / 199, 44 / 149 / 164 / 171 / 187, 44 / 179 / 195 / 199, 44 / 179 / 195 / 199 / 345, 79 / 163 / 164 / 171 / 187 / 201 / 286 / 288, 82 / 163 / 164, 82 / 163 / 164 / 171 / 187 / 201 / 203 / 208 / 286 / 288 / 320, 149 / 164 / 171 / 288, and 187 / 286. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 598 and one or more residue differences selected from 17V / 44R / 199C / 313C, 17V / 44V / 179T / 195Y / 250P / 313C / 345D, 43S / 44V / 195Y / 199C, 44R / 179T / 195Y / 199C, 44R / 179T / 195Y / 199C / 345D, 44V / 149N / 164Q / 171M / 187P, 44V / 179T / 195Y / 199C / 345D, 79E / 163D / 164Q / 171M / 187N / 201V / 286P / 288T, 82A / 163D / 164Q, 82A / 163D / 164Q / 171M / 187P / 201V / 203Q / 208I / 286P / 288T / 320V, 149N / 164Q / 171M / 288T, and 187P / 286P (relative to SEQ ID NO: 598).In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 598 and one or more residue differences selected from: S17V / G44R / T199C / L313C, S17V / G44V / L179T / H195Y / V250P / L313C / T345D, A43S / G44V / H195Y / T199C, G44R / L179T / H195Y / T199C, G44R / L179T / H195Y / T199C / T345D, G44V / S149N / S164Q / T171M / C187P, G44V / L179T / H195Y / T199C / T345D, Q79E / Q163D / S164Q / T171M / C187N / A201V / A286P / V288T, K82A / Q163D / S164Q, K82A / Q163D / S164Q / T171M / C187P / A201V / S203Q / L208I / A286P / V288T / K320V, S149N / S164Q / T171M / V288T, and C187P / A286P (relative to SEQ ID NO: 598).

[0180] In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 630 and one or more residue differences compared to SEQ ID NO: 630 at a residue position selected from 82 / 164 / 171 / 203 / 208, 135 / 163 / 164 / 201 / 203 / 208, 162, 162 / 219 / 236, 162 / 219 / 313 / 338, 162 / 236 / 342, 162 / 313 / 342, and 164 / 171 / 201 / 203 / 282. In some embodiments, the engineered polypeptide having proline hydroxylase activity with one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 630 and one or more residue differences selected from 82A / 164T / 171M / 203Q / 208I, 135P / 163D / 164Q / 201V / 203Q / 208I, 162S, 162S / 219V / 236L, 162S / 219V / 313C / 338I, 162S / 236L / 342M, 162S / 313C / 342M, and 164Q / 171M / 201V / 203Q / 282V (relative to SEQ ID NO: 630).In some embodiments, the engineered polypeptide having proline hydroxylase activity having one or more improved properties compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 630 and one or more residue differences selected from K82A / S164T / T171M / S203Q / L208I, S135P / Q163D / S164Q / A201V / S203Q / L208I, V162S, V162S / T219V / T236L, V162S / T219V / L313C / M338I, V162S / T236L / N342M, V162S / L313C / N342M, and S164Q / T171M / A201V / S203Q / L282V (relative to SEQ ID NO: 630).

[0181] As will be appreciated by those skilled in the art, in some embodiments, one or combinations of the above residue differences are selected to remain constant (i.e., maintained) as core features in the engineered proline hydroxylases, and additional residue differences at other residue positions are incorporated into the sequence to generate additional engineered proline hydroxylase polypeptides having improved properties. Accordingly, it will be understood that for any engineered proline hydroxylase comprising one or a subset of the above residue differences, the application contemplates other engineered proline hydroxylases comprising one or a subset of the residue differences, as well as additionally one or more residue differences at other residue positions disclosed herein.

[0182] As mentioned above, engineered polypeptides having proline hydroxylase activity are also capable of converting the substrate compound L-proline to the product compound trans-3-hydroxyproline. In some embodiments, engineered proline hydroxylase polypeptides are capable of converting the substrate compound L-proline to the product compound trans-3-hydroxyproline with an activity that is at least 1.2-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold or greater relative to the activity of the reference polypeptide of SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630. In some embodiments, engineered proline hydroxylase polypeptides capable of converting the substrate compound L-proline to the product compound trans-3-hydroxyproline with an activity that is at least 1.2-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold or greater relative to the activity of the reference polypeptide of SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 include an amino acid sequence having one or more features selected from improved regioselectivity, improved activity, improved specific activity, and / or improved thermostability.

[0183] In some embodiments, engineered proline hydroxylase polypeptides are capable of converting the substrate compound L-proline to the product compound trans-3-hydroxyproline with an activity that is at least 1.2-fold relative to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630, and comprise an amino acid sequence selected from the even-numbered sequences in the ranges of SEQ ID NOs: 6-658.

[0184] In some embodiments, engineered proline hydroxylase polypeptides are capable of converting the substrate compound L-proline to the product compound trans-3-hydroxyproline with an activity that is at least 2-fold relative to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630, and comprise an amino acid sequence having one or more residue differences provided herein (as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630, as appropriate).

[0185] In some embodiments, an engineered proline hydroxylase polypeptide capable of converting the substrate compound L-proline to the product compound trans-3-hydroxyproline with at least 2-fold activity relative to SEQ ID NOs: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630, comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 6-658.

[0186] In some embodiments, an engineered proline hydroxylase polypeptide is capable of converting at least 50% or more, 60% or more, 70% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, or 95% or more of the compound L-proline to the product compound trans-3-hydroxyproline in HTP assay conditions, in SFP assay conditions, or in DSP assay conditions, at a substrate loading of about 100 g / L, about 50 g / L, or about 20 g / L, in 120 h or less, 72 h or less, 48 h or less, or 24 h or less. In some embodiments, an engineered proline hydroxylase polypeptide is capable of converting at least 50% or more of the compound L-proline to the product compound trans-3-hydroxyproline in DSP assay conditions, at a substrate loading of about 20 g / L, in 24 h or less at about 25 °C.

[0187] In some embodiments, an engineered proline hydroxylase has an amino acid sequence comprising one or more residue differences compared to SEQ ID NOs: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 that increase expression of the engineered proline hydroxylase activity in a bacterial host cell, particularly E. coli.

[0188] In some embodiments, an engineered proline hydroxylase polypeptide having improved properties in the conversion of the compound L-proline to the product compound trans-3-hydroxyproline has an amino acid sequence comprising a sequence selected from the group consisting of SEQ ID NOs: 6-658.

[0189] In some embodiments, the engineered polypeptide having proline hydroxylase activity comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to one of the even numbered sequences in the ranges selected from the group consisting of SEQ ID NOs: 6-658, and the amino acid residue differences present in any of the even numbered sequences in the ranges of SEQ ID NOs: 6-658 compared to SEQ ID NOs: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 as provided in Table 4.1, Table 4.2, Table 4.3, Table 4.4, Table 5.1, Table 5.2, Table 5.3, Table 6.1, Table 7.1, Table 7.2, Table 7.3, Table 8.1, Table 8.2, Table 9.1, Table 9.2, Table 10.1, Table 10.2, Table 11.1, Table 11.2, and / or Table 12.1.

[0190] In addition to the residue positions specified above, any of the engineered proline hydroxylase polypeptides disclosed herein can also comprise other residue differences at other residue positions (i.e., residue positions other than those included in any of the even-numbered sequences in the ranges below: SEQ ID NOs: 6-658) relative to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630. The residue differences at these other residue positions can provide for additional variation of the amino acid sequence without adversely affecting the ability of the polypeptide to carry out the conversion of proline to cis-4-hydroxyproline and the conversion of the compound L-proline to the product compound trans-3-hydroxyproline. Thus, in some embodiments, in addition to the amino acid residue differences present in any of the engineered proline hydroxylase polypeptides selected from the even-numbered sequences in the ranges below: SEQ ID NOs: 6-658, the sequence can also comprise 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-14, 1-15, 1-16, 1-18, 1-20, 1-22, 1-24, 1-26, 1-30, 1-35, 1-40, 1-45, or 1-50 residue differences at other amino acid residue positions compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630. In some embodiments, the number of amino acid residue differences compared to the reference sequence can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 residue positions. In some embodiments, the number of amino acid residue differences compared to the reference sequence can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 residue positions. The residue differences at these other positions can be conservative changes or non-conservative changes. In some embodiments, the residue differences can include conservative substitutions and non-conservative substitutions compared to the naturally-occurring proline hydroxylase polypeptide of SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630.

[0191] In some embodiments, the present application also provides engineered polypeptides comprising fragments of any of the engineered proline hydroxylase polypeptides described herein that retain the functional activity and / or improved properties of engineered proline hydroxylases. Thus, in some embodiments, the present application provides polypeptide fragments capable of converting the compound L-proline to the product compound trans-3-hydroxyproline under suitable reaction conditions, wherein the fragment constitutes at least about 80%, 90%, 95%, 96%, 97%, 98%, or 99% of the full-length amino acid sequence of an engineered proline hydroxylase polypeptide of the present application, such as an exemplary engineered proline hydroxylase polypeptide selected from the even numbered sequences in the following ranges: SEQ ID NOs: 6-658.

[0192] In some embodiments, the engineered proline hydroxylase polypeptides can have an amino acid sequence that comprises a deletion in any of the exemplary engineered polypeptides comprising the engineered proline hydroxylase polypeptide sequences described herein, e.g., the even numbered sequences in the following ranges: SEQ ID NOs: 6-658. Thus, for each and every embodiment of the engineered proline hydroxylase polypeptides of the application, the amino acid sequence can comprise a deletion of one or more amino acids, 2 or more amino acids, 3 or more amino acids, 4 or more amino acids, 5 or more amino acids, 6 or more amino acids, 8 or more amino acids, 10 or more amino acids, 15 or more amino acids, or 20 or more amino acids, up to 10% of the total number of amino acids of the proline hydroxylase polypeptide, up to 20% of the total number of amino acids of the proline hydroxylase polypeptide, or up to 30% of the total number of amino acids of the proline hydroxylase polypeptide, wherein the relevant functional activity and / or improved properties of the engineered proline hydroxylase described herein are maintained. In some embodiments, the deletion can comprise 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-30, 1-35, 1-40, 1-45, or 1-50 amino acid residues. In some embodiments, the number of deletions can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 amino acid residues. In some embodiments, the deletion can comprise a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 amino acid residues.

[0193] In some embodiments, the engineered proline hydroxylase polypeptides herein can have an amino acid sequence comprising a sequence selected from the even numbered sequences in the range of SEQ ID NOs: 6-658, and optionally one or several (e.g., up to 3, 4, 5, or up to 10) amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-30, 1-35, 1-40, 1-45, or 1-50 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the substitutions can be conservative substitutions or non-conservative substitutions.

[0194] In some embodiments, the engineered proline hydroxylase polypeptides herein can have an amino acid sequence comprising a sequence selected from the even numbered sequences in the range of SEQ ID NOs: 6-658, and optionally one or several (e.g., up to 3, 4, 5, or up to 10) amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-30, 1-35, 1-40, 1-45, or 1-50 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the substitutions can be conservative substitutions or non-conservative substitutions.

[0195] In the above embodiments, suitable reaction conditions for engineering the polypeptides are provided as described in the Examples.

[0196] In some embodiments, the polypeptides of the application are fusion polypeptides, in which the engineered polypeptides are fused to other polypeptides, such as by, for example, but not limited to, antibody tags (e.g., myc epitopes), purification sequences (e.g., His tags for binding metals), and cell localization signals (e.g., secretion signals). Thus, the engineered polypeptides described herein can be used with or without being fused to other polypeptides.

[0197] It is to be understood that the polypeptides described herein are not limited to genetically encoded amino acids. In addition to genetically encoded amino acids, the polypeptides described herein can comprise, in whole or in part, naturally occurring and / or synthetic non-coded amino acids. Certain common non-coded amino acids that the polypeptides described herein can comprise include, but are not limited to: D-stereoisomers of genetically encoded amino acids; 2,3-diaminopropionic acid (Dpr); a-aminobutyric acid (Aib); ε-aminohexanoic acid (Aha); δ-aminopentanoic acid (Ava); N-methylglycine or sarcosine (MeGly or Sar); ornithine (Orn); citrulline (Cit); t-butylalanine (Bua); t-butylglycine (Bug); N-methylisoleucine (MeIle); phenylglycine (Phg); cyclohexylalanine (Cha); norleucine (Nle); naphthylalanine (Nal); 2-chlorophenylalanine (Ocf); 3-chlorophenylalanine (Mcf); 4-chlorophenylalanine (Pcf); 2-fluorophenylalanine (Off); 3-fluorophenylalanine (Mff); 4-fluorophenylalanine (Pff); 2-bromophenylalanine (Obf); 3-bromophenylalanine (Mbf); 4-bromophenylalanine (Pbf); 2-methylphenylalanine (Omf); 3-methylphenylalanine (Mmf); 4-methylphenylalanine (Pmf); 2-nitrophenylalanine (Onf); 3-nitrophenylalanine (Mnf); 4-nitrophenylalanine (Pnf); 2-cyanophenylalanine (Ocf); 3-cyanophenylalanine (Mcf); 4-cyanophenylalanine (Pcf); 2-trifluoromethylphenylalanine (Otf); 3-trifluoromethylphenylalanine (Mtf); 4-trifluoromethylphenylalanine (Ptf); 4-aminophenylalanine (Paf); 4-iodophenylalanine (Pif); 4-aminomethylphenylalanine (Pamf); 2,4-dichlorophenylalanine (Opef); 3,4-dichlorophenylalanine (Mpcf); 2,4-difluorophenylalanine (Opff); 3,4-difluorophenylalanine (Mpff); pyrid-2-ylalanine (2pAla); pyrid-3-ylalanine (3pAla); pyrid-4-ylalanine (4pAla); naphth-1-ylalanine (InAla); naphth-2-ylalanine (2nAla); thiazolylalanine (taAla); benzothienylalanine (bAla); thienylalanine (tAla); furanylalanine (fAla); homophenylalanine (hPhe); homotyrosine (hTyr); homotryptophan (hTrp); pentafluorophenylalanine (5ff); styrylalanine (sAla); anthrylalanine (aAla); 3,3-diphenylalanine (Dfa); 3-amino-5-phenylpentanoic acid (Afp); penicillamine (Pen); 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid (Tic); β-2-thienylalanine (Thi);Methionine sulfoxide (Mso); N(w)-nitroarginine (nArg); homolysine (hLys); phosphonomethylphenylalanine (pmPhe); phospho serine (pSer); phospho threonine (pThr); homoaspartic acid (hAsp); homoglutamic acid (hGlu); 1-aminocyclopent-(2 or 3)-ene-4-carboxylic acid; pipecolic acid (PA); azetidine-3-carboxylic acid (ACA); 1-aminocyclopentane-3-carboxylic acid; allylglycine (aGly); propargylglycine (pgGly); homoalanine (hAla); norvaline (nVal); homoleucine (hLeu), homovaline (hVal); homoisoleucine (hIle); homoarginine (hArg); N-acetyl lysine (AcLys); 2,4-diaminobutyric acid (Dbu); 2,3-diaminobutyric acid (Dab); N-methyl valine (MeVal); homocysteine (hCys); homoserine (hSer); hydroxyproline (Hyp), and homoproline (hPro). Additional non-coded amino acids that can be included in the polypeptides described herein will be apparent to those skilled in the art (see, e.g., the various amino acids provided in Fasman, G. H., CRC Press, Boca Raton, FL, pp. 3-70

[1989] and references cited therein, all incorporated by reference). These amino acids can be in the L- or D-configuration. CRC Practical Handbook of Biochemistry and Molecular Biology, CRC Press, Boca Raton, FL, pp. 3-70

[1989] and references cited therein). These amino acids can be in the L- or D-configuration.

[0198] One skilled in the art will recognize that amino acids or residues bearing side chain protecting groups can also constitute the polypeptides described herein. Non-limiting examples of such protected amino acids belonging to the aromatic class in this case include (protecting group listed in parentheses) but are not limited to: Arg(tos), Cys(methylbenzyl), Cys(nitrophenylsulfenyl), Glu(δ-benzyl ester), Gln(xanthyl), Asn(N-δ-xanthyl), His(bom), His(benzyl), His(tos), Lys(fmoc), Lys(tos), Ser(O-benzyl), Thr(O-benzyl), and Tyr(O-benzyl).

[0199] Conformationally restricted non-coded amino acids that can constitute the polypeptides described herein include but are not limited to N-methyl amino acids (L-configuration); 1-aminocyclopent-(2 or 3)-ene-4-carboxylic acid; pipecolic acid; azetidine-3-carboxylic acid; homoproline (hPro); and 1-aminocyclopentane-3-carboxylic acid.

[0200] In some embodiments, the engineered polypeptides can be in a variety of forms, such as, for example, isolated preparations, as substantially purified enzymes, whole cells transformed with a gene encoding the enzyme, and / or as cell extracts and / or lysates of such cells. The enzymes can be in the form of lyophilized, spray-dried, precipitated, or crude paste, as discussed further below.

[0201] In some embodiments, the engineered polypeptides can be provided on a solid support, such as a membrane, resin, solid support, or other solid phase material. The solid support can include organic polymers, such as polystyrene, polyethylene, polypropylene, polyfluoroethylene, polyethyleneoxy, and polyacrylamide, and copolymers and grafts thereof. The solid support can also be inorganic, such as glass, silica, controlled pore glass (CPG), reversed phase silica, or metals such as gold or platinum. The configuration of the solid support can be in the form of a bead, sphere, particle, pellet, gel, membrane, or surface. The surface can be flat, substantially flat, or non-flat. The solid support can be porous or non-porous, and can have swelling or non-swelling characteristics. The solid support can be configured in the form of a well, depression, or other container, vessel, feature, or location.

[0202] In some embodiments, the engineered polypeptides of the present application having proline hydroxylase activity can be immobilized on a solid support, such that the engineered polypeptide retains its improved activity, selectivity, and / or other improved properties relative to the reference polypeptide SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630. In such embodiments, the immobilized polypeptide can facilitate the biocatalytic conversion of a substrate compound or other suitable substrate to a product, and is easily retained (e.g., by retaining the polypeptide-immobilized bead) after the reaction is complete, and then reused or recycled in subsequent reactions. Such methods of immobilized enzymes allow for further efficiency and reduced cost. It is also contemplated accordingly that any of the methods using the engineered proline hydroxylase polypeptides of the present application can be carried out using the same proline hydroxylase polypeptides bound or immobilized on a solid support.

[0203] Methods for enzyme immobilization are well known in the art. Engineered peptides can be bound non-covalently or covalently. Various methods for conjugating and immobilizing enzymes onto solid supports (e.g., resins, membranes, beads, glass, etc.) are well known in the art (see, for example, Yi et al., Proc. Biochem., 42(5):895-898

[2007] ; Martin et al., Appl. Microbiol. Biotechnol., 76(4):843-851

[2007] ; Koszelewski et al., J. Mol. Cat. B: Enzymatic, 63:39-44

[2010] ; Truppo et al., Org. Proc. Res. Dev., published online: dx.doi.org / 10.1021 / op200157c; Hermanson, Bioconjugate Techniques , 2nd edition, Academic Press, Cambridge, MA

[2008] ; Mateo et al., Biotechnol. Prog., 18(3):629-34

[2002] ; and “Bioconjugation Protocols: Strategies and Methods,” in Methods in Molecular Biology Niemeyer (ed.), Humana Press, New York, NY

[2004] ; the disclosure of each of these publications is incorporated herein by reference. Solid supports that can be used to immobilize the engineered proline hydroxylase of the present invention include, but are not limited to, beads or resins comprising polymethacrylate having epoxy functional groups, polymethacrylate having amino epoxy functional groups, styrene / DVB copolymer having octadecyl functional groups, or polymethacrylate. Exemplary solid supports that can be used to immobilize the engineered proline hydroxylase of the present invention include, but are not limited to, chitosan beads, Eupergit C, and SEPABEAD (Mitsubishi), including the following different types of SEPABEAD: EC-EP, EC-HFA / S, EXA252, EXE119, and EXE120.

[0204] In some embodiments, the peptides described herein are provided in the form of a kit. The enzymes in the kit may be present alone or as more than one enzyme. The kit may also contain reagents for performing the enzymatic reaction, substrates for evaluating enzyme activity, and reagents for detecting the product. The kit may also include a reagent dispenser and instructions for using the kit.

[0205] In some embodiments, the kits of the application comprise an array containing more than one different proline hydroxylase polypeptide at different addressable locations, wherein the different polypeptides are different variants of a reference sequence each having at least one different improved enzyme property. In some embodiments, the more than one polypeptides immobilized on a solid support are configured in an array of multiple locations addressable by automated delivery of reagents or by detection methods and / or instrumentation. The array can be used to test various substrate compounds for conversion by the polypeptides. Such arrays comprising more than one engineered polypeptide and methods of use thereof are known in the art (see, e.g., WO 2009 / 008908 A2).

[0206] Polynucleotides, expression vectors, and host cells encoding engineered proline hydroxylases

[0207] In another aspect, the application provides polynucleotides encoding the engineered proline hydroxylase polypeptides described herein. The polynucleotides can be operably linked to one or more heterologous regulatory sequences that control gene expression to produce a recombinant polynucleotide capable of expressing the polypeptide. Expression constructs comprising the heterologous polynucleotides encoding the engineered proline hydroxylases are introduced into suitable host cells to express the corresponding proline hydroxylase polypeptides.

[0208] As will be apparent to the skilled person, the availability of protein sequences and knowledge of the codons corresponding to the various amino acids provides a specification of all polynucleotides capable of encoding the subject polypeptides. The degeneracy of the genetic code, wherein the same amino acid is encoded by alternative or synonymous codons, allows for the production of a vast number of nucleic acids, all of which encode the improved proline hydroxylases. Thus, once a particular amino acid sequence is known, the skilled person can produce any number of different nucleic acids by simply modifying one or more codons of the sequence in a manner that does not change the amino acid sequence of the protein. In this regard, the application expressly contemplates each and every possible variation of polynucleotides encoding the polypeptides described herein that can be produced by selecting combinations of possible codons, and all such variations are to be considered expressly disclosed herein for any polypeptide described herein, including the amino acid sequences presented in Table 4.1, Table 4.2, Table 4.3, Table 4.4, Table 5.1, Table 5.2, Table 5.3, Table 6.1, Table 7.1, Table 7.2, Table 7.3, Table 8.1, Table 8.2, Table 9.1, Table 9.2, Table 10.1, Table 10.2, Table 11.1, Table 11.2, and / or Table 12.1, and the amino acid sequences disclosed as even-numbered sequences in the range of SEQ ID NOs: 6-658 in the Sequence Listing incorporated herein by reference.

[0209] In various embodiments, codons are preferably selected to accommodate the host cell in which the protein is produced. For example, preferred codons used in bacteria are used for expression in bacteria; preferred codons used in yeast are used for expression in yeast; and preferred codons used in mammals are used for expression in mammalian cells. In some embodiments, not all codons need to be replaced to optimize codon usage of the proline hydroxylase, as the native sequence will contain preferred codons, and as it can not be necessary to use preferred codons for all amino acid residues. Thus, a codon-optimized polynucleotide encoding a proline hydroxylase can contain preferred codons at about 40%, 50%, 60%, 70%, 80%, or greater than 90% of the codon positions of the full-length coding region.

[0210] In some embodiments, the polynucleotide comprises a codon-optimized nucleotide sequence encoding a naturally occurring proline hydroxylase polypeptide amino acid sequence as represented by SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630. In some embodiments, the polynucleotide has a nucleic acid sequence comprising at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identical to a codon-optimized nucleic acid sequence encoding an even-numbered sequence in the range of SEQ ID NO: 6-658. In some embodiments, the polynucleotide has a nucleic acid sequence comprising at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identical to a codon-optimized nucleic acid sequence of an odd-numbered sequence in the range of SEQ ID NO: 5-657. The codon-optimized sequence of the odd-numbered sequence in the range of SEQ ID NO: 5-657 enhances expression of the encoded wild-type proline hydroxylase, provides an enzyme preparation capable of converting more than 80% of the compound L-proline to the product compound trans-3-hydroxyproline in a micro-DSP assay condition in vitro, and more than 45% of the compound L-proline to the product compound trans-3-hydroxyproline in a DSP assay condition. In some embodiments, the codon-optimized polynucleotide sequence can enhance expression of the proline hydroxylase by at least 1.2-fold, 1.5-fold, or 2-fold or greater compared to the naturally occurring polynucleotide sequence from ANO of the fungus sp. No. 11243.

[0211] In some embodiments, the polynucleotide is capable of hybridizing under high stringency conditions to a reference sequence selected from the odd-numbered sequences in SEQ ID NO: 3-657, or the complement thereof, and encodes a polypeptide having proline hydroxylase activity.

[0212] In some embodiments, as described above, the polynucleotide encodes an engineered polypeptide having proline hydroxylase activity having one or more improved properties as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630, wherein the polypeptide comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to a reference sequence selected from SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630, and one or more residue differences as compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 selected from the even numbered sequences in the range of SEQ ID NO: 6-658. In some embodiments, the reference amino acid sequence is selected from the even numbered sequences in the range of SEQ ID NO: 6-658. In some embodiments, the reference amino acid sequence is SEQ ID NO: 4. In some embodiments, the reference amino acid sequence is SEQ ID NO: 116. In some embodiments, the reference amino acid sequence is SEQ ID NO: 162. In some embodiments, the reference amino acid sequence is SEQ ID NO: 322. In some embodiments, the reference amino acid sequence is SEQ ID NO: 412. In some embodiments, the reference amino acid sequence is SEQ ID NO: 492. In some embodiments, the reference amino acid sequence is SEQ ID NO: 562. In some embodiments, the reference amino acid sequence is SEQ ID NO: 598. In some embodiments, the reference amino acid sequence is SEQ ID NO: 630.

[0213] In some embodiments, the polynucleotide encodes an engineered proline hydroxylase polypeptide capable of converting a substrate compound L-proline to a product compound trans-3-hydroxyproline with improved enzyme properties compared to the reference polypeptide of SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630, wherein the polypeptide comprises an amino acid sequence that is at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the reference polypeptide of any one of the even numbered sequences selected from the ranges of SEQ ID NOs: 6-658, provided that the amino acid sequence comprises any one of the sets of residue differences contained in any one of the polypeptide sequences from any one of the even numbered sequences in the ranges of SEQ ID NOs: 6-658 compared to SEQ ID NO: 4, 116, 162, 322, 412, 492, 562, 598, and / or 630 as listed in Table 4.1, Table 4.2, Table 4.3, Table 4.4, Table 5.1, Table 5.2, Table 5.3, Table 6.1, Table 7.1, Table 7.2, Table 7.3, Table 8.1, Table 8.2, Table 9.1, Table 9.2, Table 10.1, Table 10.2, Table 11.1, Table 11.2, and / or Table 12.1.

[0214] In some embodiments, the polynucleotide encoding the engineered proline hydroxylase comprises a polynucleotide sequence selected from the odd numbered sequences in the ranges of SEQ ID NOs: 5-657.

[0215] In some embodiments, the polynucleotide is capable of hybridizing to a reference polynucleotide sequence selected from the odd numbered sequences in the ranges of SEQ ID NOs: 5-657, or the complement thereof, under high stringency conditions, and encodes a polypeptide having proline hydroxylase activity with one or more improved properties described herein.

[0216] In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an engineered polypeptide having proline hydroxylase activity having one or more improved properties, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 4, and one or more residue differences as compared to SEQ ID NO: 4 at a residue position selected from 21, 28, 58 / 247, 65, 80, 85, 95, 98, 117, 120, 159, 185, 194, 199, 200, 233, 237, 243, 250, 268, 281, 282, 287, 289, 307, 324, 326, 327, 330, 338, 343, 346, and 348 or at a residue position selected from 21, 28, 45, 65, 95, 112, 117, 139, 177, 185, 199, 233, 243, 250, 281, 282, 287, 289, 307, 324, 326, 327, 335, 338, 343, and 346 or at a residue position selected from 48 / 66 / 189 / 194, 48 / 66 / 194, and 66 / 82 / 85 / 135 / 189 / 194 / 267 or at a residue position selected from 20 / 56 / 76 / 168 / 169 / 296, 20 / 56 / 232 / 294, 20 / 119 / 294 / 296, 56 / 76 / 119 / 124 / 147 / 232, 56 / 76 / 294, 76 / 168 / 232 / 294, 76 / 294 / 296, 76 / 296, 147, and 232.

[0217] In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an engineered polypeptide having proline hydroxylase activity with one or more improved properties, the engineered polypeptide comprising having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 116, and an alteration at a residue position selected from 123, 189, 195, 233, and 296 or at a residue position selected from 20 / 21 / 56, 20 / 21 / 56 / 76 / 95 / 232 / 294 / 307 / 335, 20 / 21 / 56 / 76 / 147 / 225 / 232 / 233 / 281 / 294 / 296 / 307 / 335, 20 / 21 / 56 / 95 / 147 / 281 / 294 / 307, 20 / 21 / 56 / 281 / 307, 20 / 21 / 76 / 232 / 243, 20 / 21 / 95 / 232 / 307, 20 / 21 / 95 / 281 / 294 / 296, 20 / 21 / 147 / 189 / 233 / 243 / 281 / 307, 20 / 56, 20 / 56 / 76 / 95 / 281 / 307, 20 / 56 / 76 / 147 / 294 / 296 / 307, 20 / 56 / 95 / 147 / 294, 20 / 56 / 281, 20 / 76, 20 / 76 / 95 / 281 / 294 / 296, 20 / 76 / 95 / 281 / 296 / 307, 20 / 76 / 233 / 294 / 307, 20 / 76 / 243 / 281 / 294, 21 / 76 / 147 / 233 / 294 / 307, 21 / 76 / 147 / 243 / 296 / 307 / 335, 21 / 95 / 185 / 189 / 232 / 281 / 296, 21 / 95 / 233 / 243 / 281 / 296, 21 / 95 / 294 / 296 / 307 / 335, 21 / 95 / 307, 21 / 281 / 307, 29 / 76 / 281, 56 / 76 / 95 / 232 / 243 / 281, 56 / 76 / 147 / 281 / 307, 56 / 76 / 243 / 294, 56 / 76 / 281 / 294, 56 / 76 / 296, 56 / 76 / 307, 56 / 95 / 147 / 307 / 335 / 348, 56 / 95 / 232 / 233 / 281 / 294 / 307, 56 / 95 / 243 / 281, 56 / 147 / 281, 56 / 232 / 243 / 281, 56 / 232 / 281, 56 / 232 / 281 / 294 / 296, 56 / 233 / 281 / 294 / 296, 56 / 281 / 307, 76 / 95 / 232 / 243 / 281 / 307,76 / 95 / 243 / 281 / 307 / 335, 76 / 95 / 294 / 307, 76 / 147, 76 / 147 / 233 / 243 / 294, 76 / 147 / 233 / 281 / 294 / 307, 76 / 147 / 243 / 294 / 296 / 307 / 335, 76 / 147 / 281 / 307, 76 / 189 / 296, 76 / 232 / 233 / 243 / 294 / 296 / 307, 76 / 281, 76 / 281 / 294, 76 / 294 / 296, 95 / 120, 95 / 147 / 335, 95 / 232 / 243 / 281 / 294 / 307, 95 / 232 / 281 / 294 / 296, 95 / 281 / 294 / 296, 95 / 335, 147, 147 / 225 / 232 / 243 / 281 / 296 / 307 / 335, 147 / 233 / 243 / 281 / 307, 147 / 233 / 281 / 307 / 335, 147 / 243 / 281, 147 / 307, 232 / 233 / 281 / 294 / 296 / 307, 232 / 281, 232 / 284 / 307, 233 / 243 / 281 / 296 / 307 / 335, 233 / 281 / 296 / 307, 243 / 281 / 294 / 296, 281, 281 / 294, 281 / 307, 307, and 335, or at one or more residue differences at residue positions selected from 21 / 76 / 147 / 243 / 296 / 307 / 335, 56 / 76 / 147 / 281 / 307, and 95 / 147 / 335.

[0218] In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an engineered polypeptide having proline hydroxylase activity with one or more improved properties, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 162, and one or more residue differences as compared to SEQ ID NO: 162 at a residue position selected from 2 / 85 / 123 / 237, 28 / 115 / 117 / 120 / 123 / 268 / 270 / 343 / 346 / 348, 45 / 123 / 326, 65 / 117 / 120 / 123 / 343 / 346, 85 / 123 / 281 / 282, 114 / 115 / 117 / 120 / 123 / 268 / 271 / 313 / 326 / 343 / 346, 123 / 139 / 233 / 237 / 281 / 282 / 289 / 324 / 326, and 123 / 199 / 200 / 247 / 250 / 338.

[0219] In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an engineered polypeptide having proline hydroxylase activity with one or more improved properties, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 322, and one or more residue differences as compared to SEQ ID NO: 322 at a residue position selected from 26, 54, 61, 129, 132, 149, 156, 175, 189, 201, 209, 228, 236, 248, 262, 272, 277, 291, and 345, or at a residue position selected from 25, 43, 54, 58, 61, 79, 129, 132, 143, 156, 163, 175, 179, 201, 209, 236, 248, 278, 291, 345, and 347, or at a residue position selected from 85 / 117 / 120 / 135 / 208 / 270 / 324 / 343 / 346, 85 / 117 / 120 / 135 / 208 / 281 / 282 / 289, 85 / 117 / 120 / 270 / 281 / 289, 85 / 117 / 135 / 139 / 208, and 117 / 120 / 208 / 270 / 324 / 343 / 346.

[0220] In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an engineered polypeptide having proline hydroxylase activity with one or more improved properties, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 412, and one or more residue differences compared to SEQ ID NO: 412 at a residue position selected from 47, 48, 56 / 118, 85, 95, 95 / 289, 113, 118, 118 / 247, 154, 162, 162 / 204, 164, 164 / 198 / 271, 168, 169, 187, 195, 243, 271, 275, 281, 314, 330, and 342 or at a residue position selected from 25 / 129 / 163 / 236 / 262 / 345 / 347, 120 / 156 / 175 / 179 / 201, 129 / 189 / 236 / 262 / 277 / 278, 129 / 236 / 262, 156 / 175 / 179 / 228, and 162.

[0221] In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an engineered polypeptide having proline hydroxylase activity with one or more improved properties, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 492, and one or more residue differences compared to SEQ ID NO: 492 at a residue position selected from 15, 17, 28, 29, 65, 135, 167, 177, 199, 208, 228, 235, 287, 294, 307, and 343 or at a residue position selected from 85 / 187 / 281 / 347, 85 / 187 / 347, 118 / 120 / 162 / 175 / 179 / 330, 118 / 120 / 162 / 175 / 330, 162 / 175 / 179 / 330, 175 / 228 / 330, 195 / 347, and 278 / 314 / 347.

[0222] In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an engineered polypeptide having proline hydroxylase activity with one or more improved properties, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 562, and one or more residue differences compared to SEQ ID NO: 562 at a residue position selected from 15, 40, 43, 44, 59, 79, 82, 149, 164, 179, 345, and 347 or at residue positions selected from 29 / 85 / 177 / 208 / 228 / 347, 29 / 85 / 208 / 228 / 343 / 347, 29 / 177 / 195 / 228 / 343, 29 / 208 / 228 / 278 / 294 / 347, 56 / 195 / 278, 85 / 187 / 205 / 208 / 278, 113 / 177 / 187 / 195 / 208 / 278 / 294 / 343 / 347, and 177 / 205 / 208 / 228.

[0223] In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an engineered polypeptide having proline hydroxylase activity with one or more improved properties, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 598, and one or more residue differences compared to SEQ ID NO: 598 at a residue position selected from 47, 162, 209, 219, 227, and 342 or at residue positions selected from 17 / 44 / 179 / 195 / 250 / 313 / 345, 17 / 44 / 199 / 313, 43 / 44 / 195 / 199, 44 / 149 / 164 / 171 / 187, 44 / 179 / 195 / 199, 44 / 179 / 195 / 199 / 345, 79 / 163 / 164 / 171 / 187 / 201 / 286 / 288, 82 / 163 / 164, 82 / 163 / 164 / 171 / 187 / 201 / 203 / 208 / 286 / 288 / 320, 149 / 164 / 171 / 288, and 187 / 286.

[0224] In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an engineered polypeptide having proline hydroxylase activity having one or more improved properties, the engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 630, and one or more residue differences as compared to SEQ ID NO: 630 at residue positions selected from 82 / 164 / 171 / 203 / 208, 135 / 163 / 164 / 201 / 203 / 208, 162, 162 / 219 / 236, 162 / 219 / 313 / 338, 162 / 236 / 342, 162 / 313 / 342, and 164 / 171 / 201 / 203 / 282.

[0225] In some embodiments, the polynucleotide encodes a polypeptide described herein, but has at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity at the nucleotide level to a reference polynucleotide encoding an engineered proline hydroxylase. In some embodiments, the reference polynucleotide sequence is selected from the odd numbered sequences in the range of SEQ ID NOs: 3-657.

[0226] In some embodiments, an isolated polynucleotide encoding any of the engineered proline hydroxylase polypeptides provided herein is manipulated in various ways to provide for expression of the polypeptide. In some embodiments, the polynucleotide encoding the polypeptide is provided in an expression vector in which one or more control sequences are present that control the expression of the polynucleotide and / or the polypeptide. Manipulation of the isolated polynucleotide prior to its insertion into the vector can be desirable or necessary for the purposes of the expression of the polypeptide. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art.

[0227] In some embodiments, the control sequences include, among others, a promoter, a leader sequence, polyadenylation sequence, a pre-pro peptide sequence, a signal peptide sequence, and a transcription terminator. Suitable promoters can be selected based on the host cell used, as is known in the art. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the application, include, but are not limited to, promoters obtained from the E. coli lac operon, Streptomyces coelicolor agarase gene (dagA), Bacillus subtilis levansucrase gene (sacB), Bacillus licheniformis alpha-amylase gene (amyL), Bacillus stearothermophilus maltogenic amylase gene (amyM), Bacillus amyloliquefaciens alpha-amylase gene (amyQ), Bacillus licheniformis penicillinase gene (penP), Bacillus subtilis xylA and xylB genes, and prokaryotic beta-lactamase gene (see e.g., Villa-Kamaroff et al., Proc. Natl Acad. Sci. USA 75: 3727-3731

[1978] ) as well as the tac promoter (see e.g., DeBoer et al., Proc. Natl Acad. Sci. USA 80: 21-25

[1983] ). Exemplary promoters for filamentous fungal host cells include promoters obtained from the genes for Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral alpha-amylase, Aspergillus niger acid stable alpha-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (see, e.g., WO 96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoters from the genes for Aspergillus niger neutral alpha-amylase and Aspergillus oryzae triose phosphate isomerase), and mutant, truncated, and hybrid promoters thereof.Exemplary yeast cell promoters can be derived from the genes for Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase. Other useful promoters for yeast host cells are known to those skilled in the art (see, e.g., Romanos et al., Yeast 8:423-488

[1992] ).

[0228] In some embodiments, the control sequence is a suitable transcription terminator sequence, i.e., a sequence recognized by a host cell to terminate transcription. The terminator sequence is operably linked to the 3' terminus of the nucleic acid sequence encoding the polypeptide. Any terminator that is functional in the host cell of choice can be used in the present application. Exemplary transcription terminators for filamentous fungal host cells are obtained from the genes for Aspergillus niger TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger alpha- glucosidase, and Fusarium oxysporum trypsin-like protease. Useful exemplary transcription terminators for yeast host cells are obtained from the genes for Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are known to those skilled in the art (see, e.g., Romanos et al., supra).

[0229] In some embodiments, the control sequence is a suitable leader sequence, i.e., an untranslated region of an mRNA that is important for translation by the host cell. The leader sequence is operably linked to the 5' terminus of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in the host cell of choice can be used. Useful exemplary leader sequences for filamentous fungal host cells are obtained from the genes for Aspergillus niger TAKA amylase and Aspergillus nidulans phosphoglycerate mutase. Suitable leader sequences for yeast host cells include, but are not limited to, those obtained from the genes for Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae alpha-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).

[0230] A control sequence can also be a polyadenylation sequence, a sequence that is operably linked to the 3' terminus of the nucleic acid sequence and that, when transcribed, is recognized by the host cell as a signal to add a poly(A) sequence to the 3' end of the mRNA. Any polyadenylation sequence that is functional in the host cell of choice can be used in the present application. Exemplary polyadenylation sequences for filamentous fungal host cells include, but are not limited to, those from Aspergillus niger TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger alpha-glucosidase. Useful polyadenylation sequences for yeast host cells are also known in the art (see, e.g., Guo and Sherman, Mol. Cell. Biol., 15: 5983-5990

[1995] ).

[0231] In some embodiments, the control sequence is a signal peptide coding region that codes for an amino acid sequence that is linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the secretory pathway of a cell. The 5' terminus of the coding sequence of a nucleic acid sequence can inherently contain a signal peptide coding region naturally linked in translation reading frame to the segment of the coding region that encodes the secreted polypeptide. Alternatively, the 5' terminus of the coding sequence can contain a signal peptide coding region that is foreign to the underlying nucleic acid sequence. Any signal peptide coding region that directs the expressed polypeptide into the secretory pathway of a selected host cell can be used in the expression of the engineered proline hydroxylase polypeptides provided herein. Useful signal peptide coding regions for bacterial host cells include, but are not limited to, those obtained from the genes for Bacillus NCIB 11837 maltogenic amylase, Bacillus stearothermophilus alpha-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus amyloliquefaciens neutral proteases (nprT, nprS, nprM), and Bacillus subtilis prsA. Additional signal peptides are known in the art (see, e.g., Simonen and Palva, Microbiol. Rev., 57: 109-137

[1993] ). Useful signal peptide coding regions for filamentous fungal host cells include, but are not limited to, those obtained from the genes for Aspergillus niger TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase. Useful signal peptides for yeast host cells include, but are not limited to, those from the genes for Saccharomyces cerevisiae alpha factor and Saccharomyces cerevisiae invertase.

[0232] In some embodiments, the control sequence is a propeptide coding region that encodes an amino acid sequence positioned at the amino terminus of the polypeptide. The resulting polypeptide is referred to as a "proenzyme," "propolypeptide," or in some cases a "zymogen." The propolypeptide can be converted into a mature, active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide. Propeptide coding regions include, but are not limited to, the genes for Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Saccharomyces cerevisiae alpha-factor, Rhizomucor miehei aspartic proteinase, and Myceliophthora thermophila lactase (see, e.g., WO 95 / 33836). Where both a signal peptide and a propeptide region are present at the amino terminus of the polypeptide, the propeptide region is positioned immediately adjacent to the amino terminus of the polypeptide and the signal peptide region is positioned immediately adjacent to the amino terminus of the propeptide region.

[0233] In some embodiments, regulatory sequences are also utilized. These sequences are promoters, enhancers, and / or other expression control elements that control the transcription or translation of the polypeptides relative to the growth of the host cell. Examples of control systems are those that cause gene expression to turn on or turn off in response to chemical or physical stimulation, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory sequences include, but are not limited to, the ADH2 system or GAL1 system. In filamentous fungi, suitable regulatory sequences include, but are not limited to, the TAKA alpha-amylase promoter, an Aspergillus niger glucoamylase promoter, and an Aspergillus oryzae glucoamylase promoter.

[0234] In another aspect, the present application also provides recombinant expression vectors comprising a polynucleotide encoding an engineered proline hydroxylase polypeptide and, as appropriate for the type of host cell into which it is to be introduced, one or more expression control sequences, such as promoter and terminator sequences, origins of replication, and the like. In some embodiments, the various nucleic acids and control sequences described above are combined to produce a recombinant expression vector comprising one or more convenient restriction sites for insertion or substitution of the nucleic acid sequence encoding the variant proline hydroxylase polypeptide at such sites. Alternatively, the polynucleotide sequences of the present application can be expressed by inserting the polynucleotide sequence or a nucleic acid construct comprising the polynucleotide sequence into an appropriate vector for expression. When producing an expression vector, the coding sequence is located in the vector such that it is operably linked with the appropriate control sequences for expression.

[0235] The recombinant expression vector can be any vector (e.g., a plasmid or virus) which can be conveniently subjected to recombinant DNA procedures and which can bring about the expression of the variant proline hydroxylase polynucleotide sequence. The choice of vector will often depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector can be linear or closed circular plasmids.

[0236] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extrachromosomal entity, replicates independently of chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome). The vector can include any means for assuring self-replication. In some alternative embodiments, the vector can be one which, when introduced into a host cell, is integrated into the genome and replicated together with the chromosome into which it has been integrated. Furthermore, the vector can be a single vector or plasmid or two or more vectors or plasmids, or a transposon, which together include the total DNA to be introduced into the genome of the host cell.

[0237] In some embodiments, the expression vector preferably comprises one or more selectable markers that permit easy selection of transformed cells. A "selectable marker" is a gene the product of which provides for biocide or viral resistance, resistance to heavy metals, an ability to utilize a unique nutrient supply, and the like. Examples of selectable markers for bacteria include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis, or markers that confer resistance to antibiotics such as ampicillin, kanamycin, chloramphenicol or tetracycline. Suitable markers for use in a yeast host cell include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for use in a filamentous fungal host cell include, but are not limited to, amdS (acetamidase), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase), sC (sulfate adenyltransferase), and trpC (anthranilate synthase), as well as equivalents thereof. In another aspect, the present application provides host cells comprising a polynucleotide encoding at least one engineered proline hydroxylase polypeptide of the present application operably linked to one or more control sequences for expression of the engineered proline hydroxylase in the host cell. Host cells for expression of polypeptides encoded by expression vectors of the present application are well known in the art and include, but are not limited to, bacterial cells such as Escherichia coli, Vibrio fluvialis, Streptomyces, and Salmonella typhimurium cells; fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae and Pichia pastoris [ATCC Accession No. 201178]); insect cells such as Drosophila S2 and Spodoptera Sf9 cells; animal cells such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Exemplary host cells are Escherichia coli strains (e.g., W3110 (AfhuA) and BL21).

[0238] Accordingly, in another aspect, the present application provides a method for producing an engineered proline hydroxylase polypeptide, wherein the method comprises culturing a host cell capable of expressing a polynucleotide encoding an engineered proline hydroxylase polypeptide under conditions suitable for expression of the polypeptide. In some embodiments, the method further comprises the step of isolating and / or purifying a proline hydroxylase polypeptide as described herein.

[0239] Suitable media and growth conditions for the host cells described above are well known in the art. The polynucleotide for expressing a proline hydroxylase polypeptide can be introduced into a cell by a variety of methods known in the art. Techniques include, among others, electroporation, biolistic bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion.

[0240] Engineered proline hydroxylases having the properties disclosed herein can be obtained by subjecting a polynucleotide encoding a naturally occurring or engineered proline hydroxylase polypeptide to mutagenesis and / or directed evolution methods known in the art and as described herein. An exemplary directed evolution technique is mutagenesis and / or DNA shuffling (see, e.g., Stemmer, Proc. Natl. Acad. Sci. USA 91 : 10747-10751

[1994] ; WO 95 / 22625; WO 97 / 0078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767; and U.S. Patent No. 6,537,746). Other directed evolution procedures that can be used include, among others, staggered extension process (StEP), in vitro recombination (see, e.g., Zhao et al., Nat. Biotechnol., 16:258-261

[1998] ), mutagenic PCR (see, e.g., Caldwell et al., PCR Methods Appl., 3:S136-S140

[1994] ), and cassette mutagenesis (see, e.g., Black et al., Proc. Natl. Acad. Sci. USA 93:3525-3529

[1996] ).

[0241] In some implementations, as discussed above, engineered proline hydroxylases can be obtained by subjecting the polynucleotide encoding a naturally occurring proline hydroxylases to mutagenesis and / or directed evolution. Mutagenesis can be performed using any techniques known in the art, including random mutagenesis and site-directed mutagenesis. Directed evolution can be performed using any techniques known in the art, including shuffling, to screen for improved promoter variants. Mutagenesis and directed evolution methods are well known in the art (see, for example, U.S. Patents 5,605,793, 5,811,238, 5,830,721, 5,834,252, 5,837,458, 5,928,905, 6,096,548, 6,117,679, 6,132,970, 6,165,793, 6,180,406, 6,251,674, 6,265,201, 6,277,638, 6,287,861, 6,287,86...). No. 2, No. 6,291,242, No. 6,297,053, No. 6,303,344, No. 6,309,883, No. 6,319,713, No. 6,319,714, No. 6,323,030, No. 6,326,204, No. 6,335,160, No. 6,335,198, No. 6,344,356, No. 6,352,859, No. 6,355,484, No. 6,358,740, No. 6,358,742, No. 6,365,377, No. 6,365,408, No. 6,368,8 No. 61, No. 6,372,497, No. 6,337,186, No. 6,376,246, No. 6,379,964, No. 6,387,702, No. 6,391,552, No. 6,391,640, No. 6,395,547, No. 6,406,855, No. 6,406,910, No. 6,413,745, No. 6,413,774, No. 6,420,175, No. 6,423,542, No. 6,426,224, No. 6,436,675, No. 6,444,468, No. 6,455 No. 253, No. 6,479,652, No. 6,482,647, No. 6,483,011, No. 6,484,105, No. 6,489,146, No. 6,500,617, No. 6,500,639, No. 6,506,602, No. 6,506,603, No. 6,518,065, No. 6,519,065, No. 6,521,453, No. 6,528,311, No. 6,537,746, No. 6,573,098, No. 6,576,467, No. 6,579,678, No. 6,586,U.S. Patent Nos. 6,183,772; 6,602,986; 6,605,430; 6,613,514; 6,653,072; 6,686,515; 6,703,240; 6,716,631; 6,825,001; 6,902,922; 6,917,882; 6,946,296; 6,961,664; 6,995,017; 7,024,312; 7,058,515; 7,105,297; 7,148,054; 7,220,566; 7,288,375; 7,384,387; 7,421,347; 7,430,477; 7,462,469; 7,534,564; 7,620,500; 7,620,502; 7,629,170; 7,702,464; 7,747,391; 7,747,393; 7,751,986; 7,776,598; 7,783,428; 7,795,030; 7,853,410; 7,868,138; 7,783,428; 7,873,477; 7,873,499; 7,904,249; 7,957,912; 7,981,614; 8,014,961; 8,029,988; 8,048,674; 8,058,001; 8,076,138; 8,108,150; 8,170,806; 8,224,580; 8,377,681; 8,383,346; 8,457,903; 8,504,498; 8,589,085; 8,762,066; 8,768,871; 9,593,326, and all related non-United States counterpart patents; Ling et al., Anal. Biochem., 254(2): 157-78

[1997] ; Dale et al., Meth. Mol. Biol., 57: 369-74

[1996] ; Smith, Ann. Rev. Genet., 19: 423-462

[1985] ; Botstein et al., Science, 229: 1193-1201

[1985] ; Carter, Biochem. J., 237: 1-7

[1986] ; Kramer et al., Cell, 38: 879-887

[1984] ; Wells et al., Gene, 34: 315-323

[1985] ; Minshull et al., Curr. Op. Chem. Biol.,3:284-290

[1999] ; Christians et al., Nat. Biotechnol., 17:259-264

[1999] ; Crameri et al., Nature, 391 :288-291

[1998] ; Crameri et al., Nat. Biotechnol., 15:436-438

[1997] ; Zhang et al., Proc. Nat. Acad. Sci. U.S.A., 94:4504-4509

[1997] ; Crameri et al., Nat. Biotechnol., 14:315-319

[1996] ; Stemmer, Nature, 370:389-391

[1994] ; Stemmer, Proc. Nat. Acad. Sci. USA, 91 :10747-10751

[1994] ; WO 95 / 22625; WO 97 / 0078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767; and WO 2009 / 152336, all of which are incorporated herein by reference.

[0242] In some embodiments, the enzyme clones obtained after mutagenesis treatment are screened by subjecting the enzymes to a defined temperature (or other assay conditions, such as testing the activity of the enzymes over a broad range of substrates) and measuring the amount of enzyme activity remaining after the heat treatment or other assay conditions. The clones containing polynucleotides encoding proline hydroxylase polypeptides are then sequenced to identify changes in the nucleotide sequence, if any, and the clones are used to express the enzymes in host cells. Measuring enzyme activity from expression libraries can be performed using any suitable method known in the art (e.g., standard biochemical techniques, such as HPLC analysis).

[0243] In some embodiments, the clones obtained after mutagenesis treatment can be screened for engineered proline hydroxylases having one or more desired improved enzyme properties (e.g., improved regioselectivity). Measuring enzyme activity from expression libraries can be performed using standard biochemical techniques, such as HPLC analysis and / or derivatization of the product (either pre- or post-isolation), for example, using dansyl chloride or OPA (see, e.g., Yaegaki et al., J Chromatogr. 356(1): 163-70

[1986] ).

[0244] When the sequence of an engineered polypeptide is known, a polynucleotide encoding the enzyme can be prepared by standard solid-phase methods according to known methods of synthesis. In some embodiments, fragments of up to about 100 bases can be synthesized individually and then joined (e.g., by enzymatic or chemical ligation methods or polymerase-mediated methods) to form any desired continuous sequence. For example, polynucleotides and oligonucleotides encoding portions of proline hydroxylases can be prepared by chemical synthesis according to methods known in the art (e.g., the classical phosphoramidite method of Beaucage et al., Tet. Lett. 22: 1859-69

[1981] or the method described by Matthes et al., EMBO J. 3: 801-05

[1984] ), as is commonly practiced in automated synthesis methods. According to the phosphoramidite method, oligonucleotides are synthesized (e.g., in an automated DNA synthesizer), purified, annealed, ligated, and cloned into an appropriate vector. Additionally, essentially any nucleic acid is available from any of a variety of commercial sources. In some embodiments, additional variations can be made by synthesizing oligonucleotides containing deletions, insertions, and / or substitutions and producing oligonucleotides in various permutations and combinations to produce engineered proline hydroxylases with one or more improved properties.

[0245] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising an amino acid sequence having at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to the amino acid sequence of a sequence selected from the even-numbered sequences of SEQ ID NOS: 4-658 and having one or more residue differences as compared to SEQ ID NO: 4 at a residue position selected from: 21, 28, 58 / 247, 65, 80, 85, 95, 98, 117, 120, 159, 185, 194, 199, 200, 233, 237, 243, 250, 268, 281, 282, 287, 289, 307, 324, 326, 327, 330, 338, 343, 346, and 348; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0246] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising an amino acid sequence having at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to the amino acid sequence of a sequence selected from the group consisting of even-numbered sequences of SEQ ID NOs: 4-658 and having an amino acid sequence with one or more residue differences as compared to SEQ ID NO: 4 at a residue position selected from the group consisting of 21, 28, 45, 65, 95, 112, 117, 139, 177, 185, 199, 233, 243, 250, 281, 282, 287, 289, 307, 324, 326, 327, 335, 338, 343, and 346; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0247] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising an amino acid sequence having at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to the amino acid sequence of a sequence selected from the group consisting of even-numbered sequences of SEQ ID NOs: 4-658 and having an amino acid sequence with one or more residue differences as compared to SEQ ID NO: 4 at a residue position selected from the group consisting of 48 / 66 / 189 / 194, 48 / 66 / 194, and 66 / 82 / 85 / 135 / 189 / 194 / 267; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0248] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising an amino acid sequence having at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to the amino acid sequence of a sequence selected from the group consisting of even-numbered sequences of SEQ ID NOs: 4-658 and having an amino acid sequence with one or more residue differences compared to SEQ ID NO: 4 at a residue position selected from the group consisting of: 20 / 56 / 76 / 168 / 169 / 296, 20 / 56 / 232 / 294, 20 / 119 / 294 / 296, 56 / 76 / 119 / 124 / 147 / 232, 56 / 76 / 294, 76 / 168 / 232 / 294, 76 / 294 / 296, 76 / 296, 147, and 232; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0249] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising an amino acid sequence having at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity to the amino acid sequence of a sequence selected from the group consisting of even-numbered sequences of SEQ ID NOs: 4-658 and having an amino acid sequence with one or more residue differences compared to SEQ ID NO: 116 at a residue position selected from the group consisting of: 123, 189, 195, 233, and 296; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0250] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising an amino acid sequence selected from the even-numbered sequences of SEQ ID NOs: 4-658 and having one or more residue differences compared to SEQ ID NO: 116 at a residue position selected from: 20 / 21 / 56, 20 / 21 / 56 / 76 / 95 / 232 / 294 / 307 / 335, 20 / 21 / 56 / 76 / 147 / 225 / 232 / 233 / 281 / 294 / 296 / 307 / 335, 20 / 21 / 56 / 95 / 147 / 281 / 294 / 307, 20 / 21 / 56 / 281 / 307, 20 / 21 / 76 / 232 / 243, 20 / 21 / 95 / 232 / 307, 20 / 21 / 95 / 281 / 294 / 296, 20 / 21 / 147 / 189 / 233 / 243 / 281 / 307, 20 / 56, 20 / 56 / 76 / 95 / 281 / 307, 20 / 56 / 76 / 147 / 294 / 296 / 307, 20 / 56 / 95 / 147 / 294, 20 / 56 / 281, 20 / 76, 20 / 76 / 95 / 281 / 294 / 296, 20 / 76 / 95 / 281 / 296 / 307, 20 / 76 / 233 / 294 / 307, 20 / 76 / 243 / 281 / 294, 21 / 76 / 147 / 233 / 294 / 307, 21 / 76 / 147 / 243 / 296 / 307 / 335, 21 / 95 / 185 / 189 / 232 / 281 / 296, 21 / 95 / 233 / 243 / 281 / 296, 21 / 95 / 294 / 296 / 307 / 335, 21 / 95 / 307, 21 / 281 / 307, 29 / 76 / 281, 56 / 76 / 95 / 232 / 243 / 281, 56 / 76 / 147 / 281 / 307, 56 / 76 / 243 / 294, 56 / 76 / 281 / 294, 56 / 76 / 296, 56 / 76 / 307, 56 / 95 / 147 / 307 / 335 / 348, 56 / 95 / 232 / 233 / 281 / 294 / 307, 56 / 95 / 243 / 281, 56 / 147 / 281, 56 / 232 / 243 / 281, 56 / 232 / 281, 56 / 232 / 281 / 294 / 296, 56 / 233 / 281 / 294 / 296, 56 / 281 / 307, 76 / 95 / 232 / 243 / 281 / 307, 76 / 95 / 243 / 281 / 307 / 335, 76 / 95 / 294 / 307, 76 / 147, 76 / 147 / 233 / 243 / 294, 76 / 147 / 233 / 281 / 294 / 307,76 / 147 / 243 / 294 / 296 / 307 / 335, 76 / 147 / 281 / 307, 76 / 189 / 296, 76 / 232 / 233 / 243 / 294 / 296 / 307, 76 / 281, 76 / 281 / 294, 76 / 294 / 296, 95 / 120, 95 / 147 / 335, 95 / 232 / 243 / 281 / 294 / 307, 95 / 232 / 281 / 294 / 296, 95 / 281 / 294 / 296, 95 / 335, 147, 147 / 225 / 232 / 243 / 281 / 296 / 307 / 335, 147 / 233 / 243 / 281 / 307, 147 / 233 / 281 / 307 / 335, 147 / 243 / 281, 147 / 307, 232 / 233 / 281 / 294 / 296 / 307, 232 / 281, 232 / 284 / 307, 233 / 243 / 281 / 296 / 307 / 335, 233 / 281 / 296 / 307, 243 / 281 / 294 / 296, 281, 281 / 294, 281 / 307, 307, and 335; and (b) expressing a proline hydroxylase polypeptide encoded by the polynucleotide.

[0251] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising an amino acid sequence selected from the group consisting of even-numbered sequences of SEQ ID NOs: 4-658 and having one or more residue differences compared to SEQ ID NO: 116 at a residue position selected from the group consisting of: 21 / 76 / 147 / 243 / 296 / 307 / 335, 56 / 76 / 147 / 281 / 307, and 95 / 147 / 335; and (b) expressing a proline hydroxylase polypeptide encoded by the polynucleotide.

[0252] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences as compared to SEQ ID NO: 162 at a residue position selected from 2 / 85 / 123 / 237, 28 / 115 / 117 / 120 / 123 / 268 / 270 / 343 / 346 / 348, 45 / 123 / 326, 65 / 117 / 120 / 123 / 343 / 346, 85 / 123 / 281 / 282, 114 / 115 / 117 / 120 / 123 / 268 / 271 / 313 / 326 / 343 / 346, 123 / 139 / 233 / 237 / 281 / 282 / 289 / 324 / 326, and 123 / 199 / 200 / 247 / 250 / 338; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0253] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences as compared to SEQ ID NO: 322 at a residue position selected from 26, 54, 61, 129, 132, 149, 156, 175, 189, 201, 209, 228, 236, 248, 262, 272, 277, 291, and 345; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0254] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences as compared to SEQ ID NO: 322 at a residue position selected from 25, 43, 54, 58, 61, 79, 129, 132, 143, 156, 163, 175, 179, 201, 209, 236, 248, 278, 291, 345, and 347; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0255] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences compared to SEQ ID NO: 322 at a residue position selected from: 85 / 117 / 120 / 135 / 208 / 270 / 324 / 343 / 346, 85 / 117 / 120 / 135 / 208 / 281 / 282 / 289, 85 / 117 / 120 / 270 / 281 / 289, 85 / 117 / 135 / 139 / 208, and 117 / 120 / 208 / 270 / 324 / 343 / 346; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0256] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences compared to SEQ ID NO: 412 at a residue position selected from: 47, 48, 56 / 118, 85, 95, 95 / 289, 113, 118, 118 / 247, 154, 162, 162 / 204, 164, 164 / 198 / 271, 168, 169, 187, 195, 243, 271, 275, 281, 314, 330, and 342; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0257] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences compared to SEQ ID NO: 412 at a residue position selected from: 25 / 129 / 163 / 236 / 262 / 345 / 347, 120 / 156 / 175 / 179 / 201, 129 / 189 / 236 / 262 / 277 / 278, 129 / 236 / 262, 156 / 175 / 179 / 228, and 162; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0258] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences as compared to SEQ ID NO: 492 at a residue position selected from: 15, 17, 28, 29, 65, 135, 167, 177, 199, 208, 228, 235, 287, 294, 307, and 343; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0259] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences as compared to SEQ ID NO: 492 at a residue position selected from: 15, 17, 28, 29, 65, 135, 167, 177, 199, 208, 228, 235, 287, 294, 307, and 343; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0260] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences as compared to SEQ ID NO: 492 at a residue position selected from: 15, 17, 28, 29, 65, 135, 167, 177, 199, 208, 228, 235, 287, 294, 307, and 343; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0261] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences compared to SEQ ID NO: 562 at a residue position selected from: 29 / 85 / 177 / 208 / 228 / 347, 29 / 85 / 208 / 228 / 343 / 347, 29 / 177 / 195 / 228 / 343, 29 / 208 / 228 / 278 / 294 / 347, 56 / 195 / 278, 85 / 187 / 205 / 208 / 278, 113 / 177 / 187 / 195 / 208 / 278 / 294 / 343 / 347, and 177 / 205 / 208 / 228; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0262] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences compared to SEQ ID NO: 598 at a residue position selected from: 47, 162, 209, 219, 227, and 342; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0263] Thus, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising a sequence selected from the even-numbered sequences of SEQ ID NO: 4-658 and having an amino acid sequence with one or more residue differences compared to SEQ ID NO: 598 at a residue position selected from: 17 / 44 / 179 / 195 / 250 / 313 / 345, 17 / 44 / 199 / 313, 43 / 44 / 195 / 199, 44 / 149 / 164 / 171 / 187, 44 / 179 / 195 / 199, 44 / 179 / 195 / 199 / 345, 79 / 163 / 164 / 171 / 187 / 201 / 286 / 288, 82 / 163 / 164, 82 / 163 / 164 / 171 / 187 / 201 / 203 / 208 / 286 / 288 / 320, 149 / 164 / 171 / 288, and 187 / 286; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0264] Accordingly, in some embodiments, a method for making an engineered proline hydroxylase polypeptide comprises: (a) synthesizing a polynucleotide encoding a polypeptide comprising an amino acid sequence of an even numbered sequence selected from SEQ ID NOs: 4-658 and having one or more residue differences as compared to SEQ ID NO: 630 at a residue position selected from: 82 / 164 / 171 / 203 / 208, 135 / 163 / 164 / 201 / 203 / 208, 162, 162 / 219 / 236, 162 / 219 / 313 / 338, 162 / 236 / 342, 162 / 313 / 342, and 164 / 171 / 201 / 203 / 282; and (b) expressing the proline hydroxylase polypeptide encoded by the polynucleotide.

[0265] In some embodiments of the method, the polynucleotide encodes an engineered proline hydroxylase optionally having one or a number (e.g., up to 3, 4, 5, or up to 10) of amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-30, 1-35, 1-40, 1-45, or 1-50 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the amino acid sequence optionally has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 amino acid residue deletions, insertions, and / or substitutions. In some embodiments, the substitutions can be conservative substitutions or non-conservative substitutions.

[0266] In some embodiments, any of the engineered proline hydroxylases expressed in a host cell can be recovered from the cells and / or culture medium using any one or more of the well-known techniques for protein purification, including, among others, lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography. Suitable solutions for lysis and efficient extraction of proteins from bacteria such as E. coli are commercially available (e.g., CelLytic BTM Sigma-Aldrich, St. Louis MO).

[0267] Chromatographic techniques for isolating proline hydroxylase polypeptides include, among others, reverse phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. Conditions for purifying a particular enzyme will depend in part on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, and will be apparent to one of skill in the art.

[0268] In some embodiments, affinity techniques can be used to isolate improved proline hydroxylases. For affinity chromatography purification, any antibody that specifically binds to a proline hydroxylase polypeptide can be used. To generate antibodies, various host animals including, but not limited to, rabbits, mice, rats, and the like, can be immunized by injection of a proline hydroxylase polypeptide or fragment thereof. The proline hydroxylase polypeptide or fragment can be attached to a suitable carrier, such as BSA, by means of a side chain functional group or a linker attached to a side chain functional group. In some embodiments, affinity purification can use specific ligands to which proline hydroxylases bind, such as poly(L-proline) or dye affinity columns (see, e.g., EP 0641862; Stellwagen, “Dye Affinity Chromatography,” in Methods in Molecular Biology, Vol. 218, Unit 9.2-9.2.16

[2001] ). Current Protocols in Protein Science

[0269] Methods using engineered proline hydroxylases

[0270] In some embodiments, the proline hydroxylases described herein can be used in a method for converting a suitable substrate to its hydroxylated product. Generally, a method for performing a hydroxylation reaction comprises contacting or incubating a substrate compound with a proline hydroxylase polypeptide of the present application in the presence of a co-substrate (such as a-ketoglutarate) under reaction conditions suitable for forming the hydroxylated product, as shown in Scheme 1 above.

[0271] In embodiments provided herein and illustrated in the Examples, the range of various suitable reaction conditions that can be used in the method include, but are not limited to, substrate loading, co-substrate loading, reducing agent, divalent transition metal, pH, temperature, buffer, solvent system, polypeptide loading, and reaction time. Additional suitable reaction conditions for a method of biocatalytically converting a substrate compound to a product compound using the engineered proline hydroxylase polypeptides described herein can be readily optimized by routine experimentation given the guidance provided herein, including but not limited to contacting an engineered proline hydroxylase polypeptide and a substrate compound under experimental reaction conditions of concentration, pH, temperature, and solvent conditions, and detecting a product compound.

[0272] ​Suitable reaction conditions using engineered proline hydroxylase polypeptides generally include a co-substrate used stoichiometrically for the hydroxylation reaction. Typically, the co-substrate for proline hydroxylases is a-ketoglutarate, also known as a-ketoglutaric acid and 2-oxoglutaric acid. Other analogs of a-ketoglutarate that are capable of being used as a co-substrate for proline hydroxylases can be used. An example analog that can be used as a co-substrate is a-oxoadipic acid. Because the co-substrate is used stoichiometrically, the co-substrate is present in an amount that is equal molar to the substrate compound or in a higher amount than the substrate compound (i.e., the molar concentration of the co-substrate is equal to or higher than the molar concentration of the substrate compound). In some embodiments, suitable reaction conditions can include a co-substrate molar concentration that is at least 1-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, or 5-fold or more of the molar concentration of the substrate compound. In some embodiments, suitable reaction conditions can include a co-substrate concentration, particularly a-ketoglutarate concentration, of about 0.001 M to about 2 M, 0.01 M to about 2 M, 0.1 M to about 2 M, 0.2 M to about 2 M, about 0.5 M to about 2 M, or about 1 M to about 2 M. In some embodiments, the reaction conditions include a co-substrate concentration of about 0.001 M, 0.01 M, 0.1 M, 0.2 M, 0.3 M, 0.4 M, 0.5 M, 0.6 M, 0.7 M, 0.8 M, 1 M, 1.5 M, or 2 M. In some embodiments, additional co-substrate can be added during the reaction.

[0273] The substrate compound in the reaction mixture can vary taking into account, for example, the amount of desired product compound, the effect of substrate concentration on enzyme activity, the stability of the enzyme under the reaction conditions, and the percentage conversion of substrate to product. In some embodiments, suitable reaction conditions include a substrate compound loading of at least about 0.5 g / L to about 200 g / L, 1 g / L to about 200 g / L, 5 g / L to about 150 g / L, about 10 g / L to about 100 g / L, 20 g / L to about 100 g / L, or about 50 g / L to about 100 g / L. In some embodiments, suitable reaction conditions include a substrate compound loading of at least about 0.5 g / L, at least about 1 g / L, at least about 5 g / L, at least about 10 g / L, at least about 15 g / L, at least about 20 g / L, at least about 30 g / L, at least about 50 g / L, at least about 75 g / L, at least about 100 g / L, at least about 150 g / L, or at least about 200 g / L or even more. While the values for substrate loading provided herein are based on the molecular weight of L-proline, it is also contemplated that various hydrates and salts of L-proline in equivalent molar amounts can also be used in the methods.

[0274] In carrying out the proline hydroxylase-mediated methods described herein, the engineered polypeptides can be added to the reaction mixture in the form of purified enzymes, partially purified enzymes, whole cells transformed with a gene encoding the enzyme, cell extracts and / or lysates of such cells, and / or as enzymes immobilized on a solid support. Whole cells transformed with a gene encoding an engineered proline hydroxylase, or cell extracts, lysates thereof, and isolated enzymes can be used in a variety of different forms, including solid (e.g., lyophilized, spray-dried, etc.) or semi-solid (e.g., crude paste). Cell extracts or cell lysates can be partially purified by precipitation (ammonium sulfate, polyethyleneimine, heat treatment, etc.) followed by a desalting procedure (e.g., ultrafiltration, dialysis, etc.) prior to lyophilization. Any enzyme preparation, including whole cell preparations, can be stabilized by cross-linking using known cross-linking agents such as, for example, glutaraldehyde, or immobilization to a solid phase (e.g., Eupergit C, etc.).

[0275] The genes encoding the engineered proline hydroxylase polypeptides can be transformed into host cells separately or together into the same host cell. For example, in some embodiments, one set of host cells can be transformed with a gene encoding one engineered proline hydroxylase polypeptide, and another set transformed with a gene encoding another engineered proline hydroxylase polypeptide. Both sets of transformed cells can be used together in a reaction mixture in the form of whole cells, or lysates or extracts derived therefrom. In other embodiments, a host cell can be transformed with genes encoding multiple engineered proline hydroxylase polypeptides. In some embodiments, the engineered polypeptides can be expressed in the form of secreted polypeptides, and the culture medium containing the secreted polypeptides can be used in the proline hydroxylase reaction.

[0276] In some embodiments, the improved activity and / or selectivity of the engineered proline hydroxylase polypeptides disclosed herein provide methods in which a higher percentage conversion can be achieved at a lower concentration of the engineered polypeptide. In some embodiments of the methods, suitable reaction conditions include an amount of engineered polypeptide of about 1% (w / w), 2% (w / w), 5% (w / w), 10% (w / w), 20% (w / w), 30% (w / w), 40% (w / w), 50% (w / w), 75% (w / w), 100% (w / w), or more of the substrate compound load.

[0277] In some embodiments, the engineered polypeptide is present at about 0.01 g / L to about 50 g / L; about 0.05 g / L to about 50 g / L; about 0.1 g / L to about 40 g / L; about 1 g / L to about 40 g / L; about 2 g / L to about 40 g / L; about 5 g / L to about 40 g / L; about 5 g / L to about 30 g / L; about 0.1 g / L to about 10 g / L; about 0.5 g / L to about 10 g / L; about 1 g / L to about 10 g / L; about 0.1 g / L to about 5 g / L; about 0.5 g / L to about 5 g / L; or about 0.1 g / L to about 2 g / L. In some embodiments, the proline hydroxylase polypeptide is present at about 0.01 g / L, 0.05 g / L, 0.1 g / L, 0.2 g / L, 0.5 g / L, 1 g / L, 2 g / L, 5 g / L, 10 g / L, 15 g / L, 20 g / L, 25 g / L, 30 g / L, 35 g / L, 40 g / L, or 50 g / L.

[0278] In some embodiments, the reaction conditions further include a divalent transition metal capable of serving as a cofactor in the oxidation reaction. Typically, the divalent transition metal cofactor is ferrous ion (i.e., Fe +2 ). The ferrous ion can be provided in various forms, such as ferrous sulfate (FeS04), ferrous chloride (FeCl2), ferrous carbonate (FeC03), and salts of organic acids, such as citrate, lactate, and fumarate. An exemplary source of ferrous sulfate is Mohr’s salt, (NH4)2Fe(S04)2, and is available in anhydrous and hydrated (i.e., hexahydrate) forms. While ferrous ion is the transition metal cofactor found in naturally occurring proline hydroxylases and is effective in the engineered enzymes, it is understood that other divalent transition metals capable of serving as a cofactor can be used in the method. In some embodiments, the divalent transition metal cofactor can include Mn +2 and Cr +2 . In some embodiments, the reaction conditions can include a divalent transition metal cofactor, particularly Fe +2 , at a concentration of about 0.1 mM to 10 mM, 0.1 mM to about 5 mM, 0.5 mM to about 5 mM, about 0.5 mM to about 3 mM, or about 1 mM to about 2 mM. In some embodiments, the reaction conditions include a divalent transition metal cofactor concentration of about 0.1 mM, 0.2 mM, 0.5 mM, 1 mM, 1.5 mM, 2 mM, 3 mM, 5 mM, 7.5 mM, or 10 mM. In some embodiments, higher concentrations can be used, for example, up to 50 mM or up to 100 mM of the divalent transition metal cofactor.

[0279] In some embodiments, the reaction conditions can further include a chelator capable of chelating the iron ion Fe +3Reduced to ferrous ions Fe +2 The reducing agent. In some embodiments, the reducing agent includes ascorbic acid, typically L-ascorbic acid. Although hydroxylation reactions do not require ascorbic acid, its presence enhances enzymatic activity. Without being bound by theory, it is believed that ascorbic acid maintains the enzyme-iron... +2 Form or enzyme-iron +2 Form regeneration, enzyme-iron +2 The form is the active form that mediates the hydroxylation reaction. Typically, reaction conditions may include an ascorbic acid concentration proportional to the substrate loading. In some embodiments, ascorbic acid is present at at least about 0.1, 0.2, 0.3, 0.5, 0.75, 1, 1.5, or at least 2 times the molar amount of the substrate. In some embodiments, the reducing agent, particularly L-ascorbic acid, is present at a concentration of about 0.001 M to about 0.5 M, about 0.01 M to about 0.5 M, about 0.01 M to about 0.4 M, about 0.1 M to about 0.4 M, or about 0.1 M to about 0.3 M. In some implementations, the reducing agent, particularly ascorbic acid, is present at a concentration of approximately 0.001M, 0.005M, 0.01M, 0.02M, 0.03M, 0.05M, 0.1M, 0.15M, 0.2M, 0.3M, 0.4M, or 0.5M.

[0280] In some embodiments, the reaction conditions include molecular oxygen (i.e., O2). Without being theoretically constrained, an oxygen atom from the molecular oxygen is incorporated into the substrate compound to form a hydroxylated product compound. O2 may be naturally present in the reaction solution or artificially introduced and / or added to the reaction. In some embodiments, the reaction conditions may include forced aeration (e.g., sparging) with air, O2 gas, or other O2-containing gases. In some embodiments, the O2 in the reaction can be increased by increasing the reaction pressure with O2 or an O2-containing gas. This can be achieved by conducting the reaction in a container that can be pressurized with O2. In some embodiments, the O2 gas may be purged through the reaction solution at a rate of at least 1 liter / hour (L / h), at least 2 L / h, at least 3 L / h, at least 4 L / h, at least 5 L / h, or greater. In some embodiments, the O2 gas may be purged through the reaction solution at a rate between about 1 L / h and 10 L / h, between about 2 L / h and 7 L / h, or between about 3 L / h and 5 L / h.

[0281] The pH of the reaction mixture can vary during the course of the reaction. The pH of the reaction mixture can be maintained at a desired pH or within a desired pH range. This can be accomplished by the addition of an acid or base prior to and / or during the course of the reaction. Alternatively, the pH can be controlled by the use of a buffer. Accordingly, in some embodiments, the reaction conditions include a buffer. Suitable buffers to maintain a desired pH range are known in the art and include, by way of example and not limitation, borate, phosphate, 2-(N-morpholino)ethanesulfonic acid (MES), 3-(N-morpholino)propanesulfonic acid (MOPS), acetate, triethanolamine, and 2-amino-2-hydroxymethyl-propane-l,3-diol (Tris), and the like. In some embodiments, the buffer is phosphate. In some embodiments of the method, suitable reaction conditions include a buffer (e.g., phosphate) concentration of about 0.01 M to about 0.4 M, 0.05 M to about 0.4 M, 0.1 M to about 0.3 M, or about 0.1 M to about 0.2 M. In some embodiments, the reaction conditions include a buffer (e.g., phosphate) concentration of about 0.01 M, 0.02 M, 0.03 M, 0.04 M, 0.05 M, 0.07 M, 0.1 M, 0.12 M, 0.14 M, 0.16 M, 0.18 M, 0.2 M, 0.3 M, or 0.4 M. In some embodiments, the reaction conditions include water as a suitable solvent in the absence of a buffer.

[0282] In embodiments of the method, the reaction conditions can include a suitable pH. The desired pH or desired pH range can be maintained by the use of an acid or base, a suitable buffer, or a combination of buffering and the addition of an acid or base. The pH of the reaction mixture can be controlled prior to and / or during the course of the reaction. In some embodiments, suitable reaction conditions include a solution pH of about 4 to about 10, a pH of about 5 to about 10, a pH of about 5 to about 9, a pH of about 6 to about 9, a pH of about 6 to about 8. In some embodiments, the reaction conditions include a solution pH of about 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10.

[0283] In embodiments of the methods herein, suitable temperatures can be used for reaction conditions taking into account, for example, increased reaction rates at higher temperatures and activity of the enzyme during the reaction period. Thus, in some embodiments, suitable reaction conditions include temperatures of about 10 °C to about 60 °C, about 10 °C to about 55 °C, about 15 °C to about 60 °C, about 20 °C to about 60 °C, about 20 °C to about 55 °C, about 25 °C to about 55 °C, or about 30 °C to about 50 °C. In some embodiments, suitable reaction conditions include temperatures of about 10 °C, 15 °C, 20 °C, 25 °C, 30 °C, 35 °C, 40 °C, 45 °C, 50 °C, 55 °C, or 60 °C. In some embodiments, the temperature during the enzymatic reaction can be maintained at a particular temperature throughout the course of the reaction. In some embodiments, the temperature during the enzymatic reaction can be adjusted to a temperature profile during the course of the reaction.

[0284] The methods of the present application are typically carried out in a solvent. Suitable solvents include water, aqueous buffer solutions, organic solvents, polymeric solvents, and / or co-solvent systems, which typically include aqueous, organic, and / or polymeric solvents. The aqueous solution (water or aqueous co-solvent system) can be pH-buffered or unbuffered. In some embodiments, the methods using engineered proline hydroxylase polypeptides can be carried out in an aqueous co-solvent system comprising an organic solvent (e.g., ethanol, isopropyl alcohol (IPA), dimethyl sulfoxide (DMSO), dimethylformamide (DMF), ethyl acetate, butyl acetate, 1-octanol, heptane, octane, methyl tert-butyl ether (MTBE), toluene, etc.), an ionic or polar solvent (e.g., 1-ethyl 4-methylimidazole tetrafluoroborate, 1-butyl-3-methylimidazole tetrafluoroborate, 1-butyl 3-methylimidazole hexafluorophosphate, glycerol, polyethylene glycol, etc.). In some embodiments, the co-solvent can be a polar solvent such as a polyol, dimethyl sulfoxide (DMSO), or a lower alcohol. The non-aqueous co-solvent component in the aqueous co-solvent system can be miscible with the aqueous component, providing a single liquid phase, or can be partially miscible or immiscible with the aqueous component, providing a two-liquid phase. Exemplary aqueous co-solvent systems can comprise water and one or more co-solvents selected from an organic solvent, a polar solvent, and a polyol solvent. Typically, the co-solvent components of the aqueous co-solvent system are selected such that they do not adversely inactivate the proline hydroxylase under the reaction conditions. Suitable co-solvent systems can be readily identified by measuring the enzymatic activity of a particular engineered proline hydroxylase in candidate solvent systems with a defined substrate of interest and applying an enzyme activity assay such as those described herein.

[0285] In some embodiments of the method, suitable reaction conditions include an aqueous co-solvent, wherein the co-solvent comprises about 1% to about 50% (v / v), about 1% to about 40% (v / v), about 2% to about 40% (v / v), about 5% to about 30% (v / v), about 10% to about 30% (v / v), or about 10% to about 20% (v / v) DMSO. In some embodiments of the method, suitable reaction conditions can include an aqueous co-solvent comprising about 1% (v / v), about 5% (v / v), about 10% (v / v), about 15% (v / v), about 20% (v / v), about 25% (v / v), about 30% (v / v), about 35% (v / v), about 40% (v / v), about 45% (v / v), or about 50% (v / v) DMSO.

[0286] In some embodiments, reaction conditions can include a surfactant for stabilizing or enhancing the reaction. Surfactants can include non-ionic, cationic, anionic, and / or amphiphilic surfactants. Exemplary surfactants include, for example, but are not limited to, nonylphenoxypolyethoxyethanol (NP40), Triton X-100, polyoxyethylene-stearic amide, cetyltrimethylammonium bromide, oleylamidomethylsulfate sodium, polyoxyethylene sorbitan monostearate, cetyl dimethyl amine, and the like. Any surfactant that can stabilize or enhance the reaction can be used. The concentration of surfactant used in the reaction can generally be from 0.1 mg / ml to 50 mg / ml, in particular from 1 mg / ml to 20 mg / ml.

[0287] In some embodiments, reaction conditions can include an antifoam agent that helps to reduce or prevent the formation of foam in the reaction solution, such as when the reaction solution is mixed or purged. Antifoam agents include non-polar oils (e.g., mineral oil, silicone, etc.), polar oils (e.g., fatty acids, alkyl amines, alkyl amides, alkyl sulfates, etc.), and hydrophobic agents (e.g., treated silica, polypropylene, etc.), some of which also act as surfactants. Exemplary antifoam agents include (Dow Corning), polyethylene glycol copolymers, oxy / ethoxylated alcohols, and polydimethylsiloxanes. In some embodiments, the antifoam agent can be present at about 0.001% (v / v) to about 5% (v / v), about 0.01% (v / v) to about 5% (v / v), about 0.1% (v / v) to about 5% (v / v), or about 0.1% (v / v) to about 2% (v / v). In some embodiments, the antifoam agent can be present at about 0.001% (v / v), about 0.01% (v / v), about 0.1% (v / v), about 0.5% (v / v), about 1% (v / v), about 2% (v / v), about 3% (v / v), about 4% (v / v), or about 5% (v / v) or more, as needed to facilitate the reaction.

[0288] The amounts of reactants for the hydroxylation enzyme reaction will generally vary depending on the amount of product desired and, concomitantly, the amount of proline hydroxylation enzyme substrate used. One of ordinary skill in the art will readily understand how to vary these amounts to suit the desired level of productivity and production scale.

[0289] In some embodiments, the order of addition of the reactants is not critical. The reactants can be added together at the same time into the solvent (e.g., a single-phase solvent, a biphasic aqueous co-solvent system, etc.), or alternatively, some reactants can be added separately, and some together, at different points in time. For example, the cofactor, co-substrate, proline hydroxylation enzyme, and substrate can be added first into the solvent.

[0290] Solid reactants (e.g., enzymes, salts, etc.) can be provided to the reaction in a variety of different forms, including as a powder (e.g., lyophilized, spray-dried, etc.), a solution, an emulsion, a suspension, etc. The reactants can be readily lyophilized or spray-dried using methods and equipment known to one of ordinary skill in the art. For example, a protein solution can be frozen in small aliquots at -80°C, then added to a pre-cooled lyophilization chamber, followed by application of vacuum.

[0291] For improved mixing efficiency when using an aqueous co-solvent system, the proline hydroxylation enzyme and cofactor can be added and mixed into the aqueous phase first. The organic phase can then be added and mixed into it, followed by the addition of the proline hydroxylation enzyme substrate and co-substrate. Alternatively, the proline hydroxylation enzyme substrate can be pre-mixed in the organic phase before being added to the aqueous phase.

[0292] The hydroxylation process is generally allowed to proceed until further conversion of the substrate to hydroxylated products is not significantly changed (e.g., less than 10% of the substrate is converted, or less than 5% of the substrate is converted) with reaction time. In some embodiments, the reaction is allowed to proceed until complete or near complete conversion of the substrate to product is achieved. Conversion of the substrate to product can be monitored using known methods by detecting the substrate and / or product, with or without derivatization. Suitable analytical methods include gas chromatography, HPLC, MS, and the like.

[0293] In some embodiments of the method, suitable reaction conditions include a substrate loading of at least about 5 g / L, 10 g / L, 20 g / L, 30 g / L, 40 g / L, 50 g / L, 60 g / L, 70 g / L, 100 g / L, or more, and wherein the method results in at least about 50%, 60%, 70%, 80%, 90%, 95%, or more conversion of the substrate compound to the product compound in about 48 h or less, in about 36 h or less, or in about 24 h or less.

[0294] When used in the method under suitable reaction conditions, the engineered proline hydroxylase polypeptides of the application produce an excess of the isomeric excess of trans-3- hydroxylated product over trans-4-hydroxylated product of at least 90%, 95%, 96%, 97%, 98%, 99%, or more. In some embodiments, no detectable amount of the trans-4- hydroxylated product compound is formed.

[0295] In additional embodiments of the methods of using engineered proline hydroxylase polypeptides to convert a substrate compound to a hydroxylated product compound, suitable reaction conditions can include an initial substrate loading in the reaction solution that is then contacted with the polypeptide. This reaction solution is then further supplemented with additional substrate compound, the additional substrate compound being added as a continuous or batch addition over time at a rate of at least about 1 g / L / h, at least about 2 g / L / h, at least about 4 g / L / h, at least about 6 g / L / h, or more. Thus, according to these suitable reaction conditions, the polypeptide is added to a solution having an initial substrate loading of at least about 20 g / L, 30 g / L, or 40 g / L. After the polypeptide is added, additional substrate is then continuously added to the solution at a rate of about 2 g / L / h, 4 g / L / h, or 6 g / L / h until a much higher final substrate loading of at least about 30 g / L, 40 g / L, 50 g / L, 60 g / L, 70 g / L, 100 g / L, 150 g / L, 200 g / L, or more is reached. Thus, in some embodiments of the method, suitable reaction conditions include adding the polypeptide to a solution having an initial substrate loading of at least about 20 g / L, 30 g / L, or 40 g / L, and then adding additional substrate to the solution at a rate of about 2 g / L / h, 4 g / L / h, or 6 g / L / h until a final substrate loading of at least about 30 g / L, 40 g / L, 50 g / L, 60 g / L, 70 g / L, 100 g / L, or more is reached. Such substrate-supplemented reaction conditions allow for reaching higher substrate loadings while maintaining high conversion of substrate to hydroxylated product of at least about 50%, 60%, 70%, 80%, 90%, or more of substrate conversion. In some embodiments of the method, the substrate that is added is in a solution that contains alpha-ketoglutarate in equimolar amounts or more than the additional substrate that is added.

[0296] In some embodiments of the method, the reaction using the engineered proline hydroxylase polypeptide can include the following suitable reaction conditions: (a) a substrate loading of about 60 g / L; (b) about 6 g / L of the engineered polypeptide; (c) about 1.2 molar equivalents of alpha-ketoglutarate of the substrate compound; (d) about 10 mM ascorbic acid; (e) about 4 mM FeS04; (f) a pH of about 6.8; (g) a temperature of about 20 °C; and (h) a reaction time of about 24 hours.

[0297] In some embodiments, additional reaction components or additional techniques are performed to supplement the reaction conditions. These can include taking measures to stabilize the enzyme or prevent inactivation of the enzyme, reduce product inhibition, shift the reaction equilibrium toward hydroxylated product formation.

[0298] In additional embodiments, any of the above-described methods for converting a substrate compound to a product compound can further comprise one or more steps selected from the group consisting of: extraction; isolation; purification; and crystallization of the product compound. Methods, techniques, and protocols for extracting, isolating, purifying, and / or crystallizing hydroxylated products from a biocatalytic reaction mixture produced by the above-disclosed methods are known to the ordinary skilled artisan and / or are obtainable by routine experimentation. Moreover, illustrative methods are provided in the Examples below.

[0299] Various features and embodiments of the present application are illustrated in the following representative examples, which are intended to be illustrative and not limiting.

[0300] Experiments

[0301] The following examples, including experiments and results obtained, are provided for illustrative purposes only and should not be construed as limiting the application.

[0302] In the experimental disclosure that follows, the following abbreviations have the following meanings: ppm (parts per million); M (moles / liter); mM and μΜ (millimoles / liter); nM (nanomoles / liter); mol (moles); gm and g (grams); mg (milligrams); ug and pg (micrograms); L and I (liters); ml and mL (milliliters); cm (centimeters); mm (millimeters); um and pm (micrometers); sec. (seconds); min(s) (minutes); h(s) and hr(s) (hours); U (units); MW (molecular weight); rpm (revolutions per minute); °C (Celsius); CDS (coding sequence); DNA (deoxyribonucleic acid); RNA (ribonucleic acid); NA (nucleic acid; polynucleotide); AA (amino acid; polypeptide); E. coli W3110 (a commonly used laboratory strain of E. coli, available from the Coli Genetic Stock Center [CGSC], New Haven, CT); HPLC (high pressure liquid chromatography); SDS-PAGE (sodium dodecyl sulfate polyacrylamide gel electrophoresis); PES (polyether sulfone); CFSE (carboxyfluorescein succinimidyl ester); IPTG (isopropyl β-D-l-thiogalactopyranoside); PMBS (polymyxin B sulfate); NADPH (nicotinamide adenine dinucleotide phosphate); GDH (glucose dehydrogenase); polyethylenimine (PEI); FIOPC (fold improvement over positive control); DO (dissolved oxygen); ESI (electrospray ionization); LB (Luria broth); TB (terrific broth); MeOH (methanol); HTP (high throughput); SFP (shake flask powder); DSP (downstream process powder); Athens Research (Athens Research Technology, Athens, GA); ProSpec (ProSpec Tany Technogene, East Brunswick, NJ); Sigma-Aldrich (Sigma-Aldrich, St. Louis, MO); Ram Scientific (Ram Scientific, Inc., Yonkers, NY); Pall Corp. (Pall, Corp., Pt. Washington, NY); Millipore (Millipore, Corp.Billerica MA); Difco (Difco Laboratories, BD Diagnostic Systems, Detroit, MI); Molecular Devices (Molecular Devices, LLC, Sunnyvale, CA); Kuhner (Adolf Kuhner, AG, Basel, Switzerland); Cambridge Isotope Laboratories, (Cambridge Isotope Laboratories, Inc., Tewksbury, MA); Applied Biosystems (part of Life Technologies, Corp., Grand Island, NY); Agilent (Agilent Technologies, Inc., Santa Clara, CA); Thermo Scientific (part of Thermo Fisher Scientific, Waltham, MA); Fisher (Fisher Scientific, Waltham, MA); Corning (Corning, Inc., Palo Alto, CA); Waters (Waters Corp., Milford, MA); GE Healthcare (GE Healthcare Bio-Sciences, Piscataway, NJ); Pierce (Pierce Biotechnology, now part of Thermo Fisher Scientific, Rockford, IL); Phenomenex (Phenomenex, Inc., Torrance, CA); Optimal (Optimal Biotech Group, Belmont, CA); and Bio-Rad (Bio-Rad Laboratories, Hercules, CA).

[0303] Example 1

[0304] Escherichia coli expression hosts containing recombinant proline hydroxylase genes

[0305] The initial proline hydroxylase (PH) used to generate variants of the application was obtained from the wild-type ANO sequence from the fungus sp. No. 11243 (Accession No. GAM84982). The wild-type PH protein sequence was codon-optimized for expression in E. coli and the DNA was cloned into the expression vector pCK110900 (see Figure 3 of U.S. Patent Application Publication No. 2006 / 0195947), operably linked to the lac promoter under the control of the lacI repressor. The expression vector also contains the P15a origin of replication and a chloramphenicol resistance gene. The resulting plasmid was transformed into E. coli W3110 using standard methods known in the art. Transformants were isolated by subjecting the cells to chloramphenicol selection, as known in the art (see, e.g., U.S. Patent No. 8,383,346 and WO 2010 / 144103).

[0306] Example 2

[0307] Preparation of HTP PH-containing wet cell pellets and lysates

[0308] E. coli cells containing the recombinant PH-encoding gene from a monoclonal colony were inoculated into 180 μΐ of LB containing 1% glucose and 30 μg / mL chloramphenicol (CAM) in the wells of a 96-well shallow well microtiter plate. The plate was sealed with an O2-permeable seal and the cultures were grown overnight at 30°C, 200 rpm, and 85% humidity. Then, 10 μΐ of each cell culture was transferred to a well of a 96-well deep well plate containing 390 mL TB and 30 μg / mL CAM. The deep well plate was sealed with an O2-permeable seal and incubated at 30°C, 250 rpm, and 85% humidity until an OD 600 0.6-0.8 was reached. The cell cultures were then induced with IPTG to a final concentration of 1 mM and incubated overnight at 20°C or 30°C. The cells were then pelleted by centrifugation at 4000 rpm for 10 min. The supernatant was discarded and the pellets were frozen at -80°C prior to lysis.

[0309] For lysis, 400 μΐ of lysis buffer containing 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L polymyxin B sulfate (PMBS) was added to the cell paste in each well produced as described in Example 2. The cells were lysed at room temperature for 2 hours with shaking on a bench top shaker. The plate was then centrifuged at 4000 rpm and 4°C for 15 min. The clear supernatant was then used in biocatalytic reactions to determine their level of activity.

[0310] Example 3

[0311] Preparation of freeze-dried lysates from shake flask (SF) cultures

[0312] Selected HTP cultures grown as described above were plated onto LB agar plates with 1% glucose and 30 pg / ml CAM and grown overnight at 37°C. Individual colonies from each culture were transferred to 6 ml LB with 1% glucose and 30 pg / ml CAM. Cultures were grown at 30°C, 250 rpm for 18 h and subcultured at approximately 1 :50 into 250 ml TB containing 30 pg / ml CAM to a final OD of 0.05 600 Cultures were grown at 30°C, 250 rpm for approximately 195 minutes to an OD of between 0.6-0.8 600 and induced with 1 mM IPTG. Cultures were then grown at 20°C or 30°C, 250 rpm for 20 h. Cultures were centrifuged at 4000 rpm for 20 min. The supernatant was discarded and the pellet was resuspended in 30 ml of 20 mM triethanolamine pH 7.5 and processed using a Processor system (Microfluidics) at 18,000 psi lysis. The lysate was pelleted (10,000 rpm for 60 min) and the supernatant was frozen and lyophilized to produce the shake flask (SF) enzyme.

[0313] Example 4

[0314] Improvements in the conversion of proline substrates to trans-3-hydroxyproline compared to SEQ ID NO: 4

[0315] Based on the results of the screen to select variants that convert L-proline substrates to trans-3-hydroxyproline, SEQ ID NO: 4 was selected as the parent enzyme. SEQ ID NO: 4 is identical to SEQ ID NO: 2; both of these sequences are wild-type proline hydroxylases to which an N-terminal his-tag has been added, while SEQ ID NO: 3 is a codon-optimized polynucleotide encoding a wild-type proline hydroxylase. Libraries of engineered genes were generated using well-established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). Polypeptides encoded by each gene were produced in HTP as described in Example 2 (where the proteins were expressed overnight at 20°C). For all variants, the cell pellet was lysed by adding 200 pL lysis buffer (containing 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L PMBS) and shaking on a benchtop shaker for 2 hours at room temperature. The plate was centrifuged at 4000 rpm for 15 minutes at 4°C to remove cell debris.

[0316] In 300 pL round bottom plates, 50 pL of E. coli lysate was added to 200 pL of reaction mixture (comprising 75 pL 63 g / L a-ketoglutaric acid [in 50 mM sodium phosphate pH 6.5], 50 pL 20 mM Morita salt [in 50 mM sodium phosphate pH 6.5] and 75 pL 33 g / L proline in 65 mM ascorbic acid) in each well. Plates were sealed with AirPore seals (Qiagen) and reactions were continued in Kuhner at 30 °C, 200 rpm and 85% relative humidity, with 2" swing at 18 hours overnight.

[0317] After overnight incubation, the reaction in each well was derivatized and quenched by aliquoting 25 pL of reaction mixture into a 96-well deep well plate containing 225 pL of derivatization solution (comprising 75 pL saturated sodium bicarbonate, 25 pL water and 125 pL 2.5 mg / mL FmocCl [in ACN] per well). After shaking at room temperature for 1 hr, the plate was centrifuged at 4000 rpm for one minute and 40 pL of the soluble fraction of the quenched reaction was mixed with 160 pL 1 : 1 ACN:0.5 M HC1. The derivatized and diluted samples were analyzed as described in Table 13.1. Selectivity relative to SEQ ID NO:4 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO:4. Activity relative to SEQ ID NO:4 (activity FIOP) was calculated as the ratio of the peak area of trans-3-hydroxyproline of the variant compared to the peak area of trans-3-hydroxyproline produced by SEQ ID NO:4. Results are shown in Table 4.1 and Table 4.2.

[0318]

[0319]

[0320]

[0321]

[0322] In addition to HTP analysis, a subset of selected beneficial variants from the HTP screen were also prepared at shake flask scale, as described in Example 3 (wherein the proteins were expressed overnight at 20°C). 1 mL scale reaction tests were performed on lyophilized shake flask lysate powder (SFP) under the following conditions: 10 g / L L-proline, 50 wt% proline hydroxylase variant SFP, 1.5 equivalents a-KG (a-ketoglutarate), 0.15 equivalents ascorbic acid, 4 mM iron (II) sulfate ammonium hexahydrate, 50 mM sodium phosphate pH 6.5, air, and room temperature. The reactions were run overnight and analyzed using similar methods as described above for the HTP reactions. Selectivity relative to SEQ ID NO: 4 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline: trans-4-hydroxyproline product formed by the variant compared to the ratio produced by SEQ ID NO: 4.

[0323] Results are shown in Table 4.3.

[0324]

[0325] The proline hydroxylase protein produced by SEQ ID NO: 4 is not fully stable under standard expression conditions at 30°C. As described above, for all data shown in Tables 4.1, 4.2, and 4.3, HTP and shake flask proteins were produced at 20°C expression. To select more stable and more active variants, a library of engineered genes derived from SEQ ID NO: 4 was produced using established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations) and polypeptides encoded by each gene were produced at HTP at 30°C expression. Reactions, derivatization, and analysis were performed as described above. Stability and activity relative to SEQ ID NO: 4 (stability / activity FIOP) was calculated as the ratio of the peak area of trans-3-hydroxyproline for the variant compared to the peak area of trans-3-hydroxyproline produced by SEQ ID NO: 4, where both enzymes were produced at 30°C. Results are shown in Table 4.4.

[0326]

[0327] Example 5

[0328] Improvements in conversion of proline substrate to trans-3-hydroxyproline compared to SEQ ID NO: 116

[0329] A library of engineered genes was generated using established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations) from an engineered polynucleotide (SEQ ID NO: 115) encoding a polypeptide having proline hydroxylase activity of SEQ ID NO: 116. Polypeptides encoded by each gene were produced HTP as described in Example 2 (where proteins were expressed overnight at 20°C). For all variants, cell pellets were lysed by adding 200 pL lysis buffer (containing 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L PMBS) and shaking on a benchtop shaker for 2 hours at room temperature. Plates were centrifuged at 4000 rpm for 15 minutes at 4°C to remove cell debris.

[0330] In 300 pL round bottom plates, 50 pL of E. coli lysate was added to 200 pL of reaction mixture (containing 75 pL of 63 g / L a-ketoglutaric acid [in 50 mM sodium phosphate pH 6.5], 50 pL of 20 mM Molar Salt [in 50 mM sodium phosphate pH 6.5], and 75 pL of 33 g / L proline in 65 mM ascorbic acid) in each well. Plates were sealed with AirPore seals (Qiagen) and reactions were continued in Kuhner at 30°C, 200 rpm, and 85% relative humidity, with a 2” swing at room temperature overnight (~18 hours).

[0331] Following overnight incubation, reactions were derivatized and quenched by aliquoting 25 pL of reaction mixture into 96-well deep well plates containing 225 pL of derivatization solution (containing 75 pL of saturated sodium bicarbonate, 25 pL of water, and 125 pL of 2.5 mg / mL FmocCl [in ACN] per well). After shaking at room temperature for 1 hr, plates were centrifuged for one minute at 4000 rpm and 40 pL of soluble fraction of the quenched reaction was mixed with 160 pL of 1 : 1 ACN:0.5 M HC1. Derivatized and diluted samples were analyzed as described in Table 13.1. Selectivity (selectivity FIOP) relative to SEQ ID NO: 116 was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO: 116. Results are shown in Table 5.1.

[0332]

[0333] The proline hydroxylase protein produced from SEQ ID NO: 116 was not fully stable under standard expression conditions at 30°C. To select for more stable and more active variants, a library of engineered genes derived from SEQ ID NO: 116 was generated using well-established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations) and expressed in HTP at 30°C to produce the polypeptide encoded by each gene. Reactions, derivatization, and analysis were completed as described above. Stability and activity (stability / activity FIOP) relative to SEQ ID NO: 116 was calculated as the ratio of the peak area of trans-3-hydroxyproline for the variant compared to the peak area of trans-3-hydroxyproline produced by SEQ ID NO: 116, where both enzymes were produced at 30°C. Results are shown in Table 5.2.

[0334]

[0335]

[0336]

[0337] In addition to HTP analysis, a select subset of beneficial variants from the HTP screen were also prepared at shake flask scale as described in Example 3. SFP for SEQ ID NO: 116 was produced at 20°C and variants derived from SEQ ID NO: 116 were produced at 30°C. Lyophilized shake flask lysate powder (SFP) was tested in 1 mL scale reaction tests under the following conditions: 20 g / L L-proline, 5 wt% proline hydroxylase variant SFP, 1.5 equivalents a-KG (a-ketoglutarate), 0.15 equivalents ascorbic acid, 4 mM iron (II) sulfate ammonium hexahydrate, 50 mM sodium phosphate pH 6.5, air, and room temperature. Reactions were run overnight and analyzed using similar methods as described above for HTP reactions. Stability and activity (stability / activity FIOP) relative to SEQ ID NO: 116 was calculated as the ratio of the peak area of trans-3-hydroxyproline for the variant compared to the peak area of trans-3-hydroxyproline produced by SEQ ID NO: 116, where SEQ ID NO: 116 was produced at 20°C and variants derived from SEQ ID NO: 116 were produced at 30°C. Results are shown in Table 5.3.

[0338]

[0339] Example 6

[0340] Improvements in conversion of proline substrate to trans-3-hydroxyproline compared to SEQ ID NO: 162

[0341] A library of engineered genes was generated using established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations) from an engineered polynucleotide (SEQ ID NO: 161) encoding a polypeptide having proline hydroxylase activity of SEQ ID NO: 162. Polypeptides encoded by each gene were produced HTP as described in Example 2 (where proteins were expressed overnight at 30°C). For all variants, cell pellets were lysed by adding 400 pL lysis buffer (containing 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L PMBS) and shaking on a benchtop shaker for 2 hours at room temperature. Plates were centrifuged at 4000 rpm for 15 minutes at 4°C to remove cell debris.

[0342] In 300 pL round bottom plates, 50 pL of E. coli lysate was added to 200 pL of reaction mixture (containing 75 pL of 133 g / L a-ketoglutarate [in 50 mM sodium phosphate pH 6.5], 50 pL of 20 mM Molar Salt [in 50 mM sodium phosphate pH 6.5], and 75 pL of 67 g / L proline in 65 mM ascorbic acid) in each well. Plates were sealed with AirPore seals (Qiagen) and reactions were continued in Kuhner at 30°C, 200 rpm, and 85% relative humidity, with a 2” swing at room temperature overnight (~18 hours).

[0343] Following overnight incubation, reactions were derivatized and quenched by aliquoting 25 pL of reaction mixture into 96-well deep well plates containing 225 pL of derivatization solution (containing 75 pL of saturated sodium bicarbonate, 25 pL of water, and 125 pL of 2.5 mg / mL FmocCl [in ACN] per well). After shaking at room temperature for 1 hr, plates were centrifuged for one minute at 4000 rpm and 40 pL of soluble fraction of the quenched reaction was mixed with 160 pL of 1 : 1 ACN:0.5 M HC1. Derivatized and diluted samples were analyzed as described in Table 13.1. Selectivity (selectivity FIOP) relative to SEQ ID NO: 162 was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO: 162.

[0344] In addition to HTP analysis, a subset of selected beneficial variants from the HTP screen were also prepared at shake flask scale, expressed at 30°C as described in Example 3. Lyophilized shake flask lysate powder (SFP) was tested in 1 mL scale reactions under the following conditions: 40 g / L L-proline, 5 wt% proline hydroxylase variant SFP, 1.2 equivalents a-KG (a-ketoglutarate), 25 mM ascorbic acid, 4 mM iron (II) sulfate ammonium hexahydrate, 50 mM sodium phosphate pH 6.5, air, and room temperature. Reactions were run overnight, and analyzed using similar methods described above for HTP reactions. Selectivity relative to SEQ ID NO: 162 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline: trans-4-hydroxyproline produced by the variant relative to the ratio produced by SEQ ID NO: 162. Results are shown in Table 6.1.

[0345]

[0346]

[0347] Example 7

[0348] Improvements in proline substrate conversion to trans-3-hydroxyproline relative to SEQ ID NO: 322

[0349] A library of engineered genes was generated from an engineered polynucleotide encoding a polypeptide having proline hydroxylase activity of SEQ ID NO: 322 (SEQ ID NO: 321) using established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). Polypeptides encoded by each gene were produced in HTP as described in Example 2 (with proteins expressed overnight at 30°C). For all variants, cell pellets were lysed by adding 400 pL lysis buffer (containing 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L PMBS) and shaking on a benchtop shaker at room temperature for 2 hours. Plates were centrifuged at 4000 rpm for 15 minutes at 4°C to remove cell debris.

[0350] In 300 pL round bottom plates, 50 pL of E. coli lysate was added to 200 pL of reaction mixture (containing 75 pL of 266 g / L a-ketoglutarate [in 50 mM sodium phosphate pH 6.5], 50 pL of 20 mM Molar Salt [in 50 mM sodium phosphate pH 6.5], and 75 pL of 133 g / L proline) in each well. Plates were sealed with AirPore seals (Qiagen) and reactions were continued in a Kuhner at 30°C, 200 rpm, and 85% relative humidity, with a 2” throw, overnight (~18 hours).

[0351] After overnight incubation, the reaction of each well was derivatized and quenched by aliquoting 25 pL of the reaction mixture into a 96-well deep well plate containing 225 pL of derivatization solution (75 pL of saturated sodium bicarbonate, 25 pL of water, and 125 pL of 2.5 mg / mL FmocCl [in ACN] per well). After shaking at room temperature for 1 hr, the plate was centrifuged at 4000 rpm for one minute, and 40 pL of the soluble fraction of the quenched reaction was mixed with 160 pL of 1 : 1 ACN:0.5 M HC1. The derivatized and diluted samples were analyzed as described in Table 13.1. The selectivity relative to SEQ ID NO: 322 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO: 322. The activity relative to SEQ ID NO: 322 (activity FIOP) was calculated as the ratio of the peak area of trans-3-hydroxyproline of the variant compared to the peak area of trans-3-hydroxyproline produced by SEQ ID NO: 322, and the results are shown in Tables 7.1 and 7.2, respectively.

[0352]

[0353]

[0354] In addition to the HTP analysis, a select subset of beneficial variants from the HTP screen were also prepared at shake flask scale, expressed as described in Example 3 and at 30 °C. The lyophilized shake flask lysate powder (SFP) was tested in 1 mL scale reactions under the following conditions: 40 g / L L-proline, 5 wt% proline hydroxylase variant SFP, 1.2 equivalents a-KG (a-ketoglutarate), 10 mM ascorbic acid, 4 mM iron (II) sulfate ammonium hexahydrate, 50 mM sodium phosphate pH 6.5, air, and room temperature. The reactions were run overnight and analyzed using similar methods as described above for the HTP reactions. The selectivity relative to SEQ ID NO: 322 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO: 322. The results are shown in Table 7.3.

[0355]

[0356] Example 8

[0357] Improvements in the conversion of proline substrate to trans-3-hydroxyproline compared to SEQ ID NO: 412

[0358] A library of engineered genes was generated using established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations) from an engineered polynucleotide (SEQ ID NO: 411) encoding a polypeptide having proline hydroxylase activity of SEQ ID NO: 412. Polypeptides encoded by each gene were produced HTP as described in Example 2 (where proteins were expressed overnight at 30°C). For all variants, cell pellets were lysed by adding 400 pL lysis buffer (comprising 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L PMBS) and shaking on a benchtop shaker for 2 hours at room temperature. Plates were centrifuged at 4000 rpm for 15 minutes at 4°C to remove cell debris.

[0359] In 300 pL round bottom plates, 50 pL of E. coli lysate was added to 200 pL of reaction mixture (comprising 75 pL 266 g / L a-ketoglutarate [in 50 mM sodium phosphate pH 6.5], 50 pL 20 mM Molar salt [in 50 mM sodium phosphate pH 6.5], and 75 pL 133 g / L proline in 65 mM ascorbic acid) in each well. Plates were sealed with AirPore seals (Qiagen) and reactions were continued in Kuhner at 30°C, 200 rpm, and 85% relative humidity, with a 2” swing at room temperature overnight (~18 hours).

[0360] Following overnight incubation, reactions were derivatized and quenched by aliquoting 25 pL of reaction mixture into 96-well deep well plates containing 225 pL derivatization solution (comprising 75 pL saturated sodium bicarbonate, 25 pL water, and 125 pL 2.5 mg / mL FmocCl [in ACN] per well). Following 1 hr of shaking at room temperature, plates were centrifuged for one minute at 4000 rpm and 40 pL of the soluble fraction of the quenched reaction was mixed with 160 pL 1 : 1 ACN:0.5 M HC1. Derivatized and diluted samples were analyzed as described in Table 13.1. Selectivity (selectivity FIOP) relative to SEQ ID NO: 412 was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO: 412. Results are shown in Table 8.1.

[0361]

[0362]

[0363] In addition to HTP analysis, a subset of selected beneficial variants from the HTP screen were also prepared at shake flask scale, expressed at 30°C as described in Example 3. Lyophilized shake flask lysate powder (SFP) was tested in 1 mL scale reactions under the following conditions: 60 g / L L-proline, 5 wt% proline hydroxylase variant SFP, 1.2 equivalents a-KG (a-ketoglutarate), 10 mM ascorbic acid, 4 mM iron (II) sulfate ammonium hexahydrate, 50 mM sodium phosphate pH 6.8, air, and room temperature. Reactions were run overnight, and analyzed using similar methods described above for HTP reactions. Selectivity relative to SEQ ID NO: 412 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline: trans-4-hydroxyproline produced by the variant relative to the ratio produced by SEQ ID NO: 412. Results are shown in Table 8.2.

[0364]

[0365] Example 9

[0366] Improvements in proline substrate conversion to trans-3-hydroxyproline relative to SEQ ID NO: 492

[0367] A library of engineered genes was generated from an engineered polynucleotide encoding a polypeptide having proline hydroxylase activity of SEQ ID NO: 492 (SEQ ID NO: 491) using established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). Polypeptides encoded by each gene were produced in HTP as described in Example 2 (where proteins were expressed overnight at 30°C). For all variants, cell pellets were lysed by adding 600 pL lysis buffer (containing 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L PMBS) and shaking on a benchtop shaker at room temperature for 2 hours. Plates were centrifuged at 4000 rpm for 15 minutes at 4°C to remove cell debris.

[0368] In 300 pL round bottom plates, 50 pL of E. coli lysate was added to 200 pL of reaction mixture (containing 75 pL 667 g / L a-ketoglutarate [in 50 mM sodium phosphate pH 6.5], 50 pL 20 mM Molar Salt [in 50 mM sodium phosphate pH 6.5], and 75 pL 333 g / L proline in 65 mM ascorbic acid) in each well. Plates were sealed with AirPore seals (Qiagen) and reactions were continued in Kuhner at 30°C, 200 rpm, and 85% relative humidity, with a 2” swing at room temperature overnight (~18 hours).

[0369] After overnight incubation, the reaction of each well was derivatized and quenched by aliquoting 25 pL of reaction mixture into a 96-well deep well plate containing 225 pL of derivatization solution (75 pL of saturated sodium bicarbonate, 25 pL of water, and 125 pL of 2.5 mg / mL FmocCI [in ACN] per well). After shaking at room temperature for 1 hr, the plate was centrifuged at 4000 rpm for one minute, and 40 pL of the soluble fraction of the quenched reaction was mixed with 160 pL of 1 : 1 ACN:0.5 M HC1. The derivatized and diluted samples were analyzed as described in Table 13.1. Selectivity relative to SEQ ID NO:492 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO:492. Results are shown in Table 9.1.

[0370]

[0371]

[0372] In addition to HTP analysis, a select subset of beneficial variants from the HTP screen were also prepared at shake flask scale, expressed as described in Example 3 and at 30°C. Lyophilized shake flask lysate powder (SFP) was tested in 1 mL scale reactions under the following conditions: 60 g / L L-proline, 5 wt% proline hydroxylase variant SFP, 1.2 equivalents a-KG (a-ketoglutarate), 10 mM ascorbic acid, 4 mM iron (II) sulfate ammonium hexahydrate, 50 mM sodium phosphate pH 6.8, air, and room temperature. Reactions were run overnight and analyzed using similar methods as described above for HTP reactions. Selectivity relative to SEQ ID NO:492 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO:492. Results are shown in Table 9.2.

[0373]

[0374] Example 10

[0375] Improvements in conversion of proline substrate to trans-3-hydroxyproline compared to SEQ ID NO: 562

[0376] A library of engineered genes was generated using established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations) from an engineered polynucleotide (SEQ ID NO: 561) encoding a polypeptide having proline hydroxylase activity of SEQ ID NO: 562. Polypeptides encoded by each gene were produced HTP as described in Example 2 (where proteins were expressed overnight at 30°C). For all variants, cell pellets were lysed by adding 400 pL lysis buffer (containing 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L PMBS) and shaking on a benchtop shaker for 2 hours at room temperature. Plates were centrifuged at 4000 rpm for 15 minutes at 4°C to remove cell debris.

[0377] In 300 pL round bottom plates, 50 pL of E. coli lysate was added to 200 pL of reaction mixture (containing 75 pL 400 g / L a-ketoglutarate [in 50 mM sodium phosphate pH 6.5], 50 pL 20 mM Molar salt [in 50 mM sodium phosphate pH 6.5], and 200 pL 200 g / L proline in 65 mM ascorbic acid) in each well. Plates were sealed with AirPore seals (Qiagen) and reactions were continued in Kuhner at 30°C, 200 rpm, and 85% relative humidity, with a 2” swing at room temperature overnight (~18 hours).

[0378] Following overnight incubation, the reaction in each well was derivatized and quenched by aliquoting 25 pL of the reaction into a 96-well deep well plate containing 225 pL of derivatization solution (containing 75 pL saturated sodium bicarbonate, 25 pL water, and 125 pL 2.5 mg / mL FmocCl [in ACN] per well). After shaking at room temperature for 1 hr, plates were centrifuged for one minute at 4000 rpm and 40 pL of the soluble fraction of the quenched reaction was mixed with 160 pL 1 : 1 ACN:0.5 M HC1. The derivatized and diluted samples were analyzed as described in Table 13.1. Selectivity (selectivity FIOP) was calculated relative to SEQ ID NO: 562 as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO: 562. Results are shown in Table 10.1.

[0379]

[0380]

[0381] In addition to HTP analysis, a subset of selected beneficial variants from the HTP screen were also prepared at shake flask scale, expressed at 30°C as described in Example 3. Lyophilized shake flask lysate powder (SFP) was tested in 1 mL scale reactions under the following conditions: 60 g / L L-proline, 5 wt% proline hydroxylase variant SFP, 1.2 equivalents a-KG (a-ketoglutarate), 10 mM ascorbic acid, 4 mM iron (II) sulfate ammonium hexahydrate, 50 mM sodium phosphate pH 6.8, air, and room temperature. Reactions were run overnight and analyzed using similar methods described above for HTP reactions. Selectivity relative to SEQ ID NO: 562 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline: trans-4-hydroxyproline product formed by the variant compared to the ratio produced by SEQ ID NO: 562. Results are shown in Table 10.2.

[0382]

[0383]

[0384] Example 11

[0385] Improvements in proline substrate conversion to trans-3-hydroxyproline compared to SEQ ID NO: 598

[0386] A library of engineered genes was generated from an engineered polynucleotide encoding a polypeptide having proline hydroxylase activity of SEQ ID NO: 598 (SEQ ID NO: 597) using established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations). Polypeptides encoded by each gene were produced in HTP as described in Example 2 (with proteins expressed overnight at 30°C). For all variants, cell pellets were lysed by adding 200 pL lysis buffer (containing 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L PMBS) and shaking on a benchtop shaker at room temperature for 2 hours. Plates were centrifuged at 4000 rpm for 15 minutes at 4°C to remove cell debris.

[0387] In 300 pL round bottom plates, 50 pL of E. coli lysate was added to 200 pL of reaction mixture (containing 75 pL 267 g / L a-ketoglutarate [in 50 mM sodium phosphate pH 6.5], 50 pL 20 mM Molar Salt [in 50 mM sodium phosphate pH 6.5] in 65 mM ascorbic acid, and 200 pL 133 g / L proline) in each well. Plates were sealed with AirPore seals (Qiagen) and reactions were continued in a Kuhner at 30°C, 200 rpm, and 85% relative humidity, with a 2” swing at room temperature overnight (~18 hours).

[0388] After overnight incubation, the reaction of each well was derivatized and quenched by aliquoting 25 pL of the reaction mixture into a 96-well deep well plate containing 225 pL of derivatization solution (75 pL of saturated sodium bicarbonate, 25 pL of water, and 125 pL of 2.5 mg / mL FmocCl [in ACN] per well). After shaking at room temperature for 1 hr, the plate was centrifuged at 4000 rpm for one minute, and 40 pL of the soluble fraction of the quenched reaction was mixed with 160 pL of 1 : 1 ACN:0.5 M HC1. The derivatized and diluted samples were analyzed as described in Table 13.1. Selectivity relative to SEQ ID NO:598 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO:598. Results are shown in Table 11.1.

[0389]

[0390] In addition to the HTP analysis, a select subset of beneficial variants from the HTP screen were also prepared at shake flask scale, expressed as described in Example 3 and at 30°C. The lyophilized shake flask lysate powder (SFP) was tested in 1 mL scale reactions under the following conditions: 60 g / L L-proline, 5 wt% proline hydroxylase variant SFP, 1.2 equivalents a-KG (a-ketoglutarate), 10 mM ascorbic acid, 4 mM iron (II) sulfate ammonium hexahydrate, 50 mM sodium phosphate pH 6.8, air, and room temperature. The reactions were run overnight and analyzed using similar methods as described above for the HTP reactions. Selectivity relative to SEQ ID NO:598 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO:598. Results are shown in Table 11.2.

[0391]

[0392]

[0393] Example 12

[0394] Improvements in the conversion of proline substrate to trans-3-hydroxyproline compared to SEQ ID NO:630

[0395] A library of engineered genes was generated using established techniques (e.g., saturation mutagenesis and recombination of previously identified beneficial mutations) from an engineered polynucleotide (SEQ ID NO: 629) encoding a polypeptide having proline hydroxylase activity of SEQ ID NO: 630. Polypeptides encoded by each gene were produced HTP as described in Example 2 (where proteins were expressed overnight at 30°C). For all variants, cell pellets were lysed by adding 200 pL lysis buffer (containing 50 mM sodium phosphate buffer pH 6.5, 1 g / L lysozyme, and 0.5 g / L PMBS) and shaking on a benchtop shaker for 2 hours at room temperature. Plates were centrifuged at 4000 rpm for 15 minutes at 4°C to remove cell debris.

[0396] In 300 pL round bottom plates, 50 pL of E. coli lysate was added to 200 pL of reaction mixture (containing 75 pL 267 g / L a-ketoglutarate [in 50 mM sodium phosphate pH 6.5], 50 pL 20 mM Molar salt [in 50 mM sodium phosphate pH 6.5], and 200 pL 133 g / L proline in 65 mM ascorbic acid) in each well. Plates were sealed with AirPore seals (Qiagen) and reactions were continued in Kuhner at 30°C, 200 rpm, and 85% relative humidity, with a 2” swing at room temperature overnight (~18 hours).

[0397] Following overnight incubation, reactions were derivatized and quenched by aliquoting 25 pL of reaction mixture into 96-well deep well plates containing 225 pL derivatization solution (containing 75 pL saturated sodium bicarbonate, 25 pL water, and 125 pL 2.5 mg / mL FmocCl [in ACN] per well). After shaking at room temperature for 1 hr, plates were centrifuged for one minute at 4000 rpm and 40 pL of the soluble fraction of the quenched reaction was mixed with 160 pL 1 : 1 ACN:0.5 M HC1. Derivatized and diluted samples were analyzed as described in Table 13.1. Selectivity (selectivity FIOP) relative to SEQ ID NO: 630 was calculated as the ratio of trans-3-hydroxyproline:trans-4-hydroxyproline of the product formed by the variant compared to the ratio produced by SEQ ID NO: 630.

[0398] In addition to HTP analysis, a subset of selected beneficial variants from the HTP screen were also prepared at shake flask scale, as described in Example 3 and expressed at 30°C. The lyophilized shake flask lysate powder (SFP) was tested in 1 mL scale reactions under the following conditions: 60 g / L L-proline, 5 wt% proline hydroxylase variant SFP, 1.2 equivalents a-KG (a-ketoglutarate), 10 mM ascorbic acid, 4 mM iron (II) sulfate ammonium hexahydrate, 50 mM sodium phosphate pH 6.8, air, and room temperature. Reactions were run overnight and analyzed using similar methods described above for HTP reactions. Selectivity relative to SEQ ID NO: 630 (selectivity FIOP) was calculated as the ratio of trans-3-hydroxyproline: trans-4-hydroxyproline product formed by the variant compared to the ratio produced by SEQ ID NO: 630. Results are shown in Table 12.1.

[0399]

[0400]

[0401] Example 13

[0402] Analytical detection of trans-3-hydroxyproline produced from proline

[0403] The data described in Examples 4-12 were collected using the analytical methods in Table 13.1. The methods provided herein can all be used to analyze variants produced using the application. However, it is not intended that the methods described herein are the only methods that can be suitable for analyzing the variants provided herein and / or variants produced using the methods provided herein.

[0404] Proline substrates and hydroxyproline products were analyzed as described below. Reactions were derivatized and quenched by aliquoting 25 pL of reaction mixture into a 96-well deep well plate containing 225 pL of derivatization solution (75 pL of saturated sodium bicarbonate, 25 pL of water, and 125 pL of 2.5 mg / mL FmocCl [in ACN] per well). After shaking at room temperature for 1 hr, the plate was centrifuged at 4000 rpm for one minute and 40 pL of the soluble fraction of the quenched reaction was mixed with 160 pL of 1 : 1 ACN:0.5 M HC1. The derivatized and diluted samples were analyzed as described in Table 13.1.

[0405]

[0406]

[0407] All publications, patents, patent applications and other documents cited in this text are hereby incorporated by reference in their entirety for all purposes to the same extent as if each individual publication, patent, patent application or other document were individually indicated to be incorporated by reference for all purposes.

[0408] While various specific embodiments have been shown and described in detail, it will be understood that various changes can be made without departing from the true spirit and scope of the application.

Claims

1. An engineered polypeptide having proline hydroxylase activity, wherein the engineered polypeptide differs from SEQ ID NO: 4 in amino acid residues by N194L / T.

2. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 4 in that the amino acid residues are N194L / T and 1-6 substitutions selected from A48V, Y66W, A189N, N194L, K82P, E85P, A135P and G267D.

3. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 116 in the following amino acid residues: S123T, N189S, N189A, V233A / M, L296V, or H195Y.

4. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 116 by 1-9 substitutions selected from Y20F, R21Q, S56P, H76E, G95R / P, Q232E, H294Y, V307I / L, S335M, L296I, Y147F, L243V, R281T / S, Q348K, V233A / R, A185L, N189A, Q225R, G284R, L120P, and A29T.

5. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 116 by 1-7 substitutions selected from R21Q, H76E, Y147F, L243V, L296I, V307I, S335M, G95R, S56P and R281T.

6. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 162 in that the amino acid residues are 1-8 substitutions selected from Y45S, S123T, R326G, E85L, R281T / M, L282S, G2L, Q237E, M139F, V233A, M289D, A324Q, T199A, P200V, P247L, V250Q, M338I, S65R, A117V / T, L120I / P, V343N, A346G / S, E114G, V115T, R268T, S271A, L313F, P28A, R270L, and Q348S.

7. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 322 in the following amino acid residues: N189S, V228T, S262V, V277A, G26N, D61H, A201C / T / G, L175S / V, Q236T, E132P, A129I, V272S, V156S, T345R, P291G, G54P, D248R, S149G, or C209S.

8. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 322 in the following amino acid residues: S278N, A347E, A129I, C209S, G54P / S, Q163L, T345R, A43T, E58T, D143L, H25K, V156D / S, L175V, P291G, D61H, Q79T, D248R, E132N, Q236T, E179L, or A201C.

9. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 322 in that the amino acid residues are 1-9 substitutions selected from E85L, A117V / T, L120I / P, A135S, A208E, R270L, Q324A, V343N, A346G, F139M, M281R, D289M and S282L.

10. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 412 in that the amino acid residues are substituted by 1-3 substitutions selected from H162L / V / M / A, L204S, S164D / T, A198V, S271V, I169T / C / V, S113P / R / N / H, V243Y, H195Y, V48G, F47M, R275K, G95W / A, L330G / H, C187P, F314S / A / T, S56P, A118W / V / D / P, F154L, N342R, C168V, P247A, M289V, L85P, and R281L.

11. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 412 in that the amino acid residues are 1-7 substitutions selected from H25K, A129I, Q163L, Q236T, S262V, T345R, A347E, N189S, V277A, S278N, P120V, V156S, L175V, E179L, A201G, V228A, and H162V / L.

12. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 492 in the following amino acid residues: V228T, H294T, E208L / S / M, S17C, S135T / G / N, Q167G, D235E, A29S, S177A / P / L, I307L, I15V, S65V, P28I, D287E, T199C, or V343S / T.

13. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 492 in that the amino acid residues are 1-5 substitutions selected from L85P, C187P, R281L, A347E, H195Y, N278S, F314A, A118V, P120V, H162V, L175V, L330H, V228A and E179L.

14. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 562 in the following amino acid residues: K40A, A347K, L179T, I15F, A43S, S164Q, T345D, R59L, Q79E, S149N, G44V / R, or K82A.

15. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 562 by 1-9 substitutions selected from A29S, E208S / L, V228T, N278S, H294T / Y, A347E, L85P, S177A / P, H195Y, V343T, S56P, C187P, A205S, and S113N.

16. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 598 in the following amino acid residues: V162S, T219V, F47Q, S227R, C209H, or N342M / L.

17. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 598 by 1-11 substitutions selected from S17V, G44R / V, T199C, L313C, L179T, H195Y, V250P, T345D, A43S, S149N, S164Q, T171M, C187P / N, V288T, A286P, K82A, Q163D, A201V, S203Q, L208I, K320V, and Q79E.

18. An engineered polypeptide, wherein the engineered polypeptide differs from SEQ ID NO: 630 by 1-6 substitutions selected from V162S, T219V, L313C, M338I, T236L, N342M, S135P, Q163D, S164Q / T, A201V, S203Q, L208I, K82A, T171M and L282V.

19. An engineered polypeptide, wherein the polypeptide sequence of the engineered polypeptide is an even-numbered sequence in SEQ ID NO: 6-8, 116-120 or 142-658.

20. A polynucleotide encoding an engineered polypeptide according to any one of claims 1-19.

21. The polynucleotide of claim 20, wherein the polynucleotide comprises that found in *Escherichia coli* (…). E. coli Nucleic acid sequences optimized for expression in ).

22. An expression vector comprising a polynucleotide according to any one of claims 20-21.

23. The expression vector according to claim 22 further comprises at least one control sequence, said control sequence being selected from leader sequence, polyadenylation sequence, propeptide sequence, promoter sequence, signal peptide sequence, initiation sequence, and transcription terminator.

24. A host cell comprising a polynucleotide according to any one of claims 20-21.

25. A host cell comprising an expression vector according to any one of claims 22-23.

26. The host cell according to claim 24 or 25, wherein the host cell is Escherichia coli.

27. A method for preparing engineered polypeptides, the method comprising culturing host cells according to any one of claims 24-26 under conditions suitable for expressing the polypeptides.

28. The method of claim 27, further comprising the step of isolating the engineered polypeptide.

Citation Information

Patent Citations

  • Process for producing trans-4-hydroxy-l-proline

    EP0641862A2

  • Process for production of cis-4-hydroxy-l-proline

    EP2290065A1

  • Improvement in parlor-organs

    US110900A

  • Ketoreductase polypeptides and related polynucleotides

    US20060195947A1

  • Process for production of cis-4-hydroxy-l-proline

    US20110091942A1