P450-BM3 variants with improved activity

By engineering the P450-BM3 enzyme, a polypeptide variant with high sequence identity was prepared, which solved the problem of insufficient catalytic activity of the existing enzyme towards the 1-tert-butoxycarbonylaminocyclopentanoic acid substrate, achieved efficient catalytic conversion of the substrate, and produced the (tert-butoxycarbonylamino)-cyclopentanoic acid product for drug synthesis.

CN120813367APending Publication Date: 2025-10-17CODEXIS INC
View PDF 98 Cites 0 Cited by

Patent Information

Application Number
CN202380095297.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-20
Filing Date
2023-11-21
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing P450-BM3 enzyme has insufficient catalytic activity towards the 1-tert-butoxycarbonylaminocyclopentanoic acid substrate and needs to be improved to increase its catalytic efficiency towards the substrate.

Method used

By engineering the P450-BM3 enzyme, a polypeptide variant with high sequence identity to the wild-type enzyme was prepared, and specific amino acid substitutions or sets of substitutions were introduced to improve its catalytic activity towards 1-tert-butoxycarbonylaminocyclopentanoic acid.

Benefits of technology

The improved P450-BM3 variant significantly enhances the catalytic activity towards 1-tert-butoxycarbonylaminocyclopentanoic acid and can effectively convert it into (tert-butoxycarbonylamino)-cyclopentanoic acid product for the production of active pharmaceutical ingredients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005577222140000371
    Figure BDA0005577222140000371
  • Figure BDA0005577222140000381
    Figure BDA0005577222140000381
  • Figure BDA0005577222140000382
    Figure BDA0005577222140000382
Patent Text Reader

Abstract

The present invention provides an improved P450 to BM3 variant having an improved activity. In some embodiments, the P450-BM3 variants exhibit improved activity on a 1-t-butoxycarbonylaminocyclopentanoic acid substrate. In some embodiments, the P450-BM3 variants may exhibit improved activity on the substrate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 491,247, filed on March 20, 2023, which is incorporated herein by reference in its entirety. Technical Field

[0003] The present invention provides improved P450-BM3 variants with improved activity. In some embodiments, the P450-BM3 variant exhibits improved activity towards the 1-tert-butoxycarbonylaminocyclopentanoic acid substrate.

[0004] Reference to a sequence listing, table or computer program

[0005] An official copy of the sequence listing is submitted with the specification in the form of an XML file named "CX2-233USP2_ST26.xml", created on March 17, 2023, and 4,511 kilobytes in size. The submitted sequence listing is part of the specification and is incorporated herein by reference in its entirety. Background Art

[0006] Cytochrome P450 monooxygenases ("P450s") comprise a large, widespread class of heme enzymes found ubiquitously in nature. Cytochrome P450-BM3 ("P450-BM3"), obtained from Bacillus megaterium (now also known as Priestia megaterium), catalyzes the NADPH-dependent hydroxylation of long-chain fatty acids, alcohols, and amides, as well as the epoxidation of unsaturated fatty acids (see, e.g., Narhi and Fulco, J. Biol. Chem., 261:7160-7169

[1986] ; and Capdevila et al., J. Biol. Chem., 271:2263-22671

[1996] ). P450-BM3 is unique in that its reductase (65 kDa) and monooxygenase (55 kDa) domains are fused to form a catalytically self-sufficient 120 kDa enzyme. While these enzymes have been the subject of much research, there remains a need in the art for improved P450s that exhibit high levels of enzymatic activity on a variety of substrates, including 1-tert-butoxycarbonylaminocyclopentanoic acid. Summary of the Invention

[0007] The present invention provides improved P450-BM3 variants with improved activity. In some embodiments, the P450-BM3 variants exhibit improved activity on 1-tert-butoxycarbonylaminocyclopentanoic acid substrates. The present disclosure provides a recombinant cytochrome P450-BM3 variant having at least 80% sequence identity to a polypeptide sequence comprising the sequence set forth in SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266. In some further embodiments, the recombinant cytochrome P450-BM3 variant oxidizes 1-tert-butoxycarbonylaminocyclopentanoic acid.

[0008] The present invention provides novel biocatalysts and related methods for the synthesis of (tert-butoxycarbonylamino)-cyclopentanoic acid compounds and related ketone compounds from 1-tert-butoxycarbonylaminocyclopentanoic acid. The P450-BM3 variants of the present disclosure are engineered variants of a polypeptide (SEQ ID NO: 4) which is an engineered variant of a wild-type enzyme from Bacillus megaterium (SEQ ID NO: 2). These engineered polypeptides are capable of catalyzing the conversion of 1-tert-butoxycarbonylaminocyclopentanoic acid to a (tert-butoxycarbonylamino)-cyclopentanoic acid product which can be used to produce an active pharmaceutical ingredient.

[0009] The present invention provides engineered cytochrome P450-BM3 variants comprising a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, or a functional fragment thereof, wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or substitution set, and wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266.

[0010] The present application provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 4, comprising at least one substitution or one substitution set at one or more positions selected from 32 / 83 / 88, 32 / 83 / 88 / 176, 32 / 83 / 88 / 231 / 574, 32 / 83 / 88 / 574, 52 / 83 / 88, 52 / 83 / 88 / 105, 52 / 83 / 88 / 231, 52 / 83 / 88 / 231 / 433 / 574, 52 / 83 / 88 / 433, 52 / 83 / 88 / 433 / 574, 52 / 83 / 88 / 574, 83 / 88, 83 / 88 / 105, 83 / 88 / 111, 83 / 88 / 111 / 433, 83 / 88 / 111 / 574, 83 / 88 / 231, 83 / 88 / 349, 83 / 88 / 433 / 574, and 83 / 88 / 574, wherein the positions are numbered with reference to SEQ ID NO: 4. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 32K / 83A / 88A, 32K / 83A / 88A / 176I, 32K / 83A / 88A / 231R / 574T, 32K / 83A / 88A / 574T, 52Y / 83A / 88A, 52Y / 83A / 88A / 105V, 52Y / 83A / 88A / 231R, 52Y / 83A / 88A / 231R / 433D / 574T, 52Y / 83A / 88A / 433D, 52Y / 83A / 88A / 433D / 574T, 52Y / 83A / 88A / 574T, 83A / 88A, 83A / 88A / 105L, 83A / 88A / 111Q, 83A / 88A / 111Q / 433D, 83A / 88A / 111Q / 574T, 83A / 88A / 231R, 83A / 88A / 349E, 83A / 88A / 433D / 574T, and 83A / 88A / 574T, wherein the positions are numbered with reference to SEQ ID NO: 4.In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from R32K / L83A / F88A, R32K / L83A / F88A / V176I, R32K / L83A / F88A / S231R / N574T, R32K / L83A / F88A / N574T, F52Y / L83A / F88A, F52Y / L83A / F88A / G105V, F52Y / L83A / F88A / S231R, F52Y / L83A / F88A / S231R / V433D / N574T, F52Y / L83A / F88A / V433D, F52Y / L83A / F88A / V433D / N574T, F52Y / L83A / F88A / N574T, L83A / F88A, L83A / F88A / G105L, L83A / F88A / R111Q, L83A / F88A / R111Q / V433D, L83A / F88A / R111Q / N574T, L83A / F88A / S231R, L83A / F88A / T349E, L83A / F88A / V433D / N574T, and L83A / F88A / N574T, wherein the positions are numbered with reference to SEQ ID NO: 4.

[0011] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 36, comprising at least one substitution or one substitution set selected from: 75, 75 / 374, 75 / 374 / 458 / 726, 75 / 374 / 726, 75 / 458, 75 / 458 / 726, 75 / 726, 111 / 114, 111 / 603 / 604 / 623 / 853, 111 / 623, 374 / 726, and 726, wherein positions are numbered with reference to SEQ ID NO: 36. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: 75S, 75S / 374S, 75S / 374S / 458L / 726L, 75S / 374S / 726L, 75S / 458L, 75S / 458L / 726L, 75S / 726L, 111H / 114G, 111H / 603F / 604G / 623Q / 853E, 111H / 623Q, 374S / 726L, and 726L, wherein positions are numbered with reference to SEQ ID NO: 36. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: A75S, A75S / E374S, A75S / E374S / G458L / Q726L, A75S / E374S / Q726L, A75S / G458L, A75S / G458L / Q726L, A75S / Q726L, R111H / K114G, R111H / E603F / A604G / S623Q / P853E, R111H / S623Q, E374S / Q726L, and Q726L, wherein positions are numbered with reference to SEQ ID NO: 36.

[0012] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 66, comprising at least one substitution or one substitution set selected from: 74, 75, 83, 179, 181, 182, 186, 189, 238, 267, 268, 328, 331, 355, 358, 437, and 438, wherein positions are numbered with reference to SEQ ID NO: 66. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: 74G, 75K, 75T, 83V, 179M, 179R, 181V, 182I, 182V, 186R, 189G, 189T, 238L, 267T, 268Q, 328S, 328V, 331F, 331M, 331T, 355V, 358T, 358V, 437G, 437S, and 438R, wherein positions are numbered with reference to SEQ ID NO: 66. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: Q74G, A75K, A75T, A83V, V179M, V179R, A181V, L182I, L182V, M186R, L189G, L189T, M238L, H267T, E268Q, T328S, T328V, A331F, A331M, A331T, M355V, I358T, I358V, T437G, T437S, and L438R, wherein positions are numbered with reference to SEQ ID NO: 66.

[0013] The present application also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 72, comprising at least one substitution or one substitution set selected from: 74 / 75 / 83 / 179 / 189 / 328 / 331 / 437, 74 / 75 / 268 / 328 / 331 / 358 / 437, 74 / 83 / 268 / 328 / 331 / 358 / 437, 74 / 267 / 268 / 328, 75 / 83 / 179 / 331, 75 / 83 / 179 / 268 / 328 / 331 / 437, 75 / 83 / 189 / 267 / 268, 75 / 83 / 267 / 268 / 355 / 358 / 654, 75 / 83 / 268 / 437, 75 / 268 / 328 / 331 / 358, 83 / 179 / 182 / 437, 83 / 179 / 189 / 331 / 355, 83 / 179 / 328 / 331, 83 / 179 / 355 / 358 / 437, 83 / 182 / 189 / 268 / 328 / 331 / 355 / 358 / 437, 83 / 182 / 189 / 328 / 331, 83 / 182 / 268 / 328 / 355 / 358 / 437, 83 / 189 / 267 / 268 / 358, 83 / 189 / 268 / 328 / 331, 83 / 189 / 268 / 328 / 355 / 358 / 437, 83 / 189 / 328 / 331 / 437, 83 / 267 / 268 / 328 / 331 / 355 / 358, 83 / 268, 83 / 268 / 328 / 331, 83 / 268 / 328 / 331 / 355 / 358, 83 / 268 / 328 / 331 / 358, 83 / 268 / 328 / 331 / 358 / 437, 83 / 268 / 328 / 331 / 437, 83 / 268 / 331, 83 / 331, 83 / 331 / 437, 83 / 358 / 437, 179 / 182 / 268, 179 / 189 / 331 / 437, 179 / 328 / 331, 179 / 331 / 358, 189 / 267 / 268 / 437, 189 / 268 / 328, 189 / 268 / 328 / 331 / 358 / 437, 189 / 268 / 358, 267 / 268 / 328, 267 / 268 / 328 / 331 / 355 / 358, 267 / 268 / 331 / 355 / 358 / 437, 268, 268 / 328 / 331, 268 / 328 / 355 / 358, 268 / 331 / 355 / 358,268 / 355 / 358 / 437, 268 / 358, and 328 / 331 / 358, wherein the position is numbered with reference to SEQ ID NO: 72. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 74G / 75T / 83V / 179R / 189G / 328V / 331T / 437G, 74G / 75T / 268Q / 328V / 331T / 358V / 437S, 74G / 83V / 268Q / 328S / 331M / 358V / 437G, 74G / 267T / 268Q / 328V, 75T / 83V / 179R / 189G / 331M, 75T / 83V / 179R / 268Q / 328V / 331M / 437G, 75T / 83V / 189G / 267T / 268Q, 75T / 83V / 267T / 268Q / 355V / 358V / 654L, 75T / 83V / 268Q / 437G, 75T / 268Q / 328V / 331T / 358T, 83V / 179R / 182V / 437S, 83V / 179R / 189G / 331M / 355V, 83V / 179R / 328V / 331M, 83V / 179R / 355V / 358V / 437G, 83V / 182I / 189G / 268Q / 328V / 331M / 355V / 358V / 437G, 83V / 182I / 189G / 328V / 331M, 83V / 182I / 268Q / 328V / 355V / 358V / 437S, 83V / 189G / 267T / 268Q / 358T, 83V / 189G / 268Q / 328V / 331M, 83V / 189G / 268Q / 328V / 355V / 358V / 437G, 83V / 189G / 328V / 331M / 437G, 83V / 267T / 268Q / 328V / 331M / 355V / 358T, 83V / 268Q, 83V / 268Q / 328S / 331M / 355V / 358V, 83V / 268Q / 328S / 331M / 358T / 437G, 83V / 268Q / 328S / 331T / 437G, 83V / 268Q / 328V / 331M, 83V / 268Q / 328V / 331M / 358V, 83V / 268Q / 331T, 83V / 331M, 83V / 331T / 437S, 83V / 358V / 437S, 179R / 182I / 268Q, 179R / 189G / 331M / 437S, 179R / 328V / 331M, 179R / 331M / 358V, 189G / 267T / 268Q / 437S, 189G / 268Q / 328V,189G / 268Q / 328V / 331T / 358V / 437S, 189G / 268Q / 358V, 267T / 268Q / 328V, 267T / 268Q / 328V / 331M / 355V / 358T, 267T / 268Q / 331T / 355V / 358T / 437S, 268Q, 268Q / 328V / 331M, 268Q / 328V / 355V / 358T, 268Q / 331M / 355V / 358V, 268Q / 355V / 358V / 437G, 268Q / 358V, and 328V / 331M / 358T, wherein the positions are numbered with reference to SEQ ID NO: 72. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from Q74G / A75T / A83V / V179R / L189G / T328V / A331T / T437G, Q74G / A75T / E268Q / T328V / A331T / I358V / T437S, Q74G / A83V / E268Q / T328S / A331M / I358V / T437G, Q74G / H267T / E268Q / T328V, A75T / A83V / V179R / L189G / A331M, A75T / A83V / V179R / E268Q / T328V / A331M / T437G, A75T / A83V / L189G / H267T / E268Q, A75T / A83V / H267T / E268Q / M355V / I358V / M654L, A75T / A83V / E268Q / T437G, A75T / E268Q / T328V / A331T / I358T, A83V / V179R / L182V / T437S, A83V / V179R / L189G / A331M / M355V, A83V / V179R / T328V / A331M, A83V / V179R / M355V / I358V / T437G, A83V / L182I / L189G / E268Q / T328V / A331M / M355V / I358V / T437G, A83V / L182I / L189G / T328V / A331M, A83V / L182I / E268Q / T328V / M355V / I358V / T437S, A83V / L189G / H267T / E268Q / I358T, A83V / L189G / E268Q / T328V / A331M, A83V / L189G / E268Q / T328V / M355V / I358V / T437G, A83V / L189G / T328V / A331M / T437G,A83V / H267T / E268Q / T328V / A331M / M355V / I358T, A83V / E268Q, A83V / E268Q / T328S / A331M / M355V / I358V, A83V / E268Q / T328S / A331M / I358T / T437G, A83V / E268Q / T328S / A331T / T437G, A83V / E268Q / T328V / A331M, A83V / E268Q / T328V / A331M / I358V, A83V / E268Q / A331T, A83V / A331M, A83V / A331T / T437S, A83V / I358V / T437S, V179R / L182I / E268Q, V179R / L189G / A331M / T437S, V179R / T328V / A331M, V179R / A331M / I358V, L189G / H267T / E268Q / T437S, L189G / E268Q / T328V, L189G / E268Q / T328V / A331T / I358V / T437S, L189G / E268Q / I358V, H267T / E268Q / T328V, H267T / E268Q / T328V / A331M / M355V / I358T, H267T / E268Q / A331T / M355V / I358T / T437S, E268Q, E268Q / T328V / A331M, E268Q / T328V / M355V / I358T, E268Q / A331M / M355V / I358V, E268Q / M355V / I358V / T437G, E268Q / I358V, and T328V / A331M / I358T, wherein the position is numbered with reference to SEQ ID NO: 72.

[0014] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to a reference sequence of SEQ ID NO: 198, comprising at least one substitution or a set of substitutions selected from 79, 213, and 257, wherein positions are numbered with reference to SEQ ID NO: 198. In some embodiments, the engineered polypeptide comprises at least one substitution or a set of substitutions selected from 79I, 213V, and 257Q, wherein positions are numbered with reference to SEQ ID NO: 198. In some further embodiments, the engineered polypeptide comprises at least one substitution or a set of substitutions selected from A79I, M213V, and Y257Q, wherein positions are numbered with reference to SEQ ID NO: 198.

[0015] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to a reference sequence of SEQ ID NO: 226, comprising at least one substitution or a set of substitutions selected from 315, 320, 385, 388, 391, 398, 405, 493, 497, 502, 503, 504, 541, 542, 547, 573, 576, and 577, wherein positions are numbered with reference to SEQ ID NO: 226.

[0016] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 244, comprising at least one substitution or one substitution set selected from 75 / 178 / 213 / 315, 75 / 331 / 576 / 726, 178, 178 / 179 / 213 / 437 / 497 / 573 / 576, 178 / 179 / 573 / 726, 178 / 179 / 576, 178 / 213 / 573, 178 / 213 / 726, 178 / 437, 178 / 497 / 726, 178 / 576, 178 / 726, 179 / 358, 179 / 726, 331, 331 / 358 / 391 / 437, 331 / 497, 331 / 573 / 576, 497 / 573, 573, 682, 685, 699, 701, 704, 707, 708, 726, 756, 759, 794, 796, 797, 848, 851, 862, 888, 889, 999, 1003, and 1048, wherein the positions are numbered with reference to SEQ ID NO: 244. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 75S / 178V / 213V / 315T, 75S / 331T / 576V / 726L, 178V, 178V / 179R / 213V / 437S / 497A / 573S / 576V, 178V / 179R / 573S / 726L, 178V / 179R / 576V, 178V / 213V / 573S, 178V / 213V / 726L, 178V / 437S, 178V / 497A / 726L, 178V / 576V, 178V / 726L, 179R / 358T, 179R / 726L, 331T, 331T / 358T / 391Y / 437S, 331T / 497A, 331T / 573S / 576V, 497A / 573S, 573S, 682K, 685F, 699R, 701V, 704L, 707Y, 708C, 726L, 756Q, 759T, 794C, 796H, 796N, 797H, 848T, 851G, 851V, 862L, 888Y, 889Q, 999A, 1003W, and 1048L, wherein the positions are numbered with reference to SEQ ID NO: 244.In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from the group consisting of A75S / M178V / M213V / V315T, A75S / M331T / A576V / Q726L, M178V, M178V / V179R / M213V / T437S / T497A / K573S / A576V, M178V / V179R / K573S / Q726L, M178V / V179R / A576V, M178V / M213V / K573S, M178V / M213V / Q726L, M178V / T437S, M178V / T497A / Q726L, M178V / A576V, M178V / Q726L, V179R / V358T, V179R / Q726L, M331T, M331T / V358T / F391Y / T437S, M331T / T497A, M331T / K573S / A576V, T497A / K573S, K573S, T682K, L685F, D699R, L701V, I704L, N707Y, Y708C, Q726L, L756Q, P759T, V794C, A796H, A796N, K797H, S848T, S851G, S851V, K862L, E888Y, F889Q, I999A, G1003W, and A1048L, wherein the positions are numbered with reference to SEQ ID NO: 244.

[0017] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 286, comprising at least one substitution or one substitution set selected from 75 / 331 / 701 / 726 / 851 / 1048, 75 / 331 / 726 / 999, 75 / 726 / 796 / 851 / 999, 75 / 1048, 87, 88 / 522, 89, 234, 269, 328, 330, 331, 398, 405, 408, and 411, wherein positions are numbered with reference to SEQ ID NO: 286. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 75S / 331T / 701V / 726L / 851V / 1048L, 75S / 331T / 726L / 999A, 75S / 726L / 796N / 851V / 999A, 75S / 1048L, 87I, 87V, 88S / 522R, 89S, 234F, 269G, 269P, 328T, 330G, 331S, 398R, 405A, 405S, 408R, and 411G, wherein positions are numbered with reference to SEQ ID NO: 286. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from A75S / M331T / L701V / Q726L / S851V / A1048L, A75S / M331T / Q726L / I999A, A75S / Q726L / A796N / S851V / I999A, A75S / A1048L, L87I, L87V, A88S / G522R, T89S, L234F, T269G, T269P, V328T, P330G, M331S, Q398R, Q405A, Q405S, L408R, and A411G, wherein positions are numbered with reference to SEQ ID NO: 286.

[0018] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 358 comprising at least one substitution or one substitution set selected from 75 / 269 / 707, 269, 269 / 522 / 707, 269 / 522 / 707 / 1048, 522, 522 / 726, 522 / 1048, and 726 / 1048, wherein positions are numbered with reference to SEQ ID NO: 358. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 75S / 269P / 707Y, 269P, 269P / 522G / 707Y, 269P / 522G / 707Y / 1048L, 522G, 522G / 726L, 522G / 1048L, and 726L / 1048L, wherein positions are numbered with reference to SEQ ID NO: 358. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from A75S / T269P / N707Y, T269P, T269P / R522G / N707Y, T269P / R522G / N707Y / A1048L, R522G, R522G / Q726L, R522G / A1048L, and Q726L / A1048L, wherein positions are numbered with reference to SEQ ID NO: 358.

[0019] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 358 comprising at least one substitution or one substitution set selected from 77, 170, 286, 289, 462, 547, 557, 630, 646, 651, 672, 676, 692, 775, 786, 787, 788, 814, 841, 876, 877, 888, 893, 896, 924, 941, 955, 969, 973, 982, 989, 993, and 1038, wherein the positions are numbered with reference to SEQ ID NO: 358. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 77L, 170S, 170T, 286R, 289A, 462W, 547R, 557E, 557P, 630G, 646F, 651S, 672G, 676K, 676R, 692G, 692V, 775K, 786G, 786R, 786S, 786V, 787R, 788P, 814I, 814S, 841R, 876G, 877V, 888G, 893G, 896L, 924A, 924G, 924N, 924P, 941R, 955G, 969E, 969G, 973P, 982G, 989G, 989L, 993S, and 1038Q, wherein the positions are numbered with reference to SEQ ID NO: 358. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from K77L, Q170S, Q170T, H286R, Q289A, P462W, Q547R, A557E, A557P, N630G, Q646F, A651S, E672G, P676K, P676R, E692G, E692V, P775K, E786G, E786R, E786S, E786V, K787R, Q788P, K814I, K814S, K841R, D876G, T877V, E888G, K893G, E896L, Q924A, Q924G, Q924N, Q924P, H941R, S955G, P969E, P969G, K973P, Q982G, E989G, E989L, Q993S, and E1038Q, wherein the positions are numbered with reference to SEQ ID NO: 358.

[0020] The present invention also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 410, comprising at least one substitution or a set of substitutions selected from the group consisting of: 158 / 170, 158 / 170 / 410 / 462 / 630 / 672 / 726 / 786 / 788 / 814 / 924, 158 / 170 / 410 / 462 / 630 / 672 / 726 / 786 / 788 / 814 / 924 0 / 410 / 462 / 726 / 786 / 814, 158 / 170 / 410 / 924, 158 / 170 / 630 / 726 / 786 / 788 / 814, 158 / 410, 158 / 410 / 462 / 557 / 969, 158 / 410 / 557 / 786 / 787, 158 / 410 / 814 / 924, 158 / 410 / 862, 158 / 462 / 557 / 630 / 924, 158 / 462 / 630 / 786 / 788 / 969、158 / 557、158 / 557 / 630 / 814、158 / 557 / 726 / 786 / 788 / 862、158 / 557 / 786、158 / 557 / 786 / 787 / 788、158 / 557 / 814 / 924、158 / 630 / 786 / 814、158 / 726、158 / 786 / 787 / 788、158 / 814 / 924、170 / 79, 795, 838, 871, 923, and 926, wherein positions are numbered with reference to SEQ ID NO: 410.In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 158L / 170T, 158L / 170T / 410S / 462W / 630G / 672G / 726L / 786R / 788P / 814S / 924G, 158L / 170T / 410S / 462W / 726L / 786G / 814S, 158L / 170T / 410S / 924G, 158L / 170T / 630G / 726L / 786G / 788P / 814S, 158L / 410S, 158L / 410S / 462W / 557E / 969G, 158L / 410S / 557P / 786G / 787R, 158L / 410S / 814S / 924G, 158L / 410S / 862L, 158L / 462W / 557P / 630G / 924G, 158L / 462W / 630G / 786R / 788P / 969G, 158L / 557P, 158L / 557P / 630G / 814S, 158L / 557P / 726L / 786G / 788P / 862L, 158L / 557P / 786R / 787R / 788P, 158L / 557P / 786V, 158L / 557P / 814S / 924G, 158L / 630G / 786S / 814S, 158L / 726L, 158L / 786S / 787R / 788P, 158L / 814S / 924G, 170T / 410S / 557E / 786V / 787R / 814S / 924G / 989G, 385R, 385T, 410S / 462W / 557P / 630G / 786V / 787R / 969G, 462W / 557P / 726L / 786G / 787R / 924G, 469R, 523R, 550T, 553A, 553G, 553K, 553S, 556G, 574H, 574S, 613A, 640C, 640G, 640V, 640W, 645V, 650P, 652S, 717C, 717Q, 773A, 773G, 779G, 795A, 838W, 871M, 871N, 923A, 923L, 926K, and 926R, wherein the positions are numbered with reference to SEQ ID NO: 410.In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from the group consisting of G158L / Q170T, G158L / Q170T / E410S / P462W / N630G / E672G / Q726L / E786R / Q788P / K814S / Q924G, G158L / Q170T / E410S / P462W / Q726L / E786G / K814S, G158L / Q170T / E410S / Q924G, G158L / Q170T / N630G / Q726L / E786G / Q788P / K814S, G158L / E410S, G158L / E410S / P462W / A557E / P969G, G158L / E410S / A557P / E786G / K787R, G158L / E410S / K814S / Q924G, G158L / E410S / K862L, G158L / P462W / A557P / N630G / Q924G, G158L / P462W / N630G / E786R / Q788P / P969G, G158L / A557P, G158L / A557P / N630G / K814S, G158L / A557P / Q726L / E786G / Q788P / K862L, G158L / A557P / E786R / K787R / Q788P, G158L / A557P / E786V, G158L / A557P / K814S / Q924G, G158L / N630G / E786S / K814S, G158L / Q726L, G158L / E786S / K787R / Q788P, G158L / K814S / Q924G, Q170T / E410S / A557E / E786V / K787R / K814S / Q924G / E989G, A385R, A385T, E410S / P462W / A557P / N630G / E786V / K787R / P969G, P462W / A557P / Q726L / E786G / K787R / Q924G, K469R, N523R, D550T, D553A, D553G, D553K, D553S, S556G, N574H, N574S, T613A, K640C, K640G, K640V, K640W, L645V, S650P, A652S, A717C, A717Q, V773A, V773G, V779G, L795A, V838W, E871M, E871N, E923A, E923L, Q926K, and Q926R, wherein the position is numbered with reference to SEQ ID NO: 410.

[0021] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 410 comprising at least one substitution or one substitution set selected from 51 / 851, 460, 466, 474, 597, 600, 635, 638, 655, 663, 664, 677, 694, 696, 713, 771, 783, 789, 806, 807, 840, 842, 851, 857, 860, 878, 894, 942, 947, 960, 978, 992, 1008, 1012, 1024, and 1025, wherein positions are numbered with reference to SEQ ID NO: 410. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 51H / 851Q, 460R, 466C, 474F, 474P, 597R, 597V, 600K, 635S, 638E, 638W, 655A, 655L, 655T, 663G, 664L, 677H, 677N, 694P, 696C, 713M, 771I, 783E, 783K, 783S, 789G, 806N, 807R, 807T, 840L, 842G, 851L, 857F, 860R, 878S, 894V, 942K, 947W, 960A, 960G, 960K, 978S, 992F, 992G, 1008C, 1012G, 1012M, 1024I, 1024L, 1025G, 1025T, and 1025W, wherein positions are numbered with reference to SEQ ID NO: 410.In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from the group consisting of R51H / S851Q, P460R, Q466C, K474F, K474P, N597R, N597V, D600K, N635S, D638E, D638W, P655A, P655L, P655T, F663G, S664L, G677H, G677N, S694P, Q696C, N713M, K771I, A783E, A783K, A783S, A789G, E806N, K807R, K807T, E840L, Q842G, S851L, G857F, E860R, I878S, D894V, E942K, Q947W, T960A, T960G, T960K, H978S, D992F, D992G, P1008C, A1012G, A1012M, V1024I, V1024L, S1025G, S1025T, and S1025W, wherein the position is numbered with reference to SEQ ID NO: 410.

[0022] The present invention also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 544, comprising at least one substitution or a set of substitutions selected from the group consisting of: 77 / 179 / 286 / 410 / 788 / 888, 77 / 410 / 676 / 788 / 92 4, 77 / 557 / 707 / 888, 286 / 410 / 651 / 676, 286 / 410 / 707 / 788, 286 / 410 / 888, 286 / 692 / 786 / 788, 410, 410 / 557 / 676 / 788 / 888 / 924 / 993, 410 / 557 / 692 / 788 and 410 / 646 / 651 / 788, wherein positions are numbered with reference to SEQ ID NO:544. In some embodiments, the engineered polypeptide comprises at least one substitution or a set of substitutions selected from the group consisting of: 77L / 179L / 286R / 410S / 788Q / 888G, 77L / 410S / 676K / 788Q / 924N, 77L / 557P / 707Y / 888G, 286R / 410S / 651S / 676K, 286R / 410S / 707Y / 788Q, 286R / 410S / 888G, 286R / 692V / 786R / 788Q, 410S, 410S / 557P / 676K / 788Q / 888G / 924N / 993S, 410S / 557P / 692V / 788Q and 410S / 646F / 651S / 788Q, where positions are numbered with reference to SEQ ID NO: 544. In some further embodiments, the engineered polypeptide comprises at least one substitution or a set of substitutions selected from the group consisting of K77L / V179L / H286R / E410S / P788Q / E888G, K77L / E410S / P676K / P788Q / Q924N, K77L / A557P / N707Y / E888G, H286R / E410S / A651S / P676K, H286R / E410S / N707Y / P788Q, H286R / E410S / E888G, H286R / E692V / G786R / P788Q, E410S, E410S / A557P / P676K / P788Q / E888G / Q924N / Q993S, E410S / A557P / E692V / P788Q and E410S / Q646F / A651S / P788Q, where positions are numbered with reference to SEQ ID NO: 544.

[0023] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 734, comprising at least one substitution or one substitution set selected from: 24, 53, 75, 78, 82, 88, 150, 180, 183, 257, 270, 410 / 497 / 557 / 576 / 814, 410 / 497 / 573 / 576, 410 / 497 / 814, 410 / 557 / 924, 437, and 497 / 557, wherein positions are numbered with reference to SEQ ID NO: 734. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: 24R, 24S, 53V, 75G, 78E, 82I, 88G, 88T, 88V, 150G, 180G, 183S, 257G, 257H, 257N, 257W, 270I, 270V, 410E / 497A / 557P / 576V / 814K, 410E / 497A / 573S / 576V, 410E / 497A / 814K, 410E / 557P / 924G, 437N, 437V, and 497A / 557P, wherein positions are numbered with reference to SEQ ID NO: 734. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: D24R, D24S, L53V, A75G, F78E, F82I, S88G, S88T, S88V, T150G, R180G, D183S, Q257G, Q257H, Q257N, Q257W, T270I, T270V, S410E / T497A / A557P / A576V / S814K, S410E / T497A / K573S / A576V, S410E / T497A / S814K, S410E / A557P / Q924G, S437N, S437V, and T497A / A557P, wherein positions are numbered with reference to SEQ ID NO: 734.

[0024] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 748 comprising at least one substitution or one substitution set selected from 75, 75 / 82 / 180 / 257 / 437, 75 / 82 / 257 / 268, 75 / 82 / 257 / 556 / 640, 75 / 82 / 257 / 773, 75 / 82 / 437, 75 / 180 / 183 / 257 / 268 / 270 / 385 / 437 / 556 / 613 / 652 / 923, 75 / 180 / 257 / 268 / 270 / 437 / 556 / 574 / 652, 75 / 180 / 574 / 795, 75 / 257 / 613, 75 / 257 / 640, 75 / 556 / 773, 75 / 574, 75 / 613, 82 / 613, 180 / 437 / 773 / 795, and 257 / 613 / 773 / 795, wherein the positions are numbered with reference to SEQ ID NO: 748. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 75G, 75G / 82I / 180G / 257G / 437N, 75G / 82I / 257G / 268T, 75G / 82I / 257H / 773A, 75G / 82I / 257N / 556G / 640G, 75G / 82I / 437N, 75G / 180G / 183S / 257N / 268T / 270I / 385R / 437N / 556G / 613A / 652S / 923A, 75G / 180G / 257N / 268T / 270I / 437V / 556G / 574S / 652S, 75G / 180G / 574S / 795A, 75G / 257N / 613A, 75G / 257N / 640W, 75G / 556G / 773A, 75G / 574S, 75G / 613A, 82I / 613A, 180G / 437N / 773A / 795A, and 257H / 613A / 773A / 795A, wherein the positions are numbered with reference to SEQ ID NO: 748.In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from the group consisting of A75G, A75G / F82I / R180G / Q257G / S437N, A75G / F82I / Q257G / Q268T, A75G / F82I / Q257H / V773A, A75G / F82I / Q257N / S556G / K640G, A75G / F82I / S437N, A75G / R180G / D183S / Q257N / Q268T / T270I / A385R / S437N / S556G / T613A / A652S / E923A, A75G / R180G / Q257N / Q268T / T270I / S437V / S556G / N574S / A652S, A75G / R180G / N574S / L795A, A75G / Q257N / T613A, A75G / Q257N / K640W, A75G / S556G / V773A, A75G / N574S, A75G / T613A, F82I / T613A, R180G / S437N / V773A / L795A, and Q257H / T613A / V773A / L795A, wherein the positions are numbered with reference to SEQ ID NO: 748.

[0025] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 748 comprising at least one substitution or one substitution set selected from: 45, 102, 106, 110, 111, 114, 127, 191, 193, 194, 196, 196 / 853, 198, 202, 203, 206, 210, 226, 232, 236, 237, 244, 245, 248, 254, 256, and 347, wherein positions are numbered with reference to SEQ ID NO: 748. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: 45E, 102A, 102L, 106V, 110F, 110L, 110V, 111L, 111R, 111W, 114H, 114M, 114Q, 114R, 127M, 191S, 193K, 194T, 196Q / 853L, 196W, 198S, 202Q, 203G, 206R, 210R, 210T, 226N, 232R, 236G, 236S, 237N, 237R, 244G, 244R, 244Y, 245A, 245G, 245L, 245M, 245N, 245Q, 245R, 245S, 245T, 245V, 245W, 248M, 248R, 254L, 256K, 347E, 347K, and 347S, wherein positions are numbered with reference to SEQ ID NO: 748. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: A45E, N102A, N102L, P106V, Q110F, Q110L, Q110V, H111L, H111R, H111W, G114H, G114M, G114Q, G114R, L127M, R191S, N193K, P194T, D196Q / P853L, D196W, A198S, N202Q, K203G, C206R, I210R, I210T, A226N, H232R, T236G, T236S, Q237N, Q237R, P244G, P244R, P244Y, E245A, E245G, E245L, E245M, E245N, E245Q, E245R, E245S, E245T, E245V, E245W, E248M, E248R, N254L, R256K, P347E, P347K, and P347S, wherein positions are numbered with reference to SEQ ID NO: 748.

[0026] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 828 comprising at least one substitution or one substitution set selected from 22 / 75 / 717 / 720 / 779 / 1004, 22 / 550 / 826, 22 / 616 / 717 / 1004, 22 / 717 / 795 / 799 / 826, 550 / 616 / 717 / 779, 550 / 640 / 717, 550 / 717 / 795 / 799 / 800, 616 / 717 / 720 / 799, 640, 717, 717 / 720 / 779 / 1004, 717 / 779 / 799 / 800 / 1004, 717 / 1004, 720 / 779, 779 / 1004, 800 / 1004, and 826 / 1004, wherein the positions are numbered with reference to SEQ ID NO: 828. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 22D / 75A / 717C / 720L / 779F / 1004I, 22D / 550T / 826G, 22D / 616L / 717I / 1004I, 22D / 717Q / 795V / 799I / 826G, 550T / 616L / 717I / 779F, 550T / 640D / 717I, 550T / 717C / 795V / 799I / 800Y, 616L / 717G / 720L / 799I, 640V, 717C, 717G / 720L / 779F / 1004I, 717G / 1004I, 717Q / 779F / 799I / 800Y / 1004I, 717Q / 1004I, 720L / 779F, 779F / 1004I, 800Y / 1004I, and 826G / 1004I, wherein the positions are numbered with reference to SEQ ID NO: 828.In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: N22D / G75A / A717C / G720L / V779F / S1004I, N22D / D550T / R826G, N22D / E616L / A717I / S1004I, N22D / A717Q / L795V / L799I / R826G, D550T / E616L / A717I / V779F, D550T / K640D / A717I, D550T / A717C / L795V / L799I / T800Y, E616L / A717G / G720L / L799I, K640V, A717C, A717G / G720L / V779F / S1004I, A717G / S1004I, A717Q / V779F / L799I / T800Y / S1004I, A717Q / S1004I, G720L / V779F, V779F / S1004I, T800Y / S1004I, and R826G / S1004I, wherein the positions are numbered with reference to SEQ ID NO: 828.

[0027] The present invention also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 968, comprising at least one substitution or a set of substitutions selected from the group consisting of: 3, 118 / 446, 132, 230, 285, 290, 292, 293, 295, 296, 300, 303, 305, 307, 366, 371, 372, 381, 382, ​​415, 417, 418, 424, 427, 432, 433, 446, 447, 455, 463, 465, 477, 481, 483, 484, 485, 486, 490, 501, 503, 504, 505, 506, 512, 513, 514, 515, 516, 517, 518, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535 68, 473, 473 / 790, 478, 480, 481, 506, 600 / 635 / 713 / 771 / 1025, 600 / 840 / 960, 635 / 636 / 793 / 840, 635 / 638 / 793 / 823 / 960, 635 / 713 / 1025, 635 / 771 / 894, 636 / 638 / 793 / 851, 638 / 663 / 793 / 840, 713, 771, 786 / 840 / 960, 793 / 840, 807, 840 / 960, 851 / 1024 / 1025, 851 / 1025, 960 and 1025, wherein the positions are numbered with reference to SE Q ID NO: 968.In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 3D, 3W, 118T / 446L, 132M, 132V, 230D, 230V, 285G, 285S, 290S, 292S, 293R, 293T, 295S, 296G, 296Q, 300P, 303Q, 303W, 305G, 305N, 305R, 307R, 366I, 366S, 371K, 372S, 381F, 382M, 382T, 415V, 417V, 418L, 418V, 424V, 424Y, 427T, 432V, 433D, 433K, 433L, 433R, 446S, 447L, 455R, 463A, 463V, 465Q, 468F, 473L / 790H, 473R, 478S, 480W, 481L, 506D, 506V, 600K / 635S / 713M / 771I / 1025G, 600K / 840L / 960L, 635S / 636G / 793T / 840L, 635S / 638G / 793T / 823A / 960L, 635S / 713M / 1025T, 635S / 771I / 894V, 636M / 638G / 793T / 851L, 638G / 663R / 793T / 840S, 713M, 771I, 786A / 840S / 960L, 793T / 840L, 807R, 840L / 960L, 851L / 1024I / 1025T, 851L / 1025T, 960L, 1025T, and 1025W, wherein the positions are numbered with reference to SEQ ID NO: 968.In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from I3D, I3W, A118T / V446L, E132M, E132V, Q230D, Q230V, P285G, P285S, K290S, A292S, E293R, E293T, A295S, A296G, A296Q, V300P, V303Q, V303W, S305G, S305N, S305R, K307R, T366I, T366S, D371K, V372S, E381F, N382M, N382T, L415V, M417V, M418L, M418V, F424V, F424Y, H427T, L432V, V433D, V433K, V433L, V433R, V446S, V447L, P455R, S463A, S463V, E465Q, A468F, K473L / Y790H, K473R, A478S, N480W, T481L, M506D, M506V, D600K / N635S / N713M / K771I / S1025G, D600K / E840L / T960L, N635S / S636G / Q793T / E840L, N635S / D638G / Q793T / P823A / T960L, N635S / N713M / S1025T, N635S / K771I / D894V, S636M / D638G / Q793T / S851L, D638G / F663R / Q793T / E840S, N713M, K771I, G786A / E840S / T960L, Q793T / E840L, K807R, E840L / T960L, S851L / V1024I / S1025T, S851L / S1025T, T960L, S1025T, and S1025W, wherein the positions are numbered with reference to SEQ ID NO: 968.

[0028] The present application also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 984, comprising at least one substitution or one substitution set selected from 45, 45 / 111 / 226 / 347, 45 / 853, 110 / 114, 111, 111 / 127 / 226 / 244, 111 / 194 / 244 / 347 / 853, 111 / 194 / 347, 111 / 194 / 347 / 853, 111 / 210, 111 / 210 / 347 / 853, 111 / 226, 111 / 226 / 244 / 347, 111 / 226 / 853 / 969, 111 / 244, 111 / 244 / 853, 111 / 347 / 853, 111 / 853, 114, 114 / 245, 127 / 210 / 244, 127 / 210 / 244 / 853, 127 / 210 / 347 / 853, 127 / 244, 127 / 347, 194 / 853, 210 / 853, 226 / 236 / 244 / 347 / 853, 236 / 244, 237, 237 / 245, 244 / 853, and 245, wherein the positions are numbered with reference to SEQ ID NO: 984.In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 45E, 45E / 111R / 226N / 347S, 45E / 853L, 110L / 114Q, 111L, 111L / 127M / 226N / 244Y, 111L / 210T, 111L / 244R, 111L / 244Y / 853L, 111L / 853L, 111R / 127M / 226N / 244Y, 111R / 194T / 244Y / 347S / 853L, 111R / 194T / 347S, 111R / 194T / 347S / 853L, 111R / 210T / 347S / 853L, 111R / 226N, 111R / 226N / 244Y / 347S, 111R / 347E / 853L, 111R / 853L, 111W, 111W / 226N / 853L / 969S, 114Q, 114Q / 245V, 127M / 210T / 244Y, 127M / 210T / 244Y / 853L, 127M / 210T / 347E / 853L, 127M / 244Y, 127M / 347E, 194T / 853L, 210T / 853L, 226N / 236G / 244Y / 347S / 853L, 236G / 244Y, 237R, 237R / 245V, 244Y / 853L, 245S, and 245V, wherein the positions are numbered with reference to SEQ ID NO: 984.In some further embodiments, the engineered polypeptide comprises at least one substitution or one set of substitutions selected from the group consisting of A45E, A45E / H111R / A226N / P347S, A45E / P853L, Q110L / G114Q, H111L, H111L / L127M / A226N / P244Y, H111L / I210T, H111L / P244R, H111L / P244Y / P853L, H111L / P853L, H111R / L127M / A226N / P244Y, H111R / P194T / P244Y / P347S / P853L, H111R / P194T / P347S, H111R / P194T / P347S / P853L, H111R / I210T / P347S / P853L, H111R / A226N, H111R / A226N / P244Y / P347S, H111R / P347E / P853L, H111R / P853L, H111W, H111W / A226N / P853L / P969S, G114Q, G114Q / E245V, L127M / I210T / P244Y, L127M / I210T / P244Y / P853L, L127M / I210T / P347E / P853L, L127M / P244Y, L127M / P347E, P194T / P853L, I210T / P853L, A226N / T236G / P244Y / P347S / P853L, T236G / P244Y, Q237R, Q237R / E245V, P244Y / P853L, E245S, and E245V, wherein the positions are numbered with reference to SEQ ID NO: 984.

[0029] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 984 comprising at least one substitution or one substitution set selected from: 518, 518 / 652, 519, 562, 563, 584 / 724, 586, 616, 618, 619, 621, 623, 628, 640, 653, and 666, wherein positions are numbered with reference to SEQ ID NO: 984. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: 518G, 518N / 652S, 518S, 518T, 519L, 562P, 563L, 563S, 584G / 724P, 586V, 616V, 618G, 618L, 619G, 619R, 621T, 623P, 628S, 640L, 653A, 653R, 653T, and 666R, wherein positions are numbered with reference to SEQ ID NO: 984. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: D518G, D518N / A652S, D518S, D518T, S519L, G562P, V563L, V563S, A584G / S724P, I586V, E616V, R618G, R618L, D619G, D619R, M621T, S623P, Y628S, K640L, D653A, D653R, D653T, and N666R, wherein positions are numbered with reference to SEQ ID NO: 984.

[0030] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 1160, comprising at least one substitution or one substitution set selected from 132 / 366 / 433 / 463 / 467 / 793, 132 / 366 / 467 / 661, 132 / 467 / 468 / 506 / 793, 132 / 468 / 793, 132 / 1025, 183 / 1025 / 1045 / 1048, 290 / 366 / 433 / 463 / 467, 290 / 433 / 467 / 793 / 1025, 290 / 433 / 793, 290 / 793, 290 / 1025, 366 / 433, 433, 433 / 467 / 1025, 433 / 506 / 1025, 433 / 790, 463 / 793, 473, 793, and 1025, wherein the positions are numbered with reference to SEQ ID NO: 1160. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 132G / 366I / 433K / 463A / 467K / 793T, 132G / 366I / 467G / 661D, 132G / 467K / 468F / 506V / 793T, 132G / 468F / 793T, 132G / 1025T, 183G / 1025A / 1045N / 1048R, 290T / 366I / 433K / 463A / 467K, 290T / 433K / 467K / 793T / 1025T, 290T / 433K / 793T, 290T / 793T, 290T / 1025T, 366I / 433K, 433K, 433K / 467K / 1025T, 433K / 506V / 1025T, 433K / 790H, 463A / 793T, 473L, 793T, and 1025T, wherein the positions are numbered with reference to SEQ ID NO: 1160.In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from: E132G / T366I / V433K / S463A / S467K / Q793T, E132G / T366I / S467G / G661D, E132G / S467K / A468F / M506V / Q793T, E132G / A468F / Q793T, E132G / S1025T, D183G / S1025A / D1045N / L1048R, K290T / T366I / V433K / S463A / S467K, K290T / V433K / S467K / Q793T / S1025T, K290T / V433K / Q793T, K290T / Q793T, K290T / S1025T, T366I / V433K, V433K, V433K / S467K / S1025T, V433K / M506V / S1025T, V433K / Y790H, S463A / Q793T, K473L, Q793T, and S1025T, wherein the positions are numbered with reference to SEQ ID NO: 1160.

[0031] The present application also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 1266, comprising at least one substitution or one substitution set selected from 114 / 230 / 446 / 853, 114 / 292 / 293 / 296 / 463 / 853, 132 / 183 / 366 / 467 / 661 / 1025 / 1045 / 1048, 132 / 290 / 366 / 433 / 467 / 661 / 793 / 1025, 132 / 290 / 366 / 467 / 661 / 793, 132 / 290 / 366 / 467 / 661 / 1025, 132 / 366 / 433 / 467 / 506 / 661 / 1025, 132 / 366 / 433 / 467 / 661, 132 / 366 / 433 / 467 / 661 / 790, 132 / 366 / 463 / 467 / 661 / 793, 132 / 366 / 467 / 473 / 661, 132 / 366 / 467 / 661 / 793, 366 / 467 / 468 / 506 / 661 / 793, 433 / 463 / 467 / 661 / 793, 463, 463 / 853, 689, 720, 724, 730, 769, 780, 792, 810, 817, 824, 853, 923, 926, 939, 952, 962, 968, 974, 979, 981, 995, 1006, 1015, 1017, 1022, 1027, 1031, and 1040, wherein the positions are numbered with reference to SEQ ID NO: 1266.In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from the group consisting of 114R / 230V / 446S / 853L, 114R / 292S / 293T / 296Q / 463V / 853L, 132E / 183G / 366T / 467S / 661G / 1025A / 1045N / 1048R, 132E / 290T / 366T / 433K / 467K / 661G / 793T / 1025T, 132E / 290T / 366T / 467S / 661G / 793T, 132E / 290T / 366T / 467S / 661G / 1025T, 132E / 366T / 433K / 467S / 506V / 661G / 1025T, 132E / 366T / 433K / 467S / 661G, 132E / 366T / 433K / 467S / 661G / 790H, 132E / 366T / 463A / 467S / 661G / 793T, 132E / 366T / 467S / 473L / 661G, 132E / 366T / 467S / 661G / 793T, 366T / 467K / 468F / 506V / 661G / 793T, 433K / 463A / 467K / 661G / 793T, 463V, 463V / 853L, 689G, 720Q, 724W, 730W, 769T, 780I, 792I, 792V, 810S, 817G, 824G, 853L, 923R, 923V, 926R, 939D, 952G, 962R, 968L, 974V, 979Q, 981N, 995G, 1006S, 1015L, 1017A, 1017G, 1017L, 1022S, 1022W, 1027R, 1031Q, 1040D, and 1040R, wherein the positions are numbered with reference to SEQ ID NO: 1266.G114R / A292S / E293T / A296Q / S463V / P853L, G132E / D183G / I366T / G467S / D661G / S1025A / D1045N / L1048R, G132E / K290T / I366T / V433K / G467K / D661G / Q793T / S1025T, G132E / K290T / I366T / G467S / D661G / Q793T, G132E / K290T / I366T / G467S / D661G / S1025T, G132E / I366T / V433K / G467S / M506V / D661G / S1025T, G132E / I366T / V433K / G467S / D661G, G132E / I366T / V433K / G467S / D661G / Y790H, G132E / I366T / S463A / G467S / D661G / Q793T, G132E / I366T / G467S / K473L / D661G, G132E / I366T / G467S / D661G / Q793T, I366T / G467K / A468F / M506V / D661G / Q793T, V433K / S463A / G467K / D661G / Q793T, S463V, S463V / P853L, L689G, G720Q, S724W, E730W, A769T, E780I, E792I, E792V, A810S, E817G, S824G, P853L, E923R, E923V, Q926R, S939D, N952G, H962R, M968L, T974V, V979Q, E981N, A995G, M1006S, M1015L, S1017A, S1017G, S1017L, H1022S, H1022W, A1027R, L1031Q, G1040D, and G1040R, wherein the positions are numbered with reference to SEQ ID NO: 1266.

[0032] The present disclosure also provides an engineered polypeptide comprising an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity to the reference sequence of SEQ ID NO: 1266 comprising at least one substitution or one substitution set selected from 458, 518 / 653, 519 / 628, 616 / 619, and 653, wherein the positions are numbered with reference to SEQ ID NO: 1266. In some embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from 458A, 518G / 653A, 518G / 653R, 519L / 628S, 616V / 619G, and 653R, wherein the positions are numbered with reference to SEQ ID NO: 1266. In some further embodiments, the engineered polypeptide comprises at least one substitution or one substitution set selected from G458A, D518G / D653A, D518G / D653R, S519L / Y628S, E616V / D619G, and D653R, wherein the positions are numbered with reference to SEQ ID NO: 1266.

[0033] In one embodiment, an engineered cytochrome P450-BM3 polypeptide variant is provided, the engineered cytochrome P450-BM3 polypeptide variant comprising a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to the sequence of at least one engineered cytochrome P450-BM3 variant listed in Tables 2.1, 3.1, 4.1, 5.1, 6.1, 7.1, 8.1, 9.1, 10.1, 10.2, 11.1, 11.2, 12.1, 13.1, 14.1, 14.2, 15.1, 16.1, 17.1, 18.1, 19.1, and 19.2.

[0034] In various embodiments, the engineered cytochrome P450-BM3 variant comprises at least one improved property compared to a wild-type B. megaterium cytochrome P450-BM3 or an engineered P450-BM3 variant. In some aspects, the improved property includes improved activity on a substrate. In certain aspects, the substrate includes l-tert-butoxycarbonylaminocyclopentane acid (Compound (1)). In some aspects, the improved property includes improved thermal stability or increased activity on a substrate after preincubation at 42.5 °C. In some aspects, the improved property includes improved stereoselectivity on one or more diastereomeric products.

[0035] The present invention provides engineered cytochrome P450-BM3 variants, wherein the variants are purified. The present invention also provides compositions comprising at least one engineered P450-BM3 variant.

[0036] The present invention further provides isolated recombinant polynucleotide sequences encoding the recombinant cytochrome P450-BM3 polypeptide variants provided herein. In some embodiments, the isolated recombinant polynucleotide sequence comprises SEQ ID NO: 3, 35, 65, 71, 197, 225, 243, 285, 357, 409, 533, 733, 747, 827, 967, 983, 1159, or 1265, or a functional fragment thereof. In various aspects, the polynucleotide sequence is operably linked to a control sequence. In some aspects, the polynucleotide sequence is codon-optimized. In some embodiments, the polynucleotide sequence comprises the polynucleotide sequence set forth in the odd-numbered sequences of SEQ ID NO: 3-1367.

[0037] The present invention also provides expression vectors comprising at least one polynucleotide sequence provided herein. In some further embodiments, the vector comprises at least one polynucleotide sequence operably linked to at least one regulatory sequence suitable for expressing the polynucleotide sequence in a suitable host cell. In some embodiments, the host cell is a prokaryotic or eukaryotic cell. In some further embodiments, the host cell is a prokaryotic cell. In some further embodiments, the host cell is E. coli. The present invention also provides host cells comprising the vectors provided herein.

[0038] The present invention also provides methods for producing at least one recombinant cytochrome P450-BM3 variant, the method comprising culturing a host cell provided herein under conditions such that at least one recombinant cytochrome P450-BM3 variant provided herein is produced by the host cell. In some further embodiments, the method further comprises the step of recovering at least one recombinant cytochrome P450 variant. In some aspects, the method further comprises the step of purifying the at least one engineered cytochrome P450-BM3 variant. DETAILED DESCRIPTION

[0039] The present invention provides improved P450-BM3 variants with improved activity. In some embodiments, the P450-BM3 variants exhibit improved activity on l-tert-butoxycarbonylaminocyclopentane acid. P450-BM3 enzymes exhibit the highest catalytic rate among P450 monooxygenases due to efficient electron transfer between the fused reductase and heme domain (see, e.g., Noble et al., Biochem. J., 339:371-379

[1999] ; and Munro et al., Eur. J. Biochem., 239:403-409

[2009] ). Thus, P450-BM3 is a highly desirable enzyme for manipulating biotechnological processes (see, e.g., Sawayama et al., Chem., 15:11723-11729

[2009] ; Otey et al., Biotechnol. Bioeng., 93:494-499

[2006] ; Damsten et al., Biol. Interact., 171:96-107

[2008] ; and Di Nardo and Gilardi, Int. J. Mol. Sci., 13:15901-15924). However, there remains a need in the art for P450 enzymes that exhibit activity on a variety of substrates, including l-tert-butoxycarbonylaminocyclopentane acid. The present invention provides P450-BM3 variants with improved enzymatic activity on l-tert-butoxycarbonylaminocyclopentane acid compared to the parent P450-BM3 sequence (i.e., SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266).

[0040] In some embodiments, the present invention provides P450-BM3 variants that provide improved overall conversion / activity percentage for oxidation of a substrate. In some embodiments, the present invention provides P450-BM3 variants that provide improved thermal stability in oxidation of a substrate. In some embodiments, the present invention provides P450-BM3 variants that provide improved stereoselectivity on various diastereomers in oxidation of a substrate. In particular, beneficial diversity was identified and recombined based on HTP screening results during development of the present invention.

[0041] Abbreviations and definitions:

[0042] Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, microbiology, organic chemistry, analytical chemistry, and nucleic acid chemistry described below are those well- established in the art. Such techniques are well-known and described in a number of text books and reference books. Standard techniques or modifications thereof are used for chemical syntheses and chemical analyses. All patents, patent applications, articles, and publications mentioned herein, including both those cited above and those cited below, are hereby expressly incorporated herein by reference in their entirety.

[0043] Although any suitable methods and materials known to those skilled in the art can be used in the practice of the application, certain methods and materials are described herein. It is understood that the present application is not limited to the particular methodology, protocols, and reagents described, as these can vary depending upon the context in which they are used. Accordingly, the terms defined immediately below are intended to have the meanings set forth in the specification, as a whole, and not merely in the specific examples described herein. All patents, patent applications, articles, and publications mentioned herein, including both those cited above and those cited below, are hereby expressly incorporated herein by reference in their entirety.

[0044] In addition, as used herein, the singular “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise.

[0045] Numerical ranges include the numbers defining the range. Thus, every numerical range disclosed herein is intended to include each and every number falling within the range, irrespective of the existence of any

[0046] The term “about” means an acceptable error for a particular value. In some instances, “about” means within 0.05%, 0.5%, 1.0%, or 2.0% of a given range. In some instances, “about” means within 1, 2, 3, or 4 standard deviations of a given value.

[0047] In addition, the headings provided herein are not limitations on the various aspects or embodiments of the application, which can be had by reference to the specification as a whole. Accordingly, the terms defined immediately below are intended to be merely exemplary and are not intended to limit the present application, as encompassed by the entire specification and claims. Nonetheless, for ease of reference, numerous terms are defined below.

[0048] Unless otherwise indicated, nucleic acids are written left to right in 5' to 3' orientation, and amino acid sequences are written left to right in amino to carboxy orientation, respectively.

[0049] As used herein, the term "comprising" and its cognate forms are used in the inclusive sense (i.e., equivalent to the term "including" and its corresponding cognate forms).

[0050] "EC" number refers to the enzyme nomenclature of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (NC-IUBMB). The IUBMB biochemical classification is a numerical classification system for enzymes based on the chemical reactions they catalyze.

[0051] "ATCC" refers to the American Type Culture Collection, whose biorepository includes genes and strains.

[0052] "NCBI" refers to the National Center for Biotechnology Information and the sequence databases provided therein.

[0053] As used herein, "cytochrome P450-BM3" and "P450-BM3" refer to a cytochrome P450 enzyme obtained from Bacillus megaterium (now also known as Paenibacillus megteήus) that catalyzes NADPH-dependent hydroxylation of long-chain fatty acids, alcohols and amides, as well as epoxidation of unsaturated fatty acids

[0054] "Protein," "polypeptide," and "peptide" can be used interchangeably herein to mean a polymer of at least two amino acids linked by amide bonds, without regard to length or post-translational modification (e.g., glycosylation or phosphorylation).

[0055] "Amino acid" is referred to herein by its commonly known three letter symbol or by the one letter symbol recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Likewise, nucleotides can be referred to by their universally accepted single letter codes.

[0056] The terms "engineered," "recombinant," "non-naturally occurring," and "variant," when used with respect to a cell, polynucleotide, or polypeptide, refer to a material or a material corresponding to a natural or native form of the material that has been modified in a manner that does not otherwise exist in nature, but is produced or derived from synthetic materials and / or by manipulation using recombinant techniques.

[0057] As used herein, "wild type" and "naturally occurring" refer to the form found in nature. For example, a wild type polypeptide or polynucleotide sequence is a sequence that exists in an organism that can be isolated from a source in nature and has not been intentionally modified by human manipulation.

[0058] "Encoding sequence" refers to that portion of a nucleic acid (e.g., a gene) that encodes an amino acid sequence of a protein.

[0059] The term "percent (%) sequence identity" is used herein to refer to a comparison between polynucleotides and polypeptides, and is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window can comprise additions or deletions (i.e., gaps) as compared to the reference sequence for optimal alignment of the two sequences. The percent sequence identity is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to yield the percent sequence identity. Alternatively, the percent sequence identity is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences, or aligning the nucleic acid bases or amino acid residues with gaps to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to yield the percent sequence identity. Those of skill in the art will appreciate that there are many established algorithms that can be used to align two sequences. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman (Smith and Waterman, Adv. Appl. Math., 2:482

[1981] ), by the homology alignment algorithm of Needleman and Wunsch (Needleman and Wunsch, J. Mol. Biol., 48:443

[1970] , by the search for similarity method of Pearson and Lipman (Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444

[1988] ), by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin software package), or by visual inspection. Examples of algorithms that are suitable for determining percent sequence identity and sequence similarity include, but are not limited to, the BLAST and BLAST 2.0 algorithms (see, e.g., Altschul et al., J. Mol. Biol., 215:403-410

[1990] ; and Altschul et al., 1977, Nucl. Acids Res., 3389-3402

[1977] , respectively). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information website. The algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (see, e.g., Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them.The word hits are then extended in each direction on the sequence, as far as the cumulative alignment score can be increased. For nucleotide sequences, the parameters W, T and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a wordlength (W) of 11, an expectation value (E) of 10, M=5, N=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength of 3, and expectation value (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915

[1989] ). Exemplary determination of sequence

[0060] A "reference sequence" is a defined sequence used as a basis for sequence comparison. A reference sequence can be a subset of a larger sequence, for example, a segment of a full-length gene or polypeptide sequence. Generally, a reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 residues in length, at least 50 residues in length, at least 100 residues in length, or is the full length of the nucleic acid or polypeptide. Sequence comparisons between two (or more) polynucleotides or polypeptides are typically performed by comparing sequences of the two polynucleotides or polypeptides over a "comparison window" to identify and compare local regions of sequence similarity. In some embodiments, a "reference sequence" can be based on a primary amino acid sequence, wherein the reference sequence is a sequence that can have one or more variations in the primary sequence. A "comparison window" refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acid residues wherein a sequence can be compared to a reference sequence of at least 20 contiguous nucleotides or amino acids and wherein the portion of the sequence in the comparison window can comprise up to 20% or less of the total number of additions or deletions (i.e., gaps) for optimal alignment of the two sequences. A comparison window can be longer than 20 contiguous residues, and optionally includes windows of 30, 40, 50, 100 contiguous residues, or longer.

[0061] "Corresponding to," "with reference to," or "relative to," when used in the context of numbering of a given amino acid or polynucleotide sequence, means that the numbering of the residues of a given reference sequence is specified when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue numbering or residue position of a given polymer is specified relative to the reference sequence rather than by the actual numerical position of the residue within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence, such as the amino acid sequence of an engineered P450-BM3, can be aligned with a reference sequence by introducing gaps to optimize the matching of residues between the two sequences. In these cases, the numbering of the residues in the given amino acid or polynucleotide sequence is made relative to the reference sequence to which it is aligned, notwithstanding the presence of gaps.

[0062] An "amino acid difference" or "residue difference" refers to a difference in the amino acid residue at a position of a polypeptide sequence relative to the amino acid residue at the corresponding position in a reference sequence. The position of the amino acid difference is generally referred to herein as "Xn," where n refers to the corresponding position in the reference sequence upon which the residue difference is based. For example, "a residue difference at position X93 compared to SEQ ID NO: 2" refers to a difference in the amino acid residue at the position of the polypeptide corresponding to position 93 of SEQ ID NO: 2. Thus, if the reference polypeptide of SEQ ID NO: 2 has a serine at position 93, "a residue difference at position X93 compared to SEQ ID NO: 2" refers to an amino acid substitution of any residue other than serine at said position of the polypeptide corresponding to position 93 of SEQ ID NO: 2. In most instances herein, a particular amino acid residue difference at a position is denoted as "XnY," where "Xn" designates the corresponding position as described above, and "Y" is the one-letter identifier for the amino acid found in the engineered polypeptide (i.e., the residue that differs from the reference polypeptide). In some instances (e.g., in Tables 2.1, 3.1, 4.1, 5.1, 6.1, 7.1, 8.1, 9.1, 10.1, 10.2, 11.1, 11.2, 12.1, 13.1, 14.1, 14.2, 15.1, 16.1, 17.1, 17.2, 18.1, 19.1, and 19.2), the present disclosure also provides for a particular amino acid difference denoted by the conventional notation "AnB," where A is the one-letter identifier for the residue in the reference sequence, "n" is the number of the residue position in the reference sequence, and B is the one-letter identifier for the residue substitution in the sequence of the engineered polypeptide. In some instances, a polypeptide of the present disclosure can include one or more amino acid residue differences relative to a reference sequence, which are indicated by a list of the designated positions at which there is a residue difference relative to the reference sequence. In some embodiments, where more than one amino acid can be used at a particular residue position of a polypeptide, the various amino acid residues that can be used are separated by a " / " (e.g., X307H / X307P or X307H / P). The present disclosure includes engineered polypeptide sequences comprising one or more amino acid differences, including one or both of conservative and non-conservative amino acid substitutions.

[0063] A "conservative amino acid substitution" refers to the replacement of a residue by a different residue having similar side chain. And thus generally involves the substitution of an amino acid in a polypeptide with an amino acid within the same or similar defined class of amino acids. By way of example and not limitation, amino acids with aliphatic side chains can be substituted with another aliphatic amino acid (e.g., alanine, valine, leucine, and isoleucine); amino acids with hydroxyl side chains are substituted with another amino acid with a hydroxyl side chain (e.g., serine and threonine); amino acids with aromatic side chains are substituted with another amino acid with an aromatic side chain (e.g., phenylalanine, tyrosine, tryptophan, and histidine); amino acids with basic side chains are substituted with another amino acid with a basic side chain (e.g., lysine and arginine); amino acids with acidic side chains are substituted with another amino acid with an acidic side chain (e.g., aspartic acid or glutamic acid); and / or hydrophobic or hydrophilic amino acids are replaced with another hydrophobic or hydrophilic amino acid, respectively.

[0064] A "non-conservative substitution" refers to the replacement of an amino acid in a polypeptide with an amino acid having significantly different side chain properties. Non-conservative substitutions can use amino acids between, rather than within, defined groups and affect (a) structure of the peptide backbone in the area of the substitution (e.g., as proline in a flexible region), (b) charge or hydrophobicity, or (c) the bulk of the side chain. By way of example and not limitation, exemplary non-conservative substitutions can be acid for basic or aliphatic amino acids; small for aromatic amino acids; and hydrophilic for hydrophobic amino acids.

[0065] A "deletion" refers to a modification of a polypeptide by removing one or more amino acids from a reference polypeptide. A deletion can include removal of 1 or more amino acids, 2 or more amino acids, 5 or more amino acids, 10 or more amino acids, 15 or more amino acids, or 20 or more amino acids, up to 10% of the total number of amino acids comprising the reference enzyme, or up to 20% of the total number of amino acids comprising the reference enzyme, while retaining enzyme activity and / or retaining improved properties of the engineered P450-BM3 enzyme. Deletions can be to internal portions and / or terminal portions of the polypeptide. In various embodiments, deletions can include contiguous segments or can be non-contiguous.

[0066] An "insertion" refers to a modification of a polypeptide by adding one or more amino acids to a reference polypeptide. An insertion can be in an internal portion of the polypeptide, or to the carboxyl or amino terminus. Insertions as used herein include fusion proteins known in the art. An insertion can be a contiguous segment of amino acids or separated by one or more amino acids in the naturally occurring polypeptide.

[0067] "Functional fragment" or "biologically active fragment," as used interchangeably herein, refers to a polypeptide having an amino-terminal and / or carboxyl-terminal deletion and / or internal deletion, but wherein the remaining amino acid sequence is identical to the corresponding position in the sequence to which it is compared (e.g., a full-length engineered P450-BM3 of the application) and substantially retains all of the activities of the full-length polypeptide.

[0068] "Isolated polypeptide" refers to a polypeptide that is substantially separated from other contaminants with which it is naturally associated (e.g., proteins, lipids, and polynucleotides). The term includes a polypeptide that has been removed from its naturally occurring environment or expression system (e.g., a host cell or in vitro synthesis) or purified. A recombinant P450-BM3 polypeptide can be present within a cell, in cell culture media, or prepared in various forms such as a lysate or an isolated preparation. Thus, in some embodiments, a recombinant P450-BM3 polypeptide can be an isolated polypeptide.

[0069] "Substantially pure polypeptide" refers to a composition wherein the polypeptide material is the predominant species (i.e., it is the most abundant material in the composition on a molar or weight basis) and is typically a substantially pure composition when the polypeptide of interest comprises at least about 50% (on a molar or weight basis) of the macromolecular species present. However, in some embodiments, a composition comprising P450-BM3 comprises P450-BM3 in a purity of less than 50% (e.g., about 10%, about 20%, about 30%, about 40%, or about 50%). Typically, a substantially pure P450-BM3 composition comprises about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, and about 98% or more of all macromolecular species present in the composition (on a molar or weight basis). In some embodiments, the polypeptide of interest is purified to essentially homogeneity (i.e., no detectable contaminating material by conventional detection methods) wherein the composition consists essentially of a single macromolecular species. Solvent species, small molecules (< 500 daltons), and elemental ion species are not considered macromolecular species. In some embodiments, an isolated recombinant P450-BM3 polypeptide is a substantially pure polypeptide composition.

[0070] "Improved enzyme properties" refers to an engineered P450-BM3 polypeptide that exhibits an improvement in any enzyme property as compared to a reference P450-BM3 polypeptide and / or a wild-type P450-BM3 polypeptide or another engineered P450-BM3 polypeptide. Improved properties include, but are not limited to, properties such as increased protein expression, increased thermal activity, increased thermal stability, increased pH activity, increased stability, increased enzyme activity, increased substrate specificity or affinity, increased specific activity, increased resistance to substrate or end product inhibition, increased chemical stability, improved stereoselectivity, improved chemoselectivity, improved solvent stability, increased tolerance to acidic pH, increased tolerance to proteolytic activity (i.e., decreased sensitivity to proteolysis), decreased aggregation, increased solubility, and altered temperature profile.

[0071] "Increased enzyme activity" or "enhanced catalytic activity" refers to an improved property of an engineered P450-BM3 polypeptide that can be expressed as an increase in specific activity (e.g., product produced / time / protein weight) or an increase in the percentage conversion of substrate to product (e.g., the percentage conversion of a starting amount of substrate to product using a specified amount of P450-BM3 over a specified time period) as compared to a reference P450-BM3 enzyme. Exemplary methods for determining enzyme activity are provided in the Examples. Any property related to enzyme activity can be affected, including the classical enzyme properties of K m , V max , or k cat , changes in which can result in increased enzyme activity. Improvements in enzyme activity can range from about 1.1-fold the enzyme activity of the corresponding wild-type enzyme to up to 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold, or more of the enzyme activity of a naturally occurring P450-BM3 or another engineered P450-BM3 from which the P450-BM3 polypeptide is derived.

[0072] "Conversion" refers to the enzymatic (or biotransformation) of a substrate to the corresponding product. "Percentage conversion" refers to the percentage of substrate that is converted to product over a period of time under specified conditions. Thus, the "enzyme activity" or "activity" of a P450-BM3 polypeptide can be expressed as the "percentage conversion" of substrate to product over a specified time period.

[0073] Enzymes having "generalist properties" (or "generalist enzymes") are enzymes that exhibit improved activity on a wide variety of substrates as compared to the parent sequence. Generalist enzymes do not necessarily exhibit improved activity on every possible substrate. In particular, the present application provides P450-BM3 variants having generalist properties in that they exhibit similar or improved activity relative to the parent gene on a wide variety of spatial and electronic diversity substrates. Further, the generalist enzymes provided herein are engineered to improve in a wide variety of different API-like molecules to increase the production of metabolites / products.

[0074] "Hybridization stringency" relates to hybridization conditions, such as wash conditions, in nucleic acid hybridization. Typically, hybridization reactions are performed under conditions of lower stringency, followed by washing under different but higher stringency conditions. The term "moderate stringency hybridization" refers to conditions that allow a target DNA to bind a complementary nucleic acid that has about 60% identity, preferably about 75% identity, about 85% identity, greater than about 90% identity to the target polynucleotide. Exemplary moderate stringency conditions are conditions equivalent to washing a filter in 42°C 50% formamide, 5xSSPE, 0.2% SDS, followed by 0.2xSSPE, 0.2% SDS at 42°C. "High stringency hybridization" generally refers to conditions that are about 10°C or less from the thermal melting point (Tm) of the most stable (i.e., longest) nucleic acid sequence under solution conditions in the absence of the nucleic acid to which it is hybridizing, as determined for a defined polynucleotide sequence. In some embodiments, high stringency conditions refer to those that allow hybridization only of those nucleic acid sequences that form stable hybrids at 65°C in 0.018 M NaCl (i.e., if a hybrid is not stable at 65°C in 0.018 M NaCl, it will not be stable under high stringency conditions, as contemplated herein). For example, high stringency conditions can be provided by hybridization at conditions equivalent to 42°C in 50% formamide, 5x Denhart's solution, 5x SSPE, 0.2% SDS, followed by washing in 0.1x SSPE and 0.1% SDS at 65°C. Another high stringency condition is hybridization at conditions equivalent to hybridization in 5X SSC containing 0.1% (w:v) SDS at 65°C and washing in 0.1x SSC containing 0.1% SDS at 65°C. Other high stringency hybridization conditions, as well as moderate stringency conditions, are described in the references cited above. m about 10°C or less from the thermal melting point (Tm) of the most stable (i.e., longest) nucleic acid sequence under solution conditions in the absence of the nucleic acid to which it is hybridizing, as determined for a defined polynucleotide sequence. In some embodiments, high stringency conditions refer to those that allow hybridization only of those nucleic acid sequences that form stable hybrids at 65°C in 0.018 M NaCl (i.e., if a hybrid is not stable at 65°C in 0.018 M NaCl, it will not be stable under high stringency conditions, as contemplated herein). For example, high stringency conditions can be provided by hybridization at conditions equivalent to 42°C in 50% formamide, 5x Denhart's solution, 5x SSPE, 0.2% SDS, followed by washing in 0.1x SSPE and 0.1% SDS at 65°C. Another high stringency condition is hybridization at conditions equivalent to hybridization in 5X SSC containing 0.1% (w:v) SDS at 65°C and washing in 0.1x SSC containing 0.1% SDS at 65°C. Other high stringency hybridization conditions, as well as moderate stringency conditions, are described in the references cited above.

[0075] " Codon-optimized" refers to the alteration of codons of a polynucleotide encoding a protein to those codons that are preferentially used in a particular organism, such that the encoded protein is more efficiently expressed in the organism of interest. Although the genetic code is degenerate, in that most amino acids are represented by several codons (referred to as "synonymous" or "synonymous" codons), it is well known that codon usage in a particular organism is non-random and biased towards particular codon triplets. This codon usage bias can be higher for given genes, genes of common function or ancestral origin, highly expressed proteins versus low copy number proteins, and for aggregated protein coding regions of an organism's genome. In some embodiments, the polynucleotide encoding the P450-BM3 enzyme can be codon-optimized for optimal production from a host organism selected for expression.

[0076] "Control sequences" herein refer to all components, which are necessary or advantageous for the expression of a polynucleotide and / or polypeptide of the present application. Each control sequence can be native or foreign to the nucleic acid sequence encoding the polypeptide. Such control sequences include, but are not limited to, a leader sequence, a polyadenylation sequence, a pre / pro-peptide sequence, a promoter sequence, a signal peptide sequence, an initiation sequence, and a transcription terminator. These control sequences are operably linked to the coding region of the nucleic acid sequence encoding the polypeptide. The control sequences can provide for the initiation of transcription and / or the termination of transcription and translation, and can contain sites for attachment of ribosomes and for proper initiation and processing of the mRNA. The control sequences can further contain sites for attachment of the various mRNA effector molecules, including, but not limited to, ribosomes, transcription and translation initiation complexes, and mRNA capping, splicing, editing, and / or stability factors.

[0077] "Operably linked" is defined herein as a configuration in which the control sequences are suitably positioned (i.e., in functional relationship) with respect to the polynucleotide of interest, such that the control sequences direct or modulate the expression of the polynucleotide and / or polypeptide of interest.

[0078] "Promoter sequence" refers to a nucleic acid sequence, such as a coding sequence, that is recognized by a host cell for the expression of a polynucleotide of interest. Promoter sequences contain transcriptional control sequences that mediate the expression of a polynucleotide of interest. Promoters which can be used include any nucleic acid sequence which shows transcriptional activity in the host cell of choice including mutant, truncated, and hybrid promoters, and can be derived from genes encoding proteins either homologous or heterologous to the host cell. The promoter region can be obtained from the same gene as the coding sequence or from a different gene.

[0079] "Suitable reaction conditions" refer to those conditions in the enzymatic conversion reaction solution (e.g., ranges of enzyme loading, substrate loading, temperature, pH, buffer, co-solvent, etc.) under which a P450-BM3 polypeptide of the application is capable of converting a substrate to a desired product compound. Exemplary "suitable reaction conditions" are provided herein and illustrated by the examples. "Loading," such as in "compound loading" or "enzyme loading" or "cofactor loading," refers to the concentration or amount of a component in the reaction mixture at the start of the reaction. In the context of an enzymatic conversion process, "substrate" refers to the compound or molecule acted upon by a P450-BM3 polypeptide. In the context of an enzymatic conversion process, "product" refers to the compound or molecule produced by the action of a P450-BM3 polypeptide on a substrate.

[0080] As used herein, the term "culturing" refers to growing a population of microbial cells under any suitable conditions (e.g., using liquid, gel, or solid culture media).

[0081] Recombinant polypeptides can be produced using any suitable method known in the art. A gene encoding a wild-type polypeptide of interest can be cloned in a vector, such as a plasmid, and expressed in a desired host, such as E. coli, and the like. Variants of recombinant polypeptides can be generated by various methods known in the art. In fact, various different mutagenesis techniques are well known to those skilled in the art. Moreover, mutagenesis kits are also available from many commercial molecular biology suppliers. Methods can be used to make specific substitutions at defined amino acids (site-directed), specific or random mutations in local regions of the gene (region-specific), or random mutagenesis across the entire gene (e.g., saturation mutagenesis). Numerous suitable methods are known to those skilled in the art to generate enzyme variants, including but not limited to site-directed mutagenesis of single-stranded DNA or double-stranded DNA using PCR, cassette mutagenesis, gene synthesis, error-prone PCR, shuffling, and chemical saturation mutagenesis, or any other suitable method known in the art. Non-limiting examples of methods for DNA and protein engineering are provided in U.S. Patent No. 6,117,679; U.S. Patent No. 6,420,175; U.S. Patent No. 6,376,246; U.S. Patent No. 6,586,182; U.S. Patent No. 7,747,391; U.S. Patent No. 7,747,393; U.S. Patent No. 7,783,428; and U.S. Patent No. 8,383,346. After variants are generated, they can be screened for any desired property (e.g., high or increased activity, or low or decreased activity, increased thermal activity, increased thermal stability and / or acid pH stability, etc.). In some embodiments, "recombinant P450-BM3 polypeptides" (also referred to herein as "engineered P450-BM3 polypeptides," "variant P450-BM3 enzymes," and "P450-BM3 variants") are useful.

[0082] As used herein, a "vector" is a DNA construct used to introduce a DNA sequence into a cell. In some embodiments, a vector is an expression vector operably linked to suitable control sequences capable of effecting expression of a polypeptide encoded in the DNA sequence in a suitable host. In some embodiments, an "expression vector" has a promoter sequence operably linked to a DNA sequence (e.g., a transgene) to drive expression in a host cell, and in some embodiments, also includes a transcription terminator sequence.

[0083] As used herein, the term "expression" includes any step involved in the production of a polypeptide, including but not limited to transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of a polypeptide from a cell.

[0084] As used herein, the term "production" refers to the production of a protein and / or other chemical compound by a cell. The term is intended to encompass any step involved in the production of a polypeptide, including but not limited to transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of a polypeptide from a cell.

[0085] As used herein, an amino acid or nucleotide sequence (e.g., a promoter sequence, a signal peptide, a terminator sequence, etc.) is "heterologous" to another sequence with which it is operably linked if the two sequences are not naturally associated with each other in nature.

[0086] As used herein, the terms "host cell" and "host strain" refer to a suitable host for an expression vector comprising a DNA provided herein (e.g., a polynucleotide sequence encoding a P450-BM3 variant). In some embodiments, a host cell is a prokaryotic or eukaryotic cell as known in the art that has been transformed or transfected with a vector constructed using recombinant DNA technology.

[0087] The term "analog" means a polypeptide having greater than 70% sequence identity but less than 100% sequence identity (e.g., greater than 75%, 78%, 80%, 83%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity) to a reference polypeptide. In some embodiments, an analog means a polypeptide containing one or more non-naturally occurring amino acid residues, including but not limited to homoarginine, ornithine, and norvaline, as well as naturally occurring amino acids. In some embodiments, an analog also includes one or more D-amino acid residues and non-peptide linkages between two or more amino acid residues.

[0088] The term "effective amount" means an amount that is sufficient to achieve the desired result. An effective amount can be determined by one of ordinary skill in the art through the use of routine experimentation.

[0089] The terms "isolated" and "purified" are used in reference to a molecule (e.g., an isolated nucleic acid, polypeptide, etc.) or other component that has been removed from at least one other component with which it is naturally associated. The term "purified" does not require absolute purity; rather, it is intended as a relative definition.

[0090] Engineered P450-BM3 polypeptides:

[0091] The present application provides improved P450-BM3 variants having improved activity on l-tert-butoxycarbonylamino cyclopentanoic acid.

[0092] In some embodiments, the present disclosure provides P450-BM3 variants having improved monooxygenase activity on l-tert-butoxycarbonylamino cyclopentanoic acid (Compound (1)) and having increased product conversion to l-(tert-butoxycarbonylamino)-3-hydroxycyclopentanoic acid (Compound (2)) and / or l-(tert-butoxycarbonylamino)-3-oxocyclopentanoic acid (Compound (3)) compared to the starting polypeptide, as depicted in Scheme 1.

[0093]

[0094] Scheme 1

[0095] In some embodiments, the monooxygenase (also known as oxidase) reaction uses NADPH as a cofactor, which is recycled from NADP+ by glucose dehydrogenase (GDH-105, Codexis, Inc.), as depicted above in Scheme 1. However, use of the GDH recycling system can result in a downward shift in the reaction pH. In some embodiments, phosphite dehydrogenase (PDH-102, Codexis, Inc.) is used for cofactor recycling, as shown below in Scheme 2.

[0096]

[0097] Scheme 2

[0098] The P450-BM3 variants of the present disclosure produce four diastereomers of Compound (2) depicted in Scheme 3 below. In addition, hydroxylation is also observed at position 4, resulting in an additional diastereomer represented as Compound (4) in Scheme 3 below.

[0099]

[0100] Scheme 3

[0101] In some embodiments, the engineered P450-BM3 polypeptides of the present disclosure have regioselectivity for hydroxylation at position 2 compared to a reference P450-BM3 polypeptide. In some embodiments, the engineered P450-BM3 polypeptides of the present disclosure have stereoselectivity for (1R,3S)-2 and / or (1R,3R)-2 diastereomers compared to a reference P450-BM3 polypeptide.

[0102] In some embodiments, the engineered P450-BM3 polypeptides of the present disclosure can produce a ketone product from a substrate of compound (1), as depicted below in Scheme 4.

[0103]

[0104] Scheme 4

[0105] In some embodiments, the engineered P450-BM3 polypeptides of the present disclosure have regioselectivity for oxidation to a ketone at position 3 compared to a reference P450-BM3 polypeptide.

[0106] The present disclosure provides exemplary engineered P450-BM3 polypeptides (i.e., P450-BM3 variants) having P450-BM3 activity. The tables provided in the Examples show sequence structure information correlating particular amino acid sequence features with functional activity of engineered P450-BM3 polypeptides. This structure-function correlation information is provided in the form of particular amino acid residue differences relative to a reference engineered polypeptide, as shown in the Examples. The Examples further provide experimentally determined activity data for exemplary engineered P450-BM3 polypeptides.

[0107] In some embodiments, the engineered P450-BM3 polypeptides of the application having P450-BM3 activity comprise: a) an amino acid sequence having at least 80% sequence identity to reference sequence SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266; b) amino acid residue differences at one or more amino acid positions compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266; and c) which exhibit an improved property compared to the reference sequence selected from: i) enhanced catalytic activity, ii) increased thermal stability, iii) increased tolerance to acidic pH, iv) reduced aggregation, v) increased activity on l-tert-butoxycarbonylaminocyclopentanoic acid substrate, vi) increased solubility, vii) increased regioselectivity, or viii) increased selectivity for a desired chiral product, or a combination of any of i), ii), iii), iv), v), vi), vii), or viii).

[0108] In some embodiments, the engineered P450-BM3 exhibiting improved properties has at least about 85%, at least about 88%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% amino acid sequence identity to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, and has amino acid residue differences at one or more amino acid positions compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266 (such as at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 20, or more amino acid positions compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, or to a sequence having at least 85%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more amino acid sequence identity to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266). In some embodiments, the residue differences at one or more positions compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266 will include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more conservative amino acid substitutions. In some embodiments, the engineered P450-BM3 polypeptide is a polypeptide listed in any of Tables 2.1, 3.1, 4.1, 5.1, 6.1, 7.1, 8.1, 9.1, 10.1, 10.2, 11.1, 11.2, 12.1, 13.1, 14.1, 14.2, 15.1, 16.1, 17.1, 17.2, 18.1, 19.1, and 19.2.

[0109] In some embodiments, the engineered P450-BM3 that exhibits improved properties has at least 85%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% amino acid sequence identity to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266.

[0110] In some embodiments, the engineered P450-BM3 polypeptides of the present disclosure comprise an amino acid sequence having at least 85%, 90%, 95%, or 99% sequence identity to the reference sequence of SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266 and having amino acid residue differences at one or more amino acid positions compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, wherein the engineered P450-BM3 polypeptide converts 1-tert-butoxycarbonylaminocyclopentanoic acid (Compound (1)) to 1-(tert-butoxycarbonylamino)-3-hydroxycyclopentanoic acid (Compound (2)) and / or 1-(tert-butoxycarbonylamino)-3-oxocyclopentanoic acid (Compound (3)) with an activity that is at least 1.5-fold, 2.0-fold, 5.0-fold, 10.0-fold, 50.0-fold, 100.0-fold, or 500.0-fold greater than the reference engineered P450-BM3 polypeptide of SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266 compared to the starting polypeptide.

[0111] In some embodiments, the engineered P450-BM3 polypeptides of the present disclosure comprise a reference sequence having at least 85%, 90%, 95% or 99% sequence identity to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160 or 1266 and a sequence identity to SEQ ID NO: NO:4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160 or 1266, wherein the engineered P450-BM3 polypeptide converts 1-tert-butoxycarbonylaminocyclopentanoic acid (compound (1)) into 1-(tert-butoxycarbonylamino)-3-hydroxycyclopentanoic acid (compound (2)) and / or 1-(tert-butoxycarbonylamino)-3-oxocyclopentanoic acid (compound (3)), and the like. The invention further comprises an engineered P450-BM3 polypeptide comprising at least one of the peptides of formula (I) 1 or 2 and any of the peptides of formula (I) 2 or 3, wherein the peptides are expressed as adenosine monophosphate (EPO) or adenosine monophosphate (EPO) and / or succinate-binding protein (SBP) selected from the group consisting of peptides having adenosine monophosphate (SBP) and peptides having adenosine monophosphate (SBP) or adenosine monophosphate (SBP) selected from the group consisting of peptides having adenosine monophosphate (SBP) and peptides having adenosine monophosphate (SBP) and / or succinate-binding protein ...

[0112] In some embodiments, the engineered P450-BM3 polypeptide of the present disclosure comprises an amino acid sequence having at least 85%, 90%, 95%, or 99% sequence identity to the reference sequence of SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266 and having amino acid residue differences at one or more amino acid positions compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, wherein the engineered P450-BM3 polypeptide converts l-tert-butoxycarbonylaminocyclopentanoic acid (Compound (1)) to 1-(tert-butoxycarbonylamino)-3-hydroxycyclopentanoic acid (Compound (2)) and / or l-(tert-butoxycarbonylamino)-3-oxocyclopentanoic acid (Compound (3)) with at least 5%, 20%, 30%, 50%, 75%, 80%, or 90% enantiomeric excess of Compound (2) or Compound (3) compared to the reference engineered P450-BM3 polypeptide of SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266.

[0113] In some embodiments, the engineered P450-BM3 polypeptide comprises a functional fragment of an engineered P450-BM3 polypeptide encompassed by the present disclosure. A functional fragment has at least 95%, 96%, 97%, 98%, or 99% of the activity of the engineered P450-BM3 polypeptide from which it is derived (i.e., the parent engineered P450-BM3). A functional fragment comprises at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and even 99% of the parent sequence of the engineered P450-BM3. In some embodiments, the functional fragment is truncated by fewer than 5, fewer than 10, fewer than 15, fewer than 10, fewer than 25, fewer than 30, fewer than 35, fewer than 40, fewer than 45, and fewer than 50 amino acids.

[0114] Methods of using engineered P450-BM3 polypeptide enzymes:

[0115] In some embodiments, an engineered P450-BM3 polypeptide of the application having P450-BM3 activity comprises: a) an amino acid sequence having at least 85% sequence identity to reference sequence SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, or a fragment thereof; b) amino acid residue differences at one or more amino acid positions compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266; and c) which exhibits improved activity compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266.

[0116] In some embodiments, an engineered P450-BM3 exhibiting improved activity has at least 85%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more amino acid sequence identity to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, and has amino acid residue differences at one or more amino acid positions compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, such as at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 20 or more amino acid positions compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, or to a sequence having at least 85%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more amino acid sequence identity to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266.

[0117] In some embodiments, the engineered P450-BM3 polypeptide has improved activity compared to a reference P450-BM3 polypeptide when all other assay conditions are essentially the same. In some embodiments, this activity can be measured to assess the maximum activity (e.g., k cat ) of the enzyme under conditions that monitor enzyme activity using any suitable analytical system. In other embodiments, this activity can be measured at substrate concentrations that result in one-half, one- fifth, one-tenth, or less of the maximum activity. Under either analytical method, the engineered polypeptide has an improved activity level that is about 1.0-fold, 1.5-fold, 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, or more of the enzyme activity of the reference P450-BM3. In some embodiments, the engineered P450-BM3 polypeptide has improved activity compared to a reference P450-BM3 when measured by any standard assay, including but not limited to the assays described in the Examples.

[0118] In some embodiments, the engineered P450-BM3 polypeptides described herein can be used in methods of converting l-tert-butoxycarbonylamino cyclopentanoic acid (Compound (1)) to products l-(tert-butoxycarbonylamino)-3-hydroxycyclopentanoic acid (Compound (2)) and / or l-(tert-butoxycarbonylamino)-3-oxocyclopentanoic acid (Compound (3)). Generally, methods for performing a monooxygenase reaction include contacting or incubating a substrate compound with an engineered P450-BM3 polypeptide of the application in the presence of a co-substrate, such as NADP+, under reaction conditions suitable for forming a hydroxylated product, as shown in Scheme 1 and Scheme 2 above.

[0119] In embodiments described herein and in the Examples, various ranges of suitable reaction conditions that can be used in the methods include, but are not limited to, substrate loading, reducing agent, recycling system, pH, temperature, buffer, solvent system, polypeptide loading, and reaction time. Further suitable reaction conditions for performing methods of biocatalytically converting a substrate compound to a product compound using the engineered P450-BM3 polypeptides described herein can be readily optimized according to the guidance provided herein by routine experimentation, including but not limited to contacting an engineered P450-BM3 polypeptide and a substrate compound under experimental reaction conditions of concentration, pH, temperature, and solvent conditions, and detecting the product compound.

[0120] Suitable reaction conditions for using an engineered P450-BM3 polypeptide generally include an NADP+ co-substrate, which is used stoichiometrically in the monooxygenation reaction. In general, the co-substrate for an engineered P450-BM3 polypeptide is NADP+. Other reducing agents that are capable of acting as a co-substrate for an engineered P450-BM3 polypeptide can be used. In some embodiments, suitable reaction conditions can include a co-substrate concentration, particularly from about 0.0005 M to about 2 M, 0.01 M to about 2 M, 0.1 M to about 2 M, 0.2 M to about 2 M, from about 0.5 M to about 2 M, or from about 1 M to about 2 M. In some embodiments, reaction conditions include a co-substrate concentration of about 0.0001 M, 0.001 M, 0.01 M, 0.1 M, 0.2 M, 0.3 M, 0.4 M, 0.5 M, 0.6 M, 0.7 M, 0.8 M, 1 M, 1.5 M, or 2 M, depending on the desired conversion rate. In some embodiments, additional co-substrate can be added during the course of the reaction. In some embodiments, NADP+ can be used in the reaction. + Recycling system.

[0121] The substrate compound in the reaction mixture can be varied, taking into account, for example, the desired amount of product compound, the effect of substrate concentration on enzyme activity, the stability of the enzyme under the reaction conditions, and the percentage of substrate conversion to product. In some embodiments, suitable reaction conditions include a substrate compound loading of at least about 0.5 to about 200 g / L, 1 to about 200 g / L, 5 to about 150 g / L, about 10 to about 100 g / L, 20 to about 100 g / L, or about 50 to about 100 g / L. In some embodiments, suitable reaction conditions include a substrate compound loading of at least about 0.5 g / L, at least about 1 g / L, at least about 5 g / L, at least about 10 g / L, at least about 15 g / L, at least about 20 g / L, at least about 30 g / L, at least about 50 g / L, at least about 75 g / L, at least about 100 g / L, at least about 150 g / L, or at least about 200 g / L or even greater.

[0122] In performing the engineered P450-BM3 mediated methods described herein, the enzyme can be purified, partially purified, in the form of whole cells transformed with a gene encoding the enzyme, as a cell extract and / or lysate of such cells, and / or as an enzyme immobilized on a solid support, the engineered polypeptide is added to the reaction mixture. The whole cells transformed with a gene encoding the engineered P450-BM3 enzyme or cell extracts, lysates and isolated enzymes thereof can be employed in a variety of different forms, including solid (e.g., lyophilized, spray dried, etc.) or semi-solid (e.g., a crude paste). The cell extract or cell lysate can be partially purified by precipitation (ammonium sulfate, polyethyleneimine, heat treatment, etc.) followed by a desalting procedure (e.g., ultrafiltration, dialysis, etc.) and then lyophilized. Any enzyme preparation, including whole cell preparations, can be stabilized by cross-linking using known cross-linking agents such as, for example, glutaraldehyde or immobilization to a solid phase (e.g., Eupergit C, etc.).

[0123] The genes encoding the engineered P450-BM3 polypeptides can be transformed into host cells individually or together into the same host cell. For example, in some embodiments, one set of host cells can be transformed with a gene encoding one engineered P450-BM3 polypeptide, while another set of host cells can be transformed with a gene encoding another engineered P450-BM3 polypeptide. Both sets of transformed cells can be utilized together in the reaction mixture, either in the form of whole cells or in the form of lysates or extracts derived therefrom. In other embodiments, the host cells can be transformed with genes encoding multiple engineered P450-BM3 polypeptides. In some embodiments, the engineered polypeptides can be expressed in the form of secreted polypeptides, and the culture medium containing the secreted polypeptides can be used in the P450-BM3 reaction.

[0124] In some embodiments, the improved activity and / or selectivity of the engineered P450-BM3 polypeptides disclosed herein provides methods in which a higher percentage conversion can be achieved with a lower concentration of the engineered polypeptide. In some embodiments of the methods, suitable reaction conditions include an engineered polypeptide amount of about 0.03% (w / w), 0.05% (w / w), 0.1% (w / w), 0.15% (w / w), 0.2% (w / w), 0.3% (w / w), 0.4% (w / w), 0.5% (w / w), 1% (w / w), 2% (w / w), 5% (w / w), 10% (w / w), 20% (w / w) or more, depending on the substrate compound loading and conversion desired.

[0125] In some embodiments, the engineered polypeptide is present at about 0.01 g / L to about 40 g / L; about 0.05 g / L to about 15 g / L; about 0.1 g / L to about 10 g / L; about 1 g / L to about 8 g / L; about 0.5 g / L to about 10 g / L; about 1 g / L to about 10 g / L; about 0.1 g / L to about 5 g / L; about 0.5 g / L to about 5 g / L; or about 0.1 g / L to about 2 g / L. In some embodiments, the engineered P450-BM3 polypeptide is present at about 0.01 g / L, 0.05 g / L, 0.1 g / L, 0.2 g / L, 0.5 g / L, 1 g / L, 2 g / L, 5 g / L, 10 g / L, 15 g / L, 20 g / L, 40 g / L or more.

[0126] During the course of the reaction, the pH of the reaction mixture can change. The pH of the reaction mixture can be maintained at a desired pH or within a desired pH range. This can be done by adding an acid or a base prior to and / or during the course of the reaction. Alternatively, the pH can also be controlled by using a buffer. Thus, in some embodiments, the reaction conditions include a buffer. Suitable buffers for maintaining a desired pH range are known in the art and include, for example, but are not limited to, potassium phosphate, potassium borate, phosphate, 2-(N-morpholino)ethanesulfonic acid (MES), 3-(N-morpholino)propanesulfonic acid (MOPS), acetate, triethanolamine, and 2-amino-2-hydroxymethyl-propane-l,3-diol (Tris), and the like. In some embodiments, the buffer is potassium phosphate. In some embodiments of the method, suitable reaction conditions include a buffer concentration of about 0.01 to about 0.5 M, 0.05 to about 0.4 M, 0.1 to about 0.3 M, or about 0.1 to about 0.2 M. In some embodiments, the reaction conditions include a buffer concentration of about 0.01, 0.02, 0.03, 0.04, 0.05, 0.07, 0.1, 0.12, 0.14, 0.16, 0.18, 0.2, 0.3, 0.4 M, or 0.5 M.

[0127] In some embodiments, the reaction conditions include a solvent. Any suitable solvent can be used. In some embodiments, the solvent includes acetonitrile or DMSO. In some embodiments, the reaction conditions include a solvent concentration of 2%, 5%, 20%, 15%, 20%, 25%, or 30%.

[0128] In embodiments of the methods, reaction conditions can include a suitable pH. The desired pH or the desired pH range can be maintained by use of an acid or a base, an appropriate buffer, or a combination of a buffer with an acid or a base addition. The pH of the reaction mixture can be controlled prior to and / or during the course of the reaction. In some embodiments, suitable reaction conditions include a solution pH of about 4 to about 10, a pH of about 5 to about 10, a pH of about 5 to about 9, a pH of about 6 to about 9, a pH of about 6 to about 8. In some embodiments, reaction conditions include a solution pH of about 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10.

[0129] In embodiments of the methods herein, a suitable temperature can be used for reaction conditions, for example taking into account the increase in reaction rate at higher temperatures, as well as the activity of the enzyme during the reaction period. Thus, in some embodiments, suitable reaction conditions include a temperature of about 10 °C to about 60 °C, about 10 °C to about 55 °C, about 15 °C to about 60 °C, about 20 °C to about 60 °C, about 20 °C to about 55 °C, about 25 °C to about 55 °C, or about 30 °C to about 50 °C. In some embodiments, suitable reaction conditions include a temperature of about 10 °C, 15 °C, 20 °C, 25 °C, 30 °C, 35 °C, 40 °C, 45 °C, 50 °C, 55 °C, or 60 °C. In some embodiments, the temperature during the enzymatic reaction can be maintained at a particular temperature throughout the course of the reaction. In some embodiments, the temperature during the enzymatic reaction can be adjusted according to a temperature profile during the course of the reaction.

[0130] In some embodiments, reaction conditions can include a surfactant for stabilizing or enhancing the reaction. Surfactants can include non-ionic, cationic, anionic, and / or amphiphilic surfactants. Exemplary surfactants include, for example, but are not limited to, nonylphenoxypolyethoxyethanol (NP40), Triton X-100, polyoxyethylene-stearylamine, cetyltrimethylammonium bromide, oleoyl amide sodium sulfate, polyoxyethylene-sorbitan monostearate, hexadecyl dimethylamine, etc. Any surfactant that can stabilize or enhance the reaction can be employed. The concentration of the surfactant to be employed in the reaction can generally be 0.1 to 50 mg / mL, in particular 1 to 20 mg / mL.

[0131] In some embodiments, reaction conditions can include an antifoam agent, which helps to reduce or prevent the formation of foam in the reaction solution, such as when the reaction solution is mixed or sparged. Antifoam agents include nonpolar oils (e.g., mineral, silicone, etc.), polar oils (e.g., fatty acids, alkyl amines, alkyl amides, alkyl sulfates, etc.), and hydrophobic oils (e.g., treated silica, polypropylene, etc.), some of which also function as surfactants. Exemplary antifoam agents include Y- (Dow Corning), polyethylene glycol copolymers, oxy / ethoxylated alcohols, and polydimethylsiloxanes. In some embodiments, the antifoam agent can be present at about 0.001% (v / v) to about 5% (v / v), about 0.01% (v / v) to about 5% (v / v), about 0.1% (v / v) to about 5% (v / v), or about 0.1% (v / v) to about 2% (v / v). In some embodiments, the antifoam agent can be present at about 0.001% (v / v), about 0.01% (v / v), about 0.1% (v / v), about 0.5% (v / v), about 1% (v / v), about 2% (v / v), about 3% (v / v), about 4% (v / v), or about 5% (v / v) or more as needed to facilitate the reaction.

[0132] The amounts of reactants used in the monooxygenase reaction will generally vary depending on the amount of desired product as well as the amount of indanone substrate used in conjunction. One of ordinary skill in the art will readily understand how to vary these amounts to adapt them to the desired level of productivity and production scale.

[0133] In some embodiments, the order of addition of the reactants is not important. The reactants can be added together at the same time into the solvent (e.g., a single-phase solvent, a biphasic aqueous co-solvent system, etc.), or alternatively, some reactants can be added individually while others can be added together at different points in time. For example, the cofactor, co-substrate, engineered P450-BM3 enzyme, and substrate can be added to the solvent first.

[0134] Solid reactants (e.g., enzymes, salts, etc.) can be provided to the reaction in a variety of different forms, including powders (e.g., lyophilized, spray-dried, etc.), solutions, emulsions, suspensions, etc. The reactants can be readily lyophilized or spray-dried using methods and equipment known to one of ordinary skill in the art. For example, a protein solution can be frozen at -80 °C in small aliquots and then added to a pre-chilled lyophilization chamber, followed by the application of a vacuum.

[0135] To improve mixing efficiency when using an aqueous co-solvent system, the engineered P450-BM3 polypeptide and cofactor can be added and mixed into the aqueous phase first. The organic phase can then be added and mixed, followed by the addition of the substrate and co-substrate. Alternatively, the substrate can be pre-mixed in the organic phase, which is then added to the aqueous phase.

[0136] The mono-oxygenation process is generally allowed to proceed until further conversion of substrate to product does not change significantly with reaction time (e.g., less than 10% of the substrate is converted, or less than 5% of the substrate is converted). In some embodiments, the reaction is allowed to proceed until the substrate is completely or nearly completely converted to product. Conversion of substrate to product can be monitored using known methods by detecting the substrate and / or product, with or without derivatization. Suitable analytical methods include gas chromatography, HPLC, MS, and the like.

[0137] In further embodiments of methods for converting a substrate compound to a product compound using an engineered P450-BM3 polypeptide, suitable reaction conditions can include loading an initial substrate into a reaction solution, which is then contacted with the polypeptide. The reaction solution is then further supplemented with additional substrate compound over time in a continuous or batch addition at a rate of at least about 1 g / L / h, at least about 2 g / L / h, at least about 4 g / L / h, at least about 6 g / L / h, or higher. Thus, according to these suitable reaction conditions, the polypeptide is added to a solution having an initial substrate loading of at least about 20 g / L, 30 g / L, or 40 g / L. Then, after this polypeptide addition, additional substrate is continuously added to the solution at a rate of about 2 g / L / h, 4 g / L / h, or 6 g / L / h until a much higher final substrate loading of at least about 30 g / L, 40 g / L, 50 g / L, 60 g / L, 70 g / L, 100 g / L, 150 g / L, 200 g / L, or higher is achieved. Thus, in some embodiments of the methods, suitable reaction conditions include adding the polypeptide to a solution having an initial substrate loading of at least about 20 g / L, 30 g / L, or 40 g / L, and then adding additional substrate to the solution at a rate of about 2 g / L / h, 4 g / L / h, or 6 g / L / h until a final substrate loading of at least about 30 g / L, 40 g / L, 50 g / L, 60 g / L, 70 g / L, 100 g / L, or more is achieved. Such substrate-supplementing reaction conditions allow for higher substrate loadings to be achieved while maintaining high substrate-to-product conversion of at least about 50%, 60%, 70%, 80%, 90%, or higher.

[0138] In some embodiments, NADPH is recycled to NADP+ using a recycling system. In some embodiments, the recycling system comprises glucose dehydrogenase and glucose. In some embodiments, the recycling system comprises phosphite dehydrogenase and phosphite.

[0139] In some embodiments of the methods, the reactions using engineered P450-BM3 polypeptides include the following suitable reaction conditions: (a) substrate loading of about 1-25 g / L; (b) engineered polypeptide of about 1-40 g / L; (c) 0.25-1 g / L NADP+; (d) pH of about 7-9; (e) 5%-20% solvent (acetonitrile or DMSO) 0% (we did not use co-solvent); (h) temperature of about 20°C-30°C; and (i) reaction time of about 20 hours.

[0140] In some embodiments of the methods, the reactions using engineered P450-BM3 polypeptides include the following suitable reaction conditions: (a) substrate loading of about 10 g / L; (b) engineered polypeptide of about 1.0-5.0 g / L; (c) 1 g / L NADP+; (d) pH of about 8; (e) 0% DMSO; (f) temperature of about 30°C; and (g) reaction time of about 20 hours.

[0141] In some embodiments, additional reaction components or additional techniques are performed to supplement the reaction conditions. These can include taking measures to stabilize the enzyme or prevent enzyme deactivation, reduce product inhibition, shift the reaction equilibrium toward product formation.

[0142] In further embodiments, any of the above-described methods for converting a substrate compound to a product compound can further include one or more steps selected from the group consisting of: extraction; separation; purification; and crystallization of the product compound. Methods, techniques, and protocols for extracting, separating, purifying, and / or crystallizing products from a biocatalytic reaction mixture produced by the above- described disclosed methods are known to the ordinary skilled artisan and / or are obtainable through routine experimentation. Moreover, illustrative methods are provided in the Examples below.

[0143] Various features and embodiments of the present invention are illustrated in the following representative examples, which are intended to be illustrative and not limiting.

[0144] It is further contemplated, in accordance with the guidance provided herein, that any of the exemplary engineered polypeptides can be used as a starting amino acid sequence for the synthesis of additional engineered P450-BM3 polypeptides, for example, by subsequent rounds of evolution, adding new combinations of various amino acid differences from other polypeptides and other residue positions described herein. Further improvements can be made by including amino acid differences at residue positions that remain constant throughout the early rounds of evolution.

[0145] Polynucleotides, expression vectors, and host cells encoding engineered polypeptides:

[0146] The present application provides polynucleotides encoding the engineered P450-BM3 polypeptides described herein. In some embodiments, the polynucleotides are operably linked to one or more heterologous regulatory sequences that control gene expression to produce a recombinant polynucleotide capable of expressing the polypeptide. Expression constructs containing the heterologous polynucleotides encoding the engineered P450-BM3 polypeptides can be introduced into appropriate host cells to express the corresponding P450-BM3 polypeptides.

[0147] As will be apparent to those of ordinary skill in the art, the availability of protein sequences and knowledge of the codons corresponding to various amino acids provides a description of all polynucleotides capable of encoding the subject polypeptides. The degeneracy of the genetic code, in which the same amino acid is encoded by alternative or synonymous codons, allows for the production of a very large number of nucleic acids, all of which encode the engineered P450-BM3 polypeptides. Thus, given knowledge of a particular amino acid sequence, one of ordinary skill in the art can produce many different nucleic acids by modifying the sequence of one or more codons in a manner that does not change the amino acid sequence of the protein. In this regard, the present application specifically contemplates each and every possible variation of polynucleotides that could encode a polypeptide described herein that is produced by selection of combinations based on possible codon choices, and all such variations are specifically disclosed herein, including the variants provided in Tables 2.1, 3.1, 4.1, 5.1, 6.1, 7.1, 8.1, 9.1, 10.1, 10.2, 11.1, 11.2, 12.1, 13.1, 14.1, 14.2, 15.1, 16.1, 17.1, 17.2, 18.1, 19.1, and 19.2, as well as SEQ ID NOs: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, and 1266.

[0148] In various embodiments, codons are preferably selected to suit the host cell in which the protein is produced. For example, preferred codons used in bacteria are used for expression in bacteria. Thus, codon-optimized polynucleotides encoding the engineered P450-BM3 polypeptides contain preferred codons at about 40%, 50%, 60%, 70%, 80%, or greater than 90% of the codon positions in the full-length coding region.

[0149] In some embodiments, as described above, the polynucleotide encodes an engineered polypeptide having P450-BM3 activity with the properties disclosed herein, wherein the polypeptide comprises an amino acid sequence that has at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity to the amino acid sequence of a reference sequence (e.g., SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266) or any of the variants disclosed in Tables 2.1, 3.1, 4.1, 5.1, 6.1, 7.1, 8.1, 9.1, 10.1, 10.2, 11.1, 11.2, 12.1, 13.1, 14.1, 14.2, 15.1, 16.1, 17.1, 17.2, 18.1, 19.1, and 19.2 and has one or more residue differences (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more amino acid residue positions) compared to the reference polypeptide of SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266 or the amino acid sequence of any of the variants disclosed in Tables 2.1, 3.1, 4.1, 5.1, 6.1, 7.1, 8.1, 9.1, 10.1, 10.2, 11.1, 11.2, 12.1, 13.1, 14.1, 14.2, 15.1, 16.1, 17.1, 17.2, 18.1, 19.1, and 19.2. In some embodiments, the reference sequence is selected from SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266.

[0150] In some embodiments, the polynucleotide is capable of hybridizing under high stringency conditions to a reference polynucleotide sequence selected from the group consisting of SEQ ID NO: 3, 35, 65, 71, 197, 225, 243, 285, 357, 409, 533, 733, 747, 827, 967, 983, 1159, or 1265, or the complement thereof, or a polynucleotide sequence encoding any of the variant P450-BM3 polypeptides provided herein. In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes a P450-BM3 polypeptide comprising an amino acid sequence having one or more residue differences as compared to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266.

[0151] In some embodiments, the polynucleotide is capable of hybridizing under high stringency conditions to a reference polynucleotide sequence selected from any of the polynucleotide sequences provided herein, or the complement thereof, or a polynucleotide sequence encoding any of the variant enzyme polypeptides provided herein. In some embodiments, the polynucleotide capable of hybridizing under high stringency conditions encodes an enzyme polypeptide comprising an amino acid sequence having one or more residue differences as compared to the reference sequence.

[0152] In some embodiments, isolated polynucleotides encoding any of the engineered enzyme polypeptides herein are manipulated in a variety of ways to facilitate expression of the enzyme polypeptides. In some embodiments, the polynucleotides encoding the enzyme polypeptides comprise expression vectors in which one or more control sequences are present to regulate expression of the enzyme polynucleotides and / or polypeptides. Manipulation of the isolated polynucleotides prior to their insertion into a vector can be desirable or necessary, depending on the expression vector used. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art. In some embodiments, control sequences include, among others, promoters, leader sequences, polyadenylation sequences, propeptide sequences, signal peptide sequences, and transcription terminators. In some embodiments, suitable promoters are selected based on host cell selection. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include, but are not limited to, promoters obtained from the lac operon of E. coli, the agarase gene of Streptomyces coelicolor (dagA), the levansucrase gene of Bacillus subtilis (sacB), the Bacillus licheniformis alpha-amylase gene (amyL), a Bacillus stearothermophilus maltogenic amylase gene (amyM), a Bacillus amyloliquefaciens alpha-amylase gene (amyQ), the Bacillus licheniformis penicillinase gene (penP), the Bacillus subtilis xylA and xylB genes, and a prokaryotic beta-lactamase gene (see e.g., Villa-Kamaroff et al., Proc. Natl Acad. Sci. USA 75: 3727-3731

[1978] ), as well as the tac promoter (see e.g., DeBoer et al., Proc. Natl Acad. Sci. USA 80: 21-25

[1983] ).Exemplary promoters for filamentous fungal host cells include, but are not limited to, those derived from Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral alpha-amylase, Aspergillus niger acid stable alpha-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (see, e.g., WO 96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoters from Aspergillus niger neutral alpha-amylase and Aspergillus oryzae triose phosphate isomerase), and mutant, truncated, and hybrid promoters thereof. Exemplary yeast cell promoters can be derived from Saccharomyces cerevisiae enolase (ENO-1), glucokinase (GAL1), ethanol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and triose phosphate isomerase of Saccharomyces cerevisiae. Other useful promoters for yeast host cells are known to those skilled in the art (see, e.g., Romanos et al., Yeast 8:423-488

[1992] ).

[0153] In some embodiments, the control sequence is also a transcription terminator sequence (i.e., a sequence recognized by a host cell to terminate transcription). In some embodiments, the terminator sequence is operably linked to the 3' terminus of the nucleic acid sequence encoding the enzyme polypeptide. Any suitable terminator for the selected host cell can be used in the present application. Exemplary transcription terminators for filamentous fungal host cells can be obtained from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger alpha- glucosidase, and Fusarium oxysporum trypsin-like protease. Exemplary terminators for yeast host cells can be obtained from the genes for Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are known to those skilled in the art (see, e.g., Romanos et al., supra).

[0154] In some embodiments, the control sequence is also a suitable leader sequence (i.e., an untranslated region of an mRNA important for translation by host cells) that is operably linked to the 5' terminus of the nucleic acid sequence encoding the enzyme polypeptide. In some embodiments, the leader sequence is operably linked to the 5' terminus of the nucleic acid sequence encoding the enzyme polypeptide. Any suitable leader sequence that is functional in the chosen host cell can be used in the present application. Exemplary leader sequences for filamentous fungal host cells are obtained from the genes for Aspergillus niger TAKA amylase, and Aspergillus nidulans acetamidase. Suitable leader sequences for yeast host cells are obtained from the genes for Saccharomyces cerevisiae enolase (ENO-l), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae alpha-factor, and Saccharomyces

[0155] In some embodiments, the control sequence is also a polyadenylation sequence (i.e., a sequence operably linked to the 3' terminus of the nucleic acid sequence and which, when transcribed, is recognized by the host cell as a signal to add a polyadenosine residue to the 3' end of the transcript). Any suitable polyadenylation sequence that is functional in the chosen host cell can be used in the present application. Exemplary polyadenylation sequences for filamentous fungal host cells include, but are not limited to, the A. niger TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger alpha-glucosidase genes. Useful polyadenylation sequences for yeast host cells are known (see, e.g., Guo and Sherman, Mol. Cell. Bio., 15:5983-5990

[1995] ).

[0156] In some embodiments, the control sequence is also a signal peptide (i.e., a coding region that encodes an amino acid sequence that is attached to the amino terminus of a polypeptide and directs the encoded polypeptide into the cellular secretory pathway). In some embodiments, the 5' end of a coding sequence of a nucleic acid sequence inherently contains a signal peptide coding region naturally linked in translation reading frame to the segment of the coding region that encodes the secreted polypeptide. Alternatively, in some embodiments, the 5' end of the coding sequence contains a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the secretory pathway of a selected host cell can be used to express an engineered polypeptide. Effective signal peptide coding regions for bacterial host cells are those that include, but are not limited to, those obtained from the genes for Bacillus NCIB 11837 maltogenic amylase, Bacillus stearothermophilus alpha-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus amyloliquefaciens neutral proteases (nprT, nprS, nprM), and Bacillus subtilis prsA. Other signal peptides are known in the art (see, e.g., Simonen and Palva, Microbiol. Rev., 57: 109-137

[1993] ). In some embodiments, effective signal peptide coding regions for filamentous fungal host cells include, but are not limited to, those obtained from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase. Useful signal peptides for yeast host cells include, but are not limited to, those from the genes for Saccharomyces cerevisiae alpha factor and Saccharomyces cerevisiae invertase.

[0157] In some embodiments, the control sequence is also a pre-peptide coding region that encodes an amino acid sequence located at the amino terminus of a polypeptide. The resulting polypeptide is referred to as a “preenzyme,” “prepolypeptide,” or “zymogen.” A prepolypeptide can be converted to a mature, active polypeptide by catalytic or autocatalytic cleavage of the pre-peptide from the prepolypeptide. Pre-peptide coding regions can be obtained from any suitable source, including but not limited to the genes for Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Saccharomyces cerevisiae alpha-factor, Rhizomucor miehei aspartic proteinase, and Myceliophthora thermophila lactase (see, e.g., WO 95 / 33836). When both a signal peptide and a pre-peptide region are present at the amino terminus of a polypeptide, the pre-peptide region is located near the amino terminus of the polypeptide, and the signal peptide region is located near the amino terminus of the pre-peptide region.

[0158] In some embodiments, regulatory sequences are also utilized. These sequences are useful in controlling expression of the polypeptide relative to the growth of the host cell. Examples of regulatory systems are those that allow for induction or repression of expression of a gene in response to a chemical or physical stimulus, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory systems include, but are not limited to, the ADH2 system or GAL1 system. In filamentous fungi, suitable regulatory sequences include, but are not limited to, the TAKA alpha-amylase promoter, Aspergillus niger glucoamylase promoter, and Aspergillus oryzae glucoamylase promoter.

[0159] In another aspect, the present application relates to a recombinant expression vector comprising a polynucleotide encoding an engineered enzyme polypeptide and one or more expression control regions, such as promoters and terminators, replication origins, and the like, depending on the type of host into which they are to be introduced. In some embodiments, the various nucleic acids and control sequences described herein are ligated together to produce a recombinant expression vector that comprises one or more convenient restriction sites for the insertion or substitution of a nucleic acid sequence encoding an enzyme polypeptide at such sites. Alternatively, in some embodiments, the nucleic acid sequences of the present application are expressed by inserting the nucleic acid sequence or a nucleic acid construct comprising the sequence into an appropriate expression vector. In some embodiments involving the production of an expression vector, the coding sequence is located in the vector such that the coding sequence is operably linked with the appropriate control sequences for expression.

[0160] The recombinant expression vector can be any suitable vector (e.g., a plasmid or virus) that can conveniently be subjected to recombinant DNA procedures and that brings about the expression of the enzyme polynucleotide sequence. The choice of vector depends primarily on the host cell into which the vector is to be introduced. The vector can be linear or closed circular plasmid.

[0161] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, e.g., a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome). The vector can contain any means for assuring self-replication. In some alternative embodiments, the vector is one which, when introduced into a host cell, is integrated into the genome and replicated together with the chromosome into which it has been integrated. Furthermore, in some embodiments, a single vector or plasmid, or two or more vectors or plasmids, together containing the total DNA to be introduced into the genome of the host cell, and / or a transposon, is used.

[0162] In some embodiments, the expression vector contains one or more selectable markers that permit easy selection of transformed cells. A "selectable marker" is a gene the product of which provides for biocide or viral resistance, resistance to heavy metals, prototrophy to auxotrophs, and the like. Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis, or markers that confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, or tetracycline resistance. Suitable markers for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for use in a filamentous fungal host cell include, but are not limited to, amdS (acetamidase; e.g., from Aspergillus nidulans or Aspergillus niger), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase; e.g., from Streptomyces spec.), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase; e.g., from Aspergillus nidulans or Aspergillus niger), sC (sulfate adenyltransferase), and trpC (anthranilate synthase; e.g., from Aspergillus nidulans), and equivalents thereof.

[0163] In another aspect, the present application provides a host cell comprising a recombinant polynucleotide encoding at least one engineered enzyme polypeptide of the present application operably linked to one or more control sequences for expression of the engineered enzyme in the host cell. Host cells suitable for expression of polypeptides encoded by expression vectors of the present application are well known in the art, and include, but are not limited to, bacterial cells, such as E. coli, Vibrio fluvialis, Streptomyces, and Salmonella typhimurium cells; fungal cells, such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC Accession No. 201178)); insect cells such as Drosophila S2 and Spodoptera Sf9 cells; animal cells such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Exemplary host cells also include various E. coli strains (e.g., W3110 (AfhuA) and BL21). Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis, or markers that confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, and / or tetracycline resistance.

[0164] In some embodiments, the expression vectors of the application contain one or more elements that permit the vector to be integrated into the genome of the host cell or to replicate autonomously in the cell independent of the genome. In some embodiments involving integration into the host cell genome, the vector is dependent on the nucleic acid sequence encoding the polypeptide or any other element of the vector for integration into the genome by homologous or non-homologous recombination.

[0165] In some alternative embodiments, the expression vector contains additional nucleic acid sequences for directing integration into the genome of the host cell by homologous recombination. The additional nucleic acid sequences enable the vector to integrate into the host cell genome at one or more precise locations within one or more chromosomes. To increase the likelihood of integration at the precise location, the integration element preferably contains a sufficient number of nucleotides, such as 100 to 10,000 base pairs, preferably 400 to 10,000 base pairs, and most preferably 800 to 10,000 base pairs, that are highly homologous with the corresponding target sequence to enhance the probability of homologous recombination. The integration element can be any sequence homologous to a target sequence in the genome of the host cell. Moreover, the integration element can also be a non-coding or coding nucleic acid sequence. On the other hand, the vector can integrate into the genome of the host cell by non-homologous recombination.

[0166] For autonomous replication, the vector can further comprise an origin of replication enabling the vector to replicate autonomously in the host cell in question. Examples of bacterial origins of replication are the P15A ori or the origins of replication of plasmids pBR322, pUC19, pACYC177 (which plasmid has the P15A ori), or pACYC184, which allow replication in E. coli, and the origins of replication of plasmids pUBHO, pE194, or pTA1060, which allow replication in Bacillus sp. Examples of origins of replication for use in yeast host cells are the 2 micron origin of replication, ARS1, ARS4, a combination of ARS1 and CEN3, and a combination of ARS4 and CEN6. The origin of replication can be an origin of replication with a mutation that renders the function of the origin of replication in the host cell temperature sensitive (see, e.g., Ehrlich, Proc. Natl. Acad. Sci. USA 75: 1433

[1978] ).

[0167] In some embodiments, more than one copy of a nucleic acid sequence of the application is inserted into a host cell to increase production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome, or by including a selectable marker gene with the nucleic acid sequence that can be amplified together with the nucleic acid sequence, wherein cells containing amplified copies of the selectable marker gene can be selected for by culturing the cells in the presence of a suitable selectable agent, and thereby selecting for cells that contain additional copies of the nucleic acid sequence.

[0168] Many of the expression vectors used in the application are commercially available. Suitable commercial expression vectors include, but are not limited to, p3xFLAG™ TM pCEP4 (Invitrogen), or pPoly (see, e.g., Lathe et al., Gene 57:193-201

[1987] ). Other suitable expression vectors include, but are not limited to, pBluescript II SK(-) and pBK-CMV (Stratagene), and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen), or pPoly (see, e.g., Lathe et al., Gene 57:193-201

[1987] ).

[0169] Accordingly, in some embodiments, a vector comprising a sequence encoding at least one variant engineered P450-BM3 polypeptide is transformed into a host cell to allow for propagation of the vector and expression of the variant engineered P450-BM3 polypeptide. In some embodiments, the variant engineered P450-BM3 polypeptide is post-translationally modified to remove the signal peptide, and in some cases can be cleaved after secretion. In some embodiments, the above-described transformed host cell is cultured in a suitable nutrient medium under conditions that allow the variant engineered P450-BM3 polypeptide to be expressed. Any suitable medium available for the culturing of host cells can be used in the application, including, but not limited to, minimal or complex media containing appropriate supplements. In some embodiments, the host cell is grown in an HTP medium. Suitable media can be obtained from various commercial suppliers or can be prepared according to published recipes (e.g., in the catalog of the American Type Culture Collection).

[0170] In another aspect, the present application provides host cells comprising a polynucleotide encoding an improved engineered P450-BM3 polypeptide provided herein operably linked to one or more control sequences for expression of the engineered P450-BM3 polypeptide in the host cell. Host cells for expressing engineered P450-BM3 polypeptides encoded by the expression vectors of the present application are known in the art and include, but are not limited to, bacterial cells, such as Escherichia coli, Bacillus megaterium, Lactobacillus kefiri, Streptomyces, and Salmonella typhimurium cells; fungal cells, such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC Accession No. 201178)); insect cells, such as Drosophila S2 and Spodoptera Sf9 cells; animal cells, such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Appropriate media and growth conditions for the above host cells are well known in the art.

[0171] Polynucleotides for expressing engineered P450-BM3 polypeptides can be introduced into cells by a variety of methods known in the art. Techniques include, but are not limited to, electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion. Various methods for introducing polynucleotides into cells are known to those of skill in the art.

[0172] In some embodiments, the host cell is a eukaryotic cell. Suitable eukaryotic cells include, but are not limited to, fungal cells, algal cells, insect cells, and plant cells. Suitable fungal host cells include, but are not limited to, Ascomycota, Basidiomycota, Deuteromycota, Zygomycota, Fungi Imperfecti. In some embodiments, the fungal host cell is a yeast cell and a filamentous fungal cell. Filamentous fungal host cells of the present application include all filamentous forms of the subdivisions Eumycotina and Oomycota. Filamentous fungi are characterized by a mycelial wall composed of chitin, cellulose, and other complex polysaccharides. Filamentous fungal host cells of the present application are different from yeast in morphology and have a mycelial cell wall composed of chitin and cellulose, as opposed to the cell wall of yeast cells which is mainly composed of polysaccharides.

[0173] In some embodiments of the application, the filamentous fungal host cell belongs to any suitable genus and species, including but not limited to Achlya, Acremonium, Aspergillus, Aureobasidium, Bjerkandera, Ceriporiopsis, Cephalosporium, Chrysosporium, Cochliobolus, Corynascus, Cryphonectria, Cryptococcus, Coprinus, Coriolus, Diplodia, Endothia, Fusarium, Gibberella, Gliocladium, Humicola, Hypocrea, Myceliophthora, Mucor, Neurospora, Penicillium, Podospora, Phlebia, Piromyces, Pyricularia, Rhizomucor, Rhizopus, Schizophyllum, Scytalidium, Sporotrichum, Talaromyces, Thermoascus, Thielavia, Trametes, Tolypocladium, Trichoderma, Verticillium, and / or Volvariella and / or teleomorph or anamorph and synonyms or taxonomic equivalents thereof.

[0174] In some embodiments of the application, the host cell is a yeast cell, including but not limited to a cell of the Candida, Hansenula, Saccharomyces, Schizosaccharomyces, Pichia, Kluyveromyces, or Yarrowia species. In some embodiments of the application, the yeast cell is a Hansenula polymorpha, Saccharomyces cerevisiae, Saccharomyces carlsbergensis, Saccharomyces diastaticus, Saccharomyces norbensis, Saccharomyces kluyveri, Schizosaccharomyces pombe, Pichia pastoris, Pichia finlandica, Pichia trehalophila, Pichia kodamae, Pichia membranaefaciens, Pichia opuntiae, Pichia thermotolerans, Pichia salictaria, Pichia quercuum, Pichia pijperi, Pichia stipitis, Pichia methanolica, Pichia angusta, Kluyveromyces lactis, Candida albicans, or Yarrowia lipolytica.

[0175] In some embodiments of the application, the host cell is an algal cell, such as a Chlamydomonas (e.g., C. reinhardtii) and Phormidium (P. sp. ATCC 29409).

[0176] In some other embodiments, the cell is a prokaryotic cell. Suitable prokaryotic cells include, but are not limited to, Gram-positive bacterial cells, Gram-negative bacterial cells, and Gram-stain indefinite bacterial cells. Any suitable bacterial organism can be used in the present application, including, but not limited to, Agrobacterium, Alicyclobacillus, Anabaena, Anacystis, Acinetobacter, Acidothermus, Arthrobacter, Azotobacter, Bacillus, Bifidobacterium, Brevibacterium, Butyrivibrio, Buchnera, Brassica, Campylobacter, Clostridium, Corynebacterium, Chromatium, Coprococcus, Escherichia, Enterococcus, Enterobacter, Erwinia, Fusobacterium, Faecalibacterium, Francisella, Flavobacterium, Geobacillus, Haemophilus, Helicobacter, Klebsiella, Lactobacillus, Lactococcus, Lichen, Micrococcus, Microbacterium, Mesorhizobium, Methylobacterium, Mycobacterium, Neisseria, Pantoea, Pseudomonas, Prochlorococcus, Rhodobacter, Rhodopseudomonas, Rhodopseudomonas, Ralstonia, Rhodospirillum, Rhodococcus, Scenedesmus, Streptomyces, Streptococcus, Sarcina, Saccharomonospora, Staphylococcus, Serratia, Salmonella, Shigella, Thermotoga, Tropheryma, Francisella tularensis, Tenacibaculum, Thermococcus, Thermus, Ureaplasma, Xanthomonas, Xylella, Yersinia, and Zymomonas. In some embodiments, the host cell is a species of Agrobacterium, Acinetobacter, Azotobacter, Bacillus, Bifidobacterium, Buchnera, Geobacillus, Campylobacter, Clostridium, Corynebacterium, Escherichia, Enterococcus, Erwinia, Flavobacterium, Lactobacillus, Lactococcus, Pantoea, Pseudomonas, Staphylococcus, Salmonella, Streptococcus, Streptomyces, or Zymomonas. In some embodiments, the bacterial host strain is non-pathogenic to humans. In some embodiments, the bacterial host strain is an industrial strain. A variety of bacterial industrial strains are known and suitable for use in the present application. In some embodiments of the present application, the bacterial host cell is an Agrobacterium species (e.g., A. radiobacter, A. rhizogenes, and A. rubi).In some embodiments of the application, the bacterial host cell is an Arthrobacter species (e.g., A. aurescens, A. citreus, A. globiformis, A. hydrocarboglutamicus, A. mysorens, A. nicotianae, A. paraffineus, A. protophonnlae, A. roseoparaffinus, A. sulfureus, and A. ureafaciens). In some embodiments of the application, the bacterial host cell is a Bacillus species (e.g., B. thuringensis, B. anthracis, B. megaterium, B. subtilis, B. lentus, B. circulans, B. pumilus, B. lautus, B. coagulans, B. brevis, B. firmus, B. alkaophius, B. licheniformis, B. clausii, B. stearothermophilus, B. halodurans, and B. amyloliquefaciens). In some embodiments, the host cell is an industrial Bacillus strain, including but not limited to B. subtilis, B. pumilus, B. licheniformis, B. megaterium, B. clausii, B. stearothermophilus, or B. amyloliquefaciens. In some embodiments, the Bacillus host cell is B. subtilis, B. licheniformis, B. megaterium, B. stearothermophilus, and / or B. amyloliquefaciens. In some embodiments, the bacterial host cell is a Clostridium species (e.g., C. acetobutylicum, C. tetani E88, C. lituseburense, C. saccharobutylicum, C. perfringens, and C. beijerinckii).In some embodiments, the bacterial host cell is a Corynebacterium species (e.g., C. glutamicum and C. acetoacidophilum). In some embodiments, the bacterial host cell is an Escherichia species (e.g., E. coli). In some embodiments, the host cell is E. coli W3110. In some embodiments, the bacterial host cell is an Erwinia species (e.g., E. uredovora, E. carotovora, E. ananas, E. herbicola, E. punctata, and E. terreus). In some embodiments, the bacterial host cell is a Pantoea species (e.g., P. citrea and P. agglomerans). In some embodiments, the bacterial host cell is a Pseudomonas species (e.g., P. putida, P. aeruginosa, P. mevalonii, and P. sp. D-01 10). In some embodiments, the bacterial host cell is a Streptococcus species (e.g., S. equisimiles, S. pyogenes, and S. uberis). In some embodiments, the bacterial host cell is a Streptomyces species (e.g., S. ambofaciens, S. achromogenes, S. avermitilis, S. coelicolor, S. aureofaciens, S. aureus, S. fungicidicus, S. griseus, and S. lividans). In some embodiments, the bacterial host cell is a Zymomonas species (e.g., Z. mobilis and Z. lipolytica).

[0177] Many prokaryotic and eukaryotic strains useful in the present application are readily available to the public from many culture collections, such as the American Type Culture Collection (ATCC), the German Collection of Microorganisms and Cell Cultures (DSM), the Centraalbureau voor Schimmelcultures (CBS), and the Agricultural Research Service Culture Collection Northern Regional Research Center (NRRL).

[0178] In some embodiments, host cells are genetically modified to have the feature of improving protein secretion, protein stability and / or protein expression and / or secretion required other characteristics.Genetic modification can be achieved by genetic engineering technology and / or classical microbiological technology (for example, chemical or UV mutagenesis and subsequent selection).In fact, in some embodiments, the combination of recombinant modification and classical selection technology is used to produce host cells.Use recombinant technology, the mode that the output of engineered P450-BM3 variant in host cell and / or culture medium increases can be introduced, lacked, suppressed or modified nucleic acid molecule.For example, the knockout of Alp1 function causes the cell of protease defect, and the knockout of pyr5 function causes the cell with pyrimidine defect phenotype.In a kind of genetic engineering method, homologous recombination is used to induce targeted gene modification by specific targeting gene in vivo, to suppress the expression of encoded protein.In alternative method, siRNA, antisense and / or ribozyme technology can be used for suppressing gene expression.The various methods known in the art for reducing the protein expression in cell, include but not limited to all or part of the gene of deletion coded protein and site-specific mutagenesis to destroy the expression or activity of gene product. (See, e.g., Chaveroche et al., Nucl. Acids Res., 28:22e97

[2000] ; Cho et al., Molec. Plant Microbe Interact., 19:7-15

[2006] ; Maru yama and Kitamoto, Biotechnol Lett., 30:1811-1817

[2008] ; Takahashi et al., Mol. Gen. Genom., 272:344-352

[2004] ; and You et al., Arch. Microbiol., 191:615-622

[2009] , all of which are incorporated herein by reference). Random mutagenesis followed by screening for desired mutations may also be used (See, eg, Combier et al., FEM S Microbiol. Lett., 220: 141-8

[2003] ; and Firon et al., Eukary. Cell 2: 247-55

[2003] , both incorporated by reference).

[0179] The vector or DNA construct can be introduced into the host cell using any suitable method known in the art, including but not limited to calcium phosphate transfection, DEAE-dextran mediated transfection, PEG mediated transformation, electroporation or other common techniques known in the art. In some embodiments, the E. coli expression vector pCK100900i (see U.S. Patent No. 9,714,437, which is hereby incorporated by reference) can be used.

[0180] In some embodiments, the engineered host cells of the application (i.e., "recombinant host cells") are cultured in conventional nutrient media, modified as appropriate for activating promoters, selecting transformants, or amplifying engineered P450-BM3 polynucleotides. Culture conditions, such as temperature, pH, etc., are those previously used with the host cell selected for expression, and are well known to those skilled in the art. As described above, a number of standard references and texts are available for culturing and producing cells of various origins, including bacterial, plant, animal (especially mammalian), and archaeal.

[0181] In some embodiments, cells expressing a variant engineered P450-BM3 polypeptide of the application are grown under batch or continuous fermentation conditions. Classical "batch fermentation" is a closed system in which the composition of the culture medium is set at the beginning of the fermentation and is not artificially changed during the fermentation. A variation of the batch system is "fed-batch fermentation," which can also be used in the present application. In this variation, substrate is added in increments as the fermentation proceeds. Fed-batch systems are useful when catabolite repression can inhibit the cells' metabolism and when it is desirable to have a limited amount of substrate in the medium. Batch and fed-batch fermentations are common and well known in the art. "Continuous fermentation" is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is simultaneously removed for processing. Continuous fermentation typically maintains the culture at a constant high density, with cells primarily in the log phase of growth. Continuous fermentation systems strive to maintain steady-state growth conditions. Methods for regulating nutrients and growth factors for continuous fermentation processes, as well as techniques for maximizing product formation rates, are well known in the field of industrial microbiology.

[0182] In some embodiments of the application, cell-free transcription / translation systems can be used to produce variant engineered P450-BM3 polypeptides. Several systems are commercially available, and methods are well known to those skilled in the art.

[0183] The present invention provides methods for preparing variant engineered P450-BM3 polypeptides or biologically active fragments thereof. In some embodiments, the method comprises: providing a host cell transformed with a polynucleotide encoding an amino acid sequence comprising at least about 70% (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99%) sequence identity to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160 or 1266 and comprising at least one mutation as provided herein; culturing the transformed host cell in a culture medium under conditions such that the host cell expresses the encoded variant engineered P450-BM3 polypeptide; and optionally recovering or isolating the expressed variant engineered P450-BM3 polypeptide, and / or recovering or isolating the culture medium containing the expressed variant engineered P450-BM3 polypeptide. In some embodiments, the method further provides for optionally lysing the transformed host cells after expressing the encoded engineered P450-BM3 polypeptide, and optionally recovering and / or isolating the expressed variant engineered P450-BM3 polypeptide from the cell lysate. The present invention further provides a method for preparing a variant engineered P450-BM3 polypeptide, comprising culturing host cells transformed with a variant engineered P450-BM3 polypeptide under conditions suitable for producing the variant engineered P450-BM3 polypeptide and recovering the engineered P450-BM3 polypeptide. Typically, the engineered P450-BM3 polypeptide is recovered or isolated from the host cell culture medium, the host cells, or both using protein recovery techniques well known in the art, including techniques described herein. In some embodiments, the host cells are harvested by centrifugation, destroyed by physical or chemical means, and the resulting crude extract is retained for further purification. Microbial cells used to express proteins can be destroyed by any convenient method, including but not limited to freeze-thaw cycles, ultrasonic treatment, mechanical destruction, and / or the use of cell lysis agents, as well as many other suitable methods well known to those skilled in the art.

[0184] The engineered P450-BM3 polypeptide expressed in the host cells can be recovered from the cells and / or culture medium using any one or more of the techniques known in the art for protein purification, including, among others, lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography. Suitable solutions for lysing and efficiently extracting proteins from bacteria such as E. coli are available under the trade name CelLytic B TMSigma-Aldrich). Thus, in some embodiments, the resulting polypeptides are recovered / isolated and optionally purified by any of a variety of methods known in the art. For example, in some embodiments, the polypeptides are isolated from the nutrient medium by conventional procedures including, but not limited to, centrifugation, filtration, extraction, spray-drying, evaporation, chromatography (e.g., ion-exchange, affinity, hydrophobic interaction, chromatofocusing, and size exclusion), or precipitation. In some embodiments, protein refolding steps are employed in finishing the mature protein's configuration, if necessary. Moreover, in some embodiments, high performance liquid chromatography (HPLC) is employed in the final purification step. For example, in some embodiments, methods known in the art can be used in the present application (see, e.g., Parry et al., Biochem. J., 353:117

[2001] ; and Hong et al., Appl. Microbiol. Biotechnol., 73:1331

[2007] , both of which are incorporated herein by reference). Indeed, any suitable purification method known in the art can be used in the present application.

[0185] Chromatographic techniques for isolating engineered P450-BM3 polypeptides include, but are not limited to, reverse phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. Conditions for purifying a particular enzyme will depend in part on factors known to those of skill in the art, such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, and the like.

[0186] In some embodiments, affinity techniques can be used to isolate improved engineered P450-BM3 polypeptides. For affinity chromatography purification, any antibody that specifically binds to an engineered P450-BM3 polypeptide can be used. To generate antibodies, various host animals, including but not limited to rabbits, mice, rats, and the like, can be immunized by injection of an engineered P450-BM3 polypeptide. The engineered P450-BM3 polypeptide can be attached to a suitable carrier, such as BSA, via a side chain functional group or a linker attached to a side chain functional group. Depending on the host species, and potentially the specific site of immunization, various adjuvants can be used to increase the immunological response, including but not limited to Freund's adjuvant (complete and incomplete), mineral gels such as aluminum hydroxide, surface active substances such as lysolecithin, pluronic polyols, polyanions, peptides, oil emulsions, keyhole limpet hemocyanin, dinitrophenol, and potentially useful human adjuvants such as BCG (bacilli Calmette-Guerin) and Corynebacterium parvum.

[0187] In some embodiments, the engineered P450-BM3 polypeptides are prepared and used in the form of cells expressing the enzyme, as a crude extract, or as an isolated or purified preparation. In some embodiments, the engineered P450-BM3 polypeptides are prepared as a lyophilizate, in powder form (e.g., acetone powder), or as an enzyme solution. In some embodiments, the engineered P450-BM3 polypeptides are in the form of a substantially pure preparation.

[0188] In some embodiments, the engineered P450-BM3 polypeptides are attached to any suitable solid substrate. Solid substrates include, but are not limited to, solids, surfaces, and / or membranes. Solid supports include, but are not limited to, organic polymers such as polystyrene, polyethylene, polypropylene, polyfluoroethylene, polyethylene oxide, and polyacrylamide, as well as copolymers and grafts thereof. Solid supports can also be inorganic, such as glass, silica, controlled-pore glass (CPG), reversed-phase silica, or metals such as gold or platinum. The configuration of the substrate can be in the form of a bead, sphere, particle, pellet, gel, membrane, or surface. The surface can be planar, substantially planar, or non-planar. The solid support can be porous or non-porous, and can have swelling or non-swelling properties. The solid support can be configured in the form of wells, depressions, or other containers, vessels, features, or locations. Multiple supports can be configured at different locations on an array, addressable for robotic delivery of reagents, or addressable by detection methods and / or instrumentation.

[0189] In some embodiments, immunological methods are used to purify engineered P450-BM3 polypeptide variants. In one method, antibodies raised against wild-type or engineered P450-BM3 polypeptides (e.g., against a polypeptide comprising any of SEQ ID NOs: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266, and / or variants thereof, and / or immunogenic fragments thereof) using conventional methods are immobilized on beads, mixed with cell culture media under conditions for variant engineered P450-BM3 binding, and precipitated. In a related method, immunochromatography can be used.

[0190] In some embodiments, the variant engineered P450-BM3 is expressed as a fusion protein comprising a non-enzymatic moiety. In some embodiments, the variant engineered P450-BM3 polypeptide sequence is fused to a domain that facilitates purification. As used herein, the term "domain that facilitates purification" refers to a domain that mediates purification of the polypeptide to which it is fused. Suitable purification domains include, but are not limited to, metal chelating peptides, histidine-tryptophan modules that allow purification on immobilized metal ion affinity chromatography, sequences facilitating purification on glutathione beads, hemagglutinin (HA) tags (corresponding to an epitope derived from the influenza hemagglutinin protein; see, e.g., Wilson et al., Cell 37:767

[1984] ), maltose binding protein sequences, FLAG epitopes for use with the FLAGS extension / affinity purification system (e.g., the system available from Immunex Corp), and the like. One expression vector contemplated for use in the compositions and methods described herein is for expression of a fusion protein comprising a polypeptide of the application fused to a polyhistidine tract separated by a linker peptide cleavable by enterokinase. The histidine residues facilitate purification on IMIAC (immobilized metal ion affinity chromatography; see, e.g., Porath et al., Prot. Exp. Purif., 3:263-281

[1992] ), while the enterokinase cleavage site provides a means for isolating the variant engineered P450-BM3 polypeptide from the fusion protein. The pGEX vectors (Promega) can also be used to express foreign polypeptides as fusion proteins with glutathione S-transferase (GST). In general, such fusion proteins are soluble and can be readily purified from lysed cells by adsorption to glutathione-agarose beads, followed by elution in the presence of free glutathione.

[0191] Accordingly, in another aspect, the application provides methods of producing an engineered enzyme polypeptide, wherein the method comprises culturing a host cell capable of expressing a polynucleotide encoding an engineered enzyme polypeptide under conditions suitable for expression of the polypeptide. In some embodiments, the method further comprises the step of isolating and / or purifying the enzyme polypeptide, as described herein.

[0192] Suitable media and growth conditions for host cells are well known in the art. Any suitable method for introducing a polynucleotide for expressing an enzyme polypeptide into a cell is contemplated for use in the application. Suitable techniques include, but are not limited to, electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion.

[0193] Engineered P450-BM3s having the properties disclosed herein can be obtained by subjecting polynucleotides encoding naturally occurring or engineered P450-BM3 polypeptides to mutagenesis and / or directed evolution methods known in the art and as described herein. Exemplary directed evolution techniques are mutagenesis and / or DNA shuffling (see, e.g., Stemmer, Proc. Natl. Acad. Sci. USA 91 : 10747-10751

[1994] ; WO 95 / 22625; WO 97 / 0078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767; and U.S. Patent 6,537,746). Other directed evolution procedures that can be used include, inter alia, staggered extension process (StEP), in vitro recombination (see, e.g., Zhao et al., Nat. Biotechnol., 16: 258-261

[1998] ), mutagenic PCR (see, e.g., Caldwell et al., PCR Methods Appl., 3: S136-S140

[1994] ), and cassette mutagenesis (see, e.g., Black et al., Proc. Natl. Acad. Sci. USA 93: 3525-3529

[1996] ).

[0194] For example, mutagenesis and directed evolution methods can be readily applied to polynucleotides to generate libraries of variants that can be expressed, screened, and assayed. Mutagenesis and directed evolution methods are well known in the art (see, e.g., U.S. Patent Nos. 5,605,793; 5,811,238; 5,830,721; 5,834,252; 5,837,458; 5,928,905; 6,096,548; 6,117,679; 6,132,970; 6,165,793; 6,180,406; 6,251,674; 6,265,201; 6,277,638; 6,287,861; 6,287,862; 6,291,242; 6,297,053; 6,303,344; 6,309,883; 6,319,713; 6,319,714; 6,323,030; 6,326,204; 6,335,160; 6,335,198; 6,344,356; 6,352,859; 6,355,484; 6,358,740; 6,358,742; 6,365,377; 6,365,408; 6,368,861; 6,372,497; 6,337,186; 6,376,246; 6,379,964; 6,387,702; 6,391,552; 6,391,640; 6,395,547; 6,406,855; 6,406,910; 6,413,745; 6,413,774; 6,420,175; 6,423,542; 6,426,224; 6,436,675; 6,444,468; 6,455,253; 6,479,652; 6,482,647; 6,483,011; 6,484,105; 6,489,146; 6,500,617; 6,500,639; 6,506,602; 6,506,603; 6,518,065; 6,519,065; 6,521,453; 6,528,311; 6,537,746; 6,573,098; 6,576,467; 6,579,678; 6,586,182; 6,602,986; 6,605,430; 6,613,514; 6,653,072; 6,686,515; 6,703,240; 6,716,631; 6,825,001; 6,902,922; 6,917,882; 6,946,296; 6,961,664; 6,995,017; 7,024,312; 7,058,515; 7,105,297; 7,148,054; 7,220,566; 7,288,375; 7,384,387; 7,421,347; 7,430,477; 7,462,469、7,534,564, 7,620,500, 7,620,502, 7,629,170, 7,702,464, 7,747,391, 7,747,393, 7,751,986, 7,776,598, 7,783,428, 7,795,030, 7,853,410, 7,868,138, 7,783,428, 7,873,477, 7,873,499, 7,904,249, 7,957,912, 7,981,614, 8,014,961, 8,029,988, 8,048,674, 8,058,001, 8,076,138, 8,108,150, 8,170,806, 8,224,580, 8,377,681, 8,383,346, 8,457,903, 8,504,498, 8,589,085, 8,762,066, 8,768,871, 9,593,326, 9,665,694, 9,684,771, and all related US and PCT and non-US counterpart patents; Ling et al., Anal. Biochem., 254(2): 157-78

[1997] ; Dale et al., Meth. Mol. Biol., 57: 369-74

[1996] ; Smith, Ann. Rev. Genet., 19: 423-462

[1985] ; Botstein et al., Science, 229: 1193-1201

[1985] ; Carter, Biochem. J., 237: 1-7

[1986] ; Kramer et al., Cell, 38: 879-887

[1984] ; Wells et al., Gene, 34: 315-323

[1985] ; Minshull et al., Curr. Op. Chem. Biol., 3: 284-290

[1999] ; Christians et al., Nat. Biotechnol., 17: 259-264

[1999] ; Crameri et al., Nature, 391: 288-291

[1998] ; Crameri et al., Nat. Biotechnol., 15: 436-438

[1997] ; Zhang et al., Proc. Nat. Acad. Sci. U.S.A., 94: 4504-4509

[1997] ; Crameri et al., Nat. Biotechnol., 14: 315-319

[1996] ; Stemmer, Nature, 370: 389-391

[1994] ; Stemmer, Proc. Nat. Acad. Sci. USA,91 : 10747-10751

[1994] ; WO 95 / 22625; WO 97 / 0078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651 ; WO 01 / 75767; and WO 2009 / 152336, all of which are incorporated herein by reference.

[0195] In some embodiments, enzyme clones obtained after mutagenesis treatment are screened by subjecting the enzyme to a defined temperature (or other assay condition, such as testing the activity of the enzyme on compound (1)) and measuring the amount of enzyme activity remaining after the heat treatment or other assay condition. The clones containing polynucleotides encoding P450-BM3 polypeptides are then sequenced to identify nucleotide sequence changes, if any, and used to express the enzyme in a host cell. Measuring enzyme activity from expression libraries can be performed using any suitable method known in the art (e.g., standard biochemical techniques, such as HPLC analysis).

[0196] For engineered polypeptides of known sequence, polynucleotides encoding the enzyme can be prepared according to known synthetic methods by standard solid phase methods. In some embodiments, fragments of up to about 100 bases can be synthesized individually and then joined (e.g., by enzymatic or chemical ligation methods, or polymerase-mediated methods) to form any desired contiguous sequence. For example, the polynucleotides and oligonucleotides disclosed herein can be prepared by chemical synthesis using the classic phosphoramidite method (see, e.g., Beaucage et al., Tetra. Lett., 22: 1859-69

[1981] ; and Matthes et al., EMBO J., 3: 801-05

[1984] ), as it is commonly practiced in automated synthesis methods. According to the phosphoramidite method, oligonucleotides are synthesized (e.g., in an automated DNA synthesizer), purified, annealed, ligated, and cloned into an appropriate vector.

[0197] Accordingly, in some embodiments, a method for making an engineered P450-BM3 polypeptide can comprise: (a) synthesizing a polynucleotide encoding a polypeptide comprising greater than 85% identity to an amino acid sequence of any variant provided in any of Tables 2.1, 3.1, 4.1, 5.1, 6.1, 7.1, 8.1, 9.1, 10.1, 10.2, 11.1, 11.2, 12.1, 13.1, 14.1, 14.2, 15.1, 16.1, 17.1, 17.2, 18.1, 19.1, and 19.2, as well as SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266; and (b) expressing the P450-BM3 polypeptide encoded by the polynucleotide. In some embodiments of the method, the amino acid sequence encoded by the polynucleotide can optionally have one or several (e.g., up to 3, 4, 5, or up to 10) amino acid residues deleted, inserted, and / or substituted. In some embodiments, the amino acid sequence optionally has 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-15, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-30, 1-35, 1-40, 1-45, or 1-50 amino acid residues deleted, inserted, and / or substituted. In some embodiments, the amino acid sequence optionally has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 30, 35, 40, 45, or 50 amino acid residues deleted, inserted, and / or substituted. In some embodiments, the amino acid sequence optionally has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 20, 21, 22, 23, 24, or 25 amino acid residues deleted, inserted, and / or substituted. In some embodiments, the substitutions can be conservative or non-conservative substitutions.

[0198] The foregoing and other aspects of the present application can be better understood with respect to the following non-limiting examples. The examples are provided for illustrative purposes only and are not intended to limit the scope of the present application in any way.

[0199] Examples

[0200] The following examples, including the experiments and results obtained, are provided for illustrative purposes only and are not intended to limit the scope of the present application.

[0201] In the experimental disclosure that follows, the following abbreviations are used: ppm (parts per million); M (moles); mM (millimoles); uM and μM (micromoles); nM (nanomoles); mol (moles); gm and g (grams); mg (milligrams); ug and μg (micrograms); L and I (liters); ml and mL (milliliters); cm (centimeters); mm (millimeters); um and μm (micrometers); sec. (seconds); min(s) (minutes); h(s) and hr(s) (hours); U (units); MW (molecular weight); rpm (revolutions per minute); °C (degrees Celsius); CDS (coding sequence); DNA (deoxyribonucleic acid); RNA (ribonucleic acid); NA (nucleic acid; polynucleotide); AA (amino acid; polypeptide); E. coli W3110 (a commonly used laboratory strain of E. coli, available from the E. coli Bacterial Genetic Stock Center, New Haven, CT); HPLC (high pressure liquid chromatography); SDS-PAGE (sodium dodecyl sulfate polyacrylamide gel electrophoresis); PES (polyether sulfone); CFSE (carboxyfluorescein succinimidyl ester); IPTG (isopropyl β-D-l-thiogalactopyranoside); PMBS (polymyxin B sulfate); NADPH (nicotinamide adenine dinucleotide phosphate); GDH (glucose dehydrogenase); TON (turnover number); FIOPC (fold improvement over positive control); TON (turnover number); ESI (electrospray ionization); LB (Luria broth); TB (terrific broth); MeOH (methanol); Athens Research (Athens Research Technology, Athens, GA); ProSpec (ProSpec Tany Technogene, East Brunswick, NJ); Sigma-Aldrich (Sigma-Aldrich, St. Louis, MO); Ram Scientific (Ram Scientific, Inc., Yonkers, NY); Pall Corp. (Pall, Corp., Pt. Washington, NY); Millipore (Millipore, Corp.Billerica MA); Difco (Difco Laboratories, BD Diagnostic Systems, Detroit, MI); Molecular Devices (Molecular Devices, LLC, Sunnyvale, CA); Kuhner (Adolf Kuhner, AG, Basel, Switzerland); Cambridge Isotope Laboratories, (Cambridge Isotope Laboratories, Inc., Tewksbury, MA); Applied Biosystems (part of Life Technologies, Corp., Grand Island, NY); Agilent (Agilent Technologies, Inc., Santa Clara, CA); Thermo Scientific (part of Thermo Fisher Scientific, Waltham, MA); Fisher (Fisher Scientific, Waltham, MA); Corning (Corning, Inc., Palo Alto, CA); Waters (Waters Corp., Milford, MA); GE Healthcare (GE Healthcare Bio-Sciences, Piscataway, NJ); Pierce (Pierce Biotechnology, now part of Thermo Fisher Scientific, Rockford, IL); Phenomenex (Phenomenex, Inc., Torrance, CA); Optimal (Optimal Biotech Group, Belmont, CA); and Bio-Rad (Bio-Rad Laboratories, Hercules, CA).

[0202] Example 1

[0203] Production of engineered polypeptides in pCK110900

[0204] A polynucleotide (SEQ ID NO: 3) encoding a polypeptide having monooxygenase activity (SEQ ID NO: 4) was cloned into the pCK110900 vector system (see, e.g., U.S. Patent No. 9,714,437, which is incorporated by reference herein in its entirety) and subsequently expressed in E. coli W3110 fhuA under the control of the lac promoter. This polynucleotide and related polypeptide are derived from a previously engineered B. megaterium variant (see U.S. Application No. 63 / 384,746.

[0205] In a 96-well format, single colonies were picked and grown in 190 pL of LB media containing 1% glucose and 30 pg / mL CAM at 30°C, 200 rpm, and 85% humidity. After overnight incubation, 20 pL of the grown culture was transferred to a deep well plate containing 380 pL of TB media (with 30 pg / mL CAM). The culture was grown for approximately 2.5 hours at 30°C, 250 rpm, and 85% humidity. When the optical density (OD 600 ) of the culture reached 0.4-0.6, expression of the monooxygenase gene was induced by adding IPTG to a final concentration of 1 mM. After induction, the growth was continued for 18-20 hours at 30°C, 250 rpm, and 85% humidity. The cells were harvested by centrifugation at 4,000 rpm and 4°C for 10 minutes; the supernatant was then discarded. The cell pellets were stored at -80°C until ready for use.

[0206] Prior to performing the assay, the cell pellets were thawed and resuspended in 200 or 300 pL of lysis buffer containing 1 g / L lysozyme, 0.5 g / L PMBS, and 0.025 pL / mL commercial DNAse (New England BioLabs, M0303L) in 0.1 M potassium phosphate buffer (pH 8.0). The plates were agitated on a microtiter plate shaker at room temperature for 2 hours at medium speed. The plates were then centrifuged at 4,000 rpm for 10 minutes at 4°C, and the clarified supernatant was used in the HTP assay reactions described in the following examples.

[0207] Shake flask procedures can be used to generate engineered monooxygenase shake flask powders (SFPs) that can be used in secondary screening assays and / or in the biocatalytic processes described herein. In comparison to cell lysates used in HTP assays, shake flask powder preparations of enzymes provide a more purified preparation of engineered enzymes (e.g., up to 30% of total protein) and also allow for the use of more concentrated enzyme solutions. To initiate the culture, a 10 uL aliquot of E. coli glycerol stock containing a plasmid encoding an engineered polypeptide of interest is inoculated into 8 mL of LB cell culture media containing 30 pg / mL CAM and 1% glucose. The culture is grown overnight (at least 16 hours) in a 30 °C incubator with 250 rpm shaking. The grown culture is then added to 250 mL of TB media containing 30 pg / mL CAM in a 1-L shake flask. The 250 mL culture is grown at 30 °C and 250 rpm for 3.5 hours until the OD 600 reaches 0.6-0.8. Expression of the monooxygenase gene is induced by the addition of IPTG to a final concentration of 1 mM, and growth is continued for an additional 18-20 hours. Cells are harvested by transferring the culture to a pre-weighed centrifuge bottle, followed by centrifugation at 4,000 rpm for 10 minutes at 4 °C. The supernatant is discarded, and the remaining cell pellet is weighed. In some embodiments, the cell pellet is stored at -80 °C until ready for use. To lyse, the cell pellet is resuspended in 6 mL / g wet cell weight of 25 mM potassium phosphate (pH 8.0), and lysed using a French Press® 110 L processor system (Microfluidics). Cell debris is removed by centrifugation at 10,000 rpm for 60 minutes at 4 °C. The clarified lysate is collected, frozen at -80 °C, and then lyophilized using standard methods known in the art. Lyophilization of the frozen clarified lysate provides a dry shake flask powder comprising the crude engineered polypeptide.

[0208] Example 2

[0209] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 4 to improve production of compound (2)

[0210] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 4 (SEQ ID NO: 3) were used to generate engineered polypeptides of Table 2.1. These polypeptides exhibit improved monooxygenase activity under desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) as compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 4, as described below.

[0211] Directed evolution started with the polynucleotide set forth in SEQ ID NO: 3. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0212] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 75 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 1 g / L of compound (1), 1 g / L NADPH, 1 g / L GDH-105, 3.6 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air vent seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0213] After overnight incubation, 100 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 10-fold into 1:1 acetonitrile:water for achiral LC-MS analysis according to the method outlined in Table 2.2.

[0214]

[0215]

[0216]

[0217]

[0218] Example 3

[0219] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 36 to improve production of compound (2)

[0220] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 36 (SEQ ID NO: 35) were used to generate engineered polypeptides of Table 3.1. These polypeptides exhibit improved monooxygenase activity under the desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 36, as described below.

[0221] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 35. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0222] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 5 v / v% of undiluted monooxygenase lysate (prepared as described in Example 1), 5 g / L of compound (1), 1 g / L NADPH, 1 g / L GDH-105, 8 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0223] After overnight incubation, 400 μL / well of acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 10-fold into 1:1 acetonitrile:water for achiral LC-MS analysis according to the method outlined in Table 2.2.

[0224]

[0225] Example 4

[0226] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 66 for improved production of compound (2)

[0227] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 66 (SEQ ID NO: 65) were used to generate engineered polypeptides of Table 4.1. These polypeptides exhibit improved monooxygenase activity under desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 66 as described below.

[0228] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 65. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0229] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 5 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 5 g / L compound (1), 1 g / L NADPH, 1 g / L GDH-105, 8 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air vent seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0230] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 10-fold into 1:1 acetonitrile:water for achiral LC-MS analysis according to the method in Table 2.2.

[0231]

[0232]

[0233] Example 5

[0234] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 72 to improve production of compounds (2) and (3)

[0235] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 72 (SEQ ID NO: 71) were used to generate engineered polypeptides of Table 5.1. These polypeptides exhibit improved monooxygenase activity under desired conditions, e.g., improved formation of alcohol compound (2) and / or ketone compound (3) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 72, as described below.

[0236] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 71. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2) or (3).

[0237] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μΐ per well. The reactants contained 2 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 5 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 8 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air vent seal and incubated at 30 °C with 85% humidity with shaking at 600 rpm for 18 hours.

[0238] After overnight incubation, 400 μΐ^ / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 40-fold into 1 : 1 acetonitrile: water for non-chiral LC-MS analysis.

[0239]

[0240]

[0241]

[0242] Example 6

[0243] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 198 to improve production of compound (2)

[0244] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 198 (SEQ ID NO: 197) were used to generate engineered polypeptides of Table 6.1. These polypeptides exhibit improved monooxygenase activity under desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the "backbone" amino acid sequence of SEQ ID NO: 198, as described below.

[0245] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 197. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0246] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 35 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0247] After overnight incubation, 200 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 80-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the method in Table 2.2.

[0248]

[0249] Example 7

[0250] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 226 to improve production of compound (2)

[0251] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 226 (SEQ ID NO: 225) were used to generate engineered polypeptides of Table 7.1. These polypeptides exhibit improved monooxygenase activity under desired conditions, for example, improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 226 as described below.

[0252] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 225. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0253] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 35 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0254] After overnight incubation, 400 pL / well acetonitrile was added to the reaction plate and mixed well. The plate was sealed and centrifuged at 4,000 rpm for 10 minutes. An aliquot of the supernatant was removed and further diluted 80-fold into 1 : 1 acetonitrile: water for achiral LC-MS analysis according to the method in Table 2.2.

[0255]

[0256] Example 8

[0257] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 244 to improve production of compound (2)

[0258] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 244 (SEQ ID NO: 243) were used to generate engineered polypeptides of Table 8.1. These polypeptides exhibited improved monooxygenase activity under the desired conditions, for example, improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences of even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 244, as described below.

[0259] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 243. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0260] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 pL per well. The reactants contained 5 v / v% of undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L of compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose, and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal, and incubated at 30 °C with 85% humidity at 600 rpm for 18 hours.

[0261] After overnight incubation, 400 pL / well acetonitrile was added to the reaction plate and mixed well. The plate was sealed and centrifuged at 4,000 rpm for 10 minutes. An aliquot of the supernatant was removed and further diluted 80-fold into 1 : 1 acetonitrile: water for achiral LC-MS analysis according to the method in Table 2.2.

[0262] Additional selectivity analysis was performed on selected variants. Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 40 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air vent seal and incubated at 30 °C with 85% humidity with shaking at 600 rpm for 18 hours.

[0263] After overnight incubation, 50 μL / well of a 25 μL / mL triethylamine acetonitrile solution was added to the reaction plates, followed by 50 μL / well of a 37 g / L 2-bromoacetophenone acetonitrile solution. The plates were heat sealed and incubated at 50 °C with shaking for 2 hours. After 2 hours, the plates were centrifuged at 4,000 rpm for 10 minutes. The supernatants were then filtered through hydrophilic filter plates. Aliquots of the filtrate were removed and further diluted 10-fold into 1:1 acetonitrile:water for chiral LC-MS analysis according to the method outlined in Table 8.2.

[0264]

[0265]

[0266] Derivatization was performed to facilitate chiral LC-MS analysis as outlined in Scheme 5 below.

[0267]

[0268] Scheme 5

[0269]

[0270]

[0271] Example 9

[0272] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 286 to improve production of compound (2)

[0273] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 286 (SEQ ID NO: 285) were used to generate engineered polypeptides of Table 9.1. These polypeptides exhibit improved monooxygenase activity under desired conditions, for example, improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 286 as described below.

[0274] Directed evolution started with the polynucleotide set forth in SEQ ID NO: 285. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0275] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 5 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0276] After the overnight incubation, 400 μL / well of acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 80-fold into 1:1 acetonitrile:water for achiral LC-MS analysis according to the method described in Table 2.2.

[0277] Additional selectivity assays were performed on selected variants. Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 40 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0278] After the overnight incubation, 50 μL / well of a 25 μL / mL triethylamine acetonitrile solution was added to the reaction plates, followed by 50 μL / well of a 37 g / L 2-bromoacetophenone acetonitrile solution. The plates were heat sealed and shaken at 50 °C for 2 hours. After 2 hours, the plates were centrifuged at 4,000 rpm for 10 minutes. The supernatant was then filtered through a hydrophilic filter plate. Aliquots of the filtrate were removed and further diluted 10-fold into 1:1 acetonitrile:water for chiral LC-MS analysis according to the method in Table 8.2.

[0279]

[0280]

[0281] Example 10

[0282] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 358 to improve production of compound (2)

[0283] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 358 (SEQ ID NO: 357) were used to generate engineered polypeptides of Tables 10.1 and 10.2. These polypeptides exhibited improved monooxygenase activity under the desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 358, as described below.

[0284] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 357. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0285] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 5 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0286] After the overnight incubation, 400 μL / well of acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 80-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the method described in Table 2.2.

[0287] Additional selectivity analysis was performed on selected variants. Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 40 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0288] After overnight incubation, 50 pL / well of 25 pL / mL triethylamine in acetonitrile was added to the reaction plate, followed by 50 pL / well of 37 g / L 2-bromoacetophenone in acetonitrile. The plate was heat sealed and shaken at 50 °C for 2 hours. After 2 hours, the plate was centrifuged at 4,000 rpm for 10 minutes. Then, the supernatant was filtered through a hydrophilic filter plate. An aliquot of the filtrate was taken and further diluted 10-fold into 1:1 acetonitrile:water for chiral LC-MS analysis according to the method in Table 8.2.

[0289]

[0290]

[0291] Analysis of additional variants derived from SEQ ID NO: 357 was performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 pL per well. The reactants contained 10 v / v% of undiluted monooxygenase lysate (prepared as described in Example 1 and then heated to 40 °C for 2 hours), 10 g / L of compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose, and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal, and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0292] After overnight incubation, 400 pL / well of acetonitrile was added to the reaction plate and mixed well. The plate was sealed and centrifuged at 4,000 rpm for 10 minutes. An aliquot of the supernatant was taken and further diluted 80-fold into 1:1 acetonitrile:water for achiral LC-MS analysis.

[0293]

[0294]

[0295] Example 11

[0296] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 410 to improve production of compound (2)

[0297] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 410 (SEQ ID NO: 409) were used to generate engineered polypeptides of Tables 11.1 and 11.2. These polypeptides exhibit improved monooxygenase activity under desired conditions, for example, improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 410, as described below.

[0298] Directed evolution started with the polynucleotide set forth in SEQ ID NO: 409. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0299] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 10 v / v% undiluted monooxygenase lysate (prepared as described in Example 1 and then heated to 42.5°C for 2 hours), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30°C with 85% humidity. Some variants were also analyzed under similar conditions with 5% v / v undiluted monooxygenase lysate (prepared as described in Example 1 and without a heat pre-incubation).

[0300] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 400 or 800-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the conditions in Table 2.2.

[0301]

[0302]

[0303]

[0304] Analysis of additional variants derived from SEQ ID NO: 409 was performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 5 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.25 g / L PDH-102 and were dissolved in 285 mM sodium phosphite buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30°C with 85% humidity.

[0305] The same variants were analyzed where the lysate had been heat treated. The reactants contained 40 v / v% monooxygenase lysate (prepared as described in Example 1, then diluted 4-fold into 300 mM sodium phosphite, pH 8 and heated to 42.5 °C for two hours), 10 g / L of compound (1), 1 g / L NADPH, 0.25 g / L PDH-102 and dissolved in 270 mM sodium phosphite buffer, pH 8. The reaction plates were sealed with an air vent seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0306] After overnight incubation, 400 μL / well of acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 800-fold into 1:1 acetonitrile:water for non-chiral LC-MS analysis.

[0307]

[0308]

[0309] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 544 to improve production of compound (2)

[0310] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 544 (SEQ ID NO: 543) were used to generate engineered polypeptides of Table 12.1. These polypeptides exhibit improved monooxygenase activity under desired conditions, for example, improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 544 as described below.

[0311] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 543. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0312] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 40 v / v% of 4-fold diluted monooxygenase lysate (prepared as described in Example 1 and then 4-fold diluted and heated to 46°C for 2 hours), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with gas port seals and shaken at 600 rpm for 18 hours at 30°C with 85% humidity.

[0313] After overnight incubation, 500 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 800-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the conditions in Table 2.2.

[0314]

[0315]

[0316] Example 13

[0317] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 734 to improve production of compound (2)

[0318] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 734 (SEQ ID NO: 733) were used to generate engineered polypeptides of Table 13.1. These polypeptides exhibit improved monooxygenase activity under desired conditions, for example, improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 734, as described below.

[0319] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 734. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0320] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 2 v / v% and / or 50% v / v undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C and 85% humidity.

[0321] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were taken and further diluted 400 or 800-fold into 1:1 acetonitrile:water for achiral LC-MS analysis according to the method in Table 2.2.

[0322] Additional selectivity analysis was performed for selected variants. Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 40 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.1 g / L GDH-105, 16 g / L glucose and were dissolved in 100 mM potassium phosphate buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C and 85% humidity.

[0323] After overnight incubation, 50 μL / well of a 25 μL / mL triethylamine acetonitrile solution was added to the reaction plates, followed by 50 μL / well of a 37 g / L 2-bromoacetophenone acetonitrile solution. The plates were heat sealed and shaken at 50 °C for 2 hours. After 2 hours, the plates were centrifuged at 4,000 rpm for 10 minutes. The supernatant was then filtered through a hydrophilic filter plate. Aliquots of the filtrate were taken and further diluted 50-fold into 1:1 acetonitrile:water for chiral LC-MS analysis according to the method in Table 8.2.

[0324]

[0325]

[0326] Example 14

[0327] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 748 to improve production of compound (2)

[0328] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 748 (SEQ ID NO: 747) were used to generate engineered polypeptides of Tables 14.1 and 14.2. These polypeptides exhibit improved monooxygenase activity under desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 748, as described below.

[0329] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 748. Libraries of engineered polypeptides were generated using various well-known techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0330] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 µL per well. All variants in these and other examples were screened using phosphite dehydrogenase (PDH) and phosphite instead of glucose dehydrogenase (GDH) and glucose to regenerate the NADPH cofactor. The switch was made due to the superior performance of SFP under PDH conditions. The reactants contained 1 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 10 g / L compound (1), 1 g / L NADPH, 0.25 g / L PDH-102 and were dissolved in 500 mM sodium phosphite buffer (pH 8). The reaction plates were sealed with an air hole seal and incubated at 30 °C with 85% humidity with shaking at 600 rpm for 18 hours.

[0331] After the overnight incubation, 400 µL / well of acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 800-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the conditions in Table 2.2.

[0332]

[0333]

[0334] Analysis of additional variants derived from SEQ ID NO: 747 was performed in 96- well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reaction contained 10 v / v% of undiluted monooxygenase lysate (prepared as described in Example 1), 25 g / L of compound (1), 1 g / L NADPH, 0.5 g / L PDH-102 and was dissolved in 450 mM sodium phosphite buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0335] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 100-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the conditions in Table 2.2.

[0336]

[0337]

[0338] Example 15 Evolution and screening of engineered polypeptides derived from SEQ ID NO: 828 to improve production of compound (2)

[0339] An engineered polynucleotide encoding a polypeptide having monooxygenase activity of SEQ ID NO: 828 (SEQ ID NO: 827) was used to generate engineered polypeptides of Table 15.1. These polypeptides exhibit improved monooxygenase activity under the desired conditions, for example, improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 828 as described below.

[0340] Directed evolution started with the polynucleotide set forth in SEQ ID NO: 827. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0341] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 5 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 25 g / L compound (1), 1 g / L NADPH, 0.5 g / L PDH-102 and were dissolved in 475 mM sodium phosphite buffer (pH 8). The reaction plates were sealed with gas port seals and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0342] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 800-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the conditions in Table 2.2.

[0343]

[0344]

[0345] Example 16

[0346] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 968 to improve production of compound (2)

[0347] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 968 (SEQ ID NO: 967) were used to generate engineered polypeptides of Table 16.1. These polypeptides exhibit improved monooxygenase activity under the desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 968, as described below.

[0348] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 967. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0349] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 5 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 25 g / L compound (1), 1 g / L NADPH, 0.5 g / L PDH-102 and were dissolved in 475 mM sodium phosphite buffer (pH 8). The reaction plates were sealed with gas port seals and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0350] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 800-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the conditions in Table 2.2.

[0351]

[0352]

[0353]

[0354] Example 17

[0355] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 984 to improve production of compound (2)

[0356] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 984 (SEQ ID NO: 983) were used to generate engineered polypeptides of Tables 17.1 and 17.2. These polypeptides exhibit improved monooxygenase activity under the desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 984, as described below.

[0357] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 983. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0358] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 5 (Table 17.1) or 10 v / v% (Table 17.2) undiluted monooxygenase lysate (prepared as described in Example 1), 25 g / L compound (1), 1 g / L NADPH, 0.5 g / L PDH-102 and were dissolved in 450 mM sodium phosphite buffer (pH 8). The reaction plates were sealed with an air vent seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0359] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 800-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the conditions in Table 2.2.

[0360]

[0361]

[0362]

[0363] Example 18

[0364] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 1160 to improve production of compound (2)

[0365] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 1160 (SEQ ID NO: 1159) were used to generate engineered polypeptides of Table 18.1. These polypeptides exhibit improved monooxygenase activity under the desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 1160, as described below.

[0366] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 1159. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0367] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates in duplicate with a total reaction volume of 100 μL per well. The reactants contained 5 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 25 g / L compound (1), 1 g / L NADPH, 0.5 g / L PDH-102 and were dissolved in 475 mM sodium phosphite buffer. One set of reactants was set at pH 7.5 and the other set of duplicate reactants was set at pH 8. The reaction plates were sealed with an air vent seal and incubated at 30 °C with 85% humidity with shaking at 600 rpm for 18 hours.

[0368] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 800-fold into 1:1 acetonitrile:water for analysis by achiral LC-MS according to the conditions in Table 2.2.

[0369]

[0370]

[0371] Example 19

[0372] Evolution and screening of engineered polypeptides derived from SEQ ID NO: 1266 to improve production of compound (2)

[0373] Engineered polynucleotides encoding polypeptides having monooxygenase activity of SEQ ID NO: 1266 (SEQ ID NO: 1265) were used to generate engineered polypeptides of Tables 19.1 and 19.2. These polypeptides exhibit improved monooxygenase activity under the desired conditions, e.g., improved formation of alcohol compound (2) from substrate compound (1) compared to the starting polypeptide. Engineered polypeptides having amino acid sequences with even-numbered sequence identifiers were generated from the “backbone” amino acid sequence of SEQ ID NO: 1266, as described below.

[0374] Directed evolution began with the polynucleotide set forth in SEQ ID NO: 1265. Libraries of engineered polypeptides were generated using various techniques (e.g., saturation mutagenesis, recombination of previously identified beneficial amino acid differences) and screened using HTP assays and analytical methods that measure the ability of the polypeptides to produce compound (2).

[0375] Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 10 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 25 g / L compound (1), 1 g / L NADP+, 0.5 g / L PDH-102 and were dissolved in 450 mM sodium phosphite buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0376] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 800-fold into 1:1 acetonitrile:water for achiral LC-MS analysis according to the conditions in Table 2.2.

[0377]

[0378]

[0379] Other variants derived from SEQ ID NO: 1265 were analyzed. Enzyme assays were performed in 96-well deep well (1.1 mL total volume) plates with a total reaction volume of 100 μL per well. The reactants contained 40 v / v% undiluted monooxygenase lysate (prepared as described in Example 1), 25 g / L compound (1), 1 g / L NADPH, 0.5 g / L PDH-102 and were dissolved in 300 mM sodium phosphite buffer (pH 8). The reaction plates were sealed with an air hole seal and shaken at 600 rpm for 18 hours at 30 °C with 85% humidity.

[0380] After overnight incubation, 400 μL / well acetonitrile was added to the reaction plates and mixed well. The plates were sealed and centrifuged at 4,000 rpm for 10 minutes. Aliquots of the supernatant were removed and further diluted 800-fold into 1:1 acetonitrile:water for achiral LC-MS analysis according to the conditions in Table 2.2.

[0381]

[0382] While the application has been described with reference to specific embodiments thereof, various changes in form and detail can be made thereto and equivalents thereof can be substituted without departing from the scope of the application as defined in the appended claims.

[0383] Each and every publication and patent document cited in this disclosure is hereby incorporated herein by reference for all purposes to the same extent as if each such publication or document had been specifically and individually indicated to be incorporated herein by reference. The citation of publications and patent documents is not intended as an admission that any such document is pertinent prior art, nor does it constitute an admission as to the contents or date of the same.

Claims

1. An engineered cytochrome P450-BM3 variant comprising a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 36, 4, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160 or 1266, or a functional fragment thereof, wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions, and wherein the amino acid positions of the polypeptide sequence are referenced to SEQ ID NO: NO: 36, 4, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160 or 1266.

2. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 4, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 52 / 83 / 88 / 105, 32 / 83 / 88, 32 / 83 / 88 / 176, 32 / 83 / 88 / 231 / 574, 32 / 83 / 88 / 576. 4, 52 / 83 / 88, 52 / 83 / 88 / 231, 52 / 83 / 88 / 231 / 433 / 574, 52 / 83 / 88 / 433, 52 / 83 / 88 / 433 / 574, 52 / 83 / 88 / 574, 83 / 88, 83 / 88 / 105, 83 / 88 / 111, 83 / 88 / 111 / 433, 83 / 88 / 111 / 574, 83 / 88 / 231, 83 / 88 / 349, 83 / 88 / 433 / 574, and 83 / 88 / 574, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

4.

3. The engineered cytochrome P450-BM3 variant of claim 1, wherein the polypeptide sequence is identical to SEQ ID NO:36 having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 75, 75 / 374, 75 / 374 / 458 / 726, 75 / 374 / 726, 75 / 458, 75 / 458 / 726, 75 / 726, 111 / 114, 111 / 603 / 604 / 623 / 853, 111 / 623, 374 / 726 and 726, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

36.

4. An engineered cytochrome P450-BM3 variant as described in claim 1, wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:66, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 74, 75, 83, 179, 181, 182, 186, 189, 238, 267, 268, 328, 331, 355, 358, 437 and 438, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

66.

5. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 72,and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 74 / 75 / 83 / 179 / 189 / 328 / 331 / 437, 74 / 75 / 268 / 328 / 331 / 358 / 437, 74 / 83 / 268 / 328 / 331 / 358 / 437, 74 / 267 / 268 / 328, 75 / 83 / 179 / 189 / 331, 75 / 83 / 179 / 268 / 328 / 331 / 437, 75 / 83 / 189 / 268 / 328 / 331 / 358 / 437. 58 / 654、75 / 83 / 268 / 437、75 / 268 / 328 / 331 / 358、83 / 179 / 182 / 437、83 / 179 / 189 / 331 / 355、83 / 179 / 328 / 331、83 / 179 / 355 / 358 / 437、83 / 182 / 189 / 268 / 328 / 331 / 355 / 358 / 437、83 / 182 / 189 / 328 / 331、83 / 182 / 268 / 328 / 355 / 358 / 437、83 / 189 / 267 / 268 / 358、83 / 189 / 268 / 328 / 331、83 / 189 / 268 / 328 / 355 / 358 / 437、83 / 189 / 328 / 331 / 437、83 / 267 / 268 / 328 / 331 / 355 / 358、83 / 268、83 / 268 / 328 / 331、83 / 268 / 328 / 331 / 355 / 358、83 / 268 / 328 / 331 / 358、83 / 268 / 328 / 331 / 358 / 437、83 / 268 / 328 / 331 / 437、83 / 268 / 331、83 / 331、83 / 331 / 437、83 / 358 / 437、179 / 182 / 268、179 / 189 / 331 / 437、179 / 328 / 3 31, 179 / 331 / 358, 189 / 267 / 268 / 437, 189 / 268 / 328, 189 / 268 / 328 / 331 / 358 / 437, 189 / 268 / 358, 267 / 268 / 328, 267 / 268 / 328 / 331 / 355 / 358, 267 / 268 / 331 / 355 / 358 / 437, 268, 268 / 328 / 331, 268 / 328 / 355 / 358, 268 / 331 / 355 / 358, 268 / 355 / 358 / 437, 268 / 358 and 328 / 331 / 358, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

72.

6. An engineered cytochrome P450-BM3 variant as described in claim 1, wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 198, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 79, 213 and 257, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

198.

7. An engineered cytochrome P450-BM3 variant as described in claim 1, wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:226, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 315, 320, 385, 388, 391, 398, 405, 493, 497, 502, 503, 504, 541, 542, 547, 573, 576 and 577, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

226.

8. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 244, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 75 / 178 / 213 / 315, 75 / 331 / 576 / 726, 178, 178 / 179 / 213 / 437 / 497 / 573 / 576, 178 / 179 / 573 / 726, 178 / 179 / 573 / 726. , 178 / 213 / 573, 178 / 213 / 726, 178 / 437, 178 / 497 / 726, 178 / 576, 178 / 726, 179 / 358, 179 / 726, 331, 331 / 358 / 391 / 437, 331 / 497, 331 / 573 / 576, 497 / 573, 573, 682, 685, 699, 701, 704, 707, 708, 726, 756, 759, 794, 796, 797, 848, 851, 862, 888, 889, 999, 1003, and 1048, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

244.

9. The engineered cytochrome P450-BM3 variant of claim 1, wherein the polypeptide sequence is identical to SEQ IDNO:286 has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 75 / 331 / 701 / 726 / 851 / 1048, 75 / 331 / 726 / 999, 75 / 726 / 796 / 851 / 999, 75 / 1048, 87, 88 / 522, 89, 234, 269, 328, 330, 331, 398, 405, 408 and 411, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

286.

10. An engineered cytochrome P450-BM3 variant as described in claim 1, wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:358, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 75 / 269 / 707, 269, 269 / 522 / 707, 269 / 522 / 707 / 1048, 522, 522 / 726, 522 / 1048 and 726 / 1048, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

358.

11. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 358, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 77, 170, 286, 289, 462, 547, 557, 630, 646, 651, 672, 676, 692, 775, 786, 787, 788, 814, 841, 876, 877, 888, 893, 896, 924, 941, 955, 969, 973, 982, 989, 993 and 1038, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

358.

12. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 410, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 158 / 170, 158 / 170 / 410 / 462 / 630 / 672 / 726 / 786 / 788 / 814 / 924、158 / 170 / 410 / 462 / 726 / 786 / 814、158 / 170 / 410 / 924、158 / 170 / 630 / 726 / 786 / 788 / 814、158 / 410、158 / 410 / 462 / 557 / 969、158 / 410 / 557 / 786 / 787、158 / 410 / 814 / 924、158 / 410 / 862、158 / 462 / 557 / 630 / 924、158 / 462 / 630 / 786 / 788 / 969、158 / 557、158 / 557 / 630 / 814、158 / 557 / 726 / 786 / 788 / 862、158 / 557 / 786、158 / 557 / 786 / 787 / 788、158 / 557 / 814 / 924、158 / 630 / 786 / 814、158 / 726、158 / 786 / 787 / 788、158 / 814 / 924、170 / 410 / 557 / 786 / 787 / 814 / 924 / 989, 385, 410 / 462 / 557 / 630 / 786 / 787 / 969, 462 / 557 / 726 / 786 / 787 / 924, 469, 523, 550, 553, 556, 574, 613, 640, 645, 650, 652, 717, 773, 779, 795, 838, 871, 923 and 926, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

410.

13. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 410, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 51 / 851, 460, 461 / 462. 6, 474, 597, 600, 635, 638, 655, 663, 664, 677, 694, 696, 713, 771, 783, 789, 806, 807, 840, 842, 851, 857, 860, 878, 894, 942, 947, 960, 978, 992, 1008, 1012, 1024 and 1025, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

410.

14. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 544, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 77 / 179 / 286 / 410 / 788 / 888, 77 / 41 0 / 676 / 788 / 924, 77 / 557 / 707 / 888, 286 / 410 / 651 / 676, 286 / 410 / 707 / 788, 286 / 410 / 888, 286 / 692 / 786 / 788, 410, 410 / 557 / 676 / 788 / 888 / 924 / 993, 410 / 557 / 692 / 788 and 410 / 646 / 651 / 788, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

544.

15. The engineered cytochrome P450-BM3 variant of claim 1, wherein the polypeptide sequence is identical to SEQ ID NO:734 having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 24, 53, 75, 78, 82, 88, 150, 180, 183, 257, 270, 410 / 497 / 557 / 576 / 814, 410 / 497 / 573 / 576, 410 / 497 / 814, 410 / 557 / 924, 437 and 497 / 557, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

734.

16. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 748, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 75, 75 / 82 / 180 / 257 / 437, 75 / 82 / 257 / 268, 75 / 82 / 257 / 556 / 640, 75 / 82 / 257 / 773, 75 / 82 / 437, 75 / 180 / 183 / 257 / 268 / 270 / 385 / 437 / 556 / 613 / 652 / 923, 75 / 180 / 257 / 268 / 270 / 437 / 556 / 574 / 652, 75 / 180 / 574 / 795, 75 / 257 / 613, 75 / 257 / 640, 75 / 556 / 773, 75 / 574, 75 / 613, 82 / 613, 180 / 437 / 773 / 795, and 257 / 613 / 773 / 795, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

748.

17. The engineered cytochrome P450-BM3 variant of claim 1, wherein the polypeptide sequence is identical to SEQ ID NO:748 having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 45, 102, 106, 110, 111, 114, 127, 191, 193, 194, 196, 196 / 853, 198, 202, 203, 206, 210, 226, 232, 236, 237, 244, 245, 248, 254, 256 and 347, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

748.

18. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 828, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 22 / 75 / 717 / 720 / 779 / 1004, 22 / 550 / 826, 22 / 616 / 717 / 1004 , 22 / 717 / 795 / 799 / 826, 550 / 616 / 717 / 779, 550 / 640 / 717, 550 / 717 / 795 / 799 / 800, 616 / 717 / 720 / 799, 640, 717, 717 / 720 / 779 / 1004, 717 / 779 / 799 / 800 / 1004, 717 / 1004, 720 / 779, 779 / 1004, 800 / 1004 and 826 / 1004, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

828.

19. The engineered cytochrome P450-BM3 variant of claim 1, wherein the polypeptide sequence is identical to SEQ ID NO: 968 having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 3, 118 / 446, 132, 230, 285, 290, 292, 293, 295, 296, 300, 303, 305, 307, 366, 371, 372, 381, 382, ​​415, 417, 418, 424, 427, 432, 433, 446, 447, 451, 453, 454, 455, 456, 457, 458, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492 55, 463, 465, 468, 473, 473 / 790, 478, 480, 481, 506, 600 / 635 / 713 / 771 / 1025, 600 / 840 / 960, 635 / 636 / 793 / 840, 635 / 638 / 793 / 823 / 960, 635 / 713 / 1025, 635 / 7 71 / 894, 636 / 638 / 793 / 851, 638 / 663 / 793 / 840, 713, 771, 786 / 840 / 960, 793 / 840, 807, 840 / 960, 851 / 1024 / 1025, 851 / 1025, 960 and 1025, wherein the amino acid positions of the polypeptide sequences are numbered with reference to SEQ ID NO:

968.

20. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 984, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 45, 45 / 111 / 226 / 347, 45 / 853, 110 / 114, 111, 111 / 127 / 226 / 244, 111 / 194 / 244 / 347 / 853, 111 / 194 / 347, 111 / 194 / 347 / 853, 111 / 210, 111 1 / 210 / 347 / 853, 111 / 226, 111 / 226 / 244 / 347, 111 / 226 / 853 / 969, 111 / 244, 111 / 244 / 853, 111 / 347 / 853, 111 / 853, 114, 114 / 245, 127 / 210 / 244, 127 / 210 / 244 / 853, 127 / 210 / 347 / 853, 127 / 244, 127 / 347, 194 / 853, 210 / 853, 226 / 236 / 244 / 347 / 853, 236 / 244, 237, 237 / 245, 244 / 853 and 245, wherein the amino acid positions of the polypeptide sequences are referenced to SEQ ID NO:984 is numbered.

21. An engineered cytochrome P450-BM3 variant as described in claim 1, wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:984, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 518, 518 / 652, 519, 562, 563, 584 / 724, 586, 616, 618, 619, 621, 623, 628, 640, 653 and 666, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO:

984.

22. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1160, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 132 / 366 / 433 / 463 / 467 / 793, 132 / 366 / 467 / 661, 132 / 467 / 468 / 506 / 793, 132 / 468 / 793, 132 / 1025, 183 / 1025 / 1045 / 1048, 290 / 366 / 433 / 463 / 467, 290 / 433 / 467 / 793 / 1025, 290 / 433 / 793, 290 / 793, 290 / 1025, 366 / 433, 433, 433 / 467 / 1025, 433 / 506 / 1025, 433 / 790, 463 / 793, 473, 793 and 1025, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 1160.

23. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1266, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 114 / 230 / 446 / 8 53, 114 / 292 / 293 / 296 / 463 / 853, 132 / 183 / 366 / 467 / 661 / 1025 / 1045 / 1048, 132 / 290 / 366 / 433 / 467 / 661 / 793 / 1025, 132 / 290 / 366 / 467 / 661 / 793, 132 / 290 / 366 / 467 / 661 / 1025, 132 / 366 / 433 / 467 7 / 506 / 661 / 1025、132 / 366 / 433 / 467 / 661、132 / 366 / 433 / 467 / 661 / 790、132 / 366 / 463 / 467 / 661 / 793、132 / 366 / 467 / 473 / 661、132 / 366 / 467 / 661 / 793、366 / 467 / 468 / 506 / 661 / 793、433 / 463 / 467 / 661 / 793, 463, 463 / 853, 689, 720, 724, 730, 769, 780, 792, 810, 817, 824, 853, 923, 926, 939, 952, 962, 968, 974, 979, 981, 995, 1006, 1015, 1017, 1022, 1027, 1031 and 1040, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 1266.

24. An engineered cytochrome P450-BM3 variant as described in claim 1, wherein the polypeptide sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1266, and wherein the polypeptide sequence of the engineered cytochrome P450-BM3 variant comprises at least one substitution or set of substitutions at one or more positions in the polypeptide sequence selected from the group consisting of: 458, 518 / 653, 519 / 628, 616 / 619 and 653, wherein the amino acid positions of the polypeptide sequence are numbered with reference to SEQ ID NO: 1266.

25. An engineered cytochrome P450-BM3 variant as described in claim 1, wherein the engineered cytochrome P450-BM3 variant comprises a polypeptide sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the sequence of at least one engineered cytochrome P450-BM3 variant listed in Tables 2.1, 3.1, 4.1, 5.1, 6.1, 7.1, 8.1, 9.1, 10.1, 10.2, 11.1, 11.2, 12.1, 13.1, 14.1, 14.2, 15.1, 16.1, 17.1, 17.2, 18.1, 19.1 and 19.

2.

26. The engineered cytochrome P450-BM3 variant of claim 1 , wherein the engineered cytochrome P450-BM3 variant comprises a polypeptide sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to SEQ ID NO: 4, 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160 or 1266.

27. The engineered cytochrome P450-BM3 variant of claim 1, wherein the engineered cytochrome P450-BM3 variant comprises a sequence set forth in SEQ ID NO: 36, 66, 72, 198, 226, 244, 286, 358, 410, 534, 734, 748, 828, 968, 984, 1160, or 1266.

28. An engineered cytochrome P450-BM3 variant as described in claim 1, wherein the engineered cytochrome P450-BM3 variant comprises a polypeptide sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity with the sequence of at least one engineered cytochrome P450-BM3 variant listed in the even-numbered sequences in SEQ ID NO:4-1368.

29. The engineered cytochrome P450-BM3 variant of claim 1, wherein the engineered cytochrome P450-BM3 variant comprises a polypeptide sequence set forth in at least one of the even-numbered sequences in SEQ ID NOs: 4-1368.

30. The engineered cytochrome P450-BM3 variant of any one of claims 1-29, wherein the engineered cytochrome P450-BM3 variant comprises at least one improved property compared to wild-type Bacillus megaterium cytochrome P450-BM3 or the engineered P450-BM3 variant.

31. The engineered cytochrome P450-BM3 variant of claim 30, wherein the improved property comprises improved activity towards a substrate.

32. The engineered cytochrome P450-BM3 variant of claim 31, wherein the substrate comprises 1-tert-butoxycarbonylaminocyclopentanoic acid (Compound (1)).

33. The engineered cytochrome P450-BM3 variant of claim 30, wherein the improved properties comprise improved thermal stability or increased activity towards a substrate after pre-incubation at 42.5°C.

34. The engineered cytochrome P450-BM3 variant of claim 30, wherein the improved properties comprise improved stereoselectivity for one or more diastereomeric products.

35. The engineered cytochrome P450-BM3 variant of any one of claims 1-34, wherein the engineered cytochrome P450-BM3 variant is purified.

36. A composition comprising at least one engineered cytochrome P450-BM3 variant of any one of claims 1-35.

37. A polynucleotide sequence encoding at least one engineered cytochrome P450-BM3 variant according to any one of claims 1-34.

38. A polynucleotide sequence encoding at least one engineered cytochrome P450-BM3 variant, said polynucleotide sequence comprising at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO:3, 35, 65, 71, 197, 225, 243, 285, 357, 409, 533, 733, 747, 827, 967, 983, 1159 or 1265, wherein said polynucleotide sequence of said engineered cytochrome P450-BM3 variant comprises at least one substitution at one or more positions.

39. A polynucleotide sequence encoding at least one engineered cytochrome P450-BM3, said polynucleotide sequence comprising at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 3, 35, 65, 71, 197, 225, 243, 285, 357, 409, 533, 733, 747, 827, 967, 983, 1159 or 1265 or a functional fragment thereof.

40. The polynucleotide sequence of any one of claims 37-39, wherein the polynucleotide sequence is operably linked to a control sequence.

41. The polynucleotide sequence of any one of claims 37-40, wherein the polynucleotide sequence is codon optimized.

42. The polynucleotide sequence of any one of claims 37-41, wherein the polynucleotide sequence comprises a polynucleotide sequence set forth in an odd-numbered sequence in SEQ ID NO: 3-1367.

43. An expression vector comprising at least one polynucleotide sequence according to any one of claims 37 to 42.

44. A host cell comprising at least one expression vector according to claim 43.

45. A host cell comprising at least one polynucleotide sequence according to any one of claims 37-42.

46. ​​A method for producing an engineered cytochrome P450-BM3 variant in a host cell, the method comprising culturing the host cell of claim 44 and / or 45 under suitable conditions such that at least one engineered cytochrome P450-BM3 variant is produced.

47. The method of claim 46, further comprising recovering the at least one engineered cytochrome P450-BM3 variant from the culture and / or host cell.

48. The method of claim 47, further comprising the step of purifying the at least one engineered cytochrome P450-BM3 variant.

Citation Information

Patent Citations

  • Methods for in vitro recombination

    US5605793A

  • Methods for generating polynucleotides having desired characteristics by iterative selection and recombination

    US5811238A

  • DNA mutagenesis by random fragmentation and reassembly

    US5830721A

  • End-complementary polymerase reaction

    US5834252A

  • Methods and compositions for cellular and metabolic engineering

    US5837458A