Biocatalyst and method thereof

Modified transaminase polypeptides with specific amino acid changes enhance enzyme activity and stereoselectivity, addressing the need for efficient niraparib intermediate production, achieving up to 50-fold activity improvement.

JP2025527748APending Publication Date: 2025-08-22TESARO INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025511941
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-26
Filing Date
2023-08-24
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

There is a need for improved transaminase biocatalysts that can efficiently and cost-effectively prepare amine intermediate compounds for the production of niraparib, a PARP inhibitor, while maintaining high stereoselectivity and enzymatic activity.

Method used

Modified transaminase polypeptides with specific amino acid substitutions and modifications, such as those involving X124 being isoleucine, are used to enhance enzyme activity and stereoselectivity, allowing for the production of amine intermediates through transamination reactions.

Benefits of technology

The modified transaminase variants exhibit significantly increased enzymatic activity, up to 50-fold improvement, and maintain high stereoselectivity, making them suitable for the production of niraparib intermediates with improved efficiency and cost-effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025527748000001
    Figure 2025527748000001
  • Figure 2025527748000002
    Figure 2025527748000002
  • Figure 2025527748000003
    Figure 2025527748000003
Patent Text Reader

Abstract

The present disclosure relates to transaminase biocatalysts and methods of using the biocatalysts. The present disclosure also includes methods for preparing amine intermediate compounds useful in the production of niraparib. The present disclosure also includes polynucleotides encoding the transaminase biocatalysts and host cells containing the polynucleotides.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to transaminase biocatalysts and methods of using the biocatalysts. The present disclosure also includes methods for preparing amine intermediate compounds useful in the production of niraparib. The present disclosure also includes polynucleotides encoding the transaminase biocatalysts and host cells containing the polynucleotides. [Background technology]

[0002] Poly(ADP-ribose) polymerase (PARP) is a family of enzymes involved in DNA damage repair. PARP inhibition leads to the accumulation of single-strand breaks, which subsequently lead to the accumulation of double-strand breaks. Cells with increased numbers of DSBs are more dependent on other DNA repair pathways, primarily the homologous recombination (HR) repair pathway. Therefore, cancer cells with HR deficiencies are more susceptible to PARP impairment of the BER pathway. Currently, four PARP inhibitors have been approved (e.g., niraparib, olaparib, rucaparib, and talazoparib), and more inhibitors are currently in clinical trials.

[0003] Niraparib is currently approved as a PARP inhibitor for the maintenance treatment of adult patients with advanced or recurrent epithelial ovarian, fallopian tube, or primary peritoneal cancer who have had a complete or partial response to initial platinum-based chemotherapy; and for the treatment of adult patients with advanced ovarian, fallopian tube, or primary peritoneal cancer that is associated with homologous recombination deficiency (HRD)-positive status, defined by either a deleterious or suspected deleterious BRCA mutation or genomic instability, who have been treated with three or more prior chemotherapy regimens and who have responded more than six months since their last platinum-based chemotherapy.

[0004] Transaminases, also known as aminotransferases, catalyze the transfer of an amino group, an electron pair, and a proton from the primary amine of an amino donor substrate to the carbonyl group (i.e., keto group) of an amino acceptor molecule (Shin et al., 2001, Biosci. Biotechnol. Biochem. 65:1782-1788).

[0005] The stereoselectivity of transaminases in converting ketones to the corresponding amines makes these enzymes useful in the asymmetric synthesis of optically pure amines from the corresponding keto compounds (Hohne et al., Chem. Cat. Chem. 1(1):42-51). Transaminases may also be applied to the chiral resolution of racemic amines by exploiting their ability to perform the reverse reaction, i.e., preferentially convert one enantiomer to the corresponding ketone, in a stereospecific manner, thereby yielding a mixture enriched in the other enantiomer (Koselewski et al., 2009, Org. Lett. 1 1(21):4810-2). The wild-type transaminase from Arthrobacter sp. KNK168 is an R-selective enzyme that produces R-amines from several substrates. Several studies have shown that engineered transaminase polypeptides derived from the native transaminase of Arthrobacter sp. KNK168 have been adapted to have high stability to temperature and / or organic solvents and enzymatic activity toward structurally distinct amino acceptor molecules (Savile et al., 2010, Science 329(5989):305-9).

[0006] Various manufacturing methods for preparing niraparib have been developed over the years, but there is a need for improved transaminase biocatalysts that can be used to prepare amine intermediate compounds useful in the production of niraparib, as well as new methods using biocatalysts that are simple, cost-effective, non-toxic, and commercially viable.

[0007] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in computer-readable form in XML format, which is incorporated by reference in its entirety. The XML file, created on June 27, 2023, is named "TES00050WO01.xml" and is 32,471 bytes in size. Summary of the Invention

[0008] The present invention provides modified transaminase polypeptides having transaminase activity, polynucleotides encoding the polypeptides, methods for making the polypeptides, and methods for using the polypeptides to prepare amine intermediate compounds useful in the production of niraparib.

[0009] In one aspect of the present invention, modified transaminase polypeptides are provided that contain at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 amino acid substitutions and / or modifications relative to the transaminase variant set forth in SEQ ID NO:2.

[0010] The present invention provides a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, or 95% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises the characteristic that X124 is isoleucine (I).

[0011] In a further aspect of the present invention, polynucleotides encoding the modified transaminase polypeptides of the present invention are provided. Compositions, vectors, and host cells comprising the polynucleotides are also provided.

[0012] In a further aspect, the present invention provides an asymmetric compound of formula I:

[0013] [ka]

[0014] (In the formula: R 1 is a leaving group, a halogen, a protected amino group, NO, or OH or a protected form thereof; R 2 is H; R 3 -COOR 5 , -CH2R 6 or a protected aldehyde; or R 2 and R 3 is combined

[0015] [ka]

[0016] is formed; R 4 is H or an amine protecting group; R 5 is C 1~6 Alkyl, C 3~6 Cycloalkyl, C 4~10 heterocyclyl, aryl, or heteroaryl; R 6 is a leaving group or OH or a protected form thereof) 1. A method for preparing a compound of formula II:

[0017] [ka]

[0018] (In the formula, R 1 ' is a leaving group, a halogen, a protected amino group, NO2, or OH or a protected form thereof; R 2 ' is an aldehyde or aldehyde equivalent; R 3 '-COOR5 , -CH2R 6 or a protected aldehyde; or R 2 ' and R 3 ' is a combination

[0019] [ka]

[0020] (where, * represents the attachment point) is formed) in the presence of a coenzyme and an amino donor with the modified transaminase polypeptide of any one of the embodiments disclosed herein.

[0021] In another embodiment, the present invention provides a compound of formula III, or a salt thereof:

[0022] [ka]

[0023] (In the formula, R 7 is H, an amine protecting group, C1-C6 alkyl, or tert-butyl; R 8 is H; R 9 is H, C1-C6 alkyl, or tert-butyl) A method for preparing a compound of formula I, or a salt thereof, wherein the method is produced by the method according to any one of the embodiments disclosed herein.

[0024] [ka]

[0025] (In the formula, R 1 is a halogen; R 2 is H; R 3 is -CH2R 6 and; R 4 is H or an amine protecting group; R 6 is OH or a protected form thereof) Compound IV,

[0026] [ka]

[0027] (In the formula, R 10 is H; R 11 is C1-C6 alkyl or tert-butyl) The method includes contacting the

[0028] In one aspect, the present invention provides niraparibut tosylate monohydrate of formula V:

[0029] [ka]

[0030] a method for preparing a compound of formula III, or a salt thereof, produced by the method according to any one of the embodiments disclosed herein;

[0031] [ka]

[0032] (In the formula, R 7 is H or an amine protecting group; R 8 is H; R 9 is C1-C6 alkyl or tert-butyl) with an acid. DETAILED DESCRIPTION OF THE INVENTION

[0033] Detailed Description of the Invention The present invention provides highly efficient transaminase variants that have significantly improved enzyme activity and maintain high product stereoselectivity. Such efficient transaminase variants are achieved by modifying transaminase variants. The present invention provides transaminase variants. The transaminase variants are enzymes that have transaminase activity and have at least one substitution and / or modification with respect to SEQ ID NO:2. The present invention further provides transaminase variants that contain multiple (two or more) amino acid substitutions and / or modifications with respect to SEQ ID NO:2.

[0034] In one embodiment, a modified transaminase polypeptide or functional fragment thereof is provided, comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:2.

[0035] The present invention provides a transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein X124 is I.

[0036] The present invention provides a transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein X124 is I, and the peptide has improved enzymatic properties compared to SEQ ID NO:2.

[0037] The present invention provides a transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein X124 is I, and the polypeptide has increased enzymatic activity relative to SEQ ID NO:2.

[0038] The present invention provides a transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein X124 is I and at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 is substituted with an amino acid different from the corresponding amino acid identified in SEQ ID NO:2, wherein the polypeptide has improved enzymatic properties compared to SEQ ID NO:2.

[0039] The present invention provides a transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein X124 is I, and the polypeptide has improved enzymatic properties compared to SEQ ID NO:2.

[0040] The present invention provides a transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein X124 is I and at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 is substituted with an amino acid different from the amino acid found at the corresponding amino acid position in SEQ ID NO:2, wherein the polypeptide has improved enzymatic properties compared to SEQ ID NO:2.

[0041] In one embodiment, the improved enzymatic property is increased enzymatic activity. In one embodiment, the transaminase enzyme of the invention exhibits increased activity that is at least 1.5-fold improved over the enzymatic activity of the transaminase enzyme of SEQ ID NO: 2. In one embodiment, the transaminase enzyme of the invention exhibits increased activity that is at least 2-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, or 50-fold improved over the enzymatic activity of SEQ ID NO: 2, and has a stereoselectivity of greater than 95, 96, 97, 98, or 99% ee.

[0042] In one embodiment, there is provided a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:2, and wherein X124 is I.

[0043] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X61 is C, N, or M; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R.

[0044] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X61 is C, N, or M; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R; and wherein X124 is I. In another embodiment, provided herein is a modified transaminase polypeptide or a functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X48 is substituted with R; X61 is C; X69 is C; X94 is substituted with C; X196 is substituted with R; and X297 is substituted with R. In another embodiment, provided herein is a modified transaminase polypeptide or a functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X48 is substituted with R; X61 is N; X69 is C or T; X94 is substituted with C; X196 is substituted with R; and X297 is substituted with R. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X48 is substituted with R; X61 is M; X69 is C; X94 is substituted with C; X196 is substituted with R; and X297 is substituted with R.In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X48 is substituted with R; X61 is M; X69 is C; X94 is substituted with C; X137 is substituted with T; X196 is substituted with R; and X297 is substituted with R.

[0045] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X61C substitution, an X61N substitution, or an X61M substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69C substitution or an X69T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X137T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X140Q substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X199T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X202C substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X3H substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X5I substitution.

[0046] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 amino acids have been substituted with amino acids different from those found at the corresponding amino acid positions in SEQ ID NO: 2. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises X61M and X69T substitutions. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises X61M, X69T, and X137T substitutions. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises an X61M substitution, an X69T substitution, an X137T substitution, an X140Q substitution, an X199T substitution, and an X202C substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises an X3H substitution, an X5I substitution, an X61M substitution, an X69T substitution, an X97E substitution, an X137T substitution, and an X269C substitution.

[0047] In one embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises an X61M substitution and at least one other substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises an X61M substitution and at least one amino acid substitution selected from the group consisting of: X3 substituted with H; X5 substituted with I; X48 substituted with R; X69 is C or T; X94 substituted with C; X97 substituted with E; X137 substituted with T; X140 substituted with Q; X196 substituted with R; X199 substituted with T; X202 substituted with C; X269 substituted with C or V; and X297 substituted with R.

[0048] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises a X61M substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X69T, X97E, X137T, X140Q, X199T, X202C, and X269C. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises a X61M substitution and a X3H substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises a X61M substitution and a X5I substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X61M substitution and an X97E substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X61M substitution and an X137T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X61M substitution and an X140Q substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X61M substitution and an X199T substitution.In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises a X61M substitution and a X202C substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises a X61M substitution and a X269C substitution.

[0049] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69T substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X61N, X61C, X61M, X97E, X137T, X140Q, X199T, X202C, and X269C. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69T substitution and an X3H substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69T substitution and an X5I substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69T substitution and an X97E substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69T substitution and an X137T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69T substitution and an X140Q substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69T substitution and an X199T substitution.In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises a X69T substitution and a X202C substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises a X69T substitution and a X269C substitution.

[0050] In another embodiment, provided herein is an engineered transaminase polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to an amino acid sequence set forth in an amino acid sequence selected from the group consisting of SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:16, SEQ ID NO:18, and SEQ ID NO:20. Those skilled in the art will understand that the first M in SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:16, SEQ ID NO:18, or SEQ ID NO:20 may be modified or deleted without affecting the function of the engineered transaminase polypeptide. In another embodiment, the engineered transaminase polypeptide has the formula MH x (SEQ ID NO: 21), where M is a methionine residue and H x represents a chain of x histidine residues, where x is a whole integer from 0 to 20. In certain embodiments, x is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20.

[0051] In a further aspect, there is provided a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:4, wherein the amino acid sequence comprises the characteristic that at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 is substituted with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:4.

[0052] In one embodiment, a modified transaminase polypeptide or functional fragment thereof is provided, comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:4, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:4, and wherein X124 is I.

[0053] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:4, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X61 is N or M; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R.

[0054] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X61 is N or M; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R; and wherein X124 is I. In another embodiment, provided herein is a modified transaminase polypeptide or a functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X48 is substituted with R; X61 is N; X69 is C or T; X94 is substituted with C; X196 is substituted with R; and X297 is substituted with R. In another embodiment, provided herein is a modified transaminase polypeptide or a functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X48 is substituted with R; X61 is M; X69 is C; X94 is substituted with C; X196 is substituted with R; and X297 is substituted with R.In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X48 is substituted with R; X61 is M; X69 is C; X94 is substituted with C; X137 is substituted with T; X196 is substituted with R; and X297 is substituted with R.

[0055] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X61N substitution or an X61M substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X69N substitution or an X69T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X137T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X140Q substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X199T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X202C substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X3H substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X5I substitution.

[0056] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein at least two, at least three, at least four, at least five, at least ten, or at least twenty amino acids are substituted with amino acids different from the amino acids identified at the corresponding amino acid positions in SEQ ID NO: 4. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises X61M and X69T substitutions. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises X61M, X69T, and X137T substitutions. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X61M substitution, an X69T substitution, an X137T substitution, an X140Q substitution, an X199T substitution, and an X202C substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X3H substitution, an X5I substitution, an X61M substitution, an X69T substitution, an X97E substitution, an X137T substitution, and an X269C substitution.

[0057] In one embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X61M substitution and at least one other substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X61M substitution and at least one amino acid substitution selected from the group consisting of: X3 substituted with H; X5 substituted with I; X48 substituted with R; X69 is C or T; X94 substituted with C; X97 substituted with E; X137 substituted with T; X140 substituted with Q; X196 substituted with R; X199 substituted with T; X202 substituted with C; X269 substituted with C or V, and X297 substituted with R.

[0058] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:4, wherein the amino acid sequence comprises a X61M substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X69T, X97E, X137T, X140Q, X199T, X202C, and X269C. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:4, wherein the amino acid sequence comprises a X61M substitution and a X3H substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:4, wherein the amino acid sequence comprises a X61M substitution and a X5I substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X61M substitution and an X97E substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X61M substitution and an X137T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X61M substitution and an X140Q substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X61M substitution and an X199T substitution.In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises a X61M substitution and a X202C substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises a X61M substitution and a X269C substitution.

[0059] In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:4, wherein the amino acid sequence comprises an X69T substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X61N, X61C, X61M, X97E, X137T, X140Q, X199T, X202C, and X269C. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:4, wherein the amino acid sequence comprises an X69T substitution and an X3H substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:4, wherein the amino acid sequence comprises an X69T substitution and an X5I substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X69T substitution and an X97E substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X69T substitution and an X137T substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X69T substitution and an X140Q substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises an X69T substitution and an X199T substitution.In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises a X69T substitution and a X202C substitution. In another embodiment, provided herein is a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 4, wherein the amino acid sequence comprises a X69T substitution and a X269C substitution.

[0060] In some embodiments, a single point mutation in a modified transaminase polypeptide can increase enzymatic activity by about 1.1-fold to about 10-fold, about 1.2-fold to about 7-fold, about 1.3-fold to about 5-fold, about 1.5-fold to about 2.5-fold, or about 1.5-fold to about 6-fold, or at least 2-fold, compared to the activity of the modified transaminase of SEQ ID NO: 2. The positive effect of a single point mutation can be multiplied when combined as multiple point mutations within a modified transaminase polypeptide. In some embodiments, a modified transaminase polypeptide can contain at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid substitutions. In some embodiments, the activity of the modified transaminase is increased while the stereoselectivity remains above 95, 96, 97, 98, or 99% ee. In some embodiments, the use of the modified transaminase polypeptides disclosed herein increases transaminase activity by at least about 2-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 45-fold, at least about 50-fold, or at least about 55-fold compared to the activity of the modified transaminase of SEQ ID NO: 2. In some embodiments, the use of the modified transaminase polypeptides disclosed herein increases transaminase activity by at least about 2-fold, at least about 5-fold, at least about 10-fold, at least about 15-fold, at least about 20-fold, at least about 25-fold, at least about 30-fold, at least about 35-fold, at least about 40-fold, at least about 45-fold, at least about 50-fold, or at least about 55-fold compared to the activity of the modified transaminase of SEQ ID NO: 2. In some embodiments, the control transaminase is a modified transaminase comprising the amino acid sequence of SEQ ID NO:2.

[0061] In one embodiment, a modified transaminase polypeptide or functional fragment thereof is provided, comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids F3, H5, P48, C61, V69, I94, I97, E137, I140, I196, I199, T202, P269, and S297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:2, and wherein X124 is I.

[0062] In some embodiments, the modified transaminase polypeptide of any one of the preceding embodiments is characterized in that X124 is I.

[0063] In another embodiment, provided herein is a modified transaminase polypeptide comprising an amino acid sequence set forth in an amino acid sequence selected from the group consisting of SEQ ID NO: 4, SEQ ID NO: 6, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, and SEQ ID NO: 20. One of skill in the art will understand that the first M in SEQ ID NO: 4, 6, 8, 10, 14, 16, 18, or 20 can be modified or deleted without affecting the function of the modified transaminase polypeptide.

[0064] kit In some embodiments, provided herein are compositions comprising any one of the modified transaminase polypeptides described herein.

[0065] In some embodiments, provided herein are kits comprising a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:2, wherein X124 is I; a coenzyme; and an amino donor. In some embodiments, the coenzyme is pyridoxal phosphate (PLP). In some embodiments, the amino donor is isopropylamine.

[0066] In some embodiments, provided herein are kits comprising a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X61 is C, N, or M; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R; a coenzyme; and an amino donor. In some embodiments, the coenzyme is pyridoxal phosphate (PLP). In some embodiments, the amino donor is isopropylamine.

[0067] In some embodiments, provided herein is a kit comprising a modified transaminase polypeptide comprising the amino acid sequence of SEQ ID NO: 2, or a functional fragment thereof, wherein the amino acid sequence comprises an X69T substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X61N, X61C, X61M, X97E, X137T, X140Q, X199T, X202C, and X269C; a coenzyme; and an amino donor.

[0068] In some embodiments, provided herein are kits comprising a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:4, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:4, and wherein X124 is I; a coenzyme; and an amino acid donor.

[0069] In some embodiments, provided herein are kits comprising a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:4, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X61 is N or M; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R; a coenzyme and an amino donor.

[0070] In some embodiments, provided herein are kits that include a modified transaminase polypeptide described herein, a coenzyme, and an amino donor.

[0071] nucleic acid In another aspect, provided herein are polynucleotides encoding the modified transaminase polypeptides of the present invention. In some embodiments, the polynucleotides encode a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:2, and wherein X124 is I.

[0072] In some embodiments, the polynucleotide encodes a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X61 is C, N, or M; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R; and X124 is I.

[0073] In some embodiments, the polynucleotide encodes a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X61M substitution and at least one amino acid substitution selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R.

[0074] In some embodiments, the polynucleotide encodes a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X61M substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X69T, X97E, X137T, X140Q, X199T, X202C, and X269C.

[0075] In some embodiments, the polynucleotide encodes a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69T substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X61N, X61C, X61M, X97E, X137T, X140Q, X199T, X202C, and X269C.

[0076] In some embodiments, the polynucleotide encodes a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X61M substitution and an X69T substitution.

[0077] In some embodiments, the polynucleotide encodes a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein at least two, at least three, at least four, at least five, at least ten, or at least twenty amino acids have been substituted with amino acids different from those identified at the corresponding amino acid positions in SEQ ID NO:2.

[0078] In some embodiments, the polynucleotide comprises a polynucleotide sequence having at least 80%, 85%, 90%, 95%, 99%, or 100% sequence identity to a polynucleotide sequence set forth in a polynucleotide sequence selected from the group consisting of SEQ ID NO:3, SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:15, SEQ ID NO:17, or SEQ ID NO:19.

[0079] In some embodiments, provided herein is a polynucleotide encoding any one of the foregoing embodiments of the modified transaminase polypeptide described herein.

[0080] In some embodiments, vectors comprising polynucleotides encoding the modified transaminase polypeptides of the invention are provided herein. In some embodiments, the vectors comprise a polynucleotide or functional fragment thereof encoding an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:2, and wherein X124 is I.

[0081] In some embodiments, provided herein is a vector comprising a polynucleotide encoding a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X61 is C, N, or M; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R.

[0082] In some embodiments, provided herein is a vector comprising a polynucleotide encoding a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X61M substitution and at least one amino acid substitution selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R.

[0083] In some embodiments, provided herein is a vector comprising a polynucleotide encoding a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises an X61M substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X69T, X97E, X137T, X140Q, X199T, X202C, and X269C.

[0084] In some embodiments, provided herein is a vector comprising a polynucleotide encoding a modified transaminase polypeptide or functional fragment thereof comprising the amino acid sequence of SEQ ID NO: 2, wherein the amino acid sequence comprises an X69T substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X61N, X61C, X61M, X97E, X137T, X140Q, X199T, X202C, and X269C.

[0085] In some embodiments, provided herein is a vector comprising a polynucleotide encoding a modified transaminase polypeptide of any one of the preceding embodiments described herein.

[0086] host cell In another aspect, provided herein is a host cell comprising a modified transaminase polypeptide, wherein the modified transaminase polypeptide comprises at least one amino acid substitution. In another embodiment, the host cell is Escherichia coli (E. coli). In some embodiments, the host cell comprises two or more modified transaminase polypeptides.

[0087] In another aspect, provided herein are host cells comprising a polynucleotide provided herein (e.g., a polynucleotide encoding a modified transaminase polypeptide provided herein). In some embodiments, the host cell comprises two or more polynucleotides provided herein (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleic acids).

[0088] In some embodiments, provided herein are host cells comprising a vector comprising a polynucleotide encoding a modified transaminase polypeptide of the invention. In some embodiments, the host cell comprises a vector comprising a polynucleotide encoding a modified transaminase polypeptide or functional fragment thereof, the modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:2, and wherein X124 is I.

[0089] In some embodiments, provided herein is a host cell comprising a vector containing a polynucleotide encoding a modified transaminase polypeptide or functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X61 is C, N, or M; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R.

[0090] In some embodiments, provided herein is a host cell comprising a vector comprising a polynucleotide encoding a modified transaminase polypeptide comprising the amino acid sequence of SEQ ID NO:2, or a functional fragment thereof, wherein the amino acid sequence comprises an X61M substitution and at least one amino acid substitution selected from the group consisting of: X3 is substituted with H; X5 is substituted with I; X48 is substituted with R; X69 is C or T; X94 is substituted with C; X97 is substituted with E; X137 is substituted with T; X140 is substituted with Q; X196 is substituted with R; X199 is substituted with T; X202 is substituted with C; X269 is substituted with C or V; and X297 is substituted with R.

[0091] In some embodiments, provided herein is a host cell comprising a vector comprising a polynucleotide encoding a modified transaminase polypeptide comprising the amino acid sequence of SEQ ID NO:2, or a functional fragment thereof, wherein the amino acid sequence comprises an X61M substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X69T, X97E, X137T, X140Q, X199T, X202C, and X269C.

[0092] In some embodiments, provided herein is a host cell comprising a vector comprising a polynucleotide encoding a modified transaminase polypeptide or a functional fragment thereof comprising the amino acid sequence of SEQ ID NO:2, wherein the amino acid sequence comprises an X69T substitution and at least one other amino acid substitution selected from the group consisting of X3H, X5I, X61N, X61C, X61M, X97E, X137T, X140Q, X199T, X202C, and X269C.

[0093] In some embodiments, provided herein is a host cell comprising a vector comprising a polynucleotide encoding a modified transaminase polypeptide of any one of the preceding embodiments described herein.

[0094] Synthesis method In one aspect, the present invention provides an asymmetric compound of formula I:

[0095] [ka]

[0096] (In the formula: R 1 is a leaving group, a halogen, a protected amino group, —NO2, or —OH or a protected form thereof; R 2 is H; R 3 -COOR 5 , -CH2R 6or a protected aldehyde; or R 2 and R 3 is combined

[0097] [ka]

[0098] is formed; R 4 is H or an amine protecting group; R 5 is C 1~6 Alkyl, C 3~10 Cycloalkyl, C 4~10 heterocyclyl, aryl, or heteroaryl; R 6 is a leaving group or -OH or a protected form thereof) To provide an efficient method for preparing

[0099] In one embodiment, R 1 is a leaving group. In one embodiment, R 1 is halogen. In one embodiment, R 1 is Br.

[0100] In another embodiment, R 2 and R 3 is combined

[0101] [ka]

[0102] (In the formula, R 4 is H) is formed.

[0103] In a further embodiment, R 2 and R 3 is combined

[0104] [ka]

[0105] (In the formula, R 4 is H) is formed.

[0106] In one embodiment, R 1 is Br and R 2 and R 3 is combined

[0107] [ka]

[0108] (In the formula, R 4 is H) is formed.

[0109] In another embodiment, R 2 is H and R 3 is CH2R 6 and R 4 is H and R 6 is OH.

[0110] In a further aspect, the present invention provides a method for preparing an asymmetric compound of formula I, comprising reacting any one of the modified transaminase polypeptides of the present invention with a compound of formula II:

[0111] [ka]

[0112] (In the formula, R 1 ' is a leaving group, a halogen, a protected amino group, -NO2, or -OH or a protected form thereof; R 2 ' is an aldehyde or aldehyde equivalent; R 3 '-COOR 5 , -CH2R 6or a protected aldehyde; or R 2 ' and R 3 ' is a combination

[0113] [ka]

[0114] (where, * represents the attachment point) (forming The method includes contacting the

[0115] In one embodiment, R 1 ' is a leaving group. In one embodiment, R 1 In one embodiment, R 1 ' is Br.

[0116] In one embodiment, R 2 ' and R 3 ' is combined

[0117] [ka]

[0118] (where, * represents the attachment point) is formed.

[0119] Thus, in one embodiment, an asymmetric compound of formula I:

[0120] [ka]

[0121] (In the formula, R 1 is a halogen; R 2 is H; R 3 is -CH2R6 and; R 4 is H or an amine protecting group; R 6 is OH or a protected form thereof) 1. A method for preparing a compound of formula IIa:

[0122] [ka]

[0123] with a modified transaminase polypeptide of any one of the preceding embodiments disclosed herein in the presence of a coenzyme and an amino donor to produce a compound of formula Ia:

[0124] [ka]

[0125] and optionally protecting the NH and OH groups of formula Ia to provide a compound of formula Ib:

[0126] [ka]

[0127] where PG1 is an amine protecting group and PG2 is an oxygen protecting group. In one embodiment, the method according to any one of the embodiments disclosed herein may be a method for preparing niraparib, comprising niraparib tosylate monohydrate.

[0128] In one embodiment, an asymmetric compound of formula I:

[0129] [ka]

[0130] (In the formula: R 1 is Br; R 2 is H; R 3 Ha-CH2R 6 and; R 4 is H or an amine protecting group; R 6 is OH or a protected form thereof) 1. A method for preparing a compound of formula IIa:

[0131] [ka]

[0132] with any one of the modified transaminase polypeptides of the embodiments disclosed herein in the presence of a coenzyme and an amino donor to provide a compound of Formula Ia.

[0133] [ka]

[0134] In some embodiments, the method includes optionally protecting the NH group of formula Ia to produce a compound of formula Ic:

[0135] [ka]

[0136] where PG1 is an amine protecting group. This optional step includes providing R 4 is an amine protecting group.

[0137] In some embodiments, the method further optionally comprises protecting the OH group of formula Ic to produce a compound of formula Ib:

[0138] [ka]

[0139] where PG2 is a hydroxyl protecting group. This optional step includes providing R 6 is OH in its protected form, and R 4 is an amine protecting group.

[0140] In one embodiment, the amine protecting group is selected from the group consisting of formyl, acetyl (Ac), trifluoroacetyl, benzyl (Bn), benzoyl (Bz), carbamate, benzyloxycarbonyl ("CBZ"), p-methoxybenzylcarbonyl (Moz or MeOZ), tert-butoxycarbonyl ("Boc"), trimethylsilyl ("TMS"), 2-trimethylsilyl-ethanesulfonyl ("SES"), trityl and substituted trityl groups, allyloxycarbonyl, 9-fluorenylmethyloxycarbonyl ("FMOC"), nitro-veratryloxycarbonyl ("NVOC"), p-methoxybenzyl (PMB), tosyl (Ts), 3,4-dimethoxybenzyl (DMPM), p-methoxyphenyl (PMP), 2-naphthylmethyl ether (Nap), and trichloroethyl chloroformate (Troc). In one embodiment, the amine protecting group is tert-butoxycarbonyl (“Boc”).

[0141] In one embodiment, the hydroxyl protecting group is selected from the group consisting of methyl ester, ethyl ester, acetate, propionate group, glycol ester, benzyl, trityl ether, alkyl ether, tetrahydropyranyl ether, trialkylsilyl ether, TMS, TIPS, mesylate, and allyl ether. In one embodiment, the hydroxyl protecting group is mesylate.

[0142] In one embodiment, the compound of formula II is

[0143] [ka]

[0144] is selected from the group consisting of: In one embodiment, the compound of formula I is

[0145] [ka]

[0146] is selected from the group consisting of: In one embodiment, the present invention provides an asymmetric compound

[0147] [ka]

[0148] a method for preparing a modified transaminase polypeptide of the present invention, comprising:

[0149] [ka]

[0150] in the presence of a coenzyme and an amino donor.

[0151] In one embodiment, the present invention provides an asymmetric compound

[0152] [ka]

[0153] a method for preparing a modified transaminase polypeptide of the present invention, comprising:

[0154] [ka]

[0155] in the presence of a coenzyme and an amino donor.

[0156] In one embodiment, the present invention provides an asymmetric compound

[0157] [ka]

[0158] a method for preparing a modified transaminase polypeptide of the present invention, comprising:

[0159] [ka]

[0160] in the presence of a coenzyme and an amino donor.

[0161] In one embodiment, the present invention provides an asymmetric compound

[0162] [ka]

[0163] a method for preparing a modified transaminase polypeptide of the present invention, comprising:

[0164] [ka]

[0165] in the presence of a coenzyme and an amino donor.

[0166] In one embodiment, the method of any of the preceding embodiments provides a compound of Formula I having an enantiomeric excess (ee) of at least about 95% ee, at least about 96% ee, at least about 97% ee, at least about 98% ee, or at least about 99.9% ee. In another embodiment, the transaminase-catalyzed method of a compound of Formula II described provides a compound of Formula I having an enantiomeric excess of at least 95%. In a further embodiment, the transaminase-catalyzed method of a compound of Formula II described provides a compound of Formula I having an enantiomeric excess of at least 99%. In one embodiment, the compound of Formula I is

[0167] [ka]

[0168] is selected from the group consisting of: In one embodiment, the compound of formula I is

[0169] [ka]

[0170] is selected from the group consisting of: In one embodiment, the compound of formula I is

[0171] [ka]

[0172] is. In one embodiment, the compound of formula I is

[0173] [ka]

[0174] is. In one embodiment of any of the foregoing methods disclosed herein, the modified transaminase polypeptide comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the amino acid sequence set forth in SEQ ID NO: 2, or a functional fragment thereof, wherein at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 is substituted with an amino acid different from the amino acid found at the corresponding amino acid position in SEQ ID NO: 2. In some embodiments, the modified transaminase polypeptide is selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 4, 6, 8, 10, 14, 16, 18, and 20. In one embodiment, the modified transaminase polypeptide is selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 16, 18, and 20. In some embodiments, the modified transaminase polypeptide comprises any one of the modified transaminase polypeptides of the preceding embodiments.

[0175] In one embodiment of any of the aforementioned methods disclosed herein, the coenzyme is pyridoxal phosphate (PLP). In some embodiments, the amino donor is isopropylamine.

[0176] In one aspect, the present invention provides a compound of formula III, or a salt thereof:

[0177] [ka]

[0178] (In the formula, R 7 is H, an amine protecting group, alkyl, C1-C4 alkyl, or tert-butyl; R 8 is H; R 9 is H, alkyl, C1-C4 alkyl, or tert-butyl) 1. A method for preparing a compound of formula I, or a salt thereof, produced by the method according to any one of the embodiments disclosed herein.

[0179] [ka]

[0180] (In the formula, R 1 is a halogen; R 2 is H; R 3 is -CH2R 6 and; R 4 is H or an amine protecting group; R 6 is OH or a protected form thereof) Compound IV,

[0181] [ka]

[0182] (In the formula, R 10 is H; R 11 is alkyl, C1-C4 alkyl, or tert-butyl) The method includes contacting the

[0183] In one embodiment, R 7is an amine protecting group. In one embodiment, the amine protecting group comprises tert-butyloxycarbonyl (Boc), 9-fluorenylmethyloxycarbonyl (Fmoc), carboxybenzyl group (Cbz), p-methoxybenzylcarbonyl (Moz), acetyl (Ac), benzoyl (Bz), p-methoxybenzyl (PMB), 3,4-dimethoxybenzyl (DMPM), p-methoxyphenyl (PMP), 2-naphthylmethyl ether (Nap), tosyl (Ts), or trichloroethyl chloroformate (Troc). In one embodiment, the amine protecting group is tert-butyloxycarbonyl group (Boc).

[0184] In one embodiment, R 9 is tert-butyl ( t Bu). In another embodiment, R 11 is tert-butyl ( t Bu).

[0185] In one embodiment, the contacting step comprises:

[0186] [ka] The reaction is carried out in the presence of

[0187] In another embodiment, the contacting step is in the presence of at least one agent selected from the group consisting of toluene, potassium tert-butoxide, tripotassium phosphate (KPO), copper(I) bromide, ammonium hydroxide, and heptane, or any combination thereof. In yet another embodiment, the contacting step is in the presence of a solvent. Consistent with these embodiments, the solvent includes, but is not limited to, toluene, N,N-dimethylformamide (DMF), t-butanol, dimethoxyethane (DME), acetonitrile, dichloromethane (DCM), tetrahydrofuran (THF), 2-methyltetrahydrofuran (ME-THF), isopropyl alcohol, methanol, ethanol, or any combination thereof. In some embodiments, the solvent includes toluene.

[0188] In one embodiment, the contacting step is carried out in the presence of an acid. In one embodiment, the acid comprises formic acid, acetic acid, propionic acid, butyric acid, valeric acid, caproic acid, oxalic acid, lactic acid, malic acid, citric acid, benzoic acid, carbonic acid, uric acid, taurine, p-toluenesulfonic acid, trifluoromethanesulfonic acid, aminomethylphosphonic acid, trifluoroacetic acid (TFA), phosphonic acid, sulfuric acid, nitric acid, phosphoric acid, hydrochloric acid, ethanesulfonic acid (ESA), methanesulfonic acid (MsOH), or any combination thereof. In another embodiment, the acid is methanesulfonic acid (MsOH).

[0189] In one embodiment, the compound of formula I, or a salt thereof, is

[0190] [ka]

[0191] It has the following structure. In one embodiment, the compound of formula IV, or a salt thereof, is

[0192] [ka]

[0193] It has the following structure. In one embodiment, the compound of formula III, or a salt thereof, is

[0194] [ka]

[0195] It has the following structure. In one aspect, the present invention provides niraparibut tosylate monohydrate of formula V:

[0196] [ka]

[0197] a method for preparing a compound of formula III, or a salt thereof, produced by the method according to any one of the embodiments disclosed herein;

[0198] [ka]

[0199] (In the formula, R 7 is H or an amine protecting group; R 8 is H; R 9 is C1-C4 alkyl or tert-butyl) with an acid.

[0200] In one embodiment, R 7is an amine protecting group. In one embodiment, the amine protecting group is tert-butyloxycarbonyl (Boc), 9-fluorenylmethyloxycarbonyl (Fmoc), carboxybenzyl (Cbz), p-methoxybenzylcarbonyl (Moz), acetyl (Ac), benzoyl (Bz), p-methoxybenzyl (PMB), 3,4-dimethoxybenzyl (DMPM), p-methoxyphenyl (PMP), 2-naphthylmethyl ether (Nap), tosyl (Ts), or trichloroethyl chloroformate (Troc). In one embodiment, the amine protecting group is tert-butyloxycarbonyl (Boc). In one embodiment, R 9 is tert-butyl.

[0201] In one embodiment, the acid is at least one acid selected from the group consisting of methanesulfonic acid (MsOH), formic acid, acetic acid, propionic acid, butyric acid, valeric acid, caproic acid, oxalic acid, lactic acid, malic acid, citric acid, benzoic acid, carbonic acid, uric acid, tauric acid, p-toluenesulfonic acid, trifluoromethanesulfonic acid, aminomethylphosphonic acid, trifluoroacetic acid (TFA), phosphonic acid, sulfuric acid, nitric acid, phosphoric acid, hydrochloric acid, and ethanesulfonic acid (ESA), or any combination thereof. In one embodiment, the acid is methanesulfonic acid (MsOH). In one embodiment, the acid is p-toluenesulfonic acid.

[0202] In one embodiment, the contacting step is in the presence of a solvent. In one embodiment, the solvent comprises N,N-dimethylformamide (DMF), t-butanol, dimethoxyethane (DME), acetonitrile, dichloromethane (DCM), tetrahydrofuran (THF), 2-methyltetrahydrofuran (ME-THF), isopropyl alcohol, methanol, ethanol, or any combination thereof. In another embodiment, the solvent comprises ME-THF.

[0203] In one embodiment, the method of any one of the preceding embodiments comprises reacting a compound of formula V:

[0204] [ka]

[0205] wherein the method comprises contacting a compound of formula III with an acid. In one embodiment, contacting a compound of formula III with an acid provides a compound of formula VI:

[0206] [ka]

[0207] (In the formula, R 9 is tert-butyl) Further provide.

[0208] In one embodiment, niraparibut tosylate monohydrate of formula VII:

[0209] [ka]

[0210] a method for preparing a compound of formula V, or a salt thereof, produced by the method according to any one of the preceding embodiments;

[0211] [ka]

[0212] with methanesulfonic acid or paratoluenesulfonic acid.

[0213] In one embodiment, the contacting is in the presence of a solvent. In one embodiment, the solvent is water.

[0214] In one aspect, the present invention provides i) niraparibut tosylate monohydrate of formula VII:

[0215] [ka]

[0216] and ii) less than 0.1% by weight of a compound of formula VI, or a salt thereof;

[0217] [ka]

[0218] A composition comprising: In one embodiment, the concentration of the compound of formula VI present in the composition is less than 0.09%, less than 0.08%, less than 0.07%, less than 0.06%, less than 0.05%, less than 0.04%, less than 0.03%, less than 0.02%, or less than 0.01% by weight.

[0219] In one embodiment, the compound of formula VI is present after contacting the compound of formula III with an acid according to the method for preparing niraparibut tosylate monohydrate of formula V disclosed herein.

[0220] In one embodiment, the acid is a strong acid. In one embodiment, the acid is at least one acid selected from the group consisting of methanesulfonic acid (MsOH), formic acid, acetic acid, propionic acid, butyric acid, valeric acid, caproic acid, oxalic acid, lactic acid, malic acid, citric acid, benzoic acid, carbonic acid, uric acid, tauric acid, p-toluenesulfonic acid, trifluoromethanesulfonic acid, aminomethylphosphonic acid, trifluoroacetic acid (TFA), phosphonic acid, sulfuric acid, nitric acid, phosphoric acid, hydrochloric acid, and ethanesulfonic acid (ESA), or any combination thereof. In one embodiment, the acid is methanesulfonic acid (MsOH).

[0221] In some embodiments, niraparibut tosylate monohydrate is added to a solvent. In some embodiments, niraparibut tosylate monohydrate is added to a solvent to form a crystalline form of niraparibut tosylate monohydrate. In some embodiments, the solvent is dimethyl sulfoxide (DMSO). In some embodiments, the method of any one of the preceding embodiments further comprises wet-milling the niraparibut tosylate monohydrate disclosed herein. In some embodiments, the method further comprises annealing the niraparibut tosylate monohydrate using one or more temperature cycles.

[0222] In some embodiments, the solvent comprises DMSO. In some embodiments, the method comprises contacting niraparibut tosylate monohydrate with DMSO and water. In some embodiments, a water to solvent ratio (v / v) of about 200:1 to about 1:200 is used in the contact. In some embodiments, the water to solvent ratio (v / v) is about 200:1 to about 1:200, e.g., about 200:1 to about 100:1, about 200:1 to about 10:1, about 200:1 to about 5:1, about 200:1 to about 2:1, about 200:1 to about 1:1, about 200:1 to about 1:2, about 200:1 to about 1:5, about 200:1 to about 1:10, about 200:1 to about 1:100, or about 100:1 to about 10:1. , about 100:1 to about 5:1, about 100:1 to about 2:1, about 100:1 to about 1:1, about 100:1 to about 1:2, about 100:1 to about 1:5, about 100:1 to about 1:10, about 100:1 to about 1:100, about 100:1 to about 1:200, about 10:1 to about 5:1, about 10:1 to about 2:1, about 10:1 to about 1:1, about 10:1 to about 1:2, about 10:1 to about 1:5, about 10:1 to about 1:10, about 10 :1 to about 1:100, about 10:1 to about 1:200, about 5:1 to about 2:1, about 5:1 to about 1:1, about 5:1 to about 1:2, about 5:1 to about 1:5, about 5:1 to about 1:10, about 5:1 to about 1:100, about 5:1 to about 1:200, about 2:1 to about 1:1, about 2:1 to about 1:2, about 2:1 to about 1:5, about 2:1 to about 1:10, about 2:1 to about 1:100, about 2:1 to about 1:200, about 1:1 to about 1:2 , about 1:1 to about 1:5, about 1:1 to about 1:10, about 1:1 to about 1:100, about 1:1 to about 1:200, about 1:2 to about 1:5, about 1:2 to about 1:10, about 1:2 to about 1:100, about 1:2 to about 1:200, about 1:5 to about 1:10, about 1:5 to about 1:100, about 1:5 to about 1:200, about 1:10 to about 1:100, about 1:10 to about 1:200, or about 1:100 to about 1:200. In some embodiments, the water to solvent ratio (v / v) is about 5:1 to about 1:5.

[0223] In one embodiment, an asymmetric compound of formula I:

[0224] [ka]

[0225] (In the formula: R 1 is a halogen, a protected amino group, —NO2, or —OH or a protected form thereof; R 2 is H; R 3 -COOR 5 , -CH2R 6 or a protected aldehyde; or R 2 and R 3 is combined

[0226] [ka]

[0227] is formed; R 4 is H or an amine protecting group; R 5 is C 1~6 Alkyl, C 3~6 Cycloalkyl, C 4~10 heterocyclyl, aryl, or heteroaryl; R 6 is a leaving group or -OH or a protected form thereof) Disclosed is the use of the modified transaminase polypeptide of any one of the preceding embodiments in the preparation of

[0228] In one embodiment according to the use of the foregoing embodiments, the modified transaminase polypeptide is selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 4, 6, 8, 10, 14, 16, 18, and 20. In some embodiments, the modified transaminase polypeptide is selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 16, 18, and 20.

[0229] In one embodiment, using any one of the preceding embodiments, R 2and R 3 is combined

[0230] [ka]

[0231] Form R 4 is H. In one embodiment, R 2 and R 3 is combined

[0232] [ka]

[0233] Form R 4 is H. In some embodiments, R 1 is Br and R 2 is H and R 3 Ha-CH2R 6 and R 4 is H and R 6 is OH.

[0234] definition Unless otherwise explained, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Similarly, the word "or" includes "and" unless the context clearly dictates otherwise. The term "plurality" refers to two or more. The term "at least one" refers to one or more than one.

[0235] Furthermore, numerical limitations stated with respect to concentrations or levels of substances, such as solution component concentrations or ratios thereof, and reaction conditions, such as temperature, pressure, and cycle time, are intended to be approximations. Unless otherwise specified, when a numerical range is provided, the range is inclusive, i.e., includes the endpoints.

[0236] For purposes of description herein, the abbreviations used for the genetically encoded amino acids are conventional and are as shown in Table 1.

[0237] [Table 1]

[0238] When one-letter or three-letter abbreviations are used, the amino acid may be in either the L- or D-configuration about the α-carbon (C-α) unless specifically preceded by "L" or "D" or unless otherwise apparent from the context in which the abbreviation is used. For example, "Ala" refers to alanine without specifying the configuration about the α-carbon, while "D-Ala" and "L-Ala" refer to D-alanine and L-alanine, respectively. When a peptide sequence is represented as a series of one-letter or three-letter abbreviations, the sequence is represented in the N→C direction, according to convention.

[0239] "About" or "approximately" means approximately, in the vicinity of, or in the region of. The term "about" or "approximately" also means within an acceptable contextual error range for a particular value, as determined by one of ordinary skill in the art, which will depend in part on how the value is measured, i.e., the limitations of the measurement system or the degree of precision needed for a particular purpose. When the term "about" or "approximately" is used in conjunction with a range of numerical values, it modifies that range by extending the boundaries above and below the stated numerical values.

[0240] "Acidic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that exhibits a pK value of less than about 6 when the amino acid is included in a peptide or polypeptide. Acidic amino acids typically have a negatively charged side chain at physiological pH due to loss of a hydrogen ion. Genetically encoded acidic amino acids include L-Glu (E) and L-Asp (D).

[0241] An "amino acid" or "residue" as used in the context of the polypeptides disclosed herein refers to a particular monomer at a sequence position (e.g., P5 indicates that the "amino acid" or "residue" at position 5 is proline).

[0242] An "amino acid difference" or "residue difference" or "amino acid substitution" refers to a change in an amino acid residue at a position in a polypeptide sequence relative to the amino acid residue at the corresponding position in a reference sequence. The position of an amino acid difference is generally referred to herein as "Xn," where n refers to the corresponding position in the reference sequence that is the basis for the residue difference. For example, a "residue difference at position X61 compared to SEQ ID NO:2" refers to a change in the amino acid residue at the polypeptide position corresponding to position 61 of SEQ ID NO:2. Thus, if the reference polypeptide of SEQ ID NO:2 has a tyrosine at position 61, then a "residue difference at position X61 compared to SEQ ID NO:2" is an amino acid substitution of any residue other than tyrosine at the polypeptide position corresponding to position 61 of SEQ ID NO:2. In most instances herein, a specific amino acid residue difference at a position will be designated as "XnY," where "Xn" designates the corresponding position above and "Y" is the one-letter identifier of the amino acid identified in the modified polypeptide (i.e., the residue that differs from the reference polypeptide). In some cases (e.g., Tables 7 and 8), the invention also provides specific amino acid differences, designated by the conventional designation "AnB," where A is the single-letter identifier of the residue in the reference sequence, "n" is the number of the residue position in the reference sequence, and B is the single-letter identifier of the residue substitution in the sequence of the variant polypeptide. In some cases, the polypeptides of the invention may contain one or more amino acid residue differences relative to the reference sequence, which differences are indicated by a listing of the specific positions at which the change is made relative to the reference sequence. The invention includes variant polypeptide sequences that contain one or more amino acid differences, including either conservative and non-conservative amino acid substitutions, or both.

[0243] The terms "amino acceptor" and "amine acceptor," "keto substrate," "keto," and "ketone" are used interchangeably herein and refer to a carbonyl (keto, or ketone) compound that accepts an amino group from a donor amine. An amino acceptor is a molecule of the general formula shown below:

[0244] [ka]

[0245] (In the formula, R 1 , R 2 When R is independently an alkyl, alkylaryl, or aryl group that is unsubstituted or substituted with one or more enzymatically tolerated groups. 1 is R in structure or chirality 2 In some embodiments, R 1 and R 2may be taken together to form a ring that is unsubstituted, substituted, or fused with another ring. Amino acceptors include ketocarboxylic acids and alkanones (ketones). Typical ketocarboxylic acids are α-ketocarboxylic acids, such as glyoxalic acid, pyruvic acid, oxaloacetic acid, and the like, as well as salts of these acids. Amino acceptors also include substances that are converted to amino acceptors by other enzymes or whole-cell processes, such as fumaric acid (which can be converted to oxaloacetic acid), glucose (which can be converted to pyruvic acid), lactate, maleic acid, and the like. Amino acceptors that may be used include, by way of example and without limitation, (R)-2-(3,4-dimethoxyphenetoxy)cyclohexanone, 3,4-dihydronaphthalen-1(2H)-one, 1-phenylbutan-2-one, 3,3-dimethylbutan-2-one, octan-2-one, ethyl 3-oxobutanoate, 4-phenylbutan-2-one, 1-(4-bromophenyl)ethanone, 2-methyl-cyclohexamone, 7-methoxy-2-tetralone, 1-hydroxybutan-2-one, pyruvate, acetophenone, (R)-2-(3,4-dimethoxyphenetoxy)cyclohexanone, 2-methoxy-5-fluoroacetophenone, levulinic acid, benzophenone ... acid, 1-phenylpropan-1-one, 1-(4-bromophenyl)propan-1-one, 1-(4-nitrophenyl)propan-1-one, 1-phenylpropan-2-one, 2-oxo-3-methylbutanoic acid, 1-(3-trifluoromethylphenyl)propan-1-one, hydroxypropanone, methoxyoxypropanone, 1-phenylbutan-1-one, 1-(2,5-dimethoxy-4-methylphenyl)butan-2-one, 1-(4-hydroxyphenyl)butan-3-one, 2-acetylnaphthalene, phenylpyruvic acid, 2-ketoglutaric acid, and 2-ketosuccinic acid, including both the (R) and (S) single isomers where possible.

[0246] "Amino donor" or "amine donor" refers to an amino compound that donates an amino group to an amino acceptor, thereby becoming a carbonyl species. An amino donor is a molecule of the general formula:

[0247] [ka]

[0248] (In the formula, R 3 , R 4 When R is independently an alkyl group, an alkylaryl group, or an aryl group that is unsubstituted or substituted with one or more enzymatically non-inhibitory groups. 3 is R in structure or chirality 4 In some embodiments, R 3 and R 4may be taken together to form a ring that is unsubstituted, substituted, or fused to another ring. Exemplary amino donors that may be used in embodiments of the present invention include chiral and achiral amino acids, and chiral and achiral amines.Amino donors that may be used in embodiments herein include, by way of example and not limitation, isopropylamine (also known as 2-aminopropane, and elsewhere herein referred to as "IPM"), α-phenethylamine (also known as 1-phenylethanamine), and its enantiomers (S)-1-phenylethanamine and (R)-1-phenylethanamine, 2-amino-4-phenylbutane, glycine, L-glutamic acid, L-glutamate, monosodium glutamate, L-alanine, D-alanine, D,L ... amine, L-aspartic acid, L-lysine, D,L-ornithine, β-alanine, taurine, n-octylamine, cyclohexylamine, 1,4-butanediamine (also called putrescine), 1,6-hexanediamine, 6-aminohexanoic acid, 4-aminobutyric acid, tyramine, and benzylamine, 2-aminobutane, 2-amino-1-butanol, 1-amino-1-phenylethane, 1-amino-1-(2-methoxy-5-fluorophenyl)ethane, 1-amino-1-phenylpropane, 1-amino-1-(4-hydroxyphenyl)-2-methyl ... (4-aminophenyl)propane, 1-amino-1-(4-bromophenyl)propane, 1-amino-1-(4-nitrophenyl)propane, 1-phenyl-2-aminopropane, 1-(3-trifluoromethylphenyl)-2-aminopropane, 2-aminopropanol, 1-amino-1-phenylbutane, 1-phenyl-2-aminobutane, 1-(2,5-dimethoxy-4-methylphenyl)-2-aminobutane, 1-phenyl-3-aminobutane, 1-(4-hydroxyphenyl)-3-aminobutane, 1-amino-2-methylcyclopentane, Includes 1-amino-3-methylcyclopentane, 1-amino-2-methylcyclohexane, 1-amino-1-(2-naphthyl)ethane, 3-methylcyclopentylamine, 2-methylcyclopentylamine, 2-ethylcyclopentylamine, 2-methylcyclohexylamine, 3-methylcyclohexylamine, 1-aminotetralin, 2-aminotetralin, 2-amino-5-methoxytetralin, and 1-aminoindan, including both the (R) and (S) single isomers where possible, and all possible salts of the amines.

[0249] "Aliphatic amino acid or residue" refers to a hydrophobic amino acid or residue having an aliphatic hydrocarbon side chain. Genetically encoded aliphatic amino acids include L-Ala (A), L-Val (V), L-Leu (L), and L-Ile (I).

[0250] "And / or" when used in phrases such as "A and / or B" is intended to include "A and B," "A or B," "A," and "B." Similarly, the term "and / or" when used in phrases such as "A, B, and / or C" is intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).

[0251] "Aromatic amino acid or residue" refers to a hydrophilic or hydrophobic amino acid or residue having a side chain containing at least one aromatic or heteroaromatic ring. Genetically encoded aromatic amino acids include L-Phe (F), L-Tyr (Y), and L-Trp (W). Although sometimes classified as a basic residue due to the pKa of its heteroaromatic nitrogen atom, L-His (H), or as an aromatic residue because its side chain contains a heteroaromatic ring, histidine is classified herein as a hydrophilic residue or a "constrained residue" (see below).

[0252] A "basic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that exhibits a pKa value greater than about 6 when the amino acid is included in a peptide or polypeptide. Basic amino acids typically have a positively charged side chain at physiological pH due to binding with a hydronium ion. Genetically encoded basic amino acids include L-Arg (R) and L-Lys (K).

[0253] "Chiral amines" are those of the general formula R 1 -CH(NH2)-R 2The term "amine" refers to an amine of the formula (R), and is used herein in its broadest sense to include a wide variety of aliphatic and alicyclic compounds of different mixed functional group types, characterized by the presence of a primary amino group attached to a secondary carbon atom that, in addition to a hydrogen atom (H), bears either (i) a divalent group that forms a chiral ring structure, or (ii) two substituents (other than hydrogen) that differ from each other in structure or chirality. Divalent groups that form chiral ring structures include, for example, 2-methylbutane-1,4-diyl, pentane-1,4-diyl, hexane-1,4-diyl, hexane-1,5-diyl, and 2-methylpentane-1,5-diyl. The term "amine" refers to an amine of the formula (R), and is used herein in its broadest sense to include a wide variety of aliphatic and alicyclic compounds of different mixed functional group types that are characterized by the presence of a primary amino group attached to a secondary carbon atom that bears, in addition to a hydrogen atom (H), either (i) a divalent group that forms a chiral ring structure, or (ii) two substituents (other than hydrogen) that differ from each other in structure or chirality. Divalent groups that form chiral ring structures include, for example, 2-methylbutane-1,4-diyl, pentane-1,4-diyl, hexane-1,4-diyl, hexane-1,5-diyl, and 2-methylpentane-1,5-diyl. The term "amine" refers to an amine of the formula (R), and is used herein in its broadest sense to include a wide variety of aliphatic and alicyclic compounds of different functional groups that are characterized by the presence of two different substituents (other than hydrogen) on the secondary carbon atom. 1 and R 2 ) can also vary widely and include alkyl, aralkyl, aryl, halo, hydroxy, lower alkyl, lower alkoxy, lower alkylthio, cycloalkyl, carboxy, carbalkoxy, carbamoyl, mono- and di-(lower alkyl) substituted carbamoyl, trifluoromethyl, phenyl, nitro, amino, mono- and di(lower alkyl) substituted amino, alkylsulfonyl, arylsulfonyl, alkylcarboxamido, arylcarboxamido, and the like, as well as alkyl, aralkyl, or aryl substituted by the foregoing.

[0254] As used herein, a "cofactor" refers to a non-protein compound that works in combination with an enzyme to catalyze a reaction. "Pyridoxal phosphate," "PLP," "pyridoxal-5'-phosphate," "PYP," and "P5P" are used interchangeably herein to refer to compounds that act as cofactors in transaminase reactions. In some embodiments, pyridoxal phosphate is defined by the structure 1-(4'-formyl-3'-hydroxy-2'-methyl-5'-pyridyl)methoxyphosphonic acid, CAS number [54-47-7], and pyridoxal-5'-phosphate can be generated in vivo by phosphorylation and oxidation of pyridoxol (also known as vitamin B6). In transamination reactions using transaminase enzymes, the amine group of the amino donor is transferred to the cofactor to generate a keto byproduct, while pyridoxal-5'-phosphate is converted to pyridoxamine phosphate. Pyridoxal-5'-phosphate is regenerated by reaction with a different keto compound (amino acceptor). Transfer of the amine group from pyridoxamine phosphate to the amino acceptor generates a chiral amine and regenerates the cofactor. In some embodiments, pyridoxal-5'-phosphate can be replaced by other members of the vitamin B6 family, including pyridoxine (PN), pyridoxal (PL), pyridoxamine (PM), and their phosphorylated counterparts; pyridoxine phosphate (PNP) and pyridoxamine phosphate (PMP). As used herein, "cofactor" refers to the vitamin B6 family compounds PLP, PN, PL, PM, PNP, and PMP, which may also be referred to as coenzymes.

[0255] "Coding sequence" refers to a portion of a nucleic acid (eg, a gene) that codes for the amino acid sequence of a protein.

[0256] "Conservative" amino acid substitution or mutation refers to the interchangeability of residues with similar side chains, and thus typically involves substituting an amino acid in a polypeptide with an amino acid in the same or similar defined class of amino acids. However, as used herein, in some embodiments, conservative mutations do not include the substitution of a hydrophilic residue for a hydrophilic residue, a hydrophobic residue for a hydrophobic residue, a hydroxyl-containing residue for a hydroxyl-containing residue, or a small residue for a small residue, although they may be substituted for an aliphatic residue for an aliphatic residue, a non-polar residue for a non-polar residue, a polar residue for a polar residue, an acidic residue for an acidic residue, a basic residue for a basic residue, an aromatic residue for an aromatic residue, or a constrained residue for a constrained residue. Furthermore, as used herein, A, V, L, or I may be conservatively mutated to either another aliphatic residue or another non-polar residue. Table 2 below shows exemplary conservative substitutions.

[0257] [Table 2]

[0258] "Constrained amino acid or residue" refers to an amino acid or residue that has a constrained shape. As used herein, constrained residues include L-Pro (P) and L-His (H). Histidine has a constrained shape because it has a relatively small imidazole ring. Proline also has a constrained shape because it has a five-membered ring.

[0259] "Conversion" refers to the enzymatic conversion of a substrate(s) to the corresponding product(s). "Conversion rate" refers to the percent of a substrate that is converted to a product under specific conditions within a given period of time. Thus, the "enzyme activity" or "activity" of a transaminase polypeptide can be expressed as the "percent conversion" of substrate to product.

[0260] "Corresponding," "referenced," or "relative," when used in the context of the numbering of a given amino acid sequence or polynucleotide sequence, refers to the numbering of residues in a particular reference sequence when the given amino acid sequence or polynucleotide sequence is compared to the reference sequence. Methods for comparing a sequence to a particular reference sequence are known to those of skill in the art. For example, the Needleman-Wunsch method can be used to compare any amino acid sequence or polynucleotide sequence to a reference sequence.

[0261] The term "corresponding amino acid position" is widely used and well understood by those skilled in the art. Corresponding amino acid positions can be identified by aligning amino acid sequences using any of the well-known amino acid alignment methods. For example, corresponding amino acid positions can be identified using the NCBI BLAST algorithm. "Cysteine" or L-Cys(C) is unusual in that it can form disulfide bonds with other L-Cys(C) amino acids or other sulfanyl- or sulfhydryl-containing amino acids. "Cysteine-like residues" include cysteine ​​and other amino acids containing sulfhydryl moieties available for disulfide bond formation. The ability of L-Cys(C) (and other amino acids with -SH-containing side chains) to exist within a peptide in either the reduced, free -SH or oxidized, disulfide-bridged form affects whether L-Cys(C) confers net hydrophobic or hydrophilic properties to the peptide. It will be understood that although L-Cys(C) exhibits a hydrophobicity of 0.29 according to Eisenberg's normalized consensus scale (Eisenberg et al., 1984, supra), for purposes of the present invention, L-Cys(C) will be classified in its own unique group.

[0262] "Deletion" refers to modifying a polypeptide by removing one or more amino acids from a reference polypeptide. Deletions can include removing up to 10% of the total number of amino acids comprising the reference enzyme, or up to 20% of the total number of amino acids, one or more amino acids, two or more amino acids, five or more amino acids, ten or more amino acids, fifteen or more amino acids, or twenty or more amino acids, while maintaining enzymatic activity and / or maintaining improved properties of the modified transaminase enzyme. Deletions can be directed to internal and / or terminal portions of the polypeptide. In various embodiments, deletions can include contiguous segments or can be discontinuous.

[0263] "Origin" as used herein in the context of modified transaminase enzymes identifies the original transaminase enzyme on which the modification is based and / or the gene encoding such transaminase enzyme.

[0264] As used herein, a "fragment" refers to a polypeptide that has an amino- and / or carboxy-terminal deletion, but where the remaining amino acid sequence is identical to the corresponding positions in the sequence. Fragments can be at least 14 amino acids long, at least 20 amino acids long, at least 50 amino acids long, or longer, and can represent up to 70%, 80%, 90%, 95%, 98%, and 99% or more of the full-length transaminase. As used interchangeably herein, a "functional fragment" or a "biologically active fragment" refers to a polypeptide that has an amino- and / or carboxy-terminal deletion(s) and / or an internal deletion, but where the remaining amino acid sequence is identical to the corresponding positions in the sequence to which it is compared, and that retains substantially all of the activity of the full-length polypeptide.

[0265] "Improved enzymatic properties" refer to transaminase polypeptides that exhibit improvements in any enzymatic property compared to a reference transaminase. For the modified transaminase polypeptides described herein, the comparison is generally made to a wild-type transaminase enzyme, although in some embodiments, the reference transaminase may be another improved modified transaminase. Enzymatic properties for which improvement is desired include, but are not limited to, enzymatic activity (which may be expressed as percent conversion of substrate), thermostability, solvent stability, pH-activity profile, cofactor requirements, refractoriness to inhibitors (e.g., substrate or product inhibition), stereospecificity, and stereoselectivity (including enantioselectivity). In one embodiment, a transaminase enzyme of the invention has a stereoselectivity of greater than 95, 96, 97, 98, or 99% ee.

[0266] "Increased enzymatic activity" refers to improved properties of a modified transaminase polypeptide, which may be expressed as an increase in specific activity (e.g., product produced / time / weight of protein) or an increase in percent conversion of substrate to product (e.g., percent conversion of starting amount of substrate to product in a specified time using a specified amount of transaminase) compared to a reference transaminase enzyme. Exemplary methods for determining enzymatic activity are provided in the Examples. m , V max or k catAny property associated with enzyme activity can be affected, including classical enzyme properties of k, k = k + k , and these changes can lead to increased enzyme activity. Improvements in enzyme activity can range from about 1.1-fold the enzyme activity of the corresponding transaminase enzyme set forth in SEQ ID NO:2 to 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, or more than the enzyme activity of the native transaminase or another engineered transaminase from which the transaminase polypeptide is derived. In certain embodiments, the engineered transaminase enzyme exhibits improved enzyme activity in the range of 2.0-50-fold, 2.0-100-fold, or more than the control transaminase enzyme having SEQ ID NO:2. Those skilled in the art will understand that the activity of any enzyme is diffusion-limited, whereby the catalytic turnover rate cannot exceed the diffusion rate of the substrate, including any required cofactors. Diffusion-limited or k = k + k . cat / K m The theoretical maximum value of is generally about 10 8 ~10 9 (Ms ”1 ) Therefore, any improvement in the enzymatic activity of a transaminase has an upper limit related to the diffusion rate of the substrate acted upon by the transaminase enzyme. Transaminase activity can be measured by any one of standard assays, such as monitoring changes in the spectrophotometric properties of reactants or products. Enzyme activity comparisons are made using defined enzyme preparations, defined assays under set conditions, and one or more defined substrates, as described in more detail herein.

[0267] In one embodiment, the transaminase enzymes of the invention exhibit at least a 2-fold improved increase in activity over the enzymatic activity of SEQ ID NO: 2. In one embodiment, the transaminase enzymes of the invention exhibit at least a 2-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, or 50-fold improved increase in activity over the enzymatic activity of SEQ ID NO: 2, and have a stereoselectivity of greater than 95, 96, 97, 98, or 99% ee.

[0268] "Insertion" refers to the modification of a polypeptide by the addition of one or more amino acids from a reference polypeptide. In some embodiments, improved modified transaminase enzymes include the insertion of one or more amino acids into a native transaminase polypeptide, as well as the insertion of one or more amino acids into other improved transaminase polypeptides. The insertion may be in an internal portion of the polypeptide, or at the carboxy- or amino-terminus. As used herein, insertion includes fusion proteins, which are known in the art. The insertion may be a contiguous segment of amino acids in the native polypeptide, or may be separated by one or more amino acids.

[0269] "Native" or "wild-type" refers to a form found in nature. For example, a native or wild-type polypeptide or polynucleotide sequence is a sequence that exists in an organism that can be isolated from a natural source and has not been intentionally modified by human manipulation.

[0270] "Non-natural," "recombinant," or "modified," e.g., when used in reference to a cell, nucleic acid, or polypeptide, refers to material that is produced or derived in a manner that does not occur in nature, or that is identical to that in nature but from synthetic materials and / or by manipulation using recombinant techniques, or that corresponds to the natural or native form of the material. Non-limiting examples include, among others, recombinant cells that express genes not found within the native (non-recombinant) form of the cell, or that express native genes that are otherwise expressed at different levels.

[0271] As used herein, "nucleic acid" refers to a polymeric form of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, and / or their analogs. This includes DNA, RNA, and DNA / RNA hybrids. It also includes DNA or RNA analogs, such as those containing modified backbones (e.g., peptide nucleic acid (PNA) or phosphorothioates) or modified bases. Thus, the nucleic acids of the present invention include mRNA, DNA, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, and the like. When the nucleic acid is in the form of RNA, it may or may not have a 5' cap. RNA can be small, medium, or large RNA. Small RNAs have 10 to 30 nucleotides per strand (e.g., siRNA). Medium-sized RNAs have 30 to 2,000 nucleotides per strand (e.g., non-self-replicating mRNA). Large RNAs contain at least 2,000 nucleotides per strand, for example, at least 2,500 nucleotides, at least 3,000 nucleotides, at least 4,000 nucleotides, at least 5,000 nucleotides, at least 6,000 nucleotides, at least 7,000 nucleotides, at least 8,000 nucleotides, at least 9,000 nucleotides, or at least 10,000 nucleotides per strand. The molecular weight in g / mol (or Daltons) of a single-stranded RNA molecule can be estimated using the following formula: molecular weight = (number of RNA nucleotides) × 340 g / mol. RNA may contain one or more nucleotides with modified nucleobases, in addition to any 5' cap structure. For example, RNA may contain one or more modified pyrimidine nucleobases, such as pseudouridine and / or 5-methylcytosine residues. However, in some embodiments, the RNA does not contain modified nucleobases and may also contain modified nucleotides, i.e., all nucleotides in the RNA are standard A, C, G, and U ribonucleotides (except for the 5' cap structure, which may contain 7' methylguanosine).In other embodiments, the RNA may include a 5' cap comprising a 7' methylguanosine, and the first 1, 2, or 3 5' ribonucleotides may be methylated at the 2' position of the ribose.

[0272] The nucleic acid may be in a recombinant form, i.e., a form that does not occur in nature. For example, the nucleic acid may contain one or more heterologous nucleic acid sequences (e.g., a sequence encoding another antigen and / or a regulatory sequence such as a promoter or internal ribosome entry site). The nucleic acid may be part of a vector, i.e., part of a nucleic acid designed for transduction / transfection of one or more cell types. The vector may be, for example, an "expression vector" designed for expression of a nucleotide sequence in a host cell, or a "viral vector" designed to result in the production of a recombinant virus or virus-like particle.

[0273] "Percentage of sequence identity," "percent identity," and "percent identical" are used herein to refer to a comparison between polynucleotide or polypeptide sequences, and are determined by comparing two optimally aligned sequences across a comparison window, where the portion of the polynucleotide or polypeptide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence due to optimal alignment of the two sequences. This percentage is calculated by determining the number of positions where either the same nucleic acid base or amino acid residue occurs in both sequences, or where the nucleic acid base or amino acid residue is aligned with a gap, to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Optimal alignment and determination of percent sequence identity are performed using the BLAST and BLAST 2.0 algorithms (see, e.g., Altschul et al., 1990, J. Mol. Biol. 215:403-410 and Altschul et al., 1977, Nucleic Acids Res. 3389-3402). Software for performing BLAST analyses is publicly available from the website of the National Center for Biotechnology Information.

[0274] A number of other algorithms are available that function similarly to BLAST in providing the percent identity of two sequences. Optimal alignment of sequences for comparison can be performed, for example, by the local homology algorithm of Smith and Waterman, 1981, Adv. Appl. Math. 2:482, by the homology alignment algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48:443, by the similarity search method of Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444, by computer implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin Software Package), or by visual inspection (see generally Current Protocols in Molecular Biology, FM Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)). Additionally, for sequence alignment and determination of percent sequence identity, the BESTFIT or GAP programs from the GCG Wisconsin Software package (Accelrys, Madison, WI) can be used using the default parameters provided. The ClustalW program is also suitable for determining identity.

[0275] "Protein," "polypeptide," and "peptide" are used interchangeably herein to refer to polymers of at least two amino acids covalently joined by an amide bond, regardless of length or post-translational modification (e.g., glycosylation, phosphorylation, lipidation, myristylation, ubiquitination, etc.). This definition includes D- and L-amino acids, and mixtures of D- and L-amino acids.

[0276] As used herein, "purification" or "purifying" refers to a process by which undesired components are removed from a composition or host cell or culture. Purification is a relative term and does not require the removal of all traces of undesired components from a composition. Purification can include processes such as centrifugation, dialysis, ion exchange and size exclusion chromatography, affinity purification, or precipitation. Thus, the term "purified" does not require absolute purity; rather, it is intended as a relative term. A substantially pure nucleic acid or protein preparation can be purified so that the desired nucleic acid or protein accounts for at least 50% of the total nucleic acid content of the preparation. In certain embodiments, a substantially pure nucleic acid or protein accounts for at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, or at least 95% or more of the total nucleic acid or protein content of the preparation. Immunogenic molecules, antigens, or antibodies (i.e., molecules found in nature) that have not undergone any purification steps are not suitable for use as pharmaceuticals (e.g., vaccines).

[0277] A "reference sequence" refers to a defined sequence used as the basis for sequence comparison. A reference sequence can be a subset of a larger sequence, such as a segment of a full-length gene or polypeptide sequence. Generally, a reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 residues in length, at least 50 residues in length, or the entire length of the nucleic acid or polypeptide. Because two polynucleotides or polypeptides may each contain (1) similar sequences (i.e., a portion of the complete sequence) between the two sequences and (2) additional sequences that differ between the two sequences, sequence comparison between two (or more) polynucleotides or polypeptides is typically performed by comparing the sequences of the two polynucleotides or polypeptides over a "comparison window" to identify and compare local regions of sequence similarity. In some embodiments, a "reference sequence" may be based on a primary amino acid sequence, where the reference sequence may have one or more changes in the primary sequence. For example, a "reference sequence based on SEQ ID NO: 2 having a threonine at the residue corresponding to X137" refers to a reference sequence in which the corresponding residue at X137 in SEQ ID NO: 2, which is glutamic acid (E), has been changed to threonine (T). A "comparison region" refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acid residues, where a sequence is compared to a reference sequence of at least 20 contiguous nucleotides or amino acids, and where the portion of the sequence within the comparison region may contain no more than 20 percent additions or deletions (i.e., gaps) compared to the reference sequence (which contains no additions or deletions) in order to optimally align the two sequences. A comparison region may be longer than 20 contiguous residues, and optionally include a region of 30, 40, 50, 100, or more.

[0278] "With reference to," "corresponding to," or "with respect to," when used in the context of numbering a given amino acid or polynucleotide sequence, refers to the numbering of residues in a particular reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, residue numbers or residue positions in a given polymer are specified with respect to the reference sequence, rather than the actual numerical position of the residues in the given amino acid or polynucleotide sequence. For example, a given amino acid sequence, such as that of a modified transaminase, can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, despite the presence of gaps, the numbering of residues in the given amino acid or polynucleotide sequence is done with respect to the reference sequence to which it is aligned.

[0279] A "substantially pure polypeptide" refers to a composition in which the polypeptide species is the predominant species present (i.e., it is more abundant than any other individual macromolecular species in the composition, on a molar or weight basis). A substantially purified composition generally exists when the target species constitutes at least about 50 percent of the macromolecular species present, on a molar or weight percent basis. Generally, a substantially pure transaminase composition contains about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, and about 98% or more of all macromolecular species present in the composition, on a molar or weight percent basis. In some embodiments, the target species is purified to essential homogeneity (i.e., contaminant species cannot be detected in the composition by conventional detection methods), where the composition consists essentially of a single macromolecular species. Solvent species, small molecules (<500 Daltons), and elemental ion species are not considered macromolecular species. In some embodiments, an isolated improved transaminase polypeptide is a substantially pure polypeptide composition.

[0280] "Stereoselectivity" refers to the preferential formation of one stereoisomer over another in a chemical or enzymatic reaction. Stereoselectivity can be partial, where the formation of one stereoisomer is favored over the other, or complete, where only one stereoisomer is formed. When the stereoisomers are enantiomers, stereoselectivity is referred to as enantioselectivity, which is the proportion (typically reported as a percentage) of one enantiomer in the sum of both enantiomers. Alternatively, it is commonly reported in the art as the enantiomeric excess (ee) (typically as a percentage) calculated from them according to the formula [major enantiomer - minor enantiomer] / [major enantiomer + minor enantiomer]. When the stereoisomers are diastereoisomers, the stereoselectivity is referred to as diastereoselectivity, which is the proportion of one diastereomer in a mixture of two diastereomers (typically reported as a percentage), or more commonly reported as diastereomeric excess (de). When a mixture contains more than two diastereomers, it is common to report the ratio of diastereomers or "diastereomeric ratio" rather than diastereomeric excess. Enantiomeric excess and diastereomeric excess are types of stereoisomeric excess. "Highly stereoselective" refers to a transaminase polypeptide that can convert a substrate to the corresponding chiral amine product with a stereoisomeric excess of at least about 90%, about 95%, or about 99%.

[0281] "Transaminase" or "aminotransferase" are used interchangeably herein to refer to a polypeptide having the enzymatic ability to reversibly transfer an amino group (NH), an electron pair, and a proton from a primary amine to a carbonyl group (C=O) of an acceptor molecule. As will be appreciated by those skilled in the art, an enzyme with transaminase activity may be capable of other reactions under different reaction conditions and with different substrates, e.g., dehalogenation / deamination. The present invention encompasses all such enzymes.

[0282] As used herein, transaminase includes naturally occurring (wild-type) transaminases as well as non-naturally occurring modified polypeptides produced by human engineering. In one embodiment, a modified transaminase polypeptide or functional fragment thereof is disclosed, comprising an amino acid sequence having at least 80%, 85%, 90%, or 95% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises the characteristic that X124 is isoleucine (I). In one embodiment, a modified transaminase polypeptide or functional fragment thereof is disclosed, comprising an amino acid sequence having at least 80%, 85%, 90%, or 95% sequence identity to the amino acid sequence set forth in SEQ ID NO:2, wherein the amino acid sequence comprises a substitution of at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:2, and wherein X124 is isoleucine (I).

[0283] In one embodiment, a transaminase polypeptide is disclosed comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, or 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 4, 6, 8, 10, 14, 16, 18, or 20. In one embodiment, a transaminase polynucleotide is disclosed comprising a polynucleotide sequence having at least 80%, 85%, 90%, 95%, or 100% sequence identity to the polynucleotide sequence set forth in SEQ ID NO: 3, 5, 7, 9, 13, 15, 17, or 19.

[0284] In one aspect, the present invention provides an asymmetric compound of formula I:

[0285] [ka]

[0286] (In the formula: R 1 is a leaving group, a halogen, a protected amino group, —NO2, or —OH or a protected form thereof; R 2 is H; R 3 -COOR 5 , -CH2R 6 or a protected aldehyde; or R 2 and R 3 is combined

[0287] [ka]

[0288] is formed; R 4 is H or an amine protecting group; R 5 is C 1~6 Alkyl, C 3~10 Cycloalkyl, C 4~10 heterocyclyl, aryl, or heteroaryl; R 6 is a leaving group or —OH or a protected form thereof), the method comprising the step of preparing a compound of formula II:

[0289] [ka]

[0290] (In the formula: R 1 ' is a leaving group, a halogen, a protected amino group, NO2, or OH or a protected form thereof; R 2 ' is an aldehyde or aldehyde equivalent; R 3 '-COOR 5 , -CH2R 6or a protected aldehyde; or R 2 ' and R 3 ' is a combination

[0291] [ka]

[0292] (where, * represents the attachment point) is formed) with the modified transaminase polypeptide of any one of the embodiments disclosed herein.

[0293] "Alkyl" refers to a saturated, straight-chain or branched-chain hydrocarbon moiety having the specified number of carbon atoms. The term "(C1-C6) alkyl" refers to an alkyl moiety containing 1 to 6 carbon atoms. Exemplary alkyls include, but are not limited to, methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, s-butyl, t-butyl, pentyl, and hexyl. In some embodiments, "Me" refers to a methyl group. When the term "alkyl" is used in combination with other substituents, the term "alkyl" is intended to include divalent, straight-chain or branched-chain hydrocarbon groups whose point of attachment is through the alkyl portion.

[0294] "Aryl" refers to a monocyclic or bicyclic hydrocarbon-based aromatic group. Examples of aryl include phenyl and naphthyl. Aryl groups can contain 6 to 14 carbon atoms. The term "heteroaryl" refers to a group or moiety containing an aromatic monovalent monocyclic or bicyclic group containing 5 to 10 ring atoms and at least one heteroatom independently selected from nitrogen, oxygen, and sulfur. The term also encompasses bicyclic heterocyclic aryl compounds containing an aryl ring moiety fused to a heterocycloalkyl ring moiety containing 5 to 10 ring atoms and at least one heteroatom independently selected from nitrogen, oxygen, and sulfur. Exemplary groups include, but are not limited to, furanyl, thienyl, pyrrolyl, imidazolyl, pyrazolyl, triazolyl, tetrazolyl, thiazolyl, oxazolyl, isoxazolyl, oxadiazolyl, thiadiazolyl, isothiazolyl, pyridinyl, pyridazinyl, pyrazinyl, pyrimidinyl, triazinyl, benzofuranyl, isobenzofuryl, 2,3-dihydrobenzofuryl, 1,3-benzodioxolyl, dihydrobenzodioxinyl, benzothienyl, indolizinyl, indolyl, isoindolyl, dihydroindolyl, benzimidazolyl, dihydroindolyl, benzimidazolyl, dihydroindolyl, Includes benzobenzimidazolyl, benzoxazolyl, dihydrobenzoxazolyl, benzthiazolyl, benzisothiazolyl, dihydrobenzisothiazolyl, indazolyl, imidazopyridinyl, pyrazolopyridinyl, benzotriazolyl, triazolopyridinyl, purinyl, quinolinyl, tetrahydroquinolinyl, isoquinolinyl, tetrahydroisoquinolinyl, quinoxalinyl, cinnolinyl, phthalazinyl, quinazolinyl, 1,5-naphthyridinyl, 1,6-naphthyridinyl, 1,7-naphthyridinyl, 1,8-naphthyridinyl, and pteridinyl. Examples of 5-membered "heteroaryl" groups include furanyl, thienyl, pyrrolyl, imidazolyl, pyrazolyl, triazolyl, tetrazolyl, thiazolyl, oxazolyl, isoxazolyl, oxadiazolyl, thiadiazolyl, and isothiazolyl. Examples of 6-membered "heteroaryl" groups include oxo-pyridyl, pyridinyl, pyridazinyl, pyrazinyl, and pyrimidinyl.Examples of 6,6-fused "heteroaryl" groups include quinolinyl, isoquinolinyl, quinoxalinyl, cinnolinyl, phthalazinyl, quinazolinyl, 1,5-naphthyridinyl, 1,6-naphthyridinyl, 1,7-naphthyridinyl, 1,8-naphthyridinyl, and pteridinyl. Examples of 6,5-fused "heteroaryl" groups include benzofuranyl, benzothienyl, benzimidazolyl, benzthiazolyl, indolizinyl, indolyl, isoindolyl, and indazolyl.

[0295] The term "cycloalkyl" refers to a non-aromatic, saturated, monocyclic, hydrocarbon ring containing the specified number of carbon atoms. The term "(C3-C6)cycloalkyl" refers to a non-aromatic cyclic hydrocarbon ring having from 3 to 6 ring carbon atoms. Exemplary "(C3-C6)cycloalkyl" groups useful in the present invention include cyclopropyl, cyclobutyl, cyclopentyl, and cyclohexyl. Examples of "(C3-C6)cycloalkyl(C1-C4)alkyl-" groups useful in the present invention include, but are not limited to, cyclobutylmethyl, cyclopentylmethyl, cyclohexylmethyl, cyclobutylethyl, cyclopentylethyl, and cyclohexylethyl.

[0296] "Halogen" and "halo" refer to a fluorine, chlorine, bromine, or iodine substituent.

[0297] The term "heterocycle" or "heterocyclyl" as used herein is intended to mean a 3- to 10-membered aromatic or non-aromatic heterocycle containing 1-4 heteroatoms selected from the group consisting of O, N, and S, and includes bicyclic groups. For purposes of the present invention, the term "heterocyclic" is also considered synonymous with the terms "heterocycle" and "heterocyclyl," and is understood to also have the definitions set forth herein. Thus, "heterocyclyl" thus includes the above-mentioned heteroaryls, as well as dihydro and tetrahydro analogs thereof.Further examples of "heterocyclyl" include, but are not limited to, azetidinyl, benzimidazolyl, benzofuranyl, benzofurazanyl, benzopyrazolyl, benzotriazolyl, benzothiophenyl, benzoxazolyl, carbazolyl, carbolinyl, cinnolinyl, furanyl, imidazolyl, indolinyl, indolyl, indolazinyl, indazolyl, isobenzofuranyl, isoindolyl, isoquinolyl, isothiazolyl, isoxazolyl, naphthyl, and the like. tetrahydropyridinyl, oxadiazolyl, oxooxazolidinyl, oxazolyl, oxazoline, oxopiperazinyl, oxopyrrolidinyl, oxomorpholinyl, isoxazoline, oxetanyl, pyranyl, pyrazinyl, pyrazolyl, pyridazinyl, pyridopyridinyl, pyridazinyl, pyridyl, pyrimidyl, pyrrolyl, quinazolinyl, quinolyl, quinoxalinyl, tetrahydropyranyl, tetrahydrofuranyl, tetrahydrothiopyranyl, tetrahydroisoquinolinyl, tetrazo aryl, tetrazolopyridyl, thiadiazolyl, thiazolyl, thienyl, triazolyl, 1,4-dioxanyl, hexahydroazepinyl, piperazinyl, piperidinyl, pyridin-2-onyl, pyrrolidinyl, morpholinyl, thiomorpholinyl, dihydrobenzimidazolyl, dihydrobenzofuranyl, dihydrobenzothiophenyl, dihydrobenzoxazolyl, dihydrofuranyl, dihydroimidazolyl, dihydroindolyl, dihydroisoxazolyl, dihydroisothiazolyl Heterocyclyl groups include dihydrooxadiazolyl, dihydrooxazolyl, dihydropyrazinyl, dihydropyrazolyl, dihydropyridinyl, dihydropyrimidinyl, dihydropyrrolyl, dihydroquinolinyl, dihydrotetrazolyl, dihydrothiadiazolyl, dihydrothiazolyl, dihydrothienyl, dihydrotriazolyl, dihydroazetidinyl, dihydrothiomorpholinyl, methylenedioxybenzoyl, tetrahydrofuranyl, and tetrahydrothienyl, and their N-oxides. Attachment of a heterocyclyl substituent can occur via a carbon atom or via a heteroatom.

[0298] "Hydroxy" or "hydroxyl" shall mean an --OH group.

[0299] "Leaving group" is defined as a term that can be understood by one skilled in the art, i.e., a group on a carbon at which a new bond is formed during a reaction, and upon formation of the new bond, the carbon loses a group. Typical examples of suitable leaving groups include, for example, sp 3 It is a nucleophilic substitution reaction on the hybridized carbon (S N 2 or S N 1), for example, if the leaving group is a halide, e.g., bromide, the reactant can be benzyl bromide. Another typical example of such a reaction is the nucleophilic aromatic substitution reaction (SNAr). Another example is an insertion reaction (e.g., by a transition metal) into the bond between aromatic reaction partners bearing a leaving group, followed by reductive coupling. The term "leaving group" is not limited to such mechanistic limitations. Examples of suitable leaving groups are halogens (fluorine, chlorine, bromine, or iodine), optionally substituted aryl or alkyl sulfonates, phosphonates, azides, and -S(O). 0~2 R (wherein R is, for example, an optionally substituted alkyl, an optionally substituted aryl, or an optionally substituted heteroaryl). Those skilled in the art of organic synthesis can easily identify leaving groups suitable for carrying out the desired reaction under different reaction conditions. Non-limiting characteristics and examples of leaving groups can be found, for example, in Organic Chemistry, 2nd Edition, Francis Carey (1992), pp. 328-331; Introduction to Organic Chemistry, 2nd Edition, Andrew Streitwieser and Clayton Heathcock (1981), pp. 169-171; and Organic Chemistry, 5th Edition, John McMurry, Brooks / Cole Publishing (2000), pp. 398 and 408.

[0300] A "protecting group" refers to a group of atoms that, when attached to a reactive functional group in a molecule, masks, reduces, or prevents the reactivity of the functional group. Typically, a protecting group can be selectively removed as desired during synthesis. Examples of protecting groups can be found in Wuts and Greene, "Greene's Protective Groups in Organic Synthesis," 4th Edition, Wiley Interscience (2006), and Harrison et al., Compendium of Synthetic Organic Methods, Vols. 1-8, 1971-1996, John Wiley & Sons, NY. Functional groups that can have protecting groups include, but are not limited to, hydroxy groups, amino groups, and carboxy groups.

[0301] Representative amine protecting groups include, but are not limited to, formyl, acetyl (Ac), trifluoroacetyl, benzyl (Bn), benzoyl (Bz), carbamate, benzyloxycarbonyl ("CBZ"), p-methoxybenzylcarbonyl (Moz or MeOZ), tert-butoxycarbonyl ("Boc"), trimethylsilyl ("TMS"), 2-trimethylsilyl-ethanesulfonyl ("SES"), trityl and substituted trityl groups, allyloxycarbonyl, 9-fluorenylmethyloxycarbonyl ("FMOC"), nitroveratryloxycarbonyl ("NVOC"), p-methoxybenzyl (PMB), tosyl (Ts), and the like.

[0302] Representative hydroxyl protecting groups include, but are not limited to, mesylate (SOMe), acylated hydroxyl groups (e.g., methyl and ethyl esters, acetate or propionate groups, or glycol esters), or alkylated hydroxyl groups (e.g., benzyl and trityl ethers, as well as alkyl ethers, tetrahydropyranyl ethers, trialkylsilyl ethers (e.g., TMS or TIPS groups), and allyl ethers. Other protecting groups can be found in the references cited herein.

[0303] "Protected aldehyde" is defined as a term understood by one of ordinary skill in the art, i.e., the aldehyde is protected with a group that can be converted to the unprotected aldehyde under assay conditions. Examples of protected aldehydes include, but are not limited to, acetals or hemiacetals that can be converted to the free aldehyde group by treatment with an acid (organic or inorganic), such as the acetal groups formed with polyalcohols such as propanediol or ethylene glycol, or the hemiacetal groups in sugar-related compounds such as sugars or aldose sugars, e.g., glucose or galactose. Further examples of protected aldehydes are imino groups (e.g., =NH groups) which give aldehyde groups on treatment with acid, thioacetal or dithioacetal groups (e.g., C(SR) groups, where R can be an alkyl group) which give aldehyde groups on treatment with a mercury salt, oxime groups (e.g., =NOH groups) which give aldehyde groups on treatment with acid, hydrazone groups (e.g., =N-NHR groups, where R can be an alkyl group) which give aldehyde groups on treatment with acid, and imidazolone or imidazolidine groups, or benzothiazole or dihydrobenzothiazole groups, which give aldehydes on hydrolysis, for example with acid. [Example]

[0304] Numerous modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore to be understood that, within the scope of the appended claims, those skilled in the art will recognize that the invention may be practiced otherwise than as specifically described. The exemplary embodiments and examples are not to be construed as limiting the invention.

[0305] [Example 1] Synthesis and optimization of engineered transaminase polypeptides The DNA sequence of SEQ ID NO: 1 and the polypeptide sequence of SEQ ID NO: 2 of the present disclosure correspond to the DNA sequence of SEQ ID NO: 179 and the polypeptide sequence of SEQ ID NO: 180 of WO2014 / 088984A1, published on June 12, 2014. The engineered transaminase polypeptide of SEQ ID NO:2 has the following 28 amino acid differences compared to the polypeptide sequence of wild-type Arthrobacter sp. KNK168 (GenBank Accession No.: BAK39753.1; GI:336088341): A2S; A5H; S8P, Y60F, L61Y, H62T, V65A, D81G, M94I, I96L, F122I, S124I, G136W, A169L, V199I, A209L, G215F, G217N, S223P, L269P, L273Y, T282S, A284G, P297S, I306V, and S321P.

[0306] SEQ ID NO:1 was cloned into the pCK110900 vector system (see, e.g., U.S. Patent Application Publication No. 2006 / 0195947A1) under the control of the lac promoter. This expression vector also contains a P15a origin of replication and a chloramphenicol resistance gene. The resulting plasmid was transformed into E. coli W3110 using standard methods and expressed as described in Example 2.

[0307] [Example 2] High-throughput screening to identify variants of aminotransferases from Arthrobacter species capable of converting lactol substrates to amines. A) Library preparation of mutant libraries The gene encoding the aminotransferase of Arthrobacter sp. (SEQ ID NO: 1), constructed as described in Example 1, was mutated and the population of mutated DNA molecules was used to transform a suitable E. coli host strain.

[0308] B) High-throughput propagation and expression of mutant libraries Recombinant E. coli colonies carrying genes encoding aminotransferases were picked into 96-well shallow-well microtiter plates containing 180 μL of LB broth, 1% glucose, and 30 μg / mL chloramphenicol (CAM) per well using a Q-PIX Molecular Devices robotic colony picker (Genetix USA, Inc., Boston, MA). The plates were sealed with breathable nylon seals, and the cells were grown overnight at 30°C and 85% humidity with shaking at 200 rpm. A 20 μL aliquot of this culture was then transferred to a 96-deep-well plate containing 380 μL of 2xYT broth supplemented with 100 μL of pyridoxine and 30 μg / mL CAM. The plate was again sealed with breathable nylon seals. After incubating the deep-well plate at 30°C and 85% humidity with shaking at 250 rpm for 2–3 hours, recombinant gene expression in the cultured cells was induced by adding IPTG to a final concentration of 1 mM. Plates were then incubated for 18 hours at 30°C and 85% humidity with shaking at 250 rpm. Cells were pelleted by centrifugation (3738 RCF, 10 minutes, 4°C) and pellets were frozen at -80°C for at least 2 hours.

[0309] C) High-throughput assay screening of mutant libraries to identify improved variants Antibiotic resistant transformants were selected and treated to identify those expressing aminotransferases with improved ability to carry out the following reaction under the desired reaction conditions:

[0310] [ka]

[0311] The cell pellet was thawed at room temperature for 30 to 120 minutes. The pellet was resuspended in 200 µL of lysis solution containing buffer, 1 g / L lysozyme, 0.5 g / L polymyxin B sulfate, and 0.25 g / L pyridoxal 5'-phosphate (PLP). The buffer and pH of the lysis solution varied depending on the screening conditions (see Table 3).

[0312] [Table 3]

[0313] Plates were sealed with breathable nylon seals and then vigorously shaken at room temperature for 2 hours. Cell debris was pelleted by centrifugation (3738 RCF, 10 minutes, 4°C), and the clear supernatant was either assayed directly or stored at 4°C until use. To screen for modified aminotransferases, substrates were prepared as follows: isopropylamine was adjusted to the desired pH with water, and lactol and morpholine were dissolved in DMSO by volume and diluted to the desired cosolvent conditions. For all plates, lactol and isopropylamine in DMSO were added separately to round-bottom shallow-well plates, diluted to the final concentration by adding the dissolution solution, and the reaction was initiated. Plates were heat-sealed with aluminum / polypropylene laminate heat-seal tape at 165°C for 3 seconds and incubated at 650 rpm (INFORS Thermotron) for 20 hours. Conditions for the mutants are summarized below. Reactions were prepared for the analysis described in Example 3.

[0314] [Table 4]

[0315] Variant 1 (SEQ ID NO: 4): 10% lysate, 30 g / L 3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol, 1 M isopropylamine (pH 8.5), 5% DMSO, 5 equivalents morpholine, 45°C Variant 2 (SEQ ID NO: 6): 10% lysate, 30 g / L 3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol, 1 M isopropylamine (pH 8.5), 5% DMSO, 0.05 equivalents morpholine, 45°C Variant 3 (SEQ ID NO: 8): 25% lysate, 50 g / L 3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol, 1 M isopropylamine (pH 8.5), 20% DMSO, 0.05 equivalents morpholine, 45°C Variant 4 (SEQ ID NO: 10): 20% lysate, 50 g / L 3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol, 0.38 M isopropylamine (pH 9.5), 20% DMSO, 0.05 equivalents morpholine, 45°C Variant 5 (SEQ ID NO: 12): 30% lysate, 30 g / L 3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol, 1.6 M isopropylamine (pH 9.5), 30% DMSO, 0.05 equivalents morpholine, 25°C Variant 6 (SEQ ID NO: 14): 20% lysate, 30 g / L 3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol, 1.6 M isopropylamine (pH 8.5), 30% DMSO, 0.05 equivalents morpholine, 25°C Variant 7 (SEQ ID NO: 16): 10% lysis solution, 50 g / L 3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol, 1.6 M isopropylamine (pH 8.5), 30% DMSO, 0.05 equivalents morpholine, 40°C Variant 8 (SEQ ID NO: 18): 5% lysis solution, 50 g / L 3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol, 1.6 M isopropylamine (pH 8.5), 30% DMSO, 0.05 equivalents morpholine, 40°C Variant 9 (SEQ ID NO: 20): 5% lysis solution, 50 g / L 3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol, 1.6 M isopropylamine (pH 8.5), 30% DMSO, 0.05 equivalents morpholine, 40°C

[0316] Running the reaction at a higher pH provides better buffering of the reaction, but also slows the reaction rate. Under these conditions, we were unable to determine undesirable selectivity in the mutants. Unexpectedly, when reactions were run at pH 7.5–8.5, mutants with any change at position I124 (especially I124H—see variant 5) revealed a significant loss of selectivity. Consequently, we challenged the pH tolerance screening conditions by returning to pH 8.5 and increasing the cosolvent.

[0317] [Example 3] Transaminase Powder Production - Shake Flask Method A single microbial colony of E. coli containing a plasmid encoding the aminotransferase of interest was inoculated into 50 mL of Luria-Bertoni broth containing 30 μg / mL chloramphenicol and 1% glucose. Cells were grown overnight (at least 16 hours) in a 30°C incubator with shaking at 250 rpm.

[0318] The culture was diluted into 1000 mL of 2xYT containing 30 μg / mL chloramphenicol (supplemented with 0.1 mM pyridoxine) to an approximate OD600 of 0.2 and grown at 30°C with shaking at 250 rpm. Aminotransferase expression was induced by adding isopropyl βD-thiogalactoside (IPTG) to a final concentration of 1 mM when the culture reached an OD600 of 0.6-0.8. Incubation was then continued overnight (at least 16 hours).

[0319] Cells were harvested by centrifugation (3738 RCF, 20 min, 4°C) and the supernatant was discarded. The pellet was frozen at -80°C for at least 2 hours. The pellet was then thawed and resuspended in 3 mL of 100 mM potassium phosphate buffer (pH 7.5, supplemented with 500 μM PLP) per gram of final pellet mass (e.g., 10 g of frozen pellet suspended in 30 mL of buffer).

[0320] After resuspension, the cells were filtered through a 200 μm mesh and then passed twice through a microfluidizer at 12,000 psig. Cell debris was removed by centrifugation (15,777 RCF, 40 min, 4°C). The supernatants of the clarified lysates were collected, pooled, and lyophilized to obtain a dry powder of crude aminotransferase enzyme.

[0321] [Example 4] Transaminase Powder Production - Fermentation Method An aliquot of frozen working stock (E. coli containing a plasmid carrying the aminotransferase gene of interest) was removed from the freezer and allowed to thaw at room temperature. 300 μL of this working stock was used to inoculate the primary seed stage of 250 mL of M9YE broth (1.0 g / L ammonium chloride, 0.5 g / L sodium chloride, 6.0 g / L disodium monohydrogen phosphate, 3.0 g / L potassium dihydrogen phosphate, 2.0 g / L PROCELYS SPRINGER 0251 yeast extract, 1 L deionized water) containing 30 μg / mL chloramphenicol and 1% glucose in a 1 L flask and grown at 37°C with shaking at 200 rpm. When the culture reached an OD600 of 0.5-1.0, the flask was removed from the incubator and immediately used to inoculate the secondary seed stage.

[0322] The secondary seed stage was carried out in a 5 L bench-scale fermentor using 4 L of growth medium (0.88 g / L ammonium sulfate, 0.98 g / L trisodium citrate dihydrate; 12.5 g / L dipotassium hydrogen phosphate, 6.25 g / L potassium dihydrogen phosphate, 3.3 g / L PROCELYS SPRINGER 0251 yeast extract, 0.083 g / L ferric ammonium citrate, 0.5 ml / L polypropylene glycol antifoam and 8.3 ml / L trace element solution, 1 L / L treated water). The trace element solution contained 2 g / L calcium chloride dihydrate, 2.2 g / L zinc sulfate heptahydrate, 0.5 g / L manganese sulfate monohydrate, 1 g / L copper sulfate pentahydrate, 0.1 g / L ammonium molybdate tetrahydrate, and 0.02 g / L sodium tetraborate decahydrate in 1 L / L deionized water. The growth medium was sterilized at 121°C for 40 minutes. After sterilization, 40 ml / L of feed stock solution was added (feed stock solution contained 12 g / L ammonium sulfate, 5.1 g / L magnesium sulfate heptahydrate, 500 g / L dextrose monohydrate, and 1 L / L treated water, sterilized at 121°C for 30 minutes). Fermenters were inoculated with 2 ml of primary seed at an OD600 of 0.5-1.0, supplemented with 30 μg / ml chloramphenicol, and incubated at 37°C, 300 rpm, and 0.5 vvm aeration. When the culture reached an OD600 of 0.5-1.0, the secondary seed was immediately transferred to the final stage fermentation.

[0323] Final stage fermentations were carried out at bench scale in 10 L fermentors using 6 L of growth medium (0.88 g / L ammonium sulfate, 0.98 g / L trisodium citrate dihydrate; 12.5 g / L dipotassium hydrogen phosphate, 6.25 g / L potassium dihydrogen phosphate, 3.3 g / L Procelys Springer 0251 yeast extract, 0.083 g / L ferric ammonium citrate, 0.5 ml / L polypropylene glycol antifoam and 8.3 ml / L trace element solution, 1 L / L treated water). The trace element solution contained 2 g / L calcium chloride dihydrate, 2.2 g / L zinc sulfate heptahydrate, 0.5 g / L manganese sulfate monohydrate, 1 g / L copper sulfate pentahydrate, 0.1 g / L ammonium molybdate tetrahydrate, and 0.02 g / L sodium tetraborate decahydrate in 1 L / L deionized water. The growth medium was sterilized at 121°C for 40 minutes. After sterilization, the growth medium was supplemented with 0.035 g / L pyridoxine hydrochloride and 40 ml / L of feed stock solution (12 g / L ammonium sulfate, 5.1 g / L magnesium sulfate heptahydrate, 500 g / L dextrose monohydrate, 1 L / L treated water, and sterilized at 121°C for 30 minutes).

[0324] The fermentor was inoculated with 500 ml of secondary seed at an OD600 of 0.5-1.0 and incubated at 30°C with 1.5 vvm of aeration. Dissolved oxygen was controlled at 30% by variable-speed agitation, and pH was maintained at 7.0 by adding 17.5% v / v ammonium hydroxide solution. Culture growth was maintained by adding a feed stock solution (12 g / L ammonium sulfate, 5.1 g / L magnesium sulfate heptahydrate, 500 g / L dextrose monohydrate, 1 L / L treated water, sterilized at 121°C for 30 min). After the culture reached an OD600 of 80 + / - 10, isopropyl-β-D-thiogalactoside (IPTG) was added to a final concentration of 1 mM. Fermentation was continued for an additional 18 h. At harvest, the culture was cooled to 8°C. Cells were harvested by centrifugation at 5000 G for 40 minutes in a Sorvall RC12BP centrifuge at 4° C. The harvested cell pellet was then frozen at −80° C. and stored until downstream processing and recovery.

[0325] From collection to freeze-dried powder The fermentation broth was buffered to pH 7.3-7.5 with 6.88 g / L dipotassium phosphate and 1.43 g / L monopotassium phosphate. Once the salts had dissolved and the desired pH range was reached, the broth was homogenized by mechanical lysis. Polyethyleneimine sulfate (Mn approximately 60,000, Mw 750,000) was added to the lysate to a final concentration of less than 0.6% w / w (the required amount of flocculant was determined by offline titration). The flocculant-treated lysate was clarified by centrifugation, and the solid phase was discarded, retaining the clarified lysate supernatant. The clarified lysate was concentrated approximately 10-fold by tangential flow filtration through a 30 kDa MWCO membrane. Pyridoxal 5'-phosphate (PLP) was then added to the concentrate in an amount proportional to the starting fermentation broth volume, 0.08 g PLP per L of fermentation broth. The PLP concentrate was then lyophilized. The freeze-dried material was then ground to a homogenous powder.

[0326] [Example 5] A high-throughput assay to identify variants of Arthrobacter species aminotransferases capable of converting lactol substrates to amines. A) Achiral analysis After overnight incubation, the plates were removed from the incubator, the seals removed, and the reactions diluted with an equal volume of acetonitrile. The plates were heat-sealed with aluminum / polypropylene laminate heat-seal tape at 165°C for 4 seconds, shaken for 10 minutes, and then centrifuged at 3738 RCF for 10 minutes to pellet debris. 20 μL of the 2x-diluted supernatant per well was further diluted either 10x or 20x (for 30 g / L or 50 g / L reactions, respectively) with 50% aqueous acetonitrile in a new shallow-well polypropylene plate. The plates were resealed, mixed, and analyzed by UPLC as described below.

[0327] B) Achiral UPLC method for qualitative determination of amine products The enzymatic conversion of the lactol substrate (3-(4-bromophenyl)tetrahydro-2H-pyran-2-ol) to the amine product (see Scheme 1) was determined using an Agilent 1290 UPLC equipped with an Agilent ZORBAX SB-C18 RRHD column (3.0 × 50 mm, 1.8 μm) using a gradient of 0.05% trifluoroacetic acid in water (mobile phase A) and 0.05% trifluoroacetic acid in acetonitrile (mobile phase B) at a flow rate of 2 mL / min and a column temperature of 60 °C. The gradient profile is shown below.

[0328] [Table 5]

[0329] Compound elution was monitored at 214 nm and showed that DMSO eluted at approximately 0.125 minutes, followed by the amine product at approximately 0.55 minutes and the lactol at approximately 0.89 minutes. Four minor impurity peaks eluted between the starting material and the amine product.

[0330] C) Derivatization of the amine product with MARFEY'S REAGENT: 1% MARFEY'S REAGENT (10 g / L in acetonitrile) was prepared. 20 μL per well of the 2-fold diluted supernatant prepared as described in Example 3B was transferred to a new 96-deep-well plate, followed by the addition of 30 μL of 1 M NaHCO3, followed by 200 μL of 1% MARFEY'S REAGENT. The plate was sealed and incubated at 40°C and 850 rpm (INFORS Thermotron) for 1 hour. Derivatization was quenched by adding an equal volume of 1:9 2N HCl:acetonitrile. The plate was mixed for 5 minutes and then centrifuged at 3738 RCF for 10 minutes to pellet debris. The seal was removed, and 40 μL of the supernatant was transferred to 160 μL of acetonitrile in a new shallow-well polypropylene plate. The plate was resealed, mixed, and analyzed by UPLC.

[0331] D) Chiral UPLC method for qualitative determination of amine product selectivity High-throughput quantification of product enantiomers was determined using an Agilent 1290 UPLC equipped with an Agilent ZORBAX SB-C18 RRHD column (3.0 × 50 mm, 1.8 μm) using a gradient of 0.05% trifluoroacetic acid in water (mobile phase A) and 0.05% trifluoroacetic acid in acetonitrile (mobile phase B) at a flow rate of 2 mL / min and a column temperature of 60°C. The gradient profile is shown below.

[0332] [Table 6]

[0333] The elution of the compounds was monitored at 340 nm, with the desired product eluting at approximately 4.89 minutes and the undesired product eluting at approximately 5.15 minutes.

[0334] [Table 7]

[0335] [Table 8]

[0336] Table 7 provides that all variants, except for variant 5, which has a mutation to position I124, showed increased activity while maintaining selectivity of over 99% ee. This was surprising and unexpected to the inventors of the present disclosure, as any change at position I124 (see in particular I124H-variant 5) resulted in a significant loss of selectivity. This loss of selectivity was present in all sample variants tested when the I124H substitution was introduced (see Table 9).

[0337] [Table 9]

[0338] Screening conditions for high-throughput characterization of enantioselectivity (Table 8) differed from those described in Example 2. Because reactions were not driven to completion, enantioselectivities were not quantitative and differences can be seen in the values ​​reported in Table 7. Reactions were sufficient to detect perturbations in enantioselectivity, allowing alternative variants to be selected and advanced.

[0339] [ka]

[0340] [Example 6] Chemical process. 1) Preparation of Compound 4 Stage 1A: To a stirred solution of isopropylamine hydrochloride in water at 25°C, add 1 M aqueous sodium hydroxide and adjust the pH to 8.5. Add lyophilized transaminase and stir for at least 15 minutes. Add a DMSO solution of Compound 1 to the mixture. Adjust the pH to 8.6 using 1 M aqueous sodium hydroxide. The mixture is then heated to 44°C and stirred for 21 hours. Cool the reaction to 25°C, adjust the pH to pH 2 using concentrated aqueous HCl, and stir for 1.5 hours to denature the enzyme.

[0341] Stage 1B: To the stirred solution from Stage 1A, K2HPO4 is added, followed by liquid BOC anhydride. Stirring is continued. Water is added, followed by TBME, which is then stirred for at least 10 minutes before being allowed to settle. The separated aqueous layer is removed. Cellulose is added, stirred for at least 15 minutes, and then filtered. The enzyme filter cake is washed with TBME before discarding the solids. The combined organics are washed twice with water, removing the separated aqueous layer each time. The organic solution is concentrated under vacuum to approximately 5 volumes, and then the solvent is switched to acetonitrile.

[0342] Stage 1C: DIPEA is added to the solution, cooled to 0°C, and mesyl chloride is added over 1 hour, maintaining the temperature of the contents below 5°C. After the final addition, the mixture is stirred for at least 30 minutes. Isopropanol is added, the temperature of the contents is adjusted to 20°C, and water is added over at least 10 minutes. The mixture is seeded with compound 4 and aged for 2 hours. Water is added over at least 1.5 hours, stirred for at least 2 hours, filtered, and dried. The filter cake is washed with two successive cake washes of isopropanol / water (1:1 v / v), followed by pre-chilled isopropanol and finally heptane. The cake is drained and dried in a vacuum oven at 55°C.

[0343] 2) Preparation of Compound 8 Stage 2A: Preparation of Compound 5. Toluene is charged to a reactor at 25±5° C. with stirring. Compound 7 is charged to the reactor, followed by potassium tert-butoxide. The mixture is stirred at 25±5° C. until the reaction is complete.

[0344] Stage 2B: Preparation of Compound 8. Potassium phosphate trihydrate (K3PO4), compound 6, and copper(I) bromide are charged to a reactor, followed by trans-N1,N2-dimethylcyclohexane-1,2-diamine (DMCyDA). The mixture is heated to reflux (105±10°C) and stirred at this temperature until the reaction is complete. The mixture is cooled to 45±5°C, and then ethyl acetate is charged, followed by water and 20% v / v aqueous NH3. The layers are separated, and then water and 20% v / v aqueous NH3 are added to the organic layer. The layers are separated. The organic layer is concentrated by distillation and then seeded at 45±5°C. The mixture is aged, and then heptane is added. The resulting slurry is cooled to 0±5°C and aged for at least 16 hours. The mixture is filtered, and then the cake is washed with a toluene / EtOAc / heptane mixture. The product is dried under vacuum at 60±5° C. to give compound 8.

[0345] 3) Preparation of Compound 9 Methanesulfonic acid (MsOH) is added to a slurry of compound 8 in toluene over at least 3 hours, maintaining the temperature below 40°C. The resulting biphasic solution is stirred at 38±5°C for 3 hours. The contents of the reactor are then cooled to 0±5°C. Water is added while maintaining the temperature below 30°C. The aqueous and organic phases are then allowed to separate. The aqueous phase is washed with toluene and then filtered. A portion of the solution of p-toluenesulfonic acid is added to the filtrate, followed by compound 9 seeds (prepared by the same process described herein). The remaining p-toluenesulfonic acid solution is added at 20±5°C, and the slurry is then cooled and stirred at 0±5°C for at least 4 hours. The slurry is filtered and washed four times with water / MeTHF (9:1 v / v) cake washes, and the filter cake is then dried under vacuum at NGT 40°C.

[0346] 4) Purification of Compound 9 Compound 9 is added to DMSO and heated to 33°C to dissolve. This solution is poured into water over 2 hours with stirring. After holding for 1 hour, the suspension is wet-milled. The resulting suspension is heated to 65°C, held for 7 hours, and then cooled to 5°C. The slurry is removed and washed with water. The material is dried to 40°C with periodic stirring.

[0347] An impurity (compound 10) is produced as a by-product in the stage 3 chemical reaction when methanesulfonic acid is added to a slurry of compound 8 in toluene.

[0348] [ka]

[0349] Compound 10 is controlled to be 0.10% w / w or less in the niraparibut tosylate monohydrate drug substance.

[0350] [Example 7] Structural characterization 7.1. Compound 1

[0351] [ka]

[0352] The structure of compound 1, a crystalline solid, was established by mass and NMR spectroscopy. LCMS (2.0 min ASAP): Rt = N / A, [MH] - = 255.0026 1 H NMR (chloroform-d, 700 MHz): δ (ppm) 7.41 - 7.48 (m, 3H), 7.20 (d, J=8.4 Hz, 1H), 7.14 (d, J=8.4 Hz, 2H), 6.45 - 6.46 (m, 1H), 5.20 (d, J=2.7 Hz, 1H), 4.76 (d, J=7.9 Hz, 1H), 4.10 (td, J=4.5, 2.5 Hz, 1H), 3.61 - 3.66 (m, 1H), 3.42 - 3.42 (m, 1H), 2.95 (dt, J=12.4, 3.1 Hz, 1H), 2.59 (ddd, J=11.4, 7.7, 4.1 Hz, 1H), 1.98 - 2.07 (m, 1H), 1.78 - 1.82 (m, 1H), 1.75 (br s, 1H), 1.70 - 1.77 (m, 1H), 1.68 - 1.75 (m, 1H), 1.62 - 1.66 (m, 1H).

[0353] 7.2. Compound 11

[0354] [ka]

[0355] The structure of compound 11, a crystalline solid, was established by mass and NMR spectroscopy. LCMS (5.5 min trifluoroacetic acid): Rt = 2.33 min, [M-H] + = 436.0786 1H NMR (DMSO-d6, 700 MHz): δ (ppm) 7.46 - 7.50 (m, J=8.3 Hz, 2H), 7.14 - 7.16 (m, J=8.4 Hz, 2H), 6.80 (br t, J=5.7 Hz, 1H), 4.09 - 4.14 (m, 2H), 3.12 (s, 4H), 3.02 - 3.09 (m, 1H), 2.70 - 2.77 (m, 1H), 1.67 - 1.78 (m, 1H), 1.46 - 1.56 (m, 2H), 1.36 - 1.45 (m, 1H), 1.32 (s, 9H)

[0356] 7.3. Compound 12

[0357] [ka]

[0358] The structure of compound 12, a crystalline solid, was established by mass and NMR spectroscopy. LCMS (5.5 min trifluoroacetic acid): Rt = 3.01 min, [M-H] + = 477.2869 1H NMR (chloroform-d, 700 MHz): δ (ppm) 9.90 (br s, 1H), 8.54 (s, 1H), 8.21 (d, J=7.1 Hz, 1H), 7.91 (d, J=8.3 Hz, 1H), 7.83 (d, J=8.4 Hz, 2H), 7.44 (d, J=8.4 Hz, 2H), 7.26 - 7.30 (m, 1H), 4.22 (br d, J=12.2 Hz, 1H), 4.16 (br dd, J=13.4, 1.0 Hz, 1H), 2.84 - 2.87 (m, 1H), 2.82 - 2.89 (m, 1H), 2.78 - 2.85 (m, 1H), 2.11 (br d, J=12.7 Hz, 1H), 1.86 (dt, J=13.3, 2.9 Hz, 1H), 1.72 (br dd, J=11.9, 3.0 Hz, 1H), 1.64 (dt, J=12.9, 3.7 Hz, 1H), 1.60 (s, 9H), 1.50 (s, 9H).

[0359] 7.4. Compound 13

[0360] [ka]

[0361] The structure of compound 13, a crystalline solid, was established by mass and NMR spectroscopy. LCMS (50 min trifluoroacetic acid): Rt = 17.0 min, [M-H] + = 321.1711 1H NMR (DMSO-d6, 700 MHz): δ (ppm) 7.46 - 7.50 (m, J=8.3 Hz, 2H), 7.14 - 7.16 (m, J=8.4 Hz, 2H), 6.80 (br t, J=5.7 Hz, 1H), 4.09 - 4.14 (m, 2H), 3.12 (s, 4H), 3.02 - 3.09 (m, 1H), 2.70 - 2.77 (m, 1H), 1.67 - 1.78 (m, 1H), 1.46 - 1.56 (m, 2H), 1.36 - 1.45 (m, 1H), 1.32 (s, 9H).

[0362] A brief description of arrays SEQ ID NO: 1: DNA sequence of an engineered variant of Arthrobacter transaminase (control) ATGAGCTTCTCACATGACACCCCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGCCGTCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTTATACTTCTGACGCTACCTACACCGTCTTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGGGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTATTCGTTTGATCCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTATCGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGGAGCGTGACATCACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATCGTACCGTTTGACCGCATCCGTGACGGTGTTCACCTGATGGTTGCTCAGTCAGTTCGTCGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGATCCGTGCAATTCAGGAAACCCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCTGACATCACCCCGGCTGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACTCTATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 2: Protein sequence of an engineered variant of Arthrobacter transaminase (control) MSFSHDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPPSEARISIFDQGFYTSDATYTVFHVWNGNAFRLGDHIERLFSNAESIRLIPPLTQDEVKEIALELVAKTELREAIVIVSITRGYSSTPWERDITKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLIRAIQETHDRGFELPLLLDFDNLLAEGPGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITPAELYDADEVLGCSTGGGVWPFVSVDGNSISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY SEQ ID NO: 3: DNA sequence of an engineered variant of Arthrobacter transaminase (variant 1) ATGAGCTTCTCACATGACACCCCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGCCGTCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTTGTACTTCTGACGCTACCTACACCGTCTTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGGGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTATTCGTTTGATCCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTATCGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGGAGCGTGACATCACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATCGTACCGTTTGACCGCATCCGTGACGGTGTTCACCTGATGGTTGCTCAGTCAGTTCGTCGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGATCCGTGCAATTCAGGAAACCCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCTGACATCACCCCGGCTGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACTCTATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 4: Protein sequence of an engineered variant of Arthrobacter transaminase (variant 1) MSFSHDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPPSEARISIFDQGFCTSDATYTVFHVWNGNAFRLGDHIERLFSNAESIRLIPPLTQDEVKEIALELVAKTELREAIVIVSITRGYSSTPWERDITKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLIRAIQETHDRGFELPLLLDFDNLLAEGPGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITPAELYDADEVLGCSTGGGVWPFVSVDGNSISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY SEQ ID NO: 5: DNA sequence of an engineered variant of Arthrobacter transaminase (variant 2) ATGAGCTTCTCACATGACACCCCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGAGATCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTTGTACTTCTGACGCTACCTACACCTGCTTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGGGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTTGTCGTTTGATCCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTATCGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGGAGCGTGACATCACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATCGTACCGTTTGACCGCATCCGTGACGGTGTTCACCTGATGGTTGCTCAGTCAGTTCGACGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGAGACGTGCAATTCAGGAAACCCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCCGACATCACCCCGGCTGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACCGCATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 6: Protein sequence of an engineered variant of Arthrobacter transaminase (variant 2) MSFSHDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPRSEARISIFDQGFCTSDATYTCFHVWNGNAFRLGDHIERLFSNAESCRLIPPLTQDEVKEIALELVAKTELREAIVIVSITRGYSSTPWERDITKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLRRAIQETHDRGFELPLLLDFDNLLAEGPGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITPAELYDADEVLGCSTGGGVWPFVSVDGNRISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY SEQ ID NO: 7: DNA sequence of an engineered variant of Arthrobacter transaminase (variant 3) ATGAGCTTCTCACATGACACCCCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGAGATCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTAATACTTCTGACGCTACCTACACCTGCTTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGGGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTTGTCGTTTGATCCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTATCGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGGAGCGTGACATCACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATCGTACCGTTTGACCGCATCCGTGACGGTGTTCACCTGATGGTTGCTCAGTCAGTTCGACGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGAGACGTGCAATTCAGGAAACCCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCCGACATCACCCCGGCTGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACCGCATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 8: Protein sequence of an engineered variant of Arthrobacter transaminase (variant 3) MSFSHDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPRSEARISIFDQGFNTSDATYTCFHVWNGNAFRLGDHIERLFSNAESCRLIPPLTQDEVKEIALELVAKTELREAIVIVSITRGYSSTPWERDITKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLRRAIQETHDRGFELPLLLDFDNLLAEGPGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITPAELYDADEVLGCSTGGGVWPFVSVDGNRISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY SEQ ID NO: 9: DNA sequence of an engineered variant of Arthrobacter transaminase (variant 4) ATGAGCTTCTCACATGACACCCCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGAGATCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTAATACTTCTGACGCTACCTACACCACATTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGGGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTTGTCGTTTGATCCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTATCGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGGAGCGTGACATCACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATCGTACCGTTTGACCGCATCCGTGACGGTGTTCACCTGATGGTTGCTCAGTCAGTTCGACGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGAGACGTGCAATTCAGGAAACCCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCCGACATCACCCCGGCCGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACCGCATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 10: Protein sequence of an engineered variant of Arthrobacter transaminase (variant 4) MSFSHDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPRSEARISIFDQGFNTSDATYTTFHVWNGNAFRLGDHIERLFSNAESCRLIPPLTQDEVKEIALELVAKTELREAIVIVSITRGYSSTPWERDITKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLRRAIQETHDRGFELPLLLDFDNLLAEGPGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITPAELYDADEVLGCSTGGGVWPFVSVDGNRISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY SEQ ID NO: 11: DNA sequence of an engineered variant of Arthrobacter transaminase (variant 5) ATGAGCTTCTCACATGACACCCCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGAGATCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTAATACTTCTGACGCTACCTACACCACATTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGGGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTTGTCGTTTGATCCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTCATGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGGAGCGTGACATCACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATCGTACCGTTTGACCGCATCCGTGACGGTGTTCACCTGATGGTTGCTCAGTCAGTTCGACGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGAGACGTGCAATTCAGGAAACCCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCCGACATCACCCCGGCCGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACCGCATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 12: Protein sequence of an engineered variant of Arthrobacter transaminase (variant 5) MSFSHDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPRSEARISIFDQGFNTSDATYTTFHVWNGNAFRLGDHIERLFSNAESCRLIPPLTQDEVKEIALELVAKTELREAIVHVSITRGYSSTPWERDITKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLRRAIQETHDRGFELPLLLDFDNLLAEGPGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITPAELYDADEVLGCSTGGGVWPFVSVDGNRISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY SEQ ID NO: 13: DNA sequence of an engineered variant of Arthrobacter transaminase (variant 6) ATGAGCTTCTCACATGACACCCCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGAGATCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTATGACTTCTGACGCTACCTACACCACATTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGGGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTTGTCGTTTGATCCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTATCGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGGAGCGTGACATCACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATCGTACCGTTTGACCGCATCCGTGACGGTGTTCACCTGATGGTTGCTCAGTCAGTTCGACGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGAGACGTGCAATTCAGGAAACCCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCCGACATCACCCCGGCCGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACCGCATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 14: Protein sequence of an engineered variant of Arthrobacter transaminase (variant 6) MSFSHDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPRSEARISIFDQGFMTSDATYTTFHVWNGNAFRLGDHIERLFSNAESCRLIPPLTQDEVKEIALELVAKTELREAIVIVSITRGYSSTPWERDITKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLRRAIQETHDRGFELPLLLDFDNLLAEGPGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITPAELYDADEVLGCSTGGGVWPFVSVDGNRISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY SEQ ID NO: 15: DNA sequence of an engineered variant of Arthrobacter transaminase (variant 7) ATGAGCTTCTCACATGACACCCCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGAGATCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTATGACTTCTGACGCTACCTACACCACATTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGGGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTTGTCGTTTGATCCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTATCGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGACACGTGACATCACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATCGTACCGTTTGACCGCATCCGTGACGGTGTTCACCTGATGGTTGCTCAGTCAGTTCGACGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGAGACGTGCAATTCAGGAAACCCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCCGACATCACCCCGGCCGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACCGCATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 16: Protein sequence of an engineered variant of Arthrobacter transaminase (variant 7) MSFSHDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPRSEARISIFDQGFMTSDATYTTFHVWNGNAFRLGDHIERLFSNAESCRLIPPLTQDEVKEIALELVAKTELREAIVIVSITRGYSSTPWTRDITKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLRRAIQETHDRGFELPLLLDFDNLLAEGPGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITPAELYDADEVLGCSTGGGVWPFVSVDGNRISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY SEQ ID NO: 17: DNA sequence of an engineered variant of Arthrobacter transaminase (variant 8) ATGAGCTTCTCACATGACACCCCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGAGATCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTATGACTTCTGACGCTACCTACACCACATTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGGGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTTGTCGTTTGATCCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTATTGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGACACGCGACCAAACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATTGTACCGTTTGACCGCATCCGTGACGGTGTTCATCTGATGGTTGCTCAGTCAGTTCGACGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGAGACGTGCAACCCAGGAATGTCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCCGACATCACCCCGGCCGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACCGCATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 18: Protein sequence of an engineered variant of Arthrobacter transaminase (variant 8) MSFSHDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPRSEARISIFDQGFMTSDATYTTFHVWNGNAFRLGDHIERLFSNAESCRLIPPLTQDEVKEIALELVAKTELREAIVIVSITRGYSSTPWTRDQTKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLRRATQECHDRGFELPLLLDFDNLLAEGPGGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITPAELYDADEVLGCSTGGGVWPFVSVDGNRISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY SEQ ID NO: 19: DNA sequence of an engineered variant of Arthrobacter transaminase (variant 9) ATGAGCCACTCAATTGACACACCTGAAATCGTTTACACCCACGACACCGGTCTGGACTATATCACCTACTCTGACTACGAACTGGACCCGGCTAACCCGCTGGCTGGTGGTGCTGCTTGGATCGAAGGTGCTTTCGTTCCGAGATCTGAAGCTCGTATCTCTATCTTCGACCAGGGTTTTATGACTTCTGACGCTACCTACACCACATTCCACGTTTGGAACGGTAACGCTTTCCGTCTGGGTGACCACATCGAACGTCTGTTCTCTAATGCGGAATCTTGTCGTTTGGAGCCGCCGCTGACCCAGGACGAAGTTAAAGAGATCGCTCTGGAACTGGTTGCTAAAACCGAACTGCGTGAAGCGATCGTTATTGTTTCTATCACCCGTGGTTACTCTTCTACCCCATGGACACGCGACCAAACCAAACATCGTCCGCAGGTTTACATGTATGCTGTTCCGTACCAGTGGATTGTACCGTTTGACCGCATCCGTGACGGTGTTCATCTGATGGTTGCTCAGTCAGTTCGACGTACACCGCGTAGCTCTATCGACCCGCAGGTTAAAAACTTCCAGTGGGGTGACCTGAGACGTGCAACCCAGGAATGTCACGACCGTGGTTTCGAGTTACCGCTGCTGCTGGACTTCGACAACCTGCTGGCTGAAGGTCCGGGTTTCAACGTTGTTGTTATCAAAGACGGTGTTGTTCGTTCTCCGGGTCGTGCTGCTCTGCCGGGTATCACCCGTAAAACCGTTCTGGAAATCGCTGAATCTCTGGGTCACGAAGCTATCCTGGCCGACATCACCGTGGCCGAACTGTACGACGCTGACGAAGTTCTGGGTTGCTCAACCGGTGGTGGTGTTTGGCCGTTCGTTTCTGTTGACGGTAACCGCATCTCTGACGGTGTTCCGGGTCCGGTTACCCAGTCTATCATCCGTCGTTACTGGGAACTGAACGTTGAACCTTCTTCTCTGCTGACCCCGGTACAGTAC SEQ ID NO: 20: Protein sequence of an engineered variant of Arthrobacter transaminase (variant 9) MSHSIDTPEIVYTHDTGLDYITYSDYELDPANPLAGGAAWIEGAFVPRSEARISIFDQGFMTSDATYTTFHVWNGNAFRLGDHIERLFSNAESCRLEPPLTQDEVKEIALELVAKTELREAIVIVSITRGYSSTPWTRDQTKHRPQVYMYAVPYQWIVPFDRIRD GVHLMVAQSVRRTPRSSIDPQVKNFQWGDLRRATQECHDRGFELPLLLDFDNLLAEGPGGFNVVVIKDGVVRSPGRAALPGITRKTVLEIAESLGHEAILADITVAELYDADEVLGCSTGGGVWPFVSVDGNRISDGVPGPVTQSIIRRYWELNVEPSSLLTPVQY

Claims

1. A modified transaminase polypeptide or a functional fragment thereof comprising an amino acid sequence having at least 80%, 85%, 90%, or 95% sequence identity with the amino acid sequence set forth in SEQ ID NO: 2, wherein the amino acid sequence comprises the characteristic that X124 is isoleucine (I).

2. 2. The modified transaminase polypeptide of claim 1, wherein the amino acid sequence comprises the characteristic that at least one amino acid selected from the group consisting of amino acids X3, X5, X48, X61, X69, X94, X97, X137, X140, X196, X199, X202, X269, and X297 is substituted with an amino acid different from the amino acid identified at the corresponding amino acid position in SEQ ID NO:

2.

3. The modified transaminase polypeptide of any one of claims 1 to 2, wherein the polypeptide has improved enzymatic properties compared to SEQ ID NO:

2.

4. The modified transaminase polypeptide of any one of claims 1 to 3, wherein X61 is cysteine ​​(C), asparagine (N), or methionine (M), and / or X69 is cysteine ​​(C) or threonine (T).

5. 5. The modified transaminase polypeptide of any one of claims 1 to 4, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X48 is substituted with arginine (R), X94 is substituted with C, X196 is substituted with R, X297 is substituted with R, and X137 is substituted with T.

6. 6. The modified transaminase polypeptide of any one of claims 1 to 5, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X140 is substituted with glutamine (Q), X199 is substituted with T, and X202 is substituted with C.

7. 7. The modified transaminase polypeptide of any one of claims 1 to 6, wherein the amino acid sequence comprises at least one feature selected from the group consisting of: X3 is substituted with histidine (H), X5 is substituted with I, X97 is substituted with glutamic acid (E), and X269 is substituted with C or valine (V).

8. The modified transaminase polypeptide of any one of claims 1 to 7, wherein X61 is C.

9. 8. The modified transaminase polypeptide of any one of claims 1 to 7, wherein the amino acid sequence comprises at least the following features: X48 is substituted with R, X61 is C, X69 is C, X94 is substituted with C, X196 is substituted with R, and X297 is substituted with R.

10. 8. The modified transaminase polypeptide of any one of claims 1 to 7, wherein the amino acid sequence comprises at least the following features: X48 is substituted with R, X61 is N, X69 is C, X94 is substituted with C, X196 is substituted with R, and X297 is substituted with R.

11. 8. The modified transaminase polypeptide of any one of claims 1 to 7, wherein the amino acid sequence comprises at least the following features: X48 is substituted with R, X61 is N, X69 is T, X94 is substituted with C, X196 is substituted with R, and X297 is substituted with R.

12. 8. The modified transaminase polypeptide of any one of claims 1 to 7, wherein the amino acid sequence comprises at least the following features: X48 is substituted with R, X61 is M, X69 is T, X94 is substituted with C, X196 is substituted with R, and X297 is substituted with R.

13. 8. The modified transaminase polypeptide of any one of claims 1 to 7, wherein the amino acid sequence comprises at least the following features: X48 is substituted with R, X61 is M, X69 is T, X94 is substituted with C, X137 is substituted with T, X196 is substituted with R, and X297 is substituted with R.

14. 8. The modified transaminase polypeptide of any one of claims 1 to 7, wherein the amino acid sequence comprises at least the following features: X48 is substituted with R, X61 is M, X69 is T, X94 is substituted with C, X137 is substituted with T, X140 is substituted with Q, X196 is substituted with R, X199 is substituted with T, X202 is substituted with C, and X297 is substituted with R.

15. 8. The modified transaminase polypeptide of any one of claims 1-7, wherein the amino acid sequence comprises at least the following features: X3 is substituted with H, X5 is substituted with I, X48 is substituted with R, X61 is M, X69 is T, X94 is substituted with C, X97 is substituted with E, X137 is substituted with T, X140 is substituted with Q, X196 is substituted with R, X199 is substituted with T, X202 is substituted with C, X269 is substituted with V, and X297 is substituted with R.

16. The modified transaminase polypeptide of claim 1, wherein the amino acid sequence corresponds to the sequence of SEQ ID NO: 4, 6, 8, 10, 14, 16, 18 or 20.

17. A polynucleotide encoding the modified transaminase polypeptide of any one of claims 1 to 16.

18. 18. The polynucleotide of claim 17, wherein the polynucleotide sequence corresponds to the sequence of SEQ ID NO: 3, 5, 7, 9, 13, 15, 17 or 19.

19. An expression vector comprising the polynucleotide of claim 17 or 18.

20. A host cell comprising the polynucleotide of any one of claims 17 to 18 or the expression vector of claim 19.

21. Asymmetric compounds of formula I: 【Chemical 1】 (In the formula: R 1 represents a leaving group, halogen, protected amino group, NO 2 or OH or a protected form thereof; R 2 is H; R 3 Is -COOR 5 , -CH 2 R 6 or a protected aldehyde; or R 2 and R 3 is combined 【Chemistry 2】 is formed; R 4 is H or an amine protecting group; R 5 is C 1~6 Alkyl, C 3~6 Cycloalkyl, C 4~10 heterocyclyl, aryl, or heteroaryl; R 6 is a leaving group or OH or a protected form thereof) 1. A method for preparing a compound of formula II: 【Chemistry 3】 (In the formula, R 1 ' denotes a leaving group, halogen, protected amino group, NO 2 or OH or a protected form thereof; R 2 ' is an aldehyde or aldehyde equivalent; R 3 ' is -COOR 5 , -CH 2 R 6 or a protected aldehyde; or R 2 ' and R 3 ' is a combination 【Chemistry 4】 (where, * represents the point of attachment) is formed) in the presence of a coenzyme and an amino donor with the modified transaminase polypeptide of any one of claims 1 to 16.

22. 22. The method of claim 21, wherein the method provides a compound of formula I having an enantiomeric excess of at least 95% or at least 99%.

23. 23. The method of any one of claims 21 to 22, wherein the modified transaminase polypeptide is selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 4, 6, 8, 10, 14, 16, 18, or 20.

24. 24. The method of claim 23, wherein the modified transaminase polypeptide is selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 16, 18, and 20.

25. The method according to any one of claims 21 to 24, wherein the coenzyme is pyridoxal phosphate (PLP).

26. 26. The method of any one of claims 21 to 25, wherein the amino donor is isopropylamine.

27. R 2 and R 3 is combined 【Chemistry 5】 and R 4 The method of any one of claims 21 to 26, wherein is H.

28. R 2 and R 3 is combined 【Chemistry 6】 and R 4 28. The method of claim 27, wherein: is H.

29. R 1 and R 1 29. The method of any one of claims 21 to 28, wherein each ' is Br.

30. R 2 is H and R 3 is CH 2 R 6 and R 4 is H and R 6 The method of any one of claims 21 to 29, wherein is OH.

31. The compound of formula I is 【Chemistry 7】 The method of any one of claims 21 to 30, selected from the group consisting of:

32. The compound of formula II is 【Chemistry 8】 The method according to any one of claims 21 to 31, wherein

33. The compound of formula I is 【Chemistry 9】 and the compound of formula II is 【Chemistry 10】 The method according to any one of claims 21 to 32, wherein

34. Asymmetric compounds of formula I: 【Chemistry 11】 (In the formula, R 1 is Br; R 2 is H; R 3 is -CH 2 R 6 and R 4 is H or an amine protecting group; R 6 is OH or a protected form thereof) 1. A process for preparing a compound of formula IIa: 【Chemistry 12】 with the modified transaminase polypeptide of any one of claims 1 to 16 in the presence of a coenzyme and an amino donor to produce a compound of formula Ia: 【Chemistry 13】 Optionally, the method comprises providing an NH 2 to protect the groups to give a compound of formula Ib: 【Chemistry 14】 (In the formula, PG 1 is an amine protecting group) providing Further optionally, the OH group of formula Ib can be protected to give a compound of formula Ic: 【Chemistry 15】 (In the formula, PG 2 is a hydroxyl protecting group) A method comprising the step of providing:

35. 35. The method of claim 34, wherein the amine protecting group is selected from the group consisting of formyl, acetyl (Ac), trifluoroacetyl, benzyl (Bn), benzoyl (Bz), carbamate, benzyloxycarbonyl ("CBZ"), p-methoxybenzylcarbonyl (Moz or MeOZ), tert-butoxycarbonyl ("Boc"), trimethylsilyl ("TMS"), 2-trimethylsilyl-ethanesulfonyl ("SES"), trityl and substituted trityl groups, allyloxycarbonyl, 9-fluorenylmethyloxycarbonyl ("FMOC"), nitro-veratryloxycarbonyl ("NVOC"), p-methoxybenzyl (PMB), tosyl (Ts), 3,4-dimethoxybenzyl (DMPM), p-methoxyphenyl (PMP), 2-naphthylmethyl ether (Nap), and trichloroethyl chloroformate (Troc).

36. 36. The method of claim 35, wherein the amine protecting group is tert-butoxycarbonyl ("Boc").

37. 35. The method of claim 34, wherein the hydroxyl protecting group is selected from the group consisting of methyl ester, ethyl ester, acetate, propionate group, glycol ester, benzyl, trityl ether, alkyl ether, tetrahydropyranyl ether, trialkylsilyl ether, TMS, TIPS, mesylate, and allyl ether.

38. 38. The method of claim 37, wherein the hydroxyl protecting group is mesylate.

39. 39. The method of any one of claims 34 to 38, wherein the method provides compounds of formula I having an enantiomeric excess of at least 95% or at least 99%.

40. 40. The method of any one of claims 34 to 39, wherein the modified transaminase polypeptide is selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 4, 6, 8, 10, 14, 16, 18, or 20, or wherein the modified transaminase polypeptide is selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 16, 18, and 20.

41. 41. The method of claim 40, wherein the coenzyme is pyridoxal phosphate (PLP) and / or the amino donor is isopropylamine.

42. The compound of formula Ic is 【Chemistry 16】 The method of any one of claims 34 to 41, having the structure:

43. A compound of formula III, or a salt thereof: 【Chemistry 17】 (In the formula, R 7 is an amine protecting group, alkyl, or tert-butyl; R 8 is H; R 9 is alkyl or tert-butyl) 43. A process for preparing a compound of formula I, or a salt thereof, produced by a process according to any one of claims 21 to 42. 【Chemistry 18】 (In the formula, R 1 is Br; R 2 is H; R 3 is -CH 2 R 6 and R 4 is H or an amine protecting group; R 6 is OH or a protected form thereof) with compound IV, 【Chemistry 19】 (In the formula, R 10 is H; R 11 is alkyl or tert-butyl) The method of claim 1, wherein the soluble component is a soluble component.

44. the contacting step comprises: 【Chemistry 20】 44. The method of claim 43, wherein the method is in the presence of

45. The contacting step may comprise potassium tert-butoxide, tripotassium phosphate (K 3 P.O. 4 ), copper(I) bromide, or any combination thereof.

46. 46. ​​The method of any one of claims 43 to 45, wherein the contacting step is in the presence of a solvent.

47. A compound of formula V: 【Chemical 21】 47. The method of any one of claims 43 to 46, further comprising a process for preparing: wherein said process comprises contacting a compound of formula III with an acid.

48. Contacting a compound of formula III with an acid provides a compound of formula VI 【Chemical 22】 (In the formula, R 9 is tert-butyl) 48. The method of claim 47, further comprising:

49. 49. The method of claim 48, wherein the acid is formic acid, acetic acid, propionic acid, butyric acid, valeric acid, caproic acid, oxalic acid, lactic acid, malic acid, citric acid, benzoic acid, carbonic acid, uric acid, taurine, p-toluenesulfonic acid, trifluoromethanesulfonic acid, aminomethylphosphonic acid, trifluoroacetic acid (TFA), phosphonic acid, sulfuric acid, nitric acid, phosphoric acid, hydrochloric acid, ethanesulfonic acid (ESA), methanesulfonic acid (MsOH), or any combination thereof.

50. 50. The method of claim 49, wherein the acid is methanesulfonic acid (MsOH).

51. R 7 The method of any one of claims 43 to 50, wherein is an amine protecting group.

52. 52. The method of claim 51, wherein the amine protecting group is tert-butyloxycarbonyl (Boc), 9-fluorenylmethyloxycarbonyl (Fmoc), carboxybenzyl (Cbz), p-methoxybenzylcarbonyl (Moz), acetyl (Ac), benzoyl (Bz), p-methoxybenzyl (PMB), 3,4-dimethoxybenzyl (DMPM), p-methoxyphenyl (PMP), 2-naphthylmethyl ether (Nap), tosyl (Ts), or trichloroethyl chloroformate (Troc).

53. 53. The method of claim 52, wherein the amine protecting group is a tert-butyloxycarbonyl group (Boc).

54. R 9 is tert-butyl ( t The method of any one of claims 43 to 53, wherein the compound is selected from the group consisting of methyl, ...

55. R 11 is tert-butyl ( t The method of any one of claims 43 to 54, wherein the compound is selected from the group consisting of methyl, ...

56. A compound of formula I, or a salt thereof 【Chemical 23】 The method of any one of claims 43 to 55, having the structure:

57. A compound of formula IV, or a salt thereof 【Chemistry 24】 The method of any one of claims 43 to 56, having the structure:

58. A compound of formula III, or a salt thereof 【Chemistry 25】 The method of any one of claims 43 to 57, having the structure:

59. Niraparibut tosylate monohydrate of formula VII: 【Chemical 26】 a process for preparing a compound of formula V or a salt thereof, produced by the process of any one of claims 43 to 58; 【Chemical 27】 with paratoluenesulfonic acid.

60. 60. The method of claim 59, wherein said contacting is in the presence of a solvent.

61. 61. The method of claim 60, wherein the solvent is water.

62. i) Niraparibut tosylate monohydrate of formula VII: 【Chemical Formula 28】 and ii) less than 0.1% by weight of a compound of formula VI or a salt thereof; 【Chemical 29】 A composition comprising:

63. 63. The composition of claim 62, wherein the amount of the compound of formula VI is less than 0.09 wt%, less than 0.05 wt%, or less than 0.01 wt%.

Citation Information

Patent Citations

  • Transaminase biocatalysts

    JP2012519004A

  • Process for preparing aminocyclohexyl ether compounds

    WO2012024100A2

  • Biocatalysts and methods for the synthesis of substituted lactams

    WO2013036861A1

  • Biocatalytic transamination process

    WO2014088984A1