Engineered Ketoreductase Polypeptides
Engineered ketoreductase polypeptides selectively reduce bicyclic ketones to chiral alcohols with high efficiency and selectivity, using isopropanol, eliminating DMSO and GDH/glucose recycling, enhancing product purity and yield for compounds used in depression treatment.
Patent Information
- Application Number
- JP2025540818
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-12
- Filing Date
- 2024-01-10
- Publication Date
- 2026-02-03
AI Technical Summary
Existing ketoreductases face challenges in selectively reducing bicyclic ketones to chiral alcohol products, particularly due to the U-shaped nature of these compounds, and often require cofactor recycling systems like GDH/glucose or DMSO, which can affect stereopurity and yield.
Engineered ketoreductase polypeptides that are diastereoselective in reducing bicyclic ketones with high selectivity and efficiency, using isopropanol as a reducing agent instead of DMSO, and eliminating the need for GDH/glucose cofactor recycling.
These polypeptides achieve high conversion rates and selectivity, up to 99%, without DMSO and GDH/glucose, improving product purity, diastereomeric ratio, and yield, particularly in the synthesis of compounds for treating depression.
Smart Images

Figure 2026504068000066 
Figure 2026504068000067 
Figure 2026504068000068
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of enzymology, and in particular to the field of ketoreductase enzymology. More specifically, the present invention relates to ketoreductase polypeptides with improved enzymatic activity and polynucleotide sequences encoding the improved ketoreductase polypeptides.
[0002] Sequence Listing This application has been submitted electronically in XML format and contains a Sequence Listing, which is incorporated herein by reference in its entirety. The XML copy, created on August 3, 2022, is named PAT059412-WO-PCT_SL.xml and is 866,783 bytes in size. [Background technology]
[0003] Enzymes belonging to the ketoreductase (KRED) or carbonyl reductase class (EC 1.1.1.184) are useful for the synthesis of optically active alcohols from the corresponding ketone substrates.
[0004] KREDs typically convert ketone and aldehyde substrates to the corresponding alcohol products, but can also catalyze the reverse reaction, i.e., the oxidation of alcohol substrates to the corresponding ketone / aldehyde products. The reduction of ketones and aldehydes and the oxidation of alcohols by enzymes such as KREDs require cofactors, most commonly reduced nicotinamide adenine dinucleotide (NADH) or reduced nicotinamide adenine dinucleotide phosphate (NADPH), and nicotinamide adenine dinucleotide (NAD) or nicotinamide adenine dinucleotide phosphate (NADP) for the oxidation reaction. NADH and NADPH function as electron donors, while NAD and NADP function as electron acceptors. It is commonly observed that ketoreductases and alcohol dehydrogenases accept either phosphorylated or non-phosphorylated cofactors (in their oxidized and reduced states), but not both.
[0005] KRED enzymes can be found in a wide range of bacteria and yeasts (for reviews see Kraus and Waldman, 1995, Enzyme catalysis in organic synthesis, Vols. 1 & 2 VCH Weinheim; Faber, K., 2000, Biotransformations in organic chemistry, 4th Ed. Springer, Berlin Heidelberg New York; and Hummel and Kula, 1989, Eur. J. Biochem. 184:1-13). The sequences of several KRED genes and enzymes have been reported (e.g., Candida magnoliae (GenBank accession number JC7338; GI:11360538); Candida parapsilosis (GenBank accession number BAA24528.1; GI:2815409); Sporobolomyces salmonicolor (GenBank accession number AF160799; GL6539734)).
[0006] To avoid the numerous chemical synthesis steps required to produce important compounds, ketoreductases are increasingly being utilized for the enzymatic conversion of various ketone and aldehyde substrates to chiral alcohol products. In these applications, whole cells expressing the ketoreductase can be used for biocatalytic reduction of ketones and aldehydes, or purified enzymes can be used when the presence of multiple ketoreductases in whole cells would adversely affect the stereopurity and yield of the desired product. In in vitro applications, cofactor (NADH or NADPH)-regenerating enzymes, such as glucose dehydrogenase (GDH), formate dehydrogenase, or a second ketoreductase, are used in combination with the ketoreductase. Examples of the use of ketoreductases to produce useful chemical compounds include the asymmetric reduction of 4-chloroacetoacetate esters (Zhou, 1983, J. Am. Chem. Soc. 105:5925-5926; Santaniello, J. Chem. Res. (S) 1984:132-133; U.S. Pat. Nos. 5,559,030, 5,700,670, and 5,891,685), the reduction of dioxocarboxylic acids (e.g., U.S. Pat. Nos. 6,399,313, 6,413, 6,513, 6,613, 6,713, 6,891,685), and the reduction of hydroxybenzoates (e.g., U.S. Pat. Nos. 6,399,313, 6,413, 6,513, 6,713, 6,891,685). 39); reduction of tert-butyl (S)chloro-5-hydroxy-3-oxohexanoate (e.g., U.S. Pat. No. 6,645,746 and WO 01 / 40450), reduction of pyrrolotriazine compounds (e.g., U.S. Pat. App. Pub. No. 2006 / 0286646); reduction of substituted acetophenones (e.g., U.S. Pat. No. 6,800,477); and reduction of ketothiolanes (WO 2005 / 054491).
[0007] Standard hydride reducing agents, such as lithium aluminum hydride and sodium borohydride, have been reported to reduce bicyclic ketones to the corresponding secondary alcohols. However, for compounds of structure IA, these reagents primarily produce (e.g., in a 9:1 ratio) the undesired diastereoisomer. That is, the hydride is delivered to the convex surface to generate a hydroxyl that is trans to the hydrogen at the ring junction.
[0008] Some engineered ketoreductases also have the activity of dehydrogenating secondary alcohol reducing agents, such as iPrOH. In such cases, when a secondary alcohol is used as the reducing agent, the engineered ketoreductase and secondary alcohol dehydrogenase are the same enzyme.
[0009] Therefore, it would be desirable to identify other ketoreductase enzymes that can be used to effect the conversion of various keto substrates to the corresponding chiral alcohol products, especially in areas where selective reduction along the most impeding plane is difficult due to the U-shaped nature. Furthermore, it would be desirable to identify ketoreductases that can utilize isopropanol as a stoichiometric source of reducing agent for cofactor regeneration, instead of using the well-known glucose dehydrogenase-glucose system. Summary of the Invention [Means for solving the problem]
[0010] KRED enzymes, such as those disclosed herein, have been found to be diastereoselective in the reduction of bicyclic ketone compounds. This reduction occurs through hydride delivery from the concave face of the bicyclic ketone and the introduction of a cis hydroxyl group onto the hydrogen at the ring junction. The modified KRED polypeptides of the present disclosure are surprisingly diastereoselective in the reduction of bicyclic ketone substrates to (cis) alcohol products, with selectivity of at least 96%, and exemplary polypeptides of the present disclosure have diastereoselectivity of at least 99% for the (cis) alcohol product. Notably, the modified KRED polypeptides of the present disclosure do not require the use of DMSO in the reduction reaction; instead, up to 40% by weight of isopropanol (iPrOH) can be used. By removing DMSO and increasing iPrOH in the reaction, the modified KRED polypeptides of the present disclosure enable higher conversion rates, from 85% to complete conversion. Because higher concentrations of iPrOH can be used, the modified KRED polypeptides of the present disclosure unexpectedly do not require the GDH / glucose cofactor recycling system required by other KRED enzymes in the reduction of bicyclic ketone compounds.
[0011] Provided herein are methods for preparing compounds useful for treating diseases and disorders associated with depression, such as major depressive disorder and treatment-resistant or refractory depression. In some embodiments, the compounds prepared include inhibitors of the NR2B-NMDA receptor useful for treating such diseases and disorders. In some embodiments, the compounds prepared include, for example, compounds exemplified in PCT Patent Application WO 2016 / 049165, the entire contents of which are incorporated herein. The processes described herein improve, for example, product purity, diastereomeric ratio (dr), stereoisomeric excess, and / or yield of the final product and key intermediates in its synthesis. The processes described herein will be more fully understood with reference to several reaction schemes below. In some embodiments, the processes unexpectedly result in improved product purity, improved diastereomeric ratio, improved stereoisomeric excess, and / or improved yield. Improved product purity includes, for example, improved diastereomeric purity of the reaction product.
[0012] In one aspect, the invention features an engineered ketoreductase polypeptide. In some embodiments, the engineered ketoreductase polypeptide of the invention has an identity at least about 97.7%, 97.8%, or 97.9% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 392, 394, 396, 406, 408, 420, 422, 424, 426, 428, 430, 436, 438, 462, 464, 468, 470, 474, 476, 480, 482, and 490. 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% sequence identity to the polypeptide. In some embodiments, the engineered ketoreductase polypeptides of the invention comprise: (1) an amino acid sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 54, 152, 256, or 492; and (2) with respect to said amino acid sequence, i) X17S, X17T, or X17G; ii) X18S or X18R; or iii) and ix) X198V, X198I, X198Y, or X198P, and optionally, contain one or more additional amino acid residue differences relative to the amino acid sequence. In some embodiments, an engineered ketoreductase polypeptide of the invention is a polypeptide comprising an amino acid sequence selected from Table 4 (SEQ ID NOs: 4-52), Table 5 (SEQ ID NOs: 56-186), Table 6 (SEQ ID NOs: 188-390), or Table 7 (SEQ ID NOs: 392-490).In some embodiments, the polypeptide can selectively reduce tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product).
[0013] In another aspect, the invention features an engineered ketoreductase polypeptide that can selectively reduce tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product). In some embodiments, the engineered ketoreductase polypeptides of the invention (i) comprise an amino acid sequence having at least about 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 54, 152, 256, or 492; and (ii) comprise, relative to said amino acid sequence, a substitution, deletion, addition, or insertion of one or more amino acid residues selected from Table 2, Table 4, Table 5, Table 6, or Table 7, and optionally comprise one or more additional amino acid residue differences relative to said amino acid sequence.
[0014] In some embodiments, the amino acid sequence of the modified polypeptide comprises one or more amino acid differences at positions X17, X18, X40, X56, X94, X96, X106, X145, X173, X190, X194, X196, X198, X202, and / or X206 relative to the amino acid sequence, and optionally one or more additional amino acid residue differences relative to the amino acid sequence. In some embodiments, the amino acid sequence of the modified polypeptide has one or more of the following amino acid residues added to the amino acid sequence: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X40R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X106L, X106W, X106E, X145S, X145D, X145E, X173A, X174B, X175C, X175D, X175E, X173B, X175D, X175E, X173C, X173D, X173E, X173F, X173G, X173H, X173I, X173F ... and / or X206W, and optionally one or more additional amino acid residue differences relative to said amino acid sequence. In some embodiments, the amino acid sequence of the modified polypeptide comprises one or more amino acid sequences selected from X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W, optionally comprising one or more additional amino acid residue differences relative to the amino acid sequence. In some embodiments, the amino acid sequence of the modified polypeptide comprises one or more amino acid sequences selected from X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q, optionally comprising one or more additional amino acid residue differences relative to the amino acid sequence.
[0015] In some embodiments, the amino acid sequence of the modified polypeptide comprises, relative to the amino acid sequence, (i) a residue at position X190 that is not tyrosine; (ii) a residue at position X190 that is a non-aromatic residue; (iii) a residue at position X196 that is an aliphatic residue or a small amino acid residue; (iv) a residue at position X202 that is a small amino acid residue; and / or (v) a residue at position X206 that is not methionine, and optionally comprises one or more additional amino acid residue differences relative to the amino acid sequence.
[0016] In some embodiments, the amino acid sequence of the modified polypeptide comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 392, 394, 396, 406, 408, 420, 422, 424, 426, 428, 430, 436, 438, 462, 464, 468, 470, 474, 476, 480, 482, and 490.
[0017] In some embodiments, the modified polypeptide is solvent stable. In some embodiments, the modified polypeptide reduces a substrate to a product with a conversion rate of at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99.0%, or 100%. In some embodiments, the modified polypeptide reduces a substrate to a product with a selectivity level (diastereomeric excess%) of at least about 96.0%, at least about 96.5%, at least about 97.0%, at least about 97.5%, at least about 98.0%, at least about 98.5%, at least about 99.0%, at least about 99.5%, or 100%. In some embodiments, the ability to reduce a substrate to a product is compared to a reference (e.g., parent) polypeptide. In some embodiments, the modified polypeptide has a reversed or increased diastereoselectivity for the reduction of a substrate to a product compared to the reference (e.g., parent) polypeptide. In some embodiments, the modified polypeptide has an increased level of activity (e.g., rate of conversion or desired product) compared to a reference (e.g., parent) polypeptide, having an Improvement Over Positive Control (FIOP) of greater than about 1.75, greater than about 2.00, greater than about 2.25, greater than about 2.50, greater than about 2.75, greater than about 3.00, greater than about 3.25, greater than about 3.50, greater than about 3.75, greater than about 4.00, greater than about 4.25, greater than about 4.50, greater than about 4.75, greater than about 5.00, greater than about 5.25, greater than about 5.50, greater than about 5.75, greater than about 6.00, greater than about 6.25, or greater than about 6.50. In some embodiments, the modified polypeptide is capable of reducing a substrate to a product with an FIOP conversion of greater than about 2.25, preferably greater than about 3.00, and a diastereoselectivity of greater than about 95%, preferably greater than about 97%, or more preferably greater than about 99%, compared to the reference polypeptide.
[0018] In some embodiments, the reference (e.g., parent) polypeptide is a ketoreductase peptide of wild-type Lactobacillus kefir, Lactobacillus brevis, or Lactobacillus minor. In some embodiments, the reference (e.g., parent) polypeptide is a wild-type Lactobacillus kefir ketoreductase polypeptide. In some embodiments, the reference (e.g., parent) polypeptide is a wild-type Lactobacillus kefir ketoreductase polypeptide comprising the amino acid sequence of SEQ ID NO: 492. In some embodiments, the reference (e.g., parent) polypeptide is an engineered ketoreductase polypeptide. In some embodiments, the reference (e.g., parent) polypeptide is an engineered ketoreductase polypeptide selected from SEQ ID NO: 54, SEQ ID NO: 152, or SEQ ID NO: 256.
[0019] In some embodiments, the modified polypeptide i) requires fewer cofactors; ii) does not require glucose dehydrogenase (GDH) / glucose cofactor recycling; and / or iii) does not require dimethyl sulfoxide (DMSO) in a reduction reaction to reduce a substrate to a product, compared to a reference (e.g., parent) polypeptide.
[0020] In one aspect, the invention features an engineered ketoreductase polypeptide that can selectively reduce tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product) with greater stereoselectivity (% diastereomeric excess) and / or activity (% conversion) than SEQ ID NO:54, SEQ ID NO:152, and / or SEQ ID NO:256 under suitable reaction conditions.
[0021] In some embodiments, suitable reaction conditions include one or more of the following: i) up to about 150 g / L, e.g., 50 g / L or 100 g / L, of substrate; ii) an enzyme load of less than about 5 wt%, e.g., less than about 3 wt%, e.g., less than about 1 wt%; iii) an NADP+ cofactor load of less than about 2 wt%, e.g., less than about 1 wt%, e.g., less than about 0.2 wt%, e.g., 0.1%, 0.05%, or 0.03%; iv) up to 40% (v / v) isopropanol; v) a temperature of about 15-75°C, e.g., 20-55°C, e.g., 20-45°C, e.g., 35°C or 40°C; vi) a pH of 5.0-10.0, e.g., 7.5-8.3; and vii) a reaction time of up to 30 hours, preferably 24 hours. In some embodiments, suitable reaction conditions do not require dimethyl sulfoxide (DMSO). In some embodiments, suitable reaction conditions do not require recycling of the GDH / glucose cofactor.
[0022] In some embodiments, the modified polypeptide comprises an amino acid sequence that differs from the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256 by one or more amino acid residues selected from Table 2, Table 4, Table 5, Table 6, or Table 7. In some embodiments, the modified polypeptide comprises an amino acid sequence that differs from the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256 by one or more amino acid residues selected from 17, 18, 40, 56, 94, 96, 106, 145, 173, 190, 194, 196, 198, 202, and / or 206. In some embodiments, the modified polypeptide amino acid sequence includes, relative to the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256, (i) a residue at position X190 that is not tyrosine; (ii) a residue at position X190 that is a non-aromatic residue; (iii) a residue at position X196 that is an aliphatic residue or a small amino acid residue; (iv) a residue at position X202 that is a small amino acid residue; and / or (v) a residue at position X206 that is not methionine. In some embodiments, the amino acid sequence of the modified polypeptide has the following amino acid residues compared to the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X40R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X106L, X106W, X106E, and / or X206W.In some embodiments, the amino acid sequence of the modified polypeptide comprises one or more amino acid residues selected from: i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P, compared to the amino acid sequence of SEQ ID NO: 54, SEQ ID NO: 152, or SEQ ID NO: 256. In some embodiments, the amino acid sequence of the modified polypeptide comprises one or more amino acid residues selected from X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W compared to the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256. In some embodiments, the amino acid sequence of the modified polypeptide comprises one or more amino acid residues selected from X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q compared to the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256. In some embodiments, the modified polypeptide comprises an amino acid sequence selected from: a) an amino acid sequence selected from Table 4 (SEQ ID NOs: 4-52), Table 5 (SEQ ID NOs: 56-186), Table 6 (SEQ ID NOs: 188-390), or Table 7 (SEQ ID NOs: 392-490). In some embodiments, the modified polypeptide comprises an amino acid sequence selected from Table 7.In some embodiments, the modified polypeptide has an amino acid sequence that is at least about 97.7%, 97.8%, 98.9%, 100%, 110%, 120%, 130%, 140%, 142%, 150%, 152%, 154%, 156%, 158%, 160%, 162%, 164%, 166%, 168%, 169%, 170%, 171%, 172%, 173%, 174%, 175%, 176%, 177%, 178%, 179%, 180%, 181%, 182%, 183%, 184%, 185%, 186%, 187%, 188%, 189%, 190%, 200%, 201%, 202%, 203%, 204%, 205%, In some embodiments, the amino acid sequence of the modified polypeptide comprises one or more additional amino acid residue differences compared to SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256.
[0023] In another aspect, the invention features a method for the diastereoselective reduction of a bicyclic ketone. In some embodiments, a method for the diastereoselective reduction of a bicyclic ketone is provided, the method including contacting the bicyclic ketone with a KRED under suitable reaction conditions to provide a bicyclic secondary alcohol product.
[0024] In some embodiments, the KRED is an engineered ketoreductase.
[0025] In some embodiments, the bicyclic ketone substrate has a total of 6 to 12 members.
[0026] In some embodiments, the bicyclic ketone substrate is achiral.
[0027] In some embodiments, the method comprises contacting a bicyclic ketone substrate with any of the modified polypeptides of the invention under suitable reaction conditions to obtain a bicyclic secondary alcohol product. In some embodiments, the resulting bicyclic secondary alcohol product has the formula (IB): [ka] (In the formula, The A and B rings together form a fused cycloalkyl ring, e.g., a C to C 12 represents a cycloalkyl or fused heterocyclyl ring, for example a 6- to 12-membered heterocyclyl; wherein the fused cycloalkyl or heterocyclyl is at least one R 1 occurrence of, e.g., 1 to 4 R 1 may be substituted with Each R 1 is an amine protecting group, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz), C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, C3-C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 Arylalkyl, 3-14 membered heterocyclyl, 5-20 membered heteroaryl, hydroxyl, halogen, e.g., F, C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , and -(CH2) n -C(=O)NR c R c are independently selected from; Here, alkyl, alkenyl, and alkynyl each represent one or more R a , for example, 1 to 6 R a and cycloalkyl, aryl, arylalkyl, heterocyclyl, and heteroaryl may each be substituted with one or more R b , for example, 1 to 6 R b may be substituted with; Each R a When each occurs, C3 to C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20Arylalkyl, 3-14 membered heterocyclyl, 5-20 membered heteroaryl, halogen, e.g., F, haloalkyl, e.g., C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , and -(CH2) n -C(=O)NR c R c Selected from; Each R b are independently halogens, C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , -(CH2) n -C(=O)NR c R c , C1~C 20 Alkyl, C2-C 20 Alkenyl and C2-C 20 alkynyl; Each R c When each occurs, H, C1 to C 20 Alkyl, C2-C 20 Alkenyl and C2-C 20 Alkynyl (each one or more R b , for example, 1 to 6 R b and n is 0, 1, 2, 3, 4, 5, or 6, for example, 0, 1, 2, or 3. It has the structure shown in
[0028] In some embodiments, the bicyclic ketone substrate has the formula (IA): [ka] wherein A and B are as defined above for formula (IB). It has the structure shown in
[0029] In some embodiments, the A and B rings together represent a fused 6- to 12-membered heterocyclyl containing at least one nitrogen heteroatom, and in some embodiments, the nitrogen heteroatom of the bicyclic ketone substrate is bonded to an amine protecting group, e.g., tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz).
[0030] In some embodiments, the bicyclic ketone substrate has the formula (IA)-I: [ka] (In the formula, X is NR 1a , CH2, and CH-R 1b Selected from; R 1a is an amine protecting group, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz), C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, C3-C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 selected from arylalkyl, 3- to 14-membered heterocyclyl, and 5- to 20-membered heteroaryl; R 1b is C1~C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, C3-C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 selected from arylalkyl, 3- to 14-membered heterocyclyl, and 5- to 20-membered heteroaryl; where R 1a or R 1bThe alkyl, alkenyl, and alkynyl groups each have 1 to 6 R a may be substituted with; R 1a or R 1b The cycloalkyl, aryl, arylalkyl, heterocyclyl, and heteroaryl each have 1 to 6 R b may be substituted with; Each R 1c When each occurs, C1 to C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, C3-C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 Arylalkyl, 3-14 membered heterocyclyl, 5-20 membered heteroaryl, hydroxyl, halogen, e.g., F, C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , and -(CH2) n -C(=O)NR c R c selected from the group consisting of: Each R a When each occurs, C3 to C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 Arylalkyl, 3-14 membered heterocyclyl, 5-20 membered heteroaryl, halogen, e.g., F, haloalkyl, e.g., C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , and -(CH2) n -C(=O)NR c Rc Selected from; Each R b are independently halogens, C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , -(CH2) n -C(=O)NR c R c , C1~C 20 Alkyl, C2-C 20 Alkenyl and C2-C 20 alkynyl; Each R c When each occurs, H, C1 to C 20 Alkyl, C2-C 20 Alkenyl and C2-C 20 Alkynyl (each of which has 1 to 6 R b and optionally substituted with; n is 0, 1, 2, or 3; and m is 0, 1, or 2) It is expressed as:
[0031] In some embodiments, X is NR 1a and;R 1a is C1~C 10 Alkyl, C3-C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 and m is selected from arylalkyl, and amine protecting groups, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz); and m is 0. In some embodiments, X is NR 1a and R 1a is an amine protecting group, for example, tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz).
[0032] In another aspect, the present invention relates to converting (IA)-I to (IB)-I: [ka] (In the formula, R 1a is an amine protecting group). 1a is selected from tert-butyloxycarbonyl (Boc) and N-carboxybenzyl (Cbz). In some embodiments, the method includes contacting the substrate (IA)-I with a KRED under reaction conditions suitable for reducing or converting (IA)-I to (IB)-I. In some embodiments, the KRED is an engineered ketoreductase. In some embodiments, the KRED is an engineered polypeptide described herein.
[0033] In another aspect, the invention features a method for stereoselectively reducing tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product). [ka]
[0034] In some embodiments, the method includes contacting a substrate (IA)-I, e.g., tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate, with a KRED under reaction conditions suitable for reducing or converting (IA)-I, e.g., tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate), to (IB)-I, e.g., tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product). In some embodiments, the KRED is an engineered ketoreductase. In some embodiments, the KRED is an engineered polypeptide described herein.
[0035] In some embodiments, the reaction is carried out in a solvent. In some embodiments, the solvent is selected from polar solvents, nonpolar solvents, and ionic liquids. In some embodiments, the solvent is selected from water, methanol, ethanol, n-propanol, isopropanol, isopropyl acetate, dimethyl sulfoxide, dimethylformamide, ethyl acetate, butyl acetate, 1-octanol, hexane, heptane, octane, methyl tert-butyl ether, toluene, 1-ethyl 4-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium hexafluorophosphate, glycerol, ethylene glycol, propylene glycol, and polyethylene glycol. In some embodiments, the solvent is an isopropanol solvate. In some embodiments, the reaction is carried out in up to 40% by weight isopropanol. In some embodiments, the reaction is carried out in the presence of a co-solvent. In some embodiments, the reaction is carried out in an aqueous co-solvent system. In some embodiments, the co-solvent is selected from dimethyl sulfoxide (DMSO) and an alcohol, such as methanol, ethanol, n-propanol, isopropanol, etc. In some embodiments, the reaction is not carried out in the presence of dimethyl sulfoxide (DMSO).
[0036] In some embodiments, the reaction is carried out at a temperature between 15 and 75° C. In some embodiments, the reaction is carried out at a temperature between 20 and 55° C. In some embodiments, the reaction is carried out at a temperature between 20 and 45° C., e.g., 35° C. or 40° C. In some embodiments, the reaction is carried out at a pH between 5.0 and 10.0, e.g., 7.5 and 8.3. In some embodiments, the reaction is carried out at a pH of 7.5.
[0037] In some embodiments, the bicyclic ketone substrate is present at a load concentration of up to about 150 g / L, hi some embodiments, the concentration of the bicyclic ketone substrate is at least about 5 g / L, at least about 10 g / L, at least about 20 g / L, at least about 50 g / L, at least about 100 g / L, or about 150 g / L.
[0038] In some embodiments, the polypeptide is present at a concentration of less than about 10 g / L. In some embodiments, the concentration of the polypeptide is less than about 5 g / L, e.g., less than about 5 g / L, e.g., less than about 3 g / L, e.g., less than about 1 g / L. In some embodiments, the solvent is present at a concentration of 20% to 40% v / v.
[0039] In some embodiments, the method is performed using whole cells expressing the ketoreductase enzyme, or an extract or lysate of such cells. In some embodiments, the method does not require glucose dehydrogenase (GDH) / glucose cofactor recycling.
[0040] In some embodiments, the ketoreductase is isolated and / or purified, and the reduction reaction is carried out in the presence of a cofactor for the ketoreductase. In some embodiments, the cofactor comprises nicotinamide adenine dinucleotide (NADH) or nicotinamide adenine dinucleotide phosphate (NADPH). In some embodiments, the cofactor is present at a concentration of less than about 2% by weight, e.g., less than about 1% by weight, e.g., less than about 0.2% by weight, e.g., 0.1%, 0.05%, or 0.03%.
[0041] In some embodiments, the methods result in products with selectivity (% diastereomeric excess) of greater than about 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%, and / or % conversion of at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99.0%, or 100%. In some embodiments, at least about 90% of the substrate is reduced to the product in less than about 20 hours. In some embodiments, at least about 85% of the substrate is reduced to the product in less than about 20 hours, and at least about 95% of the substrate is reduced to the product in less than about 30 hours.
[0042] In one aspect, the present invention provides a compound of formula (IC) [ka] The present invention features a method for synthesizing 6-((S)-2-((3aR,5R,6aS)-5-(2-fluorophenoxy)hexahydrocyclopenta[c]pyrrol-2(1H)-yl)-1-hydroxyethyl)pyridin-3-ol, represented by the formula (I), in free form or in pharmaceutically acceptable salt form.
[0043] In some embodiments, the method comprises a process comprising contacting a substrate of formula (IA)-I, e.g., tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate, with a KRED under reaction conditions suitable for reducing or converting (IA)-I, e.g., tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate), to a compound of formula (IB)-I, e.g., tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product). In some embodiments, the KRED is a modified KRED. In some embodiments, the KRED is a modified polypeptide described herein.
[0044] In another aspect, the invention features a method for reversing the diastereoselectivity of a ketoreductase (KRED) polypeptide in a reduction reaction from the formation of a (trans) alcohol product (e.g., (5r)-2) to the formation of a (cis) alcohol product (e.g., (5s)-2). In yet another aspect, the invention features a method for increasing the diastereoselectivity of a ketoreductase (KRED) polypeptide in a reduction reaction toward the formation of a (cis) alcohol product (e.g., (5s)-2). In some embodiments, the method includes introducing one or more amino acid differences selected from Table 2, Table 4, Table 5, Table 6, or Table 7 relative to the amino acid sequence of the KRED polypeptide, wherein the one or more amino acid differences, when used in a reduction reaction, reverse or increase the diastereoselectivity of the KRED polypeptide toward the formation of the (cis) alcohol product (e.g., (5s)-2) relative to the KRED polypeptide without the one or more amino acid differences. In some embodiments, the one or more amino acid differences are selected from the following positions: 17, 18, 40, 56, 94, 96, 106, 145, 173, 190, 194, 196, 198, 202, and / or 206 relative to the amino acid sequence of the KRED polypeptide.
[0045] In some embodiments, the one or more amino acid differences are selected from the following, relative to the amino acid sequence of the KRED polypeptide: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X40R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X106L, X106W, X106 E, X145S, X145D, X145E, X173A, X173L, X173V, X173W, X173M, X173K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W. In some embodiments, the one or more amino acid differences, compared to the amino acid sequence of the KRED polypeptide, are selected from the following: i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P. In some embodiments, the one or more amino acid differences, relative to the amino acid sequence of the KRED polypeptide, are selected from the following: X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W. In some embodiments, the one or more amino acid differences, relative to the amino acid sequence of the KRED polypeptide, are selected from the following: X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q.In some embodiments, the one or more amino acid differences, relative to the amino acid sequence of the KRED polypeptide, are selected from: (i) a residue at position X190 that is not tyrosine; (ii) a residue at position X190 that is a non-aromatic residue; (iii) a residue at position X196 that is an aliphatic residue or a small amino acid residue; (iv) a residue at position X202 that is a small amino acid residue; and / or (v) a residue at position X206 that is not methionine.
[0046] In some embodiments, the KRED polypeptide is a ketoreductase peptide of wild-type Lactobacillus kefir, Lactobacillus brevis, or Lactobacillus minor. In some embodiments, the KRED polypeptide is a wild-type Lactobacillus kefir ketoreductase polypeptide. In some embodiments, the KRED polypeptide is a wild-type Lactobacillus kefir ketoreductase polypeptide comprising the amino acid sequence of SEQ ID NO: 492. In some embodiments, the KRED polypeptide is an engineered ketoreductase polypeptide. In some embodiments, the KRED polypeptide is an engineered ketoreductase polypeptide selected from Table 4, Table 5, or Table 6. In some embodiments, the KRED polypeptide is an engineered ketoreductase polypeptide selected from SEQ ID NO: 54, SEQ ID NO: 152, and SEQ ID NO: 256.
[0047] In some embodiments, the method provides at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99.0%, or 100% conversion to the (cis) alcohol product (e.g., (5s)-2). In some embodiments, the method provides a selectivity (% diastereomeric excess) level for the (cis) alcohol product (e.g., (5s)-2) of at least about 96.0%, at least about 96.5%, at least about 97.0%, at least about 97.5%, at least about 98.0%, at least about 98.5%, at least about 99.0%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or 100%.
[0048] In a further aspect, [ka] or a salt thereof.
[0049] In a further aspect, [ka] or a salt thereof, in free form or in the form of a pharmaceutically acceptable salt thereof [ka] In some embodiments, compound (IC) is present in a diastereomeric excess (de)% of at least about 85%, at least about 96%, or at least about 99%.
[0050] In yet another aspect, the invention features a composition including any of the modified KRED polypeptides of the invention. In some embodiments, the composition further includes a structural formula of a substrate and / or product compound of the disclosure.
[0051] In a further aspect, the invention features an immobilized polypeptide. In some embodiments, the polypeptide is selected from any of the modified KRED polypeptides provided herein. In some embodiments, the polypeptide is immobilized to a solid support. In some embodiments, the polypeptide is immobilized to the solid support by chemical bonding (e.g., covalent or ionic bonding), physical adsorption, or affinity interactions. In some embodiments, the solid support is organic or inorganic and has different functional groups. Non-limiting solid supports include resins, silica, zeolites, charcoal, celite (e.g., diatomaceous earth), synthetic polymers (e.g., polymethacrylate or the anion exchange resin Amberlite), biopolymers (e.g., cellulose, chitosan, agarose, lignin, or lignocellulose), controlled-pore glass, magnetic nanoparticles, metal-organic frameworks, or DNA. In some embodiments, the polypeptide is immobilized by encapsulation in a hydrogel (e.g., alginate, chitosan, carrageenan) or matrix (e.g., polyacrylamide). In some embodiments, the polypeptides are immobilized by carrier-free immobilization, e.g., cross-linking of the enzyme in the absence of a support, achieved, for example, by precipitating the enzyme as aggregates that maintain their tertiary structure, followed by cross-linking with glutaraldehyde.
[0052] In one aspect, the invention features a polynucleotide encoding any of the modified KRED polypeptides of the invention. In some embodiments, the polynucleotide comprises a nucleic acid sequence listed in Table 4 (SEQ ID NOs: 3-51), Table 5 (SEQ ID NOs: 55-185), Table 6 (SEQ ID NOs: 187-389), or Table 7 (SEQ ID NOs: 391-489).
[0053] In another aspect, the invention features a polynucleotide encoding a modified KRED polypeptide comprising a nucleic acid sequence at least about 98%, at least about 99%, or 100% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 391, 393, 395, 405, 407, 419, 421, 423, 425, 427, 429, 435, 437, 461, 463, 467, 469, 473, 475, 479, 481, and 489.
[0054] In yet another aspect, the invention features an expression vector including any of the polynucleotides of the invention operably linked to a control sequence suitable for directing expression of the encoded polypeptide in a host cell. In some embodiments, the control sequence includes a secretion signal.
[0055] In a further aspect, the invention features a host cell containing any of the polynucleotides or expression vectors according to the invention.
[0056] In one aspect, the invention features a method for preparing an engineered ketoreductase polypeptide. In some embodiments, the method includes culturing any of the host cells of the invention under conditions suitable for gene expression, and then purifying and collecting the engineered polypeptide from the cell culture.
[0057] In another aspect, the invention features kits. In some embodiments, the kits include any of the modified KRED polypeptides provided herein. In some embodiments, the kits include any of the compositions provided herein. In some embodiments, the kits include any of the polynucleotides provided herein. In some embodiments, the kits include any of the expression vectors provided herein. In some embodiments, the kits include any of the host cells provided herein. In some embodiments, the kits include any of the modified KRED polypeptides provided herein, tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate), and a cofactor. In some embodiments, the cofactor is nicotinamide adenine dinucleotide (NADH) or nicotinamide adenine dinucleotide phosphate (NADPH). In some embodiments, the kits further include instructions for use of any of the modified polypeptides, compositions, polynucleotides, expression vectors, and / or host cells of the invention.
[0058] The compositions defined by the present invention have been isolated or prepared in connection with the examples provided below. One or more embodiments of the invention are described in detail herein. Other features and advantages of the invention will become apparent from the detailed description, examples, and claims. [Brief explanation of the drawings]
[0059] [Figure 1] Figure 1 is a schematic diagram illustrating the role of ketoreductases (KREDs) in the conversion of compound 1 to compounds (5s)-2 or (5r)-2, which utilize a KRED of the invention and a cofactor such as NAD(P)H. [Figure 2A] FIG. 2A is a graph showing the activity (% conversion) of the engineered KRED enzymes (SEQ ID NO: 152, SEQ ID NO: 256, SEQ ID NO: 426) compared to the commercial KRED enzyme (CM), the parent KRED enzyme (SEQ ID NO: 54), and the wild-type KRED enzyme (WT, SEQ ID NO: 492) at a substrate concentration of 50 g / L under condition 1 parameters ( FIG. 2A ). [Figure 2B] Figure 2B is a graph showing the selectivity (de%) of the engineered KRED enzymes (SEQ ID NO: 152, SEQ ID NO: 256, SEQ ID NO: 426) compared to the commercial KRED enzyme (CM), the parent KRED enzyme (SEQ ID NO: 54), and the wild-type KRED enzyme (WT, SEQ ID NO: 492) at a substrate concentration of 50 g / L under condition 1 parameters (Figure 2B). [Figure 3A] FIG. 3A is a graph showing the activity (% conversion) of the engineered KRED enzymes (SEQ ID NO: 152, SEQ ID NO: 256, SEQ ID NO: 426) compared to the commercial KRED enzyme (CM), the parent KRED enzyme (SEQ ID NO: 54), and the wild-type KRED enzyme (WT, SEQ ID NO: 492) at a substrate concentration of 90 g / L under condition 1 parameters ( FIG. 3A ). [Figure 3B] Figure 3B is a graph showing the selectivity (de%) of the engineered KRED enzymes (SEQ ID NO: 152, SEQ ID NO: 256, SEQ ID NO: 426) compared to the commercial KRED enzyme (CM), the parent KRED enzyme (SEQ ID NO: 54), and the wild-type KRED enzyme (WT, SEQ ID NO: 492) at a substrate concentration of 90 g / L under condition 1 parameters (Figure 3B). [Figure 4A] FIG. 4A is a graph showing the activity (% conversion) of the engineered KRED enzymes (SEQ ID NO: 152, SEQ ID NO: 256, SEQ ID NO: 426) compared to the commercial KRED enzyme (CM), the parent KRED enzyme (SEQ ID NO: 54), and the wild-type KRED enzyme (WT, SEQ ID NO: 492) under condition 2 parameters at a substrate concentration of 50 g / L ( FIG. 4A ). [Figure 4B-4C] Figures 4B and 4C are graphs showing the activity (% conversion) of the engineered KRED enzymes (SEQ ID NO: 152, SEQ ID NO: 256, SEQ ID NO: 426) compared to the commercial KRED enzyme (CM), the parent KRED enzyme (SEQ ID NO: 54), and the wild-type KRED enzyme (WT, SEQ ID NO: 492) under condition 2 parameters at substrate concentrations of 100 g / L (Figure 4B) and 150 g / L (Figure 4C). [Figure 5A-5B]Figures 5A and 5B are graphs showing the selectivity (de%) of the engineered KRED enzymes (SEQ ID NO: 152, SEQ ID NO: 256, SEQ ID NO: 426) compared to the commercially available KRED enzyme (CM), the parent KRED enzyme (SEQ ID NO: 54), and the wild-type KRED enzyme (WT, SEQ ID NO: 492) under condition 2 parameters at substrate concentrations of 50 g / L (Figure 5A) and 100 g / L (Figure 5B). [Figure 5C] Figure 5C is a graph showing the selectivity (de%) of the engineered KRED enzymes (SEQ ID NO: 152, SEQ ID NO: 256, SEQ ID NO: 426) compared to the commercial KRED enzyme (CM), the parent KRED enzyme (SEQ ID NO: 54), and the wild-type KRED enzyme (WT, SEQ ID NO: 492) under condition 2 parameters at a substrate concentration of 150 g / L (Figure 5C). DETAILED DESCRIPTION OF THE INVENTION
[0060] Bicyclic secondary alcohols are useful intermediates in the synthesis of pharmacologically active agents. They are referred to herein as (substrates) and have the following structure: [ka] tert-Butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate, represented by the formula: is particularly useful in the synthesis of compounds of formula (IC).
[0061] The compound (IC) disclosed herein is known as omfasprodil, which is an NR2B-NMDA receptor non-allosteric modulator (NAM), and has the following formula: [ka] and can be prepared as described in WO 2016 / 049165, which is incorporated herein by reference.
[0062] Evidence suggests that the NR2B negative allosteric modulators (NAMs) MK-0657 (also known as CERC-301) and CP-101,606 have a low frequency of dissociative adverse events (Garner et al. 2015; Pagnozzi et al. 1995; Preskornet et al. 2008). The relative contribution of each individual NMDAR subtype to the adverse effects of pan-NMDAR blockade is poorly understood, collectively due to the lack of selective inhibitors for the various subtypes. This suggests that achieving safe yet rapid antidepressant efficacy may be feasible with compounds that selectively inhibit NR2B receptors.
[0063] Compound (IC) or its pharmaceutically acceptable salts is a highly potent, selective, and reversible small molecular weight NR2B-NMDA receptor NAM. Compound (IC) is intended for rapid relief of depressive symptoms in patients with major depressive disorder (MDD), including treatment-resistant depression and suicide. This treatment aims to enable patients to rapidly achieve significant improvement in depressive symptoms and suicidality. Furthermore, patients with MDD accompanied by suicidality often require hospitalization for 4-5 days. Compound (I), with its rapid onset of efficacy, can reduce the number of days a patient is hospitalized (or eliminate hospitalization entirely), thus providing a greater benefit than other antidepressants, which require at least 4 weeks for patients to respond to treatment.
[0064] Compound (IC) or a pharmaceutically acceptable salt thereof is intended for the treatment of suicidality, symptoms of suicidality, including but not limited to, suicidal ideation, suicidal behavior, and self-harm, alone or in combination with a psychiatric disorder, including but not limited to, major depressive disorder. In particular, Compound (I) or a pharmaceutically acceptable salt thereof is intended for the treatment of major depressive disorder in patients with suicidal ideation.
[0065] Compound (IC) can be prepared by reducing bicyclic ketone (Int-1) with sodium borohydride in ethanol. The reduction process produces two diastereomers, 2 and 2A. Compound 2 is the undesired trans diastereoisomer and is formed as the major product. The undesired diastereomer 2 is then converted to the corresponding mesylate 18, followed by Scheme 1. N Received 2 responses. Scheme 1 [ka]
[0066] This reaction proceeds with inversion of stereochemistry at the C-5 position to give intermediate compound 54 of the appropriate stereochemistry. This method is disclosed in WO 2016 / 049165. Compound 2A is reported to be produced in approximately 6.8% yield in the reduction process. Compound 54 is then deprotected, alkylated with 2-bromo-1-(5-hydroxypyridin-2-yl)ethan-1-one, and then reduced to give compound (IC).
[0067] Thus, there is a need for improved methods for the stereoselective reduction of bicyclic ketone substrates to the corresponding alcohol products. The KRED polypeptides disclosed herein, and their use in processes for the reduction of bicyclic ketones to the corresponding secondary alcohol products, provide a solution to this problem.
[0068] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide those skilled in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them below, unless otherwise specified.
[0069] "Acidic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that exhibits a pK value of less than about 6 when the amino acid is included in a peptide or polypeptide. Acidic amino acids typically have a negatively charged side chain at physiological pH due to loss of a hydrogen ion. Genetically encoded acidic amino acids include L-Glu (E) and L-Asp (D).
[0070] "Agent" means any small molecule compound, polynucleotide, polypeptide, or fragment thereof.
[0071] "Aliphatic amino acid or residue" refers to a hydrophobic amino acid or residue having an aliphatic hydrocarbon side chain. Genetically encoded aliphatic amino acids include L-Ala (A), L-Val (V), L-Leu (L), and L-Ile (I).
[0072] "Aromatic amino acid or residue" refers to a hydrophilic or hydrophobic amino acid or residue having a side chain containing at least one aromatic or heteroaromatic ring. Genetically encoded aromatic amino acids include L-Phe (F), L-Tyr (Y), and L-Trp (W). While L-His (H) may be classified as a basic residue because of the pKa of its heteroaromatic nitrogen atom, or as an aromatic residue because its side chain contains a heteroaromatic ring, histidine is classified herein as a hydrophilic residue or "constrained residue."
[0073] A "basic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that exhibits a pK value greater than about 6 when the amino acid is included in a peptide or polypeptide. Basic amino acids typically have a positively charged side chain at physiological pH due to association with hydronium ions. Genetically encoded basic amino acids include L-Arg (R) and L-Lys (K).
[0074] "Bicyclic ketone" refers to a compound having a bicyclic skeleton characterized by at least two fused rings and at least one carbonyl group. The bicyclic ketone is preferably a fused cycloalkyl or heterocyclyl compound, such as an octahydronaphthalenone, e.g., octahydronaphthalen-1(2H)-one, hexahydropentalen-2(1H)-one, or a hexahydrocyclopenta[c]pyrrolone, e.g., hexahydrocyclopenta[c]pyrrol-5(1H)-one. Preferably, the bicyclic ketone substrate is achiral. More preferably, each ring member of the bicyclic ketone has the same number of ring atoms. The bicyclic ketone structure may be substituted with one or more substituents. The substituents may themselves be substituted. The substituents may be attached via a carbon atom or a heteroatom, e.g., N. In one embodiment, the bicyclic ketone has a total of 6 to 12 members.
[0075] "Bicyclic secondary alcohol product" refers to a product produced as a result of application of a KRED enzyme to a bicyclic ketone substrate, for example, as disclosed herein.
[0076] The terms (cis) and (trans) alcohol products refer to the stereochemical configuration of the hydroxyl group relative to the hydrogen at the ring junction of the fused bicyclic compound. A (cis) alcohol refers to a compound in which the hydroxyl group is on the same side (or face) as the hydrogen at the ring junction, and a (trans) alcohol refers to a compound in which the hydroxyl is on the opposite side (or face) of the fused bicyclic product from the hydrogen at the ring junction.
[0077] "Chiral" refers to a molecule that has the property of not being superimposable on its mirror image partner, while the term "achiral" refers to a molecule that is superimposable on its mirror image partner.
[0078] "Coding sequence" refers to a portion of a nucleic acid (eg, a gene) that encodes the amino acid sequence of a polypeptide.
[0079] "Codon optimization" refers to the alteration of codons in a protein-encoding polynucleotide to those preferentially used in a particular organism so that the encoded protein is efficiently expressed in the organism of interest. While the genetic code is degenerate in that most amino acids are represented by several codons, termed "synonymous" or "homologous" codons, it is well known that codon usage by a particular organism is not random but is biased toward certain codon triplets. This codon usage bias can be higher for a given gene, for genes of common function or ancestral origin, for highly expressed proteins and low copy number proteins, and for aggregated protein-coding regions of an organism's genome. In some embodiments, polynucleotides encoding the ketoreductase enzymes herein can be codon-optimized for optimal production from the host organism selected for expression.
[0080] The terms "preferred," "optimal," "high," or "codon usage bias," when used with respect to codons, are used interchangeably to refer to codons that are used more frequently in protein-coding regions than other codons that encode the same amino acid. Preferred codons can be determined with respect to codon usage in a single gene, a set of genes of common function or origin, highly expressed genes, codon frequency in aggregate protein-coding regions of all organisms, codon frequency in aggregate protein-coding regions of related organisms, or combinations thereof. Codons whose frequency increases with the level of gene expression are typically optimal codons for expression. Various methods are known for determining codon frequencies (e.g., codon usage, relative synonymous codon usage) and codon preference in a particular organism, including, for example, cluster analysis or correspondence analysis, and multivariate analysis using the effective number of codons used in a gene (see, for example, GCG Codon Preference, Genetics Computer Group Wisconsin Package; CodonW, John Peden, University of Nottingham; McLenemey, J.O., 1998, Bioinformatics 14:372-73; Stenico et al., 1994, Nucleic Acids Res. 222437-46; Wright, F., 1990, Gene 87:23-29). Codon usage tables are available for a growing list of organisms (see, e.g., Wada et al., 1992, Nucleic Acids Res. 20:2111-2118; Nakamura et al., 2000, Nucl. Acids Res. 28:292; Duret, et al. (ibid.); Henaut and Danchin, "Escherichia coli and Salmonella", 1996, Neidhardt, et al. Eds., ASM Press, Washington DC, pp. 2047-2066). Data sources for obtaining codon usage can rely on any available nucleotide sequence capable of encoding a protein.These datasets include nucleic acid sequences that are actually known to encode expressed proteins (e.g., complete protein-coding sequences - CDS), expressed sequence tags (ESTs), or predicted coding regions of genomic sequences (see, e.g., Mount, D., Bioinformatics: Sequence and Genome Analysis, Chapter 8, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 2001; Uberbacher, EC, 1996, Methods Enzymol. 266:259-281; Tiwari et al., 1997, Comput. Appl. Biosci. 13:263-270).
[0081] The terms "cofactor regeneration system" and "cofactor recycling system" can be used interchangeably and refer to a reaction that reduces the oxidized form of a cofactor (e.g., NADP + The term "cofactor regeneration system" refers to a series of reactions involved in the reduction of a keto substrate (from NADPH to NADPH). Cofactors oxidized by the reduction of a keto substrate catalyzed by a ketoreductase are regenerated in reduced form by a cofactor regeneration system. The cofactor regeneration system includes a source of reducing hydrogen equivalents and a stoichiometric reducing agent capable of reducing the oxidized form of the cofactor. The cofactor regeneration system may further include a catalyst, for example, an enzyme catalyst that catalyzes the reduction of the oxidized form of the cofactor by the reducing agent. NAD + or NADP + Cofactor regeneration systems for regenerating NADH or NADPH from, respectively, are known in the art and can be used in the methods described herein.
[0082] A "comparison window" refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acid residues, where a sequence can be compared to a reference sequence of at least 20 contiguous nucleotides or amino acids, and the portion of the sequence in the comparison window can contain no more than 20% additions or deletions (i.e., gaps) compared to the reference sequence (no additions or deletions) for optimal alignment of the two sequences. The comparison window can be longer than 20 contiguous residues, including windows containing 30, 40, 50, 100, or more residues.
[0083] "Conservative" amino acid substitutions or mutations refer to the interchangeability of residues with similar side chains and thus typically involve substituting an amino acid in a polypeptide with an amino acid within the same or a similar defined amino acid class. However, as used herein, conservative mutations do not include substitutions of hydrophilic residues for hydrophilic residues, hydrophobic residues for hydrophobic residues, hydroxyl-containing residues for hydroxyl-containing residues, or small residues for small residues, although conservative mutations may alternatively be substitutions of aliphatic residues for aliphatic residues, nonpolar residues for nonpolar residues, polar residues for polar residues, acidic residues for acidic residues, basic residues for basic residues, aromatic residues for aromatic residues, or constrained residues for constrained residues. Furthermore, as used herein, A, V, L, or I can be conservatively mutated to another aliphatic residue or another nonpolar residue. Table 1 below shows exemplary conservative substitutions.
[0084] [Table 1]
[0085] "Constrained amino acid or residue" refers to an amino acid or residue that has a constrained geometry. As used herein, constrained residues include L-Pro (P) and L-His (H). Histidine has a constrained geometry because it has a relatively small imidazole ring. Proline also has a constrained geometry because it has a five-membered ring.
[0086] The term "control sequences" is defined herein to include all components necessary or advantageous for expression of a polypeptide of the present disclosure. Each control sequence may be native or foreign to the nucleic acid sequence encoding the polypeptide. Such control sequences include, but are not limited to, a leader, polyadenylation sequence, propeptide sequence, promoter, signal peptide sequence, and transcription terminator. At a minimum, control sequences include a promoter, and transcriptional and translational stop signals. Control sequences may also be provided with linkers for the purpose of introducing specific restriction sites facilitating ligation of the control sequences with the coding region of the nucleic acid sequence encoding the polypeptide.
[0087] "Conversion" refers to the enzymatic reduction of a substrate to the corresponding product. "Conversion rate" refers to the proportion of a substrate that is reduced to a product under specific conditions within a given time period. Thus, the "enzyme activity" or "activity" of a ketoreductase polypeptide can be expressed as the "conversion rate" of a substrate to a product.
[0088] In this disclosure, the words "comprise," "comprising," "containing," and "having," etc., may have the meaning ascribed to them in U.S. patent law, and "consisting essentially of" or "consisting essentially of" likewise have the meaning set forth in U.S. patent law, and the term is open-ended, permitting the presence of more than what is recited, but excluding prior art embodiments, so long as the basic or novel characteristics of the recited item are not changed by the presence of more than what is recited.
[0089] A "deletion" refers to a modification in a polypeptide or polynucleotide characterized by the removal of one or more amino acids from a reference (e.g., parent) polypeptide or one or more nucleic acids from a reference (e.g., parent) polynucleotide. A deletion can include the removal of 1 or more, 2 or more, 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, or 50 or more amino acids or nucleic acids. A deletion can also include the removal of up to 10% of the total number of amino acids or nucleic acids comprising the reference enzyme, or up to 20% of the total number of amino acids or nucleic acids, while retaining enzymatic activity and / or the improved properties of the engineered ketoreductase enzyme. Deletions can be directed to internal and / or terminal portions of a polypeptide or polynucleotide. In various embodiments, a deletion can include a contiguous segment or can be discontinuous.
[0090] "Derived from" or "originating from," when used herein in reference to an engineered ketoreductase enzyme, identifies the source ketoreductase enzyme (e.g., a wild-type ketoreductase enzyme) and / or the gene encoding such ketoreductase enzyme on which the modification was based. For example, the engineered ketoreductase enzyme of SEQ ID NO:426 was obtained by artificially evolving over multiple generations the gene encoding the L. kefir ketoreductase enzyme of SEQ ID NO:492. Thus, this engineered ketoreductase enzyme is "derived from" or "originates from" the wild-type ketoreductase of SEQ ID NO:492.
[0091] With respect to a designated reference (e.g., parent) sequence, "Different from" or "differs from" refers to the differences in a given polypeptide or polynucleotide sequence when aligned with the reference (e.g., parent) sequence. Generally, the differences can be determined when the two sequences are optimally aligned. Differences include insertions, deletions, or substitutions of amino acid or nucleic acid residues compared to the reference sequence.
[0092] As used herein, "engineered ketoreductase" refers to a ketoreductase having a variant sequence produced by human manipulation (e.g., a sequence produced by directed evolution of a naturally occurring parent enzyme or by directed evolution of a variant that arose from a previously naturally occurring enzyme).
[0093] "Fragment" refers to a portion of a polypeptide or nucleic acid molecule. This portion preferably contains at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full-length sequence of the reference nucleic acid molecule or polypeptide. A fragment contains at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more nucleotides or amino acids. In some embodiments, a fragment has an amino- and / or carboxy-terminal deletion, with the remaining amino acid sequence being identical to corresponding positions in the reference sequence.
[0094] "Heterologous polynucleotide" refers to any polynucleotide that is introduced into a host cell by laboratory techniques, and includes polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into the host cell.
[0095] A "hydrophilic amino acid or residue" refers to an amino acid or residue having a side chain exhibiting a hydrophobicity of less than 0 according to the normalized consensus hydrophobicity scale of Eisenberg et al., 1984, J. Mal. Biol. 179:125-142. Genetically encoded hydrophilic amino acids include L-Thr (T), L-Ser (S), L-His (H), L-Glu (E), L-Asn (N), L-Gln (Q), L-Asp (D), L-Lys (K), and L-Arg (R).
[0096] A "hydrophobic amino acid or residue" refers to an amino acid or residue having a side chain exhibiting a hydrophobicity greater than 0 according to the normalized consensus hydrophobicity scale of Eisenberg et al., 1984, J. Mal. Biol. 179:125-142. Genetically encoded hydrophobic amino acids include L-Pro (P), L-Ile (I), L-Phe (F), L-Val (V), L-Leu (L), L-Trp (W), L-Met (M), L-Ala (A), and L-Tyr (Y).
[0097] "Hydroxyl-containing amino acid or residue" refers to an amino acid that contains a hydroxyl (-OH) moiety. Genetically encoded hydroxyl-containing amino acids include L-Ser (S), L-Thr (T), and L-Tyr (Y).
[0098] "Improved enzymatic properties" refer to a ketoreductase polypeptide that exhibits an improvement in any enzymatic property compared to a reference (e.g., parent) ketoreductase. The engineered ketoreductase polypeptides described herein are generally compared to a wild-type ketoreductase enzyme (e.g., L. kefir), although in some embodiments, the reference (e.g., parent) ketoreductase can be another improved engineered ketoreductase (e.g., SEQ ID NO: 54). Enzymatic properties for which improvement is desirable include enzyme activity (which can be expressed as percent conversion of substrate), thermostability, solvent stability, pH-activity profile, cofactor requirements, refractoriness to inhibitors (e.g., product inhibition), stereospecificity (including diastereospecificity or enantiospecificity), and stereoselectivity (including diastereoselectivity or enantioselectivity). In some embodiments, the engineered ketoreductase polypeptide exhibits increased enzymatic activity (e.g., increased % conversion). In some embodiments, the engineered ketoreductase polypeptide exhibits a reversal of stereoselectivity (e.g., a reversal of enantioselectivity or a reversal of diastereoselectivity). In some embodiments, an increase in enzyme activity is an increase in percent conversion improvement over a positive control (FIOP) or an increase in percent improvement of desired product over a positive control (FIOP). In some embodiments, an improved enzyme property is an increase in selectivity (e.g., an increase in the desired product). In some embodiments, an increase in selectivity is an increase in the proportion of the desired product or an increase in the % FIOP of the desired product.
[0099] "Increased enzymatic activity" refers to improved properties of an engineered ketoreductase polypeptide, which can be expressed as an increase in specific activity (e.g., product / time / weight protein) or an increase in the rate of conversion of substrate to product (e.g., the rate of conversion of starting amount of substrate to product in a specified period of time using a specified amount of KRED) compared to a reference (e.g., parent) ketoreductase enzyme. Exemplary methods for determining enzymatic activity are provided in the Examples. Any property associated with enzymatic activity can be affected, including the classical enzymatic properties of Km, Vmax, or kcat, and changes therein can result in increased enzymatic activity. The improved enzymatic activity can be about 1.5-fold to 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold, 500-fold, 1000-fold, 3000-fold, 5000-fold, 7000-fold, or greater than the enzymatic activity of a corresponding wild-type ketoreductase enzyme, or another engineered ketoreductase derived from a naturally occurring ketoreductase or ketoreductase polypeptide. In certain embodiments, the engineered ketoreductase enzyme exhibits improved enzymatic activity in the range of 150-3000-fold, 3000-7000-fold, or greater than 7000-fold over the enzymatic activity of the parent ketoreductase enzyme. One of skill in the art will understand that the activity of any enzyme is diffusion-limited, meaning that the catalytic turnover rate cannot exceed the diffusion rate of the substrate, including any required cofactors. The theoretical maximum diffusion limit, i.e., kca / Km, is generally about 108-109 (M-1 s-1). Therefore, improvements in the enzymatic activity of a ketoreductase have an upper limit related to the diffusion rate of the substrate acted upon by the ketoreductase enzyme. Ketoreductase activity can be measured by any one of the standard assays used to measure ketoreductases, such as the decrease in absorbance or fluorescence of NADPH due to oxidation with simultaneous reduction of the ketone to alcohol, or by-products produced in a binding assay. Enzyme activity comparisons are performed using a given enzyme preparation, a given assay under set conditions, and one or more given substrates, as described in further detail herein.Generally, when comparing lysates, one determines the number of cells and the amount of protein assayed, as well as the use of the same expression system and the same host cells, to minimize variation in the amount of enzyme produced by the host cells and present in the lysates.
[0100] "Insertion" refers to a modification to a polypeptide from a reference (e.g., parent) polypeptide by the addition of one or more amino acids. In some embodiments, improved modified oxynitrilases include insertions of one or more amino acids into naturally occurring ketoreductase polypeptides, as well as insertions of one or more amino acids into other modified ketoreductase polypeptides. Insertions can be made in an internal portion of the polypeptide, or at the carboxy- or amino-terminus. As used herein, insertions include fusion proteins, as known in the art. Insertions can be contiguous segments of amino acids or separated by one or more amino acids in naturally occurring polypeptides.
[0101] The terms "isolated" or "purified" refer to material that is free, to varying degrees, from components that normally accompany it when found in its native state. "Isolate" indicates some degree of separation from the original source or environment. "Purify" indicates a greater degree of separation than isolation. A "purified" protein has been sufficiently removed from other materials so that any impurities do not substantially affect the biological properties of the protein or cause other adverse consequences. That is, a nucleic acid or peptide of the invention is purified if it is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA technology, or substantially free of chemical precursors or other chemicals when chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purified" can mean that the nucleic acid or protein gives rise to substantially one band in an electrophoretic gel. In the case of proteins that can be modified, such as phosphorylation or glycosylation, different modifications may result in different isolated proteins that can be separately purified. "Substantially purified" refers to a composition in which a polynucleotide or polypeptide species is the predominant species present (i.e., more abundant than other individual macromolecular species in the composition, on a molar or weight basis). A substantially purified composition generally occurs when the polynucleotide or polypeptide species comprises at least about 50% of the macromolecular species present, on a molar or weight percent basis. Generally, a substantially pure ketoreductase composition comprises about 60%, 70%, 80%, 90%, 95%, 98%, or more, on a molar or weight percent basis, of all macromolecular species present in the composition. In some embodiments, the target species is purified to substantial homogeneity (i.e., contaminant species cannot be detected in the composition by conventional detection methods), and the composition consists essentially of a single macromolecular species. Solvent species, small molecules (<500 Daltons), and elemental ion species are not considered macromolecular species. In some embodiments, an isolated ketoreductase polynucleotide or polypeptide is a substantially pure polynucleotide or polypeptide composition.
[0102] An "isolated polynucleotide" refers to a nucleic acid that is free of the genes that flank it in the naturally occurring genome of the organism from which the nucleic acid molecule of the invention is derived. Thus, the term includes recombinant polynucleotides that are incorporated into a vector or that exist as separate molecules independent of other sequences. In some embodiments, any of the polynucleotides encoding the engineered ketoreductase polypeptides disclosed herein is an isolated polynucleotide.
[0103] By "isolated polypeptide" is meant a polypeptide of the invention that has been separated from components or other contaminants that naturally accompany it, such as proteins, lipids, and polynucleotides. Typically, a polypeptide is isolated when it is at least 60%, at least 75%, at least 90%, or at least 99%, by weight, free from the proteins and naturally-occurring organic molecules with which it is naturally associated. Purity can be measured by any appropriate method, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis. The term includes polypeptides that have been removed or purified from their naturally occurring environment or expression system (e.g., host cells or in vitro synthesis), or polypeptides that have been chemically synthesized. The ketoreductase enzymes disclosed herein can be present intracellularly, in cell culture medium, or prepared in various forms, such as lysates or isolated preparations. Thus, in some embodiments, the engineered ketoreductase polypeptides disclosed herein are isolated polynucleotides.
[0104] "Ketoreductase" and "KRED" are used interchangeably herein to refer to a polypeptide having the enzymatic ability to reduce a carbonyl group to its corresponding alcohol. In some embodiments, the ketoreductase enzyme reduces a substrate to a (trans) alcohol product. For example, a ketoreductase polypeptide may be capable of reducing tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (compound 1 or substrate) to tert-butyl rel-(3aR,5r,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate ((5r)-2 or product). In some embodiments, the ketoreductase enzyme reduces a substrate to a (cis) alcohol product. In certain embodiments, a ketoreductase enzyme reduces a bicyclic ketone substrate to an alcohol product, whereby the hydroxyl group is cis relative to the substituent attached to the ring junction of the bicyclic ketone substrate. For example, as disclosed herein, an engineered ketoreductase polypeptide of the invention can reduce tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (compound 1 or substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate ((5s)-2) or product). Ketoreductase polypeptides typically utilize a cofactor, such as reduced nicotinamide adenine dinucleotide (NADH) or reduced nicotinamide adenine dinucleotide phosphate (NADPH), as the reducing agent. As used herein, "ketoreductase" includes naturally occurring (wild-type) ketoreductases (e.g., L. kefir) and engineered ketoreductase polypeptides produced by human engineering that do not occur in nature (e.g., SEQ ID NO: 256).
[0105] As used herein, the term amine protecting group (PG) in the compounds of the present disclosure refers to a group that protects the relevant functional group from undesired secondary reactions, such as acylation, etherification, esterification, oxidation, solvolysis, and similar reactions. It can be removed under deprotection conditions. Depending on the protecting group used, those skilled in the art will know how to remove the protecting group to obtain the free amine NH or NH group by referring to known procedures. These include references to procedures in organic chemistry textbooks and literature, such as J.F.W.M. Comie, "Protective Groups in Organic Chemistry", Plenum Press, London and New York 1973; T.W. Greene and P.G.W. Muts, "Greene's Protective Groups in Organic Synthesis", Fourth Edition, Wiley, New York 2007; "The Peptides", Volume 3 (editors: E. Gross and J. Meienhofer), Academic Press, London and New York 1981; P.J. Kocienski, "Protecting Groups", Third Edition, Georg Thieme Verlag, Stuttgart and New York 2005; "Methoden der organischen Chemie" (Methods of Organic Chemistry), Houben Weyl, 4th edition, Volume 15 / I, Georg Thieme Verlag, Stuttgart 1974.
[0106] Preferred amine protecting groups in the compounds of the present disclosure generally include: trialkylsilyl-C1-C7 alkoxy (e.g., trimethylsilylethoxy), C1-C6 alkyl (e.g., tert-butyl), preferably C1-C4 alkyl, more preferably C1-C2 mono-, di-, or tri-substituted with aryl, preferably phenyl, or heterocyclic groups (e.g., benzyl, cumyl, benzhydryl, pyrrolidinyl, trityl, pyrrolidinylmethyl, 1-methyl-1,1-dimethylbenzyl, (phenyl)methylbenzene). Alkyl, most preferably C1 alkyl (wherein the aryl ring or heterocyclic group is unsubstituted or substituted with, for example, C1-C7 alkyl, hydroxy, C1-C7 alkoxy (e.g., para-methoxybenzyl (PMB)), C2-C8-alkanoyl-oxy, halogen, nitro, cyano, and CF3, aryl-C1-C2-alkoxycarbonyl (preferably phenyl-C1-C2-alkoxycarbonyl (e.g., benzyloxycarbonyl (Cbz), benzyloxymethyl (BOM), pivaloyloxymethyl (POM))), C1-C 10 -alkenyloxycarbonyl, C1-C6 alkylcarbonyl (e.g., acetyl or pivaloyl), C6-C 10 -arylcarbonyl; C1-C6-alkoxycarbonyl (e.g., tert-butyloxycarbonyl (Boc), methylcarbonyl, trichloroethoxycarbonyl (Troc), pivaloyl (Piv), allyloxycarbonyl), C6-C 10 -aryl C1-C6-alkoxycarbonyl (e.g., 9-fluorenylmethyloxycarbonyl (Fmoc)), allyl or cinnamyl, sulfonyl or sulfenyl, succinimidyl group, silyl group (e.g., substituted by one or more, for example two or three, residues selected from the group consisting of triarylsilyl, trialkylsilyl, triethylsilyl (TES), trimethylsilylethoxymethyl (SEM), trimethylsilyl (TMS), triisopropylsilyl, or tertbutyldimethylsilyl).
[0107] According to the present disclosure, preferred protecting groups (PG) can be selected from the group including tert-butyloxycarbonyl (Boc), benzyloxycarbonyl (Cbz), para-methoxybenzyl (PMB), 2,4-dimethoxybenzyl (DMB), methyloxycarbonyl, trimethylsilylethoxymethyl (SEM), and benzyl. The amine protecting group (PG) is preferably an acid-labile protecting group (which can be removed in the presence of an acid such as HCl or TFA), such as tert-butyloxycarbonyl (Boc), 2,4-dimethoxybenzyl (DMB), or benzyloxycarbonyl (Cbz).
[0108] The term "substituted" means that the specified group or moiety has one or more suitable substituents, and the substituents may be attached to the specified group or moiety at one or more positions. For example, an aryl substituted with a cycloalkyl may indicate that the cycloalkyl is bonded to one atom of the aryl or is fused to the aryl and shares two or more common atoms.
[0109] In the groups, radicals, or moieties defined below, the number of carbon atoms is often specified before the group, for example, C1-C8 alkyl means an alkyl group or radical having from 1 to 8 carbon atoms. Generally, for substituents containing more than one subgroup, the last named group is the point of attachment of the group, for example, "alkylaryl" means a monovalent group of formula alkyl-aryl-, while "arylalkyl" means a monovalent group of formula aryl-alkyl-.
[0110] The term "halogen" or "halo" means fluorine, chlorine, bromine, or iodine.
[0111] The term "alkyl" as used herein refers to a branched or straight-chain saturated hydrocarbon group having, for example, 1 to 20 carbon atoms, such as C1-C3 alkyl, C1-C6 alkyl, C2-C8 alkyl, C3-C8 alkyl, C1-C8 alkyl, C1-C 10 Alkyl, C1-C 20alkyl, etc. Representative examples include methyl, ethyl, propyl (e.g., prop-1-yl, prop-2-yl (or isopropyl)), butyl (e.g., 2-methylprop-2-yl (or tert-butyl), but-1-yl, but-2-yl), pentyl (e.g., pent-1-yl, pent-2-yl, pent-3-yl), 2-methylbut-1-yl, 3-methylbut-1-yl, hexyl (e.g., hex-1-yl), heptyl (e.g., hept-1-yl), octyl (e.g., oct-1-yl), nonyl (e.g., non-1-yl), and the like.
[0112] The term "alkenyl" as used herein refers to a branched or straight-chain hydrocarbon group having at least one double bond, for example, 2 to 20 carbon atoms and having at least one double bond, such as C2-C3 alkenyl, C2-C6 alkenyl, C2-C7 alkenyl, C2-C8 alkenyl, C3-C5 alkenyl, C1-C 10 -Alkenyl, C1-C 20 It represents alkenyl, etc. Representative examples are ethenyl (or vinyl), propenyl (e.g., prop-1-enyl, prop-2-enyl), butadienyl (e.g., buta-1,3-dienyl), butenyl (e.g., but-1-en-1-yl, but-2-en-1-yl), pentenyl (e.g., pent-1-en-1-yl, pent-2-en-2-yl), hexenyl (e.g., hex-1-en-2-yl, hex-2-en-1-yl), 1-ethylprop-2-enyl, 1,1-(dimethyl)prop-2-enyl, 1-ethylbut-3-enyl, 1,1-(dimethyl)but-2-enyl, and the like.
[0113] The term "alkynyl" as used herein refers to a branched or straight chain hydrocarbon group having at least one triple bond, for example, having 2 to 20 carbon atoms and having at least one triple bond, such as C2-C3 alkynyl, C2-C6 alkynyl, C2-C7 alkynyl, C2-C8 alkynyl, C3-C5 alkynyl, C1-C 10 Alkynyl, C1-C 20It represents alkynyl, etc. Representative examples are ethynyl, propynyl (e.g., prop-1-ynyl, prop-2-ynyl), butynyl (e.g., but-1-ynyl, but-2-ynyl), pentynyl (e.g., pent-1-ynyl, pent-2-ynyl), hexynyl (e.g., hex-1-ynyl, hex-2-ynyl), 1-ethylprop-2-ynyl, 1,1-(dimethyl)prop-2-ynyl, 1-ethylbut-3-ynyl, 1,1-(dimethyl)but-2-ynyl, and the like.
[0114] As used herein, the term "aryl" is intended to include monocyclic, bicyclic, or polycyclic carbocyclic aromatic rings. Representative examples include phenyl, naphthyl (e.g., naphth-1-yl, naphth-2-yl), anthryl (e.g., anthr-1-yl, anthr-9-yl), phenanthryl (e.g., phenanthr-1-yl, phenanthr-9-yl), and the like. Aryl is also intended to include monocyclic, bicyclic, or polycyclic carbocyclic aromatic rings substituted with a carbocyclic aromatic ring. Representative examples include biphenyl (e.g., biphenyl-2-yl, biphenyl-3-yl, biphenyl-4-yl), phenylnaphthyl (e.g., 1-phenylnaphth-2-yl, 2-phenylth-1-yl), and the like. Aryl is also intended to include bicyclic or polycyclic partially saturated carbocyclic rings having at least one unsaturated moiety (e.g., a benzo moiety). Representative examples are indanyl (e.g., indan-1-yl, indan-5-yl), indenyl (e.g., inden-1-yl, inden-5-yl), 1,2,3,4-tetrahydronaphthyl (e.g., 1,2,3,4-tetrahydronaphth-1-yl, 1,2,3,4-tetrahydronaphth-2-yl, 1,2,3,4-tetrahydronaphth-6-yl), 1,2-dihydronaphthyl (e.g., 1,2-dihydronaphth-1-yl, 1,2-dihydronaphth-4-yl, 1,2-dihydronaphth-6-yl), fluorenyl (e.g., fluoren-1-yl, fluoren-4-yl, fluoren-9-yl), and the like. Aryl is also intended to include bicyclic or polycyclic partially saturated carbocyclic aromatic rings containing one or two bridges. Representative examples are benzonorbornyl (e.g., benzoborn-3-yl, benzoborn-6-yl), 1,4-ethano-1,2,3,4-tetrahydronaphthyl (e.g., 1,4-ethano-1,2,3,4-tetrahydronaphth-2-yl, 1,4-ethano-1,2,3,4-tetrahydronaphth-10-yl), and the like. Aryl is also intended to include bicyclic or polycyclic partially saturated carbocyclic aromatic rings containing one or more spiro atoms.Representative examples are spiro[cyclopentane-1,1'-indan]-4-yl, spiro[cyclopentane-1,1'inden]-4-yl, spiro[piperidine-4,1'-indan]-1-yl, spiro[piperidine-3,2'-indan]-1-yl, spiro[piperidine-4,2'-indan]-1-yl, spiro[piperidine-4,1'-indan]-3'-yl, spiro[pyrrolidin-3,2'-indan]-1-yl, spiro[pyrrolidin-3,2'-indan]-3'-yl, spiro[pyrrolidin-4,1'-indan]-3' ... spiro[pyrrolidin-3,1'-(3',4'-dihydronaphthalene)]-1-yl, spiro[piperidine-3,1'-(3',4'-dihydronaphthalene)]-1-yl, spiro[piperidine-4,1'-(3',4'-dihydronaphthalene)]-1-yl, spiro[imidazolidine-4,2'-indan]-1-yl, spiro[piperidine-4,1'-inden]-1-yl, and the like.
[0115] Terms C6~C 14 Aryl should be interpreted accordingly.
[0116] Preferably, aryl refers to a monocyclic or bicyclic carbocyclic aromatic ring.
[0117] Preferred examples of aryl include, but are not limited to, phenyl and naphthyl. In one embodiment, aryl is phenyl.
[0118] As used herein, the term "heteroaryl" is intended to include monocyclic heterocyclic aromatic rings containing one or more heteroatoms selected from oxygen, nitrogen, and sulfur (O, N, and S). Representative examples are pyrrolyl, furanyl, thienyl, oxazolyl, thiazolyl, imidazolyl, pyrazolyl, isothiazolyl, isoxazolyl, triazolyl (e.g., 1,2,4-triazolyl), oxadiazolyl (e.g., 1,2,3-oxadiazolyl, 1,2,4-oxadiazolyl, 1,2,5-oxadiazolyl, 1,3,4-oxadiazolyl), thiadiazolyl (e.g., 1,2,3-thiadiazolyl, 1,2,4-thiadiazolyl, 1,2,5-thiadiazolyl, 1,3,4-thiadiazolyl), tetrazolyl, pyranyl, pyridinyl, pyridazinyl, pyrimidinyl, pyrazinyl, 1,2,3-triazinyl, 1,2,4-triazinyl, 1,3,5-triazinyl, thiadiazinyl, azepinyl, azesinyl, and the like.
[0119] Heteroaryl is also intended to include bicyclic heteroaromatic rings containing one or more heteroatoms selected from oxygen, nitrogen, and sulfur (O, N, and S). Representative examples are indolyl, isoindolyl, benzofuranyl, benzothiophenyl, indazolyl, benzopyranyl, benzimidazolyl, benzothiazolyl, benzisothiazolyl, benzoxazolyl, benzisoxazolyl, benzoxazinyl, benzotriazolyl, naphthyridinyl, phthalazinyl, pteridinyl, purinyl, quinazolinyl, cinnolinyl, quinolinyl, isoquinolinyl, quinoxalinyl, oxazolopyridinyl, isoxazolopyridinyl, pyrrolopyridinyl, furopyridinyl, thienopyridinyl, imidazopyridinyl, imidazopyrimidinyl, pyrazolopyridinyl, pyrazolopyrimidinyl, pyrazolotriazinyl, thiazolopyridinyl, thiazolopyrimidinyl, imidazothiazolyl, triazolopyridinyl, triazolopyrimidinyl, and the like.
[0120] Heteroaryl is also intended to include polycyclic heteroaromatic rings containing one or more heteroatoms selected from oxygen, nitrogen, and sulfur (O, N, and S). Representative examples are carbazolyl, phenoxazinyl, phenazinyl, acridinyl, phenothiazinyl, carbolinyl, phenanthrolinyl, and the like.
[0121] Heteroaryl is also intended to include partially saturated monocyclic, bicyclic, or polycyclic heterocyclyl containing one or more heteroatoms selected from oxygen, nitrogen, and sulfur (O, N, and S). Representative examples are imidazolinyl, indolinyl, dihydrobenzofuranyl, dihydrobenzothienyl, dihydrobenzopyranyl, dihydropyridoxazinyl, dihydrobenzodioxinyl (e.g., 2,3-dihydrobenzo[b][1,4]dioxinyl), benzodioxolyl (e.g., benzo[d][1,3]dioxole), dihydrobenzoxazinyl (e.g., 3,4-dihydro-2H-benzo[b][1,4]oxazine), tetrahydroindazolyl, tetrahydrobenzimidazolyl, tetrahydroimidazo[4,5-C]pyridyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, tetrahydroquinoxalinyl, and the like.
[0122] The heteroaryl ring structure may be optionally substituted with one or more substituents, which may themselves be substituted. The heteroaryl ring may be attached via a carbon atom or a heteroatom.
[0123] The term "5- to 20-membered heteroaryl" should be construed accordingly.
[0124] The term "monocyclic heteroaryl," as used herein, is intended to include monocyclic heteroaromatic rings as defined above.
[0125] The term "bicyclic heteroaryl," as used herein, is intended to include bicyclic heteroaromatic rings as defined above.
[0126] Examples of 5- to 20-membered heteroaryls include indolyl, imidazopyridyl, isoquinolinyl, benzoxazolonyl, pyridinyl, pyrimidinyl, pyridinonyl, benzotriazolyl, pyridazinyl, pyrazolotriazinyl, indazolyl, benzimidazolyl, quinolinyl, triazolyl (e.g., 1,2,4-triazolyl), pyrazolyl, thiazolyl, oxazolyl, isoxazolyl, pyrrolyl, oxadiazolyl (e.g., 1,2,3-oxadiazolyl, 1,2,4-oxadiazolyl, 1,2,5-oxadiazolyl, 1,3,4-oxadiazolyl), imidazolyl, pyrrolopyridinyl, tetrahydroindazolyl, and quinoxalinyl. , thiadiazolyl (e.g., 1,2,3-thiadiazolyl, 1,2,4-thiadiazolyl, 1,2,5-thiadiazolyl, 1,3,4-thiadiazolyl), pyrazinyl, oxazolopyridinyl, pyrazolopyrimidinyl, benzoxazolyl, indolinyl, isoxazolopyridinyl, dihydropyridoxazinyl, tetrazolyl, dihydrobenzodioxinyl (e.g., 2,3-dihydrobenzo[b][1,4]dioxinyl), benzodioxolyl (e.g., benzo[d][1,3]dioxole), and dihydrobenzoxazinyl (e.g., 3,4-dihydro-2H-benzo[b][1,4]oxazine).
[0127] The term "heterocyclyl," as used herein, refers to a saturated or partially saturated monocyclic or polycyclic ring containing one or more heteroatoms selected from nitrogen, oxygen, sulfur, S(=O), and S(=O)2, where there are no delocalized pi-electrons (aromaticity) shared between the ring carbons or heteroatoms. The heterocyclyl ring structure may be substituted with one or more substituents. The substituents may themselves be substituted. The heterocyclyl may be attached via a carbon atom or a heteroatom. The term polycyclic encompasses bridged, fused, and spirocyclic heterocyclyls.
[0128] Representative examples include aziridinyl (e.g., aziridin-1-yl), azetidinyl (e.g., azetidin-1-yl, azetidin-3-yl), oxetanyl, pyrrolidinyl (e.g., pyrrolidin-1-yl, pyrrolidin-2-yl, pyrrolidin-3-yl), imidazolidinyl (e.g., imidazolidin-1-yl, imidazolidin-2-yl, imidazolidin-4-yl), oxazolidinyl (e.g., oxazolidinyl ... oxazolidin-2-yl, oxazolidin-3-yl, oxazolidin-4-yl), thiazolidinyl (e.g., thiazolidin-2-yl, thiazolidin-3-yl, thiazolidin-4-yl), isothiazolidinyl, piperidinyl (e.g., piperidin-1-yl, piperidin-2-yl, piperidin-3-yl, piperidin-4-yl), homopiperidinyl (e.g., homopiperidin-1-yl, homopiperidin-1-yl, homopiperidin-2-yl, homopiperidin-3-yl, homopiperidin-4-yl), thiazolidinyl (e.g., thiazolidin-2-yl, thiazolidin-3-yl, thiazolidin-4-yl), isothiazolidinyl, piperidinyl (e.g., piperidin-1-yl, piperidin-2-yl, piperidin-3-yl, piperidin-4-yl), homopiperidinyl (e.g., homopiperidin-1-yl, homopiperidin-4-yl), homopiperidinyl (e.g., homopiperidin-1-yl, homopiperidin-4-yl), homopiperidinyl (e.g., homopiperidin-2-yl, homopiperidin-3 ... tetrahydrofuranyl (e.g., tetrahydrofuran-2-yl, tetrahydrofuran-3-yl), tetrahydrothienyl, tetrahydro-1,1-dioxothienyl, tetrahydropyranyl (e.g., 2-tetrahydropyranyl), tetrahydrothiopyranyl (e.g., 2-tetrahydrothiopyranyl), 1,4-dioxanyl, 1,3-dioxanyl, and the like. Heterocyclyl is also intended to represent a saturated 6-8 membered bicyclic ring containing one or more heteroatoms selected from nitrogen, oxygen, sulfur, S(=O), and S(=O)2.Representative examples are octahydroindolyl (e.g., octahydroindol-1-yl, octahydroindol-2-yl, octahydroindol-3-yl, octahydroindol-5-yl), decahydroquinolinyl (e.g., decahydroquinolin-1-yl, decahydroquinolin-2-yl, decahydroquinolin-3-yl, decahydroquinolin-4-yl, decahydroquinolin-6-yl), decahydroquinoxalinyl (e.g., decahydroquinoxalin-1-yl, decahydroquinoxalin-2-yl, decahydroquinoxalin-6-yl), etc. Heterocyclyl is also intended to represent a saturated 6- to 8-membered ring containing one or more heteroatoms selected from nitrogen, oxygen, sulfur, S(=O), and S(=O)2, and having one or two bridges. Representative examples are 3-azabicyclo[3.2.2]nonyl, 2-azabicyclo[2.2.1]heptyl, 3-azabicyclo[3.1.0]hexyl, 2,5-diazabicyclo[2.2.1]heptyl, atropinyl, tropinyl, quinuclidinyl, 1,4-diazabicyclo[2.2.2]octanyl, etc. Heterocyclyl is also intended to represent a 6-8 membered saturated ring containing one or more heteroatoms selected from nitrogen, oxygen, sulfur, S(=O), and S(=O)2, and having one or more spiro atoms.Representative examples include 1,4-dioxaspiro[4.5]decanyl (e.g., 1,4-dioxaspiro[4.5]decan-2-yl, 1,4-dioxaspiro[4.5]decan-7-yl), 1,4-dioxa-8-azaspiro[4.5]decanyl (e.g., 1,4-dioxa-8-azaspiro[4.5]decan-2-yl, 1,4-dioxa-8-azaspiro[4.5]decan-8-yl), 8-azaspiro[4.5]decanyl (e.g., 8-azaspiro[4.5]decan-1-yl, 8-azaspiro[4.5]decan-8-yl), 2-azaspiro[5.5]undecanyl (e.g., 2 -azaspiro[5.5]undecan-2-yl), 2,8-diazaspiro[4.5]decanyl (e.g., 2,8-diazaspiro[4.5]decan-2-yl, 2,8-diazaspiro[4.5]decan-8-yl), 2,8-diazaspiro[5.5]undecanyl (e.g., 2,8-diazaspiro[5.5]undecan-2-yl), 1,3,8-triazaspiro[4.5]decanyl (e.g., 1,3,8-triazaspiro[4.5]decan-1-yl, 1,3,8-triazaspiro[4.5]decan-3-yl, 1,3,8-triazaspiro[4.5]decan-8-yl), and the like.
[0129] The terms "6- to 12-membered heterocyclyl" and "3- to 14-membered heterocyclyl" should be construed accordingly.
[0130] As used herein, the term "cycloalkyl" means a monocyclic saturated or partially saturated carbocyclic ring, e.g., containing 3 to 10 carbon atoms, and having no delocalized pi electrons shared between ring carbons (aromaticity).
[0131] Representative examples include cyclopropenyl, cyclopropyl, cyclobutyl, cyclobutenyl, cyclopentyl, cyclohexyl, cycloheptanyl, cyclooctanyl, and the like.
[0132] The term "arylalkyl" (e.g., benzyl, phenylethyl, 3-phenylpropyl, 1-naphthylmethyl, 2-(1-naphthyl)ethyl, etc.) represents an aryl group, as defined above, attached through an alkyl chain or substituted alkyl group, as defined above, having the indicated number of carbon atoms. The term C7-C 20 Arylalkyl should be interpreted accordingly.
[0133] As used herein, the term "optional" or "optionally substituted" means that the described event or circumstance may or may not occur; for example, "optionally substituted aryl" refers to an aryl group that may be substituted or unsubstituted. This description includes both substituted and unsubstituted aryl groups.
[0134] Non-limiting exemplary naturally occurring (wild-type) ketoreductase enzymes include those derived from Lactobacillus kefir ("L. kefir"), Lactobacillus brevis ("L. brevis"), or Lactobacillus minor ("L. minor"). In some embodiments, the naturally occurring (wild-type) ketoreductase is derived from L. kefir. An exemplary nucleic acid (Accession No. QGV24812) and amino acid sequence (UniProKB Accession No. Q6WVP7) of an L. kefir ketoreductase is shown below: >ENA|QGV24812|QGV24812.1 Lactobacillus kefir SDR family NAD(P)-dependent oxidoreductase [ka] >UniProKB accession number Q6WVP7 [ka]
[0135] In some embodiments, the naturally occurring (wild-type) ketoreductase is from L. brevis. An exemplary amino acid sequence of an L. brevis ketoreductase (GENBANK Accession No. CAD66648) is shown below: >GenBank accession number CAD66648 [ka]
[0136] In some embodiments, the naturally occurring (wild-type) ketoreductase is from L. minor. An exemplary amino acid sequence of an L. minor ketoreductase (U.S. Patent Application Publication No. 2004 / 0265978) is shown below: >L. minor alcohol dehydrogenase [ka]
[0137] Non-naturally occurring engineered ketoreductases can contain one or more amino acid differences shown in Table 2, Table 4, Table 5, Table 6, or Table 7. Non-limiting examples of non-naturally occurring engineered ketoreductase polypeptides are shown in Table 4, Table 5, Table 6, or Table 7. Additional engineered ketoreductase polypeptides are described in WO 2010 / 025085 A2, which is incorporated herein by reference.
[0138] "Naturally occurring" or "wild-type" refers to a form found in nature. For example, a naturally occurring or wild-type polypeptide or polynucleotide sequence can be isolated from a natural source and is a sequence present in an organism that has not been intentionally modified by human manipulation. In some embodiments, the term "naturally occurring" or "wild-type" refers to a ketoreductase enzyme polypeptide. In some embodiments, a naturally occurring (wild-type) ketoreductase enzyme is derived from L. kefir, L. brevis, or L. minor. In some embodiments, a naturally occurring (wild-type) ketoreductase is derived from L. kefir.
[0139] A "non-conservative substitution" refers to the substitution or mutation of an amino acid in a polypeptide with an amino acid having significantly different side chain properties. A non-conservative substitution may use an amino acid between, rather than within, the defined groups listed above. In one embodiment, the non-conservative mutation affects (a) the structure of the peptide backbone in the region of substitution (e.g., proline for glycine), (b) the charge or hydrophobicity, or (c) the bulk of the side chain.
[0140] "Nonpolar amino acid or residue" refers to a hydrophobic amino acid or residue that is uncharged at physiological pH and has a side chain with a bond in which the electron pair shared by two atoms is generally held equally by each of the two atoms (i.e., the side chain is not polar). Genetically encoded nonpolar amino acids include L-Gly (G), L-Leu (L), L-Val (V), L-Ile (I), L-Met (M), and L-Ala (A).
[0141] "Operably linked" is defined herein as a configuration in which a control sequence is suitably positioned relative to the coding sequence of a DNA sequence such that the control sequence directs expression of a polynucleotide and / or polypeptide.
[0142] "Polar amino acid or residue" refers to a hydrophilic amino acid or residue that is uncharged at physiological pH but has a side chain with at least one bond in which an electron pair shared by two atoms is held more closely by one of the atoms. Genetically encoded polar amino acids include L-Asn (N), L-Gln (Q), L-Ser (S), and L-Thr (T).
[0143] "pH-stable" refers to a ketoreductase polypeptide that maintains similar activity (e.g., greater than 60% to 80%) after exposure to high or low pH (e.g., 4.5 to 6 or 8 to 12) for a period of time (e.g., 0.5 to 24 hours) compared to the unmodified enzyme.
[0144] A "promoter" is a nucleic acid sequence recognized by a host cell for expression of a coding region. Control sequences can include an appropriate promoter. A promoter contains transcriptional control sequences that mediate expression of a polypeptide. Promoters can be any nucleic acid sequence that exhibits transcriptional activity in a selected host cell, including mutant promoters, truncated promoters, and hybrid promoters, and can be derived from genes encoding extracellular or intracellular polypeptides that are homologous or heterologous to the host cell.
[0145] The terms "recombinant" and "modified" are used interchangeably herein, e.g., when used with respect to a cell, nucleic acid, or polypeptide, to refer to a material that does not otherwise occur in nature, or that is identical to it, but that has been modified in a way that is produced or derived from synthetic material and / or by manipulation using recombinant technology, or a material that corresponds to the native or inherent form of that material. Non-limiting examples include, among others, recombinant or modified ketoreductase polypeptides or recombinant cells that express genes that express native genes that are not found within the native (non-recombinant) form of the cell, or that are expressed at different levels.
[0146] "Reference" means standard or control conditions.
[0147] A "reference sequence" is a defined sequence, such as a polynucleotide or polypeptide sequence, used as a basis for sequence comparison. A reference sequence can be a subset or the entirety of a particular sequence (e.g., a segment of a full-length gene or polypeptide sequence). Generally, a reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 nucleotides or amino acid residues in length, at least 50 nucleotides or amino acid residues in length, or the full-length nucleic acid or polypeptide, or any integer in between. Because two polynucleotides or polypeptides can each contain (1) similar sequences (i.e., portions of the complete sequence) between the two sequences and (2) additional sequences that differ between the two sequences, sequence comparison between two (or more) polynucleotides or polypeptides is typically performed by comparing the sequences of the two polynucleotides or polypeptides over a "comparison window" to identify and compare local regions of sequence similarity. In some embodiments, the reference sequence is the same as the parent sequence used to generate the engineered ketoreductase polynucleotide or polypeptide. In some embodiments, the reference sequence is a wild-type (e.g., L. kefir) ketoreductase polynucleotide or polypeptide. In some embodiments, the reference sequence is a modified (e.g., SEQ ID NO: 256) ketoreductase polynucleotide or polypeptide.
[0148] As used herein, the terms "salt," "salts," or "salt form" refer to an acid addition salt or a base addition salt of the respective compound, e.g., a compound identified herein (e.g., Compound (I), or, e.g., an additional pharmaceutically active ingredient defined herein). "Salt" specifically includes "pharmaceutically acceptable salts." The term "pharmaceutically acceptable salts" refers to salts that retain the biological effectiveness and properties of the compound and are typically not biologically or otherwise undesirable. The compounds specified herein (e.g., Compound (I), or, e.g., an additional pharmaceutically active ingredient defined herein) may be capable of forming acid salts and / or base salts by virtue of the presence of amino and / or carboxyl groups, or groups similar thereto. The compounds of the present disclosure can form acid addition salts, and therefore, as used herein, the term "pharmaceutically acceptable salts of Compound (I)" refers to pharmaceutically acceptable acid addition salts of Compound (I).
[0149] Pharmaceutically acceptable acid addition salts can be formed with inorganic and organic acids.
[0150] Inorganic acids from which salts can be derived include, for example, hydrochloric acid, hydrobromic acid, sulfuric acid, nitric acid, phosphoric acid, and the like.
[0151] Examples of organic acids from which salts can be derived include acetic acid, propionic acid, glycolic acid, oxalic acid, maleic acid, malonic acid, succinic acid, fumaric acid, tartaric acid, citric acid, benzoic acid, mandelic acid, methanesulfonic acid, ethanesulfonic acid, toluenesulfonic acid, and sulfosalicylic acid.
[0152] Pharmaceutically acceptable base addition salts can be formed with inorganic and organic bases.
[0153] Inorganic bases from which salts can be derived include, for example, ammonium salts and metals from columns I to XII of the periodic table. In certain embodiments, salts are derived from sodium, potassium, ammonium, calcium, magnesium, iron, silver, zinc, and copper, with particularly suitable salts including ammonium, potassium, sodium, calcium, and magnesium salts.
[0154] Organic bases from which salts can be derived include, for example, primary, secondary, and tertiary amines, substituted amines including naturally occurring substituted amines, cyclic amines, basic ion exchange resins, etc. Particular organic amines include isopropylamine, benzathine, cholinate, diethanolamine, diethylamine, lysine, meglumine, piperazine, and tromethamine.
[0155] Pharmaceutically acceptable salts can be synthesized from basic or acidic moieties by conventional chemical methods. Generally, such salts can be prepared by reacting the free acid form of the compound with a stoichiometric amount of an appropriate base (e.g., hydroxide, carbonate, bicarbonate, etc. of Na, Ca, Mg, or K), or by reacting the free base form of the compound with a stoichiometric amount of an appropriate acid. Such reactions are usually carried out in water or an organic solvent, or a mixture of the two. Generally, the use of non-aqueous media such as ether, ethyl acetate, ethanol, isopropanol, or acetonitrile is desirable, where practicable. Additional lists of suitable salts can be found, for example, in "Remington's Pharmaceutical Sciences," 22 nd edition, Mack Publishing Company (2013); and “Handbook of Pharmaceutical Salts: Properties, Selection, and Use” by Stahl and Wermuth (Wiley-VCH, Weinheim, 2011, 2 nd edition).
[0156] "Small amino acid or residue" refers to an amino acid or residue having a side chain consisting of a total of three or fewer carbon and / or heteroatoms (excluding a carbon and hydrogen). Small amino acids or residues may be further classified as aliphatic, nonpolar, polar, or acidic according to the above definitions. Genetically encoded small amino acids include L-Ala (A), L-Val (V), L-Cys (C), L-Asn (N), L-Ser (S), L-Thr (T), and L-Asp (D).
[0157] The small amino acid L-Cys(C) is unique in that it can form disulfide bridges with other L-Cys(C) amino acids or other sulfanyl- or sulfhydryl-containing amino acids. Cysteine-like residues include cysteine and other amino acids containing sulfhydryl moieties available for disulfide bridge formation. The ability of L-Cys(C) (and other amino acids with SH-containing side chains) to exist in a peptide in their reduced, free SH, or oxidized, disulfide-bridged form influences whether L-Cys(C) contributes net hydrophobicity or net hydrophilicity to the peptide. While L-Cys(C) exhibits a hydrophobicity of 0.29 according to Eisenberg's normalized consensus scale (Eisenberg et al., 1984, supra), for purposes of this disclosure, it should be understood that L-Cys(C) is classified in its own unique group.
[0158] "Sequence identity," "percent identity," and "homology" are used interchangeably herein and refer to the similarity between amino acid or nucleic acid sequences, as determined by comparing two optimally aligned sequences over a comparison window. The portion of the polynucleotide or polypeptide sequence in the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions) due to optimal alignment of the two sequences. Sequence identity is often measured as a percentage of identity (or similarity or homology); the higher the percentage, the more similar the sequences. Homologs or variants of a given gene or protein will have a relatively high sequence identity when aligned using standard methods.
[0159] The percentage can be calculated by determining the number of positions in both sequences where an identical amino acid residue occurs to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage sequence identity. The percentage can be calculated by determining the number of positions in both sequences where an identical nucleobase or amino acid residue occurs or where the nucleobase or amino acid residue is aligned with a gap to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage sequence identity.
[0160] Those skilled in the art will appreciate that there are many established algorithms available for aligning two sequences. Sequence identity is typically measured using sequence analysis software (e.g., Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, FASTA, TFASTA, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. An exemplary approach for determining the degree of identity may use the BLAST program, e.g., -3 and e -100 A probability score between indicates a closely related sequence.
[0161] In addition, other programs and alignment algorithms are described, for example, in Smith and Waterman, 1981, Adv. Appl. Math. 2:482; Needleman and Wunsch, 1970, J. Mol. Biol. 48:443; Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444; Higgins and Sharp, 1988, Gene 73:237-244; Higgins and Sharp, 1989, CABIOS 5:151-153; Corpet et al., 1988, Nucleic Acids Research 16:10881-10890; Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444; and Altschul et al., 1994, Nature Genet.6:119-129.
[0162] The NCBI Basic Local Alignment Search Tool (BLAST (trademark)) (Altschul et al. 1990, J. Mol. Biol. 215:403-410) is readily available from various sources such as the National Center for Biotechnology Information (NCBI, Bethesda, Md.) and the Internet for use in connection with the sequence analysis programs blastp, blastn, blastx, tblastn, and tblastx. This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short word lengths W in the query sequence that match or satisfy a positive-valued threshold score T when aligned with words of the same length within the database sequences. This is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits serve as seeds to initiate a search for longer HSPs that contain them. Next, the word hits are extended in either direction along each sequence as long as the cumulative alignment score can increase. The cumulative score is calculated using the parameters M (reward score for pairs of matching residues; always >0) and N (penalty score for mismatched residues; always <0) for nucleotide sequences. For amino acid sequences, a scoring matrix is used to calculate the cumulative score. The extension of the word hits in each direction is stopped when the cumulative alignment score drops by an amount X from its achieved maximum value; when the cumulative score drops below zero due to the accumulation of one or more negative-score residues in the alignment; or when the end of either sequence is reached. The sensitivity and speed of the alignment are determined by the BLAST algorithm parameters W, T, and X. In the BLASTN program (for nucleotide sequences), a word length (W) of 11, an expectation value (E) of 10, M = 5, and N = -4 are used as defaults, and comparisons of both strands are performed.For amino acid sequences, the BLASTP program uses as default a word length of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, 1989, Proc Natl Acad Sci USA 89:10915). Exemplary sequence alignments and determination of percent sequence identity can use the BESTFIT or GAP programs using the default parameters provided in the GCG Wisconsin software package (Accelrys, Madison WI).
[0163] "Substantially identical" or "substantial identity" refers to a polypeptide or polynucleotide that exhibits at least 50% identity to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein). In some embodiments, such a sequence is at least about 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical compared to the reference sequence over a comparison window of at least 20 residues or more. Typically, the percent sequence identity is calculated by comparing the reference sequence to a sequence that contains deletions or additions totaling no more than 20% of the reference sequence over the comparison window. In certain embodiments, the term "substantial identity" as applied to a polypeptide means sharing at least about 80% sequence identity, at least about 85% sequence identity, at least about 90% sequence identity, or at least about 95% or more sequence identity (e.g., 99% sequence identity) when two polypeptide sequences are optimally aligned, such as by the programs GAP or BESTFIT using default gap weights. Polynucleotides useful in the methods of the invention include any nucleic acid molecule that encodes a polypeptide of the invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit substantial identity.
[0164] "Stereoisomer" refers to an isomer that has the same molecular formula but differs in the three-dimensional arrangement of atoms in space. Two types of stereoisomers include "enantiomers" and "diastereomers." "Enantiomers" refer to a pair of stereoisomeric compounds that have the same molecular formula but are non-superimposable mirror images of each other and typically have the same physical properties. "Diastereomers" refer to a pair of stereoisomeric compounds that have the same molecular formula but are non-superimposable, non-mirror images of each other and typically have different physical properties. Diastereomers are not considered to be enantiomers. As provided herein, the (trans) alcohol product tert-butyl rel-(3aR,5r,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate ((5r)-2) and the (cis) alcohol product tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (5s)2) are stereoisomers of each other. As shown in Figure 1, the structures of (5r)-2 and (5s)-2 are non-superimposable mirror images and are therefore also referred to herein as diastereoisomers.
[0165] "Stereoselectivity" refers to the preferential formation of one stereoisomer over another in a chemical or enzymatic reaction. Stereoselectivity can be partial, where the formation of one stereoisomer is preferred over the other, or complete, where only one stereoisomer is formed. When the stereoisomers are enantiomers, the stereoselectivity is referred to as "enantioselectivity," and is the ratio (usually reported as a percentage) of one enantiomer to the sum of both. This is commonly alternatively reported in the art (typically as a percentage) as enantiomeric excess (ee), calculated by the formula [major enantiomer - minor enantiomer] / [major enantiomer + minor enantiomer]. When the stereoisomers are diastereoisomers, the stereoselectivity is referred to as "diastereoselectivity," which is the proportion (usually reported as a percentage) of one diastereomer in a mixture of two diastereomers, and is commonly alternatively reported as diastereomeric excess (de), calculated by the formula [major diastereomer - minor diastereomer] / [major diastereomer + minor diastereomer]. Enantiomeric excess and diastereomeric excess are types of stereomeric excess. "Highly stereoselective" refers to a ketoreductase polypeptide capable of converting or reducing a substrate to the corresponding (cis) alcohol product with a diastereomeric excess of at least about 99%.
[0166] "Stereospecificity" refers to the preferential conversion of one stereoisomer over another in a chemical or enzymatic reaction. Stereospecificity can be partial, where one stereoisomer is converted in preference to the other, or complete, where only one stereoisomer is converted. If the stereoisomers are enantiomers, the stereospecificity is called enantiospecific. If the stereoisomers are diastereomers, the stereospecificity is called diastereospecific.
[0167] "Suitable reaction conditions" or "conditions suitable for reducing or converting a substrate to a product compound" refer to the conditions in a biocatalytic reaction system (e.g., enzyme load, substrate load, cofactor load, temperature, pH, buffer, cosolvent, etc.) under which a KRED polypeptide of the disclosure can convert a substrate to a desired product compound. Examples of "suitable reaction conditions" are provided in the disclosure and illustrated by the Examples.
[0168] "Reference to," "relative to," "compared to," or "corresponding to," when used in the context of numbering a given amino acid or polynucleotide sequence, refers to the numbering of residues in a particular reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, residue numbers or residue positions in a given polymer are specified with respect to the reference sequence, not by the actual numerical position of the residue within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence, such as a modified KRED, can be aligned with a reference sequence by introducing gaps to optimize residue matches between the two sequences. In such cases, despite the gaps, the numbering of residues in a given amino acid or polynucleotide sequence is done with respect to the aligned reference sequence.
[0169] Ranges provided herein are understood to be shorthand for all values within that range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or subrange from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0170] Unless otherwise stated or clear from context, as used herein, the term "or" is understood to be inclusive. Unless otherwise stated or clear from context, as used herein, the terms "a," "an," and "the" are understood to be singular or plural.
[0171] Unless otherwise specified or clear from the context, the term "about" as used herein is understood to mean within normal tolerances in the art, for example, within two standard deviations of the mean. About can be understood as within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless otherwise clear from the context, all numerical values provided herein are modified by the term about.
[0172] The recitation of a list of chemical groups in a definition of a variable herein includes defining that variable as a single group or combination of listed groups. The recitation of an embodiment for a variable or aspect herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.
[0173] Any composition or method provided herein can be combined with one or more of any of the other compositions and methods provided herein.
[0174] Ketoreductase enzymes The present disclosure provides engineered ketoreductase ("KRED") enzymes that can stereoselectively reduce defined bicyclic keto substrates to their corresponding alcohol products and have improved properties compared to naturally occurring wild-type KRED enzymes (e.g., wild-type L. kefir KRED enzyme) or engineered KRED variants thereof (e.g., SEQ ID NOs: 54, 152, or 256). Naturally occurring wild-type KRED enzymes (e.g., wild-type L. kefir ketoreductase) preferentially reduce compounds on one face of the keto group. Specifically, bicyclic ketones, such as tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (C 12 H 19 When reducing NO (MW: 225.29; referred to herein as Compound 1 or the substrate), wild-type KRED enzymes, such as the wild-type KRED enzyme from L. kefir, react with tert-butyl rel-(3aR,5r,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (C 12 H 21 It shows strong specificity for the formation of NO3 (MW: 227.30; referred to herein as the (5r)-2 or (trans) alcohol product).
[0175] However, the present disclosure provides a method for the synthesis of compound 1 with very high selectivity and specificity to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (C 12 H 21 Modified KRED enzyme polypeptides capable of reducing NO to NO (MW: 227.30; referred to herein as the (5s)-2 or (cis) alcohol product) are provided (see FIG. 1). The disclosure further provides polynucleotides encoding such modified KRED polypeptides, methods of using and producing the modified KRED polypeptides, and kits thereof.
[0176] In some embodiments, the modified enzymes described herein have one or more improved properties. Improved enzyme properties include, among others, an increase or change in enzyme activity, cofactor binding, stereoselectivity, stereospecificity, thermal stability, solvent stability, or reduced product inhibition. In some embodiments, the improved enzyme property is a reversal of stereoselectivity (e.g., a reversal of enantioselectivity or a reversal of diastereoselectivity).
[0177] In some embodiments, the improved enzyme property is a reversal or increase in diastereoselectivity. In some embodiments, the improved enzyme property is an increase in enzyme activity (e.g., an increase in conversion). In some embodiments, the improved enzyme property is an increase in selectivity (e.g., an increase in the desired product or diastereomeric excess). In some embodiments, the improved enzyme property is the ability to use fewer cofactors in the reduction reaction. In some embodiments, the improved enzyme property is the ability to not require glucose dehydrogenase (GDH) / glucose cofactor recycling in the reduction reaction. In some embodiments, the improved enzyme property is the ability to not require dimethyl sulfoxide (DMSO) in the reduction reaction.
[0178] Generally, the engineered KRED polypeptides have improved properties (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5r)-2 over (5r)-2) or activity (% conversion)) compared to a naturally occurring wild-type ketoreductase enzyme (e.g., obtained from Lactobacillus kefir ("L. kefir"; SEQ ID NO: 492), Lactobacillus brevis ("L. brevis"; SEQ ID NO: 493), or Lactobacillus minor ("L. minor"; SEQ ID NO: 494)). In some embodiments, the engineered KREDS enzyme has improved properties (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5r)-2 over (5r)-2) or activity (% conversion)) compared to a wild-type L. kefir ketoreductase (e.g., SEQ ID NO: 492). For example, in some embodiments, the engineered KRED polypeptides described herein increase the enzymatic activity for reducing a substrate (e.g., compound 1) to a product (e.g., (5s)-2) compared to a wild-type KRED enzyme (e.g., L. kefir), and / or further reverse or increase the diastereoselectivity of the (cis) alcohol product diastereomer.
[0179] In some embodiments, the engineered KRED polypeptides of the disclosure have one or more improved properties (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) compared to an engineered KRED polypeptide (e.g., SEQ ID NO: 54, 152, or 256) obtained or derived from a naturally occurring ketoreductase enzyme (e.g., L. kefir; SEQ ID NO: 492). In some embodiments, the engineered KRED polypeptides of the disclosure have improved properties (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5r)2 over (5s)-2) or activity (% conversion)) compared to a commercially available KRED enzyme (e.g., ADH-152). In some embodiments, engineered KRED polypeptides of the disclosure have improved properties (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5r)-2 over (5s)-2) or activity (% conversion)) compared to a modified KRED polypeptide selected from Table 4, Table 5, or Table 6. In some embodiments, modified KRED polypeptides of the disclosure have improved properties (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5r)-2 over (5s)-2) or activity (% conversion)) compared to a modified KRED polypeptide of SEQ ID NO: 54, 152, or 256.
[0180] In some embodiments, the modified KRED polypeptides of the invention have improved properties compared to a reference (e.g., parent) polypeptide (e.g., SEQ ID NO: 54, 152, 256, or 492). In some embodiments, the reference polypeptide is a commercially available KRED enzyme (e.g., ADH-152). In some embodiments, the reference (e.g., parent) polypeptide is a wild-type KRED polypeptide. In some embodiments, the reference (e.g., parent) polypeptide is a modified KRED variant (e.g., SEQ ID NO: 54, SEQ ID NO: 152, SEQ ID NO: 256). In some embodiments, the reference (e.g., parent) polypeptide is SEQ ID NO: 54, SEQ ID NO: 152, SEQ ID NO: 256, or SEQ ID NO: 492.
[0181] In some embodiments, the modified KRED polypeptides of the invention are improved by an increased level or rate of enzymatic activity compared to a reference (e.g., parent) polypeptide (e.g., SEQ ID NO: 54, 152, 256, or 492). In some embodiments, activity levels are measured in terms of their percent improvement (e.g., % conversion) relative to a positive control (FIOP). In some embodiments, the modified polypeptides are capable of an FIOP of greater than about 1.75, greater than about 2.00, greater than about 2.25, greater than about 2.50, greater than about 2.75, greater than about 3.00, greater than about 3.25, greater than about 3.50, greater than about 3.75, greater than about 4.00, greater than about 4.25, greater than about 4.50, greater than about 4.75, greater than about 5.00, greater than about 5.25, greater than about 5.50, greater than about 5.75, greater than about 6.00, greater than about 6.25, or greater than about 6.50 compared to the reference (e.g., parent) polypeptide as a positive control.
[0182] In some embodiments, the modified KRED polypeptides of the invention are improved with respect to their stereoselectivity (e.g., selectivity for (5s)-2 over (5r)-2) compared to a reference (e.g., parent) polypeptide (e.g., SEQ ID NO: 54, 152, 256, or 492). In particular, the modified KRED polypeptides of the invention are improved with respect to their diastereoselectivity (e.g., selectivity for (5s)-2 over (5r)-2) compared to a reference (e.g., parent) polypeptide (e.g., SEQ ID NO: 54, 152, 256, or 492). In some embodiments, the modified KRED polypeptides of the invention are improved by an increased level of selectivity with respect to their % Improvement in Production of the desired product (e.g., (5s)-2) over a positive control (FIOP) (e.g., FIOP). In some embodiments, the modified KRED polypeptide is capable of an FIOP (e.g., % of desired product (e.g., (5s)-2)) of greater than about 1.10, greater than about 1.25, greater than about 1.50, greater than about 1.75, greater than about 2.00, greater than about 2.25, or greater than about 2.50 compared to a reference (e.g., parent) polypeptide as a positive control.
[0183] In some embodiments, the engineered KRED polypeptides described herein have an amino acid sequence with one or more amino acid differences compared to a reference (e.g., parent) amino acid sequence (e.g., a wild-type KRED (e.g., L. kefir) or an engineered KRED (e.g., SEQ ID NO: 54)) that results in improved properties of the enzyme (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) for a defined keto substrate. In some embodiments, the engineered KRED polypeptides described herein have an amino acid sequence with one or more amino acid differences compared to a wild-type KRED (e.g., L. kefir) that results in improved properties of the enzyme (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) for a defined keto substrate. In some embodiments, the modified KRED polypeptides described herein have an amino acid sequence with one or more amino acid differences compared to a modified KRED (e.g., SEQ ID NO: 54, 152, or 256) that results in improved properties of the enzyme for a defined keto substrate (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)).
[0184] In some embodiments, the modified KRED polypeptides described herein may contain one or more of the observed mutation characteristics shown in Table 2. In some embodiments, the modified KRED polypeptides contain one or more amino acid differences shown in Table 4, Table 5, Table 6, and Table 7. In some embodiments, the modified KRED polypeptides contain one or more amino acid differences shown in Table 2 and one or more amino acid differences shown in Table 4, Table 5, Table 6, and Table 7.
[0185] [Table 2]
[0186] [Table 3]
[0187] In certain embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are based on mutating one or more of the following residues relative to a reference (e.g., parent) amino acid sequence (e.g., a wild-type L. kefir ketoreductase or an engineered variant thereof): X17, X18, X40, X56, X94, X96, X106, X145, X173, X190, X194, X196, X198, X202, and / or X206. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are based on mutating one or more of the following residues relative to a reference (e.g., parent) amino acid sequence (e.g., a wild-type L. Kefir ketoreductase or an engineered variant thereof): X94, X96, X190, X196, X202, and / or X206. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are based on mutation of one or more of the following residues relative to a reference (e.g., parent) amino acid sequence (e.g., a wild-type L. Kefir ketoreductase or an engineered variant thereof): (i) mutation of X190 to a residue that is not tyrosine; (ii) mutation of X190 to a non-aromatic residue; (iii) mutation of X196 to an aliphatic or small amino acid residue; (iv) mutation of X202 to an amino acid residue; and / or (v) mutation of X206 to a residue that is not methionine.
[0188] In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are based on one or more of the following mutations relative to a reference (e.g., parent) amino acid sequence (e.g., a wild-type L. kefir ketoreductase or an engineered variant thereof): X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X40R, X56V, X56S, X56T, X56D, X94A, X94B, X94C, X94D, X94E, X94F, X94G, X94H ... 94V, X94L, X94I, X94W, X96Y, X106I, X106V, X106L, X106W, X106E, X145S, X145D, X145E, X173A, X173L, X173V, X173W, X173M, X173K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are based on mutating one or more of the following residues relative to a reference (e.g., parent) amino acid sequence (e.g., a wild-type L. kefir ketoreductase or an engineered variant thereof): i) X17S, X17 T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P.In some embodiments, the improved properties (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are based on one or more of the following mutations relative to a reference (e.g., parent) amino acid sequence (e.g., a wild-type L. kefir ketoreductase or an engineered variant thereof): X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W. In some embodiments, the improved properties (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are based on one or more of the following mutations relative to a reference (e.g., parent) amino acid sequence (e.g., a wild-type L. kefir ketoreductase or an engineered variant thereof): X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q.
[0189] In some embodiments, engineered KRED polypeptides of the disclosure having improved properties (e.g., a reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are derived from a wild-type ketoreductase (e.g., L. kefir, L. brevis, or L. minor). In some embodiments, engineered KRED polypeptides of the disclosure have an amino acid sequence having one or more amino acid differences compared to the wild-type L. kefir KRED polypeptide (SEQ ID NO: 492). In some embodiments, engineered KRED polypeptides of the disclosure have an amino acid sequence having one or more amino acid differences as set forth in Table 2, Table 4, Table 5, Table 6, and Table 7 compared to an amino acid sequence having at least about 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% similarity to SEQ ID NO: 492. In some embodiments, a modified KRED polypeptide of the present disclosure has an amino acid sequence having one or more amino acid differences as shown in Table 2, Table 4, Table 5, Table 6, and Table 7 compared to the amino acid sequence of SEQ ID NO:492.
[0190] In some embodiments, modified KRED polypeptides of the disclosure have improved properties (e.g., reversed diastereoselectivity) compared to the wild-type KRED polypeptide of SEQ ID NO:492. In certain embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 492: X17, X18, X40, X56, X94, X96, X106, X145, X173, X190, X194, X196, X198, X202, and / or X206. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 492: X94, X96, X190, X196, X202, and / or X206.In some embodiments, the improved properties of the modified KRED polypeptides of the disclosure (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are based on mutation of one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 492: (i) mutation of X190 to a residue that is not tyrosine; (ii) mutation of X190 to a non-aromatic residue; (iii) mutation of X196 to an aliphatic or small amino acid residue; (iv) mutation of X202 to an amino acid residue; and / or (v) mutation of X206 to a residue that is not methionine. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 492: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H ... 0R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X1 06L, X106W, X106E, X145S, X145D, X145E, X173A, X173L, X173V, X173W, X173M, X17 3K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W.In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are achieved by comparing the following with an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO:492: based on one or more mutations: i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 492: X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 492: X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q.
[0191] In some embodiments, engineered KRED polypeptides of the disclosure having improved properties (e.g., diastereoselectivity (i.e., selectivity for (5r)-2 over (5s)-2) or reversed or increased activity (conversion &)) are derived from engineered variants of wild-type (e.g., L. kefir) ketoreductases (e.g., SEQ ID NOs: 54, 152, 256). In some embodiments, polypeptides of the disclosure have an amino acid sequence with one or more amino acid differences compared to an engineered variant of a wild-type L. kefir ketoreductase (e.g., SEQ ID NOs: 54, 152, 256). In some embodiments, modified KRED polypeptides of the disclosure have an amino acid sequence that has one or more amino acid differences compared to an engineered variant that comprises an amino acid sequence having at least about 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to an amino acid sequence of Table 4, Table 5, Table 6, or Table 7. In some embodiments, modified KRED polypeptides of the disclosure have an amino acid sequence that has one or more amino acid differences compared to an engineered variant that comprises an amino acid sequence of Table 4, Table 5, Table 6, or Table 7. In some embodiments, modified KRED polypeptides of the disclosure have an amino acid sequence that has one or more amino acid differences compared to an engineered variant that comprises an amino acid sequence having at least about 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256. In some embodiments, a modified KRED polypeptide of the present disclosure has an amino acid sequence that has one or more amino acid differences compared to a modified variant comprising the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256.
[0192] In some embodiments, the modified KRED polypeptide has an amino acid sequence that has one or more amino acid differences selected from Table 2, Table 4, Table 5, Table 6, and Table 7 compared to the modified KRED polypeptide (e.g., SEQ ID NO:2, SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256). In some embodiments, the modified KRED polypeptide has an amino acid sequence that has one or more amino acid differences selected from Table 2 and one or more amino acid differences selected from Table 4, Table 5, Table 6, and Table 7 compared to the modified KRED polypeptide (e.g., SEQ ID NO:2, SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256). In some embodiments, the modified KRED polypeptides of the present disclosure contain one or more amino acid differences selected from Table 2, Table 4, Table 5, Table 6, and Table 7 compared to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256. In some embodiments, modified KRED polypeptides of the disclosure comprise one or more amino acid differences selected from Table 2 and one or more amino acid differences selected from Table 4, Table 5, Table 6, or Table 7 compared to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256. In some embodiments, modified KRED polypeptides of the disclosure comprise one or more amino acid differences selected from Table 2 and one or more amino acid differences selected from Table 4, Table 5, Table 6, or Table 7 compared to SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256. In some embodiments, modified KRED polypeptides of the disclosure comprise one or more amino acid differences selected from Table 2 and one or more amino acid differences selected from Table 4, Table 5, Table 6, or Table 7 compared to SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256.
[0193] In some embodiments, modified KRED polypeptides of the disclosure have improved properties compared to the modified KRED polypeptide of SEQ ID NO: 54. In some embodiments, modified KRED polypeptides of the disclosure can selectively reduce tert-butyl rel-(3aR,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product) with greater stereoselectivity (% diastereomeric excess) and / or activity (% conversion) than SEQ ID NO: 54 under suitable reaction conditions.
[0194] In certain embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 54: X17, X18, X40, X56, X94, X96, X106, X145, X173, X190, X194, X196, X198, X202, and / or X206. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 54: X94, X96, X190, X196, X202, and / or X206. In some embodiments, the improved properties of the modified KRED polypeptides of the present disclosure (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) are based on mutation of one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 54: (i) mutation of X190 to a residue that is not tyrosine; (ii) mutation of X190 to a non-aromatic residue; (iii) mutation of X196 to an aliphatic or small amino acid residue; (iv) mutation of X202 to an amino acid residue; and / or (v) mutation of X206 to a residue that is not methionine.In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 54: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H ... 0R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X1 06L, X106W, X106E, X145S, X145D, X145E, X173A, X173L, X173V, X173W, X173M, X17 3K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are achieved by comparing the following with an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 54: based on one or more mutations: i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P.In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 54: X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 54: X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q.
[0195] In some embodiments, modified KRED polypeptides of the disclosure have improved properties compared to the modified KRED polypeptide of SEQ ID NO: 152. In some embodiments, modified KRED polypeptides of the disclosure can selectively reduce tert-butyl rel-(3aR,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product) with greater stereoselectivity (% diastereomeric excess) and / or activity (% conversion) than SEQ ID NO: 152 under suitable reaction conditions.
[0196] In certain embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical to SEQ ID NO: 152: X17, X18, X40, X56, X94, X96, X106, X145, X173, X190, X194, X196, X198, X202, and / or X206. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 152: X94, X96, X190, X196, X202, and / or X206. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the present disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 152: (i) mutation of X190 to a residue that is not tyrosine; (ii) mutation of X190 to a non-aromatic residue; (iii) mutation of X196 to an aliphatic or small amino acid residue; (iv) mutation of X202 to an amino acid residue; and / or (v) mutation of X206 to a residue that is not methionine.In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 152: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X4 0R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X1 06L, X106W, X106E, X145S, X145D, X145E, X173A, X173L, X173V, X173W, X173M, X17 3K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are achieved by comparing the following with an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 152: based on one or more mutations: i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P.In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 152: X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 152: X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q.
[0197] In some embodiments, modified KRED polypeptides of the disclosure have improved properties compared to the modified KRED polypeptide of SEQ ID NO: 256. In some embodiments, modified KRED polypeptides of the disclosure can selectively reduce tert-butyl rel-(3aR,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product) with greater stereoselectivity (% diastereomeric excess) and / or activity (% conversion) than SEQ ID NO: 256 under suitable reaction conditions.
[0198] In certain embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 256: X17, X18, X40, X56, X94, X96, X106, X145, X173, X190, X194, X196, X198, X202, and / or X206. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 256: X94, X96, X190, X196, X202, and / or X206. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on mutating one or more of the following residues to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 256: (i) mutation of X190 to a residue that is not tyrosine; (ii) mutation of X190 to a non-aromatic residue; (iii) mutation of X196 to an aliphatic or small amino acid residue; (iv) mutation of X202 to an amino acid residue; and / or (v) mutation of X206 to a residue that is not methionine.In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 256: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H ... 0R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X1 06L, X106W, X106E, X145S, X145D, X145E, X173A, X173L, X173V, X173W, X173M, X17 3K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion) of the modified KRED polypeptides of the disclosure are achieved by comparing the following with an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 256: based on one or more mutations: i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P.In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 256: X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W. In some embodiments, the improved properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)) of the modified KRED polypeptides of the disclosure are based on one or more of the following mutations to an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to SEQ ID NO: 256: X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q.
[0199] Exemplary modified KRED polypeptides of the disclosure include, but are not limited to, modified KRED polypeptides comprising an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, or 100% identical, to an amino acid sequence listed in Table 4, Table 5, Table 6, or Table 7. In some embodiments, a modified KRED polypeptide of the disclosure comprises an amino acid sequence listed in Table 4, Table 5, Table 6, or Table 7. In some embodiments, a modified KRED polypeptide of the disclosure comprises an amino acid sequence listed in Table 7.
[0200] In some embodiments, the engineered ketoreductase polypeptides of the invention have an identity at least about 97.7%, 97.8%, 97.9%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9 ...9%, 98.1%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 98.9%, 98.9%, 98.1%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 98.9%, 98.9%, 98.1%, 98.1%, and / or 100% identical to the amino acid sequence of the present invention. In some embodiments, a modified KRED polypeptide of the disclosure comprises an amino acid sequence selected from the group consisting of: SEQ ID NOs: 392, 394, 396, 406, 408, 420, 422, 424, 426, 428, 430, 436, 438, 462, 464, 468, 470, 474, 476, 480, 482, and 490.
[0201] In some embodiments, the engineered KRED polypeptides may have one or more additional amino acid residue differences compared to a reference (e.g., parent) polypeptide (e.g., a wild-type L. Kefir ketoreductase (e.g., SEQ ID NO: 492) or an engineered variant thereof (e.g., SEQ ID NO: 54, SEQ ID NO: 152, SEQ ID NO: 256). These differences may be amino acid insertions, deletions, substitutions, or any combination of such changes. In some embodiments, the amino acid sequence differences may include non-conservative, conservative, and combinations of non-conservative and conservative amino acid substitutions. In some embodiments, the amino acid differences may include conservative substitutions as set forth in Table 1. The various amino acid residue locations at which such changes can occur are described herein.
[0202] In some embodiments, the modified KRED polypeptide is derived from a naturally occurring KRED containing one or more mutations corresponding to any of the mutations set forth herein. One of skill in the art would be able to identify any homologous proteins and the corresponding residues in their respective encoding nucleic acids by methods well known in the art, such as sequence alignment and determination of homologous residues. Thus, one of skill in the art would be able to generate mutations in any naturally occurring KRED (e.g., having homology to L. kefir) that correspond to any of the mutations described herein. For example, one of skill in the art would be able to generate mutations in L. brevis or L. minor that correspond to mutations in L. kefir (e.g., mutations set forth in Tables 2, 4, 5, 6, and 7).
[0203] Polynucleotides encoding engineered ketoreductase enzymes The present disclosure provides polynucleotides encoding the engineered KRED enzymes disclosed herein. The polynucleotides can be operably linked to a promoter or one or more heterologous regulatory sequences that control gene expression to create a recombinant polynucleotide capable of expressing the polypeptide. In some embodiments, the polypeptides utilize codons optimized for a particular desired expression system. An expression construct containing a heterologous polynucleotide encoding an engineered KRED polypeptide can be introduced into a suitable host cell to express the corresponding ketoreductase polypeptide.
[0204] Codons corresponding to various amino acids are well known in the art. Thus, the availability of a polypeptide sequence provides one of skill in the art with a description of all polynucleotides capable of encoding the subject polypeptide. The degeneracy of the genetic code, in which the same amino acid is coded for by alternative or synonymous codons, allows for the creation of a vast number of nucleic acids, all of which will encode the modified KRED enzymes disclosed herein. Thus, once a particular amino acid sequence is determined, one of skill in the art can generate any number of different polynucleotides by modifying one or more codons in a manner that does not alter the amino acid sequence of the protein. In this regard, the present disclosure specifically contemplates all possible variations of polynucleotides that can be made by selecting combinations based on possible codon choices, and all such variations are considered to be specifically disclosed for any polypeptide disclosed herein, including the amino acid sequences presented herein (e.g., Table 4, Table 5, Table 6, or Table 7).
[0205] In one embodiment, a polynucleotide of the disclosure encodes any of the modified KRED polypeptides disclosed herein that include amino acids with one or more amino acid differences compared to a reference (e.g., parent) amino acid sequence (e.g., a wild-type KRED (e.g., an L. kefir ketoreductase) or modified KRED amino acid sequence (e.g., SEQ ID NO: 54)). In some embodiments, a polynucleotide of the disclosure encodes any of the modified KRED polypeptides described herein that are derived from a wild-type ketoreductase (e.g., L. kefir, L. brevis, L. minor). In some embodiments, a polynucleotide of the disclosure encodes any of the modified KRED polypeptides described herein that include an amino acid sequence with one or more amino acid differences compared to a wild-type L. kefir KRED polypeptide.
[0206] In some embodiments, an engineered ketoreductase polypeptide of the invention comprises an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 54, 152, 256, or 492; and is selected from the group consisting of i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X 56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P.
[0207] In certain embodiments, polynucleotides of the present disclosure encode an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 54, 152, 256, or 492, and that contains, relative to the amino acid sequence, a substitution, deletion, addition, or insertion of one or more amino acid residues selected from Table 2, Table 4, Table 5, Table 6, or Table 7. In some embodiments, polynucleotides of the disclosure encode amino acids having one or more amino acid differences selected from X17, X18, X40, X56, X94, X96, X106, X145, X173, X190, X194, X196, X198, X202, and / or X206 relative to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492. In some embodiments, polynucleotides of the disclosure encode amino acids having one or more amino acid differences selected from X94, X96, X190, X196, X202, and / or X206 relative to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492.In some embodiments, polynucleotides of the disclosure encode amino acids that have one or more amino acid differences selected from: (i) X190 to a residue that is not tyrosine; (ii) X190 to a non-aromatic residue; (iii) X196 to an aliphatic or small amino acid residue; (iv) X202 to a small amino acid residue; and / or (v) X206 to a residue that is not methionine, relative to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492.
[0208] In some embodiments, the polynucleotides of the disclosure comprise any of the following amino acid sequences: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X40R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94R, X94S, X94T, X94R ... and encoding amino acids having one or more amino acid differences selected from W, X96Y, X106I, X106V, X106L, X106W, X106E, X145S, X145D, X145E, X173A, X173L, X173V, X173W, X173M, X173K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W. In some embodiments, polynucleotides of the disclosure encode amino acids having one or more amino acid differences selected from X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W.
[0209] In some embodiments, a polynucleotide of the disclosure encodes an amino acid sequence selected from Table 4, Table 5, Table 6, or Table 7. In some embodiments, a polynucleotide of the disclosure encodes an amino acid sequence listed in Table 7. In some embodiments, a polynucleotide of the disclosure has at least about 97.7%, 97.8%, 97.9%, 98.0%, 99.1%, 100.0%, 101.0%, 102.0%, 103.0%, 104.0%, 105.0%, 106.0%, 107.0%, 108.0%, 109.0%, 1109.0%, 1111%, 1120%, 1121%, 1130%, 1131%, 1140%, 1142%, 1143%, 1144%, 1150%, 1151%, 1152%, 1153%, 1154%, 1155%, 1160%, 1161%, 1162%, 1163%, 1164%, 1165%, 1166%, 1167%, 1168%, 1170%, 1171%, 1172%, 1173%, 1174%, 1175%, 1176%, 1177%, 1178%, 1179%, 1180%, 1181%, 1182%, 1183%, 1184%, 1185%, 1186%, 1187%, 1188%, 1189%, 1190%, 1191%, 1192%, 1193%, 1194%, 1195%, 1196%, 1197%, 1198%, 1199%, 1200%, 1 or 100% identical to any of the modified KRED polypeptides disclosed herein. In some embodiments, a polynucleotide of the disclosure encodes any of the modified KRED polypeptides disclosed herein comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 392, 394, 396, 406, 408, 420, 422, 424, 426, 428, 430, 436, 438, 462, 464, 468, 470, 474, 476, 480, 482, and 490.
[0210] In some embodiments, a KRED polynucleotide of this disclosure comprises an amino acid sequence listed in Table 4, Table 5, Table 6, or Table 7. In some embodiments, a KRED polynucleotide of this disclosure comprises a nucleic acid sequence listed in Table 7. In some embodiments, a KRED polynucleotide of this disclosure has an identity of at least about 97.7%, 97.8%, 97.9 ... In some embodiments, the KRED polynucleotide comprises a nucleic acid sequence selected from SEQ ID NOs: 391, 393, 395, 405, 407, 419, 421, 423, 425, 427, 429, 435, 437, 461, 463, 467, 469, 473, 475, 479, 481, or 489.
[0211] In various embodiments, codons are preferably selected to be compatible with the host cell in which the protein is produced. For example, codons preferred for bacterial usage are used to express genes in bacteria, codons preferred for yeast usage are used to express genes in yeast, and codons preferred for mammalian usage are used to express genes in mammalian cells. By way of example, a polynucleotide may be codon-optimized for expression in Escherichia coli ("E. coli") but encodes the otherwise naturally occurring L. kefir KRED.
[0212] In certain embodiments, because the native sequence will contain preferred codons, and preferred codon usage is not required for every amino acid residue, it is not necessary to replace every codon to optimize the codon usage of a KRED. As a result, a codon-optimized polynucleotide encoding a KRED enzyme may contain preferred codons at more than about 40%, 50%, 60%, 70%, 80%, or 90% of the codon positions in the full-length coding region.
[0213] In various embodiments, an isolated polynucleotide encoding a modified KRED polypeptide of the present disclosure can be manipulated in a variety of ways to provide for expression of the polypeptide. Depending on the expression vector, it may be desirable or necessary to manipulate the isolated polynucleotide prior to insertion into the vector. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art. Guidance is provided in Sambrook et al., 2001, Molecular Cloning: A Laboratory Manual, 3rd Ed., Cold Spring Harbor Laboratory Press; and Current Protocols in Molecular Biology, Ausubel, F. Eds., Greene Pub. Associates, 1998 (updated 2006).
[0214] In bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include those derived from the Escherichia coli (E. coli) lac operon, the Streptomyces coelicolor agarase gene (dagA), the Bacillus subtilis levansucrase gene (sacB), the Bacillus licheniformis alpha-amylase gene (amyL), the Bacillus stearothermophilus maltogenic amylase gene (amyM), the Bacillus amyloliquefaciens alpha-amylase gene (amyQ), the Bacillus licheniformis alpha-amylase gene (amyQ), the Bacillus Examples of promoters that can be used include promoters obtained from the Bacillus licheniformis penicillinase gene (penP), the Bacillus subtilis xylA and xylB genes, and prokaryotic beta-lactamase genes (Villa-Kamaroff et al., 1978, Proc. Natl. Acad. Sci. USA 75:3727-3731), and the tac promoter (DeBoer et al., 1983, Proc. Natl. Acad. Sci. USA 80:21-25).
[0215] In filamentous fungal host cells, suitable promoters for directing the transcription of the nucleic acid constructs of the disclosure include those encoding Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger acid-stable alpha-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans Examples of promoters that can be used include promoters obtained from the genes for Aspergillus nidulans acetamidase and Fusarium oxysporum trypsin-like protease (WO 96 / 00787), and the NA2-tpi promoter (a hybrid of the promoters from the genes for Aspergillus niger neutral alpha-amylase and Aspergillus oryzae triosephosphate isomerase), as well as mutant, truncated, and hybrid promoters thereof.
[0216] In yeast hosts, useful promoters can be those from Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GALI), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase. Other useful promoters for yeast host cells are described in Romanos et al., 1992, Yeast 8:423-488.
[0217] The control sequence may also be a suitable transcription terminator sequence, i.e., a sequence recognized by a host cell to terminate transcription. The terminator sequence is operably linked to the 3' end of the nucleic acid sequence encoding the polypeptide. Any terminator that functions in the selected host cell may be used in the present invention. For example, exemplary transcription terminators for filamentous fungal host cells can be obtained from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger alpha-glucosidase, and Fusarium oxysporum trypsin-like protease. Exemplary terminators for yeast host cells can be obtained from the genes for Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYCl), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are described in Romanos et al., 1992 (ibid.).
[0218] The control sequence may also be a suitable leader sequence, a nontranslated region of an mRNA that is important for translation by the host cell. The leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in the host cell of choice can be used. Exemplary leaders for filamentous fungal host cells are obtained from the genes for Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase. Suitable leaders for yeast host cells are obtained from the genes for Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae alpha-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).
[0219] The regulatory sequence may also be a polyadenylation sequence, i.e., a sequence operably linked to the 3' end of a nucleic acid sequence that, upon transcription, is recognized by a host cell as a signal for the addition of polyadenosine residues to the transcribed mRNA. Any polyadenylation sequence that functions in the selected host cell may be used in the present invention. Exemplary polyadenylation sequences for filamentous fungal host cells may be those from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger alpha-glucosidase. Useful polyadenylation sequences for yeast host cells are described in Guo and Sherman, 1995, Mol Cell Bio 15:5983-5990.
[0220] The control sequence may also be a signal peptide coding region that encodes an amino acid sequence linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the cell's secretory pathway. The 5' end of the coding sequence of a nucleic acid sequence may inherently contain a signal peptide coding region naturally linked in translation reading frame with the segment of the coding region that encodes the secreted polypeptide. Alternatively, the 5' end of the coding sequence may contain a signal peptide coding region that is foreign to the coding sequence. A foreign signal peptide coding region may be required when the coding sequence does not naturally contain a signal peptide coding region.
[0221] Alternatively, the foreign signal peptide coding region may simply replace the natural signal peptide coding region in order to enhance secretion of the polypeptide. However, any signal peptide coding region which directs the expressed polypeptide into the secretory pathway of a host cell of choice may be used in the present invention.
[0222] Useful signal peptide coding regions in bacterial host cells are those obtained from the genes for Bacillus NCIB 11837 maltogenic amylase, Bacillus stearothermophilus alpha-amylase, Bacillus licheniformis beta-lactamase, Bacillus stearothermophilus neutral protease (nprT, nprS, nprM), and Bacillus subtilis prsA. Additional signal peptides are described in Simonen and Palva, 1993, Microbiol Rev 57:109-137.
[0223] Signal peptide coding regions effective in filamentous fungal host cells can be signal peptide coding regions obtained from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase.
[0224] Useful signal peptides for yeast host cells can be those obtained from the genes for Saccharomyces cerevisiae alpha-factor and Saccharomyces cerevisiae invertase. Other useful signal peptide coding regions are described in Romanos et al., 1992 (supra).
[0225] The control sequence may also be a propeptide coding region, which encodes an amino acid sequence positioned at the amino terminus of a polypeptide. The resulting polypeptide is known as a proenzyme or propolypeptide (or, in some cases, a zymogen). Propolypeptides are generally inactive and can be converted to a mature, active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide. Propeptide coding regions can be obtained from the genes for Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Saccharomyces cerevisiae alpha-factor, Rhizomucor miehei aspartic proteinase, and Myceliophthora thermophila lactase (WO 95 / 33836).
[0226] When both the signal peptide and propeptide regions are present at the amino terminus of a polypeptide, the propeptide region is located adjacent to the amino terminus of the polypeptide and the signal peptide region is located adjacent to the amino terminus of the propeptide region.
[0227] It may also be desirable to add regulatory sequences that allow regulation of polypeptide expression relative to the growth of the host cell. Examples of regulatory sequences are those that turn gene expression on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory systems include, for example, the ADH2 system or the GALI system. In filamentous fungi, suitable regulatory sequences include the TAKA alpha-amylase promoter, the Aspergillus niger glucoamylase promoter, and the Aspergillus oryzae glucoamylase promoter.
[0228] Other examples of regulatory sequences are those that allow for gene amplification. In eukaryotic systems, these include the dihydrofolate reductase gene, which is amplified in the presence of methotrexate, and the metallothionein genes, which are amplified with heavy metals. In these cases, the nucleic acid sequence encoding a KRED polypeptide of the invention would be operably linked to the regulatory sequence.
[0229] Thus, in another embodiment, the present disclosure is also directed to recombinant expression vectors comprising a polynucleotide encoding a modified KRED polypeptide or variant thereof and, depending on the type of host into which they are to be introduced, one or more expression control regions, such as a promoter and terminator, an origin of replication, etc. The various nucleic acid and control sequences described above can be ligated together to create recombinant expression vectors that may contain one or more convenient restriction sites to allow for the insertion or substitution of a nucleic acid sequence encoding a polypeptide at such sites.
[0230] Alternatively, the nucleic acid sequences of the present disclosure can be expressed by inserting the nucleic acid sequence or a nucleic acid construct containing the sequence into an appropriate expression vector. In creating an expression vector, a coding sequence is placed in the vector so that it is operably linked to suitable control sequences for expression.
[0231] A recombinant expression vector may be any vector (e.g., a plasmid or virus) that can be conveniently used in recombinant DNA procedures and that can bring about the expression of a polynucleotide sequence. The vector is typically selected based on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector may be a linear or a closed circular plasmid.
[0232] The expression vector may be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity whose replication is independent of chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. The vector may contain any means for ensuring self-replication. Alternatively, the vector may be a vector that, when introduced into a host cell, integrates into the genome and replicates together with the chromosome into which it is integrated. Furthermore, a single vector or plasmid, or two or more vectors or plasmids, may be used that together contain all the DNA to be introduced into the genome of the host cell.
[0233] The expression vectors of the present invention preferably contain one or more selectable markers that allow for easy selection of transformed cells. A selectable marker is a gene the product of which confers biocide or viral resistance, resistance to heavy metals, prototrophy for auxotrophs, etc. Examples of bacterial selectable markers are the dal genes from Bacillus subtilis or Bacillus licheniformis, or markers that confer antibiotic resistance, such as ampicillin resistance, kanamycin resistance, chloramphenicol resistance (see Example 1), or tetracycline resistance. Suitable markers for yeast host cells are ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3.
[0234] Selectable markers for use in filamentous fungal host cells include, but are not limited to, amdS (acetamidase), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase), sC (sulfate adenyltransferase), and trpC (anthranilate synthase), and equivalents thereof. Embodiments for use in Aspergillus cells include the amdS and pyrG genes of Aspergillus nidulans or Aspergillus oryzae, and the bar gene of Streptomyces hygroscopicus.
[0235] The expression vectors of the present invention preferably contain elements that allow for integration of the vector into the genome of the host cell or for autonomous replication of the vector within the cell independent of the genome. For integration into the host cell genome, the vector may rely on the nucleic acid sequence encoding the polypeptide or any other element of the vector to integrate the vector into the genome by homologous or non-homologous recombination.
[0236] Alternatively, the expression vector may contain additional nucleic acid sequences to direct integration into the host cell genome by homologous recombination. The additional nucleic acid sequences enable the vector to integrate into the host cell genome at a precise chromosomal location. To increase the likelihood of integration at a precise location, the integration element preferably contains a sufficient number of nucleic acids (e.g., 100 to 10,000 base pairs, preferably 400 to 10,000 base pairs, most preferably 800 to 10,000 base pairs) that are highly homologous to the corresponding target sequence and increase the probability of homologous recombination. The integration element may be any sequence that is homologous to the target sequence in the host cell genome. Furthermore, the integration element may be a non-coding or coding nucleic acid sequence. Alternatively, the vector may be integrated into the host cell genome by non-homologous recombination.
[0237] For autonomous replication, the vector may further comprise an origin of replication that enables the vector to replicate autonomously in the intended host cell. Examples of bacterial origins of replication are the pISA ori, or the origins of replication of plasmids pBR322, pUC19, pACYC177 (which contain the pISA ori), or pACYC184, which are replicable in E. coli, and pUB110, pEI94, pTAI060, or pAMP1, which are replicable in Bacillus. Examples of origins of replication for use in yeast host cells are the 2 micron origin of replication, ARS1, ARS4, a combination of ARS1 and CEN3, and a combination of ARS4 and CEN6. The origin of replication may also have a mutation that makes it temperature-sensitive to function in the host cell (see, e.g., Ehrlich, 1978, Proc Natl Acad Sci. USA 75:1433).
[0238] More than one copy of a nucleic acid sequence of the invention may be inserted into a host cell to increase production of the gene product. Increasing the copy number of a nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene along with the nucleic acid sequence; cells containing an amplified copy of the selectable marker gene, and thereby additional copies of the nucleic acid sequence, can be selected by culturing the cells in the presence of an appropriate selection agent.
[0239] Many expression vectors for use in the present invention are commercially available. A suitable commercially available expression vector includes the p3xFLAG™ expression vector from Sigma Aldrich Chemicals, St. Louis, MO, which contains a CMV promoter and hGH polyadenylation site for expression in mammalian host cells, and the pBR322 origin of replication and ampicillin resistance marker for amplification in E. coli. Other suitable expression vectors are pBluescriptII SK(-) and pBK-CMV, available from Stratagene, LaJolla, CA, as well as plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen), or pPoly (Lathe et al., 1987, Gene 57:193-201).
[0240] Host cells and expression vectors for expression of ketoreductase polypeptides The present disclosure further provides host cells comprising any of the polynucleotides and / or expression vectors described herein. The host cells can be used for expression and isolation of the modified KRED enzymes described herein, or alternatively, can be used directly to convert a substrate (compound 1) to a product ((5s)-2).
[0241] In some embodiments, the host cell comprises a polynucleotide encoding a modified KRED polypeptide operably linked to one or more control sequences for expression of the KRED enzyme in the host cell. Host cells for use with KRED polypeptides encoded by expression vectors of the present disclosure are well known in the art and include, but are not limited to, bacterial cells such as E. coli, L. kefir, L. brevis, L. minor, Streptomyces, and Salmonella typhimurium; fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC Accession No. 201178)); insect cells such as Drosophila S2 and Spodoptera Sf9 cells; animal cells such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Suitable culture media and growth conditions for the above-described host cells are well known in the art.
[0242] Polynucleotides for KRED expression can be introduced into cells by various methods known in the art. Non-limiting techniques include electroporation, bioparticle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion. Various methods for introducing polynucleotides into cells will be apparent to those skilled in the art.
[0243] An exemplary host cell is E. coli W3110. This expression vector was generated by operably linking a polynucleotide encoding a modified KRED polypeptide to plasmid pCKl 10900 (the vector shown as Figure 3 in U.S. Patent Application Publication No. 2006 / 0195947, incorporated herein by reference). The expression vector also contained a P15a origin of replication and a chloramphenicol (CAM) resistance gene. Cells containing the polynucleotide of interest were isolated in E. coli W3110 by subjecting the cells to chloramphenicol selection.
[0244] Methods for Producing Engineered Ketoreductase Polypeptides The present disclosure provides methods for generating or producing any of the modified KRED polynucleotides and polypeptides disclosed herein. Generally, methods for generating or producing a modified KRED polynucleotide or polypeptide provided herein involve introducing one or more differences (e.g., substitutions, deletions, additions, or insertions) into the nucleic acid or amino acid sequence of the KRED polypeptide.
[0245] In some embodiments, a naturally occurring KRED enzyme that catalyzes a reduction reaction is obtained (or derived) for use as a parent polynucleotide sequence (e.g., from L. kefir, L. minor, or L. brevis) to make the engineered KRED polynucleotides and polypeptides of the disclosure. In some embodiments, the parent polynucleotide sequence is obtained (or derived) from a naturally occurring ketoreductase enzyme selected from L. kefir, L. minor, or L. brevis. In some embodiments, the parent polynucleotide sequence is obtained (or derived) from L. kefir (e.g., SEQ ID NO: 492). In some embodiments, an engineered variant of a naturally occurring KRED enzyme that catalyzes a reduction reaction is obtained (or derived) for use as a parent polynucleotide sequence (e.g., SEQ ID NO: 54, 152, or 256) to make the engineered KRED polynucleotides and polypeptides of the disclosure. In some embodiments, the parent polynucleotide sequences are obtained (or derived) from a modified KRED variant selected from Table 4, Table 5, or Table 6. In some embodiments, the parent polynucleotide sequences are obtained (or derived) from a modified KRED variant selected from SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256.
[0246] In some embodiments, parent polynucleotide sequences are codon-optimized to enhance expression of the KRED polypeptide in a particular host cell. As an example, a parent polynucleotide sequence encoding a wild-type KRED polypeptide in L. kefir was constructed from oligonucleotides prepared based on the known polypeptide sequence of the L. kefir KRED sequence (available under UniProKB accession number Q6WVP7). The parent polynucleotide sequence designated as SEQ ID NO: 492 was codon-optimized for expression in E. coli, and this codon-optimized polynucleotide was cloned into an expression vector, placing expression of the ketoreductase gene under the control of the lac promoter and lacI repressor gene. Clones expressing active ketoreductase in E. coli were identified, and the genes were sequenced to confirm their identity. The parent sequence was then utilized as the starting point for the majority of experiments and for the construction of a library of engineered KREDs (e.g., SEQ ID NO: 54, SEQ ID NO: 152, or SEQ ID NO: 256) evolved from the L. kefir ketoreductase.
[0247] The engineered KRED polypeptides herein can be obtained by subjecting a parent polynucleotide encoding a naturally occurring ketoreductase (e.g., L. kefir) or variant thereof (e.g., SEQ ID NOs: 54, 152, 256) to mutagenesis and / or directed evolution. Generally, the engineered KRED polypeptides generated or produced as described herein have an amino acid sequence that has one or more amino acid differences compared to the parent KRED, e.g., a wild-type KRED (e.g., L. kefir) or an engineered KRED variant (e.g., SEQ ID NOs: 54, 152, or 256).
[0248] In some embodiments, modified KRED polynucleotides of the disclosure can be obtained by subjecting a parent polynucleotide encoding an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 54, 152, 256, or 492 to mutagenesis to introduce into said amino acid sequence one or more substitutions, deletions, additions, or insertions of amino acid residues selected from Table 2, Table 4, Table 5, Table 6, or Table 7. In some embodiments, modified KRED polynucleotides of the disclosure can be obtained by subjecting a parent polynucleotide encoding an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 54, 152, 256, or 492 to mutagenesis to introduce into said amino acid sequence one or more amino acid residue substitutions, deletions, additions, or insertions of one or more amino acid residues selected from Table 2 and one or more amino acid residues selected from Table 4, Table 5, Table 6, or Table 7.
[0249] In some embodiments, modified KRED polypeptides of the disclosure can be obtained by introducing one or more amino acid differences selected from X17, X18, X40, X56, X94, X96, X106, X145, X173, X190, X194, X196, X198, X202, and / or X206 to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492. In some embodiments, modified KRED polypeptides of the disclosure can be obtained by introducing one or more amino acid differences selected from X94, X96, X190, X196, X202, and / or X206 to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492. In some embodiments, modified KRED polypeptides of the disclosure can be obtained by introducing, to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492, one or more amino acid differences selected from: (i) X190 to a residue that is not tyrosine; (ii) X190 to a non-aromatic residue; (iii) X196 to an aliphatic residue or a small amino acid residue; (iv) X202 to a small amino acid residue; and / or (v) X206 to a residue that is not methionine.
[0250] In some embodiments, the modified KRED polypeptides of the disclosure include any of the following amino acid sequences: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X40R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94R, X94S, X94T, X94R, X94T ... and X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W. In some embodiments, the modified KRED polypeptides of the disclosure are prepared by subjecting a parent polynucleotide encoding an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492 to mutagenesis, where the parent polynucleotide encodes an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492, and adding to said amino acid sequence: i) X17S, X17T, or can be obtained by introducing one or more amino acid differences selected from X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P.In some embodiments, modified KRED polypeptides of the disclosure can be obtained by introducing one or more amino acid differences selected from X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492. In some embodiments, modified KRED polypeptides of the disclosure can be obtained by introducing one or more amino acid differences selected from X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q to an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, or SEQ ID NO:492.
[0251] In some embodiments, the modified KRED polypeptide generated or produced by any of the methods provided herein has an amino acid sequence selected from the group consisting of SEQ ID NOs: 392, 394, 396, 406, 408, 420, 422, 424, 426, 428, 430, 436, 438, 462, 464, 468, 470, 474, 476, 480, 482, and 490. and having an amino acid sequence that is at least about 97.7%, 97.8%, 97.9%, 98.0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% identical to In some embodiments, the modified KRED polypeptide generated or produced by any of the methods provided herein has an amino acid sequence selected from the group consisting of SEQ ID NOs: 392, 394, 396, 406, 408, 420, 422, 424, 426, 428, 430, 436, 438, 462, 464, 468, 470, 474, 476, 480, 482, and 490.
[0252] In some embodiments, a modified KRED polypeptide generated or produced by any of the methods provided herein has an amino acid sequence selected from Table 4 (SEQ ID NOs: 4-52), Table 5 (SEQ ID NOs: 56-186), Table 6 (SEQ ID NOs: 188-390), or Table 7 (SEQ ID NOs: 392-490). In some embodiments, a modified KRED polypeptide produced by any of the methods provided herein has an amino acid sequence selected from Table 7 (SEQ ID NOs: 392-490).
[0253] As discussed herein, mutagenesis and / or directed evolution methods are well known in the art. Exemplary directed evolution techniques are mutagenesis and / or DNA shuffling, as described in Stemmer, 1994, Proc Natl Acad Sci USA 91:10747-10751; WO 95 / 22625; WO 97 / 0078; WO 97 / 35966; WO 98 / 27230, WO 00 / 42651; WO 01 / 75767, and U.S. Pat. No. 6,537,746. Other directed evolution procedures that can be used include stepwise extension (StEP), in vitro recombination (Zhao et al., 1998, Nat. Biotechnol. 16:258-261), mutagenic PCR (Caldwell et al., 1994, PCR Methods Appl. 3:S136-S140), and cassette mutagenesis (Black et al., 1996, Proc Natl Acad Sci USA 93:3525-3529), among others.
[0254] Clones obtained after the mutagenesis treatment are screened for modified KREDs with desired improved enzyme properties (e.g., reversal of diastereoselectivity, increased selectivity (i.e., (5s)-2 over (5r)-2), increased activity (% conversion), etc.). Measurement of the enzyme activity of the expression library is performed by measuring the conversion of NADH or NADPH to NAD. + or NADP +This can be performed using standard biochemical techniques, in which the rate of decrease in NADH or NADPH concentration is monitored (by a decrease in absorbance or fluorescence) as the ketoreductase converts the ketone substrate to the corresponding hydroxyl group. In this reaction, NADH or NADPH is consumed (oxidized) by the ketoreductase as it reduces the ketone substrate to the corresponding hydroxyl group. The rate of decrease in NADH or NADPH concentration, measured by the decrease in absorbance or fluorescence per unit time, indicates the relative (enzymatic) activity of the KRED polypeptide in a fixed amount of lysate (or lyophilized powder produced therefrom). If the desired improved enzyme property is thermostability, enzyme activity can be measured after subjecting the enzyme preparation to a predetermined temperature and measuring the amount of enzyme activity remaining after heat treatment. Clones containing polynucleotides encoding KRED are then isolated, sequenced to identify nucleotide sequence changes (if any), and used to express the enzyme in host cells.
[0255] In some embodiments, multiple rounds of mutagenesis can be performed to screen for modified KREDs with desired improved enzymatic properties (e.g., reversal or increase in diastereoselectivity (i.e., selectivity for (5s)-2 over (5r)-2) or activity (% conversion)). For example, KRED polypeptide variants (e.g., SEQ ID NO: 54) derived from L. kefir were selected to have favorable initial activity / selectivity for forming compound (5s)-2, and these variants were used as "backbones" for further rounds of codon optimization and directed evolution. Further rounds of directed evolution can then be performed using genes encoding the most improved polypeptides (e.g., SEQ ID NO: 152; SEQ ID NO: 256) from each round as parent backbone sequences for subsequent rounds of evolution to identify exemplary improved modified KRED polypeptide sequences.
[0256] If the sequence of the modified polypeptide is known, a KRED enzyme-encoding polypeptide can be prepared by standard solid-phase methods according to known synthesis methods. In some embodiments, fragments of up to about 100 bases can be synthesized separately and then joined (e.g., by enzymatic or chemical ligation or polymerase-mediated methods) to form any desired contiguous sequence. For example, polynucleotides and oligonucleotides of the present invention can be prepared by chemical synthesis using, for example, the classical phosphoramidite method described in Beaucage et al., 1981, Tet Lett 22:1859-69, or the method described in Matthes et al., 1984, EMBO J. 3:801-05, as typically performed in automated synthesis methods. According to the phosphoramidite method, oligonucleotides are synthesized, for example, in an automated DNA synthesizer, purified, annealed, ligated, and cloned into an appropriate vector. Additionally, virtually any nucleic acid can be obtained from a variety of commercial sources, such as The Midland Certified Reagent Company, Midland, TX, The Great American Gene Company, Ramona, CA, ExpressGen Inc. Chicago, IL, Operon Technologies Inc., Alameda, CA, and many others.
[0257] The engineered KRED enzymes expressed in host cells can be recovered from the cells and / or culture medium using one or more of the well-known techniques of lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography, among others. A solution suitable for lysis and highly efficient extraction of proteins from bacteria, such as E. coli, is commercially available from Sigma-Aldrich of St. Louis, MO, under the trade name CelLytic B™.
[0258] Chromatographic techniques for isolating KRED polypeptides include, among others, reverse-phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. Conditions for purifying a particular enzyme will depend in part on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and will be apparent to those skilled in the art.
[0259] In some embodiments, affinity techniques can be used to isolate the engineered KRED enzyme. Affinity chromatography purification can utilize antibodies that specifically bind to the KRED polypeptide. For antibody production, various host animals (e.g., rabbits, mice, rats, etc.) can be immunized by injecting the compound. The compound can be coupled to a suitable carrier, such as BSA, via a side chain functional group or a linker attached to the side chain functional group. Depending on the host species, various adjuvants can be used to enhance the immune response, including, but not limited to, Freund's (complete and incomplete), inorganic gels such as aluminum hydroxide, surfactants such as lysolecithin, pluronic polyols, polyanions, peptides, oil emulsions, keyhole limpet hemocyanin, dinitrophenol, and potentially useful human adjuvants such as BCG (Bacillus Calmette-Guerin) and Corynebacterium parvum.
[0260] In some embodiments, the methods provided herein can generate or produce engineered KREDs with reversed stereoselectivity (e.g., diastereoselectivity (i.e., preference for (5s)-2 over (5r)-2)). In particular, the methods provided herein reverse the diastereoselectivity of a KRED polypeptide. For example, wild-type L. kefir KRED (e.g., SEQ ID NO: 492) exhibits specificity for the formation of the (trans) alcohol product (i.e., (5r)-2) when reducing a substrate (i.e., compound 1). However, upon introduction of one or more mutations provided by the methods of the present disclosure, the resulting engineered KRED polypeptide exhibits reversed selectivity and strong specificity for the formation of the (cis) alcohol product (i.e., the (5s)-2) diastereomer).
[0261] In some embodiments, a parent KRED polynucleotide or polypeptide used in a method for reversing diastereoselectivity is a wild-type KRED (e.g., L. kefir) that exhibits specificity for the formation of the (trans) alcohol product (i.e., (5r)-2). In some embodiments, the wild-type KRED is L. kefir, L. brevis, or L. minor. In some embodiments, the wild-type KRED polypeptide is L. kefir (e.g., SEQ ID NO: 492). Additional parent KRED polynucleotides or polypeptides for use in a method for reversing diastereoselectivity include engineered variants of wild-type KRED that exhibit specificity for the formation of the (trans) alcohol product (i.e., (5r)-2).
[0262] In some embodiments, the methods provided herein can also produce or generate engineered ketoreductase polypeptides with increased stereoselectivity (e.g., diastereoselectivity (i.e., preference for (5s)-2 over (5r)-2)). In particular, the methods provided herein increase the stereoselectivity or diastereoselectivity of a KRED polypeptide toward the formation of the (cis) alcohol product (i.e., (5s)-2). When reducing a substrate (i.e., compound 1), a KRED polypeptide (e.g., an engineered variant of L. kefir (e.g., SEQ ID NO: 54)) may exhibit low specificity toward the formation of the (cis) alcohol product (i.e., (5s)-2). However, upon introduction of one or more mutations provided by the methods of the present disclosure, the engineered KRED polypeptide exhibits a reversal of selectivity and strong specificity toward the formation of the (cis) alcohol product (i.e., the (5s)-2 diastereomer).
[0263] In some embodiments, the parent KRED polynucleotide or polypeptide used in the methods for increasing stereoselectivity is an engineered variant of a wild-type L. kefir KRED polypeptide (e.g., SEQ ID NO: 54, 152, 256). In some embodiments, the engineered variant for use in the methods for increasing stereoselectivity is selected from the amino acid sequences of Table 4, Table 5, or Table 6. In some embodiments, the engineered variant for use in the methods for increasing stereoselectivity is selected from the amino acid sequences of SEQ ID NO: 54, SEQ ID NO: 152, or SEQ ID NO: 256. In some embodiments, the methods provided herein increase the level of stereoselectivity (e.g., % diastereomeric excess) of the KRED polypeptide toward the formation of the (cis) alcohol product (i.e., (5s)-2) by at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the methods provided herein increase the level of stereoselectivity (e.g., % diastereomeric excess) of a KRED polypeptide toward the formation of the (cis) alcohol product (i.e., (5s)-2) to at least about 96%. In some embodiments, the methods provided herein increase the level of stereoselectivity (e.g., % diastereomeric excess) of a KRED polypeptide toward the formation of the (cis) alcohol product (i.e., (5s)-2) to at least about 99%.
[0264] In some embodiments, the methods provided herein can also generate or produce engineered ketoreductase polypeptides with increased activity (e.g., % conversion). In particular, the methods provided herein increase the activity of a KRED polypeptide that converts a substrate (e.g., compound 1) to a product (e.g., (5s)-2). When reducing a substrate (e.g., compound 1) to a product (e.g., (5s)-2), a KRED polypeptide (e.g., a wild-type or engineered variant (e.g., SEQ ID NO: 54)) may exhibit a low conversion rate. However, upon introduction of one or more mutations provided by the methods of the present disclosure, the engineered KRED polypeptide exhibits an increased conversion rate of a substrate (e.g., compound 1) to a product (e.g., (5s)-2).
[0265] In some embodiments, a parent KRED polynucleotide or polypeptide used in a method for increasing activity (e.g., % conversion) is a wild-type KRED (e.g., L. kefir). In some embodiments, the wild-type KRED is L. kefir, L. brevis, or L. minor. In some embodiments, the wild-type KRED is L. kefir (e.g., SEQ ID NO: 492). Additional parent KRED polynucleotides or polypeptides used in a method for increasing activity include engineered variants of wild-type KRED (e.g., SEQ ID NO: 54, 152, or 256). In some embodiments, the modified variant is selected from Table 4, Table 5, or Table 6. In some embodiments, the modified variant is selected from SEQ ID NO: 54, SEQ ID NO: 152, or SEQ ID NO: 256. In some embodiments, the methods provided herein increase the activity (e.g., % conversion) of the modified KRED to at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99.0%, or 100% % conversion. In some embodiments, the methods provided herein improve the activity (e.g., conversion rate or desired product) of the modified KRED compared to the parent KRED, having an Improvement Over Positive Control (FIOP) of greater than about 1.75, greater than about 2.00, greater than about 2.25, greater than about 2.50, greater than about 2.75, greater than about 3.00, greater than about 3.25, greater than about 3.50, greater than about 3.75, greater than about 4.00, greater than about 4.25, greater than about 4.50, greater than about 4.75, greater than about 5.00, greater than about 5.25, greater than about 5.50, greater than about 5.75, greater than about 6.00, greater than about 6.25, or greater than about 6.50. In some embodiments, the methods provided herein improve the activity (e.g., conversion rate or desired product) of the modified KRED compared to the parent KRED, having an FIOP of greater than about 2.25, preferably greater than 3.00. In some embodiments, the modified polypeptides are capable of reducing substrates to products with greater than about 2.25 FIOP conversion and greater than about 99% diastereoselectivity compared to the parent KRED.
[0266] Any of the modified KRED polypeptides used in any of the methods provided herein can be derived from a different naturally occurring KRED by introducing one or more mutations corresponding to any of the mutations provided herein. One of skill in the art would be able to identify any homologous proteins and the corresponding residues in their respective encoding nucleic acids by methods well known in the art, such as sequence alignment and determination of homologous residues. Thus, one of skill in the art would be able to generate mutations in any naturally occurring KRED (e.g., having homology with L. kefir) that correspond to any of the mutations described herein. For example, one of skill in the art would be able to generate mutations in L. brevis or L. minor that correspond to mutations in L. kefir (e.g., mutations described in Tables 2, 4, 5, 6, and 7).
[0267] Any of the modified KRED polypeptides used in the methods herein may have one or more additional modifications. Modifications may include substitutions, deletions, and insertions. The substitutions may be non-conservative, conservative, or a combination of non-conservative and conservative substitutions.
[0268] Methods of using engineered ketoreductase polypeptides and compounds prepared therewith The present disclosure provides methods for using any of the modified KRED polynucleotides, polypeptides, or compositions thereof provided herein. In some embodiments, the modified KREDs of the present disclosure can be used in the form of whole cells, crude extracts, isolated enzymes, or purified enzymes. In some embodiments, the modified KREDs of the present disclosure can be used alone or in immobilized form (e.g., immobilized on a resin).
[0269] In the methods of the disclosure, modified KRED polypeptides of the disclosure are used to selectively catalyze the reduction of a bicyclic ketone substrate. In some embodiments, the methods include contacting or incubating a bicyclic ketone with a KRED polypeptide disclosed herein under reaction conditions suitable for reducing or converting the substrate to a product compound.
[0270] In the methods of the present disclosure, the bicyclic ketone substrate is a compound of formula (IA): [ka] (In the formula, The A and B rings together form a fused cycloalkyl ring, e.g., a C to C 12 represents a cycloalkyl or fused heterocyclyl ring, for example a 6- to 12-membered heterocyclyl; wherein the fused cycloalkyl or heterocyclyl is at least one R 1 , e.g., 1 to 4 R 1 may be substituted with Each R 1 is an amine protecting group, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz), C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, C3-C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 Arylalkyl, 3-14 membered heterocyclyl, 5-20 membered heteroaryl, hydroxyl, halogen, e.g., F, C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , and -(CH2) n -C(=O)NR c R c are independently selected from; Here, alkyl, alkenyl, and alkynyl each represent one or more R a , for example, 1 to 6 R a may be substituted with Cycloalkyl, aryl, arylalkyl, heterocyclyl, and heteroaryl each have one or more R b , for example, 1 to 6 R b may be substituted with; Each R a When each occurs, C3 to C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 Arylalkyl, 3-14 membered heterocyclyl, 5-20 membered heteroaryl, halogen, e.g., F, haloalkyl, e.g., C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , and -(CH2) n -C(=O)NR c R c Selected from; Each R b are independently halogens, C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , -(CH2) n -C(=O)NR c R c , C1~C 20 Alkyl, C2-C 20 Alkenyl and C2-C 20 alkynyl; Each R c When each occurs, H, C1 to C 20 Alkyl, C2-C20 Alkenyl and C2-C 20 Alkynyl (each one or more R b , for example, 1 to 6 R b and n is 0, 1, 2, 3, 4, 5, or 6, for example, 0, 1, 2, or 3. It is expressed as:
[0271] In one embodiment, the A and B rings together represent a fused 6- to 12-membered heterocyclyl containing at least one nitrogen heteroatom. In a further embodiment, the nitrogen heteroatom of the bicyclic ketone substrate is bonded to an amine protecting group, e.g., tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz).
[0272] In one embodiment, the bicyclic ketone substrate is represented by formula (IA)-I [ka] (In the formula, X is NR 1a , CH2, and CH-R 1b Selected from; R 1a is an amine protecting group, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz), C1-C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, C3-C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 selected from arylalkyl, 3- to 14-membered heterocyclyl, and 5- to 20-membered heteroaryl; R 1b is C1~C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, C3-C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20selected from arylalkyl, 3- to 14-membered heterocyclyl, and 5- to 20-membered heteroaryl; where R 1a or R 1b The alkyl, alkenyl, and alkynyl groups each have 1 to 6 R a may be substituted with; R 1a or R 1b The cycloalkyl, aryl, arylalkyl, heterocyclyl, and heteroaryl each have 1 to 6 R b may be substituted with; Each R 1c When each occurs, C1 to C 20 Alkyl, C2-C 20 Alkenyl, C2-C 20 Alkynyl, C3-C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 Arylalkyl, 3-14 membered heterocyclyl, 5-20 membered heteroaryl, hydroxyl, halogen, e.g., F, C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , and -(CH2) n -C(=O)NR c R c selected from the group consisting of: Each R a When each occurs, C3 to C 10 Cycloalkyl, C6-C 14 Aryl, C7-C 20 Arylalkyl, 3-14 membered heterocyclyl, 5-20 membered heteroaryl, halogen, e.g., F, haloalkyl, e.g., C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c, -(CH2) n -C(=O)R c , and -(CH2) n -C(=O)NR c R c Selected from; Each R b are independently halogens, C1-C 20 Haloalkyl, e.g., -CF3, -OR c , -NR c R c , -(CH2) n COOR c , -(CH2) n -C(=O)R c , -(CH2) n -C(=O)NR c R c , C1~C 20 Alkyl, C2-C 20 Alkenyl and C2-C 20 alkynyl; Each R c When each occurs, H, C1 to C 20 Alkyl, C2-C 20 Alkenyl and C2-C 20 Alkynyl (each of which has 1 to 6 R b and optionally substituted with; n is 0, 1, 2, or 3; and m is 0, 1, or 2) It is expressed as:
[0273] In one embodiment of (IA)-I, X is NR 1a and R 1a is an amine protecting group, for example, tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz).
[0274] In a further embodiment of (IA)-I, X is NR 1a and R 1a is C1~C 10 Alkyl, C3-C 10 Cycloalkyl, C6-C14 Aryl, C7-C 20 arylalkyl, and amine protecting groups, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz); and m=0.
[0275] In a further embodiment of (IA)-I, X is NR 1a and;R 1a is an amine protecting group, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz); and m is 0.
[0276] In the methods of the disclosure, a modified KRED polypeptide is used to generate (IA) (substrate): [ka] (5s)-IA (product): [ka] can selectively catalyze the reduction of
[0277] The engineered KRED polypeptide is used to generate (IA)-I (substrate): [ka] (5s)-IA-I (product): [ka] can selectively catalyze the reduction of
[0278] In the methods of the disclosure, modified KRED polypeptides can be used to selectively catalyze the reduction of compounds of formula (IA)-I to (IB)-I: [ka] (In the formula, R1a is an amine protecting group). In certain embodiments, R 1a is selected from tert-butyloxycarbonyl (Boc) and N-carboxybenzyl (Cbz).
[0279] In the disclosed methods, modified KRED polypeptides are used to synthesize Compound 1 (substrate): tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole 2(1H)-carboxylate: [ka] (5s)-2(product) tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate: [ka] The compound selectively catalyzes the reduction of HCl to HCl.
[0280] In some embodiments, a method for selectively reducing a compound of Formula (IA) or Formula (IA)-I (substrate) to (5s)-IA or (5s)-IA-I (product), e.g., compound I to (5s)-2, comprises contacting or incubating the substrate with a KRED polypeptide disclosed herein under reaction conditions suitable for reducing or converting the substrate to the product compound. In some embodiments, the substrate is converted to the (cis)-alcohol product (5s)-IA or (5s)-IA-I, e.g., (5s)-2, or to the corresponding (trans)-alcohol product (5r)-IA or (5r)-IA-I, e.g., (5r)-2: [ka] to a diastereomeric excess of about 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%.
[0281] In some embodiments, a method for selectively reducing compound 1 (substrate) to (5s)-2 (product) comprises contacting or incubating the substrate with a modified KRED polypeptide disclosed herein under reaction conditions suitable for reducing or converting the substrate to a product compound. In some embodiments, a method for selectively reducing compound 1 (substrate) to (5s)-2 (product) comprises contacting or incubating the substrate with a modified KRED polypeptide disclosed herein under reaction conditions suitable for reducing or converting the substrate to a product compound. In some embodiments, the substrate is contacted ... [ka] In some embodiments, the substrate is reduced to the (cis) alcohol product ((5r)-2) in a diastereomeric excess of greater than about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% relative to the corresponding (cis) alcohol product ((5s)-2). In some embodiments, the substrate is reduced to the (cis) alcohol product ((5r)-2) in a diastereomeric excess of greater than about 96% relative to the corresponding (trans) alcohol product ((5s)-2). In some embodiments, the substrate is reduced to the (cis) alcohol product ((5r)-2) in a diastereomeric excess of greater than about 99% relative to the corresponding (trans) alcohol product ((5s)-2).
[0282] In some embodiments, a method for selectively reducing compound 1 (substrate) to (5s)-2 (product) comprises contacting or incubating the substrate with a modified KRED polypeptide disclosed herein under reaction conditions suitable for reducing or converting the substrate to a product compound with greater selectivity (% diastereomeric excess) and / or activity (% conversion) than a reference (e.g., parent) sequence (e.g., SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, SEQ ID NO:492). In some embodiments, a method for selectively reducing compound 1 (substrate) to (5s)-2 (product) comprises contacting or incubating the substrate with a modified KRED polypeptide disclosed herein under reaction conditions suitable for reducing or converting the substrate to a product compound with greater selectivity (% diastereomeric excess) and / or activity (% conversion) than SEQ ID NO:54, SEQ ID NO:152, SEQ ID NO:256, and / or SEQ ID NO:492.
[0283] As known to those skilled in the art, ketoreductase-catalyzed reduction reactions typically require a cofactor. The reduction reactions catalyzed by the engineered KRED enzymes described herein also typically require a cofactor, although many embodiments of the engineered ketoreductases require far fewer cofactors than reactions catalyzed by wild-type KRED enzymes. As used herein, the term "cofactor" refers to a non-protein compound that acts in combination with the ketoreductase enzyme. Suitable cofactors for use with the engineered ketoreductase enzymes described herein include NADP + (nicotinamide adenine dinucleotide phosphate), NADPH (NADP + reduced form), NAD + (nicotinamide adenine dinucleotide) and NADH (NAD + Typically, the reduced form of the cofactor is added to the reaction mixture. The reduced NAD(P)H form can be optionally converted to oxidized NAD(P) using a cofactor regeneration system. +In some embodiments, the KRED enzyme polypeptide used in the methods provided herein is isolated and / or purified, and the reduction reaction is carried out in the presence of a cofactor. In some embodiments, the cofactor is nicotinamide adenine dinucleotide (NADH) or nicotinamide adenine dinucleotide phosphate (NADPH).
[0284] In some embodiments, suitable reaction conditions include NADP + , NADPH, NAD + and NADH at a concentration of about 0.05 g / L to about 0.1 g / L, about 0.1 g / L to about 10 g / L, about 0.2 g / L to about 5 g / L, or about 0.5 g / L to about 2.5 g / L. In some embodiments, the cofactor is NADH or NADPH. Thus, in some embodiments, suitable reaction conditions may include the cofactor NADH or NADPH at a concentration of about 0.05 g / L to about 10 g / L, 0.1 g / L to about 10 g / L, about 0.2 g / L to about 5 g / L, or about 0.5 g / L to about 2.5 g / L. In some embodiments, reaction conditions include about 10 g / L or less, about 5 g / L or less, about 2.5 g / L or less, about 1.0 g / L or less, about 0.5 g / L or less, or about 0.05 g / L or less. In some embodiments, suitable reaction conditions for the methods provided herein include about 0.05 g / L to about 1.0 g / L NADP+. In some embodiments, suitable reaction conditions for the methods provided herein include about 0.03 wt% to about 0.2 wt% cofactor (e.g., NADP+). In some embodiments, suitable reaction conditions for the methods provided herein include at least about 0.2 wt% cofactor (e.g., NADP+). In some embodiments, suitable reaction conditions for the methods provided herein include less than about 2 wt% cofactor (e.g., NADP+).
[0285] In some embodiments of the process (e.g., when whole cells or lysates are used), the cofactor is naturally present in the cell extract and does not need to be supplemented. In some embodiments of the process (e.g., when partially purified or purified aldolases are used), the process can further include adding a cofactor to the enzyme reaction mixture. In some embodiments, the cofactor is added at the beginning of the reaction and / or additional cofactor is added during the reaction.
[0286] The term "cofactor regeneration system" refers to a reaction that reduces the oxidized form of a cofactor (e.g., NADP + The term "cofactor regeneration system" refers to a series of reactions involved in the reduction of a keto substrate (from NADPH to NADPH). Cofactors oxidized by the reduction of a keto substrate catalyzed by a ketoreductase are regenerated in reduced form by a cofactor regeneration system. The cofactor regeneration system includes a source of reducing hydrogen equivalents and a stoichiometric reducing agent capable of reducing the oxidized form of the cofactor. The cofactor regeneration system may further include a catalyst, for example, an enzyme catalyst that catalyzes the reduction of the oxidized form of the cofactor by the reducing agent. NAD + or NADP + Cofactor regeneration systems for regenerating NADH or NADPH from, respectively, are known in the art and can be used in the methods described herein.
[0287] Exemplary suitable cofactor regenerating systems that can be used include, but are not limited to, glucose and glucose dehydrogenase, formate and formate dehydrogenase, glucose-6-phosphate and glucose-6-phosphate dehydrogenase, secondary (e.g., isopropanol) alcohol and secondary alcohol dehydrogenase, phosphite and phosphite dehydrogenase, molecular hydrogen and hydrogenase, etc. These systems use NADP as a cofactor. + / NADPH or NAD + / NADH. Electrochemical regeneration using hydrogenase can also be used as a cofactor regeneration system. See, e.g., U.S. Pat. Nos. 5,538,867 and 6,495,023, both of which are incorporated herein by reference. Chemical cofactor regeneration systems including a metal catalyst and a reducing agent (e.g., molecular hydrogen or formic acid) are also suitable. See, e.g., WO 2000 / 053731, which is incorporated herein by reference.
[0288] The terms "glucose dehydrogenase" and "GDH" are used interchangeably herein and refer to the enzymes that bind D-glucose and NAD + or NADP + to gluconate and NADH or NADPH, respectively. + or NADP + The following equation (1) shows the NAD catalyzed by glucose dehydrogenase. + or NADP + The reduction of by glucose is described. [ka]
[0289] Glucose dehydrogenases suitable for use in practicing the methods described herein include both naturally occurring and non-naturally occurring glucose dehydrogenases. Genes encoding naturally occurring glucose dehydrogenases have been reported in the literature. For example, the Bacillus subtilis 61297 GDH gene was reported to be expressed in E. coli and exhibit the same physicochemical properties as the enzyme produced in its native host (Vasantha et al., 1983, Proc. Natl. Acad. Sci. USA 80:785). The gene sequence of the B. subtilis GDH gene corresponding to GenBank accession number M12276 was reported by Lampel et al., 1986, J. Bacteriol. 166:238-243, and in a modified form by Yamane et al., 1996, Microbiology 142:3047-3056 under GenBank accession number D50453. Naturally occurring GDH genes also include those encoding GDH from B. cereus ATCC14579 (Nature, 2003, 423:87-91; Genbank Accession No. AE01 7013) and B. megaterium (Eur. J. Biochem., 1988, 174:485-490, Genbank Accession No. X12370; J. Ferment. Bioeng., 1990, 70:363-369, Genbank Accession No. 01216270). Glucose dehydrogenase from Bacillus sp. is presented in WO 2005 / 018579 as SEQ ID NOs: 10 and 12 (encoded by polynucleotide sequences corresponding to SEQ ID NOs: 9 and 11, respectively), the disclosure of which is incorporated herein by reference.
[0290] Non-naturally occurring glucose dehydrogenases can be generated, for example, by mutagenesis, directed evolution, etc. GDH enzymes with suitable activity, whether naturally occurring or not, can be readily identified using the assay described in Example 4 of WO 2005 / 018579 (the disclosure of which is incorporated herein by reference). Exemplary non-naturally occurring glucose dehydrogenases are set forth in WO 2005 / 018579 as SEQ ID NOS: 62, 64, 66, 68, 122, 124, and 126. The polynucleotide sequences encoding them are set forth in WO 2005 / 018579 as SEQ ID NOS: 61, 63, 65, 67, 121, 123, and 125, which sequences are incorporated herein by reference. Additional non-naturally occurring glucose dehydrogenases suitable for use in the reduction reactions catalyzed by the ketoreductases disclosed herein are disclosed in U.S. Patent Application Publication Nos. 2005 / 0095619 and 2005 / 0153417, the disclosures of which are incorporated herein by reference.
[0291] The glucose dehydrogenases used in the ketoreductase-catalyzed reduction reactions described herein have a catalytic activity of at least about 10 μmol / min / mg, and sometimes at least about 10 μmol / min / mg, in the assay described in Example 4 of WO 2005 / 018579. 2 μmol / min / mg or approximately 10 3 μmol / min / mg, maximum approximately 10 4 It can exhibit activity of μmol / min / mg or higher.
[0292] As disclosed herein and illustrated in the Examples, the present disclosure contemplates a range of suitable reaction conditions that may be used in the processes herein, including, but not limited to, pH, temperature, buffer, solvent system, substrate load, product stereoisomers, e.g., mixtures of diastereoisomers, polypeptide load, cofactor load, pressure, and reaction time. Additional suitable reaction conditions for carrying out the methods of enzymatically converting substrate compounds to product compounds using the modified KRED polypeptides described herein can be readily optimized by routine experimentation, including, but not limited to, contacting the modified KRED polypeptide with substrate compounds under experimental reaction conditions of various concentrations, pH, temperatures, and solvent conditions, and detecting the product compounds, for example, using methods described in the Examples provided herein.
[0293] The ketoreductase-catalyzed reduction reactions described herein are generally carried out in a solvent. Suitable solvents include water, organic solvents (e.g., ethyl acetate, isopropyl acetate, butyl acetate, methanol, ethanol, n-propanol, isopropanol, dimethyl sulfoxide, dimethylformamide, 1-octanol, hexane, heptane, octane, methyl t-butyl ether (MTBE), toluene, glycerol, polyethylene glycol, etc.), and ionic liquids (e.g., 1-ethyl 4-methylimidazolium tetrafluoroborate, l-butyl 3-methylimidazolium tetrafluoroborate, l-butyl 3-methylimidazolium hexafluorophosphate, etc.). In some embodiments, an aqueous solvent, such as water and an aqueous co-solvent system, is used.
[0294] In some embodiments, the solvent is present at a concentration of 0 to 500 g / L. In some embodiments, the solvent is present at a concentration of 100 to 200 g / L. In some embodiments, the solvent is present at a concentration of 0% to 100% v / v. In some embodiments, the solvent is present at a concentration of 20% to 40% v / v.
[0295] Exemplary aqueous co-solvent systems include water and one or more organic solvents. Generally, the organic solvent component of the aqueous co-solvent system is selected so as not to completely inactivate the ketoreductase enzyme. Suitable co-solvent systems can be readily identified by measuring the enzymatic activity of a particular engineered ketoreductase with a given substrate of interest in a candidate solvent system using an enzyme activity assay, such as those described herein.
[0296] The organic solvent component of the aqueous co-solvent system may be miscible with the aqueous component to provide a single liquid phase, or may be partially miscible or immiscible with the aqueous component to provide two liquid phases. Generally, when an aqueous co-solvent system is used, it is selected to be biphasic, with water dispersed in the organic solvent or vice versa.
[0297] Generally, when an aqueous co-solvent system is utilized, it is desirable to select an organic solvent that can be easily separated from the aqueous phase. In general, the ratio of water to organic solvent in the co-solvent system typically ranges from about 90:10 to about 10:90 (v / v) organic solvent to water, and from 80:20 to 20:80 (v / v) organic solvent to water. The co-solvent system may be preformed prior to addition to the reaction mixture, or may be formed in situ in the reaction vessel.
[0298] In some embodiments, the organic solvent in the aqueous co-solvent system is selected from ethyl acetate, isopropyl acetate, butyl acetate, methanol, ethanol, N-propanol, isopropanol, dimethyl sulfoxide, dimethylformamide, 1-octanol, hexane, heptane, octane, methyl t-butyl ether (MTBE), toluene, glycerol, polyethylene glycol, etc., and ionic liquids (e.g., 1-ethyl-4-methyl-imidazolium tetrafluoroborate, l-butyl-3-methylimidazolium tetrafluoroborate, l-butyl-3-methylimidazolium hexafluorophosphate, etc.). In some embodiments, the aqueous co-solvent system is water / isopropanol. In certain embodiments, isopropanol is present at a concentration of 20% to 40% v / v.
[0299] The aqueous solution (water or aqueous co-solvent system) can be pH buffered or unbuffered. Generally, the reduction can be carried out at a pH of about 10 or less, usually in the range of about 5 to about 10. In some embodiments, the reduction is carried out at a pH of about 9 or less, usually in the range of about 5 to about 9. In some embodiments, the reduction is carried out at a pH of about 8 or less, often in the range of about 5 to about 8, usually in the range of about 6 to about 8. The reduction can also be carried out at a pH of about 7.8 or less, or 7.5 or less. Alternatively, the reduction can be carried out at a neutral pH, i.e., about 7.
[0300] During the reduction reaction, the pH of the reaction mixture may change. The pH of the reaction mixture may be maintained at a desired pH or within a desired pH range by adding an acid or base during the course of the reaction. Alternatively, the pH may be controlled by using an aqueous solvent containing a buffer. Suitable buffers for maintaining a desired pH range are known in the art and include, for example, phosphate buffer, triethanolamine buffer, etc. A combination of buffering and the addition of an acid or base may also be used.
[0301] When a glucose / glucose dehydrogenase cofactor regeneration system is used, the co-production of gluconic acid (pKa = 3.6), as represented by Equation (1), will lower the pH of the reaction mixture if the resulting aqueous gluconic acid solution is not otherwise neutralized. The pH of the reaction mixture can be maintained at the desired level by standard buffering techniques, in which a buffer neutralizes gluconic acid up to its buffering capacity, or by adding a base during the conversion process. A combination of buffering and base addition can also be used. Suitable buffers for maintaining the desired pH range are described above. Suitable bases for neutralizing gluconic acid include organic bases, such as amines, alkoxides, and inorganic bases, such as hydroxide salts (e.g., NaOH), carbonate salts (e.g., NaHCO), bicarbonate salts (e.g., KCO), and basic phosphate salts (e.g., KHPO, NaPO). The addition of base during the conversion process can be performed manually while monitoring the pH of the reaction mixture, or, more conveniently, by using an automatic titrator as a pH meter. A combination of partial buffering capacity and base addition can also be used for process control.
[0302] When base addition is used to neutralize the gluconic acid released during the ketoreductase-catalyzed reduction reaction, the progress of the conversion can be monitored by the amount of base added to maintain the pH. Typically, base added to an unbuffered or partially buffered reaction mixture throughout the course of the reduction is added in aqueous solution.
[0303] In some embodiments, the cofactor regeneration system can include formate dehydrogenase. The terms "formate dehydrogenase" and "FDH" are used interchangeably herein to refer to the synthesis of formate and NAD. + or NADP + to carbon dioxide and NADH or NADPH, respectively. + or NADP +"Cofactor-dependent enzymes" refers to formate dehydrogenases suitable for use as cofactor regeneration systems in the ketoreductase-catalyzed reduction reactions described herein include both naturally occurring and non-naturally occurring formate dehydrogenases. Formate dehydrogenases include those corresponding to SEQ ID NO: 70 (Pseudomonas sp.) and SEQ ID NO: 72 (Candida boidinii), which are encoded by polynucleotide sequences corresponding to SEQ ID NOs: 69 and 71, respectively, in WO 2005 / 018579 (the disclosure of which is incorporated herein by reference). Formate dehydrogenases used in the methods described herein, whether naturally occurring or non-naturally occurring, have a cofactor regeneration rate of at least about 1 μmol / min / mg, sometimes at least about 10 μmol / min / mg, or at least about 10 μmol / min / mg. 2 μmol / min / mg, maximum approximately 10 3 It may exhibit activity of μmol / min / mg or higher and can be easily screened using, for example, the assay described in Example 4 of WO 2005 / 018579.
[0304] As used herein, the term "formic acid" refers to formate anion (HC0), formic acid (HC0H), and mixtures thereof. Formic acid can be provided in the form of a salt, typically an alkali or ammonium salt (e.g., HC0Na, KHCONH, etc.), in the form of formic acid, typically an aqueous formic acid solution, or in the form of a mixture thereof. Formic acid is a moderate acid. In aqueous solutions within a few pH units of its pKa (aqueous solution of pKa = 3.7), formic acid is a hydroxyl group of HCl, ... - At pH values above about pH 4, formic acid exists primarily as HCO2 -When the formic acid is provided as formic acid, the reaction mixture is typically buffered or reduced in acidity by adding a base to provide the desired pH, typically to a pH of about 5 or greater. Bases suitable for neutralizing the formic acid include, but are not limited to, organic bases such as amines, alkoxides, and the like, and inorganic bases such as hydroxide salts (e.g., NaOH), carbonate salts (e.g., NaHCO), bicarbonate salts (e.g., KCO), basic phosphate salts (e.g., KHPO, NaPO), and the like.
[0305] Formic acid is mainly HCO2 - At pH values above about pH 5, the following equation (2) shows that NAD exists as + or NADP + The formate dehydrogenase-catalyzed reduction of [ka]
[0306] When formate and formate dehydrogenase are used as the cofactor regenerating system, the pH of the reaction mixture can be maintained at the desired level by standard buffering techniques in which the buffer releases protons up to the buffering capacity provided, or by adding an acid during the course of the conversion. Suitable acids to add during the reaction to maintain pH include organic acids, such as carboxylic acids, sulfonic acids, phosphonic acids, and the like; inorganic acids, such as hydrohalic acids (e.g., hydrochloric acid), sulfuric acid, phosphoric acid, and the like; acid salts, such as dihydrogen phosphate (e.g., KH2PO4), bisulfate (e.g., NaHSO4), and the like. Some embodiments utilize formic acid, thereby maintaining both the formate concentration and the pH of the solution.
[0307] When acid addition is used to maintain the pH during the reduction reaction using a formate / formate dehydrogenase cofactor regenerating system, the progress of the conversion can be monitored by the amount of acid added to maintain the pH. Typically, acids added to unbuffered or partially buffered reaction mixtures throughout the course of the conversion are added in aqueous solution.
[0308] The terms "secondary alcohol dehydrogenase" and "sADH" are used interchangeably herein and refer to the enzymes that react with D-glucose and NAD + or NADP + to gluconate and NADH or NADPH, respectively. + or NADP + The following formula (3) shows the NAD-dependent enzyme by secondary alcohols, exemplified by isopropanol. + or NADP + The reduction of [ka]
[0309] Secondary alcohol dehydrogenases suitable for use as cofactor regeneration systems in the ketoreductase-catalyzed reduction reactions described herein include both naturally occurring and non-naturally occurring secondary alcohol dehydrogenases, including the known alcohol dehydrogenases of Thermoanerobium brockii, Rhodococcus etythropolis, Lactobacillus kefir, Lactobacillus minor, and Lactobacillus brevis, as well as modified alcohol dehydrogenases derived from them. The secondary alcohol dehydrogenase used in the methods described herein, whether naturally occurring or non-naturally occurring, has a catalysis of at least about 1 μmol / min / mg, sometimes at least about 10 μmol / min / mg, or at least about 10 μmol / min / mg. 2 μmol / min / mg, maximum approximately 10 3 It can exhibit activity of μmol / min / mg or higher.
[0310] Suitable secondary alcohols include lower secondary alkanols and aryl-alkylcarbinols. Examples of lower secondary alcohols include isopropanol, 2-butanol, 3-methyl-2-butanol, 2-pentanol, 3-pentanol, 3,3-dimethyl-2-butanol, etc. In one embodiment, the secondary alcohol is isopropanol. Suitable aryl-alkylcarbinols include unsubstituted and substituted 1-arylethanols.
[0311] In one embodiment, when oxidation of isopropanol to acetone is used to regenerate NADH / NADPH, the reaction can be carried out at reduced pressure so that acetone is removed from the reaction mixture.
[0312] When secondary alcohols and secondary alcohol dehydrogenase are used as cofactor regenerating systems, the resulting NAD + Alternatively, NADP+ is reduced by the coupled oxidation of a secondary alcohol to a ketone by a secondary alcohol dehydrogenase. Some engineered ketoreductases also have the activity of dehydrogenating a secondary alcohol reducing agent. In some embodiments where a secondary alcohol is used as the reducing agent, the engineered ketoreductase and the secondary alcohol dehydrogenase are the same enzyme.
[0313] When using a cofactor regeneration system to perform embodiments of the ketoreductase-catalyzed reduction reactions described herein, the oxidized or reduced form of the cofactor can be initially provided. As described above, the cofactor regeneration system converts the oxidized cofactor to its reduced form, which is then utilized in the reduction of the ketoreductase substrate.
[0314] In some embodiments, a cofactor regeneration system is not used. For reduction reactions performed without a cofactor regeneration system, the cofactor is added to the reaction mixture in reduced form. In some embodiments, reaction conditions suitable for the methods provided herein do not require recycling of the GDH / glucose cofactor.
[0315] In some embodiments, the method is carried out using whole cells expressing the ketoreductase enzyme, or an extract or lysate of such cells. When the process is carried out using whole cells of a host organism, the whole cells may naturally provide the cofactor. Alternatively, or in combination, the cells may naturally or recombinantly provide glucose dehydrogenase.
[0316] When carrying out the stereoselective reduction reactions described herein, the engineered KRED enzyme and any enzymes, including an optional cofactor regeneration system, can be added to the reaction mixture in the form of purified enzymes, whole cells transformed with genes encoding the enzymes, and / or cell extracts and / or lysates of such cells. The genes encoding the engineered ketoreductase enzyme and the optional cofactor regeneration enzyme can be transformed into host cells separately or together into the same host cell. For example, in some embodiments, one set of host cells can be transformed with genes encoding the engineered KRED enzyme and another set can be transformed with genes encoding the cofactor regeneration enzyme. Both sets of transformed cells can be utilized together in the reaction mixture in the form of whole cells or in the form of lysates or extracts derived therefrom. In other embodiments, host cells can be transformed with genes encoding both the engineered ketoreductase enzyme and the cofactor regeneration enzyme.
[0317] Whole cells transformed with genes encoding the engineered KRED enzyme and / or optional cofactor-regenerating enzyme, or cell extracts and / or lysates thereof, can be used in a variety of different forms, including solid (e.g., lyophilized, spray dried, etc.) or semi-solid (e.g., a crude paste).
[0318] Cell extracts or cell lysates may be partially purified by precipitation (ammonium sulfate, polyethyleneimine, heat treatment, etc.) followed by desalting procedures (e.g., ultrafiltration, dialysis, etc.) prior to lyophilization. Any cell preparation may be stabilized by cross-linking with known cross-linking agents such as glutaraldehyde, or by fixation to a solid phase material (e.g., Eupergit C, resin, etc.).
[0319] In any of the process embodiments disclosed herein, the reaction is carried out under suitable reaction conditions as described herein, and the engineered ketoreductase polypeptide is immobilized on a solid support, such as a membrane, resin, solid support, or other solid phase material. The solid support can be composed of organic polymers such as microcrystalline cellulose, polystyrene, polyethylene, polypropylene, polyfluoroethylene, polyethyleneoxy, polymethacrylate, and polyacrylamide, as well as copolymers and grafts thereof. The solid support can also be Celite (diatomaceous earth), glass, silica, controlled pore glass (CPG), reverse-phase silica, or metals such as gold or platinum. The solid support can be in the form of beads, spheres, particles, granules, gels, membranes, or surfaces. The surfaces can be planar, substantially planar, or non-planar. The solid support can be porous or non-porous, and can be swellable or non-swellable. The solid support can be configured in the form of a well, depression, or other container, vessel, feature, or location. Useful solid supports for immobilizing engineered ketoreductase enzymes for carrying out reactions include, but are not limited to, beads or resins such as polymethacrylates, e.g., epoxy-functionalized polymethacrylates, amino-functionalized polymethacrylates, polymethacrylates, styrene / DVB copolymers, or octadecyl-functionalized polymethacrylates. In certain embodiments, the solid support is a bead or resin comprising polymethacrylate.
[0320] Exemplary solid supports include, but are not limited to, chitosan beads, Eupergit C, IB-150, IB-350, IB-C435, IB-A369, IB-A161, IB-A171, IBS500, IB-S861, SEPABEADS (Mitsubishi) such as Sepabeads EC-EP, Sepabeads EC-HFA, Sepabeads EC-HG, Sepabeads EC-BU, Sepabeads EC-OD, Sepabeads EC-CM, Sepabeads EC-IDA, Sepabeads EC-EA, Sepabeads EC-HA, Sepabeads EC-QA, Sepabeads EXE, Sepabeads EXA, Dilbeads-TA, Amberzyme Oxirane, Amberlite XAD-7HP, Amberlite FPA98Cl, Amberlite IRA958Cl, Amberlite IRA67, Amberlite FPA90Cl, Amberlite FPA40Cl, Amberlite XAD18, Accurel EP100, ECR8206F / 5730, ECR8206 / 5803, ECR8206M / 5749, ReliZyme EP403, ReliZyme EP113, Lewatit VP OC 1600, Diaion WA20, Diaion WA21J, Diaion WA30, Dowex 66, Diaion HPA-25L, Lewatit VP OC 1064 MD PH, Lewatit VP OC 1163, Lifetech ECR8304F, Lifetech ECR8309F, Lifetech ECR8315F, Lifetech ECR8204F, Lifetech ECR8285, Lifetech ECR1090M, Lifetech ECR1030M, Lifetech ECR8806M, Chromalite(MAM2 / F)D6591, Chromalite MIDA / M, Chromalite MIDA / M / Fe, Chromalite MIDA / M / Co, Chromalite MIDA / M / Ni, Chromalite MIDA / M / Cu, and Chromalite Examples include MIDA / M / Zn.
[0321] In any of the embodiments of the processes disclosed herein in which the modified polypeptide is expressed in the form of a secreted polypeptide, the culture medium containing the secreted polypeptide can be used in the processes herein.
[0322] In any of the process embodiments disclosed herein, solid reactants (e.g., enzymes, salts, etc.) can be provided to the reaction in a variety of different forms, including powders (e.g., lyophilized, spray-dried, etc.), solutions, emulsions, suspensions, etc. Reactants can be readily lyophilized or spray-dried using methods and equipment known to those of skill in the art. For example, a protein solution can be frozen in aliquots at -80°C and then placed in a pre-cooled lyophilization chamber, after which a vacuum is applied. After removing water from the sample, the temperature is raised to 4°C, typically for 2 hours, after which the vacuum is released and the lyophilized sample is removed.
[0323] The amounts of reactants used in the reduction reaction will generally vary depending on the amount of desired product and the concomitant amount of ketoreductase substrate used. The following guidelines can be used to determine the amount of ketoreductase, cofactor, and optional cofactor regeneration system to use. Generally, the keto substrate can be used at a concentration of about 5 to 150 grams / liter, using about 50 mg to about 5 g of ketoreductase and about 10 mg to about 150 mg of cofactor.
[0324] In some embodiments, suitable reaction conditions for the methods provided herein include about 5 g / L to about 150 g / L of substrate. In some embodiments, suitable reaction conditions for the methods provided herein use substrate concentrations of at least about 5 g / L, at least about 10 g / L, at least about 20 g / L, at least about 50 g / L, at least about 100 g / L, or at least about 150 g / L. In some embodiments, the ketoreductase concentration is less than about 5 g / L. In some embodiments, the ketoreductase concentration is at least 3 g / L. In some embodiments, suitable reaction conditions for the methods provided herein include a ketoreductase load of at least about 1 wt.%. In some embodiments, suitable reaction conditions for the methods provided herein include a ketoreductase load of 3 wt.%. In some embodiments, suitable reaction conditions for the methods provided herein include a ketoreductase load of about 0.8 wt.% and 150 g / L of ketone substrate.
[0325] Those skilled in the art will readily understand how to vary these amounts to suit the desired level of productivity and production scale. Appropriate amounts of the optional cofactor regeneration system can be readily determined by routine experimentation based on the amount of cofactor and / or ketoreductase utilized. Generally, reducing agents (e.g., glucose, formic acid, and isopropanol) are utilized at levels greater than equimolar levels of the ketoreductase substrate to achieve substantially complete or near complete conversion of the ketoreductase substrate.
[0326] In any of the process embodiments disclosed herein, the order of addition of reactants is not important. The reactants may be added together to a solvent (e.g., a one-phase solvent, a two-phase aqueous co-solvent system, etc.) at the same time, or some of the reactants may be added separately and some may be added together at different times. For example, the cofactor regenerating system, cofactor, ketoreductase, and ketoreductase substrate may be added to the solvent first.
[0327] When an aqueous co-solvent system is used, to improve mixing efficiency, the cofactor regenerating system, ketoreductase, and cofactor may be added to the aqueous phase first and mixed. The organic phase may then be added and mixed, followed by the ketoreductase substrate. Alternatively, the ketoreductase substrate may be premixed with the organic phase and then added to the aqueous phase.
[0328] Suitable conditions for carrying out the ketoreductase-catalyzed reduction reactions described herein include a wide variety of conditions that can be readily optimized by routine experimentation, including, but not limited to, contacting the engineered ketoreductase enzyme and substrate at an experimental pH and temperature, and detecting the product using, for example, the methods described in the Examples provided herein.
[0329] The ketoreductase-catalyzed reduction is typically carried out at a temperature ranging from about 15° C. to about 75° C. In some embodiments, the reaction is carried out at a temperature ranging from about 20° C. to about 55° C. In yet other embodiments, the reaction is carried out at a temperature ranging from about 20° C. to about 45° C. In some embodiments, the reaction is carried out at 40° C. The reaction may also be carried out under ambient conditions.
[0330] The reduction reaction generally proceeds until the reduction of the substrate is substantially complete or nearly complete. In some embodiments, isopropanol (iPrOH) is used as a solvent. In some embodiments, suitable reaction conditions for the methods provided herein include up to 40% (v / v) iPrOH. Optionally, acetone formed upon oxidation of iPrOH is removed. Further, the addition of iPrOH can promote completion of the reaction. The reduction of the substrate to the product can be monitored using known methods by detecting the substrate and / or the product. Suitable methods include gas chromatography, HPLC, and the like. The conversion yield of the alcohol reduction product produced in the reaction mixture is generally greater than about 50%, but can also be greater than about 60%, 70%, 80%, or 90%, and often greater than about 97%. In some embodiments, the conversion yield of the alcohol reduction product produced in the reaction mixture is generally greater than about 85%, with up to complete conversion.
[0331] Whether the method is carried out with whole cells, cell extracts, or purified ketoreductase enzymes, a single ketoreductase enzyme can be used, or a mixture of two or more ketoreductase enzymes can be used.
[0332] Suitable reaction conditions can include a combination of reaction parameters that provide for the biocatalytic conversion of a substrate compound to its corresponding product compound. Thus, in some embodiments of the process, the combination of reaction parameters includes one or more of the following: i) substrate load, e.g., a Compound 1 load of about 5 g / L to about 150 g / L; ii) an engineered enzyme polypeptide concentration of at least about 3 g / L, or at least about 1% by weight, or less than about 3% by weight; iii) cofactor / cofactor load, e.g., an NADP+ cofactor load of about 10 mg to about 150 mg, or at least about 0.1% by weight NADP+ cofactor, or less than about 2% by weight NADP+ cofactor; iv) a cosolvent concentration of about 20% (v / v) to about 60% (v / v) cosolvent concentration (e.g., up to 40% (v / v) isopropanol); v) a temperature of about 15-75°C, e.g., 20-55°C, e.g., 20-45°C, e.g., 40°C; vi) a pH of 5.0-10.0, e.g., 7.5; and vii) a reaction time of up to 24 hours.
[0333] The combination of reaction parameters can also affect reaction time. In some embodiments, at least about 90% of the substrate is reduced to product in less than about 20 hours. In some embodiments, at least about 95% of the substrate is reduced to product in less than about 24 hours.
[0334] The method for carrying out an enzymatic reaction may further include a step of isolating the product of the enzymatic reaction. In particular, this step is carried out after the enzymatic reaction is completed. The product is particularly separated from one or more components of the reaction mixture, particularly from substantially all other components. For example, the product is separated from remaining substrates, by-products, enzymes, and / or organic solvents. Isolation of the product can be achieved by means and techniques known in the art, including, for example, evaporation of the solvent, condensation or crystallization, as well as filtration, phase separation, chromatographic separation, etc.
[0335] The present disclosure also provides a compound of formula (IC): [ka] 1. A method for synthesizing 6-((S)-2-((3aR,5R,6aS)-5-(2-fluorophenoxy)hexahydrocyclopenta[c]pyrrol-2(1H)-yl)-1-hydroxyethyl)pyridin-3-ol in free form or in a pharmaceutically acceptable salt form, comprising: contacting the compound of formula (IA)-I with a KRED polypeptide according to the present disclosure to obtain a compound of formula (IB)-I. [ka] (In the formula, R 1a is an amine protecting group. 1a is selected from tert-butyloxycarbonyl (Boc) and N-carboxybenzyl (Cbz).
[0336] In certain embodiments, the process includes contacting Compound 1 with a KRED polypeptide according to the disclosure to produce tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate: [ka] The method includes the step of obtaining:
[0337] In a further aspect, the formula [ka] or a salt thereof.
[0338] In a further aspect, in the preparation of 6-((S)-2-((3aR,5R,6aS)-5-(2-fluorophenoxy)hexahydrocyclopenta[c]pyrrol-2(1H)-yl)-1-hydroxyethyl)pyridin-3-ol in free form or in pharmaceutically acceptable salt form, [ka] or a salt thereof.
[0339] Compound (IC) can be synthesized from tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (or (5s)-2) by synthetic procedures known in the art. In particular, compound (IC) can be synthesized from tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate according to the methods and procedures disclosed in WO 2016 / 049165 A1.
[0340] Compositions, Kits, and Administration The present disclosure provides compositions comprising any of the modified KRED polypeptides, polynucleotides, expression vectors, and / or host cells described herein. Pharmaceutical compositions described herein may be prepared by any method known in the art of pharmacology. Generally, such preparative methods include the step of bringing into association the "active ingredient" (e.g., modified KRED polypeptide) with the carrier and / or one or more other accessory ingredients. Compositions may be prepared, packaged, and / or sold in bulk. The relative amounts of the active ingredient and / or any additional ingredients in compositions of the invention will vary depending on use.
[0341] It will also be understood that any of the compositions described herein can include one or more additional agents. In some embodiments, the composition includes a substrate compound of the structural formula (e.g., tert-butyl rel-(3aR,6as)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (Compound 1)) and / or a product compound (e.g., (5s)-2). For example, the composition can include the structural formula of Compound 1 and / or a compound of (5s)-2. In some embodiments, the composition can include a cofactor such as NAD(P)H.
[0342] Kits are also encompassed by this disclosure. Kits of the invention can include any of the modified KRED polypeptides, polynucleotides, expression vectors, and / or host cells described herein, or compositions thereof. In some embodiments, the kit includes a modified KRED polypeptide, a substrate (e.g., tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (Compound 1)), and a cofactor. In some embodiments, the cofactor is nicotinamide adenine dinucleotide (NADH) or nicotinamide adenine dinucleotide phosphate (NADPH). In some embodiments, the kit includes an immobilized KRED polypeptide.
[0343] The kits may be useful for practicing any of the methods described herein, such as producing and / or using any of the modified KRED polypeptides, polynucleotides, expression vectors, and / or host cells described herein, or compositions thereof. In some embodiments, the kits are used in methods for reducing tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product). In some embodiments, the kits are used in methods for reversing the diastereoselectivity of a ketoreductase polypeptide. In some embodiments, the kits are used in methods for increasing the diastereoselectivity of a ketoreductase polypeptide.
[0344] The kits provided herein can include one or more containers (e.g., vials, ampoules, bottles, syringes, and / or dispenser packages, or other suitable containers). The kits provided herein can also include written instructions for making or using any of the modified KRED polypeptides, polynucleotides, expression vectors, and / or host cells, or compositions thereof, described herein.
[0345] The practice of the present invention employs, unless otherwise indicated, conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology, which are well within the skill of the art. Such techniques are fully explained in the literature, e.g., "Molecular Cloning: A Laboratory Manual," second edition (Sambrook, 1989); "Oligonucleotide Synthesis" (Gait, 1984); "Animal Cell Culture" (Freshney, 1987); "Methods in Enzymology" and "Handbook of Experimental Immunology" (Weir, 1996); "Gene Transfer Vectors for Mammalian Cells" (Miller and Calos, 1987); "Current Protocols in Molecular Biology" (Ausubel, 1987); "PCR: The Polymerase Chain Reaction" (Mullis, 1994); and "Current Protocols in Immunology" (Coligan, 1991). These techniques are applicable to the production of the polynucleotides and polypeptides of the invention and therefore may be considered in making and practicing the invention. Techniques that are particularly useful for certain embodiments are discussed in the following sections.
[0346] The following examples are put forward so as to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the assay, screening, and treatment methods of the present invention, and are not intended to limit the scope of what the inventors regard as their invention. [Example]
[0347] Example 1: Description of expression strains A polypeptide from a previous ketoreductase (KRED) panel (SEQ ID NO: 2) was selected to have suitable initial activity / selectivity for the formation of compound (5s)-2, and variants thereof were used as the "scaffold" for the first round of evolution. A second polypeptide from the previous ketoreductase panel (SEQ ID NO: 54) was found to be more suitable for evolution with respect to its selectivity for the formation of compound (5s)-2 at low conversion rates. This variant was used as the scaffold for the second round of evolution.
[0348] The parent genes of the KREDs (round 1 and round 2 backbones) used to generate the variants of the invention were codon-optimized for expression in E. coli and cloned into the vector pCKI 10900 (shown as Figure 3 in U.S. Patent Application Publication No. 2006 / 0195947, which is incorporated herein by reference). The expression vector also contained a P15a origin of replication and a chloramphenicol (CAM) resistance 20 gene. The resulting plasmids were transformed into E. coli W3110 (fhu-) using standard methods.
[0349] Multiple rounds of directed evolution of the KRED gene in the pCKI 10900 plasmid were performed, using the gene encoding the most improved polypeptide from each round as the parent backbone sequence for subsequent evolutionary rounds. The resulting exemplary engineered ketoreductase polypeptide sequences, as well as the specific mutations and relative activities of the invention, are listed in Tables 4, 5, 6, and 8.
[0350] Example 2: Preparation of cell pellets (from growth to lysis) E. coli W3110 (fhu-) cells were transformed with the pCKI 10900 plasmid containing the KRED-encoding gene, plated on LB agar plates containing 1% glucose and 30 μg / mL CAM, and grown overnight at 37°C. Monoclonal colonies were picked and inoculated into 180 μL of LB containing 1% glucose and 30 μg / mL CAM in 96-well shallow-well microtiter plates. The plates were sealed with O2-permeable seals, and the cultures were grown overnight at 30°C, 200 rpm, and 85% RH. 10 μL of cell culture was then transferred to each well of a 96-well deep-well plate containing 390 μL of Terrific Broth (TB) and 30 μg / mL CAM. The deep-well plates were sealed with O2-permeable seals, and the optical density at 600 nm (OD 600 The cells were incubated at 30°C, 250 rpm, and 85% humidity until the pH reached 0.6-0.8. Expression of the KRED gene was induced by adding isopropyl-β-D-thiogalactoside (IPTG) to a final concentration of 1 mM, and the cells were incubated overnight at 30°C, 250 rpm, and 85% humidity. The cells were then pelleted at 4000 rpm to a pH of 10. The supernatant was discarded, and the pellet was frozen at -80°C before lysis.
[0351] Example 3: Lysis and clarification Lysate preparation Frozen pellets prepared as described in Example 2 were lysed in 200 μL of lysis buffer containing 100 mM TEoA (triethanolamine chloride) or NaPi (sodium phosphate) buffer, pH 7.0, 1 mg / mL lysozyme, 0.5 mg / mL PMBS, 0.2 U / mL DNAse, and 10 mM MgSO. The resulting lysis mixture was stirred at room temperature for 2.5 hours. The plate was then centrifuged at 4000 rpm and 4°C for 10 minutes. The supernatant was then used as the clarified lysate in biocatalysis reactions in the experiments described below to determine activity levels.
[0352] Example 4: Preparation of shake flask powder A shake flask procedure was used to generate the modified polypeptide powder used in the high-throughput activity assay. A single microbial colony of E. coli containing a plasmid with the KRED gene of interest was used to inoculate 50 mL of Luria-Bertani (LB) broth containing 30 μg / mL chloramphenicol and 1% glucose. Cells were grown overnight (at least 16 hours) in a 30°C incubator with shaking at 250 rpm. The culture was diluted into 250 mL of TB in a 1 L flask and incubated at an OD of 0.2. 600 The culture was grown at 30°C and 250 rpm until the OD 600 When the RI was 0.6-0.8, expression of the ketoreductase gene was induced with 1 mM IPTG, and incubation was continued overnight (at least 16 hours). Cells were harvested by centrifugation (5000 rpm, 15 minutes, 4°C), and the supernatant was discarded. Cells were resuspended in 50 mL of cold (4°C) 100 mM triethanolamine (chloride) buffer, pH 7.0, and harvested by centrifugation as described above. The washed cells were resuspended in 30 mL of cold triethanolamine (chloride) buffer and passed through a microfluidizer. Cell debris was removed by centrifugation (9000 rpm, 45 minutes, 4°C). The cleared lysate supernatant was collected and stored at -20°C. The frozen cleared lysate was lyophilized to obtain a dry powder of crude KRED enzyme.
[0353] Example 5: Analytical methods for activity and selectivity evaluation An HPLC method with UV detection was developed to analyze the conversion of compound 1 to (5s)-2 and (5r)-2 (see Figure 1). The conversion is expressed as a percentage and is calculated as follows:
number
[0354] The product compounds (5s)-2 and (5r)-2 were diastereomers and, as a result, were well separated by the analytical methods described. The diastereomeric excess (de), expressed as a percentage, was calculated as follows:
number
[0355] [Table 4]
[0356] Example 6: Evaluation of a series of KREDs for the reduction of Compound 1 Each well contained 40 mM triethanolamine (TEoA) buffer, pH 7.0, 5 g / L of Compound 1, and 1 g / L of NADP. + A series of ketoreductases (KREDs) were screened using a 250 μL reaction volume in a deep-well Costar plate containing 20% (v / v) isopropanol, and 40% (v / v) enzyme lysate. The sealed plate was incubated at room temperature with vigorous shaking at 850 rpm in an Infors incubator. After 24 h of incubation, the reaction was quenched by adding 750 μL of acetonitrile. The plate was resealed and vigorously shaken at 800 rpm on a plate shaker at room temperature for 10 min, followed by centrifugation at 4000 g for 10 min. 200 μL of the supernatant was transferred to a microtiter plate (V-bottom Greiner plate), sealed with aluminum foil, and subjected to HPLC analysis using the method described in Example 5. The KRED having SEQ ID NO:2 (nucleic acid sequence of SEQ ID NO:1) was identified as the best-performing enzyme for producing compound (5s)-2.
[0357] Example 7: Improvement of KRED on SEQ ID NO: 2 for the diastereoselective generation of compound (5s)-2 (Round 1) The modified KRED, SEQ ID NO:2 (nucleic acid sequence, SEQ ID NO:1), was selected as the parent enzyme for the first round of directed evolution. A library of modified genes was generated using well-established techniques (e.g., saturation mutagenesis and recombination of potentially beneficial mutations identified in variants screened in an enzyme panel). The enzymes encoded by each gene were produced in an HTP as described in Example 2, and clarified lysates were used to generate clarified lysates as described in Example 3.
[0358] Each 200 μL reaction was prepared in a 96-well deep-well format (2 mL volume) containing 0.625–2.5% (v / v) clarified lysate, 10 g / L Compound 1, 20% (v / v) isopropanol, 40 mM TEoA buffer (pH 7.0), 0.1 g / L NADP + The plates were sealed and agitated at 30°C on an Infors shaker at 600 rpm for 24 hours.
[0359] The reaction was quenched by adding 800 μL of acetonitrile to each well of the reaction plate, which was then sealed and shaken for 15 minutes. The plate was then centrifuged at 4000 rpm for 10 minutes, and 200 μL of the supernatant was transferred to an analytical plate and subjected to achiral HPLC analysis using the method described in Example 5.
[0360] The activity of each mutant was calculated as the percent conversion of the formed product based on peak area [((5s)-2 + (5r)-2) / (compound 1 + (5s)-2 + (5r)-2) * 100]. The selectivity of each mutant was calculated as the diastereomeric excess (de) of the formed desired product (5s)-2 based on peak area [((5s)-2 - (5r)-2) / ((5s)-2 + (5r)2) * 100]. Furthermore, the amount of the desired product (5s)-2 was reported as the HPLC peak area, since this is the amount obtained when conversion and selectivity are convolved. For all reported values, the percent improvement relative to SEQ ID NO:2 was calculated. Because the parent SEQ ID NO:2 had a negative %de% value, the reciprocal of the percent improvement was reported for clarity in order to report the improvement in enzyme activity toward the desired product (5s)-2 as a positive value greater than 1. The results are shown in Table 4.
[0361] [Table 5]
[0362] [Table 6]
[0363] The most potent variants were prepared as shake flask powders and evaluated under two conditions: 10 g / L of Compound 1 containing 20% isopropanol (v / v) or 20 g / L of Compound 1 containing 40% isopropanol (v / v). Reactions were prepared in a volume of 200 μL, and both conditions contained 1 g / L of NADP +and 0.1 mM TeoA buffer (pH 7.0). Two-fold serial dilutions of the enzyme shake flask powder, with the highest final concentration of 4 g / L, were added to the reaction mixture. The reaction was incubated at room temperature with shaking at 600 rpm for 24 hours, then quenched with acetonitrile and centrifuged at 4000 rpm for 10 minutes. 100 μL of the supernatant was transferred to a Greiner microtiter plate and subjected to HPLC analysis. Conversion and diastereomeric excess (de%) were calculated as described above. The de% dose curve showed increased formation of the desired product with increasing enzyme concentration, suggesting kinetic resolution of compound (5s)-2.
[0364] Example 8: Evaluation of a series of KREDs for the reduction of Compound 1 in low concentration lysates A series of ketoreductases (KREDs) were screened to identify exemplary enzymes as in Example 6. Screening was performed in a 250 μL reaction volume in a Costar plate, with each well containing 40 mM triethanolamine (TEoA) buffer (pH 7.0), 5 g / L Compound 1, 1 g / L NADP + The enzyme lysate was diluted with 20% isopropanol (v / v), and 40% (v / v) enzyme lysate, 5% (v / v) lysate, or 0.625% (v / v) lysate. The sealed plates were incubated at room temperature with vigorous shaking at 850 rpm in an Infors incubator. After 24 h of incubation, the reaction was quenched by adding 750 μL of acetonitrile. The plates were resealed and incubated at room temperature for 10 min with vigorous shaking, followed by centrifugation at 3220 g for 2 min. 200 μL of the supernatant was transferred to a microtiter plate (V-bottom Greiner plate), sealed with aluminum foil, and subjected to HPLC analysis using the method described in Example 5. Conversion and diastereomeric excess were calculated at all lysate concentrations tested, and a constant de% was used as the primary selection criterion to identify the optimal enzyme. KRED having SEQ ID NO: 54 (nucleic acid sequence SEQ ID NO: 53) was identified as the enzyme with the highest potency for producing compound (5s)-2 at various lysate concentrations.
[0365] Example 9: Improvement of KRED to SEQ ID NO: 54 for the diastereoselective generation of compound (5s)-2 (Round 2) SEQ ID NO:54 (nucleic acid sequence, SEQ ID NO:53) was selected as the parent enzyme for the second round of directed evolution. A library of modified genes was generated using well-established techniques (e.g., saturation mutagenesis and recombination of mutations identified as potentially beneficial in previous rounds of evolution). The enzymes encoded by each gene were produced in an HTP as described in Example 2, and clarified lysates were used to generate clarified lysates as described in Example 3.
[0366] Each 200 μL reaction was prepared in a 96-well deep-well format (2 mL volume) containing 10% (v / v) clarified lysate, 20 g / L Compound 1, 40% (v / v) isopropanol, 40 mM TEoA buffer (pH 7.5), 1 g / L NADP + The plates were sealed and agitated at 30°C on an Infors shaker at 600 rpm for 24 hours.
[0367] The reaction was quenched by adding 800 μL of acetonitrile to each well of the reaction plate, which was then sealed and shaken for 15 minutes. The plate was then centrifuged at 4000 rpm for 10 minutes, and 200 μL of the supernatant was transferred to an analytical plate and subjected to achiral HPLC analysis using the method described in Example 5. Conversion and diastereomeric excess (de%) were calculated based on peak areas as described above. For all reported values, the improvement relative to SEQ ID NO: 54 was calculated. The results are shown in Table 5.
[0368] [Table 7]
[0369] [Table 8]
[0370] [Table 9]
[0371] The most potent variants were prepared as shake flask powders and evaluated under two conditions: 20 g / L Compound 1 or 50 g / L Compound 1 containing 40% isopropanol (v / v). Reactions were prepared in a volume of 200 μL, and both conditions contained 1 g / L NADP + and 0.1 mM KPi buffer (pH 7.5). Two-fold serial dilutions of the enzyme shake flask powder, with the highest final concentration of 20 g / L, were added to the reaction mixture. The reaction was incubated at room temperature with shaking at 230 rpm for 24 hours, then quenched with acetonitrile and centrifuged at 4000 rpm for 10 minutes. 100 μL of the supernatant was transferred to a Greiner microtiter plate and subjected to HPLC analysis.
[0372] Example 10: Improvement of KRED to SEQ ID NO: 152 for the diastereoselective generation of compound (5s)-2 (Round 3) SEQ ID NO: 152 (nucleic acid sequence, SEQ ID NO: 151) was selected as the parent enzyme for the third round of directed evolution. A library of modified genes was generated using well-established techniques (e.g., recombination of mutations identified as potentially beneficial in previous evolutionary rounds). The enzymes encoded by each gene were produced by HTP as described in Example 2, and clarified lysates were produced as described in Example 3.
[0373] Each 100 μL reaction was prepared in a 96-well deep-well format (2 mL volume) containing 10% (v / v) clarified lysate, 150 g / L Compound 1, 40% (v / v) isopropanol, 40 mM NaPi buffer (pH 7.5), 0.1 g / L NADP + The plates were sealed and agitated at 40°C on an Infors shaker at 600 rpm for 24 hours.
[0374] The reaction was quenched by adding 900 μL of acetonitrile to each well of the reaction plate, which was then sealed and shaken for 15 minutes. The plate was then centrifuged at 4000 rpm for 10 minutes, and 70 μL of the supernatant was transferred to an analytical plate containing 140 μL of acetonitrile and subjected to achiral HPLC analysis using the method described in Example 5. The conversion and diastereomeric excess (de%) were calculated based on the peak areas as described above. The improvement in the conversion percentage relative to SEQ ID NO: 152 was calculated. The results are shown in Table 6.
[0375] [Table 10]
[0376] [Table 11]
[0377] [Table 12]
[0378] [Table 13]
[0379] Example 11: Improvement of KRED to SEQ ID NO: 256 for the diastereoselective generation of compound (5s)-2 (Round 4) After three rounds of directed evolution, a KRED enzyme was identified with increased selectivity and greater than 95% activity (% conversion) for the (5s)-2 product (Table 8). However, a fourth round of directed evolution was performed to further improve the diastereoselective production of the (5s)-2 product by improving the activity of the enzyme and decreasing the enzyme concentration to achieve the same level of conversion and selectivity.
[0380] SEQ ID NO:256 (nucleic acid sequence, SEQ ID NO:255) was selected as the parent enzyme for the fourth round of directed evolution. A library of modified genes was generated using well-established techniques (e.g., saturation mutagenesis and recombination of potentially beneficial mutations identified in previous rounds of evolution). The enzymes encoded by each gene were produced in an HTP as described in Example 2, and clarified lysates were used to generate clarified lysates as described in Example 3.
[0381] Each 200 μL reaction was prepared in a 96-well deep-well format (2 mL volume) containing 2.5% (v / v) clarified lysate, 150 g / L Compound 1, 40% (v / v) isopropanol, 40 mM NaPi buffer (pH 7.5), 0.05 g / L NADP + The plates were sealed and agitated at 30°C on an Infors shaker at 600 rpm for 24 hours.
[0382] The reaction was quenched by adding 800 μL of acetonitrile to each well of the reaction plate, sealed, and shaken for 15 minutes. The plate was then centrifuged at 4000 rpm for 10 minutes, and 200 μL of the supernatant was transferred to an analytical plate and subjected to achiral HPLC analysis using the method described in Example 5. The stereomeric excess (de%) of the desired product was calculated based on the peak area as described above. The percent conversion was calculated as an improvement relative to SEQ ID NO: 256. The results are shown in Table 7.
[0383] [Table 14]
[0384] [Table 15]
[0385] [Table 16]
[0386] After the fourth round of evolution, KRED enzymes with increased selectivity and greater than 95% activity (% conversion) for the (5s)-2 product were identified (Table 8).
[0387] [Table 17]
[0388] Example 12: Evaluation of KRED activity and selectivity The activity and selectivity of the modified KRED enzymes SEQ ID NO:152, SEQ ID NO:256, and SEQ ID NO:426 were evaluated against the commercially available KRED enzyme (CM; Johnson Matthey (ADH-152)), the parent KRED enzyme (SEQ ID NO:54), and the wild-type (WT; SEQ ID NO:492) by performing enzyme dose curves in a 96-well plate format under several test conditions (Figures 2-5).
[0389] Two test conditions were used to measure activity (conversion %) and selectivity (de%). The parameters were selected so that the enzyme could demonstrate maximum performance. The enzyme was evaluated based on the maximum conversion and selectivity achieved, regardless of the evaluation conditions. Specific evaluation parameters are shown in Table 9.
[0390] [Table 18]
[0391] Under conditions 1 and 2, the modified KRED enzymes SEQ ID NO:152, SEQ ID NO:256, and SEQ ID NO:426 exhibited better activity (Figures 2A, 3A, 4A) and selectivity (Figures 2B, 3B, 5A) compared to the commercially available enzyme, the parent KRED, and the wild-type. The exemplary enzyme SEQ ID NO:426 exhibited the highest activity (Figures 2A, 3A, 4A) and selectivity (Figures 2B, 3B, 5A). Improved performance of SEQ ID NO:426 was also observed under screening conditions with high substrate concentrations (Figures 4A-4C, 5A-5C). Specifically, the exemplary enzyme SEQ ID NO:426 achieved 97% conversion and 99.6% DE at 50 g / L substrate and 5% w / w enzyme loading, and 93% conversion and 98.8% DE at 100 g / L substrate and 1% w / w enzyme loading. The commercial enzyme was inactive under condition 2 without the cofactor recycling enzyme GDH / glucose (Figures 4A and 5A), indicating that the cofactor recycling enzyme is a prerequisite for activity. The WT control showed no activity under either condition.
[0392] Example 13: Immobilization of KREDs Preparation of immobilization buffer To prepare a phosphate buffer solution (0.05 M pH = 7.5) containing magnesium sulfate (0.002 M), NaHPO 12H O (230.18 g), NaHPO 2H O (16.58 g), and MgSO (3.61 g) were dissolved in water (15 L). 12.2 L of this solution was used for the immobilization process. The buffer concentration can range from 0.01 to 0.10 M. Sodium phosphate salts may also be substituted with potassium phosphate salts.
[0393] Preparation of a 0.5% solution of glutaraldehyde in phosphate buffer To prepare a 0.5% glutaraldehyde solution in phosphate buffer, 25% glutaraldehyde in water (24.5 mL) was mixed with 0.05 M pH 7.5 phosphate buffer (1200 mL). A total of 1160 mL of this solution was used in the immobilization process. The concentration of glutaraldehyde can range from 0.2 to 2. As an alternative to glutaraldehyde, other bifunctional linkers may be used to attach enzyme residues to the resin.
[0394] Preactivation of enzyme carriers The immobilization support was selected after several screenings of amino- and epoxy-functionalized resins from commercial suppliers with various particle sizes, pore diameters, and hydrophobicities.
[0395] The amino-functionalized resin Lifetech ECR8304F (325 g) was washed three times with 580 mL of immobilization buffer and incubated with 1160 mL of 0.5% glutaraldehyde solution at room temperature for 2 hours. The preactivated resin was then washed four times with 1160 mL of immobilization buffer.
[0396] Immobilization of KRED To the preactivated resin, a solution of KRED (16 g) in immobilization buffer (1160 mL) was added, and the mixture was incubated at room temperature for 18 h. The immobilized enzyme was washed with immobilization buffer (1160 mL), 0.5 M NaCl solution (2 × 1160 mL), and immobilization buffer (2 × 1160 mL). This procedure yielded approximately 400 g of wet immobilized KRED, representing an immobilization yield of 68-69% and 66% recovery of enzyme activity.
[0397] The enzyme / resin ratio can range from 10 to 100 mg enzyme / g resin. The enzyme / resin ratio (approximately 50 mg enzyme / g resin) was selected to maximize the recovery of enzyme activity after the immobilization process and the specific activity of the final immobilized product. This can be further optimized; lower enzyme / resin ratios may result in higher activity recovery while having less impact on the specific activity of the final immobilized KRED. A wider range of temperatures and time periods can be included. The volume of wash solution and number of rinses were adapted to the procedure proposed by the resin supplier but can be further optimized. The addition of the cofactor NADP during incubation was evaluated in the DoE as a measure to increase enzyme activity retention and can be used as an optional additive.
[0398] Biocatalytic reactions To prepare a phosphate buffer solution containing magnesium sulfate (0.1 M pH = 7.5), NaHPO·12H2O (152 g), NaHPO·2H2O (11.7 g), and magnesium sulfate (6.3 g) were dissolved in water (5 L, of which 2 kg was used in the reaction). The pH was adjusted to 7.4-7.6 using 3 N NaOH or 29% phosphoric acid. Buffer concentrations ranging from 0.01 to 1.00 M can be used.
[0399] The biocatalysis procedure was performed three times, with the recovered immobilized enzyme used in the second and third batches. Substrate (1.00 kg) was dissolved in iPrOH (1.95 kg). Water (2.5 kg) and phosphate buffer (1.9 kg) were added, and the pH was controlled and adjusted as needed (pH range = 7.5-8.3). NADPNa (2 g) was added as a solution in phosphate buffer (100 g). Immobilized KRED (380 g) was added as a slurry in a 1:1 (v / v) mixture of water and isopropanol (875 mL:875 mL). The enzyme vessel was then rinsed with a 1:1 (v / v) mixture of water and isopropanol (625 mL:625 mL) and placed in the reaction vessel. The pH was controlled and adjusted as needed (target pH range = 7.5-8.3). The mixture was warmed to 35°C and stirred for 20 h. The azeotrope of acetone, isopropanol, and water was distilled off (2 L) and replaced with a 1:1 (w / w) mixture of water and isopropanol (1 kg:1 kg). The mixture was stirred for an additional 10 h at IT = 35 °C. The temperature range, reaction time, and NADP loading amount can be further optimized.
[0400] Workup / Isolation: The reaction mixture was filtered through a Pall K900 depth filter, and the immobilized enzyme was washed with a 1:1 (v / v) mixture of water and isopropanol (3 x 1 L:1 L). The filtrates were combined, and isopropanol (10 L) was evaporated under reduced pressure at JT = 50 °C, while a portion of the distilled volume was replaced by the addition of water (5 kg). The precipitated product was stirred at 25 °C for 18 h.
[0401] The product was isolated by filtration. The filter cake was washed with water (2 kg). The wet product g was dried at 60°C and full vacuum to give the product as a white solid (843-895 g, 84-89% yield, range of 3 batches).
[0402] By using immobilized enzyme, clear filtration on Cellflock was avoided. This significantly reduced filtration time (from 83 minutes for a 2.5 kg batch using free enzyme to 3-9 minutes for three 1 kg batches using immobilized KRED). All three batches met specifications, with yield and purity values comparable to the non-immobilized enzyme process. Using a total of three batches of enzyme, the amount of KRED was reduced from 1% (w / w) to 0.5% (w / w). The amount of distilled and added water, if modified, depended on the starting concentration.
[0403] Enzyme storage between batches The immobilized enzyme was stored at 2–8 °C as a slurry in a 1:1 (v / v) mixture of water and isopropanol (1 L:1 L) and used in a total of three batches, with a maximum storage time of 5 weeks between two batches. The amount of 1:1 water:isopropanol solution added to the reaction vessel in the second and third batches was adapted based on the volume of the stored enzyme slurry.
[0404] To preserve the immobilized enzyme between batches, various stabilization solutions were screened over a one-month period. Ultimately, 1:1 isopropanol / water was selected because the immobilized enzyme retained 100% of its initial activity after one month of storage in this solution. The use of 1:1 isopropanol / water eliminated the need for rinsing the immobilized KRED enzyme between batches. Furthermore, the use of 1:1 isopropanol / water prevented microbial growth during storage. SEM images (not shown) of immobilized KRED beads stored in aqueous buffer for several months revealed bacterial growth on the surface, whereas no colonies were observed on immobilized KRED beads stored in 1:1 isopropanol / water.
[0405] Immobilization parameters The following parameters were investigated: Enzyme carrier: Amino-functionalized, different pore sizes, different providers: Relizyme EA113, Purolite Lifetech ECR8309F, and ECR8304F Epoxy functionalization, hydrophobicity, and pore size differences: Purolite Lifetech ECR8204F and ECR8285 The Lifetech ECR8304 was selected to test the following parameters: Enzyme / carrier ratio (mg enzyme / carrier): 10-100mg / g ·Enzyme solution concentration: 2.5~50mg / mL Glutaraldehyde percentage used as a linker between the enzyme and the amino-functionalized carrier: 0.2-2% NADP sodium salt as an additive to increase enzyme activity retention: 0-3 mg / mL (in enzyme solution)
[0406] equivalent The disclosures of all patents, patent applications, and publications cited herein are incorporated herein by reference in their entirety. While the present invention has been disclosed with reference to certain specific embodiments, it will be apparent that further embodiments and variations of the present invention may be devised by those skilled in the art without departing from the true spirit and scope of the present invention. It is intended that the appended claims be construed to include all such embodiments and equivalent variations.
Claims
1. 1. An engineered ketoreductase polypeptide, said polypeptide comprising: a) a sequence encoding an amino acid sequence selected from the group consisting of SEQ ID NOs: 392, 394, 396, 406, 408, 420, 422, 424, 426, 428, 430, 436, 438, 462, 464, 468, 470, 474, 476, 480, 482, and 490, with at least about 97.7%, 97.8%, 97.9%, 98.0%, 99.1%, 100.0%, 101.0%, 102.0%, 103.0%, 104.0%, 105.0%, 106.0%, 107.0%, 108.0%, 109.0%, 1109.0%, 1111.0%, 112.0%, 113.0%, 114.0%, 115.0%, 116.0%, 117.0%, 118.0%, 119.0%, 120.0%, 121.0%, 122.0%, 123.0%, 124.0%, 125.0%, 126.0%, 127.0%, 128.0%, 129.0%, 130.0%, 131.0%, 132.0%, 133.0%, 134.0%, 135.0%, 136.0%, 137.0%, 138.0%, 139.0%, 140.0%, 141.0%, 142.0%, 143.0%, 144.0%, 145.0%, 146.0%, 147.0%, 148.0%, a polypeptide comprising an amino acid sequence having 8.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or 100% sequence identity; b) (1) an amino acid sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 54, 152, 256 or 492; and (2) with respect to said amino acid sequence, i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; v) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; viii) X196G; and ix) X198V, X198I, X198Y, or X198P and containing one or more amino acid differences selected from Optionally, the amino acid sequence includes one or more additional amino acid residue differences. a polypeptide comprising the amino acid sequence; or c) a polypeptide comprising an amino acid sequence selected from Table 4 (SEQ ID NOS: 4-52), Table 5 (SEQ ID NOS: 56-186), Table 6 (SEQ ID NOS: 188-390), or Table 7 (SEQ ID NOS: 392-490). The engineered ketoreductase polypeptide is selected from any of the following:
2. 2. The modified polypeptide of claim 1, wherein the polypeptide is capable of selectively reducing tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product).
3. 1. An engineered ketoreductase polypeptide capable of selectively reducing tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product), said polypeptide comprising: (i) a polypeptide selected from the group consisting of SEQ ID NOs: 54, 152; , 256, or 492; and (ii) a substitution, deletion, addition, or insertion of one or more amino acid residues selected from Table 2, Table 4, Table 5, Table 6, or Table 7 relative to said amino acid sequence, and optionally, one or more additional amino acid residue differences relative to said amino acid sequence.
4. 4. The modified polypeptide of any one of claims 1 to 3, wherein the amino acid sequence of the modified polypeptide comprises one or more amino acid differences at positions X17, X18, X40, X56, X94, X96, X106, X145, X173, X190, X194, X196, X198, X202, and / or X206 relative to the amino acid sequence, and optionally one or more additional amino acid residue differences relative to the amino acid sequence.
5. 5. The modified polypeptide of any one of claims 1 to 4, wherein the amino acid sequence of the modified polypeptide comprises, relative to the amino acid sequence: (i) a residue at position X190 that is not tyrosine; (ii) a residue at position X190 that is a non-aromatic residue; (iii) a residue at position X196 that is an aliphatic residue or a small amino acid residue; (iv) a residue at position X202 that is a small amino acid residue; and / or (v) a residue at position X206 that is not methionine, and optionally comprises one or more additional amino acid residue differences relative to the amino acid sequence.
6. The amino acid sequence of the modified polypeptide may have any of the following amino acid residues added to the amino acid sequence: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X40R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X106L, X106W, X106E, X145S, X145D, X145E, X173A, X173L, X173V, X173W, X1 6. The modified polypeptide of any one of claims 1 to 5, comprising one or more of the following amino acid residue differences relative to the amino acid sequence: X73M, X173K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W, and optionally one or more additional amino acid residue differences relative to the amino acid sequence.
7. 7. The modified polypeptide of any one of claims 1 to 6, wherein the amino acid sequence of the modified polypeptide comprises one or more amino acid residue differences selected from X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W relative to the amino acid sequence, and optionally one or more additional amino acid residue differences relative to the amino acid sequence.
8. 8. The modified polypeptide of any one of claims 1 to 7, wherein the amino acid sequence of the modified polypeptide comprises one or more amino acid residue differences selected from X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A and X206Q relative to the amino acid sequence, and optionally one or more additional amino acid residue differences relative to the amino acid sequence.
9. 9. The modified polypeptide of any one of claims 1 to 8, wherein the amino acid sequence of the modified polypeptide comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 392, 394, 396, 406, 408, 420, 422, 424, 426, 428, 430, 436, 438, 462, 464, 468, 470, 474, 476, 480, 482, and 490.
10. The modified polypeptide of any one of claims 1 to 9, wherein the modified polypeptide is solvent-stable.
11. 11. The modified polypeptide of any one of claims 2 to 10, wherein the modified polypeptide reduces a substrate to a product with a conversion of at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99.0%, or 100%.
12. 12. The modified polypeptide of any one of claims 2 to 11, wherein the modified polypeptide reduces a substrate to a product with a selectivity level (% diastereomeric excess) of at least about 96.0%, at least about 96.5%, at least about 97.0%, at least about 97.5%, at least about 98.0%, at least about 98.5%, at least about 99.0%, at least about 99.5%, or 100%.
13. 13. The modified polypeptide of any one of claims 2 to 12, wherein the modified polypeptide does not require glucose dehydrogenase (GDH) / glucose cofactor recycling to reduce a substrate to a product.
14. The modified polypeptide of any one of claims 2 to 13, wherein the ability to reduce a substrate to a product is compared to a reference polypeptide.
15. 15. The modified polypeptide of claim 14, wherein the modified polypeptide has a reversed or increased diastereoselectivity for the reduction of a substrate to a product compared to the reference polypeptide.
16. 16. The modified polypeptide of claim 14 or 15, wherein the modified polypeptide has an increased level of activity (e.g., percent conversion) compared to the reference polypeptide and has a percent improvement over positive control (FIOP) of greater than about 1.75, greater than about 2.00, greater than about 2.25, greater than about 2.50, greater than about 2.75, greater than about 3.00, greater than about 3.25, greater than about 3.50, greater than about 3.75, greater than about 4.00, greater than about 4.25, greater than about 4.50, greater than about 4.75, greater than about 5.00, greater than about 5.25, greater than about 5.50, greater than about 5.75, greater than about 6.00, greater than about 6.25, or greater than about 6.
50.
17. 17. The modified polypeptide of any one of claims 14 to 16, wherein the modified polypeptide is capable of reducing a substrate to a product with an FIOP conversion of greater than about 2.25, preferably greater than about 3.00, and a diastereoselectivity of greater than about 95%, preferably greater than about 97%, or more preferably greater than about 99%, compared to the reference polypeptide.
18. 18. The modified polypeptide of any one of claims 14 to 17, wherein the reference polypeptide is a ketoreductase peptide of wild-type Lactobacillus kefir, Lactobacillus brevis, or Lactobacillus minor.
19. 19. The modified polypeptide of any one of claims 14 to 18, wherein the reference polypeptide is a wild-type Lactobacillus kefir ketoreductase polypeptide.
20. 20. The modified polypeptide of any one of claims 14-19, wherein the reference polypeptide is a wild-type Lactobacillus kefir ketoreductase polypeptide comprising the amino acid sequence of SEQ ID NO:
492.
21. The modified polypeptide of any one of claims 14 to 17, wherein the reference polypeptide is an engineered ketoreductase polypeptide.
22. 22. The engineered polypeptide of any one of claims 14 to 17 and 21, wherein the reference polypeptide is an engineered ketoreductase polypeptide selected from SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:
256.
23. 23. The modified polypeptide of any one of claims 14 to 17, 21 and 22, wherein the modified polypeptide i) requires fewer cofactors; ii) does not require glucose dehydrogenase (GDH) / glucose cofactor recycling; and / or iii) does not require dimethyl sulfoxide (DMSO) in a reduction reaction to reduce a substrate to a product, compared to the reference polypeptide.
24. An engineered ketoreductase polypeptide that, under suitable reaction conditions, can selectively reduce tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product) with greater stereoselectivity (% diastereomeric excess) and / or activity (% conversion) than SEQ ID NO:54, SEQ ID NO:152, and / or SEQ ID NO:
256.
25. 25. The modified polypeptide of claim 24, wherein the suitable reaction conditions comprise one or more of the following: i) up to about 150 g / L, such as 50 g / L or 100 g / L of substrate; ii) an enzyme load of less than about 5 wt%, such as less than about 3 wt%, for example less than about 1 wt%; iii) an NADP+cofactor load of less than about 2 wt%, such as less than about 1 wt%, for example less than about 0.2 wt%, for example 0.1%, 0.05% or 0.03%; iv) up to 40% (v / v) isopropanol; v) a temperature of about 15-75°C, such as 20-55°C, for example 20-45°C, for example 30°C or 40°C; vi) a pH of 5.0-10.0, such as a pH of 7.5-8.3; and vii) a reaction time of up to 30 hours, preferably 24 hours.
26. 26. The modified polypeptide of claim 24 or 25, wherein the suitable reaction conditions do not require dimethyl sulfoxide (DMSO).
27. 27. The modified polypeptide of any one of claims 24 to 26, wherein the suitable reaction conditions do not require recycling of the GDH / glucose cofactor.
28. 28. The modified polypeptide of any one of claims 24 to 27, comprising an amino acid sequence that differs from the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256 at one or more amino acid residues selected from Table 2, Table 4, Table 5, Table 6, or Table 7, and optionally contains one or more additional amino acid residue differences relative to said amino acid sequence.
29. 29. The modified polypeptide of any one of claims 24 to 28, wherein the modified polypeptide comprises an amino acid sequence that differs from the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256 at one or more amino acid residues selected from 17, 18, 40, 56, 94, 96, 106, 145, 173, 190, 194, 196, 198, 202, and / or 206, and optionally includes one or more additional amino acid residue differences relative to the amino acid sequence.
30. 30. The modified polypeptide of any one of claims 24-29, wherein the amino acid sequence of the modified polypeptide comprises, relative to the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256, (i) a residue at position X190 that is not tyrosine; (ii) a residue at position X190 that is a non-aromatic residue; (iii) a residue at position X196 that is an aliphatic residue or a small amino acid residue; (iv) a residue at position X202 that is a small amino acid residue; and / or (v) a residue at position X206 that is not methionine, and optionally comprises one or more additional amino acid residue differences relative to the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:
256.
31. The amino acid sequence of the modified polypeptide may be any of the following with respect to the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X40R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X106L, X106W, X106E, X145S, X145D, X145E, X173A, X173L, X173V, X173W, X17 31. The modified polypeptide of any one of claims 24 to 30, comprising one or more amino acid residues selected from: 3M, X173K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W, and optionally comprising one or more additional amino acid residue differences relative to said amino acid sequence.
32. The amino acid sequence of the modified polypeptide is, relative to the amino acid sequence of SEQ ID NO: 54, SEQ ID NO: 152, or SEQ ID NO: 256: i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; iv) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M or X173K; vii) X194A, X194V, X194W, or X194T; vii) X196G; and viii) X198V, X198I, X198Y, or X198P and comprising one or more amino acid residues selected from 32. The modified polypeptide of any one of claims 24 to 31, optionally comprising one or more additional amino acid residue differences relative to the amino acid sequence.
33. 33. The modified polypeptide of any one of claims 24 to 32, wherein the amino acid sequence of the modified polypeptide comprises one or more amino acid residues selected from X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W relative to the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256, and optionally comprises one or more additional amino acid residue differences relative to the amino acid sequence.
34. 34. The modified polypeptide of any one of claims 24 to 33, wherein the amino acid sequence of the modified polypeptide comprises one or more amino acid residues selected from X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q relative to the amino acid sequence of SEQ ID NO:54, SEQ ID NO:152, or SEQ ID NO:256, and optionally one or more additional amino acid residue differences relative to the amino acid sequence.
35. The modified polypeptide comprises: a) an amino acid sequence selected from Table 4 (SEQ ID NOs: 4-52), Table 5 (SEQ ID NOs: 56-186), Table 6 (SEQ ID NOs: 188-390), or Table 7 (SEQ ID NOs: 392-490), preferably an amino acid sequence selected from Table 7; or b) an amino acid sequence selected from the group consisting of SEQ ID NOs: 392, 394, 396, 406, 408, 420, 422, 424, 426, 428, 430, 436, 438, 462, 464, 468, 470, 474, 476, 480, 482, and 490, with at least about 97.7%, 97.8%, 97.9%, 98. an amino acid sequence having 0%, 98.1%, 98.2%, 98.3%, 98.4%, 98.5%, 98.6%, 98.7%, 98.8%, 98.9%, 99.0%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% sequence identity 35. The modified polypeptide of any one of claims 24 to 34, comprising an amino acid sequence selected from:
36. A method for the diastereoselective reduction of a bicyclic ketone, comprising contacting the bicyclic ketone with a ketoreductase (KRED) under suitable reaction conditions to provide a bicyclic secondary alcohol product.
37. 37. The method of claim 36, wherein the KRED is an engineered ketoreductase.
38. 36. A method for diastereoselectively reducing a bicyclic ketone substrate, the method comprising contacting the bicyclic ketone substrate with a modified polypeptide of any one of claims 1 to 35 under suitable reaction conditions to provide a bicyclic secondary alcohol product.
39. 39. The method of any one of claims 36 to 38, wherein the bicyclic ketone substrate has a total of 6 to 12 members.
40. 40. The method of any one of claims 36 to 39, wherein the bicyclic ketone substrate is achiral.
41. The bicyclic secondary alcohol product has formula (IB): 【Chemistry 1】 (In the formula, The A and B rings together form a fused cycloalkyl ring, e.g., C 6 ~C 12 represents a cycloalkyl or a fused heterocyclyl ring, for example a 6- to 12-membered heterocyclyl; wherein the fused cycloalkyl or heterocyclyl is at least one R 1 occurrences of, for example, 1 to 4 R 1 may be substituted with Each R 1 is an amine protecting group, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz), C 1 ~C 20 Alkyl, C 2 ~C 20 Alkenyl, C 2 ~C 20 Alkynyl, C 3 ~C 10 Cycloalkyl, C 6 ~C 14 Aryl, C 7 ~C 20 Arylalkyl, 3- to 14-membered heterocyclyl, 5- to 20-membered heteroaryl, hydroxyl, halogen, e.g., F, C 1 ~C 20 Haloalkyl, e.g., —CF 3 , -OR c , -NR c R c , -(CH 2 ) n COOR c , -(CH 2 ) n -C(=O)R c , and -(CH 2 ) n -C(=O)NR c R c are independently selected from wherein the alkyl, alkenyl, and alkynyl each represent one or more R a , for example, 1 to 6 R a may be substituted with The cycloalkyl, aryl, arylalkyl, heterocyclyl, and heteroaryl each may be one or more R b , for example, 1 to 6 R b optionally substituted with; Each R a are independently generated when they occur, 3 ~C 10 Cycloalkyl, C 6 ~C 14 Aryl, C 7 ~C 20 arylalkyl, 3- to 14-membered heterocyclyl, 5- to 20-membered heteroaryl, halogen, such as F, haloalkyl, such as C 1 ~C 20 Haloalkyl, e.g., —CF 3 , -OR c , -NR c R c , -(CH 2 ) n COOR c , -(CH 2 ) n -C(=O)R c , and -(CH 2 ) n -C(=O)NR c R c Selected from: Each R b each occurrence independently represents a halogen, C 1 ~C 20 Haloalkyl, e.g., —CF 3 , -OR c , -NR c R c , -(CH 2 ) n COOR c , -(CH 2 ) n -C(=O)R c , -(CH 2 ) n -C(=O)NR c R c , C 1 ~C 20 Alkyl, C 2 ~C 20 Alkenyl, and C 2 ~C 20 alkynyl; Each R c are independently H, C when they occur. 1 ~C 20 Alkyl, C 2 ~C 20 Alkenyl, and C 2 ~C 20 Alkynyl (each of which may contain one or more R b , for example, 1 to 6 R b and n is 0, 1, 2, 3, 4, 5, or 6.
41. The method of any one of claims 36 to 40, having the structure shown in
42. The bicyclic ketone substrate has the formula (IA): 【Chemistry 2】 42. The method of any one of claims 36 to 41, having the structure shown in
43. 43. The method of claim 41 or 42, wherein the A and B rings together represent a fused 6- to 12-membered heterocyclyl containing at least one nitrogen heteroatom.
44. 44. The method of claim 43, wherein the nitrogen heteroatom of the bicyclic ketone substrate is bound to an amine protecting group, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz).
45. The bicyclic ketone substrate has the formula (IA)-I 【Transformation 3】 (In the formula, X is N-R 1a , C.H. 2 , and CH-R 1b Selected from: R 1a is an amine protecting group, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz), C 1 ~C 20 Alkyl, C 2 ~C 20 Alkenyl, C 2 ~C 20 Alkynyl, C 3 ~C 10 Cycloalkyl, C 6 ~C 14 Aryl C 7 ~C 20 selected from arylalkyl, 3- to 14-membered heterocyclyl, and 5- to 20-membered heteroaryl; R 1b is C 1 ~C 20 Alkyl, C 2 ~C 20 Alkenyl, C 2 ~C 20 Alkynyl, C 3 ~C 10 Cycloalkyl, C 6 ~C 14 Aryl, C 7 ~C 20 selected from arylalkyl, 3- to 14-membered heterocyclyl, and 5- to 20-membered heteroaryl; Here, R 1a or R 1b The alkyl, alkenyl, and alkynyl groups each have 1 to 6 R a optionally substituted with; R 1a or R 1b The cycloalkyl, aryl, arylalkyl, heterocyclyl, and heteroaryl each have 1 to 6 R b may be substituted with; Each R 1c are independently generated when they occur, 1 ~C 20 Alkyl, C 2 ~C 20 Alkenyl, C 2 ~C 20 Alkynyl, C 3 ~C 10 Cycloalkyl, C 6 ~C 14 Aryl, C 7 ~C 20 Arylalkyl, 3- to 14-membered heterocyclyl, 5- to 20-membered heteroaryl, hydroxyl, halogen, e.g., F, C 1 ~C 20 Haloalkyl, e.g., —CF 3 , -OR c , -NR c R c , -(CH 2 ) n COOR c , -(CH 2 ) n -C(=O)R c , and -(CH 2 ) n -C(=O)NR c R c selected from the group consisting of: Each R a are independently generated when they occur, 3 ~C 10 Cycloalkyl, C 6 ~C 14 Aryl, C 7 ~C 20 arylalkyl, 3- to 14-membered heterocyclyl, 5- to 20-membered heteroaryl, halogen, such as F, haloalkyl, such as C 1 ~C 20 Haloalkyl, e.g., —CF 3 , -OR c , -NR c R c , -(CH 2 ) n COOR c , -(CH 2 ) n -C(=O)R c , and -(CH 2 ) n -C(=O)NR c R c Selected from: Each R b each occurrence independently represents a halogen, C 1 ~C 20 Haloalkyl, e.g., —CF 3 , -OR c , -NR c R c , -(CH 2 ) n COOR c , -(CH 2 ) n -C(=O)R c , -(CH 2 ) n -C(=O)NR c R c , C 1 ~C 20 Alkyl, C 2 ~C 20 Alkenyl, and C 2 ~C 20 alkynyl; Each R c are independently H, C when they occur. 1 ~C 20 Alkyl, C 2 ~C 20 Alkenyl, and C 2 ~C 20 Alkynyl (each of which has 1 to 6 R b and optionally substituted with n is 0, 1, 2, or 3; and m is 0, 1, or 2. The method according to any one of claims 36 to 44, wherein
46. X is N-R 1a and R 1a is C 1 ~C 10 Alkyl, C 3 ~C 10 Cycloalkyl, C 6 ~C 14 Aryl, C 7 ~C 20 46. The method of claim 45, wherein m is selected from aryl, alkyl, and amine protecting groups, such as tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz); and m=0.
47. X is N-R 1a and R 1a is an amine protecting group, for example, tert-butyloxycarbonyl (Boc), N-carboxybenzyl (Cbz).
48. A method for stereoselectively reducing a compound of formula (IA)-I to (IB)-I, 【Chemistry 4】 (In the formula, R 1a is an amine protecting group), contacting a substrate of formula (IA)-I with a ketoreductase (KRED) under reaction conditions suitable for reducing or converting (IA)-I to (IB)-I.
49. 49. The method of claim 48, wherein the KRED is an engineered ketoreductase.
50. The method of claim 48 or 49, wherein the KRED is a modified polypeptide according to any one of claims 1 to 35.
51. 51. The method of any one of claims 48 to 50, wherein the amine protecting group is tert-butyloxycarbonyl (Boc) or N-carboxybenzyl (Cbz).
52. 1. A method for the stereoselective reduction of tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate) to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (product), comprising: 【Transformation 5】 contacting the substrate tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate with a ketoreductase (KRED) under reaction conditions suitable for reducing or converting tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate to tert-butyl rel-(3aR,5s,6aS)-5-hydroxyhexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate.
53. 53. The method of claim 52, wherein the KRED is a modified KRED.
54. The method of claim 52 or 53, wherein the KRED is a modified polypeptide according to any one of claims 1 to 35.
55. 55. The method of any one of claims 36 to 54, wherein the reaction is carried out in a solvent.
56. 56. The method of claim 55, wherein the solvent is selected from a polar solvent, a non-polar solvent, and an ionic liquid.
57. 57. The method of claim 55 or 56, wherein the solvent is selected from water, methanol, ethanol, n-propanol, isopropanol, isopropyl acetate, dimethyl sulfoxide, dimethylformamide, ethyl acetate, butyl acetate, 1-octanol, hexane, heptane, octane, methyl tert-butyl ether, toluene, 1-ethyl 4-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium hexafluorophosphate, glycerol, ethylene glycol, propylene glycol, and polyethylene glycol.
58. 58. The method of any one of claims 55 to 57, wherein the solvent is isopropanol.
59. 59. The method of any one of claims 36 to 58, wherein the reaction is carried out in up to 40% by weight of isopropanol.
60. 60. The method of any one of claims 36 to 59, wherein the reaction is carried out in the presence of a co-solvent.
61. 61. The method of any one of claims 36 to 60, wherein the reaction is carried out in an aqueous co-solvent system.
62. 62. The method of claim 60 or 61, wherein the co-solvent is selected from dimethyl sulfoxide (DMSO) and alcohols such as methanol, ethanol, n-propanol, isopropanol.
63. 62. The method of any one of claims 36 to 61, wherein the reaction is not carried out in the presence of dimethyl sulfoxide (DMSO).
64. 64. A process according to any one of claims 36 to 63, wherein the reaction is carried out at a temperature of from 15 to 75°C, such as from 20 to 55°C, for example from 20 to 45°C, such as 35°C or 40°C.
65. 65. A method according to any one of claims 36 to 64, wherein the reaction is carried out at a pH of from 5.0 to 10.0, for example from 7.5 to 8.
3.
66. 66. The method of any one of claims 36 to 65, wherein the concentration of the bicyclic ketone is at most about 150 g / L, e.g., at least about 5 g / L, at least about 10 g / L, at least about 20 g / L, at least about 50 g / L, at least about 100 g / L, or about 150 g / L.
67. 67. The method of any one of claims 36 to 66, wherein the concentration of the polypeptide is less than about 10 g / L, such as less than about 5 g / L, such as less than about 3 g / L, such as less than about 1 g / L.
68. 68. The method of any one of claims 36 to 67, wherein the solvent is present at a concentration of 20% to 40% v / v.
69. 69. The method of any one of claims 36 to 68, performed using whole cells expressing the ketoreductase enzyme, or an extract or lysate of such cells.
70. 70. The method of any one of claims 36 to 69, wherein glucose dehydrogenase (GDH) / glucose cofactor recycling is not required.
71. 71. The method of any one of claims 36 to 70, wherein the ketoreductase is isolated and / or purified and the reduction reaction is carried out in the presence of a cofactor for the ketoreductase.
72. 72. The method of claim 71, wherein the cofactor comprises nicotinamide adenine dinucleotide (NADH) or nicotinamide adenine dinucleotide phosphate (NADPH).
73. 73. The method of claim 71 or 72, wherein the cofactor is present at a concentration of less than about 2 wt%, such as less than about 1 wt%, for example less than about 0.2 wt%, such as about 0.1 wt%, about 0.05 wt%, or about 0.03 wt%.
74. 74. The method of any one of claims 36 to 73, resulting in a product with a selectivity (% diastereomeric excess) of greater than about 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%, and / or a % conversion of at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99.0%, or 100%.
75. 75. The method of any one of claims 36-74, wherein at least about 85% of the substrate is reduced to the product in less than about 20 hours, and at least about 95% of the substrate is reduced to the product in less than about 30 hours.
76. Compounds of formula (IC) 【Transformation 6】 A method of synthesizing in free form or in pharmaceutically acceptable salt form, comprising: Formula (IA)-I 【Transformation 7】 (substrate) with a ketoreductase (KRED) under suitable reaction conditions to produce a compound of formula (IB)-I 【Transformation 8】 Compound (product) of formula (wherein R 1a a method comprising generating an amine protecting group, such as tert-butyloxycarbonyl (Boc) or N-carboxybenzyl (Cbz).
77. 77. The method of claim 76, wherein the KRED is a modified KRED.
78. 78. The method of claim 76 or 77, wherein the KRED is defined by any one of claims 1 to 35.
79. 1. A method for reversing or increasing the diastereoselectivity of a ketoreductase (KRED) polypeptide toward the formation of a (cis) alcohol product (e.g., (5s)-2) in a reduction reaction, comprising introducing one or more amino acid differences selected from Table 2, Table 4, Table 5, Table 6, or Table 7, relative to the amino acid sequence of the KRED polypeptide, wherein the one or more amino acid differences, when used in a reduction reaction, reverse or increase the diastereoselectivity of the KRED polypeptide toward the formation of a (cis) alcohol product (e.g., (5s)-2), relative to the KRED polypeptide lacking the one or more amino acid differences.
80. 80. The method of claim 79, wherein the one or more amino acid differences are selected from the following positions: 17, 18, 40, 56, 94, 96, 106, 145, 173, 190, 194, 196, 198, 202, and / or 206 relative to the amino acid sequence of the KRED polypeptide.
81. The one or more amino acid differences, as compared to the amino acid sequence of the KRED polypeptide, are as follows: X17S, X17T, X17G, X17A, X17R, X18G, X18S, X18R, X40H, X40R, X56V, X56S, X56T, X56D, X94A, X94V, X94L, X94I, X94W, X96Y, X106I, X106V, X106L, X106W, X106E, X145S, X145D, X145E , X173A, X173L, X173V, X173W, X173M, X173K, X173D, X190A, X194A, X194V, X194L, X194W, X194D, X194E, X194T, X194P, X196G, X198V, X198I, X198Y, X198D, X198P, X202A, X202G, X206Q, X206N, and / or X206W.
82. The one or more amino acid differences, compared to the amino acid sequence of the KRED polypeptide, are i) X17S, X17T, or X17G; ii) X18S or X18R; iii) X56S, X56T, or X56D; iv) X94I; iv) X106I, X106V, X106L, or X106W; vi) X173A, X173W, X173M, or X173K; vii) X194A, X194V, X194W, or X194T; vii) X196G; and viii) X198V, X198I, X198Y, or X198P 82. The method of any one of claims 79 to 81, selected from:
83. 83. The method of any one of claims 79 to 82, wherein the one or more amino acid differences are selected from the following, compared to the amino acid sequence of the KRED polypeptide: X94I, X96Y, X190A, X196G, X202A, X202G, and / or X206W.
84. 84. The method of any one of claims 79 to 83, wherein the one or more amino acid differences are selected from the following, compared to the amino acid sequence of the KRED polypeptide: X18G, X40H, X56V, X94A, X106E, X145E, X145, X173D, X194P, X198D, X202A, and X206Q.
85. The method of any one of claims 79 to 84, wherein the one or more amino acid differences, compared to the amino acid sequence of the KRED polypeptide, are selected from the following: (i) a residue at position X190 that is not tyrosine; (ii) a residue at position X190 that is a non-aromatic residue; (iii) a residue at position X196 that is an aliphatic residue or a small amino acid residue; (iv) a residue at position X202 that is a small amino acid residue; and / or (v) a residue at position X206 that is not methionine.
86. 86. The method of any one of claims 79 to 85, wherein the KRED polypeptide is a ketoreductase peptide of wild-type Lactobacillus kefir, Lactobacillus brevis, or Lactobacillus minor.
87. 87. The method of any one of claims 79 to 86, wherein the KRED polypeptide is a wild-type Lactobacillus kefir ketoreductase polypeptide.
88. 88. The method of any one of claims 79-87, wherein the KRED polypeptide is a wild-type Lactobacillus kefir ketoreductase polypeptide comprising the amino acid sequence of SEQ ID NO:
492.
89. 86. The method of any one of claims 79 to 85, wherein the KRED polypeptide is an engineered ketoreductase polypeptide.
90. 90. The method of claim 89, wherein the KRED polypeptide is an engineered ketoreductase polypeptide selected from Table 4, Table 5, or Table 6.
91. 90. The method of claim 89, wherein the KRED polypeptide is an engineered ketoreductase polypeptide selected from SEQ ID NO:54, SEQ ID NO:152, and SEQ ID NO:
256.
92. 92. The method of any one of claims 79-91, resulting in a percent conversion to the (cis)alcohol product (e.g., (5s)-2) of at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99.0%, or 100%.
93. 93. The method of any one of claims 79-92, wherein a selectivity (% diastereomeric excess) level of the (cis) alcohol product (e.g., (5s)-2) of at least about 96.0%, at least about 96.5%, at least about 97.0%, at least about 97.5%, at least about 98.0%, at least about 98.5%, at least about 99.0%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or 100% is obtained.
94. A composition comprising a modified polypeptide according to any one of claims 1 to 35.
95. 95. The composition of claim 94, further comprising structural formulas of substrate and / or product compounds.
96. 36. The modified polypeptide of any one of claims 1 to 35, wherein the polypeptide is immobilized on a solid support, encapsulated in a hydrogel (e.g., alginate, chitosan, carrageenan) or matrix (e.g., polyacrylamide), or by carrier-free immobilization, such as cross-linking.
97. 97. The modified polypeptide of claim 96, wherein the polypeptide is immobilized on the solid support by chemical bonding (e.g., covalent or ionic bonding), physical adsorption or affinity interaction.
98. 98. The modified polypeptide of claim 96 or 97, wherein the solid support is organic or inorganic.
99. 99. The modified polypeptide of any one of claims 96-98, wherein the solid support is selected from a resin, silica, zeolite, charcoal, celite (e.g., diatomaceous earth), synthetic polymer (e.g., polymethacrylate or the anion exchange resin Amberlite), biopolymer (e.g., cellulose, chitosan, agarose, lignin, or lignocellulose), controlled pore glass, magnetic nanoparticles, metal-organic frameworks, or DNA.
100. A polynucleotide encoding the modified polypeptide of any one of claims 1 to 35.
101. 101. The polynucleotide of claim 100, comprising a nucleic acid sequence listed in Table 4 (SEQ ID NOs: 3-51), Table 5 (SEQ ID NOs: 55-185), Table 6 (SEQ ID NOs: 187-389), or Table 7 (SEQ ID NOs: 391-489).
102. or 100% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 391, 393, 395, 405, 407, 419, 421, 423, 425, 427, 429, 435, 437, 461, 463, 467, 469, 473, 475, 479, 481, and 489. A polynucleotide encoding a modified KRED polypeptide, comprising a nucleic acid sequence.
103. 103. An expression vector comprising the polynucleotide of any one of claims 100 to 102 operably linked to a control sequence suitable for directing expression of the encoded polypeptide in a host cell.
104. 104. The expression vector of claim 103, wherein the control sequence comprises a secretion signal.
105. A host cell comprising a polynucleotide according to any one of claims 100 to 102, or an expression vector according to claim 103 or 104.
106. 106. A method for producing an engineered ketoreductase polypeptide, comprising culturing the host cell of claim 105 under conditions suitable for gene expression, and then purifying and recovering the engineered polypeptide from the cell culture.
107. 104. A kit comprising a modified polypeptide according to any one of claims 1 to 35, a composition according to any one of claims 94 or 95, a polynucleotide according to any one of claims 100 to 102, an expression vector according to claim 103 or 104, and / or a host cell according to claim 105.
108. A kit comprising the modified polypeptide of any one of claims 1 to 35, tert-butyl rel-(3aR,6aS)-5-oxohexahydrocyclopenta[c]pyrrole-2(1H)-carboxylate (substrate), and a cofactor.
109. 109. The kit of claim 108, wherein the cofactor is nicotinamide adenine dinucleotide (NADH) or nicotinamide adenine dinucleotide phosphate (NADPH).
110. 110. The kit of any one of claims 107 to 109, further comprising instructions for use of the modified polypeptide, composition, polynucleotide, expression vector, and / or host cell. 【Request Item 111】 【Chemistry 9】 or a salt thereof.
112. 112. The compound of claim 111, which is present in a diastereomeric excess (de)% of at least about 85%, at least about 96%, or at least about 99%. 【Request Item 113】 【Chemistry 10】 113. Use of a compound according to claim 111 or 112 or a salt thereof in the preparation of the compound in free form or in a pharmaceutically acceptable salt form.