Ketoreductase enzymes for the synthesis of 1,3-diol-substituted indenes
A ketoreductase enzyme with enhanced properties, developed through directed evolution, addresses the need for efficient conversion of ketones to chiral alcohols, achieving high stereoselectivity and stability for the synthesis of 1,3-indanediols and related pharmaceuticals.
Patent Information
- Application Number
- JP2025500026
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-08
- Filing Date
- 2023-07-05
- Publication Date
- 2025-07-23
Smart Images

Figure 2025523627000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to ketoreductase enzymes useful in biocatalysts and synthetic processes involving the reduction of ketones to chiral alcohols. Such enzymes can be particularly useful in the preparation of 1,3-diol substituted indanes.
[0002] Reference to Electronically Submitted Sequence Listing This application includes a sequence listing submitted electronically in XML format, which is hereby incorporated by reference in its entirety. The file name of the XML file created on October 31, 2022 is 25557-WO-PCT_SL.XML, and the size is 26,015 bytes.
Background Art
[0003] Enzymes are polypeptides that help (often by several orders of magnitude) accelerate the chemical reactions of living cells. Without enzymes, most biochemical reactions would be too slow to even carry out life processes. Enzymes exhibit great specificity and are not permanently altered by their involvement in the reaction. Since enzymes do not change during the reaction, when used as a catalyst for a desired chemical transformation, enzymes are particularly cost-effective.
[0004] Ketoreductase, also known as alcohol dehydrogenase, is an enzyme reductant that is a specific class of enzymes that catalyze the selective reduction of ketones to chiral alcohols. Enzymes belonging to the ketoreductase or carbonyl reductase class are useful in the synthesis of optically active alcohols. Ketoreductase enzymes can selectively convert ketone substrates or aldehyde substrates to the corresponding chiral alcohol products, and these enzymes can convert alcohols to the corresponding ketones or aldehydes in the reverse reaction. The enzymatic reduction of ketones and aldehydes requires the involvement of a cofactor that can act as an electron donor, and the enzymatic oxidation of alcohols requires the involvement of a cofactor that can act as an electron acceptor.
[0005] Ketoreductase enzymes are well-known in nature, and numerous genes encoding ketoreductase enzymes and ketoreductase enzyme sequences have been reported. See, for example, Candida magnoliae (Genbank accession number JC7338; GI:11360538), Candida parapsilosis (Genbank accession number 10 BAA24528.1; GI:2815409), Sporobolomyces salmonicolor (Genbank accession number AF160799; GI:6539734), and Rhodococcus erythropoli (Genbank accession number AAN73270.1; GI:34776951).
[0006] Ketoreductase enzymes are being used more and more frequently to provide alternative synthetic routes to important compounds. For example, Kosjek, B. et al. disclosed the asymmetric synthesis of the chiral precursor 4,4-dimethoxy-2H-pyran-3-ol using ketoreductase. Organic Process Research & Development, 2008, 12.584 - 588. When used, ketoreductase enzymes can be provided as purified enzymes or as whole cells expressing the desired ketoreductase. Given their promise for improved synthetic routes, there remains a need to identify additional ketoreductase enzymes that can be used to carry out specific chemical transformations to prepare specific chiral alcohols. SUMMARY OF THE INVENTION
[0007] The present disclosure relates to a ketoreductase enzyme that can convert a ketone to a chiral alcohol, particularly on a substituted indane backbone. In some embodiments, the ketoreductase enzyme of the subject matter described herein can convert hydroxyindanone to diastereomerically pure 1,3-indanediol useful in the synthesis of belzutifan, 3-[[(1S,2S,3R)-2,3-difluoro-2,3-dihydro-1-hydroxy-7-(methylsulfonyl)-1H-indene-4-yl]oxy]-5-fluorobenzonitrile. This agent is a hypoxia-inducible factor inhibitor and has recently been approved by the US Food and Drug Administration for the treatment of adult patients with von Hippel-Lindau disease who require treatment for related renal cell carcinoma, central nervous system hemangioblastoma, or pancreatic neuroendocrine tumors and do not require immediate surgery.
[0008] Further embodiments describe a process for preparing the subject ketoreductase enzyme and a process for using the subject ketoreductase enzyme.
[0009] Other embodiments, aspects, and features of the invention will be further described in or will become apparent from the following description, examples, and appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0010]
Figure 1
[0011] Definitions Specific technical and scientific terms are specifically defined as follows. Unless specifically defined elsewhere in this specification, all other technical and scientific terms used in this specification shall have the meanings commonly understood by those skilled in the art to which this disclosure pertains. Nevertheless, unless otherwise specified, the following definitions apply throughout the specification and the claims. Chemical names, common names, and chemical structures may be used interchangeably to describe the same structure.
[0012] As used in this specification and throughout the disclosure, the following terms shall be understood to have the following meanings, unless otherwise indicated: This disclosure also encompasses isotopically labeled compounds that are identical to those listed herein, except that one or more atoms are replaced by atoms having an atomic mass or mass number different from the atomic mass or mass number typically found in nature. Examples of isotopes that can be incorporated into the compounds of the invention include isotopes of hydrogen, carbon, nitrogen, oxygen, phosphorus, fluorine, chlorine, and iodine, such as, respectively 2 H, 3 H, 11 C, 13 C, 14 C, 15 N, 18 O, 17 O, 31 P, 32 P, 35 S, 18 F, 36 Cl and 123 I.
[0013] Certain isotopically labeled compounds (e.g., 3 H and 14 C labeled ones) are useful in compound and / or substrate tissue distribution assays. Tritiation (i.e., 3 H) and carbon-14 (i.e., 14C) Isotopes are particularly preferred due to their ease of preparation and detectability. Isotope substitution at the site where epimerization occurs slows down or reduces the epimerization process, thereby allowing the more active or effective form of the compound to be retained for a longer period. Isotope-labeled compounds, particularly those containing isotopes with a longer half-life (T 1 / 2 > 1 day), can generally be prepared by procedures similar to those disclosed in the schemes and / or examples below herein by using an appropriate isotope-labeled reagent in place of the non-isotope-labeled reagent.
[0014] The compounds of this specification may contain one or more stereocenters and may exist as racemates, racemic mixtures, single enantiomers, mixtures of diastereomers, and individual diastereomers. Depending on the nature of the various substituents on the molecule, additional asymmetric centers may be present. Each such asymmetric center independently generates two optical isomers, and all possible optical isomers and diastereomers in the mixture, as well as pure or partially purified compounds, are included in this disclosure. Any formula, structure, or name of a compound described herein that does not specify a particular stereochemistry is meant to encompass any and all of the above existing isomers and mixtures thereof in any proportion. When the stereochemistry is specified, this disclosure is meant to encompass that particular isomer in pure form or as part of a mixture with other isomers in any proportion.
[0015] Mixtures of diastereomers can be separated into individual diastereomers based on physicochemical differences by methods well known to those skilled in the art, such as chromatography and / or fractional crystallization. Enantiomers can be separated by converting the enantiomer mixture into a diastereomer mixture by reaction with an appropriate optically active compound (e.g., a chiral auxiliary such as a chiral alcohol or Mosher's acid chloride), separating the diastereomers, and converting the individual diastereomers into the corresponding pure enantiomers (e.g., by hydrolysis). Enantiomers can also be separated by the use of a chiral HPLC column.
[0016] All stereoisomers (e.g., geometric isomers, optical isomers, etc.) of the disclosed compounds (including salts, solvates, and prodrug salts, solvates, and esters thereof), such as enantiomeric forms (which may exist even in the absence of chiral carbons), rotamers, atropisomers, and diastereomeric forms, which may exist due to chiral carbons on various substituents, are contemplated within the scope of the present disclosure. Individual stereoisomers of the compounds may, for example, be substantially free of other isomers, or may be, for example, in racemic form or mixed with all or other selected stereoisomers. Chiral centers can have the S or R configuration as defined by the IUPAC 1974 Recommendations.
[0017] The present disclosure further includes the compounds and synthetic intermediates in all of their isolated forms. For example, the above compounds are intended to encompass all forms of the compounds such as any solvates, hydrates, stereoisomers, and tautomers thereof.
[0018] "Ketoreductase" and "KRED" are used interchangeably herein and refer to a polypeptide having the enzymatic ability to reduce a carbonyl group to its corresponding alcohol. More specifically, the disclosed ketoreductase polypeptide can stereoselectively reduce fluorohydroxyindanone (6) (below) to fluorodiol (7) (below).
Chemical formula
[0019] Polypeptides typically utilize the cofactor reduced nicotinamide adenine dinucleotide (NADH) or reduced nicotinamide adenine dinucleotide phosphate (NADPH) as a reducing agent. Ketoreductases as used herein include naturally occurring (wild-type) ketoreductases, as well as non-naturally occurring engineered polypeptides produced by human manipulation.
[0020] "Protein", "polypeptide", and "peptide" are used interchangeably herein to refer to a polymer of at least two amino acids covalently linked by amide bonds, regardless of length or post-translational modification (e.g., glycosylation or phosphorylation, lipidation, myristoylation, ubiquitination, etc.). This definition includes d-amino acids and l-amino acids, as well as mixtures of d-amino acids and l-amino acids, and polymers containing d-amino acids and l-amino acids, and mixtures of d-amino acids and l-amino acids. Proteins, polypeptides, and peptides can include tags such as histidine tags and should not be included when determining the percentage of sequence identity.
[0021] The "amino acid" or "residue" used in connection with the polypeptides disclosed herein refers to a particular monomer at a sequence position. Amino acids are referred to herein by either their generally known three-letter symbols or the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Similarly, nucleotides can be referred to by their generally accepted one-letter codes.
[0022] Abbreviations used for genetically encoded amino acids are conventional and are alanine (Ala or A), arginine (Arg or R), asparagine (Asn or N), aspartic acid (Asp or D), cysteine (Cys or C), glutamic acid (Glu or E), glutamine (Gln or Q), histidine (His or H), isoleucine (Ile or I), leucine (Leu or L), lysine (Lys or K), methionine (Met or M), phenylalanine (Phe or F), proline (Pro or P), serine (Ser or S), threonine (Thr or T), tryptophan (Trp or W), tyrosine (Tyr or Y), and valine (Val or V).
[0023] The abbreviations used for genetically encoding nucleosides are conventional and are as follows: adenosine (A); guanosine (G); cytidine (C); thymidine (T); and uridine (U). Unless otherwise specified, the abbreviated nucleosides can be either ribonucleosides or 2'-deoxyribonucleosides. A nucleoside can be specified as being either a ribonucleoside or a 2'-deoxyribonucleoside, either individually or as a group. When a nucleic acid sequence is shown as a series of single-letter abbreviations, the sequence is shown in the 5' to 3' direction according to common convention, and phosphates are not shown.
[0024] As used herein in the context of an enzyme, "derived from" identifies the originating enzyme on which the enzyme is based and / or the gene encoding such an enzyme. For example, the ketoreductase enzyme of SEQ ID NO: 2 was obtained by artificially evolving the gene encoding the ketoreductase enzyme of SEQ ID NO: 1 over multiple generations. Accordingly, this evolved ketoreductase enzyme is "derived from" the ketoreductase of SEQ ID NO: 1.
[0025] A "reference sequence" refers to a defined sequence used as a basis for sequence comparison. A reference sequence can be a subset of a larger sequence, such as a segment of a full-length gene or polypeptide sequence. Generally, a reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 residues in length, at least 50 residues in length, or the full length of a nucleic acid or polypeptide. Since two polynucleotides or polypeptides can each (1) contain sequences that are similar between the two sequences (i.e., a portion of the complete sequence) and (2) further contain sequences that differ between the two sequences, sequence comparison between two (or more) polynucleotides or polypeptides is typically performed by comparing the sequences of the two polynucleotides over a "comparison window" to identify and compare local regions of sequence similarity.
[0026] In some embodiments, a "reference sequence" can be based on a primary amino acid sequence, and the reference sequence can be a sequence that has one or more variations from the primary sequence. For example, a reference sequence "based on SEQ ID NO: 1 having proline at the residue corresponding to X190" refers to a reference sequence in which the corresponding residue of X190 in SEQ ID NO: 1 has been changed to proline.
[0027] A "hydrophilic amino acid or residue" refers to an amino acid or residue having a side chain that exhibits a hydrophobicity of less than 0 according to the normalized consensus hydrophobicity scale of Eisenberg et al., 1984, J. Mol. Biol. 179: 125-142. Genetically encoded hydrophilic amino acids include l-Thr (T), l-Ser (S), l-His (H), l-Glu (E), l-Asn (N), l-Gln (Q), l-Asp (D), l-Lys (K), and l-Arg (R).
[0028] An "acidic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that exhibits a pK value of less than about 6 when the amino acid is included in a peptide or polypeptide. Acidic amino acids typically have a negatively charged side chain at physiological pH due to the loss of a hydrogen ion. Genetically encoded acidic amino acids include l-Glu (E) and l-Asp (D).
[0029] A "basic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that exhibits a pKa value of greater than about 6 when the amino acid is included in a peptide or polypeptide. Basic amino acids typically have a positively charged side chain at physiological pH due to the association with a hydronium ion. Genetically encoded basic amino acids include l-Arg (R) and l-Lys (K).
[0030] "Polar amino acid or residue" refers to a hydrophilic amino acid or residue that is uncharged at physiological pH and has a side chain with at least one bond in which a pair of electrons shared in common by two atoms is held more closely by one of the atoms. Genetically encoded polar amino acids include l-Asn (N), l-Gln (Q), l-Ser (S), and l-Thr (T).
[0031] "Hydrophobic amino acid or residue" refers to an amino acid or residue having a side chain that exhibits a hydrophobicity greater than 0 according to the normalized consensus hydrophobicity scale of Eisenberg et al., 1984, J. Mol. Biol. 179:125-142. Genetically encoded hydrophobic amino acids include l-Pro (P), l-Ile (I), l-Phe (F), l-Val (V), l-Leu (L), l-Trp (W), l-Met (M), l-Ala (A), and l-Tyr (Y).
[0032] "Aromatic amino acid or residue" refers to a hydrophilic or hydrophobic amino acid or residue having a side chain that includes at least one aromatic or heteroaromatic ring. Genetically encoded aromatic amino acids include l-Phe (F), l-Tyr (Y), l-His (H), and l-Trp (W). l-His (H) histidine is also classified herein as a hydrophilic residue or a constrained residue.
[0033] As used herein, "constrained amino acid or residue" refers to an amino acid or residue having a constrained geometric shape. Here, constrained residues include l-Pro (P) and l-His (H). Histidine has a constrained geometric shape because it has a relatively small imidazole ring. Proline also has a constrained geometric shape because it has a five-membered ring.
[0034] "Nonpolar amino acid or residue" refers to a hydrophobic amino acid or residue that is uncharged at physiological pH and has a side chain with a bond in which the pair of electrons commonly shared by two atoms is generally held equally by each of the two atoms (i.e., the side chain is not polar). Genetically encoded nonpolar amino acids include l-Gly (G), l-Leu (L), l-Val (V), l-Ile (I), l-Met (M), and l-Ala (A).
[0035] As used herein, "aliphatic amino acid or residue" refers to a hydrophobic amino acid or residue having an aliphatic hydrocarbon side chain. Genetically encoded aliphatic amino acids include l-Ala (A), l-Val (V), l-Leu (L), and l-Ile (I).
[0036] The ability of l-Cys (C) (and other amino acids with -SH-containing side chains) to be present in a peptide in either the reduced free -SH or oxidized disulfide-bridged form affects whether l-Cys (C) contributes to the net hydrophobicity or hydrophilicity of the peptide. l-Cys (C) exhibits a hydrophobicity of 0.29 according to Eisenberg's normalized consensus scale (Eisenberg et al., 1984, supra), but for the purposes of the present disclosure, it should be understood that l-Cys (C) is classified into its own unique group. Note that cysteine (or "l-Cys" or "[C]") is unique in that it can form disulfide bridges with other l-Cys (C) amino acids or other sulfanyl- or sulfhydryl-containing amino acids. "Cysteine-like residues" include cysteine and other amino acids containing a sulfhydryl moiety available for the formation of disulfide bridges.
[0037] As used herein, "small amino acid or residue" refers to an amino acid or residue having a side chain composed of three or fewer total carbons and / or heteroatoms (excluding α-carbon and hydrogen). Small amino acids or residues can be further classified as aliphatic, nonpolar, polar or acidic small amino acids or residues according to the above definition. Genetically encoded small amino acids include l-Ala (A), l-Val (V), l-Cys (C), l-Asn (N), l-Ser (S), l-Thr (T) and l-Asp (D).
[0038] "Hydroxyl-containing amino acid or residue" refers to an amino acid containing a hydroxyl (-OH) moiety. Genetically encoded hydroxyl-containing amino acids include l-Ser (S), l-Thr (T) and l-Tyr (Y).
[0039] As used herein, "conservative amino acid substitution" refers to the substitution of a residue by a different residue having a similar side chain, and thus typically includes substitution of an amino acid in a polypeptide by an amino acid within the same or a similar defined amino acid class of amino acids. By way of example and not limitation, in some embodiments, an amino acid having an aliphatic side chain is substituted by another aliphatic amino acid (e.g., alanine, valine, leucine and isoleucine), an amino acid having a hydroxyl side chain is substituted by another amino acid having a hydroxyl side chain (e.g., serine and threonine), an amino acid having an aromatic side chain is substituted by another amino acid having an aromatic side chain (e.g., phenylalanine, tyrosine, tryptophan and histidine), an amino acid having a basic side chain is substituted by another amino acid having a basic side chain (e.g., lysine and arginine), an amino acid having an acidic side chain is substituted by another amino acid having an acidic side chain (e.g., aspartic acid and glutamic acid), and / or a hydrophobic or hydrophilic amino acid is substituted by another hydrophobic or hydrophilic amino acid, respectively.
[0040] As used herein, "non-conservative substitution" refers to the substitution of an amino acid in a polypeptide with an amino acid having significantly different side chain characteristics. Non-conservative substitutions can use amino acids between, rather than within, defined groups, and can affect (a) the structure of the peptide backbone in the region of substitution (e.g., proline for glycine), (b) the charge or hydrophobicity, or (c) the majority of the side chain. By way of example and not limitation, exemplary non-conservative substitutions can be acidic amino acids substituted with basic or aliphatic amino acids; aromatic amino acids substituted with small amino acids; and hydrophilic amino acids substituted with hydrophobic amino acids.
[0041] As used herein, "deletion" refers to the modification of a polypeptide by the removal of one or more amino acids from a reference polypeptide. Deletions can include the removal of one or more amino acids, two or more amino acids, five or more amino acids, ten or more amino acids, fifteen or more amino acids, or twenty or more amino acids, up to a maximum of 10% of the total number of amino acids, or up to a maximum of 20% of the total number of amino acids, while retaining the enzymatic activity of the evolved enzyme and / or retaining improved properties. Deletions can be directed to internal and / or terminal portions of the polypeptide. In various embodiments, deletions can include contiguous segments or can be discontinuous. Deletions are typically indicated by "-" in the amino acid sequence.
[0042] As used herein, "insertion" refers to the modification of a polypeptide by the addition of one or more amino acids to a reference polypeptide. Insertions can be internal to the polypeptide or at the carboxy or amino terminus. Insertions as used herein include fusion proteins known in the art. Insertions can be contiguous segments of amino acids or can be separated by one or more of the amino acids in a naturally occurring polypeptide.
[0043] The term "amino acid substitution set" or "substitution set" refers to a group of amino acid substitutions within a polypeptide sequence as compared to a reference sequence. A substitution set can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more amino acid substitutions.
[0044] The terms "functional fragment" and "biologically active fragment" are used interchangeably herein to refer to a polypeptide that has (one or more) amino-terminal and / or carboxy-terminal deletions and / or internal deletions, but the remaining amino acid sequence is identical to the corresponding positions in the sequence to which it is being compared and retains substantially all of the activity of the full-length polypeptide.
[0045] As used herein, "isolated polypeptide" refers to a polypeptide that is substantially separated from other contaminants that naturally accompany it (e.g., proteins, lipids, and polynucleotides). This term encompasses polypeptides that have been removed or purified from their natural environment or expression system (e.g., within a host cell or via in vitro synthesis). Recombinant polypeptides may be present within a cell, in the cell culture medium, or may be prepared in various forms such as lysates or isolated preparations. Thus, in some embodiments, a recombinant polypeptide can be an isolated polypeptide.
[0046] As used herein, "substantially pure polypeptide" or "purified protein" refers to a composition in which the polypeptide species is the dominant species present (i.e., is more abundant than any other individual macromolecular species in the composition on a molar or weight basis), and generally, when the subject species comprises at least about 50 percent of the macromolecular species present in molar or weight percent, is a substantially purified composition. However, in some embodiments, the enzyme-containing composition comprises an enzyme at a purity of less than 50% (e.g., about 10%, about 20%, about 30%, about 40%, or about 50%). Generally, a substantially pure enzyme or polypeptide composition comprises about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, and about 98% or more of the molar or weight percent of all macromolecular species present in the composition. In some embodiments, the subject species is purified to an essential homogeneity in which the composition consists essentially of a single macromolecular species (i.e., contaminating species cannot be detected in the composition by conventional detection methods). Solvent species, small molecules (less than 500 daltons), and elemental ion species are not considered macromolecular species. In some embodiments, an isolated recombinant polypeptide is a substantially pure polypeptide composition.
[0047] "Improved enzyme properties" refers to an enzyme that exhibits an improvement in any enzyme property as compared to a reference enzyme. For the enzymes described herein, generally a wild-type enzyme is compared, but in some embodiments, the reference enzyme can be another improved enzyme. Enzyme properties for which improvement may be desired include, but are not limited to, enzyme activity (which can be expressed in terms of substrate conversion rate), thermal stability, pH activity profile, cofactor requirement, resistance to inhibitors (e.g., product inhibition), stereospecificity, and stereoselectivity (including enantioselectivity).
[0048] "Increase in enzyme activity" refers to an improved property of an enzyme that can be represented by an increase in specific activity (e.g., product generated / time / weight protein) or an increase in the conversion rate of a substrate to a product (e.g., the conversion rate of the starting amount of a substrate to a product over a specified period using a specified amount of enzyme) compared to a reference enzyme. Exemplary methods for determining enzyme activity are provided in the Examples. K m , V max , or k cat Any property related to enzyme activity, including classical enzyme properties of, can be affected, and the change can result in an increase in enzyme activity. The improvement in enzyme activity can be about 1.5 to 2 times the enzyme activity of the corresponding wild-type enzyme. 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold, 500-fold, 1000-fold, 3000-fold, 5000-fold, 7000-fold or more enzyme activity compared to other enzymes from which a naturally occurring enzyme or polypeptide is derived. In certain embodiments, the enzyme exhibits improved enzyme activity in the range of 150 to 3000-fold, 3000 to 7000-fold, or more than 7000-fold compared to the enzyme activity of the parental enzyme. It is understood by those skilled in the art that the activity of any enzyme is diffusion-limited such that the catalytic turnover rate cannot exceed the diffusion rate of the substrate, including any necessary cofactors. The theoretically maximum value of the diffusion limit, i.e., k cat / K m is generally about 10 8 ~10 9 (M -1 s -1) Therefore, the improvement of enzyme activity has an upper limit related to the diffusion rate of the substrate acted upon by the enzyme. Enzyme activity can be measured by any one of the standard assays used to measure kinase activity, or by a binding assay with a nucleoside phosphorylase enzyme that can catalyze the reaction between a polypeptide product and a nucleoside base to obtain a nucleoside, or by any of the conventional methods for assaying chemical reactions including but not limited to HPLC, HPLC-MS, UPLC, UPLC-MS, TLC, and NMR. The comparison of enzyme activities is performed using a defined preparation of the enzyme, a defined assay under set conditions, and one or more defined substrates, as described in more detail herein. Generally, when comparing lysates, the number of cells and the amount of protein assayed are determined in the same manner as the use of the same expression system and the same host cell in order to minimize variations in the amount of enzyme produced by the host cell and present in the lysate.
[0049] As used herein, a "vector" is a DNA construct for introducing a DNA sequence into a cell. In some embodiments, the vector is an expression vector operably linked to appropriate control sequences for expression in an appropriate host of a polypeptide encoded by the DNA sequence. In some embodiments, an "expression vector" has a promoter sequence operably linked to a DNA sequence (e.g., a transgene) for driving expression in a host cell, and in some embodiments, also includes a transcription terminator sequence.
[0050] As used herein, the term "expression" includes any process involved in the production of a polypeptide, including but not limited to transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses the secretion of the polypeptide from the cell.
[0051] As used herein, the term "produce" refers to the production of proteins and / or other compounds by a cell. This term is intended to encompass any process involved in the production of a polypeptide, including but not limited to transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, this term also encompasses the secretion of a polypeptide from a cell.
[0052] As used herein, an amino acid or nucleotide sequence (e.g., a promoter sequence, a signal peptide, a terminator sequence, etc.) is "heterologous" to another sequence to which it is operably linked if the two sequences are not related in nature. For example, a "heterologous polynucleotide" is any polynucleotide introduced into a host cell by laboratory techniques, and this term includes polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into the host cell.
[0053] As used herein, the terms "host cell" and "host strain" refer to a host suitable for an expression vector containing the DNA provided herein (e.g., a polynucleotide encoding a variant). In some embodiments, the host cell is a prokaryotic or eukaryotic cell that has been transformed or transfected with a vector constructed using recombinant DNA techniques known in the art.
[0054] The term "analog" means a polypeptide having greater than 70% sequence identity but less than 100% sequence identity (e.g., greater than 75%, 78%, 80%, 83%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity) to a reference polypeptide. In some embodiments, "analog" means a polypeptide that contains one or more non-naturally occurring amino acid residues, including but not limited to homoarginine, ornithine, and norvaline, as well as naturally occurring amino acids. In some embodiments, an analog also includes one or more d-amino acid residues and non-peptide bonds between two or more amino acid residues.
[0055] As used herein, the term "EC" number refers to the enzyme nomenclature of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (NC-IUBMB). The IUBMB biochemical classification is a numerical classification system for enzymes based on the chemical reactions catalyzed by the enzymes.
[0056] As used herein, "ATCC" refers to the American Type Culture Collection, whose biological repository collection includes genes and strains.
[0057] As used herein, "NCBI" refers to the National Center for Biological Information and the sequence databases provided therein.
[0058] A "coding sequence" refers to the portion of a nucleic acid (e.g., a gene) that encodes the amino acid sequence of a protein.
[0059] "Naturally occurring" or "wild-type" refers to the form found in nature. For example, a naturally occurring or wild-type polypeptide or polynucleotide sequence can be isolated from a natural source and is a sequence that exists in an organism that has not been intentionally modified by human manipulation, with the sole exception that a wild-type polypeptide or polynucleotide sequence identified herein may contain tags such as a histidine tag and should not be included when determining the percentage of sequence identity. In this specification, a "wild-type" polypeptide or polynucleotide sequence may be denoted as "WT".
[0060] "Recombinant", when used with respect to, for example, a cell, nucleic acid, or polypeptide, refers to a material that has been modified in a manner that would not otherwise exist in nature, or is identical thereto, but is produced or derived from synthetic materials and / or by manipulation using recombinant techniques, or a material corresponding to the natural or native form of the material. Non-limiting examples include, inter alia, recombinant cells that express a gene not found within the natural (non-recombinant) form of the cell, or that express a native gene that would otherwise be expressed at a different level.
[0061] "Percentage of sequence identity", "percent identity", and "identity percent" are used herein to refer to a comparison between polynucleotide sequences or polypeptide sequences, and are determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide sequence or polypeptide sequence within the comparison window may include additions or deletions (i.e., gaps) as compared to the reference sequence for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences, or the number of positions at which the nucleic acid base or amino acid residue aligned with a gap matches, dividing the number of matched positions by the total number of positions within the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Optimal alignment and determination of percent sequence identity are performed using the BLAST and BLAST 2.0 algorithms (see, e.g., Altschul et al., 1990, J. Mol. Biol. 215:403-410; and Altschul et al., 1977, Nucleic Acids Res. 3389-3402). Software for performing BLAST analyses is publicly available through the website of the National Center for Biotechnology Information.
[0062] Briefly, BLAST analysis involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which match or exceed a certain positive threshold score T when aligned with words of the same length in the database sequences. T is called the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits serve as seeds to initiate a search to find longer HSPs that contain them. The word hits are then extended in both directions along each sequence as long as the cumulative alignment score can be increased. The cumulative score is calculated for nucleotide sequences using parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for a mismatch residue; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. The extension of the word hit in each direction stops when the cumulative alignment score drops by an amount X from its maximum achieved value. The cumulative score becomes zero or less due to the accumulation of one or more negative-score residue alignments, or reaches the end of either sequence. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses, by default, a word length (W) of 11, an expectation value (E) of 10, M = 5, N = -4, and comparison of both strands. For amino acid sequences, the BLASTP program uses, by default, a word length (W) of 3, an expectation value (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, 1989, Proc. Natl. Acad. Sci. USA 89:10915).
[0063] A number of other algorithms are available that function in a similar manner to BLAST in providing percent identity for two arrays. Optimal alignment of sequences for comparison can be conducted, for example, by the local homology algorithms of Smith and Waterman, 1981, Adv. Appl. Math. 2:482, by the homology alignment algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48:443, by the search for similarity method of Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444, computer implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin software package), or by visual inspection (generally see Current Protocols in Molecular Biology, F. M. Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)). Further, determination of sequence alignment and percent sequence identity can be made using the BESTFIT or GAP programs of the GCG Wisconsin Software package (Accelrys, Madison, Wisconsin) with the default parameters provided.
[0064] "Substantial identity" refers to a polynucleotide or polypeptide sequence having at least 80% sequence identity, preferably at least 85% sequence identity, more preferably at least 89% sequence identity, more preferably at least 95% sequence identity, even more preferably at least 99% sequence identity when compared to a reference sequence over a comparison window of at least 20 residue positions, often over a window of at least 30 to 50 residues. The percentage of sequence identity is calculated by comparing the reference sequence with a sequence that includes deletions or additions that total 20% or less of the reference sequence over the comparison window. In certain embodiments applied to polypeptides, the term "substantial identity" means that two polypeptide sequences share at least 80% sequence identity, preferably at least 89% sequence identity, more preferably at least 95% sequence identity or more (e.g., 99% sequence identity) when optimally aligned by a program such as GAP or BESTFIT using the default gap weights. Preferably, the non-identical residue positions differ by conservative amino acid substitutions.
[0065] When used in reference to the numbering of a given amino acid or polynucleotide sequence, "corresponding to", "referring to", or "relative to" refers to the numbering of the residues of a particular reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue numbers or residue positions of a given polymer are designated with respect to the reference sequence, rather than by the actual numerical positions of the residues within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, gaps are present, but the numbering of the residues in the given amino acid or polynucleotide sequence is done with respect to the reference sequence to which it is aligned.
[0066] "Stereoselectivity" refers to the preferential formation of one stereoisomer over another in a chemical or enzymatic reaction. Stereoselectivity can be partial, where the formation of one stereoisomer is favored over the other, or complete, where only one stereoisomer is formed. When the stereoisomers are enantiomers, the stereoselectivity is called enantioselectivity, which is the ratio of one enantiomer in the total of both (typically reported as a percentage). This is generally alternatively reported in the art as the enantiomeric excess (EE), calculated therefrom according to the formula [major enantiomer - minor enantiomer] / [major enantiomer + minor enantiomer] (typically as a percentage). When the stereoisomers are diastereoisomers, the stereoselectivity is called diastereoselectivity, which is the ratio of one diastereomer in a mixture of two diastereomers (typically reported as a percentage), and is generally alternatively reported as the diastereomeric excess (DE). Enantiomeric excess and diastereomeric excess are types of stereoisomeric excess.
[0067] "Highly stereoselective" refers to a chemical or enzymatic reaction that can convert a substrate to its corresponding product with a stereoisomeric excess of at least about 85%.
[0068] "Chemoselectivity" refers to the preferential formation of one product over another in a chemical or enzymatic reaction.
[0069] "Conversion" refers to the enzymatic conversion of a substrate to its corresponding product. "Conversion rate" refers to the proportion of substrate that is converted to product within a certain period under specific conditions. Thus, for example, the "enzymatic activity" or "activity" of a polypeptide can be expressed as the "conversion rate" of substrate to product.
[0070] "Chiral alcohol" refers to an amine of the general formula R 1 -CH(OH)-R 2 wherein R 1 and R 2are not the same and are used herein in their broadest sense and include a wide variety of aliphatic and cycloaliphatic compounds of different and mixed functional types, characterized by the presence of a primary hydroxyl group bonded to a secondary carbon atom having either (i) a divalent group forming a chiral cyclic structure or (ii) two substituents (other than hydrogen) that differ in structure or chirality. Examples of the divalent group forming a chiral cyclic structure include 2-methylbutane-1,4-diyl, pentane-1,4-diyl, hexane-1,4-diyl, hexane-1,5-diyl, 2-methylpentane-1,5-diyl. The two different substituents (R 1 and R 2 ) on the secondary carbon atom can also vary widely and can include alkyl, aralkyl, aryl, halo, hydroxy, lower alkyl, lower alkoxy, lower alkylthio, cycloalkyl, carboxy, carboalkoxy, carbamoyl, mono- and di(lower alkyl)-substituted carbamoyl, trifluoromethyl, phenyl, nitro, amino, mono- and di(lower alkyl)-substituted amino, alkylsulfonyl, arylsulfonyl, alkylcarboxamide, arylcarboxamide, etc., and alkyl, aralkyl or aryl substituted by the foregoing.
[0071] Immobilized enzyme preparations have several recognized advantages. They can, for example, confer a storage life on the enzyme preparation, improve reaction stability, enable stability in organic solvents, and assist in protein removal from the reaction stream. "Stable" refers to the ability of the immobilized enzyme to retain its structural conformation and / or its activity in a solvent system containing an organic solvent. A stable immobilized enzyme loses less than 10% of its activity per hour in a solvent system containing an organic solvent. A stable immobilized enzyme loses less than 9% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 8% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 7% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 6% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 5% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 4% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 3% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 2% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 1% of its activity per hour in a solvent system containing an organic solvent.
[0072] "Thermal stability" refers to a polypeptide that maintains similar activity (e.g., exceeding 60% - 80%) after exposure to high temperatures (e.g., 40°C - 80°C) for a certain period (e.g., 0.5 hours - 24 hours) compared to the untreated enzyme.
[0073] "Solvent stability" refers to a polypeptide that maintains a similar activity (e.g., exceeding 60% - 80%) after being exposed to solvents (such as isopropyl alcohol, tetrahydrofuran, 2-methyltetrahydrofuran, acetone, toluene, butyl acetate, methyl tert-butyl ether, etc.) at various concentrations (e.g., 5% - 99%) for a certain period (e.g., 0.5 hours to 24 hours) compared to the untreated enzyme.
[0074] "pH stability" refers to a polypeptide that maintains a similar activity (e.g., exceeding 60% - 80%) after being exposed to high or low pH (e.g., 4.5 - 6 or 8 - 12) for a certain period (e.g., 0.5 hours to 24 hours) compared to the untreated enzyme.
[0075] "Thermal stability and solvent stability" refers to a polypeptide that has both thermal stability and solvent stability.
[0076] As used herein, the terms "biocatalysis", "biocatalytic", "in vivo conversion", and "biosynthesis" refer to the use of an enzyme to perform a chemical reaction on an organic compound.
[0077] The term "effective amount" means an amount sufficient to bring about the desired result. One of ordinary skill in the art can determine an effective amount using routine experimentation.
[0078] The terms "isolated" and "purified" are used to refer to a molecule (e.g., an isolated nucleic acid, polypeptide, etc.) or other component that has been removed from at least one other component to which it is naturally bound. The term "purified" does not require absolute purity but is rather intended as a relative definition.
[0079] "Control sequences" are defined herein as comprising all components necessary or advantageous for the expression of the polynucleotide and / or polypeptide of interest. Each control sequence may be native or foreign to the nucleic acid sequence encoding the polypeptide. Such control sequences include, but are not limited to, a leader, polyadenylation sequence, propeptide sequence, promoter, signal peptide sequence, and transcription terminator. At a minimum, the control sequences include a promoter, as well as transcription and translation stop signals. Linkers may be provided in the control sequences for the purpose of introducing specific restriction sites that facilitate ligation of the control sequences to the coding region of the polynucleotide of interest, for example, a nucleic acid sequence encoding a polypeptide.
[0080] "Operably linked" is defined herein as a configuration in which a control sequence is positioned at an appropriate position relative to the polynucleotide sequence (i.e., in a functional relationship) such that the control sequence directs the expression of the polynucleotide and / or the polypeptide encoded by the polynucleotide.
[0081] "Promoter sequence" is a nucleic acid sequence recognized by a host cell for the expression of a polynucleotide. The control sequences may include an appropriate promoter sequence. The promoter sequence includes a transcriptional control sequence that mediates the expression of the polynucleotide. A promoter can be any nucleic acid sequence (including mutant promoters, truncated promoters, and hybrid promoters) that exhibits transcriptional activity in the selected host cell and can be obtained from a gene encoding an extracellular or intracellular polypeptide that is homologous or heterologous to the host cell.
[0082] Exemplary methods and materials are described herein, but methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure. The materials, methods, and examples are illustrative only and are not intended to be limiting. [Table 1]
[0083] Keto-reductase The present disclosure relates to a keto-reductase enzyme that can reduce a ketone to a chiral alcohol, particularly on a substituted indane backbone. In embodiments, the keto-reductase enzyme is capable of the following conversion; [Chemical formula]
[0084] In certain embodiments, the keto-reductase enzyme is utilized in combination with an electrophilic fluorinating agent in a transformation. [Chemical formula]
[0085] In embodiments, the keto-reductase enzyme described herein has an amino acid sequence having one or more amino acid differences compared to a reference amino acid sequence of a commercially available non-wild-type keto-reductase, resulting in improved properties of the enzyme with respect to a defined keto-substrate.
[0086] The keto-reductase enzyme described herein is the product of directed evolution from a commercially available keto-reductase (SEQ ID NO: 1, Codexis, Inc.), identified by screening a collection of Codexis enzymes, and has the amino acid sequence shown below.
[0087] TIFF2025523627000006.tif79153
[0088] In embodiments, the keto-reductase enzyme of the present disclosure may exhibit improvements such as increased enzyme activity, stereoselectivity, stereospecificity, thermal stability, solvent stability, or decreased product inhibition, compared to the keto-reductase enzyme of SEQ ID NO: 1.
[0089] In some embodiments, the ketoreductase enzyme of the present disclosure may demonstrate an improvement in the rate of enzyme activity, i.e., the rate of converting substrate to product. In some embodiments, the ketoreductase polypeptide can convert substrate to product at a rate of at least 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 25-fold, 50-fold, 100-fold, 150-fold, 200-fold, 400-fold, 1000-fold, 3000-fold, 5000-fold, 7000-fold or more than 7000-fold the rate shown by the enzyme of SEQ ID NO: 1.
[0090] In some embodiments, such a ketoreductase polypeptide can also convert substrate to product with at least about 80% percent enantiomeric excess. In some embodiments, such a ketoreductase polypeptide can also convert substrate to product with at least about 90% percent enantiomeric excess. In some embodiments, such a ketoreductase polypeptide can also convert substrate to product with at least about 99% percent enantiomeric excess.
[0091] In some embodiments, the ketoreductase polypeptide is highly stereoselective, and the polypeptide can reduce substrate to product with a stereoisomeric excess of about 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8% or more than 99.9%.
[0092] In some embodiments, the improved ketoreductase polypeptide of the present disclosure is a polypeptide comprising an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 2 having the amino acid sequence shown below.
[0093] TIFF2025523627000007.tif81153
[0094] In some embodiments, the improved ketoreductase polypeptide of the present disclosure is based on the sequence set forth in SEQ ID NO: 2 and may include an amino acid sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the reference sequence of SEQ ID NO: 2.
[0095] These differences between these variants and SEQ ID NO: 2 can be insertions, deletions, substitutions of amino acids, or any combination of such changes. In some embodiments, the differences in the amino acid sequences can include non-conservative, conservative, and combinations of non-conservative and conservative amino acid substitutions.
[0096] In some embodiments, the ketoreductase polypeptide is a polypeptide selected from SEQ ID NOs: 2, 8, 9, 10, 11, 12, 13, 14, 15, and 16. In certain embodiments, the ketoreductase polypeptide is a polypeptide selected from SEQ ID NOs: 2, 14, 15, and 16.
[0097] In certain embodiments, the improved ketoreductase polypeptide of the present disclosure, wherein the amino acid sequence consists of SEQ ID NO: 2. In certain embodiments, the improved ketoreductase polypeptide of the present disclosure consists of SEQ ID NO: 2.
[0098] In one aspect, the improved ketoreductase polypeptide of the present disclosure is based on an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 2 and the following conditions: (a) Amino acid (aa) residue 2 of SEQ ID NO: 2 is other than alanine, or (b) Aa residue 11 of SEQ ID NO: 2 is other than glutamic acid, at least one of which is satisfied.
[0099] In certain embodiments, both conditions (a) and (b) are met in the improved ketoreductase polypeptides of the present disclosure. In certain embodiments, aa residue 2 of SEQ ID NO: 2 is lysine in the improved ketoreductase polypeptides of the present disclosure. In certain embodiments, aa residue 11 of SEQ ID NO: 2 is alanine in the improved ketoreductase polypeptides of the present disclosure.
[0100] In certain embodiments of this aspect, the improved ketoreductase polypeptides of the present disclosure are based on amino acid sequences having at least 85%, 90%, 95% or 98% sequence identity with SEQ ID NO: 2.
[0101] The applicant has discovered that polypeptides having aa residues other than alanine residues at the position corresponding to position 2 of SEQ ID NO: 2 are expressed better from their parent polynucleotides. Similarly, the applicant has found that polypeptides having aa residues other than glutamic acid residues at the position corresponding to position 11 of SEQ ID NO: 2 are also expressed better from their parent polynucleotides.
[0102] Further embodiments provide host cells comprising the polynucleotides and / or expression vectors described herein. The host cell may be E. coli, or a different organism such as L. brevis. The host cell can be used for the expression and isolation of the ketoreductase enzyme described herein, or can be used directly for the conversion of a substrate to a stereoisomeric product.
[0103] Whether the method is carried out using whole cells, cell extracts or purified ketoreductase enzyme, a single ketoreductase enzyme may be used, or a mixture of two or more ketoreductase enzymes may be used.
[0104] Embodiments relate to ketoreductase enzymes capable of selectively preparing chiral alcohols in the synthesis of indanediols, particularly fluoro-substituted indanediols. In embodiments, the ketoreductase enzyme is capable of the following conversions: [Chem.]
[0105] In certain embodiments, the ketoreductase enzyme, in combination with an electrophilic fluorinating agent, enables the following conversion: [Chem.]
[0106] Desirable enzyme characteristics for improvement include, but are not limited to, enzyme activity, thermal stability, pH activity profile, cofactor requirements, insensitivity to inhibitors (e.g., product inhibition), stereospecificity, stereoselectivity, and solvent stability. The improvement can relate to a single enzyme characteristic such as enzyme activity, or a combination of different enzyme characteristics such as enzyme activity and stereoselectivity.
[0107] Table 1 below provides a list of SEQ ID NOs disclosed herein having the relevant activities. The following sequences are based on the ketoreductase sequence of SEQ ID NO:1, unless otherwise specified. In Table 1 below, each row enumerates a SEQ ID NO. The column enumerating the number of mutations (i.e., residue changes) refers to the number of amino acid substitutions compared to the ketoreductase sequence of SEQ ID NO:1. In the Activity column of Table 1, + indicates 30 - 50% (v / v) of untreated lysate, ++ indicates 15 - 30% (v / v) of untreated lysate, +++ indicates 15 - 30% of iPrOH-treated lysate, ++++ indicates 10 - 15% of iPrOH-treated lysate, and +++++ indicates 5 - 10% of iPrOH-treated lysate. In the Selectivity column of the table, + indicates an overall diastereomer ratio < 10:1, ++ indicates a total dr of 10:1 - 25:1, +++ indicates a total dr of 25:1 - 50:1, ++++ indicates a total dr of 50:1 - 99:1, and +++++ indicates a total dr > 99:1. In the Solvent Tolerance column of the table, + indicates 0 - 12.5% (v / v) of total organic solvent, ++ indicates 12.5 - 18% (v / v) of total organic solvent, +++ indicates 18 - 25% (v / v) of total organic solvent, ++++ indicates 25 - 40% (v / v) of total organic solvent, and +++++ indicates the presence of 40 - 50% (v / v) of total organic solvent. In the Volumetric Productivity column of the table, + indicates 0 - 20 g / L of purified substrate, ++ indicates 20 - 40 g / L of crude substrate, +++ indicates 40 - 50 g / L of crude substrate, ++++ indicates 50 - 60 g / L of crude substrate, and +++++ indicates 60 - 80 g / L of crude substrate. In the iPrOH Utilization column of the table, + indicates 7.5 - 10% (v / v), ++ indicates 11 - 13% (v / v), +++ indicates 13.5% (v / v) iPrOH containing 1.5% (v / v) acetone, ++++ indicates 15% (v / v) iPrOH, and +++++ indicates 16 - 18% (v / v) iPrOH.
[0108] The improvement level of the properties listed in Table 1 was determined by setting the reaction under increasingly difficult conditions, including an increase in the organic solvent (both composition and amount) and an increase in the substrate concentration. The performance of each enzyme variant was measured by observing the amount of product formed with an Agilent UPLC instrument, typically after an overnight reaction as described in Example 2. When the screening focused on the improvement of diastereoselectivity, the amount of each product formed was measured with an Agilent SFC instrument, also described in Example 2.
Table 2
[0109] TIFF2025523627000011.tif231152
[0110] TIFF2025523627000012.tif188152
[0111] TIFF2025523627000013.tif120153
[0112] A polynucleotide encoding a ketoreductase In another aspect, the present disclosure provides a polynucleotide encoding a ketoreductase polypeptide disclosed herein. The polynucleotide can be operably linked to one or more heterologous regulatory sequences that control gene expression to create a recombinant polynucleotide capable of expressing the polypeptide. An expression construct containing the heterologous polynucleotide encoding the ketoreductase can be introduced into a suitable host cell to express the corresponding ketoreductase polypeptide.
[0113] Due to knowledge of the codons corresponding to various amino acids, the availability of the protein sequence provides an explanation of all the polynucleotides that can encode the subject. The degeneracy of the genetic code, where the same amino acid is encoded by alternative or synonymous codons, allows for the creation of a vast number of nucleic acids, all of which encode the improved ketoreductase enzymes disclosed herein. Thus, having identified a particular amino acid sequence, one of ordinary skill in the art can create any number of different nucleic acids by simply modifying the sequence of one or more codons in a manner that does not change the amino acid sequence of the protein. In this regard, the present disclosure specifically contemplates every possible variation of polynucleotides that can be made by selecting combinations based on possible codon choices, and all such variations should be considered to be specifically disclosed for any polypeptide disclosed herein.
[0114] In various embodiments, the codons are preferably selected to be compatible with the host cell in which the protein is being produced. For example, the preferred codons used in bacteria are those used for gene expression in bacteria, the preferred codons used in yeast are those used for expression in yeast, and the preferred codons used in mammals are those used for expression in mammalian cells. By way of example, the polynucleotide of SEQ ID NO: 3 (see below) encoding SEQ ID NO: 2 is codon-optimized for expression in E. coli.
[0115] In another aspect, the Applicant has discovered that specific changes in the triplet codons in the polynucleotide encoding the ketoreductase polypeptide of the present disclosure result in an increase in the expression of the polypeptide. These changes typically occur in the codons encoding amino acid residues towards the N-terminus of the polypeptide. For example, with respect to the polynucleotide of SEQ ID NO: 17 (see below) encoding SEQ ID NO: 1, the Applicant notes that changing the codons encoding the alanine residue (GCT) at position 2, the lysine residue (AAA) at position 3, the isoleucine residue (ATC) at position 4, and the glutamic acid residue (GAA) at position 11 of the polypeptide results in an increase in the expression of the polypeptide. In some embodiments, the change in the triplet codon at such a position does not result in a change in the amino acid residue of the encoded polypeptide. In other embodiments, the change in the triplet codon at such a position results in a change in the amino acid residue of the encoded polypeptide. Thus, in one embodiment, the present disclosure provides a polynucleotide encoding a polypeptide having at least 80% or 90% identity with SEQ ID NO: 2, with the following conditions: (a) The triplet codon encoding the amino acid residue at position 2 of the polypeptide is other than GCT, (b) The triplet codon encoding the amino acid residue at position 3 of the polypeptide is other than AAA, (c) The triplet codon encoding the amino acid residue at position 4 of the polypeptide is other than ATC, or (d) The triplet codon encoding the amino acid residue at position 11 of the polypeptide is other than GAA, wherein at least one of the above is satisfied.
[0116] In some embodiments, at least two of the conditions (a)-(d) are satisfied. In certain embodiments, at least three of the conditions (a)-(d) are satisfied. In certain embodiments, all four of the conditions (a)-(d) are satisfied.
[0117] In certain embodiments, because the native sequence contains preferred codons and the use of preferred codons may not be required for all amino acid residues, it is not necessary to replace all codons to optimize the codon usage of ketoreductase. As a result, a codon-optimized polynucleotide encoding a ketoreductase enzyme can contain preferred codons at about 40%, 50%, 60%, 70%, 80%, or more than 90% of the codon positions in the full-length coding region.
[0118] In various embodiments, an isolated polynucleotide encoding a modified ketoreductase polypeptide can be engineered in various ways to provide for expression of the polypeptide. Manipulation of the isolated polynucleotide prior to insertion into a vector can be desirable or necessary depending on the expression vector. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art. Guidance is provided in Sambrook et al., 2001, Molecular Cloning: A Laboratory Manual, 3rd Ed., Cold Spring Harbor Laboratory Press; and Current Protocols in Molecular Biology, Ausubel, F. ed., Greene Pub. Associates, 1998, updates to 2006.
[0119] In some embodiments, an isolated polynucleotide encoding any of the ketoreductase polypeptides herein is engineered in various ways to promote the expression of the ketoreductase polypeptide. In some embodiments, a polynucleotide encoding a ketoreductase polypeptide comprises an expression vector having one or more control sequences for regulating the expression of the ketoreductase polynucleotide and / or polypeptide. Manipulation of the isolated polynucleotide prior to insertion into the vector may be desirable or necessary depending on the expression vector utilized. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art. In some embodiments, control sequences include, among others, a promoter, a leader sequence, a polyadenylation sequence, a propeptide sequence, a signal peptide sequence, and a transcription terminator. In some embodiments, an appropriate promoter is selected based on host cell selection.In the case of bacterial host cells, suitable promoters for directing the transcription of the nucleic acid constructs of the present disclosure include the Escherichia coli lac operon, the Streptomyces coelicolor agarase gene (dagA), the Bacillus subtilis levansucrase gene (sacB), the Bacillus licheniformis alpha-amylase gene (amyL), the Bacillus stearothermophilus maltose-forming amylase gene (amyM), the Bacillus amyloliquefaciens alpha-amylase gene (amyQ), the Bacillus licheniformis penicillinase gene (penP), the Bacillus subtilis xylA and xylB genes, and promoters obtained from prokaryotic beta-lactamase genes (see, for example, Villa-Kamaroff et al., Proc. Natl Acad. Sci. USA 75:3727-3731
[1978] ), as well as the tac promoter (see, for example, DeBoer et al., Proc. Natl Acad. Sci. USA 80:21-25
[1983] ), but are not limited thereto.Exemplary promoters for filamentous fungal host cells include those obtained from the genes of Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral α - amylase, Aspergillus niger acid - stable α - amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase and Fusarium oxysporum trypsin - like protease (see, e.g., WO 96 / 00787), as well as the NA2 - tpi promoter (a hybrid of the promoters from the genes of Aspergillus niger neutral alpha - amylase and Aspergillus oryzae triose phosphate isomerase), and mutant promoters, truncated promoters, and hybrid promoters thereof, but are not limited thereto. Exemplary yeast cell promoters can be derived from the genes of Saccharomyces cerevisiae enolase (ENO - 1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde - 3 - phosphate dehydrogenase (ADH2 / GAP) and Saccharomyces cerevisiae 3 - phosphoglycerate kinase. Other useful promoters for yeast host cells are known in the art (see, e.g., Romanos et al., Yeast 8:423 - 488
[1992] ).
[0120] In some embodiments, the control sequence is also an appropriate transcription terminator sequence (i.e., a sequence recognized by the host cell to terminate transcription). In some embodiments, the terminator sequence is operably linked to the 3' end of the nucleic acid sequence encoding the enzyme polypeptide. Any suitable terminator that is functional in the selected host cell is used in the present invention. Exemplary transcription terminators for filamentous fungal host cells can be obtained from the genes of Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger alpha-glucosidase, and Fusarium oxysporum trypsin-like protease. Exemplary terminators for yeast host cells can be obtained from the genes of Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are known in the art (see, e.g., Romanos et al., supra).
[0121] In some embodiments, the control sequence is also a suitable leader sequence (i.e., the untranslated region of the mRNA that is important for translation by the host cell). In some embodiments, the leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding ketoreductase. Any suitable leader sequence that is functional in the selected host cell is used in the present invention. Exemplary leaders for filamentous fungal host cells are obtained from the genes of Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase. Leaders suitable for yeast host cells are obtained from the genes of Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae alpha-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).
[0122] In some embodiments, the control sequence is also a polyadenylation sequence (i.e., a sequence that is operably linked to the 3' end of the nucleic acid sequence and is recognized by the host cell as a signal to add polyadenosine residues to the transcribed mRNA when transcribed). Any suitable polyadenylation sequence that is functional in the selected host cell can be used in the present invention. Exemplary polyadenylation sequences for filamentous fungal host cells include, but are not limited to, the genes of Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger alpha-glucosidase. Polyadenylation sequences useful for yeast host cells are known (see, for example, Guo and Sherman, Mol. Cell. Biol., 15:5983-5990
[1995] ).
[0123] In some embodiments, the control sequence is also a signal peptide (i.e., a coding region that encodes an amino acid sequence linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the secretory pathway of a cell). In some embodiments, the 5' end of the coding sequence of the nucleic acid sequence essentially comprises a signal peptide coding region that is naturally linked in frame with a segment of the coding region encoding the secreted polypeptide. Alternatively, in some embodiments, the 5' end of the coding sequence comprises a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the secretory pathway of the optimal host cell is used for the expression of the engineered (one or more) polypeptides. Effective signal peptide coding regions for bacterial host cells include, but are not limited to, those obtained from the genes of Bacillus NClB 11837 maltogenic amylase, Bacillus stearothermophilus alpha-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus stearothermophilus neutral protease (nprT, nprS, nprM), and Bacillus subtilis prsA. Further signal peptides are known in the art (see, for example, Simonen and Palva, Microbiol. Rev., 57:109-137
[1993] ). In some embodiments, effective signal peptide coding regions for filamentous fungal host cells include, but are not limited to, those obtained from the genes of Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase.Useful signal peptides for yeast host cells include, but are not limited to, those derived from the Saccharomyces cerevisiae alpha-factor and the gene of Saccharomyces cerevisiae invertase.
[0124] In some embodiments, regulatory sequences are also utilized. These sequences facilitate the regulation of polypeptide expression as compared to the growth of the host cell. Examples of regulatory systems are those that turn gene expression on or off in response to chemical or physical stimuli, including the presence of regulatory compounds. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory systems include, but are not limited to, the ADH2 system or the GAL1 system. In filamentous fungi, suitable regulatory sequences include, but are not limited to, the TAKA alpha-amylase promoter, the Aspergillus niger glucoamylase promoter, and the Aspergillus oryzae glucoamylase promoter.
[0125] In another aspect, the present invention relates to a recombinant expression vector comprising a polynucleotide encoding a ketoreductase polypeptide and one or more expression regulatory regions such as a promoter, a terminator, an origin of replication, etc., depending on the type of host into which they are introduced. In some embodiments, one or more convenient restriction sites are included to allow the joining together of the various nucleic acids and control sequences described herein and the insertion or substitution at such sites of a nucleic acid sequence encoding an enzyme polypeptide to produce a recombinant expression vector. Alternatively, in some embodiments, the nucleic acid sequences of the present invention are expressed by inserting a nucleic acid sequence or a nucleic acid construct containing the sequence into a suitable vector for expression. In some embodiments involving the production of an expression vector, the coding sequence is placed into the vector such that the coding sequence is operably linked to a suitable control sequence for expression.
[0126] A recombinant expression vector can be any suitable vector (e.g., a plasmid or a virus) that is conveniently used in recombinant DNA procedures and can result in the expression of an enzyme polynucleotide sequence. The choice of vector typically depends on the compatibility of the vector with the host cell into which it is introduced. The vector can be a linear or closed circular plasmid.
[0127] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extrachromosomal entity whose replication is independent of chromosomal replication, e.g., a plasmid, an extrachromosomal element, a minichromosome or an artificial chromosome). The vector can include any means for ensuring self-replication. In some alternative embodiments, the vector is a vector that, when introduced into a host cell, is integrated into the genome and replicated together with the integrated chromosome(s). Further, in some embodiments, a single vector or plasmid, or two or more vectors or plasmids that together contain all the DNA introduced into the genome of the host cell, and / or a transposon are utilized.
[0128] In some embodiments, the expression vector contains one or more selectable markers that enable easy selection of the transformed cells. A "selectable marker" is a gene whose product provides, for example, biocide or virus resistance, resistance to heavy metals, prototrophy for auxotrophs, etc. Examples of bacterial selectable markers include, but are not limited to, the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers conferring antibiotic resistance such as ampicillin, kanamycin, chloramphenicol or tetracycline resistance. Markers suitable for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1 and URA3. Selectable markers for use in filamentous fungal host cells include amdS (acetamidase; e.g., from A. nidulans or A. oryzae), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase; e.g., from S. hygroscopicus), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase; e.g., from A. nidulans or A. oryzae), sC (sulfate adenylyltransferase), and trpC (anthranilate synthase), and equivalents thereof, but are not limited thereto.
[0129] In another aspect, the present invention provides a host cell comprising at least one polynucleotide encoding at least one ketoreductase of the present disclosure, wherein the (one or more) polynucleotide(s) is operably linked to one or more control sequences for the expression of at least one ketoreductase in the host cell. Host cells suitable for use in expressing the polypeptides encoded by the expression vectors of the present invention are well known in the art and include bacterial cells such as Escherichia coli, Vibrio fluvialis, Streptomyces and Salmonella typhimurium cells, fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC accession number 201178)), but are not limited thereto. Exemplary host cells include various Escherichia coli strains (e.g., W3110 (ΔfhuA) and BL21). Examples of bacterial selection markers include, but are not limited to, the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers conferring antibiotic resistance such as ampicillin, kanamycin, chloramphenicol and / or tetracycline resistance.
[0130] In some embodiments, the expression vector of the present invention comprises one or more elements that enable integration of the vector into the genome of the host cell or autonomous replication of the vector in cells independent of the genome. In some embodiments involving integration into the host cell genome, the vector relies on nucleic acid sequences encoding polypeptides or any other element of the vector for integration of the vector into the genome by homologous or non-homologous recombination.
[0131] In some alternative embodiments, the expression vector contains additional nucleic acid sequences for directing integration by homologous recombination into the genome of the host cell. The additional nucleic acid sequences enable the vector to be integrated into the host cell genome at an exact (one or more) position(s) in the (one or more) chromosome(s). To enhance the likelihood of integration at an exact position, the integration element preferably contains a sufficient number of nucleotides, such as 100 base pairs to 10,000 base pairs, preferably 400 base pairs to 10,000 base pairs, most preferably 800 base pairs to 10,000 base pairs, etc., which are highly homologous to the corresponding target sequence to enhance the probability of homologous recombination. The integration element can be any sequence homologous to the target sequence in the genome of the host cell. Further, the integration element can be a non-coding nucleic acid sequence or a coding nucleic acid sequence. On the other hand, the vector can be integrated into the genome of the host cell by non-homologous recombination.
[0132] In the case of autonomous replication, the vector can further contain an origin of replication that enables the vector to replicate autonomously in the host cell in question. Examples of bacterial origins of replication are P15A ori, or the origins of replication of plasmids pBR322, pUC19, pACYCl77 (containing P15A ori) or pACYC184 (containing P15A ori), which enable replication in Escherichia coli, and pUB110, pE194, or pTA1060 enable replication in Bacillus. Examples of origins of replication for use in yeast host cells are the 2 micron origin of replication, ARS1, ARS4, the combination of ARS1 and CEN3, and the combination of ARS4 and CEN6. The origin of replication can have a mutation that makes its function temperature-sensitive in the host cell (see, for example, Ehrlich, Proc. Natl. Acad. Sci. USA 75:1433
[1978] ).
[0133] In some embodiments, two or more copies of the nucleic acid sequences of the invention are inserted into a host cell to increase the production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene in the nucleic acid sequence, and cells containing amplified copies of the selectable marker gene and thereby additional copies of the nucleic acid sequence can be selected by culturing the cells in the presence of an appropriate selectable agent.
[0134] Many of the expression vectors used in the present invention are commercially available. Suitable commercially available expression vectors include, but are not limited to, the pET E. coli T7 expression vectors (Millipore Sigma) and p3xFLAGTM expression vectors (Sigma-Aldrich Chemicals) of Novagen®. Other suitable expression vectors include, but are not limited to, pBluescriptII SK(-) and pBK-CMV (Stratagene), and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen) or pPoly (see, for example, Lathe et al., Gene 57:193-201
[1987] ).
[0135] Accordingly, in some embodiments, a vector comprising an array encoding at least one variant ketoreductase is transformed into a host cell to enable growth of the vector and expression of the variant(s) ketoreductase. In some embodiments, the transformed host cell is cultured in a suitable nutrient medium under conditions that enable expression of the variant(s) ketoreductase. Any suitable medium useful for culturing the host cell, e.g., but not limited to, minimal medium or complex medium containing appropriate supplements, is used in the present invention. In some embodiments, the host cell is grown in HTP medium. Suitable media are available from various commercial suppliers or can be prepared according to published recipes (e.g., in the catalog of the American Type Culture Collection).
[0136] Host Cells for Expression of Ketoreductase In another aspect, the present disclosure provides a host cell comprising a polynucleotide encoding an improved ketoreductase polypeptide of the present disclosure, the polynucleotide being operably linked to one or more control sequences for the expression of the ketoreductase enzyme in the host cell. Host cells for use in the expression of the ketoreductase polypeptide encoded by the expression vectors of the present invention are well known in the art and include bacterial cells such as Escherichia coli, B. subtilis, B. licheniformis, B. megaterium, B. stearothermophilus, B. amyloliquefaciens, Lactobacillus kejir, Lactobacillus brevis, Lactobacillus minor, Streptomyces and Salmonella typhimurium; cells, fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC accession number 201178)), but are not limited thereto. Suitable culture media and growth conditions for the above host cells are well known in the art.
[0137] The polynucleotide for the expression of ketoreductase can be introduced into cells by various methods known in the art. Techniques include, inter alia, electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection and protoplast fusion. Various methods for introducing polynucleotides into cells will be apparent to those skilled in the art.
[0138] In some embodiments of the present invention, the filamentous fungal host cell is of any suitable genus and species including, but not limited to, Achlya, Acremonium, Aspergillus, Aureobasidium, Bjerkandera, Ceriporiopsis, Cephalosporium, Chrysosporium, Cochliobolus, Corynascus, Cryphonectria, Cryptococcus, Coprinus, Coriolus, Diplodia, Endothis, Fusarium, Gibberella, Gliocladium, Humicola, Hypocrea, Myceliophthora, Mucor, Neurospora, Penicillium, Podospora, Phlebia, Piromyces, Pyricularia, Rhizomucor, Rhizopus, Schizophyllum, Scytalidium, Sporotrichum, Talaromyces, Thermoascus, Thielavia, Trametes, Tolypocladium, Trichoderma, Verticillium, and / or Volvariella, and / or teleomorphs, or anamorphs, and synonyms, basionyms, or taxonomic equivalents thereof.
[0139] In some embodiments of the present invention, the host cell is a yeast cell including, but not limited to, cells of Candida species, Hansenula species, Saccharomyces species, Schizosaccharomyces species, Pichia species, Kluyveromyces species or Yarrowia species. In some embodiments of the present invention, the yeast cell is Hansenula polymorpha, Saccharomyces cerevisiae, Saccharomyces carlsbergensis, Saccharomyces diastaticus, Saccharomyces norbensis, Saccharomyces kluyveri, Schizosaccharomyces pombe, Pichia pastoris, Pichia finlandica, Pichia trehalophila, Pichia kodamae, Pichia membranaefaciens, Pichia opuntiae, Pichia thermotolerans, Pichia salictaria, Pichia quercuum, Pichia pijperi, Pichia stipitis, Pichia methanolica, Pichia angusta, Kluyveromyces lactis, Candida albicans, or Yarrowia lipolytica.
[0140] In some other embodiments, the host cell is a prokaryotic cell. Suitable prokaryotic cells include, but are not limited to, Gram-positive, Gram-negative, and Gram-variable bacterial cells. Any suitable bacterial organism used in the present invention includes Agrobacterium, Alicyclobacillus, Anabaena, Anacystis, Acinetobacter, Acidothermus, Arthrobacter, Azobacter, Bacillus, Bifidobacterium, Brevibacterium, Butyrivibrio, Buchnera, Campestris, Camplyobacter, Clostridium, Corynebacterium, Chromatium, Coprococcus, Escherichia, Enterococcus, Enterobacter, Erwinia, Fusobacterium, Faecalibacterium, Francisella, Flavobacterium, Geobacillus, Haemophilus, Helicobacter, Klebsiella, Lactobacillus, Lactococcus, Ilyobacter, Micrococcus, Microbacterium, Mesorhizobium, Methylobacterium, Methylobacterium, Mycobacterium,Neisseria, Pantoea, Pseudomonas, Prochlorococcus, Rhodobacter, Rhodopseudomonas, Rhodospirillum, Roseburia, Rhodococcus, Scenedesmus, Streptomyces, Streptococcus, Synecoccus, Saccharomonospora, Staphylococcus, Serratia, Salmonella, Shigella, Thermoanaerobacterium, Tropheryma, Tularensis, Temecula, Thermosynechococcus, Thermococcus, Ureaplasma, Xanthomonas, Xylella, Yersinia, and Zymomonas, but are not limited thereto. In some embodiments, the host cell is Agrobacterium, Acinetobacter, Azotobacter, Bacillus, Bifidobacterium, Brucella, Geobacillus, Campylobacter, Clostridium, Corynebacterium, Escherichia, Enterococcus, Erwinia, Flavobacterium, Lactobacillus, Lactococcus, Pantoea, Pseudomonas, Staphylococcus, Salmonella, Streptococcus, or Zymomonas. In some embodiments, the bacterial host strain is non-pathogenic to humans. In some embodiments, the bacterial host strain is an industrial strain. A number of bacterial industrial strains are known and suitable for the present invention. In some embodiments of the present invention, the bacterial host cell is an Agrobacterium species (e.g.,A. radiobacter, A. rhizogenes, and A. rubi. In some embodiments of the present invention, the bacterial host cell is an Arthrobacter species (e.g., A. aurescens, A. citreus, A. globiformis, A. hydrocarboglutamicus, A. mysorens, A. nicotianae, A. parafineus, A. protophonniae, A. roseoparqffinus, A. sulfureus, and A. ureafaciens). In some embodiments of the present invention, the bacterial host cell is a Bacillus species (e.g., B. thuringensis, B. anthracis, B. megaterium, B. subtilis, B. lentus, B. circulans, B. pumilus, B. lautus, B. coagulans, B. brevis, B. firmus, B. alkaophius, B. licheniformis, B. clausii, B. stearothermophilus, B. halodurans, and B. amyloliquefaciens). In some embodiments, the host cell is an industrial Bacillus strain including, but not limited to, B. subtilis, B. pumilus, B. licheniformis, B. megaterium, B. clausii, B. stearothermophilus, or B. amyloliquefaciens. In some embodiments, the Bacillus genus host cell is B. subtilis, B. licheniformis, B. megaterium, B. stearothermophilus, and / or B. amyloliquefaciens. In some embodiments, the bacterial host cell is a Clostridium species (e.g., C. acetobutylicum,C. tetani E88, C. lituseburense, C. saccharobutylicum, C. perfringens, and C. beijerinckii). In some embodiments, the bacterial host cell is a Corynebacterium species (e.g., C. glutamicum and C. acetoacidophilum). In some embodiments, the bacterial host cell is an Escherichia species (e.g., E. coli). In some embodiments, the host cell is E. coli W3110. In some embodiments, the host is E. coli BL21 or BL21(DE3). In some embodiments, the bacterial host cell is an Erwinia species (e.g., E. uredovora, E. carotovora, E. ananas, E. herbicola, E. punctata, and E. terreus). In some embodiments, the bacterial host cell is a Pantoea species (e.g., P. citrea and P. agglomerans). In some embodiments, the bacterial host cell is a Pseudomonas species (e.g., P. putida, P. aeruginosa, P. mevalonii, and P. sp. D-0l10). In some embodiments, the bacterial host cell is a Streptococcus species (e.g., S. equisimiles, S. pyogenes, and S. uberis). In some embodiments, the bacterial host cell is a Streptomyces species (e.g., S. ambofaciens, S. achromogenes, S. avermitilis, S. coelicolor, S. aureofaciens, S. aureus, S. fungicidicus,S. griseus and S. lividans. In some embodiments, the bacterial host cell is a Zymomonas species (e.g., Z. mobilis and Z. lipolytica).
[0141] Many of the prokaryotic and eukaryotic strains used in the present invention are readily available from a number of culture collections such as the American Type Culture Collection (ATCC), Deutsche Sammlung von Mikroorganismen und Zellkulturen GmbH (DSM), Centraalbureau Voor Schimmelcultures (CBS), and the Agricultural Research Service Patent Culture Collection, Northern Regional Research Center (NRRL).
[0142] In some embodiments, the host cells are genetically modified to have properties that improve protein secretion, protein stability and / or other properties desirable for protein expression and / or secretion. The genetic modification can be achieved by genetic engineering techniques and / or classical microbiological techniques (e.g., chemical or UV mutagenesis and subsequent selection). Indeed, in some embodiments, a combination of recombinant modification and classical selection techniques is used to produce the host cells. Recombinant techniques can be used to introduce, delete, inhibit or modify nucleic acid molecules so as to result in an increase in the yield of the (one or more) ketoreductase variants in the host cell and / or in the culture medium. In one genetic engineering approach, homologous recombination is used to induce targeted gene modification by specifically targeting genes in vivo to suppress the expression of the encoded proteins. In an alternative approach, siRNA, antisense and / or ribozyme technologies are used to inhibit gene expression. A variety of methods for reducing protein expression in cells are known in the art, including, but not limited to, deletion of all or part of the gene encoding the protein and site-directed mutagenesis to disrupt the expression or activity of the gene product. (See, e.g., Chaveroche et al., Nucl. Acids Res., 28:22 e97
[2000] ; Cho et al., Molec. Plant Microbe Interact., 19:7-15
[2006] ; Maruyama and Kitamoto, Biotechnol. Lett., 30:1811-1817
[2008] ; Takahashi et al., Mol. Gen. Genom., 272:344-352
[2004] ; and You et al., Arch. Microbiol., 191:615-622
[2009] , all of which are incorporated herein by reference).Random mutagenesis, followed by screening for the desired mutation, is also used (see, e.g., Combier et al., FEMS Microbiol. Lett., 220:141-8
[2003] ; and Firon et al., Eukary. Cell 2:247-55
[2003] , both of which are incorporated by reference).
[0143] Introduction of the vector or DNA construct into the host cell can be achieved using any suitable method known in the art, including but not limited to calcium phosphate transfection, DEAE-dextran mediated transfection, PEG mediated transformation, electroporation, or other common techniques known in the art.
[0144] In some embodiments, the engineered host cells of the invention (i.e., "recombinant host cells") are cultured in a conventional nutrient medium that has been appropriately modified for activation of the promoter, selection of transformants, or amplification of the ketoreductase polynucleotide. Culture conditions such as temperature, pH, etc. are those previously used with the host cells selected for expression and are well known to those skilled in the art. As noted above, many standard references and texts are available for the culture and production of many cells, including cells of bacterial, plant, animal (especially mammalian), and archaeal origin.
[0145] In some embodiments, cells expressing the ketoreductase of the invention are grown under batch or continuous fermentation conditions. Classical "batch fermentation" is a closed system in which the composition of the medium is set at the start of fermentation and is not subject to artificial change during fermentation. A variation of the batch system is "fed-batch fermentation", which is also used in the present invention. In this variant, substrate is gradually added as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit cell metabolism and it is desirable to have a limited amount of substrate in the medium. Batch and fed-batch fermentations are common and well known in the art. "Continuous fermentation" is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the culture at a constant high density at which the cells are mainly in the logarithmic growth phase. Continuous fermentation systems strive to maintain steady-state growth conditions. Methods for regulating nutrients and growth factors for continuous fermentation processes, as well as techniques for maximizing product formation rates, are well known in the field of industrial microbiology.
[0146] Two or more copies of the nucleic acid sequence of the invention may be inserted into a host cell to increase the production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene in the nucleic acid sequence, and cells containing amplified copies of the selectable marker gene and thereby additional copies of the nucleic acid sequence can be selected by culturing the cells in the presence of an appropriate selectable agent.
[0147] In some embodiments of the invention, cell-free transcription and translation systems are used for the production of (one or more) ketoreductases. Several systems are commercially available and this method is well known to those skilled in the art.
[0148] Methods for the evolution of ketoreductase In some embodiments, to produce the ketoreductase of the present disclosure, the ketoreductase enzyme that catalyzes the reduction reaction is obtained (or derived) from E. coli. In some embodiments, the parental polynucleotide sequence is codon-optimized to enhance the expression of ketoreductase in a particular host cell. The parental polynucleotide sequence, designated SEQ ID NO: 3, encoding SEQ ID NO: 2, was codon-optimized for expression in E. coli, the codon-optimized polynucleotide was cloned into an expression vector, and the expression of the ketoreductase gene was placed under the control of the T7 promoter. The T7 polymerase required to express the gene of interest is under the control of the lac promoter, and both the gene of interest and the T7 polymerase are subject to lacI repression. The presence of IPTG activates T7 polymerase production, eliminates repression, and results in the expression of the ketoreductase protein. Clones expressing active ketoreductase in E. coli were identified and the genes were sequenced to confirm their identity.
[0149] The ketoreductase of the present disclosure can be obtained by subjecting the polynucleotide encoding the parental sequence to mutagenesis and / or directed evolution methods. Exemplary directed evolution techniques are described in Stemmer, 1994, Proc. Natl. Acad. Sci. USA 91:10747-10751; WO 95 / 22625; WO 97 / 20078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767 and U.S. Pat. No. 6,537,746, mutagenesis and / or DNA shuffling. Other directed evolution procedures that can be used include, inter alia, the staggered extension process (StEP), in vitro recombination (Zhao et al., 1998, Nat. Biotechnol. 16:258-261), mutagenic PCR (Caldwell et al., 1994, PCR Methods Appl. 3:S136-S140), and cassette mutagenesis (Black et al., 1996, Proc. Natl. Acad. Sci. USA 93:3525-3529).
[0150] The clones obtained after mutagenesis treatment are screened for a ketoreductase having the desired improved enzyme properties. Measurement of enzyme activity from the expression library can be carried out using standard chemical analysis techniques (such as UPLC-MS) for measuring substrates and products, as well as standard biochemical techniques for monitoring the rate of decrease (through the decrease in absorbance or fluorescence) when the NADH or NADPH concentration is converted to NAD+ or NADP+. In this reaction, since the ketoreductase reduces the ketone substrate to the corresponding hydroxyl group, NADH or NADPH is consumed (oxidized) by the ketoreductase. The rate of decrease in the NADH or NADPH concentration, measured by the decrease in absorbance or fluorescence per unit time, indicates the relative (enzyme) activity of the ketoreductase polypeptide in a fixed amount of lysate (or lyophilized powder prepared therefrom). If the desired improved enzyme property is thermal stability, the enzyme activity can be measured after subjecting the enzyme preparation to a defined temperature and measuring the amount of enzyme activity remaining after heat treatment. Next, clones containing the polynucleotide encoding the ketoreductase are isolated, sequenced (if any) to identify nucleotide sequence changes, and used for expressing the enzyme in host cells.
[0151] When the sequence of a polypeptide is known, the polynucleotide encoding the enzyme can be prepared by standard solid-phase methods according to known synthetic methods. In some embodiments, fragments up to about 100 bases in length can be synthesized individually and then ligated (e.g., by enzymatic or chemical ligation methods, or polymerase-mediated methods) to form any desired contiguous sequence. For example, the polynucleotides and oligonucleotides of the present invention can be prepared by chemical synthesis using, for example, the classical phosphoramidite method described in Beaucage et al., 1981, Tet. Lett. 22:1859-69, or the method described in Matthes et al., 1984, EMBO J. 3:801-05, typically performed by an automated synthesis method. According to the phosphoramidite method, oligonucleotides are synthesized, purified, annealed, ligated, and cloned into an appropriate vector, for example, using an automated DNA synthesizer. Furthermore, essentially any nucleic acid can be obtained from any of a variety of commercial sources such as The Midland Certified Reagent Company (Midland, Texas), The Great American Gene Company (Ramona, California), ExpressGen Inc. (Chicago, Illinois), Operon Technologies Inc. (Alameda, California), and the like.
[0152] The ketoreductase enzyme expressed in the host cell can be recovered from the cell and / or the culture medium using any one or more of the well-known techniques for protein purification, including, inter alia, lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography. A solution suitable for lysis and high-efficiency extraction of proteins from bacteria such as E. coli is commercially available under the trade name CelLytic B® from Sigma-Aldrich, St. Louis, Missouri.
[0153] As chromatography techniques for isolating ketoreductase polypeptides, there are, among others, reverse-phase chromatography, high-performance liquid chromatography, ion-exchange chromatography, gel electrophoresis, and affinity chromatography. The conditions for purifying a specific enzyme depend, in part, on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and will be apparent to those skilled in the art.
[0154] In some embodiments, affinity techniques can be used to isolate improved ketoreductase enzymes. For affinity chromatography purification, the protein sequence can be tagged with a recognition sequence to enable purification. Common tags include cell-binding domains, poly-His tags, di-His chelates, FLAG tags, and many others that will be apparent to those skilled in the art. Antibodies can also be used as affinity purification reagents. Any antibody that specifically binds to the ketoreductase polypeptide can be used.
[0155] Process for using ketoreductase The ketoreductase enzymes described herein can catalyze the reduction of substrate-substituted indanone compounds such as fluoro-hydroxyindanone (6)
Chemical formula
[0156] to fluoro-diol (7)
Chemical formula
[0157]
[0158] In certain embodiments, the ketoreductase enzymes described herein are hydroxyindanone (5)
Chemical formula
[0159] It can be used immediately after the electrophilic fluorination to provide fluorodiol (7).
[0160] In some embodiments, the process for preparing fluorodiol (7) comprises contacting hydroxyindanone (5) with a fluorinating agent under acidic conditions to obtain fluoro-hydroxyindanone (6), and contacting fluoro-hydroxyindanone (6) with a ketoreductase disclosed herein under reaction conditions suitable for reducing or converting fluoro-hydroxyindanone (6) to fluorodiol (7). Fluorodiol (7) is an intermediate for the synthesis of belzutifan (WELIREG). Thus, in the process for preparing belzutifan, this process can include the step of converting hydroxyindanone (5) to fluorodiol (7) using a ketoreductase disclosed herein. In some embodiments, the product has a diastereomeric excess of greater than about 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8% or 99.9% relative to the corresponding (1R) alcohol product.
[0161] As described herein, in some embodiments, the ketoreductase can comprise an amino acid sequence that is at least about 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical when compared to a reference sequence comprising the sequence of SEQ ID NO: 2. In some embodiments, these ketoreductase polypeptides can have one or more modifications to the amino acid sequence of SEQ ID NO: 2. The modifications can include substitutions, deletions, and insertions. The substitutions can be non-conservative substitutions, conservative substitutions, or a combination of non-conservative and conservative substitutions.
[0162] In some embodiments of the method for reducing a substrate to a product, the substrate is reduced to the product with a diastereomeric excess of greater than 99%, and the ketoreductase polypeptide comprises a sequence corresponding to SEQ ID NO: 2.
[0163] In another embodiment of this method for reducing a substrate to a product, when carried out using a substrate in excess of about 60 g / L and a polypeptide of less than about 5 g / L, at least about 95% of the substrate is converted to the product in less than about 24 hours, and the polypeptide comprises an amino acid sequence corresponding to SEQ ID NO: 2.
[0164] As is known to those skilled in the art, ketoreductase-catalyzed reduction reactions typically require cofactors. The reduction reactions catalyzed by the ketoreductase enzymes described herein typically also require cofactors, although many embodiments of the engineered ketoreductases require far fewer cofactors than the reactions catalyzed by wild-type ketoreductase enzymes. As used herein, the term "cofactor" refers to a non-protein compound that acts in combination with a ketoreductase enzyme. Suitable cofactors for use with the engineered ketoreductase enzymes described herein include, but are not limited to, NADP+ (nicotinamide adenine dinucleotide phosphate), NADPH (the reduced form of NADP+), NAD+ (nicotinamide adenine dinucleotide), and NADH (the reduced form of NAD+). Generally, the reduced form of the cofactor is added to the reaction mixture. The reduced NAD(P)H form may be regenerated from the oxidized NAD(P)+ form using a cofactor regeneration system.
[0165] The term "cofactor regeneration system" refers to a series of reactants involved in the reaction that reduces the oxidized form of a cofactor (e.g., from NADP+ to NADPH). The cofactor oxidized by the ketoreductase-catalyzed reduction of a keto substrate is regenerated in its reduced form by the cofactor regeneration system. The cofactor regeneration system is a source of reducing hydrogen equivalents and includes a stoichiometric reducing agent capable of reducing the oxidized form of the cofactor. The cofactor regeneration system may further include a catalyst, such as an enzyme catalyst that catalyzes the reduction of the oxidized form of the cofactor by the reducing agent. In some embodiments, the ketoreductase of the present disclosure itself may function as an enzyme catalyst that catalyzes the reduction of the oxidized form of the cofactor by the reducing agent. Thus, in such embodiments, the ketoreductase serves a dual catalytic role. Cofactor regeneration systems that regenerate NADH or NADPH from NAD+ or NADP+ respectively are known in the art and can be used in the methods described herein.
[0166] Suitable exemplary cofactor regeneration systems that can be used include, but are not limited to, glucose and glucose dehydrogenase, formic acid and formic acid dehydrogenase, glucose-6-phosphate and glucose-6-phosphate dehydrogenase, secondary (e.g., isopropanol) alcohols and secondary alcohol dehydrogenase, phosphorous acid and phosphorous acid dehydrogenase, molecular hydrogen and hydrogenase, etc. These systems can be used in combination with either NADP+ / NADPH or NAD+ / NADH as a cofactor. As the cofactor regeneration system, electrochemical regeneration using hydrogenase may also be used. See, for example, U.S. Patent Nos. 5,538,867 and 6,495,023 (both of which are incorporated herein by reference). Chemical cofactor regeneration systems containing a metal catalyst and a reducing agent (e.g., molecular hydrogen or formate) are also suitable. See, for example, International Publication No. WO 2000 / 053731, which is incorporated herein by reference.
[0167] The terms "glucose dehydrogenase" and "GDH" are used interchangeably herein to refer to NAD+ or NADP+-dependent enzymes that catalyze the respective conversions of D-glucose and NAD+ or NADP+ to gluconic acid and NADH or NADPH.
[0168] Glucose dehydrogenases suitable for use in the practice of the methods described herein include both naturally occurring glucose dehydrogenases and non-naturally occurring glucose dehydrogenases. Genes encoding naturally occurring glucose dehydrogenases have been reported in the literature. For example, the Bacillus subtilis 61297 GDH gene was expressed in Escherichia coli and reported to exhibit the same physicochemical properties as the enzyme produced in its native host (Vasantha et al., 1983, Proc. Natl. Acad. Sci. USA 80:785). The gene sequence of the B. subtilis GDH gene corresponding to Genbank accession number M12276 was reported by Lampel et al., 1986, J. Bacterial. 166:238-243 and reported in a corrected form by Yamane et al., 1996, Microbiology 142:3047-3056 as Genbank Acc. No. D50453. Naturally occurring GDH genes also include those encoding GDH from B. cereus ATCC 14579 (Nature, 2003, 423:87-91; Genbank accession number AE017013) and B. megaterium (Eur. J. Biochem., 1988, 174:485-490, Genbank accession number X12370; J. Ferment. Bioeng., 1990, 70:363-369, Genbank accession number GI216270). Glucose dehydrogenases from Bacillus species are provided in PCT Publication WO 2005 / 018579, the disclosure of which is incorporated herein by reference.
[0169] Non-naturally occurring glucose dehydrogenases can be produced using known methods such as mutagenesis, directed evolution, etc. GDH enzymes having appropriate activity can be readily identified by those skilled in the art, including using the assays described in Example 4 of International Publication No. WO 2005 / 018579, the disclosure of which is incorporated herein by reference, whether they are naturally occurring or non-naturally occurring.
[0170] The ketoreductase-catalyzed reduction reaction described herein is generally carried out in a solvent. Suitable solvents include water, organic solvents (e.g., ethyl acetate, butyl acetate, 1-octanol, heptane, octane, methyl t-butyl ether (MTBE), toluene, etc.), and ionic liquids (e.g., 1-ethyl-4-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium hexafluorophosphate, etc.). In some embodiments, an aqueous solvent comprising water and an aqueous co-solvent system is used.
[0171] Exemplary aqueous co-solvent systems have water and one or more organic solvents. Generally, the organic solvent component of the aqueous co-solvent system is selected so as not to completely inactivate the ketoreductase enzyme. Suitable co-solvent systems can be readily identified by measuring the enzyme activity of a specific engineered ketoreductase enzyme using a defined substrate of interest in a candidate solvent system using an enzyme activity assay such as those described herein.
[0172] The organic solvent component of the aqueous cosolvent system is miscible with the aqueous component and may provide a single liquid phase, or is partially miscible or immiscible with the aqueous component and may provide two liquid phases. Generally, when an aqueous cosolvent system is used, it is selected to be biphasic, with water dispersed in the organic solvent or vice versa. Generally, when utilizing an aqueous cosolvent system, it is desirable to select an organic solvent that can be easily separated from the aqueous phase. Generally, the ratio of water to organic solvent in the cosolvent system is typically in the range of about 90:10 to about 10:90 (volume / volume) of organic solvent to water, or 80:20 to 20:80 (volume / volume) of organic solvent to water. The cosolvent system may be preformed before being added to the reaction mixture or may be formed in situ within the reaction vessel.
[0173] The aqueous solvent (water or an aqueous cosolvent system) can be pH buffered or unbuffered. Generally, reduction can be carried out at a pH in the range of about 10 or less, usually about 5 to about 10. In some embodiments, reduction is carried out at a pH of about 9 or less, usually in the range of about 5 to about 9. In some embodiments, reduction is carried out at a pH of about 8 or less, often in the range of about 5 to about 8, usually in the range of about 6 to about 8. Reduction can be carried out at a pH of about 7, 8 or less, or 7.5 or less. Alternatively, reduction may be carried out at a neutral pH, i.e., about 7.
[0174] During the course of the reduction reaction, the pH of the reaction mixture can change. The pH of the reaction mixture may be maintained within the desired pH or desired pH range by adding an acid or base during the course of the reaction. Alternatively, the pH can be controlled by using an aqueous solvent containing a buffer. Suitable buffers for maintaining the desired pH range are known in the art and include, for example, phosphate buffers, triethanolamine buffers, etc. Combinations of buffering and the addition of an acid or base can also be used.
[0175] When a glucose / glucose dehydrogenase cofactor regeneration system is used, the coproduction of gluconic acid (pKa = 3.6) as represented by formula (1) will lower the pH of the resulting aqueous gluconic acid solution if it is not neutralized in another way. The pH of the reaction mixture may be maintained at a desired level by standard buffering techniques, and the buffer neutralizes the gluconic acid up to the provided buffering capacity or neutralizes the gluconic acid by adding a base simultaneously with the conversion process. A combination of buffering and base addition may also be used. Suitable buffers for maintaining the desired pH range are described above. Bases suitable for neutralizing gluconic acid are organic bases such as amines, alkoxides, etc., and inorganic bases such as hydroxide salts (e.g., NaOH), carbonate salts (e.g., NaHCO3), bicarbonate salts (e.g., K2CO3), basic phosphate salts (e.g., K2HPO4, Na3PO4), etc. Adding a base simultaneously with the conversion process may be done manually while monitoring the pH of the reaction mixture, or more conveniently, by using an automatic titrator as a pH-stat. A combination of partial buffering capacity and base addition can also be used for process control.
[0176] When using base addition to neutralize the gluconic acid released during the ketoreductase-catalyzed reduction reaction, the progress of the conversion may be monitored by the amount of base added to maintain the pH. Typically, the base added to the unbuffered or partially buffered reaction mixture during the reduction process is added in an aqueous solution.
[0177] In some embodiments, the cofactor regeneration system can include formate dehydrogenase. The terms "formate dehydrogenase" and "FDH" are used interchangeably herein to refer to NAD+ or NADP+-dependent enzymes that catalyze the conversion of formate and NAD+ or NADP+ to carbon dioxide and NADH or NADPH, respectively. Suitable formate dehydrogenases for use as a cofactor regeneration system in the ketoreductase-catalyzed reduction reactions described herein include both naturally occurring formate dehydrogenases and non-naturally occurring formate dehydrogenases. Examples of formate dehydrogenases include those corresponding to SEQ ID NO: 70 (Pseudomonas species) and SEQ ID NO: 72 (Candida boidinii) of PCT Publication WO 2005 / 018579, which are encoded by polynucleotide sequences corresponding to SEQ ID NO: 69 and SEQ ID NO: 71 of PCT Publication WO 2005 / 018579, the disclosures of which are incorporated herein by reference. The formate dehydrogenase used in the methods described herein, whether naturally occurring or non-naturally occurring, can exhibit an activity of at least about 1 μmol / min / mg, sometimes at least about 10 μmol / min / mg, or at least about 10 2 μmol / min / mg, and can have a maximum of about 10 3 μmol / min / mg or more, and can be readily screened for activity in the assay described in Example 4 of PCT Publication WO 2005 / 018579.
[0178] As used herein, the term "formate" refers to formate anion (HCO2 - ), formic acid (HCO2H), and mixtures thereof. Formate can be provided in the form of salts, typically alkali salts or ammonium salts (e.g., HCO2Na, KHCO2, NH4HCO2, etc.), formic acid, typically an aqueous solution of formic acid, or mixtures thereof. Formic acid is a moderately strong acid. In an aqueous solution within a few pH units of its pKa (pKa = 3.7 in water), formate is present at equilibrium as HCO2 -and exists as both HCO2H. At pH values above about pH 4, the formate is mainly HCO2 - and exists as. When formate is provided as formic acid, the reaction mixture is typically buffered or made less acidic by adding a base to provide a desired pH, typically of about pH 5 or higher. Bases suitable for neutralizing formic acid include, but are not limited to, organic bases such as amines, alkoxides, etc., and inorganic bases. When using a glucose / glucose dehydrogenase cofactor regeneration system, the coproduction of gluconic acid (pKa = 3.6) will lower the pH of the resulting aqueous gluconic acid solution if it is not neutralized in some other way. The pH of the reaction mixture may be maintained at the desired level by standard buffering techniques, and the buffer may neutralize gluconic acid up to the buffering capacity provided or neutralize gluconic acid by adding a base simultaneously with the conversion process. Combinations of buffering and base addition may also be used. Suitable buffers for maintaining the desired pH range are described above. Bases suitable for neutralization are, for example, hydroxide salts (e.g., NaOH), carbonates (e.g., NaHCO3), bicarbonates (e.g., K2CO3), basic phosphates (e.g., K2HPO4, Na3PO4), etc.
[0179] When formic acid and formate dehydrogenase are used as a cofactor regeneration system, the pH of the reaction mixture can be maintained at the desired level by standard buffering techniques, and the buffer can release protons up to the buffering capacity provided or release protons by adding an acid simultaneously with the conversion process. Acids suitable for addition during the reaction to maintain the pH include organic acids such as carboxylic acids, sulfonic acids, phosphonic acids, etc., mineral acids such as hydrohalic acids (e.g., hydrochloric acid), sulfuric acid, phosphoric acid, etc., acidic salts such as dihydrogen phosphates (e.g., KH2PO4), bisulfates (e.g., NaHSO4), etc. Some embodiments utilize formic acid, thereby maintaining both the formate concentration and the pH of the solution.
[0180] When using acid addition to maintain pH during the reduction reaction using a formic acid / formate dehydrogenase cofactor regeneration system, the progress of the conversion may be monitored by the amount of acid added to maintain the pH. Typically, the acid added to the unbuffered or partially buffered reaction mixture during the conversion is added in an aqueous solution.
[0181] When carrying out the embodiments of the ketoreductase-catalyzed reduction reaction described herein using a cofactor regeneration system, either the oxidized or reduced cofactor may be provided first. As described above, the cofactor regeneration system converts the oxidized cofactor to its reduced form, which is then utilized for the reduction of the ketoreductase substrate.
[0182] In some embodiments, a cofactor regeneration system is not used. For reduction reactions carried out without using a cofactor regeneration system, the cofactor is added to the reaction mixture in its reduced form.
[0183] In some embodiments, if the process is carried out using the whole cells of a host organism, the whole cells may naturally provide the cofactor. Alternatively, or in combination, the cells may naturally or recombinantly provide glucose dehydrogenase.
[0184] When performing the stereoselective reduction reaction described in this specification, a ketoreductase enzyme and any enzyme containing an optional cofactor regeneration system may be added to the reaction mixture in the form of a purified enzyme, whole cells transformed with the gene(s) encoding the enzyme, and / or cell extracts and / or lysates of such cells. The gene(s) encoding the ketoreductase enzyme and any cofactor regeneration enzyme can be transformed into the host cell separately or together into the same host cell. For example, in some embodiments, one set of host cells can be transformed with the gene(s) encoding the ketoreductase enzyme, and another set can be transformed with the gene(s) encoding the cofactor regeneration enzyme. The transformed cells of both sets can be used together in the reaction mixture in the form of whole cells or in the form of lysates or extracts derived therefrom. In other embodiments, the host cell can be transformed with the gene(s) encoding both the ketoreductase enzyme and the cofactor regeneration enzyme.
[0185] Whole cells transformed with the gene(s) encoding the ketoreductase enzyme and / or any cofactor regeneration enzyme, or cell extracts and / or lysates thereof, can be used in a variety of different forms including solids (e.g., freeze-dried, spray-dried, etc.) or semi-solids (e.g., crude paste).
[0186] Cell extracts or cell lysates can be partially purified by precipitation (ammonium sulfate, polyethyleneimine, heat treatment, etc.) followed by a desalting procedure (e.g., ultrafiltration, dialysis, etc.) prior to freeze-drying. Any cell preparation may also be stabilized by cross-linking using a known cross-linking agent.
[0187] In some embodiments, a ketoreductase enzyme with improved purity may be desired. In such embodiments, the clarified cell lysate used to obtain the ketoreductase enzyme may be pretreated with isopropanol to bring the volume percentage of isopropanol to 25% to 30%. The isopropanol-treated lysate may be incubated at 30 °C for 1 hour to overnight. The isopropanol-treated lysate can then be centrifuged to remove the pellet. The supernatant may then be transferred to a Petri dish and frozen at -80 °C for at least 2 hours. The sample may then be lyophilized using a standard automated protocol. As shown in Figure 1, gel electrophoresis using sodium dodecyl sulfate (SDS, also known as lauryl sulfate) and polyacrylamide gel (also known as SDS-PAGE) shows the removal of proteins insoluble in isopropyl alcohol (iPrOH), and a purified ketoreductase is obtained. In Figure 1, lane 1 (indicated as standard) is a marker, lane 2 (indicated as P012024-B07) is a crude enzyme preparation, lane 3 (indicated as B07, 30% IPA / 30% IPA / 30C / 1 hour) is a crude enzyme preparation treated with a 30% iPrOH solution for 1 hour and centrifuged, and lane 4 (indicated as 4 pellet - B07 / 30% IPA) is the solid from centrifugation. As seen in Figure 1, lane 3 shows the removal of many bands from the enzyme preparation when treated with iPrOH. Lane 4 shows the protein that precipitated and was removed by centrifugation.
[0188] Suitable conditions for performing the ketoreductase-catalyzed reduction reaction described herein include, but are not limited to, contacting the ketoreductase enzyme and the substrate operated at the experimental pH and temperature, and a wide variety of conditions that can be easily optimized by routine experiments including detecting the product using, for example, the methods described in the examples provided herein.
[0189] Ketoreductase-catalyzed reduction is typically carried out at a temperature in the range of about 15 °C to about 75 °C. For some embodiments, the reaction is carried out at a temperature in the range of about 20 °C to about 55 °C. In still other embodiments, the reaction is carried out at a temperature in the range of about 20 °C to about 45 °C. The reaction may be carried out under ambient conditions.
[0190] The reduction reaction is generally allowed to proceed until substantially complete or nearly complete reduction of the substrate is obtained. The reduction of the substrate to the product can be monitored using known methods by detecting the substrate and / or the product. Suitable methods include gas chromatography, HPLC, and the like. The conversion yield of the alcohol reduction product formed in the reaction mixture is generally greater than about 50%, may be greater than about 60%, may be greater than about 70%, may be greater than about 80%, may be greater than 90%, and often greater than about 97%.
[0191] [Examples] Example 1: Preparation of Enzyme E. coli cultures each having a plasmid encoding a ketoreductase enzyme that can be represented by the amino acid sequences shown in SEQ ID NOs: 1, 2, 4 to 16 above were serially diluted to 10 -4 10 -5 and 10 -6 using Luria-Bertani broth (cell culture medium) as a diluent. 100 μL of each dilution was spread on a Petri dish containing LB agar supplemented with 50 μg / mL kanamycin. The plates were placed in an incubator at 30 °C overnight.
[0192] 200 μL / well of Luria-Bertani broth (cell culture medium) (500 mL LB + 50 μg / mL kanamycin) was dispensed into a labeled 96-well shallow well plate. The shallow well plate was loaded into the plate stacker of the colony picker. An agar plate containing colonies that were sufficiently diluted so that most of the colonies were isolated from each other (known to those skilled in the art as single colonies) was picked into the respective wells of the shallow well plate. The colonies were grown overnight at 200 rpm, 30 °C, and 85% RH.
[0193] 390 μL of terrific broth (TB) growth medium (commercially available from ThermoFisher Scientific under catalog number A1374301) (TB + 50 μg / mL kanamycin) was aliquoted into a labeled 96-well deep well subculture plate. 13 μL of the overnight growth culture was transferred from each well of the master shallow well plate to the corresponding labeled deep well subculture plate. The plate was sealed with a breathable film and the plate was shaken at 250 rpm, 30 °C, and 85% RH for 2 - 2.5 hours. After shaking, the optical density (OD 600 , optical density at a wavelength of 600 nm) of at least one plate was measured for growth. When the OD 600 of this plate was in the range of 0.4 - 0.8, the deep well plate was induced with 4 μL per well of 1 M IPTG solution. The plate was resealed and incubated at 250 rpm, 30 °C, and 85% RH for 18 - 20 hours with shaking.
[0194] After incubation, all deep well plates were centrifuged at 4 °C and 4000 rpm for 15 minutes. After centrifugation, the supernatant was discarded. The plate was heat-sealed and stored at -80 °C.
[0195] The cell plate was removed from -80 °C storage and thawed at room temperature. 100 mM potassium phosphate pH 8.0, 1 mg / mL lysozyme, 0.50 mg / mL polymyxin B sulfate (PMBS), 3 units / mL DNase I, 4 mM MgCl2, and 1 mg / mL NADP+ A lysis buffer was prepared. 200 μL of the lysis buffer was dispensed into each well. The lysis mixture was shaken on a plate shaker at room temperature at 1000 rpm for 1 - 1.5 hours. Then, the lysis mixture was centrifuged at 4°C and 4000 rpm for 15 minutes to prepare an enzyme - containing lysate solution (in the supernatant). In some examples, the lysate solution was further incubated with isopropanol in a 25 - 30% isopropanol solution for 1 - 5 hours. After incubation, the lysate was centrifuged as described above.
[0196] Example 2: Ketoreductase reaction in a well plate
Chemical formula
[0197] A reaction buffer was prepared by resuspending a solid - flow substrate (the impure product (6) of the previous chemical fluorination reaction, see Example 5 below) to a concentration of 75 - 100 g / L in a 16:16:10:58 (volume / volume) acetonitrile:methanol:isopropanol:potassium phosphate buffer pH 8.0 solution. Then, the pH was further adjusted to 8.0 using an aqueous sodium hydroxide solution.
[0198] 80 μL of the reaction buffer was added to a 0.3 mL round - bottom well plate, followed by 20 μL of the enzyme - containing lysate solution of Example 1. The plate was heat - sealed and shaken overnight at 35°C and 1000 rpm.
[0199] After shaking overnight, 30 μL from the reaction mixture was added to a new round-bottom well plate containing 240 μL of acetonitrile. This mixture was aged for 1 hour, at which point 30 μL was added on top of a water filter stack (a filter plate on top of a round-bottom plate containing 0.20 μM hydrophilic PTFE, commercially available from Millipore MSRLN2250) containing an additional 240 μL of 20% (v / v) acetonitrile in water. The plate was then centrifuged at 4 °C and 4000 rpm for 3 minutes. The filter plate was removed and the clear solution in the receiving plate was heat-sealed.
[0200] The filtered solution was analyzed by ultra-high performance liquid chromatography (UPLC) using a high-throughput screening method to monitor the substrate and substrate degradation, as well as the product peak areas. UPLC was performed on an Agilent instrument equipped with a Waters HSS T3 1.8 μm, 2.1 × 75 mm column at a flow rate of 1 mL / min over 1.1 minutes using an isocratic method of 14% CH3CN + 0.1% TFA / H2O + 0.1% TFA. The two diastereomers of the starting material eluted at 0.53 minutes and 0.66 minutes, respectively, the diastereomer of the desired product (7) eluted at 0.46 minutes, and the diastereomers of the undesired products (7-1, 7-2, 7-3) eluted at 0.48 minutes, 0.88 minutes, and 0.47 minutes, respectively. Diastereoselectivity was determined by supercritical fluid chromatography using an Agilent instrument with a Daicel CHIRALPAK IG-3 3.0 μm, 50 × 4.6 mm column. Mobile phase A was supercritical CO2 and mobile phase B was isopropanol. The flow rate was 2.5 mL / min with a linear gradient of 18 - 37% B over 1.5 minutes, 37% B for 0.2 minutes, a linear gradient back to 18% B over 0.05 minutes, and an equilibration time of 0.15 minutes at 18% B (total time 1.9 minutes). The two diastereomers of the starting material eluted at 0.79 minutes and 1.6 minutes, the diastereomer of the desired product (7) eluted at 1.1 minutes, and the diastereomers of the undesired products (7-1, 7-2, 7-3) eluted at 1.3 minutes, 0.9 minutes, and 0.6 minutes, respectively.
[0201] Example 3: Enzyme Preparation in a Shaking Flask Inoculate 10 μL of E. coli cells into 5 mL of Luria - Bertani broth (cell culture medium) (250 mL LB + 50 μg / ml kanamycin + 1% glucose), which had been stored at -80 °C in 20% glycerol and dispensed into labeled 15 mL cell culture tubes, and has plasmids encoding ketoreductase enzymes that can be represented by the amino acid sequences shown in SEQ ID NOs: 1, 2, and 4 - 16 below. Seal the cell culture tubes and incubate with shaking at 30 °C and 250 rpm for 20 - 24 hours.
[0202] After overnight growth, an overnight growth culture (2 - 5 mL of cell culture (having an initial OD of 0.2)) was added to 250 mL of terrific broth (TB) growth medium (commercially available from ThermoFisher Scientific under catalog number A1374301) (TB + 50 μg / ml kanamycin) to bring the final volume to 250 mL. Shake the flask at 250 rpm and 30 °C for 3 - 4 hours. After shaking, measure the OD for growth until the OD reaches 0.4 - 0.6. 600 At this point, add 1 mM IPTG (250 μL of 1 M IPTG) to the culture to induce expression and grow the culture at 250 rpm and 30 °C for 20 - 24 hours. 600 600
[0203] After the additional growth period, transfer the culture to a centrifuge bottle of known weight and centrifuge at 4 °C and 4000 rpm for 20 minutes. After centrifugation, discard the supernatant and weigh the remaining cell pellet in the bottle. Calculate the weight of the cell pellet by subtracting the weight of the known bottle, and resuspend the cell pellet in 5 volumes of 50 mM sodium phosphate buffer (pH = 7) at a volume 5 times the weight of the cell pellet.
[0204] Cells from the resuspended cell pellet were lysed using a microfluidizer, the cell lysate was recovered, and centrifuged at 10,000 rpm for 60 minutes at 4°C. The clarified lysate was further treated with isopropanol in a 25 - 30% isopropanol solution for 1 - 5 hours, with occasional additional treatment. After incubation, the lysate was centrifuged as described above. The clarified supernatant was transferred to a Petri dish and frozen at -80°C for approximately 2 hours. The sample was lyophilized using a standard automated protocol.
[0205] Example 4: Ketoreductase Reaction 0.90 g of NADP was added to 60 mL of an aqueous solution of 200 mM K2HPO4 +It was dissolved, and subsequently 2.4 g of the enzyme powder prepared in Example 3 was added to the solution. 135 mL of isopropanol was added to the solution, and the pH was adjusted to 8.0 with 5N aqueous NaOH solution. The above enzyme solution was added to a quenched chemical fluorination reaction mixture containing 60 g of the fluoroketone substrate (6). The temperature was set at 33 °C, and the reaction mixture was stirred for 18 hours. The enzyme reaction mixture was quenched with 900 mL of ethyl acetate, cooled to 20 °C, and filtered. The biphasic mixture obtained from the filtrate was separated, and the organic layer was washed with 40% (weight / volume) aqueous (NH4)2SO4 solution (2 × 240 mL), and then washed with 25% (weight / volume) aqueous K2HPO4 solution (2 × 120 mL). The volume of the organic matter was reduced to 900 mL by distillation and mixed with 500 mL of ethyl acetate. The volume was reduced to 900 mL again, mixed with 200 mL of ethyl acetate, and concentrated to 900 mL by the third distillation. CUNO#5 (7.8 g) was added to the saturated solution, and the mixture was stirred at room temperature for 45 minutes and then filtered through CELITE. The filtrate was distilled to 600 mL at 60 °C and then cooled to 30 °C over 45 minutes. The supersaturated solution was stirred at 350 rpm, and the product crystals (600 mg) were seeded. Toluene (518 mL) was added over 4 hours, and the mixture was distilled to 360 mL at 29 - 33 °C over 1.5 hours and then cooled to 20 °C over 2 hours. After aging at 20 °C for 14 hours, the slurry was filtered, and the solid was washed with 120 mL of 1:4 (volume / volume) ethyl acetate:toluene and then 120 mL of toluene. The wet cake was placed in a vacuum at 45 °C and dried for 1 day to obtain fluorodiol (7).
[0206] Example 5: Preparation of (1S,2R,3S)-2,4-difluoro-7-(methylsulfonyl)-2,3-dihydro-1H-indene-1,3-diol (7)
Chemical formula
[0207] Acetonitrile (2.5 L), methanol (2.5 L), (5) (1.0 kg, 1.0 equivalent), methanesulfonic acid (118 g, 0.3 equivalent), and SELECTFLUOR (1.596 kg, 1.1 equivalents) were charged into a 5-gallon Hastelloy C-276 reactor. The vessel was pressurized and purged with N2 a total of 5 times, and then the resulting mixture was stirred and heated to 60 °C for 16 hours. Methyl acetoacetate (95 g, 0.2 equivalent) and water (369 g, 5 equivalents) were charged, and the batch was further aged at 60 °C for 2 hours. The batch was then cooled to 20 °C, and 0.2 M K2HPO4 (7.75 L) was added. 50 wt% NaOH (344 g, 1.05 equivalents) was added over 15 minutes to neutralize the acidic solution (final pH = 6.3). The mixture was stirred and then transferred to a high-density polyethylene carboy. The reaction mixture was then transferred to a 30-L glass reactor under vacuum. Isopropanol (2.25 L) was added to the reaction, followed by a solution containing KRED and NADP (1.5 L of 0.2 M K2HPO4 containing 40 g of KRED and 15 g of NADP). The pH of the solution was adjusted by adding 5 N NaOH (288 g, 0.3 equivalents) over 25 minutes (final pH = 8.1). The solution was stirred and heated to 34 °C for 17 hours. The reaction volume was reduced to 10 L by batch concentration. When the batch volume reached 10 L, D.I. H2O (1 L) was added, and the batch was concentrated to a total volume of 10 L. K2CO3-treated CELITE (44 g, 0.044 weight / weight) and (NH4)2SO4 (5.8 kg, 5.8 weight / weight) were charged into the reactor, the batch was stirred, and heated to 50 °C for 2 hours. EtOAc (15 L) was charged at 50 °C and mixed for 30 minutes. The reaction was then cooled to 20 °C and aged for 30 minutes. The slurry was filtered through a 2-inch filter pot lined with polypropylene and filter paper, and the waste cake was washed with EtOAc (2 × 5 L). The filtrates were combined in a 50-L glass reactor, the aqueous phase was drained and discarded. The organic layer was washed successively with 40% (weight / weight) (NH4)2SO4 (2 × 4 L), 25% (weight / weight) K2HPO4 (2 L), and 50% (weight / weight) K2HPO4 (2 L). The resulting organic matter was concentrated to 15 L, and fresh EtOAc was added until the water content was less than 0.3 wt%, and distilled at a constant volume.Transfer the batch to a 50 L round-bottom flask, add CUNO #5 (130 g, 0.13 wt / wt), and stir the mixture for 30 minutes under ambient conditions. Filter the slurry through a 10-inch filter pot lined with polypropylene and filter paper, and wash the filter cake with EtOAc (2 × 2 L). Transfer the filtrate to a 20 L glass container, concentrate it to 10 L at 60 °C, then cool it to 30 °C and seed with (7) (10 g, 0.01 wt / wt). Age the resulting slurry at 30 °C for 1 hour, then add toluene (8.5 L) over 4 hours and age for an additional 9 hours. Concentrate the mixture to 6 L, then add fresh toluene (2 L) and distill at a constant volume. Cool the slurry to 20 °C for 30 minutes, age for 1 hour, and filter. Wash the product cake with 1:4 EtOAc / toluene (2 L) and toluene (2 L), then dry in a 40 °C vacuum oven to obtain (7) as an off-white solid (950 g, 85% yield, 96.7 area % by LC, 96.3 wt%). 1 H NMR (599.90 MHz, DMSO-d6) δ 7.92 (ddd, J = 8.6, 4.7 and 0.7 Hz, 1H, CH), 7.46 (t, J = 8.9 Hz, 1H, CH), 6.14 (d, J = 7.1 Hz, 1H, OH), 5.96 (d, J = 6.9 Hz, 1H, OH), 5.56 (dd, J = 6.8, 5.2 and 3.2 Hz, 1H, CH), 5.40 (ddd, J = 14.1, 7.0 and 5.2 Hz, 1H, CH), 4.89 (dt, J = 51.1 and 5.2 Hz, 1H, CH), 3.31 (s, 3H, CH3) ppm. 13 C{ 1 H} NMR (150.85 MHz, DMSO-d6) δ 162.31 (d, J CF = 258.7 Hz, CF), 142.85 (dd, J CF = 6.0 and 2.9 Hz, C), 133.89 (d, J CF = 3.4 Hz, C), 132.20 (d, J CF = 8.9 Hz, C), 130.53 (dd, J CF = 16.4 and 10.1 Hz, CH), 117.25 (d, J CF = 21.4 Hz, CH), 97.82 (d, J CF= 194.0 Hz, CH), 73.17 (d, J CF = 25.2 Hz, CH), 68.95 (d, J CF = 17.9 Hz, CH), 44.93 (s, CH3) ppm. 19 F NMR (564.47 MHz, DMSO-d6) δ -111.51 (dd, J HF = 9.9 and 4.6 Hz, 1F), -203.88 (dd, J HF = 51.0 and 13.9 Hz, 1F) ppm. Array TIFF2025523627000019.tif62152
[0208] TIFF2025523627000020.tif229153
[0209] TIFF2025523627000021.tif229151
[0210] TIFF2025523627000022.tif230151
[0211] TIFF2025523627000023.tif229151
[0212] TIFF2025523627000024.tif229153
[0213] TIFF2025523627000025.tif229153
[0214] TIFF2025523627000026.tif230152
[0215] TIFF2025523627000027.tif229151
[0216] TIFF2025523627000028.tif230151
[0217] TIFF2025523627000029.tif229151
[0218] TIFF2025523627000030.tif46151
[0219] It will be understood that the various ones of the above-described and other features and functions, or alternatives thereto, can desirably be combined in many other different systems or applications. Also, various presently unforeseen or unexpected alternatives, modifications, variations, or improvements therein may be made later by those skilled in the art, which are also intended to be encompassed by the following claims.
Claims
1. A polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:
2.
2. The polypeptide according to claim 1, wherein the amino acid sequence has at least 95% sequence identity with SEQ ID NO:
2.
3. The polypeptide according to claim 1, wherein the amino acid sequence has at least 98% sequence identity with SEQ ID NO:
2.
4. The polypeptide according to claim 1, wherein the amino acid sequence consists of SEQ ID NO:
2.
5. The polypeptide according to claim 1, which consists of SEQ ID NO:
2.
6. The following conditions: (a) Amino acid (aa) residue 2 of SEQ ID NO: 2 is other than alanine, or (b) Aa residue 11 of SEQ ID NO: 2 is other than glutamic acid, The polypeptide according to claim 1, which satisfies at least one of the above.
7. A polynucleotide encoding the polypeptide according to any one of claims 1 to 6.
8. The polynucleotide according to claim 7, wherein the polynucleotide contains SEQ ID NO:
3.
9. The following conditions: (a) The triplet codon encoding the amino acid residue at position 2 of the polypeptide is other than GCT, (b) The triplet codon encoding the amino acid residue at position 3 of the polypeptide is other than AAA, (c) The triplet codon encoding the amino acid residue at position 4 of the polypeptide is other than ATC, or (d) The triplet codon encoding the amino acid residue at position 11 of the polypeptide is other than GAA, The polynucleotide according to claim 7, wherein at least one of the above is satisfied.
10. An expression vector comprising the polynucleotide according to any one of claims 7 to 9, operably linked to one or more control sequences suitable for directing the expression of the encoded polypeptide in a host cell.
11. The expression vector according to claim 10, wherein the control sequence contains a promoter.
12. The expression vector according to claim 11, wherein the promoter contains an E. coli promoter.
13. A host cell comprising the expression vector according to claim 11.
14. The host cell according to claim 12, wherein the host cell is E. coli.
15. A process for precipitating a protein from a cell lysate containing a polypeptide having an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 2, the process comprising treating the cell lysate containing the polypeptide to obtain a cell lysate composition containing more than 20% isopropanol. **Claim 16** The process according to claim 15, wherein the volume percent of isopropanol in the cell lysate composition is about 25%.