Engineered enzymes for preparing hydroxylated indanone intermediates useful in the synthesis of belzutifan

The FoPip4H enzyme from Fusarium oxysporum c8D enables efficient, environmentally friendly synthesis of optically pure belzutifan intermediates by converting indanone substrates with high enantiomeric excess, addressing the limitations of existing methods.

JP2025523626APending Publication Date: 2025-07-23MERCK SHARP & DOHME LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025500017
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-08
Filing Date
2023-07-05
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Existing processes for preparing belzutifan intermediates rely on hazardous reagents and transition metal catalysts, and there is a need for efficient, environmentally friendly methods to produce optically pure intermediates for belzutifan synthesis.

Method used

Utilization of a modified FoPip4H enzyme from Fusarium oxysporum c8D to stereospecifically convert 4-fluoro-7-(methylsulfonyl)-2,3-dihydro-1H-inden-1-one to (R)-4-fluoro-3-hydroxy-7-(methylsulfonyl)-2,3-dihydro-1H-inden-1-one with high enantiomeric excess, using a process that includes α-ketoglutaric acid and a reducing agent.

Benefits of technology

The process achieves high enantiomeric excess in producing optically pure intermediates for belzutifan synthesis with reduced environmental impact and minimal use of hazardous reagents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523626000001
    Figure 2025523626000001
  • Figure 2025523626000002
    Figure 2025523626000002
  • Figure 2025523626000003
    Figure 2025523626000003
Patent Text Reader

Abstract

The present disclosure provides an enzyme derived from the fungal Fusarium oxysporum c8D (the "FoPip4H enzyme") having improved properties compared to a naturally occurring wild-type enzyme, including the ability to hydroxylate a specific substituted indanone to provide optically pure 3-hydroxyindanone. Also provided are a polynucleotide encoding the FoPip4H enzyme, a host cell capable of expressing the FoPip4H enzyme, and a process for using the FoPip4H enzyme to synthesize (R)-4-fluoro-3-hydroxy-7-(methylsulfonyl)-2,3-dihydro-1H-inden-1-one, a useful intermediate in the synthesis of belzutifan.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the provision of hydroxylation of an enzyme derived from Fusarium oxysporum c8D, which is useful in biocatalysts and synthetic processes involving the oxidation of indanone to provide chiral alcohols. Such an enzyme is particularly useful for providing optically pure 3-hydroxyindanone, which is a useful intermediate in the preparation of belzutifan.

[0002] Reference to an electronically submitted sequence listing This application includes a sequence listing submitted electronically in XML format, which is hereby incorporated by reference in its entirety. The file name of the XML file created on October 31, 2022 is 25554-WO-PCT_SL.XML, and the size is 24,536 bytes.

Background Art

[0003] Enzymes are polypeptides that play a role in accelerating (often by several orders of magnitude) the chemical reactions of living cells. Without enzymes, most biochemical reactions would be too slow to even carry out life processes. Enzymes exhibit great specificity and are not permanently altered by their involvement in the reaction. Since enzymes do not change during the reaction, they are particularly cost-effective when used as catalysts for desired chemical conversions.

[0004] M. Hibi et al. disclose a protein derived from Fusarium oxysporum c8D that can hydroxylate pipecolic acid. Appl Environ Microbiol. 2016 Apr 1;82(7):2070-2077. The literature discloses that this enzyme is part of the Fe(II) / α-ketoglutarate-dependent dioxygenase superfamily. Furthermore, it is disclosed that 4-hydroxyamino acids such as (4S)-hydroxy-L and D-pipecolic acid can be efficiently produced with high optical purity.

[0005] Recently, the novel HIF-2α inhibitor 3-[(1S,2S,3R)-2,3-difluoro-1-hydroxy-7-methylsulfonyl-indan-4-yl]oxy-5-fluoro-benzonitrile (hereinafter, belzutifan) has received approval from the US Food and Drug Administration for the treatment of adult patients with von Hippel-Lindau (VHL) disease who require treatment for related renal cell carcinoma (RCC), central nervous system (CNS) hemangioblastoma, or pancreatic neuroendocrine tumor (pNET) and do not require immediate surgery.

Chemical formula

[0006] In recent clinical studies, an objective response was confirmed in 49% of RCC patients with VHL disease who were administered belzutifan, and the size of renal tumors decreased in most patients. In addition, 30% of CNS hemangioblastoma patients showed a response after treatment with belzutifan, and 91% of pancreatic neuroendocrine tumor patients showed a response after treatment. Belzutifan may serve to complement alternative or surgical treatments in patients with VHL disease. The researchers hypothesized that belzutifan may delay or eliminate the need for serial surgeries that can impose significant complications on such patients. (Jonasch, E. et al., N Engl J Med 2021;385:2036-2046).

[0007] In the study of belzutifan for the treatment of RCC, the agent demonstrated excellent in vitro potency, pharmacokinetic profile, and in vivo efficacy in a mouse model and showed promising outcomes in patients with progressive RCC (Xu, Rui, et al., J. Med. Chem. 62:6876-6893 (2019)). In a recent clinical study of patients who had previously been treated for progressive clear cell RCC, the confirmed objective response rate was 25%, and the median progression-free survival was 14.5 months. (Choueri, T.K. et al., Nature Medicine vol.27, 802-805 (2021)). For the treatment of patients suffering from VHL disease and its therapeutic effect, as well as its potential in the treatment of RCC patients, an efficient process is needed to prepare large quantities of belzutifan to support commercial supply and ongoing clinical research. U.S. Patent Application Publication No. 2022 / 0881407 discloses a method for preparing certain substituted indanes containing belzutifan. U.S. Provisional Patent Application No. 63 / 191,356, filed on May 21, 2021, discloses the crystalline forms of certain synthetic intermediates and specific processes for isolating such forms that are advantageous for the preparation of belzutifan.

[0008] The above publications and applications describe useful and expandable processes for preparing belzutifan, but further processes for preparing compounds, especially those that minimize the use of hazardous reagents and transition metal catalysts, are beneficial. In particular, it would be desirable to explore new transformations that provide optically pure intermediates with minimal environmental impact by using biocatalysts such as enzymes. The present disclosure provides such enzymes and processes for preparing useful intermediates for the preparation of belzutifan.

Prior Art Documents

Patent Documents

[0009]

Patent Document 1

Non-Patent Documents

[0010]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

[0011] The present disclosure relates to an enzyme derived from Fusarium oxysporum c8D (hereinafter, "FoPip4H enzyme") that can hydroxylate a substituted indanone to provide an optically pure alcohol. In some embodiments, the FoPip4H enzyme of the subject matter described herein can stereospecifically convert 4-fluoro-7-(methylsulfonyl)-2,3-dihydro-1H-inden-1-one (II) to the stereoisomerically pure (R)-4-fluoro-3-hydroxy-7-(methylsulfonyl)-2,3-dihydro-1H-inden-1-one (I), which is useful for the synthesis of belzutifan, 3-[[(1S,2S,3R)-2,3-difluoro-2,3-dihydro-1-hydroxy-7-(methylsulfonyl)-1H-inden-4-yl]oxy]-5-fluorobenzonitrile.

[0012] Further embodiments describe a process for preparing the subject FoPip4H enzyme and a process for using the subject FoPip4H enzyme.

[0013] In one aspect, the present disclosure provides a polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2. In some embodiments, the amino acid sequence has at least 95% sequence identity with SEQ ID NO: 2. In certain embodiments, the amino acid sequence has at least 98% sequence identity with SEQ ID NO: 2. In particular embodiments, the amino acid sequence consists of SEQ ID NO: 2. In certain embodiments, the polypeptide consists of SEQ ID NO: 2.

[0014] In another aspect, the present disclosure provides a polynucleotide encoding any of the foregoing polypeptides. In certain embodiments, the polynucleotide comprises SEQ ID NO: 3.

[0015] In another aspect, the present disclosure provides an expression vector comprising any of the foregoing polynucleotides operably linked to one or more control sequences suitable for directing the expression of the encoded polypeptide in a host cell. In some embodiments, the control sequence comprises a promoter. In certain embodiments, the promoter comprises an E. coli promoter.

[0016] In another aspect, the present disclosure provides a host cell comprising any of the foregoing expression vectors. In certain embodiments, the host cell is E. coli.

[0017] In another aspect, the present disclosure provides a process for preparing a compound according to formula (I)

Chemical formula

Chemical formula

[0018] In some embodiments, the process further comprises a reducing agent selected from the group consisting of L-cysteine, ascorbic acid, dithiothreitol, D-cysteine, L-homocysteine, and D-cysteine ethyl ester.

[0019] ​In certain embodiments, the process is carried out using a buffer selected from the group consisting of phosphate buffer, 2-morpholinoethanesulfonic acid, bis-tris, PIPES, citrate, bicine, and TEOA.

[0020] In some embodiments, the process further comprises an iron salt selected from the group consisting of Mohr's salt ((NH4)2Fe(SO4)2·6H2O) and iron chloride.

[0021] Other embodiments, aspects, and features of the invention will be further described in or will become apparent from the following description, examples, and appended claims. **DETAILED DESCRIPTION OF THE INVENTION**

[0022] Definitions Certain technical and scientific terms are specifically defined below. Unless otherwise specifically defined elsewhere in this document, all other technical and scientific terms used herein shall have the meaning commonly understood by one of ordinary skill in the art to which this disclosure pertains. Nevertheless, unless otherwise specified, the following definitions apply throughout this specification and the appended claims. Chemical names, common names, and chemical structures may be used interchangeably to describe the same structure.

[0023] As used throughout this specification and the entire disclosure, the following terms shall be understood to have the following meanings unless otherwise indicated.

[0024] This disclosure also encompasses isotopically labeled compounds that are identical to those recited herein except for the fact that one or more atoms are replaced by atoms having an atomic mass or mass number different from the atomic mass or mass number typically found in nature. Examples of isotopes that can be incorporated into the compounds of the invention include, respectively 2 H, 3 H, 11 C, 13 C, 14 C, 15 N, 18 O,17 O, 31 P, 32 P, 35 S, 18 F, 36 Cl, and 123 isotopes of hydrogen, carbon, nitrogen, oxygen, phosphorus, fluorine, and chlorine, and iodine are included.

[0025] Certain isotopically labeled compounds (e.g., 3 H and 14 C labeled ones) are useful in compound and / or substrate tissue distribution assays. Tritiation (i.e., 3 H) and carbon-14 (i.e., 14 C) isotopes are particularly preferred due to the ease of their preparation and detectability. Isotope substitution at the site where epimerization occurs slows down or decreases the epimerization process, whereby a more active or more effective form of the compound can be retained for a longer period. Isotopically labeled compounds, particularly those containing isotopes with a longer half-life (T 1 / 2 > 1 day), can generally be prepared by following procedures similar to those disclosed in the following schemes and / or examples herein, using appropriate isotopically labeled reagents in place of non-isotopically labeled reagents.

[0026] The compounds of this specification may contain one or more stereocenters and can exist as racemates, racemic mixtures, single enantiomers, mixtures of diastereomers, and individual diastereomers. Depending on the nature of the various substituents on the molecule, additional asymmetric centers may be present. Each such asymmetric center independently gives rise to two optical isomers, and all possible optical isomers and diastereomers in mixtures and as pure or partially purified compounds are included in this disclosure. Any formula, structure, or name of a compound described herein that does not specify a particular stereochemistry is meant to encompass any and all of the above existing isomers, and mixtures thereof in any proportion. When the stereochemistry is specified, this disclosure is meant to encompass that particular isomer, in pure form or as part of a mixture with other isomers in any proportion.

[0027] Mixtures of diastereomers can be separated into the individual diastereomers based on their physicochemical differences by methods well known to those skilled in the art, such as chromatography and / or fractional crystallization. Enantiomers can be separated by reacting the enantiomeric mixture to form a mixture of diastereomers with a suitable optically active compound (e.g., a chiral auxiliary such as a chiral alcohol or Mosher's acid chloride), separating the diastereomers, and converting the individual diastereomers to the corresponding pure enantiomers (e.g., by hydrolysis). Enantiomers can also be separated by the use of chiral HPLC columns.

[0028] All stereoisomers (e.g., geometric isomers and optical isomers, etc.) of the disclosed compounds (including salts and solvates of the compounds, as well as salts, solvates, and esters of prodrugs), such as those that may exist due to chiral carbons on various substituents, including enantiomeric forms (which may or may not have chiral carbons), rotational isomeric forms, atropisomers, and diastereomeric forms, are contemplated within the scope of the present disclosure. Individual stereoisomers of the compounds may, for example, be substantially free of other isomers or may be, for example, in racemic form or mixed with all other or other selected stereoisomers. Chiral centers can have the S or R configuration as defined by the IUPAC 1974 Recommendations.

[0029] The present disclosure further includes all those isolated forms of the compounds and synthetic intermediates. For example, the above compounds are intended to encompass all forms of the compounds, such as any solvates, hydrates, stereoisomers, and tautomers thereof.

[0030] "Protein", "polypeptide", and "peptide" are used interchangeably herein to denote a polymer of at least two amino acids covalently linked by amide bonds, regardless of length or post-translational modifications (e.g., glycosylation or phosphorylation, lipidation, myristoylation, ubiquitination, etc.). This definition includes polymers of d-amino acids and l-amino acids, as well as mixtures of d-amino acids and l-amino acids. Proteins, polypeptides, and peptides may include tags such as histidine tags, which should not be included when determining the percentage of sequence identity.

[0031] As used in the context of the polypeptides disclosed herein, "amino acid" or "residue" refers to a specific monomer at a sequence position. Amino acids are referred to herein by either the commonly known three-letter symbols or the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Similarly, nucleotides can be referred to by their generally accepted one-letter codes.

[0032] The abbreviations used for the genetically encoded amino acids are conventional and are, as follows, alanine (Ala or A), arginine (Arg or R), asparagine (Asn or N), aspartic acid (Asp or D), cysteine (Cys or C), glutamic acid (Glu or E), glutamine (Gln or Q), histidine (His or H), isoleucine (Ile or I), leucine (Leu or L), lysine (Lys or K), methionine (Met or M), phenylalanine (Phe or F), proline (Pro or P), serine (Ser or S), threonine (Thr or T), tryptophan (Trp or W), tyrosine (Tyr or Y), and valine (Val or V).

[0033] The abbreviations used for the nucleosides that are used to genetically encode are conventional and are, as follows, adenosine (A), guanosine (G), cytidine (C), thymidine (T), and uridine (U). Unless otherwise specified, the abbreviated nucleosides can be either ribonucleosides or 2'-deoxyribonucleosides. Nucleosides can be specified as being either ribonucleosides or 2'-deoxyribonucleosides, either individually or as an aggregate. When a nucleic acid sequence is presented as a series of one-letter abbreviations, the sequence is presented in the 5' to 3' direction according to common convention, and the phosphate is not shown.

[0034] As used herein in the context of enzymes, "derived from" identifies the originating enzyme on which the enzyme was based and / or the gene encoding such an enzyme. For example, the FoPip4H enzyme of SEQ ID NO: 2 was obtained by artificially evolving the gene encoding the wild-type (wt) FoPip4H enzyme of SEQ ID NO: 1 over multiple generations. Thus, this evolved FoPip4H enzyme is "derived from" the FoPip4H of SEQ ID NO: 1.

[0035] A "reference sequence" refers to a defined sequence used as a basis for sequence comparison. A reference sequence can be a subset of a larger sequence, e.g., a segment of a full-length gene or polypeptide sequence. Generally, a reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 residues in length, at least 50 residues in length, or the full length of a nucleic acid or polypeptide. Since two polynucleotides or polypeptides can each (1) contain a similar sequence (i.e., a portion of the complete sequence) between the two sequences and (2) further contain sequences that differ between the two sequences, sequence comparison between two (or more) polynucleotides or polypeptides is typically performed by comparing the sequences of the two polynucleotides over a "comparison window" to identify and compare local regions of sequence similarity.

[0036] In some embodiments, a "reference sequence" can be based on a primary amino acid sequence, and the reference sequence is a sequence that can have one or more changes in the primary sequence. For example, a "reference sequence based on SEQ ID NO: 1 having leucine at the residue corresponding to X96" refers to a reference sequence in which the corresponding residue of X96 in SEQ ID NO: 1 has been changed to leucine.

[0037] "Hydrophilic amino acid or residue" refers to an amino acid or residue having a side chain exhibiting a hydrophobicity less than zero according to the normalized consensus hydrophobicity scale of Eisenberg et al., 1984, J. Mol. Biol. 179: 125-142. Genetically encoded hydrophilic amino acids include l-Thr (T), l-Ser (S), l-His (H), l-Glu (E), l-Asn (N), l-Gln (Q), l-Asp (D), l-Lys (K), and l-Arg (R).

[0038] "Acidic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain exhibiting a pK value of less than about 6 when the amino acid is included in a peptide or polypeptide. Acidic amino acids typically have a side chain that is negatively charged at physiological pH due to the loss of a hydrogen ion. Genetically encoded acidic amino acids include l-Glu (E) and l-Asp (D).

[0039] "Basic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain exhibiting a pKa value of greater than about 6 when the amino acid is included in a peptide or polypeptide. Basic amino acids typically have a side chain that is positively charged at physiological pH due to association with a hydronium ion. Genetically encoded basic amino acids include l-Arg (R) and l-Lys (K).

[0040] "Polar amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that is not charged at physiological pH but has at least one bond in which an electron pair shared jointly by two atoms is held more closely by one of the atoms. Genetically encoded polar amino acids include l-Asn (N), l-Gln (Q), l-Ser (S), and l-Thr (T).

[0041] "Hydrophobic amino acid or residue" refers to an amino acid or residue having a side chain that exhibits hydrophobicity greater than zero according to the normalized consensus hydrophobicity scale of Eisenberg et al., 1984, J. Mol. Biol. 179: 125-142. Genetically encoded hydrophobic amino acids include l-Pro (P), l-Ile (I), l-Phe (F), l-Val (V), l-Leu (L), l-Trp (W), l-Met (M), l-Ala (A), and l-Tyr (Y).

[0042] "Aromatic amino acid or residue" refers to a hydrophilic or hydrophobic amino acid or residue having a side chain that includes at least one aromatic or heteroaromatic ring. Genetically encoded aromatic amino acids include l-Phe (F), l-Tyr (Y), l-His (H), and l-Trp (W). l-His (H) histidine is also classified herein as a hydrophilic residue or a constrained residue.

[0043] As used herein, "constrained amino acid or residue" refers to an amino acid or residue having a constrained geometric shape. As used herein, constrained residues include l-Pro (P) and l-His (H). Histidine has a relatively small imidazole ring and thus has a constrained geometric shape. Proline also has a five-membered ring and thus has a constrained geometric shape.

[0044] "Nonpolar amino acid or residue" refers to a hydrophobic amino acid or residue having a side chain that is not charged at physiological pH but has a bond in which an electron pair shared jointly by two atoms is generally held equally by each of the two atoms (i.e., the side chain is not polar). Genetically encoded nonpolar amino acids include l-Gly (G), l-Leu (L), l-Val (V), l-Ile (I), l-Met (M), and l-Ala (A).

[0045] As used herein, "aliphatic amino acid or residue" refers to a hydrophobic amino acid or residue having an aliphatic hydrocarbon side chain. Genetically encoded aliphatic amino acids include l-Ala (A), l-Val (V), l-Leu (L), and l-Ile (I).

[0046] The ability of l-Cys (C) (and other amino acids with -SH-containing side chains) to be present in a peptide in either the reduced free SH or oxidized disulfide bridge form affects whether l-Cys (C) confers net hydrophobic or net hydrophilic properties to the peptide. l-Cys (C) exhibits a hydrophobicity of 0.29 according to the normalized consensus scale of Eisenberg (Eisenberg et al., 1984, supra), but for the purposes of the present disclosure, it should be understood that l-Cys (C) is classified into its own unique group. It should be noted that cysteine (or "l-Cys" or "[C]") is unique in that it can form disulfide bridges with other l-Cys (C) amino acids or other sulfanyl- or sulfhydryl-containing amino acids. "Cysteine-like residues" include cysteine and other amino acids containing a sulfhydryl moiety available for the formation of disulfide bridges.

[0047] As used herein, "small amino acid or residue" refers to an amino acid or residue having a side chain composed of three or fewer total carbons and / or heteroatoms (excluding the α-carbon and hydrogen). Small amino acids or residues can be further classified as aliphatic, nonpolar, polar, or acidic small amino acids or residues according to the above definition. Genetically encoded small amino acids include l-Ala (A), l-Val (V), l-Cys (C), l-Asn (N), l-Ser (S), l-Thr (T), and l-Asp (D).

[0048] As used herein, "polynucleotide" and "nucleic acid" refer to two or more nucleotides covalently linked to each other. A polynucleotide can be composed entirely of ribonucleotides (i.e., RNA), entirely of 2'-deoxyribonucleotides (i.e., DNA), or a mixture of ribonucleotides and 2'-deoxyribonucleotides. Nucleosides will typically be joined together by standard phosphodiester bonds, but a polynucleotide can contain one or more non-standard bonds. A polynucleotide can be single-stranded or double-stranded, or a polynucleotide can contain both single-stranded and double-stranded regions. Further, a polynucleotide will typically be composed of naturally occurring coding nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), but can contain one or more modified and / or synthetic nucleobases such as, for example, inosine, xanthine, hypoxanthine, etc. In some embodiments, such modified or synthetic nucleobases are nucleobases that encode an amino acid sequence.

[0049] "Hydroxyl-containing amino acid or residue" refers to an amino acid containing a hydroxyl (-OH) moiety. Genetically encoded hydroxyl-containing amino acids include l-Ser (S), l-Thr (T), and l-Tyr (Y).

[0050] As used herein, "conservative amino acid substitution" refers to the substitution of a residue by a different residue having a similar side chain, and thus typically involves substitution of an amino acid in a polypeptide by an amino acid within the same or a similar defined class of amino acids. By way of example and not limitation, in some embodiments, an amino acid having an aliphatic side chain is substituted with another aliphatic amino acid (e.g., alanine, valine, leucine, and isoleucine), an amino acid having a hydroxyl side chain is substituted with another amino acid having a hydroxyl side chain (e.g., serine and threonine), an amino acid having an aromatic side chain is substituted with another amino acid having an aromatic side chain (e.g., phenylalanine, tyrosine, tryptophan, and histidine), an amino acid having a basic side chain is substituted with another amino acid having a basic side chain (e.g., lysine and arginine), an amino acid having an acidic side chain is substituted with another amino acid having an acidic side chain (e.g., aspartic acid and glutamic acid), and / or a hydrophobic or hydrophilic amino acid is replaced with another hydrophobic or hydrophilic amino acid, respectively.

[0051] As used herein, "non-conservative substitution" refers to substituting an amino acid in a polypeptide with an amino acid having significantly different side chain characteristics. Non-conservative substitutions may use amino acids from defined groups rather than within defined groups and affect (a) the structure of the peptide backbone in the region of substitution (e.g., proline instead of glycine), (b) the charge or hydrophobicity, or (c) the bulk of the side chain. By way of example and not limitation, exemplary non-conservative substitutions can be an acidic amino acid substituted with a basic or aliphatic amino acid, an aromatic amino acid substituted with a small amino acid, and a hydrophilic amino acid substituted with a hydrophobic amino acid.

[0052] As used herein, "deletion" refers to a modification of a polypeptide by the removal of one or more amino acids from a reference polypeptide. A deletion can involve the removal of one or more amino acids, two or more amino acids, five or more amino acids, ten or more amino acids, fifteen or more amino acids, or twenty or more amino acids, up to a maximum of 10% of the total number of amino acids, or up to a maximum of 20% of the total number of amino acids, from the reference enzyme while retaining the enzyme activity of the evolved enzyme and / or retaining improved properties. The deletion can target the internal portion and / or the terminal portion of the polypeptide. In various embodiments, the deletion can include a continuous segment or can be discontinuous. Deletions are typically indicated by "-" in the amino acid sequence.

[0053] As used herein, "insertion" refers to a modification of a polypeptide by the addition of one or more amino acids to a reference polypeptide. An insertion can be internal to the polypeptide or at the carboxy or amino terminus. Insertions as used herein include fusion proteins known in the art. An insertion can be a continuous segment of amino acids or can be separated by one or more of the amino acids in a naturally occurring polypeptide.

[0054] The term "amino acid substitution set" or "substitution set" refers to a group of amino acid substitutions within a polypeptide sequence as compared to a reference sequence. A substitution set can have one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, or more amino acid substitutions.

[0055] The terms "functional fragment" and "biologically active fragment" are used interchangeably herein to refer to a polypeptide that has (one or more) deletions at the amino and / or carboxy terminus and / or internal deletions, but the remaining amino acid sequence is identical to the corresponding positions in the sequence to which it is being compared and retains substantially all of the activity of the full-length polypeptide.

[0056] As used herein, "isolated polypeptide" refers to a polypeptide that is substantially separated from other contaminants (e.g., proteins, lipids, and polynucleotides) that naturally accompany it. This term encompasses polypeptides that have been removed or purified from their natural environment or expression system (e.g., those within a host cell or by in vitro synthesis). Recombinant polypeptides may be present within cells, in cell culture media, or prepared in various forms such as lysates or isolated preparations. Thus, in some embodiments, a recombinant polypeptide can be an isolated polypeptide.

[0057] As used herein, "substantially pure polypeptide" or "purified protein" refers to a composition in which the polypeptide species is the dominant species present (i.e., is more abundant than any other individual macromolecular species in the composition on a molar or weight basis), and generally, a composition is substantially purified when the target species comprises at least about 50 percent of the macromolecular species present on a molar or weight percent basis. However, in some embodiments, an enzyme-containing composition contains an enzyme with a purity of less than 50% (e.g., about 10%, about 20%, about 30%, about 40%, or about 50%). Generally, a substantially pure enzyme or polypeptide composition comprises about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, and about 98% or more of all macromolecular species present in the composition on a molar or weight percent basis. In some embodiments, the target species is purified until it is essentially homogeneous (i.e., contaminant species cannot be detected in the composition by conventional detection methods), where the composition consists essentially of a single macromolecular species. Solvent species, small molecules (<500 daltons), and elemental ion species are not considered macromolecular species. In some embodiments, an isolated recombinant polypeptide is a substantially pure polypeptide composition.

[0058] "Improved enzyme properties" refers to an enzyme that exhibits an improvement in any enzyme property as compared to a reference enzyme. For the enzymes described herein, the comparison is generally made to the wild-type enzyme, although in some embodiments, the reference enzyme can be another improved enzyme. Enzyme properties for which improvement may be desirable include, but are not limited to, enzyme activity (which can be expressed in terms of the conversion rate of a substrate), thermal stability, pH activity profile, cofactor requirements, insensitivity to inhibitors (e.g., product inhibition), stereospecificity, and stereoselectivity (including enantioselectivity).

[0059] "Increased enzyme activity" refers to an improved property of an enzyme, which can be expressed as an increase in specific activity (e.g., product produced / time / protein weight) or an increase in the conversion rate of a substrate to a product as compared to a reference enzyme (e.g., the conversion rate of an initial amount of substrate to product over a specified period using a specified amount of enzyme). Exemplary methods for determining enzyme activity are provided in the Examples. K m , V max , or k cat Any property related to enzyme activity, including classical enzyme properties of, can potentially be affected such that the change leads to an increase in enzyme activity. Improvement in enzyme activity can be from about 1.5-fold to up to 2-fold the enzyme activity of the corresponding wild-type enzyme. 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold, 500-fold, 1000-fold, 3000-fold, 5000-fold, 7000-fold, or more enzyme activity of a naturally occurring enzyme or another enzyme derived from a polypeptide. In specific embodiments, the enzyme exhibits improved enzyme activity in the range of 150 - 3000-fold, 3000 - 7000-fold, or more than 7000-fold higher than the enzyme activity of the parental enzyme. It is understood by those skilled in the art that the activity of any enzyme is diffusion-limited such that the catalytic turnover rate cannot exceed the diffusion rate of the substrate including any required cofactors. The theoretically maximum value of the diffusion limit, i.e., k cat / K m is generally about 10 8 ~10 9 (M -1 s -1) Thus, any improvement in enzyme activity will have an upper limit related to the diffusion rate of the substrate on which the enzyme acts. Enzyme activity can be measured by any one of the standard assays used to measure kinase activity, or by a binding assay using a nucleoside phosphorylase enzyme that can catalyze the reaction between a polypeptide product and a nucleoside base to obtain a nucleoside, or by any of the conventional methods for assaying chemical reactions including but not limited to HPLC, HPLC-MS, UPLC, UPLC-MS, TLC, and NMR. The comparison of enzyme activities is performed using a defined preparation of the enzyme, an assay under defined set conditions, and one or more defined substrates, as described in more detail herein. Generally, when comparing lysates, not only the same expression system and the same host cell are used to minimize variations in the amount of enzyme produced by the host cell and present in the lysate, but also the number of cells and the amount of protein to be assayed are determined.

[0060] As used herein, a "vector" is a DNA construct for introducing a DNA sequence into a cell. In some embodiments, the vector is an expression vector operably linked to suitable control sequences that can effect the expression of a polypeptide encoded in the DNA sequence in a suitable host. In some embodiments, an "expression vector" has a promoter sequence operably linked to a DNA sequence (e.g., a transgene) to drive expression in a host cell, and in some embodiments, also includes a transcription terminator sequence.

[0061] As used herein, the term "expression" includes any step involved in the production of a polypeptide, including but not limited to transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses the secretion of a polypeptide from a cell.

[0062] As used herein, the term "produce" refers to the production of proteins and / or other compounds by a cell. This term is intended to encompass any step involved in the production of a polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, this term also encompasses the secretion of a polypeptide from a cell.

[0063] As used herein, an amino acid or nucleotide sequence (e.g., a promoter sequence, signal peptide, terminator sequence, etc.) is "heterologous" to another sequence to which it is functionally linked if the two sequences are not related in nature. For example, a "heterologous polynucleotide" is any polynucleotide introduced into a host cell by experimental techniques, and this term includes polynucleotides that are removed from a host cell, subjected to experimental manipulation, and then reintroduced into the host cell.

[0064] As used herein, the terms "host cell" and "host strain" refer to a host suitable for an expression vector containing the DNA provided herein (e.g., a polynucleotide encoding a variant). In some embodiments, the host cell is a prokaryotic or eukaryotic cell that has been transformed or transfected with a vector constructed using recombinant DNA techniques known in the art.

[0065] The term "analog" means a polypeptide that has greater than 70% sequence identity with a reference polypeptide, but less than 100% sequence identity (e.g., greater than 75%, 78%, 80%, 83%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity). In some embodiments, "analog" means a polypeptide containing one or more non-natural amino acid residues, including, but not limited to, homoarginine, ornithine, and norvaline, as well as naturally occurring amino acids. In some embodiments, the analog also includes one or more d-amino acid residues and non-peptide bonds between two or more amino acid residues.

[0066] As used herein, the term "EC" number refers to the enzyme nomenclature of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (NC-IUBMB). The biochemical classification of the IUBMB is a numerical classification system of enzymes based on the chemical reactions catalyzed by the enzymes.

[0067] As used herein, "ATCC" refers to the American Type Culture Collection, and its biological repository collection includes genes and strains.

[0068] As used herein, "NCBI" refers to the National Center for Biological Information and the sequence databases provided therein.

[0069] "Coding sequence" refers to the portion of a nucleic acid (e.g., a gene) that encodes the amino acid sequence of a protein.

[0070] "Naturally occurring" or "wild-type" refers to the form found in nature. For example, a naturally occurring or wild-type polypeptide or polynucleotide sequence is a sequence that can be isolated from a natural source and is present in an organism that has not been intentionally modified by human manipulation, with the sole exception that a wild-type polypeptide or polynucleotide sequence identified herein may contain tags such as a histidine tag that should not be included when determining the percentage of sequence identity. In this specification, a "wild-type" polypeptide or polynucleotide sequence may be designated as "WT".

[0071] "Recombinant", when used with respect to, for example, a cell, nucleic acid, or polypeptide, refers to a material that has been modified in a manner that would not otherwise occur in nature or that is identical thereto but produced or derived from synthetic materials and / or by means of an operation using recombinant techniques, or a material corresponding to the natural or native form of the material. Non-limiting examples include, inter alia, recombinant cells that express a gene not found within the natural (non-recombinant) form of the cell or that express a native gene that would otherwise be expressed at a different level.

[0072] "Percentage of sequence identity", "identity rate", and "identity ratio" are used herein to refer to a comparison between polynucleotide sequences or polypeptide sequences and are determined by comparing two optimally aligned sequences over a comparison window, wherein a portion of the polynucleotide sequence or polypeptide sequence within the comparison window may include additions or deletions (i.e., gaps) as compared to the reference sequence for optimal alignment of the two sequences. The percentage is calculated by determining either the number of positions at which the identical nucleic acid bases or amino acid residues occur in both sequences or the number of positions at which the nucleic acid bases or amino acid residues are aligned with gaps, obtaining the number of matched positions, dividing the number of matched positions by the total number of positions within the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Optimal alignment and determination of the sequence identity rate are performed using the BLAST and BLAST 2.0 algorithms (see, e.g., Altschul et al., 1990, J. Mol. Biol. 215:403-410, and Altschul et al., 1977, Nucleic Acids Res. 3389-3402). Software for performing BLAST analysis is publicly available on the website of the National Center for Biotechnology Information of the United States.

[0073] Briefly described, BLAST analysis involves first identifying short words of length W in the query sequence that match or exceed a positive-valued threshold score T when aligned with words of the same length within the database sequences. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits serve as seeds to initiate a search to find longer HSPs that contain them. The word hits are then extended in both directions along each sequence as long as the cumulative alignment score can be increased. The cumulative score is calculated for nucleotide sequences using parameters M (reward score for a pair of matching residues, always >0) and N (penalty score for non-matching residues, always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. The extension of the word hits in each direction stops when the cumulative alignment score drops by an amount X from its maximum achieved value, when the cumulative score becomes 0 or less due to the accumulation of one or more negative-score residue alignments, or when the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses, by default, a word length (W) of 11, an expectation value (E) of 10, M = 5, N = -4, and comparison of both strands. For amino acid sequences, the BLASTP program uses, by default, a word length (W) of 3, an expectation value (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, 1989, Proc. Natl. Acad. Sci. USA 89:10915).

[0074] A number of other algorithms are available that function in a similar manner to BLAST in providing an identity rate for two arrays. Optimal alignment of the arrays for comparison can be performed, for example, by the local homology algorithm of Smith and Waterman, 1981, Adv. Appl. Math. 2:482, the homology algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48:443, the similarity search method of Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444, computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin software package), or by visual inspection (generally, see Current Protocols in Molecular Biology, F. M. Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)). In addition, for sequence alignment and determination of sequence identity rates, the BESTFIT or GAP programs of the GCG Wisconsin software package (Accelrys, Madison, Wisconsin) can be employed using the provided default parameters.

[0075] "Substantial identity" refers to a polynucleotide or polypeptide sequence having at least 80 percent sequence identity, preferably at least 85 percent sequence identity, more preferably at least 89 percent sequence identity, more preferably at least 95 percent sequence identity, still more preferably at least 99 percent sequence identity when compared to a reference sequence over a comparison window of at least 20 residue positions, often over a window of at least 30 - 50 residues. The percentage of sequence identity is calculated by comparing the reference sequence to a sequence that includes no more than 20 percent deletion or addition of the reference sequence over the total length of the reference sequence over the comparison window. In a specific embodiment applied to polypeptides, the term "substantial identity" means that two polypeptide sequences share at least 80 percent sequence identity, preferably at least 89 percent sequence identity, more preferably at least 95 percent sequence identity, or more (e.g., 99 percent sequence identity) when optimally aligned by a program such as GAP or BESTFIT using default gap weights. Preferably, the non-identical residue positions differ by conservative amino acid substitutions.

[0076] When used in the context of the numbering of a given amino acid or polynucleotide sequence, "corresponding to", "with respect to" or "relative to" refers to the numbering of the residues of a particular reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue numbers or residue positions of a given polymer are assigned with respect to the reference sequence, rather than by the actual numerical position of the residues within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, gaps are present, but the numbering of the residues within the given amino acid or polynucleotide sequence is done with respect to the reference sequence to which it is aligned.

[0077] "Stereoselectivity" refers to the preferential formation of one stereoisomer over another in a chemical or enzymatic reaction. Stereoselectivity can be partial, where the formation of one stereoisomer is favored over the other, or complete, where only one stereoisomer is formed. When the stereoisomers are enantiomers, the stereoselectivity is called enantioselectivity and is the proportion of one enantiomer in the total (typically reported as a percentage). Alternatively, it is generally reported in the art as the enantiomeric excess (EE) calculated therefrom according to the formula [major enantiomer - minor enantiomer] / [major enantiomer + minor enantiomer] (typically as a percentage). When the stereoisomers are diastereoisomers, the stereoselectivity is called diastereoselectivity and is the proportion of one diastereomer in a mixture of two diastereomers (typically reported as a percentage), or is generally reported as the diastereomeric excess (DE). Enantiomeric excess and diastereomeric excess are types of stereoisomeric excess.

[0078] "Highly stereoselective" refers to a chemical or enzymatic reaction that can convert a substrate to its corresponding product with a stereoisomeric excess of at least about 85%.

[0079] "Chemoselectivity" refers to the preferential formation of one product over another in a chemical or enzymatic reaction.

[0080] "Conversion" refers to the enzymatic conversion of a substrate to its corresponding product. "Conversion rate" refers to the percentage of substrate that is converted to product within a certain period under specific conditions. Thus, for example, the "enzymatic activity" or "activity" of a polypeptide can be expressed as the "conversion rate" of a substrate to a product.

[0081] "Chiral alcohol" refers to an amine of the general formula R 1 -CH(OH)-R 2 wherein R 1 and R 2are not the same and are used herein in their broadest sense and include a wide variety of aliphatic and cycloaliphatic compounds of different and mixed functional types, and in addition to hydrogen atoms, are characterized by the presence of a primary hydroxyl group bonded to a secondary carbon atom having either (i) a divalent group forming a chiral ring structure, or (ii) two substituents (other than hydrogen) that differ from each other in structure or chirality. Divalent groups forming a chiral cyclic structure include, for example, 2-methylbutane-1,4-diyl, pentane-1,4-diyl, hexane-1,4-diyl, hexane-1,5-diyl, 2-methylpentane-1,5-diyl. The two different substituents (R 1 and R 2 ) on the secondary carbon atom can also vary widely and include alkyl, aralkyl, aryl, halo, hydroxy, lower alkyl, lower alkoxy, lower alkylthio, cycloalkyl, carboxy, carboalkoxy, carbamoyl, mono- and di-(lower alkyl) substituted carbamoyl, trifluoromethyl, phenyl, nitro, amino, mono- and di-(lower alkyl) substituted amino, alkylsulfonyl, arylsulfonyl, alkylcarboxamide, arylcarboxamide, etc., as well as alkyl, aralkyl, or aryl substituted by those described above.

[0082] Immobilized enzyme preparations have several recognized advantages. They can, for example, impart a shelf life to enzyme preparations, improve reaction stability, enable stability in organic solvents, and assist in the removal of proteins from the reaction stream. "Stable" refers to the ability of immobilized enzymes to retain their structural conformation and / or their activity in a solvent system containing an organic solvent. A stable immobilized enzyme loses less than 10% of its activity per hour in a solvent system containing an organic solvent. A stable immobilized enzyme loses less than 9% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 8% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 7% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 6% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 5% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme has less than 4% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 3% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 2% of its activity per hour in a solvent system containing an organic solvent. Preferably, a stable immobilized enzyme loses less than 1% of its activity per hour in a solvent system containing an organic solvent.

[0083] "Thermostable" refers to a polypeptide that maintains similar activity (e.g., more than 60% - 80%) after being exposed to a high temperature (e.g., 40°C - 80°C) for a certain period (e.g., 0.5 h - 24 h) compared to the untreated enzyme.

[0084] "Solvent-stable" refers to a polypeptide that maintains a similar activity (e.g., more than 60% - 80%) after being exposed to solvents (such as isopropyl alcohol, tetrahydrofuran, 2-methyltetrahydrofuran, acetone, toluene, butyl acetate, methyl tert-butyl ether, etc.) at various concentrations (e.g., 5% - 99%) for a certain period (e.g., 0.5 h - 24 h) compared to the untreated enzyme.

[0085] "pH-stable" refers to a polypeptide that maintains a similar activity (e.g., more than 60% - 80%) after being exposed to high or low pH (e.g., 4.5 - 6 or 8 - 12) for a certain period (e.g., 0.5 h - 24 h) compared to the untreated enzyme.

[0086] "Thermostable and solvent-stable" refers to a polypeptide that has both thermostability and solvent stability.

[0087] As used herein, the terms "biocatalyst", "biocatalysis", "in vivo conversion", and "biosynthesis" refer to the use of an enzyme to effect a chemical reaction on an organic compound.

[0088] The term "effective amount" means an amount sufficient to produce the desired result. One of ordinary skill in the art may determine an effective amount using routine experimentation.

[0089] The terms "isolated" and "purified" are used to refer to a molecule (e.g., an isolated nucleic acid, polypeptide, etc.) or other component that has been removed from at least one other component to which it is naturally bound. The term "purified" does not require absolute purity; rather, it is intended as a relative definition.

[0090] "Control array" is defined herein as including all components necessary or advantageous for the expression of a polynucleotide and / or polypeptide of interest. Each control array may be native or foreign to the nucleic acid sequence encoding the polypeptide. Such control arrays include, but are not limited to, a leader, a polyadenylation sequence, a propeptide sequence, a promoter, a signal peptide sequence, and a transcription terminator. At a minimum, the control arrays include a promoter, as well as transcription and translation stop signals. A linker may be provided in the control arrays for the purpose of introducing specific restriction sites that facilitate ligation of the control arrays to the coding region of a polynucleotide of interest, such as a nucleic acid sequence encoding a polypeptide.

[0091] "Functionally linked" is defined herein as a configuration in which the control array is placed in an appropriate position relative to the polynucleotide sequence such that the control array directs the expression of the polynucleotide and / or polypeptide encoded by the polynucleotide (i.e., in a functional relationship).

[0092] "Promoter sequence" is a nucleic acid sequence recognized by a host cell for the expression of a polynucleotide. The control array may include an appropriate promoter sequence. The promoter sequence encompasses transcriptional control sequences that mediate the expression of the polynucleotide. A promoter can be any nucleic acid sequence that exhibits transcriptional activity in a selected host cell, including mutant promoters, truncated promoters, and hybrid promoters, and can be obtained from a gene encoding an extracellular or intracellular polypeptide homologous or heterologous to the host cell.

[0093] The "co-substrate" of the FoPip4H enzyme refers to an α-ketoglutaric acid analog and a co-substrate analog that can replace α-ketoglutaric acid in the hydroxylation of an indanone substrate analog. Co-substrate analogs include, but are not limited to, 2-oxoadipic acid (e.g., Majamaa et al., 1985, Biochem. J. 229:127-133).

[0094] The "FoPip4H enzyme" or "FoPip4H polypeptide" refers to a polypeptide having the enzymatic ability to oxidize a cyclic substrate to provide an alcohol-substituted cyclic structure. More specifically, the disclosed FoPip4H polypeptide can stereoselectively hydroxylate a substituted indanone of formula (II) to a substituted 3-hydroxyindanone of formula (I) (shown above). As used herein, the FoPip4H enzyme includes naturally occurring (wild-type) hydroxylases and non-naturally occurring engineered polypeptides produced by human manipulation.

[0095] Exemplary methods and materials are described herein, but methods and materials similar or equivalent to those described herein can also be used in the practice or testing of this disclosure. The materials, methods, and examples are illustrative only and not intended to be limiting.

[0096] [Table 1] TIFF2025523626000005.tif12156

[0097] FoPip4H enzyme This disclosure relates to a FoPip4H enzyme that can hydroxylate a substituted indanone to provide an optically pure alcohol. In embodiments, the FoPip4H enzyme is capable of the following conversions: [Chemical Formula]

[0098] In certain embodiments, the FoPip4H enzyme described herein has an amino acid sequence having one or more amino acid differences that result in improved properties of the enzyme with respect to a defined indanone substrate, compared to the reference amino acid sequence of wild-type FoPip4H.

[0099] The FoPip4H enzyme described herein is a product of directed evolution from wild-type FoPip4H c8D (SEQ ID NO: 1) having the following amino acid sequence (including the added histidine tag), identified by screening a panel of enzymes referenced in the literature: MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKKRGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQANPNRGYSFVGQENVANISGYEKGLGPLKTRDIKETVDFGSANDELVDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAENEFRILHYPAIPASELADGTATRIAEHTDFGTITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSIAYFAKPNRSASLFPLKEFIEEGVPCKYEDVTAWEWNNRRIEKLFSAEAKA (SEQ ID NO: 1).

[0100] In embodiments, the FoPip4H enzyme of the present disclosure may demonstrate improvements such as an increase in enzyme activity, stereoselectivity, stereospecificity, thermal stability, solvent stability, a decrease in product inhibition, or a decrease in peroxidation, compared to the FoPip4H enzyme of SEQ ID NO: 1.

[0101] In some embodiments, the FoPip4H enzyme of the present disclosure may demonstrate an improvement in the rate of enzyme activity, i.e., the rate of conversion of an indanone substrate to a product. In some embodiments, the FoPip4H polypeptide can convert a substrate to a product at a rate of at least 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 25-fold, 50-fold, 100-fold, 150-fold, 200-fold, 400-fold, 1000-fold, 3000-fold, 5000-fold, 7000-fold, or more than 7000-fold the rate exhibited by the enzyme of SEQ ID NO: 1.

[0102] In some embodiments, such FoPip4H polypeptides can also convert indanone substrates to products with an enantiomeric excess of at least 60%. In some embodiments, such FoPip4H polypeptides can also convert substrates to products with an enantiomeric excess of at least 90%. In some embodiments, such FoPip4H polypeptides can also convert substrates to products with an enantiomeric excess of at least about 99%.

[0103] In some embodiments, the FoPip4H polypeptide is highly enantioselective, and the polypeptide can reduce substrates to products with an enantiomeric excess greater than about 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%.

[0104] In some embodiments, the improved FoPip4H polypeptide of the present disclosure is a polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2 having the amino acid sequence shown below: MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGVVKLINHPIPSESIHELFAQTKRFFNLPLETKMLAKHPEQALPARGYAFVGQENVANISGYEKGLPPLKTRDIKETVDFGSANDEKYDNLWVPEEELPGFRSFMEGFYELAFKTEMQILEALAIALGVSPDHLKSLHNRAHSELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDGVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTWPPSIKDGDGSEVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPPKYEDLTFEEWNNRRIEKLFSAEAKA (SEQ ID NO: 2)

[0105] In some embodiments, the improved FoPip4H polypeptide of the present disclosure is based on the sequence set forth in SEQ ID NO: 2 and can include an amino acid sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the reference sequence of SEQ ID NO: 2.

[0106] These differences between these variants and SEQ ID NO: 2 can be insertions, deletions, substitutions of amino acids, or any combination of such changes. In some embodiments, the differences in the amino acid sequence can include non-conservative, conservative, and combinations of non-conservative and conservative amino acid substitutions.

[0107] In a particular embodiment, an improved FoPip4H polypeptide of the present disclosure, wherein the amino acid sequence consists of SEQ ID NO: 2. In a specific embodiment, the improved FoPip4H polypeptide of the present disclosure consists of SEQ ID NO: 2.

[0108] Further embodiments provide a host cell comprising a polynucleotide and / or expression vector described herein. The host cell may be Escherichia coli (E. coli) or a different organism such as Lactobacillus brevis (L. brevis). The host cell can be used for the expression and isolation of the FoPip4H enzyme described herein or can be used directly for the conversion of a substrate to a stereoisomeric product.

[0109] Whether this method is carried out using whole cells, cell extracts, or purified FoPip4H enzyme, a single FoPip4H enzyme may be used, or a mixture of two or more FoPip4H enzymes may be used.

[0110] Table 1 below provides a list of FoPip4H polypeptides (identified by the SEQ ID Nos. disclosed herein) and their related properties. The following sequences are based on the FoPip4H sequence of SEQ ID No. 1, unless otherwise specified. In the following table, each row lists the SEQ ID No. The column listing the number of mutations (i.e., residue changes) refers to the number of amino acid substitutions compared to the FoPip4H sequence of SEQ ID No. 1.

[0111] In the column of the thermal stability of the table, + indicates T50 < 25 °C, ++ indicates 25 °C < T50 < 45 °C, and +++ indicates T50 > 45 °C. In the column of the conversion of the table, + indicates < 10% conversion of the substrate to the product, ++ indicates 10% - 60% conversion, and +++ indicates > 60% conversion. The conversion analysis was performed based on the assumptions of 25 wt.% enzyme and 10 g / L substrate. In the column of the selectivity of the table, + indicates < 60% enantiomeric excess (ee), ++ indicates 60% - 90% ee, and +++ indicates > 90% ee. In the column of the peroxidation of the table, +++ indicates > 15% peroxidation. ++ indicates 5% - 15% peroxidation. + indicates < 5% peroxidation.

[0112] [Table 2] TIFF2025523626000008.tif209153TIFF2025523626000009.tif74152

[0113] Polynucleotide encoding the FoPip4H enzyme In another aspect, the present disclosure provides a polynucleotide encoding a FoPip4H polypeptide disclosed herein. The polynucleotide can be operably linked to one or more heterologous regulatory sequences that control gene expression to produce a recombinant polynucleotide capable of expressing the polypeptide. An expression construct containing the heterologous polynucleotide encoding FoPip4H can be introduced into a suitable host cell to express the corresponding FoPip4H polypeptide.

[0114] Since the codons corresponding to various amino acids are known, if a protein sequence is available, an explanation of all the polynucleotides that can encode the subject can be obtained. Due to the degeneracy of the genetic code in which the same amino acid is encoded by alternative or synonymous codons, a very large number of nucleic acids can be produced, and all of them encode the improved FoPip4H enzyme disclosed herein. Thus, by identifying a specific amino acid sequence, one of ordinary skill in the art may be able to generate any number of different nucleic acids by simply modifying the sequence of one or more codons in a way that does not change the amino acid sequence of the protein. In this regard, the present disclosure specifically contemplates every possible variation of polynucleotides that can be produced by selecting combinations based on possible codon choices, and all such variations should be considered specifically disclosed for any polypeptide disclosed herein.

[0115] In various embodiments, the codons are preferably selected to be compatible with the host cell in which the protein is being produced. For example, the preferred codons used in bacteria are those used for gene expression in bacteria, the preferred codons used in yeast are those used for expression in yeast, and the preferred codons used in mammals are those used for expression in mammalian cells. By way of example, the polynucleotide of SEQ ID NO: 3 is codon-optimized for expression in E. coli.

[0116] In certain embodiments, the native sequence will contain preferred codons, and since the use of preferred codons may not be required for all amino acid residues, it is not necessary to replace all codons to optimize the codon usage frequency of the FoPip4H enzyme. As a result, the codon-optimized polynucleotide encoding the FoPip4H enzyme may contain preferred codons at about 40%, 50%, 60%, 70%, 80%, or more than 90% of the codon positions in the full-length coding region.

[0117] In various embodiments, the isolated polynucleotide encoding the improved FoPip4H polypeptide can be manipulated in a variety of ways to provide for the expression of the polypeptide. Manipulation of the isolated polynucleotide prior to insertion into a vector may be desirable or necessary depending on the expression vector. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art. Guidance is provided in Sambrook et al., 2001, Molecular Cloning: A Laboratory Manual, 3 rd Ed., Cold Spring Harbor Laboratory Press, and Current Protocols in Molecular Biology, Ausubel, F. ed., Greene Pub. Associates, 1998, updates to 2006

[0118] In some embodiments, an isolated polynucleotide encoding any of the FoPip4H polypeptides herein is engineered in various ways to promote the expression of the FoPip4H polypeptide. In some embodiments, a polynucleotide encoding a FoPip4H polypeptide comprises an expression vector in which one or more control sequences for regulating the expression of the FoPip4H polynucleotide and / or polypeptide are present. Manipulation of the isolated polynucleotide prior to insertion into the vector may be desirable or necessary depending on the expression vector utilized. Techniques for modifying polynucleotides and nucleic acid sequences using recombinant DNA methods are well known in the art. In some embodiments, control sequences include, among others, a promoter, a leader sequence, a polyadenylation sequence, a propeptide sequence, a signal peptide sequence, and a transcription terminator. In some embodiments, a suitable promoter is selected based on host cell selection.For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include, but are not limited to, the Escherichia coli lac operon, the Streptomyces coelicolor agarase gene (dagA), the Bacillus subtilis levansucrase gene (sacB), the Bacillus licheniformis α-amylase gene (amyL), the Bacillus stearothermophilus maltogenic amylase gene (amyM), the Bacillus amyloliquefaciens α-amylase gene (amyQ), the Bacillus licheniformis penicillinase gene (penP), the Bacillus subtilis xylA and xylB genes, and promoters obtained from prokaryotic β-lactamase genes (see, for example, Villa-Kamaroff et al., Proc. Natl Acad. Sci. USA 75:3727-3731

[1978] ), as well as the tac promoter (see, for example, DeBoer et al., Proc. Natl Acad. Sci. USA 80:21-25

[1983] ).Exemplary promoters for filamentous fungal host cells include those obtained from the genes of Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral α - amylase, Aspergillus niger acid - stable α - amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin - like protease (see, e.g., WO 96 / 00787), as well as the NA2 - tpi promoter (a hybrid of the promoters from the genes of Aspergillus niger neutral α - amylase and Aspergillus oryzae triose phosphate isomerase), and their variant promoters, truncated promoters, and hybrid promoters, but are not limited thereto. Exemplary yeast cell promoters can be derived from the genes of Saccharomyces cerevisiae enolase (ENO - 1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde - 3 - phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3 - phosphoglycerate kinase. Other useful promoters for yeast host cells are known in the art (see, e.g., Romanos et al., Yeast 8:423 - 488

[1992] ).

[0119] In some embodiments, the control sequence is also a suitable terminator sequence (i.e., a sequence recognized by the host cell to terminate transcription). In some embodiments, the terminator sequence is operably linked to the 3' end of the nucleic acid sequence encoding the enzyme polypeptide. Any suitable terminator that is functional in the selected host cell is used in the present invention. Exemplary transcription terminators for filamentous fungal host cells can be obtained from the genes of Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger α-glucosidase, and Fusarium oxysporum trypsin-like protease. Exemplary terminators for yeast host cells can be obtained from the genes of Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are known in the art (see, e.g., Romanos et al. supra).

[0120] In some embodiments, the control sequence is also a suitable leader sequence (i.e., the untranslated region of the mRNA important for translation by the host cell). In some embodiments, the leader sequence is operably linked to the 5' end of the nucleic acid sequence encoding FoPip4H. Any suitable leader sequence that is functional in the selected host cell is used in the present invention. Exemplary leaders for filamentous fungal host cells are obtained from the genes of Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase. Leaders suitable for yeast host cells are obtained from the genes of Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae α-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).

[0121] In some embodiments, the control sequence is also a polyadenylation sequence (i.e., a sequence that is operably linked to the 3' end of a nucleic acid sequence and that, when transcribed, is recognized by the host cell as a signal for adding polyadenosine residues to the transcribed mRNA). Any suitable polyadenylation sequence that is functional in the selected host cell can be used in the present invention. Exemplary polyadenylation sequences for filamentous fungal host cells include, but are not limited to, the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger α-glucosidase. Polyadenylation sequences useful for yeast host cells are known (see, e.g., Guo and Sherman, Mol. Cell. Biol., 15:5983-5990

[1995] ).

[0122] In some embodiments, the control sequence is also a signal peptide (i.e., an amino acid sequence that is encoded and attached to the amino terminus of a polypeptide and that encodes a polypeptide that directs the encoded polypeptide into the secretory pathway of a cell). In some embodiments, the 5' end of the coding sequence of the nucleic acid sequence essentially contains a signal peptide coding region that is naturally linked in-frame with a segment of the coding region encoding the secreted polypeptide. Alternatively, in some embodiments, the 5' end of the coding sequence contains a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the secretory pathway of the selected host cell is used for the expression of the engineered (one or more) polypeptides. Effective signal peptide coding regions for bacterial host cells include, but are not limited to, those obtained from the genes of Bacillus NClB 11837 maltogenic amylase, Bacillus stearothermophilus α-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis β-lactamase, Bacillus stearothermophilus neutral protease (nprT, nprS, nprM), and Bacillus subtilis prsA. Additional signal peptides are known in the art (see, for example, Simonen and Palva, Microbiol. Rev., 57:109-137

[1993] ). In some embodiments, effective signal peptide coding regions for filamentous fungal host cells include, but are not limited to, those obtained from the genes of Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase. Signal peptides useful for yeast host cells include, but are not limited to, those derived from the genes of Saccharomyces cerevisiae α-factor and Saccharomyces cerevisiae invertase.

[0123] In some embodiments, regulatory sequences are also utilized. These sequences facilitate the regulation of polypeptide expression with respect to the growth of the host cell. Examples of regulatory systems are those that turn gene expression on or off in response to chemical or physical stimuli, including the presence of regulatory compounds. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory systems include, but are not limited to, the ADH2 system or the GAL1 system. In filamentous fungi, suitable regulatory sequences include, but are not limited to, the TAKA α-amylase promoter, the Aspergillus niger glucoamylase promoter, and the Aspergillus oryzae glucoamylase promoter.

[0124] In another aspect, the present invention is directed to a recombinant expression vector comprising a polynucleotide encoding a FoPip4H polypeptide and one or more expression regulatory regions, such as a promoter and a terminator, an origin of replication, etc., depending on the type of host into which it is to be introduced. In some embodiments, the various nucleic acids and control sequences described herein are ligated together to produce a recombinant expression vector containing one or more convenient restriction sites, allowing the insertion or substitution of nucleic acid sequences encoding enzyme polypeptides at such sites. Alternatively, in some embodiments, the nucleic acid sequences of the present invention are expressed by inserting the nucleic acid sequence or nucleic acid construct containing the sequence into an appropriate vector for expression. In some embodiments involving the production of an expression vector, the coding sequence is positioned within the vector such that the coding sequence is operably linked to appropriate control sequences for expression.

[0125] A recombinant expression vector can be any suitable vector (e.g., a plasmid or a virus) that is conveniently used in recombinant DNA procedures and can result in the expression of an enzyme polynucleotide sequence. The choice of vector typically depends on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector can be a linear or circular plasmid.

[0126] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extrachromosomal entity whose replication is replicated independently of chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome). The vector can include any means for ensuring self-replication. In some alternative embodiments, the vector is a vector that is integrated into the genome and replicated with the integrated chromosome(s) when introduced into the host cell. Further, in some embodiments, a single vector or plasmid, or two or more vectors or plasmids that together contain all the DNA introduced into the genome of the host cell, and / or a transposon are utilized.

[0127] In some embodiments, the expression vector contains one or more selectable markers that enable easy selection of the transformed cells. A "selectable marker" is a gene whose product provides, for example, biocide or virus resistance, resistance to heavy metals, and prototrophy for auxotrophs. Examples of bacterial selectable markers include, but are not limited to, the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers conferring antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, or tetracycline resistance. Markers suitable for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for use in filamentous fungal host cells include, but are not limited to, amdS (acetamidase, e.g., from A. nidulans or A. oryzae), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase, e.g., from S. hygroscopicus), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase, e.g., from A. nidulans or A. oryzae), sC (sulfate adenylyltransferase), and trpC (anthranilate synthase), and their equivalents.

[0128] In another aspect, the present invention provides a host cell comprising at least one polynucleotide encoding at least one FoPip4H of the present disclosure, wherein the (one or more) polynucleotide(s) is / are operably linked to one or more control sequences for the expression of at least one FoPip4H in the host cell. Host cells suitable for use in the expression of polypeptides encoded by the expression vectors of the present invention are well known in the art and include, but are not limited to, bacterial cells such as Escherichia coli, Vibrio fluvialis, Streptomyces, and Salmonella typhimurium cells, and fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC accession number 201178)). Exemplary host cells also include various Escherichia coli strains (e.g., W3110(ΔfhuA) and BL21). Examples of bacterial selectable markers include the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers conferring antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, and / or tetracycline resistance, but are not limited thereto.

[0129] In some embodiments, the expression vectors of the present invention include (one or more) elements that enable integration of the vector into the genome of the host cell or autonomous replication of the vector intracellularly independent of the genome. In some embodiments involving integration into the host cell genome, the vector relies on a nucleic acid sequence encoding a polypeptide or any other element of the vector for integration of the vector into the genome by homologous or non-homologous recombination.

[0130] In some alternative embodiments, the expression vector contains additional nucleic acid sequences to direct integration by homologous recombination into the genome of the host cell. The additional nucleic acid sequences enable the vector to be integrated into the host cell genome at an exact (one or more) location(s) in the (one or more) chromosome(s). To enhance the likelihood of integration at the exact location, the integration element preferably contains a sufficient number of nucleotides such as 100 to 10,000 base pairs, preferably 400 to 10,000 base pairs, most preferably 800 base pairs to 10,000 base pairs, which are highly homologous to the corresponding target sequence to enhance the probability of homologous recombination. The integration element can be any sequence homologous to the target sequence in the genome of the host cell. Further, the integration element can be a non-coding nucleic acid sequence or a coding nucleic acid sequence. On the other hand, the vector can be integrated into the genome of the host cell by non-homologous recombination.

[0131] In the case of autonomous replication, the vector can further contain an origin of replication that enables the vector to replicate autonomously in the host cell in question. Examples of bacterial origins of replication are P15A ori, or the plasmids pBR322, pUC19, pACYCl77 (containing P15A ori), or pACYC184 (containing P15A ori) that enable replication in Escherichia coli, and the origins of replication of pUB110, pE194, or pTA1060 that enable replication in Bacillus. Examples of origins of replication for use in yeast host cells are the 2 micron origin of replication, ARS1, ARS4, the combination of ARS1 and CEN3, and the combination of ARS4 and CEN6. The origin of replication can have a mutation that makes its function temperature-sensitive in the host cell (see, for example, Ehrlich, Proc. Natl. Acad. Sci. USA 75:1433

[1978] ).

[0132] In some embodiments, two or more copies of the nucleic acid sequences of the invention are inserted into a host cell to increase the production of the gene product. The increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene in the nucleic acid sequence, and cells containing the amplified copies of the selectable marker gene and thereby additional copies of the nucleic acid sequence can be selected by culturing the cells in the presence of an appropriate selectable agent.

[0133] Many of the expression vectors used in the present invention are commercially available. Suitable commercially available expression vectors include, but are not limited to, the (registered trademark) pET E. coli T7 expression vectors (Millipore Sigma) from Novagen, and the p3xFLAGTM (trademark) expression vectors (Sigma-Aldrich Chemicals). Other suitable expression vectors include, but are not limited to, pBluescriptII SK(-) and pBK-CMV (Stratagene), and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen), or pPoly (see, for example, Lathe et al., Gene 57:193-201

[1987] ).

[0134] Accordingly, in some embodiments, a vector comprising a sequence encoding at least one variant FoPip4H is transformed into a host cell to enable propagation of the vector and expression of the variant(s) FoPip4H. In some embodiments, the transformed host cell is cultured in a suitable nutrient medium under conditions that enable expression of the variant(s) FoPip4H. Any suitable medium useful for culturing the host cell, e.g., but not limited to, minimal medium or complex medium containing appropriate supplements, is used in the present invention. In some embodiments, the host cell is grown in HTP medium. Suitable media are available from various commercial suppliers or can be prepared according to published recipes (e.g., those found in the catalog of the American Type Culture Collection).

[0135] Host cells for expression of the FoPip4H enzyme In another aspect, the present disclosure provides a host cell comprising a polynucleotide encoding an improved FoPip4H polypeptide of the present disclosure, the polynucleotide being operably linked to one or more control sequences for the expression of the FoPip4H enzyme in the host cell. Host cells for use in the expression of the FoPip4H polypeptide encoded by the expression vectors of the present invention are well known in the art and include, but are not limited to, bacterial cells such as Escherichia coli, B. subtilis, B. licheniformis, B. megaterium, B. stearothermophilus, B. amyloliquefaciens, Lactobacillus kejir, Lactobacillus brevis, Lactobacillus minor, Streptomyces and Salmonella typhimurium cells, and fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC accession number 201178)). Suitable culture media and growth conditions for the above host cells are well known in the art.

[0136] The polynucleotide for the expression of the FoPip4H polypeptide can be introduced into cells by various methods known in the art. Techniques include, among others, electroporation, biolistic microparticle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion. Various methods for introducing polynucleotides into cells will be apparent to those skilled in the art.

[0137] In some embodiments of the present invention, the filamentous fungal host cell is any suitable genus and species including, but not limited to, Achlya, Acremonium, Aspergillus, Aureobasidium, Bjerkandera, Ceriporiopsis, Cephalosporium, Chrysosporium, Cochliobolus, Corynascus, Cryphonectria, Cryptococcus, Coprinus, Coriolus, Diplodia, Endothis, Fusarium, Gibberella, Gliocladium, Humicola, Hypocrea, Myceliophthora, Mucor, Neurospora, Penicillium, Podospora, Phlebia, Piromyces, Pyricularia, Rhizomucor, Rhizopus, Schizophyllum, Scytalidium, Sporotrichum, Talaromyces, Thermoascus, Thielavia, Trametes, Tolypocladium, Trichoderma, Verticillium, and / or Volvariella, and / or their teleomorphs or anamorphs, and their synonyms, basionyms, or taxonomic equivalents.

[0138] In some embodiments of the present invention, the host cell is a yeast cell including, but not limited to, cells of Candida, Hansenula, Saccharomyces, Schizosaccharomyces, Pichia, Kluyveromyces, or Yarrowia species. In some embodiments of the present invention, the yeast cell is Hansenula polymorpha, Saccharomyces cerevisiae, Saccharomyces carlsbergensis, Saccharomyces diastaticus, Saccharomyces norbensis, Saccharomyces kluyveri, Schizosaccharomyces pombe, Pichia pastoris, Pichia finlandica, Pichia trehalophila, Pichia kodamae, Pichia membranaefaciens, Pichia opuntiae, Pichia thermotolerans, Pichia salictaria, Pichia quercuum, Pichia pijperi, Pichia stipitis, Pichia methanolica, Pichia angusta, Kluyveromyces lactis, Candida albicans, or Yarrowia lipolytica.

[0139] In some other embodiments, the host cell is a prokaryotic cell. Suitable prokaryotic cells include, but are not limited to, gram-positive bacterial cells, gram-negative bacterial cells, and gram-variable bacterial cells. In the present invention, Agrobacterium, Alicyclobacillus, Anabaena, Anacystis, Acinetobacter, Acidothermus, Arthrobacter, Azobacter, Bacillus, Bifidobacterium, Brevibacterium, Butyrivibrio, Buchnera, Campestris, Camplyobacter, Clostridium, Corynebacterium, Chromatium, Coprococcus, Escherichia, Enterococcus, Enterobacter, Erwinia, Fusobacterium, Faecalibacterium, Francisella, Flavobacterium, Geobacillus, Haemophilus, Helicobacter, Klebsiella, Lactobacillus, Lactococcus, Ilyobacter, Micrococcus, Microbacterium, Mesorhizobium, Methylobacterium, Methylobacterium, Mycobacterium, Neisseria, Pantoea,Any suitable bacterial organism including, but not limited to, Pseudomonas, Prochlorococcus, Rhodobacter, Rhodopseudomonas, Rhodospirillum, Rhodococcus, Scenedesmus, Streptomyces, Streptococcus, Synecoccus, Saccharomonospora, Staphylococcus, Serratia, Salmonella, Shigella, Thermoanaerobacterium, Tropheryma, Tularensis, Temecula, Thermosynechococcus, Thermococcus, Ureaplasma, Xanthomonas, Xylella, Yersinia, and Zymomonas is used. In some embodiments, the host cell is of the species Agrobacterium, Acinetobacter, Azotobacter, Bacillus, Bifidobacterium, Brucella, Geobacillus, Campylobacter, Clostridium, Corynebacterium, Escherichia, Enterococcus, Erwinia, Flavobacterium, Lactobacillus, Lactococcus, Pantoea, Pseudomonas, Staphylococcus, Salmonella, Streptococcus, Streptomyces, or Zymomonas. In some embodiments, the bacterial host strain is non-pathogenic to humans. In some embodiments, the bacterial host strain is an industrial strain. A number of bacterial industrial strains are known and suitable for the present invention. In some embodiments of the present invention, the bacterial host cell is of the Agrobacterium species (e.g., A. radiobacter,A. rhizogenes, and A. rubi). In some embodiments of the present invention, the bacterial host cell is an Agrobacterium species (e.g., A. aurescens, A. citreus, A. globiformis, A. hydrocarboglutamicus, A. mysorens, A. nicotianae, A. paraffineus, A. protophonniae, A. roseoparqffinus, A. sulfureus, and A. ureafaciens). In some embodiments of the present invention, the bacterial host cell is a Bacillus species (e.g., B. thuringensis, B. anthracis, B. megaterium, B. subtilis, B. lentus, B. circulans, B. pumilus, B. lautus, B. coagulans, B. brevis, B. firmus, B. alkaophius, B. licheniformis, B. clausii, B. stearothermophilus, B. halodurans, and B. amyloliquefaciens). In some embodiments, the host cell is an industrial Bacillus strain including but not limited to B. subtilis, B. pumilus, B. licheniformis, B. megaterium, B. clausii, B. stearothermophilus, or B. amyloliquefaciens. In some embodiments, the Bacillus host cell is B. subtilis, B. licheniformis, B. megaterium, B. stearothermophilus, and / or B. amyloliquefaciens. In some embodiments, the bacterial host cell is a Clostridium species (e.g.,C. acetobutylicum, C. tetani E88, C. lituseburense, C. saccharobutylicum, C. perfringens, and C. beijerinckii). In some embodiments, the bacterial host cell is a Corynebacterium species (e.g., C. glutamicum and C. acetacidophilum). In some embodiments, the bacterial host cell is an Escherichia species (e.g., E. coli). In some embodiments, the host cell is E. coli W3110. In some embodiments, the host is E. coli BL21 or BL21(DE3). In some embodiments, the bacterial host cell is an Erwinia species (e.g., E. uredovora, E. carotovora, E. ananas, E. herbicola, E. punctata, and E. terreus). In some embodiments, the bacterial host cell is a Pantoea species (e.g., P. citrea and P. agglomerans). In some embodiments, the bacterial host cell is a Pseudomonas species (e.g., P. putida, P. aeruginosa, P. mevalonii, and P. sp. D-0l 10). In some embodiments, the bacterial host cell is a Streptococcus species (e.g., S. equisimiles, S. pyogenes, and S. uberis). In some embodiments, the bacterial host cell is a Streptomyces species (e.g., S. ambofaciens, S. achromogenes, S. avermitilis, S. coelicolor, S. aureofaciens, S. aureus,S. fungicidicus, S. griseus, and S. lividans). In some embodiments, the bacterial host cell is a Zymomonas species (e.g., Z. mobilis and Z. lipolytica).

[0140] Many of the prokaryotic and eukaryotic strains used in the present invention are readily available from several culture collections such as the American Type Culture Collection (ATCC), Deutsche Sammlung von Mikroorganismen und Zellkulturen GmbH (DSM), Centraalbureau Voor Schimmelcultures (CBS), and Agricultural Research Service Patent Culture Collection, Northern Regional Research Center (NRRL).

[0141] In some embodiments, the host cell is genetically modified to have properties that improve protein secretion, protein stability, and / or other properties desirable for the expression and / or secretion of the protein. The genetic modification can be achieved by genetic engineering techniques and / or classical microbiological techniques (e.g., chemical or UV mutagenesis and subsequent selection). Indeed, in some embodiments, a combination of recombinant modification and classical selection techniques is used to produce the host cell. Recombinant techniques can be used to introduce, delete, inhibit, or modify nucleic acid molecules so as to increase the yield of the (one or more) FoPip4H variants in the host cell and / or in the culture medium. In one genetic engineering approach, homologous recombination is used to induce targeted genetic modification by specifically targeting genes in vivo to suppress the expression of the encoded protein. In an alternative approach, siRNA, antisense, and / or ribozyme technologies are used to inhibit gene expression. A variety of methods for reducing the expression of a protein in a cell are known in the art, including but not limited to deletion of all or part of the gene encoding the protein and site-directed mutagenesis to disrupt the expression or activity of the gene product. (See, for example, Chaveroche et al., Nucl. Acids Res., 28:22 e97

[2000] , Cho et al., Molec. Plant Microbe Interact., 19:7-15

[2006] , Maruyama and Kitamoto, Biotechnol. Lett., 30:1811-1817

[2008] , Takahashi et al., Mol. Gen. Genom., 272:344-352

[2004] , and You et al., Arch. Microbiol., 191:615-622

[2009] , all of which are incorporated herein by reference).Random mutagenesis, followed by screening for the desired mutations, is also used (see, for example, Combier et al., FEMS Microbiol. Lett., 220:141-8

[2003] , and Firon et al., Eukary. Cell 2:247-55

[2003] , both of which are incorporated by reference).

[0142] Introduction of the vector or DNA construct into the host cell can be accomplished using any suitable method known in the art, including, but not limited to, calcium phosphate transfection, DEAE-dextran-mediated transfection, PEG-mediated transformation, electroporation, or other common techniques known in the art.

[0143] In some embodiments, the engineered host cells of the invention (i.e., "recombinant host cells") are cultured in a conventional nutrient medium that has been appropriately modified for activation of the promoter, selection of transformants, or amplification of the FoPip4H polynucleotide. Culture conditions such as temperature and pH are those previously used with the host cells selected for expression and are well known to those of skill in the art. As noted above, many standard references and texts are available for the culture and production of many cells, including bacteria, plants, animals (especially mammals), and archaebacterial-derived cells.

[0144] In some embodiments, cells expressing the FoPip4H enzyme of the present disclosure are grown under batch or continuous fermentation conditions. Classical "batch fermentation" is a closed system where the composition of the medium is set at the start of fermentation and undergoes no artificial changes during fermentation. A variation of the batch system, also used in the present invention, is "fed-batch fermentation." In this variation, substrate is gradually added as fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit cell metabolism and it is desirable to have a limited amount of substrate in the medium. Batch and fed-batch fermentations are common and well-known in the art. "Continuous fermentation" is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is simultaneously removed for processing. Continuous fermentation generally maintains the culture at a constant high density where the cells are primarily in the logarithmic growth phase. Continuous fermentation systems strive to maintain steady-state growth conditions. Methods for regulating nutrients and growth factors for continuous fermentation processes, as well as techniques for maximizing product formation rates, are well-known in the field of industrial microbiology.

[0145] Two or more copies of the nucleic acid sequence of the present invention may be inserted into a host cell to increase the production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene in the nucleic acid sequence. Cells containing amplified copies of the selectable marker gene and thereby additional copies of the nucleic acid sequence can be selected by culturing the cells in the presence of an appropriate selectable agent.

[0146] In some embodiments of the present invention, cell-free transcription and translation systems are used for the production of the FoPip4H polypeptide. Several systems are commercially available and this method is well-known to those skilled in the art.

[0147] Methods for evolving the FoPip4H enzyme In some embodiments, to generate the FoPip4H polypeptide of the present disclosure, the FoPip4H enzyme that catalyzes the reduction reaction is obtained (or extracted) from E. coli. In some embodiments, the parental polynucleotide sequence is codon-optimized to enhance the expression of the FoPip4H polypeptide in a particular host cell. The parental polynucleotide sequence designated as SEQ ID NO: 3 was codon-optimized for expression in E. coli, the codon-optimized polynucleotide was cloned into an expression vector, and the expression of the FoPip4H gene was placed under the control of the T7 promoter. The T7 polymerase required to express the gene of interest is under the control of the lac promoter, and both the gene of interest and the T7 polymerase are subject to lacI repression. The presence of IPTG activates the expression of T7 polymerase, eliminates the repression, and results in the production of the FoPip4H gene. Clones expressing active FoPip4H in E. coli were identified and the genes were sequenced to confirm their identity.

[0148] The FoPip4H polypeptide of the present disclosure can be obtained by subjecting the polynucleotide encoding the parental sequence to mutagenesis and / or directed evolution methods. Exemplary directed evolution techniques are mutagenesis and / or DNA shuffling as described in Stemmer, 1994, Proc. Natl. Acad. Sci. USA 91:10747-10751, WO 95 / 22625, WO 97 / 20078, WO 97 / 35966, WO 98 / 27230, WO 00 / 42651, WO 01 / 75767, and U.S. Pat. No. 6,537,746. Other directed evolution procedures that can be used include, inter alia, staggered extension process (StEP), in vitro recombination (Zhao et al., 1998, Nat. Biotechnol. 16:258-261), mutagenic PCR (Caldwell et al., 1994, PCR Methods Appl. 3:S136-S140), and cassette mutagenesis (Black et al., 1996, Proc. Natl. Acad. Sci. USA 93:3525-3529).

[0149] Screen the clones obtained after mutagenesis treatment to find FoPip4H polypeptides having the desired improved enzyme properties. Measurement of enzyme activity from the expression library can be carried out using standard chemical analysis techniques for measuring substrates and products, such as UPLC-MS. If the desired improved enzyme property is thermostability, the enzyme activity can be measured after subjecting the enzyme preparation to a defined temperature and measuring the amount of enzyme activity remaining after heat treatment. Next, clones containing the polynucleotide encoding the FoPip4H polypeptide are isolated, sequenced to identify nucleotide sequence changes (if any), and used to express the enzyme in host cells.

[0150] When the polypeptide sequence is known, the polynucleotide encoding the enzyme can be prepared by standard solid-phase methods according to known synthetic methods. In some embodiments, fragments of up to about 100 bases can be synthesized individually and then ligated (e.g., by enzymatic or chemical ligation methods, or polymerase-mediated methods) to form any desired contiguous sequence. For example, the polynucleotides and oligonucleotides of the present invention can be prepared by chemical synthesis using, for example, the classical phosphoramidite method described in Beaucage et al., 1981, Tet. Lett. 22:1859-69, or the method described in Matthes et al., 1984, EMBO J. 3:801-05, typically implemented by an automated synthesis method. According to the phosphoramidite method, oligonucleotides are synthesized, purified, annealed, ligated, and cloned into an appropriate vector, for example, using an automated DNA synthesizer. In addition, essentially any nucleic acid can be obtained from any of a variety of commercial sources, such as The Midland Certified Reagent Company, Midland, Texas, The Great American Gene Company, Ramona, California, ExpressGen Inc., Chicago, Illinois, Operon Technologies Inc., Alameda, California, and many others.

[0151] The FoPip4H enzyme expressed in host cells can be recovered from cells and / or the medium using any one or more of the well-known techniques for protein purification, including, among others, lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography. Solutions suitable for lysis and high-efficiency extraction of proteins from bacteria such as E. coli are commercially available under the trade name CelLytic B® from Sigma-Aldrich, St. Louis, Missouri.

[0152] As chromatography techniques for isolating the FoPip4H polypeptide, there are, inter alia, reverse phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. The conditions for purifying a specific enzyme will depend, in part, on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and will be apparent to those skilled in the art.

[0153] In some embodiments, affinity techniques may be used to isolate an improved FoPip4H enzyme. In the case of affinity chromatography purification, the protein sequence can be tagged with a recognition sequence to enable purification. Common tags include cellulose binding domains, polyHis tags, di-His chelates, FLAG tags, and many other tags that will be apparent to those skilled in the art. Antibodies can also be used as affinity purification reagents. Any antibody that specifically binds to the FoPip4H polypeptide may be used.

[0154] Processes using the FoPip4H enzyme The FoPip4H enzyme described herein catalyzes the hydroxylation of substituted indanone substrate (II)

Chemical formula

Chemical formula

[0155] Hydroxyindanone (I) is an intermediate for the synthesis of belzutifan (WELIREG). Thus, in a process for preparing belzutifan, the process can include the step of converting substituted indanone (II) to substituted hydroxyindanone (I) using the FoPip4H enzyme disclosed herein.

[0156] In the embodiments and examples shown in this specification, various ranges of suitable reaction conditions that can be used in the process include, but are not limited to, substrate loading, co-substrate loading, reducing agent, divalent transition metal, pH, temperature, buffer, solvent system, polypeptide loading, and reaction time. Further suitable reaction conditions for carrying out a process for biocatalytic conversion of a substrate compound to a product compound using the engineered FoPip4H polypeptide described herein can be readily optimized by routine experimentation including, but not limited to, contacting the engineered FoPip4H polypeptide with the substrate compound under experimental reaction conditions of concentration, pH, temperature, and solvent conditions, and detecting the product compound, taking into account the guidance provided herein.

[0157] Suitable reaction conditions for using the engineered FoPip4H polypeptide typically include a co-substrate that is stoichiometrically used in the hydroxylation reaction. Generally, the co-substrate for the FoPip4H enzyme is α-ketoglutaric acid, also known as α-ketoglutarate. Other analogs of α-ketoglutaric acid that can function as a co-substrate for the FoPip4 enzyme can be used. An exemplary analog that can function as a co-substrate is 2-oxoadipic acid. Since the co-substrate is used stoichiometrically, the co-substrate is present in an equimolar or greater amount than the molar concentration of the substrate compound, i.e., the molar concentration of the co-substrate is equal to or greater than the molar concentration of the substrate compound. In some embodiments, suitable reaction conditions can include a co-substrate molar concentration that is at least 1-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, or 5-fold, or more than the molar concentration of the substrate compound.

[0158] The substrate compound in the reaction mixture can be varied, for example, considering the amount of the desired product compound, the influence of the substrate concentration on the enzyme activity, the stability of the enzyme under the reaction conditions, and the conversion rate of the substrate to the product. In some embodiments, suitable reaction conditions include a substrate compound loading of 0.5 to 200 g / L, 1 to 200 g / L, 5 to 100 g / L, or 10 to 50 g / L. The values of the substrate loading provided herein are based on the molecular weight of the substituted indanone (II), although various hydrates and salts of the equivalent molar amount of substituted indanone (II) are also considered to be usable in this process.

[0159] When carrying out the FoPip4H enzyme-mediated process described herein, the engineered polypeptide can be added to the reaction mixture in the form of a purified enzyme, a partially purified enzyme, the whole cells transformed with the gene(s) encoding the enzyme, the cell extract and / or lysate of such cells, and / or the enzyme immobilized on a solid support. The whole cells transformed with the gene(s) encoding the engineered FoPip4H enzyme or their cell extracts, lysates, and isolated enzymes can be employed in a variety of different forms, including solids (e.g., freeze-drying and spray-drying, etc.) or semi-solids (e.g., crude paste). The cell extract or cell lysate is partially purified (e.g., by ammonium sulfate precipitation, polyethyleneimine, or heat treatment), followed by a desalting procedure (e.g., ultrafiltration and dialysis, etc.) before freeze-drying. Any of the enzyme preparations (including whole cell preparations) can be stabilized, for example, by cross-linking using a known cross-linking agent such as glutaraldehyde, or immobilization on a solid phase (e.g., Eupergit C, etc.).

[0160] The improved activity and / or stereoselectivity of the engineered FoPip4H polypeptides disclosed herein provides a process that can achieve higher conversion rates with lower concentrations of the engineered polypeptides. In some embodiments of this process, suitable reaction conditions include an amount of engineered polypeptide of 1% (w / w) to 100% (w / w), 1% (w / w) to 30% (w / w), 2.5% (w / w) to 20% (w / w), or 2.5% (w / w) to 10% (w / w). In certain embodiments of the process, suitable reaction conditions include an amount of engineered polypeptide of 1% (w / w), 2% (w / w), 5% (w / w), 7.5% (w / w), 10% (w / w), 20% (w / w), 30% (w / w), 40% (w / w), 50% (w / w), 75% (w / w), 100% (w / w), or more of the substrate compound load.

[0161] In some embodiments, the engineered polypeptide is present at 0.01 g / L to 50 g / L, 0.1 g / L to 40 g / L, 1 g / L to 40 g / L, or 2 g / L to 10 g / L.

[0162] In some embodiments, the reaction conditions also include a divalent transition metal that can function as a cofactor in the oxidation reaction. Generally, the divalent transition metal cofactor is ferrous ion, i.e., Fe +2 . Ferrous ion can be provided in various forms such as ferrous sulfate (FeSO4), ferrous chloride (FeCl2), ferrous carbonate (FeCO3), and salts of organic acids such as citrate, lactate, and fumarate. An exemplary source of ferrous sulfate is Mohr's salt, which is ammonium ferrous sulfate (NH4)2Fe(SO4)2 and is available in anhydrous and hydrated (i.e., hexahydrate) forms. Ferrous ion is a transition metal cofactor commonly found in naturally occurring FoPip4 enzymes and functions efficiently in the engineered enzyme, but it should be understood that other divalent transition metals that can act as cofactors can be used in the process. In some embodiments, the divalent transition metal cofactor is Mn +2 and Cr +2It may include. In some embodiments, the reaction conditions are a divalent transition metal cofactor at a concentration of 0.1 mM to 100 mM, 0.5 mM to 80 mM, 10 mM to 60 mM, or 40 to 60 mM, particularly Fe +2 It may include.

[0163] In some embodiments, the reaction conditions may further include a reducing agent capable of reducing the second iron ion Fe +3 to the first iron ion Fe +2 . In some embodiments, the reducing agent is L-cysteine, ascorbic acid, dithiothreitol, D-cysteine, L-homocysteine, or D-cysteine ethyl ester. In some embodiments, the reducing agent includes cysteine, typically L-cysteine. Cysteine is not required for the hydroxylation reaction, but the enzyme activity is enhanced in its presence. Without being bound by theory, cysteine is thought to maintain or regenerate the enzyme-Fe +2 form that mediates the hydroxylation reaction. Generally, the reaction conditions may include a corresponding cysteine concentration proportional to the substrate load. In some embodiments, cysteine is present at at least 0.1-fold, 0.2-fold, 0.3-fold, or at least 0.5-fold the molar amount of the substrate. In some embodiments, the reducing agent, particularly L-cysteine, is at a concentration of 1 mM to 100 mM, 5 mM to 80 mM, or 50 to 70 mM.

[0164] In some embodiments, the reaction conditions include molecular oxygen, i.e., O2. Without being bound by theory, one atom of oxygen from molecular oxygen is incorporated into the substrate compound to form the hydroxylation product compound. O2 may be naturally present in the reaction solution or may be artificially introduced and / or supplemented to the reaction. In some embodiments, the reaction conditions may include forced aeration (e.g., sparging) with air, O2 gas, or other O2-containing gas. In some embodiments, the O2 during the reaction can be increased by increasing the pressure of the reaction with O2 or an O2-containing gas. This can be done by performing the reaction in a vessel that can be pressurized with O2 gas.

[0165] During the reaction process, the pH of the reaction mixture can change. The pH of the reaction mixture can be maintained within a desired pH or a desired pH range. This can be done by adding an acid or a base before and / or during the reaction. Alternatively, a buffer solution may be used to control the pH. Thus, in some embodiments, the reaction conditions include a buffer solution. Suitable buffer solutions for maintaining the desired pH range are known in the art and include, by way of example and not limitation, borate, phosphate, 2-(N-morpholino)ethanesulfonic acid (MES), 3-(N-morpholino)propanesulfonic acid (MOPS), acetate, triethanolamine, and 2-amino-2-hydroxymethyl-propane-1,3-diol (Tris). In some embodiments, the buffer solution is phosphate. In some embodiments of the process, suitable reaction conditions include a buffer (e.g., phosphate) concentration of about 0.001 to about 0.2 M, 0.003 to about 0.1 M, or 0.005 to about 0.05 M. In some embodiments, the reaction conditions include a buffer (e.g., phosphate) concentration of 0.001, 0.002, 0.003, 0.004, 0.005, or 0.008 M. In some embodiments, the reaction conditions include water as a suitable solvent in the absence of a buffer solution.

[0166] In embodiments of the process, the reaction conditions can include a suitable pH. The desired pH or desired pH range can be maintained by using an acid or a base, a suitable buffer solution, or a combination of buffer and acid or base addition. The pH of the reaction mixture can be controlled before and / or during the reaction. In some embodiments, suitable reaction conditions include a solution pH of about 4 to about 10, a pH of about 5 to about 10, a pH of about 5 to about 9, a pH of about 6 to about 9, a pH of about 6 to about 8. In some embodiments, the reaction conditions include a solution pH of about 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10.

[0167] In the process embodiments of this specification, for example, a suitable temperature can be used as a reaction condition in consideration of the increase in reaction rate at a higher temperature and the activity of the enzyme during the reaction period. Therefore, in some embodiments, the suitable reaction conditions include a temperature of 10°C to 30°C, for example, 25°C to 30°C. In some embodiments, the temperature during the enzyme reaction can be maintained at a specific temperature throughout the reaction process. In some embodiments, the temperature during the enzyme reaction can be adjusted over the temperature profile during the reaction process.

[0168] The processes of the present disclosure are generally carried out in a solvent. Suitable solvents include water, aqueous buffer solutions, organic solvents, polymer solvents, and / or cosolvent systems, which generally include aqueous solvents, organic solvents, and / or polymer solvents. The aqueous solvent (water or an aqueous cosolvent system) may or may not be pH buffered. In some embodiments, the process using the engineered FoPip4H polypeptide is carried out in an aqueous cosolvent system containing an organic solvent (e.g., ethanol, isopropanol (i-PrOH), dimethyl sulfoxide (DMSO), dimethylformamide (DMF), ethyl acetate, butyl acetate, 1-octanol, heptane, octane, methyl t-butyl ether (MTBE), and toluene, etc.), an ionic or polar solvent (e.g., 1-ethyl-4-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium tetrafluoroborate, 1-butyl-3-methylimidazolium hexafluorophosphate, glycerol, and polyethylene glycol, etc.). In some embodiments, the cosolvent can be a polar solvent such as a polyol, dimethyl sulfoxide (DMSO), or a lower alcohol. The non-aqueous cosolvent component of the aqueous cosolvent system may be miscible with the aqueous component and provide a single liquid phase, or may be partially miscible or immiscible with the aqueous component and provide two liquid phases. Exemplary aqueous cosolvent systems can include water and one or more cosolvents selected from organic solvents, polar solvents, and polyol solvents. Generally, the cosolvent component of the aqueous cosolvent system is selected so as not to unfavorably inactivate the FoPip4H enzyme under the reaction conditions. Suitable cosolvent systems can be readily identified by measuring the enzyme activity of a particular engineered FoPip4H enzyme using a defined substrate of interest in a candidate solvent system using an enzyme activity assay such as those described herein.

[0169] In one embodiment, the process using the FoPip4H polypeptide can be carried out in a cosolvent system of water and 1-octanol. The ratio of water to octanol can be 5:1 to 50:1, such as 10:1 to 40:1 or 20:1 to 40:1.

[0170] In some embodiments, the reaction conditions may include an antifoaming agent, which helps to reduce or prevent the formation of bubbles in the reaction solution, such as when the reaction solution is mixed or sparged. Antifoaming agents include non-polar oils (e.g., mineral, silicone, etc.), polar oils (e.g., fatty acids, alkyl amines, alkyl amides, alkyl sulfates, etc.), and hydrophobic substances (e.g., treated silica, polypropylene, etc.), some of which also function as surfactants. Exemplary antifoaming agents include Glanapon 2000 Konz, Y-30 (registered trademark) (Dow Corning), polyglycol copolymers, oxy / ethoxylated alcohols, and polydimethylsiloxane. In some embodiments, the antifoaming agent may be present at 0.001% (v / v) to about 1% (v / v), 0.003 to 0.02% (v / v), or 0.005 to 0.01% (v / v).

[0171] In some embodiments, the reaction conditions may include a surfactant to stabilize or enhance the reaction. The surfactant may include non-ionic, cationic, anionic, and / or amphiphilic surfactants. Exemplary surfactants include, by way of example and not limitation, nonylphenoxypolyethoxyethanol (NP40), Triton X-100, polyoxyethylene-stearylamine, cetyltrimethylammonium bromide, sodium oleylamide sulfate, polyoxyethylene-sorbitan monostearate, hexadecyldimethylamine, and the like. Any surfactant that can stabilize or enhance the reaction may be employed. The concentration of the surfactant employed in the reaction may generally be 0.1 to 50 mg / ml, particularly 1 to 20 mg / ml.

[0172] The amount of reactants used in the hydroxylase reaction generally varies depending on the amount of the desired product and the amount of FoPip4H substrate employed simultaneously. Those skilled in the art will readily understand how to vary these amounts to suit the desired level of productivity and production scale.

[0173] [Examples]

[0174] [Example 1] Enzyme preparation Escherichia coli cultures each having a plasmid encoding the FoPip4 enzyme, which can be represented by the amino acid sequences shown in SEQ ID NOs: 1, 2, and 4 to 19 below, were serially diluted with Luria-Bertani broth (cell culture medium) as a diluent to 10 -4 , 10 -5 , and 10 -6 . 100 μL of each dilution was spread on a Petri dish containing LB agar medium supplemented with 30 μg / mL kanamycin and 1% (w / v) glucose. The plates were placed in an incubator at 37 °C overnight.

[0175] 200 μL / well of Luria-Bertani broth (cell culture medium) (500 mL LB + 30 μg / mL kanamycin + 1% (w / v) glucose) was dispensed into labeled 96-well shallow-well plates. The shallow-well plates were loaded into the plate stacker of a colony picker. An agar medium plate containing colonies that were sufficiently diluted so that most of the colonies were isolated from each other (known to those skilled in the art as single colonies) was sampled into the respective wells of the shallow-well plates. The colonies were grown overnight at 200 rpm, 30 °C, and 85% RH.

[0176] 390 μL of terrific broth (TB) growth medium (commercially available from ThermoFisher Scientific as catalog number A1374301) (TB + 50 μg / mL kanamycin) was dispensed into labeled 96-well deep-well subculture plates. 13 μL of the overnight-grown culture was transferred from each well of the master shallow-well plate to the corresponding labeled deep-well subculture plate. The plates were sealed with breathable film and shaken at 250 rpm, 30 °C, and 85% RH for 2 - 2.5 h. After shaking, the optical density (OD 600 , optical density at a wavelength of 600 nm) of at least one plate was measured to examine growth. The OD 600When it was in the range of 0.4 to 0.8, the deep well plate was induced with an induction medium (2.2 mM IPTG, final concentration 0.2 mM) of 40 μL / well. The plate was resealed and incubated at 250 rpm, 30 °C, and 85% RH for 18 - 20 h with shaking.

[0177] Alternatively, instead of the TB growth medium using the induction medium, Studier induction ZYM - 5052 medium was also used for expression. 390 μL of ZYM - 5052 growth medium (commercially available from Teknova as catalog number 3S2000) was dispensed into a labeled 96 - well deep well sub - culture plate. 13 μL of an overnight - grown culture was transferred from each well of the master shallow well plate to the corresponding labeled deep well sub - culture plate. The plate was sealed with a breathable film and shaken at 250 rpm, 30 °C, and 85% RH for 20 - 22 h.

[0178] After incubation, all deep well plates were centrifuged at 4 °C, 4000 rpm for 15 min. After centrifugation, the supernatant was discarded. The plates containing the cell pellets were heat - sealed and stored at - 80 °C.

[0179] The plates containing the cell pellets were taken out from the - 80 °C storage location and thawed at RT. A lysis buffer of 10 mM potassium phosphate pH 7.0, 0.5 mg / mL lysozyme, 0.25 mg / mL polymyxin B sulfate (PMBS), 1 unit / mL DNase I, and 4 mM MgSO4 was prepared. 400 μL of the lysis buffer was dispensed into each well. The lysis mixture was shaken at 1000 rpm on a plate shaker at RT for 1.5 - 2 h. Then, the lysis mixture was centrifuged at 4 °C, 4000 xg for 10 min to prepare an enzyme - containing lysate solution (in the supernatant).

[0180] [Example 2] Hydroxylase Reaction in Well Plates Using a Chemspeed automated solid dispenser, 5 mg of the ground substrate was dispensed into each well of a 1 mL round-bottom well plate. A reaction buffer was prepared by mixing 448 mM of α-ketoglutaric acid, 63 mM of L-cysteine, and 49 mM of mol salt in 10 mM potassium phosphate and adjusted to a final pH of 6.5.

[0181] 115 μL of the reaction buffer was added to a 1 mL round-bottom well plate. Then, 10 μL of the enzyme-containing lysate of Example 1 was added. The plate was sealed with an air-permeable seal and shaken overnight at 30 °C, 250 rpm, and 85% RH.

[0182] For the thermal stability test, the same reaction apparatus as above was used, but 50 μL of the enzyme-containing lysate solution of Example 1 was first transferred to a Biorad hard-shell PCR plate, and the enzyme was heat-treated at T50 for 10 minutes, and then 10 μL of the heat-denatured lysate was transferred to the reaction plate to initiate the reaction. The heat-treated supernatant was assayed for residual activity compared to the untreated supernatant.

[0183] After shaking overnight, 250 μL of DMSO was added to each well of the reaction plate. The plate was heat-sealed and shaken at 1000 rpm for 10 minutes at RT. 180 μL of a 50% ACN / water mixture was dispensed onto a filter stack (a filter plate on top of a round-bottom plate with 0.20 μM hydrophilic PTFE commercially available from Millipore MSRLN2250). The reaction plate was unsealed. 10 μL of the reaction mixture from the reaction plate was transferred to the corresponding filter plate with the receiving plate below. Then, the plate was centrifuged at 4 °C, 4000 rpm for 3 min. The filter plate was removed, and the clear solution in the receiving plate was heat-sealed.

[0184] The filtered solution was analyzed by ultra-high performance liquid chromatography (UPLC) using a high-throughput screening method to observe the substrate, product, and the peak area of the peroxide of the product. The structure of the peroxide product is shown below.

Chemical formula

[0185] UPLC was performed using the following gradient method on an Acquity UPLC

[0001] HSS T3 1.8 μm 2.1×50 mm column, with mobile phase component A being water containing 0.1% TFA and component B being ACN containing 0.1% TFA. From 0 to 0.7 min, 12% B; from 0.7 to 0.8 min, a gradient from 12% to 95% B; from 0.8 to 1.2 min, held at 95% B; from 1.2 to 1.25 min, a gradient from 95% to 12% B; and from 1.25 to 1.5 min, held at 12% B. The column temperature was set at 40 °C, the flow rate was set at 0.75 mL / min, and absorbances at 254 and 290 nm were measured. The starting material eluted at 1.1 min, the desired product eluted at 0.61 min, and the undesired peroxide product eluted at 0.48 min.

[0186] In addition to UPLC analysis, plate reader analysis was also used as a method for measuring peroxidation. After an overnight reaction, the plates were centrifuged at 1000 xg for 10 minutes. 190 μL of water was dispensed into each well of a Greiner 96-well black plate. 10 μL of the overnight reaction product was transferred to the Greiner plate together with water. Absorbances at 300 and 400 nm were measured, and the ratio of 300 to 400 nm was used as a surrogate for the relative amount of the desired product relative to the undesired peroxidation.

[0187] [Example 3] Enzyme Preparation in a Shaker Flask 10 μL of E. coli cells, each containing a plasmid encoding a FoPip4H enzyme that can be represented by the amino acid sequences shown below for SEQ ID NOs: 1, 2, and 4 - 19, were inoculated into 5 mL of Luria-Bertani broth (cell culture medium) (250 mL LB + 30 μg / mL kanamycin + 1% glucose), which had been dispensed into 15 mL cell culture tubes that were stored at -80 °C in 20% glycerol and labeled. The cell culture tubes were sealed and incubated with shaking at 30 °C and 250 rpm for 20 - 24 h.

[0188] After overnight growth, an overnight growth culture (2 - 5 mL of cell culture (having an initial OD of 0.2)) was added to 250 mL of terrific broth (TB) growth medium (commercially available from ThermoFisher Scientific under catalog number A1374301) (TB + 30 μg / mL kanamycin) to bring the final volume to 250 mL. The flask was shaken at 30 °C and 250 rpm for 3 - 4 h. After shaking, OD was measured for growth until it reached 0.4 - 0.6. 600 At this point, 0.2 mM IPTG (50 μL of 1 M IPTG) was added to the culture to induce expression, and the culture was grown at 30 °C and 250 rpm for 20 - 24 h. 600 After further growth period, the culture was transferred to a centrifuge bottle of known weight and centrifuged at 4 °C and 4000 rpm for 20 min. After centrifugation, the supernatant was discarded and the remaining cell pellet in the bottle was weighed. The weight of the cell pellet was calculated by subtracting the weight of the known bottle, and the cell pellet was resuspended in 5 mL of 10 mM potassium phosphate buffer (pH = 7) per gram of wet cell pellet. 600 Cells from the resuspended cell pellet were lysed using a microfluidizer, the cell lysate was recovered and centrifuged at 4 °C and 10000 rpm for 60 min. The clarified supernatant was transferred to a Petri dish and frozen at -80 °C for about 2 h. The sample was lyophilized using a standard automated protocol.

[0189]

[0190]

[0191] [Example 4] Hydroxylase Reaction, Conversion, and Selectivity Determination in Vial Using the selected variants of the FoPip4H enzyme, the hydroxylase reaction was carried out in a vial with the indanone substrate of formula (II) to further evaluate the conversion and selectivity (i.e., enantioselectivity) of the variants. To a 50 mL vial, α-ketoglutaric acid (2.8 g, 19.2 mmol), molybdate (0.86 g, 2.2 mmol), and cysteine (0.34 g, 2.8 mmol) were added. The mixture was dissolved in 25 mL of water, the pH was adjusted to 6.0 with 5N NaOH, and the reaction solution was obtained by bringing the final volume to 44.5 mL with water. To a new 50 mL vial, the indanone substrate of formula (II) (0.2 g, 0.88 mmol) and the FoPip4H enzyme (15 mg, 7.5 wt%, from Example 3) were added to obtain a solid vial. 4.45 mL of the reaction solution was transferred to the solid vial, and 0.3 mL of 1-octanol was added. The solution was stirred for 24 hours before sampling.

[0192] Conversion was determined by comparing the peak areas of the product and starting material by Agilent UPLC equipped with a Waters Aquity HSS T3 column. The mobile phase was 0.1% phosphoric acid in water and acetonitrile. The product and starting material were identified by comparison with pure standards.

[0193] Enantioselectivity (ee) was determined by comparing the peak areas of the enantiomers separated on a Waters SFC column. The SFC column was equipped with a Chiralpak AD-3 column, and methanol and CO2 were used as the mobile phase. The enantiomers were determined by comparison with pure standards.

[0194] [Example 5] Determination of T50 15 mg of each selected FoPip4H enzyme was dissolved in 10 mM potassium phosphate buffer (pH 6.0). Separate solutions were held at temperatures in the range of 40 - 65 °C for 1 hour and then centrifuged to clarify the mixture. Then, the solutions were used to set up the assay described in Example 2.

[0195] [Example 6] Preparation of (R)-4-Fluoro-3-hydroxy-7-(methylsulfonyl)-2,3-dihydro-1H-inden-1-one (I) [Chemical formula]

[0196] Water (34.5 L, 23 vol) was added to a 100 L reactor and adjusted to 20 - 30 °C. After adding α-ketoglutaric acid (2.11 kg, 14.46 mol, 2.20 eq), the solution was sparged with N2 until the dissolved oxygen (dO) level reached <2%. L-Cysteine (0.26 kg, 2.10 mol, 0.32 eq) was added and the pH was adjusted to 6.8 with 10N NaOH. Mohr's salt (0.64 kg, 1.64 mol, 0.25 eq) was added, followed by 5N NaOH to adjust the pH to 6.0. 1-Octanol (1.13 L, 0.75 vol) was added, followed by the addition of a lyophilized fermentation powder containing FoPip4H (0.11 kg, 7.5 wt%), and then sulfone (II) (1.5 kg, 6.57 mol, 1.00 eq) was added. The reaction dO level was set to 100%, and the suspension was stirred for 24 - 48 h. Then, sulfuric acid was added to adjust the pH to 4.7. To this mixture, (NH4)2(SO4) (9.0 kg, 68.1 mol) and 50% MeCN in toluene (42 L, 28 vol) were added. Then, CELITE (3.0 kg) was added, and the mixture was heated at 45 °C for 2 h and then cooled to 25 °C. The mixture was filtered to separate the organic layer. The filter cake was washed twice with 25% MeCN in toluene (15 L, 10 vol). The aqueous layer was back-extracted twice with the cake washing solvent, and the combined organic layers were washed with water (1.1 L, 0.75 vol). The organic layer was concentrated under vacuum to 10 vol at 50 °C, then stirred at 55 °C for 2 h, cooled to 25 °C over 4 h, and further stirred for 10 h. The slurry was filtered and washed three times with 5% MeCN in toluene (1.5 L, 1 vol). The cake was dried under vacuum to obtain 1.39 kg of hydroxysulfone (I) (99 wt%, 5.63 mol, 86% yield). 11H NMR (599.90 MHz, DMSO-d6) δ 8.11 (dd, J = 8.4 and 4.4 Hz, 1H, CH), 7.78 (t, J = 8.5 Hz, 1H, CH), 5.97 (d, J = 7.3 Hz, 1H, OH), 5.46 (td, J = 7.1 and 2.2 Hz, 1H, CH), 3.42 (s, 3H, CH3), 3.20 (dd, J = 18.8 and 6.8 Hz, 1H, CH’H”), 2.59 (dd, J = 18.8 and 2.2 Hz, 1H, CH’H”) ppm. 13 C{ 1 1H}NMR (150.85 MHz, DMSO-d6) δ 200.17 (s, C=O), 162.85 (d, J CF = 259.9 Hz, CF), 145.05 (d, J CF = 19.2 Hz, C), 136.02 (d, J CF = 5.1 Hz, C), 133.24 (d, J CF = 4.0 Hz, C), 132.32 (d, J CF = 8.5 Hz, CH), 121.39 (d, J CF = 21.0 Hz, CH), 63.85 (s, CH), 47.22 (s, CH2), 42.65 (s, CH3) ppm. 19 19F NMR (564.47 MHz, DMSO-d6) δ -110.86 (dd, J HF = 8.6 and 4.5 Hz, 1F) ppm.

[0197] Sequence: MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKKRGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQANPNRGYSFVGQENVANISGYEKGLGPLKTRDIKETVDFGSANDELVDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAENEFRILHYPAIPASELADGTATRIAEHTDFGTITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSIAYFAKPNRSASLFPLKEFIEEGVPCKYEDVTAWEWNNRRIEKLFSAEAKA (SEQ ID NO: 1)

[0198] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGVVKLINHPIPSESIHELFAQTKRFFNLPLETKMLAKHPEQALPARGYAFVGQENVANISGYEKGLPPLKTRDIKETVDFGSANDEKYDNLWVPEEELPGFRSFMEGFYELAFKTEMQILEALAIALGVSPDHLKSLHNRAHSELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDGVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTWPPSIKDGDGSEVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPPKYEDLTFEEWNNRRIEKLFSAEAKA (SEQ ID NO: 2)

[0199]

[0200] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKKRGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQANPNRGYLFVGQENVANISGYEKGLGPLKTRDIKETVDFGSANDELVDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAENEFRILHYPAIPASELADGTATRIAEHTDFGTITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSIAYFAKPNRSASLFPLKEFIEEGVPCKYEDVTAWEWNNRRIEKLFSAEAKA (SEQ ID NO: 4)

[0201] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKKRGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQANPARGYLFVGQENVANISGYEKGLGPLKTRDIKETVDFGSANDELVDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAENELRILHYPAIPASELADGTATRIAEHTDFGTITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSIAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAWEWNNRRIEKLFSAEAKA (SEQ ID NO: 5)

[0202] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKKRGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQALPARGYLFVGQENVANISGYEKGLGPLKTRDIKETVDFGSANDELYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAENELRILHYPAIPASELADGTATRIAEHTDFGTITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSIAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAWEWNNRRIEKLFSAEAKA(SEQ ID NO: 6)

[0203] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKKRGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQALPARGYLFVGQENVANISGYEKGLGPLKTRDIKETVDFGSANDELYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAENELRILHYPAIPASELADGTATRIAEHTDFGTITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSIAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAEEWNNRRIEKLFSAEAKA(SEQ ID NO: 7)

[0204] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQALPARGYLFVGQENVANISGYEKGLSPLKTRDIKETVDFGSANDELYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAENELRILHYPAIPASELADGTATRIAEHTDFGTITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAEEWNNRRIEKLFSAEAKA (SEQ ID NO: 8)

[0205] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQALPARGYLFVGQENVANISGYEKGLSPLKTRDIKETVDFGSANDELYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAHNELRILHYPAIPASELADGTATRIAEHTDFGTITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAEEWNNRRIEKLFSAEAKA (SEQ ID NO: 9)

[0206] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQPLPARGYLFVGQENVANISGYEKGLSPLKTRDIKETVDFGSANDELYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAHNELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAEEWNNRRIEKLFSAEAKA(SEQ ID NO: 10)

[0207] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGGVKLINHPIPSTSIHELFAQTKRFFNLPLETKMLAKHPPQALPARGYSFVGQENVANISGYEKGLSPLKTRDIKETVDFGSANDELYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAHNELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAEEWNNRRIEKLFSAEAKA(SEQ ID NO: 11)

[0208] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGGVKLINHPIPSESIHELFAQTKRFFNLPLETKMLAKHPPQALPARGYSFVGQENVANISGYEKGLSPLKTRDIKETVDFGSANDELYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAHNELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAEEWNNRRIEKLFSAEAKA(SEQ ID NO: 12)

[0209] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGGVKLINHPIPSESIHELFAQTKRFFNLPLETKMLAKHPPQALPARGYAFVGQENVANISGYEKGLSPLKTRDIKETVDFGSANDELYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAHNELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDSVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKAGDGEQVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAEEWNNRRIEKLFSAEAKA(SEQ ID NO: 13)

[0210] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGGVKLINHPIPSESIHELFAQTKRFFNLPLETKMLAKHPPQALPARGYAFVGQENVANISGYEKGLPPLKTRDIKETVDFGSANDEKYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAHNELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDGVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTYPPSIKDGDGEQVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPCKYEDVTAEEWNNRRIEKLFSAEAKA(SEQ ID NO: 14)

[0211] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGVVKLINHPIPSESIHELFAQTKRFFNLPLETKMLAKHPPQALPARGYAFVGQENVANISGWEKGLPPLKTRDIKETVDFGSANDEKYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAHNELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDGVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTWPPSIKDGDGTQVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPCKYEDLTAEEWNNRRIEKLFSAEAKA(SEQ ID NO: 15)

[0212] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGVVKLINHPIPSESIHELFAQTKRFFNLPLETKMLAKHPEQALPARGYAFVGQENVANISGYEKGLPPLKTRDIKETVDFGSANDEKYDNLWVPEEELPGFRSFMEGFYELAFKTEMQLLEALAIALGVSPDHLKSLHNRAHSELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDGVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTWPPSIKDGDGSEVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPPKYEDLTAEEWNNRRIEKLFSAEAKA(SEQ ID NO: 16)

[0213] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKARGVVKLINHPIPSESIHELFAQTKRFFNLPLETKMLAKHPEQALPARGYAFVGQENVANISGYEKGLPPLKTRDIKETVDFGSANDEKYDNLWVPEEELPGFRSFMEGFYELAFKTEMQILEALAIALGVSPDHLKSLHNRAHSELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDGVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTWPPSIKDGDGSEVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPPKYEDLTAEEWNNRRIEKLFSAEAKA(SEQ ID NO: 17)

[0214] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKAKGVVKLINHPIPAESIKELFAQTKRFFNLPLETKMLAKHPEQALPARGYAFVGQENVANISGYEKGLPPLKTRDIKETVDFGSANDEKYDNLWVPEEELPGFRSFMEGFYELAFKTEMQILEALAIALGVSPDHLKSLHNGAHSELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDGVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFKAACHRVTWPPSIKDGDGSEVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPPKYEDLTFEEWNNRRIEKLFSAEAKA(SEQ ID NO: 18)

[0215] MGSHHHHHHHHGSAALNADTLDMSLFFGTPSQKQDFCDSLLRLLKAKGVVKLINHPIPAESIKELFAQTKRFFNLPLETKMLAKHPEQALPARGYAFVGQENVANISGYEKGLPPLKTRDIKETVDFGSANDEKYDNLWVPEEELPGFRSFMEGFYELAFKTEMQILEALAIALGVSPDHLKSLHNSAHSELRILHYPAIPASELADGTATRIAEHTDFGSITMLFQDGVGGLQVEDQENLGTFNNVESASPTDIILNIGDSLQRLTNDTFRAACHRVTWPPSIKDGDGSEVIPERYSVAYFVKPNRSASLFPLKEFIEEGVPPKYEDLTFEEWNNRRIEKLFSAEAKA(SEQ ID NO: 19)

[0216] It will be understood that the various things described above as well as other features and functions, or alternatives thereof, can desirably be combined in many other different systems or applications. Also, various alternatives, modifications, variations, or improvements thereof that are presently unforeseen or unexpected may subsequently be made by those skilled in the art, and these are also intended to be encompassed by the following claims.

Claims

1. A polypeptide comprising an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

2.

2. The polypeptide according to claim 1, wherein the amino acid sequence has at least 95% sequence identity with SEQ ID NO:

2.

3. The polypeptide according to claim 1, wherein the amino acid sequence has at least 98% sequence identity with SEQ ID NO:

2.

4. The polypeptide according to claim 1, wherein the amino acid sequence consists of SEQ ID NO:

2.

5. The polypeptide according to claim 1, consisting of SEQ ID NO:

2.

6. A polynucleotide encoding the polypeptide according to any one of claims 1 to 5.

7. The polynucleotide according to claim 6, wherein the polynucleotide comprises SEQ ID NO:

3.

8. An expression vector comprising the polynucleotide according to any one of claims 6 and 7, which is operably linked to one or more control sequences suitable for directing the expression of the encoded polypeptide in a host cell.

9. The expression vector according to claim 8, wherein the control sequence comprises a promoter.

10. The expression vector according to claim 9, wherein the promoter comprises an E. coli promoter.

11. A host cell comprising the expression vector according to claim 8.

12. The host cell according to claim 11, wherein the host cell is E. coli.

13. A process for preparing a compound according to formula (I) with at least 60% ee, the process comprising contacting indanone according to formula (II) 【Chemical Formula 1】 in the presence of the polypeptide according to claim 1 with α-ketoglutaric acid to provide the compound of formula (I). [Chemical 2] with Process.

14. The process according to claim 13, further comprising a reducing agent selected from the group consisting of L-cysteine, ascorbic acid, dithiothreitol, D-cysteine, L-homocysteine, and D-cysteine ethyl ester.

15. The process according to claim 14, further comprising a buffer selected from the group consisting of phosphate buffer, 2-morpholinoethanesulfonic acid, bis-tris, PIPES, citrate, bicine, and TEOA.

16.

17. Mohr's salt ((NH 4 ), 2 Fe(SO 4 ), 2 6H 2 O) and the process according to claim 14, further comprising an iron salt selected from the group consisting of iron chloride. The process according to claim 13, wherein the FoPip4 hydroxylase comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:

2.

18. ​ The process according to claim 13, wherein the compound of formula (I) is prepared with an ee of at least 95%.

Citation Information

Patent Citations

  • Penicillin v amide hydrolase gene from fusarium oxysporum

    JP1996266290A

  • Pipecolinic acid position-4 hydroxylase and method for producing 4-hydroxyamino acid using same

    WO2015115398A1

  • US2022/0881407