Carboxyesterase biocatalysts

Engineered carboxylesterase enzymes from A. acidocaldarius esterase 2 overcome the limitations of conventional amide synthesis by enabling efficient, direct conversion of esters to amides in aqueous and alcohol solvents, enhancing production efficiency and reducing waste.

JP2025111490APending Publication Date: 2025-07-30GLAXOSMITHKLINE INTPROP DEV LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025062083
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-12-13
Filing Date
2025-04-03
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Conventional amide synthesis methods suffer from low atom economy, generation of unwanted by-products, lack of enantioselectivity and chemoselectivity, use of toxic coupling reagents, and limited substrate and solvent tolerance, making them inefficient and environmentally harmful.

Method used

Development of engineered carboxylesterase enzymes derived from A. acidocaldarius esterase 2 with enhanced amidation activity, allowing direct synthesis of amides from ester and amine precursors in aqueous and alcohol solvents, reducing the need for stoichiometric activators and organic solvents.

Benefits of technology

The engineered carboxylesterase enzymes achieve high conversion rates of ester substrates to amides under mild conditions, improving efficiency and reducing waste, while enabling large-scale production of valuable pharmaceutical intermediates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111490000058
    Figure 2025111490000058
  • Figure 2025111490000059
    Figure 2025111490000059
  • Figure 2025111490000060
    Figure 2025111490000060
Patent Text Reader

Abstract

To provide improved carboxyesterase biocatalysts, and to provide methods of using the biocatalysts to make amides.SOLUTION: Provided are carboxyesterase polypeptides having a specific sequence or comprising a functional fragment thereof, comprising the characteristic that residues corresponding to the specific amino acid sequence of the carboxyesterase polypeptide are selected from nonpolar residues, aromatic residues, and aliphatic residues. Also provided are polynucleotides encoding the carboxyesterase enzymes, host cells capable of expressing the engineered carboxyesterase enzymes, and methods for producing commercially valuable amides using the engineered carboxyesterase enzymes. Also provided are amides produced using the engineered carboxyesterase enzymes.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an improved carboxylesterase biocatalyst and a method for producing amides using the biocatalyst. Background of the Invention

[0002] The formation of amide bonds is one of the most common reactions in organic synthesis, and amides are commonly found in active pharmaceutical ingredients, bioactive molecules, synthetic polymers, peptides, and proteins. A study by the Novartis Institute for BioMedical Research found that the formation of amide bonds and related acylation chemistry accounted for 21.3% of all chemical reactions performed in pharmaceutical synthesis over the past 40 years (Schneider, et al., J. Med. Chem. 59, 4385-4402, 2016). Conventional amide synthesis methods use carboxylic acid and amine substrates and require stoichiometric coupling reagents. Since this reaction proceeds through highly reactive activated intermediates, unwanted side reactions can occur, leading to the formation of unwanted by-products such as urea. This results in a costly amide formation method with low atom economy and a significant amount of (metal-containing and often toxic) waste generated. Other drawbacks of conventional amide synthesis methods include the lack of enantioselectivity and chemoselectivity, the use of volatile or toxic coupling reagents, and the need to protect other functional groups present in the reactants.

[0003] Chemical catalytic approaches have been developed that eliminate the need for stoichiometric coupling reagents, thus improving atom economy and reducing the amount of waste generated (reviewed in Pattabiraman and Bode, Nature, 480, 471-479, 2011 and de Figueiredo, et al., Chemical Reviews, 12029-12122, 2016). Boronic acid-catalyzed reactions are the oldest approach to chemical amidation, in which transient activation of carboxylic acids by aryl boronic acids enables catalytic amide formation. However, these methods suffer from limited substrate scope, often require high temperatures, and in addition have extremely low solvent tolerance, limiting their broader use. More recent studies have involved the use of metal-catalyzed amidation, in which metal salts are used as Lewis acids for transient activation of carboxylic acids for the purpose of assisting amidation. To date, these studies have many of the same drawbacks as boronic acid-catalyzed reactions, requiring high temperatures, catalyst loading, and having limited solvent scope and substrate tolerance. Also explored were redox-based methods using either N-heterocyclic carbenes (NHCs) or metal catalysts that enable oxidative conversion of alcohols, aldehydes, ketones, or nitriles to their corresponding amides. Unfortunately, both metal catalysts and NHC catalysts are expensive, are themselves extremely toxic, often require noxious co-solvents, and generally have low functional group tolerance.

[0004] Lipases have been used as biocatalysts to form amide bonds in organic solvents by directly activating ester starting materials and then coupling them with amines. Advantageously, these enzymes generally have high enantioselectivity and thermal stability (for reviews, see Gotor, Bioorg Med Chem, 7, 2189-2197, 1999). However, most of the lipases currently under investigation appear to have narrow substrate specificities and furthermore must be used in dry organic solvents to prevent unwanted hydrolysis. This specificity problem is particularly pronounced in the synthesis of tertiary amides, and very few enzymes have been shown to be even slightly active (studied in van Pelt, Green, Chem. 13, 1791-1798, 2011).

[0005] To overcome such limitations, directed evolution methods are generally used in which enzyme variants are expressed and studied in a high-throughput manner. However, these enzymes are often derived from Pseudomonas or Bacillus strains and cannot be easily expressed in laboratory strains such as Escherichia coli (E. coli) or Saccharomyces cerevisiae (S. cerevisiae) where robust genetic engineering techniques exist. Furthermore, the need for dry solvents and molecular sieves makes directed evolution methods extremely difficult due to both the high water content of cell lysates and the technical difficulty of drying hundreds of reactions in parallel. BRIEF DESCRIPTION OF THE DRAWINGS

[0006]

Figure 1

Figure 2

[0007] In view of the limitations of the prior art, the inventors have developed a series of mutants from the wild-type carboxylesterase enzyme A. acidocaldarius esterase 2 (SEQ ID NO: 2) with high thermotolerance, and such mutant enzymes have an improved amidation activity that is more than 785,000 times that of the wild-type enzyme. Due to this dramatically changed activity, these mutant enzymes have substantial water and alcohol tolerance and enable large-scale direct synthesis of amides from simple ester and amine precursors. This direct synthesis strategy shortens the amide synthesis by 1-2 chemical steps, reduces the use of organic solvents, eliminates the use of stoichiometric activators, and overall shows a dramatic improvement in the chemical level of the art.

[0008] The present disclosure provides, in particular, polypeptides, polynucleotides encoding such polypeptides, and methods of using such polypeptides for biocatalytic conversion of ethyl oxazole-5-carboxylate (the "ester substrate") to (4-isopropylpiperazin-1-yl)(oxazol-5-yl)methanone (the "product") in the presence of 1-isopropylpiperazine (the "amine substrate"). The present disclosure also provides, in particular, polypeptides, polynucleotides encoding such polypeptides, and methods of using such polypeptides for biocatalytic conversion of ethyl oxazole-5-carboxylate (the "ester substrate") to ((2S,6R)-2,6-dimethylmorpholino)(oxazol-5-yl)methanone (another product) in the presence of the amine substrate cis-2,6-dimethylmorpholine. These products are starting materials for pharmaceuticals that are of interest in the development for the treatment of chronic obstructive pulmonary disease (COPD). Specifically, these products can be used to synthesize phosphoinositide 3-kinase δ inhibitors (PI3Kδ inhibitors), which are drug species used to treat inflammation, autoimmune diseases, and cancer. The compositions of the present invention can also be used as starting materials for other pharmaceutical species.

[0009] As shown in Table 3, the wild-type polypeptide, carboxylesterase enzyme A. acidocaldarius esterase 2 (SEQ ID NO: 2), acts on ester substrates with extremely low efficiency (less than 1% of the substrate is converted to the product), but the engineered carboxylesterase (E.C. 3.1.1) of the present disclosure can readily convert ester substrates to products in the presence of amine substrates. Thus, in one aspect, the present disclosure relates to an improved carboxylesterase that can convert ethyl oxazole-5-carboxylate, an ester substrate, to (4-isopropylpiperazin-1-yl)(oxazol-5-yl)methanone, the product, to a measurable level of up to about 0.1% conversion by analytical techniques such as HPLC-UV absorbance in the presence of 1-isopropylpiperazine "amine substrate".

[0010] In some embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to the amino acid sequence shown in SEQ ID NO: 4, or a functional fragment thereof, and the improved carboxylesterase amino acid sequence comprises the feature that the residue corresponding to X198 is selected from nonpolar residues, aromatic residues, and aliphatic residues. Guidance on the selection of the various amino acid residues that may be present at the designated residue positions is provided in the following detailed description and claims.

[0011] In some embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence corresponding to the amino acid sequence shown in SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0012] In some embodiments, the improved carboxylesterase polypeptide consists of the amino acid sequence shown in SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0013] In some embodiments, the present disclosure provides a carboxylesterase polypeptide comprising the amino acid sequence shown in SEQ ID NO: 122. In another embodiment, the present disclosure provides a polynucleotide sequence encoding the carboxylesterase polypeptide sequence shown in SEQ ID NO: 122. In yet another embodiment, the present disclosure provides a polynucleotide encoding a carboxylesterase polypeptide, the polynucleotide comprising the polynucleotide sequence shown in SEQ ID NO: 121. In yet another embodiment, the polynucleotide encodes a carboxylesterase polypeptide consisting of the polynucleotide sequence shown in SEQ ID NO: 121.

[0014] In another aspect, the improved carboxylesterase polypeptide can be used in a method for producing an amide, (a) R1-COOR2 [where R1 is sp with 0 to 3 alkyl substituents 3 selected from carbon and aromatic rings, and R2 is selected from a methyl group; an ethyl group; and a C1-C6 alkyl chain], an ester of the structure; (b) an amine substrate; (c) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity or more with the amino acid sequence shown in SEQ ID NO: 4, or a functional fragment thereof, an improved carboxylesterase polypeptide (the carboxylesterase polypeptide amino acid sequence thereof includes the feature that the residue corresponding to X198 of SEQ ID NO: 4 is selected from non-polar residues, aromatic residues, and aliphatic residues); and (d) a component containing a solvent are combined. Detailed Description of the Invention

[0015] The present disclosure provides a highly efficient biocatalyst that can mediate the conversion including amidation of a specific amide group acceptor, for example, the synthesis of a compound of formula III. The biocatalyst is an engineered amidation polypeptide that can convert ethyl oxazole-5-carboxylate (the "ester substrate"), which is a substrate of formula I, into (4-isopropylpiperazin-1-yl)(oxazol-5-yl)methanone (the "product"), which is a product of formula III, in the presence of an amine substrate of formula II (1-isopropylpiperazine).

Chemical Formula

[0016] In certain embodiments, the engineered carboxylesterase is derived from a native carboxylesterase from A. acidocaldarius esterase 2 and is a carboxylesterase enzyme that catalyzes the hydrolysis of esters via the formation and breakdown of an acyl-enzyme intermediate amine. The carboxylesterase of SEQ ID NO: 4 differs from the native enzyme derived from A. acidocaldarius esterase 2 (SEQ ID NO: 2), which is a wild-type carboxylesterase, in having a substitution of leucine (L) for glutamate (E) at residue position X198 and having measurable activity against the ester substrate ethyl oxazole-5-carboxylate (Formula I). The carboxylesterase of SEQ ID NO: 4 has been engineered to mediate the efficient conversion of the ester substrate of Formula I to the product of Formula III in the presence of an amine substrate such as 1-isopropylpiperazine (Formula II). This conversion can be carried out under mild conditions (30 °C, high conversion %), which makes this method applicable to the large-scale production of the amides of Formulas III and V.

[0017] Abbreviations and Definitions As used herein, the abbreviations used for genetically encoded amino acids are conventional and are as shown in Table 1.

[0018]

Table 1

[0019] When using three-letter abbreviations, if "L" or "D" is not specifically prefixed, or if it is not clear from the context in which the abbreviation is used, the amino acid can be either L- or D-type with respect to the α-carbon (Cα). For example, "Ala" represents alanine without specifying the configuration with respect to the α-carbon, while "D-Ala" and "L-Ala" represent D-alanine and L-alanine, respectively. When using one-letter abbreviations, capital letters represent L-type amino acids with respect to the α-carbon, and lowercase letters represent D-type amino acids with respect to the α-carbon. For example, "A" represents L-alanine and "a" represents D-alanine. Peptide sequences are represented as a series of one-letter or three-letter abbreviations (or a mixture thereof), and these sequences are shown in the N→C direction according to convention.

[0020] Technical and scientific terms used herein have the meanings generally understood by those skilled in the art, unless otherwise specifically defined. Thus, the following terms are intended to have the following meanings. All U.S. patents and published U.S. patent applications cited herein are hereby expressly incorporated by reference in their entirety, including all sequences disclosed within such patents and patent applications.

[0021] "Acid by-product" or "hydrolysis by-product" refers to a carboxylic acid resulting from the reaction of an ester substrate and water as a result of carboxylesterase enzyme. The acid by-product is a molecule of general formula (3) where R3 is -H. R1 is as described above. [Chemical formula]

[0022] "Acidic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain with a pK value of less than about 6 when the amino acid is included in a peptide or polypeptide. Acidic amino acids generally have a negatively charged side chain at physiological pH due to the loss of a hydrogen ion. Generally, the encoded acidic amino acids include L-Glu (E) and L-Asp (D).

[0023] "Alkyl" is intended to include an alkyl group having a specified length and having either a straight-chain or branched-chain configuration. Exemplary alkyl groups are methyl, ethyl, propyl, isopropyl, butyl, sec-butyl, tert-butyl, pentyl, isopentyl, hexyl, isohexyl, etc. The alkyl group is unsubstituted or substituted with 1 to 3 groups independently selected from halogen, hydroxy, carboxy, aminocarbonyl, amino, C l-4 alkoxy, and C l-4 alkylthio.

[0024] "Amide" is intended to mean a functional group containing a carbonyl group linked to a nitrogen atom. Amide also refers to any compound containing an amide functional group. Amides are derived from carboxylic acids and amines.

[0025] "Amidate" or "amidation" is generally intended to mean the formation of an amide functional group from a carbonyl-containing compound that results from the reaction of a carboxylic acid and an amine functional group and is also formed here from an ester and an amine functional group.

[0026] "Amine" is intended to mean a functional group containing an sp3 hybridized nitrogen atom. Amine also means any compound containing an amine functional group.

[0027] "Amine substrate" refers to an amino compound that can replace the alcohol side chain of an ester substrate under the action of carboxylesterase, thereby producing an amide product. The amine substrate is a molecule of general formula (5), wherein each of R3 and R4, when independent, is an alkyl or an aryl group that is unsubstituted or substituted with at least one enzymatically non-inhibiting group. When the groups R3 and R4 are taken together, they may form a ring that is unsubstituted, substituted, or fused to another ring. Typical amine substrates that can be used with the present invention include, but are not limited to, cyclic piperazinyl or morpholino moieties, as well as primary amines or aromatic amines. With respect to the present disclosure, the amine substrate includes, inter alia, 1-isopropylpiperazine, which is a compound of formula II, and cis-2,6-dimethylmorpholine, which is a compound of formula IV.

Chemical formula

[0028] "Amino acid" or "residue", when used with respect to the polypeptides disclosed herein, refers to a particular monomer at a certain sequence position (e.g., P5 indicates that the "amino acid" or "residue" at position 5 of SEQ ID NO: 2 is proline).

[0029] "Amino acid difference" or "residue difference" refers to a change in the residue at a specified position of a polypeptide sequence when compared to a reference sequence. The position of a polypeptide sequence where a particular amino acid or amino acid change ("residue difference") is present may be described herein as "Xn" or "position n", where n refers to the residue position relative to the reference sequence.

[0030] For example, the residue difference at the position of X8 where the reference array has serine refers to the change of the residue at the position of X8 to any residue other than serine. As disclosed herein, the enzyme can include one or more residue differences relative to the reference array, in which case the multiple residue differences are generally indicated by a list of the specified positions that have changed relative to the reference array (e.g., "one or more residue differences when compared to SEQ ID NO: 4 at the following residue positions: X27, X30, X35, X37, X57, X75, X103, X185, X207, X208, X271, X286, or X296").

[0031] A specific substitution mutation that is the substitution of a specific residue of the reference array with a specified different residue can be represented by the conventional notation "X(number)Y", where X is the one-letter identifier of the residue of the reference array, "number" is the residue position in the reference array, and Y is the one-letter identifier of the residue substitution in the engineered array.

[0032] "Aliphatic amino acid or residue" refers to a hydrophobic amino acid or residue having an aliphatic hydrocarbon side chain. Genetically encoded aliphatic amino acids include L-Ala (A), L-Val (V), L-Leu (L), and L-Ile (I).

[0033] "Aromatic amino acid or residue" refers to a hydrophilic or hydrophobic amino acid or residue having a side chain containing at least one aromatic or heteroaromatic ring. Genetically encoded aromatic amino acids include L-Phe (F), L-Tyr (Y), and L-Trp (W). Due to the pKa of its heteroaromatic nitrogen atom, L-His (H) may be classified as a basic residue or as an aromatic residue since its side chain contains a heteroaromatic ring, but herein, histidine is classified as a hydrophilic residue or as a "constrained residue" (see below).

[0034] "Aryl" is intended to mean an aromatic group including phenyl and naphthyl. "Aryl" is unsubstituted or substituted with fluoro, hydroxy, trifluoromethyl, amino, Cl-4 alkyl, and C l-4 It is substituted with 1 to 5 substituents independently selected from alkoxy.

[0035] The term "basic amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain with a pKa value greater than about 6 when the amino acid is included in a peptide or polypeptide. Basic amino acids generally have a side chain that is positively charged at physiological pH for association with hydronium ions. Genetically encoded basic amino acids include L-Arg (R) and L-Lys (K).

[0036] "Carboxylesterase" is used to refer to a polypeptide having the enzymatic ability to interconvert the side chains of an ester substrate (1) and a donor compound (2), and to convert the ester substrate (1) into its corresponding ester product (3) and a free alcohol-type ester side chain (4).

Chemical formula

[0037] "Codon-optimized" refers to changing the codons of a polynucleotide encoding a protein to those that are preferentially used in the target organism so that the protein encoded is efficiently expressed in that particular organism. The genetic code is degenerate in that most amino acids are represented by several codons called "synonymous" codons, but the codon usage frequency by a particular organism is not random and is known to be biased towards certain codon triplets. This codon usage bias is seen to be higher for a given gene, genes of common function or ancestral origin, high-expression proteins versus low-copy-number proteins, and the clustered protein-coding regions of an organism's genome. In some embodiments, the polynucleotide encoding the carboxylesterase enzyme can be codon-optimized for optimal production from the host organism selected for expression.

[0038] A "comparison window" refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acid residues, wherein the sequence can be compared to a reference sequence of at least 20 contiguous nucleotides or amino acids, and the positions of the sequence within the comparison window may include additions or deletions (i.e., gaps) of 20 percent or less as compared to the reference sequence (not including additions or deletions) for optimal alignment of the two sequences. The comparison window may be a contiguous residue longer than 20, and may optionally include windows of 30, 40, 50, 100, or longer.

[0039] A "conservative" amino acid substitution or mutation refers to the interchangeability of residues having similar side chains, and thus generally includes substitutions of an amino acid within the same or a similar, defined class of amino acids of the polypeptide's amino acids. However, as used herein, in some embodiments, a conservative mutation may be an aliphatic residue to aliphatic residue, nonpolar residue to nonpolar residue, polar residue to polar residue, acidic residue to acidic residue, basic residue to basic residue, aromatic residue to aromatic residue, or constrained residue to constrained residue substitution, provided that the conservative mutation does not include substitution of a hydrophilic residue to a hydrophilic residue, a hydrophobic residue to a hydrophobic residue, a hydroxyl-containing residue to a hydroxyl-containing residue, or a small residue to a small residue when it could otherwise be such. Further, as used herein, A, V, L, or I may also be conservatively mutated to another aliphatic residue or another nonpolar residue. Table 2 below shows examples of conservative substitutions.

[0040] [Table 2]

[0041] "Constrained amino acid or residue" refers to an amino acid or residue having a constrained geometry. As used herein, the constrained residues include L-Pro (P) and L-His (H). Histidine has a constrained geometry because it has a relatively small imidazole ring. Proline also has a constrained geometry because it has a five-membered ring.

[0042] As used herein, "control sequence" is defined to include all components necessary or advantageous for the expression of a polynucleotide and / or polypeptide of the present disclosure. Each control sequence may be native or foreign to the polynucleotide of interest. Such control sequences include, but are not limited to, leader, polyadenylation sequence, propeptide sequence, promoter, signal peptide sequence, and transcription terminator.

[0043] "Conversion" refers to the enzymatic conversion of a substrate to the corresponding product.

[0044] "Corresponding to", "referring to", or "relative to", when used with respect to the numbering of a given amino acid sequence or polynucleotide sequence, refers to the residue number of the explicitly stated reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue number or residue position of a given polymer is indicated with respect to the reference sequence, rather than by the actual number position of the residue within the given amino acid or polynucleotide sequence which is not the reference sequence. For example, a given amino acid sequence, such as the amino acid sequence of an engineered carboxylesterase, can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, gaps are present, but the residue numbers within the given amino acid or polynucleotide sequence are assigned with respect to the reference sequence to which it is aligned.

[0045] "Cysteine" or L-Cys (C) is unique in that it can form disulfide bridges with other L-Cys (C) amino acids or other sulfanyl-containing or sulfhydryl-containing amino acids. "Cysteine-like residues" include cysteine and other amino acids containing a sulfhydryl moiety available for the formation of disulfide bridges. The ability of L-Cys (C) (and other amino acids with -SH-containing side chains) to exist within a peptide either in its reduced free -SH or oxidized disulfide-bridged form affects whether L-Cys (C) confers net hydrophobicity or hydrophilicity to the peptide. L-Cys (C) exhibits a hydrophobicity of 0.29 according to the normalized consensus scale of Eisenberg et al. (1984, supra), but for the purposes of the present disclosure, L-Cys (C) should be understood to be classified in its own group.

[0046] "Deletion" refers to the modification of a polypeptide by removing one or more amino acids from a reference polypeptide. Deletions can include the removal of one or more amino acids, two or more amino acids, five or more amino acids, ten or more amino acids, fifteen or more amino acids, or twenty or more amino acids, up to 10% of the total number of amino acids, up to 20% of the total number of amino acids, or up to 30% of the total number of amino acids of the polypeptide, while retaining enzyme activity and / or retaining the improved properties of the engineered carboxylesterase enzyme. Deletions can be to internal and / or terminal portions of the polypeptide. In various embodiments, deletions may comprise contiguous segments or may be discontinuous.

[0047] "Derived from" when used herein with respect to an engineered enzyme identifies the originating enzyme and / or the gene encoding the enzyme on which the engineering is based. For example, the engineered carboxylesterase enzyme of SEQ ID NO: 4 was obtained by mutagenesis of the carboxylesterase of SEQ ID NO: 2. Thus, this engineered carboxylesterase enzyme of SEQ ID NO: 4 "is derived from" the polypeptide of SEQ ID NO: 2.

[0048] As used herein, "engineered carboxylesterase" refers to a carboxylesterase-type protein that has been systematically modified by insertion of a new amino acid into its reference sequence, deletion of an amino acid present in its reference sequence, or mutagenesis by random mutagenesis followed by selection of mutants having certain characteristics, or by intentional introduction of a specific amino acid into the protein sequence, through mutation of an amino acid in its reference sequence to another amino acid.

[0049] "Ester" is intended to mean a functional group containing a carbonyl group linked to an oxygen atom and then to a carbon atom. Ester also refers to any compound containing an ester functional group. Esters are derived from carboxylic acids and alcohols.

[0050] "Ester substrate" specifically refers to a compound of formula (1) containing an ester that reacts with an engineered carboxylesterase. With respect to the present disclosure, examples of ester substrates include, among others, ethyl oxazole-5-carboxylate, which is a compound of formula I.

Chemical formula

[0051] As used herein, "fragment" refers to a polypeptide having an amino-terminal deletion and / or a carboxy-terminal deletion, but the remaining amino acid sequence is identical to the corresponding position of that sequence. Fragments can be at least 14 amino acids in length, at least 20 amino acids in length, at least 50 or more amino acids in length, and can be up to 70%, 80%, 90%, 95%, 98%, and 99%, or more, of the full-length carboxylesterase polypeptide, for example, the polypeptide of SEQ ID NO: 4.

[0052] As used interchangeably herein, "functional fragment" or "biologically active fragment" refers to a polypeptide that has amino-terminal and / or carboxy-terminal and / or internal deletions, but the remaining amino acid sequence is identical to the corresponding position of the sequence to which it is compared (e.g., the engineered full-length T. fusca enzyme of the present invention), and retains substantially all of the activities of the full-length polypeptide.

[0053] "Halogen" is intended to include fluorine, chlorine, bromine, and iodine, which are halogen atoms.

[0054] "Heterologous" polynucleotide refers to any polynucleotide introduced into a host cell by laboratory techniques, including polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into the host cell.

[0055] "Hybridization stringency" relates to hybridization conditions such as washing conditions in nucleic acid hybridization. Generally, hybridization reactions are carried out under lower stringency conditions, followed by washing under various, but higher, stringency conditions. "Moderately stringent hybridization" refers to conditions under which a target DNA can bind to a complementary nucleic acid having about 60% identity, preferably about 75% identity, about 85% identity with the target DNA, or having more than about 90% identity with the target polynucleotide. Example moderately stringent conditions correspond to hybridization in 50% formamide, 5× Denhardt's solution, 5× saline-sodium phosphate-EDTA (SSPE), 0.2% sodium dodecyl sulfate (SDS) at 42°C, followed by washing in 0.2× SSPE, 0.2% SDS at 42°C. "High stringency hybridization" generally refers to the thermal melting point T determined under solution conditions for a defined polynucleotide sequence. mRefers to conditions that are approximately 10 °C lower. In some embodiments, high stringency conditions refer to conditions that allow hybridization only to nucleic acid sequences that form stable hybrids at 65 °C in 0.018 M NaCl (i.e., if the hybrid is not stable at 65 °C in 0.018 M NaCl, it is not stable under high stringency conditions as contemplated herein). High stringency conditions can be provided, for example, by hybridization at 42 °C in 50% formamide, 5× Denhardt's solution, 5× SSPE, 0.2% SDS and subsequent washing at 65 °C in 0.1× SSPE and 0.1% SDS. Another high stringency condition is hybridization at a condition corresponding to 65 °C in 5× SSC containing 0.1% (w:v) SDS and washing at 65 °C in 0.1× SSC containing 0.1% SDS. Other high stringency hybridization conditions, as well as moderately stringent conditions, are described in the references cited above.

[0056] "Hydrophilic amino acid or residue" refers to an amino acid or residue having a side chain that exhibits a hydrophobicity less than 0 according to the normalized consensus hydrophobicity scale of Eisenberg et al., 1984, J. Mol. Biol. 179:125-142. Genetically encoded hydrophilic amino acids include L-Thr (T), L-Ser (S), L-His (H), L-Glu (E), L-Asn (N), L-Gln (Q), L-Asp (D), L-Lys (K) and L-Arg (R). "Hydrophobic amino acid or residue" refers to an amino acid or residue having a side chain that exhibits a hydrophobicity greater than 0 according to the normalized consensus hydrophobicity scale of Eisenberg et al., 1984, J. Mol. Biol. 179:125-142. Genetically encoded hydrophobic amino acids include L-Pro (P), L-Ile (I), L-Phe (F), L-Val (V), L-Leu (L), L-Trp (W), L-Met (M), L-Ala (A) and L-Tyr (Y).

[0057] "Hydroxyl-containing amino acid or residue" refers to an amino acid containing a hydroxyl (-OH) moiety. Genetically encoded hydroxyl-containing amino acids include L-Ser (S), L-Thr (T), and L-Tyr (Y).

[0058] "Improved enzyme properties" refers to any enzyme property that is better or more desirable for a particular purpose as compared to the properties found in a reference enzyme. For the engineered carboxylesterase polypeptides described herein, comparison is generally made to the wild-type carboxylesterase enzyme, although in some embodiments, the reference carboxylesterase can be another improved, engineered carboxylesterase. Enzyme properties that can be improved include, but are not limited to, enzyme activity (which can be expressed as the percentage of substrate conversion over a defined period of time), thermal stability, solvent stability, pH activity profile, cofactor requirement, resistance to inhibitors (e.g., product inhibition), stereospecificity, and suppression of acid byproduct production.

[0059] "Increase in enzyme activity" or "increase in activity" refers to an improved property of an engineered enzyme that can be represented by an increase in specific activity (e.g., amount of product produced / time / protein weight) or an increase in the percentage of substrate conversion to product (e.g., the percentage of conversion of a starting amount of substrate to product over a defined period of time using a specified amount of carboxylesterase) when compared to a reference enzyme. Exemplary methods for measuring enzyme activity are shown in the Examples. The K that can lead to an increase in enzyme activity m , V max or k catIt can affect any property related to enzyme activity, including the classical enzyme properties. The improvement of enzyme activity can range from about 1.5-fold of the enzyme activity of the corresponding wild-type or engineered enzyme to 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, 1000-fold, 10,000-fold, 100,000-fold, 500,000-fold, 785,000-fold or more of the enzyme activity of a natural enzyme (e.g., carboxylesterase) or another engineered enzyme from which an enzyme showing increased activity is derived. In certain embodiments, the engineered carboxylesterase enzymes of the present disclosure exhibit improved enzyme activity in the range of 1.5 to 50-fold, 1.5 to 100-fold or more than that of the parental carboxylesterase enzyme (i.e., the wild-type or engineered carboxylesterase from which they are derived). One skilled in the art will understand that the activity of any enzyme is diffusion-limited so that it cannot exceed the diffusion rate of the substrate, including any coenzymes required for the catalytic turnover rate. The theoretical maximum value of the diffusion limit is generally about 10 8 ~10 9 (M -1 s -1 ). Therefore, any improvement in the enzyme activity of carboxylesterase has an upper limit with respect to the diffusion rate of the substrate on which the carboxylesterase enzyme acts. Carboxylesterase activity can be measured by any of the standard assays used to measure carboxylesterase, such as changes in substrate or product concentration, or changes in amine substrate concentration. Comparison of enzyme activities is performed using defined enzyme production methods, defined assays under certain set conditions, and one or more defined substrates, as described in more detail herein. Generally, when enzymes in cell lysates are compared, the same expression system and the same host cells are used, and the number of cells and the amount of protein assayed are determined to minimize fluctuations in the amount of enzyme produced by the host cell and present in the lysate.

[0060] "Insertion" refers to the modification of a polypeptide by the addition of one or more amino acids to a reference polypeptide. In some embodiments, an improved, engineered carboxylesterase enzyme comprises the insertion of one or more amino acids into a native carboxylesterase polypeptide as well as the insertion of one or more amino acids into another improved carboxylesterase polypeptide. The insertion may be in an internal portion of the polypeptide or may be to the carboxy terminus or amino terminus. An insertion, as used herein, includes fusion proteins as are known in the art. The insertion may be a continuous segment of amino acids or may be separated by one or more of the amino acids of the native polypeptide.

[0061] "Isolated polypeptide" refers to a polypeptide substantially separated from other contaminants that naturally accompany it, such as proteins, lipids, and polynucleotides. This term encompasses polypeptides removed or purified from their natural environment or expression system (e.g., host cell or in vitro synthesis). An improved carboxylesterase enzyme can be present intracellularly, can be present in a cell medium, or can be prepared in various forms such as a lysate or an isolated preparation. Thus, in some embodiments, an improved carboxylesterase enzyme can be an isolated polypeptide.

[0062] "Non-conservative substitution" refers to the substitution or mutation of an amino acid within a polypeptide with an amino acid having significantly different side chain properties. Non-conservative substitutions can use amino acids not within the defined groups listed above, but between groups. In one embodiment, non-conservative mutations affect (a) the substitution area of the structure of the peptide backbone (e.g., proline instead of glycine); (b) charge or hydrophobicity; or (c) the bulk of the side chain.

[0063] "Non-polar amino acid" or "non-polar residue" refers to a hydrophobic amino acid or residue having a side chain with a bond in which there is no charge at physiological pH and the electron pair shared by two atoms is generally held equally by each of the two atoms (i.e., the side chain is non-polar). Genetically encoded non-polar amino acids include L-Gly (G), L-Leu (L), L-Val (V), L-Ile (I), L-Met (M), and L-Ala (A).

[0064] "Conversion rate %" refers to the percentage of substrate converted to product within a period under the specified conditions. Thus, for example, "enzyme activity" or the "activity" of a carboxylesterase polypeptide can be expressed as the "conversion rate %" from substrate to product.

[0065] "Percent sequence identity", "percent identity", and "percent identical to" are used herein to refer to a comparison between polynucleotide sequences or polypeptide sequences, and are determined by comparing two optimally aligned sequences in a comparison window, where a portion of the polynucleotide sequence or polypeptide sequence in the comparison window may include additions or deletions (i.e., gaps) as compared to a reference sequence for optimal alignment of the two sequences. The percentage is determined by counting the number of positions at which the same nucleic acid base or amino acid residue is found in both sequences, or the nucleic acid bases or amino acid residues are aligned using gaps to determine the number of positions that match, and the number of matching positions is divided by the total number of positions in the comparison window and multiplied by 100 to obtain the percentage of sequence identity. Determination of optimal alignment and percent sequence identity is performed using the BLAST and BLAST 2.0 algorithms (see, e.g., Altschul, et al., 1990, J. Mol. Biol. 215: 403-410 and Altschul, et al., 1977, Nucleic Acids Res. 3389-3402). Software for performing BLAST analyses is publicly available at the website of the National Center for Biotechnology Information.

[0066] Briefly stated, BLAST analysis involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which are either identical or meet a positive threshold score T when aligned with words of the same length in the database sequences. T is referred to as the neighborhood word score threshold (Altschul, et al, supra). These initial neighborhood word hits serve as seeds to initiate a search to find longer HSPs that contain them. The word hits are extended in both directions along the sequences as long as the cumulative alignment score can be increased. For nucleotide sequences, the cumulative score is calculated using parameters M (reward score for pairs of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. The extension of the word hits in each direction is terminated when the cumulative alignment score decreases by an amount X from its maximum achieved value; when the cumulative score becomes zero or less due to the accumulation of one or more negative-score residue alignments; or when the end of either sequence is reached. The parameters W, T, and X of the BLAST algorithm determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses a word length (W) of 11, an expectation value (E) of 10, M = 5, N = -4, and comparison of both strands as default. For amino acid sequences, the BLASTP program uses a word length (W) of 3, and an expectation value (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff (1989) Proc. Natl. Acad. Sci. USA 89:10915) as default.

[0067] Many other algorithms are available that function similarly to BLAST in obtaining the percent identity of two arrays. The optimal alignment of sequences can be achieved, for example, by the local homology algorithm of Smith and Waterman, 1981, Adv. Appl. Math. 2:482, by the homology alignment algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48:443, by the similarity search method of Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444, by computer implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin software package), or by visual inspection (see generally Current Protocols in Molecular Biology, F. M. Ausubel, et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)). In addition, the BESTFIT or GAP programs of the GCG Wisconsin software package (Accelrys, Madison WI) can be used with the provided default parameters for sequence alignment and determination of percent sequence identity.

[0068] "pH stable" refers to a polypeptide that maintains equivalent activity (e.g., greater than 60% - 80%) compared to the untreated enzyme after exposure to low or high pH (e.g., 4.5 - 6 or 8 - 12) for a defined period (e.g., 0.5 - 24 hours).

[0069] "Polar amino acid or residue" refers to a hydrophilic amino acid or residue having a side chain that has no charge at physiological pH but has at least one bond in which a pair of electrons shared by two atoms is held more closely by one of those atoms. Genetically encoded polar amino acids include L-Asn (N), L-Gln (Q), L-Ser (S), and L-Thr (T).

[0070] "Preferred, optimal, high codon usage bias codons" are used interchangeably to refer to codons that are used at a higher frequency in protein-coding regions other than codons that encode the same amino acid. Preferred codons can be determined with respect to codon usage frequency in a single gene, codon usage frequency in a set of genes with a common function or origin, codon usage frequency in highly expressed genes, codon frequency in the aggregated protein-coding regions of an entire organism, codon frequency in the aggregated protein-coding regions of related organisms, or combinations thereof. Codons whose frequency increases with the level of gene expression are generally optimal codons for expression. To determine codon frequencies (e.g., codon usage frequency, relative synonymous codon usage frequency) and codon selection in a particular organism, various methods are known, including, for example, cluster analysis or correspondence analysis, and multivariate analysis using the effective number of codons used in a gene (see GCG CodonPreference, Genetics Computer Group Wisconsin Package; CodonW, John Peden, University of Nottingham; McInerney, J. O, 1998, Bioinformatics 14:372-73; Stenico, et al., 1994, Nucleic Acids Res. 222437-46; Wright, F., 1990, Gene 87:23-29). Codon usage tables are becoming available for an increasing number of organisms (see, for example, Wada, et al., 1992, Nucleic Acids Res. 20:2111-2118; Nakamura, et al., 2000, Nucl. Acids Res. 28:292; Duret, et al., supra; Henaut and Danchin, “Escherichia coli and Salmonella,” 1996, Neidhardt, et al., eds., ASM Press, Washington D.C., p. 2047-2066).The data source for obtaining codon usage frequencies may rely on any available nucleotide sequence that can encode a protein. These datasets include nucleic acid sequences that are actually known to encode expressed proteins (e.g., full protein coding sequences - CDS), expressed sequence tags (ESTs), or predicted coding regions of genomic sequences (see, e.g., Mount, D., Bioinformatics: Sequence and Genome Analysis, Chapter 8, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 2001; Uberbacher, E. C., 1996, Methods Enzymol. 266:259-281; Tiwari, et al.., 1997, Comput. Appl. Biosci. 13:263-270).

[0071] "Protein", "polypeptide", and "peptide" are used interchangeably herein to refer to a polymer of at least two amino acids covalently linked by amide bonds, regardless of length or post-translational modification (e.g., glycosylation, phosphorylation, lipidation, myristoylation, ubiquitination, etc.). This definition includes D-amino acids and L-amino acids, as well as mixtures of D-amino acids and L-amino acids.

[0072] The term "reference sequence" refers to a defined sequence to which another (e.g., altered) sequence is compared. The reference sequence may be a subset of a larger sequence, e.g., a segment of a full-length gene sequence or a full-length polypeptide sequence. Generally, the reference sequence is at least 20 nucleotides or amino acid residues in length, at least 25 residues in length, at least 50 residues in length, or the full length of the nucleic acid or polypeptide. Since two polynucleotides or polypeptides each comprise a similar sequence (i.e., part of the complete sequence) between the two sequences and (2) may further comprise sequences that are different from each other between the two sequences, sequence comparison between two (or more) polynucleotides or polypeptides is generally performed by comparing the sequences of the two polynucleotides in a comparison window to identify and compare local regions of sequence similarity.

[0073] The term "reference sequence" is not intended to be limited to a wild-type sequence and may include engineered or altered sequences. For example, in some embodiments, the "reference sequence" may be an amino acid sequence that has been previously engineered or altered. For example, a "reference sequence based on SEQ ID NO: 2 having a glycine residue at position X12" refers to a reference sequence corresponding to SEQ ID NO: 2 having a glycine residue at X12 (the unaltered form of SEQ ID NO: 2 has an aspartate at X12).

[0074] "Small amino acid" or "small residue" refers to an amino acid or residue having a side chain composed of a total of 3 or fewer carbon and / or heteroatoms (excluding the α-carbon and hydrogen). Small amino acids or residues can be further classified as aliphatic, nonpolar, polar, or acidic small amino acids or residues according to the above definition. Genetically encoded small amino acids include L-Ala (A), L-Val (V), L-Cys (C), L-Asn (N), L-Ser (S), L-Thr (T), and L-Asp (D).

[0075] "Solvent stability" or "solvent stability" refers to a polypeptide that maintains equivalent (e.g., greater than 60% - 80%) activity after exposure to solvents (e.g., isopropyl alcohol, dimethyl sulfoxide, tetrahydrofuran, 2-methyltetrahydrofuran, acetone, toluene, butyl acetate, methyl tert-butyl ether, acetonitrile, etc.) at various concentrations (e.g., 5 - 99%) for a certain period (e.g., 0.5 - 24 hours) compared to the untreated enzyme.

[0076] "Substantially identical" refers to a polynucleotide sequence or polypeptide sequence that has at least 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 91, 93, 94, 95, 96, 97, 98, 99 or more percent sequence identity when compared to a reference sequence in a comparison window of at least 20 residues, often in a window of at least 30 - 50 residues. The percentage of sequence identity is calculated by comparing the reference sequence with a sequence that contains no more than 20 percent deletion or addition in total in the comparison window. In a particular embodiment applied to polypeptides, the term "substantially identical" means that two polypeptide sequences, when optimally aligned, have at least 80 percent sequence identity, preferably at least 89 percent sequence identity, at least 95 percent sequence identity or more (e.g., 99 percent sequence identity), such as by using the GAP or BESTFIT program with a default gap weight. Preferably, the non-identical residue positions differ by conservative amino acid substitutions.

[0077] "Substantially pure polypeptide" refers to a composition in which the polypeptide species is the predominant species present (i.e., it is more abundant in the composition than any other individual macromolecular species, based on moles or weight), and generally, a substantially purified composition where the target species occupies at least about 50 percent of the macromolecular species present, in mole or weight percent. Generally, a substantially pure carboxylesterase composition occupies about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, or about 99% or more of all the macromolecular species present in the composition, in mole or weight percent. In some embodiments, the target species is purified to essentially homogeneous (i.e., contaminating species cannot be detected in the composition by conventional detection methods), and this composition consists essentially of a single macromolecular species. Solvent species and elemental ion species are not considered macromolecular species. In some embodiments, the isolated, improved carboxylesterase polypeptide is a substantially pure polypeptide composition.

[0078] "Substrate", as used herein, refers to a carboxylesterase-reactive compound consisting of a compound containing an ester (1), an amine (5), or an alcohol (2). With respect to the present disclosure, substrates for carboxylesterase include, inter alia, compounds of Formula I and compounds of Formula II as further described herein.

[0079] "Thermostability" or "heat stability" is used interchangeably to refer to a polypeptide that is resistant to inactivation when exposed to a set of temperature conditions (e.g., 40 - 80 °C) for a certain period of time (e.g., 0.5 - 24 hours), compared to the untreated enzyme, and thus retains a certain level of residual activity (e.g., exceeding 60% - 80%) after exposure to high temperatures.

[0080] As used herein, "solvent-stable" refers to the ability of a polypeptide to maintain equivalent activity (e.g., greater than 60% - 80%) after exposure to various concentrations (e.g., 5 - 99%) of solvents (e.g., isopropyl alcohol, tetrahydrofuran, 2-methyltetrahydrofuran, acetone, toluene, butyl acetate, methyl tert-butyl ether, etc.) for a certain period (e.g., 0.5 - 24 hours) compared to the untreated enzyme.

[0081] Detailed Description of Embodiments of the Invention In an embodiment herein, the engineered carboxylesterase has an improved ability to convert the ester substrate ethyl oxazole-5-carboxylate to the product (4-isopropylpiperazin-1-yl)(oxazol-5-yl)methanone in the presence of the amine substrate 1-isopropylpiperazine compared to the wild-type carboxylesterase enzyme A. acidocaldarius esterase 2 (SEQ ID NO: 2). The carboxylesterases, including those described herein, are self-standing enzymes lacking chemical cofactors and are water-soluble enzymes that can be formulated as a lytic enzyme, as an enzyme immobilized on a resin, or as a lyophilized powder in the presence of one or more salts.

[0082] In some embodiments, the improvement in enzyme activity relates to another engineered carboxylesterase, such as the polypeptide of SEQ ID NO: 4. The improvement in activity towards an ester substrate can be represented by an increase (e.g., conversion rate %) in the amount of substrate converted to product by the engineered enzyme compared to a reference enzyme (e.g., wild type, SEQ ID NO: 2) under defined conditions. The improvement in activity can include an increase in the rate of product formation that results in an increase in the conversion of an ester substrate to product in the presence of an amine substrate under defined conditions for a defined period. The increase in activity (e.g., increase in conversion rate % and / or conversion rate) can also be characterized by converting the same amount of substrate to the same amount of product with a lesser amount of enzyme. The amount of product can be evaluated by various techniques, such as separation of the reaction mixture (e.g., by chromatography) and detection of the separated product by UV absorbance or tandem mass spectrometry (MS / MS) (see, e.g., Example 4). An example of defined reaction conditions for comparison with the activity of SEQ ID NO: 2 is about 40 g / L ethyl oxazole-5-carboxylate (ester substrate), 44 g / L 1-isopropylpiperazine (amine substrate), and 20 g / L of a carboxylesterase polypeptide corresponding to an amino acid sequence selected from SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 92, or 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126, which enzyme is produced in the presence of sodium sulfate and acts in the presence of about 10 g / L to about 20 g / L water in methyl isobutyl ketone (MIBK) as described hereinafter in the description of the reaction conditions for the carboxylesterases listed in Table 3. The defined reaction conditions for comparison with a particular engineered carboxylesterase are also shown in the description of the carboxylesterases listed in Table 3 and the corresponding description of Example 7.In some embodiments, the engineered carboxylesterase has at least 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 30-fold, 40-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold, 300-fold, 400-fold, 500-fold, 1000-fold, 1,500-fold, 2,000-fold, 10,000-fold, 100,000-fold, 500,000-fold, 785,000-fold, or more than the activity of the polypeptide of SEQ ID NO: 2 under defined reaction conditions.

[0083] In some embodiments, the improved enzyme activity is also associated with other improvements in enzyme properties. In some embodiments, the improvement in enzyme properties relates to thermal stability, such as at 60 °C or higher.

[0084] In some embodiments, the improved enzyme activity is associated with an improvement in solvent stability, such as when carried out at 98% (v / v) in methyl isobutyl ketone (MIBK) or tert-butyl methyl ether (TBME).

[0085] In some embodiments, the engineered carboxylesterase polypeptide of the present disclosure has an activity greater than the activity of the polypeptide of SEQ ID NO: 2 in the presence of an amine substrate, such as 1-isopropylpiperazine, and can convert the ester substrate ethyl oxazole-5-carboxylate to the product (4-isopropylpiperazin-1-yl)(oxazol-5-yl)methanone, and comprises an amino acid sequence that is at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical, or more than identical, to the reference sequence of SEQ ID NO: 2, or a functional fragment thereof.

[0086] In some embodiments, the engineered carboxylesterase polypeptide of the present disclosure has greater activity than the polypeptide of SEQ ID NO: 2 in the presence of an amine substrate such as 1-isopropylpiperazine and can convert the ester substrate ethyl oxazole-5-carboxylate to the product (4-isopropylpiperazin-1-yl)(oxazol-5-yl)methanone, and comprises an amino acid sequence that is at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical or more identical to a reference sequence listed in Table 3, such as SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126, or a functional fragment thereof.

[0087] In some embodiments, the engineered carboxylesterase polypeptide comprises an amino acid sequence having one or more residue differences when compared to a carboxylesterase reference sequence. The residue differences can be non-conservative substitutions, conservative substitutions, or a combination of non-conservative and conservative substitutions. With respect to the description of residue differences and residue positions, the carboxylesterases provided herein can be described with reference to the amino acid sequence of a native carboxylesterase such as A. acidocaldarius esterase 2 (SEQ ID NO: 2), the carboxylesterase of SEQ ID NO: 2, or an engineered carboxylesterase polypeptide such as the polypeptide of SEQ ID NO: 4. For the description herein, the position of an amino acid residue within a reference sequence is determined beginning with the start methionine (M) residue in the carboxylesterase (i.e., M indicates residue position 1), but it will be understood by those skilled in the art that this start methionine residue can be removed by biological processing mechanisms, such as in a host cell or an in vitro translation system, to produce a mature protein lacking the start methionine residue.

[0088] In some embodiments, the residue differences can be found in one or more of the following residue positions: X2, X7, X9, X10, X19, X20, X22, X27, X28, X29, X30, X33, X34, X35, X36, X37, X38, X46, X48, X54, X57, X66, X75, X85, X86, X87, X96, X103, X139, X160, X176, X181, X183, X185, X188, X190, X197, X198, X205, X207, X208, X212, X216, X248, X249, X255, X263, X266, X270, X271, X278, X280, X286, X290 and X296. In some embodiments, the residue differences or combinations thereof are related to the improvement of enzyme properties. In some embodiments, the carboxylesterase polypeptide may further have 1 to 2, 1 to 3, 1 to 4, 1 to 5, 1 to 6, 1 to 7, 1 to 8, 1 to 9, 1 to 10, 1 to 11, 1 to 12, 1 to 13, 1 to 14, 1 to 15, 1 to 16, 1 to 17, 1 to 18, 1 to 19, 1 to 20, 1 to 22, 1 to 24, 1 to 26, 1 to 28, 1 to 30, 1 to 35, 1 to 40, 1 to 45, 1 to 50, 1 to 55, 1 to 60, or 1 to 62 residue differences at residue positions other than the specific positions represented by the above "Xn". In some embodiments, the number of differences is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 35, 40, 45, 50, 55, 60, or 62 residue differences at other amino acid residue positions. In some embodiments, the residue differences at other residue positions comprise substitutions by conservative amino acid residues.

[0089] In an embodiment of the present specification, the difference in residues at the residue positions that affect substrate binding to carboxylesterase when compared with SEQ ID NO: 2 enables the application of the ester substrate of structural formula (I) further described below, in particular, the ethyl ester substrate oxazole-5-carboxylic acid. Without being bound by theory, at least two regions, namely, a first substrate binding region and a second substrate binding region, interact with different structural elements of the ester substrate. The first binding region comprises residues X85, X185, X214, X215 and X254, and the second binding region comprises residue positions X30, X33, X34, X37, X82, X205, X210, X283, X286 and X287, but the positions of X83, X84, X155, X156, X206, X214 and X282 overlap at these two sites. These residues were determined by X-ray crystallography. Thus, the carboxylesterase polypeptide herein has a difference in one or more residues at the residue positions comprising X30, X33, X34, X37, X85, X185, X205, and X286. In some embodiments, the carboxylesterase polypeptide herein has a difference in at least two or more, or three or more residues at the specified residue positions related to substrate binding.

[0090] In other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence that is at least 80% identical to the amino acid sequence set forth in SEQ ID NO:4, and the improved carboxylesterase amino acid sequence is characterized in that the residue corresponding to X198 is selected from nonpolar residues, aromatic residues, and aliphatic residues. In yet other embodiments, the improved carboxylesterase polypeptide comprises the feature that X198 is selected from F, L, I, Y, and M. In some embodiments, the improved carboxylesterase polypeptide may comprise an amino acid sequence that, when compared to the sequence of SEQ ID NO:4, includes differences of one or more residues at residue positions corresponding to X27, X30, X35, X37, X57, X66, X75, X103, X207, X208, X271, X286, and X296. Guidelines regarding the selection of the various amino acid residues that may be present at the indicated residue positions are set forth in the following detailed description.

[0091] In other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence that is at least 80% identical to the amino acid sequence set forth in SEQ ID NO:4, and this amino acid sequence is such that the residue corresponding to X27 is a restricted residue; the residue corresponding to X30 is an aliphatic residue; the residue corresponding to X35 is selected from basic residues and polar residues; the residue corresponding to X37 is selected from aliphatic residues and polar residues; the residue corresponding to X57 is a nonpolar residue; the residue corresponding to X75 is selected from basic residues and polar residues; the residue corresponding to X103 is selected from nonpolar and aromatic residues; the residue corresponding to X185 is selected from aliphatic residues, nonpolar residues, and aromatic residues; the residue corresponding to X207 is selected from acidic residues and polar residues; the residue corresponding to X208 is selected from aliphatic residues, basic residues, and polar residues; the residue corresponding to X271 is selected from acidic residues and polar residues; the residue corresponding to X286 is selected from aliphatic residues, polar residues, and small residues; and the residue corresponding to X296 is selected from aliphatic residues and basic residues, and includes at least one feature selected therefrom.

[0092] In yet another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence that is at least 80% identical to the amino acid sequence shown in SEQ ID NO: 4, and this amino acid sequence has the following characteristics: X27 is P; X30 is selected from I, L, and V; X35 is H; X37 is selected from I, L, T, and V; X57 is M; X75 is R; X103 is selected from F, M, and W; X185 is selected from F, I, and M; X207 is E; X208 is selected from R, L, and H; X271 is D; X286 is selected from M, V, and G; and X296 is selected from V, L, and R. In some embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence having the following characteristics: X35 is selected from basic residues and polar residues; and X185 is selected from polar residues and aliphatic residues. In another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence having the following characteristics: X35 is H; and X185 is selected from F, I, and M.

[0093] In some embodiments, the improved carboxylesterase polypeptide comprises a residue difference at at least one residue position selected from X9, X19, X34, X35, X37, X46, X48, X66, X87, X103, X139, X190, X207, X216, X263, X271, X278, and X296 when compared to the amino acid sequence set forth in SEQ ID NO: 24. In other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising at least one of the following characteristics: the residue corresponding to X9 is an aromatic residue; the residue corresponding to X19 is selected from basic and polar residues; the residue corresponding to X34 is selected from constrained, acidic, and polar residues; the residue corresponding to X35 is selected from polar residues; the residue corresponding to X46 is an aliphatic residue; the residue corresponding to X48 is an aliphatic residue; the residue corresponding to X66 is an aliphatic residue; the residue corresponding to X87 is selected from aliphatic and small residues; the residue corresponding to X103 is selected from aromatic residues; the residue corresponding to X139 is a basic residue; the residue corresponding to X190 is an aromatic residue; the residue corresponding to X207 is a basic residue; the residue corresponding to X216 is selected from aromatic, basic, and polar residues; the residue corresponding to X263 is selected from aliphatic and polar residues; the residue corresponding to X271 is selected from acidic and polar residues; the residue corresponding to X278 is selected from aliphatic and aromatic residues; and the residue corresponding to X296 is selected from aliphatic and basic residues.

[0094] In yet other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence that is at least 80% identical to the amino acid sequence set forth in SEQ ID NO: 24, the amino acid sequence comprising at least one characteristic selected from the group consisting of: X9 is Y; X19 is R; X34 is selected from E, N or P; X35 is S; X37 is T; X46 is selected from I, L or V; X48 is L; X66 is V; X87 is A; X103 is selected from W or F; X139 is R; X190 is Y; X207 is E; X216 is selected from N and W; X263 is selected from T and A; X271 is D; X278 is selected from W and L; and X296 is selected from V, L, and R. In yet other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising the following characteristics: X9 is an aromatic residue and X87 is an aliphatic residue. In another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising the following characteristics: X9 is Y and X87 is A.

[0095] In other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising at least one residue position selected from X20, X28, X29, X30, X33, X34, X188, X216, and X286, which comprises a residue different from the amino acid sequence shown in SEQ ID NO: 54. In another embodiment, the improved carboxylesterase polypeptide is characterized in that X20 is selected from aliphatic residues and basic residues; the residue corresponding to X28 is selected from acidic residues, polar residues, and restricted residues; the residue corresponding to X29 is selected from acidic residues and polar residues; the residue corresponding to X30 is an aliphatic residue; the residue corresponding to X33 is an aromatic residue; the residue corresponding to X34 is a small residue; the residue corresponding to X188 is selected from small residues and aromatic residues; the residue corresponding to X216 is a polar residue; and the residue corresponding to X286 is selected from aliphatic residues, small residues, non-polar residues, and polar residues, and comprises an amino acid sequence comprising at least one of these features.

[0096] In some embodiments, the improved carboxylesterase polypeptide comprises the amino acid sequence shown in SEQ ID NO: 54, which comprises at least one specific mutation selected from the following: X20 is selected from I and R; X28 is selected from D, P, and S; X29 is D; X30 is V; X33 is W; X34 is G; X188 is selected from G and F; X216 is N; and X286 is selected from S, M, V, G, and A. In another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence characterized in that X216 is a polar residue. In yet another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence characterized in that X216 is N.

[0097] In other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising a residue at one residue position selected from X10, X20, X22, X28, X30, X33, X36, X37, X46, X66, X75, X103, X197, X263, X266, X280, and X290 that is different from the amino acid sequence set forth in SEQ ID NO: 68. In another embodiment, the improved carboxylesterase polypeptide has the residue corresponding to X10 being an aliphatic residue; the residue corresponding to X20 being selected from an aliphatic residue and a basic residue; the residue corresponding to X22 being an aromatic residue; the residue corresponding to X28 being selected from an acidic residue, a polar residue, and a constrained residue; the residue corresponding to X30 being an aliphatic residue; the residue corresponding to X33 being an aromatic residue; the residue corresponding to X36 being an aliphatic residue or an aromatic residue; the residue corresponding to X37 being an aromatic residue or a small residue; the residue corresponding to X46 being a basic residue; the residue corresponding to X66 being a polar residue; the residue corresponding to X75 being a basic residue; the residue corresponding to X103 being an aromatic residue; the residue corresponding to X197 being an aliphatic residue; the residue corresponding to X263 being a basic residue; the residue corresponding to X266 being a polar residue; the residue corresponding to X280 being selected from an aliphatic residue and a polar residue; and the residue corresponding to X290 being selected from an aliphatic residue and an aromatic residue.

[0098] In still other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence that is at least 80% identical to the amino acid sequence set forth in SEQ ID NO: 68, the amino acid sequence comprising at least one characteristic selected from: X10 is selected from L and M; X20 is selected from I and R; X22 is W; X28 is selected from D, P, and S; X30 is V; X33 is W; X36 is selected from F, I, and M; X37 is selected from G and Y; X46 is R; X66 is T; X75 is R; X103 is W; X197 is L; X263 is R; X266 is T; X280 is selected from M and T; and X290 is selected from W and I. In another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising at least one characteristic selected from: X30 is an aliphatic residue; X33 is an aromatic residue; X75 is a basic residue; and X103 is an aromatic residue. In yet another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising the following characteristics: X30 is V; X33 is W; X75 is R; and X103 is W.

[0099] In other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising a residue at one residue position selected from X28, X38, X46, X54, X66, X75, X85, X86, X96, X160, X176, X183, X188, X205, X212, X248, X249, X255, X270, and X286, which is different from the amino acid sequence shown in SEQ ID NO: 68. In another embodiment, the improved carboxylesterase polypeptide is such that the residue corresponding to X28 is selected from acidic residues, polar residues, small residues, and constrained residues; the residue corresponding to X38 is selected from aliphatic residues and basic residues; the residue corresponding to X46 is selected from acidic residues and basic residues; the residue corresponding to X54 is selected from acidic residues and polar residues; the residue corresponding to X66 is a polar residue; the residue corresponding to X75 is a basic residue; the residue corresponding to X85 is selected from aromatic residues or basic residues and small residues; the residue corresponding to X86 is a polar residue; the residue corresponding to X96 is selected from nonpolar residues and aliphatic residues; the residue corresponding to X160 is selected from polar residues and constrained residues; the residue corresponding to X176 is selected from aliphatic residues, aromatic residues, or basic residues and nonpolar residues; the residue corresponding to X183 is a nonpolar residue; the residue corresponding to X188 is selected from aromatic residues and small residues; the residue corresponding to X205 is an aromatic residue; the residue corresponding to X212 is an acidic residue; the residue corresponding to X248 is an aliphatic residue; the residue corresponding to X249 is an aromatic residue; the residue corresponding to X255 is a polar residue; the residue corresponding to X270 is selected from aliphatic residues and polar residues; and the residue corresponding to X286 is selected from aliphatic residues, nonpolar residues, small residues, and polar residues, comprising an amino acid sequence comprising at least one characteristic selected therefrom.

[0100] In yet another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence that is at least 80% identical to the amino acid sequence set forth in SEQ ID NO: 68, and this amino acid sequence is such that X28 is selected from C, D, S, H, P, G, and R; X38 is selected from E and L; X46 is selected from K, R, and Q; X54 is selected from R, Q, and S; X66 is selected from L, T, and V; X75 is R; X85 is selected from G and H; X86 is T; X96 is selected from M and L; X160 is selected from T and P; X176 is selected from M, L, and H; X183 is Q; X188 is selected from G and F; X205 is F; X212 is D; X248 is V; X249 is W; X255 is N; X270 is selected from N and L; and X286 is selected from M, V, G, N, and S. In another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence that includes at least one feature selected from the group consisting of X28 being a polar residue, X38 being a basic residue, and X85 being a small residue. In yet another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence that includes the following features: X28 is C; X38 is E; and X85 is G.

[0101] In other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising a residue different from the amino acid sequence set forth in SEQ ID NO: 100 at one residue position selected from X7, X22, X36, X38, X46, X54, X66, and X75. In another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence having at least one characteristic selected from the group consisting of: the residue corresponding to X7 is an aliphatic residue; the residue corresponding to X22 is selected from an aliphatic residue and an aromatic residue; the residue corresponding to X36 is selected from a polar residue and a nonpolar residue; the residue corresponding to X38 is an aromatic residue; the residue corresponding to X46 is selected from a polar residue and a basic residue; the residue corresponding to X54 is selected from a polar residue and a basic residue; the residue corresponding to X66 is a polar residue; and the residue corresponding to X75 is selected from a basic residue and a nonpolar residue.

[0102] In yet other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence that is at least 80% identical to the amino acid sequence set forth in SEQ ID NO: 100, the amino acid sequence having at least one characteristic selected from the group consisting of: X7 is L; X22 is selected from W and L; X36 is selected from T and M; X38 is W; X46 is selected from K and Q; X54 is selected from S, Q, and K; X66 is selected from G and T; and X75 is selected from M and R. In another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence having at least one characteristic selected from the group consisting of: X36 is a polar residue; X38 is an aromatic residue; and X75 is a basic residue. In yet another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence having the following characteristics: X36 is T; X38 is W; and X75 is R.

[0103] In other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence comprising a residue different from the amino acid sequence shown in SEQ ID NO: 114 at one residue position selected from X2, X181, and X286. In another embodiment, the improved carboxylesterase polypeptide is characterized in that the residue corresponding to X2 is selected from aliphatic residues, basic residues, polar residues, and aromatic residues; the residue corresponding to X181 is a basic residue; and the residue corresponding to X286 is selected from polar residues and nonpolar residues.

[0104] In still other embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence that is at least 80% identical to the amino acid sequence shown in SEQ ID NO: 114, and this amino acid sequence is characterized in that X2 is selected from L, Q, R, and H; X181 is Q; and X286 is selected from C and S. In another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence characterized in that X286 is a nonpolar residue. In yet another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence characterized in that X286 is C.

[0105] In some embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence corresponding to the amino acid sequence shown in SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0106] In a further aspect, the present disclosure provides polynucleotides encoding each of the above-described engineered, engineered carboxylesterase polypeptides. In some embodiments, these polynucleotides can be part of an expression vector having one or more control sequences for the expression of the carboxylesterase polypeptide. In another embodiment, the polynucleotide corresponds to any one of the nucleotide sequences set forth in SEQ ID NO: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, or 125.

[0107] In another aspect, the present disclosure provides a host cell comprising a polynucleotide encoding an engineered carboxylesterase or an expression vector capable of expressing an engineered carboxylesterase. In some embodiments, the host cell can be a bacterial host cell such as Escherichia coli (E. coli). These host cells can be used for the expression and isolation of the engineered carboxylesterase enzymes described herein, or the host cells can also be used directly for the conversion of ester substrates to products. In some embodiments, the engineered amides can be used individually or as combinations of various engineered amides in the form of whole cells, crude extracts, isolated polypeptides, or purified polypeptides.

[0108] One skilled in the art will recognize that during the production of an enzyme, post-translational modifications can occur, particularly depending on the cell line used and the specific amino acid sequence of the enzyme. For example, such post-translational modifications can include cleavage of a specific leader sequence, addition of various sugar moieties in various glycosylation and phosphorylation patterns, deamidation, oxidation, disulfide bond scrambling, isomerization, C-terminal lysine clipping, and N-terminal glutamine cyclization. The present invention encompasses the use of engineered carboxylesterase enzymes that have undergone, or are likely to have undergone, one or more post-translational modifications. Thus, the engineered carboxylesterases of the present invention can include those that have undergone post-translational modifications such as those described herein.

[0109] Deamidation is an enzymatic reaction that mainly converts asparagine (N) to isoaspartic acid (isoaspartate) and aspartic acid (aspartate) (D) in an approximate ratio of 3:1. Thus, this deamidation reaction is related to the isomerization of aspartate (D) to isoaspartate. Both the deamidation of asparagine and the isomerization of aspartate involve an intermediate succinimide. To a much lesser extent, deamidation can also occur with glutamine residues.

[0110] Oxidation can occur during production and storage (i.e., in the presence of oxidizing conditions) and results in covalent modification of the protein induced directly by reactive oxygen species or indirectly by reaction with secondary by-products of oxidative stress. Oxidation mainly occurs with methionine residues, but can also occur with tryptophan residues and free cysteine residues.

[0111] Disulfide bond scrambling can occur during production and under basic storage conditions. Under certain circumstances, disulfide bonds can break or form inappropriately, resulting in unpaired cysteine residues (-SH). These free (unpaired) sulfhydryl (-SH) groups can promote shuffling.

[0112] The N-terminal glutamine (Q) and glutamate (glutamic acid) (E) in the engineered carboxylesterase may form pyroglutamate (pGlu) via cyclization. Most pGlu formation occurs during manufacture, but it may also be formed non-enzymatically depending on the pH and temperature of the processing and storage conditions.

[0113] C-terminal lysine clipping is an enzymatic reaction catalyzed by carboxypeptidase and is commonly seen in enzymes. As a variation of this process, removal of lysine from enzymes derived from recombinant host cells is included.

[0114] In the present invention, it is not known that the post-translational modifications and changes in the above primary amino acid sequence cause a significant change in the activity of the engineered carboxylesterase enzyme.

[0115] Table 3 below shows examples of carboxylesterase polypeptides engineered therein, listing two accession numbers in each column, with the odd numbers representing nucleotide sequences encoding amino acid sequences shown by the even numbers. The residue differences are based on comparison with the reference sequence of SEQ ID NO: 2, which is a carboxylesterase corresponding to wild-type A. acidocaldarius esterase 2 as shown in Example 6.In the activity column, the levels of activity enhancement (i.e., "+", "++", "+++", etc.) were defined as follows: "-" represents a substrate-to-product conversion rate of less than 1%, indicating a conversion rate of 0.9% or less (175 μL of lysate, in TBME, 100 mM ester, 100 mM isopropylpiperazine, 2% water); "+" indicates at least 1.1 to 80 times the activity of SEQ ID NO: 2 and not exceeding the activity of SEQ ID NO: 4 (175 μL of lysate, in TBME, 100 mM ester, 100 mM isopropylpiperazine, 2% water); "++" indicates at least 1.1 to 11 times the activity of SEQ ID NO: 4 and not exceeding the activity of SEQ ID NO: 18 (150 μL of lysate, in MIBK, 200 mM ester, 200 mM isopropylpiperazine, 2% water); "+++" indicates at least 1.1 to 5 times the activity of SEQ ID NO: 18 and not exceeding the activity of SEQ ID NO: 54 (120 μL of lysate, in MIBK, 300 mM ester, 300 mM isopropylpiperazine, 2% water); "++++" indicates at least 1.5 to 2 times the activity of SEQ ID NO: 54 and not exceeding the activity of SEQ ID NO: 68 (90 μL of lysate, in MIBK, 300 mM ester, 300 mM isopropylpiperazine, 2% water); "+++++" indicates at least 1.1 to 2.0 times the activity of SEQ ID NO: 68 (50 μL of lysate, in MIBK, 300 mM ester, 300 mM isopropylpiperazine, 2% water); "$" indicates at least 1.1 to 2 times the activity of SEQ ID NO: 68 and not exceeding the activity of SEQ ID NO: 100 (50 μL of lysate, in MIBK, 354 mM ester, 425 mM isopropylpiperazine, 2% water); "$$" indicates at least 1.1 to 5 times the activity of SEQ ID NO: 100 and not exceeding the activity of SEQ ID NO: 114 (50 μL of lysate, in MIBK, 354 mM ester, 425 mM isopropylpiperazine, 2% water); "$$$" indicates at least 1.1 to 2 times the activity of SEQ ID NO: 114 (50 μL of lysate, in MIBK, 354 mM ester, 354 mM isopropylpiperazine, 2% water).In each case, the activity was determined using various amounts of the lysate, as described in Example 5, by adding the lysate to a multiwell freeze-drying and activity screen and then reacting it with the substrate at the indicated concentration in a volume of 200 μL for 16 hours in the indicated solvent system.

[0116]

Table 3

[0117] As described above, in some embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence having at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity, or more than that, with a reference sequence of SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126. In some embodiments, the improved carboxylesterase polypeptide may have a difference of 1 - 2, 1 - 3, 1 - 4, 1 - 5, 1 - 6, 1 - 7, 1 - 8, 1 - 9, 1 - 10, 1 - 11, 1 - 12, 1 - 13, 1 - 14, 1 - 15, 1 - 16, 1 - 17, 1 - 18, 1 - 19, 1 - 20, 1 - 22, 1 - 24, 1 - 26, 1 - 28, 1 - 30, 1 - 35, 1 - 40, 1 - 45, 1 - 50, 1 - 55, 1 - 60, or 1 - 62 residues when compared with the carboxylesterase shown in SEQ ID NO: 2. In some embodiments, the number of residue differences may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 35, 40, 45, 50, 55, 60, or 62 differences when compared with SEQ ID NO: 2.

[0118] In some embodiments, the improved carboxylesterase polypeptide comprises an amino acid sequence that is at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to a reference sequence based on SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126, provided that the improved carboxylesterase amino acid sequence comprises any one of the sets of residue differences included in any one of the polypeptide sequences listed in Table 3 when compared to SEQ ID NO: 2. In some embodiments, the improved carboxylesterase polypeptide may further have differences of 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-14, 1-15, 1-16, 1-18, 1-20, 1-22, 1-24, 1-26, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, or 1-62 residues at other amino acid residue positions when compared to the reference sequence. In some embodiments, the number of differences may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 18, 20, 22, 24, 26, 30, 35, 40, 45, 50, 55, or 62 residue differences at other residue positions. In some embodiments, the residue differences at other residue positions comprise substitutions by conservative amino acid residues.

[0119] In some embodiments, an improved carboxylesterase polypeptide capable of converting the ester substrate ethyl oxazole-5-carboxylate to a product level detectable by HPLC-UV at 230 nm in the presence of an amine substrate in water-saturated MIBK comprises an amino acid sequence selected from the sequences of SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0120] In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is 2-fold, 5-fold, 10-fold, 20-fold, 25-fold, 50-fold, 75-fold, 100-fold, 1000-fold, 10,000-fold, 100,000-fold, 500,000-fold, 785,000-fold or more than that of the polypeptide of SEQ ID NO: 2. In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is 50- to 100-fold or more than that of the polypeptide of SEQ ID NO: 2 and comprises an amino acid sequence corresponding to SEQ ID NO: 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0121] In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.1 to 5 times or more than that of the polypeptide of SEQ ID NO: 24. In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.1 to 5 times or more than that of the polypeptide of SEQ ID NO: 24, and comprises an amino acid sequence corresponding to the sequence of SEQ ID NO: 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0122] In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.1 to 5 times or more than that of the polypeptide of SEQ ID NO: 54. In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.1 to 5 times or more than that of the polypeptide of SEQ ID NO: 54, and comprises a sequence corresponding to the sequence of SEQ ID NO: 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0123] In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.1 to 6-fold or more than that of the polypeptide of SEQ ID NO: 68. In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.1 to 5-fold or more than that of the polypeptide of SEQ ID NO: 68, and comprises an amino acid sequence corresponding to the sequence of SEQ ID NO: 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0124] In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.1 to 5-fold or more than that of the polypeptide of SEQ ID NO: 100. In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.7-fold or more than that of the polypeptide of SEQ ID NO: 100, and comprises an amino acid sequence corresponding to the sequence of SEQ ID NO: 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0125] In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.1 to 2-fold or more than that of the polypeptide of SEQ ID NO: 114. In some embodiments, the engineered carboxylesterase polypeptide can convert an ester substrate to a product with an activity that is about 1.1 to 5-fold or more than that of the polypeptide of SEQ ID NO: 114, and comprises an amino acid sequence corresponding to the sequence of SEQ ID NO: 122, 124, or 126.

[0126] In some embodiments, the engineered carboxylesterase polypeptide can comprise a deletion in a specific amino acid residue of the engineered carboxylesterase polypeptide described herein. Thus, for each embodiment of the carboxylesterase polypeptide of the present disclosure, the deletion can comprise 1 or more amino acids, 2 or more amino acids, 3 or more amino acids, 4 or more amino acids, 5 or more amino acids, 6 or more amino acids, 8 or more amino acids, 10 or more amino acids, 15 or more amino acids, or 20 or more amino acids of the carboxylesterase polypeptide, up to 5% of the total number of amino acids, up to 10% of the total number of amino acids, up to 20% of the total number of amino acids, or up to 30% of the total number of amino acids, as long as the functional activity of the carboxylesterase activity is maintained. In some embodiments, the deletion can comprise 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-14, 1-15, 1-16, 1-18, 1-20, 1-22, 1-24, 1-26, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, or up to 1-62 amino acid residues. In some embodiments, the number of deletions can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 18, 20, 22, 24, 26, 30, 35, 40, 45, 50, 55, 60, or up to 62 amino acids. As described herein, the carboxylesterase polypeptide of the present disclosure can be in the form of a fusion polypeptide in which the carboxylesterase polypeptide is fused to other polypeptides such as, by way of example and not limitation, antibody tags (e.g., myc epitope), purification sequences (e.g., His tag for binding to metal), and cell localization signals (e.g., secretion signal). Thus, the carboxylesterase polypeptide can be used with or without being fused to other polypeptides.

[0127] The polypeptides described herein are not limited to genetically encoded amino acids. In addition to genetically encoded amino acids, the polypeptides described herein can be composed, in whole or in part, of natural and / or synthetic non-coded amino acids. Some specific non-coded amino acids that are commonly encountered and can constitute the polypeptides described herein include, but are not limited to, D-stereoisomers of genetically encoded amino acids; 2,3-diaminopropionic acid (Dpr); α-aminoisobutyric acid (Aib); ε-aminohexanoic acid (Aha); δ-aminovaleric acid (Ava); N-methylglycine or sarcosine (MeGly or Sar); ornithine (Orn); citrulline (Cit); t-butylalanine (Bua); t-butylglycine (Bug); N-methylisoleucine (MeIle); phenylglycine (Phg); cyclohexylalanine (Cha); norleucine (Nle); naphthylalanine (Nal); 2-chlorophenylalanine (Ocf); 3-chlorophenylalanine (Mcf); 4-chlorophenylalanine (Pcf); 2-fluorophenylalanine (Off); 3-fluorophenylalanine (Mff); 4-fluorophenylalanine (Pff); 2-bromophenylalanine (Obf); 3-bromophenylalanine (Mbf); 4-bromophenylalanine (Pbf); 2-methylphenylalanine (Omf); 3-methylphenylalanine (Mmf); 4-methylphenylalanine (Pmf); 2-nitrophenylalanine (Onf); 3-nitrophenylalanine (Mnf); 4-nitrophenylalanine (Pnf); 2-cyanophenylalanine (Ocf); 3-cyanophenylalanine (Mcf); 4-cyanophenylalanine (Pcf); 2-trifluoromethylphenylalanine (Otf); 3-trifluoromethylphenylalanine (Mtf); 4-trifluoromethylphenylalanine (Ptf); 4-aminophenylalanine (Paf); 4-iodophenylalanine (Pif); 4-aminomethylphenylalanine (Pamf); 2,4-dichlorophenylalanine (Opef); 3,4-dichlorophenylalanine (Mpcf); 2,4-difluorophenylalanine (Opff); 3,4-difluorophenylalanine (Mpff);Pyrid-2-ylalanine (2pAla); Pyrid-3-ylalanine (3pAla); Pyrid-4-ylalanine (4pAla); Naphth-1-ylalanine (1nAla); Naphth-2-ylalanine (2nAla); Thiazolylalanine (taAla); Benzothienylalanine (bAla); Thienylalanine (tAla); Furylalanine (fAla); Homophenylalanine (hPhe); Homotyrosine (hTyr); Homotryptophan (hTrp); Pentafluorophenylalanine (5ff); Styrylalanine (sAla); Anthrylalanine (aAla); 3,3-Diphenylalanine (Dfa); 3-Amino-5-phenylpentanoic acid (Afp); Penicillamine (Pen); 1,2,3,4-Tetrahydroisoquinoline-3-carboxylic acid (Tic); β-2-Thienylalanine (Thi); Methionine sulfoxide (Mso); N(w)-Nitroarginine (nArg); Homolysine (hLys); Phosphonomethylphenylalanine (pmPhe); Phosphoserine (pSer); Phosphothreonine (pThr); Homolaspartic acid (hAsp); Homoglutamic acid (hGlu); 1-Aminocyclopent-(2 or 3)-ene-4-carboxylic acid; Piperic acid (PA), Azetidine-3-carboxylic acid (ACA); 1-Aminocyclopentane-3-carboxylic acid; Allylglycine (aOly); Propargylglycine (pgGly); Homolalanine (hAla); Norvaline (nVal); Homoleucine (hLeu), Homovaline (hVal); Homoisoleucine (hIle); Homoarginine (hArg); N-Acetyllysine (AcLys); 2,4-Diaminobutyric acid (Dbu); 2,3-Diaminobutyric acid (Dab); N-Methylvaline (MeVal); Homocysteine (hCys); Homoserine (hSer);It contains hydroxyproline (Hyp) and homoproline (hPro). Additional non-coded amino acids that can constitute the polypeptides described herein will be apparent to those skilled in the art (see, for example, Fasman, 1989, CRC Practical Handbook of Biochemistry and Molecular Biology, CRC Press, Boca Raton, FL, at pp. 3-70 and the various amino acids shown in the references cited therein). These amino acids can be of the L or D type.;

[0128] Those skilled in the art will recognize that amino acids or residues with side chain protecting groups can also constitute the polypeptides described herein. Non-limiting examples of such protected amino acids (in this case belonging to the aromatic category) (where the protecting groups are listed in parentheses) include, but are not limited to, Arg(tos), Cys(methylbenzyl), Cys(nitropyridinesulfenyl), Glu(δ-benzyl ester), Gln(xanthyl), Asn(N-δ-xanthyl), His(bom), His(benzyl), His(tos), Lys(fmoc), Lys(tos), Ser(O-benzyl), Thr(O-benzyl), and Tyr(O-benzyl).

[0129] Conformationally constrained non-coded amino acids that can constitute the polypeptides described herein include, but are not limited to, N-methyl amino acid (L-type); 1-aminocyclopenta-(2 or 3)-ene-4-carboxylic acid; pipecolic acid; azetidine-3-carboxylic acid; homoproline (hPro); and 1-aminocyclopentane-3-carboxylic acid.

[0130] As described above, the various modifications introduced into the native polypeptide to create the engineered carboxylesterase enzyme can be aimed at affecting specific properties of the enzyme such as activity, specificity for its substrate, and thermal stability.

[0131] In another aspect, the present disclosure provides polynucleotides encoding improved carboxylesterase polypeptides. These polynucleotides may be operably linked to one or more heterologous regulatory sequences that control gene expression to create recombinant polynucleotides capable of expressing their carboxylesterase polypeptides. Expression constructs containing heterologous polynucleotides encoding engineered carboxylesterases can be introduced into appropriate host cells to express the corresponding carboxylesterase polypeptides.

[0132] Due to knowledge of the codons corresponding to various amino acids, the availability of the protein sequence provides an account of all the polynucleotides capable of encoding the subject. The degeneracy of the genetic code, where the same amino acid is encoded by another codon or synonymous codon, allows for the creation of a vast number of nucleic acids, all of which encode the improved carboxylesterase polypeptides disclosed herein. Thus, once a particular amino acid sequence is identified, one of ordinary skill in the art can create any number of different nucleic acids by simply modifying the sequence of one or more codons without changing the amino acid sequence of the protein. In this regard, the present disclosure specifically contemplates each possible variant of the polynucleotide that can be created by selecting combinations based on possible codon choices, and all such variants should be considered specifically disclosed with respect to any of the polypeptides disclosed herein, including the amino acid sequences shown in Table 3.

[0133] In some embodiments, these polynucleotides can be selected and / or engineered to comprise codons selected to be compatible with the host cell in which the protein is to be produced. For example, preferred codons used in bacteria are used for gene expression in bacteria; preferred codons used in yeast are used for expression in yeast; and preferred codons used in mammals are used for expression in mammalian cells. Since it is not necessary to replace all codons to optimize the codon usage frequency of carboxylesterase (e.g., the native sequence may have preferred codons, and the use of preferred codons is not required for all amino acid residues), a polynucleotide having codons optimized to encode a carboxylesterase polypeptide can comprise preferred codons at positions of more than about 40%, 50%, 60%, 70%, 80%, or 90% of the full-length coding region.

[0134] In some embodiments, this polynucleotide encodes a carboxylesterase polypeptide comprising an amino acid sequence that is at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical, or more identical, to the reference sequence of SEQ ID NO: 4, or a functional fragment thereof, and this polypeptide can convert an ester substrate in the presence of an amine substrate with an activity improved as compared to the activity of the carboxylesterase of SEQ ID NO: 2 from A. acidocaldarius esterase 2.

[0135] In some embodiments, this polynucleotide encodes a carboxylesterase polypeptide comprising an amino acid sequence having at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity with a polypeptide comprising the amino acid sequence corresponding to SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126, or a functional fragment thereof, and this polypeptide has at least one improved property in converting the ester substrate ethyl oxazole-5-carboxylate to the product (4-isopropylpiperazin-1-yl)(oxazol-5-yl)methanone in the presence of the amine substrate 1-isopropylpiperazine. In some embodiments, the encoded carboxylesterase polypeptide has an activity equal to or greater than the activity of the polypeptide of SEQ ID NO: 2.

[0136] In some embodiments, the polynucleotide encodes a carboxylesterase polypeptide comprising an amino acid sequence that is at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to a reference sequence based on SEQ ID NO: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126, or a functional fragment thereof, provided that the improved carboxylesterase amino acid sequence comprises any one of the sets of residue differences included in any one of the polypeptide sequences listed in Table 3 when compared to SEQ ID NO: 2.

[0137] In some embodiments, the polynucleotide encoding the improved carboxylesterase polypeptide is selected from SEQ ID NO: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, or 125.

[0138] In some embodiments, the polynucleotide can hybridize under high stringency conditions with a polynucleotide comprising SEQ ID NO: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, or 125, or its complement, and a polynucleotide that hybridizes under high stringency encodes a carboxylesterase polypeptide that can convert a product with equal or greater activity than the polypeptide of SEQ ID NO: 2 in the presence of an amine substrate.

[0139] In some embodiments, these polynucleotides encode the polypeptides described herein and have at least about 80% sequence identity, about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity at the nucleotide level with a reference polynucleotide encoding the engineered carboxylesterase described herein. In some embodiments, the reference polynucleotide is selected from SEQ ID NO: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 121, 123, or 125.

[0140] In some embodiments, the carboxylesterase polypeptide comprises the amino acid sequence set forth in SEQ ID NO: 122. In another embodiment, the present disclosure provides a polynucleotide sequence encoding the carboxylesterase polypeptide sequence set forth in SEQ ID NO: 122. In yet another embodiment, the present disclosure provides a polynucleotide encoding a carboxylesterase polypeptide, the polynucleotide comprising the polynucleotide sequence set forth in SEQ ID NO: 121. In yet another embodiment, the carboxylesterase polypeptide consists of the polypeptide sequence set forth in SEQ ID NO: 122. In another embodiment, the carboxylesterase polypeptide consists of residues 2-310 of SEQ ID NO: 122.

[0141] Improved carboxylesterases and polynucleotides encoding such polypeptides can be produced using methods conventional to those skilled in the art. As noted above, the wild-type carboxylesterase enzyme A. acidocaldarius esterase 2, from which the parental sequence SEQ ID NO: 2 is derived, the native amino acid sequence and corresponding polynucleotide encoding it are available in WO02 / 057411 (see SEQ ID NO: 10). In some embodiments, the parental polynucleotide sequence is codon-optimized to enhance the expression of carboxylesterase in the specified host cell. Engineered carboxylesterases can be obtained by subjecting the polynucleotide encoding the native carboxylesterase to mutagenesis and / or directed evolution methods. Examples of directed evolution techniques are mutagenesis and / or DNA shuffling as described in Stemmer, 1994, Proc Natl Acad Sci USA 91:10747-10751; WO95 / 22625; WO97 / 0078; WO97 / 35966; WO98 / 27230; WO00 / 42651; WO01 / 75767; and U.S. Patent No. 6,537,746.

[0142] Other possible directed evolution methods include, among others, the staggered extension process (StEP), in vitro recombination (Zhao, et al., 1998, Nat. Biotechnol. 16:258-261), mutagenic PCR (Caldwell, et al., 1994, PCR Methods Appl. 3:S136-S140), and cassette mutagenesis (Black, et al., 1996, Proc Natl Acad Sci USA 93:3525-3529). Mutagenesis and directed evolution techniques useful for the purposes of this specification are also described in the following references: Ling, et al., 1997, “Approaches to DNA mutagenesis: an overview,” Anal. Biochem. 254(2):157-78; Dale, et al., 1996, “Oligonucleotide-directed random mutagenesis using the phosphorothioate method,” Methods Mol. Biol. 57:369-74; Smith, 1985, “In vitro mutagenesis,” Ann. Rev. Genet. 19:423-462; Botstein, et al., 1985, “Strategies and applications of in vitro mutagenesis,” Science 229:1193-1201; Carter, 1986, “Site-directed mutagenesis,” Biochem. J. 237:1-7; Kramer, et al., 1984, “Point Mismatch Repair,” Cell 38:879-887; Wells, et al., 1985, “Cassette mutagenesis: an efficient method for generation of multiple mutations at defined sites,” Gene 34:315-323; Minshull, et al., 1999, “Protein evolution by molecular breeding,” Curr Opin Chem Biol 3:284-290; Christians, et al., 1999, “Directed evolution of thymidine kinase for AZT phosphorylation using DNA family shuffling,” Nature Biotech 17:259-264; Crameri, et al., 1998, “DNA shuffling of a family of genes from diverse species accelerates directed evolution,” Nature 391:288-291; Crameri, et al., 1997, “Molecular evolution of an arsenate detoxification pathway by DNA shuffling,” Nature Biotech 15:436-438; Zhang, et al., 1997, “Directed evolution of an effective fructosidase from a galactosidase by DNA shuffling and screening,” Proc Natl Acad Sci USA 94:45-4-4509; Crameri, et al., 1996, “Improved green fluorescent protein by molecular evolution using DNA shuffling,’ Nature Biotech 14:315-319; and Stemmer, 1994, “Rapid evolution of a protein in vitro by DNA shuffling,” Nature 370:389-391。.

[0143] In some embodiments, the clones obtained after mutagenesis treatment are screened for carboxylesterases having the desired improved enzyme properties. Measurement of carboxylesterase enzyme activity from an expression library can be performed using standard techniques such as detection of the product by separation of the product (e.g., by HPLC) and measurement of the UV absorbance of the separated substrate and product and / or detection using tandem mass spectrometry (e.g., MS / MS). An assay example is described in Example 4 below. The rate of increase in the amount of the desired product per unit time indicates the relative (enzymatic) activity of the carboxylesterase polypeptide in a fixed amount of lysate (or lyophilized powder made therefrom). If the desired improvement in enzyme properties is thermostability, the enzyme activity can be measured by measuring the amount of enzyme activity remaining after heat treatment after exposing the enzyme preparation to a defined temperature. Clones containing the polynucleotide encoding the desired carboxylesterase are then isolated, sequenced to identify changes (if any) in the nucleotide sequence, and used to express the enzyme in a host cell.

[0144] In other embodiments well known in the art, the enzymes may be genetically diversified while maintaining their target activity, such as by the technique of neutral drift.

[0145] If the sequence of the engineered polypeptide is known, the polynucleotide encoding the enzyme can be made by standard solid-phase methods according to known synthetic methods. In some embodiments, fragments up to about 100 bases are synthesized individually and then ligated (e.g., by enzymatic or chemical litigation methods, or polymerase-mediated methods) to form any desired contiguous sequence. For example, the polynucleotides and oligonucleotides of the present invention can be made by chemical synthesis using, for example, the classical phosphoramidite method described by Beaucage, et al., 1981, Tet Lett 22:1859-69 or the method described by Matthes, et al., 1984, EMBO J. 3:801-05, as is typically carried out by automated synthetic methods. According to the phosphoramidite method, oligonucleotides are synthesized, purified, annealed, ligated, and cloned into an appropriate vector, for example, in an automated DNA synthesizer. Furthermore, essentially any nucleic acid can be obtained from any of The Great American Gene Company, Ramona, CA, ExpressGen Inc, Chicago, IL, Operon Technologies Inc, Alameda, CA, and many other various commercial sources.

[0146] The engineered carboxylesterase enzyme expressed in the host cell can be recovered from the cells and / or culture medium using any one or more of the well-known techniques for protein purification, including, among other things, lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography. A solution suitable for lysis and high-efficiency extraction of proteins from bacteria such as E. coli is commercially available from Sigma-Aldrich, St. Louis, MO under the trademark CelLytic BTM.

[0147] As chromatography techniques for the isolation of carboxylesterase polypeptides, there are included, inter alia, reverse phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. The conditions for purifying a specific enzyme vary to some extent depending on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and will be apparent to those skilled in the art. In some embodiments, the engineered carboxylesterase can be expressed as a fusion protein with a purification tag such as a His tag having an affinity for metal or an antibody tag for binding to an antibody, for example, a myc epitope tag.

[0148] In some embodiments, affinity techniques can be used to purify the improved carboxylesterase enzyme. For affinity chromatography purification, any antibody that specifically binds to the carboxylesterase polypeptide can be used. For the production of antibodies, various host animals including, but not limited to, rabbits, mice, rats, etc. can be immunized by injecting the engineered polypeptide. This polypeptide can be conjugated to a suitable carrier such as BSA by a side chain functional group or a linker attached to the side chain functional group. Depending on the host species, various adjuvants including, but not limited to, Freund's (complete and incomplete), inorganic gels such as aluminum hydroxide, lysolecithin, pluronic polyol, polyanion, peptide, oil emulsion, keyhole limpet hemocyanin, surfactants such as dinitrophenol, and potentially useful human adjuvants such as BCG (bacilli Calmette Guerin) and Corynebacterium parvum can be used to enhance the immune response.

[0149] In a further aspect, the improved carboxylesterase polypeptide described herein can be used in a method for amidating a specific amide group acceptor (e.g., an ester substrate) in the presence of an amine substrate.

[0150] In some embodiments, the improved carboxylesterase is (a) an ester having the structure R1-COOR2 [where R1 is selected from sp 3 carbon and aromatic rings and has 0 to 3 alkyl substituents, and R2 is selected from a methyl group; an ethyl group; and a C1-C6 alkyl chain]; (b) an amine substrate; (c) an improved carboxylesterase polypeptide; and (d) a solvent which can be used in a method for producing an amide by combining the components containing them.

[0151] In one embodiment, the solvent is an organic solvent containing up to 3 molar equivalents of water relative to the ester substrate in an amount of about 0.5% (vol / vol) to about 3% (vol / vol).

[0152] In some embodiments of this method, the improved carboxylesterase is selected from SEQ ID NOs: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, and 126.

[0153] In other embodiments, the improved carboxylesterase polypeptide is (a) R1-COOR2 [where R1 is sp having 0 to 3 alkyl substituents 3An ester substrate having a structure selected from carbon and aromatic rings and wherein R2 is selected from a methyl group; an ethyl group; and a C1-C6 alkyl chain; (b) an amine substrate; (c) the improved carboxylesterase polypeptide described above; and (d) a component containing a solvent, which can be used in a method for producing an amide molecule. In another embodiment of this method, an organic solvent is used. In one embodiment, the organic solvent is selected from toluene; 2-methyltetrahydrofuran; tetrahydrofuran; dimethylacetamide; methyl isobutyl ketone (MIBK); dichloromethane; tert-butyl methyl ether; cyclopentyl methyl ether; methylcyclohexane; dichloromethane; acetonitrile; methyl ethyl ketone; isopropyl acetate; ethanol; isopropanol; ethyl acetate; heptane; xathane; and 2-methyltetrahydrofuran (2-Me-THF); and water. In yet another embodiment, the organic solvent contains up to 3 molar equivalents of water relative to the ester substrate in an amount of about 0.5% (vol / vol) to about 3% (vol / vol). In another embodiment of this method, the carboxylesterase polypeptide in step (c) is produced in the presence of a salt to stabilize its physical form during the reaction. In yet another embodiment, the salt is added as an additional reaction component. In one embodiment of this method, the ester has the formula: [Chemical formula] is ethyl oxazole-5-carboxylate; the amine substrate has the formula: [Chemical formula] is 1-isopropylpiperazine; the amide has the formula: [Chemical formula] is (4-isopropylpiperazin-1-yl)(oxazol-5-yl)methanone.

[0154] In yet another embodiment of this method, the ester has the formula: [ka] and the amine substrate is ethyl oxazole-5-carboxylate having the formula: [ka] and the amide is cis-2,6-dimethylmorpholine having the formula: [ka] The compound is ((2S,6R)-2,6-dimethylmorpholino)(oxazol-5-yl)methanone, having the formula:

[0155] In another embodiment, the reactants in this method comprise about 50 g / L of ethyl oxazole-5-carboxylate, 44 g / L of 1-isopropylpiperazine, and about 25 g / L of SEQ ID NOs: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74 , 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126, wherein the carboxylesterase is produced in the presence of sodium sulfate and is functional in the presence of about 10 g / L to about 20 g / L of water in MIBK.

[0156] In some embodiments, the invention is an amide made by these methods using the improved carboxylesterase. In another embodiment, the invention is an amide of the formula: [ka] 4-Isopropylpiperazin-1-yl)(oxazol-5-yl)methanone having the formula, and the amide of formula III is prepared by the above method.

[0157] In another embodiment, the invention provides an amide of the formula:

Chemical formula

[0158] Compounds of formula I (Astatech, AS-23210), formula II (Oakwood Products Inc, OAK-008910), and formula IV (Oakwood Products Inc, OAK-091224) were obtained from commercial suppliers.

[0159] In some embodiments, the method comprises contacting or incubating an ethyl oxazole-5-carboxylate ester substrate with an improved carboxylesterase in the presence of an amine substrate under suitable reaction conditions to convert the ester substrate to the product (4-isopropylpiperazin-1-yl)(oxazol-5-yl)methanone with a conversion rate and / or activity of about 50 to about 785,000-fold or more that of SEQ ID NO: 2. Examples of polypeptides comprise amino acid sequences corresponding to SEQ ID NO: 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, or 126.

[0160] In some embodiments of the above method, the reaction solvent for carrying out this method is selected from methyl isobutyl ketone (MIBK), toluene, tert-butyl methyl ether (TBME) or 2-methyltetrahydrofuran (2-Me-THF).

[0161] In some embodiments of the above method, the enzyme preparation for the reaction contains a salt selected from one of potassium phosphate (KPi), potassium sulfate, or sodium sulfate.

[0162] In some embodiments, the reaction conditions for carrying out this method can include a temperature of about 15°C to about 30°C. In one embodiment, the amine substrate used in this method can be a chiral amine or an achiral amine. The achiral amine substrate is not limited to a specific stereoisomer in its reaction and thus has the advantage of not requiring the amine substrate as much. Various suitable amine substrates can be used, and by way of example, but not limited to, 1-isopropylpiperazine and cis-2,6-dimethylmorpholine can be mentioned.In some embodiments, other amine substrates can also be used, in particular, α-phenethylamine (also called 1-phenylethylamine), and its enantiomers (S)-1-phenylethylamine and (R)-1-phenylethylamine, 2-amino-4-phenylbutane, glycine, L-glutamic acid, L-glutamate, monosodium glutamate, L-aspartic acid, L-lysine, L-ornithine, β-alanine, taurine, n-octylamine, cyclohexylamine, 1,4-butanediamine, 1,6-hexanediamine, 6-aminohexanoic acid, 4-aminobutyric acid, tyramine, and benzylamine, 2-aminobutane, 2-amino-1-butanol, 1-amino-1-phenylethane, 1-amino-1-(2-methoxy-5-fluorophenyl)ethane, 1-amino-1-phenylpropane, 1-amino-1-(4-hydroxyphenyl)propane, 1-amino-1-(4-bromophenyl)propane, 1-amino-1-(4-nitrophenyl)propane, 1-phenyl-2-aminopropane, 1-(3-trifluoromethylphenyl)-2-aminopropane, 2-aminopropanol, 1-amino-1-phenylbutane, 1-phenyl-2-aminobutane, 1-(2,5-dimethoxy-4-methylphenyl)-2-aminobutane, 1-phenyl-3-aminobutane, 1-(4-hydroxyphenyl)-3-aminobutane, 1-amino-2-methylcyclopentane, 1-amino-3-methylcyclopentane, 1-amino-2-methylcyclohexane, 1-amino-1-(2-naphthyl)ethane, 3-methylcyclopentylamine, 2-methylcyclopentylamine, 2-ethylcyclopentylamine, 2-methylcyclohexylamine, 3-methylcyclohexylamine, 1-aminotetralin, 2-aminotetralin, 2-amino-5-methoxytetralin, and 1-aminoindane, and, if possible, both the (R) and (S) single enantiomers are included.

[0163] In some embodiments, a method for converting ethyl oxazole-5-carboxylate, an ester substrate, comprises contacting about 36 mL / L of the ester substrate in MIBK in the presence of 43 mL / L of 1-isopropylpiperazine at a temperature of about 30 °C with about 20 g / L of the carboxylesterase described herein, and at least 80%, 85%, 90%, 92%, 94%, 96%, or 98% or more of the ester substrate is converted to product in 24 hours. In some embodiments, the carboxylesterase polypeptide capable of performing the above reaction comprises an amino acid sequence corresponding to SEQ ID NO: 122.

[0164] In some embodiments, the above method may further comprise the step of isolating the compound of formula III or the compound of formula V from the reaction solvent.

[0165] Also provided herein are carboxylesterase and substrate / product compositions. In some embodiments, these compositions may comprise a compound of formula III or a compound of formula V, and the improved carboxylesterase of the present disclosure.

[0166] Any one or more of the improved, engineered carboxylesterases may be part of this composition.

Examples

[0167] Various features and embodiments of the present disclosure are shown in the following representative examples, which are illustrative and not intended to be limiting.

[0168] Example 1: Acquisition of the Acidicaldarius esterase 2 wild-type carboxylesterase gene and construction of an expression vector The carboxylesterase (CE) coding gene was designed for expression in E. coli based on the reported amino acid sequence of A. acidocaldarius esterase 2 (SEQ ID NO: 2), which is a carboxylesterase, and the codon optimization algorithm described in Example 1 of patent application US2008 / 0248539. After synthesizing oligonucleotides individually, they were ligated using oligonucleotides generally composed of 42 nucleotides. Next, this gene was cloned under the control of the lac promoter in the expression vector pCK110900 (shown in Figure 3 of US application 2006 / 195947, both of which are hereby incorporated by reference in their entirety for all purposes). This expression vector also contains the P15a origin of replication and the chloramphenicol resistance gene. The resulting plasmid was transformed into E. coli W3110 using standard methods. The codon-optimized gene and the encoded polypeptide are shown as SEQ ID NOs: 1 and 2 in Table 3 and the following Sequence Listing, respectively.

[0169] Similarly, the genes encoding the engineered carboxylesterases of the present disclosure listed in Table 3 (SEQ ID NOs: 3 - 94) were cloned into the vector pCK110900 for expression in E. coli W3110.

[0170] Example 2: Production of carboxylesterase powder - shake flask method A single microbial colony of Escherichia coli containing a plasmid encoding the carboxylesterase of interest was inoculated into 50 mL of Luria Bertoni culture medium containing 30 μg / mL of chloramphenicol and 1% glucose. The cells were grown overnight (at least 16 hours) with shaking at 250 rpm at 30 °C in an incubator. The culture was diluted with 1000 mL of Terrific culture medium containing 30 μg / mL of chloramphenicol to an OD600 of approximately 0.2 and grown with shaking at 250 rpm at 30 °C. Expression of the carboxylesterase gene was induced by adding isopropyl β-D-thiogalactoside (IPTG) to a final concentration of 1 mM when the OD600 of the culture reached 0.6 - 0.8, and then incubation was continued overnight (at least 16 hours). The cells were harvested by centrifugation (3738 RCF, 20 minutes, 4 °C), and the supernatant was discarded. The pellet was frozen at -80 °C for 2 hours. Next, the pellet was thawed and resuspended in 3 mL of sodium sulfate buffer (consisting of 15 g / L of anhydrous sodium sulfate in water) per 1 gram of the final pellet mass (for example, 10 g of frozen pellet was suspended in 30 mL of sodium sulfate buffer). The cell debris was removed by centrifugation (15,777 RCF, 40 min, 4 °C). The supernatant of the clear lysate was collected, pooled, and lyophilized to obtain a dry powder of the crude carboxylesterase enzyme.

[0171] Example 3: Production of carboxylesterase powder - fermentation method An aliquot of the frozen working stock (E. coli containing the plasmid with the carboxylesterase gene of interest) was taken out of the freezer and thawed at room temperature. In a 1 L flask, 300 μL of this working stock was inoculated into the primary seed stage of 250 mL of M9YE culture medium (1.0 g / L ammonium chloride, 0.5 g / L sodium chloride, 6.0 g / L disodium hydrogen phosphate, 3.0 g / L potassium dihydrogen phosphate, 2.0 g / L Tastone-154 yeast extract, 1 L / L deionized water) containing 30 μg / ml chloramphenicol and 1% glucose. It was grown with shaking at 220 rpm at 26 °C. When the OD600 of the culture reached 0.5 - 1.0, the flask was taken out of the incubator and immediately used to inoculate the secondary seed stage.

[0172] The secondary seed stage was carried out using 4 L of growth medium (0.88 g / L ammonium sulfate, 0.98 g / L sodium citrate; 12.5 g / L dipotassium hydrogen phosphate trihydrate, 6.25 g / L potassium dihydrogen phosphate, 3.3 g / L Springer 0251 yeast extract, 0.083 g / L ferric ammonium citrate, 0.5 mL / L antifoaming agent, and 8.3 ml / L trace element solution containing 2 g / L calcium chloride dihydrate, 2.2 g / L zinc sulfate heptahydrate, 0.5 g / L manganese sulfate monohydrate, 1 g / L cuprous sulfate heptahydrate, 0.1 g / L ammonium molybdate tetrahydrate and 0.02 g / L sodium tetraborate) sterilized at 120 °C for 40 minutes in a bench-scale 5 L fermenter. 2 ml of the primary seed with OD0.5 - 1.0 was inoculated into the fermenter and incubated at 30 °C, 300 rpm and aeration of 0.5 vvm. When the OD600 of the culture reached 0.5 - 1.0 OD600, the secondary seed was immediately transferred to the final stage of fermentation.

[0173] The final stage of fermentation was carried out in a 10 L fermenter at bench scale, sterilized at 121 °C for 40 minutes, and after sterilization, 6 L of growth medium supplemented with 20 g / L of glucose monohydrate, 0.48 g / L of ammonium chloride, and 0.204 g / L of magnesium sulfate heptahydrate (0.88 g / L of ammonium sulfate, 0.98 g / L of sodium citrate; 12.5 g / L of dipotassium hydrogen phosphate trihydrate, 6.25 g / L of potassium dihydrogen phosphate, 3.3 g / L of Springer 0251 yeast extract, 0.083 g / L of ferric ammonium citrate, 0.5 mL / L of antifoaming agent, and 8.3 mL / L of trace element solution containing 2 g / L of calcium chloride dihydrate, 2.2 g / L of zinc sulfate heptahydrate, 0.5 g / L of manganese sulfate monohydrate, 1 g / L of cuprous sulfate heptahydrate, 0.1 g / L of ammonium molybdate tetrahydrate, and 0.02 g / L of sodium tetraborate). 500 mL of secondary seed with an OD600 of 0.5 - 1.0 was inoculated into the fermenter and incubated at 30 °C with aeration at 1.6 vvm. Dissolved oxygen was controlled at 30% by a variable shaking speed of 300 - 950 rpm. The pH was maintained at 7.0 by addition of 20% v / v ammonium hydroxide. The growth of the culture was maintained by adding a feed solution containing 500 g / L of glucose monohydrate, 12 g / L of ammonium chloride, and 5.1 g / L of magnesium sulfate heptahydrate.

[0174] After the culture reached an OD600 of 80 ± 10, the expression of carboxylesterase was induced by adding isopropyl-β-D-thiogalactoside (IPTG) to a final concentration of 1 mM, and fermentation was continued for an additional 24 hours. Next, the culture was cooled to 8 °C and maintained at this temperature until harvested. The cells were harvested by centrifugation at 4 °C, 5000 G for 40 minutes in a Sorvall RC12BP centrifuge. The harvested cell pellet was then frozen at -80 °C and stored until downstream processing and regeneration as described below.

[0175] The pellets were frozen at -80 °C for 2 hours. Next, the pellets were thawed and resuspended in 3 mL of sodium sulfate buffer (consisting of 15 g / L of anhydrous sodium sulfate in water) per 1 gram of the final pellet mass (for example, 10 g of frozen pellets were suspended in 30 mL of sodium sulfate buffer). After resuspension, the cells were filtered through a 200-μm mesh and then passed twice through a microfluidizer at 12000 psig. The cell residue was removed by centrifugation (15,777 RCF, 40 min, 4 °C). The supernatant of the clear lysate was collected, pooled, and lyophilized to obtain a dry powder of the crude carboxylesterase enzyme. The carboxylesterase powder was stored at -80 °C.

[0176] Example 4: High-throughput analysis method for identifying mutants of A. acidicaldarius esterase 2 capable of converting ester substrates to amides under aqueous conditions UPLC method for determining the conversion of ester substrate I to amide III : The enzymatic conversion of the ester substrate of formula I (commercially available, CAS number 118994-89-1) to the amide of formula III was determined using an Agilent 1290 UPLC equipped with an Agilent Zorbax RRHD Eclipse Plus Phenyl-Hexyl column (3.0 × 50 mm, 1.8 μm) with a gradient of 5 mM NH4Ac in water (mobile phase A) and acetonitrile (mobile phase B) at a flow rate of 2 mL / min and a column temperature of 60 °C. Starting from a ratio of 99.9:0.1 of A:B, the method followed a 0.25-min hold, then a 0.05-min gradient to 80:20 A:B, then a 0.5-min gradient to 60:40 A:B, then a 0.1-min purge gradient to 0:100 A:B, a 0.2-min hold at 0:100 A:B, and a 0.1-min gradient to 99.9:0.1 A:B, and finally a 0.3-min hold at 99.9:0.1 A:B. The elution of the compounds was monitored at 210 nm and 230 nm. The ester eluted at 0.56 min, the amide eluted at 0.52 min, and the acid by-product of the reaction eluted as a narrow peak near the solvent front at 0.14 min.

[0177] UPLC method for determining the conversion of the ester substrate of formula I to the amide of formula III:The enzymatic conversion of the ester substrate of formula I to the amide of formula III was determined using an Agilent 1290 UPLC equipped with an Agilent Zorbax SB-C18 RRHD column (3.0×50 mm, 1.8 μm) at a flow rate of 2 mL / min and a column temperature of 60 °C, using a gradient of 0.05% TFA in water (mobile phase A) and 0.05% TFA in acetonitrile (mobile phase B). Starting with a 99.9:0.10 ratio of A:B, the method followed a 0.25 min hold, then a 0.25 min gradient to 80:20 A:B, then a 0.1 min gradient to 100:0 A:B, then a 0.1 min hold, then a 0.1 min gradient to 99.9:0.1 A:B, then a 0.2 min hold. Elution of the compounds was monitored at 210 nm and 230 nm. The ester eluted at 0.53 min, the amide eluted at 0.23 min, and the acid by-product of the reaction eluted as a narrow peak near the solvent front at 0.2 min, with an injection volume of 1 μL.

[0178] UPLC method for determining the conversion of the ester substrate of formula I to the amide of formula V: The enzymatic conversion of the ester substrate of formula I to the amide of formula V was determined using an Agilent 1290 UPLC equipped with an Agilent Zorbax SB-C18 column (3.0×50 mm, 1.8 μm) at a flow rate of 1.5 mL / min and a column temperature of 60 °C, using a gradient of 0.05% TFA in water (mobile phase A) and 0.05% TFA in acetonitrile (mobile phase B). Starting with an 80:20 ratio of A:B, the method followed a 0.9 min hold, then a 0.1 min gradient to 0:100 A:B, then a 0.1 min gradient to 80:20 A:B, then a 0.4 min hold. Elution of the compounds was monitored at 210 nm and 230 nm. The ester eluted at 0.5 min, the amide eluted at 0.39 min, and the acid by-product of the reaction eluted as a narrow peak near the solvent front at 0.14 min.

[0179] Example 5: High-throughput screening for identifying mutants of A. acidicaldarius esterase 2 capable of converting ester substrates to amides The gene encoding A. acidocaldarius esterase 2 (SEQ ID NO: 2) constructed as described in Example 1 was mutagenized using the following method, and a suitable E. coli host strain was transformed with the resulting population of altered DNA molecules. Antibiotic-resistant transformants were selected and processed to identify those expressing a carboxylesterase with improved ability to convert an ester substrate of formula I to a compound of formula (III) and (V), respectively, in the presence of an amine substrate of either formula (II) or (IV). Cell selection, growth, induction of expression of the carboxylesterase mutant enzyme, and collection of cell pellets are described below.

[0180] Recombinant E. coli colonies having the gene encoding carboxylesterase were picked into 96-well shallow-bottom microtiter plates containing 180 μL of LB culture medium, 1% glucose, and 30 μg / mL chloramphenicol (CAM) per well using a Q-PIX molecular devices robotic colony picker (Genetix USA, Inc., Boston, MA). The cells were grown overnight at 30 °C with shaking at 200 rpm. Next, a 20 μL aliquot of this culture was transferred to 96-well deep-bottom plates containing 380 μL of TB culture medium and 30 μg / mL CAM. After incubating these deep-bottom plates at 30 °C for 2 - 3 h with shaking at 250 rpm, expression of the recombinant gene in the cultured cells was induced by adding IPTG to a final concentration of 1 mM. These plates were then incubated at 30 °C for 18 h with shaking at 250 rpm.

[0181] The cells were pelleted by centrifugation (3738 RCF, 10 minutes, 4 °C), resuspended in 200 μL of lysis buffer, and lysed by shaking at room temperature for 2 hours. For the lyophilization screening conditions, the lysis buffer contained 100 mM sodium sulfate (14.2 g / L), 1 mg / mL lysozyme, 500 μg / mL polymyxin B sulfate (PMBS), and 12.5 U / mL benzonase. For the aqueous screening conditions, the lysis buffer contained 10 mM potassium phosphate, pH 7.0, 1 mg / mL lysozyme, and 500 μg / mL PMBS. After sealing these plates with an air-permeable nylon seal, they were shaken vigorously at room temperature for 2 hours. The cell debris was pelleted by centrifugation (3738 RCF, 10 minutes, 4 °C), and the clear supernatant was assayed directly and stored at 4 °C until use.

[0182] In the screening under semi-aqueous conditions using the engineered carboxylesterase at the initial stage, 120 μL aliquots of the substrate solution (720 mL / L DMSO, 90 mL / L 200-proof ethanol, 19.67 mL / L isopropyl-piperazine, 23.5 mL / L ester substrate, and 8.5 mL / L 6N HCl) were added to each well of a Costar deep-well plate, and then 80 μL of the recovered lysate supernatant was added using a Biomek FX robotic instrument (Beckman Coulter, Fullerton, CA). A solution with a final pH of 9.0 was obtained, containing 100 mM ester substrate, 100 mM isopropyl-piperazine, 45% DMSO, and 5% EtOH. After heat-sealing these plates with an aluminum / polypropylene laminate heat-seal tape at 165 °C for 4 seconds, they were shaken at 50 °C overnight (at least 16 hours). The reaction was quenched by adding 200 μL of acetonitrile using Biomex FX. The plates were resealed, shaken for 5 minutes, and then centrifuged at 3738 RCF for 10 minutes. 20 μL of the substrate sample was transferred to a shallow-well polypropylene plate (Costar #3365) containing 180 μL of 75% acetonitrile in water, sealed, shaken for 10 minutes, and then analyzed as described in Example 4.

[0183] In the screening under organic conditions using the initially engineered carboxylesterase, 150 μL aliquots of the recovered lysate supernatant were added to an aluminum 96-well rack (F158359, Unchained Labs, Pleasanton, CA) with a 1 mL glass vial insert (S11168, Unchained Labs, Pleasanton, CA) inserted. Next, this device was lyophilized, gently warmed to room temperature, and 4 μL of distilled water was added using a Multidrop Combi Reagent Dispenser (Thermo Scientific, Waltham, MA), followed by 200 μL of an organic substrate solution (10.93 mL / L isopropyl-piperazine, 23.5 mL / L ester substrate) in tert-butyl methyl ether (tBME). Next, these plates were sealed using a metal rack lid (F158424, Unchained Labs, Pleasanton, CA) attached by a Teflon sheet (S11690-2, Unchained Labs, Pleasanton, CA) laminated under two rubber gaskets (S13086, Unchained Labs, Pleasanton, CA) and seven screws (C151943-050, Unchained Labs, Pleasanton, CA). Next, these constructs were incubated at 50 °C overnight (at least 16 hours) with shaking. The reaction was quenched by adding 200 μL of isopropyl alcohol by Biomek FX, sealed, shaken for 10 minutes, and centrifuged at 235 RCF for 2 minutes to sediment the remaining solids. Next, 200 μL of the sample was transferred to a Costar deep-well plate containing 200 μL of isopropanol in each well, heat-sealed at 165 °C for 4 seconds with an aluminum / polypropylene laminate heat-seal tape, shaken for 10 minutes, and then centrifuged at 3738 RCF for 10 minutes. 40 μL of the substrate sample was transferred to a shallow-bottom polypropylene plate (Costar #3365) containing 160 μL of isopropyl alcohol, sealed, shaken for 10 minutes, and then analyzed as described in Example 4.

[0184] For screening under organic conditions using the engineered carboxylesterase at the late stage, 120 μL aliquots of the recovered lysate supernatant were added to an aluminum 96-well rack (F158359, Unchained Labs, Pleasanton, CA) with a 1 mL glass vial insert (S11168, Unchained Labs, Pleasanton, CA) inserted. Next, this device was lyophilized, gently warmed to room temperature, and 200 μL of an organic substrate solution (53.2 mL / L isopropyl-piperazine, 43 mL / L ester substrate) in methyl isobutyl ester (MIBK) was added, followed by 4 μL of distilled water using a Multidrop Combi Reagent Dispenser (Thermo Scientific, Waltham, MA). Next, these plates were sealed using a metal rack lid (F158424, Unchained Labs, Pleasanton, CA) attached by a Teflon sheet (S11690-2, Unchained Labs, Pleasanton, CA) laminated under one rubber gasket (S13086, Unchained Labs, Pleasanton, CA) and five screws (C151943-050, Unchained Labs, Pleasanton, CA). Next, these constructs were incubated at 15 °C overnight (at least 16 hours) with shaking. The reaction products were removed from the incubator and centrifuged at 235 RCF. 20 μL of the supernatant sample was transferred to a shallow-bottom polypropylene plate (Costar#3365) containing 180 μL of isopropanol (containing 5 g / L naphthalene) per well, heat-sealed at 165 °C for 4 seconds with an aluminum / polypropylene laminate heat-seal tape, shaken for 10 minutes, and then centrifuged at 3738 RCF for 10 minutes. 10 μL of the substrate sample was transferred to a shallow-bottom polypropylene plate (Costar#3365) containing 190 μL of isopropanol, sealed, shaken for 10 minutes, and then analyzed as described in Example 4.

[0185] Example 6: Amidation of the ester substrate of formula I and the amine substrate of formula II in methyl isobutyl ester (MIBK) by an engineered carboxylesterase derived from A. acidicaldarius esterase 2 The improved late-stage carboxylesterase described in Table 3 was evaluated on a preparative scale in MIBK as follows. 10 mg of lyophilized enzyme powder was added to a 1.5 mL HPLC vial together with 15 μL of distilled water. Next, 485 μL of substrate solution (42.93 mL of isopropylpiperazine / L MIBK, 36.4 mL of ester / L MIBK) was added and the vials were sealed. The reaction was shaken at 50 °C and 850 rpm for 16 h on an Eppendorf Thermomixer C heating vial shaker. The reaction was quenched by the addition of 500 μL of isopropanol. Next, 50 μL of the sample was transferred to a shallow-bottom polypropylene plate (Costar #3365) containing 150 μL of isopropanol, shaken for 10 min, and then centrifuged at 3738 RCF for 10 min. 10 μL of the substrate sample was transferred to a shallow-bottom polypropylene plate (Costar #3365) containing 190 μL of isopropanol, sealed, shaken for 10 min, and then analyzed as described in Example 4.

[0186] The improved late-stage carboxylesterase described in Table 3 was evaluated on a preparative scale in MIBK as follows. 11.25 mg of lyophilized enzyme powder was added to a 1.5 mL HPLC vial. Next, 750 μL of substrate solution (53.2 mL of isopropylpiperazine / L MIBK, 43 mL of ester / L MIBK) was added together with 15 μL of distilled water and the vials were sealed. The reaction was shaken at 15 °C and 850 rpm for 16 h on an Eppendorf Thermomixer C heating vial shaker. The reaction was quenched by removing 20 μL as described in the late-stage engineered carboxylesterase screening in Example 5. Table 3 shows the sequence numbers corresponding to the carboxylesterase mutants tested by this method as well as the number of amino acid residue differences from the A. acidocaldarius esterase 2 wild-type carboxylesterase (SEQ ID NO: 2).

[0187] Example 7: Amidation of ester substrate I and amine substrate II in methyl isobutyl ester (MIBK) by an engineered carboxylesterase derived from A. acidicaldarius esterase The following examples illustrate a gram-scale method used to increase the conversion of the compound ethyl oxazole-2-carboxylate of formula I, which is an ester substrate, and the compound 1-isopropylpiperazine of formula II, which is an amine substrate. This method utilizes improved liquid mixing at a large scale to increase the conversion of the substrate to the product. According to relevant monitoring, capture and isolation exceeding 70% of the total yield of the product are possible.

[0188] The large-scale reaction method contains the following reaction components.

Table 4

[0189] Method To a 2 L CLR reactor equipped with an overhead stirrer, 340 mL of MIBK was added and set to stir at ambient temperature. Next, after adding 11 mL of water, it was gently heated to 30 °C. 14 g of carboxylesterase was then added, followed by 20 mL of MIBK wash solution. Next, 23.62 g of 1-isopropylpiperazine was added, followed by 20 mL of MIBK wash solution. Finally, 20 g of ethyl oxazole-2-carboxylate was added, followed by the final 20 mL of MIBK wash solution. After 16 hours, the conversion rate exceeded 90%.

[0190] The remaining solid was filtered and washed with 80 mL of MIBK to isolate the adsorbed material. Then, the wash solutions were pooled, washed with 20 mL of 10% w / w sodium chloride, stirred, and separated. The organic phase was recovered, concentrated to 60 mL under vacuum, and then cooled to 15 °C, at which point crystallization began. 70 mL of n-heptane was added over 10 minutes to allow crystallization to proceed completely over 2 hours, and then the crystalline material was collected by filtration. The crystalline product was washed twice with 80 mL of 1:4 MIBK:n-heptane and dried under vacuum at 50 °C. By this method, a total yield (w / w) of 66% was obtained when evaluated by the mass of the product.

[0191] Example 8: Amidation of ester substrate I and amine substrate IV in methyl isobutyl ester (MIBK) by an engineered carboxylesterase derived from A. acidicaldarius esterase The operationally active carboxylesterase was used to perform a reaction on a vial scale for the production of the amide of formula V. 25 mg of lyophilized enzyme powder was added to a 1.5 mL HPLC vial together with 10 μL of distilled water. Next, 490 μL of substrate solution (26.72 mL of (2S,6R)-2,6-dimethylmorpholine / L MIBK, 24.27 mL of ester / L MIBK) was added, and the vials were sealed. The reaction mixture was shaken at 50 °C and 800 rpm for 16 h on a heated vial shaker. A 100 μL aliquot of each reaction solution was added to a shallow-bottom polypropylene plate (Costar #3365) containing 100 μL of isopropanol, sealed, shaken for 10 min, and then centrifuged at 3738 RCF for 10 min. Next, a 10 μL aliquot of the supernatant was diluted in a shallow-bottom polypropylene plate (Costar #3365) with 190 μL of acetonitrile containing 25% water, sealed, and then shaken at 850 rpm for 5 min. The reaction was analyzed as described in Example 4. In all cases observed, the enzyme activity was in exact agreement with the enzyme activity of the reactions described in Example 6 and Table 3.

[0192] Although various specific embodiments have been illustrated and described, it will be understood that various changes can be made without departing from the spirit and scope of the invention.

[0193] [Table 5] JPEG2025111490000022.jpg239161 JPEG2025111490000023.jpg236161 JPEG2025111490000024.jpg237161 JPEG2025111490000025.jpg231161 JPEG2025111490000026.jpg235161 JPEG2025111490000027.jpg235161 JPEG2025111490000028.jpg236161 JPEG2025111490000029.jpg237161 JPEG2025111490000030.jpg228161 JPEG2025111490000031.jpg240161 JPEG2025111490000032.jpg237161 JPEG2025111490000033.jpg236161 JPEG2025111490000034.jpg236161 JPEG2025111490000035.jpg238161 JPEG2025111490000036.jpg236161 JPEG2025111490000037.jpg237161 JPEG2025111490000038.jpg233161 JPEG2025111490000039.jpg236161 JPEG2025111490000040.jpg233161 JPEG2025111490000041.jpg237161 JPEG2025111490000042.jpg234161 JPEG2025111490000043.jpg235161 JPEG2025111490000044.jpg234161 JPEG2025111490000045.jpg233161 JPEG2025111490000046.jpg233161 JPEG2025111490000047.jpg235161 JPEG2025111490000048.jpg234161 JPEG2025111490000049.jpg233161 JPEG2025111490000050.jpg231161 JPEG2025111490000051.jpg234161

Claims

1. A carboxylesterase polypeptide comprising an amino acid sequence that is at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical to the amino acid sequence shown in SEQ ID NO: 4, or more than that, or a functional fragment thereof, wherein the amino acid sequence of the carboxylesterase polypeptide includes the feature that the residue corresponding to X198 in SEQ ID NO: 4 is selected from nonpolar residues, aromatic residues, and aliphatic residues.

2. The carboxylesterase polypeptide according to claim 1, wherein the amino acid sequence of the carboxylesterase polypeptide includes the feature that X198 is selected from F, L, I, Y, and M.

3. The carboxylesterase polypeptide according to claim 1 or 2, wherein the amino acid sequence of the carboxylesterase polypeptide includes residues different from the amino acid sequence shown in SEQ ID NO: 4 at at least one residue position selected from X27, X30, X35, X37, X57, X75, X103, X185, X207, X208, X271, X286, and X296.

4. The amino acid sequence of the carboxylesterase polypeptide is such that the residue corresponding to X27 is a constrained residue; the residue corresponding to X30 is an aliphatic residue; the residue corresponding to X35 is selected from basic residues and polar residues; the residue corresponding to X37 is selected from aliphatic residues and polar residues; the residue corresponding to X57 is a nonpolar residue; the residue corresponding to X75 is selected from basic residues and polar residues; the residue corresponding to X103 is selected from aliphatic residues and aromatic residues; the residue corresponding to X185 is selected from nonpolar residues, aliphatic residues, and aromatic residues; the residue corresponding to X207 is selected from acidic residues and polar residues; the residue corresponding to X208 is selected from aliphatic residues, basic residues, and polar residues; the residue corresponding to X271 is selected from acidic residues and polar residues; the residue corresponding to X286 is selected from nonpolar residues, aliphatic residues, polar residues, and small residues; and the residue corresponding to X296 is selected from aliphatic residues and basic residues The carboxylesterase polypeptide according to claim 1, 2, or 3, comprising at least one feature selected from

5. The carboxylesterase polypeptide according to claim 3 or 4, wherein the amino acid sequence of the carboxylesterase polypeptide comprises at least one feature selected from: X27 is P; X30 is selected from I, L, and V; X35 is H; X37 is selected from I, L, T, and V; X57 is M; X75 is R; X103 is selected from F, M, and W; X185 is selected from F, I, and M; X207 is E; X208 is selected from R, L, and H; X271 is D; X286 is selected from M, V, and G; and X296 is selected from V, L, and R.

6. The carboxylesterase polypeptide according to claim 3 or 4, wherein the amino acid sequence of the carboxylesterase polypeptide comprises the following features: X35 is selected from basic residues and polar residues; and X185 is selected from polar residues and aliphatic residues.

7. The carboxylesterase polypeptide according to claim 3, 4, 5, or 6, wherein the amino acid sequence of the carboxylesterase polypeptide comprises the following features: X35 is H; and X185 is selected from F, I, and M.

8. The carboxylesterase polypeptide according to claim 1, wherein the amino acid sequence of the carboxylesterase polypeptide comprises a residue different from the amino acid sequence shown in SEQ ID NO: 24 at at least one residue position selected from X9, X19, X34, X35, X37, X46, X48, X66, X87, X103, X139, X190, X207, X216, X263, X271, X278, and X296.

9. The amino acid sequence of the carboxylesterase polypeptide has the following features: The residue corresponding to X9 is an aromatic residue; The residue corresponding to X19 is selected from basic residues and polar residues; The residue corresponding to X34 is selected from constrained residues, acidic residues, and polar residues; The residue corresponding to X35 is selected from polar residues; The residue corresponding to X46 is an aliphatic residue; the residue corresponding to X48 is an aliphatic residue; The residue corresponding to X66 is an aliphatic residue; The residue corresponding to X87 is selected from aliphatic residues and small residues; The residue corresponding to X103 is selected from aromatic residues; The residue corresponding to X139 is a basic residue; the residue corresponding to X190 is an aromatic residue; The residue corresponding to X207 is a basic residue; The residue corresponding to X216 is selected from aromatic residues, basic residues, and polar residues; The residue corresponding to X263 is selected from aliphatic residues and polar residues; the residue corresponding to X271 is selected from acidic residues and polar residues; The residue corresponding to X278 is selected from aliphatic residues and aromatic residues; and the residue corresponding to X296 is selected from aliphatic residues and basic residues The carboxylesterase polypeptide according to claim 8, comprising at least one of the above.

10. The amino acid sequence of the carboxylesterase polypeptide has the following characteristics: X9 is Y; X19 is R; X34 is selected from E, N, and P; X35 is S; X37 is T; X46 is selected from I, L, and V; X48 is L; X66 is V; X87 is A; X103 is selected from W and F; X139 is R; X190 is Y; X207 is E; X216 is selected from N and W; X263 is selected from T and A; X271 is D; X278 is selected from W and L; and X296 is selected from V, L, and R. The carboxylesterase polypeptide according to claim 8 or 9, comprising at least one of the above.

11. The amino acid sequence of the carboxylesterase polypeptide has the following characteristics: X9 is an aromatic residue, and X87 is an aliphatic residue. The carboxylesterase polypeptide according to claim 8 or 9.

12. The amino acid sequence of the carboxylesterase polypeptide has the following characteristics: X9 is Y; and X87 is A. The carboxylesterase polypeptide according to claim 8, 9, 10, or 11.

13. The carboxylesterase polypeptide according to claim 1, wherein the amino acid sequence of the carboxylesterase polypeptide contains a residue different from the amino acid sequence represented by SEQ ID NO: 54 at at least one residue position selected from X20, X28, X29, X30, X33, X34, X188, X216, and X286.

14. The amino acid sequence of the carboxylesterase polypeptide is such that the residue corresponding to X20 is selected from an aliphatic residue and a basic residue; the residue corresponding to X28 is selected from an acidic residue, a polar residue, and a restricted residue; the residue corresponding to X29 is selected from an acidic residue and a polar residue; the residue corresponding to X30 is an aliphatic residue; the residue corresponding to X33 is an aromatic residue; the residue corresponding to X34 is a small residue; the residue corresponding to X188 is selected from a small residue and an aromatic residue; the residue corresponding to X216 is a polar residue; and the residue corresponding to X286 is selected from an aliphatic residue, a small residue, a nonpolar residue, and a polar residue The carboxylesterase polypeptide according to claim 13, comprising at least one feature selected from the above.

15. The carboxylesterase polypeptide according to claim 12, 13, or 14, wherein the amino acid sequence of the carboxylesterase polypeptide contains at least one feature selected from the following: X20 is selected from I and R; X28 is selected from D, P, and S; X29 is D; X30 is V; X33 is W; X34 is G; X188 is selected from G and F; X216 is N; and X286 is selected from S, M, V, G, and A.

16. The carboxylesterase polypeptide according to claim 13 or 14, wherein the amino acid sequence of the carboxylesterase polypeptide contains the following feature: X216 is a polar residue.

17. The carboxylesterase polypeptide according to claim 13, 14, 15, or 16, wherein the amino acid sequence of the carboxylesterase polypeptide contains the following feature: X216 is N.

18. The carboxylesterase polypeptide according to claim 1, wherein the amino acid sequence of the carboxylesterase polypeptide comprises a residue different from the amino acid sequence represented by SEQ ID NO: 68 at at least one residue position selected from X10, X20, X22, X28, X30, X33, X36, X37, X46, X66, X75, X103, X197, X263, X266, X280, and X290.

19. The amino acid sequence of the carboxylesterase polypeptide is such that the residue corresponding to X10 is an aliphatic residue; the residue corresponding to X20 is selected from an aliphatic residue and a basic residue; the residue corresponding to X22 is an aromatic residue; the residue corresponding to X28 is selected from an acidic residue, a polar residue, and a constrained residue; the residue corresponding to X30 is an aliphatic residue; the residue corresponding to X33 is an aromatic residue; the residue corresponding to X36 is an aliphatic residue or an aromatic residue; the residue corresponding to X37 is an aromatic residue or a small residue; the residue corresponding to X46 is a basic residue; the residue corresponding to X66 is a polar residue; the residue corresponding to X75 is a basic residue; the residue corresponding to X103 is an aromatic residue; the residue corresponding to X197 is an aliphatic residue; the residue corresponding to X263 is a basic residue; the residue corresponding to X266 is a polar residue; the residue corresponding to X280 is selected from an aliphatic residue and a polar residue; and the residue corresponding to X290 is selected from an aliphatic residue and an aromatic residue The carboxylesterase polypeptide according to claim 18, comprising at least one feature selected from the above.

20. The amino acid sequence of the carboxylesterase polypeptide is selected such that X10 is selected from L and M; X20 is selected from I and R; X22 is W; X28 is selected from D, P, and S; X30 is V; X33 is W; X36 is selected from F, I, and M; X37 is selected from G and Y; X46 is R; X66 is T; X75 is R; X103 is W; X197 is L; X263 is R; X266 is T; X280 is selected from M and T; and X290 is selected from W and I, and includes at least one feature selected therefrom, the carboxylesterase polypeptide according to claim 18 or 19.

21. The amino acid sequence of the carboxylesterase polypeptide includes at least one feature selected such that X30 is an aliphatic residue, X33 is an aromatic residue, X75 is a basic residue, and X103 is an aromatic residue, the carboxylesterase polypeptide according to claim 18 or 19.

22. The amino acid sequence of the carboxylesterase polypeptide includes the following features: X30 is V; X33 is W; X75 is R; and X103 is W, the carboxylesterase polypeptide according to claim 18, 19, 20, or 21.

23. The amino acid sequence of the carboxylesterase polypeptide includes a residue different from the amino acid sequence shown in SEQ ID NO: 68 at at least one residue position selected from X28, X38, X46, X54, X66, X75, X85, X86, X96, X160, X176, X183, X188, X205, X212, X248, X249, X255, X270, and X286, the carboxylesterase polypeptide according to claim 1.

24. The amino acid sequence of the carboxylesterase polypeptide is such that the residue corresponding to X28 is selected from acidic residues, polar residues, small residues, and constrained residues; the residue corresponding to X38 is selected from aliphatic residues and basic residues; the residue corresponding to X46 is selected from acidic residues and basic residues; the residue corresponding to X54 is selected from acidic residues and polar residues; The residue corresponding to X66 is a polar residue; The residue corresponding to X75 is a basic residue; The residue corresponding to X85 is selected from an aromatic residue or a basic residue and a small residue; The residue corresponding to X86 is a polar residue; the residue corresponding to X96 is selected from a non-polar residue and an aliphatic residue; The residue corresponding to X160 is selected from a polar residue and a constrained residue; The residue corresponding to X176 is selected from an aliphatic residue, an aromatic residue or a basic residue and a non-polar residue; the residue corresponding to X183 is a non-polar residue; The residue corresponding to X188 is selected from an aromatic residue and a small residue; The residue corresponding to X205 is an aromatic residue; the residue corresponding to X212 is an acidic residue; The residue corresponding to X248 is an aliphatic residue; The residue corresponding to X249 is an aromatic residue; The residue corresponding to X255 is a polar residue; The residue corresponding to X270 is selected from an aliphatic residue and a polar residue; and The residue corresponding to X286 is selected from an aliphatic residue, a non-polar residue, a small residue and a polar residue The carboxylesterase polypeptide according to claim 23, comprising at least one feature selected therefrom.

25. The amino acid sequence of the carboxylesterase polypeptide is such that X28 is selected from C, D, S, H, P, G and R; X38 is selected from E and L; X46 is selected from K, R and Q; X54 is selected from R, Q, and S; X66 is selected from L, T and V; X75 is R; X85 is selected from G and H; X86 is T; X96 is selected from M and L; X160 is selected from T and P; X176 is selected from M, L and H; X183 is Q; X188 is selected from G and F; X205 is F; X212 is D; X248 is V; X249 is W; X255 is N; X270 is selected from N and L; and X286 is selected from M, V, G, N and S. The carboxylesterase polypeptide according to claim 23 or 24, comprising at least one feature selected therefrom.

26. The carboxylesterase polypeptide according to claim 23, 24, or 25, comprising an amino acid sequence comprising at least one feature selected from the group consisting of X28 being a polar residue, X38 being a basic residue, and X85 being a small residue.

27. The carboxylesterase polypeptide according to claim 23, 24, 25, or 26, comprising an amino acid sequence comprising the following features: X28 is C; X38 is E; and X85 is G.

28. A carboxylesterase polypeptide comprising an amino acid sequence comprising a residue different from the amino acid sequence represented by SEQ ID NO: 100 at one residue position selected from X7, X22, X36, X38, X46, X54, X66, and X75.

29. The residue corresponding to X7 is an aliphatic residue; the residue corresponding to X22 is selected from an aliphatic residue and an aromatic residue; the residue corresponding to X36 is selected from a polar residue and a nonpolar residue; the residue corresponding to X38 is an aromatic residue; the residue corresponding to X46 is selected from a polar residue and a basic residue; the residue corresponding to X54 is selected from a polar residue and a basic residue; the residue corresponding to X66 is a polar residue; and the residue corresponding to X75 is selected from a basic residue and a nonpolar residue The carboxylesterase polypeptide according to claim 28, comprising an amino acid sequence comprising at least one feature selected therefrom.

30. The carboxylesterase polypeptide according to claim 28 or 29, wherein the amino acid sequence comprises at least one feature selected from the group consisting of X7 being L; X22 being selected from W and L; X36 being selected from T and M; X38 being W; X46 being selected from K and Q; X54 being selected from S, Q, and K; X66 being selected from G and T; and X75 being selected from M and R.

31. The carboxylesterase polypeptide according to claim 28, 29, or 30, comprising an amino acid sequence comprising at least one feature selected from the group consisting of X36 being a polar residue, X38 being an aromatic residue, and X75 being a basic residue. In yet another embodiment, the improved carboxylesterase polypeptide comprises an amino acid sequence having the following characteristics: X36 is T; X38 is W; and X75 is R.

32. The carboxylesterase polypeptide according to claim 1, comprising an amino acid sequence comprising a residue different from the amino acid sequence represented by SEQ ID NO: 114 at one residue position selected from X2, X181, and X286.

33. The carboxylesterase polypeptide according to claim 32, comprising an amino acid sequence comprising at least one characteristic selected from the group consisting of: the residue corresponding to X2 is selected from an aliphatic residue, a basic residue, a polar residue, and an aromatic residue; the residue corresponding to X181 is a basic residue; and the residue corresponding to X286 is selected from a polar residue and a nonpolar residue.

34. The carboxylesterase polypeptide according to claim 32 or 33, wherein the amino acid sequence comprises at least one characteristic selected from the group consisting of: X2 is selected from L, Q, R, and H; X181 is Q; and X286 is selected from C and S.

35. The carboxylesterase polypeptide according to claim 32, 33, or 34, comprising an amino acid sequence comprising at least one characteristic selected from the group consisting of: X286 is a nonpolar residue.

36. The carboxylesterase polypeptide according to claim 32, 33, 34, or 35, comprising an amino acid sequence having the following characteristic: X286 is C.

37. The carboxylesterase polypeptide according to claim 1, wherein the amino acid sequence corresponds to the amino acid sequence represented by any one of SEQ ID NOs: 122, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 124, or 126.

38. A carboxylesterase polypeptide comprising the amino acid sequence represented by SEQ ID NO:

122.

39. A polynucleotide encoding the polypeptide according to any one of claims 1 to 38.

40. A composition comprising at least one engineered carboxylesterase according to any one of claims 1 to 39.

41. The polynucleotide according to claim 39, corresponding to any one of the nucleotide sequences represented by SEQ ID NO: 121, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, 115, 117, 119, 123, or 125.

42. A polynucleotide sequence encoding the carboxylesterase polypeptide sequence represented by SEQ ID NO:

122.

43. A polynucleotide encoding a carboxylesterase polypeptide comprising the polynucleotide sequence represented by SEQ ID NO:

121.

44. A method for producing an amide, comprising: (a) R 1 -COOR 2 [R 1 has 0 to 3 alkyl substituents and is selected from sp 3 carbon and aromatic rings, and R 2 is a methyl group; an ethyl group; and is selected from 1-6 carbon alkyl chains] ester substrate of the structure; (b) an amine substrate; (c) a carboxylesterase polypeptide according to any one of claims 1 to 38; and (d) a solvent combining the components containing them.

45. The method according to claim 44, wherein the carboxylesterase polypeptide in step (c) is produced in the presence of a salt.

46. The method according to claim 45, wherein the salt is selected from sodium sulfate; potassium sulfate; lithium sulfate; sodium phosphate; and potassium phosphate.

47. The method according to claim 44, wherein the solvent is an organic solvent selected from toluene; 2-methyltetrahydrofuran; tetrahydrofuran; dimethylacetamide; methyl isobutyl ketone (MIBK); dichloromethane; tert-butyl methyl ether; cyclopentyl methyl ether; methylcyclohexane; dichloromethane; acetonitrile; methyl ethyl ketone; isopropyl acetate; ethanol; isopropanol; ethyl acetate; heptane; xanthane; and 2-methyltetrahydrofuran (2-Me-THF); and water.

48. The method according to claim 44, 45, 46, or 47, wherein the organic solvent contains up to 3 molar equivalents of water relative to the ester substrate in an amount of about 0.5% (vol / vol) to about 3% (vol / vol).

49. The ester substrate has the formula: 【Chemical 1】 Ethyl oxazole-5-carboxylate having, and the amine substrate has the formula: [Chemical Formula 2] 1-Isopropylpiperazine having, and the amide has the formula: 【Chemical Formula 3】 (4-Isopropylpiperazin-1-yl)(oxazol-5-yl)methanone having, the method according to claim 44, 45, 46, 47, or 48.

50. The ester substrate has the formula: 【Chemical Formula 4】 Ethyl oxazole-5-carboxylate having, and the amine substrate has the formula: [Chemical Formula 5] Cis-2,6-dimethylmorpholine having, and the amide has the formula: 【Chemical Formula 6】 ((2S,6R)-2,6-Dimethylmorpholino)(oxazol-5-yl)methanone having, the method according to claim 44, 45, 46, 47, 48, or 49.

51. The reaction comprises about 40 g / L of ethyl oxazole-5-carboxylate, about 44 g / L of 1-isopropylpiperazine, and about 20 g / L of a carboxylesterase polypeptide corresponding to an amino acid sequence selected from SEQ ID NOs: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, and 126, and the carboxylesterase polypeptide is produced in the presence of sodium sulfate and acts in the presence of about 10 g / L to about 20 g / L of water in MIBK, the method according to claim 44.

52. An amide produced by the method according to claim 44, 45, 46, 47, 48, 49, 50, or 51.

Citation Information

Patent Citations

  • Carboxylesterase Biocatalyst

    JP7401435B2

  • Hydrolases, nucleic acids encoding them and methods for improving paper strength

    WO2006096834A2

  • Esterases and related nucleic acids and methods

    WO2007092314A2