Engineered carboxylesterase enzymes and methods for their use in macrocyclization of non-canonical tetrapeptides

Engineered carboxylesterases address the limitations of chemical and enzymatic methods by providing high-yield, regioselective macrocyclization of oligopeptides, particularly tetrapeptides, suitable for large-scale production of therapeutic peptides.

WO2026019621A1PCT designated stage Publication Date: 2026-01-22MERCK SHARP & DOHME LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/037058
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-15
Filing Date
2025-07-10
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing methods for synthesizing macrocyclic peptides face challenges such as low specificity, formation of impurities and undesired byproducts, and inefficiencies in chemical approaches, while enzymatic methods struggle with scalability, robustness, and tolerance to process conditions.

Method used

Engineered carboxylesterases with improved activity, enantioselectivity, and reduced product hydrolysis are developed through directed evolution, enabling the biocatalytic synthesis of macrocyclic peptides with high yields and regioselectivity.

Benefits of technology

The engineered carboxylesterases achieve efficient macrocyclization of oligopeptides, including tetrapeptides with non-canonical amino acids, demonstrating enhanced activity, thermostability, and solvent tolerance, suitable for large-scale production of therapeutic peptides.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025037058_22012026_PF_FP_ABST
    Figure US2025037058_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides engineered carboxylesterase enzymes having improved activity in catalyzing cyclization of short polypeptides, as compared to a naturally occurring wild-type carboxylesterase enzyme. These variants catalyze the cyclization of linear tetrapeptides with high regioselectivity. Also provided are polynucleotides and expression vectors encoding the carboxylesterase enzymes and methods of using the carboxylesterase enzymes to generate cyclic tetrapeptides containing non-canonical amino acids.
Need to check novelty before this filing date? Find Prior Art

Description

ENGINEERED CARBOXYLESTERASE ENZYMES AND METHODS FOR THEIR USEIN MACROCYCLIZATION OF NON-CANONICAL TETRAPEPTIDESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 671,563 filed July 15, 2024, the entire contents of which are incorporated by reference herein.REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY

[0002] The contents of the electronic sequence listing (25920-WO-PCT_SL.xmh Size: 69,000 bytes; and Date of Creation: April 4, 2025) are herein incorporated by reference in their entirety.FIELD OF THE INVENTION

[0003] The present invention relates to engineered carboxylesterases, useful in biocatalytic and synthetic processes through regiospecific amide bond formation, including for the generation of macrocyclic tetrapeptides. Such enzymes may be particularly useful in synthetic processes that may be used as part of the preparation of isopropyl ((1 lS,12S,13S,95,12S)-9-amino-12-((l-(6- aminohexyl)-5-fluoro- l / 7-indol-3-yl)methyl)-4, 10, 13-trioxo-2-oxa-5, 11 -di aza- 1(3,1 )-py rrolidina- 7(l,3)-benzenacyclotridecaphane-12-carbonyl)-L-threoninate.BACKGROUND OF THE INVENTION

[0004] Therapeutic peptides are gaining increasing interest from the pharmaceutical industry due to their potential to specifically target proteins through surface interactions which could be distinct from traditional small molecule modalities. Therapeutic peptides, however, suffer from limitations such as susceptibility7to proteolytic cleavage and / or low cell permeability7, which leads to lower clinical efficacy. Macrocyclization of linear peptides through ligation between N and C termini or sidechains has become a promising strategy to address these challenges. Macrocyclic peptides were shown to be more stable, resistant to proteolytic cleavage, and permeable through cell membrane, making them attractive drug candidates with advantageous properties from both small-molecule and peptide drugs. See Vinogradov et al., J. Am. Chem. Soc. 2019, 141, 10, 4167-4181 for a complete review.

[0005] The synthesis of macrocyclic peptides, however, is difficult to control through traditional chemical methods of amide formation, resulting in lack of specificity and formation of impurities and undesired byproducts. For macrocyclization, an entropically7unfavored precyclization confirmation must be formed before the desired cyclization can occur. In addition,attempts to macrocyclize oligopeptides can result in a range of bond configurations including head-to-tail, head-to-sidechain, sidechain-to-sidechain, and tail-to-sidechain. Given the multiple chemical approaches from cross-coupling and photochemistry that have been explored for chemoselective macrocyclization of peptides, enzyme-mediated macrocyclization has been of great interest due to its potential efficiency and high regio- and chemo-selectivity. Moreover, enzymatic reactions are often performed under mild conditions and eliminate the need for toxic solvent, holding economic and environmental advantages over traditional chemical synthesis.

[0006] The recent development of biocatalysts has provided a promising solution for the increasing need for a more sustainable manufacturing of chemicals and medicine. However, using enzymes at industrial scale applications has often faced challenges such as scalability, robustness, and tolerance to relevant process conditions, which could be addressed by protein engineering.

[0007] Therefore, there is a need in the art for enzymes that can cyclize oligopeptides at scale while maintaining desired regiochemistry, reduced by-product formation and high yields.SUMMARY OF THE INVENTION

[0008] The present disclosure provides, inter alia, engineered polypeptides (e.g., carboxylase enzymes) capable of catalyzing a macrocyclization through amide bond formation in oligopeptides. The present disclosure further provides polynucleotides encoding these engineered polypeptides and expression vectors containing these polynucleotides, and host cells containing these expression vectors. Further disclosed are methods of producing these engineered polypeptides, and methods of synthesizing cyclic oligopeptides using these polypeptides. The carboxylesterase enzymes described herein are capable of macrocyclizing oligopeptides, such as linear tetrapeptides. Such enzymes may be useful in the preparation of intermediates in processes to generate complex biological compounds, such as macrocyclic peptides.

[0009] A unique family of hydrolases with promiscuous acyltransferase activities was recently identified. Members from this family, i.e.. family VIII carboxylesterases, have performed efficient catalysis of amide formation on multiple substrates over a hydrolysis reaction, suggesting their potential for enzymatic peptide macrocyclization. Chemical synthesis of macrocyclic peptides has long posed challenges due to the complexity of the molecules and potentially high cost. In contrast, biocatalysis, which harnesses the power of enzymes, offers an appealing alternative. Enzymes can facilitate reactions under environmentally-friendly conditions, leading to high yields and specificity. Nonetheless, use of hydrolases for theindustrial synthesis of macrocyclic peptides has been limited due to challenges in low activity, poor enantioselectivity, and product hydrolysis of wild-type enzymes.

[0010] The present disclosure provides engineered carboxylesterases having improved activity', enantioselectivity, and reduced product hydrolysis that satisfy needs in the art. The present disclosure is based, at least in part, on the identification, engineering and improvement of multiple enzymes to enable the biocatalytic synthesis of a cyclic peptide containing non- canonical amino acids. Directed evolution and ultra-high-throughput experimentation were applied to explore the sequence space of multiple enzy mes. The diverse approach employed enabled the efficient identification of enzyme variants with enhanced activity, stability, and specificity required for enzymatic cascade synthesis. The Examples demonstrate the effectiveness of integrating different strategies in protein engineering challenges and unlocks new opportunities for the production of therapeutic cyclic peptides.

[0011] The present disclosure provides engineered nucleic acid and protein variants of the wildtype carboxylesterase defined by the sequences of SEQ ID NOs: 1 and 2. respectively, identified from the public genome database (GenBank: MBG69902.1) that exhibit the activity to catalyze macrocyclization of the tetrapeptide, isopropyl ((2S',31S -l-((S)-2-((1S -2-amino-3-(3- (aminomethyl)phenyl)propanamido)-3-(l-(6-aminohexyl)-5-fluoro-l / / -indol-3-yl)propanoyl)-3- (2-(te / 7-butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate, with desired regioselectivity. Through iterative rounds of direction evolution of these sequences, enzy mes were engineered with improved activity and more than 100-fold increased tolerance to organic co-solvent relative to the wild-type carboxylesterase. The engineered enzy mes exhibit an increased thermostability and avoid the undesired hydrolysis of substrate by more than 10-fold. These properties enable the biocatalytic synthesis of isopropyl ((1 LS'.125.135,9N12>S')-9-ammo- 12-((l-(6-aminohexyl)-5-fluoro-l / f-indol-3-yl)methyl)-4,10,13-trioxo-2-oxa-5,l l-diaza-l(3,l)- pyrrolidina-7(l,3)-benzenacyclotridecaphane-12-carbonyl)-L-threoninate at scale.

[0012] The wild-type carboxylesterase enzyme (SEQ ID NO: 2) was identified as a starting point to catalyze macrocyclization of a tetrapeptide of pharmaceutical interest; however, it exhibited extremely low activity and generated undesired hydrolysis products. These effects were more pronounced in the presence of co-solvent acetonitrile. Low activity and selectivity, along with low tolerance to organic co-solvent, limits any practical applications for this enzyme. Through rounds of protein engineering, improved carboxylesterase variants, including the carboxylesteraseRMBB variant having the polypeptide sequence of SEQ ID NO: 28, have been developed with the improved activity, tolerance to organic co-solvent, thermostability, and reduced hydrolysis by-product formation for the conversion of isopropyl ((2<S',3S)-l-((S)-2-((lS)-2-amino-3-(3-(aminomethyl)phenyl)propanamido)-3-(l-(6-aminohexyl)-5-fluoro-l / / -indol-3- yl)propanoyl)-3-(2-(ter / -butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate to isopropyl ((1 lS,12iS',13S',9S,121S)-9-amino-12-((l-(6-aminohexyl)-5-fluoro-177-indol-3-yl)methyl)-4,10,13- trioxo-2-oxa-5.11 -diaza- 1(3,1 )-py rrolidina-7 ( 1 ,3)-benzenacy clotridecaphane- 12-carbony 1)-L- threoninate. The later compound is a macrocyclic tetrapeptide having non-canonical (or modified) amino acids.

[0013] To enhance desired properties, 32 mutations were introduced into the amino acid sequence of SEQ ID NO: 2 during directed evolution to yield the carboxylesteraseR MBB (SEQ ID NO: 28) selected after the fourteenth round of evolution. The incorporation of these mutations improved the enzyme in multiple parameters, including activity (>1000-fold increase relative to SEQ ID NO: 2), tolerance to organic solvent (>80-fold increase), thermostability (>20-fold increase), and desired product to hydrolysis product ratio (>10-fold increase). These characteristics are all crucial for successful large-scale manufacturing of macrocyclic peptides.

[0014] Thus, in some aspects, provided herein are engineered polypeptides comprising an amino acid sequence having at least 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity’ to any one of the amino acid sequences disclosed in Table 15. In one aspect, provided are engineered polypeptides comprising an amino acid sequence having at least 92.5%, 95%, 96%, 98%. or 99% sequence identity to any one of SEQ ID NOs: 22, 24, 26, and 28.

[0015] In some embodiments, the polypeptides comprise an amino acid sequence that comprises 100 or more consecutive amino acids in common with any one of SEQ ID NOs: 22, 24, 26, and 28, such as a stretch of at least 300, 325, 350, 375, or 400 consecutive amino acids of any one of SEQ ID NOs: 22. 24. 26. and 28.

[0016] In some embodiments, the disclosed polypeptides further comprise one or more affinity tags, such as a hexa-histidine tag.

[0017] In some embodiments, the polypeptides exhibit activity in catalyzing a macrocyclization of a linear tetrapeptide. These polypeptides may be carboxylesterases known as a ‘‘macrocyclase” or “NF-macrocyclase.” These polypeptides may exhibit activity in catalyzing a reaction in which Compound (1)is converted to Compound (2)The disclosed polypeptides may exhibit greater or improved activity, regioselectivity, co-solvent tolerance, and / or thermostability in catalyzing the above reaction relative to a polypeptide having the amino acid sequence of SEQ ID NO: 2.

[0018] In some aspects, provided herein are expression vectors comprising any of the disclosed polynucleotides operably linked to one or more control sequences suitable for directing expression of the encoded polypeptide in a host cell. The control sequence may comprise a promoter, such as an E. coll promoter.

[0019] Further provided herein are host cells comprising any of the disclosed expression vectors and / or polynucleotides. In some embodiments, the host cell is E. coli.

[0020] Further provided are methods of producing a polypeptide comprising culturing any of the disclosed host cells, e g., under conditions and for a time suitable for expression of the polypeptide, and optionally, recovering the polypeptide and / or isolating the polypeptide.

[0021] In additional aspects, processes are described for preparing the disclosed engineered carboxylesterase enzymes and processes for using them to catalyze macrocyclizations and synthesis schemes involving such macrocyclizations. Because the disclosed enzymes are evolved using particular selection pressures such as regioselectivity, reduced by-product formation, and thermostability, the disclosed enzy mes exhibit these properties in a manner such that the disclosed synthetic processes exhibit these properties.

[0022] Other embodiments, aspects and features of the present invention are either further described in or will be apparent from the ensuing description, examples and appended claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] FIG. 1 shows that an exemplary macrocyclase variant was expressed robustly during production of shake flask powders. The enzyme having the amino acid sequence of SEQ ID NO: 10 was over-expressed compared to other host cell proteins under different inducer (IPTG) concentrations.

[0024] FIG. 2 shows limited expression of the macrocyclase under 5-liter (5L) fermentation with complex medium. Over-expression of the enzyme of SEQ ID NO: 10 was not observed under 5L fermentation with complex medium at multiple time points after IPTG induction.

[0025] FIG. 3 shows that the macrocyclase over-expresses under 5L fermentation with chemical defined medium (CDM). Over-expression of the enzyme of SEQ ID NO: 10 was observed under 5L fermentation with CDM at multiple time points after IPTG induction.

[0026] FIG. 4A shows a summary of the directed evolution scheme, rounds 1 to 8. The tier positions were as follows: tierl (I159Y, Y154M, N156P, G170W, L172C, L199G, M368A, M154N, G158E, R330S. P160L. I162G. A272H, A165P, T265S, and D266L); tier2 (N132R, K352T, R133M, M327Y, R17Y, F134Q, and P325S); tier3 (T138R, KI 84V, A191D, and S350D); and others (VI 151, T140P, C147R, and F373Y).

[0027] FIG. 4B shows the change of activity of the root enzyme variant of each round of evolution.

[0028] FIG. 4C shows a modeled structure of the wild-type ECE20. Positions of mutations listed in FIG. 4A were highlighted in with light grey circles. Docked product (compound 2) is highlighted in dark grey.

[0029] FIG. 5A shows a summary of evolution through round 9 to 13. The tier positions were as follows: tierl (G199S, L266D, H166W, G281S, A368L, G162I, and H272A); tier2 (Q134F); tier3 (L343F, F343L); and others (T10K).

[0030] FIG. 5B shows the activity (dots) and selectivity' (columns) change of the root enzy me variant of each round of evolution. The activity was in relative comparison to the enzyme having the amino acid sequence of SEQ ID NO: 28.

[0031] FIG. 5C shows a modeled structure of the wild-type ECE20. Positions of mutations listed in FIG. 5A are highlighted in light grey. Docked product (compound 2) is highlighted in dark grey.DETAILED DESCRIPTION

[0032] The present disclosure provides engineered esterase enzymes, and in particular engineered carboxylesterase enzymes. The disclosure provides engineered polypeptides derived from a wild-type esterase isolated from a Roseibacillus bacterial species, referred to herein as “ECE20 variant(s)”, “carboxylesterase”, “macrocyclase”, or “NF-macrocyclase”, which was identified by a BLASTP search using a previously reported promiscuous wild-type hydrolase EstCEl (see Muller et al. Discovery and Design of Family VIII Carboxylesterases as Highly Efficient Acyltransferases. Angew. Chem. Int. Ed., 2021, 60, 2013-2017, which is incorporatedherein by reference), as is described in Example 1. EstCEl is a family VIII carboxylesterase that catalyzes acyl transfer (transacylation) with high efficiency, and in particular catalyzes the irreversible amidation and carbamoylation of amines in water. Another member of the family, ECE20, is a homolog of EstCEl. The disclosed engineered polypeptides exhibit improved enzyme catalysis properties relative to wild-type ECE20, including improved enzyme activity for the cyclization of oligopeptides through amide bond formation, which properties were discovered through iterative rounds of directed evolution.

[0033] The emergence of new therapeutic modalities requires the development of complementary’ tools for their efficient syntheses. Non-natural peptides have gained attention in the pharmaceutical industry due to their high selectivity, efficacy, and safety profiles. However, their widespread application has been hindered by the high costs of synthesis and the unique chemistries involved. Enzy mes present a promising solution to supplement existing chemical approaches, as they offer high specificity and operate under mild and environmentally safer reaction conditions. This disclosure provides enzymatic strategies for the synthesis of non-natural peptides without the need for protecting group manipulations utilizing carboxylesterases, including in sequential or simultaneous enzymatic cascades.

[0034] In some aspects, the carboxylesterase enzymes described herein are the result of directed evolution of a carboxylesterase having the amino acid sequence of SEQ ID NO: 2. Such enzymes are capable of cyclizing tetrapeptides containing non-canonical amino acids. Wild-type ECE20 was shown to exhibit low activity in cyclizing a linear tetrapeptide containing non-canonical (modified) try ptophan, proline and phenylalanine residues (i.e., compound 1) to compound 2, the cyclized version of this tetrapeptide. In particular. a His-tagged variant of ECE20 wherein a signal peptide region was removed (having the sequence of SEQ ID NO: 2) exhibited very low amide bond formation activity for compound 1, even under high substrate (lysate) loading, and further generated the undesired hydrolysis by-product compound 3. It was therefore concluded that directed evolution of this polypeptide was necessary to substantially enhance this activity. As described in the Examples, carboxylesterases capable of catalyzing the desired amide bond formation between modified proline and modified phenylalanine amino acids within the substrate were engineered via structure-guided semi-rational directed evolution using techniques such as single-site-saturation mutagenesis (SSM) and combinatorial library generation.

[0035] For example, the Examples describe the generation of multiple variants of ECE20, referred to as “macrocyclases” and “NF-macrocyclases”, that exhibit more than 1000-fold enhanced activity and unsurpassed regioselectivity in catalyzing synthesis of a cyclized tetrapeptide containing modified Tryptophan-Proline-Threonine-Phenylalanine (SEQ ID NO: 44)residues. relative to a wild-type ECE20 construct containing a N-terminal 6xHis affinity tag (SEQ ID NO: 41) and lacking an N-terminal 15-amino acid signal peptide region (SEQ ID NO: 2). The macrocyclase variants described herein provided high regioselectivity for this cyclization (i.e., regioisomeric excess of at least about 99%), which was demonstrated on up to 100-gram scale. Additional advantageous properties of the macrocyclase variants are described herein, including solvent tolerance and thermostability, demonstrate their synthetic utility. Examples of such macrocyclase variants are the polypeptides having amino acid sequences SEQ ID NOs: 26 and 28.

[0036] Other macrocyclase variants disclosed herein (e.g., the variants of SEQ ID NOs: 22 and 24) also have improved enzyme properties, including improved enzyme activity, compared to wild-type ECE20. Therefore, the present disclosure provides several engineered carboxylesterase polypeptides (e.g., macrocyclases) derived from ECE20 that enable the biocatalytic synthesis of oligopeptides containing non-canonical amino acids at scale.

[0037] The engineered polypeptides (macrocyclases) provided herein are carboxylesterase enzymes that catalyze amide bond formation between amino acid residues within a linear tetrapeptide, and thus achieve a macrocyclization. These carboxylesterases are derived from family VIII, of which EstCEl and its homolog ECE20 are members, having a protein fold of |3- lactamase, class C (EC 3.1.1), having a catalytic triad S65, K68, and Y171 (see Muller et al. (2021), and Elend etal. Applied And Environmental Microbiology. 2006, 72(5):3637-3645, herein incorporated by reference). ECE20 was selected as a starting point for evolution because it was shown to promote transacylation over hydrolysis. The desired reaction for the directed evolution of the Examples was transacylation, rather than hydrolysis, of the modified proline of Compound 1.

[0038] In some aspects, the disclosed carboxylesterase enzymes may be used in a multi-step synthesis scheme. In some embodiments, the disclosed enzymes are used in multi-step scheme in a one-pot process, such as a one-pot reaction involving one or more carboxylesterases or a one- pot reaction involving one or more ligases. In some embodiments, the disclosed enzymes may be used in a multi-step synthesis with a tryptophan ligase and a phenylalanine ligase, such as an engineered Trp-ligase and an engineered Phe-ligase. This multi-step synthesis may generate a cyclic tetrapeptide product (e.g., compound 2) from one or more monomers, such as from a linear dipeptide or linear tripeptide, or a combination thereof. In some aspects, the disclosed enz mes exhibit improved activity in an enzymatic cascade relative to wild-type enzyme. In some aspects, these enzymes exhibit improved regioselectivity, solvent tolerance, and / or thermostability in an enzymatic cascade relative to wild-ty pe enzy me.

[0039] The present disclosure also provides polynucleotides and expression vectors encoding the engineered polypeptides of the present disclosure. The present disclosure also provides host cells comprising these polynucleotides or expression vectors, such as E. coll host cells. The host cells can be used for the expression and isolation of the carboxylesterase enzy mes described herein, or, alternatively, they can be used directly for the conversion of the substrate to product.

[0040] Further, the disclosure provides methods of generating the carboxylesterase polypeptides. Further, the disclosure provides methods of macrocyclizing oligopeptides using the engineered carboxylesterases of the present disclosure. In various embodiments, the disclosure provides methods of synthesizing a tetrapeptide (e.g., a tetrapeptide that includes a modified tryptophan, a modified phenylalanine and / or a modified proline) using the engineered carboxylesterases.

[0041] In some aspects, the multiple steps of these methods may be run simultaneously in cascades such that isolations and / or purifications of intermediates is eliminated. Advantageously, this eliminates the waste typically generated in multi-step chemical processes, and shortens the time required for manufacturing.Definitions

[0042] Listed below are definitions of various terms used herein. These definitions apply to the terms as they are used throughout this specification and claims, unless otherwise limited in specific instances, either individually or as part of a larger group.

[0043] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, and peptide chemistry are those well-known and commonly employed in the art.

[0044] As used herein, the articles “a” and “an” refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element. Furthermore, use of the term “including” as well as other forms, such as “include,” “includes,” and “included,” is not limiting.

[0045] As used herein, the term “about” in quantitative terms refers to plus or minus 10% of the value it modifies (rounded up to the nearest whole number if the value is not sub-dividable, such as a number of molecules or nucleotides).

[0046] All ranges disclosed herein are inclusive of the recited endpoint and independently combinable (for example, the range of “10-14 rounds” is inclusive of the endpoints, 10 rounds and 14 rounds, and all intermediate values). The endpoints of the ranges and any values disclosedherein are not limited to the precise range or value; they are sufficiently imprecise to include values approximating these ranges and / or values.

[0047] As used herein, the term “comprising” may include the embodiments “consisting of’ and “consisting essentially of.” The terms “comprise(s),” “include(s),” “having,” “has,” “may,” “contain(s),” and variants thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that require the presence of the named ingredients / steps and permit the presence of other ingredients / steps. However, such description should be construed as also describing compositions or processes as “consisting of and “consisting essentially of the enumerated components, which allows the presence of only the named components or compounds, along with any acceptable carriers or fluids, and excludes other components or compounds.

[0048] “Derived from” as used herein in the context of enzy mes, identifies the originating enzyme, and / or the gene encoding such enzyme, upon which the enzyme was based. For example, the disclosed carboxylase enzymes are “derived from” the wild-type Roseibacillus sp. esterase enzyme having the amino acid sequence of SEQ ID NO: 40 (UniProt ID: A0A2E5D605_9BACT). As used herein in the context of an amino acid residue, a “derivative” refers to a modified amino acid that is derived from a canonical L-amino acid.

[0049] As used herein, “polynucleotide” and “nucleic acid” refer interchangeably to two or more nucleotides that are covalently linked together. The polynucleotide may be wholly comprised of ribonucleotides (i.e., RNA), wholly comprised of 2' deoxyribonucleotides (i.e., DNA), or comprised of mixtures of ribo- and 2' deoxyribonucleotides. While the nucleosides will ty pically be linked together via standard phosphodiester linkages, the polynucleotides may include one or more non-standard linkages. The polynucleotide may be single-stranded or double-stranded, or the polynucleotide may include both single-stranded regions and doublestranded regions. Moreover, while a polynucleotide will typically be composed of the naturally occurring encoding nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), it may include one or more modified and / or synthetic nucleobases. such as, for example, inosine, xanthine, hypoxanthine.

[0050] As used herein, the terms “protein,” “polypeptide,” and “peptide” are used interchangeably herein to denote a polymer of at least two amino acids covalently linked by an amide bond, regardless of length or post-translational modification (e.g, glycosylation, phosphorylation, lipidation, myristoylation, ubiquitination, and the like). Included within this definition are D- and L-amino acids, and mixtures of D- and L-amino acids, as well as polymers comprising D- and L-amino acids, and mixtures of D- and L-amino acids. Further includedwithin this definition are linear, branched and cyclic peptides. Proteins, polypeptides, and peptides may include a tag (e g., an affinity tag), such as a histidine tag, and / or a signal peptide region. As used herein, a “polypeptide” may encode an enzyme. The disclosure contemplates the use of polypeptides as substrates, products, and intermediates in an enzymatic reaction.

[0051] As used herein, the terms “amino acid” or “residue” as used in context of the polypeptides disclosed herein refers to the specific monomer at a sequence position of a polypeptide molecule. Amino acids are referred to herein by either their commonly know n three- letter symbols or by the one-letter symbols recommended by International Union of Pure and Applied Chemistry (IUPAC) - International Union of Biochemistry (IUB) Biochemical Nomenclature Commission. This term encompasses modified (i.e., covalently or non-covalently modified), or non-canonical or non-standard, amino acids. For instance, this term encompasses covalently modified try ptophan, covalently modified proline, covalently modified phenylalanine, and covalently modified threonine amino acid molecules. As used herein, the terms “non- canonical amino acid” and “non-standard amino acid” refer to any amino acid other than the 20 naturally occurring L-amino acids, such as covalently modified amino acids and D-amino acids.

[0052] “Mutation” refers to any change in a polypeptide or polynucleotide sequence, and encompasses any number (i.e., one or more) of substitutions, deletions, insertions, and / or rearrangements present in a sequence compared to a reference sequence.

[0053] As used herein with respect to amino acid sequences, a “substitution” refers to a difference in the amino acid residue at a position of a polypeptide sequence relative to the amino acid residue at a corresponding position in a reference sequence. In some instances, the present disclosure provides specific amino acid differences denoted by the conventional notation “AnB,” where A is the single letter identifier of the residue in the reference sequence, n is the number of the residue position in the reference sequence, and B is the single letter identifier of the residue substitution in the sequence of the engineered polypeptide.

[0054] The term “amino acid substitution set” or “substitution set” refers to a group of amino acid substitutions in a polypeptide sequence, as compared to a reference sequence. For example, a substitution set may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-20, 20-25 or more than 25 amino acid substitutions.

[0055] “Corresponding to” or “relative to” w hen used in the context of the numbering of a given amino acid or polynucleotide sequence refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue number or residue position of a given polymer is designated with respect to the reference sequence rather than by the actual numericalposition of the residue within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, although the gaps are present, the numbering of the residue in the given amino acid or polynucleotide sequence is made with respect to the reference sequence to which it has been aligned.

[0056] As used herein, “isolated polypeptide” refers to a composition in which the polypeptide is substantially separated from other contaminants that naturally accompany it (e.g., protein, lipids, and polynucleotides). The term embraces polypeptides that have been removed or purified from their naturally occurring environment or expression system (e.g., within a host cell or via in vitro synthesis). The recombinant polypeptides may be present within a cell, present in the cellular medium, or prepared in various forms, such as lysates or isolated preparations. As such, in some embodiments, the recombinant polypeptides can be an isolated polypeptide.

[0057] As used herein, a “carboxylesterase” is a polypeptide having an enzymatic capability of catalyzing the hydrolysis of esters, amides, thioesters, or carbamates. Carboxylesterases are a subtype of esterase. In some embodiments, a carboxylesterase can act as a transacylase. In some embodiments, a carboxylesterase is capable of catalyzing the formation of an amide bond between an amine and an ester, including amine and ester moieties present in a single molecule. In various embodiments, the disclosed carboxylesterases are capable of catalyzing a cyclization of a linear tetrapeptide through amide bond formation. Carboxylesterases as used herein include naturally occurring (wild-type) polypeptides as well as non-naturally occurring engineered polypeptides generated by human manipulation.

[0058] As used herein, a “macrocyclase” is an engineered variant of carboxylesterase that exhibits high activity in catalyzing a cyclization of a linear tetrapeptide through amide bond formation.

[0059] “Improved enzy me property” refers to any property of an enzy me that exhibits an improvement relative to a reference enzyme. For the enzymes described herein, the reference enzyme is a wild-type enzyme or another engineered enzyme. For example, in various embodiments, the reference enzyme is an affinity-tagged variant of a wild-type enzyme that has been codon-optimized (e.g., an E. co / z-codon-optimized, 6xHis-tagged variant of the ECE20 enzyme). Enzyme properties for which improvement may be desirable include, but are not limited to, enzy matic activity (which may be expressed in terms of percent conversion of the substrate), thermal stability (or thermostability), stability under high ammonia concentration, soluble expression, reduced by-product generation, higher substrate loading, pH activity profile (pH tolerance), phosphate tolerance (or higher phosphate loading), cosolvent tolerance, improvedactivity in an enzymatic cascade, reduced cofactor requirements, refractoriness to inhibitors (e.g., product inhibition), regioselectivity, stereospecificity, and stereoselectivity (including enantioselectivity).

[0060] ‘‘Increased enzymatic activity’' refers to an improved property of the enzymes that is represented by an increase in specific activity (e.g., product produced / time / weight protein) or an increase in percent conversion of the substrate to the product (e g., percent conversion of starting amount of substrate to product in a specified time period using a specified amount of enzyme) as compared to a reference enzy me. Exemplary' methods to determine enzy me activity' are provided in the Examples. Any property relating to enzyme activity’ may be affected, including the classical enzyme properties of Km, Nmax, or kCflr, changes of which can lead to increased enzymatic activity. Improvements in enzyme activity can be from about 1.5 times the enzy matic activity of the corresponding wild- type enzyme, to as much as 100 times, 150 times, 200 times, 500 times, 1000 times, 3000 times, 5000 times, 7000 times, 1000 times, 1500 times, 2000 times. 5000 times, 10000 times, 15000 times, 20000 times, 22500 times, 25000 times, or more enzymatic activity than the reference enzyme, e g., a naturally occurring enzyme or another enzyme from which the polypeptides were derived. In some examples, the enzyme exhibits improved enzymatic activity in the range of 100 to 3000 times, 3000 to 7000 times, or more than 7000 times greater than that of the parent enzyme. In some examples, the enzyme exhibits improved enzymatic activity in the range of 1000 to 10000 times, 10000 to 20000 times, 20000 to 25000 times, or more than 25000 times greater than that of the reference enzyme. It is understood by the skilled artisan that the activity' of any enzyme is diffusion limited such that the catalytic turnover rate cannot exceed the diffusion rate of the substrate, including any required cofactors. The theoretical maximum of the diffusion limit, or k«« / Km. is generally about 108to 109(M^s-1). Hence, any improvements in the enzyme activity7will have an upper limit related to the diffusion rate of the substrates acted on by the enzyme. Enzyme activity' can be measured by any suitable approach, e.g., an enzyme activity' assay or by any of the traditional methods for assaying chemical reactions, including but not limited to high-performance liquid chromatography (HPLC), ultra high-performance liquid chromatography (UHPLC), HPLC-mass spectrometry’ (MS), ultra-performance liquid chromatography (UPLC), UPLC-MS, thin-layer chromatography (TLC), and nuclear magnetic resonance (NMR). Comparisons of enzy me activities may be made using a defined preparation of enzyme, a defined assay under a set condition, and one or more defined substrates, as further described in detail herein. Generally, when lysates generated after lysing cells expressing the enzyme are compared, the numbers of cells and the amount of protein assayed are determined, as well as use of identical expressionsystems (i.e.. identical host cells) to minimize variations in amount of enzyme produced by the host cells and present in the lysates.

[0061] As used herein, “by-products” refer to undesired products of a catalytic reaction, such as the hydrolysis reactions or undesired macrocyclization reactions catalyzed by the disclosed enzymes. In some embodiments, this term encompasses the product of an undesired hydrolysis reaction, such as the hydrolysis of a terminal ester to an alcohol. This term further encompasses undesired regioisomers. Examples of undesired by-products include compound 3 shown in Table A.

[0062] As used herein, a “vector” is a DNA construct for introducing a DNA sequence into a cell. In some embodiments, the vector is an expression vector that is operably linked to a suitable control sequence capable of effecting the expression of the polypeptide encoded by the polynucleotide (e.g., DNA) sequence in a suitable host cell. In some embodiments, an “expression vector” has a promoter sequence operably linked to the polynucleotide (e.g.. DNA) sequence (e.g.. transgene) to drive expression in a host cell, and in some embodiments, also comprises a transcription terminator sequence.

[0063] As used herein with respect to polypeptides, the terms “expression” and “production” includes any step involved in the production of a polypeptide (i.e., enzyme-encoding polypeptide) including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of the polypeptide from a cell.

[0064] As used herein, an amino acid or nucleotide sequence (e g., a promoter sequence, signal peptide, terminator sequence, and the like) is “heterologous” to another sequence with which it is operably linked if the two sequences are not associated in nature. For example, a “heterologous polynucleotide” is any polynucleotide that is introduced into a host cell by laboratory techniques, and the term includes polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into a host cell.

[0065] As used herein, the term “host cell” refers to a suitable host for an expression vector comprising a polynucleotide (e.g., DNA) provided herein (e.g., a polynucleotide encoding a carboxylesterase polypeptide disclosed herein). In some embodiments, the host cells are prokary otic or eukary otic cells that have been transformed or transfected with vectors constructed using recombinant DNA techniques as known in the art. In some embodiments, the host cells are E. coli cells.

[0066] “Coding sequence” refers to that portion of a polynucleotide (e.g., a gene) that encodes an amino acid sequence of a polypeptide.

[0067] "Naturally occurring’7or " wild-type’’ refers to a form found in nature. For example, a naturally occurring or wild-type polypeptide or polynucleotide sequence is a sequence present in an organism that can be isolated from a source in nature and that has not been intentionally modified by human manipulation, with the sole exception that wild-type polypeptide or polynucleotide sequences as identified herein may include a tag, such as a histidine tag (6xHis tag). Herein, '‘wild-type” polypeptide or polynucleotide sequences may be denoted “WT”.

[0068] “Operably linked” is defined herein as a configuration in which a control sequence is appropriately placed at a position relative to a polynucleotide sequence (i.e., in a functional relationship) such that the control sequence directs the expression of the polynucleotide and / or a polypeptide encoded by the polynucleotide.

[0069] A “promoter sequence” is a nucleic acid sequence that is recognized by a host cell for expression of a polynucleotide. The control sequence may comprise an appropriate promoter sequence. The promoter sequence contains transcriptional control sequences, which mediate the expression of the polynucleotide. The promoter may be any nucleic acid sequence which shows transcriptional activity in the host cell of choice including mutant, truncated, and hybrid promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides either homologous or heterologous to the host cell.

[0070] The terms “engineered.” “recombinant,” “variant,” and “non-naturally occurring,” when used with reference to, e.g., a polynucleotide, polypeptide, or cell, refers to a material, or a material corresponding to the natural or native form of the material, that has been modified in a manner that would not otherwise exist in nature. Non-limiting examples include, among others, recombinant cells expressing genes that are not found within the native (non-recombinant) form of the cell or express native genes that are otherwise expressed at a different level.

[0071] As used herein, the terms “percent identity” and '‘percent identical” refer to comparisons between polynucleotide sequences or polypeptide sequences, and are determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which either the identical nucleic acid base or amino acid residue occurs in both sequences or a nucleic acid base or amino acid residue is aligned with a gap to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity. Determination of optimal alignment and percent sequence identity' can be performed using theBLAST and BLAST 2.0 algorithms (see e.g., Altschul et al., 1990. J. Mol. Biol. 215: 403-410; and Altschul et al., 1977, Nucleic Acids Res. 3389-3402). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information website.

[0072] Briefly, the BLAST analyses involve first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as, the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T. and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) of 10, M = 5, N = -4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff. 1989, Proc. Natl. Acad. Sci. USA 89: 10915).

[0073] Numerous other algorithms are available that function similarly to BLAST in providing percent identify for two sequences. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman, 1981, Adv. Appl. Math. 2:482. by the homology alignment algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48:443, by the search for similarity method of Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444, by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin Software Package), or by visual inspection (see generally, Current Protocols in Molecular Biology’, F. M. Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)). Additionally, determination of sequence alignment and percentsequence identity can employ the BESTF1T or GAP programs in the GCG Wisconsin Software package (Accelrys. Madison WI), using default parameters provided.

[0074] As used herein, the terms “stereoselectivity'” and “stereospecificity” refer to the preferential formation in a chemical or enzymatic reaction of one stereoisomer over another. Stereoselectivity can be partial, where the formation of one stereoisomer is favored over the other, or it may be complete where only one stereoisomer is formed. When the stereoisomers are enantiomers, the stereoselectivity is referred to as enantioselectivity, the fraction (typically reported as a percentage) of one enantiomer in the sum of both. It is commonly alternatively reported in the art (typically as a percentage) as the enantiomeric excess (EE) calculated therefrom according to the formula [major enantiomer - minor enantiomer] / [major enantiomer + minor enantiomer]. Where the stereoisomers are diastereomers, the stereoselectivity is referred to as diastereoselectivity, the fraction (typically reported as a percentage) of one diastereomer in a mixture of two diastereomers, commonly alternatively reported as the diastereomeric excess (DE). Enantiomeric excess and diastereomeric excess are types of stereomeric excess.

[0075] “Regioselectivity” refers to the preferential formation in a chemical or enzymatic reaction of one regioisomer over another. Regioselectivity can be partial, where the formation of one regioisomer is favored over the other, or it may be complete where only one regioisomer is formed. “Highly regioselective” refers to a chemical or enzymatic reaction that is capable of converting a substrate to its corresponding product with at least about 85% regioisomeric excess. “Regioisomers” encompass diastereomers and enantiomers. Regioselectivity may result when an enzyme demonstrates preferential catalysis of a chemical moiety' at a single position or configuration in a substrate molecule than at other positions.

[0076] “Chemoselectivity” refers to the preferential formation in a chemical or enzymatic reaction of one product over another.

[0077] “Conversion” refers to the enzy matic transformation of a substrate to the corresponding product. “Percent conversion” refers to the percent of the substrate that is converted to the product within a period of time under specified conditions. Thus, for example, the “enzymatic activity” or “activity” of a polypeptide can be expressed as “percent conversion” of the substrate to the product.

[0078] As used herein, the terms “thermostability” and “thermal stability” refer to the ability of an enzyme to maintain the same or similar level of activity in catalyzing a reaction (more than 60% to 80% of product conversion, for example) after exposure to elevated temperatures (e.g., 37 °C to 50 °C, or 37 °C to 80 °C) for a period of time (e.g., 0.5 h to 24 h) relative to the corresponding enzyme at room temperature. Thermostability may be measured using a heatchallenge test, e.g.. measuring enzyme activity for about 60 minutes after exposure to elevated temperatures (e.g., 35 °C, 37 °C, or 39-40 °C).

[0079] As used herein, the terms “biocatalysis,” “biocatalytic,” “biotransformation,” and “biosynthesis” refer to the use of enzymes to perform chemical reactions on organic compounds.

[0080] As used herein, “polyphosphate” refers to an oligomer of two or more phosphate molecules (e.g., ions) covalently linked together.

[0081] As used herein, the terms “co-solvent tolerance” and “solvent tolerance” refer to the capability of an enzyme to exhibit activity under a range of co-solvent concentrations, such as organic co-solvents, that are relevant to the chemical process occurring in an aqueous medium (e.g., an aqueous buffer such as HEPES or sodium phosphate buffer). Exemplary organic cosolvents are acetonitrile, acetone, dimethyl sulfoxide, hexane and heptane. In particular embodiments, the organic co-solvent is acetonitrile (ACN). A co-solvent may or may not be present with a primary aqueous solvent, in any desired reaction medium. Enzyme co-solvent tolerance can be measured by measuring the activity of the enzyme under low co-solvent concentrations (e.g., acetonitrile absent) relative to activity under higher co-solvent concentrations (e.g., >15% or >20% v / v acetonitrile); a co-solvent tolerant enzyme displays minimal difference in activity between low and higher co-solvent concentrations.

[0082] As used herein, the terms “substrate loading” and “substrate loads” refer to the concentration of substrate present in a reaction vessel following a contacting of substrate to enzyme, e.g., during a synthesis reaction. Many natural enzymes exhibit little to no activity at high substrate loads. As used herein, the term “high substrate loading” refers to a concentration of substrate between 5 and 75 mM. Example substrate concentration ranges for “high substrate loading” include 1-10 mM, 5-15 mM. 10-20 mM. 20-30 mM. 30-40 mM, 5-45 mM, 15-45 mM. 25-45 mM, 40-50 mM, 50-60 mM, 60-75 mM, 5-75 mM, 15-75 mM, 25-75 mM, or 50-75 mM. “Low substrate loading” may refer to concentrations of less than 5 mM.

[0083] As used herein, the term “cascade” refers to the use of several enzymes, such as two or more engineered enzymes that each catalyze a unique step, to catalyze the steps of a multi-step reaction simultaneously in a one-pot process.Carboxylesterase polypeptides

[0084] In some embodiments, provided herein are carboxylesterase enzymes that are capable of catalyzing the formation of an amide bond within modified Phenylalanine-Tryptophan-Proline- Threonine tetrapeptides, such as Compound 1 as depicted in Table A, to cyclize the molecule. Inparticular embodiments, the disclosed enzymes selectively catalyze transacylation between the modified proline and modified phenylalanine residues, over hydrolysis of the modified proline.

[0085] In particular embodiments, provided herein are engineered variants of ECE20 carboxylesterase (Roseibacilhis). The carboxylesterase enzy mes described herein are the product of directed evolution from SEQ ID NO: 2 (encoded by SEQ ID NO: 1), a wild-type hydrolase from a Roseibacillus bacterial species, referred to herein as '‘ECE20”, “carboxylesterase”, “macrocyclase”, or “NF-macrocyclase”, which was identified by a BLAS TP search using a previously reported promiscuous wild-type hydrolase EstCEl described in Elend et al. Applied And Environmental Microbiology, 2006. In some embodiments, provided are engineered variants that have been evolved through successive rounds of directed evolution. In some embodiments, provided are engineered variants that have been evolved through 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 successive rounds of directed evolution. In some embodiments, provided are engineered variants that have been evolved through 5-10 rounds, 10-14 rounds, 13 rounds, or more than 14 successive rounds of directed evolution.

[0086] The polynucleotide used for directed evolution was codon optimized for production in the E. coll host and a sequence encoding a short histidine affinity tag (6xHis (SEQ ID NO: 41)) was added at the N-terminal encoding region. The amino acid sequence of this tag and short linker is MHHHHHH (SEQ ID NO: 37). In addition, the polynucleotides encoding the first 15 amino acids of the carboxylesterase coding sequence were removed, as they were predicted to be a signal peptide region and were excluded to promote cytoplasmic expression within the host. This construct is encoded by polynucleotide of SEQ ID NO: 1. The polypeptide sequence of this His-tagged ECE20 is set forth below in SEQ ID NO: 2:MHHHHHHLSTARKPSSRELPVADPATVGMSRDKLQLVGDKVQSL1RENR1AGAS VMVTRKGKIVYSESFGLRDIENEKPMESDTIFRIYSMTKAVTSVAAMMLVERGEL HLEDAVSKYLPEFKNAKVWKDENRFPPKTPTTIKDLLCHTSGYSYGNIGIPAIDEA HKENGSLLETIPLRKFCRKAALIPMAFEPGERWLYGISTDLLGAVVEQVSDMTLD RFFQSQIFTPLGMVDTGFVIPKDKRTRLAAAYDSDKKGTLKRRTTDLFTYSAKTR MLS GGGGL ASTIRDYTRFLQMMVNGGELHGHRLLQKTTVEEMTRNHLS GP AMP IRFPGNLRHGTGFGLGFSVKVSQKDWNQAGRLGEYGWGGMLSTHFWISPADQL VVVTMEQTFPFDFLLEDALKPLIYNSIE (SEQ ID NO: 2)

[0087] In some embodiments, the carboxylesterase enzymes of the disclosure may exhibit improvements relative to the carboxylesterase enzyme of SEQ ID NO: 2, such as increases in enzyme activity, regioselectivity, thermostability, phosphate tolerance, organic co-solvent tolerance, and reduction in by-product formation. These improvements can relate to a singleenzyme property, such as enzymatic activity, or a combination of different enzyme properties, such as enzymatic activity and thermostability.

[0088] Any of the below-described properties may be measured in accordance with techniques known in the art. Any of these properties may be measured following performance of the reaction in an aqueous medium, such as sodium phosphate buffer, in the presence or absence of cosolvent. Any of these properties may be measured following performance of the reaction in a medium containing cofactor magnesium ion (Mg2+).

[0089] The present disclosure provides numerous exemplary carboxylesterase enzymes capable of generating tetrapeptides, such as compound 1. Those exemplary polypeptides were evolved from SEQ ID NO: 2 and exhibit improved properties, particularly improved activity in the conversion of linear tetrapeptides to cyclic tetrapeptides. These exemplary engineered carboxylesterase enzy mes having tetrapeptide cyclization activity have amino acid sequences that include one or more residue differences as compared to SEQ ID NO: 2. as depicted in the accompanying sequence listing in Table 15.

[0090] As such, these variants may comprise amino acid sequences having at least 80%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 92.5%, 93%, 94%, 95%, 96%, 97%, 97.5%, 98% or 99% sequence identity to SEQ ID NO: 2. In various embodiments, these variants comprise amino acid sequences having at least 92.5%, 94%, 95%, 96%, 98%, or 99% identity to any one of SEQ ID NOs: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32 and 34. In some embodiments, these variants comprise amino acid sequences comprising any one of SEQ ID NOs: 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30. 32 and 34.

[0091] In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 92.5%, 94%, 95%. 96%. 98%. or 99% sequence identity to any one of SEQ ID NOs: 22, 24, 26, and 28. In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 30, 32, and 34. In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 38 and 39. In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 92.5% sequence identity to any one of SEQ ID NOs: 22, 24, 26, and 28.

[0092] In some embodiments, the polypeptide comprises an amino acid sequence having at least 92.5% sequence identity to SEQ ID NO: 22. In some embodiments, the polypeptide comprises an amino acid sequence having at least 92.5% sequence identity to SEQ ID NO: 24. In some embodiments, the polypeptide comprises an amino acid sequence having at least 92.5%sequence identity to SEQ ID NO: 26. In some embodiments, the enzyme polypeptide comprises an amino acid sequence having at least 92.5% sequence identity to SEQ ID NO: 28.

[0093] Polypeptides comprising any of SEQ ID NOs: 22, 24, 26, and 28 are provided. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 22. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 24. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 26. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 28. Polypeptides consisting essentially of any of SEQ ID NOs: 22, 24, 26, and 28 are provided.

[0094] Polypeptides comprising either of SEQ ID NOs: 38 and 39 are provided.

[0095] In some aspects, provided are engineered polypeptides comprising an amino acid sequence that comprises a stretch of at least 100, 150, 200, 250, 300, 325, 350, 375, or 400 consecutive amino acids of any one of SEQ ID NOs: 22, 24, 26, and 28. In some embodiments, the polypeptides comprise an amino acid sequence that comprises a stretch of at least 100, 150, or 200 or more consecutive amino acids of any one of SEQ ID NOs: 22, 24, 26, and 28. In some aspects, provided are engineered polypeptides comprising an amino acid sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-25, 25-35, 35-50, 50-75, or more than 75 amino acids relative to the amino acid sequence of any one of SEQ ID NOs 22, 24, 26, and 28. The polypeptides may comprise an amino acid sequence that differs by 1. 2, 3, 4, or 5 amino acids relative any one of SEQ ID NOs: 22, 24, 26, and 28.

[0096] In some embodiments, the polypeptides comprise an amino acid sequence that comprises a stretch of at least 200, 300, 400 or more than 400 consecutive amino acids of SEQ ID NO: 28. In some aspects, provided are engineered polypeptides comprising an amino acid sequence that differs by 1. 2. 3, 4, 5. 6, 7, 8. 9. 10. 11. 12. 13. 14. 15. 15-25, 25-35. 35-50, 50-75, or more than 75 amino acids relative to the amino acid sequence of SEQ ID NO: 28.

[0097] In some embodiments, engineered carboxylesterase polypeptides are provided that comprise an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 2, wherein the polypeptide comprises at least 2, 3, 4. 5, 6, 7. 8, 9, 10, or more than 10 mutations at positions selected from 10, 17, 115, 132, 133, 138, 140, 147, 154, 156, 157, 158, 159, 169, 162, 165, 166, 170, 172, 184, 187, 191, 199, 246, 265, 281, 325, 327, 330, 350, 352, 368, and 373, relative to SEQ ID NO: 2. In some embodiments, at least 5 mutations are present at positions selected from 10, 17. 115,132, 133, 138. 140, 147, 154, 156, 157. 158, 159, 169, 162, 165, 166, 170, 172, 184, 187, 191, 199, 246, 265, 281, 325, 327, 330, 350, 352, 368, and 373, relative to SEQ ID NO: 2.

[0098] In some embodiments, the polypeptides comprise the following substitutions relative to SEQ ID NO: 2: I159Y, Y154M, N156P, G170W, L172C, L199G, N132R, K352T, M368A. In some embodiments, the polypeptides further comprise the following substitutions: R133M, T138R, M154N, K184V, A191D, M327Y, VI 151, T140P, C147R, G158E, F373Y. In some embodiments, the polypeptides further comprise the following substitutions: R330S, P160L. I162G, A272H, S350D, R17Y, F134Q, A165P, T265S, D266L, P325S.

[0099] In some embodiments, the polypeptides comprise mutations at positions 17, 115, 132, 133, 138, 140, 147, 154, 156, 158, 159, 160, 162, 165, 166, 170, 172, 184, 191, 199, 265, 272, 281, 325. 327, 330, 343, 350, 352. 368, and 373. In particular embodiments, the polypeptides comprise the following substitution set: R17Y, VI 151, N132R, R133M, T138R, T140P, C147R, Y154N, N156P, G158E, I159Y, P160L, I162G, A165P, H166W, G170W, L172C, K184V, A191D, L199S, T265S, A272H, G281S, P325S, M327Y, R330S, L343F, S350D, K352T, M368A, and F373Y.

[0100] In some embodiments, the polypeptides comprise mutations at positions 17. 115, 132, 133, 138, 140, 147, 154, 156, 158, 159, 160, 162, 165, 166, 170, 172, 184, 191, 199, 265, 272, 281, 325, 327, 330, 343, 350, 352, 368, and 373. In particular embodiments, the polypeptides comprise the following substitution set: R17Y, VI 151, N132R, R133M, T138R, T140P, C147R, Y154N, N156P, G158E, I159Y, P160L. I162G, A165P. H166W, G170W, L172C, K184V, A191D, L199S, T265S, A272H, G281S, P325S, M327Y, R330S, L343F, S350D, K352T, M368L, and F373Y.

[0101] In some embodiments, the polypeptides comprise mutations at positions 17, 115, 132, 133, 138. 140, 147, 154, 156, 157. 158, 159, 160, 162, 165, 166, 170. 172, 184, 187, 191, 199. 265, 272. 281, 325, 327. 330, 343. 350, 352, 368. and 373. In particular embodiments, the polypeptides comprise the following substitution set: R17Y, V115L, N132R, R133M, T138R, T140P, C147R, Y154N, N156P, I157F, G158E, I159Y, P160L, I162G, A165P, H166W, G170L, L172F, KI 84V, L187Q, A191D, L199S, T265S, A272H, G281S, P325S. M327Y, R330S, L343F. S350D, K352T, M368L. and F373Y.

[0102] In some embodiments, the polypeptides comprise mutations at positions 10, 17, 1 15, 132, 133, 138, 140, 147, 154, 156, 157, 158, 159, 160, 162, 165, 166, 170, 172, 184, 187, 191, 199, 246, 265, 281, 325, 327, 330, 350, 352, 368, and 373. In particular embodiments, the polypeptides comprise the following substitution set: T10K, R17Y. V115L, N132R, R133M, T138R, T140P, C147R, Y154N, N156P, I157F, G158E, I159Y, P160L, A165P, H166W, G170L, L172F, KI 84V, L187Q, A191D, L199S, T246H, T265S, G281S, P325S, M327Y, R330S, S350D, K352T, M368L, and F373Y.

[0103] In various embodiments, the polypeptides comprise an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 2, wherein the polypeptide comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 substitutions selected from T10K, R17Y, V115L, N132R, R133M, T138R, T140P, C147R. Y154N, N156P, I157F, G158E, I159Y, P160L, I162G, A165P. H166W, G170L. L172F. K184V, L187Q, A191D, L199S, T246H, T265S, G281S. P325S. M327Y, R330S, S350D, K352T, M368L, and F373Y, relative to SEQ ID NO: 2.

[0104] In some embodiments, the disclosed polypeptides comprise one or more amino acid substitution sets set forth in Table 1.

[0105] In various embodiments, the polypeptides have esterase activity. In various embodiments, the polypeptides have carboxylesterase activity7. In various embodiments, the polypeptides have acyltransferase activity. In various embodiments, the polypeptides exhibit activity in catalyzing a cyclization of a linear tetrapeptide.

[0106] In various embodiments, the polypeptides may be isolated.

[0107] In some embodiments, the polypeptides comprise an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs 22, 24, 26, and 28, wherein the polypeptide contains 0, 1 or 2 amino acid residues that differ from amino acids 147-172 of SEQ ID NO: 24.

[0108] In some embodiments, the carboxylesterase enzymes of the disclosure exhibit improved activity on substrate compound 1 relative to the carboxylesterase of SEQ ID NO: 2, in the production of cyclic tetrapeptide compound 2. In some embodiments, the carboxylesterase enzymes of the disclosure exhibit improved activity in converting compound 1 into compound 2 relative to the carboxylesterase of SEQ ID NO: 2.

[0109] In some embodiments, the disclosed engineered carboxylesterases able to generate a percent conversion of substrate to product (e g., converting compound 1 into compound 2) of at least about 10%, at least about 40%, at least about 60%, at least about 80%, at least about 90%, at least about 92%, at least about 95%, at least about 96%. or at least about 98%. In some embodiments, the disclosed carboxylesterases can generate a percent conversion of at least about 80%. In some embodiments, the disclosed carboxylesterases can generate a percent conversion of at least about 90%. In some embodiments, the disclosed carboxylesterases can generate a percent conversion of at least about 95%.

[0110] In some embodiments, the engineered carboxylesterase enzymes catalyze the formation of Compound 2 with at least about 5-fold, 10-fold, 50-fold, 100-fold, 250-fold, 500- fold, 1000- fold, or more than 1000-fold the activity7compared to SEQ ID NO: 2, under suitable reaction conditions. In some embodiments, a greater than 50-fold increase relative to SEQ ID NO: 2 isexhibited. In particular embodiments, a greater than 1000-fold increase relative to SEQ ID NO: 2 is exhibited.[OHl] In some embodiments, the engineered carboxylesterases of the disclosure exhibit improved regioselectivity in the production of compound 2 relative to the wild-type ECE20 carboxylesterase of SEQ ID NO: 2. In some embodiments, the engineered carboxylesterases can form the desired product 2 with regioisomeric ratios of at least 1 : 10, 10: 1, 50: 1 and more than 100: 1, relative to the carboxylesterase of SEQ ID NO: 2. These enzymes generate undesired regioisomers of Compound 2 in substantially reduced amounts relative to SEQ ID NO: 2.

[0112] In some embodiments, the carboxylesterases of the disclosure exhibit improved generation of undesired by-products (e.g., compound 3) relative to SEQ ID NO: 2. In particular embodiments, the engineered carboxylesterases can form the desired product compound 2 over by-product compound 3 in ratios of at least 10: 1 (10-fold), 25:1 (25-fold), 50: 1 (50-fold), 99: 1, 100: 1, or more than 100: 1. In some embodiments, a greater than 10-fold product to by-product ratio is exhibited. In particular embodiments, a 99: 1 product to by-product ratio is exhibited.

[0113] In some embodiments, the carboxylesterase enzymes of the disclosure exhibit improved thermostability' relative to SEQ ID NO: 2. The carboxylesterase enzymes of the disclosure may exhibit higher activity at low temperatures relative to SEQ ID NO: 2. In particular embodiments, the disclosed enzymes exhibit a greater than 20-fold increase in thermostability relative to SEQ ID NO: 2. In some embodiments, the engineered carboxylesterase enzymes have been improved for thermostability and can maintain at least 50% of its enzymatic activity at temperatures of at least about 35 °C, 37°C, 39°C, 42°C, 43°C, 50°C or higher, when subjected to a heat challenge test for about 60 minutes. In particular embodiments, the disclosed enzymes are thermally stable at 50°C or higher.

[0114] In various embodiments, the disclosed engineered carboxylesterases exhibit improved kinetics of the reaction, i.e., reduced time of reaction, relative to the ECE20 enzyme of SEQ ID NO: 2. In some embodiments, these enzymes are able to generate a conversion of product of at least about 80% or at least about 90% in a reaction time less than about 96 hours, about 48 hours, about 20 hours, about 6 hours or less than 6 hours. In some embodiments, these enzymes are able to generate a conversion of product of at least about 80% or at least about 90% in a reaction time less than about 20 hours, about 6 hours, or less than 6 hours at substrate loading concentrations of 1-10 mM, 5-15 mM, 10-20 mM, 20-30 mM, 30-40 mM, 5-45 M, 15-45 mM, 25-45 mM, 40-50 mM, 50-60 mM, 60-75 mM, 5-75 mM, 15-75 mM, 25-75 mM, or 50-75 mM.

[0115] In some embodiments, the disclosed carboxylesterases exhibit increased soluble expression, e.g., increased recombinant / soluble expression in E. coll, relative to the carboxylesterase of SEQ ID NO: 2.

[0116] In some embodiments, the carboxylesterases of the disclosure exhibit improved tolerance to organic co-solvents relative to SEQ ID NO: 2. In some embodiments, the disclosed carboxyl esterase enzymes of the disclosure exhibit improved tolerance to acetonitrile co-sol vent (e.g., concentrations of acetonitrile in buffer of 15%, 20%, 25%, or 30%), relative to SEQ ID NO: 2. In particular embodiments, the disclosed enzymes exhibit a greater than 80-fold increase in tolerance of ACN co-solvent relative to SEQ ID NO: 2.

[0117] In some embodiments, any of the disclosed carboxylesterase polypeptides further comprise a tag (e.g., an affinity tag). Any suitable tag may be used, e.g., a 6xHis tag or an 8xHis tag, a FLAG tag, a fluorescent protein tag (e.g., green fluorescent protein (GFP), yellow fluorescent protein (YFP), or red fluorescent protein (RFP)), a hemagglutinin (HA) tag, an ALFA-tag. a V5-tag. a Myc-tag. a SPOT-tag, a T7-tag, or an NE-tag. In some embodiments, the affinity tag is a His tag. In some embodiments, the affinity tag comprises the amino acid sequence of MHHHHHH (SEQ ID NO: 37). In some embodiments, the affinity tag comprises an 8xHis tag (SEQ ID NO: 42). In some embodiments, the polypeptide comprises an N-terminal methionine residue, and the epitope tag is inserted immediately following the N-terminal methionine residue, e.g., relative to a reference sequence.

[0118] In some embodiments, the disclosed Carboxylesterase polypeptides do not contain an affinity tag. In some embodiments, the disclosed carboxylesterase polypeptides do not contain a histidine tag, such as an 8xHis tag. For example, the carboxylesterase polypeptide may comprise either of SEQ ID NO: 38 and 39, which correspond to the sequences of SEQ ID NOs: 26 and 28 without a histidine tag, respectively.Polynucleotides Encoding Carboxylesterases

[0119] In another aspect, the present disclosure provides polynucleotides encoding the carboxylesterase enzymes disclosed herein. The polynucleotides may be operatively linked to one or more heterologous regulatory' sequences that control gene expression to create a recombinant vector capable of expressing the polypeptide. Expression vectors containing a heterologous or engineered polynucleotide encoding the carboxylesterase can be introduced into appropriate host cells (e.g., E. coli cells) to express the corresponding carboxylesterase polypeptide. The disclosed polynucleotides are derived from the polynucleotide sequence of SEQ ID NO: 1, which encodes a wild-type carboxylesterase enzy me that lacks the signal peptideregion. and which has been codon-optimized for E. colt expression and further contains a 3' 6xHis tag. SEQ ID NO: 1 has a 74.6% sequence identity to the wild type carboxylesterase- encoding polynucleotide sequence (GenBank ID: PBAIO 1000020.1), provided in Muller et al. Angew. Chem. Int. Ed., 2021.

[0120] In various embodiments, provided herein are engineered polynucleotides that comprise nucleic acid sequences having at least 75%, 80%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 92.5%, 93%, 94%, 95%, 96%, 97%, 97.5%, 98% or 99% sequence identity to SEQ ID NO: 1. In various embodiments, these variants comprise nucleic acid sequences having at least 75%, 80%, 85%. 90%, 92.5%, 94%, 95%. 96%. 98%. or 99% identity to any one of SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, and 35. In some embodiments, these variants comprise nucleic acid sequences having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% identity' to any one of SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, and 27. In some embodiments, these variants comprise nucleic acid sequences comprising any one of SEQ ID NOs: 3. 5, 7, 9. 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31. and 35. In some examples, the polynucleotide sequence of any of SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, and 35 is modified to no longer code for an N-terminal 6xHis tag having the amino acid sequence of SEQ ID NO: 37. In some examples, the polynucleotide sequence of any of SEQ ID NOs: 3, 5, 7, 9. 11. 13. 15. 17, 19, 21, 23, 25, 27, 29, 31, and 35 is modified to have a signal peptide region-encoding sequence.

[0121] In some embodiments are provided engineered polynucleotides comprising a nucleic acid sequence having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 21, 23, 25, and 27. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 27. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity' to SEQ ID NO: 25. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 23. In some embodiments, the enzyme polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity’ to SEQ ID NO: 21.

[0122] Polynucleotides comprising any of SEQ ID NOs: 21, 23, 25, and 27 are provided. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 27. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 25. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 23. In some embodiments, the enzy me comprises the nucleic acid sequence of SEQ ID NO: 21.

[0123] In some aspects, provided are engineered polynucleotides compnsing an nucleic acid sequence that comprises a stretch of at least 100, 150, 200, 250, 300, 325, 350, 500, 650, 800, 1000, 1100, 1200, or 1225 consecutive nucleotides of any one of SEQ ID NOs 21, 23, 25, and 27. In some aspects, provided are engineered polynucleotides comprising an nucleic acid sequence that comprises a stretch of at least 250, 350. 500, 650. 800, 1000. 1100, 1200, or 1225 consecutive nucleotides of any one of SEQ ID NOs 21, 23, 25, and 27. In some aspects, provided are engineered polynucleotides comprising an nucleic acid sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-25, 25-35, 35-50, 50-75, or more than 75 nucleotides relative to the sequence of any one of SEQ ID NOs 21, 23, 25, and 27. The polynucleotides may comprise an nucleic acid sequence that differs by 1, 2, 3, 4, or 5 nucleic acids relative to SEQ ID NO: 25 or 27.

[0124] In various embodiments, the disclosed polynucleotides are codon-optimized for expression in a particular organism, such as E. coll. Because of the knowledge of the codons corresponding to the various amino acids, availability of a protein sequence provides a description of all the polynucleotides capable of encoding the subject. The degeneracy of the genetic code, where the same amino acids are encoded by alternative or synonymous codons allows an extremely large number of nucleic acids to be made, all of which encode the improved Carboxylesterase enzymes disclosed herein. Thus, having identified a particular amino acid sequence, those skilled in the art could make any number of different nucleic acids by simply modifying the sequence of one or more codons in a way that does not change the amino acid sequence of the protein. In this regard, the present disclosure specifically contemplates each and every possible variation of polynucleotides that could be made by selecting combinations based on the possible codon choices, and all such variations are to be considered specifically disclosed for any polypeptide disclosed herein.

[0125] In various embodiments, the codons are preferably selected to fit the host cell in which the protein is being produced. For example, preferred codons used in bacteria are used to express the gene in bacteria; preferred codons used in yeast are used for expression in yeast; and preferred codons used in mammals are used for expression in mammalian cells. By way of example, the polynucleotide of SEQ ID NO: 1 have been codon optimized for expression in Escherichia coli to afford SEQ ID NO: 3.

[0126] In certain embodiments, all codons need not be replaced to optimize the codon usage of the carboxylesterase enzyme since the natural sequence will comprise preferred codons and because use of preferred codons may not be required for all amino acid residues. Consequently, codon optimized polynucleotides encoding the carboxylesterase enzymes may contain preferredcodons at about 40%, 50%, 60%, 70%, 80%, or greater than 90% of codon positions of the full- length coding region.

[0127] In various embodiments, an isolated polynucleotide encoding an improved carboxylesterase polypeptide may be manipulated in a variety of ways to provide for expression of the polypeptide. Manipulation of the isolated polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector. The techniques for modifying polynucleotides and nucleic acid sequences utilizing recombinant DNA methods are well know n in the art. Guidance is provided in Sambrook et al., 2001, MOLECULAR CLONING: A LABORATORY MANUAL, 3rd Ed.. Cold Spring Harbor Laboratory Press: and CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, Ausubel. F. ed., Greene Pub. Associates, 1998, updates to 2006.

[0128] In some embodiments, an isolated polynucleotide encoding any of the carboxylesterase polypeptides herein is manipulated in a variety’ of ways to facilitate expression of the carboxylesterase polypeptide. In some embodiments, the polynucleotides encoding the carboxylesterase polypeptides comprise expression vectors where one or more control sequences is present to regulate the expression of the carboxylesterase polypeptides. Manipulation of the isolated polynucleotide prior to its insertion into a vector may be desirable or necessary' depending on the expression vector utilized. Techniques for modifying polynucleotides and nucleic acid sequences utilizing recombinant DNA methods are well known in the art. In some embodiments, the control sequences include among others, promoters, leader sequences, polyadenylation sequences, propeptide sequences, signal peptide sequences, and transcription terminators. In some embodiments, suitable promoters are selected based on the host cell selection. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure, include, but are not limited to, promoters obtained from the E. coll lac operon. In addition, suitable promoters may include Streptomyces coelicolor agarase gene (dagA), Bacillus subtilis levansucrase gene (sacB), Bacillus licheniformis alphaamylase gene (amyL), Bacillus stearothermophilus maltogenic amylase gene (amyM). Bacillus amyloliquefaciens alpha-amylase gene (amyQ), Bacillus licheniformis penicillinase gene (penP), Bacillus subtilis xylA and xylB genes, and prokaryotic beta-lactamase gene (See e.g., Villa- Kamaroff et al., PROC. NATL ACAD. SCI. USA 75: 3727-3731

[1978] ), as well as the tac promoter (See e.g.. DeBoer et al., PROC. NATL ACAD. SCI. USA 80: 21-25

[1983] ).

[0129] In some embodiments, the control sequence is also a suitable transcription terminator sequence (i.e., a sequence recognized by a host cell to terminate transcription). In some embodiments, the terminator sequence is operably linked to the 3’ terminus of the nucleic acidsequence encoding the enzyme polypeptide. Any suitable terminator that is functional in the host cell of choice finds use in the present invention.

[0130] In some embodiments, the control sequence is also a suitable leader sequence (i.e., a non-translated region of an mRNA that is important for translation by the host cell). In some embodiments, the leader sequence is operably linked to the 5' terminus of the nucleic acid sequence encoding the carboxylesterase polypeptide. Any suitable leader sequence that is functional in the host cell of choice find use in the present invention. Exemplary leaders for E. coli will encode a ribosome binding site.

[0131] In some embodiments, the control sequence is a signal peptide region (i.e., a coding region that codes for an amino acid sequence linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the cell’s secretory pathway). In some embodiments, the 5' end of the coding sequence of the nucleic acid sequence inherently contains a signal peptide coding region naturally linked in translation reading frame with the segment of the coding region that encodes the secreted polypeptide. Alternatively, in some embodiments, the 5' end of the coding sequence contains a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the secretory pathway of a host cell of choice finds use for expression of the engineered polypeptide(s). Effective signal peptide coding regions for bacterial host cells are the signal peptide coding regions include, but are not limited to, those obtained from the genes for Roseibacillus , Bacillus NC1B 11837 maltogenic amylase, Bacillus stearothermophilus alphaamylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus stearothermophilus neutral proteases (nprT, nprS. nprM), and Bacillus subtilis prsA.

[0132] In some embodiments, regulatory sequences are also utilized. These sequences facilitate the regulation of the expression of the polypeptide relative to the growth of the host cell. Examples of regulatory' systems are those that cause the expression of the gene to be turned on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory systems include, but are not limited to, the ADH2 system or GALI system. In filamentous fungi, suitable regulatory sequences include, but are not limited to, the TAKA alpha-amylase promoter, Aspergillus niger glucoamylase promoter, and Aspergillus oryzae glucoamylase promoter.

[0133] In another aspect, the present disclosure provides a recombinant expression vector comprising a polynucleotide encoding a carboxylesterase polypeptide, and one or more expression regulating regions such as a promoter and a terminator, a replication origin, etc.,depending on the type of hosts into which they are to be introduced. In some embodiments, the various nucleic acid and control sequences described herein are joined together to produce expression vectors that include one or more convenient restriction sites to allow for insertion or substitution of the nucleic acid sequence encoding the enzyme polypeptide at such sites. Alternatively, in some embodiments, the nucleic acid sequence of the present invention is expressed by inserting the nucleic acid sequence or a nucleic acid construct comprising the sequence into an appropriate vector for expression. In some embodiments involving the creation of the expression vector, the coding sequence is located in the vector so that the coding sequence is operably linked with the appropriate control sequences for expression.

[0134] The disclosed expression vector may be any suitable vector (e.g, a plasmid or virus), that can be conveniently subjected to recombinant DNA procedures and bring about the expression of the enzyme polynucleotide sequence. The choice of the vector typically depends on the compatibility of the vector with the host cell into which the vector is to be introduced. The vectors may be linear or closed circular plasmids.

[0135] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extra-chromosomal entity, the replication of which is independent of chromosomal replication, such as a plasmid, an extra-chromosomal element, a minichromosome, or an artificial chromosome). The vector may contain any means for assuring self-replication. In some alternative embodiments, the vector is one in which, when introduced into the host cell, it is integrated into the genome and replicated together with the chromosome(s) into which it has been integrated. Furthermore, in some embodiments, a single vector or plasmid, or two or more vectors or plasmids which together contain the total DNA to be introduced into the genome of the host cell, and / or a transposon is utilized.

[0136] In some embodiments, the expression vector contains one or more selectable markers, which permit easy selection of transformed cells. A “selectable marker’’ is a gene, the product of which provides for biocide or viral resistance, resistance to heavy metals, prototrophy to auxotrophs, and the like. Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis , or markers, which confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol or tetracycline resistance. Suitable markers for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2. MET3, TRP1, and URA3. Selectable markers for use in filamentous fungal host cells include, but are not limited to, amdS (acetamidase; e.g, from A. nidulans or A. orzyae), argB (ornithine carbamoyltransferases), bar (phosphinothricin acetyltransferase; e.g, from S. hygroscopicus), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5 '-phosphatedecarboxylase; e.g., from A. nidulans or A. orzyae). sC (sulfate adenyltransferase), and trpC (anthranilate synthase), as well as equivalents thereof.

[0137] In some alternative embodiments, the expression vectors contain additional nucleic acid sequences for directing integration by homologous recombination into the genome of the host cell. The additional nucleic acid sequences enable the vector to be integrated into the host cell genome at a precise location(s) in the chromosome(s). To increase the likelihood of integration at a precise location, the integrational elements preferably contain a sufficient number of nucleotides, such as 100 to 10,000 base pairs, preferably 400 to 10,000 base pairs, and most preferably 800 to 10.000 base pairs, which are highly homologous with the corresponding target sequence to enhance the probability of homologous recombination. The integrational elements may be any sequence that is homologous with the target sequence in the genome of the host cell. Furthermore, the integrational elements may be non-encoding or encoding nucleic acid sequences. On the other hand, the vector may be integrated into the genome of the host cell by non-homologous recombination.

[0138] For autonomous replication, the vector may further comprise an origin of replication enabling the vector to replicate autonomously in the host cell in question. Examples of bacterial origins of replication are ColEl ori, P15A ori, and the origins of replication of plasmids pBR322, pUC19, pACYC177 (which contains the P15A ori), or pACYC184 (which contains the P15A ori) permitting replication in E. coll, and pUBl 10, pE194, or pTA1060 permitting replication in Bacillus. The origin of replication may be one having a mutation which makes its functioning temperature-sensitive in the host cell (see e.g.. Ehrlich, Proc. Natl. Acad. Sci. USA 75: 1433

[1978] ), or a kanamycin resistance selection marker.

[0139] In some embodiments, more than one copy of a nucleic acid sequence of the present invention is inserted into the host cell to increase production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene with the nucleic acid sequence where cells containing amplified copies of the selectable marker gene, and thereby additional copies of the nucleic acid sequence, can be selected for by cultivating the cells in the presence of the appropriate selectable agent.

[0140] Many of the expression vectors for use in the present invention are commercially available. Suitable commercial expression vectors include, but are not limited to. Novagen’s pET® E. coll T7 expression vectors (Millipore Sigma) and the p3xFLAGTM™ expression vectors (Sigma- Aldrich Chemicals). In some embodiments, a pET30a expression vector is used. Other suitable expression vectors include, but are not limited to, pBluescriptll SK(-) and pBK-CMV (Stratagene). and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen) or pPoly (See e.g., Lathe et al., Gene 57:193-201

[1987] ).

[0141] Thus, in some embodiments, a vector comprising a sequence encoding at least one variant carboxylesterase is transformed into a host cell in order to allow propagation of the vector and expression of the variant carboxylesterases. In some embodiments, the transformed host cell described above is cultured in a suitable nutrient medium under conditions permitting the expression of the variant carboxylesterases. Any suitable medium useful for culturing the host cells finds use in the present invention, including, but not limited to minimal or complex media containing appropriate supplements. In some embodiments, host cells are grown in HTP media. Suitable media are available from various commercial suppliers or may be prepared according to published recipes (e.g., in catalogues of the American Type Culture Collection).Host Cells for Expression of Carboxylesterases

[0142] In another aspect, the present disclosure provides a host cell comprising at least one polynucleotide encoding at least one carboxyl esterase of the present disclosure, the polynucleotide being operatively linked to one or more control sequences for expression of the at least one carboxylesterase in the host cell. Host cells suitable for use in expressing the carboxylesterase polypeptides encoded by the expression vectors of the present disclosure are well known in the art and include but are not limited to, bacterial cells, such as E. coli, Vibrio fluvialis, Streptomyces and Salmonella typhimurium cells; fungal cells, such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC Accession No. 201178)). In various embodiments, the disclosed host cells are E. coli cells. Exemplar)’ host cells include various E. coli strains (e.g. W3110 (AfhuA) and BL21). Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis , or markers, which confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, and or tetracycline resistance. Appropriate culture mediums and growth conditions for the abovedescribed host cells are well known in the art.

[0143] Bacterial host cells for use in expressing the polypeptides encoded by the expression vectors of the present disclosure are well know n in the art and include but are not limited to, E. coli, B. subtilis, B. licheniformis, B. megaterium, B. stearothermophilus, B. amyloliquefaciens , Klebsiella aerogenes, Lactobacillus kejir, Lactobacillus brevis. Lactobacillus minor. Streptomyces and Salmonella typhimurium cells. In some embodiments, the host cell or cell line is Escherichia coli BL21 or BL21(DE3).

[0144] Many prokaryotic strains that find use in the present disclosure are readily available to the public from a number of culture collections such as American Type Culture Collection (ATCC), Deutsche Sammlung von Mikroorganismen und Zellkulturen GmbH (DSM), Centraalbureau Voor Schimmelcultures (CBS), and Agricultural Research Sendee Patent Culture Collection. Northern Regional Research Center (NRRL).

[0145] In some embodiments, host cells are genetically modified to have characteristics that improve protein secretion, protein stability and / or other properties desirable for expression and / or secretion of a protein. Genetic modification can be achieved by genetic engineering techniques and / or classical microbiological techniques (e.g., chemical or UV mutagenesis and subsequent selection). Indeed, in some embodiments, combinations of recombinant modification and classical selection techniques are used to produce the host cells. Using recombinant technology, nucleic acid molecules can be introduced, deleted, inhibited or modified, in a manner that results in increased yields of carboxylesterase variant(s) within the host cell and / or in the culture medium. In one genetic engineering approach, homologous recombination is used to induce targeted gene modifications by specifically targeting a gene in vivo to suppress expression of the encoded protein. In alternative approaches, siRNA, antisense and / or ribozy me technology find use in inhibiting gene expression. A variety of methods are known in the art for reducing expression of protein in cells, including, but not limited to deletion of all or part of the gene encoding the protein and site-specific mutagenesis to disrupt expression or activity of the gene product. (See e.g., Chaveroche et al., NUCL. ACIDS RES., 28:22 e97

[2000] ; Cho et al., MOLEC. PLANT MICROBE INTERACT., 19:7-15

[2006] ; Maruyama and Kitamoto, BIOTECHNOL LETT., 30: 1811-1817

[2008] ; Takahashi et al., MOL. GEN. GENOM., 272: 344-352

[2004] ; and You et al., Arch. Microbiol. ,191:615-622

[2009] . all of which are incorporated by reference herein).Random mutagenesis, followed by screening for desired mutations also finds use (See e.g., Combier et al., FEMS MICROBIOL. LETT., 220:141-8

[2003] ; and Firon et al., EUKARY. CELL 2:247-55

[2003] , both of which are incorporated by reference).

[0146] Introduction of a vector or polynucleotides for expression of the carboxylesterase into a host cell can be accomplished using any suitable method known in the art, including but not limited to calcium phosphate transfection, DEAE-dextran mediated transfection, PEG-mediated transformation, electroporation, or other common techniques known in the art, which include electroporation, biolistic particle bombardment, liposome mediated transfection, calcium chloride transfection, and protoplast fusion.

[0147] In some embodiments, the engineered host cells (i.e., ‘‘recombinant host cells”) of the present disclosure are cultured in conventional nutrient media modified as appropriate foractivating promoters, selecting transformants, or amplifying the carboxylesterase polynucleotide. Culture conditions, such as temperature, pH and the like, are those previously used with the host cell selected for expression, and are well-known to those skilled in the art. As noted, many standard references and texts are available for the culture and production of many cells, including cells of bacterial origin.

[0148] In some embodiments, cells expressing the carboxylesterase of the disclosure are grown under batch or continuous fermentation conditions. Classical “batch fermentation” is a closed system, wherein the compositions of the medium are set at the beginning of the fermentation and is not subject to artificial alternations during the fermentation. A variation of the batch system is a “fed-batch fermentation” that also finds use in the present invention. In this variation, the substrate is added in increments as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit the metabolism of the cells and where it is desirable to have limited amounts of substrate in the medium. Batch and fed-batch fermentations are common and well known in the art. “Continuous fermentation” is an open system where a defined fermentation medium is added continuously to a bioreactor and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the cultures at a constant high density where cells are primarily in log phase growth. Continuous fermentation systems strive to maintain steady state growth conditions. Methods for modulating nutrients and growth factors for continuous fermentation processes as well as techniques for maximizing the rate of product formation are well know n in the art of industrial microbiology'.

[0149] More than one copy of a nucleic acid sequence of the present invention may be inserted into the host cell to increase production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene with the nucleic acid sequence where cells containing amplified copies of the selectable marker gene, and thereby additional copies of the nucleic acid sequence, can be selected for by cultivating the cells in the presence of the appropriate selectable agent.

[0150] In some embodiments of the present disclosure, cell-free transcription and translation systems find use in producing the carboxylesterases. Several systems are commercially available, and the methods are well-known to those skilled in the art.Methods of Producing Carboxylesterases

[0151] In some embodiments, the carboxylesterases of the present disclosure are obtained, evolved, or derived from a bacterial wild-type enzyme. Evolution (e.g., directed evolution) may be used to identify polypeptides of the present disclosure. For example, in some embodiments, to make a carboxylesterase polypeptide of the present disclosure, a carboxylesterase-encoding polynucleotide was obtained (or derived) from Roseibacillus sp. In some embodiments, this parent polynucleotide sequence is codon-optimized to enhance expression of the carboxylesterase in a specified host cell. In the Examples, a parental wild-type polynucleotide sequence was codon optimized for A’, coli expression, and its 3' signal peptide-encoding sequence removed, to afford SEQ ID NO: 1. This carboxylesterase sequence was cloned into an expression vector, placing the expression of the carboxylesterase gene under the control of the lac promoter under control of the lac repressor. Clones expressing the active carboxylesterase in E. coli were identified, and the genes were sequenced to confirm their identity.

[0152] The carboxylesterase of the disclosure may be obtained by subjecting the polynucleotide encoding the parent sequence to mutagenesis and / or directed evolution (DE) methods. Any DE technique may be used, including cry stal structure-guided library design, single-site-saturation mutagenesis (SSM) library generation, or combinatorial library generation, or a combination of these. An exemplary directed evolution technique is mutagenesis and / or DNA shuffling as described in Stemmer, 1994, Proc. Natl. Acad. Sci. USA 91 : 10747-10751; WO 95 / 22625; WO 97 / 20078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767 and U.S. Pat. No. 6,537,746. Other directed evolution procedures that can be used include, among others, staggered extension process (StEP), in vitro recombination (Zhao et al.. 1998, Nat. Biotechnol. 16:258- 261), mutagenic PCR (Caldwell et al., 1994, PCR METHODS APPL. 3:S136-S140), and cassette mutagenesis (Black et al., 1996, Proc. Natl. Acad. Sci. USA 93:3525-3529).

[0153] Exemplary crystal structure-guided library design techniques may make use of computational modeling and / or may be rational or semi-rational. Examples of such techniques include molecular dynamics (MD) simulations, including in silico MD, Rosetta de novo design (e.g., trRosetta), Schrodinger toolbox and Glide modeling. TrRosetta is a deep neural net-based protein structure prediction method, offers a path forw ard for predicting 3D structures when there is a lack of close homologue structures, which takes advantage of multiple sequence alignment (MSA), where derived features from the MSA are fed into a deep neural network to predict interresidue geometries, including distance and orientations. Glide modeling (Friesner et al. J. Med. Chem. 2004, 47(7), 1739-1749) may be used to perform substrate docking or product docking.

[0154] The clones obtained following mutagenesis treatment are screened for carboxylesterases exhibiting one or more desired improved enzy me properties. These desired improved enzyme properties may be selected by applying each property7, one-by-one as a selection pressure during iterative rounds of evolution. For instance, improved enzy me activity may be applied as a selection pressure during one or more rounds of evolution. Additional selection pressures that may be applied include improved thermostability, co-solvent tolerance, and reduced formation of undesired by-products. For example, improved thermostability was applied as a selection pressure during evolution of the disclosed enzymes, as described in the Examples. As another example, reduced formation of hydrolysis by-product compound 3 was applied as a selection pressure during evolution of the disclosed enzymes. Co-solvent tolerance may be applied as a selection pressure to mimic potential industrial process conditions.

[0155] Measuring enzyme activity from the expression libraries can be performed using standard chemistry analytical techniques for measuring substrates and products such as ultraperformance liquid chromatography-mass spectrometry (UPLC-MS). The reaction may also be run under conditions where the carboxylesterase is the yield-limiting catalyst, such that a doubling or halving in concentration of the carboxylesterase will result in a doubling or halving of the yield of product observed at a given timepoint.

[0156] Where the improved enzy me property desired is thermostability, enzyme activity may be measured after subjecting the enzyme preparations to a defined temperature and measuring the amount of enzy me activity7remaining after heat treatments. Any suitable approach may be used, e.g., differential scanning colorimetry (DSC) a biochemical assay, or spectroscopy. Clones containing a polynucleotide encoding carboxylesterases are then isolated, sequenced to identify the nucleotide sequence changes (if any), and used to express the enzyme in a host cell.

[0157] Where the sequence of the polypeptide is known, the polynucleotides encoding the enzyme can be prepared by standard solid-phase methods, according to known synthetic methods. In some embodiments, fragments of up to about 100 bases can be individually- synthesized, then joined (e.g., by enzymatic or chemical litigation methods, or polymerase mediated methods) to form any desired continuous sequence. For example, polynucleotides and oligonucleotides of the invention can be prepared by chemical synthesis using, e.g., the classical phosphoramidite method described by Beaucage et al., 1981, Tet. Lett. 22: 1859-69, or the method described by Matthes et al., 1984, EMBO J. 3:801-05, e.g., as it is typically practiced in automated synthetic methods. According to the phosphoramidite method, oligonucleotides are synthesized, e.g., in an automatic DNA synthesizer, purified, annealed, ligated and cloned in appropriate vectors. In addition, essentially any nucleic acid can be obtained from any of avariety of commercial sources, such as Integrated DNA Technologies (IDT). Genscript, The Midland Certified Reagent Company, Midland, Tex., The Great American Gene Company, Ramona, Calif., ExpressGen Inc. Chicago, Ill., Operon Technologies Inc., Alameda, Calif., and many others.

[0158] Carboxylesterase enzymes expressed in a host cell can be recovered and isolated from the cells and / or the culture medium using any one or more of the well-known techniques for protein purification, including, among others, lysozyme treatment, sonication, filtration, saltingout, ultra-centrifugation, and chromatography. Suitable solutions for lysing and the high efficiency extraction of proteins from bacteria, such as E. coli, are commercially available under the trade name CELLYTIC B® from Sigma- Aldrich.

[0159] Chromatographic techniques for isolation of the carboxylesterase polypeptide include, among others, reverse phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. Conditions for purifying a particular enzyme will depend, in part, on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and will be apparent to those having skill in the art.

[0160] In some embodiments, affinity techniques may be used to isolate the improved carboxylesterase enzymes. For affinity chromatography purification, the protein sequence can be tagged with a recognition sequence to enable purification. Common tags include cellulose- binding domains, poly His-tags (e.g., 6xHis tags), di-His chelates, FLAG-tags and many others that will be apparent to those having skill in the art. Antibodies can also be used as affinity purification reagents. Any antibody that specifically binds the carboxylesterase polypeptide may be used.Methods of Generating Cyclized Oligopeptides

[0161] Provided herein are methods (or processes) of generating cyclized oligopeptides via one or more carboxylesterase-catalyzed reactions. These methods comprise performing a reaction wi th any of the disclosed engineered polypeptides. These methods may comprise the conversion of a linear tetrapeptide into a cyclized tetrapeptide (through amide bond formation and acyl transfer). In various aspects, the product of these methods is a cyclic tetrapeptide.

[0162] Any of these methods may be performed in a medium containing cofactor magnesium ion (Mg2+), for instance in a magnesium phosphate or magnesium chloride salt. Any of these methods may be performed in a medium containing sodium phosphate, a lysozyme and / or polymyxin B sulfate. In some embodiments, the medium contains co-solvent acetonitrile, e.g.,15%. 20%. 25%. or 30% acetonitrile. In some embodiments, the medium is maintained at a neutral pH, e g., a pH of 7.4 or a pH of 7. In some embodiments, the medium is maintained at a temperature of 20 °C.

[0163] In some aspects, provided are methods for preparing compound 2. the method comprising contacting compound 1 in the presence of Mg21and any of the disclosed polypeptides. These methods may produce compound 3 in reduced amounts relative to a corresponding method in which compound 1 is contacted with the polypeptide of SEQ ID NO: 2.

[0164] In some embodiments, the disclosed methods exhibit higher (1) co-solvent tolerance and / or (2) thermostability relative to a corresponding method in which Compound 1 is contacted with the polypeptide of SEQ ID NO: 2.EXAMPLESABBREVIATIONSExample 1Synthesis, Expression, and Assay of carboxyesterases with macrocyclization activity

[0165] This example describes methods to synthesize, codon optimize, and assay carboxylesterase enzymes, and the composition of optimized enzymes.Gene synthesis and construction of expression strain:

[0166] The wild-type carboxylesterase polypeptide UniProt ID: A0A2E5D605 9BACT belonging to an unclassified Roseibacillus species) was identified by a BLASTP search using a previously reported promiscuous wild-type hydrolase EstCEl (Muller et al., Angew. Chem. Int. Ed. (2021) 60, 2013). Using the DeepSig algorithm (Savojardo et al., Bioinformatics. 2018, 34, 1690), the first 15 amino acids were predicted to be a signal peptide region of carboxylesterase and were therefore excluded from the design for gene synthesis to promote cytoplasmic expression within E. coli. The encoding carboxylesterase polynucleotide without the signal peptide encoding region was codon optimized for expression in E. coli and synthesized with a sequence encoding an N-terminal 6xHistidine tag as SEQ ID NO: 1. SEQ ID NO: 1 is 74.6% identical to the wild-type enzy me ECE20-encoding polynucleotide sequence (GenBank ID: PBAI01000020. 1). SEQ ID NO: 1 was cloned into pET30 vector under the control of T7 promoter for expression and transformed into E. coli strain (BL21(DE3)).HTP Growth, Expression, and Lysate Preparation:

[0167] The transformants expressing SEQ ID NO:1 as carboxylesterase polypeptide (SEQ ID NO: 2) were picked and grown in Luria-Bertani Broth medium with kanamycin (30 pg / ml) and 1% w / v glucose in 96-well plates with shaking (200 rpm) at 30 °C overnight. Subsequently, the overnight cultures were diluted for expression to optical density of 0.05, measured at 600 nm (ODeoo), with auto-induction medium (ZYM-5052, Teknova) and antibiotic (kanamycin 30 pg / ml) and grown for 18-20 hours at 30 °C. Post-induction, cells were centrifuged and cell pellet was resuspended in the lysis buffer (50 mM sodium phosphate buffer pH 7.4, 0.5 mg / mL lysozyme, 0.5 mg / mL polymixin B sulfate, 1 U / mL DNAse I, 1 mM magnesium sulfate) at 25 °C with shaking at 1000 rpm for 1 hour. The cell lysate was clarified by centrifugation (4000 x g, 15 minutes). The clarified cell lysates were used in the following well-plate enzy matic reactions. In parallel to the dilution of overnight cultures for expression, bacterial cultures were diluted with 25% v / v glycerol and frozen at -80 °C for storage. Overnight cultures were also used as PCR templates to amplify the coding region for the carboxylesterase, which were then subjected to sequencing to confirm or determine the sequence identity of the carboxylesterase in each well of the 96-well plates.Assay method for Carboxylesterase activity:

[0168] For a 50 pL volume reaction (scheme below), 25 pL of 50 mM sodium phosphate buffer (pH 7.4) was mixed with 5 pL of substrate (linear tetrapeptide, compound 1, 5 mg / mL dissolved in 50% DMSO) and 20 pL of clarified lysate. Reaction mixtures were incubated at 30 °C with shaking at 600 rpm for 18 hours. The reactions were quenched with 150 pL of acetonitrile and mixed thoroughly. Samples were further filtered through 0.22 pm filters before UHPLC-UV analysis.Reaction scheme 1:Compound NG: ' Compound NO 2 Compound NO 3Because the engineered ECE20 carboxylesterase variants demonstrate macrocyclizing activity, they are also referred to as “macrocyclase” or ■■NF-macrocyclase."UHPLC-UV analytical method for carboxylesterase activity:

[0169] The conversion of compound 1 to compound 2 in the enzymatic reaction was determined by UHPLC-UV using a Thermo Hypersil Gold PFP column (3.0 x 50 mm, 1.9 pm particle size). Mobile phase 0.1% difluoroacetic acid (A) / acetonitrile (B) with the gradient program described below, injection volume 1 pL, flow rate 1 mL min'1, detection wavelength 280 nm, and column temperature of 40 °C. Retention time of the substrate (compound 1): 1.093 minutes; Retention time of the desired product (compound 2): 0.876 minutes; Retention time of the hydrolysis product (Compound 3): 0.697 minutes.UHPLC gradient ProgramIdentification of starting point for evolution:

[0170] Activity was detected for the ECE20 enzyme of SEQ ID NO: 2 catalyzing macrocyclization of tetrapeptides (conversion of compound 1 to compound 2) with the desired regiochemistry. However, the conversion was very low even under high percentage of cellular lysate loading, limiting industrial applications. In addition, SEQ ID NO: 2 also generates a hydrolysis product (compound 3), reducing the reaction yield. Therefore, it was determined that protein engineering should be initiated to improve the activity and selectivity (acyl transfer over hydrolysis) of SEQ ID NO: 2.Table A, Compound structures:

[0171] Substrate compound 1 is a modified Phenylalanine-Tryptophan-Proline-Threonine tetrapeptide (SEQ ID NO: 43) that contains modified proline, tryptophan, and phenylalanine residues. Product compound 2 is a cyclized version of compound 1. Both compounds have a molecular weight of 793.94 g / mol.Directed evolution strategy' and summary of results

[0172] The activity (e.g., conversion, total turnover number) of the wild-type carboxylesterase enzyme, SEQ ID NO: 2, was insufficient for industrial macrocyclization of tetrapeptide.Therefore, a directed evolution strategy was developed to engineer SEQ ID NO: 2 to improve the activity, thermostability, organic solvent tolerance, and selectivity (acyl transfer over hydrolysis). To prepare for high-throughput screening of enzyme variant libraries, protein expression conditions were first optimized to maximize soluble protein expression. Enzy matic reactions (Reaction scheme 1) were also scaled down to the 96-well microtiter plate format. A 5-minute UHPLC method was developed to quantify the conversions to enable library-scale screening of carboxylesterase variants, as described above.

[0173] In the directed evolution campaign, single-site-saturation mutagenesis (SSM) libraries were designed, built, expressed, and screened for desired properties relative to the starting enzyme (Table 1). SSM libraries in each round of evolution were generated using splicing byoverlap extension (SOEing) PCR methods described in Ho et al., Gene 1989, 77(1), 51-59. In short, mutations at the designated positions w ere incorporated by degenerate oligonucleotides through overlapping PCRs. The full-length gene-of-interest region was further amplified and assembled into the expression vector (pET30a) by Gibson Assembly (Gibson et al. Nat. Methods 2009, 6(5). 343-345). The combinatorial libraries were built following the instructions from the QUIKCHANGE® Lightning Multi Site-Directed Mutagenesis kit (Agilent Technologies). Both mutagenesis libraries were transformed into the BL21(DE3) E.coli strains by electroporation and plated on LB agar plates with selection (1% w / v glucose, 30 pg / mL kanamycin).Table 1. Summary of Carboxylesterase evolution

[0174] To prioritize the sites for SSM library designs, a structural model of SEQ ID NO: 2 was built using the Protein Data Bank (PDB) template 4IVK, which had 52.5% sequence similarity. Schrodinger toolbox was used to design the homology model, which was further refined by adding hydrogens using PROPKA (Olsson et al. J. Chem. Theory Comput. 2011. 7(2), 525-537) and by running a restrained minimization to converge heavy atoms to a maximum root mean square deviation of 0.30 A. Substrate docking was performed using Glide (Friesner et al. J. Med. Chem. 2004, 47(7), 1739-1749) and the best pose was selected after visually inspecting a pool ofdocked poses rank ordered by their docking energy scores and filtering them based on the distances between the reaction site and the catalytic residues.

[0175] The protein homology7model was partitioned into four regions: active site, surfacel, surface , and the remaining residues. The active site shell residues were selected using a distance criterion, i.e., all residues within 10.5 A of the docked substrate, resulting in 96 sites. The remaining residues were rank ordered by their solvent accessibilities and grouped into sets of 96 sites. The top two sets were classified as surface sitel (positions with solvent accessible surface area > 64 A) and site2, (positions with solvent accessible surface area <64 A and > 8 A) consisting of high surface accessibility7sites, whereas the remaining were classified as the remaining residues. In the initial 4 rounds, each tier per round was sequentially targeted to screen all the residues beginning with the active site proximal, followed by surface residues, and finally to the remaining residues. In subsequent rounds, residues were prioritized to target either based on structural insights or MD (molecular dynamics) simulation (see Table 1). Beneficial mutations identified from these SSM libraries were recombined into combinatorial libraries for subsequent screening in the following rounds. The best variant served as the backbone for the next round of mutagenesis. Additional details regarding the rounds of evolution are provided below in Examples 3 to 14.

[0176] From round 1 to round 3 of evolution (Table 1), a series of mutations (I159Y, Y154M, N156P, G170W, L172C, L199G, N132R, K352T, M368A) were identified relative to SEQ ID NO: 2 that improved product formation under screening conditions by more than 50-fold. As the activity continued to improve through evolution, acetonitrile was included as a co-solvent to engineer variants with higher tolerance to organic solvent, which reflects potential industrial process conditions. In rounds 4 and 5. additional mutations (R133M, T138R. M154N, KI 84V. A191D, M327Y, VI 151, T140P, C147R, G158E, F373Y) were identified relative to SEQ ID NO: 8, which improved product formation under organic solvent exposure by more than 80-fold. From round 6 to round 8, heat-treatment of enzy mes prior to activity assays was also included to engineer more thermostable variants. Mutations (R330S, P160L, I162G, A272H. S350D, R17Y. F134Q, A165P, T265S, D266L, P325S) were further identified relative to SEQ ID NO: 12, which improved product formation under screening conditions by more than 20-fold. From round 9 to round 13, improvement of acyltransferase activity7over hydrolysis activity was prioritized to increase the overall yield of desired product Compound 2. Mutations (G199S, Q134F, L266D. H166W, G281S, L343F, A368L, I115L, I157F, W170L, C172F, L187Q, T10K, G162I, T246H, H272A, F343L) were further identified relative to SEQ ID NO: 18, which improved selectivity ([Compound 2] / [Compound 3]) under screening conditions by more than 10-fold (see Table9. 1). Ultimately, this evolution campaign led to the engineering of SEQ ID NO: 28, which had improved activity, organic solvent tolerance, thermo-stability and selectivity in comparison to the wild-type parent (SEQ ID NO: 2).Example 2 (Rdl evolution of SEQ ID NO: 2)Enzyme variants of SEQ ID NO: 2

[0177] In this example, engineering of SEQ ID NO: 2 for improved activity' and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites were subjected to mutagenesis resulting in library 1.1 (Example 1, Table 1). These libraries were plated to form single colonies, which were picked, grown, expressed, and screened using the high-throughput grow th and expression methods described in Example 1. The high-throughput assay conditions are described below.HTP analytical methods for enzyme activity in round 1:

[0178] The conversion in the enzymatic reaction was determined by the method described in Example 1 or using a Waters Acquity BEH Cl 8 column (2.1 x 50 mm, 1.7 pm particle size). Mobile phase (A) 0.1% (v / v) formic acid in water, (B) 0.1% (v / v) formic acid in acetonitrile, 1 pL injection volume, flow rate 0.75 mL min1, detection wavelength of 210 nm and 280 nm at 55 °C column temperature. Retention time of the substrate (Compound 1): 2.893 minutes; Retention time of the desired product (Compound 2): 2.624 minutes; Retention time of the undesired hydrolysis product (Compound 3): 2.021 minutes.Gradient for achiral HPLC method.HTP activity assay for Carboxylesterase variants:

[0179] For a 50 pL volume reaction. 30 pL of the 50mM Sodium phosphate buffer (pH 7.4) was mixed with 10 pL tetrapeptide (Compound 1, 5 mg / mL dissolved in 50% DMSO) and 10 pL of cellular lysate. Reaction mixtures were incubated at 30°C with shaking at 600 rpm for 18 hours.Production of Shake Flask Powders (SFP):

[0180] Based on the analysis of the HTP assay results using the methods described above, variants of SEQ ID NO: 1 / 2 were selected for larger scale analysis and further characterization. About 10 pL of glycerol stock or a freshly streaked single colony for each variant was inoculated into 30 mL Luria Broth media with 30 gg / mL kanamycin and 1% w / v glucose in a 125 mL baffled flask and incubated at 30 °C with shaking (250 rpm) for 18 hours. The next day, the overnight saturated culture was diluted to an initial optical density of 0.05 (measured at 600 nm (ODeoo)) with Terrific Broth supplemented with 30 pg / mL kanamycin in 250 mL volume and grown at 37 °C with shaking (250 rpm). When the growth reached ODeoo of 0.5, IPTG was added to the culture to a final concentration of 1 mM to induce protein production and the culture was further incubated for 20 hours at 30 °C. Cells were then collected by centrifugation and resuspended in 5 mL of buffer (20mM Triethanolamine, pH 7.5) per gram of cell pellet mass. The cell pellet was resuspended by shaking (250 rpm) at 8 °C for 25 minutes. Cells were lysed by the microfl uidizer (LM-10) at the pressure of 16,000 psi. The resulting lysate was clarified by centrifugation at 22,000xg for 45 minutes at 4 °C, and the supernatant was subsequently frozen and lyophilized to generate enzy me powders.

[0181] Engineered polypeptides with >2-fold conversion relative to the parent polypeptide are listed in Table 2.1, and the activities were measured with the shake flask powder samples.Table 2.1 Variants and Conversion

[0182] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 2 and defined as: “+” = conversion at least 2-fold relative to reference polypeptide

[0183] Variants with mutations I159Y and G155P produced more macrocyclic peptide product (Compound 2) from the tetrapeptide substrate (compound 1), and these engineered carboxylesterase enzymes provide new biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations I159Y (SEQ ID NO: 4), had the preferable activity. Thus, the encoding polynucleotide (SEQ ID NO: 3) was selected for further directed evolution.Example 3 (Rd2 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 4

[0184] In this example, engineering of SEQ ID NO: 4 for improved activity toward macrocyclization and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites and dynamic loop regions were subjected to mutagenesis individually and combinatorically resulting in libraries 2.1 and 2.2 (Example 1, Table 1). These libraries were plated to form single colonies, which were picked, grown, and screened using the high- throughput growth and expression methods described in Example 1 and the HTP assay and analytical method described in Example 2.HTP assay for Carboxylesterase activity:

[0185] Engineered polypeptides with >5-fold conversion relative to the parent polypeptide are listed in Table 3.1, and the activities were measured with shake flask powder samples produced as described in Example 2.Table 3.1 Variant and Conversion

[0186] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 4 and defined as: “+” = conversion at least 5-fold relative to reference polypeptide

[0187] SEQ ID NO: 6 with mutations Y154M, N156P, G170W, L172C, L199G relative to SEQ ID NO: 4 produced more macrocyclic peptide (compound 2) from the tetrapeptide (compound 1), and this engineered carboxylesterase enzyme was useful as a biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations Y154M, N156P, G170W, L172C, L199G (SEQ ID NO: 6), had the highest activity in the assayed library. Thus, the encoding polynucleotide (SEQ ID NO: 5) was selected for further engineering.Example 4 (Rd3 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 6

[0188] In this example, engineering of SEQ ID NO: 6 for improved activity' toward macrocyclization and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites were subjected to mutagenesis individually and combinatorically resulting inlibraries 3.1 and 3.2 (Example 1, Table 1). These libraries were plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 1. The high-throughput assay conditions are described below. The production of shake flask powders and EITP analytical methods were described in Example 2.HTP assay for Carboxylesterase activity:

[0189] For a 50 pL volume reaction, 30 pL of the 50mM Sodium phosphate buffer (pH 7.4) was mixed with 10 pL substrate (tetrapeptide Compound 1 at 25 mg / mL dissolved in 25% v / v DMSO in water) and 10 pL of clarified cellular lysate. Reaction mixtures w ere incubated at 40 °C with shaking at 600 rpm for 18 hours.

[0190] Engineered polypeptides with >3-fold conversion relative to the parent polypeptide are listed in Table 4. 1 as measured with the shake flask powder samples.Table 4.1 Variant and Conversion

[0191] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 6 and defined as: “+” = conversion at least 3-fold relative to reference polypeptide

[0192] Engineered variant SEQ ID No: 32 with mutations N132R, N132K, R133M, G158E, I162W, L187K, K352T, M368A produced more macrocyclic peptide (compound 2) from the tetrapeptide (compound 1), and this engineered carboxylesterase enzyme provided a new biocatalytic reagent for use in new7methods for the macrocyclization reaction of polypeptides. The variant with mutations N132R, K352T, M368A (SEQ ID NO: 8) had the highest activity' in the library. Thus, the encoding polynucleotide (SEQ ID NO: 7) was selected for further directed evolution.Example 5 (Rd4 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 8

[0193] In this example, engineering of SEQ ID NO: 8 for improved activity’ toyvard macrocyclization yvith organic co-solvent ACN and the resulting improved variants are described. Engineering via directed evolution yvas carried out by constructing libraries of variant genes in which positions associated with solvent accessibility and active sites were subjected to mutagenesis individually and combinatorically resulting in libraries 4.1 and 4.2 (Example 1,Table 1). These libraries were plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 1. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 2.HTP assay for Carboxylesterase activity:

[0194] For a 50 pL volume reaction, 40 pL of assay mixtures (75 mM sodium phosphate buffer pH 7.4, 6.25 mg / mL of tetrapeptide (Compound 1), and 18.75% v / v of acetonitrile) was added with 10 pL of clarified cellular lysate, resulting in 15% v / v acetonitrile concentration in the reactions. Reaction mixtures were incubated at 20 °C with shaking at 600 rpm for 18 hours.

[0195] Engineered polypeptides with >7-fold conversion relative to the parent (reference) polypeptide are listed in Table 5.1 as measured with shake flask powder samples.Table 5.1 Variants and Conversion

[0196] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 8 and defined as: “+” = conversion at least 7-fold relative to reference polypeptide

[0197] Engineered variant Seq ID No: 10 with mutations R133M. T138R, M154N, KI 84V, Al 9 ID, M327Y produced more macrocyclic peptide (compound 2) from the tetrapeptide (compound 1) with the organic co-solvent, and this engineered carboxylesterase enzyme provide a new biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations R133M, T138R, M154N, KI 84V, A191D, M327Y (SEQ ID NO: 10), had the highest activity. Thus, the encoding polynucleotide (SEQ ID NO: 9) w as selected for further directed evolution.Optimization on the fermentation of the Carboxylesterase

[0198] In order to utilize the macrocyclase in macrocyclization of peptides at scale, the production of the macrocyclase at reduced cost is desirable. Therefore, the expression of SEQ ID NO: 10 was examined under fermentation conditions in E.coli (BL21(DE3)). SEQ ID NO: 10 was fermented at 5L scale with complex medium. While well -expressed at shake flask scale (see FIG. 1), the expression of the macrocyclase was poor in the 5L fermentation (see FIG. 2). Therefore, the fermentation conditions were optimized by switching the medium from complex medium to chemical defined medium (CDM). Fermentation with the optimized conditions resulted in ahigher expression level of SEQ ID NO: 10 (see FIG. 3). which is useful for large-scale production of engineered carboxylesterase for industrial applications.Example 6 (Rd5 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 10

[0199] In this example, engineering of SEQ ID NO: 10 for improved activity toward macrocyclization with organic co-solvent ACN and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with improved activity with or without co-solvent from previous rounds were subjected to mutagenesis combinatorically. resulting in library 5. 1 (Example 1, Table 1). This library was plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 1. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 2.Assay for Carboxylesterase activity:

[0200] In the high-throughput screening, 40 pL of assay mixture (225 mM sodium phosphate buffer, 112.5 mM HEPES, 31.2 mM magnesium chloride, 6.25 mg / mL of tetrapeptide Compound 1, 43.75% v / v acetonitrile in water, pH 7) was mixed with 10 pL of 10-fold diluted clarified cellular lysate, resulting in 35% v / v acetonitrile in the reactions. Reactions were incubated at 20 °C with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate concentration in 200 mM sodium phosphate buffer pH 7 and 15% v / v acetonitrile with 15 % (wt enzyme / wt substrate) of carboxylesterase at 30°C with shaking at 600 rpm for 12 hours.

[0201] The engineered polypeptide with >8-fold conversion relative to the parent polypeptide under these conditions is listed in Table 6.1, as measured with the shake flask powder samples.Table 6.1 Variants and Conversion

[0202] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 10 and defined as: “+” = conversion at least 8-fold relative to reference polypeptide

[0203] Engineered variant SEQ ID NO: 12 with mutations VI 151, T140P, C147R, G158E, F373Y produced more macrocyclic peptide (compound 2) from the tetrapeptide (compound 1) with the organic co-solvent, and this engineered carboxylesterase enzyme provided a newbiocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides with organic co-solvent. The variant with mutations VI 151, T140P, C147R, G158E, F373Y (SEQ ID NO: 12), had the highest activity. Thus, the encoding polynucleotide (SEQ ID NO: 11) was selected for further directed evolution.Example 7 (Rd6 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 12

[0204] In this example, engineering of SEQ ID NO: 12 for improved activity toward macrocyclization with organic co-solvent (solvent tolerance) and heat treatment (thermostability) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which 96 positions associated with solvent accessible / surface residues were subjected to mutagenesis resulting in library 6.1 (Example 1, Table 1). This library was plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 1. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 2.Assay for Carboxylesterase activity:

[0205] In the high-throughput screening, 40 pL of assay mixture (225 mM sodium phosphate buffer, 112.5 mM HEPES, 31.2 mM magnesium chloride, 6.25 mg / mL of tetrapeptide Compound 1, 25% v / v acetonitrile, pH 7) was mixed with 10 pL of 10-fold diluted clarified cellular lysate that had been pre-incubated at 46.2 °C for 1 hour, resulting in 20% v / v acetonitrile in the reactions. Reaction mixtures were incubated at 20 °C with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate concentration in 200 mM sodium phosphate buffer pH 7, 30% v / v acetonitrile, 100 mM HEPES, and 25 mM magnesium chloride with 5% wt enzyme / wt substrate at 30 °C with shaking at 600 rpm for 14 hours.

[0206] Engineered polypeptides with >2-fold conversion relative to the parent polypeptide are listed in Table 7. 1, as measured with the shake flask powder samples.Table 7.1 Variants and Conversion

[0207] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 12 and defined as:‘’+” = conversion at least 2-fold relative to reference polypeptide

[0208] Engineered variants with SEQ ID NO: 14 with mutation R330S produced more macrocyclic peptide (compound 2) from the tetrapeptide (compound 1), and this engineered carboxylesterase enzyme serves as a new biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides in the presence of organic co-solvent and at various temperatures. The variant with mutations R330S (SEQ ID NO: 14), had the highest activity in library 6.1. Thus, the encoding polynucleotide (SEQ ID NO: 13) was selected for further directed evolution.Example 8 (Rd7 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 14

[0209] In this example, engineering of SEQ ID NO: 14 for improved activity toward macrocyclization and improved stability (e.g.. stability in the presence of organic co-solvent and elevated temperature) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with solvent accessible surface residues were subjected to mutagenesis individually and combinatorically resulting in libraries 7.1 and 7.2 (Example 1, Table 1). These libraries were plated to form single colonies, which were picked, grown, and screened using the high- throughput grow th and expression methods described in Example 1. The high-throughput assay conditions are described below'. The production of shake flask powders and HTP analy tical methods were described in Example 2.Assay for Carboxylesterase activity:

[0210] In the high-throughput screening, 40 pL of assay mixtures (225 mM sodium phosphate buffer, 112.5 mM HEPES, 31.2 mM magnesium chloride, 6.25 mg / mL of tetrapeptide Compound 1, 25% v / v acetonitrile, pH 7) was added with 10 pL of 10-fold diluted cellular lysate that was pre-incubated at 50 °C for 1 hour, resulting in 20% v / v acetonitrile in the reactions. Reactions were incubated at 20 °C with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate concentration in 200 mM sodium phosphate buffer pH 7, 15% v / v acetonitrile, 100 mM HEPES, and 25 mM magnesium chloride with 0.25% wt enz me / wl substrate at 30 °C with shaking at 600 rpm for 14 hours.

[0211] Engineered polypeptides with >4-fold conversion relative to the parent polypeptide are listed in Table 8.1, and the activities were measured with the shake flask powder samples.Table 8.1 Variants and Conversion

[0212] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 14 and defined as: ”+” = conversion at least 4-fold relative to reference polypeptide.

[0213] Engineered variants SEQ ID NO: 16 and 34 with mutations R17D, P160L, I162G, A272H, S350D produced more macrocyclic peptide (compound 2) from the tetrapeptide (compound 1), and these engineered carboxylesterase enzymes are new biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides with organic co-solvent and at high temperature. The variant with mutations P160L, I162G, A272H, S350D (SEQ ID NO: 16), had the highest activity in libraries 7.1 and 7.2. Thus, the encoding polynucleotide (SEQ ID NO: 15) was selected for further directed evolution.Example 9 (Rd8 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 16

[0214] In this example, engineering of SEQ ID NO: 16 for improved activity7toward macrocyclization and improved stability (e.g., stability in the presence of organic co-solvent and elevated temperature) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with improvements in activity or stability7in previous rounds were subjected to mutagenesis combinatorically resulting in library 8.1 (Example 1, Table 1). This library was plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 1. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 2.Assays for Carboxylesterase activity:

[0215] In the high-throughput screening, 40 pL of assay mixtures (225 mM sodium phosphate buffer, 112.5 mM HEPES, 31.2 mM magnesium chloride, 6.25 mg / mL of tetrapeptide (Compound 1), 37.5% v / v acetonitrile, pH 7) was added with 10 pL of 10-fold diluted clarified cellular lysate that was pre-incubated at 48 °C for 1 hour, resulting in 30% v / v of acetonitrileconcentration in the reactions. Reactions were incubated at 20 °C with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate concentration in 200 mM sodium phosphate buffer pH 7, 20% v / v acetonitrile, 100 mM HEPES, and 25 mM magnesium chloride with 0.5% wt enzyme / wt substrate of enzyme at 30 °C with shaking at 600 rpm for 14 hours.

[0216] Engineered polypeptides with >2-fold conversion relative to the parent polypeptide are listed in Table 9.1, and the activities were measured with the shake flask powder samples.Table 9.1 Variant and Conversion

[0217] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 16 and defined as: ”+” = conversion at least 2-fold relative to reference polypeptide.

[0218] Engineered variant SEQ ID NO: 18 with mutations R17Y, F134Q, A165P, T265S, D266L, P325S produced more macrocyclic peptide (compound 2) from the tetrapeptide (compound 1), and this engineered carboxylesterase enzyme provides a new biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides with organic co-solvent and at various temperatures. The variant with mutations R17Y, F134Q, A1 5P, T265S, D266L, P325S (SEQ ID NO: 18), had the highest activity in library 8.1. Thus, the encoding polynucleotide (SEQ ID NO: 17) was selected for further directed evolution.Example 10 (Rd9 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 18

[0219] In this example, engineering of SEQ ID NO: 18 for improved selectivity toward macrocyclization relative to hydrolysis and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites were subjected to mutagenesis combinatorically resulting in library 9.1 (Example 1, Table 1). This library was plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 1. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 2.Assays for Carboxylesterase activity:

[0220] In the high-throughput screening, 50 pL reactions (200 mM sodium phosphate buffer, 100 mM HEPES, 25 mM magnesium chloride, 25 mM tetrapeptide Compound 1, pH 7, 0.4% v / v clarified cellular lysates) were incubated at 20 °C with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate in 200 mM sodium phosphate buffer pH 7, 100 mM HEPES, and 25 mM magnesium chloride with 0.06 % wt enzyme / wt substrate at 30 °C with shaking at 600 rpm for 14 hours.

[0221] Engineered polypeptides with >1.2-fold relative to the parent polypeptide are listed in Table 10.1. and the activities were measured as shake flask powder samples.Table 10.1 Variants and Conversion

[0222] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 18 and defined as: ”+” = conversion at least 1.2-fold relative to the reference polypeptide

[0223] Engineered variant SEQ ID NO: 20 with mutations Q134F, G199S, L266D produced more macrocyclic peptide (compound 2) from the tetrapeptide (compound 1), and this engineered carboxylesterase enzyme provides a new biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides with industrially relevant reaction conditions. The variant with mutations Q134F, G199S, L266D (SEQ ID NO: 20), had the highest activity in library 9.1. Thus, the encoding polynucleotide (SEQ ID NO: 19) was selected for further directed evolution.Example 11 (Rd 10 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 20

[0224] In this example, engineering of SEQ ID NO: 20 for improved selectivity against hydrolysis of tetrapeptide (Compound 1) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions that were identified as beneficial in previous rounds were subjected to mutagenesis combinatorically and positions that had not been targeted in previous evolution rounds were subjected to mutagenesis individually, resulting in libraries 10.1 and 10.2. respectively (Example 1, Tablet). These libraries were plated to form single colonies, whichwere picked, grown, and screened using the high-throughput growth and expression methods described in Example 1. The production of shake flask powders and HTP analytical methods were described in Example 2.Assays for Carboxylesterase activity:

[0225] In the high-throughput screening, 50 pL reactions (200 mM sodium phosphate buffer, 100 mM HEPES, 25 mM magnesium chloride, 25 mM tetrapeptide compound 1, pH 7, 0.3% v / v clarified cellular lysates) were incubated at 15 °C with shaking at 600 rpm for 3 hours. For assays with SFP samples, reactions were performed at 37.5 mM tetrapeptide in 200 mM sodium phosphate buffer pH 7.5. 100 mM HEPES and with 0. 15 g / L of enzyme at 20°C with shaking at 600 rpm for 30 minutes.

[0226] Engineered polypeptides with >1.5-fold selectivity, defined as the ratio of the conversion % of substrate to the product (compound 2) over the conversion % of substrate to the by-product (compound 3), relative to the parent polypeptide, are listed in Table 11.1. The selectivity was measured with the shake flask powder samples.Table 11.1 Variants and Conversion

[0227] Levels of increased conversion w ere determined relative to the reference polypeptide of SEQ ID NO: 20 and defined as: "‘+” = selectivity at least 1.5-fold relative to reference polypeptide

[0228] Engineered variant SEQ ID NO: 22 with mutations H166W, G281S, L343F produced less hydrolysis by-products (compound 3) from the tetrapeptide (compound 1). and this engineered carboxylesterase enzyme provides a more selective biocatalytic reagent for use in new methods for the macrocyclization of polypeptides. The variant with mutations H166W, G281S, L343F (SEQ ID NO: 22), had the highest selectivity in libraries 10.1 and 10.2. Thus, the encoding polynucleotide (SEQ ID NO: 21) was selected for further directed evolution.Example 12 (Rd 11 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 22

[0229] In this example, engineering of SEQ ID NO: 22 for improved selectivity toward macrocyclization against hydrolysis of tetrapeptide (compound 1) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructinglibraries of variant genes in which positions associated with active sites and surface residues were subjected to mutagenesis individually resulting in libraries 11.1 and 11.2. These libraries were plated to form single colonies, which were picked, grown, and screened using the high- throughput growth and expression methods described in Example 1 and below. The high- throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described below and in Example 2.HTP Growth, Expression, and Lysate Preparation:

[0230] The colonies from libraries 11.1 and 11.2 were picked and grown in Luria-Bertani Broth medium with kanamycin (50 pg / mL) in 96 deep-well plates with shaking (400 rpm) at 30 °C overnight. Subsequently, the overnight cultures were diluted 1: 100 with Terrific Broth medium containing antibiotic (kanamycin 50 pg / mL) and grown to cellular optical density of 0.7, measured at 600 nm (ODeoo). Protein production w as induced by adding IPTG to 1 mM final concentration, and cells were grown for 20 hours with shaking (450 rpm) at 30 °C. After induction and expression, plates were centrifuged and pellets were resuspended in the lysis buffer (100 mM EIEPES buffer pH 7.5, 1 mg / mL lysozyme, 0.5 mg / mL polymixin B sulfate, 3 U / mL DNase I, 4 mM magnesium sulfate) at 25°C with shaking at 800 rpm for 2 hours. The cell lysate was clarified by centrifugation (4000 x g, 15 minutes).HTP Analytical method for Carboxylesterase activity:

[0231] The conversion in the enzymatic reaction was determined using a 2-minute UPLC-MS method in the primary screening and a longer 5-minute method in the retest screening. In the 2- minute method, an Acquity BEH Cl 8 column (2.1 x 50 mm, 1.7 pm particle size) was used, with mobile phase (A) 0.1% (v / v) formic acid in water. (B) 0.1% (v / v) formic acid in acetonitrile. 1 pl injection volume, flow rate 0.8 mL min’1, 55 °C column temperature, detection by MS detector (SIM ion: m / z=406.80 for the substrate (compound 1) and by-product (compound 3); SIM ion: m / z=397.80 for the product (compound 2)). Retention time of the substrate (Compound 1): 1.257 minutes; Retention time of the desired product (Compound 2): 1.13 minutes; Retention time of the undesired hydrolysis product (compound 3): 0.695 minutes.Gradient program

[0232] In the 5-minute method, an Acquity BEH Cl 8 column (2.1 x 50 mm, 1.7 pm particle size) was used, with mobile phase (A) 0.1% (v / v) formic acid in water, (B) 0.1% (v / v) formic acid in acetonitrile, 1 pl injection volume, flow rate 0.75 mL min'1, 55 °C column temperature, detection by MS detector (SIM ion: m / z=406.80 for the substrate (compound 1) and by-product (compound 3); SIM ion: m / z=397.80 for the product (compound 2)). Retention time of the substrate (compound 1): 3.1 minutes; Retention time of the desired product (compound 2): 2:8 minutes; Retention time of the undesired hydrolysis product (compound 3): 2.3 minutes.Gradient programAssays for Carboxylesterase activity:

[0233] In the high-throughput screening, 100 pL reactions (300 mM sodium phosphate buffer. 100 mM HEPES, 75 mM magnesium chloride, 25 mM tetrapeptide compound 1, pH 7, 0.5% v / v clarified cellular lysates) were incubated at 20 °C with shaking at 1000 rpm for 2 hours. For assays with SFP samples, reactions were performed at 25 mM substrate concentration in 300 mM sodium phosphate buffer pH 7.5, 100 mM HEPES, 75 mM magnesium chloride, and with 0. 1 g / L of enzyme at 20°C with shaking at 1000 rpm for 2 hours.Production of Shake Flask Powders (SFP):

[0234] Based on the analysis of the HTP assay results using the methods described above, variants of SEQ ID NO: 21 / 22 were selected for larger scale analysis and further characterization. About 10 pL of glycerol stock or a freshly streaked single colony for each variant was inoculated in 7.5 mL of Luria-Bertani Broth medium with kanamycin (50 pg / ml) in 50 mL centrifugal tube at 30 °C with shaking (250 rpm) for 20 hours. The overnight culture was diluted 1 : 100 with 100 mL of Terrific Broth medium containing antibiotics (kanamycin 50 pg / ml) in IL flask and grew until the cellular optical density reached 0.7. measured at 600 nm (ODeoo). Protein production was induced by adding IPTG to ImM final concentration for 20 hours with shaking (250 rpm) at 30 °C. Cells were then collected by centrifugation and resuspended in 20 mL lysis solution (100 mM HEPES buffer pH 7.5) per 100 mL of growihcultures. Cells were lysed by ultrasonication at 500 W for 15 min on ice (performed at 2 seconds of sonication with 4 seconds of intervals). The resulting lysate was clarified by centrifugation at 4 °C, and the supernatant was subsequently frozen and lyophilized to generate the enzyme powders.Analytical method for Carboxylesterase activity from SFP samples:

[0235] The conversion in the enzymatic reaction with SFP samples was determined using an Acquity BEH C18 column (2.1 x 50 mm, 1.7 pm particle size) was used, with mobile phase (A) 0.1% (v / v) TFA in water, (B) acetonitrile, 1 pl injection volume, flow rate 0.7 mL min'1, 55 °C column temperature, detection by MS detector (SIM ion: m / z=406.80 for the substrate (compound 1) and by-product (compound 3); SIM ion: m / z=397.80 for the product (compound 2)). Retention time of the substrate (compound 1): 5.6 minutes; Retention time of the desired product (compound 2): 4.7 minutes; Retention time of the undesired hydrolysis product (compound 3): 3.8 minutes.Gradient program

[0236] Engineered polypeptides with >2-fold selectivity, defined as the ratio of the conversion % of substrate to the product (compound 2) over the conversion % of substrate to the by-product (compound 3), relative to the parent polypeptide, are listed in Table 12.1. The selectivity was measured with the shake flask powder samples.Table 12.1 Variants and Conversion

[0237] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 22 and defined as: “+’■ = selectivity at least 2-fold relative to reference polypeptide.

[0238] Carboxylesterase variants with mutations A368L or A368H produced less hydrolysis by-products (compound 3) from the tetrapeptide (compound 1), and these engineeredcarboxylesterase enzymes provide more selective biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations A368L (SEQ ID NO: 24), had the highest selectivity. Thus, the encoding polynucleotide (SEQ ID NO: 23) was selected for further directed evolution.Example 13 (Rdl2 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 24

[0239] In this example, engineering of SEQ ID NO: 24 for improved selectivity against hydrolysis of tetrapeptide (compound 1) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites were subjected to mutagenesis combinatorically resulting in library 12. 1 and surface residues were subjected to mutagenesis individually resulting in library 12.2. These libraries were plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 12. The production of shake flask pow ders and HTP analytical methods were described in Example 12.Assays for Carboxylesterase activity:

[0240] In the high-throughput screening, 50 pL reactions (300 mM sodium phosphate buffer, 100 mM HEPES, 75 mM magnesium chloride, 25 mM tetrapeptide compound 1, pH 7, 0.5% v / v clarified cellular lysates) were incubated at 20 °C with shaking at 1000 rpm for 2 hours. For assays with SFP samples, reactions were performed at 25 mM tetrapeptide (compound 1) in 300 mM sodium phosphate buffer pH 7.5, 100 mM HEPES. 75 mM magnesium chloride, and with 0. 1 g / L of enzyme at 20 °C with shaking at 1000 rpm for 2 hours.

[0241] Engineered polypeptides with >1.5-fold selectivity7, defined as the ratio of the conversion % of substrate to the product (compound 2) over the conversion % of substrate to the by-product (Compound 3), relative to the parent polypeptide, are listed in Table 13.1. The selectivity was measured with the shake flask powder samples.Table 13.1 Variants and Selectivity

[0242] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 24 and defined as: "‘+” = selectivity at least 1.5-fold relative to reference polypeptide.

[0243] Variants with mutations II 15L. I157F, W170L, C172F, L187Q produced less hydrolysis by-products (compound 3) from the tetrapeptide (compound 1). and these engineered carboxylesterase enzymes provide more selective biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations II 15L, I157F, W170L, C172F, L187Q (SEQ ID NO: 26), had the highest selectivity . Thus, the encoding polynucleotide (SEQ ID NO: 25) was selected for further directed evolution.Example 14 (Rd 13 evolution of Carboxylesterase)Enzyme variants of SEQ ID NO: 26

[0244] In this example, engineering of SEQ ID NO: 26 for improved activity while maintaining selectivity against hydrolysis of tetrapeptide (compound 1) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites and surface residues and improved performance in previous rounds were subjected to mutagenesis combinatorically, resulting in library 13. 1. These libraries were plated to form single colonies, which were grown and screened using the high-throughput growth and expression methods described in Example 12. The HTP analytical methods are described in Example 12.Assays for Carboxylesterase activity:

[0245] In the high-throughput screening, 50 pL reactions (300 mM sodium phosphate buffer, 100 mM HEPES, 75 mM magnesium chloride, 25 mM tetrapeptide with (compound 1), pH 7, 0.2% v / v clarified cellular lysates) were incubated at 20 °C with shaking at 1000 rpm for 2 hours. Candidates with improved selectivity7were picked for growth and expression again. Activity was re-examined in triplicate using the same conditions as the primary' screening.

[0246] Engineered polypeptides with >2-fold activity relative to SEQ ID NO: 26 are listed in Table 14.1. The selectivity shown was measured with the retest samples.Table 14.1 Variants and Selectivity

[0247] Levels of increased conversion were determined relative to the reference polypeptide of SEQ ID NO: 26 and defined as: “+’ = activity at least 2-fold relative to reference polypeptide

[0248] Variants with mutations T10K. G162I. T246H, H272A, F343L produced more desired product (compound 2) from the compound tetrapeptide (compound 1) and only low level of the hydrolysis by-product (compound 3), and these engineered carboxylesterase enzymes provide more selective biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations T10K. G162I. T246H, H272A, F343L (SEQ ID NO: 28), had the highest activity and excellent selectivity (>99: 1 product: by-product ratio). Thus, the encoding polynucleotide (SEQ ID NO: 27) was selected as the final variant.Discussion

[0249] Following 13 rounds of evolution, carboxylesterase variant SEQ ID NO: 28 emerged, which contains 32 mutations (T10K, R17Y, V115L, N132R, R133M, T138R, T140P, C147R, Y154N, N156P, I157F, G158E, I159Y, P160L, A165P, H166W, G170L, L172F, KI 84V, L187Q, A191D, L199S, T246H, T265S, G281S, P325S, M327Y, R330S, S350D, K352T, M368L, and F373Y) relative to reference wild-type ECE20 sequence (SEQ ID NO: 2). Among the 32 mutations, 15 mutations are active site residues, 7 mutations are the surface 1 residues (positions with solvent accessible surface area > 64 A), 5 mutations are the surface 2 residues (positions with solvent accessible surface area <64 A and > 8 A), and 4 are residues not located in the 3 preceding groups (see FIGs. 4A and 5A). Consistent with the primary focus on activity' improvement in the first 3 rounds of evolution, beneficial mutations identified were highly enriched in active site residues. As organic co-solvent tolerance and thermostability were incorporated as selection pressures in subsequent rounds (see FIG. 4B), more beneficial mutations from the surface area of the carboxylesterase were observed. To rationalize the effect of the mutations introduced, a model of carboxylesterase was constructed using the same homology template 4IVK and superimposed it on the wild-type model. The model indicates that ECE20 has a shallow active site pocket near the enzyme surface. The majority of beneficial mutations identified up to round 8 are in the flexible loop regions of this enzyme (FIG. 4C), suggesting potential roles of these loops to stabilize substrates in the catalytic pocket.

[0250] In later rounds of evolution, the selection pressure focus shifted to improving the selectivity (reducing hydrolysis by-product (compound 3)) of the macrocyclase, mutations were identified (G199S, L266D, G281S, L343F) that are in close proximity to the substrate and increase the selectivity of this enzy me while reducing overall activity (see FIG. 5B), possibly through direct interaction with the substrate. Subsequently, compensatory’ mutations were identified that are more distant to the active site to restore the activity while maintaining achieved selectivity (FIG. 5C). Notably, many mutations identified in round 9 to 13 are in recurrent positions from earlier rounds and either mutated back to the wild-type residues or anothersubstitution (FIG. 5B, * labeled). The enrichment of mutations in recurrent positions highlights not only the importance of these residues in the function of the macrocyclase but also capability to fine-tune the enzyme performances through evolution and screening strategies.

[0251] Overall, polypeptide sequences having SEQ ID NOs: 22, 24, 26 and 28 emerged following 13 rounds of evolutionary pressures as the best-performing macrocyclase variants. These engineered polypeptides show improved (1) activity for the product of compound 2, (2) selectivity for compound 2 over undesired hydrolysis by-product compound 3 and undesired regioisomers of compound 2, (3) thermostability, and (4) co-solvent tolerance relative to ECE20 (SEQ ID NO: 2). These characteristics are important for successful large-scale manufacturing of non-canonical phenylalanine amino acids.

[0252] The disclosed subject matter is not to be limited in scope by the specific embodiments and examples described herein. Indeed, various modifications of the disclosure in addition to those described will become apparent to those skilled in the art from the foregoing description and accompanying figures. Such modifications are intended to fall within the scope of the appended claims.

[0253] All references (e.g., publications or patents or patent applications) cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual reference (e.g.. publication or patent or patent application) was specifically and individually indicated to be incorporated by reference in its entirety for all purposes. Other embodiments are within the following claims.

Claims

CLAIMSWhat is claimed is:

1. An engineered polypeptide comprising an amino acid sequence having at least 92.5%. 94%. 95%. 96%. 98%. or 99% sequence identity to any one of SEQ ID NOs:

22. 24, 26, and 28.

2. The polypeptide according to claim 1, wherein the polypeptide comprises the sequence of SEQ ID NO: 22.

3. The polypeptide according to claim 1, wherein the polypeptide comprises the sequence of SEQ ID NO: 24.

4. The polypeptide according to claim 1, wherein the polypeptide comprises the sequence of SEQ ID NO: 26.

5. The polypeptide according to claim 1, wherein the polypeptide comprises the sequence of SEQ ID NO: 28.

6. An engineered polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 2, wherein the polypeptide comprises at least 2 at positions selected from 10.

17. 115,132, 133, 138. 140, 147, 154, 156, 157. 158, 159, 169, 162, 165, 166. 170, 172, 184. 187, 191. 199, 246, 265. 281, 325. 327, 330. 350, 352, 368. and 373, relative to SEQ ID NO: 2.

7. The polypeptide of claim 6, wherein the polypeptide comprises at least 5 mutations at positions selected from 10, 17, 115,132, 133, 138, 140, 147. 154, 156. 157, 158, 159, 169, 162, 165, 166, 170, 172, 184, 187, 191, 199, 246, 265, 281, 325, 327, 330, 350, 352, 368, and 373, relative to SEQ ID NO: 2.

8. The polypeptide of claim 6 or 7, wherein the polypeptide comprises any of the following sets of mutations relative to SEQ ID NO: 2: a) mutations at positions 17, 115, 132, 133, 138, 140, 147, 154, 156, 158, 159, 160, 162, 165, 166, 170, 172, 184, 191, 199, 265, 272, 281, 325, 327, 330, 343, 350, 352, 368, and 373;b) mutations at positions 17. 115, 132. 133, 138.

140. 147, 154. 156, 158. 159, 160, 162. 165, 166, 170, 172, 184, 191, 199, 265, 272, 281, 325, 327, 330, 343, 350, 352, 368, and 373; c) mutations at positions 17, 115, 132, 133, 138, 140, 147, 154, 156, 157, 158, 159, 160, 162, 165. 166, 170, 172, 184, 187, 191, 199, 265, 272, 281, 325, 327. 330, 343, 350, 352, 368, and 373; or d) mutations at positions 10, 17, 115, 132, 133, 138, 140, 147, 154, 156, 157, 158, 159, 160, 162, 165, 166, 170, 172, 184, 187, 191, 199, 246, 265, 281, 325, 327, 330, 350, 352, 368, and 373.

9. The polypeptide of any one of claims 6-8, wherein the polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 2, wherein the polypeptide comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 substitutions selected from T10K, R17Y. VI 15L, N132R, R133M, T138R, T140P, C147R. Y154N, N156P, I157F, G158E, I159Y. P160L, I162G, A165P, H166W, G170L, L172F, K184V. L187Q, A191D, L199S, T246H, T265S, G281S, P325S, M327Y, R330S, S350D, K352T, M368L, and F373Y, relative to SEQ ID NO: 2.

10. The polypeptide of any one of claims 6-9, wherein the polypeptide comprises any of the following substitution sets relative to SEQ ID NO: 2: a) a substitution set of R17Y, V115I, N132R, R133M, T138R, T140P, C147R, Y154N, N156P, G158E, I159Y, P160L, I162G, A165P, H166W, G170W, L172C, KI 84V, A191D, L199S. T265S. A272H, G281S, P325S, M327Y, R330S, L343F, S350D, K352T, M368A, and F373Y; b) a substitution set of R17Y, VI 151, N132R, R133M, T138R, T140P, C147R, Y154N, N156P, G158E, I159Y, P160L, I162G, A165P, H166W, G170W, L172C, KI 84V, A191D, L199S, T265S. A272H, G281S, P325S, M327Y, R330S, L343F, S350D, K352T, M368L, and F373Y; c) a substitution set of R17Y, V115L, N132R, R133M, T138R, T140P, C 147R, Y154N, N156P, I157F, G158E, I159Y, P160L, I162G, A165P, H166W, G170L, L172F, KI 84V, L187Q, A191D, L199S, T265S, A272H, G281S, P325S, M327Y, R330S, L343F, S350D, K352T, M368L, and F373Y; or d) a substitution set of T10K, R17Y, VI 15L, N132R, R133M, T138R, T140P, C147R, Y154N, N156P, I157F, G158E, I159Y, P160L, A165P, H166W, G170L, L172F, KI 84V,L187Q, A191D, L199S. T246H. T265S. G281S, P325S, M327Y. R330S, S350D, K352T.M368L, and F373Y.

11. An engineered polypeptide comprising an amino acid sequence that comprises a stretch of at least 100, 150, 200, 250, 300. 325, 350. 375, or 400 consecutive amino acids of any one of SEQ ID NOs: 22, 24, 26, and 28.

12. An engineered polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 22, 24, 26, and 28, wherein the polypeptide contains 0, 1 or 2 amino acid residues that differ from amino acids 147-172 of SEQ ID NO: 24.

13. The polypeptide of any one of claims 1-12, further comprising an affinity tag.

14. The polypeptide of any one of claims 1-13, wherein the polypeptide is an esterase.

15. The polypeptide of any one of claims 1, 11, and 12, wherein the polypeptide comprises the sequence of SEQ ID NO: 38 or 39.

16. The polypeptide of any one of claims 1-14, wherein the polypeptide exhibits activity in catalyzing a cyclization of a linear tetrapeptide.

17. The polypeptide of any one of claims 1-16, wherein polypeptide exhibits activity in catalyzing a reaction in which Compound (1)is converted to Compound (2)18. The polypeptide of claim 17, wherein the polypeptide exhibits greater activity in catalyzing the reaction relative to a polypeptide having the amino acid sequence of SEQ ID NO: 2.

19. The polypeptide of claim 17 or 18, wherein the polypeptide exhibits greater regioselectivity in catalyzing the reaction relative to a polypeptide having the amino acid sequence of SEQ ID NO: 2.

20. The polypeptide of any one of claims 17-19, wherein the polypeptide exhibits higher co-solvent tolerance and / or thermostability relative to a polypeptide having the amino acid sequence of SEQ ID NO: 2.

21. A polynucleotide encoding the polypeptide of any one of claims 1-20.

22. A polynucleotide comprising a nucleic acid sequence having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 33, and 35.

23. The polynucleotide of claim 21 or 22, wherein the polynucleotide comprises the sequence of any one of SEQ ID NOs:

5. 7, 9, 11, 13, 15, 17, 19, 21.

23.

25.

27.

29. 33, and 35.

24. A polynucleotide comprising a nucleic acid sequence having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 21, 23, 25, and 27.

25. The polynucleotide of claim 21 or 24, wherein the polynucleotide comprises the sequence of any one of SEQ ID NOs: 21, 23, 25, and 27.

26. The polynucleotide of any one of claims 21-25, wherein the polynucleotide is codon-optimized for expression in E. coli.

27. An expression vector comprising the polynucleotide of any one of claims 21-25, operably linked to one or more control sequences suitable for directing expression of the encoded polypeptide in a host cell.

28. The expression vector of claim 27, wherein the control sequence comprises a promoter.

29. The expression vector of claim 28. wherein the promoter comprises an E. coli promoter.

30. A host cell comprising the expression vector of any one of claims 27-29.

31. The host cell of claim 30, wherein the host cell is E. coli.

32. A method of producing a polypeptide, the method comprising culturing the host cell of claim 30 or 31 under conditions and for a time suitable for expression of the polypeptide, optionally further comprising recovering the polypeptide, and / or isolating the polypeptide.

Citation Information

Patent Citations

  • Enzymes and formulations for broad-specificity decontamination of chemical and biological warfare agents

    WO2008036061A2