Engineered ATP-dependent ligases for the synthesis of oligopeptides

Engineered ATP-dependent ligases overcome the limitations of wild-type enzymes by enhancing substrate scope and selectivity, enabling efficient synthesis of tripeptides and tetrapeptides without protecting groups, suitable for complex peptide production.

WO2026019620A1PCT designated stage Publication Date: 2026-01-22MERCK SHARP & DOHME LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/037055
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-15
Filing Date
2025-07-10
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing ATP-dependent amino acid ligases have limited substrate scope and are hindered by high costs and complex chemistries, limiting their use in large-scale synthesis of oligopeptides, particularly those containing non-canonical amino acids, with a need for enzymes that can form oligopeptides in a protecting group-free manner with high selectivity and yield.

Method used

Engineered ATP-dependent amino acid ligases, derived from wild-type enzymes via directed evolution, exhibit improved activity, regioselectivity, and thermostability, enabling the synthesis of tripeptides and tetrapeptides through amide bond formation without protecting groups, using sequential or simultaneous enzymatic cascades.

Benefits of technology

The engineered ligases demonstrate enhanced substrate tolerance, reduced by-product formation, and increased thermostability, facilitating the efficient synthesis of complex peptides like active pharmaceutical ingredients, reducing the number of synthetic steps and waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025037055_22012026_PF_FP_ABST
    Figure US2025037055_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides engineered ATP-dependent ligase enzymes having improved properties as compared to a naturally occurring wild-type ATP-dependent ligase enzyme, including the catalysis of synthesis of oligopeptides, such as tripeptides and tetrapeptides. Also provided are polynucleotides encoding the ligase enzyme variants, methods of generating the ligase variants and methods of using the ligase variants to synthesize oligopeptides.
Need to check novelty before this filing date? Find Prior Art

Description

ENGINEERED ATP-DEPENDENT LIGASES FOR THE SYNTHESIS OF OLIGOPEPTIDESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 671,560 filed July 15, 2024, the entire contents of which are incorporated by reference herein.REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY

[0002] The contents of the electronic sequence listing (25869-WO-PCT_SL.xml; Size: 625,738 bytes; and Date of Creation: June 2, 2025) are herein incorporated by reference in their entirety.FIELD

[0003] The present invention relates to ATP-dependent amino acid ligases, useful in biocatalytic and synthetic processes involving amino acid ligations for the generation of tripeptides and tetrapeptides. Such enzymes may be particularly useful in synthetic processes that may be used as part of the preparation of isopropyl ((2S,3S)-l-((S)-2-((S)-2-amino-3-(3- (aminomethyl)phenyl)propanamido)-3-(l-(6-arninohexyl)-5-fluoro-17f-indol-3-yl)propanoyl)-3- (2-(7ert-butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate or intermediates formed during preparation of such compounds.BACKGROUND

[0004] Enzymes are protein molecules that function to catalyze many biological processes, often by several orders of magnitude. Enzymes in nature display a high degree of efficiency and great specificity for their cognate substrates, making them ideal catalysts for carrying biological functions with high fidelity. Because they are not modified during the reactions in which they participate, enzymes may be cost effectively used as catalysts for desired chemical transformations. Enzymes are being increasingly harnessed by biocatalysis, providing green approaches to industrial processes by decreasing the number of synthetic steps and waste, allowing for production of value-added synthetic intermediates and products with lower cost.

[0005] The emergence of new therapeutic modalities requires the development of complementary tools for their efficient syntheses. Non-natural peptides have gained attention in the pharmaceutical industry due to their high selectivity, efficacy, and safety profiles. However, their widespread application has been hindered by the high costs of synthesis and the unique chemistries involved. Enzy mes present a promising solution to supplement existing chemical approaches, as they offer high specificity' and operate under mild and environmentally safer reaction conditions.

[0006] ATP-dependent amino acid ligase enzymes are commonly found in nature in the ribosome and non-ribosomal peptide synthesis pathways and can catalyze the formation of amide bonds through acyl-adenylated intermediate generation (see Ogasawara et al., Chem. Eur. J. 2017, 23. 10714-10724). These enzymes include ATP-grasp ligases, which generate in the presence of ATP activated acyl-phosphate intermediates that are subsequently coupled to an extending peptide chain through a nucleophilic attack. This biocatalytic transformation is particularly attractive for chemical processes because it does not involve the use of protecting group manipulations commonly found in traditional chemical synthesis of peptides.

[0007] ATP-dependent amino acid ligase enzymes may hold promise to fulfill the increasing needs for synthesizing different non-canomcal amino acids and oligopeptides containing non- canonical amino acids in the chemical manufacturing and pharmaceutical industries.Nonetheless, use of these enzymes for the large-scale synthesis of oligopeptides has been limited, at least in part due to the limited substrate scope of these enzymes.

[0008] Therefore, there is a need in the art for enzymes that can form oligopeptides in a protecting group-free manner while maintaining desired regiochemistry (i.e., regioselectivity), high selectivity for the desired product (i.e., reduced by-product formation), and high yields. There is a particular need for ATP-dependent amino acid ligases that have improved activity, regioselectivity and ability to catalyze oligopeptide-forming reactions in a multistep cascade.SUMMARY

[0009] The present disclosure provides, inter alia, engineered polypeptides (e.g., ATP- dependent amino acid ligase enzymes) capable of catalyzing amide bond formation resulting in natural or nonnatural oligopeptides. The present disclosure further provides polynucleotides encoding these engineered polypeptides and expression vectors containing these polynucleotides, and host cells containing these expression vectors. Further disclosed are methods of producing these engineered polypeptides, and methods of synthesizing oligopeptides using these polypeptides. The ATP-dependent amino acid ligase enzymes described herein are capable of coupling (i.e., ligating) natural or nonnatural amino acids to the same amino acids, to other amino acids, or to oligopeptides, such as linear dipeptides and tripeptides. These enzy mes catalyze the activation of amino acid carboxylates to acylphosphate intermediates for nucleophilic attack by dipeptide and tripeptide nucleophiles. The products of these ligation reactions, i.e.. linear tripeptides and tetrapeptides, respectively, ultimately may be used in the synthesis of larger peptides. In some aspects, the disclosed ATP-dependent amino acid ligase enzymes are engineered polypeptides that may be useful in the preparation of compounds such as isopropyl((21S',3<S’)-l-(( ’)-2-amino-3-(l-(6-aminohexyl)-5-fluoro-17 / -indol-3-yl)propanoyl)-3-(2-(tert- butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate and isopropyl ((2S,3<S -1 -((S)-2-((S)-2-amino-3-(3-(aminomethyl)phenyl)propanamido)-3-(l-(6-aminohexyl)-5-fluoro-177-indol-3- yl)propanoyl)-3-(2-(tert-butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate.

[0010] Thus far, ATP-dependent ligases have found limited use in commercial process applications given their limited substrate scope and substrate concentration tolerance. However, in view of this selectivity, this class of enzymes present as excellent catalysts to use in protecting-group free chemistries, simplifying the number of synthetic steps involved in the synthesis of peptides. Herein, the use of protein engineering to improve enzyme activity, selectivity, phosphate tolerance and thermostability of several ATP-dependent ligases of the ATP-grasp superfamily (EC 6.3.2.X, acid-amino-acid ligases or peptide synthases) is demonstrated. In particular, engineering of wild-type ATP-grasp ligases having enhanced activity to form tripeptides and tetrapeptides en route to an isopropyl ((2S,35)-l-((S)-2-((S)-2-amino-3- (3-(aminomethyl)phenyl)propanamido)-3-(l-(6-aminohexyl)-5-fluoro-17 / -indol-3-yl)propanoyl)-3-(2-( / e / 7-butoxy)-2 -oxoethoxy )pyrrolidine-2-carbonyl)-E-threoninate product is demonstrated. Such enzymes maybe useful in the preparation of intermediates in processes to generate complex compounds with biological activities, for example active pharmaceutical ingredients.

[0011] As such, in some embodiments, provided are engineered polypeptides useful for converting a natural or nonnatural amino acid and a natural or nonnatural dipeptide into a natural or nonnatural tripeptide, in the presence of ATP. In additional embodiments, provided are engineered polypeptides useful for catalyzing the conversion of a natural or nonnatural amino acid and a natural or nonnatural tripeptide into a natural or nonnatural tetrapeptide, in the presence of ATP. In various embodiments, the natural or nonnatural tripeptide and tetrapeptide products of these reactions may be useful in the synthesis of more complex peptides, such as macrocyclic peptides that are active pharmaceutical ingredients.

[0012] This disclosure is based, at least in part, on the engineering of a wild-type ATP-grasp ligase enzyme via directed evolution over several iterative rounds to generate engineered enzymes that are capable of ligating (1) a natural or nonnatural dipeptide and natural or nonnatural amino acid into a tripeptide, and (2) a natural or nonnatural tripeptide and natural nor nonnatural amino acid into a tetrapeptide with high yields. Because these enzymes are engineered using particular screening designs to identify variants with improved regioselectivity, reduced by-product formation, and improved thermostability, the disclosed engineered polypeptides (i.e., enzymes) exhibit these properties and / or catalyze their syntheses in a manner such that the synthetic processes exhibit these properties. The disclosed engineered enzymes exhibit severalunique properties during catalysis of the disclosed syntheses, and as such are advantageous over existing corresponding enzymes.

[0013] Additional embodiments describe processes for preparing the subject ATP-dependent amino acid ligase enzymes and processes for using the subject ATP-dependent amino acid ligase enzymes to catalyze ligation couplings and synthesis schemes involving such couplings.

[0014] Thus, in some aspects, provided herein are engineered polypeptides comprising an amino acid sequence having at least 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of the amino acid sequences disclosed in Table 7. In one aspect, provided are engineered polypeptides comprising an amino acid sequence having at least 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 40, 42, 44, 46, 48, and 50. In some aspects, these polypeptides have at least 98%, at least 99%, or 100% identity to any one of SEQ ID NOs: 40, 42, 44, 46, 48, and 50.

[0015] In some embodiments, the polypeptides comprise an amino acid sequence that comprises 100 or more consecutive amino acids in common with any one of SEQ ID NOs: 40. 42, 44, 46, 48, and 50, such as a stretch of at least 100, 150, 200, 250, 300, 325, 350, or 375 consecutive amino acids of any one of SEQ ID NOs: 40, 42, 44, 46, 48, and 50.

[0016] In another aspect, provided are engineered polypeptides comprising an amino acid sequence having at least 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 108, 1 10, 112, 114, and 116. In some aspects, these polypeptides have at least 98%, at least 99%, or 100% identity to any one of SEQ ID NOs: 108, 110, 112, 114, and 116.

[0017] In some embodiments, the polypeptides comprise an amino acid sequence that comprises 100 or more consecutive amino acids in common with any one of SEQ ID NOs: 108, 110, 112. 114, and 116, such as a stretch of at least 100. 150. 200, 250. 300, 325. 350, or 375 consecutive amino acids of any one of SEQ ID NOs: 108, 110, 112, 114, and 116.

[0018] In some embodiments, the disclosed polypeptides further comprise one or more affinity tags, such as a hexa-histidine tag.

[0019] In some embodiments, the polypeptides exhibit activity in catalyzing a ligation of an amino acid and a dipeptide to generate a tripeptide. These polypeptides may be an ATP- dependent ligase known as a “Trp-ligase.” These polypeptides may exhibit activity in catalyzing a reaction in which nonnatural dipeptide Compound (2c)is contacted with nonnatural amino acid Compound (Id)in the presence of ATP, Mg2+, a polyphosphate, and an enzyme utilizing a polyphosphate to synthesize ATP from ADP, to generate nonnatural tripeptide Compound (3d)

[0020] In some embodiments, the polypeptides exhibit activity in catalyzing a ligation of an amino acid and a tripeptide to generate a tetrapeptide. These polypeptides may be an ATP- dependent ligase known as a "Phe-ligase.” or In some embodiments, the polypeptide exhibits activity in catalyzing a reaction in which nonnatural tripeptide Compound (3d)is contacted with nonnatural amino acid Compound (4f)(4f), in the presence of ATP, Mg2+, a polyphosphate, and an enzyme utilizing a polyphosphate to synthesize ATP from ADP, to generate nonnatural tetrapeptide Compound (5k)

[0021] In some embodiments, the engineered polypeptides exhibit higher (e.g., significantly higher) (1) phosphate concentration tolerance. (2) substrate concentration tolerance and catalytic turnover rate, (3) thermostability, and / or (4) pH tolerance relative to the polypeptide having the reference sequence of SEQ ID NO: 4.

[0022] In some aspects, provided herein are polynucleotides encoding any of the disclosed engineered polypeptides. Provided are polynucleotides having a nucleic acid sequence that has at least 80%, 85%, 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 39, 41, 43, 45, 47, and 49. In some embodiments, the polynucleotides comprise the sequence of any one of SEQ ID NOs: 39, 41, 43, 45, 47, and 49. Further provided are polynucleotides having a nucleic acid sequence that has at least 80%, 85%, 90%, 92.5%, 95%, 96%. 98%. or 99% sequence identity to any one of SEQ ID NOs: 107, 109. I l l, 113, and 115. In some embodiments, the polynucleotides comprise the sequence of any one of SEQ ID NOs: 107, 109, 111, 113, and 115. Further provided are polynucleotides having a nucleic acid sequence that has at least 80%, 85%, 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 126-130. In some embodiments, the polynucleotides comprise the sequence of any one of SEQ ID NOs: 126-130.

[0023] Provide herein are expression vectors comprising any of the disclosed polynucleotides operably linked to one or more control sequences suitable for directing expression of the encoded polypeptide in a host cell. The control sequence may comprise a promoter, such as an E. coli promoter.

[0024] Further provided herein are host cells comprising any of the disclosed expression vectors and / or polynucleotides. In some embodiments, the host cell is E. coli.

[0025] Further provided are methods of producing a polypeptide comprising culturing any of the disclosed host cells, recovering the polypeptide, and isolating the polypeptide.

[0026] Other embodiments, aspects and features of the present invention are either further described in or will be apparent from the ensuing description, examples, and appended claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] FIG. 1 shows a general chemical reaction scheme depicting the synthesis of a tetrapeptide of Formulas V (i.e., one of the compounds among 5e-5k) (SEQ ID NO: 132)from the modified amino acids of Formula I (one of the compounds among la-ld) and Formula IV (compound 4e or 41) and modified dipeptide of Formula II (one of the compounds among 2a-2c), via an intermediate tripeptide of Formula III (one of the compounds among 3a-3f). using engineered ligase enzymes Trp-ligase and Phe-ligase, together with ATP, magnesium, the phosphate donor polyphosphate (PolyP), and a polyphosphate kinase enzyme (PPK).

[0028] FIG. 2 shows a chemical reaction scheme depicting Phe-ligase catalyzed-formation of desired and undesired regioisomers 5j (SEQ ID NO: 138) and 5j-regio from the contacting of compound 4f and compound 3c).DETAILED DESCRIPTION

[0029] The present disclosure provides engineered ATP-dependent ligase enzymes. As is described in the Examples, the present disclosure provides engineered polypeptides derived from Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1, SEQ ID NO: 2). These engineered polypeptides exhibit improved enzy me catalysis properties relative to wild-type SEQ ID NO: 2, including improved enzyme activity' for the synthesis of oligopeptides from (i) natural or nonnatural dipeptide nucleophiles and natural or nonnatural amino acid carboxylates and (ii) natural or nonnatural tripeptide nucleophiles and natural or nonnatural amino acid carboxylates, which properties were engineered through iterative rounds of directed evolution. This disclosure also provides enzy matic strategies for the synthesis of non-natural peptides without the need for protecting group manipulations utilizing ATP-dependent ligases in sequential or simultaneous enzymatic cascades.

[0030] In some aspects, the ATP-dependent amino acid ligase enzymes described herein are the result of engineering via directed evolution from an ATP-dependent amino acid ligase from SEQ ID NO: 4. Such enzymes are capable of activating modified tryptophan carboxylates of Formula I to form an acylphosphate intermediate enabling the nucleophilic attack of modified proline-threonine dipeptide nucleophiles of Formula II to form a tetrahedral intermediate that collapses to yield an amide bond, producing modified tryptophan-proline-threonine tripeptides of Formula III, as depicted in Table 1. Wild-type ligase SEQ ID NO: 2 exhibited only trace activity in catalyzing amide bond formation between canonical (natural) tryptophan and proline-threonine dipeptide (reaction A in Table 1) or canonical phenylalanine and canonical tryptophan-proline-threonine tripeptide (reaction E in Table 2). In particular, a His-tagged variant of wild-type ligase SEQ ID NO: 4 did not exhibit any detectable activity' for the desired proline-modified threonine derivative dipeptide and modified tryptophan substrates — or for that matter, any compounds from entries B, C and D in Table 1. It was therefore concluded that engineering the substrate scope of this polypeptide via directed evolution was necessary to introduce and substantially enhance this activity'. As described in the Examples, ATP-dependent ligases capable of catalyzing one of two desired amide bond forming steps in the elaboration of a dipeptide to tetrapeptide were engineered via structure-guided semi-rational directed evolution using techniques such as single- site-saturation mutagenesis (SSM) and combinatorial library generation. Since the wild-type enzyme did not recognize the modified amino acid substrates, protein engineering involved engineering the protein to expand substrate scope (i.e., a substrate walk) as well as engineering along two trajectories (see Tables 1 and 2) to achieve desired regiospecificity and selectivity for each of the two amide bond formations, particularly in a cascade context. Successful engineering of final ligase variants enabled tetrapeptide formation from dipeptide in a cascade reaction with an inexpensive phosphate source. Acknowledging the expense associated with stoichiometric ATP, an ATP recycling system was incorporated into the cascade using a phosphate source and catalytic AMP to regenerate ATP.

[0031] For example, the Examples describe the generation of multiple variants of SEQ ID NO: 3 / 4 (Trajectory 1), referred to as “Trp-ligases,” exhibiting more than 20,000-fold enhanced activity and regioselectivity in catalyzing the synthesis of a modified tryptophan-proline- threonine tripeptide, relative to SEQ ID NO: 4 or wild-type SEQ ID NO: 2. The Trp-ligase variants described herein provided unsurpassed regioselectivity (regioisomeric excess of about 100%), which was demonstrated on up to 100-gram scale. Additional advantageous properties of the Trp-ligase variants are described herein, including reduced by-product formation and increased thermostability, demonstrating their synthetic utility. Examples of such Trp-ligase variants are the polypeptides of SEQ ID NOs: 40, 42, and 44.

[0032] Additionally, the Examples describe the generation of multiple variants of SEQ ID NO: 3 / 4 referred to as “Phe-ligases” (Trajectory 2), exhibiting more than 107-fold improved activity and regioselectivity in catalyzing synthesis of a modified try ptophan-proline-threonine-phenylalanine tetrapeptide, relative to SEQ ID NO: 4 or wild-type SEQ ID NO:2. The Phe-ligase variants described herein provided unsurpassed regioselectivity (regioisomeric excess of at least about 99.9%), which was demonstrated on up to 100-gram scale. Additional advantageous properties of the Phe-ligase variants are described herein, including reduced by-product formation, and thermostability, demonstrating their synthetic utility. Examples of such Phe-ligase variants are the polypeptides of SEQ ID NOs: 108, 110, and 112.

[0033] Other ATP -dependent ligase variants disclosed herein (e.g., the Trp-ligase variants of SEQ ID NOs: 46, 48, and 50 and the Phe-ligase variants of SEQ ID NOs: 114 and 116) also have improved enzyme properties, including improved enzyme activity, compared to wild-type SEQ ID NO: 2 or SEQ ID NO: 4. Therefore, the present disclosure provides several engineered ATP- dependent ligase polypeptides derived from SEQ ID NO: 1 / 2 that enable the biocatalytic synthesis of oligopeptides containing modified or non-canonical amino acids at scale.

[0034] The engineered polypeptides provided herein are ATP-grasp enzymes that catalyze amide bond formation between amino acids, and thus achieve a ligation. Enzymatic transformations involving amide bond formation between a carboxylic acid and an amine require the activation of the carboxylic acid group to make the overall reaction thermodynamically feasible. For ATP-grasp enzymes, this energy source is provided by ATP, such as through hydrolytic coupling of ATP to ADP and phosphate. The latter route can be catalyzed by enzymes in the ATP-grasp family, which activate a carboxylic acid-containing substrate with ATP to form a high-energy acyl-phosphate intermediate prior to condensation with a co-substrate nucleophile. Structurally, these enzy mes typically have a fold comprised of two a + (3 domains responsible for “grasping” a molecule of ATP in the active site. See Goswami & van Lanen, Mol. BioSyst. (2015) 11, 338-353, herein incorporated by reference.

[0035] In some aspects, the disclosed enzy mes may be used in a multi-step synthesis scheme. In some embodiments, the disclosed enzymes are used in multi-step scheme in a one-pot process, which is also referred to herein as a one-pot cascade, to generate a tetrapeptide product (such as compound 5k) from one or more monomers, or from a dipeptide and one or more monomers (see scheme depicted in FIG. 1). Provided herein are methods to generate oligopeptides in one or more enzymatic cascades, including in a single one-pot reaction or tw o one-pot reactions involving (1) Trp-ligases and (2) Phe-ligases. The disclosed enzymes w ere engineered in a cascade context to avoid substrate and product inhibition of the other enzymes in the cascade and to improve selectivity and reduce by-product formation. As such, in some aspects, these enzymes exhibit improved activity in an enzymatic cascade relative to wild- type enzyme. In some aspects,these enzymes exhibit improved regioselectivity, product inhibition, and / or thermostability in an enzymatic cascade relative to wild-type enzyme.

[0036] The present disclosure also provides polynucleotides and expression vectors encoding the engineered ATP -dependent ligase polypeptides of the present disclosure. The present disclosure also provides host cells comprising these polynucleotides or expression vectors, such as E. coll host cells. The host cells can be used for the expression and isolation of the ATP- dependent amino acid ligase enzymes described herein, or, alternatively, they can be used directly for the conversion of the substrate to product.

[0037] Further, the disclosure provides methods of generating the ATP-dependent ligase polypeptides. Further, the disclosure provides methods of synthesizing oligopeptides using the engineered ATP-dependent ligases of the present disclosure. In various embodiments, the disclosure provides methods of synthesizing a tripeptide (e.g., a tripeptide that includes a modified tryptophan and / or a modified proline) using the engineered ATP-dependent ligases. In various embodiments, the disclosure provides methods of synthesizing a tetrapeptide (e.g.. a tetrapeptide that includes a modified tryptophan, a modified phenylalanine, and / or a modified proline) using the engineered ATP-dependent ligases.

[0038] In some aspects, the multiple steps of these methods may be run simultaneously in cascades such that isolations and / or purifications of intermediates is eliminated. Advantageously, this eliminates the waste ty pically generated in multi-step chemical processes, and shortens the time required for manufacturing.Definitions

[0039] Listed below are definitions of various terms used herein. These definitions apply to the terms as they are used throughout this specification and claims, unless otherwise limited in specific instances, either individually or as part of a larger group.

[0040] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, and peptide chemistry are those well-known and commonly employed in the art.

[0041] As used herein, the articles “a” and “an” refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element. Furthermore, use of the term “including” as well as other forms, such as “include,” “includes,” and “included,” is not limiting.

[0042] As used herein, the term “about” in quantitative terms refers to plus or minus 10% of the value it modifies (rounded up to the nearest whole number if the value is not sub-dividable, such as a number of molecules or nucleotides).

[0043] All ranges disclosed herein are inclusive of the recited endpoint and independently combinable (for example, the range of “from 50 mg to 500 mg” is inclusive of the endpoints, 50 mg and 500 mg, and all the intermediate values). The endpoints of the ranges and any values disclosed herein are not limited to the precise range or value; they are sufficiently imprecise to include values approximating these ranges and / or values.

[0044] As used herein, the term “comprising” may include the embodiments “consisting of’ and “consisting essentially of.” The terms “comprise(s),” “include(s),” “having,” “has,” “may,” “contain(s),” and variants thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that require the presence of the named ingredients / steps and permit the presence of other ingredients / steps. However, such description should be construed as also describing compositions or processes as “consisting of’ and “consisting essentially of’ the enumerated components, which allows the presence of only the named components or compounds, along with any acceptable carriers or fluids, and excludes other components or compounds.

[0045] “Derived from” as used herein in the context of enzymes, identifies the originating enzyme, and / or the gene encoding such enzyme, upon which the enzyme was based. For example, the disclosed tryptophan ATP-dependent ligase enzymes (or Trp-ligases) are “derived from” the wild-type ligase enzy me of SEQ ID NO: 1 / 2 (or ligase of SEQ ID NO: 3 / 4 that is codon-optimized). The disclosed phenylalanine ATP-dependent ligase enzymes (or Phe-ligases) are “derived from” the wild-type ligase enzyme of SEQ ID NO: 1 / 2 (as well as the codon- optimized ligase of SEQ ID NO: 3 / 4). As used herein in the context of an amino acid residue, a “derivative” refers to a modified amino acid that is derived from a canonical L-amino acid.

[0046] As used herein, “polynucleotide” and “nucleic acid’ refer interchangeably to two or more nucleotides that are covalently linked together. The polynucleotide may be wholly comprised of ribonucleotides (i.e., RNA), wholly comprised of 2' deoxyribonucleotides (i.e., DNA), or comprised of mixtures of ribo- and 2' deoxyribonucleotides. While the nucleosides will typically be linked together via standard phosphodiester linkages, the polynucleotides may include one or more non-standard linkages. The polynucleotide may be single-stranded or double-stranded, or the polynucleotide may include both single-stranded regions and doublestranded regions. Moreover, while a polynucleotide will typically be composed of the naturally occurring encoding nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), it mayinclude one or more modified and / or synthetic nucleobases. such as, for example, inosine, xanthine, hypoxanthine, etc. In some embodiments, such modified or synthetic nucleobases are nucleobases encoding amino acid sequences.

[0047] As used herein, the terms “protein,” “polypeptide,” and “peptide” are used interchangeably herein to denote a polymer of at least two amino acids covalently linked by an amide bond, regardless of length or post-translational modification (e.g., glycosylation, phosphorylation, lipidation, myristoylation, ubiquitination, and the like). Included within this definition are D- and L-amino acids, and mixtures of D- and L-amino acids, as well as polymers comprising D- and L-amino acids, and mixtures of D- and L-amino acids. Proteins, polypeptides, and peptides may include a tag (e.g., an affinity tag), such as a histidine tag, and / or a signal peptide. As used herein, a “polypeptide” may encode an enzyme. The disclosure contemplates the use of polypeptides as reactants, products, and intermediates in an enzymatic reaction.

[0048] As used herein, the terms “amino acid” or “residue” as used in context of the polypeptides disclosed herein refers to the specific monomer at a sequence position of a polypeptide molecule. Amino acids are referred to herein by either their commonly known three- letter symbols or by the one-letter symbols recommended by International Union of Pure and Applied Chemistry (IUPAC) - International Union of Biochemistry (IUB) Biochemical Nomenclature Commission. This term encompasses modified (i.e., covalently or non-covalently modified), or non-canonical or non-standard, amino acids. For instance, this term encompasses covalently modified tryptophan, covalently modified proline, covalently modified phenylalanine, and covalently modified threonine amino acid molecules. As used herein, the terms “non- canonical amino acid” and “non-standard amino acid” refer to any amino acid other than the 20 naturally occurring L-amino acids, such as covalently modified amino acids and D-amino acids.

[0049] “Mutation” refers to any change in a polypeptide or polynucleotide sequence, and encompasses any number (i.e., one or more) of substitutions, deletions, insertions, and / or rearrangements present in a sequence compared to a reference sequence.

[0050] As used herein with respect to amino acid sequences, a “substitution” refers to a difference in the amino acid residue at a position of a polypeptide sequence relative to the amino acid residue at a corresponding position in a reference sequence. In some instances, the present disclosure provides specific amino acid differences denoted by the conventional notation “AnB,” where A is the single letter identifier of the residue in the reference sequence, n is the number of the residue position in the reference sequence, and B is the single letter identifier of the residue substitution in the sequence of the engineered polypeptide.

[0051] The term "amino acid substitution set” or “substitution set” refers to a group of amino acid substitutions in a polypeptide sequence, as compared to a reference sequence. For example, a substitution set may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-20, 20-25 or more than 25 amino acid substitutions.

[0052] “Corresponding to,” “reference to” or “relative to” when used in the context of the numbering of a given amino acid or polynucleotide sequence refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue number or residue position of a given polymer is designated with respect to the reference sequence rather than by the actual numerical position of the residue within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, although the gaps are present, the numbering of the residue in the given amino acid or polynucleotide sequence is made with respect to the reference sequence to which it has been aligned.

[0053] As used herein, “isolated polypeptide” refers to a composition in which the polypeptide is substantially separated from other contaminants that naturally accompany it (e.g., protein, lipids, and polynucleotides). The term embraces polypeptides that have been removed or purified from their naturally occurring environment or expression system (e.g., within a host cell or via in vitro synthesis). The recombinant polypeptides may be present within a cell, present in the cellular medium, or prepared in various forms, such as lysates or isolated preparations. As such, in some embodiments, the recombinant polypeptides can be an isolated polypeptide.

[0054] As used herein, a “ligase” is a polypeptide having an enzymatic capability of ligating, or coupling together, two molecules into an oligomer in which the molecules are linked by a bond, such as an amide bond. In various embodiments, the ligases of the disclosure are ATP-dependent ligases, in that they utilize the cofactor ATP as an activating agent. In some embodiments, the disclosed ligases are capable of catalyzing a ligation reaction between an amino acid and a di- or tripeptide. Ligases as used herein include naturally occurring (wild type) polypeptides as well as non-naturally occurring engineered polypeptides generated by human manipulation.

[0055] “Improved enzy me property” refers to any property of an enzy me that exhibits an improvement relative to a reference enzy me. For the enzymes described herein, the reference enzyme is a wild-type enzyme or another engineered enzyme. For example, in various embodiments, the reference enzyme is an affinity -tagged variant of a wild-type enzyme that has been codon-optimized (e.g., SEQ ID NO: 3, an E. co / Lcodon-optimized variant of SEQ ID NO: 1 which encodes SEQ ID NO: 4, a 6xHis-tagged variant set forth as SEQ ID NO: 2). Enzymeproperties for which improvement may be desirable include, but are not limited to. enzymatic activity (which may be expressed in terms of percent conversion of the substrate or total turnover number or product formed over time), thermal stability (or thermostability), stability7under high ammonia concentration, soluble expression, reduced by-product generation, higher substrate concentration, pH activity profile (pH tolerance), phosphate tolerance (or higher phosphate loading), cosolvent tolerance, improved activity in an enzymatic cascade, reduced cofactor requirements, refractoriness to inhibitors (e.g., product inhibition), regioselectivity, stereospecificity, and stereoselectivity (including enantioselectivity).

[0056] "‘Increased enzymatic activity7” refers to an improved property of the enzymes that is represented by an increase in specific activity (e.g., product produced / time / weight protein) or an increase in percent conversion of the substrate to the product (e.g., percent conversion of starting amount of substrate to product in a specified time period using a specified amount of enzyme) as compared to a reference enzyme. Exemplary7methods to determine enzyme activity are provided in the Examples. Any property relating to enzyme activity7may be affected, including the classical enzyme properties of K™, VOTm, or kcat, changes of which can lead to increased enzymatic activity. Improvements in enzy me activity can be from about 1.5 times the enzy matic activity of the corresponding wild-type enzyme, to as much as 2 times. 5 times, 10 times, 20 times, 25 times, 50 times, 75 times, 100 times, 150 times. 200 times, 500 times, 1000 times, 3000 times, 5000 times, 7000 times or more enzymatic activity than the reference enzyme, e.g., a naturally occurring enzyme or another enzyme from which the poly peptides were derived. In some examples, the enzyme exhibits improved enzymatic activity in the range of 100 to 3000 times, 3000 to 7000 times, or more than 7000 times greater than that of the parent enzyme. It is understood by the skilled artisan that the activity of any enzyme is diffusion limited such that the catalytic turnover rate cannot exceed the diffusion rate of the substrate, including any required cofactors. The theoretical maximum of the diffusion limit, or kcat / Km, is generally about 108to 109(M“1s’1). Hence, any improvements in the enzy me activity will have an upper limit related to the diffusion rate of the substrates acted on by the enzyme. Enzyme activity can be measured by any suitable approach, e.g., an enzyme activity assay or by any of the traditional methods for assaying chemical reactions, including but not limited to high-performance liquid chromatography (HPLC), HPLC-mass spectrometry (MS), ultra-performance liquid chromatography (UPLC), UPLC-MS, and nuclear magnetic resonance (NMR). Comparisons of enzyme activities may be made using a defined preparation of enzyme, a defined assay under a set condition, and one or more defined substrates, as further described in detail herein. Generally, when lysates are compared, the numbers of cells or the amount of protein assayed are determinedas well as use of identical expression systems and identical host cells to minimize variations in amount of enzyme produced by the host cells and present in the lysates.

[0057] As used herein, “by-products” refer to undesired products of a catalytic reaction, such as a catalytic ligation reaction. In some embodiments, this term encompasses the product of undesired ligation reactions, such as the ligation of two molecules of identical amino acid substrate. This term further encompasses undesired regioisomers. Examples of undesired products include compound 5j -regio.

[0058] As used herein, a “vector” is a DNA construct for introducing a DNA sequence into a cell. In some embodiments, the vector is an expression vector that is operably linked to a suitable control sequence capable of effecting the expression of the polypeptide encoded by the polynucleotide (e.g., DNA) sequence in a suitable host. In some embodiments, an “expression vector” has a promoter sequence operably linked to the polynucleotide (e.g., DNA) sequence (e.g., transgene) to drive expression in a host cell, and in some embodiments, also comprises a transcription terminator sequence.

[0059] As used herein with respect to polypeptides, the terms “expression” and “production” includes any step involved in the production of a polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of the polypeptide from a cell.

[0060] As used herein, an amino acid or nucleotide sequence (e.g., a promoter sequence, signal peptide, terminator sequence, and the like) is “heterologous” to another sequence with which it is operably linked if the two sequences are not associated in nature. For example, a “heterologous polynucleotide” is any polynucleotide that is introduced into a host cell by laboratory’ techniques, and the term includes polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into a host cell.

[0061] As used herein, the terms “host cell” and “host strain” refer to suitable hosts for an expression vector comprising a polynucleotide (e.g., DNA) provided herein (e.g., a polynucleotide encoding a ligase polypeptide disclosed herein). In some embodiments, the host cells are prokaryotic or eukaryotic cells that have been transformed or transfected with vectors constructed using recombinant DNA techniques as known in the art. In some embodiments, suitable host cells are E. coli cells.

[0062] “Coding sequence” refers to that portion of a polynucleotide (e.g.. a gene) that encodes an ammo acid sequence of a polypeptide.

[0063] “Naturally occurring” or “wild-ty pe” refers to a form found in nature. For example, a naturally occurring or wild-ty pe polypeptide or polynucleotide sequence is a sequence present inan organism that can be isolated from a source in nature and that has not been intentionally modified by human manipulation, with the sole exception that wild-type polypeptide or polynucleotide sequences as identified herein may include a tag, such as a histidine tag (6xHis tag (SEQ ID NO: 131)). Herein, “wild-type"’ polypeptide or polynucleotide sequences may be denoted “WT.”

[0064] “Operably linked” is defined herein as a configuration in which a control sequence is appropriately placed at a position relative to a polynucleotide sequence (i.e., in a functional relationship) such that the control sequence directs the expression of the polynucleotide and / or a polypeptide encoded by the polynucleotide.

[0065] A “promoter sequence” is a nucleic acid sequence that is recognized by a host cell for expression of a polynucleotide. The control sequence may comprise an appropriate promoter sequence. The promoter sequence contains transcriptional control sequences, which mediate the expression of the polynucleotide. The promoter may be any nucleic acid sequence which shows transcriptional activity in the host cell of choice including mutant, truncated, and hybrid promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides either homologous or heterologous to the host cell.

[0066] The terms “engineered,” “recombinant,” “variant,"’ and “non-naturally occurring,"’ when used with reference to. e.g., a polynucleotide, polypeptide, or cell, refers to a material, or a material corresponding to the natural or native form of the material, that has been modified in a manner that would not otherwise exist in nature. Non-limiting examples include, among others, recombinant cells expressing genes that are not found within the native (non-recombinant) form of the cell or express native genes that are otherwise expressed at a different level.

[0067] As used herein, the terms “percent identity,” and “percent identical” refer to comparisons between polynucleotide sequences or polypeptide sequences and are determined by comparing two optimally aligned sequences over a comparison window', wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (i.e.. gaps) as compared to the reference sequence for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at w'hich either the identical nucleic acid base or amino acid residue occurs in both sequences or a nucleic acid base or amino acid residue is aligned with a gap to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity. Determination of optimal alignment and percent sequence identity can be performed using the BLAST and BLAST 2.0 algorithms (see e.g., Altschul et al., 1990, J. Mol. Biol. 215: 403-410;and Altschul et al.. 1977, Nucleic Acids Res. 3389-3402). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information website.

[0068] Briefly, the BLAST analyses involve first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as, the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T. and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation I of 10, M = 5, N = -4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation (I) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, 1989, Proc. Natl. Acad. Sci. USA 89: 10915).

[0069] Numerous other algorithms are available that function similarly to BLAST in providing percent identify for two sequences. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman, 1981, Adv. Appl. Math. 2:482, by the homology alignment algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48:443. by the search for similarity method of Pearson and Lipman. 1988, Proc. Natl. Acad. Sci. USA 85:2444, by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin Software Package), or by visual inspection (see generally, Current Protocols in Molecular Biology', F. M. Ausubel et al., eds., Current Protocols, ajoint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)). Additionally, determination of sequence alignment and percent sequence identify' can employ the BESTFIT or GAP programs in the GCG Wisconsin Software package (Accelrys, Madison WI), using default parameters provided.

[0070] As used herein, the terms "stereoselectivity’ and “stereospecificity” refer to the preferential formation in a chemical or enzymatic reaction of one stereoisomer over another. Stereoselectivity can be partial, where the formation of one stereoisomer is favored over the other, or it may be complete where only one stereoisomer is formed. When the stereoisomers are enantiomers, the stereoselectivity is referred to as enantioselectivity, the fraction (typically reported as a percentage) of one enantiomer in the sum of both. It is commonly alternatively reported in the art (typically as a percentage) as the enantiomeric excess (EE) calculated therefrom according to the formula [major enantiomer - minor enantiomer] / [major enantiomer + minor enantiomer]. Where the stereoisomers are diastereomers, the stereoselectivity is referred to as diastereoselectivity, the fraction (typically reported as a percentage) of one diastereomer in a mixture of two diastereomers, commonly alternatively reported as the diastereomeric excess (DE). Enantiomeric excess and diastereomeric excess are types of stereomeric excess.

[0071] “Regioselectivity” refers to the preferential formation in a chemical or enzymatic reaction of one regioisomer over another. Regioselectivity can be partial, where the formation of one regioisomer is favored over the other, or it may be complete where only one regioisomer is formed. “Highly regioselective” refers to a chemical or enzymatic reaction that is capable of converting a substrate to its corresponding product with at least about 85% regioisomeric excess. “Regioisomers” encompass diastereomers and enantiomers. Regioselectivity may result when an enzyme demonstrates preferential catalysis of a chemical moiety at a single position or configuration in a substrate molecule than at other positions.

[0072] “Chemoselectivity” refers to the preferential formation in a chemical or enzy matic reaction of one product over another.

[0073] “Conversion” refers to the enzymatic transformation of a substrate to the corresponding product. “Percent conversion” refers to the percent of the substrate that is converted to the product within a period of time under specified conditions. Thus, for example, the “enzy matic activity” or “activity” of a polypeptide can be expressed as “percent conversion” of the substrate to the product.

[0074] As used herein, the terms “thermostability” and “thermal stability” refer to the ability of an enzyme to maintain the same or similar level of activity in catalyzing a reaction (more than 60% to 80% of product conversion, for example) after exposure to elevated temperatures (e.g., 37 °C to 50 °C, or 37 °C to 80 °C) for a period of time (e.g., 0.5 h to 24 h) relative to the corresponding enzyme at room temperature. Thermostability may be measured using a heat challenge test, e.g., measuring enzyme activity for about 60 minutes after exposure to elevated temperatures (e.g., 35 °C, 37 °C, or 39-40°C).

[0075] As used herein, the terms “biocatalysis,” “biocatalytic.” “biotransformation.” and “biosynthesis” refer to the use of enzymes to perform chemical reactions on organic compounds.

[0076] As used herein, “phosphate tolerance” refers to the capability of an enzy me to exhibit activity under a range of phosphate salt concentrations that are relevant to the process, e.g., a high phosphate salt concentration. Enzyme phosphate tolerance can be measured by measuring the activity of the enzyme under low phosphate concentrations (e g., 1-50 mM) relative to activity under higher phosphate concentrations (e.g., >50 mM); a phosphate tolerant enzyme displays minimal difference in activity between low and higher phosphate concentrations.

[0077] As used herein, “pH tolerance” refers to the capability of an enzyme to exhibit activity under a range of pH values during a reaction. It may refer to a maintained activity in the presence of a pH reduced in the presence of inorganic phosphate molecules. “pH stable” or “pH tolerant” refers to a polypeptide that maintains similar activity7(more than e.g., 60% to 80%) after exposure to high or low pH (e.g., 4.5 to 6 or 8 to 12) for a period of time (e.g., 0.5h to 24h) relative to an untreated enzyme.

[0078] As used herein, “polyphosphate” refers to an oligomer of two or more phosphate molecules (e.g., ions) covalently linked together.

[0079] As used herein, the terms “co-solvent tolerance” and “solvent tolerance” refer to the capability of an enzyme to exhibit activity under a range of co-solvent concentrations, such as organic co-solvents, that are relevant to the chemical process occurring in an aqueous medium (e.g., an aqueous buffer such as HEPES). Exemplary organic co-solvents are acetonitrile, acetone, dimethyl sulfoxide (DMSO), hexane and heptane. A co-solvent may or may not be present with a primary aqueous solvent, in any desired reaction medium. Enzyme co-solvent tolerance can be measured by measuring the activity of the enzyme under low co-solvent concentrations (e.g., no acetonitrile) relative to activity under higher co-solvent concentrations (e.g., >5% or >10% v / v acetonitrile); a co-solvent tolerant enzy me displays minimal difference in activity between low and higher co-solvent concentrations.

[0080] As used herein, the terms “substrate loading” and “substrate loads” and “substrate concentration” refer to the concentration of substrate present in a reaction vessel following a contacting of substrate to enzy me, e.g., during a synthesis reaction. Many natural enzymes exhibit little to no activity at high substrate concentrations. As used herein, the term “high substrate loading” refers to a concentration of substrate between 5 and 75 mM. Example substrate concentration ranges for “high substrate loading” include 1-10 mM, 5-15 mM, 10-20 mM, 20-30 mM, 30-40 mM, 5-45 mM, 15-45 mM, 25-45 mM, 40-50 mM, 50-60 mM, 60-75mM, 5-75 mM, 15-75 mM, 25-75 mM, or 50-75 mM. “Low substrate loading” may refer to concentrations of less than 5 mM.

[0081] As used herein, the term “product inhibition” refers to a reduction in enzy matic rate in converting substrate to product observed when a product molecule resides too long in the active site of an enzyme following its synthesis, which prevents the enzyme from proceeding to release the product and bind anew substrate molecule. The degree of product inhibition can be measured by determining the enzymatic rate over a range of product concentrations. As used herein, “reduced product inhibition” refers to a restoration of enzy matic rate in converting substrate to product in the presence of product molecules. A directed evolution screen can aim to reduce product inhibition by performing reactions in the presence of increased concentrations of substrate or product. This term may be synonymous with “refractoriness to inhibitors” and “reduced competitive inhibition.” As used herein, in the context of a one-pot cascade, this term encompasses a reduction in enzymatic activity observed when a substrate or product of a parallel reaction resides in the active site or another binding site of an enzyme that catalyzes a different reaction, preventing said enzyme from performing catalysis.

[0082] As used herein, the term “cascade” or “enzyme cascade” or “enzy matic cascade” or “multistep cascade” refers to the use of multiple enzymes, such as two or more wild-type or engineered enzymes that each catalyze a unique bond forming or bond breaking chemical transformation, to catalyze the steps of a multi-step reaction and or recycling of cofactors or cosubstrates in a one-pot process or through process, simultaneously or sequentially.Trajectory 1 : Engineered ATP-dependent amino acid ligases for tripeptide formation

[0083] The ATP-dependent ammo acid ligase enzymes described herein are the product of directed evolution from BAD 1200, a wild-type ATP-dependent amino acid ligase from Bifidobacterium adolescentis described in literature by Arai et al., Biosci. Biotechnol. Biochem. 2010, 74. 1572-1577. This protein is encoded by a polynucleotide of SEQ ID NO: 1 and a polypeptide sequence as set forth below in SEQ ID NO: 2:MKVLLLQQPKSFSNYPKWIEEIQERFDCLEVMVFTSNDRAAHHSWPSSVIKEIEV SDYSSDSATAKFFDIVRKFKPDRIVSSSEEDVLRVAEARSLFGIPGLQHELALSCRD KVTMKQSALDAGLKIIPYTTCQGFGDIISAFDRWETVVLKPRWGAGSAGITILHSK DDLPALATKPEFIRNVHSNQYYLEEYCSGSVYHVDVVYINSGSILISPSRYLVPPL DFEKQNTGSVMLDENGADYSELLRLTKQLIASFNDQTIPNVMHIEFYKNETGDFV FGEMAARRGGGLIKQELAAAYGIDQSKANFLLELGLVDADANITRSSQYGILLETAGLNWPKEKEIPDWAVLESVGKKKGIAHNSVDSDRKFL1SGKNESEI1QRSNYLIN SQYNE (SEQ ID NO: 2)

[0084] The polynucleotide of the ATP-dependent amino acid ligase wild type enzy mes described herein and subsequently used for directed evolution was codon optimized for production in the E. coll host and a sequence encoding a short histidine affinity tag (6xHis) was added at the 5’ terminus of the gene. The polypeptide sequence of this tag and short linker is MHHHHHHGS (SEQ ID NO: 117). This codon-optimized construct is encoded by polynucleotide of SEQ ID NO: 3. The polypeptide sequence of this His-tagged ligase is set forth below in SEQ ID NO: 4:MHHHHHHGSMKVLLLQQPKSFSNYPKWIEEIQERFDCLEVMVFTSNDRAAHHS WPSSVIKEIEVSDYSSDSATAKFFDIVRKFKPDRIVSSSEEDVLRVAEARSLFGIPG LQHELALSCRDKVTMKQSALDAGLKIIPYTTCQGFGDIISAFDRWETVVLKPRWG AGSAGITILHSKDDLPALATKPEFIRNVHSNQYYLEEYCSGSVYHVDVVYINSGSI LISPSRYLVPPLDFEKQNTGSVMLDENGADYSELLRLTKQLIASFNDQTIPNVMHI EFYKNETGDFVFGEMAARRGGGLIKQELAAAYGIDQSKANFLLELGLVDADANI TRSSQYGILLETAGLNWPKEKEIPDWAVLESVGKKKGIAHNSVDSDRKFLISGKN ESEIIQRSNYLINSQYNE (SEQ ID NO. 4)

[0085] In some aspects, ATP-dependent amino acid ligase enzymes of the disclosure may demonstrate improvements relative to the ATP-dependent ligase enzyme of SEQ ID NO: 4, such as improvements in enzy me activity, regioselectivity, thermostability, phosphate tolerance, cosolvent tolerance, reduction of by-products and reduction in product inhibition for formation of tripeptides in the reactions shown in Table 1. The improvements can relate to a single enzyme property, such as enzymatic activity, or a combination of different enzyme properties, such as enzymatic activity and regioselectivity.

[0086] Any of the below-described properties may be measured in accordance with techniques known in the art. Any of these properties may be measured following performance of the reaction in an aqueous medium, such as HEPES buffer, in the presence or absence of cosolvent. Any of these properties may be measured following performance of the reaction in a medium containing reagents ATP, magnesium ion (Mg2+), and / or sodium polyphosphate. Any of these properties may be measured following performance of the reaction in a medium containing sodium polyphosphate and a polyphosphate kinase (PPK), such as wild-type PPK enzymes published in literature (e.g., Tavanti et al. Green Chem. 23:828-837, 2021) and those found in public databases (e.g., Uniprot ID: A0A3D5XRJ5_9FIRM, Genbank ID: CCX64920. 1), including e.g.,PPK22 or PPK.12. Suitable reaction conditions under which these improved properties of the engineered ATP-dependent amino acid ligases are further described in Examples 4-8.

[0087] In some embodiments, the ATP-dependent amino acid ligase enzymes of the disclosure exhibit improved activity on substrate compound 2c (a dipeptide) relative to the engineered ligase of SEQ ID NO: 4. in the production of tripeptide compound 3d. In some embodiments, the ATP-dependent amino acid ligase enzymes of the disclosure exhibit improved activity on compound Id relative to the ligase of SEQ ID NO: 4, in the production of compound 3d. In some embodiments, the ATP-dependent amino acid ligase enzymes of the disclosure exhibit improved activity in converting compound 2c into compound 3d relative to the ligase of SEQ ID NO: 4.

[0088] In some embodiments, the ligase enzymes of the disclosure may demonstrate improvements in the rate of enzymatic activity, i.e., the rate of converting the substrate to the product. In some embodiments, the ligase polypeptides are capable of converting the substrate to the product at a rate that is at least 1.5-times, 2-times, 3-times, 4-times, 5-times, 10-times, 25- times, 50-times, 100-times, 150-times. 200-times. 400-times, 1000-times, 3000-times, 5000- times, 7000-times or more than 7000-times the rate exhibited by the enzyme of SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent amino acid ligases catalyze the formation of products depicted in Formula III (Table 1) with at least about 1.1-fold, 5-fold, 10-fold, 100- fold or 1.000-fold or more the activity of SEQ ID NO: 4, under suitable reaction conditions.

[0089] The present disclosure provides numerous exemplary ATP-dependent amino acid ligases (Trp-ligases) capable of generating the tripeptides show n in Table 1. Those exemplary polypeptides were evolved from SEQ ID NO: 4 and exhibit improved properties, particularly improved activity’ in the conversion of various amino acid nucleophiles and amino acid electrophiles, including the conversion of compounds la and 2a to tripeptide compound 3a, the conversion of compound lb and 2b to tripeptide compound 3b, the conversion of compound 1c and 2b to tripeptide compound 3c and compound Id and 2c to tripeptide compound 3d, as depicted in Table 1, Reactions A-D. These exemplary engineered ATP-dependent Trp-ligase enzymes having tripeptide formation activity have amino acid sequences that include one or more residue differences as compared to SEQ ID NO: 4 as depicted in the accompanying sequence listing with sequence identifiers SEQ ID NOs: 6-50. In various embodiments, these variants comprise amino acid sequences having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%. or 99% sequence identity’ to SEQ ID NO: 4. In various embodiments, these variants comprise amino acid sequences having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 97%, 98%, or 99% identity to any one of SEQ ID NOs: 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, or 50. In some embodiments, these variants comprise aminoacid sequences comprising any one of SEQ ID NOs: 6, 8. 10. 12. 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, or 50.

[0090] In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 92.5%, 93%, 94%, 95%, 96%. 97%. 97.5%. 98%. or 99% sequence identity to any one of SEQ ID NOs: 40. 42, 44, 46, 48, and 50. In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 40, 42, and 44. Polypeptides comprising any of SEQ ID NOs: 40, 42, 44, 46, 48, and 50 are provided. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 40. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 42. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 44. In some embodiments, the enzyme polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 48. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 50. Polypeptides consisting essentially of any of SEQ ID NOs: 40, 42, 44, 46, 48, and 50 are provided.

[0091] In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 40. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 42. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 44. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 48. In some embodiments, the polypeptide comprises the amino acid sequence of SEQ ID NO: 50.

[0092] In some aspects, provided are engineered polypeptides comprising an amino acid sequence that comprises a stretch of at least 100, 150, 200, 250, 300, 325, 350, or 375 consecutive amino acids of any one of SEQ ID NOs: 40, 42, 44, 46, 48, and 50. In some embodiments, the polypeptides comprise an amino acid sequence that comprises a stretch of at least 100, 150, or 200 or more consecutive amino acids of any one of SEQ ID NOs: 40, 42, and 44. In some aspects, provided are engineered polypeptides comprising an amino acid sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9. 10, 11, 12, 13, 14, 15, 15-25, 25-35, 35-50, 50-75, or more than 75 amino acids relative to the amino acid sequence of any one of SEQ ID NOs: 40, 42, 44, 46, 48, and 50. The polypeptides may comprise an amino acid sequence that differs by 1, 2, 3, 4, or 5 amino acids relative any one of SEQ ID NOs: 40, 42, and 44.

[0093] In some embodiments, such engineered Trp-ligase polypeptides are also capable of converting the substrate to the product with a percent regioisomeric excess of at least about 80%, 85%, 90%, 95%, 96%, 98%, 99%, or more than 99%. In some embodiments, such ligase polypeptides are also capable of converting the substrate to the product with a percent regioisomeric excess of at least about 85%. In some embodiments, such ligase polypeptides are also capable of converting the substrate to the product with a percent regioisomeric excess of at least about 95% or about 99%. In some embodiments, the ligase polypeptide is highly regioselective, wherein the polypeptide can catalyze formation of the product in greater than about 99%. 99.1%. 99.2%. 99.3%. 99.4%. 99.5%. 99.6%. 99.7%. 99.8%. or 99.9% regioisomeric excess.

[0094] In some embodiments, the engineered ATP-dependent Trp-ligase enzymes of the disclosure exhibit improved regioselectivity in the production of tripeptide compound 3d relative to the ligase of SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent amino acid ligases catalyzing the formation of products of Formula III depicted in Table 1 can also form a regioisomer of product 3d. In particular embodiments, the Trp-ligase enzymes catalyzing the formation of products of Formula III can form the desired product 3d with regioisomeric ratios of at least 1 : 10, 10: 1, 50: 1 and more than 100: 1, relative to the ligase SEQ ID NO: 4. The disclosed Trp-ligases may exhibit regioselectivity for compound 3d over undesired by-products such as dimers, trimers and other oligomers of compounds according to Formula I or IV (i.e., oligomers of compound 4e or compound 41) and a by-product of ligation of compound 3d and compound Id. As such, the disclosed Phe-ligases may have activity' that generates oligomers of compound 4e or compound 4f and / or a product of ligation of compound 3d and compound Id in reduced amounts relative to the ligase of SEQ ID NO: 4.

[0095] In some embodiments, the engineered ATP-dependent Trp-ligases catalyzing the formation of products of Formula III may also catalyze the formation of by-products such as dimers, trimers and oligomers of substrates with Formula I. Any of the disclosed engineered polypeptides (ligases) may generate a ratio of product over by-products of at least about 500: 1, 1000: 1, 1500: 1, 2000: 1, 2500: 1, or more than 2500: 1. In some embodiments, these enzymes generate by-products at a level of less than 0.5%, thus exhibiting a desired ratio of product over by-products of over 2000: 1. Any of the disclosed engineered polypeptides (ligases) may generate fewer by-products than the ligase of SEQ ID NO: 4. In particular embodiments, the disclosed engineered ligases have a desired ratio of product over by-products of at least about 1 : 10, 10: 1, 30: 1, 50: 1, 90: 1 or more, relative to SEQ ID NO: 4. In some embodiments, the engineered ATP- dependent amino acid ligase enzy mes cataly zing the formation of products of Formula III canform the desired product 3d over dimers of Id with a selectivity of at least 1: 1. 1 : 10. 10: 1. 30: 1. 50: 1, 90: 1, or more relative to SEQ ID NO: 4.

[0096] In some embodiments the engineered ATP-dependent ligases depicted in Table 1 are engineered to generate products of Formula III at substrate loading concentrations of at least about 10 mM, 25 mM, 50 mM, or more than 50 mM with a percent conversion of at least about 0.001%, about 0.01%, about 0.1%, about 1%, about 10%, about 40%, about 60%, about 80% or at least about 95%. In some embodiments the engineered ATP-dependent Trp-ligases are engineered to generate these products at substrate concentrations (i.e., substrate loadings) of about 10-20 mM, 20-30 mM, 30-40 mM, 25-45 mM, 40-50 mM, 50-60 mM, or 50-75 mM. In some embodiments, the disclosed ligase enzymes exhibit activities or efficiencies in the presence of higher substrate loads in an enzymatic cascade relative to the corresponding activities of SEQ ID NO: 4.

[0097] In various embodiments, the disclosed engineered Trp-ligases exhibit improved kinetics of the reaction, i.e., reduced time of reaction, relative to the enzyme of SEQ ID NO: 4. In some embodiments, these enzy mes are able to generate a conversion of product of at least about 80% or at least about 90% in a reaction time less than about 96 hours, about 48 hours, about 20 hours, about 6 hours, or less than 6 hours. In some embodiments, these enzymes are able to generate a conversion of product of at least about 80% or at least about 90% in a reaction time less than about 20 hours, about 6 hours, or less than 6 hours at substrate concentrations of at least about 10 mM, 25 mM, 50 mM, or more than 50 mM in a reaction time less than about 96 hours, about 48 hours, about 20 hours, about 10 hours, or less.

[0098] In some embodiments, the Trp-ligase enzymes of the disclosure exhibit improved thermostability relative to SEQ ID NO: 4. The ATP-dependent amino acid hgase enzymes of the disclosure may7exhibit higher activity at low temperatures relative to SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent amino acid ligase enzy mes catalyzing the formation of products of Formula V depicted in Table 2 have been improved for thermostability and can maintain at least 50% of its enzymatic activity at temperatures of at least about 35 °C. 37°C, 39°C, 42°C, 43°C, 50°C or more degrees compared to SEQ ID NO: 4 when subjected to a heat challenge test for about 60 minutes.

[0099] In some embodiments, the engineered ATP-dependent amino acid ligases catalyzing the formation of products of Formula III depicted in Table 1 have improved phosphate tolerance of at least about 1.2-fold, 2-fold, or 3-fold or more, relative to the hgase of SEQ ID NO: 4.Table 1. Trp-ligase catalyzed ATP-dependent amino acid ligation reactions for tripeptide formation (Trajectory 1)

[0100] Compound Id is a modified fluorinated tryptophan that contains a 6-aminohexyl moiety and may be referred to herein as “aFTrpT “FTrp” or “Linker Trpf Compound Id has a molecular weight of 321 .40. Compound 2c is a Proline-Threonine derivative dipeptide that contains a modified proline residue and modified threonine residue, having a molecular weight of 388.46. Compound 3d is a modified tryptophan-proline-threonine tripeptide that contains modified proline, threonine, and tryptophan residues. This tripeptide retains the 6-aminohexyl tail or ■’linker."Trajectory72: Engineered ATP-dependent amino acid ligases for tetrapeptide formation

[0101] In some embodiments, provided herein are ATP-dependent ligase enzymes that are capable of activating phenylalanine analog carboxylic acids of Formula IV via reaction with ATP to form acylphosphate intermediates for nucleophilic attack by tryptophan-proline-threonine analog nucleophiles of Formula III to produce phenylalanine-tryptophan-proline-threonine analog tetrapeptides of Formula V as depicted in Table 2. In particular embodiments, SEQ ID NO: 4 showed no detectable activity for substrates illustrated in Reactions F. G, H. I. J and K in Table 2 below, and therefore it was determined that protein engineering of the ligase was necessary'.

[0102] In particular embodiments, provided herein are engineered variants of SEQ ID NO: 2 / 4 BAD1200 ligase Bifidobacterium adolescentis'). In some embodiments, provided are engineered variants that have been engineered through successive rounds of directed evolution. In some embodiments, provided are engineered variants that have been engineered through 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 successive rounds of directed evolution. In some embodiments,provided are engineered variants that have been evolved through 10-15, 15-20. 20-25, or more than 25 successive rounds of directed evolution.

[0103] In some embodiments, the ATP-dependent amino acid ligase enzymes of the disclosure may exhibit improvements relative to the ATP-dependent ligase enzyme of SEQ ID NO: 4, such as increases in enzyme activity, regioselectivity, thermostability, phosphate tolerance, organic cosolvent tolerance, reduction in by-product formation, and reduction in product inhibition for tetrapeptide formation in the reactions shown in Table 2. These improvements can relate to a single enzyme property', such as enzy matic activity, or a combination of different enzyme properties, such as enzymatic activity and thermostability.

[0104] Any of the below-described properties may be measured in accordance with techniques know n in the art. Any of these properties may be measured following performance of the reaction in an aqueous medium, such as HEPES buffer, in the presence or absence of cosolvent. Any of these properties may be measured following performance of the reaction in a medium containing reagents ATP, magnesium ion (Mg2+). and / or sodium polyphosphate. Any of these properties may be measured following performance of the reaction in a medium containing sodium polyphosphate and a polyphosphate kinase (PPK), such as wild-type PPK22 or PPK12. Suitable reaction conditions under which the above-stated improved properties of the engineered ATP- dependent amino acid ligases are further described in Examples 4-8.

[0105] The present disclosure provides numerous exemplary ATP-dependent amino acid ligase enzymes capable of generating tetrapeptides shown in Table 2. Those exemplary polypeptides were engineered from SEQ ID NO: 4 and exhibit improved properties, particularly improved activity in the conversion of various amino acid nucleophiles and amino acid electrophiles, including the conversion of compounds 4e and 3a to the tetrapeptide compound 5e, conversion of compounds 4f and 3a to the tetrapeptide compound 5f, conversion of compounds 4f and 3e to tetrapeptide compound 5g, conversion of compounds 4f and 3b to tetrapeptide compound 5h, conversion of compounds 4f and 3f to tetrapeptide compound 5i, conversion of compounds 4f and 3c to tetrapeptide compound 5j. and compounds 4f and 3d to tetrapeptide compound 5k as depicted in Table 2, Entries E-K. These exemplary engineered ATP-dependent amino acid ligase enzymes having tetrapeptide formation activity comprise amino acid sequences that include one or more residue differences as compared to SEQ ID NO: 4 as depicted in the accompanying sequence listing with sequence identifiers SEQ ID NOs: 52-116 and 126-130.

[0106] In various embodiments, these variants comprise amino acid sequences having at least about 80%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 92.5%, 93%, 94%, 95%, 96%, 97%, 97.5%, 98% or 99% sequence identity to SEQ ID NO: 4. In various embodiments, thesevariants comprise amino acid sequences having at least about 80%. 85%. 90%. 92.5%. 94%. 95%, 96%, 98%, or 99% identity to any one of SEQ ID NOs: 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, and 116. In some embodiments, these variants comprise amino acid sequences comprising any one of SEQ ID NOs: 52. 54. 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78. 80. 82. 84. 86. 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, and 116.

[0107] In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least about 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 108, 110, 112, 114, and 116. In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 90% sequence identity' to any one of SEQ ID NOs: 108, 110, and 112. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 108. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity’ to SEQ ID NO: 110. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 112. In some embodiments, the enzy me polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 114. In some embodiments, the enzyme polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 116.

[0108] Polypeptides comprising any of SEQ ID NOs: 108, 110, 1 12, 114, and 116 are provided. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 108. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 110. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 112. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 114. In some embodiments, the enzyme comprises the amino acid sequence of SEQ ID NO: 116.Polypeptides consisting essentially of any of SEQ ID NOs: 108, 110, 112, 114, and 116 are provided.

[0109] In some aspects, provided are engineered polypeptides comprising an amino acid sequence that comprises a stretch of at least 100, 150, 200, 250, 300, 325, 350, or 375 consecutive amino acids of any one of SEQ ID NOs: 108, 110, 112, 114, and 116. In some embodiments, the polypeptides comprise an amino acid sequence that comprises a stretch of at least 100, 150, or 200 or more consecutive amino acids of any one of SEQ ID NOs: 108. 110, and 1 12. In some aspects, provided are engineered polypeptides comprising an amino acid sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-25, 25-35, 35-50, 50-75, or more than 75 amino acids relative to the amino acid sequence of any one of SEQ ID NOs: 108,110, 112. 114. and 116. The polypeptides may comprise an amino acid sequence that differs by 1, 2, 3, 4, or 5 amino acids relative any one of SEQ ID NOs: 108, 110, and 112.

[0110] In some embodiments, the ATP-dependent amino acid ligase enzy mes (Phe-ligases) of the disclosure exhibit improved activity on substrate compound 3d relative to the ligase of SEQ ID NO: 4, in the production of tetrapeptide compound 5k. In some embodiments, the ATP- dependent amino acid ligase enzymes of the disclosure exhibit improved activity on compound 4f relative to the ligase of SEQ ID NO: 4, in the production of compound 5k. In some embodiments, the ATP-dependent amino acid ligase enzymes of the disclosure exhibit improved activity in converting compound 3d into compound 5k relative to the ligase of SEQ ID NO: 4.

[0111] In some embodiments, the disclosed ligases are engineered to generate a percent conversion of substrate to product (e.g., compound 2c into compound 3d or converting compound 3d into compound 5k) of at least about 0.01%, about 0.1%, at least about 1%, at least about 10%. at least about 40%, at least about 60%, at least about 80% or at least about 90%. In some embodiments, the disclosed ligases can generate a percent conversion of at least about 80%. In some embodiments, the disclosed ligases can generate a percent conversion of at least about 90%.

[0112] In some embodiments, the engineered ATP-dependent Phe-ligase enzy mes catalyze the formation of products of Formula V (Table 2) and exhibit at least about 5-fold. 10-fold, 100-fold, 1,000, 10,000-fold or more the activity compared to SEQ ID NO: 4, under suitable reaction conditions.

[0113] In some embodiments, the engineered Phe-ligases of the disclosure exhibit improved regioselectivity in the production of tetrapeptide compound 5f relative to the wild-type ligase of SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent Phe-ligases catalyzing the formation of products of Formula V depicted in Table 2 can also form a regioisomer of product 5k. In particular embodiments, the engineered ATP-dependent amino acid ligase enzymes catalyzing the formation of products of Formula V can form the desired product 5k with regioisomeric ratios of at least 1: 10, 10: 1, 50: 1 and more than 100: 1, relative to the ligase SEQ ID NO: 4.

[0114] In some embodiments, the Phe-ligases of the disclosure exhibit improved specificity in the form of reduced generation of undesired by-products (e.g., dimers of the tripeptide, dimers of the tetrapeptides, or dimers and trimers of compound 3d) relative to SEQ ID NO: 4. In some embodiments, the engineered Phe-ligases catalyzing the formation of products of Formula V can form by-products from substrate Id, such as dimers, when present in solution. In particular embodiments, the engineered ATP-dependent amino acid ligase enzymes catalyzing theformation of products of Formula V can form the desired product 5k over dimers of Id with a selectivity of at least 10: 1, 25: 1, 50: 1, 99: 1, 100: 1, or more than 100: 1 relative to SEQ ID NO: 4. In some embodiments, the Phe-ligases catalyzing the formation of products of Formula V may generate a by-product from the ligation of substrate Id to 5k when present in solution. In particular embodiments, the engineered ATP-dependent amino acid ligase enzymes catalyzing the formation of products of Formula V are able to form the desired product 5k over a ligation product of substrate Id to 5k, with a selectivity of at least 2: 1, 5: 1, 10: 1, 20: 1, 50: 1 or more than 100: 1, relative to the ligase of SEQ ID NO: 4. The disclosed Phe-ligases may exhibit regioselectivity for compound 5k over undesired by-products such as dimers, trimers and other oligomers of compounds according to Formula IV (i.e. , oligomers of compound 4e or compound 41) and a product of ligation of compound 5k to compound Id. As such, the disclosed Phe-ligases may have activity' that generates oligomers of Compound 4e or Compound 4f, or a product of ligation of 5k to Id, in reduced amounts relative to the ligase of SEQ ID NO: 4.

[0115] In some embodiments, the Phe-ligase enzymes of the disclosure exhibit improved phosphate tolerance relative to SEQ ID NO: 4. In some embodiments, the engineered ATP- dependent Phe-ligases catalyzing the formation of products of Formula V have improved phosphate tolerance of at least about 1.5-fold, 2.5-fold, 5-fold, or more, relative to SEQ ID NO: 4. In some embodiments, the disclosed ligase enzymes exhibit improved phosphate tolerance in an enzy matic cascade relative to and SEQ ID NO: 4.

[0116] In some embodiments, the Phe-ligase enzymes of the disclosure exhibit improved thermostability relative to SEQ ID NO: 4. The ATP-dependent amino acid ligase enzymes of the disclosure may exhibit higher activity at low temperatures relative to SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent amino acid ligase enzymes catalyzing the formation of products of Formula V have been improved for thermostability and can maintain at least 50% of its enzy matic activity' at temperatures of at least about 35 °C, 37 °C, 39 °C, 42 °C, 43 °C, 50 °C or more degrees compared to SEQ ID NO: 4 when subjected to a heat challenge test for about 60 minutes.

[0117] In some embodiments, the Phe-ligase enzymes of the disclosure exhibit activities or efficiencies in the presence of higher substrate loads relative to the corresponding activity of SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent amino acid ligases catalyzing the formation of products of Formula V are able to catalyze the production (i.e., ligation) of tetrapeptide products of Formula V (e.g., tetrapeptide compound 5k) at substrate concentrations of at least about 1 mM, 5 mM, 10 mM, 25 mM, 50 mM, 50 mM to 75 mM, or more than 75 mM. In some embodiments, these engineered ligases are engineered to catalyze the production of thesetetrapeptide products at substrate concentrations of 1-10 mM. 5-15 mM, 10-20 mM, 20-30 mM. 30-40 mM, 5-45 mM, 15-45 mM, 25-45 mM, 40-50 mM, 50-60 mM, 60-75 mM, 5-75 mM, 15- 75 mM, 25-75 mM, or 50-75 mM.

[0118] In various embodiments, the disclosed engineered Phe-ligases exhibit improved kinetics in the reaction, i.e., reduced time of reaction, relative to the ligase enzyme of SEQ ID NO: 4. In some embodiments, these enzymes are able to generate a conversion of substrate of at least about 80% or at least about 90% in a reaction time less than about 96 hours, about 48 hours, about 20 hours, about 6 hours, or less than 6 hours. In some embodiments, these enzy mes are able to generate a conversion of substrate of at least about 80% or at least about 90% in a reaction time less than about 20 hours, about 6 hours, or less than 6 hours at substrate concentrations of 1-10 mM, 5-15 mM, 10-20 mM, 20-30 mM, 30-40 mM, 5-45 mM, 15-45 mM, 25-45 mM, 40-50 mM, 50-60 mM, 60-75 mM, 5-75 mM, 15-75 mM, 25-75 mM, or 50-75 mM.

[0119] In some embodiments, the disclosed engineered Phe-ligases exhibit activities or efficiencies in the presence of higher substrate loads in an enzymatic cascade relative to the corresponding activities of SEQ ID NO: 4.

[0120] In some embodiments, the disclosed Trp-ligase and Phe-ligase enzy mes of the disclosure exhibit improved tolerance to organic co-solvents relative to SEQ ID NO: 4. In some embodiments, the disclosed Trp-ligase and Phe-ligase enzymes of the disclosure exhibit improved tolerance to acetonitrile co-solvent (e.g., concentrations of acetonitrile in buffer of 15%, 20%, 25%, or 30%), relative to SEQ ID NO: 4.

[0121] In some embodiments, the disclosed Phe-ligase enzymes of the disclosure exhibit improved pH tolerance relative to SEQ ID NO: 4. In some embodiments, the disclosed Phe-ligase enzymes of the disclosure exhibit improvements (i.e.. reductions) in product inhibition, for instance at high substrate loads, relative to SEQ ID NO: 4.

[0122] In some embodiments, any of the disclosed Phe-ligase enzymes of the disclosure exhibit improved activity' in the presence of increased phosphate loads relative to SEQ ID NO: 4.

[0123] In some embodiments, any of the disclosed Trp-ligase and Phe-ligase polypeptides further comprises a tag (e g., an affinity tag). Any suitable tag may be used, e.g., a 6xHis tag or an 8xHis tag (SEQ ID NO: 133), a FLAG tag, a fluorescent protein tag (e.g., green fluorescent protein (GFP), yellow fluorescent protein (YFP), or red fluorescent protein (RFP)). a hemagglutinin (HA) tag. an ALFA-tag, a V5-tag, a Myc-tag, a SPOT-tag, a T7-tag, or an NE-tag. In some embodiments, the affinity tag is a His tag. In some embodiments, the affinity tag comprises the amino acid sequence of MHHHHHHGS (SEQ ID NO: 117). In some embodiments, the affinity7tag comprises the amino acid sequence of GSHHHHHHHHSG (SEQID NO: 118) (an 8xHis tag). In some embodiments, the polypeptide comprises an N-terminal methionine residue, and the epitope tag is inserted immediately following the N-terminal methionine residue, e.g., relative to a reference sequence.

[0124] In some embodiments, the disclosed Trp-ligase and Phe-ligase polypeptides do not contain an affinity tag. In some embodiments, the disclosed Trp-ligase and / or Phe-ligase polypeptides do not contain a histidine tag, such as an 8xHis tag. For example, the Trp-ligase polypeptide may comprise any of SEQ ID NOs: 119-121, which correspond to the sequences of SEQ ID NOs: 40, 42, and 44 without a histidine tag, respectively. As another example, the Phe- ligase polypeptide may comprise any of SEQ ID NOs: 122-124, which correspond to the sequences of SEQ ID NOs: 110, 112, and 114 without a histidine tag, respectively.Table 2. Phe-ligase catalyzed ATP-dependent amino acid ligation reactions for tetrapeptide formation (Trajectory 2)

[0125] Compound 4f is a modified phenylalanine ((S)-2-amino-3-(3- (aminomethyl)phenyl)propanoic acid or aminomethylphenylalanine) that may be referred to herein as '‘aPhe” (molecular weight of 194.23). Compound 3d is the product of the abovedescribed reaction D shown in Table 1, i.e., a tryptophan-proline-threonine analog tripeptide that contains modified try ptophan, proline, and threonine residues. Compound 5k is a modified Phenylalanine-Tryptophan-Proline-Threonine tetrapeptide (SEQ ID NO: 132) that contains modified phenylalanine, tryptophan, proline, and threonine residues. This tetrapeptide has a molecular weight of 793.94.In some embodiments compound 5k can be generated from a cascade reaction by the addition of compound 4f, compound Id and compound 2c with a corresponding ATP-dependent amino acid ligase enzymes that forms tripeptides (Trp-ligase from any one of SEQ ID NOs: 24-50 and a corresponding ATP-dependent amino acid ligase that forms tetrapeptides (Phe-ligase from any one of SEQ ID NOs: 76-116) as shown in FIG. 1, together with Mg2+, ATP, and an ATPrecy cling system including a phosphate donor such as polyphosphate and an enzyme regenerating ATP such as polyphosphate kinase.

[0126] In some embodiments, the ATP-dependent amino acid ligases described herein demand equimolar concentrations of ATP or the use of an ATP regeneration system via activated phosphate sources for the formation of tripeptides and tetrapeptides. Exemplary ATP regeneration systems are well known to those skilled in the art and can comprise the addition of sub-stoichiometric concentrations of adenosine diphosphate or adenosine triphosphate in the presence of propionyl phosphate and an acetate kinase. In some embodiments the ATP recycling system can be comprised of a wild-type polyphosphate kinase (PPK). polyphosphate, and a cofactor such as adenosine monophosphate, adenosine diphosphate or adenosine triphosphate.

[0127] In some embodiments, the wild-type polyphosphate kinase is PPK12. (See Tavanti et al., Green Chem. (2021) 23, 828-837, which is herein incorporated by reference.) In some embodiments, the wild-type polyphosphate kinase is PPK22 (see id.). In some embodiments, the polynucleotides encoding these wild-type polyphosphate kinases are codon optimized for expression in E. coll. In some embodiments the polynucleotides encoding these wild-type polyphosphate kinases further encode affinity tags, for example a polyhistidine tag.

[0128] In some aspects, provided herein are compositions comprising any of the disclosed engineered polypeptides. In some embodiments, provided are stable preparations of any of the disclosed engineered Trp-ligase and Phe-ligase enzymes. In some embodiments, these preparations comprise immobilized enzymes. Immobilized enzyme preparations have a number of recognized advantages. They can confer shelf life to enzyme preparations, they can improve reaction stability, they can enable stability in organic solvents, they can aid in protein removal from reaction streams, as examples. ■’Stable’’ refers to the ability of the immobilized enzymes to retain their structural conformation and / or their activity in a solvent system that contains organic solvents.

[0129] ’‘Solvent stable'’ refers to a polypeptide that maintains similar activity (more than e.g., 60% to 80%) after exposure to varying concentrations (e.g.. 5% to 99%) of solvent (isopropyl alcohol, tetrahydrofuran, 2-methyltetrahydrofuran, acetone, toluene, butylacetate, methyl tert- butylether, etc.) for a period of time (<?.g, 0.5 h to 24 h) compared to an untreated enzy me. Solvent stable enzy mes may lose less than 10% activity per hour in a solvent system that contains organic solvents. Solvent stable enzymes may lose less than 9%, 8%, 7%, 6%, or 5% activity per hour in a solvent system that contains organic solvents. The disclosed preparations may comprise one or more diluents or carriers.

[0130] In some embodiments, the disclosed enzyme preparations are thermostable. In some embodiments, the disclosed enzyme preparations are solvent stable. In some embodiments, the disclosed enzyme preparations are thermostable and solvent stable. In some embodiments, the disclosed enzyme preparations are pH tolerant and / or phosphate tolerant. In some embodiments, the disclosed enzyme preparations are thermostable and solvent stable, pH tolerant, and phosphate tolerant.

[0131] Tables 3, 4, and 5 below provide exemplary engineered ATP-dependent ligase polypeptides with the ability to form tripeptides (Table 3) and tetrapeptides (Tables 4 and 5). Each row lists two SEQ ID Nos. with the odd number referring to the nucleotide (“nt”) sequence encoding the amino acid (“aa”) sequence provided by the even number. The residue differences are based on comparison to a reference sequence such as of SEQ ID NO: 4, an ATP-dependent amino acid ligase derived from Bifidobacterium adolescentis that differs from the naturally occurring enzyme BAD1200 (SEQ ID NO: 2) in that it contains a short hexahistidine (6xHis) tag at the N-terminus and in that its encoding gene is codon optimized for E. coli expression. The column listing the number of mutations (i.e., residue changes) refers to the number of amino acid substitutions as compared to the ATP-dependent amino acid ligase enzyme of SEQ ID NO: 4.

[0132] Tables 3, 4, and 5 below provide a list of the SEQ ID NOs disclosed herein with associated absolute conversions with shake flask powder as described in the examples. Conversions were defined as follows: “+” indicates less than 10% conversion of substrate to product observed for the desired reactions (listed as Reaction A through Reaction J), “++” indicates between 10 and 60% conversion and “+++” indicates more than 60% conversion.Empty fields indicate substrates were not tested under the desired conditions.Table 3. Activity of ATP-dependent amino acid ligases for tripeptide formationTable 4. Activity of ATP-dependent amino acid ligases for tetrapeptide formation(Reactions E-I)Table 5. Activity of ATP-dependent Phe-ligases for tetrapeptide formation (Reactions I-K)Polvnucleotides Encoding ATP-dependent amino acid ligases

[0133] In another aspect, the present disclosure provides polynucleotides encoding the ATP- dependent amino acid ligase enzymes disclosed herein. The polynucleotides may be operatively linked to one or more heterologous regulatory sequences that control gene expression to create a recombinant vector capable of expressing the polypeptide. Expression vectors containing a heterologous or engineered polynucleotide encoding the ATP-dependent amino acid ligase can be introduced into appropriate host cells (e g., E. coll cells) to express the corresponding ATP- dependent amino acid ligase polypeptide. The disclosed polynucleotides are derived from the polynucleotide sequence of SEQ ID NO: 3, which encodes wild-type BAD 1200 enzy me that hasbeen codon-optimized for expression in E. coii. and further encodes an N-terminal 6xHis tag. The wild-type BAD1200-encoding sequence is provided as SEQ ID NO: 1.

[0134] In various embodiments, provided herein are engineered polynucleotides that comprise nucleic acid sequences having at least about 75%, 80%, 84%, 85%, 86%, 87%, 88%. 89%. 90%. 91%. 92%. 92.5%. 93%. 94%. 95%. 96%. 97%. 97.5%. 98% or 99% sequence identity to SEQ ID NO: 3. In various embodiments, these variants comprise nucleic acid sequences having at least about 75%, 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% identity7to any one of SEQ ID NOs: 53, 55, 57, 59. 61. 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91. 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, and 115. In some embodiments, these variants comprise nucleic acid sequences comprising any one of SEQ ID NOs: 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, and 115. In some examples, the polynucleotide sequence of any of SEQ ID NOs: 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83. 85. 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109. I l l, 113, and 115 is modified to no longer code for an affinity tag. In some examples any of these polynucleotide sequences do not code for an N-terminal 6xHis tag, such as the tag having the amino acid sequence of SEQ ID NO: 117.

[0135] In some embodiments are provided engineered polynucleotides comprising a nucleic acid sequence having at least 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 39, 41, 43, 45, 47, and 49. In some embodiments are provided engineered polynucleotides comprising a nucleic acid sequence having at least 85% or 90% sequence identity to any one of SEQ ID NOs: 39, 41, and 43. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 39. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity' to SEQ ID NO: 41. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 43. In some embodiments, the enzy me polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 45. In some embodiments, the enzyme polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 47. In some embodiments, the enzyme polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 49.

[0136] Polynucleotides comprising any of SEQ ID NOs: 39, 41, 43, 45. 47. and 49 are provided. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 39. In some embodiments, the enzy me comprises the nucleic acid sequence of SEQ ID NO: 41. In some embodiments, the enzy me comprises the nucleic acid sequence of SEQ ID NO: 43.

[0137] In some aspects, provided are engineered polynucleotides comprising a nucleic acid sequence that comprises a stretch of at least 100, 150, 200, 250, 300, 325, 350, 500, 650, 800, 1000, 1100, 1200, or 1225 consecutive nucleotides of any one of SEQ ID NOs: 39, 41, 43, 45, 47, and 49. In some aspects, provided are engineered polynucleotides comprising an nucleic acid sequence that comprises a stretch of at least 250, 350. 500, 650. 800, 1000. 1100, 1200, or 1225 consecutive nucleotides of any one of SEQ ID NOs: 39, 41, 43, 45, 47, and 49. In some aspects, provided are engineered polynucleotides comprising an nucleic acid sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-25, 25-35, 35-50, 50-75, or more than 75 nucleotides relative to the sequence of any one of SEQ ID NOs: 39. 41, 43, 45, 47, and 49. The polynucleotides may comprise a nucleic acid sequence that differs by 1, 2, 3, 4, or 5 nucleic acids relative any one of SEQ ID NOs: 39, 41, and 43.

[0138] In some embodiments are provided polynucleotides comprising a nucleic acid sequence having at least 85%, 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 107, 109. I l l, 113. and 115. In some embodiments are provided engineered polynucleotides comprising a nucleic acid sequence having at least 85% or 90% sequence identity to any one of SEQ ID NOs: 107, 109, and 111. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 85% sequence identity to SEQ ID NO: 107. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 85% sequence identity to SEQ ID NO: 109. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 85% sequence identity to SEQ ID NO: 111. In some embodiments, the enzy me polynucleotide comprises a nucleic acid sequence having at least 85% sequence identity to SEQ ID NO: 113. In some embodiments, the enzyme polynucleotide comprises a nucleic acid sequence having at least 85% sequence identity to SEQ ID NO: 115.

[0139] Polynucleotides comprising any of SEQ ID NOs: 107, 109, 111, 113, and 115 are provided. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 107. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 109. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 111.

[0140] In some aspects, provided are engineered polynucleotides comprising a nucleic acid sequence that comprises a stretch of at least 250, 350, 500, 650, 800, 1000, 1100, 1200, or 1225 consecutive nucleotides of any one of SEQ ID NOs: 107, 109, 111, 113, 115. 125, 126, 127, 128, 129, and 130. In some aspects, provided are engineered polynucleotides comprising an nucleic acid sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-25, 25-35, 35-50, 50-75, or more than 75 nucleotides relative to the sequence of any one of SEQ ID NOs: 107, 109,111, 113. 115, 125, 126. 127, 128. 129, and 130. The polynucleotides may comprise a nucleic acid sequence that differs by 1, 2, 3, 4, or 5 nucleic acids relative any one of SEQ ID NOs: 107, 109, and 111.

[0141] In various embodiments, the disclosed polynucleotides are codon-optimized for expression in a particular organism, such as E. coli. Because of the knowledge of the codons corresponding to the various amino acids, availability of a protein sequence provides a description of all the polynucleotides capable of encoding the subject. The degeneracy of the genetic code, where the same amino acids are encoded by alternative or synonymous codons allows an extremely large number of nucleic acids to be made, all of which encode the improved ATP-dependent amino acid ligase enzymes disclosed herein. Thus, having identified a particular amino acid sequence, those skilled in the art could make any number of different nucleic acids by simply modifying the sequence of one or more codons in a way that does not change the amino acid sequence of the protein. In this regard, the present disclosure specifically contemplates each and every possible variation of polynucleotides that could be made by selecting combinations based on the possible codon choices, and all such variations are to be considered specifically disclosed for any polypeptide disclosed herein.

[0142] In various embodiments, the codons are preferably selected to fit the host cell in which the protein is being produced. For example, preferred codons used in bacteria are used to express the gene in bacteria; preferred codons used in yeast are used for expression in yeast; and preferred codons used in mammals are used for expression in mammalian cells. By way of example, the polynucleotide of SEQ ID NO: 1 have been codon optimized for expression in Escherichia coli to yield SEQ ID NO: 3.

[0143] In certain embodiments, all codons need not be replaced to optimize the codon usage of the ATP-dependent amino acid ligase enzyme since the natural sequence will comprise preferred codons and because use of preferred codons may not be required for all amino acid residues. Consequently, codon optimized polynucleotides encoding the ATP-dependent amino acid ligase enzymes may contain preferred codons at about 40%. 50%. 60%. 70%. 80%. or greater than 90% of codon positions of the full-length coding region.

[0144] In various embodiments, an isolated polynucleotide encoding an improved ATP- dependent amino acid ligase polypeptide may be manipulated in a variety of ways to provide for expression of the polypeptide. Manipulation of the isolated polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector. The techniques for modifying polynucleotides and nucleic acid sequences utilizing recombinant DNA methods are well known in the art. Guidance is provided in Sambrook et al., 2001, MOLECULAR CLONING:A LABORATORY MANUAL. 3rd Ed., Cold Spring Harbor Laboratory Press; and CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, Ausubel. F. ed., Greene Pub. Associates, 1998, updates to 2006.

[0145] In some embodiments, an isolated polynucleotide encoding any of the ATP-dependent amino acid ligase polypeptides herein is manipulated in a variety of ways to facilitate expression of the ATP-dependent amino acid ligase polypeptide. In some embodiments, the polynucleotides encoding the ATP-dependent amino acid ligase polypeptides comprise expression vectors where one or more control sequences is present to regulate the expression of the ATP-dependent amino acid ligase polypeptides. Manipulation of the isolated polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector utilized. Techniques for modifying polynucleotides and nucleic acid sequences utilizing recombinant DNA methods are well known in the art. In some embodiments, the control sequences include among others, promoters, leader sequences, poly adenylation sequences, propeptide sequences, signal peptide sequences, and transcription terminators. In some embodiments, suitable promoters are selected based on the host cell selection. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure, include, but are not limited to, promoters obtained from the E. coll lac operon. In addition, suitable promoters may include Streptomyces coelicolor agarase gene (dagA). Bacillus subtilis levansucrase gene (sacB), Bacillus licheniformis alpha-amylase gene (amyL), Bacillus stearothermophilus maltogenic amylase gene (amyM), Bacillus amyloliquefaciens alpha-amylase gene (amyQ), Bacillus licheniformis penicillinase gene (penP), Bacillus subtilis xylA and xylB genes, and prokary otic beta-lactamase gene (See e.g, Villa-Kamaroff et al., PROC. NATL ACAD. SCI. USA 75: 3727-3731

[1978] ), as well as the tac promoter (See e.g., DeBoer et al., PROC. NATL ACAD. SCI. USA 80: 21- 25

[1983] ).

[0146] In some embodiments, the control sequence is also a suitable transcription terminator sequence (i.e., a sequence recognized by a host cell to terminate transcription). In some embodiments, the terminator sequence is operably linked to the 3' terminus of the nucleic acid sequence encoding the enzyme polypeptide. Any suitable terminator that is functional in the host cell of choice finds use in the present disclosure.

[0147] In some embodiments, the control sequence is also a suitable leader sequence (i.e., a non-translated region of an mRNA that is important for translation by the host cell). In some embodiments, the leader sequence is operably linked to the 5' terminus of the nucleic acid sequence encoding the ATP-dependent amino acid ligase polypeptide. Any suitable leadersequence that is functional in the host cell of choice find use in the present disclosure. Exemplary leaders for E. coll will encode a ribosome binding site.

[0148] In some embodiments, the control sequence is a signal peptide (i.e., a coding region that codes for an amino acid sequence linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the cell’s secretory pathway). In some embodiments, the 5' end of the coding sequence of the nucleic acid sequence inherently contains a signal peptide coding region naturally linked in translation reading frame with the segment of the coding region that encodes the secreted polypeptide. Alternatively, in some embodiments, the 5' end of the coding sequence contains a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the secretory pathway of a host cell of choice finds use for expression of the engineered polypeptide(s). Effective signal peptide coding regions for bacterial host cells are the signal peptide coding regions include, but are not limited to, those obtained from the genes for Bacillus NC1B 11837 maltogenic amylase, Bacillus stearothermophilus alpha-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis betalactamase, Bacillus stearothermophilus neutral proteases (nprT, nprS, nprM), and Bacillus subtilis prsA. Further signal peptides are known in the art (See e.g, Simonen and Palva, Microbiol. Rev., 57: 109-137, 1993). In some embodiments, effective signal peptide coding regions for filamentous fungal host cells include, but are not limited to. the signal peptide coding regions obtained from the genes for Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase. Useful signal peptides for yeast host cells include, but are not limited to, those from the genes for Saccharomyces cerevisiae alpha-factor and Saccharomyces cerevisiae invertase.

[0149] In some embodiments, regulator}’ sequences are also utilized. These sequences facilitate the regulation of the expression of the polypeptide relative to the growth of the host cell. Examples of regulatory' systems are those that cause the expression of the gene to be turned on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory systems include, but are not limited to, the ADH2 system or GALI system. In filamentous fungi, suitable regulatory' sequences include, but are not limited to, the TAKA alpha-amylase promoter, Aspergillus niger glucoamylase promoter, and Aspergillus oryzae glucoamylase promoter.

[0150] In another aspect, the present disclosure is directed to a recombinant expression vector comprising a polynucleotide encoding ATP-dependent amino acid ligase polypeptide, and one ormore expression regulating regions such as a promoter and a terminator, a replication origin, etc., depending on the type of hosts into which they are to be introduced. In some embodiments, the various nucleic acid and control sequences described herein are joined together to produce expression vectors that include one or more convenient restriction sites to allow for insertion or substitution of the nucleic acid sequence encoding the enzyme polypeptide at such sites. Alternatively, in some embodiments, the nucleic acid sequence of the present disclosure is expressed by inserting the nucleic acid sequence or a nucleic acid construct comprising the sequence into an appropriate vector for expression. In some embodiments involving the creation of the expression vector, the coding sequence is inserted in the vector such that the coding sequence is operably linked with the appropriate control sequences for expression.

[0151] The disclosed expression vector may be any suitable vector (e.g., a plasmid or virus), that can be conveniently subjected to recombinant DNA procedures and bring about the expression of the enzyme polynucleotide sequence. The choice of the vector typically depends on the compatibility of the vector with the host cell into which the vector is to be introduced. The vectors may be linear or closed circular plasmids.

[0152] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extra-chromosomal entity, the replication of which is independent of chromosomal replication, such as a plasmid, an extra-chromosomal element, a minichromosome, or an artificial chromosome). The vector may contain any means for assuring self-replication. In some alternative embodiments, the vector is one in which, when introduced into the host cell, it is integrated into the genome and replicated together with the chromosome(s) into which it has been integrated. Furthermore, in some embodiments, a single vector or plasmid, or two or more vectors or plasmids which together contain the total DNA to be introduced into the genome of the host cell, and / or a transposon is utilized.

[0153] In some embodiments, the expression vector contains one or more selectable markers, which permit easy selection of transformed cells. A “selectable marker'’ is a gene, the product of which provides for biocide or viral resistance, resistance to heavy metals, prototrophy to auxotrophs, and the like. Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis , or markers, which confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol or tetracycline resistance. Suitable markers for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2. MET3, TRP1, and URA3. Selectable markers for use in filamentous fungal host cells include, but are not limited to, amdS (acetamidase; e.g., from A. nidulans or A. orzyae), argB (ornithine carbamoyltransferases), bar (phosphinothricin acetyltransferase; e.g, from S. hygroscopicus),hph (hygromycin phosphotransferase), niaD (nitrate reductase). pyrG (orotidine-5'-phosphate decarboxylase; e.g., from A. nidulans or A. orzyae), sC (sulfate adenyltransferase), and trpC (anthranilate synthase), as well as equivalents thereof.

[0154] In some alternative embodiments, the expression vectors contain additional nucleic acid sequences for directing integration by homologous recombination into the genome of the host cell. The additional nucleic acid sequences enable the vector to be integrated into the host cell genome at a precise location(s) in the chromosome(s). To increase the likelihood of integration at a precise location, the integrational elements preferably contain a sufficient number of nucleotides, such as 100 to 10.000 base pairs, preferably 400 to 10,000 base pairs, and most preferably 800 to 10,000 base pairs, which are highly homologous with the corresponding target sequence to enhance the probability of homologous recombination. The integrational elements may be any sequence that is homologous with the target sequence in the genome of the host cell. Furthermore, the integrational elements may be non-encoding or encoding nucleic acid sequences. On the other hand, the vector may be integrated into the genome of the host cell by non-homologous recombination.

[0155] For autonomous replication, the vector may further comprise an origin of replication enabling the vector to replicate autonomously in the host cell in question. Examples of bacterial origins of replication are ColEl ori, P15A ori. and the origins of replication of plasmids pBR322, pUC19, pACYC177 (which contains the P15A ori), or pACYC184 (which contains the P15A ori) permitting replication in E. coli. and pUBl 10, pE194, or pTA1060 permitting replication in Bacillus. The origin of replication may be one having a mutation which makes its functioning temperature-sensitive in the host cell (see e.g., Ehrlich, Proc. Natl. Acad. Set. USA 75: 1433

[1978] ), or a kanamycin resistance selection marker.

[0156] In some embodiments, more than one copy of a nucleic acid sequence of the present disclosure is inserted into the host cell to increase production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene with the nucleic acid sequence where cells containing amplified copies of the selectable marker gene, and thereby additional copies of the nucleic acid sequence, can be selected for by cultivating the cells in the presence of the appropriate selectable agent.

[0157] Many of the expression vectors for use in the present disclosure are commercially available. Suitable commercial expression vectors include, but are not limited to. Novagen’s® pET E. coli T7 expression vectors (Millipore Sigma) and the p3xFLAG™ expression vectors (Sigma- Aldrich Chemicals). In some embodiments, a pET30a expression vector is used. Othersuitable expression vectors include, but are not limited to. pBluescriptll SK(-) and pBK-CMV (Stratagene), and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen) or pPoly (e.g., Lathe et al., Gene 57: 193-201

[1987] ).

[0158] Thus, in some embodiments, a vector comprising a sequence encoding at least one variant ATP -dependent amino acid ligase is transformed into a host cell in order to allow propagation of the vector and expression of the variant ATP-dependent amino acid ligase(s). In some embodiments, the transformed host cell described above is cultured in a suitable nutrient medium under conditions permitting the expression of the variant ATP-dependent amino acid ligase(s). Any suitable medium useful for culturing the host cells finds use in the present disclosure, including, but not limited to minimal or complex media containing appropriate supplements. In some embodiments, host cells are grown in HTP media. Suitable media are available from various commercial suppliers or may be prepared according to published recipes (e.g., in catalogues of the American Type Culture Collection).Host Cells for Expression of ATP-dependent amino acid ligases

[0159] In another aspect, the present disclosure provides a host cell comprising at least one polynucleotide encoding at least one ATP-dependent amino acid ligase of the present disclosure, the polynucleotide being operatively linked to one or more control sequences for expression of the at least one ATP-dependent amino acid ligase in the host cell. Host cells suitable for use in expressing the ATP-dependent amino acid ligase polypeptides encoded by the expression vectors of the present disclosure are well know n in the art and include but are not limited to, bacterial cells, such as E. coli. Vibrio fluvialis, Streptomyces and Salmonella typhimurium cells; fungal cells, such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC Accession No. 201178)). In various embodiments, the disclosed host cells are E. coli cells. Exemplary host cells include various E. coli strains (e.g, W3110 (AfhuA) and BL21 [DE3]). Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis . or markers, which confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol, and or tetracycline resistance. Appropriate culture mediums and growth conditions for the above-described host cells are well known in the art.

[0160] Bacterial host cells for use in expressing the polypeptides encoded by the expression vectors of the present disclosure are well known in the art and include but are not limited to, E. coli, B. subtilis, B. licheniformis, B. megaterium, B. stearothermophilus, B. amyloliquefaciens , Klebsiella aerogenes, Lactobacillus kejir, Lactobacillus brevis, Lactobacillus minor,Streptomyces and Salmonella typhimurium cells. In some embodiments, the host cell or cell line is Escherichia coli BL21 or BL21(DE3).

[0161] Many prokary otic strains that find use in the present disclosure are readily available to the public from a number of culture collections such as American Type Culture Collection (ATCC). Deutsche Sammlung von Mikroorganismen und Zellkulturen GmbH (DSM), Centraalbureau Voor Schimmelcultures (CBS), and Agricultural Research Sendee Patent Culture Collection, Northern Regional Research Center (NRRL).

[0162] In some embodiments, host cells are genetically modified to have characteristics that improve protein secretion, protein stability- and / or other properties desirable for expression and / or secretion of a protein. Genetic modification can be achieved by genetic engineering techniques and / or classical microbiological techniques (e.g., chemical or UV mutagenesis and subsequent selection). Indeed, in some embodiments, combinations of recombinant modification and classical selection techniques are used to produce the host cells. Using recombinant technology-, nucleic acid molecules can be introduced, deleted, inhibited, or modified, in a manner that results in increased yields of ATP-dependent amino acid ligase variant(s) within the host cell and / or in the culture medium. In one genetic engineering approach, homologous recombination is used to induce targeted gene modifications by specifically targeting a gene in vivo to suppress expression of the encoded protein. In alternative approaches, siRNA, antisense and / or ribozyme technology find use in inhibiting gene expression. A variety of methods are known in the art for reducing expression of protein in cells, including, but not limited to deletion of all or part of the gene encoding the protein and site-specific mutagenesis to disrupt expression or activity of the gene product, (e.g., Chaveroche et al.. NUCL. ACIDS RES., 28:22 e97

[2000] ; Cho et al., MOLEC. PLANT MICROBE INTERACT., 19:7-15

[2006] ; Maruyama and Kitamoto, BIOTECHNOL LETT., 30: 1811- 1817

[2008] ; Takahashi et al., MOL. GEN. GENOM., 272: 344-352

[2004] ; and You et al., Arch. Microbiol. ,191:615-622

[2009] , all of which are incorporated by reference herein). Random mutagenesis, followed by screening for desired mutations also finds use (e.g., Combier et al., FEMS MICROBIOL. LETT.. 220: 141-8

[2003] ; and Firon et al., EUKARY. CELL 2:247-55

[2003] , both of which are incorporated by reference).

[0163] Introduction of a vector or polynucleotides for expression of the ATP-dependent amino acid ligase into a host cell can be accomplished using any suitable method known in the art, including but not limited to calcium phosphate transfection. DEAE-dextran mediated transfection, PEG-mediated transformation, electroporation, or other common techniques known in the art, which include electroporation, biolistic particle bombardment, liposome mediated transfection, calcium chloride transfection, and protoplast fusion.

[0164] In some embodiments, the engineered host cells (i.e., “recombinant host cells”) of the present disclosure are cultured in conventional nutrient media modified as appropriate for activating promoters, selecting transformants, or amplifying the ATP-dependent amino acid ligase polynucleotide. Culture conditions, such as temperature, pH and the like, are those previously used with the host cell selected for expression, and are well-known to those skilled in the art. As noted, many standard references and texts are available for the culture and production of many cells, including cells of bacterial origin.

[0165] In some embodiments, cells expressing the ATP-dependent amino acid ligase of the disclosure are grown under batch, fed-batch, or continuous fermentation conditions. Classical “batch fermentation” is a closed system, wherein the compositions of the medium are set at the beginning of the fermentation and is not subject to artificial alternations during the fermentation. A variation of the batch system is a “fed-batch fermentation” that also finds use in the present disclosure. In this variation, the substrate is added in increments as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit the metabolism of the cells and w here it is desirable to have limited amounts of substrate in the medium. Batch and fed- batch fermentations are common and well known in the art. “Continuous fermentation” is an open system where a defined fermentation medium is added continuously to a bioreactor and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the cultures at a constant high density where cells are primarily in log phase growth. Continuous fermentation systems strive to maintain steady state growth conditions. Methods for modulating nutrients and growth factors for continuous fermentation processes as well as techniques for maximizing the rate of product formation are well known in the art of industrial microbiology.

[0166] More than one copy of a nucleic acid sequence of the present disclosure may be inserted into the host cell to increase production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene with the nucleic acid sequence where cells containing amplified copies of the selectable marker gene, and thereby additional copies of the nucleic acid sequence, can be selected for by cultivating the cells in the presence of the appropriate selectable agent.

[0167] In some embodiments of the present disclosure, cell-free transcription and translation systems find use in producing the ATP-dependent amino acid ligases. Several systems are commercially available, and the methods are well-known to those skilled in the art.Methods of Producing ATP dependent ligases

[0168] In some embodiments, the ATP-dependent ligases of the present disclosure are obtained, evolved, or derived from a bacterial wild-type enzyme. Protein engineering (e.g., protein design, semi-rational engineering, structure-guided engineering, directed evolution) may be used to identify polypeptides of the present disclosure. For example, in some embodiments, to make a ligase polypeptide of the present disclosure, a polynucleotide sequence encoding a ligase polypeptide may be obtained (or derived) from the genome of Bifidobacterium adolescentis. In some embodiments, the parent polynucleotide sequence is codon-optimized to enhance expression of the ATP-dependent amino acid ligase in a specified host cell. The parental polynucleotide sequence, designated as SEQ ID NO: 1, was codon optimized for E. coli expression to afford SEQ ID NO: 3. The codon optimized ATP-dependent amino acid ligase was cloned into an expression vector, placing the expression of the ATP-dependent amino acid ligase gene under the control of the lac promoter under control of the lac repressor. Clones expressing the active ATP-dependent amino acid ligase in E. coli were identified, and the genes sequenced to confirm their identity.

[0169] The ATP-dependent amino acid ligase of the disclosure may be obtained by subjecting the polynucleotide encoding the parent sequence to mutagenesis and / or directed evolution (DE) methods. Any DE technique may be used, including crystal structure-guided library design, modelling-guided library design, single-site-saturation mutagenesis (SSM) library generation, or combinatorial library generation, or a combination of these. An exemplary directed evolution technique is mutagenesis and / or DNA shuffling as described in Stemmer, 1994, Proc. Natl. Acad. Sci. USA 91 : 10747-10751; WO 95 / 22625; WO 97 / 20078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767 and U.S. Pat. No. 6.537.746. Other directed evolution procedures that can be used include, among others, staggered extension process (StEP), in vitro recombination (Zhao et al., 1998, Nat. Biotechnol. 16:258-261), mutagenic PCR (Caldwell et al., 1994, PCR METHODS APPL. 3:S136-S140), and cassette mutagenesis (Black et al., 1996, Proc. Natl. Acad. Sci. USA 93:3525-3529).

[0170] Exemplary crystal stmcture-guided library design techniques may make use of computational modeling and / or may be rational or semi-rational. Examples of such techniques include molecular dynamics (MD) simulations, including in silico MD, Rosetta de novo design (e.g., trRosetta), and Glide modeling. trRosetta is a deep neural net-based protein structure prediction method, offers a path forward for predicting 3D structures when there is a lack of close homologue structures, which takes advantage of multiple sequence alignment (MSA),where derived features from the MSA are fed into a deep neural network to predict inter-residue geometries, including distance and orientations.

[0171] The clones obtained following mutagenesis treatment are screened to identify ATP- dependent amino acid ligases exhibiting one or more desired improved enzyme properties. These desired improved enzyme properties may be selected by applying each property, one-by-one as a selection pressure during iterative rounds of evolution. For instance, improved enzyme activity may be applied as a selection pressure during one or more rounds of evolution. Additional selection pressures that may be applied include improved thermostability, reduced product inhibition, co-solvent tolerance, broadened pH tolerance, increased substrate concentration tolerance, increased phosphate concentration tolerance, and reduced formation of undesired byproducts. For instance, improved thermostability was applied as a selection pressure during evolution of the disclosed Trp-ligase enzymes, as described in the Examples. For instance, reduced formation of oligomers of compound 4f and oligomers of compound 5k was applied as a selection pressure during evolution of the disclosed Phe-ligase enzymes.

[0172] Measuring enzy me activity from the expression libraries can be performed using standard chemistry' analytical techniques for measuring substrates and products such as UPLC- MS. In this reaction, amino acid residues such as the ones described in Tables 1 and 2 can be ligated into tripeptides and tetrapeptides in the presence of an ATP-dependent amino acid ligases and stochiometric amounts of ATP. In some embodiments, an ATP regeneration system consisting of a polyphosphate kinase a phosphate donor such as polyphosphate and an adenosine cofactor such as AMP, ADP or ATP can be employed in ATP-dependent amino acid ligase reactions. The reaction may also be run under conditions where the ATP-dependent amino acid ligase is the yield-limiting catalyst, such that a doubling or halving in concentration of the ATP- dependent amino acid ligase will result in a doubling or halving of the yield of product observed at a given timepoint.

[0173] Where the improved enzy me property desired is thermostability, enzyme activity may be measured after subjecting the enzyme preparations to a defined temperature and measuring the amount of enzyme or enzyme activity remaining after heat treatments. Any suitable approach may be used, e.g., differential scanning calorimetry (DSC), a biochemical assay, or spectroscopy. Clones containing a polynucleotide encoding ATP-dependent amino acid ligases are then isolated, sequenced to identify the nucleotide sequence changes (if any), and used to express the enzyme in a host cell.

[0174] Where the sequence of the polypeptide is known, the polynucleotides encoding the enzy me can be prepared by standard solid-phase methods, according to known syntheticmethods. In some embodiments, fragments of up to about 100 bases can be individually synthesized, then joined (e.g., by enzymatic or chemical litigation methods, or polymerase mediated methods) to form any desired continuous sequence. For example, polynucleotides and oligonucleotides of the disclosure can be prepared by chemical synthesis using, e.g, the classical phosphoramidite method described by Beaucage et al., 1981, TET. LETT. 22: 1859-69, or the method described by Matthes et al., 1984, EMBO J. 3:801-05, e.g, as it is typically practiced in automated synthetic methods. According to the phosphoramidite method, oligonucleotides are synthesized, e.g., in an automatic DNA synthesizer, purified, annealed, ligated and cloned in appropriate vectors. In addition, essentially any nucleic acid can be obtained from any of a variety of commercial sources, such as Twist Bioscience, South San Francisco, Calif; Genscript, Piscataway, New Jersey; Azenta Life Sciences, South Plainfield, New Jersey; Biomatik, Wilmington, Delaware; Integrated DNA Technologies, Coralville, Iowa; and many others.

[0175] ATP-dependent amino acid ligase enzymes expressed in a host cell can be recovered and isolated from the cells and / or the culture medium using any one or more of the well-known techniques for protein purification, including, among others, lysozyme treatment, sonication, microfluidizer lysis, filtration, salting-out, ultra-centrifugation, and chromatography. Suitable solutions for lysing and the high efficiency extraction of proteins from bacteria, such as E. coli, are commercially available under the trade name CELLYTIC B® from Sigma-Aldrich.

[0176] Chromatographic techniques for isolation of the ATP-dependent amino acid ligase polypeptide include, among others, reverse phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. Conditions for purifying a particular enzyme will depend, in part, on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and will be apparent to those having skill in the art.

[0177] In some embodiments, affinity techniques may be used to isolate the improved ATP- dependent amino acid ligase enzymes. For affinity chromatography purification, the protein sequence can be tagged with a recognition sequence to enable purification. Common tags include cellulose-binding domains, poly His-tags (e.g., 6xHis tags), di-His chelates, FLAG-tags and many others that will be apparent to those having skill in the art. Antibodies can also be used as affinity purification reagents. Any antibody that specifically binds the ATP-dependent ligase polypeptide may be used.Methods of Generating Oligopeptides

[0178] Provided herein are methods (or processes) of generating oligopeptides via one or more ATP-dependent amino acid ligase-catalyzed reactions. These methods comprise performing a reaction with any of the disclosed engineered polypeptides. These methods may comprise the conversion of a dipeptide into a tripeptide (through a ligation reaction). These methods may comprise the conversion of a tripeptide into a tetrapeptide (through a ligation reaction). In various aspects, the ultimate product of these methods is a tetrapeptide.

[0179] Any of these methods may be performed in a medium containing cofactors ATP, magnesium ion (Mg2+), and / or sodium polyphosphate. In some embodiments, sub-stoichiometric amounts of ADP and / or ATP is used. Any of these methods may be performed in a medium containing sodium polyphosphate and a polyphosphate kinase (PPK). In some embodiments, wild-type PPK22 or PPK12 is employed. In some embodiments, both PPK22 and PPK12 are employed. In some embodiments, the medium contains co-solvent acetonitrile, e.g., 15%, 20%, 25%. or 30% acetonitrile. In some embodiments, the medium is maintained at a neutral pH, e.g.. a pH of 7.4.

[0180] In some aspects, provided are methods for preparing Compound 3d, the method comprising contacting Compound 2c with Compound Id in the presence of ATP, Mg2, polyphosphate, a polyphosphate kinase, and any of the disclosed polypeptides. These methods may produce oligomers of Compound la, Compound lb, Compound 1c, or Compound Id, or in the context of a cascade reaction with a Phe-ligase and a compound of Formula IV, may produce a product of ligation of Compound 5k to Compound Id in reduced amounts relative to a corresponding method in which Compound 2c and Compound Id are contacted with the polypeptide of SEQ ID NO: 4.

[0181] In some aspects, provided are methods for preparing Compound 5k, the method comprising contacting Compound 3d with Compound 4f, in the presence of ATP, Mg2+, polyphosphate, a polyphosphate kinase, and any of the disclosed polypeptides. In some embodiments, these methods produce oligomers of Compound 4e or Compound 4f in reduced amounts relative to a corresponding method in which Compound 3d and Compound 4f are contacted with the polypeptide SEQ ID NO: 4.

[0182] In some embodiments, the disclosed methods exhibit higher (1) phosphate tolerance, (2) substrate loading capacity. (3) thermostability, and / or (4) pH tolerance, relative to a corresponding method in which Compound 2c and Compound Id are contacted, or in which Compound 3d and Compound 4f are contacted, with the polypeptide of SEQ ID NO: 4.

[0183] Any of these methods may be performed in a medium containing cofactors ATP, magnesium ion (Mg2+), and / or propionyl phosphate. In some embodiments, these methods are performed in a medium in the presence of ADP or ATP in the presence of propionyl phosphate and an acetate kinase. In some embodiments, wild-type or commercially available acetate kinase is employed.

[0184] In some embodiments, the disclosed methods are performed in the presence of an ATP regeneration system, e.g., sources of activated phosphates. This ATP regeneration system may comprise, or consist of, a polyphosphate kinase, a phosphate donor such as polyphosphate, and an adenosine cofactor such as AMP. ADP, or ATP.

[0185] In some embodiments, the steps of the disclosed methods may each be performed in isolation of the others, with purification of each intermediate between steps. In other embodiments, the enzy mes and substrates can all be combined into one vessel (one-pot synthesis) and allowed to react simultaneously. In other embodiments, above-disclosed dipeptide and / or tripeptide intermediates can be synthesized chemically, and subsequently reacted with appropriate Phe-ligase and / or Trp-ligase enzymes in cascades to yield tripeptide or tetrapeptide products. In some aspects, the disclosed method steps may be run simultaneously in cascades such that isolations and / or purifications of intermediates is eliminated.

[0186] Whether carrying out the method with whole cells, cell extracts or purified ATP- dependent amino acid ligase enzymes, a single ATP-dependent amino acid ligase enzyme may be used or, alternatively, mixtures of two or more ATP-dependent amino acid ligase enzymes may be used.EXAMPLESABBREVIATIONSExample 1: Construct optimization and gene synthesis

[0187] A DNA sequence encoding wild-type ATP-dependent ligase polypeptide from Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1, SEQ ID NO: 2) identified from a literature report (See Arai et al.. Biosci. Biotechnol. Biochem. 2010 for reference) was codon optimized for A. coli expression and chemically synthesized to yield reference nucleic acid sequence SEQ ID NO: 3. An oligonucleotide sequence encoding an N-terminal hexahistidine tag (SEQ ID NO: 117) was included as part of the sequence. The gene was cloned into a pET30a(+) E. coli expression vector containing a T7 promoter, a ColEl origin of replication, and a KanR selection marker, sequence verified and transformed into BL21(DE3) E. coli cells. Likewise, genes encoding engineered ATP-dependent ligases were cloned into a pET30a vector and transformed into BL21(DE3) cells.Example 2; High throughput growth, expression, and enzyme-containing lysate preparation for 96-well-plate reactions

[0188] The transformants from Example 1 were selected on LB-agar Petri plates supplemented with 30 pg / mL kanamycin and 1% w / v glucose. The transformants expressing SEQ ID NO: 3 as ATP-dependent ligase (SEQ ID No: 4) were picked and grow n in Luria-Bertani Broth medium with kanamycin (30 pg / rnL) and 1% w / v glucose in 96-well plates with shaking (200 rpm) at 30 °C overnight. The following day, the optical density at 600 nm (ODeoo) of the saturated culture was measured. Cell cultures were then diluted to an ODeoo of 0.05 in -390 pL of ZYM-5052 autoinduction media supplemented with 30 pg / mL of kanamycin in a 96-deep-well plate. Proteinexpression proceeded for 20 h at 30 °C and 250 rpm shaking. Cells were pelleted by centrifugation at 4000 x g for 15 min, and the supernatant was discarded. Pelleted cells were frozen / thawed once and then lysed in 0.2 mL / well of 50 mM HEPES pH 7.5 buffer, 1 mg / mL lysozyme, 0.5 mg / mL polymyxin B sulfate, 1 mM MgCh and 1 U / mL DNasel endonuclease for 2 h at room temperature with shaking at 1000 rpm. Cell lysate was clarified by centrifugation at 4000 x g for 15 minutes, and the clarified lysates (supernatants) were used in subsequent wellplate reactions.Example 3: Enzyme Preparation for Vial and Larger-scale reactions

[0189] Twenty microliters of a glycerol stock of BL21(DE3) containing an ATP -dependent ligase gene of interest in pET30a(+) was inoculated in 25 mL of LB broth supplemented with 30 pg / mL kanamycin and 1% (w / v) glucose in a 125 mL baffled flask. Cell cultures were grown for 18 h at 30 °C and 250 rpm. The following day, 1 L of TB media supplemented with 30 pg / mL kanamycin was subcultured with the overnight saturated culture diluted to an initial OD of 0.05. Cells were grown at 30 °C with shaking at 250 rpm until ODeoo reached 0.6-0.8. Protein production was initiated with the addition of 1 mM IPTG and cells w ere cultured for an additional 20 h at 30 °C. Biomass was pelleted by centrifugation at 4000 x g for 15 min, and the spent media was discarded. Cell pellets were frozen / thawed once and resuspended in 50 mM triethanolamine-HCl pH 7.5 (5 mL per gram of pellet). Cell suspensions were shaken at 18 °C for 30 minutes, after which cells were disrupted by high-pressure homogenization (16,000 PSI). The resulting lysate was clarified by centrifugation at 10,000 x g for 45 minutes at 4°C. The clarified lysate was frozen and lyophilized. The lyophilized powder (shake flask powder, SFP) was used in subsequent vial (small-scale) and larger scale reactions as detailed in the following Examples.Example 4: Engineering of polypeptides derived from SEQ ID NO: 4 for formation of tripeptide product of Reaction A

[0190] ATP-dependent ligase enzymes derived from SEQ ID NO: 4 were engineered via directed evolution from the polynucleotide having the nucleotide sequence of SEQ ID NO: 3 against a selection pressure for activity (conversion of substrate to product 3a). Libraries of engineered polypeptides were generated by homology model-guided semi-rational directed evolution using site saturation mutagenesis and combinatorial libraries.Directed evolution strategy'

[0191] In the directed evolution campaign, single-site-saturation mutagenesis (SSM) libraries were built and screened for increased conversion relative to the starting enzyme. SSM libraries ineach round of evolution were generated using gene splicing by overlap extension (SOEing) PCR methods described in Ho et al.. Gene 1989, 77(1), 51-59. In short, mutations at the designated positions were incorporated by degenerate oligos through overlapping PCRs. The full-length gene-of-interest region was further amplified and assembled into the expression vector (pET30a) by Gibson Assembly (Gibson et al. Nat. Methods 2009. 6(5), 343-345). The combinatorial libraries were built following the instructions from the QUIKCHANGE® Lightning Multi Site- Directed Mutagenesis kit (Agilent Technologies). Both mutagenesis libraries were transformed into the BL21(DE3) E.coli strains by electroporation and plated on LB agar plates with selection (1% w / v glucose, 30 pg / mL kanamycin).

[0192] To prioritize the sites for SSM library designs, a structural homology model of wild- tj pe BAD1200 was built using the Protein Data Bank (PDB) template 4WD3, which had 23% sequence similarity'. Schrodinger toolbox was used to design the homology model, which was further refined by adding hydrogens using PROPKA (Olsson et al. J. Chem. Theory Comput. 2011, 7(2). 525-537) at pH 10 and running a restrained minimization to converge heavy atoms to a maximum root mean square deviation of 0.30A. Because of the low sequence similarity, the model was improved using trRosetta. trRosetta’s top 5 predicted structures based on Rosetta score in REU were visually inspected and loops consisting of positions 123-234 were low confidence. trRosetta models were energy minimized using the FastRelax algorithm. The magnesium ions and the ATP binding sites are conserved within the family and were placed according to the homology model (template PDB ID: 42d3). Substrate docking and product docking was performed using Glide (Friesner et al. J. Med. Chem. 2004, 47(7), 1739-1749), and the best pose was selected after visually inspecting a pool of docked poses rank ordered by their docking energy scores and filtering them based on the distances between the reaction site and the catalytic residues.

[0193] The model was partitioned into five regions: active site, surfacel, surface2, core, and “leftover.” The active site shell residues were selected using a distance criterion, i.e., all residues within 15. 1 A of the docked substrate, resulting in 96 sites. The remaining residues were rank ordered by their solvent accessibilities and grouped into sets of 96 sites. The top two sets were classified as surface sitel (positions with solvent accessible surface area > 73 A) and site2, (positions with solvent accessible surface area <=73 A and > 17 A) consisting of higher solvent accessibility sites, whereas the remaining were classified as the core. A fifth region was identified as “leftover” sites consisting of residues not covered in any of the other three library designs to have a complete coverage of the protein for SSM. Iterations of targeting active site, surface, core and leftover residues were employed across rounds of protein engineering.Beneficial mutations identified from these SSM libraries were recombined into combinatorial libraries for subsequent screenings. The best variant served as the backbone for the next round of mutagenesis.HTP activity assay for ATP-dependent ligase variants:

[0194] ATP-dependent ligase enzymes derived from SEQ ID NO: 4 (affinity tagged-BAD1200) were screened in a 96-well microtiter plate format using the following conditions: In a final volume of 200 pL / well 10 mM substrate la, 40 mM substrate 2a, 100 mM Tris-HCl pH 8, 12.5 mM MgCh, 1 mM ATP, 25 mM propionyl phosphate, 0.3 g / L acetate kinase were incubated in the presence of 10 % (v / v) enzyme lysates derived from SEQ ID NO: 4. Reactions were incubated at 30°C for 20 h shaking at 600 rpm. The following day, reactions were quenched by adding 100 pL of reaction to 100 pL of ACN. Quenched reactions were filtered and formation of product 3a was analyzed by LC-MS.HTP analytical methods for enzyme activity for formation of tripeptide product of Reaction A:

[0195] For each of the Examples below, except where otherwise indicated, LC-MS analysis was conducted using the following method:_

[0196] For the particular LC-MS analysis of this Example, the following mobile phase, flow rate and gradient were used:Mobile Phase: MP A: H2O (0.1% v / v TFA), MP B: ACN (0.1% v / v TFA)Flow rate: 0.75 mL / minInj ection volume: 1 pLGradient program:Retention time 3a: 0.441 min (UV @220 nm)MS Parameters:Drying gas flow(V / min) 12Nebulizer pressure(psi) 35 Drying gas Temp (C) 350 Monitored (3a) at MS-SIM Pos m / z 403.20

[0197] Library variants catalyzing the highest conversion to product 3a were expressed at shake flask scale as described in Example 3 and screened under similar conditions as described above for high throughput, with the exception that 2 g / L of shake flask powder was used to determine formation of product 3a. Engineered ATP-dependent ligase polypeptides having the amino acid sequences of SEQ ID NOs: 6, 8, 10 and 12 were identified from several rounds of evolution that identified variants (or "‘hits”) exhibiting improvements in enzymatic activity in substrates of Reaction A (see Table 1). Specifically, from a 96-position site saturation mutagenesis (SSM) library, SEQ ID NO: 5 / 6, with S92K relative to SEQ ID NO: 3 / 4, demonstrated the highest production of 3a.

[0198] In a subsequent round of evolution starting from SEQ ID NO: 5 / 6, three separate 96- position SSMs targeting active site, surface, and core as well as a combinatorial library drawing from diversity observed in homologous sequences were constructed, expressed, and screened for activity, resulting in the identification of SEQ ID NO: 7 / 8, encoding V374A relative to SEQ ID NO: 5 / 6, as catalyzing the highest conversion to 3a. In the following round of evolution starting from SEQ ID NO: 7 / 8. a combinatorial library sampling beneficial diversity from the previous round was constructed, expressed, and screened for activity, resulting in the identification of SEQ ID NO: 9 / 10, with K19H, C37P, V325H, and H371L relative to SEQ ID NO: 7 / 8. Two 96- position SSM libraries were then built, expressed, and screened from SEQ ID NO: 9 / 10 targeting the active site and core regions, and SEQ ID NO: 11 / 12, S68L relative to SEQ ID NO: 9 / 10, was identified has having increased conversion toward 3a and the ability to catalyze Reaction B (Table 1). The engineered ATP-dependent ligase polypeptide having the amino acid sequence of SEQ ID NO: 12 also exhibited improved enzy matic activity' on the substrates of Reaction B and thus was selected for further evolution.

[0199] In this Example and Examples 5 and 6. as part of the enzyme engineering strategy’ to develop a Trp-ligase (Trajectory 1), a substrate walk was performed from substrates la and 2a, to substrates lb and 2b, substrates 1c and 2c, and finally substrates Id and 2d (see Table 1). A “substrate walk” refers to an approach taken to engineer an enzy me for activity on a desired substrate when that particular enzy me exhibits little to no activity for such substrate by engineering both the substrate and enzyme in parallel. As an example, starting from a substrate for which some enzyme activity is observed (e.g., a substrate more closely resembling the enzy me’s native or natural substrate), bulkier substituents are added to the substrate along withmutations in the enzyme active site until variants with improved activity on the bulkier substrate are found. The process is repeated until activity is obtained on the desired substrate. The “substrate walk” technique is not limited to the use of bulkier substituents (other examples include changing electronics around the substrate nucleophiles or electrophiles).Example 5; Engineering of Trp-ligase polypeptides derived from SEO ID NO: 12 for formation of tripeptide product of Reaction B (Formula 3b)

[0200] ATP-dependent ligase enzymes derived from SEQ ID NO: 12 were engineered via directed evolution from the polynucleotide SEQ ID NO: 11. In this round of engineering, three libraries were constructed, expressed, and screened: a combinatorial library sampling beneficial diversity identified in previous rounds, a 96-position SSM targeting the core region, and a combinatorial library sampling natural diversity from homologs. ATP-dependent ligase enzy mes derived from SEQ ID NO: 12 were screened in a 96- well-plate format using the following conditions: in a final volume of 200 pL / well, 10 mM substrate lb, 40 rnM substrate 2b, 100 mM Tris-HCl pH 8, 12.5 mM MgCE, 1 mM ATP, 25 mM propionyl phosphate, and 0.3 g / L acetate kinase were incubated in the presence of 5 % (v / v) clarified enzyme lysates. Reactions were incubated at 30 °C for 20 h with shaking at 600 rpm. The following day, reactions were quenched by adding 100 pL of reaction to 100 pL of ACN. Quenched reactions were filtered and formation of product 3b was analyzed by LC-MS using the column and conditions described in Example 4 with the following mobile phase, flow rate and gradient:Mobile Phase: MP A: H2O (10 mM ammonium formate pH 8.5), MP B: ACNFlow rate: 0.6 mL / minGradient program: |_Retention time 3b: 0.856 min @UV 220 nmMS Parameters:Drying gas flow(V / min) 12Nebulizer pressure(psi) 35 Dry ing gas Temp (C) 350 Monitored (3b) at MS-SIM Pos m / z 423.20

[0201] Library variants catalyzing the highest conversion to product 3b were expressed at shake flask scale as described in Example 3 and screened under similar conditions as described above for high throughput, with the exception that 2 g / L of shake flask powder was used to determine formation of product 3b. The engineered ATP-dependent ligase polypeptide having the amino acid sequence of SEQ ID NO: 14, having F21E, S22A, E236L, Y280L, S400Q, and Q401A mutations relative to SEQ ID NO: 12, exhibited improved enzy matic activity7(i.e., conversion) with Reaction B. The enzy me polypeptides having the amino acid sequences of SEQ ID NO: 12 and 14 also exhibited improved enzymatic activity in Reaction C. Polypeptide of SEQ ID NO: 12 exhibited better enzymatic activity in Reaction D and was therefore selected for further evolution.Example 6; Engineering of Trp-ligase polypeptides derived from SEQ ID NO: 12 for formation of tripeptide product of Reaction C

[0202] ATP-dependent ligase enzymes derived from SEQ ID NO: 12 were engineered via directed evolution from the polynucleotide SEQ ID NO: 11. In this round of engineering, four libraries were constructed, expressed, and screened: two combinatorial library7sampling beneficial diversity identified in previous rounds, two 96-position SSMs targeting tier 1 (active site) and tier 2 (surface) residues. ATP-dependent ligase enzymes derived from SEQ ID NO: 12 were screened in a 96-well-plate format using the following conditions: in a final volume of 100 pL / well, 10 mM substrate 1c, 25 mM substrate 2c, 100 mM HEPES pH 8, 10 % (v / v) DMSO, 25 mM MgCh, 20 mM ATP were incubated in the presence of 7.5 % (v / v) clarified enzyme lysates. Reactions were incubated at 30 °C for 20 h with shaking at 600 rpm. The following day. reactions were quenched by adding 100 pL of reaction to 100 pL of ACN. Quenched reactions were filtered and formation of product 3c was analyzed by LC-MS using the column and conditions described in Example 4 with the following mobile phase, flow rate and gradient:Mobile phase: MP A: Water (0.1% difluoroacetic acid), MP B: ACN (0.1% difluoroacetic acid)Flow rate: 0.6 mL / min, BEH C18 Column 2.1X50 mmGradient program:Retention time 3c: 0.950 min@ UV 220 nmMS Parameters:Dry ing gas flow(V / min) 12Nebulizer pressure(psi) 35Dry ing gas Temp (C) 350Monitored (3c) at MS-SIM Pos-693.30

[0203] Library variants catalyzing the highest conversion to product 3c were expressed at shake flask scale and assayed under similar conditions as described above for high throughput, with the exception that 2 g / L of shake flask powder was used to determine formation of product 3c. The engineered ATP-dependent ligase polypeptide having the amino acid sequence of SEQ ID NO: 16, having A49E, Al 66V, I354Y, and L360D mutations relative to SEQ ID NO: 12, exhibited improved enzy matic activity' on (i.e., conversion of) substrates of Reactions C and D to tripeptide products and thus SEQ ID NO: 16 and its encoding polynucleotide SEQ ID NO: 15 were selected for further engineering.Example 7; Engineering of Trp-ligase polypeptides derived from SEQ ID NO: 16 for formation of tripeptide product of Reaction D

[0204] ATP-dependent ligase enzymes derived from SEQ ID NO: 16 were engineered via directed evolution from the polynucleotide SEQ ID NO: 15. In this round of engineering, two libraries were constructed, expressed, and screened: a combinatorial library' sampling beneficial diversity identified in previous rounds, a 96-position SSMs targeting tier 3 core residues. ATP- dependent ligase enzymes derived from SEQ ID NO: 16 were screened in a 96-well-plate format using the following conditions: in a final volume of 50 pL / well, 10 mM substrate Id, 25 mM substrate 2c, 100 mM HEPES pH 8, 5 % (v / v) DMSO, 25 mM MgCh, 20 mM ATP were incubated in the presence of 10 % (v / v) clarified enzy me lysates. Reactions were incubated at 30 °C for 20 h with shaking at 1000 rpm. The following day, reactions were quenched by adding100 pL of ACN into 1000 pL of reaction. Quenched reactions were filtered and formation of product 3d was analyzed by LC-MS using the column and conditions described in Example 4 with the following mobile phase, flow rate and gradient:Mobile phase: MP A: Water (0.1% difluoroacetic acid). MP B: ACN (0.1% difluoroacetic acid)Flow rate: 0.6 mL / minGradient program:Retention time 3d: 1.25 min@UV 220 nmS Parameters:Drying gas flow(VZmin) 12 Nebulizer pressure(psi) 35 Dry ing gas Temp (C) 350 Monitored (3d) at MS-SIM Pos-691.84

[0205] Library variants catalyzing the highest conversion to product 3d were expressed at shake flask scale and assayed under similar conditions as described above for high throughput, with the exception that 2 g / L of shake flask powder was used to determine formation of product 3d. The engineered ATP-dependent ligase polypeptide having the amino acid sequence of SEQ ID NO: 18. having K92L, F235S, N239Y. T240H, D360E, and K379M mutations relative to SEQ ID NO: 16, exhibited improved enzymatic activity on (i.e., conversion ol) Reaction D to form the tripeptide product and thus SEQ ID NO: 18 and its encoding polynucleotide SEQ ID NO: 17 were selected for further engineering.Example 8; Engineering of Trp-ligase polypeptides derived from SEQ ID NO: 18, 20, 22, 26, 28, 32, 34, 38, and 40 for formation of tripentide product of Reaction D

[0206] SEQ ID NO: 18 was further engineered by directed evolution to identify hits that exhibited 1) improved activity, 2) substrate load, 3) phosphate tolerance, 4) co-solvent tolerance, and / or 5) reduced by-product formation. The best hit (i.e., engineered variant) from each round of evolution was used as the backbone in the following round of evolution, as described below.Assay conditions and evolution pressures are provided in Table 6 as follows:Table 6: Assay conditions and evolution pressures

[0207] The ligase having the sequence of SEQ ID NO: 40, indicated in the final row of the Table 6 as emerging from a final round of evolution, contains the following 47 substitutions relative to the reference sequence of SEQ ID NO 4: L15M, Q17H, K19Q, F21L, C37P, N46S, A49H, S68G, S91G, S92L, L111V, A166V. G170S, A182E, S205N, L229M, V230R, F235R, K237I, N239H, T240H, N248G. A264G, F266L, L300I, A307L, V325H, A327E, L341I, E342M, A344E, I354Y, L360E, V363D, K366T, G368E, A370N, H371S, S373G, V374A, D375W, S376W, K379M, S388E, S394Q, N395R, and Q401W.

[0208] SEQ ID NOs: 40, 42, 44, 46, 48, and 50 emerged from 21 rounds of engineering via directed evolution as the best-performing ATP-dependent ligase variants. The engineered polypeptides of SEQ ID NOs: 42, 44, 46, 48, and 50 constitute variants evolved from SEQ ID NO: 38 in a final round of evolution for which activity and by-product reduction selection pressures were applied. To engineer phosphate tolerance, 100 mM KPi was added to the screen for variants of SEQ ID NO: 19 / 20. To engineer co-solvent tolerance and selectivity against substrate 4f, these components were added to the screen for variants of SEQ ID NO: 21 / 22. To engineer specificity against pentapeptide formation from ligation of 4f and 5k, substrate 5k was added to the screen of variants of SEQ ID NO: 33 / 34. These engineered polypeptides show improved (1) activity for the formation of tripeptide product Compound 3d; (2) regioselectivity' for Compound 3d over dimers, trimers and other oligomers of compounds according to Formula I (i.e., oligomers of Compound la, Compound lb, Compound 1c, or Compound Id) and over aproduct of ligation of compound 5k to Compound Id in a tetrapeptide-forming cascade context;(3) phosphate tolerance; (4) substrate loading tolerance; (5) thermostability; and / or (6) byproduct formation reduction relative to BAD 1200 (SEQ ID NO: 4), such as dipeptide byproducts formed in a cascade context from a compound of Formula I and a compound of Formula IV. or a tetrapeptide by-product formed from a dimerization of a compound of Formula I ligated to a dipeptide of Formula II. For example, the Trp-ligase of SEQ ID NO: 40 exhibited about 23,000-fold activity level improvement relative to SEQ ID NO: 4, generating compound 3d with 100% regioselectivity and by-products at a level of less than 0.5% of all products.Example 9: Engineering of Phe-ligase polypeptides derived from SEO ID NO: 4 for formation of tetrapeptide product of Reaction E

[0209] ATP-dependent ligase enzymes derived from SEQ ID NO: 4 were evolved from the polynucleotide of SEQ ID NO: 3. In the second trajectory, to engineer Phe-ligase variants, a substrate walk was performed among the substrates of Formula III and Formula IV from substrates 3a and 4e, to substrates 3a and 4f, to substrate 3e and 4f, to substrates 3b and 4f, to substrates 3f and 4f, to substrate 3c and 4f, and finally to substrate 3d and 4f (see Table 2).

[0210] Libraries of engineered polynucleotides based on SEQ ID NO: 3 were generated using site saturation mutagenesis of active site, core, and surface residues and expressed as described in Example 2. ATP-dependent Phe-ligase enzymes derived from SEQ ID NO: 4 were screened in a 96-well-plate format to identity hits for improved activity in catalyzing the formation of tetrapeptide compound 5e using the following conditions: In a final volume of 150 pL / well 10 mM substrate 4e, 10 mM substrate 3a. 100 mM Tris-HCl pH 8. 12.5 mM MgCh, 1 mM ATP, 25 mM propionyl phosphate, and 0.3 mg / mL acetate kinase were incubated in the presence of 10% (v / v) enzyme lysate. Reactions were incubated at 30 °C for 20 h with shaking at 1000 rpm. The following day, reactions were quenched by transferring 20 pL of the reaction into 180 pL 80% v / v ACN. Quenched reactions were filtered through a 0.2 pm filter, and formation of product 5e was analyzed by LC-MS using the following mobile phase, flow rate and gradient:Mobile phase: MP A: Water (0.1% trifluoroacetic acid), MP B: ACN (0.1% trifluoroacetic acid) Flow Rate: 0.75 mL / minRetention time 5e: 0.816 min

[0211] Variants catalyzing the highest conversion to product 5e were further scaled up to shake flasks and screened under similar conditions as shown above with the exception that 2 g / L of polypeptide shake flask powder was used to determine formation of product 5e. The polypeptide having the amino acid sequence of SEQ ID NO: 52 was identified as a variant exhibiting significant improvements in enzymatic activity for reactions E. F and G and thus used as a backbone for further evolution.Example 10: Engineering of Phe-ligase polypeptides derived from SEQ ID NO; 52 for formation of tetrapeptide of Reaction F

[0212] In the next round of engineering, a library was built from SEQ ID NO: 51 that combinatorically sampled beneficial diversity identified from the Example 9 SSM libraries. This library was built, expressed, and screened in a 96-well-plate format to identify hits for improved activity for the formation of product 5f using the following conditions: In a final volume of 200 pL / well 10 mM substrate 4f. 10 mM substrate 3a. 100 mM Tris-HCl pH 8. 12.5 mM MgCh. 1 rnM ATP, 25 mM propionyl phosphate, 0.3 mg / rnL acetate kinase were incubated in the presence of 5 % (v / v) enzyme lysate. Reactions were incubated at 30 °C for 20 h shaking at 1000 rprn. The following day, reactions were quenched by transferring 10 pL of reaction to 190 pL 80% v / v ACN. Quenched reactions were filtered and formation of product 5f analyzed by LC-MS using the following mobile phase, flow rate and gradient:Mobile phase: MP A: Water (0.1% trifluoroacetic acid), MP B: ACN (0.1% trifluoroacetic acid)Flow Rate: 0.75 mL / minRetention time 5f: 1.701 min

[0213] Variants catalyzing the highest conversion to product 5f were further scaled to shake flasks and screened under the following conditions: 50 mM 4f, 50 mM 3a, 1 mM ATP, 60 mM propionyl phosphate, 25 mM MgCh, 0.3 mg / mL acetate kinase, 100 mM Tris-HCl pH 8.0, 2 mg / mL ligase shake flask powder. Reactions were shaken at 25 °C for 20 h. After overnight incubation, reactions were quenched with 500 pL ACN, filtered, and formation of product 5f analyzed by LC.

[0214] The polypeptide having the amino acid sequence of SEQ ID NO: 58 was identified as a variant having improved enzymatic activity for product formation on Reactions F and G at higher substrate loads and was thus selected as the backbone for further evolution.Example 11: Engineering of Phe-ligase polypeptides derived from SEQ ID NO: 58 for formation of tetrapeptide of Reaction G

[0215] Using SEQ ID NO: 57 as the template, a library combinatorically sampling previously identified beneficial diversity and a 96-position SSM library targeting the active site region were constructed, expressed, and screened. These ATP-dependent Phe-ligase variants of SEQ ID NO: 58 were screened to identify hits improved for activity under the following conditions: In 96- well-plate and a final volume of 50 pL / well 40 mM substrate 4f, 40 mM substrate 3e, 100 mM Tris-HCl pH 8, 25 mM MgCb, 1 mM ADP, 40 mM polyphosphate, 0.1 mg / mL PPK.12 were incubated in the presence of 1.25% (v / v) enzyme lysate. Reactions were incubated at 30 °C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 20 pL of reaction to 180 pL 80% v / v ACN. Quenched reactions were filtered and formation of product 5g analyzed by LC-MS using the method described in Example 10, except that the Product 5g retention time was 1.639 min.

[0216] Variants catalyzing the highest conversion to product 5g were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 pL, 50 mM 4f, 50 mM 3e, 5 mM ATP, 60 mM polyphosphate. 25 mM MgCh, 0. 1 mg / mL PPK.12, 100 mM Tris- HCl pH 8.0 were mixed with 2 mg / mL shake flask powder of ATP-dependent ligase derived from SEQ ID NO: 58. Reactions were incubated at 25 °C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quench with 500 pL ACN, filtered, and analyzed by LC.The polypeptide having the amino acid sequence of SEQ ID NO: 60 was identified as a variant exhibiting improvements in enzy matic activity in Reactions G, H and I and thus selected for further evolution.Example 12: Engineering of Phe-ligase polypeptides derived from SEO ID NO: 60 for formation of tetrapeptide of Reaction H

[0217] Using SEQ ID NO: 59 as the template, two 96-position SSM libraries targeting the core and surface regions were constructed, expressed, and screened. These ATP-dependent Phe- ligase variants of SEQ ID NO: 60 were screened to identify hits improved for activity under the following conditions: In 96-well-plate and a final volume of 50 pL / well 40 mM substrate 4f, 40 mM substrate 3b, 100 mM Tris-HCl pH 8, 25 mM MgCh, 1 mM ADP, 40 mM polyphosphate, 0.1 mg / mL PPK12 were incubated in the presence of 2% (v / v) enzyme lysate. Reactions were incubated at 30 °C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 20 pL of reaction to 180 pL 50% ACN. Quenched reactions were filtered and formation of product 5h analyzed by LC-MS using the following mobile phase, flow rate and gradient:Mobile phase: MP A: Water (0.1% formic acid), MP B: ACN (0.1% formic acid) Flow Rate: 0.75 mL / minRetention time 5h: 1 .022 min

[0218] Variants catalyzing the highest conversion to product 5h were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 pL. 50 mM 4f, 50 mM 3b, 5 mM ATP, 60 mM polyphosphate, 25 mM MgCb, 0.1 mg / mL PPK12, 100 mM Tris- HCl pH 8.0, were mixed with 2 mg / mL shake flask of enzy mes derived from SEQ ID NO: 60. Reactions were incubated at 25 °C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quenched with 500 pL ACN. filtered, and analyzed by LC.

[0219] The polypeptide of amino acid sequence SEQ ID NO: 62 was identified as a variant having improvements in enzymatic activity for Reaction I, thus selected for further evolution.Example 13: Engineering of Phe-ligase polypeptides derived from SEQ ID NO; 62 for formation of tetrapeptide of Reaction I

[0220] Using SEQ ID NO: 61 as the template, a library combinatorically sampling beneficial diversity from previous rounds and a 96-position SSM library targeting the active site region were constructed, expressed, and screened. These ATP -dependent Phe-ligase variants of SEQ ID NO: 62 were screened to identify hits improved for activity under the following conditions: In 96-well-plate and a final volume of 50 pL / well 20 mM substrate 4f, 20 mM substrate 3f, 100 mM Tris-HCl pH 8. 25 mM MgCh, 1 m ADP, 40 mM polyphosphate, 0.1 mg / mL PPK.12 were incubated in the presence of 2.5% (v / v) enzyme lysate. Reactions were incubated at 30 °C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 20 pL of reaction to 180 pL 50 % ACN. Quenched reactions were filtered and formation of product 5i analyzed by LC-MS using the following mobile phase, flow rate and gradient:Mobile phase: Water (0.1% formic acid), B: ACN (0.1% formic acid) Flow Rate: 0.75 mL / minRetention time 5i: 0.714 min

[0221] Variants catalyzing the highest conversion to product 5i were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 pL, 50 mM 4f, 50 mM 3f. 5 mM ATP, 60 mM polyphosphate, 25 mM MgCb. 0.1 mg / mL PPK12, 100 mM Tris- HCl pH 8.0, were mixed with 2 mg / mL shake flask of enzymes derived from SEQ ID NO 62. Reactions were incubated at 25°C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quenched with 500 pL MeCN, filtered, and analyzed by LC.

[0222] The polypeptide having amino acid sequence SEQ ID NO: 66 was identified as a variant having improvements in enzymatic activity for Reactions I and J. Upon further assessment of the LC-MS chromatogram of the ligase having the amino acid sequence of SEQ ID NO: 66, it was observed that this enzy me forms two distinct regioisomers depicted in the scheme of FIG. 2 whensubstrates from Reaction J are present (compounds 5j and 5j-regio). SEQ ID NO: 66 was selected for further evolution.Example 14: Engineering of Phe-ligase polypeptides derived from SEO ID NO; 66 for formation of tetrapeptide of Reaction J

[0223] When engineered Phe-ligase enzymes were assayed with reaction J (substrates 3c and 41), the undesired product 5j-regio was observed as the major product. Additionally, no activity was observed for reaction K (substrates 3d and 41). To improve regios declivity by accommodating the noncanonical components of substrates 3c and 3d, 5j and 5j-regio products were modeled in the Phe-ligase active site followed by MD simulations to identify residues that: i) participate in strong interactions with undesired product, ii) are not critical for desired product binding, iii) are not directly interacting with the products, but are involved in the enzy me function. To capture the effects of allostery, residues far from the active site that either i) underwent significant conformational changes in the 5j or 5j-regio docking compared to compound 5k simulations, or ii) had become rigid or flexible (based on RMSF measure) compared to 5k simulations were identified. S372 and V374 were discovered to undergo significant conformational changes when comparing the desired and undesired regioisomer simulations. Based on this analysis, an SSM library the residues described above was designed, constructed, expressed, and screened for activity on reaction J.

[0224] Using SEQ ID NO: 65 as the template, a library combinatorically sampling beneficial diversity from previous rounds and a 96-position SSM library targeting the active site region were constructed, expressed, and screened. These ATP-dependent Phe-ligase variants of SEQ ID NO: 66 were screened to identify hits improved for f) activ ity and 2) reduction in undesired regioisomer (5j-regio) formation under the following conditions: In 96-well-plate and a final volume of 50 pL / well 10 mM substrate 4f, 10 mM substrate 3c, 100 mM Tris-HCl pH 8, 25 mM MgCh, 2 m ADP, 40 mM polyphosphate, 0. 1 mg / mL PPK12 w ere incubated in the presence of 20 % (v / v) enzyme lysate. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 25 pL of reaction to 125 pL 80 % ACN. Quenched reactions were filtered and formation of product 5j w as analyzed by LC-MS using the following method:Instrument: Agilent 1260 equipped with an Agilent 6130 Quadrupole MS Column: Acquity UPLC HSS Cyano 1.8 pm, 2.1 x 50 mm Column Temp: 45°C Flow Rate: 0.65 mL / minInjection volume: 1 pLDetector: UV at 210 nm and MS (for low conversion variants)Gradient: A: 5 mM ammonium carbonate in water, B: ACNRetention time 5j: 0.910 min (desired regioisomer), 5j-regio: 1.137 (undesired regioisomer)

[0225] A cluster of mutations between positions 364-376 provided the best improvements in selectivity ranging from 20-25-fold improvement compared to SEQ ID NO: 66, while maintaining modest improvements in activity (~1.4x). These mutations sit in a flexible loop and sen e to restrict enzyme motion such that the enzyme cannot accommodate the undesired 5j- regio. Mutations A19R, R17W, and H166W provided the largest improvements in activity (2- fold, 2.5-fold and 6-fold respectively). These mutations rearrange the binding site of the enzyme resulting in weaker binding of substrates positioned to form 5j-regio and more favorable binding of substrates positioned to form 5j.

[0226] Variants catalyzing the highest conversion to product 5j were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 pL. 50 mM 4f, 50 mM 3c, 5 mM ATP, 60 mM polyphosphate, 25 mM MgCb, 0.1 mg / mL PPK12, 100 mM Tris- HC1 pH 8.0, were mixed with 2 mg / mL shake flask of enzymes derived from SEQ ID NO 66. Reactions were incubated at 25 °C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quench with 500 pL ACN, filtered, and analyzed by LC.

[0227] SEQ ID NO: 68 was identified as a variant exhibiting improvements in both enzymatic activity and regioselectivity for Reaction K and thus selected for further evolution.Example 15: Engineering of Phe-ligase polypeptides derived from SEQ ID NO; 68 for formation of tetrapeptide of Reaction K

[0228] Using SEQ ID NO: 67 as the template, a library combinatorically sampling beneficial diversity from previous rounds and a 96-position SSM library' targeting the active site region were constructed, expressed, and screened.

[0229] These ATP-dependent Phe-ligase variants of SEQ ID NO: 68 were screened to identify hits improved for activity and regioselectivity under the following conditions: In 96-well-plate and a final volume of 50 pL / well 10 mM substrate 4f, 10 mM substrate 3d, 100 mM Tris-HCl pH 8, 25 mM MgCh, 2 mM ADP. 40 mM polyphosphate, 0. 1 mg / mL PPK12 were incubated in the presence of 10 % (v / v) enzyme lysate. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 25 pL of reaction to 125 pL 80 % ACN. Quenched reactions were filtered and formation of product 5k analyzed by LC-MS using the following flow rate and gradient:Flow Rate: 0.80 mL / minRetention time 5k: 0.900 min

[0230] Notably, a Y372N and K379R double mutant (SEQ ID NO: 69 / 70) provided about 90x improvement in activity.

[0231] Variants catalyzing the highest conversion to product 5k were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 pL. 25 mM 4f, 25 mM 3d, 5 mM ATP, 60 mM polyphosphate, 25 mM MgCE, 0.1 mg / mE PPK12, 100 mM Tris- HCl pH 8.0, w ere mixed with 2 mg / mL shake flask of enzy es derived from SEQ ID NO 58. Reactions were incubated at 25 °C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quenched with 500 pL ACN. filtered, and analyzed by LC.

[0232] SEQ ID NO: 70 was identified as a variant exhibiting improvements in enzymatic activity' and regioselectivity for Reaction K and thus selected for further evolution.Example 16: Engineering of Phe-ligase polypeptides derived from SEQ ID NO; 70 and beyond for formation of tetrapeptide product of Reaction K

[0233] SEQ ID NO: 70 was further evolved sequentially to identify hits that exhibited 1) improved activity, 2) substrate load, 3) phosphate tolerance, 4) co-solvent tolerance, and 5)reduced product and competitive inhibition. Assay conditions and evolution pressures are provided in Table 6 as follows.

[0234] Phe-ligase SEQ ID NO: 64 was purified and crystallized resulting in an X-ray cry stal structure solved at 1.85 A resolution. Using this structure, the Phe-ligase polypeptide was repartitioned into four regions: active site, surfacel. surface2. core. The active site shell residues were selected using a distance criterion. The remaining residues were rank ordered by their solvent accessibilities and grouped into sets of 96 sites. The top two sets were classified as surface sitel and site2, consisting of higher solvent accessibility sites, whereas the remaining were classified as the core. These regions were used for site saturation mutagenesis libraries in subsequent rounds of engineering.

[0235] In the context of a one-pot cascade of Trp-ligase and Phe-ligase to generate tripeptide and tetrapeptide products, respectively , it was found that Id, the desired substrate for Trp-ligase, competed with 4f for Phe-ligase binding, resulting in the formation of undesired products including Id dimers and tetramers consisting of Id ligated to 3c. Molecular dynamics (MD) simulations were performed to identify residues in the Phe-ligase active site (SEQ ID NO: 72) that promoted Id binding, as well as positions outside of the active site where there was significant protein motion upon Id binding. The design criteria were to identify Phe-ligase residues that i) participate in strong interactions with undesired product of Id ligated to 3c, ii) are not critical for desired product 5k binding, and iii) sites that are outside the immediate active sites but show strong motions when bound with the undesired Id substrate. These positions were targeted in an SSM library. Mutations at positions 173, 236 and 237 are predicted to interact with Id and ATP. Specifically, He 173 interacts vxi th the 6.5-ring system in the ATP which results in tighter binding with ATP and allows for slight rearrangement in the active site to accommodate the larger desired substrate. Glu 236 and Lys 237 sit in a loop and interact with hexylamine linker portion of Id. Positions 153, 403 and 404 are predicted from the protein fluctuation analysis to be involved in long-range allostery and were observed to improve Phe-ligase activity by 1.2- to 14- fold over SEQ ID NO: 72.

[0236] Two approaches were taken to improve thermostability of Phe-ligase SEQ ID NO: 96. First, in silico site saturation mutagenesis was performed in Rosetta to identify point mutations that had ddG < -2 REU (Rosetta energy unit). These predicted stabilizing mutations were experimentally tested using a combinatorial library. The best combinatorial variant improved thermostability 17-fold compared to the SEQ ID NO: 96 when variants were subjected to a heat challenge of 37 °C for 1 h while maintaining enzy matic activity7in the absence of a heat challenge. Second, MD simulations were performed at 295K and at 330K, each with the apoprotein and the product and co-factor-bound enzyme. Residues that had major differences (>5 ) in RMSF between high and low temperature MD simulations and showed weakened interactions to the substrate over the course of MD simulations were targeted with site saturation mutagenesis. The best mutants improved thermostability 1.5-3-fold compared to SEQ ID NO: 96 when variants were subjected to a heat challenge of 37 °C for 1 h while maintaining enzymatic activity. Notably, A164S sits at the dimer interface and the serine mutant results in stronger polar interactions at the interface, lowering the dimer interface energy by -25 REU compared with the parental enzyme interface. A cluster of thermostabilizing mutations w ere found on a loop that became increasingly mobile during high temperature MD at residues 350, 352, 359, 363. 367, and 399. P53V mutation sits in a loop that connects two beta strands and the valine variant makes the secondary structure more rigid.Table 6. Engineering of Phe-ligase polynucleotides and polypeptides starting from SEQ ID NO; 69 / 70 for tetrapeptide formation

[0237] The Phe-ligase having the sequence of SEQ ID NO: 107 / 108, indicated in the final row of the above table as emerging from a final round of evolution, contains the following 66 substitutions relative to the reference sequence of SEQ ID NO 4: Q16Y, Q17R, K19G, F21D, P25G, C37H, A50E. H51P, H52P. E61V, S65D, S68Y. F83W, S91G, E94D. R120I. S128K. D131K, G144C, F152L, W164A, G165L, A166H, A169Q, P181E, A182R, A184C, K186S, C204L, S205E, S207P, V208Q, G219N, V230C, D234N, K237T, Q238D, A250T, S253Q, R257P, A264Y, T284S, G299S, E304S, G310N, V325L, A329F, S334Q, Y337N, K350F, V359Y, E361H, K365E. I369L, H371Q, D377K, K379R. S383F. S388E, I391L. S394C. Y396M, I398T, S400W, Q401N, and Y402G.

[0238] SEQ ID NOs: 108, 110, 112, 114, and 116 emerged from 29 rounds of evolutionary pressures as the best-performing ATP-dependent Phe-ligase variants. The engineered polypeptides of SEQ ID NOs: 108, 110, 112, 114. and 116 constitute variants evolved from SEQ ID NO: 105 / 106 in a final round of evolution for which thermostability, substrate affinity, and by-product reduction selection pressures were applied.

[0239] To further improve expression of the Phe-ligase of SEQ ID NO: 108, five codon- optimized versions of SEQ ID NO: 107, i.e.. SEQ ID NOs: 126, 127, 128, 129, and 130, were generated. Each of these five polynucleotides, along with SEQ ID NO: 107, was fermented at 5L scale in BL21(DE3) cells, lysed, flocculated, clarified, and lyophilized. 0.5 g / L of expressed Phe- ligase (SEQ ID NO: 108) was charged to 0.5 mL reactions containing 100 mM HEPES (pH 7.5), 3% TERGITOL, 7.5 mM 4f, 7.5 mM 3d, 5 mM AMP, 75 mM MgCh, 30 mM polyphosphate, 0.2 g / L PPK22, 25 °C. 500 rpm, 1 h. Among these codon-optimized sequences, SEQ ID NO: 127 resulted in expression leading to the highest conversion to desired product.

[0240] These engineered polypeptides show improved 1) activity for the production of Compound 5k, 2) selectivity for Compound 5k over dimers, trimers and / or oligomers of compounds of Formula IV (i.e., Compound 4e and Compound 41), 3) thermostability, 4) reaction kinetics, 5) pH tolerance, 6) phosphate tolerance, 7) reduction in product inhibition and competitive inhibition, and / or 8) substrate loading capacity relative to BAD1200 (SEQ ID NO: 4).Example 17: Production of Phe-ligase polypeptides

[0241] Escherichia coJi BL21(DE3) cells expressing SEQ ID NO: 108 were fermented with Chemically Defined Media (CDM). In brief, a 1 mL aliquot of glycerol stock of the transformed cells was added to 400 mL LB media supplemented with 50 pg / mL kanamycin to generate a seedculture. After incubation at 37 °C and 220 rpm for about 4 to 5 h. the OD600 reached about 2.5 and the seed culture was transferred to a fermenter containing 25 L of CDM supplemented with about 10 g / L glycerol and about 50 pg / mL kanamycin. The agitation rate was set to 400 rpm, with 30 L / min airflow, and the culture was grown at 37 °C. Aliquots of feeding media containing 700 g / L glycerol were added when dissolved oxygen (DO) spiked. The DO was maintained between 20% and 40%, by varying the agitation rate, by feeding pure oxygen, and by changing the amount of feed media. The pH was maintained at 7.0 by adding a solution of 30% ammonium hydroxide. When OD600 reached from 60 AU to 80 AU, protein expression was induced by adding IPTG to a final concentration of 0.2 mM, and the culture temperature was decreased to 25 °C. The cells were harvested when the culture reached a stationary phase at 24 h and OD600 of about 120 AU. After centrifugation, the cells were lysed in 50 mM triethanolamine buffer at pH 7.5 using a homogenizer and at a pressure of 800 bar. Polyethylenimine was added to a final concentration of 0.3% wt / v, and the suspension was incubated for 60 min at room temperature and 300 rpm. followed by centrifugation. The supernatant was concentrated by tangential flow filtration using a 30 kDa membrane. Finally, the concentrated solution was lyophilized to obtain an enzyme powder comprising the ligase with SEQ ID NO: 108. Similar fermentations have been performed with media supplemented with 10 g / L glucose. The feed media composition has also been varied from 43% w / v to 70% w / v glucose. The feed media has also been supplemented with phosphate salts such as ammonium phosphate, ammonium phosphate monobasic, ammonium phosphate dibasic, sodium phosphate, sodium phosphate monobasic, sodium phosphate dibasic, potassium phosphate, potassium phosphate monobasic, potassium phosphate dibasic, and mixtures thereof.Example 18: Production of Trp-ligase polypeptides

[0242] Escherichia coli BL21(DE3) cells transformed with a plasmid encoding for a ligase with SEQ ID NO: 40 were fermented in Chemically Defined Media (CDM). In brief, a 1 mL aliquot of glycerol stock of the transformed cells was added to 400 mL LB media supplemented with 50 pg / mL kanamycin to generate a seed culture. After incubation at 37 °C and 220 rpm for 4 to 5 h, the OD600 reached about 2.5 and the seed culture was transferred to a fermenter containing 25 L of CDM supplemented with 10 g / L glucose and 50 pg / mL kanamycin. The agitation rate was set at 200 rpm, the airflow at 30 L / min, and the temperature at 37 °C. Aliquots of feeding media containing 700 g / L glucose were added when dissolved oxygen (DO) spiked. The DO was maintained between 20% and 40%, by varying the agitation rate, by feeding pure oxygen, and by changing the amount of feed media. The pH was maintained at about 7.0 by adding a solution of30% ammonium hydroxide. When OD600 reached from 60 AU to 80 AU, protein expression was induced by adding IPTG to a final concentration of 0.2 mM, and the culture temperature was decreased to 25 °C. The cells were harvested when the culture reached a stationary phase at 30 h and OD600 of 170 AU. After centrifugation, the cells were lysed in 50 mM triethanolamine buffer at pH 7.5 using a homogenizer and at a pressure of 800 bar. Polyethylenimine was added to a final concentration 0.3% wt / v was added, and the suspension was incubated for 60 min at room temperature and 300 rpm, followed by centrifugation. The supernatant was concentrated by tangential flow filtration using a 10 kDa membrane. Finally, the concentrated solution was lyophilized to obtain an enzyme powder comprising the ligase with SEQ ID NO: 40. This fermentation has also been performed with CDM media supplemented with about 10 g / L glycerol. The feed media composition has also been varied from 43% w / v to 70% w / v glucose. The feed media has also been supplemented with phosphate salts such as ammonium phosphate, ammonium phosphate monobasic, ammonium phosphate dibasic, sodium phosphate, sodium phosphate monobasic, sodium phosphate dibasic, potassium phosphate, potassium phosphate monobasic, potassium phosphate dibasic, and mixtures thereof.

[0243] The disclosed subject matter is not to be limited in scope by the specific embodiments and examples described herein. Indeed, various modifications of the disclosure in addition to those described will become apparent to those skilled in the art from the foregoing description and accompanying figures. Such modifications are intended to fall within the scope of the appended claims.

[0244] All references (e.g., publications or patents or patent applications) cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual reference (e.g.. publication or patent or patent application) was specifically and individually indicated to be incorporated by reference in its entirety for all purposes. Other embodiments are within the following claims.

[0245] The sequences corresponding to SEQ ID NOs: 1-130 are set forth below in Table 7.Table 7: Sequences

Claims

CLAIMSWhat Is Claimed Is:

1. An engineered polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NO: 40, 42, 44, 46, 48, and 50.

2. The polypeptide according to claim 1, having at least 90% sequence identity to SEQ ID NO: 40.

3. The polypeptide according to claim 1, having at least 90% sequence identity to SEQ ID NO: 42.

4. The polypeptide according to claim 1, having at least 90% sequence identity to SEQ ID NO: 44.

5. The polypeptide according to claim 1, having at least 90% sequence identity to SEQ ID NO: 46.

6. The polypeptide according to claim 1, having at least 90% sequence identity to SEQ ID NO: 48.

7. The polypeptide according to claim 1, having at least 90% sequence identity to SEQ ID NO: 50.

8. The polypeptide according to claim 1, wherein the polypeptide comprises the sequence of SEQ ID NO: 40.

9. The polypeptide according to claim 1, wherein the polypeptide comprises the sequence of SEQ ID NO: 42.

10. The polypeptide according to claim 1, wherein the polypeptide comprises the sequence of SEQ ID NO: 44.

11. An engineered polypeptide comprising an amino acid sequence that comprises a stretch of at least 100 consecutive amino acids of any one of SEQ ID NOs: 40, 42, 44, 46, 48, and 50.

12. An engineered polypeptide comprising an amino acid sequence having at least 85% sequence identity to any one of SEQ ID NOs: 108, 110, 112, 1 14, and 116.

13. The polypeptide according to claim 12, having at least 90% sequence identity to SEQ ID NO: 108.

14. The polypeptide according to claim 12, having at least 90% sequence identity to SEQ ID NO: 110.

15. The polypeptide according to claim 12, having at least 90% sequence identity’ to SEQ ID NO: 112.

16. The polypeptide according to claim 12, having at least 90% sequence identity to SEQ ID NO: 114.

17. The polypeptide according to claim 12, having at least 90% sequence identity to SEQ ID NO: 116.

18. The polypeptide according to claim 12. wherein the polypeptide comprises the sequence of SEQ ID NO: 108.

19. The polypeptide according to claim 12, wherein the polypeptide comprises the sequence of SEQ ID NO: 110.

20. The polypeptide according to claim 12, wherein the polypeptide comprises the sequence of SEQ ID NO: 112.

21. An engineered polypeptide comprising an amino acid sequence that comprises a stretch of at least 100 consecutive amino acids of any one of SEQ ID NOs: 108, 110, 112, 114, and 116.

22. The polypeptide of any one of claims 1-21, wherein the polypeptide further comprises an affinity tag.

23. The polypeptide of any one of claims 1-22, wherein the polypeptide is an ATP-dependent ligase.

24. The polypeptide of any one of claims 1-11 and 23, wherein the polypeptide exhibits activity in catalyzing a ligation of an amino acid and a dipeptide to generate a tripeptide.

25. The polypeptide of any one of claims 1-11, 23 and 24, wherein the polypeptide exhibits activity in catalyzing a reaction in which Compound (2c)in the presence of ATP, Mg2+, and a polyphosphate, to generate Compound (3d)26. The polypeptide of claim 25, wherein the polypeptide exhibits greater regioselectivity in catalyzing the reaction than a polypeptide having the amino acid sequence of SEQ ID NO: 4.

27. The polypeptide of claim 25 or 26, wherein the catalysis activity generates oligomers of Compound la, Compound lb, Compound 1c, or Compound Id, or a product of ligation of Compound 3d to Compound Id, in reduced amounts relative to the polypeptide of SEQ ID NO: 4.

28. The polypeptide of any one of claims 12-23, wherein the polypeptide exhibits activity in catalyzing a ligation of an amino acid and a tripeptide to generate a tetrapeptide.

29. The polypeptide of any one of claims 12-23 and 28, wherein the polypeptide exhibits activity in catalyzing a reaction in which Compound (3d)is contacted with Compound (4f)(4f), in the presence of ATP, Mg2+, and a polyphosphate, to generate Compound (5k)30. The polypeptide of claim 29, wherein the polypeptide exhibits greater regioselectivity in catalyzing the reaction than a polypeptide having the amino acid sequence of SEQ ID NO: 4.

31. The polypeptide of claim 29 or 30, wherein the catalysis activity generates oligomers of Compound 4e or Compound 4f, or a product of ligation of Compound 5k to Compound Id, in reduced amounts relative to the polypeptide of SEQ ID NO: 4.

32. The polypeptide of any one of claims 1-31, wherein the polypeptide exhibits higher (1) phosphate tolerance, (2) substrate loading capacity. (3) thermostability, and / or (4) pH tolerance relative to the polypeptide of SEQ ID NO: 4.

33. The polypeptide of any one of claims 1-4 and 11, wherein the polypeptide comprises the sequence of any one of SEQ ID NOs: 119-121.

34. The polypeptide of any one of claims 12-15 and 21, wherein the polypeptide comprises the sequence of any one of SEQ ID NOs: 122-124.

35. A polynucleotide encoding the polypeptide of any one of claims 1-1 1.

36. A polynucleotide comprising a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 39, 41, 43, 45, 47, and 49.

37. The polynucleotide of claim 35 or 36, wherein the polynucleotide comprises the sequence of any one of SEQ ID NOs: 39, 41, 43, 45, 47, and 49.

38. A polynucleotide encoding the polypeptide of any one of claims 12-23.

39. A polynucleotide comprising a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 107, 109. I l l, 113, 115, 126, 127. 128, 129. and 130.

40. The polynucleotide of claim 38 or 39, wherein the polynucleotide comprises the sequence of any one of SEQ ID NOs: 107, 109, 111, 113, 115. 126, 127, 128, 129, and 130.

41. The polynucleotide of any one of claims 38-40, wherein the polynucleotide comprises the sequence of SEQ ID NO: 127.

42. The polynucleotide of any one of claims 35-41, wherein the polynucleotide is codon-optimized for expression in E. coli.

43. An expression vector comprising the polynucleotide of any one of claims 35-42, operably linked to one or more control sequences suitable for directing expression of the encoded polypeptide in a host cell.

44. The expression vector of claim 43, wherein the control sequence comprises a promoter.

45. The expression vector of claim 44, wherein the promoter comprises an E. coli promoter.

46. A host cell comprising the expression vector of any one of claims 43-45.

47. The host cell of claim 46, wherein the host cell is E. coli.

48. A method of producing a polypeptide comprising culturing the host cell of claim 46 or 47, recovering the polypeptide, and isolating the polypeptide.

Citation Information

Patent Citations

  • Peptide production method

    US20100248307A1

  • Mutant enzyme, use thereof and process for preparing tripeptide by using enzymatic method

    US20230407286A1