Synthesis of cyclic oligopeptides
The enzymatic process using engineered ATP-dependent amino acid ligases and carboxylesterases addresses inefficiencies in traditional oligopeptide synthesis by reducing steps and waste, achieving high-yield, regioselective macrocyclic peptide production.
Patent Information
- Application Number
- PCT/US2025/037066
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-15
- Filing Date
- 2025-07-10
- Publication Date
- 2026-01-22
AI Technical Summary
Traditional methods for synthesizing oligopeptides are inefficient, requiring multiple steps, generating waste, and producing undesired byproducts, while enzyme-mediated macrocyclization faces challenges like scalability, robustness, and maintaining desired regiochemistry.
An enzymatic process using engineered ATP-dependent amino acid ligases and carboxylesterases to catalyze amide bond formation in tripeptides and tetrapeptides, reducing steps and waste, and ensuring high yields and regioselectivity.
The process achieves efficient, scalable synthesis of macrocyclic peptides with reduced by-product formation and improved regioselectivity, suitable for large-scale manufacturing.
Smart Images

Figure US2025037066_22012026_PF_FP_ABST
Abstract
Description
SYNTHESIS OF CYCLIC OLIGOPEPTIDES CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 671,551 filed July 15, 2024, the entire contents of which are incorporated by reference herein. REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY
[0002] The contents of the electronic sequence listing (25971-WO-PCT_SL.xml; Size: 281,970 bytes; and Date of Creation: January 30, 2025) are herein incorporated by reference in their entirety. FIELD
[0003] The present disclosure relates to a process for making macrocyclic oligopeptides from amino acids, dipeptides, and / or tripeptides using engineered ATP-dependent amino acid ligase and carboxylesterase enzymes. BACKGROUND
[0004] Traditional approaches to large-scale oligopeptide synthesis generally involve many synthetic steps leading to increased waste and time. For example, chemical amide synthesis relies on protecting group strategies and requires isolations between each step and are often accompanied by generation of stoichiometric byproducts of chemical coupling reagents.
[0005] Enzymes are being increasingly harnessed by biocatalysis, providing green approaches to industrial processes by decreasing the number of synthetic steps and waste, allowing for production of value-added synthetic intermediates and products with lower cost.
[0006] ATP-dependent amino acid ligase enzymes are commonly found in nature in the ribosome and non-ribosomal peptide synthesis pathways and can catalyze the formation of amide bonds through acyl-adenylated intermediate generation (see Ogasawara et al., Chem. Eur. J. 2017, 23, 10714-10724). These enzymes include ATP-grasp ligases, which generate in the presence of ATP activated acyl-phosphate intermediates that are subsequently coupled to an extending peptide chain through a nucleophilic attack. This biocatalytic transformation is particularly attractive for chemical processes because it does not involve the use of protecting group manipulations commonly found in traditional chemical synthesis of peptides. However, use of these enzymes for the large-scale synthesis of oligopeptides has been limited, at least in part due to the limited substrate scope of these enzymes.
[0007] Therapeutic peptides are gaining increasing interest from the pharmaceutical industry due to their potential to specifically target proteins through surface interactions which could be distinct from traditional small-molecule modalities. Macrocyclization of linear peptides through ligation between N and C termini or sidechains has become a promising strategy to address limitations of therapeutic peptides such as susceptibility to proteolytic cleavage and / or low cell permeability, which leads to lower clinical efficacy. However, the synthesis of macrocyclic peptides is difficult to control through traditional chemical methods of amide formation, resulting in lack of specificity and formation of impurities and undesired byproducts. For macrocyclization, an entropically unfavored pre-cyclization conformation must be formed before the desired cyclization can occur. In addition, attempts to macrocyclize oligopeptides can result in a range of bond configurations including head-to-tail, head-to-sidechain, sidechain-to- sidechain, and tail-to-sidechain. Given the multiple chemical approaches from cross-coupling and photochemistry that have been explored for chemoselective macrocyclization of peptides, enzyme-mediated macrocyclization has been of great interest due to its potential efficiency and high regio- and chemo-selectivity. Moreover, enzymatic reactions are often performed under mild conditions and eliminate the need for toxic solvent, holding economic and environmental advantages over traditional chemical synthesis.
[0008] The recent development of biocatalysts has provided a promising solution for the increasing need for a more sustainable manufacturing of chemicals and medicine. However, using enzymes at industrial scale applications has often faced challenges such as scalability, robustness, and tolerance to relevant process conditions, which could be addressed by protein engineering.
[0009] Although traditional macrocyclic peptide synthesis methods are able to make oligopeptides described herein, they are not suitable for large-scale manufacturing due to: (1) the need for multiple steps using protecting group strategies that require isolation between each step and that lead to undesired waste and / or stoichiometric byproducts; (2) failure to maintain desired regiochemistry (i.e., regioselectivity), (3) low yields; and (4) lengthy time commitment for manufacturing.
[0010] Therefore, there is a need in the art for processes that can generate oligopeptides, including cyclized oligopeptides in large scale while maintaining desired regiochemistry, reduced by-product formation and high yields. SUMMARY
[0011] This disclosure relates to an enzymatic route for making macrocyclic peptides that reduces the number of synthetic steps and amount of waste generated. An embodiment of this disclosure relates to an enzymatic process where only one desired amide bond is formed at each step. In various embodiments, provided is a process for using tripeptide and tetrapeptide intermediates in the presence of engineered carboxylesterase enzymes in the synthesis of complex peptides such as macrocyclic peptides and / or complex peptides that are active agents.
[0012] An embodiment of this disclosure relates to a process for making macrocyclic peptides from tripeptides and / or tetrapeptides, using engineered polypeptides (e.g., carboxylase enzymes) capable of catalyzing a macrocyclization through amide bond formation in oligopeptides. An aspect of this embodiment of the process is realized when the desired regioselectivity is obtained.
[0013] Other embodiments of the disclosure relate to processes for generating tripeptides and / or tetrapeptides via one or more ATP-dependent amino acid ligase-catalyzed reactions. These processes comprise performing a reaction with any of the disclosed engineered ligase polypeptides in Table 7. These processes may comprise the conversion of a dipeptide into a tripeptide (through a ligation reaction) using any of the disclosed engineered ligase polypeptides in Table 7. These processes may comprise the conversion of a tripeptide into a tetrapeptide (through a ligation reaction) using any of the disclosed engineered ligase polypeptides in Table 7. In various aspects, the ultimate product of ligation reactions is a tetrapeptide having non- canonical (or modified) amino acids.
[0014] Another embodiment of this disclosure relates to coupling (i.e., ligating) oligopeptides, such as linear tripeptides and tetrapeptides from activated amino acid electrophiles and dipeptide and tripeptide nucleophiles using engineered ATP-dependent amino acid ligase enzymes for use as intermediates in the synthesis of macrocyclic peptides. Still another embodiment of this disclosure relates to a catalytic ligase approach wherein the stoichiometric byproduct of the ligase sequence is simple phosphate.
[0015] Other embodiments, aspects and features of the present invention are either further described in or will be apparent from the ensuing description, examples, and appended claims. DETAILED DESCRIPTION
[0016] The present disclosure relates to an efficient and scalable enzymatic process for synthesizing macrocyclic peptides of Formula 1’ or salt, hydrate, and / or solvate thereof:25971 Comprising the steps of,1) combining a NH2R4with an engineered carboxylesteraseof Formula 1’, and 2) isolating the compound of Formula 1’; wherein R1and R5are independently selected from hydrogen, C1-10 alkyl, aryl, and heteroaryl; R2is selected from hydrogen, C1-6 alkyl, C2-6 alkenyl, said alkyl and alkenyl optionally substituted with 1 to 3 groups of R, R is selected from NH2, OH, OC1-6alkyl, OC2-6alkenyl, and NHC1-6alkyl, N(C1-6 alkyl)2; R3is selected from hydrogen and halogen, R4is selected from hydrogen, OH, -OC1-6alkyl, -O(CH2)nCOOR5, -O(CH2)nC(O)SR5, and -O(CH2)nC(O)NHR5and n is 1 to 3.
[0017] Another embodiment of the disclosure is directed to a process for preparing Compound 3k’ comprising the steps of a) combining the compound of Formula 3d’: R4H OHor salt thereof, with Compound 1f’:or salt thereof, in the presence of anenzyme, adding a polyphosphate kinase, adenosine phosphate, inorganic magnesium salt, and a phosphate donor, to generate a compound of Formula 3k’ and b) isolating 3k’, wherein R1, R2, R3and R4are as described herein and R6is selected from hydrogen and (CH2)nNH2. An aspect of this embodiment is realized when Compound 3k’ is a nonnatural tetrapeptide. Another aspect of this embodiment is realized when the ligase enzymes act in the presence of kinase enzyme that utilizes a polyphosphate to synthesize ATP from ADP and / or AMP, to generate Compound (3k’). Compound 3k’ is a modified Phenylalanine-Proline- Threonine tetrapeptide, relative to SEQ ID NO: 4 or wild-type SEQ ID NO:2.
[0018] Another embodiment of the disclosure is directed to a process for preparing Compound 3d’ by a) combining the compound of Formula 2c’or a salt thereof, with a compound of Formula 1d’ or salt thereof, in the presence ofpolyphosphate kinase, adenosine phosphate, an inorganic salt of magnesium, and a phosphate donor, to generate a compound of Formula 3d’ and b) isolating 3d’. An aspect of this embodiment is realized when Compound 2c’ is a nonnatural dipeptide and Compound 1d’ is a nonnatural amino acid. Another aspect of thisembodiment is realized when the ligase enzymes act in the presence of kinase enzyme that utilizes a polyphosphate to synthesize ATP from ADP and / or AMP, to generate Compound (3d’). Compound 3d’ is a modified Tryptophan-Proline-Threonine tripeptide, relative to SEQ ID NO: 4 or wild-type SEQ ID NO: 2. Another aspect of this embodiment is realized when the Compound of 3d’ is a nonnatural tripeptide.
[0019] Another embodiment of the disclosure is directed to a process for preparing Compound of Formula 1’, or a salt, hydrate, and / or solvate thereof,comprising the steps of a) combining 1f’, 2c’ and 1d’ or saltTrp-ligase enzymes, polyphosphate kinase, adenosine phosphate, an inorganic salt of magnesium, and a phosphate donor, to produce a compound of Formula 3k’ NH2Meb) adding an engineeredprovide a compound of Formula 1’, andc) isolating the compound of Formula 1’. An aspect of this embodiment is realized when Compound 2c’ is a nonnatural dipeptide, Compound 1d’ is a nonnatural amino acid, Compound of 3d’ is a nonnatural tripeptide and Compound 3k’ is a nonnatural tetrapeptide. Another aspect of this embodiment is realized when the ligase enzymes act in the presence of kinase enzyme that utilizes a polyphosphate to synthesize ATP from ADP and / or AMP, to generate Compound 3k’.
[0020] Another embodiment of the disclosure is directed to a process for preparing Compound 1, or salt, hydrate, and / or solvate thereof, comprising the steps ofa) combining 1f, 2c and 1d or saltTrp-ligase enzymes, polyphosphate kinase selected from wild-type PPK22 and PPK12; adenosine phosphate selected from ATP, ADP, AMP, or combination thereof; inorganic magnesium salt selected from magnesium chloride hexahydrate, and magnesium sulfate; and phosphate donor selected from sodium hexametaphosphate, sodium polyphosphate, and salts or analogs thereof to produce compound 3kb) adding an provide Compound 1, andc) isolating Compound is realized when Compound 2c is a nonnatural dipeptide, Compound 1d is a nonnatural amino acid, Compound 3k is a nonnatural tetrapeptide and Compound 1 is a nonnatural macrocyclic peptide. Another aspect of this embodiment is realized when the ligase enzyme act in the presence of kinase enzyme that utilizes a polyphosphate to synthesize ATP from ADP and / or AMP, to generate Compound 3k.
[0021] The present disclosure provides processes that utilizes engineered polypeptides derived from Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1, SEQ ID NO: 2). These engineered polypeptides exhibit improved enzyme catalysis properties relative to wild-type SEQ ID NO: 2, including improved enzyme activity for the synthesis of oligopeptides from (i) natural or nonnatural dipeptide nucleophiles and natural or nonnatural amino acid carboxylates and (ii) natural or nonnatural tripeptide nucleophiles and natural or nonnatural amino acid carboxylates, which properties were engineered through iterative rounds of directed evolution.
[0022] An embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the pH is maintained at about 6.5 to about 9.0, preferably about 7.0 to about 8.0 using a base. An aspect of this embodiment is realized when the base is any base that will maintain the desired pH, for example NaOH, KOH, NH4OH, and the like, preferably NaOH.
[0023] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the magnesium ion is derived from any magnesium ion (Mg2+)source known to those skilled in the art, e.g., magnesium chloride hexahydrate, or magnesium sulfate.
[0024] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the polyphosphate is any polymeric phosphate anhydride of chain length equal to or greater than 4. Examples of polyphosphate are sodium hexametaphosphate, sodium polyphosphate, and the like.
[0025] The disclosed methods may be performed in the presence of an ATP regeneration system, e.g., sources of activated phosphates. This ATP regeneration system may comprise, or consist of, a polyphosphate kinase, a phosphate donor such as polyphosphate and an adenosine cofactor such as AMP, ADP or ATP or a combination thereof. Thus, in some embodiments, the process for making a compound of Formula 1’ or Compound 1 is realized when sub- stoichiometric amounts of adenosine phosphate selected from AMP, ADP, ATP, or a combination thereof in the presence of a polyphosphate kinase (PPK), and a phosphate donor such as polyphosphate is used. Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the adenosine phosphate is ATP.
[0026] In some embodiments, the process for making a compound of Formula 1’ or Compound 1 is realized when the PPK is any PPK that generates some ATP as a part of its equilibrium speciation of adenosine phosphates. In some embodiments, the PPK is wild-type PPK22 or PPK12.
[0027] In one embodiment, when an engineered ligase enzyme is used the process step is performed in a medium containing cofactors ATP, inorganic magnesium salt, and / or propionyl phosphate. In some embodiments, when an engineered ligase enzyme is used the process step is performed in a medium in the presence of ADP or ATP in the presence of propionyl phosphate and an acetate kinase. In some embodiments, the acetate kinase is wild-type acetate kinase ACK- 101.
[0028] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when a nonionic surfactant is added. An aspect of this step is realized when the nonionic surfactant is selected from Tergitol™, Triton X-100®, and / or reduced Triton. Another aspect of this step is realized when the amount of nonionic surfactant added is from about 0.5% to about 5% v / v. Still another aspect of this step is realized when the amount of nonionic surfactant added is from about 1% to about 5% v / v. Yet another aspect of this step is realized when the amount of nonionic surfactant added is from about 2% to about 3% v / v.
[0029] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the temperature ranges from about 15^C to about 35^C, about 20^C to about 30^C, and preferably 25^C.
[0030] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the engineered ligase enzyme used to make 3d’ or 3d is derived from wild-type ligase Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1, SEQ ID NO: 2. Another embodiment of the process is realized when the engineered ligase enzyme is an ATP-dependent amino acid ligase enzyme known as“Trp-ligase”. An aspect of this process is realized when the engineered ATP-dependent amino acid Trp-ligase enzymes are selected from engineered polypeptides comprising an amino acid sequence having at least 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 40, 42, 44, 46, 48, and 50. In some aspects, the engineered ATP-dependent amino acid Trp-ligase polypeptides have at least 98%, at least 99%, or 100% identity to any one of SEQ ID NOs: 40, 42, 44, 46, 48, and 50. Another aspect of this process is realized when the engineered ATP-dependent amino acid Trp-ligase is selected from polypeptide sequence SEQ ID NOs: 40, 42, 44, 46, 48, and 50. In a further aspect, the engineered ATP-dependent Trp-ligase enzyme is SEQ ID NO: 40. In a further aspect, the engineered Trp-ligase enzyme is SEQ ID NO: 42. In a further aspect, the engineered ATP-dependent Trp-ligase enzyme is SEQ ID NO: 44. In a further aspect, the engineered Trp-ligase enzyme is SEQ ID NO: 46. In a further aspect, the engineered ATP-dependent Trp-ligase enzyme is SEQ ID NO: 48. In a further aspect, the engineered Trp-ligase enzyme is SEQ ID NO: 50. The reactions can also be performed with enzymes immobilized on a solid support, i.e., via His tags on metal affinity resins.
[0031] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the engineered ligase enzyme used to make 3k’ or 3k is derived from wild-type ligase Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1, SEQ ID NO: 2. Another embodiment of the process is realized when the engineered ligase enzyme is an ATP-dependent amino acid ligase enzyme known as “Phe-ligase”. An aspect of this process is realized when the engineered ATP-dependent amino acid Phe-ligase enzymes are selected from engineered polypeptides comprising an amino acid sequence having at least about 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 108, 110, 112, 114, and 116. In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 108, 110, and 112. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 108. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 110. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 112. In some embodiments, the enzyme polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 114. In some embodiments, the enzyme polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 116.
[0032] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the reaction in is aged from about 12 hours to about 72 hours,25971 preferably from about 36 hours to about 42 hours. An aspect of this embodiment is realized when the aging is accompanied with stirring during part or the entire time of aging.
[0033] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when oligomers of the Substrate Compound of Formula I (e.g., Compound 1a, Compound 1b, Compound 1c, or Compound 1d), or a product of ligation of Compound 3k’ or 3k to Compound 1d’ or 1d, are produced in reduced amounts relative to a corresponding method in which Compound (2c’ or 2c) and Compound (1d’ or 1d) are contacted with the polypeptide of SEQ ID NO: 4.
[0034] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when oligomers of the Substrate Compound of Formula IV (e.g., Compound 1e or Compound 1f), are produced in reduced amounts relative to a corresponding method in which Compound (3d’ or 3d) and Compound (1f’ or 1f) are contacted with the polypeptide of SEQ ID NO: 4.
[0035] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when steps 1), 1a), 2) or 2a) exhibits higher (1) phosphate tolerance, (2) substrate loading capacity, (3) thermostability, and / or (4) pH tolerance, relative to a corresponding method in which Compound (2c) and Compound (1d) are contacted, or in which Compound (3d) and Compound (1f) are contacted, with the polypeptide of SEQ ID NO: 4.
[0036] The engineered carboxylesterase enzyme in step 3) is derived from a wild-type esterase obtained from a Roseibacillus bacterial species, referred to as “carboxylesterase” ( see Muller et al., Discovery and Design of Family VIII Carboylesterases a Highly Efficient Acyltransferases, Angew. Chemie Int. Ed., 2021, 60, 2013-2017, which is incorporated herein by reference).
[0037] An embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the engineered carboxylesterase enzyme used in the process is derived from engineered nucleic acid and protein variants of the wild-type carboxylesterase defined by sequences SEQ ID Nos: 131 and 132, respectively, identified from the public genome database (GenBank: MBG69902.1). The process utilizes engineered carboxylesterase enzymes to prepare the compound of Formula 1’ or Compound 1 by macrocyclization between an amine on phenylalanine and tert-butyl ester appendage on proline of Compound 3k’ or 3k.
[0038] In an aspect of this embodiment, the engineered carboxylesterase enzyme used in the macrocyclization step comprises an amino acid sequence having at least 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of the amino acid sequences disclosed in Table 15. In an aspect of this embodiment, the engineered carboxylesterase enzyme used in the macrocyclization process comprises an amino acid sequence having at least 92.5%, 95%, 96%,25971 98%, or 99% sequence identity to any one of SEQ ID NOs: 152, 154, 156, and 158. Another aspect of this embodiment is realized when the engineered carboxylesterase has the polypeptide sequence of SEQ ID NOs: 152, 154, 156, and 158. Another aspect of this embodiment is realized when the engineered carboxylesterase, has the polypeptide sequence of SEQ ID NO: 152. Another aspect of this embodiment is realized when the engineered carboxylesterase, has the polypeptide sequence of SEQ ID NO: 154. Another aspect of this embodiment is realized when the engineered carboxylesterase, has the polypeptide sequence of SEQ ID NO: 156. Another aspect of this embodiment is realized when the engineered carboxylesterase, has the polypeptide sequence of SEQ ID NO: 158. In some embodiments, the polypeptides comprise an amino acid sequence that comprises 100 or more consecutive amino acids in common with any one of SEQ ID NOs: 152, 154, 156, and 158, such as a stretch of at least 300, 325, 350, 375, or 400 consecutive amino acids of any one of SEQ ID NOs SEQ ID NOs: 152, 154, 156, and 158. In some embodiments, the disclosed polypeptides further comprise one or more affinity tags, such as a hexa-histidine tag.
[0039] In some embodiments, the polypeptides exhibit activity in catalyzing a macrocyclization of a linear tetrapeptide. These polypeptides may be carboxylesterases known as a “macrocyclase”.
[0040] In some embodiments, the carboxylesterase enzymes may be used in a multi-step synthesis with a tryptophan ligase and a phenylalanine ligase, such as an engineered Trp-ligase and an engineered Phe-ligase. In an embodiment, this multi-step synthesis may generate a cyclic tetrapeptide product (e.g., Compound 1) from one or more monomers, such as from a linear dipeptide or linear tripeptide, or a combination thereof. In various embodiments, the disclosure provides methods of synthesizing a tetrapeptide (e.g., a tetrapeptide that includes a modified tryptophan, a modified phenylalanine and / or a modified proline) using the engineered carboxylesterases. In some aspects, the carboxylesterase enzymes exhibit improved activity in an enzymatic cascade relative to wild-type enzyme. In some aspects, the carboxylesterase enzymes exhibit improved regioselectivity, solvent tolerance, and / or thermostability in an enzymatic cascade relative to wild-type enzyme.
[0041] An embodiment of this disclosure is realized when about 0.1 to about 1 equivalent, preferably about 0.3 to about 0.5 equivalent, of diatomaceous earth relative to amino acid monomer 1d hydrobromide salt [Trp-Linker.HBr] is added before or during charge of the carboxylesterase enzyme. An example of diatomaceous earth is CELITE® 545.
[0042] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the reactions are performed in an aqueous environment. Another25971 embodiment, the reactions are performed in an aqueous solution optionally containing organic co-solvent such as MeCN, DMSO DMAc, DMF or a mixture thereof.
[0043] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 is realized when the reactions are performed in the presence of a buffer, for example HEPES.
[0044] Another embodiment of the process for making a compound of Formula 1’ or Compound 1 relates to an enzymatic process where only one desired amide bond is formed at each step. An aspect of this embodiment of the process is realized when the desired regioselectivity is obtained.
[0045] In other aspects, a process for making a compound of Formula 1’ or Compound 1 is realized where the enzymes can be used in a multi-step synthesis scheme. In other aspects, a process for making a compound of Formula 1’ or Compound 1 is realized where the enzymes are used in multi-step scheme in a one-pot process, such as a one-pot reaction involving one or more ligases and carboxylesterases. In other aspects, a process for making a compound of Formula 1’ or Compound 1 is realized where the ATP-dependent amino acid ligase enzymes are used in a cascade scheme, such as a one-pot reaction involving one or more carboxylesterase enzymes.
[0046] In other aspects, preparation of linear tripeptide and / or linear tetrapeptide intermediates in a multi-step scheme in a one-pot process, such as a one-pot reaction involving one or more amino ATP-dependent acid ligase enzymes is disclosed. In other aspects, a process is disclosed where the ATP-dependent amino acid ligase enzymes are used in a cascade scheme, such as a one-pot reaction involving one or more amino acid ligase enzymes.
[0047] Another embodiment relates to a process where each step in the synthesis of tripeptide and tetrapeptide intermediates is performed in isolation of the other steps and resulting intermediates are purified and isolated between steps. In other embodiments, the enzymes and substrates in steps to make the tripeptide and tetrapeptide can all be combined into one vessel (one pot synthesis) and allowed to react simultaneously. In other embodiments, dipeptide 2c’, 2c and / or tripeptide 3d’, or 3d intermediates can be synthesized chemically, and subsequently reacted with appropriate engineered ATP-dependent ligase enzymes (e.g., Phe-ligase and / or Trp- ligase enzymes) in cascades to yield a tetrapeptide product. In some aspects, the steps to make the tripeptide and tetrapeptide may be run simultaneously in cascades such that isolations and / or purifications of intermediates is eliminated.
[0048] Whether carrying out the ligation steps with whole cells, cell extracts or purified ATP- dependent amino acid ligase enzymes, a single ATP-dependent amino acid ligase enzyme may be25971 used or, alternatively, mixtures of two or more ATP-dependent amino acid ligase enzymes may be used.
[0049] In other aspects, a macrocyclization process is disclosed where the engineered carboxylesterase enzymes and ATP-dependent amino acid ligase enzymes can be used in a multi- step synthesis scheme. In other aspects, a macrocyclization process is disclosed where the engineered carboxylesterase enzymes and ATP-dependent amino acid ligase enzymes are used in multi-step scheme in a one-pot process, such as a one-pot reaction. In other aspects, a macrocyclization process is disclosed where the engineered carboxylesterase enzymes and ATP- dependent amino acid ligase enzymes are used in a cascade scheme, such as a one-pot reaction.
[0050] In other aspects, the disclosure relates to an enzymatic process for making a compound of Formula 1’ or Compound 1 where each step is run simultaneously in cascades thereby eliminating one or more isolation and / or purification steps as well as waste typically generated in such operations. An aspect of this embodiment relates to an enzymatic process wherein time required for manufacturing is shortened. Another aspect of this embodiment relates to a one pot reaction wherein purification and isolation are only required to provide Formula 1’ or Compound 1. An aspect of this embodiment relates to is an enzymatic process that eliminates waste. An aspect of this embodiment relates to a catalytic ligase approach to amide bond formation that avoids generation of the complex stoichiometric byproducts of chemical coupling reagents.
[0051] The present disclosure further relates to an efficient and scalable enzymatic process for synthesizing macrocyclic peptide process for synthesizing isopropyl ((11S,12S,13S,9S,12S)-9- amino-12-((1-(6-aminohexyl)-5-fluoro-1H-indol-3-yl)methyl)-4,10,13-trioxo-2-oxa-5,11-diaza- 1(3,1)-pyrrolidina-7(1,3)-benzenacyclotridecaphane-12-carbonyl)-L-threoninate [(Compound 1), or salt, hydrate and / or solvate thereof, comprising: 1a”) adding a carboxylesterase enzyme to an aqueous solution containing isopropyl ((2S,3S)-1- ((S)-2-((S)-2-amino-3-(3-(aminomethyl)phenyl)propanamido)-3-(1-(6-aminohexyl)-5-fluoro-1H- indol-3-yl)propanoyl)-3-(2-(tert-butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate 3k resulting in a slurry, and 2a”)filtering the slurry to provide Compound 1, wherein the reaction conditions, including the substrates and enzymes, for this step are described herein. An aspect of this embodiment of the disclosure is realized when Compound 1 is a cyclic tetrapeptide having non-canonical (or modified) amino acids. An aspect of this embodiment of the process is realized when Compound 1 is a cyclic tetrapeptide having the desired regioselectivity.
[0052] The present disclosure further provides efficient and scalable enzymatic processes for preparing intermediates used in the manufacturing of a compound of Formula 1’, Compound 1,25971 or a salt, hydrate, and / or solvate thereof. In an embodiment, the disclosure relates to a process for preparing isopropyl ((2S,3S)-1-((S)-2-((S)-2-amino-3-(3-(aminomethyl)phenyl)propanamido)-3- (1-(6-aminohexyl)-5-fluoro-1H-indol-3-yl)propanoyl)-3-(2-(tert-butoxy)-2- oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate (Compound 3k) comprising: combining, isopropyl ((2S,3S)-1-((S)-2-amino-3-(1-(6-aminohexyl)-5-fluoro-1H-indol-3- yl)propanoyl)-3-(2-(tert-butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate (Compound 3d) with Compound 1f: in the presence of an engineered ligase kinase, adenosine phosphate, aninorganic salt of magnesium, and a polyphosphate, to generate compound 3k conditions, including the substrates and enzymes, for this step are described herein.
[0053] In another embodiment, the disclosure relates to a process for preparing isopropyl ((2S,3S)-1-((S)-2-amino-3-(1-(6-aminohexyl)-5-fluoro-1H-indol-3-yl)propanoyl)-3-(2-(tert- butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate (Compound 3d) comprising: combining a compound 2c Formula 2c O OtBu with a compound of Formula 1d25971 in the presence of an engineered ligase enzyme, polyphosphate kinase, adenosine phosphate, an inorganic salt of magnesium, and a polyphosphate, to generate compound 3d, wherein the reaction conditions, including the substrates and enzymes, for this step are described herein. See USSN 63 / 668,869, filed 07 / 09 / 2024; incorporated herein in its entirety for illustration of how to make 2c.
[0054] An embodiment of the process for making compounds of Formula 1’ or Compound 1, or salt, hydrate, and / or solvate thereof is realized when R1and R4are independently selected from C1-10alkyl, aryl, and heteroaryl. An aspect of this embodiment is realized when R1is selected from methyl, ethyl, isopropyl, n-propyl, n-butyl, sec-butyl, isobutyl, tertbutyl, cyclohexyl, cyclopentyl, cyclobutyl, cyclopropyl, benzyl, phenyl, naphthyl, pyridyl, and the like; R2is selected from hydrogen, C1-6 alkyl, C2-6 alkenyl, C2-6 alkynyl, and (CH2)nR, said alkyl, alkenyl and alkynyl optionally substituted with 1 to 3 groups of R and R is selected from NH2, OH, OC1-6alkyl, and OC2-6 alkenyl; and R3is selected from hydrogen and halogen.
[0055] Another aspect of the disclosure is realized by a compound of structural Formula 1a or 1’: Oor a pharmaceutically acceptable salt, hydrate, and / or solvate thereof, wherein R1and R5are independently selected from hydrogen, C1-10 alkyl, aryl, and heteroaryl; R2is selected from hydrogen, C1-6alkyl, C2-6alkenyl, said alkyl and alkenyl optionally substituted with 1 to 3 groups of R, R is selected from NH2, OH, OC1-6 alkyl, OC2-6 alkenyl; R3is selected from hydrogen and halogen, R4is selected from hydrogen, OH, -OC1-6alkyl, -O(CH2)nCOOR5, and n is 1 to 3.
[0056] A subembodiment of Formula 1a and 1’ is realized when R1is selected from hydrogen, ethyl, propyl, isopropyl, butyl, pentyl and hexyl. An aspect of this subembodiment is realized when R1is isopropyl.
[0057] Another embodiment of Formula 1a or Formula 1’ is realized when R2is selected from hydrogen, -(CH2)nOH, -(CH2)nNH2, -(CH2)nNHCH3, -(CH2)nNHCH2CH3, and n is 1 to 6. A subembodiment of this aspect is realized when R2is selected from -(CH2)nOH, and -(CH2)nNH2.25971 Another subembodiment of this aspect is realized when R2is -(CH2)nOH. Another subembodiment of this aspect is realized when R2is -(CH2)nNH2.
[0058] Another embodiment of Formula 1a or Formula 1’ is realized when R3is hydrogen. Another embodiment of Formula 1a or Formula 1’ is realized when R3is halogen selected from chlorine, fluorine, bromine and iodine, preferably fluorine.
[0059] Another embodiment of Formula 1a or Formula 1’ is realized when R1is selected from hydrogen, ethyl, propyl, isopropyl, butyl, pentyl and hexyl; R2is selected from hydrogen, - (CH2)nOH, -(CH2)nNH2, -(CH2)nNHCH3, -(CH2)nNHCH2CH3, and n is 1 to 6; and R3is hydrogen or halogen. An aspect of this embodiment is realized when R1is isopropyl; R2is selected from - (CH2)nOH, and -(CH2)nNH2, n is 1 to 6; and R3is hydrogen or halogen.
[0060] Another aspect of the disclosure is realized by a compound of structural Formula 3kaor 3k’:wherein R1, R2and R3are as described herein.
[0061] Another aspect of the disclosure is realized by a compound of structural Formula 3daor 3d’:wherein R1, R2and R3are as described herein.
[0062] Non-limiting examples of the compounds of this disclosure are:25971 isopropyl ((11S,12S,13S,9S,12S)-9-amino-12-((1-(6-aminohexyl)-5-fluoro-1H-indol-3- yl)methyl)-4,10,13-trioxo-2-oxa-5,11-diaza-1(3,1)-pyrrolidina-7(1,3)-benzenacyclotridecaphane- 12-carbonyl)-L-threoninate, isopropyl ((2S,3S)-1-((S)-2-((S)-2-amino-3-(3-(aminomethyl)phenyl)propanamido)-3-(1-(6- aminohexyl)-5-fluoro-1H-indol-3-yl)propanoyl)-3-(2-(tert-butoxy)-2-oxoethoxy)pyrrolidine-2- carbonyl)-L-threoninate, and isopropyl ((2S,3S)-1-((S)-2-amino-3-(1-(6-aminohexyl)-5-fluoro-1H-indol-3-yl)propanoyl)-3-(2- (tert-butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate, or a pharmaceutically acceptable or a pharmaceutically acceptable salt, hydrate, and / or solvate thereof. Definitions
[0063] Listed below are definitions of various terms used herein. These definitions apply to the terms as they are used throughout this specification and claims, unless otherwise limited in specific instances, either individually or as part of a larger group.
[0064] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, and peptide chemistry are those well-known and commonly employed in the art.
[0065] As used herein, the articles “a” and “an” refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element. Furthermore, use of the term “including” as well as other forms, such as “include,” “includes,” and “included,” is not limiting.
[0066] As used herein, the term “about” in quantitative terms refers to plus or minus 10% of the value it modifies (rounded up to the nearest whole number if the value is not sub-dividable, such as a number of molecules or nucleotides).
[0067] All ranges disclosed herein are inclusive of the recited endpoint and independently combinable (for example, the range of “from 50 mg to 500 mg” is inclusive of the endpoints, 50 mg and 500 mg, and all the intermediate values). The endpoints of the ranges and any values disclosed herein are not limited to the precise range or value; they are sufficiently imprecise to include values approximating these ranges and / or values.
[0068] As used herein, the term “comprising” may include the embodiments “consisting of” and “consisting essentially of.” The terms “comprise(s),” “include(s),” “having,” “has,” “may,” “contain(s),” and variants thereof, as used herein, are intended to be open-ended transitional25971 phrases, terms, or words that require the presence of the named ingredients / steps and permit the presence of other ingredients / steps. However, such description should be construed as also describing compositions or processes as “consisting of” and “consisting essentially of” the enumerated components, which allows the presence of only the named components or compounds, along with any acceptable carriers or fluids, and excludes other components or compounds.
[0069] “Derived from” as used herein in the context of enzymes, identifies the originating enzyme, and / or the gene encoding such enzyme, upon which the enzyme was based. For example, the disclosed tryptophan ATP-dependent ligase enzymes (or Trp-ligases) are “derived from” the wild-type ligase enzyme of SEQ ID NO: 1 / 2 (or ligase of SEQ ID NO: 3 / 4 that is codon-optimized). The disclosed phenylalanine ATP-dependent ligase enzymes (or Phe-ligases) are “derived from” the wild-type ligase enzyme of SEQ ID NO: 1 / 2 (as well as the codon- optimized ligase of SEQ ID NO: 3 / 4). As used herein in the context of an amino acid residue, a “derivative” refers to a modified amino acid that is derived from a canonical L-amino acid.
[0070] As used herein, the terms “protein,” “polypeptide,” and “peptide” are used interchangeably herein to denote a polymer of at least two amino acids covalently linked by an amide bond, regardless of length or post-translational modification (e.g., glycosylation, phosphorylation, lipidation, myristoylation, ubiquitination, and the like). Included within this definition are D- and L-amino acids, and mixtures of D- and L-amino acids, as well as polymers comprising D- and L-amino acids, and mixtures of D- and L-amino acids. Further included within this definition are linear, branched and cyclic peptides. Proteins, polypeptides, and peptides may include a tag (e.g., an affinity tag), such as a histidine tag and / or a signal peptide region. As used herein, a “polypeptide” may encode an enzyme. The disclosure contemplates the use of polypeptides as substrates, products, and intermediates in an enzymatic reaction.
[0071] As used herein, the terms “amino acid” or “residue” as used in context of the polypeptides disclosed herein refers to the specific monomer at a sequence position of a polypeptide molecule. Amino acids are referred to herein by either their commonly known three- letter symbols or by the one-letter symbols recommended by International Union of Pure and Applied Chemistry (IUPAC) - International Union of Biochemistry (IUB) Biochemical Nomenclature Commission. This term encompasses modified (i.e., covalently or non-covalently modified), or non-canonical or non-natural, amino acids. For instance, this term encompasses covalently modified tryptophan, covalently modified proline, covalently modified phenylalanine, and covalently modified threonine amino acid molecules.25971
[0072] “Mutation” refers to any change in a polypeptide or polynucleotide sequence, and encompasses any number (i.e., one or more) of substitutions, deletions, insertions, and / or rearrangements present in a sequence compared to a reference sequence.
[0073] As used herein with respect to amino acid sequences, a “substitution” refers to a difference in the amino acid residue at a position of a polypeptide sequence relative to the amino acid residue at a corresponding position in a reference sequence. In some instances, the present disclosure provides specific amino acid differences denoted by the conventional notation “AnB,” where A is the single letter identifier of the residue in the reference sequence, n is the number of the residue position in the reference sequence, and B is the single letter identifier of the residue substitution in the sequence of the engineered polypeptide.
[0074] The term “amino acid substitution set” or “substitution set” refers to a group of amino acid substitutions in a polypeptide sequence, as compared to a reference sequence. For example, a substitution set may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-20, 20-25 or more than 25 amino acid substitutions.
[0075] “Corresponding to,” “reference to” or “relative to” when used in the context of the numbering of a given amino acid or polynucleotide sequence refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. In other words, the residue number or residue position of a given polymer is designated with respect to the reference sequence rather than by the actual numerical position of the residue within the given amino acid or polynucleotide sequence. For example, a given amino acid sequence can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences. In these cases, although the gaps are present, the numbering of the residue in the given amino acid or polynucleotide sequence is made with respect to the reference sequence to which it has been aligned.
[0076] As used herein, “isolated polypeptide” refers to a composition in which the polypeptide is substantially separated from other contaminants that naturally accompany it (e.g., protein, lipids, and polynucleotides). The term embraces polypeptides that have been removed or purified from their naturally occurring environment or expression system (e.g., within a host cell or via in vitro synthesis). The recombinant polypeptides may be present within a cell, present in the cellular medium, or prepared in various forms, such as lysates or isolated preparations. As such, in some embodiments, the recombinant polypeptides can be an isolated polypeptide.
[0077] As used herein, a “ligase” is a polypeptide having an enzymatic capability of ligating, or coupling together, two molecules into an oligomer in which the molecules are linked by a bond, such as an amide bond. In various embodiments, the ligases of the disclosure are ATP-dependent25971 ligases, in that they utilize the cofactor ATP as an activating agent. In some embodiments, the disclosed ligases are capable of catalyzing a ligation reaction between an amino acid and a di- or tripeptide. Ligases as used herein include naturally occurring (wild type) polypeptides as well as non-naturally occurring engineered polypeptides generated by human manipulation.
[0078] As used herein, a “carboxylesterase” is a polypeptide having an enzymatic capability of forming an amide bond between an amine and an ester, including amine and ester moieties present in a single molecule. Carboxylesterases are a subtype of esterase. In various embodiments, the disclosed carboxylesterases are capable of catalyzing a cyclization of a linear tetrapeptide through amide bond formation. Carboxylesterases as used herein include naturally occurring (wild-type) polypeptides as well as non-naturally occurring engineered polypeptides generated by human manipulation.
[0079] As used herein, a “macrocyclase” is an engineered variant of carboxylesterase that exhibits high activity in catalyzing a cyclization of a linear tetrapeptide through amide bond formation.
[0080] “Improved enzyme property” refers to any property of an enzyme that exhibits an improvement relative to a reference enzyme. For the enzymes described herein, the reference enzyme is a wild-type enzyme or another engineered enzyme. For example, in various embodiments, the reference enzyme is an affinity-tagged variant of a wild-type enzyme that has been codon-optimized (e.g., SEQ ID NO: 3, an E. coli-codon-optimized variant of SEQ ID NO:1 which encodes SEQ ID NO: 4, a 6xHis-tagged variant SEQ ID NO: 2). Enzyme properties for which improvement may be desirable include, but are not limited to, enzymatic activity (which may be expressed in terms of percent conversion of the substrate or total turnover number or product formed over time), thermal stability (or thermostability), stability under high ammonia concentration, soluble expression, reduced by-product generation, higher substrate concentration, pH activity profile (pH tolerance), phosphate tolerance (or higher phosphate loading), cosolvent tolerance, improved activity in an enzymatic cascade, reduced cofactor requirements, refractoriness to inhibitors (e.g., product inhibition), regioselectivity, stereospecificity, and stereoselectivity (including enantioselectivity).
[0081] “Increased enzymatic activity” refers to an improved property of the enzymes that is represented by an increase in specific activity (e.g., product produced / time / weight protein) or an increase in percent conversion of the substrate to the product (e.g., percent conversion of starting amount of substrate to product in a specified time period using a specified amount of enzyme) as compared to a reference enzyme. Exemplary methods to determine enzyme activity are provided in the Examples. Any property relating to enzyme activity may be affected, including the25971 classical enzyme properties of Km, Vmax, or kcat, changes of which can lead to increased enzymatic activity. Improvements in enzyme activity can be from about 1.5 times the enzymatic activity of the corresponding wild-type enzyme, to as much as 2 times.5 times, 10 times, 20 times, 25 times, 50 times, 75 times, 100 times, 150 times, 200 times, 500 times, 1000 times, 3000 times, 5000 times, 7000 times, 10000 times, 15000 times, 20000 times, 22500 times, 25000 times, or more enzymatic activity than the reference enzyme, e.g., a naturally occurring enzyme or another enzyme from which the polypeptides were derived. In some examples, the enzyme exhibits improved enzymatic activity in the range of 100 to 3000 times, 3000 to 7000 times, or more than 7000 times greater than that of the parent enzyme. It is understood by the skilled artisan that the activity of any enzyme is diffusion limited such that the catalytic turnover rate cannot exceed the diffusion rate of the substrate, including any required cofactors. The theoretical maximum of the diffusion limit, or kcat / Km, is generally about 108to 109(M-1s-1). Hence, any improvements in the enzyme activity will have an upper limit related to the diffusion rate of the substrates acted on by the enzyme. Enzyme activity can be measured by any suitable approach, e.g., an enzyme activity assay or by any of the traditional methods for assaying chemical reactions, including but not limited to high-performance liquid chromatography (HPLC), HPLC-mass spectrometry (MS), ultra-performance liquid chromatography (UPLC), UPLC-MS, thin-layer chromatography (TLC), and nuclear magnetic resonance (NMR). Comparisons of enzyme activities may be made using a defined preparation of enzyme, a defined assay under a set condition, and one or more defined substrates, as further described in detail herein. Generally, when lysates are compared, the numbers of cells or the amount of protein assayed are determined as well as use of identical expression systems and identical host cells to minimize variations in amount of enzyme produced by the host cells and present in the lysates.
[0082] As used herein, “by-products” refer to undesired products of a catalytic reaction, such as a catalytic ligation reaction or the macrocyclization reaction catalyzed by the disclosed carboxylesterase enzymes. In some embodiments, this term encompasses the product of undesired ligation reactions, such as the ligation of two molecules of identical amino acid substrate. This term further encompasses undesired regioisomers. Examples of undesired products include compound 3j-regio. In other embodiments, this term encompasses the product of an undesired hydrolysis reaction, such as the hydrolysis of a terminal ester to an alcohol in the macrocyclization reaction. This term further encompasses undesired regioisomers. Examples of undesired by-products include compound 3 shown in Table A.
[0083] As used herein, a “vector” is a DNA construct for introducing a DNA sequence into a cell. In some embodiments, the vector is an expression vector that is operably linked to a suitable25971 control sequence capable of effecting the expression of the polypeptide encoded by the polynucleotide (e.g., DNA) sequence in a suitable host. In some embodiments, an “expression vector” has a promoter sequence operably linked to the polynucleotide (e.g., DNA) sequence (e.g., transgene) to drive expression in a host cell, and in some embodiments, also comprises a transcription terminator sequence.
[0084] As used herein with respect to polypeptides, the terms “expression” and “production” includes any step involved in the production of a polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, and post-translational modification. In some embodiments, the term also encompasses secretion of the polypeptide from a cell.
[0085] As used herein, an amino acid or nucleotide sequence (e.g., a promoter sequence, signal peptide, terminator sequence, and the like) is “heterologous” to another sequence with which it is operably linked if the two sequences are not associated in nature. For example, a “heterologous polynucleotide” is any polynucleotide that is introduced into a host cell by laboratory techniques, and the term includes polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into a host cell.
[0086] As used herein, the terms “host cell” and “host strain” refer to suitable hosts for an expression vector comprising a polynucleotide (e.g., DNA) provided herein (e.g., a polynucleotide encoding a ligase polypeptide disclosed herein). In some embodiments, the host cells are prokaryotic or eukaryotic cells that have been transformed or transfected with vectors constructed using recombinant DNA techniques as known in the art. In some embodiments, suitable host cells are E. coli cells.
[0087] “Coding sequence” refers to that portion of a polynucleotide (e.g., a gene) that encodes an amino acid sequence of a polypeptide.
[0088] “Naturally occurring” or “wild-type” refers to a form found in nature. For example, a naturally occurring or wild-type polypeptide or polynucleotide sequence is a sequence present in an organism that can be isolated from a source in nature and that has not been intentionally modified by human manipulation, with the sole exception that wild-type polypeptide or polynucleotide sequences as identified herein may include a tag, such as a histidine tag (6xHis tag). Herein, “wild-type” polypeptide or polynucleotide sequences may be denoted “WT”.
[0089] “Operably linked” is defined herein as a configuration in which a control sequence is appropriately placed at a position relative to a polynucleotide sequence (i.e., in a functional relationship) such that the control sequence directs the expression of the polynucleotide and / or a polypeptide encoded by the polynucleotide.25971
[0090] A “promoter sequence” is a nucleic acid sequence that is recognized by a host cell for expression of a polynucleotide. The control sequence may comprise an appropriate promoter sequence. The promoter sequence contains transcriptional control sequences, which mediate the expression of the polynucleotide. The promoter may be any nucleic acid sequence which shows transcriptional activity in the host cell of choice including mutant, truncated, and hybrid promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides either homologous or heterologous to the host cell.
[0091] The terms “engineered,” “recombinant,” “variant,” and “non-naturally occurring,” when used with reference to, e.g., a polynucleotide, polypeptide, or cell, refers to a material, or a material corresponding to the natural or native form of the material, that has been modified in a manner that would not otherwise exist in nature. Non-limiting examples include, among others, recombinant cells expressing genes that are not found within the native (non-recombinant) form of the cell or express native genes that are otherwise expressed at a different level.
[0092] As used herein, the terms “percent identity,” and “percent identical” refer to comparisons between polynucleotide sequences or polypeptide sequences, and are determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which either the identical nucleic acid base or amino acid residue occurs in both sequences or a nucleic acid base or amino acid residue is aligned with a gap to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity. Determination of optimal alignment and percent sequence identity can be performed using the BLAST and BLAST 2.0 algorithms (see e.g., Altschul et al., 1990, J. Mol. Biol.215: 403-410; and Altschul et al., 1977, Nucleic Acids Res.3389-3402). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information website.
[0093] Briefly, the BLAST analyses involve first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as, the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequencefor as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) of 10, M = 5, N = -4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, 1989, Proc. Natl. Acad. Sci. Usa 89:10915).
[0094] Numerous other algorithms are available that function similarly to BLAST in providing percent identity for two sequences. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman, 1981, Adv. Appl. Math.2:482, by the homology alignment algorithm of Needleman and Wunsch, 1970, J. Mol. Biol.48:443, by the search for similarity method of Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 85:2444, by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the GCG Wisconsin Software Package), or by visual inspection (see generally, Current Protocols in Molecular Biology, F. M. Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)). Additionally, determination of sequence alignment and percent sequence identity can employ the BESTFIT or GAP programs in the GCG Wisconsin Software package (Accelrys, Madison WI), using default parameters provided.
[0095] As used herein, the terms “stereoselectivity” and “stereospecificity” refer to the preferential formation in a chemical or enzymatic reaction of one stereoisomer over another. Stereoselectivity can be partial, where the formation of one stereoisomer is favored over the other, or it may be complete where only one stereoisomer is formed. When the stereoisomers are enantiomers, the stereoselectivity is referred to as enantioselectivity, the fraction (typically reported as a percentage) of one enantiomer in the sum of both. It is commonly alternatively reported in the art (typically as a percentage) as the enantiomeric excess (EE) calculated therefrom according to the formula [major enantiomer – minor enantiomer] / [major enantiomer + minor enantiomer]. Where the stereoisomers are diastereomers, the stereoselectivity is referred toas diastereoselectivity, the fraction (typically reported as a percentage) of one diastereomer in a mixture of two diastereomers, commonly alternatively reported as the diastereomeric excess (DE). Enantiomeric excess and diastereomeric excess are types of stereomeric excess.
[0096] “Regioselectivity” refers to the preferential formation in a chemical or enzymatic reaction of one regioisomer over another. Regioselectivity can be partial, where the formation of one regioisomer is favored over the other, or it may be complete where only one regioisomer is formed. “Highly regioselective” refers to a chemical or enzymatic reaction that is capable of converting a substrate to its corresponding product with at least about 85% regioisomeric excess. “Regioisomers” encompass diastereomers and enantiomers. Regioselectivity may result when an enzyme demonstrates preferential catalysis of a chemical moiety at a single position or configuration in a substrate molecule than at other positions.
[0097] “Chemoselectivity” refers to the preferential formation in a chemical or enzymatic reaction of one product over another.
[0098] “Conversion” refers to the enzymatic transformation of a substrate to the corresponding product. “Percent conversion” refers to the percent of the substrate that is converted to the product within a period of time under specified conditions. Thus, for example, the “enzymatic activity” or “activity” of a polypeptide can be expressed as “percent conversion” of the substrate to the product.
[0099] As used herein, the terms “thermostability” and “thermal stability” refer to the ability of an enzyme to maintain the same or similar level of activity in catalyzing a reaction (more than 60% to 80% of product conversion, for example) after exposure to elevated temperatures (e.g., 37 °C to 50 °C, or 37 °C to 80 °C) for a period of time (e.g., 0.5 h to 24 h) relative to the corresponding enzyme at room temperature. Thermostability may be measured using a heat challenge test, e.g., measuring enzyme activity for about 60 minutes after exposure to elevated temperatures (e.g., 35 °C, 37 °C, or 39-40°C).
[0100] As used herein, the terms “biocatalysis,” “biocatalytic,” “biotransformation,” and “biosynthesis” refer to the use of enzymes to perform chemical reactions on organic compounds.
[0101] As used herein, “phosphate tolerance” refers to the capability of an enzyme to exhibit activity under a range of phosphate molecule concentrations that are relevant to the process, e.g., a high phosphate salt concentration.
[0102] As used herein, “pH tolerance” refers to the capability of an enzyme to exhibit activity under a range of pH values during a reaction. It may refer to a maintained activity in the presence of a pH reduced in the presence of inorganic phosphate molecules. “pH stable” refers to a polypeptide that maintains similar activity (more than e.g., 60% to 80%) after exposure to high or25971 low pH (e.g., 4.5 to 6 or 8 to 12) for a period of time (e.g., 0.5h to 24h) relative to an untreated enzyme.
[0103] As used herein, “polyphosphate” refers to an oligomer of two or more phosphate molecules (e.g., ions) covalently linked together.
[0104] As used herein, the terms “co-solvent tolerance” and “solvent tolerance” refer to the capability of an enzyme to exhibit activity under a range of co-solvent concentrations, such as organic co-solvents, that are relevant to the chemical process occurring in an aqueous medium (e.g., an aqueous buffer such as HEPES or sodium phosphate buffer). Exemplary organic co- solvents are acetonitrile, acetone, dimethyl sulfoxide, hexane and heptane. In particular embodiments, the organic co-solvent is acetonitrile (ACN). A co-solvent may or may not be present with a primary aqueous solvent, in any desired reaction medium. Enzyme co-solvent tolerance can be measured by measuring the activity of the enzyme under low co-solvent concentrations (e.g., acetonitrile absent) relative to activity under higher co-solvent concentrations (e.g., >15% or >20% v / v acetonitrile); a co-solvent tolerant enzyme displays minimal difference in activity between low and higher co-solvent concentrations.
[0105] As used herein, the terms “substrate loading”, “substrate loads” and “substrate concentration” refer to the concentration of substrate present in a reaction vessel following a contacting of substrate to enzyme, e.g., during a synthesis reaction. Many natural enzymes exhibit little to no activity at high substrate concentration. As used herein, the term “high substrate loading” refers to a concentration of substrate between 5 and 75 mM. Example substrate concentration ranges for “high substrate loading” include 1-10 mM, 5-15 mM, 10-20 mM, 20-30 mM, 30-40 mM, 5-45 mM, 15-45 mM, 25-45 mM, 40-50 mM, 50-60 mM, 60-75 mM, 5-75 mM, 15-75 mM, 25-75 mM, or 50-75 mM. “Low substrate loading” may refer to concentrations of less than 5 mM.
[0106] As used herein, the term “product inhibition” refers to a reduction in enzymatic activity observed when a product molecule resides too long in the active site of an enzyme following its synthesis, causing it to stall, which prevents the enzyme from proceeding to diffuse toward and contact a new substrate molecule. Product inhibition may be observed at high substrate loads. As used herein, “reduced product inhibition” refers to a restoration in enzymatic activity following an interference with the active site residence time of product molecules. A reduced product inhibition may be a selection pressure applied in a directed evolution screen of an enzyme exhibiting product inhibition. This term may be synonymous with “refractoriness to inhibitors” and “reduced competitive inhibition”. As used herein, in the context of a one-pot cascade, this term encompasses a reduction in enzymatic activity observed when a substrate or product of a25971 parallel reaction resides in the active site or another binding site of an enzyme that catalyzes a different reaction, preventing said enzyme from performing catalysis.
[0107] As used herein, the term “cascade” or “enzyme cascade” or “enzymatic cascade” or “multistep cascade” refers to the use of multiple enzymes, such as two or more wild-type or engineered enzymes that each catalyze a unique bond forming or bond breaking chemical transformation, to catalyze the steps of a multi-step reaction and or recycling of cofactors or cosubstrates in a one-pot process or through process, simultaneously or sequentially.
[0108] The phrase "pharmaceutically acceptable" is employed herein to refer to those compounds, materials, compositions, and / or dosage forms which are, within the scope of sound medical judgment, suitable for use in contact with the tissues of human beings and animals without excessive toxicity, irritation, allergic response, or other problem or complication, commensurate with a reasonable benefit / risk ratio.
[0109] The salts of the compounds of disclosure may be pharmaceutically acceptable salts or non-pharmaceutical salts useful in the preparation of the compounds according to the disclosure.
[0110] As used herein, "pharmaceutically acceptable salts" refer to derivatives wherein the parent compound is modified by making acid or base salts thereof. Salts in the solid form may exist in more than one crystal structure and may also be in the form of hydrates. Examples of pharmaceutically acceptable salts include, but are not limited to, mineral or organic acid salts of basic residues such as amines; alkali or organic salts of acidic residues such as carboxylic acids; and the like. The pharmaceutically acceptable salts include the conventional non-toxic salts or the quaternary ammonium salts of the parent compound formed, for example, from non-toxic inorganic or organic acids. For example, such conventional non-toxic salts include those derived from inorganic acids such as formic, hydrochloric, hydrobromic, sulfuric, sulfamic, phosphoric, nitric and the like; and the salts prepared from organic acids such as acetic, propionic, succinic, glycolic, stearic, lactic, malic, tartaric, citric, ascorbic, pamoic, maleic, hydroxymaleic, phenylacetic, glutamic, benzoic, salicylic, sulfanilic, 2-acetoxybenzoic, fumaric, toluenesulfonic, methanesulfonic, ethane disulfonic, oxalic, isethionic, and the like. Salts derived from inorganic bases include aluminum, ammonium, calcium, copper, ferric, ferrous, lithium, magnesium, manganic salts, manganous, potassium, sodium, zinc, and the like.
[0111] When the compound of the present disclosure is basic, salts may be prepared from pharmaceutically acceptable non-toxic acids, including inorganic and organic acids. Such acids include acetic, benzenesulfonic, benzoic, camphorsulfonic, citric, ethanesulfonic, fumaric, gluconic, glutamic, hydrobromic, hydrochloric, isethionic, lactic, maleic, malic, mandelic, methanesulfonic, mucic, nitric, pamoic, pantothenic, phosphoric, succinic, sulfuric, tartaric, p-25971 toluenesulfonic acid, and the like. In one aspect of the disclosure the salts are citric, hydrobromic, hydrochloric, maleic, phosphoric, sulfuric, fumaric, and tartaric acids. Similarly, the salts of the acidic compounds are formed by reactions with the appropriate inorganic or organic base.
[0112] An embodiment of this disclosure is realized when the compounds and / or intermediates are salts are selected from oxalate, hydrochloride, hydrobromide, tosylate, acetic acid, mesylate, glycolate, oxalate, 4-hydroxybenzoate, and di-tolyl tartrate salts. As aspect of this embodiment is realized when the salt is oxalate, hydrochloride or hydrobromide. Another aspect of this embodiment is realized when Compound 1 is crystallized as bisoxalate salt.
[0113] The present disclosure provides processes using numerous exemplary ATP-dependent amino acid ligases (Trp-ligases) to generate the tripeptides shown in Table 1. Those exemplary polypeptides were evolved from SEQ ID NO: 4 and exhibit improved properties, particularly improved activity in the conversion of various amino acid nucleophiles and amino acid electrophiles, including the conversion of compounds 1a and 2a to tripeptide compound 3a, the conversion of compound 1b and 2b to tripeptide compound 3b, the conversion of compound 1c and 2b to tripeptide compound 3c and compound 1d and 2c to tripeptide compound 3d, as depicted in Table 1, Reactions A-D. These exemplary engineered ATP-dependent Trp-ligase enzymes having tripeptide formation activity have amino acid sequences that include one or more residue differences as compared to SEQ ID NO: 4 as depicted in the accompanying sequence listing with sequence identifiers SEQ ID NOs: 6-50. In various embodiments, these variants comprise amino acid sequences having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to SEQ ID NO: 4. In various embodiments, these variants comprise amino acid sequences having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 97%, 98%, or 99% identity to any one of SEQ ID NOs: 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, or 50. In some embodiments, these variants comprise amino acid sequences comprising any one of SEQ ID NOs: 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, or 50.
[0114] In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 92.5%, 93%, 94%, 95%, 96%, 97%, 97.5%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 40, 42, 44, 46, 48, and 50.
[0115] In some embodiments, use of the engineered ATP-dependent amino acid Trp-ligase enzymes of the disclosure provides improved regioselectivity in the production of tetrapeptide from tripeptide compound 3d relative to the wild-type Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1) SEQ ID NO: 4. In some25971 embodiments, the engineered ATP-dependent amino acid ligases catalyzing the formation of products of Formula VI^ depicted in Table 2 can also form a regioisomer of product 3d. In particular embodiments, the Trp-ligase enzymes catalyzing the formation of products of Formula VI can form the desired product 3d with regioisomeric ratios of at least 1:10, 10:1, 50:1 and more than 100:1, relative to the ligase SEQ ID NO: 4. The disclosed Trp-ligases may exhibit regioselectivity for compound 3d over undesired by-products such as dimers, trimers and other oligomers of compounds according to Formula IV^ (i.e., oligomers of compound 1e or compound 1f) and a by-product of ligation of compound 3d and compound 1d. As such, the disclosed Phe-ligases may have activity that generates oligomers of compound 1e or compound 1f and / or a product of ligation of compound 3d and compound 1d in reduced amounts relative to the ligase of SEQ ID NO: 4.
[0116] In some embodiments, the engineered ATP-dependent Trp ligases catalyzing the formation of products of Formula III^ may also catalyze the formation of by-products such as dimers, trimers and oligomers of substrates with Formula I’ or Compound 1. Any of the disclosed engineered polypeptides (ligases) may generate a ratio of product over by-products of at least about 500:1, 1000:1, 1500:1, 2000:1, 2500:1, or more than 2500:1. In some embodiments, these enzymes generate by-products at a level of less than 0.5%, thus exhibiting a desired ratio of product over by-products of over 2000:1. Any of the disclosed engineered polypeptides (ligases) may generate fewer by-products than the ligase of SEQ ID NO: 4. In particular embodiments, the disclosed engineered ligases have a desired ratio of product over by- products of at least about 1:10, 10:1, 30:1, 50:1, 90:1 or more, relative to SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent amino acid ligase enzymes catalyzing the formation of products of Formula III^ can form the desired product 3d over dimers of 1d with a selectivity of at least 1:1, 1:10, 10:1, 30:1, 50:1, 90:1 or more relative to SEQ ID NO: 4.
[0117] In some embodiments use of the engineered ATP-dependent ligases depicted in Table 1 are able to generate products of Formula III^ at substrate loading concentrations of at least about 10 mM, 25 mM, 50 mM or more than 50 mM with a percent conversion of at least about 0.001%, about 0.01%, about 0.1%, about 1%, about 10%, about 40%, about 60%, about 80% or at least about 95%. In some embodiments the engineered ATP-dependent Trp-ligases are able to generate these products at substrate loading concentrations of about 10-20 mM, 20-30 mM, 30- 40 mM, 25-45 mM, 40-50 mM, 50-60 mM, or 50-75 mM. In some embodiments, the disclosed ligase enzymes exhibit activities or efficiencies in the presence of higher substrate loads in an enzymatic cascade relative to the corresponding activities of SEQ ID NO: 4.25971
[0118] In various embodiments, the disclosed engineered Trp-ligases exhibit improved kinetics of the reaction, i.e., reduced time of reaction, relative to the Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1, SEQ ID NO: 2) enzyme of SEQ ID NO: 4. In some embodiments, these enzymes are able to generate a conversion of product of at least about 80% or at least about 90% in a reaction time less than about 96 hours, about 48 hours, about 20 hours, about 6 hours or less than 6 hours. In some embodiments, these enzymes are able to generate a conversion of product of at least about 80% or at least about 90% in a reaction time less than about 20 hours, about 6 hours, or less than 6 hours at substrate loading concentrations of at least about 10 mM, 25 mM, 50 mM or more than 50 mM in a reaction time less than about 96 hours, about 48 hours, about 20 hours, about 10 hours or less.
[0119] In some embodiments, use of the Trp-ligase enzymes of the disclosure provides improved thermostability relative to SEQ ID NO: 4. The ATP-dependent amino acid ligase enzymes of the disclosure may exhibit higher activity at low temperatures relative to SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent amino acid ligase enzymes catalyzing the formation of products of Formula VI^ depicted in Table 2 have been improved for thermostability and can maintain at least 50% of its enzymatic activity at temperatures of at least about 35 °C, 37°C, 39°C, 42°C, 43°C, 50°C or more degrees compared to SEQ ID NO: 4 when subjected to a heat challenge test for about 60 minutes.
[0120] In some embodiments, use of the engineered ATP-dependent amino acid ligases catalyzing the formation of products of Formula III^ depicted in Table 1 provides improved phosphate tolerance of at least about 1.2-fold, 2-fold or 3-fold or more, relative to the ligase of SEQ ID NO: 4.25971 Table 1. Trp-ligase catalyzed ATP-dependent amino acid ligation reactions for tripeptide formation Substrate Reaction Compound of Substrate Compound Product Compound of ^ ^ ^25971 Substrate Reaction Compound of Substrate Compound Product Compound of ^ ^ ^ety and may be referred to herein as “aFTrp”, “FTrp” or “Linker Trp”. Compound 1d has a molecular weight of 321.40. Compound 2c is a Proline-Threonine derivative dipeptide that contains a modified proline residue and modified threonine residue, having a molecular weight of 388.46. Compound 3d is a modified Tryptophan-Proline-Threonine tripeptide that contains modified proline, threonine, and tryptophan residues. This tripeptide retains the 6-aminohexyl tail or “linker”.25971
[0122] In some embodiments, provided herein are ATP-dependent ligase enzymes that are capable of activating phenylalanine analog carboxylic acids of Formula IV^ via reaction with ATP to form acylphosphate intermediates for nucleophilic attack by tryptophan-proline-threonine analog nucleophiles of Formula III to produce phenylalanine-tryptophan-proline-threonine analog tetrapeptides of Formula V^as depicted in Table 2. In particular embodiments, SEQ ID NO: 4 showed no detectable activity for substrates illustrated in Reactions F, G, H, I, J and K in Table 2 below, and therefore it was determined that protein engineering of the ligase was necessary.
[0123] In some embodiments, the ATP-dependent amino acid phe-ligase enzymes of the disclosure may exhibit improvements relative to the ATP-dependent ligase enzyme of SEQ ID NO: 4, such as increases in enzyme activity, regioselectivity, thermostability, phosphate tolerance, organic co-solvent tolerance, reduction in by-product formation, and reduction in product inhibition for tetrapeptide formation in the reactions shown in Table 2. These improvements can relate to a single enzyme property, such as enzymatic activity, or a combination of different enzyme properties, such as enzymatic activity and thermostability.
[0124] Any of the below-described properties may be measured in accordance with techniques known in the art. Any of these properties may be measured following performance of the reaction in an aqueous medium, such as HEPES buffer, in the presence or absence of cosolvent. Any of these properties may be measured following performance of the reaction in a medium containing reagents ATP, inorganic magnesium salt, and / or sodium polyphosphate. Any of these properties may be measured following performance of the reaction in a medium containing sodium polyphosphate and a polyphosphate kinase (PPK), such as wild-type PPK enzymes published in literature and those found in public databases (e.g., Uniprot ID: A0A3D5XRJ5_9FIRM, Genbank ID: CCX64920.1), including e.g., PPK22 (Uniprot accession ID: R5CAF0) and PPK12 (Uniprot accession ID: A0A3D5XRJ5). Suitable reaction conditions under which the above-stated improved properties of the engineered ATP-dependent amino acid ligases are further described in Examples 4-8.
[0125] The present disclosure provides numerous exemplary ATP-dependent amino acid Phe- ligase enzymes capable of generating tetrapeptides shown in Table 2. Those exemplary polypeptides were engineered from SEQ ID NO: 4 and exhibit improved properties, particularly improved activity in the conversion of various amino acid nucleophiles and amino acid electrophiles, including the conversion of compounds 4e and 3a to the tetrapeptide compound 5e, conversion of compounds 4f and 3a to the tetrapeptide compound 5f, conversion of compounds 4f and 3e to tetrapeptide compound 5g, conversion of compounds 4f and 3b to tetrapeptide25971 compound 5h, conversion of compounds 4f and 3f to tetrapeptide compound 5i, conversion of compounds 4f and 3c to tetrapeptide compound 5j, and compounds 4f and 3d to tetrapeptide compound 5k as depicted in Table 2, Entries E-K. These exemplary engineered ATP-dependent amino acid ligase enzymes having tetrapeptide formation activity comprise amino acid sequences that include one or more residue differences as compared to SEQ ID NO: 4 as depicted in the accompanying sequence listing with sequence identifiers SEQ ID NOs: 52-116 and 126-130.
[0126] In various embodiments, these variants comprise amino acid sequences having at least about 80%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 92.5%, 93%, 94%, 95%, 96%, 97%, 97.5%, 98% or 99% sequence identity to SEQ ID NO: 4. In various embodiments, these variants comprise amino acid sequences having at least about 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% identity to any one of SEQ ID NOs: 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, and 116. In some embodiments, these variants comprise amino acid sequences comprising any one of SEQ ID NOs: 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 88, 90, 92, 94, 96, 98, 100, 102, 104, 106, 108, 110, 112, 114, and 116.
[0127] In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least about 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 108, 110, 112, 114, and 116. In some embodiments are provided engineered polypeptides comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 108, 110, and 112. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 108. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 110. In some embodiments, the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 112. In some embodiments, the enzyme polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 114. In some embodiments, the enzyme polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 116.
[0128] In some embodiments, the disclosed engineered Phe-ligases and / or Trp-ligases are able to generate a percent conversion of substrate to product (e.g., compound 2c into compound 3d or converting compound 3d into compound 3k) of at least about 0.01%, about 0.1%, at least about 1%, at least about 10%, at least about 40%, at least about 60%, at least about 80% or at least about 90%. In some embodiments, the disclosed ligases can generate a percent conversion of at least about 80%. %. In some embodiments, the disclosed ligases can generate a percent conversion of at least about 90%.25971
[0129] In some embodiments, the engineered ATP-dependent Phe-ligase enzymes catalyze the formation of products depicted in Formula VI^ (Table 2) exhibit at least about 5-fold, 10-fold, 100-fold, 1,000, 10,000-fold or more the activity compared to SEQ ID NO: 4, under suitable reaction conditions.
[0130] In some embodiments, the engineered Phe-ligases of the disclosure exhibit improved regioselectivity in the production of tetrapeptide compound 3f relative to the wild-type Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1) SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent Phe-ligases catalyzing the formation of products of Formula VI^ depicted in Table 2 can also form a regioisomer of product 3k. In particular embodiments, the engineered ATP-dependent amino acid ligase enzymes catalyzing the formation of products of Formula VI^ can form the desired product 3k with regioisomeric ratios of at least 1:10, 10:1, 50:1 and more than 100:1, relative to the ligase SEQ ID NO: 4.
[0131] In some embodiments, use of the Phe-ligases of the disclosure provides improved generation of undesired by-products (e.g., dimers of the tripeptide, dimers of the tetrapeptides, or dimers and trimers of compound 3d); improved thermostability relative, and improved phosphate tolerance relative to SEQ ID NO: 4.
[0132] In some embodiments, the engineered Phe-ligases catalyzing the formation of products of Formula VI^ can form by-products of substrate 1d such as dimers when present in solution. In particular embodiments, the engineered ATP-dependent amino acid ligase enzymes catalyzing the formation of products of Formula VI^ can form the desired product 3k over dimers of 1d with a selectivity of at least 1:5, 1:1, 10:1, 25:1, 50:1, or more than 100:1 relative to SEQ ID NO: 4.
[0133] In some embodiments the Phe-ligases catalyzing the formation of products of Formula VI^ may generate a by-product from the ligation of substrate 1d to 3k when present in solution.
[0134] In some embodiments, the Phe-ligase enzymes of the disclosure exhibit activities or efficiencies in the presence of higher substrate loads relative to the corresponding activity of SEQ ID NO: 4. In some embodiments, the engineered ATP-dependent amino acid ligases catalyzing the formation of products of Formula VI^ are able to catalyze the production (i.e., ligation) of tetrapeptide products of Formula VI^ (e.g., tetrapeptide compound 3k) at substrate loading concentrations of at least about 1 mM, 5 mM, 10 mM, 25 mM, 50 mM, 50 mM to 75 mM, or more than 75 mM. In some embodiments, these engineered ligases are able to catalyze the production of these tetrapeptide products at substrate loading concentrations of 1-10 mM, 5-1525971 mM, 10-20 mM, 20-30 mM, 30-40 mM, 5-45 mM, 15-45 mM, 25-45 mM, 40-50 mM, 50-60 mM, 60-75 mM, 5-75 mM, 15-75 mM, 25-75 mM, or 50-75 mM.
[0135] In some embodiments, use of the disclosed Trp-ligase and Phe-ligase enzymes of the disclosure provide improved tolerance of organic co-solvents, pH tolerance, improved activity in the presence of increased phosphate loads, improved activity in the presence of increased phosphate loads, and / or improvements (i.e., reductions) in product inhibition, for instance at high substrate loads, relative to SEQ ID NO: 4.
[0136] In some embodiments, any of the disclosed Trp-ligase and Phe-ligase polypeptides further comprises a tag (e.g., an affinity tag). Any suitable tag may be used, e.g., a 6xHis tag or an 8xHis tag, a FLAG tag, a fluorescent protein tag (e.g., green fluorescent protein (GFP), yellow fluorescent protein (YFP), or red fluorescent protein (RFP)), a hemagglutinin (HA) tag, an ALFA-tag, a V5-tag, a Myc-tag, a SPOT-tag, a T7-tag, or an NE-tag. In some embodiments, the affinity tag is a His tag. In some embodiments, the affinity tag comprises the amino acid sequence of MHHHHHHGS (SEQ ID NO: 117). In some embodiments, the affinity tag comprises the amino acid sequence of GSHHHHHHHHSG (SEQ ID NO: 118) (an 8xHis tag). In some embodiments, the polypeptide comprises an N-terminal methionine residue, and the epitope tag is inserted immediately following the N-terminal methionine residue, e.g., relative to a reference sequence.
[0137] In some embodiments, the disclosed Trp-ligase and Phe-ligase polypeptides do not contain an affinity tag. In some embodiments, the disclosed Trp-ligase and / or Phe-ligase polypeptides do not contain a histidine tag, such as an 8xHis tag.25971 Table 2. Phe-ligase Catalyzed ATP-dependent amino acid ligation reactions for tetrapeptide formation Reacti on ID Substrate Compound Substrate Compound Product Compound ^ ^ ^25971 Reacti on ID Substrate Compound Substrate Compound Product Compound ^ ^ ^25971 Reacti on ID Substrate Compound Substrate Compound Product Compound ^ ^ ^propanoic acid or aminomethylphenylalanine) that may be referred to herein as “aPhe” (molecular weight of 194.23). Compound 3d is the product of the above-described reaction D shown in Table 1, i.e., a Tryptophan-Proline-Threonine derivative tripeptide that contains modified proline, threonine, and tryptophan residues. Compound 3k is a modified Phenylalanine- Tryptophan-Proline-Threonine tetrapeptide that contains modified proline, tryptophan, and phenylalanine residues. This tetrapeptide has a molecular weight of 793.94.
[0139] In some embodiments compound 3k can be generated from a cascade reaction by the addition of compound 1f, compound 1d and compound 2c with the corresponding ATP- dependent amino acid ligase enzymes that form tripeptides (Trp-ligase from any one of SEQ ID NOs: 24-50 and the corresponding ATP-dependent amino acid ligase that form tetrapeptides (Phe-ligase from any one of SEQ ID NOs: 76-116) as shown in Formula VIII, below.25971
[0140] In some embodiments, the ATP-dependent amino acid ligases described herein demand equimolar concentrations of ATP or the use of an ATP regeneration system via activated phosphate sources for the formation of tripeptides and tetrapeptides. Exemplary ATP regeneration systems are well known to those skilled in the art and can comprise the addition of sub-stoichiometric concentrations of adenosine diphosphate or adenosine triphosphate in the presence of propionyl phosphate and an acetate kinase. In some embodiments the ATP recycling system can be comprised of a wild-type polyphosphate kinase (PPK), polyphosphate, and a cofactor such as adenosine monophosphate, adenosine diphosphate or adenosine triphosphate.
[0141] In some embodiments, the wild-type polyphosphate kinase is PPK12 (Uniprot accession ID: A0A3D5XRJ5, See Tavanti et al., Green Chem. (2021) 23, 828-837, which is herein incorporated by reference.) In some embodiments, the wild-type polyphosphate kinase is PPK22 (Uniprot accession ID: R5CAF0). In embodiments, provided are stable preparations of any of the disclosed engineered Trp ligase and Phe ligase enzymes. In some embodiments, these preparations comprise immobilized enzymes. Immobilized enzyme preparations have a number of recognized advantages. They can confer shelf life to enzyme preparations, they can improve reaction stability, they can enable stability in organic solvents, they can aid in protein removal from reaction streams, as examples. “Stable” refers to the ability of the immobilized enzymes to retain their structural conformation and / or their activity in a solvent system that contains organic solvents. “Solvent stable” refers to a polypeptide that maintains similar activity (more than e.g., 60% to 80%) after exposure to varying concentrations (e.g., 5% to 99%) of solvent (isopropyl alcohol, tetrahydrofuran, 2-methyltetrahydrofuran, acetone, toluene, butylacetate, methyl tert- butylether, etc.) for a period of time (e.g., 0.5h to 24h) compared to an untreated enzyme. Stable (e.g., solvent stable) immobilized enzymes lose less than 10% activity per hour in a solvent system that contains organic solvents. Stable immobilized enzymes may lose less than 9%, 8%, 7%, 6%, 5%, activity per hour in a solvent system that contains organic solvents. The disclosed preparations may comprise one or more diluents or carriers.
[0142] In some embodiments, the disclosed enzyme preparations are thermostable. In some embodiments, the disclosed enzyme preparations are solvent-stable. In some embodiments, the disclosed enzyme preparations are thermostable and solvent stable. In some embodiments, the disclosed enzyme preparations are pH tolerant and / or phosphate tolerant. In some embodiments, the disclosed enzyme preparations are thermostable and solvent stable, pH tolerant and phosphate tolerant.
[0143] Tables 3, 4 and 5 below provide exemplary engineered ATP-dependent ligase polypeptides with the ability to form tripeptides (Table 3) and tetrapeptides (Tables 4 and 5).25971 Each row lists two SEQ ID NOs, with the odd number referring to the nucleotide (“nt”) sequence encoding the amino acid (“aa”) sequence provided by the even number. The residue differences are based on comparison to a reference sequence such as of SEQ ID NO: 4, an ATP-dependent amino acid ligase derived from Bifidobacterium adolescentis that differs from the naturally occurring enzyme Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1) SEQ ID NO: 2 in that it contains a short hexahistidine (6xHis) tag at the N-terminus and in that its encoding gene is codon optimized for E. coli expression. The column listing the number of mutations (i.e., residue changes) refers to the number of amino acid substitutions as compared to the ATP-dependent amino acid ligase enzyme of SEQ ID NO: 4.
[0144] Tables 3, 4 and 5 below provide a list of the SEQ ID NOs disclosed herein with associated absolute conversions with shake flask powder as described in the examples. Conversions were defined as follows: “+” indicates less than 10% conversion of substrate to product observed for the desired reactions (listed as Reaction A through Reaction J), “++” indicates between 10 and 60% conversion and “+++” indicates more than 60% conversion. Empty fields indicate substrates were not tested under the desired conditions. Table 3. Activity of ATP-dependent amino acid ligases for tripeptide formation Number of SEQ ID amino acid Reaction Reaction C Reaction D -25971 Number of SEQ ID amino acid Reaction Ami Reaction C Reaction D NO no acid differences relative to changes A Reaction B Cnvr Cnvrs-25971 Number of SEQ ID amino acid Reaction Ami Reaction C Reaction D NO no acid differences relative to changes A Reaction B Cnvr Cnvrs-25971 Number of SEQ ID amino acid Reaction Ami Reaction C Reaction D NO no acid differences relative to changes A Reaction B Cnvr Cnvrs-25971 Number of SEQ ID amino acid Reaction Amino acid d Reaction C Reaction D NO ifferences relative to changes A Reaction B Cnvr Cnvrs-I) Number of I -25971 Table 5. Activity of ATP-dependent Phe-ligases for tetrapeptide formation (Reactions I-K) Number of SEQ ID amino acid Amino acid dif Reaction I Reaction J NO ferences relative to SEQ ID changes C C Reaction K on25971 Number of SEQ ID amino acid Amino acid Reaction I Reaction J NO differences relative to SEQ ID changes Cnvr Cnvr Reaction K on25971 Number of SEQ ID amino acid Amino acid Reaction I Reaction J NO differences relative to SEQ ID changes Cnvr Cnvr Reaction K on25971 Number of SEQ ID amino acid Amino acid Reaction I Reaction J NO differences relative to SEQ ID changes Cnvr Cnvr Reaction K on25971 Number of SEQ ID amino acid Amino acid differences relative to SEQ ID chang Reaction I Reaction J NO es C nv r C nv r Reaction K onPolynucleotides Encoding ATP-dependent amino acid ligases
[0145] In various embodiments, engineered Phe-ligase polynucleotides that comprise nucleic acid sequences having at least 75%, 80%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 92.5%, 93%, 94%, 95%, 96%, 97%, 97.5%, 98% or 99% sequence identity to SEQ ID NO: 3 are used in the process. In various embodiments, these variants comprise nucleic acid sequences having at least 75%, 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% identity to any one of SEQ ID NOs: 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, and 115. In some embodiments, these variants25971 comprise nucleic acid sequences comprising any one of SEQ ID NOs: 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, and 115. In some examples, the polynucleotide sequence of any of SEQ ID NOs: 53, 55, 57, 59, 61, 63, 65, 67, 69, 71, 73, 75, 77, 79, 81, 83, 85, 87, 89, 91, 93, 95, 97, 99, 101, 103, 105, 107, 109, 111, 113, and 115 is modified to no longer code for an N-terminal 6xHis tag having the amino acid sequence of SEQ ID NO: 117.
[0146] In some embodiments are provided engineered Trp-ligase polynucleotides comprising a nucleic acid sequence having at least 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 39, 41, 43, 45, 47, and 49.
[0147] In some embodiments, Trp-ligase polynucleotides comprising any of SEQ ID NOs: 39, 41, 43, 45, 47, and 49 are used. In some aspects, provided are engineered polynucleotides comprising an nucleic acid sequence that comprises a stretch of at least 100, 150, 200, 250, 300, 325, 350, or 375 consecutive nucleotides of any one of SEQ ID NOs 39, 41, 43, 45, 47, and 49. In some aspects, provided are engineered Trp-ligase polynucleotides comprising an nucleic acid sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-25, 25-35, 35-50, 50-75, or more than 75 nucleotides relative to the sequence of any one of SEQ ID NOs 39, 41, 43, 45, 47, and 49. The polynucleotides may comprise an nucleic acid sequence that differs by 1, 2, 3, 4, or 5 nucleic acids relative any one of SEQ ID NOs 39, 41, and 43.
[0148] In some embodiments engineered Phe-ligase polynucleotides comprising a nucleic acid sequence having at least 85%, 90%, 92.5%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 107, 109, 111, 113, and 115 are used.
[0149] In some aspect, provided are engineered Phe-ligase polynucleotides comprising an nucleic acid sequence that comprises a stretch of at least 100, 150, 200, 250, 300, 325, 350, or 375 consecutive nucleotides of any one of SEQ ID NOs 107, 109, 111, 113, and 115. In some aspect, provided are engineered Phe-ligase polynucleotides comprising an nucleic acid sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-25, 25-35, 35-50, 50-75, or more than 75 nucleotides relative to the sequence of any one of SEQ ID NOs 107, 109, 111, 113, and 115. The polynucleotides may comprise a nucleic acid sequence that differs by 1, 2, 3, 4, or 5 nucleic acids relative any one of SEQ ID NOs 107, 109, and 111. Methods of Producing ATP dependent ligases
[0150] In some embodiments, the ATP-dependent ligases of the present disclosure are obtained, evolved, or derived from a bacterial wild-type enzyme. Protein engineering (e.g., protein design, semi-rational engineering, structure-guided engineering, directed evolution) may be used to identify polypeptides of the present disclosure. For example, in some embodiments, to25971 make a ligase polypeptide of the present disclosure a polynucleotide sequence encoding a ligase polypeptide may be obtained (or derived) from the genome of Bifidobacterium adolescentis. In some embodiments, the parent polynucleotide sequence is codon-optimized to enhance expression of the ATP-dependent amino acid ligase in a specified host cell. The parental polynucleotide sequence, designated as SEQ ID NO: 1, was codon optimized for E. coli expression to afford SEQ ID NO: 3. The codon optimized ATP-dependent amino acid ligase was cloned into an expression vector, placing the expression of the ATP-dependent amino acid ligase gene under the control of the lac promoter under control of the lac repressor. Clones expressing the active ATP-dependent amino acid ligase in E. coli were identified, and the genes sequenced to confirm their identity.
[0151] The ATP-dependent amino acid ligase of the disclosure may be obtained by subjecting the polynucleotide encoding the parent sequence to mutagenesis and / or directed evolution (DE) methods. Any DE technique may be used, including crystal structure-guided library design, single-site-saturation mutagenesis (SSM) library generation, or combinatorial library generation, or a combination of these. An exemplary directed evolution technique is mutagenesis and / or DNA shuffling as described in Stemmer, 1994, Proc. Natl. Acad. Sci. USA 91:10747-10751; WO 95 / 22625; WO 97 / 20078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767 and U.S. Pat. No.6,537,746. Other directed evolution procedures that can be used include, among others, staggered extension process (StEP), in vitro recombination (Zhao et al., 1998, Nat. Biotechnol. 16:258-261), mutagenic PCR (Caldwell et al., 1994, PCR METHODS APPL.3:S136-S140), and cassette mutagenesis (Black et al., 1996, Proc. Natl. Acad. Sci. USA 93:3525-3529). See also Attorney Docket No.25869 (Engineered ATP-Dependent Ligases for the Synthesis of Oligopeptides) and Attorney Docket No.25920 (Engineered Carboxylesterase Enzymes And Methods For Their Use In Macrocyclization Of Non-Canonical Tetrapeptides), both filed contemporaneously with the instant disclosure and both incorporated by reference in its entirety.
[0152] Exemplary crystal structure-guided library design techniques may make use of computational modeling and / or may be rational or semi-rational. Examples of such techniques include molecular dynamics (MD) simulations, including in silico MD, Rosetta de novo design (e.g., trRosetta), and Glide modeling. TrRosetta is a deep neural net-based protein structure prediction method, offers a path forward for predicting 3D structures when there is a lack of close homologue structures, which takes advantage of multiple sequence alignment (MSA), where derived features from the MSA are fed into a deep neural network to predict inter-residue geometries, including distance and orientations.25971
[0153] The clones obtained following mutagenesis treatment are screened for amino acid ligase activity having a desired improved enzyme property. Measuring enzyme activity from the expression libraries can be performed using standard chemistry analytical techniques for measuring substrates and products such as UPLC-MS. In this reaction, amino acid residues such as the ones described in Tables 1 and 2 can be ligated into tripeptides and tetrapeptides in the presence of an ATP-dependent amino acid ligases and stochiometric amounts of ATP. In some embodiments, an ATP regeneration system consisting of a polyphosphate kinase a phosphate donor such as polyphosphate and an adenosine cofactor such as AMP, ADP or ATP can be employed in ATP-dependent amino acid ligase reactions. The reaction may also be run under conditions where the ATP-dependent amino acid ligase is the yield-limiting catalyst, such that a doubling or halving in concentration of the ATP-dependent amino acid ligase will effect a doubling or halving of the yield of product observed at a given timepoint.
[0154] Where the improved enzyme property desired is thermostability, enzyme activity may be measured after subjecting the enzyme preparations to a defined temperature and measuring the amount of enzyme activity remaining after heat treatments. Any suitable approach may be used, e.g., differential scanning colorimetry (DSC) a biochemical assay, or spectroscopy. Clones containing a polynucleotide encoding ATP-dependent amino acid ligases are then isolated, sequenced to identify the nucleotide sequence changes (if any), and used to express the enzyme in a host cell.
[0155] Where the sequence of the polypeptide is known, the polynucleotides encoding the enzyme can be prepared by standard solid-phase methods, according to known synthetic methods. In some embodiments, fragments of up to about 100 bases can be individually synthesized, then joined (e.g., by enzymatic or chemical litigation methods, or polymerase mediated methods) to form any desired continuous sequence. For example, polynucleotides and oligonucleotides of the invention can be prepared by chemical synthesis using, e.g., the classical phosphoramidite method described by Beaucage et al., 1981, TET. LETT.22:1859-69, or the method described by Matthes et al., 1984, EMBO J.3:801-05, e.g., as it is typically practiced in automated synthetic methods. According to the phosphoramidite method, oligonucleotides are synthesized, e.g., in an automatic DNA synthesizer, purified, annealed, ligated and cloned in appropriate vectors. In addition, essentially any nucleic acid can be obtained from any of a variety of commercial sources, such as Twist Bioscience, South San Francisco, Calif.; Genscript, Piscataway, New Jersey; Azenta Life Sciences, South Plainfield, New Jersey; Biomatik, Wilmington, Delaware; Integrated DNA Technologies, Coralville, Iowa; and many others.25971
[0156] Chromatographic techniques for isolation of the ATP-dependent amino acid ligase polypeptide include, among others, reverse phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. Conditions for purifying a particular enzyme will depend, in part, on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and will be apparent to those having skill in the art.
[0157] In some embodiments, affinity techniques may be used to isolate the improved ATP- dependent amino acid ligase enzymes. For affinity chromatography purification, the protein sequence can be tagged with a recognition sequence to enable purification. Common tags include cellulose-binding domains, poly His-tags (e.g., 6xHis tags), di-His chelates, FLAG-tags and many others that will be apparent to those having skill in the art. Antibodies can also be used as affinity purification reagents. Any antibody that specifically binds the ATP-dependent ligase polypeptide may be used. Abbreviations NaPi Sodium phosphate HTS High throughput screening25971 IPTG isopropyl β-D-1-thiogalactopyranoside Combi Combinatorial library generation HEPES 4 2 h d th l 1 i i th lf i id
[0158] A DNA sequence encoding wild-type ATP-dependent ligase polypeptide from Bifidobacterium adolescentis (SEQ ID NO: 1) identified from a literature report (See Arai et al., Biosci. Biotechnol. Biochem.2010 for reference) was codon optimized for E. coli expression and chemically synthesized to afford reference nucleic acid sequence SEQ ID NO: 3. An N-terminal hexahistidine tag was included as part of the sequence. The gene was cloned into a pET30a(+) E. coli expression vector containing a LacI promoter, a ColE1 origin of replication and a KanR selection marker. Cloned gene was propagated into DH5a strain, sequence verified and transformed into BL21(DE3) E. coli cells. Likewise, genes of engineered ATP-dependent ligases were cloned into a pET30a vector and transformed into BL21(DE3) cells. Example 2: Enzyme Preparation for well-plate Reactions
[0159] A glycerol stock from BL21(DE3) cells containing an ATP-dependent ligase gene of interest in pET30a(+) was inoculated into LB-agar 96-well plates supplemented with 30 µg / mL kanamycin and 1% (w / v) glucose and grown overnight at 30°C. The following day, individual colonies were inoculated into a 96-well plate containing 0.2 mL per well of LB broth supplemented with 30 µg / mL of kanamycin and 1 % (w / v) glucose. Plates were incubated at 30°C shaking at 200 rpm overnight. The following day, the optical density at 600 nm (OD600) of the saturated culture was measured. Cells were then diluted to an OD600of 0.05 in ~390 µL of ZYM-5052 autoinduction media supplemented with 30 µg / mL of kanamycin in a 96-deep well plate. Protein expression proceeded for 20 h at 30°C and 250 rpm. Cells were pelleted by centrifugation at 4000 g for 15 min, and the supernatant was discarded. Pelleted cells were frozen / thawed once and then lysed in 0.2 mL / well of 50 mM HEPES pH 7.5 buffer, 1 mg / mL lysozyme, 0.5 mg / mL polymyxin B sulfate, 1 mM MgCl2 and 1 U / mL DNaseI endonuclease for 2 h at room temperature shaking at 1000 rpm. Cell debris was clarified by centrifugation at 4000 x g for 15 minutes, and the clarified lysate (supernatant) used in subsequent well-plate reactions.25971 Example 3: Enzyme Preparation for Vial and Larger-scale reactions
[0160] Twenty microliters of a glycerol stock of BL21(DE3) containing the ATP-dependent ligase gene of interest in pET30a(+) was inoculated in 25 mL of LB broth supplemented with 30 µg / mL kanamycin and 1% (w / v) glucose. Cells were grown for 18 h at 30°C and 250 rpm. The following day, 1 L of TB media supplemented with 30 µg / mL kanamycin was subcultured with the overnight saturated culture to an initial OD of 0.05. Cells were grown at 30°C shaking at 250rpm until OD600reached 0.6-0.8. Protein production was initiated with the addition of 1 mM IPTG and cells were cultured for 20 h at 30°C. Biomass was pelleted by centrifugation at 4000 g for 15 min, and the supernatant was discarded. Cell pellets were frozen / thawed once and resuspended in 50 mM triethanolamine-HCl pH 7.5 (5 mL per gram of pellet). Cell suspensions were shaken at 18°C for 30 minutes, after which cells were disrupted by high-pressure homogenization (16,000 PSI). The resulting lysate was clarified by centrifugation at 10,000 x g for 45 minutes at 4°C. The clarified lysate was frozen and lyophilized. Powders were used in subsequent vial (small-scale) reactions as detailed in the following Examples. Example 4: Evolution and screening of polypeptides derived from SEQ ID NO: 4 for tripeptide formation of product of Reaction A
[0161] ATP-dependent ligase enzymes derived from SEQ ID NO: 4 were evolved from the polynucleotide having the nucleotide sequence of SEQ ID NO: 3 against a selection pressure for activity (conversion of substrate to product 3a), which encodes the polypeptide of SEQ ID NO: 4. Libraries of engineered polypeptides were generated by structure-guided directed evolution using site saturation mutagenesis and combinatorial libraries. Directed evolution strategy
[0162] In the directed evolution campaign, single-site-saturation mutagenesis (SSM) libraries were built and screened for increased conversion relative to the starting enzyme. SSM libraries in each round of evolution were generated using gene splicing by overlap extension (SOEing) PCR methods described in Hiraga et al. ACS Synth. Biol.2021, 10(2), 357-370. In short, mutations at the designated positions were incorporated by degenerate NNK oligonucleotides through overlapping PCRs (using Mutation Maker). The full-length gene-of-interest region was further amplified and assembled into the expression vector (pET30a) by Gibson Assembly (Gibson et al. Nat. Methods 2009, 6(5), 343-345). The combinatorial libraries were built following the instructions from the QUIKCHANGE® Lightning Multi Site-Directed Mutagenesis kit (Agilent Technologies). Both mutagenesis libraries were transformed into the BL21(DE3) E.coli strains25971 by electroporation and plated on LB agar plates with selection (1% glucose, 30 μg / ml Kanamycin).
[0163] ATP-dependent ligase enzymes derived from SEQ ID NO: 4 (affinity tagged-BAD1200) were screened in a 96-well microtiter plate format using the following conditions: In a final volume of 200 µL / well 10 mM substrate 1a, 40 mM substrate 2a, 100 mM Tris-HCl pH 8, 12.5 mM MgCl2, 1 mM ATP, 25 mM propionyl phosphate, 0.3 g / L ACK-101 were incubated in the presence of 10 % (v / v) enzyme lysates derived from SEQ ID NO: 4. Reactions were incubated at 30°C for 20 h shaking at 600 rpm. The following day, reactions were quenched by adding 100 µL of reaction to 100 µL of ACN. Quenched reactions were filtered and formation of product 3a was analyzed by LC-MS. For each of the Examples below, except where otherwise indicated, LC-MS analysis was conducted using the following method: Instrument: Agilent 1260 equipped with an Agilent 6130 Quadrupole MS Column: Ac uit BEH C18 130Å 17 m 21x50mm (SKU: 186002350)
[0006] or e par cu ar C- S ana ys s o s xampe, e o ow ng mobile phase, flow rate and gradient were used: Mobile Phase: MP A: H2O (0.1% TFA), MP B: ACN (0.1% TFA) Flow rate: 0.75 mL / min Time (min) A% B% 0 84 16 n
[0165] Variants catalyzing the highest conversion to product 3a were further scaled up to shake flasks as described in Example 3 and screened under similar conditions as described above with the exception that 2 g / L of polypeptide shake flask powder was used to determine formation of product 3a. Engineered polypeptides having the amino acid sequences of SEQ ID NOs: 6, 8, 10 and 12 were identified from several rounds of evolution that identified variants (or “hits”) exhibiting improvements in enzymatic activity in substrates of Reaction A (see Table 1).Polypeptide having the amino acid sequence of SEQ ID NO: 12 also exhibited improved enzymatic activity on the substrates of Reaction B and thus selected for further evolution. Example 5: Evolution and screening of polypeptides derived from SEQ ID NO: 12 for tripeptide formation of product of Reaction B (Formula 3b)
[0166] ATP-dependent ligase enzymes derived from SEQ ID NO: 12 were evolved from the polynucleotide having the nucleotide sequence of SEQ ID NO: 11. ATP-dependent ligase enzymes derived from SEQ ID NO: 12 were screened in a 96-well-plate format using the following conditions: In a final volume of 200 µL / well 10 mM substrate 1b, 40 mM substrate 2b, 100 mM Tris-HCl pH 8, 12.5 mM MgCl2, 1 mM ATP, 25 mM propionyl phosphate, 0.3 g / L ACK-101 were incubated in the presence of 5 % (v / v) enzyme lysates derived from SEQ ID NO: 12. Reactions were incubated at 30°C for 20 h shaking at 600 rpm. The following day, reactions were quenched by adding 100 µL of reaction to 100 µL of ACN. Quenched reactions were filtered and formation of product 3b was analyzed by LC-MS using the following mobile phase, flow rate and gradient: Mobile Phase: MP A: H2O (10 mM ammonium formate pH 8.5), MP B: ACN Flow rate: 0.6 mL / min Time (min) A% B% 0 90 10 n
[0167] Variants catalyzing the highest conversion to product 3b were further scaled up to shake flasks and screened under similar conditions as shown above with the exception that 2 g / L of polypeptide shake flask powder was used to determine formation of product 3b. The polypeptide having the amino acid sequence of SEQ ID NO: 14 and was identified for improvements in enzymatic activity in substrates of Formula B. The enzyme polypeptide having the amino acid sequence of SEQ ID NO: 12 and 14 also exhibited improved enzymatic activity in Reaction C. Polypeptide of SEQ ID NO.12 exhibited better enzymatic activity in Reaction C and selected for further evolution.Example 6: Evolution and screening of polypeptides derived from SEQ ID NO: 12 for tripeptide formation of product of Reaction C
[0168] ATP-dependent ligase enzymes derived from SEQ ID NO: 12 were evolved from the polynucleotide having the nucleotide sequence of SEQ ID NO: 11. ATP-dependent ligase enzymes derived from SEQ ID NO: 12 were screened in a 96-well-plate format using the following conditions: In a final volume of 100 µL / well 10 mM substrate 1c, 25 mM substrate 2c, 100 mM HEPES pH 8, 10 % (v / v) DMSO, 25 mM MgCl2, 20 mM ATP were incubated in the presence of 7.5 % (v / v) enzyme lysates derived from SEQ ID NO: 12. Reactions were incubated at 30°C for 20 h shaking at 600 rpm. The following day, reactions were quenched by adding 100 µL of reaction to 100 µL of ACN. Quenched reactions were filtered and formation of product 3c was analyzed by LC-MS using the following mobile phase, flow rate and gradient: Mobile phase: MP A: Water (0.1% difluoroacetic acid), MP B: ACN (0.1% difluoroacetic acid) Flow rate: 0.6 mL / min Time, min %A %B n
[0169] Variants catalyzing the highest conversion to product 3c were further scaled up to shake flasks and screened under similar conditions as shown above, with the exception that 2 g / L of polypeptide shake flask powder was used to determine formation of product 3c. The polypeptide having the amino acid sequence of SEQ ID NO: 16 was identified for improvements in enzymatic activity on substrates of Reactions C and D and thus selected for further evolution. Example 7: Evolution and screening of polypeptides derived from SEQ ID NO: 16 for tripeptide formation of compounds of Reaction D
[0170] ATP-dependent ligase enzymes derived from SEQ ID NO: 16 were evolved from the polynucleotide having the nucleotide sequence of SEQ ID NO: 15. ATP-dependent ligase enzymes derived from SEQ ID NO: 16 were screened in a 96-well-plate format using the following conditions: In a final volume of 50 µL / well 10 mM substrate 1d, 25 mM substrate 2c, 100 mM HEPES pH 8, 5 % (v / v) DMSO, 25 mM MgCl2, 20 mM ATP were incubated in the25971 presence of 10 % (v / v) enzyme lysates derived from SEQ ID NO: 12. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 100 µL of reaction to 100 µL of ACN. Quenched reactions were filtered and formation of product 3d was analyzed by LC-MS using the following mobile phase, flow rate and gradient: Mobile phase: MP A: Water (0.1% difluoroacetic acid), MP B: ACN (0.1% difluoroacetic acid) Flow rate: 0.8 mL / min Time, min %A %B
[0171] Variants catalyzing the highest conversion to product 3d were further scaled up to shake flasks and screened under similar conditions as shown above with the exception that 2 g / L of polypeptide shake flask powder was used to determine formation of product 3d. The polypeptide having the amino acid sequence of SEQ ID NO: 18 was identified for improvements in enzymatic activity and thus selected for further evolution. Example 8: Evolution and screening of polypeptides derived from SEQ ID NO: 18 and beyond for tripeptide formation of products of Reaction D
[0172] SEQ ID NO: 18 was further evolved sequentially to identify hits that exhibited 1) improved activity, 2) substrate load, 3) phosphate tolerance, 4) co-solvent tolerance, and / or 5) reduced by-product formation. The best hits i.e., engineered variants from each round of evolution was used as the backbone in the following round of evolution, as described below. Assay conditions and evolution pressures are provided in Table 5 as follows: Table 5. Further evolution pressures applied to variants derived from SEQ ID NO: 4 for tripeptide formation25971 4,196,197,198,199,200,232,233,234,235,236,237,23 8,239,240,241,276,291,292,293,294,295,296,297,29 8,299,300,301,302,330,331,332,333,334,335,336,3325971 Combi: L14V,M15L,S45Q,S56D,S56H,K75C,K75S,H113L ,W164S,Y239H,K315F,L319E,Y354F,S334R25971
[0173] The ligase having the sequence of SEQ ID NO: 40, indicated in the final row of the above table as emerging from a final round of evolution, contains the following 47 substitutions relative to the reference sequence of SEQ ID NO 4: L15M, Q17H, K19Q, F21L, C37P, N46S, A49H, S68G, S91G, S92L, L111V, A166V, G170S, A182E, S205N, L229M, 3fV230R, F235R, K237I, N239H, T240H, N248G, A264G, F266L, L300I, A307L, V325H, A327E, L341I, E342M, A344E, I354Y, L360E, V363D, K366T, G368E, A370N, H371S, S373G, V374A, D375W, S376W, K379M, S388E, S394Q, N395R, and Q401W.
[0174] SEQ ID NOs 40, 42, 44, 46, 48, and 50 emerged from 21 rounds of engineering via directed evolution as the best-performing ATP-dependent Trp-ligase variants. The engineered polypeptides of SEQ ID NOs: 42, 44, 46, 48, and 50 constitute variants evolved from SEQ ID NO: 40 in a final round of evolution for which activity and by-product reduction selection pressures were applied. To engineer phosphate tolerance, 100 mM KPi was added to the screen for variants of SEQ ID NO: 19 / 20. To engineer co-solvent tolerance and selectivity against substrate 4f, these components were added to the screen for variants of SEQ ID NO: 21 / 22. To engineer specificity against pentapeptide formation, substrate 5k was added to the screen of variants of SEQ ID NO: 33 / 34. These engineered polypeptides show improved 1) activity for the tripeptide product of Compound 3d, 2) regioselectivity for Compound 3d over dimers, trimers and other oligomers of compounds according to Formula I (i.e., oligomers of Compound 1a, Compound 1b, Compound 1c, or Compound 1d) and a product of ligation of compound 3k to Compound 1d in a tetrapeptide-forming cascade context;, (3) phosphate tolerance;, (4) substrate loading tolerance; (5) thermostability; and / or (6) by-product formation reduction relative to SEQ ID NO: 4, such as dipeptide by-products formed in a cascade context from a compound of Formula I and a compound of Formula IV, or a tetrapeptide by-product formed from a dimerization of a compound of Formula I ligated to a dipeptide of Formula II. For example, the Trp-ligase of SEQ ID NO: 40 exhibited about 23,000-fold activity level improvement relative to SEQ ID NO: 4, generating compound 3d with 100% regioselectivity and by-products at a level of less than 0.5% of all products. Example 9: Engineering of Phe-ligase polypeptides derived from SEQ ID NO: 4 for formation of tetrapeptide product of Reaction E
[0175] ATP-dependent ligase enzymes derived from SEQ ID NO: 4 were evolved from the polynucleotide of SEQ ID NO: 3. In the second trajectory, to engineer Phe-ligase variants, a substrate walk was performed among the substrates of Formula V from substrate 2d, to substrate 2e, to substrate 2e, to substrate 2g, to substrate 2h, and finally to substrate 3d (see Table 2).25971
[0176] Libraries of engineered polypeptides were generated using site saturation mutagenesis and combinatorial libraries. ATP-dependent ligase enzymes derived from SEQ ID NO: 4 were screened in a 96-well-plate format to identify hits for improved activity in catalyzing the formation of tetrapeptide compound 1 using the following conditions: In a final volume of 150 µL / well 10 mM substrate 1e, 10 mM substrate 2d, 100 mM Tris-HCl pH 8, 12.5 mM MgCl2, 1 mM ATP, 25 mM propionyl phosphate, 0.3 mg / mL ACK-101 were incubated in the presence of 10 % (v / v) enzyme lysate. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by transferring 20 µL of the reaction to 180 µL 80% ACN. Quenched reactions were filtered through a 0.2 µm filter, and formation of product 3e was analyzed by LC-MS using the following mobile phase, flow rate and gradient: Mobile phase: MP A: Water (0.1% trifluoroacetic acid), MP B: ACN (0.1% trifluoroacetic acid) Flow Rate: 0.75 mL / min Time, min %A %B n
[0177] Variants catalyzing the highest conversion to product 3e were further scaled up to shake flasks and screened under similar conditions as shown above with the exception that 2 g / L of polypeptide shake flask powder was used to determine formation of product 3e. The polypeptide having the amino acid sequence of SEQ ID NO: 52 was identified as a variant exhibiting significant improvements in enzymatic activity for reactions E, F and G and thus used as a backbone for further evolution. Example 10: Engineering of Phe-ligase polypeptides derived from SEQ ID NO: 52 for formation of tetrapeptide of Reaction F
[0178] In the next round of engineering, a library was built from SEQ ID NO: 51 that combinatorically sampled beneficial diversity identified from the Example 9 SSM libraries. This library was built, expressed, and screened in a 96-well-plate format to identify hits for improved activity for the formation of product 3f using the following conditions: In a final volume of 20025971 µL / well 10 mM substrate 1f, 10 mM substrate 2d, 100 mM Tris-HCl pH 8, 12.5 mM MgCl2, 1 mM ATP, 25 mM propionyl phosphate, 0.3 mg / mL acetate kinase were incubated in the presence of 5 % (v / v) enzyme lysate. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by transferring 10 µL of reaction to 190 µL 80% ACN. Quenched reactions were filtered and formation of product 3f analyzed by LC-MS using the following mobile phase, flow rate and gradient: Mobile phase: MP A: Water (0.1% trifluoroacetic acid), MP B: ACN (0.1% trifluoroacetic acid) Flow Rate: 0.75 mL / min Time, min %A %B the highest conversion to product 3f were further scaled to shakeflasks and screened under the following conditions: 50 mM 1f, 50 mM 2d, 1 mM ATP, 60 mM propionyl phosphate, 25 mM MgCl2, 0.3 mg / mL acetate kinase, 100 mM Tris-HCl pH 8.0, 2 mg / mL ligase shake flask powder. Reactions were shaken at 25°C for 20 h. After overnight incubation, reactions were quenched with 500 µL MeCN, filtered, and formation of product 3f analyzed by LC.
[0180] The polypeptide having the amino acid sequence of SEQ ID NO: 58 was identified as a variant having improved enzymatic activity for product formation on Reactions F and G at higher substrate loads and thus selected as the backbone for further evolution. Example 11: Engineering of Phe-ligase polypeptides derived from SEQ ID NO: 58 for formation of tetrapeptide of Reaction G
[0181] Using SEQ ID NO: 57 as the template, a library combinatorically sampling previously identified beneficial diversity and a 96-position SSM library targeting the active site region were constructed, expressed, and screened. These ATP-dependent Phe-ligase variants of SEQ ID NO: 58 were screened to identify hits improved for activity under the following conditions: In 96- well-plate and a final volume of 50 µL / well 40 mM substrate 1f, 40 mM substrate 2e, 100 mM Tris-HCl pH 8, 25 mM MgCl2, 1 mM ADP, 40 mM polyphosphate, 0.1 mg / mL PPK12 were25971 incubated in the presence of 1.25 % (v / v) enzyme lysate. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 20 µL of reaction to 180 µL 80 % ACN. Quenched reactions were filtered and formation of product 3g analyzed by LC-MS using the method described in Example 10, except that the Product 3g retention time was 1.639 min.
[0182] Variants catalyzing the highest conversion to product 3g were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 µL, 50 mM 1f, 50 mM 2e, 5 mM ATP, 60 mM polyphosphate, 25 mM MgCl2, 0.1 mg / mL PPK12, 100 mM Tris- HCl pH 8.0 were mixed with 2 mg / mL shake flask powder of ATP-dependent ligase derived from SEQ ID NO: 58. Reactions were incubated at 25°C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quench with 500 µL MeCN, filtered, and analyzed by LC.
[0183] The polypeptide having the amino acid sequence of SEQ ID NO: 60 was identified as a variant exhibiting improvements in enzymatic activity in Reactions G, H and I and thus selected for further evolution. Example 12: Engineering of Phe-ligase polypeptides derived from SEQ ID NO: 60 for formation of tetrapeptide of Reaction H
[0184] Libraries encoding ATP-dependent ligase enzymes derived from SEQ ID NO: 60 were screened to identify hits improved for activity under the following conditions: In 96-well-plate and a final volume of 50 µL / well 40 mM substrate 1f, 40 mM substrate 2f, 100 mM Tris-HCl pH 8, 25 mM MgCl2, 1 mM ADP, 40 mM polyphosphate, 0.1 mg / mL PPK12 were incubated in the presence of 2 % (v / v) enzyme lysate. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 20 µL of reaction to 180 µL 50 % ACN. Quenched reactions were filtered and formation of product 3h analyzed by LC-MS using the following mobile phase, flow rate and gradient: Mobile phase: MP A: Water (0.1% formic acid), MP B: ACN (0.1% formic acid) Flow Rate: 0.75 mL / min Time, i A B n
[0185] Variants catalyzing the highest conversion to product 3h were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 µL, 50 mM 1f, 5025971 mM 2f, 5 mM ATP, 60 mM polyphosphate, 25 mM MgCl2, 0.1 mg / mL PPK12, 100 mM Tris- HCl pH 8.0, were mixed with 2 mg / mL shake flask of enzymes derived from SEQ ID NO 60. Reactions were incubated at 25°C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quenched with 500 µL MeCN, filtered, and analyzed by LC.
[0186] The polypeptide of amino acid sequence SEQ ID NO: 62 was identified as a variant having improvements in enzymatic activity for Reaction I, thus selected for further evolution. Example 13: Engineering of Phe-ligase polypeptides derived from SEQ ID NO: 62 for formation of tetrapeptide of Reaction I
[0187] Libraries encoding ATP-dependent ligase enzymes derived from SEQ ID NO: 62 were screened to identify hits improved for activity under the following conditions: In 96-well-plate and a final volume of 50 µL / well 20 mM substrate 1f, 20 mM substrate 2g, 100 mM Tris-HCl pH 8, 25 mM MgCl2, 1 mM ADP, 40 mM polyphosphate, 0.1 mg / mL PPK12 were incubated in the presence of 2.5 % (v / v) enzyme lysate. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 20 µL of reaction to 180 µL 50 % ACN. Quenched reactions were filtered and formation of product 3i analyzed by LC-MS using the following mobile phase, flow rate and gradient: Mobile phase: Water (0.1% formic acid), B: ACN (0.1% formic acid) Flow Rate: 0.75 mL / min Time, min %A %B
[0188] Variants catalyzing the highest conversion to product 3i were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 µL, 50 mM 1f, 50 mM 2g, 5 mM ATP, 60 mM polyphosphate, 25 mM MgCl2, 0.1 mg / mL PPK12, 100 mM Tris- HCl pH 8.0, were mixed with 2 mg / mL shake flask of enzymes derived from SEQ ID NO 62. Reactions were incubated at 25°C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quenched with 500 µL MeCN, filtered, and analyzed by LC.
[0189] The polypeptide having amino acid sequence SEQ ID NO: 66 was identified as a variant having improvements in enzymatic activity for Reactions I and J. Upon further assessment of the LC-MS chromatogram of the ligase having the amino acid sequence of SEQ ID NO: 66, it was25971 observed that this enzyme forms two distinctive regioisomers depicted below in Formula IX when substrates from Reaction J are present (compounds 3j and 3j-regio). SEQ ID NO: 66 was selected for further evolution.formation of tetrapeptide of Reaction J
[0190] Using SEQ ID NO: 65 as the template, a library combinatorically sampling beneficial diversity from previous rounds and a 96-position SSM library targeting the active site region were constructed, expressed, and screened. This SSM was designed to target residues where modification was anticipated to better accommodate the modified tryptophan and threonine residues found in substrate 3c. These ATP-dependent Phe-ligase variants of SEQ ID NO: 66 were screened to identify hits improved for 1) activity and 2) reduction in undesired regioisomer (3j-regio) formation under the following conditions: In 96-well-plate and a final volume of 50 µL / well 10 mM substrate 1f, 10 mM substrate 2h, 100 mM Tris-HCl pH 8, 25 mM MgCl2, 2 mM ADP, 40 mM polyphosphate, 0.1 mg / mL PPK12 were incubated in the presence of 20 % (v / v) enzyme lysate. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 25 µL of reaction to 125 µL 80 % ACN. Quenched reactions were filtered and formation of product 3j was analyzed by LC-MS using the following method: Instrument: Agilent 1260 equipped with an Agilent 6130 Quadrupole MS Column: Acquity UPLC HSS Cyano 1.8 μm, 2.1 x 50 mm Column Temp: 45°C Flow Rate: 0.65 mL / min Injection volume: 1 μL Detector: UV at 210 nm and MS (for low conversion variants) Gradient: A: 5 mM ammonium carbonate in water, B: ACN25971 Time, min %A %B 0 80 20(desired regioisomer), 1.137 (undesired regioisomer)
[0191] Variants catalyzing the highest conversion to product 3j were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 µL, 50 mM 1f, 50 mM 2h, 5 mM ATP, 60 mM polyphosphate, 25 mM MgCl2, 0.1 mg / mL PPK12, 100 mM Tris- HCl pH 8.0, were mixed with 2 mg / mL shake flask of enzymes derived from SEQ ID NO 66. Reactions were incubated at 25°C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quench with 500 µL MeCN, filtered, and analyzed by LC.
[0192] SEQ ID NO: 68 was identified as a variant exhibiting improvements in both enzymatic activity and regioselectivity for Reaction K and thus selected for further evolution. Example 15: Engineering of Phe-ligase polypeptides derived from SEQ ID NO: 68 for formation of tetrapeptide of Reaction K
[0193] Using SEQ ID NO: 67 as the template, a library combinatorically sampling beneficial diversity from previous rounds and a 96-position SSM library targeting the active site region were constructed, expressed, and screened. These ATP-dependent Phe-ligase variants of SEQ ID NO: 68 were screened to identify hits improved for activity and regioselectivity under the following conditions: In 96-well-plate and a final volume of 50 µL / well 10 mM substrate 1f, 10 mM substrate 2i, 100 mM Tris-HCl pH 8, 25 mM MgCl2, 2 mM ADP, 40 mM polyphosphate, 0.1 mg / mL PPK12 were incubated in the presence of 10 % (v / v) enzyme lysate. Reactions were incubated at 30°C for 20 h shaking at 1000 rpm. The following day, reactions were quenched by adding 25 µL of reaction to 125 µL 80 % ACN. Quenched reactions were filtered and formation of product 3k analyzed by LC-MS using the following flow rate and gradient: Flow Rate: 0.80 mL / min Time,2.75 80 20 3.00 80 20n
[0194] Variants catalyzing the highest conversion to product 3k were further scaled to shake flasks and screened under the following conditions: In a final volume of 400 µL, 25 mM 1f, 25 mM 2i, 5 mM ATP, 60 mM polyphosphate, 25 mM MgCl2, 0.1 mg / mL PPK12, 100 mM Tris- HCl pH 8.0, were mixed with 2 mg / mL shake flask of enzymes derived from SEQ ID NO 58. Reactions were incubated at 25°C for 20 h shaking at 400 rpm. After overnight incubation, reactions were quenched with 500 µL MeCN, filtered, and analyzed by LC.
[0195] SEQ ID NO: 70 was identified as a variant exhibiting improvements in enzymatic activity and regioselectivity for Reaction K and thus selected for further evolution. Example 16: Secondary directed evolution strategy: Designs of libraries targeting regioselectivity, cascade product inhibition, and thermostability
[0196] Prior to and during execution of directed evolution for activity (product formation), additional SSM libraries were generated to provide structure-guided mutagenesis of Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1, SEQ ID NO: 2 residues that would promote regioselectivity, reduced product inhibition in enzyme cascade, and thermostability.
[0197] A challenge of a one-pot enzymatic cascade is substrate selectivity / specificity. Individually discovered and engineered enzymes need to display selectivity and ability to work in the presence of multiple substrates, simultaneously. In the context of a one-pot cascade of the two enzymatic reactions to generate tripeptide and tetrapeptide products, respectively, it was discovered that FTrp (a substrate of the first enzymatic reaction in the pot) is a competitive inhibitor against aPhe (a substrate of the second enzymatic reaction in the pot). To overcome this challenge, molecular dynamics (MD) simulations were performed to identify residues in the active site that associate strongly with FTrp, as well as positions outside of the active site where there is significant protein motion when FTrp is bound. MD simulations provide molecular-level details to capture positions far from the active site that can modulate enzyme activity and selectivity.
[0198] To identify such positions, simulations modeled the competitive inhibitor and the product it would form, in addition to desired product in the binding site. A MD simulation wasperformed to capture the dynamics of enzyme and the products. The design criteria were to identify residues that i) participate in strong interactions with undesired product ii) are not critical for desired product binding, and iii) sites that are outside the immediate active sites but show strong motions when bound with the undesired substrate. These positions were targeted in an SSM library. Mutations at positions 173, 236 and 237 are predicted to interact with FTrp and ATP. Specifically, Ile 173 interacts with the 6,5-ring system in the ATP which results in tighter binding with ATP (interaction energy of -3.5 kcal / mol vs. -1.5 kcal / mol in the wt variant) and allows for slight rearrangement in the active site to accommodate the larger desired substrate. Glu 236 and Lys 237 sit in a loop and interact with hexylamine linker portion of FTrp. Positions 153, 403 and 404 predicted from the protein fluctuation analysis are involved in long-range allostery and improved enzymatic activity by 1.2- to 14-fold over the parental backbone.
[0199] Finally, to improve thermostability of the enzyme, an MD-based SSM library was designed in which two separate MD simulations at 295K and at 330K were performed. Protein positions that had major differences (>5Å) in root-mean square fluctuations (RMSF) between high and low temperature MD simulations and showed weakened interactions to the substrate over the course of MD simulations are the sites that were targeted for an SSM library. This SSM library was screened by subjecting the variants to a heat challenge of 37°C. The screening yielded several mutations showing 1.5- to 3-fold improvement in thermostability over parental backbone. Notably, A164S sits at the dimer interface (Fig 5d) and the serine mutant results in stronger polar interactions at the interface, lowering the dimer interface energy by -25 REU compared with the parental enzyme interface. This sheds light on the importance of considering domain motions when designing for thermal stability. A cluster of these hits sit in a loop that became increasingly mobile during high temperature MD: Protein positions 350, 352, 359, 363, 367 and 399. P53V mutation sits in a loop that connects two beta strands and the valine variant makes the secondary structure more rigid. A combinatorial library built from the improved hits from the MD-based SSM improved activity by an additional ~2-fold.
[0200] An assessment of the T50 of each of the backbones chosen from this MD / SSM approach found an overall 5°C gained in thermostability while simultaneously improving enzymatic activity. Example 17: Engineering of Phe-ligase polypeptides derived from SEQ ID NO: 70 and beyond for formation of tetrapeptide product of Reaction K
[0201] SEQ ID NO: 70 was further evolved sequentially to identify hits that exhibited 1) improved activity, 2) substrate load, 3) phosphate tolerance, 4) co-solvent tolerance, and 5)25971 reduced product inhibition. Assay conditions and evolution pressures are provided in Table 6 as follows. Table 6. Engineering of Phe-ligase polynucleotides and polypeptides from SEQ ID NO: 69 / 70 for tetrapeptide formation Backbone Starting selected for evolution add’l , NO25971 10,25,27,32,35,38,39,41,50,52,54,55,57,58 ,61,62,70,71,78,79,83,86,97,104,105,106,1 10,112,113,117,123,124,128,133,140,141,25971 Combi: S91G,I136A,I136F,Y139W,W164A,L229 N,L229R,K237F,N282K,N282P,E283I,E25971 ,191,218,219,235,236,238,247,248,250,25 3,257,260,264,267,268,270,284,303,310,3 15,326,327,328,330,332,334,346,350,351,25971 Combi: G25P,D47V,S65D,S65E,A164S,L165G,G 205D,G205E,G205R,V207P,V230T,V23025971
[0202] The Phe-ligase having the sequences of SEQ ID NO: 107 / 108, indicated in the final row of the above table as emerging from a final round of evolution in Example 17, contains the following 66 substitutions relative to the reference sequence of SEQ ID NO 4: Q16Y, Q17R, K19G, F21D, P25G, C37H, A50E, H51P, H52P, E61V, S65D, S68Y, F83W, S91G, E94D, R120I, S128K, D131K, G144C, F152L, W164A, G165L, A166H, A169Q, P181E, A182R, A184C, K186S, C204L, S205E, S207P, V208Q, G219N, V230C, D234N, K237T, Q238D, A250T, S253Q, R257P, A264Y, T284S, G299S, E304S, G310N, V325L, A329F, S334Q, Y337N, K350F, V359Y, E361H, K365E, I369L, H371Q, D377K, K379R, S383F, S388E, I391L, S394C, Y396M, I398T, S400W, Q401N, and Y402G.
[0203] SEQ ID NOs 108, 110, 112, 114, and 116 emerged from 29 rounds of evolutionary pressures as the best-performing ATP-dependent Phe-ligase variants. The engineered polypeptides of SEQ ID NOs: 108, 110, 112, 114, and 116 constitute variants evolved from SEQ ID NO: 105 / 106 in a final round of evolution for which activity and by-product reduction selection pressures were applied. These engineered polypeptides may exhibit improved 1) activity for the product of Compound 3k, 2) selectivity for Compound 3k over dimers, trimers and / or oligomers of compounds of Formula IV (i.e., Compound 1e and Compound 1f), 3) thermostability, 4) reaction kinetics, 5) pH tolerance, 6) phosphate tolerance, (7) reduction in product inhibition, and / or (8) substrate loading capacity relative to wild-type Bifidobacterium adolescentis (GenBank: AP009256.1, SEQ ID NO: 1; encoding GenBank: BAF39981.1, (SEQ ID NO: 4). Polynucleotides Encoding Carboxylesterases
[0204] In various embodiments, provided herein are engineered polynucleotides that comprise nucleic acid sequences having at least 75%, 80%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 92.5%, 93%, 94%, 95%, 96%, 97%, 97.5%, 98% or 99% sequence identity to SEQ ID NO: 131. In various embodiments, these variants comprise nucleic acid sequences having at least 75%, 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% identity to any one of SEQ ID NOs: 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, and 165. In some embodiments, these variants comprise nucleic acid sequences having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% identity to any one of SEQ ID NOs: 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, and 157. In some embodiments, these variants comprise nucleic acid sequences comprising any one of SEQ ID NOs: 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, and 165. In some examples, the polynucleotide sequence of any of SEQ ID NOs: 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155,25971 157, 159, 161, and 165 is modified to no longer code for an N-terminal 6xHis tag having the amino acid sequence of SEQ ID NO: 167. In some examples, the polynucleotide sequence of any of SEQ ID NOs: 133, 135, 137, 139, 141, 143, 145, 147, 149, 151, 153, 155, 157, 159, 161, and 165 is modified to have a signal peptide region-encoding sequence.
[0205] In some embodiments are provided engineered polynucleotides comprising a nucleic acid sequence having at least 80%, 85%, 90%, 92.5%, 94%, 95%, 96%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 151, 153, 155, and 157. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 157. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 155. In some embodiments, the polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 153. In some embodiments, the enzyme polynucleotide comprises a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 151.
[0206] Polynucleotides comprising any of SEQ ID NOs: 151, 153, 155, and 157 are provided. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 157. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 155. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 153. In some embodiments, the enzyme comprises the nucleic acid sequence of SEQ ID NO: 151.
[0207] In some aspects, provided are engineered polynucleotides comprising an nucleic acid sequence that comprises a stretch of at least 100, 150, 200, 250, 300, 325, 350, 500, 650, 800, 1000, 1100, 1200, or 1225 consecutive nucleotides of any one of SEQ ID NOs 151, 153, 155, and 157. In some aspects, provided are engineered polynucleotides comprising a nucleic acid sequence that comprises a stretch of at least 250, 350, 500, 650, 800, 1000, 1100, 1200, or 1225 consecutive nucleotides of any one of SEQ ID NOs 151, 153, 155, and 157. In some aspects, provided are engineered polynucleotides comprising a nucleic acid sequence that differs by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-25, 25-35, 35-50, 50-75, or more than 75 nucleotides relative to the sequence of any one of SEQ ID NOs 151, 153, 155, and 157. The polynucleotides may comprise a nucleic acid sequence that differs by 1, 2, 3, 4, or 5 nucleic acids relative to SEQ ID NO: 155 or 157.
[0208] In various embodiments, the disclosed polynucleotides are codon-optimized for expression in a particular organism, such as E. coli. Because of the knowledge of the codons corresponding to the various amino acids, availability of a protein sequence provides a description of all the polynucleotides capable of encoding the subject. The degeneracy of the genetic code, where the same amino acids are encoded by alternative or synonymous codons25971 allows an extremely large number of nucleic acids to be made, all of which encode the improved Carboxylesterase enzymes disclosed herein. Thus, having identified a particular amino acid sequence, those skilled in the art could make any number of different nucleic acids by simply modifying the sequence of one or more codons in a way that does not change the amino acid sequence of the protein. In this regard, the present disclosure specifically contemplates each and every possible variation of polynucleotides that could be made by selecting combinations based on the possible codon choices, and all such variations are to be considered specifically disclosed for any polypeptide disclosed herein.
[0209] In various embodiments, the codons are preferably selected to fit the host cell in which the protein is being produced. For example, preferred codons used in bacteria are used to express the gene in bacteria; preferred codons used in yeast are used for expression in yeast; and preferred codons used in mammals are used for expression in mammalian cells. By way of example, the polynucleotide of SEQ ID NO: 131 have been codon optimized for expression in Escherichia coli to afford SEQ ID NO: 133.
[0210] In certain embodiments, all codons need not be replaced to optimize the codon usage of the carboxylesterase enzyme since the natural sequence will comprise preferred codons and because use of preferred codons may not be required for all amino acid residues. Consequently, codon optimized polynucleotides encoding the carboxylesterase enzymes may contain preferred codons at about 40%, 50%, 60%, 70%, 80%, or greater than 90% of codon positions of the full- length coding region.
[0211] In various embodiments, an isolated polynucleotide encoding an improved carboxylesterase polypeptide may be manipulated in a variety of ways to provide for expression of the polypeptide. Manipulation of the isolated polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector. The techniques for modifying polynucleotides and nucleic acid sequences utilizing recombinant DNA methods are well known in the art. Guidance is provided in Sambrook et al., 2001, MOLECULAR CLONING: A LABORATORY MANUAL, 3rd Ed., Cold Spring Harbor Laboratory Press; and CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, Ausubel. F. ed., Greene Pub. Associates, 1998, updates to 2006.
[0212] In some embodiments, an isolated polynucleotide encoding any of the carboxylesterase polypeptides herein is manipulated in a variety of ways to facilitate expression of the carboxylesterase polypeptide. In some embodiments, the polynucleotides encoding the carboxylesterase polypeptides comprise expression vectors where one or more control sequences is present to regulate the expression of the carboxylesterase polypeptides. Manipulation of the25971 isolated polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector utilized. Techniques for modifying polynucleotides and nucleic acid sequences utilizing recombinant DNA methods are well known in the art. In some embodiments, the control sequences include among others, promoters, leader sequences, polyadenylation sequences, propeptide sequences, signal peptide sequences, and transcription terminators. In some embodiments, suitable promoters are selected based on the host cell selection. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure, include, but are not limited to, promoters obtained from the E. coli lac operon. In addition, suitable promoters may include Streptomyces coelicolor agarase gene (dagA), Bacillus subtilis levansucrase gene (sacB), Bacillus licheniformis alpha- amylase gene (amyL), Bacillus stearothermophilus maltogenic amylase gene (amyM), Bacillus amyloliquefaciens alpha-amylase gene (amyQ), Bacillus licheniformis penicillinase gene (penP), Bacillus subtilis xylA and xylB genes, and prokaryotic beta-lactamase gene (See e.g., Villa- Kamaroff et al., PROC. NATL ACAD. SCI. USA 75: 3727-3731
[1978] ), as well as the tac promoter (See e.g., DeBoer et al., PROC. NATL ACAD. SCI. USA 80: 21-25
[1983] ).
[0213] In some embodiments, the control sequence is also a suitable transcription terminator sequence (i.e., a sequence recognized by a host cell to terminate transcription). In some embodiments, the terminator sequence is operably linked to the 3' terminus of the nucleic acid sequence encoding the enzyme polypeptide. Any suitable terminator that is functional in the host cell of choice finds use in the present invention.
[0214] In some embodiments, the control sequence is also a suitable leader sequence (i.e., a non-translated region of an mRNA that is important for translation by the host cell). In some embodiments, the leader sequence is operably linked to the 5' terminus of the nucleic acid sequence encoding the carboxylesterase polypeptide. Any suitable leader sequence that is functional in the host cell of choice find use in the present invention. Exemplary leaders for E. coli will encode a ribosome binding site.
[0215] In some embodiments, the control sequence is a signal peptide region (i.e., a coding region that codes for an amino acid sequence linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the cell’s secretory pathway). In some embodiments, the 5ʹ end of the coding sequence of the nucleic acid sequence inherently contains a signal peptide coding region naturally linked in translation reading frame with the segment of the coding region that encodes the secreted polypeptide. Alternatively, in some embodiments, the 5ʹ end of the coding sequence contains a signal peptide coding region that is foreign to the coding sequence. Any suitable signal peptide coding region that directs the expressed polypeptide into the25971 secretory pathway of a host cell of choice finds use for expression of the engineered polypeptide(s). Effective signal peptide coding regions for bacterial host cells are the signal peptide coding regions include, but are not limited to, those obtained from the genes for Roseibacillus, Bacillus NClB 11837 maltogenic amylase, Bacillus stearothermophilus alpha- amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus stearothermophilus neutral proteases (nprT, nprS, nprM), and Bacillus subtilis prsA.
[0216] In some embodiments, regulatory sequences are also utilized. These sequences facilitate the regulation of the expression of the polypeptide relative to the growth of the host cell. Examples of regulatory systems are those that cause the expression of the gene to be turned on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include, but are not limited to, the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory systems include, but are not limited to, the ADH2 system or GAL1 system. In filamentous fungi, suitable regulatory sequences include, but are not limited to, the TAKA alpha-amylase promoter, Aspergillus niger glucoamylase promoter, and Aspergillus oryzae glucoamylase promoter.
[0217] In another aspect, the present disclosure provides a recombinant expression vector comprising a polynucleotide encoding a carboxylesterase polypeptide, and one or more expression regulating regions such as a promoter and a terminator, a replication origin, etc., depending on the type of hosts into which they are to be introduced. In some embodiments, the various nucleic acid and control sequences described herein are joined together to produce expression vectors that include one or more convenient restriction sites to allow for insertion or substitution of the nucleic acid sequence encoding the enzyme polypeptide at such sites. Alternatively, in some embodiments, the nucleic acid sequence of the present invention is expressed by inserting the nucleic acid sequence or a nucleic acid construct comprising the sequence into an appropriate vector for expression. In some embodiments involving the creation of the expression vector, the coding sequence is located in the vector so that the coding sequence is operably linked with the appropriate control sequences for expression.
[0218] The disclosed expression vector may be any suitable vector (e.g., a plasmid or virus), that can be conveniently subjected to recombinant DNA procedures and bring about the expression of the enzyme polynucleotide sequence. The choice of the vector typically depends on the compatibility of the vector with the host cell into which the vector is to be introduced. The vectors may be linear or closed circular plasmids.
[0219] In some embodiments, the expression vector is an autonomously replicating vector (i.e., a vector that exists as an extra-chromosomal entity, the replication of which is independent of25971 chromosomal replication, such as a plasmid, an extra-chromosomal element, a minichromosome, or an artificial chromosome). The vector may contain any means for assuring self-replication. In some alternative embodiments, the vector is one in which, when introduced into the host cell, it is integrated into the genome and replicated together with the chromosome(s) into which it has been integrated. Furthermore, in some embodiments, a single vector or plasmid, or two or more vectors or plasmids which together contain the total DNA to be introduced into the genome of the host cell, and / or a transposon is utilized.
[0220] In some embodiments, the expression vector contains one or more selectable markers, which permit easy selection of transformed cells. A “selectable marker” is a gene, the product of which provides for biocide or viral resistance, resistance to heavy metals, prototrophy to auxotrophs, and the like. Examples of bacterial selectable markers include, but are not limited to, the dal genes from Bacillus subtilis or Bacillus licheniformis, or markers, which confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol or tetracycline resistance. Suitable markers for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for use in filamentous fungal host cells include, but are not limited to, amdS (acetamidase; e.g., from A. nidulans or A. orzyae), argB (ornithine carbamoyltransferases), bar (phosphinothricin acetyltransferase; e.g., from S. hygroscopicus), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase; e.g., from A. nidulans or A. orzyae), sC (sulfate adenyltransferase), and trpC (anthranilate synthase), as well as equivalents thereof.
[0221] In some alternative embodiments, the expression vectors contain additional nucleic acid sequences for directing integration by homologous recombination into the genome of the host cell. The additional nucleic acid sequences enable the vector to be integrated into the host cell genome at a precise location(s) in the chromosome(s). To increase the likelihood of integration at a precise location, the integrational elements preferably contain a sufficient number of nucleotides, such as 100 to 10,000 base pairs, preferably 400 to 10,000 base pairs, and most preferably 800 to 10,000 base pairs, which are highly homologous with the corresponding target sequence to enhance the probability of homologous recombination. The integrational elements may be any sequence that is homologous with the target sequence in the genome of the host cell. Furthermore, the integrational elements may be non-encoding or encoding nucleic acid sequences. On the other hand, the vector may be integrated into the genome of the host cell by non-homologous recombination.
[0222] For autonomous replication, the vector may further comprise an origin of replication enabling the vector to replicate autonomously in the host cell in question. Examples of bacterial25971 origins of replication are ColE1 ori, P15A ori, and the origins of replication of plasmids pBR322, pUC19, pACYCl77 (which contains the P15A ori), or pACYC184 (which contains the P15A ori) permitting replication in E. coli, and pUB110, pE194, or pTA1060 permitting replication in Bacillus. The origin of replication may be one having a mutation which makes its functioning temperature-sensitive in the host cell (see e.g., Ehrlich, Proc. Natl. Acad. Sci. USA 75:1433
[1978] ), or a kanamycin resistance selection marker.
[0223] In some embodiments, more than one copy of a nucleic acid sequence of the present invention is inserted into the host cell to increase production of the gene product. An increase in the copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene with the nucleic acid sequence where cells containing amplified copies of the selectable marker gene, and thereby additional copies of the nucleic acid sequence, can be selected for by cultivating the cells in the presence of the appropriate selectable agent.
[0224] Many of the expression vectors for use in the present invention are commercially available. Suitable commercial expression vectors include, but are not limited to, Novagen’s pET® E. coli T7 expression vectors (Millipore Sigma) and the p3xFLAGTMTMexpression vectors (Sigma-Aldrich Chemicals). In some embodiments, a pET30a expression vector is used. Other suitable expression vectors include, but are not limited to, pBluescriptII SK(-) and pBK- CMV (Stratagene), and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen) or pPoly (See e.g., Lathe et al., Gene 57:193-201
[1987] ).
[0225] Thus, in some embodiments, a vector comprising a sequence encoding at least one variant carboxylesterase is transformed into a host cell in order to allow propagation of the vector and expression of the variant carboxylesterases. In some embodiments, the transformed host cell described above is cultured in a suitable nutrient medium under conditions permitting the expression of the variant carboxylesterases. Any suitable medium useful for culturing the host cells finds use in the present invention, including, but not limited to minimal or complex media containing appropriate supplements. In some embodiments, host cells are grown in HTP media. Suitable media are available from various commercial suppliers or may be prepared according to published recipes (e.g., in catalogues of the American Type Culture Collection). Methods of Producing Carboxylesterases
[0226] In some embodiments, the carboxylesterases of the present disclosure are obtained, evolved, or derived from a bacterial wild-type enzyme. Evolution (e.g., directed evolution) may be used to identify polypeptides of the present disclosure. For example, in some embodiments, to make a carboxylesterase polypeptide of the present disclosure, a carboxylesterase-encoding25971 polynucleotide was obtained (or derived) from Roseibacillus sp. In some embodiments, this parent polynucleotide sequence is codon-optimized to enhance expression of the carboxylesterase in a specified host cell. In the Examples, a parental wild-type polynucleotide sequence was codon optimized for E. coli expression, and its 3’ signal peptide-encoding sequence removed, to afford SEQ ID NO: 131. This carboxylesterase sequence was cloned into an expression vector, placing the expression of the carboxylesterase gene under the control of the lac promoter under control of the lac repressor. Clones expressing the active carboxylesterase in E. coli were identified, and the genes were sequenced to confirm their identity.
[0227] The carboxylesterase of the disclosure may be obtained by subjecting the polynucleotide encoding the parent sequence to mutagenesis and / or directed evolution (DE) methods. Any DE technique may be used, including crystal structure-guided library design, single-site-saturation mutagenesis (SSM) library generation, or combinatorial library generation, or a combination of these. An exemplary directed evolution technique is mutagenesis and / or DNA shuffling as described in Stemmer, 1994, Proc. Natl. Acad. Sci. USA 91:10747-10751; WO 95 / 22625; WO 97 / 20078; WO 97 / 35966; WO 98 / 27230; WO 00 / 42651; WO 01 / 75767 and U.S. Pat. No. 6,537,746. Other directed evolution procedures that can be used include, among others, staggered extension process (StEP), in vitro recombination (Zhao et al., 1998, Nat. Biotechnol.16:258- 261), mutagenic PCR (Caldwell et al., 1994, PCR METHODS APPL.3:S136-S140), and cassette mutagenesis (Black et al., 1996, Proc. Natl. Acad. Sci. USA 93:3525-3529).
[0228] Exemplary crystal structure-guided library design techniques may make use of computational modeling and / or may be rational or semi-rational. Examples of such techniques include molecular dynamics (MD) simulations, including in silico MD, Rosetta de novo design (e.g., trRosetta), Schrodinger toolbox and Glide modeling. TrRosetta is a deep neural net-based protein structure prediction method, offers a path forward for predicting 3D structures when there is a lack of close homologue structures, which takes advantage of multiple sequence alignment (MSA), where derived features from the MSA are fed into a deep neural network to predict inter- residue geometries, including distance and orientations. Glide modeling (Friesner et al. J. Med. Chem.2004, 47(7), 1739-1749) may be used to perform substrate docking.
[0229] The clones obtained following mutagenesis treatment are screened for carboxylesterases exhibiting one or more desired improved enzyme properties. These desired improved enzyme properties may be selected by applying each property, one-by-one as a selection pressure during iterative rounds of evolution. For instance, improved enzyme activity may be applied as a selection pressure during one or more rounds of evolution. Additional selection pressures that may be applied include improved thermostability, co-solvent tolerance, and reduced formation of25971 undesired by-products. For example, improved thermostability was applied as a selection pressure during evolution of the disclosed enzymes, as described in the Examples. As another example, reduced formation of hydrolysis by-product Compound X was applied as a selection pressure during evolution of the disclosed enzymes. Co-solvent tolerance may be applied as a selection pressure to mimic potential industrial process conditions.
[0230] Measuring enzyme activity from the expression libraries can be performed using standard chemistry analytical techniques for measuring substrates and products such as ultraperformance liquid chromatography-mass spectrometry (UPLC-MS). The reaction may also be run under conditions where the carboxylesterase is the yield-limiting catalyst, such that a doubling or halving in concentration of the carboxylesterase will result in a doubling or halving of the yield of product observed at a given timepoint.
[0231] Where the improved enzyme property desired is thermostability, enzyme activity may be measured after subjecting the enzyme preparations to a defined temperature and measuring the amount of enzyme activity remaining after heat treatments. Any suitable approach may be used, e.g., differential scanning colorimetry (DSC) a biochemical assay, or spectroscopy. Clones containing a polynucleotide encoding carboxylesterases are then isolated, sequenced to identify the nucleotide sequence changes (if any), and used to express the enzyme in a host cell.
[0232] Where the sequence of the polypeptide is known, the polynucleotides encoding the enzyme can be prepared by standard solid-phase methods, according to known synthetic methods. In some embodiments, fragments of up to about 100 bases can be individually synthesized, then joined (e.g., by enzymatic or chemical litigation methods, or polymerase mediated methods) to form any desired continuous sequence. For example, polynucleotides and oligonucleotides of the invention can be prepared by chemical synthesis using, e.g., the classical phosphoramidite method described by Beaucage et al., 1981, Tet. Lett.22:1859-69, or the method described by Matthes et al., 1984, EMBO J.3:801-05, e.g., as it is typically practiced in automated synthetic methods. According to the phosphoramidite method, oligonucleotides are synthesized, e.g., in an automatic DNA synthesizer, purified, annealed, ligated and cloned in appropriate vectors. In addition, essentially any nucleic acid can be obtained from any of a variety of commercial sources, such as Integrated DNA Technologies (IDT), Genscript, The Midland Certified Reagent Company, Midland, Tex., The Great American Gene Company, Ramona, Calif., ExpressGen Inc. Chicago, Ill., Operon Technologies Inc., Alameda, Calif., and many others.
[0233] Carboxylesterase enzymes expressed in a host cell can be recovered and isolated from the cells and / or the culture medium using any one or more of the well-known techniques for25971 protein purification, including, among others, lysozyme treatment, sonication, filtration, salting- out, ultra-centrifugation, and chromatography. Suitable solutions for lysing and the high efficiency extraction of proteins from bacteria, such as E. coli, are commercially available under the trade name CELLYTIC B® from Sigma-Aldrich.
[0234] Chromatographic techniques for isolation of the carboxylesterase polypeptide include, among others, reverse phase chromatography, high performance liquid chromatography, ion exchange chromatography, gel electrophoresis, and affinity chromatography. Conditions for purifying a particular enzyme will depend, in part, on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, etc., and will be apparent to those having skill in the art.
[0235] In some embodiments, affinity techniques may be used to isolate the improved carboxylesterase enzymes. For affinity chromatography purification, the protein sequence can be tagged with a recognition sequence to enable purification. Common tags include cellulose- binding domains, poly His-tags (e.g., 6xHis tags), di-His chelates, FLAG-tags and many others that will be apparent to those having skill in the art. Antibodies can also be used as affinity purification reagents. Any antibody that specifically binds the carboxylesterase polypeptide may be used. Methods of Generating Cyclized Oligopeptides
[0236] Provided herein are methods (or processes) of generating cyclized oligopeptides via one or more carboxylesterase-catalyzed reactions. These methods comprise performing a reaction with any of the disclosed engineered polypeptides. These methods may comprise the conversion of a linear tetrapeptide into a cyclized tetrapeptide (through amide bond formation and acyl transfer). In various aspects, the product of these methods is a cyclic tetrapeptide.
[0237] Any of these methods may be performed in a medium containing cofactor magnesium ion (Mg2+), for instance in a magnesium phosphate or magnesium chloride salt. Any of these methods may be performed in a medium containing sodium phosphate, a lysozyme and / or polymyxin B sulfate. In some embodiments, the medium contains co-solvent acetonitrile, e.g., 15%, 20%, 25%, or 30% acetonitrile. In some embodiments, the medium is maintained at a neutral pH, e.g., a pH of 7.4 or a pH of 7. In some embodiments, the medium is maintained at a temperature of 20 °C.
[0238] In some aspects, provided are methods for preparing Compound 1, the method comprising contacting compound 3k in the presence of Mg2+and any of the disclosed polypeptides. These methods may produce Compound 3 reduced amounts relative to a25971 corresponding method in which compound 3k is contacted with the polypeptide of SEQ ID NO: 132.
[0239] In some embodiments, the disclosed methods exhibit higher (1) co-solvent tolerance and / or (2) thermostability relative to a corresponding method in which Compound 3k is contacted with the polypeptide of SEQ ID NO: 132. Example 18 Synthesis, Expression, and Assay of carboxyesterases with macrocyclization activity
[0240] This example describes methods to synthesize, codon optimize, and assay carboxylesterase enzyme and activity, and the composition of optimized enzymes. Gene synthesis and construction of expression strain:
[0241] The wild-type carboxylesterase polypeptide UniProt ID: A0A2E5D605_9BACT belonging to an unclassified Roseibacillus species) was identified by a BLASTP search using a previously reported promiscuous wild-type hydrolase EstCE1 (Müller et al., Angew. Chem. Int. Ed. (2021) 60, 2013). Using the DeepSig algorithm (Savojardo et al., Bioinformatics, 2018, 34, 1690), the first 15 amino acids were predicted to be a signal peptide region of carboxylesterase and were therefore excluded from the design for gene synthesis to promote cytoplasmic expression within E. coli. The encoding carboxylesterase polynucleotide without the signal peptide encoding region was codon optimized for expression in E. coli and synthesized with a sequence encoding an N-terminal 6xHistidine tag as SEQ ID NO: 131. SEQ ID NO: 131 is 74.6% identical to the wild-type enzyme CARBOXYLESTERASE-encoding polynucleotide sequence (GenBank ID: PBAI01000020.1). SEQ ID NO: 131 was cloned into pET30 vector under the control of T7 promoter for expression and transformed into E. coli strain (BL21(DE3)). HTP Growth, Expression, and Lysate Preparation:
[0242] The transformants expressing SEQ ID NO:131 as carboxylesterase polypeptide (SEQ ID NO: 132) were picked and grown in Luria-Bertani Broth medium with kanamycin (30 μg / ml) and 1% w / v glucose in 96-well plates with shaking (200 rpm) at 30 °C overnight. Subsequently, the overnight cultures were diluted for expression to optical density of 0.05, measured at 600 nm (OD600), with auto-induction medium (ZYM-5052, Teknova) and antibiotic (kanamycin 30 μg / ml) and grown for 18-20 hours at 30 °C. Post-induction, cells were centrifuged and cell pellet was resuspended in the lysis buffer (50 mM sodium phosphate buffer pH 7.4, 0.5 mg / mL lysozyme, 0.5 mg / mL polymixin B sulfate, 1 U / mL DNAse I, 1 mM magnesium sulfate) at 25 °C with shaking at 1000 rpm for 1 hour. The cell lysate was clarified by centrifugation (4000 x g, 15 minutes). The clarified cell lysates were used in the following well-plate enzymatic reactions. In parallel to the dilution of overnight cultures for expression, bacterial cultures were diluted with25971 25% v / v glycerol and frozen at -80 ℃ for storage. Overnight cultures were also used as PCR templates to amplify the coding region for the carboxylesterase, which were then subjected to sequencing to confirm or determine the sequence identity of the carboxylesterase in each well of the 96-well plates. Assay method for Carboxylesterase activity:
[0243] For a 50 µL volume reaction (scheme below), 25 µL of 50 mM sodium phosphate buffer (pH 7.4) was mixed with 5 µL of substrate (linear tetrapeptide, compound 3k, 5 mg / mL dissolved in 50% DMSO) and 20 µL of clarified lysate. Reaction mixtures were incubated at 30 ℃ with shaking at 600 rpm for 18 hours. The reactions were quenched with 150 µL of acetonitrile and mixed thoroughly. Samples were further filtered through 0.22 µm filters before UHPLC-UV analysis. Reaction scheme 1:
[0244] The conversion of compound 3k to compound 1 in the enzymatic reaction was determined by UHPLC-UV using a Thermo Hypersil Gold PFP column (3.0 x 50 mm, 1.9 µm particle size). Mobile phase 0.1% difluoroacetic acid (A) / acetonitrile (B) with the gradient program described below, injection volume 1 µL, flow rate 1 mL min-1, detection wavelength 280 nm, and column temperature of 40 °C. Retention time of the substrate (Compound 3k): 1.093 minutes; Retention time of the desired product (Compound 1): 0.876 minutes; Retention time of the hydrolysis product (Compound X): 0.697 minutes. UHPLC gradient Program Time [min] Mobile phase A Mobile phase B25971 Identification of starting point for evolution:
[0245] Activity was detected for the carboxylesterase enzyme of SEQ ID NO: 132 catalyzing macrocyclization of tetrapeptides (conversion of Compound 3k to Compound 1) with the desired regiochemistry. However, the conversion was very low even under high percentage of cellular lysate loading, limiting industrial applications. In addition, SEQ ID NO: 132 also generates a hydrolysis product (Compound X), reducing the reaction yield. Therefore, it was determined that protein engineering should be initiated to improve the activity and selectivity (acyl transfer over hydrolysis) of SEQ ID NO: 132. Table A, Compound structures: ID Structure Note Compound Linear tetrapeptidetetrapeptide that contains modified proline, tryptophan, and phenylalanine residues. Product Compound 1 is a cyclized version of Compound 3k. Both compounds have a molecular weight of 793.94 g / mol.25971 Directed evolution strategy and summary of results
[0247] The activity (e.g., conversion, total turnover number) of the wild-type carboxylesterase enzyme, SEQ ID NO: 132, was insufficient for industrial macrocyclization of tetrapeptide. Therefore, a directed evolution strategy was developed to engineer SEQ ID NO: 132 to improve the activity, thermostability, organic solvent tolerance, and selectivity (acyl transfer over hydrolysis). To prepare for high-throughput screening of enzyme variant libraries, protein expression conditions were first optimized to maximize soluble protein expression. Enzymatic reactions (Reaction scheme 1) were also scaled down to the 96-well microtiter plate format. A 5- minute UHPLC method was developed to quantify the conversions to enable library-scale screening of carboxylesterase variants, as described above.
[0248] In the directed evolution campaign, single-site-saturation mutagenesis (SSM) libraries were designed, built, expressed, and screened for desired properties relative to the starting enzyme (Table 1.1). SSM libraries in each round of evolution were generated using splicing by overlap extension (SOEing) PCR methods described in Ho et al., Gene 1989, 77(1), 51-59. In short, mutations at the designated positions were incorporated by degenerate oligonucleotides through overlapping PCRs. The full-length gene-of-interest region was further amplified and assembled into the expression vector (pET30a) by Gibson Assembly (Gibson et al. Nat. Methods 2009, 6(5), 343-345). The combinatorial libraries were built following the instructions from the QUIKCHANGE® Lightning Multi Site-Directed Mutagenesis kit (Agilent Technologies). Both mutagenesis libraries were transformed into the BL21(DE3) E.coli strains by electroporation and plated on LB agar plates with selection (1% w / v glucose, 30 μg / mL kanamycin). Table 1.1. Summary of Carboxylesterase evolution Round Screening conditions Library Target- ed Mutagenesis residues Selection -25971 Round Screening conditions Library Target- Selection designs ed Mutagenesis residues residue pressure - - - - -25971 Round Screening conditions Library Target- Selection designs ed Mutagenesis residues residue pressure - - - , - ,25971 Round Screening conditions Library Target- e Selection designs d Mutagenesis residues residue pressure - , - , ty - ty -25971 Round Screening conditions Library Target- Selection designs ed Mutagenesis residues residue pressure ty - ty - ty -25971 ound Screening conditions Li Target- R brary ed Muta Selection designs genesis residues residue pressure ty - ty - ty -
[0249] To prioritize the sites for SSM library designs, a structural model of SEQ ID NO: 132 was built using the Protein Data Bank (PDB) template 4IVK, which had 52.5% sequence similarity. Schrodinger toolbox was used to design the homology model, which was further refined by adding hydrogens using PROPKA (Olsson et al. J. Chem. Theory Comput.2011, 7(2), 525-537) and by running a restrained minimization to converge heavy atoms to a maximum root mean square deviation of 0.30 Å. Substrate docking was performed using Glide (Friesner et al. J. Med. Chem.2004, 47(7), 1739-1749) and the best pose was selected after visually inspecting a pool of docked poses rank ordered by their docking energy scores and filtering them based on the distances between the reaction site and the catalytic residues.
[0250] The protein homology model was partitioned into four regions: active site, surface1, surface2, and the remaining residues. The active site shell residues were selected using a distancecriterion, i.e., all residues within 10.5 Å of the docked substrate, resulting in 96 sites. The remaining residues were rank ordered by their solvent accessibilities and grouped into sets of 96 sites. The top two sets were classified as surface site1 (positions with solvent accessible surface area > 64 Å) and site2, (positions with solvent accessible surface area <=64 Å and > 8 Å) consisting of high surface accessibility sites, whereas the remaining were classified as the remaining residues. In the initial 4 rounds, each tier per round was sequentially targeted to screen all the residues beginning with the active site proximal, followed by surface residues, and finally to the remaining residues. In subsequent rounds, residues were prioritized to target either based on structural insights or MD (molecular dynamics) simulation (see Table 1.1). Beneficial mutations identified from these SSM libraries were recombined into combinatorial libraries for subsequent screening in the following rounds. The best variant served as the backbone for the next round of mutagenesis. Additional details regarding the rounds of evolution are provided below in Examples 20 to 31.
[0251] From round 1 to round 3 of evolution (Table 1.1), a series of mutations (I159Y, Y154M, N156P, G170W, L172C, L199G, N132R, K352T, M368A) were identified relative to SEQ ID NO: 132 that improved product formation under screening conditions by more than 50-fold. As the activity continued to improve through evolution, acetonitrile was included as a co-solvent to engineer variants with higher tolerance to organic solvent, which reflects potential industrial process conditions. In rounds 4 and 5, additional mutations (R133M, T138R, M154N, K184V, A191D, M327Y, V115I, T140P, C147R, G158E, F373Y) were identified relative to SEQ ID NO: 138, which improved product formation under organic solvent exposure by more than 80-fold. From round 6 to round 8, heat-treatment of enzymes prior to activity assays was also included to engineer more thermostable variants. Mutations (R330S, P160L, I162G, A272H, S350D, R17Y, F134Q, A165P, T265S, D266L, P325S) were further identified relative to SEQ ID NO: 142, which improved product formation under screening conditions by more than 20-fold. From round 9 to round 13, improvement of acyltransferase activity over hydrolysis activity was prioritized to increase the overall yield of desired product Compound 1. Mutations (G199S, Q134F, L266D, H166W, G281S, L343F, A368L, I115L, I157F, W170L, C172F, L187Q, T10K, G162I, T246H, H272A, F343L) were further identified relative to SEQ ID NO: 148, which improved selectivity ([Compound 1] / [Compound X]) under screening conditions by more than 10-fold (see Table 9.1). Ultimately, this evolution campaign led to the engineering of SEQ ID NO: 158, which had improved activity, organic solvent tolerance, thermo-stability and selectivity in comparison to the wild-type parent (SEQ ID NO: 132).Example 19 (Rd1 evolution of SEQ ID NO: 132) Enzyme variants of SEQ ID NO: 132
[0252] In this example, engineering of SEQ ID NO: 132 for improved activity and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites were subjected to mutagenesis resulting in library 1.1 (Example 18, Table 1.1). These libraries were plated to form single colonies, which were picked, grown, expressed, and screened using the high-throughput growth and expression methods described in Example 18. The high-throughput assay conditions are described below. HTP analytical methods for enzyme activity in round 1:
[0253] The conversion in the enzymatic reaction was determined by the method described in Example 1 or using a Waters Acquity BEH C18 column (2.1 x 50 mm, 1.7 μm particle size). Mobile phase (A) 0.1% (v / v) formic acid in water, (B) 0.1% (v / v) formic acid in acetonitrile, 1 µL injection volume, flow rate 0.75 mL min-1, detection wavelength of 210 nm and 280 nm at 55 °C column temperature. Retention time of the substrate (Compound 3k): 2.893 minutes; Retention time of the desired product (Compound 1): 2.624 minutes; Retention time of the undesired hydrolysis product (Compound X): 2.021 minutes. Gradient for achiral HPLC method. Time [min] Mobile phase A Mobile phase B 0.00 95 5HTP activity assay for Carboxylesterase variants:
[0254] For a 50 µL volume reaction, 30 µL of the 50mM Sodium phosphate buffer (pH 7.4) was mixed with 10 µL tetrapeptide (Compound 1, 5 mg / mL dissolved in 50% DMSO) and 10 µL of cellular lysate. Reaction mixtures were incubated at 30℃ with shaking at 600 rpm for 18 hours. Production of Shake Flask Powders (SFP):
[0255] Based on the analysis of the HTP assay results using the methods described above, variants of SEQ ID NO: 131 / 132 were selected for larger scale analysis and further characterization. About 10 μL of glycerol stock or a freshly streaked single colony for each25971 variant was inoculated into 30 mL Luria Broth media with 30 μg / mL kanamycin and 1% w / v glucose in a 125 mL baffled flask and incubated at 30 °C with shaking (250 rpm) for 18 hours. The next day, the overnight saturated culture was diluted to an initial optical density of 0.05 (measured at 600 nm (OD600)) with Terrific Broth supplemented with 30 μg / mL kanamycin in 250 mL volume and grown at 37 °C with shaking (250 rpm). When the growth reached OD600 of 0.5, IPTG was added to the culture to a final concentration of 1 mM to induce protein production and the culture was further incubated for 20 hours at 30 °C. Cells were then collected by centrifugation and resuspended in 5 mL of buffer (20mM Triethanolamine, pH 7.5) per gram of cell pellet mass. The cell pellet was resuspended by shaking (250 rpm) at 8 °C for 25 minutes. Cells were lysed by the microfluidizer (LM-10) at the pressure of 16,000 psi. The resulting lysate was clarified by centrifugation at 22,000xg for 45 minutes at 4 °C, and the supernatant was subsequently frozen and lyophilized to generate enzyme powders.
[0256] Engineered polypeptides with >2-fold conversion relative to the parent polypeptide are listed in Table 2.1, and the activities were measured with the shake flask powder samples. Table 2.1 Variants and Conversion SEQ ID NOs (nucleotide Amino acid differences Increased conversion (a) sequence (nt) / amino acid ofSEQ ID NO: 132 and defined as: + = conversion at least 2-fold relative to reference polypeptide
[0258] Variants with mutations I159Y and G155P produced more macrocyclic peptide product (Compound 1) from the tetrapeptide substrate (Compound 3k), and these engineered carboxylesterase enzymes provide new biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations I159Y (SEQ ID NO: 134), had the preferable activity. Thus, the encoding polynucleotide (SEQ ID NO: 133) was selected for further directed evolution. Example 20 (Rd2 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 134
[0259] In this example, engineering of SEQ ID NO: 134 for improved activity toward macrocyclization and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated25971 with active sites and dynamic loop regions were subjected to mutagenesis individually and combinatorically resulting in libraries 2.1 and 2.2 (Example 18, Table 1.1). These libraries were plated to form single colonies, which were picked, grown, and screened using the high- throughput growth and expression methods described in Example 18 and the HTP assay and analytical method described in Example 19. HTP assay for Carboxylesterase activity:
[0260] Engineered polypeptides with >5-fold conversion relative to the parent polypeptide are listed in Table 3.1, and the activities were measured with shake flask powder samples produced as described in Example 19. Table 3.1 Variant and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased conversion (a) SEQ ID NO: 135 / 136 Y154M, N156P, G170W, +of SEQ ID NO: 134 and defined as: “+” = conversion at least 5-fold relative to reference polypeptide.
[0262] SEQ ID NO: 136 with mutations Y154M, N156P, G170W, L172C, L199G relative to SEQ ID NO: 134 produced more macrocyclic peptide (Compound 1) from the tetrapeptide (Compound 3k), and this engineered carboxylesterase enzyme was useful as a biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations Y154M, N156P, G170W, L172C, L199G (SEQ ID NO: 136), had the highest activity in the assayed library. Thus, the encoding polynucleotide (SEQ ID NO: 135) was selected for further engineering. Example 21 (Rd3 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 136
[0263] In this example, engineering of SEQ ID NO: 136 for improved activity toward macrocyclization and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites were subjected to mutagenesis individually and combinatorically resulting in libraries 3.1 and 3.2 (Example 18, Table 1.1). These libraries were plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 18. The high-throughput assay conditions are described below.25971 The production of shake flask powders and HTP analytical methods were described in Example 19. HTP assay for Carboxylesterase activity:
[0264] For a 50 µL volume reaction, 30 µL of the 50mM Sodium phosphate buffer (pH 7.4) was mixed with 10 µL substrate (tetrapeptide Compound 3k at 25 mg / mL dissolved in 25% v / v DMSO in water) and 10 µL of clarified cellular lysate. Reaction mixtures were incubated at 40 ℃ with shaking at 600 rpm for 18 hours.
[0265] Engineered polypeptides with >3-fold conversion relative to the parent polypeptide are listed in Table 4.1 as measured with the shake flask powder samples. Table 4.1 Variant and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased conversion (a) SEQ ID NO: 137 / 138 N132R, K352T, M368A +of SEQ ID NO: 136 and defined as: “+” = conversion at least 3-fold relative to reference polypeptide.
[0267] Engineered variant SEQ ID No: 162 with mutations N132R, N132K, R133M, G158E, I162W, L187K, K352T, M368A produced more macrocyclic peptide (Compound 1) from the tetrapeptide (Compound 3k), and this engineered carboxylesterase enzyme provided a new biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations N132R, K352T, M368A (SEQ ID NO: 138) had the highest activity in the library. Thus, the encoding polynucleotide (SEQ ID NO: 137) was selected for further directed evolution. Example 22 (Rd4 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 138
[0268] In this example, engineering of SEQ ID NO: 138 for improved activity toward macrocyclization with organic co-solvent ACN and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with solvent accessibility and active sites were subjected to mutagenesis individually and combinatorically resulting in libraries 4.1 and 4.2 (Example 18, Table 1.1). These libraries were plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 18.25971 The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 19. HTP assay for Carboxylesterase activity:
[0269] For a 50 µL volume reaction, 40 µL of assay mixtures (75 mM sodium phosphate buffer pH 7.4, 6.25 mg / mL of tetrapeptide (Compound 3k), and 18.75% v / v of acetonitrile) was added with 10 µL of clarified cellular lysate, resulting in 15% v / v acetonitrile concentration in the reactions. Reaction mixtures were incubated at 20 ℃ with shaking at 600 rpm for 18 hours.
[0270] Engineered polypeptides with >7-fold conversion relative to the parent (reference) polypeptide are listed in Table 5.1 as measured with shake flask powder samples. Table 5.1 Variants and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased conversion (a) SEQ ID NO: 139 / 140 R133M, T138R, M154N, +of SEQ ID NO: 138 and defined as: “+” = conversion at least 7-fold relative to reference polypeptide
[0272] Engineered variant Seq ID No: 140 with mutations R133M, T138R, M154N, K184V, A191D, M327Y produced more macrocyclic peptide (Compound 1) from the tetrapeptide (Compound 3k) with the organic co-solvent, and this engineered carboxylesterase enzyme provide a new biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations R133M, T138R, M154N, K184V, A191D, M327Y (SEQ ID NO: 140), had the highest activity. Thus, the encoding polynucleotide (SEQ ID NO: 139) was selected for further directed evolution. Optimization on the fermentation of the Carboxylesterase
[0273] In order to utilize carboxylesterase in macrocyclization of peptides at scale, the production of carboxylesterase at reduced cost is desirable. Therefore, we set out to examine the fermentation conditions in E.coli (BL21(DE3)) during round 5 of evolution. SEQ ID NO: 140 was fermented at 5L scale with complex medium. Inconsistent with the expression results at the shake flask scale, the expression of carboxylesterase was not over-expressed at 5L fermentation. Leaky expression of proteins without addition of inducers has been reported in bacterial fermentation with complex medium. Unregulated protein over-expression in cells may lead to potential metabolic burden and toxicity for growth. Therefore, we optimized the fermentation conditions by switching the medium from complex medium to chemical defined medium (CDM),25971 which suppresses leaky protein expression in cells. Fermentation with the optimized conditions resulted in a higher expression level of carboxylesterase, which is useful for large scale production of carboxylesterase for industrial applications. Example 23 (Rd5 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 140
[0274] In this example, engineering of SEQ ID NO: 140 for improved activity toward macrocyclization with organic co-solvent ACN and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with improved activity with or without co-solvent from previous rounds were subjected to mutagenesis combinatorically, resulting in library 5.1 (Example 18, Table 1.1). This library was plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 18. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 19. Assay for Carboxylesterase activity:
[0275] In the high-throughput screening, 40 µL of assay mixture (225 mM sodium phosphate buffer, 112.5 mM HEPES, 31.2 mM magnesium chloride, 6.25 mg / mL of tetrapeptide Compound 3k, 43.75% v / v acetonitrile in water, pH 7) was mixed with 10 µL of 10-fold diluted clarified cellular lysate, resulting in 35% v / v acetonitrile in the reactions. Reactions were incubated at 20 ℃ with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate concentration in 200 mM sodium phosphate buffer pH 7 and 15% v / v acetonitrile with 15 % (wt enzyme / wt substrate) of carboxylesterase at 30℃ with shaking at 600 rpm for 12 hours.
[0276] The engineered polypeptide with >8-fold conversion relative to the parent polypeptide under these conditions is listed in Table 6.1, as measured with the shake flask powder samples. Table 6.1 Variants and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased conversion (a)e e s o c ease co e s o e e e e e e a e o e e e e ce po ypep e of SEQ ID NO: 140 and defined as: “+” = conversion at least 8-fold relative to reference polypeptide.25971
[0278] Engineered variant SEQ ID NO: 142 with mutations V115I, T140P, C147R, G158E, F373Y produced more macrocyclic peptide (Compound 1) from the tetrapeptide (Compound 3k) with the organic co-solvent, and this engineered carboxylesterase enzyme provided a new biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides with organic co-solvent. The variant with mutations V115I, T140P, C147R, G158E, F373Y (SEQ ID NO: 142), had the highest activity. Thus, the encoding polynucleotide (SEQ ID NO: 141) was selected for further directed evolution. Example 24 (Rd6 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 142
[0279] In this example, engineering of SEQ ID NO: 142 for improved activity toward macrocyclization with organic co-solvent (solvent tolerance) and heat treatment (thermostability) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which 96 positions associated with solvent accessible / surface residues were subjected to mutagenesis resulting in library 6.1 (Example 18, Table 1.1). This library was plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 18. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 19. Assay for Carboxylesterase activity:
[0280] In the high-throughput screening, 40 µL of assay mixture (225 mM sodium phosphate buffer, 112.5 mM HEPES, 31.2 mM magnesium chloride, 6.25 mg / mL of tetrapeptide Compound 3k, 25% v / v acetonitrile, pH 7) was mixed with 10 µL of 10-fold diluted clarified cellular lysate that had been pre-incubated at 46.2 ℃ for 1 hour, resulting in 20% v / v acetonitrile in the reactions. Reaction mixtures were incubated at 20 ℃ with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate concentration in 200 mM sodium phosphate buffer pH 7, 30% v / v acetonitrile, 100 mM HEPES, and 25 mM magnesium chloride with 5% wt enzyme / wt substrate at 30 ℃ with shaking at 600 rpm for 14 hours.
[0281] Engineered polypeptides with >2-fold conversion relative to the parent polypeptide are listed in Table 7.1, as measured with the shake flask powder samples.25971 Table 7.1 Variants and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased conversion (a) SEQ ID NO: 143 / 144 R330S + ofSEQ ID NO: 142 and defined as:“+” = conversion at least 2-fold relative to reference polypeptide
[0283] Engineered variants with SEQ ID NO: 144 with mutation R330S produced more macrocyclic peptide (Compound 1) from the tetrapeptide (Compound 3k), and this engineered carboxylesterase enzyme serves as a new biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides in the presence of organic co-solvent and at various temperatures. The variant with mutations R330S (SEQ ID NO: 144), had the highest activity in library 6.1. Thus, the encoding polynucleotide (SEQ ID NO: 143) was selected for further directed evolution. Example 25 (Rd7 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 144
[0284] In this example, engineering of SEQ ID NO: 144 for improved activity toward macrocyclization and improved stability (e.g., stability in the presence of organic co-solvent and elevated temperature) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with solvent accessible surface residues were subjected to mutagenesis individually and combinatorically resulting in libraries 7.1 and 7.2 (Example 18, Table 1.1). These libraries were plated to form single colonies, which were picked, grown, and screened using the high- throughput growth and expression methods described in Example 18. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 19. Assay for Carboxylesterase activity:
[0285] In the high-throughput screening, 40 µL of assay mixtures (225 mM sodium phosphate buffer, 112.5 mM HEPES, 31.2 mM magnesium chloride, 6.25 mg / mL of tetrapeptide Compound 3k, 25% v / v acetonitrile, pH 7) was added with 10 µL of 10-fold diluted cellular lysate that was pre-incubated at 50 ℃ for 1 hour, resulting in 20% v / v acetonitrile in the reactions. Reactions were incubated at 20 ℃ with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate concentration in 200 mM sodium25971 phosphate buffer pH 7, 15% v / v acetonitrile, 100 mM HEPES, and 25 mM magnesium chloride with 0.25% wt enzyme / wt substrate at 30 ℃ with shaking at 600 rpm for 14 hours.
[0286] Engineered polypeptides with >4-fold conversion relative to the parent polypeptide are listed in Table 8.1, and the activities were measured with the shake flask powder samples. Table 8.1 Variants and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased conversion (a) SEQ ID NO: 145 / 146 P160L, I162G, A272H, + ofSEQ ID NO: 144 and defined as: “+” = conversion at least 4-fold relative to reference polypeptide.
[0288] Engineered variants SEQ ID NO: 146 and 164 with mutations R17D, P160L, I162G, A272H, S350D produced more macrocyclic peptide (Compound 1) from the tetrapeptide (Compound 3k), and these engineered carboxylesterase enzymes are new biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides with organic co-solvent and at high temperature. The variant with mutations P160L, I162G, A272H, S350D (SEQ ID NO: 146), had the highest activity in libraries 7.1 and 7.2. Thus, the encoding polynucleotide (SEQ ID NO: 145) was selected for further directed evolution. Example 26 (Rd8 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 146
[0289] In this example, engineering of SEQ ID NO: 146 for improved activity toward macrocyclization and improved stability (e.g., stability in the presence of organic co-solvent and elevated temperature) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with improvements in activity or stability in previous rounds were subjected to mutagenesis combinatorically resulting in library 8.1 (Example 18, Table 1.1). This library was plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 18. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 19.25971 Assays for Carboxylesterase activity:
[0290] In the high-throughput screening, 40 µL of assay mixtures (225 mM sodium phosphate buffer, 112.5 mM HEPES, 31.2 mM magnesium chloride, 6.25 mg / mL of tetrapeptide (Compound 3k), 37.5% v / v acetonitrile, pH 7) was added with 10 µL of 10-fold diluted clarified cellular lysate that was pre-incubated at 48 ℃ for 1 hour, resulting in 30% v / v of acetonitrile concentration in the reactions. Reactions were incubated at 20 ℃ with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate concentration in 200 mM sodium phosphate buffer pH 7, 20% v / v acetonitrile, 100 mM HEPES, and 25 mM magnesium chloride with 0.5% wt enzyme / wt substrate of enzyme at 30 ℃ with shaking at 600 rpm for 14 hours.
[0291] Engineered polypeptides with >2-fold conversion relative to the parent polypeptide are listed in Table 9.1, and the activities were measured with the shake flask powder samples. Table 9.1 Variant and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased conversion (a) SEQ ID NO: 148 R17Y, F134Q, A165P, + ofQ : an e ne as: = convers on a eas - o re a ve o re erence polypeptide.
[0293] Engineered variant SEQ ID NO: 148 with mutations R17Y, F134Q, A165P, T265S, D266L, P325S produced more macrocyclic peptide (Compound 1) from the tetrapeptide (Compound 3k), and this engineered carboxylesterase enzyme provides a new biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides with organic co-solvent and at various temperatures. The variant with mutations R17Y, F134Q, A165P, T265S, D266L, P325S (SEQ ID NO: 148), had the highest activity in library 8.1. Thus, the encoding polynucleotide (SEQ ID NO: 147) was selected for further directed evolution. Example 27 (Rd9 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 148
[0294] In this example, engineering of SEQ ID NO: 148 for improved selectivity toward macrocyclization relative to hydrolysis and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites were subjected to mutagenesis combinatorically resulting in library 9.1 (Example 18, Table 1.1). This library was plated to form single colonies,25971 which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 18. The high-throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described in Example 19. Assays for Carboxylesterase activity:
[0295] In the high-throughput screening, 50 µL reactions (200 mM sodium phosphate buffer, 100 mM HEPES, 25 mM magnesium chloride, 25 mM tetrapeptide Compound 3k, pH 7, 0.4% v / v clarified cellular lysates) were incubated at 20 ℃ with shaking at 600 rpm for 18 hours. For assays with SFP samples, reactions were performed at 5 g / L substrate in 200 mM sodium phosphate buffer pH 7, 100 mM HEPES, and 25 mM magnesium chloride with 0.06 % wt enzyme / wt substrate at 30 ℃ with shaking at 600 rpm for 14 hours.
[0296] Engineered polypeptides with >1.2-fold relative to the parent polypeptide are listed in Table 10.1, and the activities were measured as shake flask powder samples. Table 10.1 Variants and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased conversion (a) SEQ ID NO: 149 / 150 Q134F, G199S, L266D +of SEQ ID NO: 148 and defined as: “+” = conversion at least 1.2-fold relative to the reference polypeptide.
[0298] Engineered variant SEQ ID NO: 150 with mutations Q134F, G199S, L266D produced more macrocyclic peptide (Compound 1) from the tetrapeptide (Compound 3k), and this engineered carboxylesterase enzyme provides a new biocatalytic reagent for use in new methods for the macrocyclization reaction of polypeptides with industrially relevant reaction conditions. The variant with mutations Q134F, G199S, L266D (SEQ ID NO: 150), had the highest activity in library 9.1. Thus, the encoding polynucleotide (SEQ ID NO: 149) was selected for further directed evolution. Example 28 (Rd10 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 150
[0299] In this example, engineering of SEQ ID NO: 150 for improved selectivity against hydrolysis of tetrapeptide (Compound 1) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions that were identified as beneficial in previous rounds were subjected to25971 mutagenesis combinatorically and positions that had not been targeted in previous evolution rounds were subjected to mutagenesis individually, resulting in libraries 10.1 and 10.2, respectively (Example 18, Table1.1). These libraries were plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 18. The production of shake flask powders and HTP analytical methods were described in Example 19. Assays for Carboxylesterase activity:
[0300] In the high-throughput screening, 50 µL reactions (200 mM sodium phosphate buffer, 100 mM HEPES, 25 mM magnesium chloride, 25 mM tetrapeptide compound 3k, pH 7, 0.3% v / v clarified cellular lysates) were incubated at 15 ℃ with shaking at 600 rpm for 3 hours. For assays with SFP samples, reactions were performed at 37.5 mM tetrapeptide in 200 mM sodium phosphate buffer pH 7.5, 100 mM HEPES and with 0.15 g / L of enzyme at 20℃ with shaking at 600 rpm for 30 minutes.
[0301] Engineered polypeptides with >1.5-fold selectivity, defined as the ratio of the conversion % of substrate to the product (Compound 1) over the conversion % of substrate to the by-product (Compound X), relative to the parent polypeptide, are listed in Table 11.1. The selectivity was measured with the shake flask powder samples. Table 11.1 Variants and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased selectivity (a) SEQ ID NO: 151 / 152 H166W G281S L343F +of SEQ ID NO: 150 and defined as: “+” = selectivity at least 1.5-fold relative to reference polypeptide.
[0303] Engineered variant SEQ ID NO: 152 with mutations H166W, G281S, L343F produced less hydrolysis by-products (Compound X) from the tetrapeptide (compound 3k), and this engineered carboxylesterase enzyme provides a more selective biocatalytic reagent for use in new methods for the macrocyclization of polypeptides. The variant with mutations H166W, G281S, L343F (SEQ ID NO: 152), had the highest selectivity in libraries 10.1 and 10.2. Thus, the encoding polynucleotide (SEQ ID NO: 151) was selected for further directed evolution. Example 29 (Rd11 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 152
[0304] In this example, engineering of SEQ ID NO: 152 for improved selectivity toward macrocyclization against hydrolysis of tetrapeptide (Compound 3k) and the resulting improved25971 variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites and surface residues were subjected to mutagenesis individually resulting in libraries 11.1 and 11.2. These libraries were plated to form single colonies, which were picked, grown, and screened using the high- throughput growth and expression methods described in Example 18 and below. The high- throughput assay conditions are described below. The production of shake flask powders and HTP analytical methods were described below and in Example 19. HTP Growth, Expression, and Lysate Preparation:
[0305] The colonies from libraries 11.1 and 11.2 were picked and grown in Luria-Bertani Broth medium with kanamycin (50 μg / mL) in 96 deep-well plates with shaking (400 rpm) at 30 °C overnight. Subsequently, the overnight cultures were diluted 1:100 with Terrific Broth medium containing antibiotic (kanamycin 50 μg / mL) and grown to cellular optical density of 0.7, measured at 600 nm (OD600). Protein production was induced by adding IPTG to 1 mM final concentration, and cells were grown for 20 hours with shaking (450 rpm) at 30 °C. After induction and expression, plates were centrifuged and pellets were resuspended in the lysis buffer (100 mM HEPES buffer pH 7.5, 1 mg / mL lysozyme, 0.5 mg / mL polymixin B sulfate, 3 U / mL DNase I, 4 mM magnesium sulfate) at 25 °C with shaking at 800 rpm for 2 hours. The cell lysate was clarified by centrifugation (4000 x g, 15 minutes). HTP Analytical method for Carboxylesterase activity:
[0306] The conversion in the enzymatic reaction was determined using a 2-minute UPLC-MS method in the primary screening and a longer 5-minute method in the retest screening. In the 2- minute method, an Acquity BEH C18 column (2.1 x 50 mm, 1.7 μm particle size) was used, with mobile phase (A) 0.1% (v / v) formic acid in water, (B) 0.1% (v / v) formic acid in acetonitrile, 1 µl injection volume, flow rate 0.8 mL min-1, 55 °C column temperature, detection by MS detector (SIM ion: m / z=406.80 for the substrate (compound 3k) and by-product (compound X); SIM ion: m / z=397.80 for the product (Compound 1)). Retention time of the substrate (Compound 3k): 1.257 minutes; Retention time of the desired product (Compound 1): 1.13 minutes; Retention time of the undesired hydrolysis product (compound X): 0.695 minutes. Gradient program time [min] Mobile phase A Mobile phase B25971
[0307] In the 5-minute method, an Acquity BEH C18 column (2.1 x 50 mm, 1.7 μm particle size) was used, with mobile phase (A) 0.1% (v / v) formic acid in water, (B) 0.1% (v / v) formic acid in acetonitrile, 1 µl injection volume, flow rate 0.75 mL min-1, 55 °C column temperature, detection by MS detector (SIM ion: m / z=406.80 for the substrate (compound 3k) and by-product (compound X); SIM ion: m / z=397.80 for the product (Compound 1)). Retention time of the substrate (compound 3k): 3.1 minutes; Retention time of the desired product (compound 1): 2:8 minutes; Retention time of the undesired hydrolysis product (compound X): 2.3 minutes. Gradient program time [min] Mobile phase A Mobile phase B 0 95 5ssays o a o yes e ase ac y
[0308] In the high-throughput screening, 100 µL reactions (300 mM sodium phosphate buffer, 100 mM HEPES, 75 mM magnesium chloride, 25 mM tetrapeptide Compound 3k, pH 7, 0.5% v / v clarified cellular lysates) were incubated at 20 ℃ with shaking at 1000 rpm for 2 hours. For assays with SFP samples, reactions were performed at 25 mM substrate concentration in 300 mM sodium phosphate buffer pH 7.5, 100 mM HEPES, 75 mM magnesium chloride, and with 0.1 g / L of enzyme at 20℃ with shaking at 1000 rpm for 2 hours. Production of Shake Flask Powders (SFP):
[0309] Based on the analysis of the HTP assay results using the methods described above, variants of SEQ ID NO: 151 / 152 were selected for larger scale analysis and further characterization. About 10 μL of glycerol stock or a freshly streaked single colony for each variant was inoculated in 7.5 mL of Luria-Bertani Broth medium with kanamycin (50 μg / ml) in 50 mL centrifugal tube at 30 °C with shaking (250 rpm) for 20 hours. The overnight culture was diluted 1:100 with 100 mL of Terrific Broth medium containing antibiotics (kanamycin 50 μg / ml) in 1L flask and grew until the cellular optical density reached 0.7, measured at 600 nm (OD600). Protein production was induced by adding IPTG to 1mM final concentration for 20 hours with shaking (250 rpm) at 30 °C. Cells were then collected by centrifugation and resuspended in 20 mL lysis solution (100 mM HEPES buffer pH 7.5) per 100 mL of growth25971 cultures. Cells were lysed by ultrasonication at 500 W for 15 min on ice (performed at 2 seconds of sonication with 4 seconds of intervals). The resulting lysate was clarified by centrifugation at 4 °C, and the supernatant was subsequently frozen and lyophilized to generate the enzyme powders. Analytical method for Carboxylesterase activity from SFP samples:
[0310] The conversion in the enzymatic reaction with SFP samples was determined using an Acquity BEH C18 column (2.1 x 50 mm, 1.7 μm particle size) was used, with mobile phase (A) 0.1% (v / v) TFA in water, (B) acetonitrile, 1 µl injection volume, flow rate 0.7 mL min-1, 55 °C column temperature, detection by MS detector (SIM ion: m / z=406.80 for the substrate (Compound 3k) and by-product (Compound X); SIM ion: m / z=397.80 for the product (Compound 1)). Retention time of the substrate (Compound 3k): 5.6 minutes; Retention time of the desired product (Compound 1): 4.7 minutes; Retention time of the undesired hydrolysis product (compound X): 3.8 minutes. Gradient program time [min] Mobile phase A Mobile phase B 0.0 90 10
[0311] Engineered polypeptides with >2-fold selectivity, defined as the ratio of the conversion % of substrate to the product (Compound 1) over the conversion % of substrate to the by-product (compound X), relative to the parent polypeptide, are listed in Table 12.1. The selectivity was measured with the shake flask powder samples. Table 12.1 Variants and Conversion SEQ ID NOs (nt / aa): Amino acid differences Increased selectivity (a) S ID NO 153154 A368L +of SEQ ID NO: 142 and defined as: “+” = selectivity at least 2-fold relative to reference polypeptide.
[0313] Carboxylesterase variants with mutations A368L or A368H produced less hydrolysis by-products (compound X) from the tetrapeptide (Compound 3k), and these engineered carboxylesterase enzymes provide more selective biocatalytic reagents for use in new methods25971 for the macrocyclization reaction of polypeptides. The variant with mutations A368L (SEQ ID NO: 154), had the highest selectivity. Thus, the encoding polynucleotide (SEQ ID NO: 153) was selected for further directed evolution. Example 30 (Rd12 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 154
[0314] In this example, engineering of SEQ ID NO: 154 for improved selectivity against hydrolysis of tetrapeptide (Compound 3k) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites were subjected to mutagenesis combinatorically resulting in library 12.1 and surface residues were subjected to mutagenesis individually resulting in library 12.2. These libraries were plated to form single colonies, which were picked, grown, and screened using the high-throughput growth and expression methods described in Example 12. The production of shake flask powders and HTP analytical methods were described in Example 12. Assays for Carboxylesterase activity:
[0315] In the high-throughput screening, 50 µL reactions (300 mM sodium phosphate buffer, 100 mM HEPES, 75 mM magnesium chloride, 25 mM tetrapeptide Compound 3k, pH 7, 0.5% v / v clarified cellular lysates) were incubated at 20 ℃ with shaking at 1000 rpm for 2 hours. For assays with SFP samples, reactions were performed at 25 mM tetrapeptide (Compound 3k) in 300 mM sodium phosphate buffer pH 7.5, 100 mM HEPES, 75 mM magnesium chloride, and with 0.1 g / L of enzyme at 20 ℃ with shaking at 1000 rpm for 2 hours.
[0316] Engineered polypeptides with >1.5-fold selectivity, defined as the ratio of the conversion % of substrate to the product (Compound 1) over the conversion % of substrate to the by-product (Compound X), relative to the parent polypeptide, are listed in Table 13.1. The selectivity was measured with the shake flask powder samples. Table 13.1 Variants and Selectivity SEQ ID NOs (nt / aa): Amino acid differences Increased selectivity (a)p yp p of SEQ ID NO: 154 and defined as: “+” = selectivity at least 1.5-fold relative to reference polypeptide.25971
[0318] Variants with mutations I115L, I157F, W170L, C172F, L187Q produced less hydrolysis by-products (compound X) from the tetrapeptide (compound 3k), and these engineered carboxylesterase enzymes provide more selective biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations I115L, I157F, W170L, C172F, L187Q (SEQ ID NO: 156), had the highest selectivity. Thus, the encoding polynucleotide (SEQ ID NO: 155) was selected for further directed evolution. Example 31 (Rd13 evolution of Carboxylesterase) Enzyme variants of SEQ ID NO: 156
[0319] In this example, engineering of SEQ ID NO: 156 for improved activity while maintaining selectivity against hydrolysis of tetrapeptide (compound 3k) and the resulting improved variants are described. Engineering via directed evolution was carried out by constructing libraries of variant genes in which positions associated with active sites and surface residues and improved performance in previous rounds were subjected to mutagenesis combinatorically, resulting in library 13.1. These libraries were plated to form single colonies, which were grown and screened using the high-throughput growth and expression methods described in Example 12. The HTP analytical methods were described in Example 12. Assays for Carboxylesterase activity:
[0320] In the high-throughput screening, 50 µL reactions (300 mM sodium phosphate buffer, 100 mM HEPES, 75 mM magnesium chloride, 25 mM tetrapeptide with (compound 1), pH 7, 0.2% v / v clarified cellular lysates) were incubated at 20 ℃ with shaking at 1000 rpm for 2 hours. Candidates with improved selectivity were picked for growth and expression again. Activity was re-examined in triplicate using the same conditions as the primary screening.
[0321] Engineered polypeptides with >2-fold activity relative to SEQ ID NO: 156 are listed in Table 14.1. The selectivity shown was measured with the retest samples. Table 14.1 Variants and Selectivity SEQ ID NOs (nt / aa): Amino acid differences Increased selectivity (a) E ID N 1 T1 K 1 2I T24 H fSEQ ID NO: 156 and defined as: “+” = activity at least 2-fold relative to reference polypeptide
[0322] Variants with mutations T10K, G162I, T246H, H272A, F343L produced more desired product (compound 1) from the compound tetrapeptide (compound 3k) and only low level of the hydrolysis by-product (compound X), and these engineered carboxylesterase enzymes provide25971 more selective biocatalytic reagents for use in new methods for the macrocyclization reaction of polypeptides. The variant with mutations T10K, G162I, T246H, H272A, F343L (SEQ ID NO: 158), had the highest activity and excellent selectivity (>99:1 product: by-product ratio). Thus, the encoding polynucleotide (SEQ ID NO: 157) was selected as the final variant.
[0323] Overall, polypeptide sequences having SEQ ID NOs: 152, 154, 156 and 158 emerged following 14 rounds of evolutionary pressures as the best-performing macrocyclase variants. These engineered polypeptides show improved (1) activity for the product of compound 1, (2) selectivity for compound 1 over undesired hydrolysis by-product compound X and undesired regioisomers of compound 1, (3) thermostability, and (4) co-solvent tolerance relative to carboxylesterase (SEQ ID NO: 132). These characteristics are important for successful large- scale manufacturing of non-canonical phenylalanine amino acids. Reaction Scheme A’:added 15 mL of NMP and 5.20 g (54.0 mmol) of solid sodium tert-butoxide. To this solution was added 3.00 g (13.5 mmol) of solid 5-fluorotryptophan (Compound 3) and the sides of the flask rinsed with 6 mL of NMP. To the mixture was added 0.073 mL of water and the mixture was warmed to 45 ºC for 1.5 h. In a separate flask was added 5.2 g (14.2 mmol) of 6-Bromohexan-1- aminium 4-methylbenzenesulfonate (Compound 2A) and 10 mL of NMP and the mixture was warmed to 45 ºC to give a homogeneous solution which was added to the above mixture. The25971 mixture is aged for 1 h, and then at this point and the pH of the mixture was adjusted to between 4-9 by the addition of 48% aqueous HBr solution. The mixture is then cooled to below 5 °C to allow for crystal growth of Compound 1d. The temperature is returned to 20 °C and then is diluted with 30 mL of ½ saturated aqueous NaBr which was prepared by dissolving 49.0 g of NaBr in 100 mL of water. The temperature of the reaction mixture was raised between 50-55 ºC, and the resulting slurry was aged at 50 °C overnight, and then cooled to 5 °C over a period of 5 h. The slurry was filtered and washed with 8:11-PrOH / water and dried under vacuum at 55 °C overnight to give Compound 1d: mp 246 ºC (DSC);1H NMR (DMSO-d6, 500 MHz) δ 7.87 (br s, 5H), 7.41 (dd, 1H, J = 8.9 and 4.4 Hz), 7.37 (dd, 1H, J = 10.1 and 2.5 Hz), 7.28 (s, 1H), 6.94 (td, 1H, J = 9.2 and 2.5 Hz), 4.12 (t, 2H, J = 6.4 Hz), 3.55 (t, 2H, J = 5.5 Hz), 3.30 (br s, 2H), 3.10 (qd, 2H, J = 15.0 and 5.5 Hz), 2.71 (m, 2H), 1.73 (m, 2H), 1.49 (m, 2H), 1.37-1.21 (m, 2H), 1.13 (m, 2H);13C NMR (DMSO-d6, 125 MHz) δ 171.5, 156.9 (d, J = 231.0 Hz), 132.7, 129.7, 128.1 (d, J = 9.8 Hz), 110.6 (d, J = 10.0 Hz), 109.0 (d, J = 26.4 Hz), 108.5 (d, J = 4.8 Hz), 103.7 (d, J = 22.0 Hz), 54.5, 45.4, 38.5, 29.4, 26.8, 26.6, 25.6, 25.3;19F NMR (DMSO-d6, 471 MHz) δ -125.6. HRMS Cacld. For C17H15FNO2: 322.1931 [M + H]. Found: 322.1925 [M + H]. See USSN 63 / 611847, filed 12 / 19 / 23; incorporated herein in its entirety. Intermediate Example 32bBatch procedure:
[0325] A 1-L pyrex bottle was charged with L-3-cyanophenylalanine (Compound 4g: See (Knittel et al., 1990 Pept. Res.3:176-181; US Patent Publication US2006 / 0142305; Chien et al., 2018 J. Med. Chem.61:7358-7373) (34.00 g, 1.0 equiv.), isopropanol (245 mL), water (163 mL) and 12 M aqueous HCl solution (29.6 mL), and stirred until all solids had dissolved. The obtained solution was transferred into a 1-L autoclave followed by adding 10% Pd / C catalyst (1.7 g). The autoclave was sealed and purged with nitrogen gas three times. Next, it was purged with hydrogen gas three times, the reaction pressure set to 400 psi and the temperature to 25oC, before aging the reaction overnight. When the reaction was complete, the autoclave was vented and the atmosphere inerted by purging three times with nitrogen gas. Agitation was stopped and the solution discharged into a 1-L Pyrex bottle. The reaction mixture was then filtered with a25971 funnel under vacuum to obtain a solution containing Compound 1f. See also USSN 63 / 639279, filed 4 / 26 / 24; incorporated herein in its entirety. Intermediate Example 33 (2S,3S)-3-(2-(tert-butoxy)-2-oxoethoxy)-2-(((2S,3R)-3-hydroxy-1-isopropoxy-1-oxobutan-2- yl)carbamoyl)pyrrolidine (2c)
[0326] To a 3-necked benzyl (2S,3S)-3-(2-(tert-butoxy)-2-oxoethoxy)-2-(((2S,3R)-3-hydroxy-1-isopropoxy-1-oxobutan-2- yl)carbamoyl)pyrrolidine-1-carboxylate (~40.0 g, 1:1:0.04 MTBE / iPrOH / water, 390 mL) and 5% Pd / C (2 g). A balloon of H2was attached, and the headspace of the flask was flushed thoroughly with H2. The reaction flask was placed under H2 atmosphere, and the reaction mixture was then heated to 55°C. The reaction mixture was aged for 4 h. After aging, the reaction mixture was poured onto a pad of diatomaceous earth (CELITE® 545, available from Millipore Sigma, 20 g) that had been wetted with iPrOH:MTBE 1:1. The filtrate was collected, and the filter pad was washed with 1:1 iPrOH / MTBE (40 mL). The wash was combined with the filtrate, and the combined liquids were concentrated to provide (2S,3S)-3-(2-(tert-butoxy)-2-oxoethoxy)-2- (((2S,3R)-3-hydroxy-1-isopropoxy-1-oxobutan-2-yl)carbamoyl)pyrrolidine. See USSN 63 / 668,869, filed 07 / 09 / 2024; incorporated herein in its entirety. Example 32 – General synthesis of Compound 3d using Trp-ligase in the presence of monomers of 1d and 1f25971
[0327] To a 2-dram vial, 4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid (71.5 mg), dipeptide 2c (oxalate salt (111.3 mg)), amino acid monomer 1f hydrochloride salt [Amino- Phe.HCl] (64.3 mg), magnesium chloride hexahydrate (61 mg), and adenosine 5'-monophosphate monohydrate [AMP] (5.9 mg) was dissolved in 2.8 mL water. Sodium hydroxide (5 N solution in water) was then added to adjust pH to 7.5. Next, sodium hexametaphosphate (122.4 mg), amino acid monomer 1d hydrobromide salt [Trp-Linker.HBr] (99.5 mg), tergitol 15-s-9 (90 mg) was added. Sodium hydroxide (5 N solution in water) was then added to adjust pH to 8.0. To initiate the reaction, 30 mL PPK22 (20 g / L) and 30 mL amino acid ligase Trp-ligase SEQ ID NO: 36 (50 g / L) was added. The reaction mixture was agitated at 20-25 °C for ca.24 hrs to give Tripeptide 3d as a slurry. MS [M+H]+C35H55FN5O8, 692.4029; Found: 692.4030.
[0328] In this reaction, Trp-Linker.HBr can be replaced by Trp-Linker.HCl. Dipeptide 2d (ProThr.oxalate) can be replaced by ProThr.Tartrate, ProThr.Etidronate and ProThr.Diphenyl phosphate. Other ProThr.Oxalate alternatives include the tBu ester moiety, carboxyamide, and free carboxylic acid, which may be used in place of the isopropyl ester on the threonine of ProThr.Oxalate. Example 33 - Synthesis of Compound 3k using Phe-ligasestock solution at 20 °C. PPK22 (20 mg) lyophilized powder was dissolved in 1 mL 0.1 M HEPES (pH 7.5) to make a 20 mg / mL stock solution. Phe-ligase SEQ ID NO.102 (50mg) lyophilized powder was dissolved in 1 mL 0.1 M HEPES (pH 7.5) to make a 50 mg / mL stock solution.1 M HEPES (pH 8.0) was purchased from Thermo Fisher.
[0330] To a 2-mL glass vial, 4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid (71.5 mg), monomer 1f hydrochloride salt (Amino-Phe.HCl, 64.3 mg), magnesium chloride hexahydrate (61 mg), and adenosine 5'-monophosphate monohydrate (5.9 mg) was dissolved in 2.8 mL water. Sodium hydroxide (5 N solution in water) was then added to adjust pH to 7.7. Next, 150 mL25971 tripeptide (150 mM) was added, followed by sodium hexametaphosphate (9.2 mg) and tergitol 15-s-9TM(90 mg). To initiate the reaction, 30 mL PPK22 (20 g / L) and 60 mL Phe-ligase SEQ ID NO.102 (50 g / L) was added. The reaction mixture was agitated at 20-25 °C for ca.6 hrs to give Tetrapeptide 3k as a slurry, MS [M+H]+C45H67FN7O9, 868.4979; Found: 868. Example 34 - Synthesis of Compound 11f bishydrochloride salt (19.86 kg) were dissolved in 930 L water at 25 °C. Sodium Hydroxide (10 N solution in water) was then added to adjust pH to 7.5. Dipeptide 2c oxalate salt (38.20 kg) was then added and dissolved. Sodium hydroxide (10 N solution in water) was again added to adjust pH back to 7.5. Sodium hexametaphosphate (42.22 kg), adenosine 5'-monophosphate monohydrate (2.52 kg), tergitol 15-S-9 (27.6 kg), compound 1d hydrobromide salt (29.84 kg) and magnesium chloride hexahydrate (18.70 kg) were sequentially added. Sodium hydroxide (10 N solution in water) was then added to adjust pH to 8.0. PPK22 (0.20 kg), Trp-ligase SEQ ID NO.36 (0.30 kg) and Phe-ligase SEQ ID NO.102 (1.8 kg) were then sequentially added. During the course of the reaction, the pH was adjusted at ca.10 h (target pH 7.5) and ca.20 h (target pH 7.9) with sodium hydroxide (10 N solution in water). The resulting reaction mixture was agitated at 25 °C for 48 h to give tetrapeptide 3k as a slurry. Analysis confirmed residual tripeptide 3d ≤ 4.0 area% and residual compound 1d ≤ 0.1 area% by UPLC, the pH was adjusted to 8.0 with sodium hydroxide (10 N solution in water). Carboxyesterase SEQ ID NO.156 (0.20 kg) was25971 added, and the resulting reaction mixture was agitated at 25 °C for 12 h, yielding Compound 1 as a slurry: MS [M+H]+C41H57FN7O8794.4; Found: 794.4. Example 35 – Compound 1 crystallization procedure
[0332] After ageing the mixture in Example 34 at 25 °C for 19 h, phosphoric acid (85%, 40.3 kg) was charged to the vessel to adjust pH to 3.74. The reaction was then aged for 69 h at 20 °C before the pH was adjusted to 7.44 with NaOH (10N, 88.6 kg). Acid washed Celite 545 (29.00 kg) and 1-butanol (411.1 kg) were then charged to the vessel and stirred for 2 h. The resulting mixture was then filtered through a pressure filter and the filtrate was collected. The vessel was rinsed with additional 1-butanol (102.7 kg) which was then filtered through the waste cake and collected. The combined filtrates were then charged back to the vessel, stirred for 30 min and allowed to settle for 2 h. The aqueous layer was then cut to waste and potassium bicarbonate (15 % w / v in water, 723.2 kg) was added to the vessel. After stirring for 30 mins, the mixture was allowed to settle for 1 h and the aqueous layer was cut to waste. Water (603.4 kg) was added to the vessel and the resulting mixture was stirred for 30 min and allowed to settle for 1 h. The aqueous layer was cut to waste and oxalic acid (2.40 kg) was added to the vessel and aged for 2 h to dissolve. The organic layer was drummed and sampled for assay yield (43.2 kg macrocyclic peptide Compound 1 as a freebase, 78.9% assay yield).
[0333] The resulting solution was then transferred through an inline 10 micron filter to a second vessel for crystallization. At 25 °C, a solution of oxalic acid was charged (126.9 kg of 3.69 wt% oxalic acid in 99.5:0.5 MTBE:water, v:v) to bring total oxalic acid equivalents in batch to 1.45 with respect to the compound 1 freebase. The resulting solution was seeded with macrocyclic peptide 1 as bis-oxalate salt (0.22 kg, 0.5 wt%) and aged for 1 h then warmed to 35 °C and aged 30 min further. Oxalic acid solution (13.4 kg of 3.69 wt% oxalic acid in 99.5:0.5 MTBE:water, v:v) was charged over 2 h at 35 °C, then cooled slowly to 20 °C over 4 h. After ageing at 20 °C for 1 h, the slurry was heated to 40 °C, aged 30 mins, then cooled to 35 °C and aged for 125971 h. Oxalic acid solution (20.1 kg of 3.69 wt% oxalic acid in 99.5:0.5 MTBE:water, v:v) was charged over 34 h at 35 °C, then cooled to 20 °C over 3 h. After ageing at 20 °C for 30 min, the slurry was heated to 40 °C and aged 30 mins. The slurry was then returned to 35 °C and aged for 1 h. Oxalic acid solution (73.6 kg of 3.69 wt% oxalic acid in 99.5:0.5 MTBE:water, v:v) was charged over 8 h at 35 °C, then cooled to 10 °C over 6 h. After ageing for 3 days at 10 °C, the slurry was then warmed to 15 °C and transferred to an agitated filter dryer and filtered. A displacement wash of 1.75:1 v:v 1-butanol:MTBE (66.0 kg) was charged and the cake was deliquored. A second displacement wash of MTBE (93.3 kg) was charged and the cake was again deliquored. A slurry wash of wet MTBE (186.6 kg of 99:1 MTBE:water v:v) was charged and the batch was agitated for 2 hours followed by deliquoring.
[0334] The cake was then dried under a flow of nitrogen. A slurry wash of wet MTBE (186.6 kg of 99:1 MTBE:water v:v) was charged and the batch was agitated for 2 h followed by deliquoring. A final displacement wash of the cake was then executed with wet MTBE (93.3 kg of 99:1 MTBE:water v:v). After deliquoring, the cake was dried under a flow of nitrogen with intermittent agitation at 25 °C. The desired macrocyclic peptide Compound 1 bisoxalate salt was isolated as a solid (52.2 kg).
[0335] 1H NMR (600 MHz, 6:4 acetonitrile-d3:D2O) δ 7.87 (1H, d, J = 7.9 Hz), 7.55 (1H, dd, J = 8.3, 4.3 Hz), 7.39-7.34 (3H, m), 7.25-7.23 (2H, m), 7.16 (1H, s), 6.98 (1H, td, J = 9.3, 2.5 Hz), 6.88 (1H, s), 5.00 (1H, hept, J = 6.2 Hz), 4.91 (1H, dd, J = 7.6, 5.6 Hz), 4.72 (1H, dd, J = 14.7, 8.3 Hz), 4.58 (1H, s), 4.28-4.27 (2H, m), 4.23 (1H, dd, J = 9.5, 3.7 Hz), 4.18 (1H, d, J = 15.8 Hz), 4.17 (1H, dd, J = 14.7, 4.0 Hz), 4.17 (1H, s), 4.10 (1H, d, J = 16.1 Hz), 4.07 (2H, t, J = 7.4 Hz), 3.97 (1H, dd, J = 14.7, 4.0 Hz), 3.71 (1H, q, J = 9.4 Hz), 3.27 (1H, dd, J = 13.8, 3.6 Hz), 3.10-3.03 (4H, m), 2.88-2.85 (2H, m), 1.96 (2H, m), 1.74 (2H, pent, J = 7.4 Hz), 1.55 (2H, pent, J = 7.6 Hz), 1.33-1.31 (2H, m), 1.26-1.24 (2H, m), 1.25 (3H, d, J = 6.2 Hz), 1.23 (3H, d, J = 6.2 Hz), 1.16 (3H, d, J = 6.2 Hz).13C NMR (150 MHz, 6:4 acetonitrile-d3:D2O) δ 171.80, 171.79, 171.60, 170.18, 168.03, 166.41 (4 oxalate carbons), 158.40 (d, JC-F = 232.2 Hz), 139.68, 134.72, 133.73, 130.70, 130.42, 129.66, 129.38, 128.97 (d, JC-F= 10.0 Hz), 128.79, 111.78 (d, JC-F= 10.0 Hz), 110.40 (d, JC-F = 26.5 Hz), 108.72 (d, JC-F = 4.4 Hz), 104.21 (d, JC-F = 23.2 Hz), 81.62, 70.94, 68.16,65.19, 59.51, 54.42, 52.01, 46.79, 45.79, 42.74, 40.23, 37.36, 31.33, 30.59, 29.13, 27.60, 26.79, 26.43, 21.93, 21.86, 20.28. MS [M+H]+C41H57FN7O8794.4; Found: 794.4. See Attorney Docket 25970, filed contemporaneously with the instant application and incorporated herein in its entirety.25971
[0336] Compound 1 can be isolated from aqueous reaction streams by extraction with a non- miscible alcohol, and can be crystallized as a tosylate, hydrochloride, acetic acid, mesylate, glycolate, oxalate, 4-hydroxybenzoate, or di-tolyl tartrate salt.
[0337] Example 36 – Preparation of Compound 1 Ot-Bu 1 wt%-Trp-Ligase, 6 wt%-Phe-Ligase O O OH 0.6 wt%-PPK22, H NOOH O AMPO Ni-PrO NOH H ONPr
[0338]
[0339] acid (HEPES, 21.90. kg) and compound 1f bishydrochloride salt (19.86 kg) was dissolved in 920 L water at 25 °C. Sodium hydroxide (10 N solution in water) was then added to adjust pH to 7.6. Dipeptide 2c oxalate salt (38.16 kg) was then added and dissolved. Sodium hydroxide (10 N solution in water) was again added to adjust pH back to 7.4. Sodium hexametaphosphate (42.24 kg), Adenosine 5'-monophosphate monohydrate (2.52 kg), Tergitol 15-S-9 (27.6 kg), Compound 1d hydrobromide salt (29.9 kg) and magnesium chloride hexahydrate (18.72 kg) were sequentially added. Sodium hydroxide (10 N solution in water) was then added to adjust pH to 8.0. Polyphosphate kinase PPK22 (0.2 kg), ligase Trp-Ligase SEQ ID NO.36 (0.3 kg) and Phe-ligase SEQ ID NO.102 (1.84 kg) were then sequentially added. During the course of the reaction, the pH was adjusted at10 h (target pH 7.5) and 20 h (target pH 8.0) with sodium hydroxide (10 N solution in water). The resulting reaction mixture was agitated at 25 °C for 48 h to give tetrapeptide 3k as a slurry. The pH was adjusted to 7.8 with sodium hydroxide (10 N solution in water). Carboxylesterase SEQ ID NO.155 (0.2 kg ( 0.6 wt%)) was added, and the resulting reaction mixture was agitated at 25 °C for 8 h, yielding the macrocylic peptide compound 1 as a slurry. MS [M+H]+C41H57FN7O8794.4; Found: 794.4.25971 Example 37Phosphoric acid (85%, 40.3 kg) was charged to the vessel to adjust pH to 3.6. The reaction was then aged for 4 h at 20 °C before the pH was adjusted to 7.44 with NaOH (10N, 88.6 kg). Acid washed Celite 545 (29.0 kg) and 1-butanol (11 kg) were then charged to the vessel and stirred for 2 h. The resulting mixture was then filtered through a pressure filter and the filtrate was collected. The vessel was rinsed with additional 1-butanol (102.7 kg) which was then filtered through the waste cake and collected. The combined filtrates were then charged back to the vessel, stirred for 0.5 h and allowed to settle for 2 h. The aqueous layer was then cut to waste and potassium bicarbonate (15 % w / v in water, 723.3 kg) was added to the vessel. After stirring for 30 mins, the mixture was allowed to settle for 1 h and the aqueous layer was cut to waste. Water (603.4 kg) was added to the vessel and the resulting mixture was stirred for 30 mins and allowed to settle for 2 h. The aqueous layer was cut to waste and oxalic acid (2.42 kg) was added to the vessel and aged for 2 h to dissolve. The organic layer was drummed off.
[0341] The resulting solution was then transferred through an inline 10 micron filter to a second vessel for crystallization. At 25 °C, a solution of oxalic acid was charged (124.4 kg of 3.69 wt% oxalic acid in 99.5:0.5 MTBE:water, v:v). The resulting solution was seeded with macrocyclic peptide 1 as bis-oxalate salt (0.22 kg, 0.5 wt%) and aged for 30 min then warmed to 35 °C and aged 30 min further. Oxalic acid solution (19.8 kg of 3.69 wt% oxalic acid in 99.5:0.5 MTBE:water, v:v) was charged over 3 h at 35 °C, then cooled slowly to 20 °C over 4 h. After ageing at 20 °C for 1 h, the slurry was heated to 40 °C, aged 30 min, then cooled to 35 °C and aged for 1 h. Oxalic acid solution (20.1 kg of 3.69 wt% oxalic acid in 99.5:0.5 MTBE:water, v:v) was charged over 3 h at 35 °C, then cooled to 20 °C over 3 h. After ageing at 20 °C for 30 min, the slurry was heated to 40 °C and aged 30 min. The slurry was then returned to 35 °C and aged for 1 h. Oxalic acid solution (72.6 kg of 3.69 wt% oxalic acid in 99.5:0.5 MTBE:water, v:v) was charged over 8 h at 35 °C, aged at 35 °C for 1 h, then cooled to 10 °C over 6 h. After25971 ageing for 9 h at 10 °C, the slurry was then warmed to 15 °C and transferred to an agitated filter dryer and filtered. A displacement wash of 1-butanol (100.8 kg) was charged and the cake was deliquored. A second displacement wash of MTBE (92.1 kg) was charged and the cake was again deliquored. A slurry wash of wet MTBE (184.2 kg of 99:1 MTBE:water v:v) was charged and the batch was agitated for 2 h followed by deliquoring. A slurry wash of wet MTBE (184.2 kg of 99:1 MTBE:water v:v) was charged and the batch was agitated for 2 h followed by deliquoring. A displacement wash of wet MTBE (92.1 kg of 99:1 MTBE:water v:v) was charged and the batch was deliquored.
[0342] The cake was then dried under a flow of nitrogen. A slurry wash of wet MTBE (184.2 kg of 99:1 MTBE:water v:v) was charged and the batch was agitated for 2 h followed by deliquoring. A final displacement wash of the cake was then executed with wet MTBE (92.1 kg of 99:1 MTBE:water v:v). After deliquoring, the cake was dried under a flow of nitrogen with intermittent agitation at 25 °C to provide the desired macrocyclic peptide product 1 bisoxalate salt (48.8 kg).
[0343] The disclosed process is not to be limited in scope by the specific embodiments and examples described herein. Indeed, various modifications of the disclosure in addition to those described will become apparent to those skilled in the art from the foregoing description and accompanying figures. Such modifications are intended to fall within the scope of the appended claims.
[0344] All references (e.g., publications or patents or patent applications) cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual reference (e.g., publication or patent or patent application) was specifically and individually indicated to be incorporated by reference in its entirety for all purposes. Other embodiments are within the following claims. Table 7 Sequences SEQ ID Sequence NO: T T A T A T T C C C A A25971 SEQ ID Sequence NO: AAGCAAGGCCAACTTCCTATTGGAACTTGGTCTTGTCGATGCGGATGCGAATATTACGCGATCAT A A A D I I D I G A G C T T TT T C A A G T A G T A C S IP S HI G G A G C T T TT T C A A G T A G T A C S IP S25971 SEQ ID Sequence NO: VYHVDVVYINSGSILISPSRYLVPPLDFEKQNTGSVMLDENGADYSELLRLTKQLIASFNDQTIPNVMHI G G A G C T T TT T C A A G T A G T A T S IP S HI G G A G C T T TT T C A A G T A G T A C S IP S HI G G A G C T25971 SEQ ID Sequence NO: CCGGGTCTGCAGCATGAACTGGCTTTAAGCTGCCGCGATAAGGTGACCATGAAACAGAGCGCTTT TT T C A A G T A G T A C L IP S HI G C A G C T T TT T C A G T A G T A C Y II G A G A G C T T TT T C A A G T25971 SEQ ID Sequence NO: GAAACCGGTGACTTTGTGTTTGGTGAGATGGCAGCACGTCGTGGCGGCGGTTTAATCAAACAAGA G T A C L IP S HI G G A G C T T TT T C A T G T A G T A C L IP S H G A G C T T TT T C A T G T A G T T T L IP25971 SEQ ID Sequence NO: YTTCQGFGDIISAFDRWETVVLKPRWGVGSAGITILHSKDDLPALATKPEFIRNVHSNQYYLEEYCSGS H G A G C T T TT T C A T G T A G T T T L IP S H G A G C T T TT T C A T G T A G T A T L IP S H G A G C25971 SEQ ID Sequence NO: GCATCGTGAGCAGCCTCGAAGAAGATGTGCTGCGTGTGGCCGAAGCCCGTAGTCTGTTTGGTATT T TT T C A T G A G A A Y II G M A G A G C T T TT T C A T G A G A A Y II G M G A G C T T TT T C A T G25971 SEQ ID Sequence NO: CAACTGATCTGTAGCCTGAATGATCAGACCATTCCTAATGTGATGCATATTGAATTTTACAAAAAT A T A A Y II G M A G A G C T T TT T C A T G T A T A A L I G M E G A G C T T TT T C A C G T A T A A25971 SEQ ID Sequence NO: 34 MHHHHHHGSMKVLLMQHPHSLSNYPKWIEEIQERFDPLEVMVFTSSDRHAHHSWPSSVIKEIEVSDYL II G M E G A G C T T TT T C A C G T A T A A T I G M E G A G C T T T T C A C G A G A A A G I G M E G A25971 SEQ ID Sequence NO: CCAGCAGCGATCGCCACGCCCATCACAGCTGGCCGAGCAGCGTGATTAAAGAGATCGAGGTGAG C T T T T C A C G A G A A T KI G M E G A G C T T T T C A C G A G C T T KI G M G G A G C T T T T C A25971 SEQ ID Sequence NO: CAGCGGCAGCATTTTAATTAGCCCGAGCCGCTATATGCGTCCTCCGCTGGACCGCGAGATACAGC G A G A A T KI G M E G A G C T T T T C A C G A G C T T KI G M G G A G C T T T T C A C G A G A A25971 SEQ ID Sequence NO: ATCAATTCGAATGGCGCGTGGTCGGACCGCATGTTTCTGATTAGCGGCAAAAACGAAGAGGAAAT KI G M E G A G C T T T T C A C G A G A A T KI G M E G A G C T T T T C A A G T A G T A C Y II G25971 SEQ ID Sequence NO: 53 ATGCATCATCATCACCACCATGGCAGCATGAAAGTGCTGTTACTGCAGCGGCCGAAGAGCGACAG A G C T T T T C A A G T A G T A C S P S HI G G A G C T T T T C A A G T A G T A C S P S HI G G A G C T T T T25971 SEQ ID Sequence NO: TACACAGTAAAGACGATCTGCCGGCTTTAGCCACCAAACCGGAGTTCATCCGCAATGTGCATAGC A A G T A G TT T C S P S HI G G A G C T T T T C A A G T A A G C A S P S H A G A G C T T T T C A A G T G A G25971 SEQ ID Sequence NO: GCCGAAAGAGAAAGAGATCCCGGATTGGGCAGTTCTGCATAGCGTTGGCCAAAAGAAGGGTATC A S P S H A G A T G C A C T A C A C G G A G C A E P S H A G A G C T T T T C A A G T G A G C A S P S H25971 SEQ ID Sequence NO: IEFYKNETGDFVFGEMAARRGGSLIKQGLAAAYGIDQSKANFLLELGLLDADANITRSSQWGILLETA G A G C T T T T C A A G T G A T T A Y P S H A G A G C T T T T C A A G T G A T T A Y P S H A G A G C T T25971 SEQ ID Sequence NO: AGATGCCGGTCTGAAAATTATTCCGTATACCACATGCCAAGGTTTCGGCGATATCATCAGCGCCTT T C A A G T G A T T A Y IP S H A G A G C T T T T C A A G T G A T T A Y IP S H G A G C T T T T C A A G T C25971 SEQ ID Sequence NO: ACTGGCCGCCGCCTATGGCATTGATCAAAGCAAAGCAAATTTTTTACTGGAACTGGGTTTACTTGA T T A Y IP S H G A G C T T T T G A T A A G G T A C Y IP S H G A G C T T T T G A T A A G G T A C Y IP S25971 SEQ ID Sequence NO: VYHVDVVYINSGSILISPSRYLVPPLNFEKDNTGSVMLDENGADYQELLRLTKQLIASFNDQTIPNVMH G A T G C A C T A C A C G G T G C T Y II G M A G A T G C A C T A C A C G G T G C T Y IP S H G G A T G C25971 SEQ ID Sequence NO: CGGGTCTGCAGCATGAACTGGCTTTAAGCTGCATCGATAAGGTGACCATGAAACAGAGCGCTTTA C T A C C A A C G G G C Y IP S HI G G A T G C A T TT A C C A A C G G G C Y I G A G A T G C A C A C C A A25971 SEQ ID Sequence NO: AAGCGGTGACTTTGTGTTCGGTGAGATGGCAGCACGTCGTGGCGGCTCTCTGATCAAACAATCGC G G G C Y I G A G A T G C A C A C C A A C G G G T Y I GS HI L G A T G C A C T A C C A A C G G G T Y I25971 SEQ ID Sequence NO: PYTTCQGFGDIISAFDRWETVVLKPRALHGSMGITILHSKDDLEALATSPEFIRNVHSNQYYLEEYLLGS HI L G A T G C A C T A C C A A C G G G T Y I GS HI L G A T G C A T T C A A C G G T G T T Y I G A G A T G25971 SEQ ID Sequence NO: CATCGTGAGCGGTAGCGAAGATGATGTGCTGCGTGTGGCCGAAGCCCGTAGTCTGTTTGGTATTC A T T C A A C G G T G T T Y I A G A T G C A T T C A A C G G T G T T KI G H G G A T G C A C T A C C A25971 SEQ ID Sequence NO: ACTGATCTATAGCTTTAATGATCAGACCATTCCTAATGTGATGCATATTGAATTTTACAAAAATGA C G G G T KI G H G G A T G C A T C A A C G G T G T T Y K G H G G A T G C A T C A A C G G T G T T25971 SEQ ID Sequence NO: 108 MHHHHHHGSMKVLLLYRPGSDSNYGKWIEEIQERFDHLEVMVFTSNDRAEPPSWPSSVIKVIEVDDY K G H G G A T G C A T C A A C G G T G T T Y K G H G G A T G C A T C A A C G G T G T T K G H G G A25971 SEQ ID Sequence NO: CCAGCAATGATCGCGCCGAACCTCCTTCGTGGCCGAGCTCGGTGATTAAAGTGATCGAGGTGGAT G C A T C A A C G G T G T T Y K G H G G A T G C A T C A A C G G T G T T Y K G A D DI I D P D DI25971 SEQ ID Sequence NO: ISAFDRWETVVLKPRWGVGSASITILHSKDDLPELATKPEFIRNVHSNQYYLEEYCNGSVYHVDVVYI D P D DI N F P F G YI DF D F G YI D P D D IN F D G A T G C A T C A A C G G T G T T T C G C A G A A25971 SEQ ID Sequence NO: CCGGCAGTGTCATGCTGGACGAAAACGGTACGGATTACCAAGAATTGCTGACGCTGACCAAACA A C G G C C G A C G C A T A C C A A T G G C A C G C C G T A T A G A TT C C G G A C C G A A C C G C C25971 SEQ ID Sequence NO: TGATCTATAGCTTTAACGACCAGACCATCCCAAATGTGATGCATATCGAGTTCTTCAAAAACGAA G C C G T C G C C A G C A A A G T G GSEQ ID Sequence T G C C G A T C T A T G C G C G A G Y D L T IR F25971 Table 15: Sequences SEQ ID Sequence NO: T G C C G A T A G G G A A A A G C A Y D W T PI D T G C C G A T C T A T G G T T G T Y D R F T G C25971 Table 15: Sequences SEQ ID Sequence NO: C G A T C T A T G G T T G T Y D W T PI D T G C C G A T T A T G G T T G T Y D R FP T G C C G A25971 Table 15: Sequences SEQ ID Sequence NO: C T A T G G T T G T Y E T PI D T G C C G A C T A T G G T T G T Y E T PI D T G C C G A C A25971 Table 15: Sequences SEQ ID Sequence NO: T A T G C G C G A G Y E W T PI D G A G C A G C A G G A A A A G C A Y E W T I D G A G C A G C A G G25971 Table 15: Sequences SEQ ID Sequence NO: A A A A G C A Y E W T PI D G A G C A G C A G G A A A A G C A Y E W T I D G A G C A G C A G G A A A25971 Table 15: Sequences SEQ ID Sequence NO: A G C A Y E W T I D G A G C A G C A G G T T A A G C A G Y E W T I D G A G C A G C A T T T A C T T G G G25971 Table 15: Sequences SEQ ID Sequence NO: C R R Y PF T G C C G A T A G G G A A A A G C A Y D L T IR F T G C C G A T A C T A T G G T T G T25971 Table 15: Sequences SEQ ID Sequence NO: Y D R R F G A G C A G C A G G C G C G A G Y E W T PI D G A G C A G C A G G A A A A G C A Y E W25971 Table 15: Sequences SEQ ID Sequence NO: T I D D R L H H K D R L S R L
Claims
1. 25971 WHAT IS CLAIMED IS:
1. A process for preparing Compound 1’ O O N HOO or a salt hydrate, and / or solvatecomprising the steps of 1) combining a compound of formula 3k’:or salt, hydrate, and / or solvate thereof, with an engineered carboxylesterase enzyme to provide a compound of Formula 1’, and 2) isolating the compound of Formula 1’ or salt, hydrate, and / or solvate thereof; wherein R1and R5are independently selected from hydrogen, C1-10 alkyl, aryl, and heteroaryl; R2is selected from hydrogen, C1-6 alkyl, C2-6 alkenyl, said alkyl and alkenyl optionally substituted with 1 to 3 groups of R, R is selected from NH2, OH, OC1-6alkyl, OC2-6alkenyl, and NHC1-6alkyl, N(C1-6 alkyl)2; R3is selected from hydrogen and halogen, R4is selected from hydrogen, OH, -OC1-6alkyl, -O(CH2)nCOOR5, -O(CH2)nC(O)SR5, and -O(CH2)nC(O)NHR5and n is 1 to 6.
2. The process of Claim 1, wherein the carboxylesterase enzyme is selected from one of SEQ ID NOs: 152, 154, 156 and 158.25971 3. The process of Claim 1, further comprising the step of stirring the resulting solution of step 1 at least 12 hours.
4. The process of Claim 1, wherein the pH is maintained at about 6.5 to about 8.
0.
5. The process of Claim 1, further comprising the steps of preparing Compound 3k’ by a) combining the compound of Formula 3d’:or salt thereof, with Compound 1f’: or salt, hydrate, and / or solvate thereof,of a Phe-ligase enzyme, adding a polyphosphate kinase, adenosine phosphate, magnesium ion (Mg2+), and a phosphate donor, to generate a compound of Formula 3k’, salt, hydrate, and / or solvate and b) isolating 3k’, wherein R6is selected from hydrogen and (CH2)nNH2.
6. The process of Claim 5 wherein the Phe-ligase enzyme is selected from SEQ ID NOs: 108, 110, 112, 114, and 116.
7. The process of Claim 5 wherein the polyphosphate kinase is selected from wild-type PPK22 and PPK12; adenosine phosphate is selected from ATP, ADP, AMP, or combination thereof; magnesium ion is selected from magnesium chloride hexahydrate, and magnesium sulfate; and phosphate donor is selected from sodium hexametaphosphate, and sodium polyphosphate.25971 8. The process of Claim 5 wherein the temperature is maintained from about 15^C to about 35^C.
9. The process of Claim 5 wherein the reaction is aged from about 12 hours to about 72 hours, with or without stirring.
10. The process of claim 5 wherein a nonionic surfactant selected from Tergitol™, Triton X- 100®, and / or reduced Triton is added.
11. The process of Claim 1, further comprising the steps of preparing Compound 3d’ by a) combining the compound of Formula 2c’or salt, hydrate, and / or solvate thereof, with a compound of Formula 1d’ R3Oor salt, hydrate, and / or solvate thereof, in the presence of a Trp-ligase enzyme, polyphosphate kinase, adenosine phosphate, an inorganic salt of magnesium, and a phosphate donor, to generate a compound of Formula 3d’, or salt thereof, and b) isolating 3d’.
12. The process of claim 11 wherein a nonionic surfactant selected from Tergitol™, Triton X-100®, and / or reduced Triton is added.
13. The process of Claim 11 wherein the Trp-ligase enzyme is selected from SEQ ID NOs: 40, 42, 44, 46, 48, and 50.25971 14. The process of Claim 11 wherein the polyphosphate kinase is selected from wild-type PPK22 and PPK12; adenosine phosphate is selected from ATP, ADP, AMP, or combination thereof; magnesium ion is selected from magnesium chloride hexahydrate, and magnesium sulfate; and phosphate donor is selected from sodium hexametaphosphate, and sodium polyphosphate.
15. The process of Claim 11 wherein the temperature is maintain from about 15^C to about 35^C.
16. The process of Claim 11 wherein the reaction is aged for at least 12 hours with agitation.
17. A process for preparing Compound 1’, or a salt, hydrate, and / or solvate thereof comprising the steps of a) combining 1f’, 2c’ and 1d’ orTrp-ligase enzymes, polyphosphate kinase, adenosine phosphate, an inorganic salt of magnesium, and a phosphate donor, to produce a compound of Formula 3k’, or salt thereof,25971b) adding an engineered carboxylesterase enzyme to provide a compound of Formula 1’, and c) isolating the compound of Formula 1’, or salt, hydrate and / or solvate thereof, wherein R1and R5are independently selected from hydrogen, C1-10 alkyl, aryl, and heteroaryl; R2is selected from hydrogen, C1-6alkyl, C2-6alkenyl, said alkyl and alkenyl optionally substituted with 1 to 3 groups of R, R is selected from NH2, OH, OC1-6 alkyl, OC2-6 alkenyl, and NHC1-6 alkyl, N(C1-6alkyl)2; R3is selected from hydrogen and halogen, R4is selected from hydrogen, OH, -OC1-6 alkyl, -O(CH2)nCOOR5, -O(CH2)nC(O)SR5, and -O(CH2)nC(O)NHR5, R6is selected from hydrogen and (CH2)nNH2,and n is 1 to 6.
18. The process of claim 17 wherein a nonionic surfactant selected from Tergitol™, Triton X-100®, and / or reduced Triton is added.
19. The process of Claim 17 wherein the Trp-ligase enzymes are selected from SEQ ID NOs: 40, 42, 44, 46, 48, and 50 and the Phe-ligase enzymes are 108, 110, 112, 114, and 116.
20. The process of Claim 17 wherein the polyphosphate kinase is selected from wild-type PPK22 and PPK12; adenosine phosphate is selected from ATP, ADP, AMP, or combination thereof; magnesium ion is selected from magnesium chloride hexahydrate, and magnesium sulfate; and phosphate donor is selected from sodium hexametaphosphate, and sodium polyphosphate.
21. The process of Claim 17 wherein the carboxylesterase enzyme is selected from one of SEQ ID NOs: 152, 154, 156 and 158.25971 22. The process of Claim 17 wherein the temperature is maintained from about 15^C to about 35^C.
23. The process of Claim 17 wherein the reaction is aged at least 12 hours with agitation.
24. The process of Claim 17 wherein diatomaceous earth is optionally added prior to addition of the carboxylesterase enzyme.
25. A process for preparing Compound 1: or a salt hydrate, and / ora) combining 1f, 2c and 1d orkinase selected from wild-type PPK22 and PPK12; adenosine phosphate selected from ATP, ADP, AMP, or combination thereof; magnesium ion selected from magnesium chloride hexahydrate, and magnesium sulfate; and phosphate donor is selected from sodium hexametaphosphate, and sodium polyphosphate, to produce compound 3k, or salt hydrate, and / or solvate thereof,25971 O NH2OtBu b) adding an Compound 1, andc) isolating a 26. The process of claim 25 wherein a nonionic surfactant selected from Tergitol™, Triton X-100®, and / or reduced Triton is added.
27. The process of Claim 25 wherein the Trp-ligase enzymes is selected from SEQ ID NOs: 40, 42, 44, 46, 48, and 50 and Phe-ligase enzymes are selected from 108, 110, 112, 114, and 116.
28. The process of Claim 25 wherein the carboxylesterase enzyme is selected from one of SEQ ID NOs: 152, 154, 156 and 158.
29. The process of Claim 25 wherein the temperature is maintained from about 15^C to about 35^C.
30. The process of Claim 25 wherein the reaction is aged at least 12 hours with agitation.
31. The process of Claim 25 wherein diatomaceous earth is optionally added prior to addition of the carboxylesterase enzyme.
32. A compound represented by structural Formula 1a:25971or a pharmaceutically acceptable salt, hydrate, and / or solvate thereof, wherein R1is selected from hydrogen, C1-10 alkyl, aryl, and heteroaryl; R2is selected from hydrogen, C1-6 alkyl, C2-6alkenyl, said alkyl and alkenyl optionally substituted with 1 to 3 groups of R, R is selected from NH2, OH, OC1-6 alkyl, OC2-6 alkenyl, and NHC1-6 alkyl, N(C1-6 alkyl)2; R3is selected from hydrogen and halogen.
33. The compound according to claim 32, or a pharmaceutically acceptable salt thereof, wherein R1is selected from hydrogen, ethyl, propyl, isopropyl, butyl, pentyl and hexyl; R2is selected from hydrogen, -(CH2)nOH, -(CH2)nNH2, -(CH2)nNHCH3, -(CH2)nNHCH2CH3, and n is 1 to 6; and R3is hydrogen or halogen.
34. A compound represented by structural Formula 3ka: NH2or a pharmaceutically acceptable salt, hydrate, and / or solvate thereof, wherein wherein R1and R5are independently selected from hydrogen, C1-10 alkyl, aryl, and heteroaryl; R2is selected from hydrogen, C1-6 alkyl, C2-6 alkenyl, said alkyl and alkenyl optionally substituted with 1 to 3 groups of R, R is selected from NH2, OH, OC1-6alkyl, OC2-6alkenyl, and NHC1-6alkyl, N(C1-6alkyl)2;25971 R3is selected from hydrogen and halogen, R4is selected from hydrogen, OH, -OC1-6alkyl, - O(CH2)nCOOR5, -O(CH2)nC(O)SR5, and -O(CH2)nC(O)NHR5, and n is 1 to 6.
35. A compound represented by structural Formula 3da:or a pharmaceutically acceptable thereof, wherein R1and R5are independently selected from hydrogen, C1-10alkyl, aryl, and heteroaryl; R2is selected from hydrogen, C1-6 alkyl, C2-6 alkenyl, said alkyl and alkenyl optionally substituted with 1 to 3 groups of R, R is selected from NH2, OH, OC1-6alkyl, OC2-6alkenyl, and NHC1-6alkyl, N(C1-6alkyl)2; R3is selected from hydrogen and halogen, R4is selected from hydrogen, OH, -OC1-6 alkyl, - O(CH2)nCOOR5, -O(CH2)nC(O)SR5, and -O(CH2)nC(O)NHR5, and n is 1 to 6.
36. A compound which is isopropyl ((11S,12S,13S,9S,12S)-9-amino-12-((1-(6-aminohexyl)- 5-fluoro-1H-indol-3-yl)methyl)-4,10,13-trioxo-2-oxa-5,11-diaza-1(3,1)-pyrrolidina-7(1,3)- benzenacyclotridecaphane-12-carbonyl)-L-threoninate, isopropyl ((2S,3S)-1-((S)-2-((S)-2-amino-3-(3-(aminomethyl)phenyl)propanamido)-3-(1-(6- aminohexyl)-5-fluoro-1H-indol-3-yl)propanoyl)-3-(2-(tert-butoxy)-2-oxoethoxy)pyrrolidine-2- carbonyl)-L-threoninate, isopropyl ((2S,3S)-1-((S)-2-amino-3-(1-(6-aminohexyl)-5-fluoro-1H- indol-3-yl)propanoyl)-3-(2-(tert-butoxy)-2-oxoethoxy)pyrrolidine-2-carbonyl)-L-threoninate, or a pharmaceutically acceptable salt, hydrate, and / or solvate thereof.
37. A composition comprising any one of the compounds of Claims 32 to 36 and an acceptable salt thereof.
38. The compound or composition according to any one of claims 32 to 37 for use in the preparation of a medicament used in the treatment of hypercholesterolemia and / or coronary heart disease.
Citation Information
Patent Citations
Designing an enzymatic peptide fragment condensation strategy
US20180187231A1
PCSK9 antagonist compounds
WO2019246349A1