Compositions and methods for the generation of peptide macrocycles

WO2025213103A8PCT designated stage Publication Date: 2026-05-15THE TRUSTEES OF PRINCETON UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE TRUSTEES OF PRINCETON UNIV
Filing Date
2025-04-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing methods for synthesizing peptide macrocycles using ATP-grasp enzymes are limited to specific substrate structures, lacking flexibility to produce diverse macrocyclic and branched peptides, which hinders drug discovery and manufacturing advancements.

Method used

A recombinant strategy utilizing an ATP-grasp enzyme to synthesize a variety of macrocyclic peptide structures by designing peptide substrates with specific donor and acceptor residues, enabling the formation of ester, amide, and thioester linkages, and employing native chemical ligation to create branched and cyclic peptides.

Benefits of technology

Enables the production of diverse macrocyclic and branched peptides, enhancing the potential for drug discovery and manufacturing by leveraging the substrate flexibility of ATP-grasp enzymes.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Disclosed herein are compositions and methods for the generation of peptide macrocycles from peptide substrates for ATP-grasp enzymes. The peptide substrates contain sufficient linkage cites to produce monocyclic, bicyclic, tricyclic, or multicyclic peptides comprising macrocyclic linkages selected from ω-ester linkages, ω-amide linkages, ω-thioester linkages, and combinations thereof.
Need to check novelty before this filing date? Find Prior Art

Description

5391.1039001 COMPOSITIONS AND METHODS FOR THE GENERATION OF PEPTIDE MACROCYCLES RELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Application No.63 / 574,627, filed on April 4, 2024. The entire teachings of the above application are incorporated herein by reference. GOVERNMENT SUPPORT

[0002] This invention was made with government support under Grant No. GM107036awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND

[0003] Cyclic peptides form the basis of several existing drugs and are a promisingmodality for future drug discovery. The cyclization of peptides often renders them more stable to proteolysis than their linear counterparts. Peptide cyclization can also serve to prepay some of the entropic penalty incurred upon binding a target, e.g., by engineering the conformational constraints required for the peptide to bind the target. Peptide cyclization can be achieved via synthetic methods and is also a hallmark of many peptidic natural products in which the cyclization is carried out by enzymes. The natural product superfamily of ribosomally synthesized and post-translationally modified peptides (RiPPs) includes many examples of cyclic peptides. One such RiPP family is the graspetides, in which peptide cyclization is achieved via ester and amide crosslinks between pairs of peptide side chains. The formation of these crosslinks is catalyzed by an adenosine triphosphate (ATP)-grasp enzyme. Several genome mining studies revealed thousands of putative graspetides encoded in bacterial genome. Along with recent experimental studies across several classes of graspetides, their highly diverse sequence patterns implicate a wide variety of unique, complex structures of macrocyclic and multi-macrocyclic peptides. While widespread across different phyla of bacteria, ATP-grasp enzymes have shown a high degree of substrate fidelity, with little to no room for synthesizing peptides whose structures differ from their natural substrates. - 1 - 4144725.v15391.1039001

[0004] Accordingly, methods, compositions, and other tools are needed to utilize thespecific functionality of ATP-grasp enzymes, for example, in applications such as improving the manufacture of existing drugs and discovering new drugs. SUMMARY

[0005] The present disclosure demonstrates that an ATP-grasp enzyme synthesizes awide variety of macrocyclic peptide structures from variants of its natural substrate. Additionally, a recombinant strategy to leverage this enzyme to produce various branched and cyclic peptides is disclosed.

[0006] In one aspect, the disclosure provides a peptide substrate for an ATP-graspenzyme, the peptide substrate comprising at least one core peptide unit. Each core peptide unit comprises a stem and a loop. A stem comprises at least one pair of macrocyclic linkage sites, and each pair of macrocyclic linkage sites comprises a donor residue and an acceptor residue. A loop comprises at least 4 amino acid residues.

[0007] In some embodiments, the donor residue of at least one pair of macrocycliclinkage sites is selected from serine (Ser; S), valine (Val; V), cysteine (Cys; C), and Lysine (Lys; K). In some embodiments, at least one pair of macrocyclic linkage sites is a pair of esterification sites. In some embodiments, the donor residue of a pair of esterification sites is Ser or Thr. In some embodiments, at least one pair of macrocyclic linkage sites is a pair of amidation sites. In some embodiments, the donor residue of a pair of amidation sites is Lys. In some embodiments, at least one pair of macrocyclic linkage sites is a pair of thioesterification sites. In some embodiments, the donor residue of a pair of thioesterification sites is Cys.

[0008] In some embodiments, the acceptor residue of at least one pair of macrocycliclinkage sites is selected from glutamic acid (Glu; E) and asparagine (Asn; N). In some embodiments, the acceptor residue of at least one pair of macrocyclic linkage sites is aspartic acid (Asp; D).

[0009] In some embodiments, the stem comprises an N-terminal portion connected to anN-terminal end of the loop and a C-terminal portion connected to a C-terminal end of the loop. In some embodiments, the donor residue of each pair of macrocyclic linkage sites is on the N-terminal portion of the stem. In some embodiments, the acceptor residue of each pair of macrocyclic linkage sites is on the C-terminal portion of the stem. - 2 - 4144725.v15391.1039001

[0010] In some embodiments, the stem comprises at least two pairs of macrocycliclinkage sites. In some embodiments, the stem comprises at least three pairs of macrocyclic linkage sites. In some embodiments, each donor residue is separated from an adjacent donor residue by 1-4 amino acid residues. In some embodiments, each acceptor residue is separated from an adjacent acceptor residue by 1-4 amino acid residues.

[0011] In some embodiments, the N-terminal portion of the stem of at least one corepeptide unit comprises an amino acid sequence selected from SEQ ID NOs: 82-101. In some embodiments, the C-terminal portion of the stem of at least one core peptide unit comprises an amino acid sequence selected from SEQ ID NOs: 102-114.

[0012] In some embodiments, the loop comprises at least 8 amino acid residues. In someembodiments, the loop comprises a peptide of interest or a protein of interest. In some embodiments, the loop further comprises at least one linker connecting the peptide of interest or protein of interest to the stem. In some embodiments, the protein of interest is a fluorescent protein.

[0013] In some embodiments, the peptide substrate comprises at least two core peptideunits, wherein adjacent core peptide units are connected by a linker. In some embodiments, the peptide substrate comprises at least three core peptide units.

[0014] In some embodiments, the peptide substrate further comprises a leader peptideconnected to an N-terminal end of a first core peptide unit. In some embodiments, the peptide substrate further comprises a tail peptide connected to a C-terminal end of a last core peptide unit. In some embodiments, the peptide substrate further comprises a detectable moiety. In some embodiments, the detectable moiety comprises a protein tag or a hydrazide.

[0015] In some embodiments, the peptide substrate is expressed in a host cell. In someembodiments, the peptide substrate and the ATP-grasp enzyme are co-expressed in a host cell. In some embodiments, the host cell is an E. coli cell.

[0016] In some embodiments, the ATP-grasp enzyme is a graspetide synthase. In someembodiments, the ATP-grasp enzyme is encoded in the genome of Thermobifida fusca. In some embodiments, the ATP-grasp enzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232).

[0017] In another aspect, the disclosure provides a system for generating a macrocycle,comprising a peptide substrate and an ATP-grasp enzyme.

[0018] In some embodiments, the macrocycle comprises one or more macrocycliclinkages. Each macrocyclic linkage is between a donor residue and an acceptor residue of a - 3 - 4144725.v15391.1039001 pair of macrocyclic linkage sites. In some embodiments, the one or more macrocyclic linkages are selected from ω-ester linkages, ω-amide linkages, ω-thioester linkages, and a combination of the foregoing. In some embodiments, the macrocyclic linkages are installed by the ATP-grasp enzyme.

[0019] In yet another aspect, the disclosure provides a polynucleotide encoding a peptidesubstrate disclosed herein. In some embodiments, a sequence of the polynucleotide that encodes the loop of at least one core peptide unit of the peptide substrate comprises a restriction site.

[0020] In yet another aspect, the disclosure provides a vector comprising a polynucleotideencoding a peptide substrate disclosed herein. In some embodiments, the vector further comprises an additional polynucleotide encoding an ATP-grasp enzyme. In some embodiments, the ATP-grasp enzyme is a graspetide synthase. In some embodiments, the ATP-grasp enzyme is encoded in the genome of Thermobifida fusca. In some embodiments, the ATP-grasp enzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232).

[0021] In yet another aspect, the disclosure provides a host cell comprising apolynucleotide encoding a peptide substrate disclosed herein, or a vector that comprises the polynucleotide encoding the peptide substrate. In some embodiments, the vector may further comprise an additional polynucleotide encoding an ATP-grasp enzyme. In some embodiments, the host cell further comprises an additional vector encoding an ATP-grasp enzyme. In some embodiments, the ATP-grasp enzyme is a graspetide synthase. In some embodiments, the ATP-grasp enzyme is encoded in the genome of Thermobifida fusca. In some embodiments, the ATP-grasp enzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232).

[0022] In yet another aspect, the disclosure provides a kit comprising a system disclosedherein for generating a macrocycle, a polynucleotide disclosed herein encoding the peptide substrate, a vector disclosed herein comprising the polynucleotide, or a host cell disclosed herein comprising a polynucleotide or vector disclosed herein.

[0023] In yet another aspect, the disclosure provides a method for generating amacrocycle, the method comprising contacting the peptide substrate with an ATP-grasp enzyme.

[0024] In some embodiments, the ATP-grasp enzyme is a graspetide synthase. In someembodiments, the ATP-grasp enzyme is encoded in the genome of Thermobifida fusca. In - 4 - 4144725.v15391.1039001 some embodiments, the ATP-grasp enzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232).

[0025] In yet another aspect, the disclosure provides a method for generating a head-to-tail cyclized macrocycle, the method comprising contacting the peptide substrate with an ATP-grasp enzyme, thereby producing a thioester-bearing macrocycle; and contacting the thioester-bearing macrocycle with trypsin.

[0026] In some embodiments, the peptide substrate comprises at least two cysteine (Cys)donor residues on an N-terminal portion of the stem and at least one acceptor residue on a C- terminal portion of the stem. In some embodiments, the N-terminal portion of the stem comprises SEQ ID NO: 101. In some embodiments, the C-terminal portion of the stem comprises SEQ ID NO: 113.

[0027] In some embodiments, the ATP-grasp enzyme is a graspetide synthase. In someembodiments, the ATP-grasp enzyme is encoded in the genome of Thermobifida fusca. In some embodiments, the ATP-grasp enzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232).

[0028] In yet another aspect, the disclosure provides a method for generating a branched,cyclic, or multicyclic peptide, the method comprising contacting the peptide substrate with an ATP-grasp enzyme, thereby producing a thioester-bearing macrocycle; and performing native chemical ligation (NCL) on the thioester-bearing macrocycle. In some embodiments, the NCL is performed with free cysteine. In some embodiments, NCL is performed with 2-(4- sulfanylphenyl)acetic acid (MPAA).

[0029] In some embodiments, the peptide substrate comprises at least two cysteine (Cys)donor residues on an N-terminal portion of the stem and at least one acceptor residue on a C- terminal portion of the stem. In some embodiments, the N-terminal portion of the stem comprises SEQ ID NO: 101. In some embodiments, the C-terminal portion of the stem comprises SEQ ID NO: 113.

[0030] In some embodiments, the ATP-grasp enzyme is a graspetide synthase. In someembodiments, the ATP-grasp enzyme is encoded in the genome of Thermobifida fusca. In some embodiments, the ATP-grasp enzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232). - 5 - 4144725.v15391.1039001 BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The foregoing will be apparent from the following more particular description ofexample embodiments, as illustrated in the accompanying drawings in which like reference characters refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating embodiments.

[0032] FIGs. 1A-F. Assessing the tolerance of the pre-fuscimiditide (ThfA) structure toamino acid substitutions. (FIG.1A) Schematics of the core peptide of pre-fuscimiditide (SEQ ID NO: 1), showing the demarcation for the esterification sites, the stem, and the loop. (FIGs. 1B-D) Deconvoluted mass spectra of ThfA variants with amino acid substitutions targeting the (FIG.1B) esterification sites, (FIG.1C) the stem macrocycle, and (FIG.1D) the loop macrocycle. (FIGs.1E-F) Schematics of the core peptide structures of the ThfA variants with loop substituted as (FIG.1E) GSG (SEQ ID NO: 15) and (FIG.1F) (GSSG)6(SEQ ID NO: 9) peptide sequences, as suggested by tandem mass spectrometry (MS / MS).

[0033] FIGs. 2A-B. Mass spectrometry analysis of the tryptic core fragment of the ThfB-modified ThfA D18E variant after hydrazinolysis (hydrazine represented by Z). All identified ions are listed in Table 7. (FIG.2A) MS1 analysis showing +14 Da shift in unesterified hydrazinolyzed species, +0 Dalton (Da) shift in singly esterified hydrazinolyzed species, and a lack of doubly esterified species expected with a -18 Da shift. (FIG.2B) Ion-ladder interpretation (top; SEQ ID NO: 18) of MS2 analysis (bottom) shows detected b-ions and γ- ions in black at corresponding fragmentation loci.

[0034] FIGs. 3A-B. Mass spectrometry analysis of the tryptic core fragment of the ThfB-modified ThfA D22E variant after hydrazinolysis. All identified ions are listed in Table 8.

[0035] FIG. 4. Mass spectrometry analysis of the ThfB-modified ThfA R5A and E20Avariants. Deconvoluted mass spectra of the R5A and E20A variants.

[0036] FIG. 5. Mass spectrometry analysis of the ThfB-modified ThfA variants (SEQ IDNOs: 21-24) for changing the size of the stem macrocycle. (Left) Cartoon showing the sequence and intended ester connectivity of the variants. That of the wild type (WT) is shown at the far left as a reference. (Right) Deconvoluted mass spectra of the ThfB-modified variants as whole proteins.

[0037] FIGs. 6A-B. (FIG. 6A) Extracted ion chromatograms (EICs) of the massesindicated for the ThfB-modified ThfA GSG-loop variant. (FIG.6B) Mass spectrometry analysis of the ThfB-modified ThfA GSG-loop variant (ion-ladder interpretation, SEQ ID - 6 - 4144725.v15391.1039001 NO: 25). All identified MS / MS ions are listed in Table 9. The black dashed line indicates an ester linkage.

[0038] FIGs. 7A-D. (FIG. 7A) Extracted ion chromatograms (EICs) of the massesindicated for the ThfB-modified ThfA (GSSG)6-loop variant. (FIGs.7B-D) Mass spectrometry analysis of the ThfB-modified ThfA (GSSG)6-loop variant (ion-ladder interpretations: FIG.7B, SEQ ID NOs: 26-27; FIG.7C, SEQ ID NOs: 27-28; FIG.7D, SEQ ID NO: 9). All identified ions from the MS / MS spectra in FIGs.7B-D are listed in Tables 10- 12, respectively.

[0039] FIG. 8. Mass spectrometry analysis of the ThfB-modified ThfA variants withamino acid substitutions at esterification sites.

[0040] FIG. 9. Mass spectrometry analysis of the ThfB-modified ThfA variants withfurther deletions from loop truncation. (Left) Cartoon showing the sequence and intended ester connectivity of the variants. That of the wild type (WT) is shown at the far left as a reference. (Right) Deconvoluted mass spectra of the ThfB-modified variants as whole proteins. All variants exhibit no modification, remaining as linear peptides.

[0041] FIGs. 10A-G. Fluorescent protein (FP) inserted into the fuscimiditide loop. (FIG.10A) Schematics representing the general amino acid (aa) sequence of FP-tethered ThfA constructs. LDP, leader peptide. (FIG.10B) Deconvoluted mass spectra of (left) mRuby2, (middle) ffGFP, (right) mTurquoise2-inserted (top) ThfA and (bottom) ThfB-modified ThfA (mThfAB). The peak marked by an asterisk symbol (*) represents the species with a mature chromophore and no ester. (FIG.10C) MS spectra and (FIG.10D) extracted ion chromatograms of tryptic fragments of the stem region. The peaks correspond to the 1018.42 Da, 1534.67 Da, and 1552.68 Da fragments 1-3, left-to-right in FIG.10C and top-to-bottom in FIG.10D, with ω-ester linkage connectivity shown in FIG.10E, FIG.10F, and FIG.10G, respectively.

[0042] FIG. 11. Deconvoluted mass spectra of ffGFP-inserted and mTurquoise2-insertedmThfABafter dithiothreitol (DTT) treatment. The left spectra correspond to ffGFP-inserted mThfAB, and the right spectra correspond to mTurquoise2-inserted mThfAB. The top spectra obtained before DTT treatment are identical to those in FIG.10B. They are present in FIG. 11 for comparison against the mass spectra obtained after DTT treatment, shown as the bottom spectra. DH, dehydration.

[0043] FIG. 12. Fluorescence comparison between mRuby2, mRuby2-inserted ThfA(ThfA-mRuby2), and mRuby2-inserted mThfAB(mThfAB-mRuby2). Relative fluorescence - 7 - 4144725.v15391.1039001 was calculated by dividing the measured fluorescence value by molar concentration and then normalizing by the highest calculated value. The height of each bar represents the average of 4 replicates, and the lengths of the error bars represent the standard deviations. The differences in relative fluorescence between mRuby2, ThfA-mRuby2, and mThfAB-mRuby2 are not statistically significant.

[0044] FIG. 13. Tandem mass spectrometry analysis of the 1018.42 Da tryptic fragmentfrom ffGFP-inserted mThfAB(1) (ion-ladder interpretation, SEQ ID NO: 46). All identified ions are listed in Table 13.

[0045] FIG. 14. Tandem mass spectrometry analysis of the 1534.67 Da tryptic fragmentfrom mTurquoise2-inserted mThfAB (2) (ion-ladder interpretation, SEQ ID NOs: 27, 46). All identified ions are listed in Table 14.

[0046] FIG. 15. Tandem mass spectrometry analysis of the 1552.67 Da tryptic fragmentfrom mRuby2-inserted mThfAB(3) (ion-ladder interpretation, SEQ ID NOs: 27, 46). All identified ions are listed in Table 15.

[0047] FIGs. 16A-B. Installing a third ω-ester linkage in fuscimiditide. The vertical linesin the cartoons represent the expected ω-ester linkages. In the deconvoluted mass spectra, a peak corresponding to -54 Da shift from the unmodified mass indicates the desired third ester crosslink.6 and 8 (FIG.16A) exhibit the desired three-fold modifications by the ATP-grasp enzyme.

[0048] FIG. 17. Sequence alignment of the C-terminal region of ThfA, T. alba putativegraspetide precursor, and 6. Solid-lined boxes represent identical matches between ThfA’s ester-forming residues and the putative graspetide precursor from T. alba, which were preserved in 6. Shaded boxes represent other identical matches. Dashed-lined boxes represent the intended third ester linkage in 6. based on T. alba putative graspetide precursor.

[0049] FIG. 18. Extracted ion chromatograms (EICs) of the tryptic fragments of 9. Thefirst EIC (mass-to-charge ratio (m / z) = 771.62) at the top corresponds to the 29 aa long triply esterified core peptide. The second EIC (m / z = 650.80) corresponds to the 25 aa long doubly esterified core peptide. The third EIC (m / z = 655.30) corresponds to the 25 aa long singly esterified core peptide. The last EIC (m / z = 659.80) at the bottom corresponds to the 25 aa long linear core peptide. Peaks annotated with an asterisk symbol (*) were further investigated with MS / MS experiments (refer to FIGs.19-22).

[0050] FIG. 19. Tandem mass spectrometry analysis of the singly esterified core peptideof 9 (ion-ladder interpretation, SEQ ID NO: 61). All identified ions are listed in Table 16. - 8 - 4144725.v15391.1039001

[0051] FIG. 20. Tandem mass spectrometry analysis of the major product of the doublyesterified core peptide of 9 (ion-ladder interpretation, SEQ ID NO: 61). All identified ions are listed in Table 17.

[0052] FIG. 21. Tandem mass spectrometry analysis of the minor product of the doublyesterified core peptide of 9 (ion-ladder interpretation, SEQ ID NO: 61). All identified ions are listed in Table 18.

[0053] FIG. 22. Tandem mass spectrometry analysis of the triply esterified core peptideof 9 (ion-ladder interpretation, SEQ ID NO: 52). All identified ions are listed in Table 19.

[0054] FIGS. 23A-C. ThfA variants harboring multivalent core peptides. Cartoonsrepresent the intended design of the multivalent core peptides, with black lines representing the intended placements of the ω-ester linkages. Deconvoluted mass spectra are shown. (FIG. 23A) Variants with divalent core peptides testing different linker sequences. (FIG.23B) Variant with trivalent, bis-macrocyclic core peptide. (FIG.23C) Variant with trivalent, monocyclic core peptide.

[0055] FIG. 24. Mass spectrometry analysis of the trivalent, monocyclic core peptidefrom trypsin digest. (Top) Extracted ion chromatogram (EIC) of the 63 aa long tryptic fragment with m / z value of 1270.18 and z = 5, matching the expected mass of 6345.91 Da. (Bottom) Mass spectrum extracted from the peak shown above. The vertical dotted grey lines mark the expected monoisotopic masses for the triply, doubly, and singly esterified and linear 63 aa long tryptic fragment.

[0056] FIGs. 25A-C. Tandem mass spectrometry analysis of the trivalent, monocycliccore peptide. Ion-ladder interpretation (FIG.25A) and MS / MS spectrum (FIG.25B) of the trivalent and monocyclic core peptide (SEQ ID NO: 69). All identified ions are listed in Table 20.

[0057] FIGs. 26A-B. Formation of side chain-to-side chain linkages in pre-fuscimiditide.(FIG.26A) Crosslinks in ThfA catalyzed by ThfB. The ThfB enzyme natively installs ω-ester crosslinks between Thr and Asp (left); the possibility of ω-amide or ω-thioester crosslinks was also investigated. (FIG.26B) Deconvoluted mass spectra of ThfB-modified ThfA variants designed to harbor an ω-amide (top) or ω-thioester linkage (bottom). Two-fold dehydration observed across the variants suggests the nonnative ω-amide or ω-thioester linkage formation.

[0058] FIG. 27. MS / MS of the doubly cross-linked core peptide from the ThfB-modifiedThfA T3K variant after ester-selective hydrazinolysis (ion-ladder interpretation, SEQ ID NO: - 9 - 4144725.v15391.1039001 71). The quadruply protonated, net singly dehydrated, and net singly hydrazinolyzed precursor ion with m / z value of 601.5 (labeled as M4+-4Da) was subjected to 16.9 volt (V) collision induced dissociation (CID). All identified ions are listed in Table 21. Limited fragmentation was observed with most fragment ions exhibiting relative intensity (RI; with respect to the tallest peak) less than 30%, probably due to the macrocyclic nature of the peptide. b2, y19, y20, and y21were the only regular b and y ions present, and all y ions predominantly exhibited a net mass shift of -4 Da, corresponding to the presence of a hydrazide and either a cross-link or an acylium ion putatively emerged from a-cleavage of a likely Lys3-Asp22 ω-amide linkage. A pentapeptide internal ion corresponding to PGQSD (SEQ ID NO: 116) with hydrazide was detected, though at a high error rate of 15.3 parts per million error (ppm). The presence of this ion further corroborates the Thr7-Asp18 ω-ester linkage present prior to ester-selective hydrazinolysis. The Z in the peptide sequence represents the expected aspartyl hydrazide residue.

[0059] FIG. 28. MS / MS of the singly cross-linked intermediate core peptide obtainedfrom the ThfB-modified ThfA T3K variant (ion-ladder interpretation, SEQ ID NO: 72). This core peptide was obtained by trypsin cleavage at Lys3, resulting in a 19 aa, truncated core peptide. The triply protonated, net singly dehydrated precursor ion with m / z value of 711.5 (labeled as M3+1DH) was subjected to 25V collision induced dissociation (CID). All identified ions are listed in Table 22. Limited fragmentation was observed with most fragment ions exhibiting relative intensity (RI; with respect to the tallest peak) less than 10%, probably due to the macrocyclic nature of the peptide. All b ions in this experiment exhibited a -18 Da shift from putative ω-ester McLafferty rearrangement, and y ion fragmentation was not observed past Asp18, suggesting Thr7-Asp18 ω-ester linkage. The fragment ion labeled as [M-K]2+1DH corresponds to the doubly protonated, singly dehydrated ion corresponding to the precursor ion with an internal lysine residue (likely Lys13) cleaved out of the peptide. [M-GQK]2+1DH corresponds to a similar ion with the tripeptide GQK cleaved out.

[0060] FIG. 29. MS / MS of the first core peptide obtained from the ThfB-modified ThfAT7K variant (ion-ladder interpretation, SEQ ID NO: 73). The triply protonated, net singly dehydrated precursor ion with m / z value of 797.0 (labeled as M3+1DH) was subjected to 30V collision induced dissociation. All identified ions are listed in Table 23. Limited fragmentation was observed with fragment ions exhibiting relative intensity (RI; with respect to the tallest peak) less than 20%, probably due to the macrocyclic nature of the peptide. All y ions identified in this spectrum harbored the AGTMR pentapeptide (SEQ ID NO: 27) cross- - 10 - 4144725.v15391.1039001 linked with Asp22, suggesting an internal trypsin cleavage event at Arg5 with Thr3-Asp22 ω- ester linkage present. An ion with mass corresponding to a dehydrated AGTMR pentapeptide (SEQ ID NO: 27) was present, putatively emerged from McLafferty rearrangement of the Thr3-Asp22 ω-ester. No b ions were detected, along with no further y ion fragmentation past Asp18, suggesting the presence of the Lys7-Asp18 ω-amide linkage.

[0061] FIG. 30. MS / MS of the second core peptide obtained from the ThfB-modifiedThfA T7K variant (ion-ladder interpretation, SEQ ID NO: 73). The triply protonated, net singly dehydrated precursor ion with m / z value of 797.0 (labeled as M3+1DH) was subjected to 30V collision induced dissociation (CID). All identified ions are listed in Table 24. Unlike the MS / MS experiment of the first core peptide (FIG.27), the CID of this peptide yielded many high-intensity fragment ions, exhibiting a lower extent of macrocyclization. All y ions identified in this spectrum harbored the AGTMR pentapeptide (SEQ ID NO: 27) cross-linked with Asp22, suggesting an internal trypsin clevage event at Arg5 with Thr3-Asp22 ω-ester linkage present. An ion with mass corresponding to a dehydrated AGTMR pentapeptide (SEQ ID NO: 27) was present, putatively emerged from McLafferty rearrangement of the Thr3-Asp22 ω-ester. Additionally, all y ions harboring Lys13 and Asp18 exhibited a -18 Da mass shift, with no b and y ion fragmentation detected between Lys13 and Asp18, suggesting the presence of the Lys7-Asp18 ω-amide linkage.

[0062] FIG. 31. Iodoacetamide-alkylation of the ThfB-modified ThfA T3C variant. Thereaction was performed on the whole protein and then digested with trypsin, yielding a leader peptide fragment represented as M1 (left; SEQ ID NO: 74) and the core peptide fragment harboring two ThfB-installed cross-links represented as M2 (right; SEQ ID NO: 75). Hence, in this experiment, M1 served as an internal positive control for alkylation. Here, the extracted ion chromatograms (EICs) of the alkylated species and non-alkylated species of M1 (targeting m / z values of 668.32 and 649.32, respectively) and M2 (targeting m / z values of 601.28 and 587.02, respectively) are compared. M1 was alkylated from the reaction, consistent with the presence of one free cysteine. On the other hand, M2 remained non- alkylated, suggesting Cys3 to be cross-linked.

[0063] FIG. 32. Iodoacetamide-alkylation of the ThfB-modified ThfA T7C variant. Thereaction was performed on the whole protein and then digested with trypsin, yielding a leader peptide fragment represented as M1 (left; SEQ ID NO: 74) and the core peptide fragment harboring two ThfB-installed cross-links represented as M2 (right; SEQ ID NO: 76). Hence, in this experiment, M1 served as an internal positive control for alkylation. Here, the - 11 - 4144725.v15391.1039001 extracted ion chromatograms (EICs) of the alkylated species and non-alkylated species of M1 (targeting m / z values of 668.32 and 649.32, respectively) and M2 (targeting m / z values of 601.28 and 587.02, respectively) are compared. The bottom left peak corresponds to a doubly protonated species with a monoisotopic m / z value of 647.76, unrelated to M1. The fourth isotope of this species coincides with the target m / z value of 649.32. Similarly, the first peak in the bottom right EIC trace corresponds to a quadruply protonated species with m / z value of 586.51, different than the targeted quadruply protonated M2 with m / z value of 587.00, which corresponds to the second peak. Based on the observed species related to M1 and M2, the presence of alkylated M1 is consistent with the presence of one free cysteine, and the lack of alkylation of M2 suggests Cys7 to be cross-linked.

[0064] FIGs. 33A-C. Native chemical ligation on the putative Cys3-Asp22 linkage of thepre-fuscimiditide T3C variant, generating a branched peptide. (FIG.33A) Proposed scheme of the native chemical ligation of cysteine on the pre-fuscimiditide variant containing Thr7- Asp18 ω-ester and Cys3-Asp22 ω-thioester (13). Adding 4-mercaptophenylacetic acid (MPAA) resulted in quantitative conversion to the product containing the isopeptide-bonded cysteine (14). (FIG.33B) Mass spectra for before the reaction (0 hours (h)), after the reaction without MPAA (16 h), with MPAA (16 h +MPAA), and with MPAA and tris(2- carboxyethyl)phosphine (TCEP) added prior to liquid chromatography-mass spectrometry (LC-MS) analysis (16 h +MPAA reduced). The mass peaks in the left part of the spectra correspond to 13, and those at the right correspond to 14 or disulfide-bonded 14. (FIG.33C) The MS / MS fragmentation of 14 with the added Cys shown at the far right of the ladder sequence (ion-ladder interpretation, SEQ ID NO: 77). All detected ions are listed in Table 25. All y-ion fragments contained the ligated cysteine, and y20, y21, and y22exhibited a -18 Da shift due to the T7-D18 ω-ester linkage. The b7and b8ions arose from the McLafferty rearrangement.

[0065] FIG. 34. Extracted ion chromatograms (EICs) for the native chemical ligation ofcysteine on to pre-fuscimiditide T3C variant. Top EICs correspond to the pre-fuscimiditide T3C variant (13), and bottom EICs correspond to the Cys-ligated product (14). The EICs on the left, middle, and right were extracted from the LC-MS experiments on a purified stock of 13 prior to mixing with other reaction components, after the reaction without MPAA, and after the reaction with MPAA followed by further reduction with TCEP, respectively. - 12 - 4144725.v15391.1039001

[0066] FIG. 35. Deconvoluted mass spectrum of ThfB-modified ThfA A1C / T3C variant.The protein exhibited a -38 Da mass shift corresponding to the presence of two crosslinks installed by ThfB and one disulfide bond.

[0067] FIGs. 36A-C. Intramolecular macrocyclic rearrangement of the pre-fuscimiditideA1C / T3C variant. (FIG.36A) Trypsin cleaved the leader peptide and revealed Cys1 at the N- terminus. Subsequently, intramolecular native chemical ligation occurred, with the Cys1 thiol attacking the Cys3-Asp22 ω-thioester followed by the N-S acyl shift. This yielded the head- to-tail cyclized bis-macrocyclic structure 15, with an isopeptide bond formed between the Cys1 amine and the side chain of Asp22. (FIG.36B) Two-dimensional nuclear magnetic resonance (2D NMR) experiments supporting the Cys1-Asp22 isopeptide bond. The structure on the left summarizes selected total correlation spectroscopy (TOCSY) and nuclear Overhauser effect spectroscopy (NOESY) signals. Right: overlay of TOCSY (thick line) and NOESY (thin line) showing the correlations that establish the isopeptide bond between the Cys1 amide proton (HN) and the Asp22 beta protons (Hb). (FIG.36C) The solution structure of 15. The Thr7-Asp18 ω-ester linkage and the Cys1-Asp22 isopeptide bond are shown.

[0068] FIG. 37. Mass spectra of the head-to-tail cyclized pre-fuscimiditide A1C / T3Cvariant (16) undergoing iodoacetamide reaction. The top spectrum represents the peptide prior to the reaction, and the bottom spectrum represents the peptide after the reaction. The peptide exhibited a quantitative conversion to a species with +114 Da mass shift corresponding to two-fold alkylation consistent with the presence of two free thiol groups. M corresponds to the monoisotopic mass of the doubly dehydrated pre-fuscimiditide A1C / T3C variant.

[0069] FIG. 38. Mass spectra of the head-to-tail cyclized pre-fuscimiditide A1C / T3Cvariant (16) undergoing ester-selective hydrazinolysis reaction. The top spectrum represents the peptide prior to the reaction, and the bottom spectrum represents the peptide after the reaction. The reacted peptide predominantly exhibited a +32 Da mass shift corresponding to one-fold ester hydrazinolysis. A negligible amount of two-fold hydrazinolyzed product was observed, consistent with the presence of only one ester. M corresponds to the monoisotopic mass of the doubly dehydrated pre-fuscimiditide A1C / T3C variant.

[0070] FIGs. 39A-B. MS / MS spectrum of the head-to-tail cyclized pre-fuscimiditideA1C / T3C variant after ester-selective hydrazinolysis (ion-ladder interpretations, SEQ ID NOs: 78-80). Collision-induced dissociation was performed at 35V collision energy, selecting m / z = 804 as the precursor ion. The injected sample was a crude reaction mixture from ester- - 13 - 4144725.v15391.1039001 selective hydrazinolysis performed on the head-to-tail cyclized pre-fuscimiditide A1C / T3C variant (16). The MS / MS spectrum was extracted from 7.07-7.27 min retention time range. The nomenclature system created by Ngoka and Gross (1999)39is used here to label the bnJZfragment ions, where n represents the number of amino acid residues present in the fragment ion, J represents the residue preceding (i.e. N-terminal to) the cleavage site, and Z represents the residue following (i.e. C-terminal to) the cleavage site. The fragment ions containing Asp18 exhibited +14 Da mass shift, indicating the presence of the hydrazide at that residue. b17KP, b19KP, and b21KPions contain Asp22 and Cys1, confirming the head-to-tail cyclization of the peptide. Table 26 lists the major fragment ions detected and labeled in the MS / MS spectrum shown above and their error values.

[0071] FIGs. 40A-B. Aspartimidylation of the pre-fuscimiditide A1C / T3C variant as afurther structural modification. (FIG.40A) Proposed scheme of Cys1-Asp22 aspartimidylation catalyzed by a protein L-isoaspartyl O-methyltransferase (PIMT). The bonds that comprise the original Cys1-Asp22 isoAsp are bolded to track how it is transformed during the PIMT-catalyzed aspartimidylation. The enzyme is known to recognize and methylate the β amino acid isoAsp using S-adenosyl L-methionine (SAM) as the methyl group donor. Subsequently, the covalent bond between the Cys1 amide and Asp22 α-methylester carbonyl carbon was formed, yielding the aspartimide. (FIG.40B) Mass spectra of the PIMT-modified pre-fuscimiditide A1C / T3C variant, (top) reacted for 20 h with PIMT and (bottom) incubated for 1 h in absence of PIMT after the post-reaction high performance liquid chromatography (HPLC) purification. The spectra display the peaks corresponding to +3 charged species of the starting peptide (middle, 0), the O-methylated peptide (right, +14 Da), and the aspartimidylated peptide (left, -18 Da).

[0072] FIG. 41. TOCSY spectrum of the head-to-tail cyclized pre-fuscimiditideA1C / T3C variant. The peak assignments are listed in Table 27.

[0073] FIG. 42. NOESY spectrum of the head-to-tail cyclized pre-fuscimiditideA1C / T3C variant with a mixing time of 150 milliseconds (ms). The peak assignments are listed in Table 27.

[0074] FIG. 43. NOESY spectrum of the head-to-tail cyclized pre-fuscimiditideA1C / T3C variant with a mixing time of 700 ms. The peak assignments are listed in Table 27.

[0075] FIG. 44. Comparison between the proton chemical shifts of the head-to-tailcyclized pre-fuscimiditide A1C / T3C variant (16) and pre-fuscimiditide. The absolute value of differences in chemical shifts for each respective proton of the two peptides were averaged by - 14 - 4144725.v15391.1039001 residue, except Cys1 / Ala1 and Cys3 / Thr3. The average differences in proton chemical shifts were lower in the loop region (Tyr8-Ser17) and higher in the stem region (Gly2, Met4-Thr7, Asp18-Asp22).

[0076] FIG. 45. Comparison between the top NMR structures of pre-fuscimiditide andthe head-to-tail cyclized A1C / T3C variant (16). The top structure of pre-fuscimiditide was fetched from Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) ID No.: 7LI2. The cartoons on the left show the main chain and ω-ester connectivity. In the top NMR structures on the right, the arrows indicate crosslinking of the peptide sequence (N-to-C terminal).

[0077] FIG. 46. Top 20 NMR structures of the head-to-tail cyclized pre-fuscimiditideA1C / T3C variant (16). The structures were energy-minimized using the MMFF94 force field42from molecular editor Avogadro41. DETAILED DESCRIPTION General

[0078] Unless defined otherwise, all technical and scientific terms used herein have thesame meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0079] As used herein, the indefinite articles “a” and “an” should be understood to mean“at least one” unless clearly indicated to the contrary.

[0080] The phrase “and / or”, as used herein, should be understood to mean “either orboth” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases.

[0081] It should also be understood that, unless clearly indicated to the contrary, in anymethods described herein that include more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited.

[0082] Unless otherwise indicated or otherwise evident from the context andunderstanding of one of ordinary skill in the art, values that are expressed as ranges can assume any specific value or subrange within the stated ranges in various embodiments, unless the context clearly dictates otherwise. - 15 - 4144725.v15391.1039001 Abbreviations

[0083] Selected Amino Acids:Arginine (Arg, R) Aspartic acid (Asp, D) Glutamic acid (Glu, E) Glutamine (Gln, Q) Glycine (Gly, G) Lysine (Lys, K) Serine (Ser, S) Threonine (Thr, T) Valine (Val, V)

[0084] MS / MS: tandem mass spectrometry, comprising first (MS1) and second (MS2)stages / spectra

[0085] RiPPs: ribosomally synthesized and post-translationally modified peptides

[0086] ThfA: Thermobifida fusca pre-fuscimiditide; a graspetide; substrate for ThfB

[0087] ThfB: Thermobifida fusca ATP-grasp enzyme (graspetide synthetase)Peptide Substrates

[0088] The natural product superfamily of ribosomally synthesized and post-translationally modified peptides (RiPPs) includes many examples of cyclic peptides. One such RiPP family is the graspetides, which are characterized by the presence of ester or amide side chain-side chain linkages resulting in peptide macrocycles. The term “graspetides” derives from the ATP-grasp enzymes that install the side chain-side chain linkages in the peptides. Some graspetides are also known as microviridins, ω-ester linked peptides (OEPs), or both. In some embodiments, a graspetide comprises a stem-loop structure, e.g., a hairpin structure.

[0089] Pre-fuscimiditide or fuscimiditide precursor (RefSeq Accession No.WP_011292231.1; SEQ ID NO: 283), referred to as ThfA herein, is a graspetide encoded in the genome of Thermobifida fusca. ThfA comprises a doubly dehydrated core peptide containing two ω-ester linkages between Thr7-Asp18 and Thr3-Asp22, forming the bis- macrocyclic structure comprising the stem and loop macrocycles.

[0090] MSTAVTDAFPLGRDENRNDQVTEWRPFGMRYGVQPTPIPVPLSDTKYDPDQQVLVVADGQPCAKIERAGTMRVTYPDGQKPGQSDVEKD (SEQ ID NO: 283) - 16 - 4144725.v15391.1039001

[0091] Without being bound by theory, it is understood that O-methyltransferase ThfM,an enzyme with homology to the protein repair catalyst protein L-isoaspartyl methyltransferase (PIMT), methylates an Asp residue of ThfA. The methyl ester is subsequently attacked by the adjacent backbone to form an aspartimide, thereby producing fuscimiditide. Fuscimiditide comprises a stem-loop structure peptide macrocycle with two sidechain-sidechain ester linkages (e.g., the two ω-ester linkages between Thr7-Asp18 and Thr3-Asp22 of pre-fuscimiditide) and an aspartimide modification within the backbone of the loop macrocycle. As used herein, mThfABrefers to ThfA modified by ThfB.

[0092] Certain aspects of the invention disclosed herein are drawn to a peptide substratefor an ATP-grasp enzyme, the peptide substrate comprising at least one core peptide unit. Each core peptide unit comprises a stem and a loop. As used herein, the terms “stem” and “loop” refer to structures of a modified peptide substrate (e.g., a macrocycle) or to corresponding regions of an unmodified peptide substrate.

[0093] “Stem” refers to the portion of a peptide substrate that harbors macrocycliclinkage sites. A stem comprises at least one pair of macrocyclic linkage sites, and each pair of macrocyclic linkage sites comprises a donor residue and an acceptor residue.

[0094] The stem comprises an N-terminal portion connected to an N-terminal end of theloop, as well as a C-terminal portion connected to a C-terminal end of the loop. In some embodiments, the donor residue of each pair of macrocyclic linkage sites is on the N-terminal portion of the stem. In some embodiments, the acceptor residue of each pair of macrocyclic linkage sites is on the C-terminal portion of the stem. Upon modification by an ATP-grasp enzyme, the donor and acceptor residues of a pair of macrocyclic linkage sites are cross- linked.

[0095] In some embodiments disclosed herein, a peptide substrate comprises an aminoacid sequence with at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to ThfA (SEQ ID NO: 1).

[0096] In some embodiments disclosed herein, a peptide substrate is a ThfA variant thatcomprises one or more amino acid substitutions, insertions, deletions, or a combination thereof, relative to wild-type ThfA. In some embodiments, a ThfA variant comprises one or more substitutions at an ester-forming residue (e.g., positions 3, 7, 18, 22), at a residue that forms part of the stem structure of wild-type ThfA (e.g., positions 4-6 and 19-21), at a residue - 17 - 4144725.v15391.1039001 that forms part of the loop structure of wild-type ThfA (e.g., positions 8-17), or a combination of the foregoing.

[0097] In some embodiments, a ThfA variant comprises a T3S or T7S substitution (e.g.,as in SEQ ID NO: 2 or 3, respectively) and is modified by ThfB to form at least a product exhibiting two-fold dehydration. In some embodiments, a ThfA variant comprises a D18E or D22E substitution (e.g., as in SEQ ID NO: 4 or 5, respectively) and yields at least a singly dehydrated product upon modification by ThfB. In some embodiments, a ThfA variant comprises amino acid substitutions in the stem macrocycle and yields a mixture of unmodified, singly dehydrated, and doubly dehydrated product upon modification by ThfB. In some embodiments, amino acid substitutions in the stem macrocycle comprise one or more of M4G, R5S, V6G, V19G, E20S, and K21G. In some embodiments, amino acid substitutions in the stem macrocycle comprise M4G, R5S, and V6G (e.g., as in SEQ ID NO: 6). In some embodiments, amino acid substitutions in the stem macrocycle comprise V19G, E20S, and K21G (e.g., as in SEQ ID NO: 7). Table 1. Examples of Peptide Substrate Variants

[0098] In some embodiments, the donor residue of at least one pair of macrocycliclinkage sites is selected from serine (Ser; S), valine (Val; V), threonine (Thr; T), cysteine (Cys; C), and Lysine (Lys; K).

[0099] In some embodiments, at least one pair of macrocyclic linkage sites is a pair ofesterification sites. In some embodiments, the donor residue of a pair of esterification sites is - 18 - 4144725.v15391.1039001 Ser, Val, or Thr. Wild-type ThfA comprises two pairs of esterification sites wherein the donor residue is Thr.

[0100] In some embodiments, at least one pair of macrocyclic linkage sites is a pair ofamidation sites. In some embodiments, the donor residue of a pair of amidation sites is Lys. Wild-type ThfA does not comprise amidation sites.

[0101] In some embodiments, at least one pair of macrocyclic linkage sites is a pair ofthioesterification sites. In some embodiments, the donor residue of a pair of thioesterification sites is Cys. Wild-type ThfA does not comprise thioesterification sites. Peptide substrates comprising at least one pair of thioesterification sites are useful, for example, in generating thioester-bearing macrocycles that can then be used as the basis for generating branched, multi-macrocyclic, and heat-to-tail cyclized peptides and proteins.

[0102] In some embodiments, the acceptor residue of at least one pair of macrocycliclinkage sites is selected from glutamic acid (Glu; E) and asparagine (Asn; N). In some embodiments, the acceptor residue of at least one pair of macrocyclic linkage sites is aspartic acid (Asp; D), as in wild-type ThfA.

[0103] In some embodiments, the N-terminal portion of the stem of at least one corepeptide unit comprises an amino acid sequence having at least, e.g., 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more, sequence identity to any one of SEQ ID NOs: 82-101. In some embodiments, the N-terminal portion of the stem of at least one core peptide unit comprises an amino acid sequence selected from SEQ ID NOs: 82-101.

[0104] In some embodiments, the C-terminal portion of the stem of at least one corepeptide unit comprises an amino acid sequence having at least, e.g., 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more, sequence identity to any one of SEQ ID NOs: 102-114. In some embodiments, the C-terminal portion of the stem of at least one core peptide unit comprises an amino acid sequence selected from SEQ ID NOs: 102-114. Table 2. Example Stem Sequences.- 19 - 4144725.v15391.1039001N-ter, N-terminal portion; C-ter, C-terminal portion

[0105] In some embodiments, the stem comprises at least two or three pairs ofmacrocyclic linkage sites. In such embodiments, each donor residue is separated from an adjacent donor residue by 1-4 amino acid residues. Similarly, in such embodiments, each acceptor residue is separated from an adjacent acceptor residue by 1-4 amino acid residues. Examples of the separation of adjacent donor or acceptor residues are found in FIGs. S4, 3A, and 3B. Table 3. Examples of Peptide Substrates Comprising Three Pairs of Macrocyclic Linkage Sites- 20 - 4144725.v15391.1039001-int, intermediate macrocycle

[0106] In some embodiments, a peptide substrate comprises a change to the sequence,length, or both sequence and length of the loop region corresponding to amino acid residues 8-17 of wild-type ThfA (see SEQ ID NO: 1). In some embodiments, a peptide substrate comprises a loop region that is about 0-3 or at least about 24 amino acid residues in length, and yields at least a singly dehydrated product upon modification by ThfB. In some embodiments, a peptide substrate comprises a loop region that is about 4-24 amino acid residues in length, and yields at least a doubly dehydrated product upon modification by ThfB. In some embodiments, a loop comprises at least 4 amino acid residues. In some embodiments, the loop comprises at least 8 amino acid residues.

[0107] In some embodiments, a peptide substrate comprises a large loop region, e.g., aloop region comprising a partial or complete peptide or protein. In some embodiments, the loop comprises a peptide of interest or a protein of interest. In some embodiments, the loop further comprises at least one linker connecting the peptide of interest or protein of interest to the stem. In some embodiments, a linker connecting the peptide of interest or protein of interest to the stem comprises SEQ ID NO: 56. In some embodiments, the peptide or protein comprises a fluorescent protein. In some embodiments, the fluorescent protein comprises mRuby2, ffGFP, or mTurquoise2. In some embodiments, a peptide substrate comprising a large loop region comprises an N-terminal region comprising SEQ ID NOs: 44. In some embodiments, a peptide substrate comprising a large loop region comprises a C-terminal region comprising SEQ ID NOs: 45. - 21 - 4144725.v15391.1039001

[0108] RAGTMRVTGSGSGS (SEQ ID NO: 44)

[0109] KGSGSGSDVEKD (SEQ ID NO: 45)

[0110] GSGSGS (SEQ ID NO: 56)

[0111] In some embodiments, a peptide substrate comprises any one of SEQ ID NOs: 8-17 and yields at least a singly dehydrated product upon modification by ThfB. In some embodiments, a peptide substrate comprises any one of SEQ ID NOs: 9-14 and yields at least a doubly dehydrated product upon modification by ThfB.

[0112] In some embodiments, a peptide substrate comprises amino acid rearrangements.In some embodiments, the rearrangements result in lengthening the stem region. In some embodiments, a peptide substrate comprises one or more rearrangements selected from T7Y / Y8T, S17D / D18S, G2T / T3G, V6T / T7V, D18V / V19D, T3M / M4T, and K21D / D22K. In some embodiments, a peptide substrate comprises T7Y / Y8T / S17D / D18S (e.g., as in SEQ ID NO: 21), G2T / T3G (e.g., as in SEQ ID NO: 22), V6T / T7V / D18V / V19D (e.g., as in SEQ ID NO: 23), or T3M / M4T / K21D / D22K (e.g., as in SEQ ID NO: 24) and yields at least a singly dehydrated product upon modification by ThfB. In some embodiments, a peptide substrate comprises G2T / T3G and yields at least a doubly dehydrated product upon modification by ThfB. - 22 - 4144725.v15391.1039001 Table 5. Examples of Stem-Variant Peptide Substrates

[0113] In some embodiments, the peptide substrate comprises at least two core peptideunits, and wherein adjacent core peptide units are connected by a linker. In some embodiments, the peptide substrate comprises at least three core peptide units, as in SEQ ID NOs: 67-68. In some embodiments, each core peptide unit is a variant of the wild type ThfA core peptide unit (SEQ ID NO: 62). As used herein, “wild type” refers to the canonical amino acid sequence as found in nature. As those of skill in the art would appreciate, a nucleic acid sequence can be modified, e.g., for codon optimization in a host cell (e.g., bacteria, yeast, and plant host cells).

[0114] AGTMRVTYPDGQKPGQSDVEKDSGTMRVTYPDGQKPGQSDVEKDSGTMRVTYPDGQKPGQSDVEKD (SEQ ID NO: 67)

[0115] AGVMRVTYPDGQKPGQSDVEANGPGVMAVTYPDGQKPGQSDVEANGPGVMAVTYPDGQKPGQSDVEAN (SEQ ID NO: 68)

[0116] TMRVTYPDGQKPGQSDVEKD (SEQ ID NO: 62)

[0117] In some embodiments, the linker connecting core peptide units comprises any oneof SEQ ID NOs: 63-66. Table 6: Example Core-to-core Linkers

[0118] In some embodiments, the peptide substrate further comprises a leader peptideconnected to an N-terminal end of a first core peptide unit. In some embodiments, the peptide substrate further comprises a tail peptide connected to a C-terminal end of a last core peptide - 23 - 4144725.v15391.1039001 unit. In some embodiments, the peptide substrate further comprises a detectable moiety. In some embodiments, the detectable moiety comprises a protein tag or a hydrazide. In some embodiments, the detectable moiety comprises a fluorescent protein (e.g., in addition to a loop comprising a fluorescent protein).

[0119] In some embodiments, the peptide substrate is expressed in a host cell. In someembodiments, the peptide substrate and the ATP-grasp enzyme are co-expressed in a host cell. In some embodiments, the host cell is an E. coli cell. ATP-Grasp Enzymes

[0120] Members of the adenosine triphosphate (ATP)-grasp enzyme superfamilycomprise an ATP-binding site called the ATP-grasp fold. Two domains of the ATP-grasp fold grasp an ATP, which is then used in catalysis of ligation reactions. In some embodiments, an ATP-grasp enzyme activates acceptor amino acid residues for attack by nucleophilic side chains of donor amino acid residues, thereby forming crosslinks between pairs of peptide side chains. In some embodiments, an ATP-grasp enzyme activates Glu or Asp amino acid residues for attack by Ser, Thr, or Lys, thereby forming omega (ω)-ester linkages and ω-amide linkages. In some embodiments, an ATP-grasp enzyme forms ω- thioester linkages. As used herein, the terms “ω-ester linkage” and “ester-linkage” are used interchangeably. As used herein, the terms “ω-amide linkage” and “amide-linkage” are used interchangeably. As used herein, the terms “ω-thioester linkage” and “thioester-linkage” are used interchangeably. As used herein, “ω” generally refers the involvement of a last atom in an amino acid residue side chain in a linkage.

[0121] In some embodiments, an ATP-grasp enzyme is a graspetide synthase, e.g., anenzyme that installs one or more side chain-side chain linkages in a substrate (e.g., a graspetide precursor peptide), thereby producing a graspetide. In some embodiments, one or more linkages installed by an ATP-grasp enzyme result in a macrocycle, i.e., a large molecule (e.g., a peptide or protein) comprising a cyclic structure. In some embodiments, such linkages are called macrocyclic linkages.

[0122] In some embodiments, the ATP-grasp enzyme is a graspetide synthase. In someembodiments, the ATP-grasp enzyme is encoded in the genome of Thermobifida fusca. ThfB is an example of a graspetide synthase. ThfB (RefSeq Accession No. WP_193587235.1, SEQ ID NO: 115; SEQ ID NO: 232) is encoded in the genome of the thermophilic actinobacterium - 24 - 4144725.v15391.1039001 Thermobifida fusca and installs macrocyclic linkages in its substrates, e.g., pre-fuscimiditide (ThfA) and fuscimiditide.

[0123] MTVLILTNPFDITADDVILRLTEHGVPVVRLDPADFPQQVVLHSEIGGNGWTGTLTTPHRILDLSTVTGIWYRRPRKFRLPAQMSQAEYEFAATEARRGFGGIINSLTG WINHPSAIGRAEYKPYQLHHAVQAGLNVPRTLITNDPKQAKGWCARVGDVVYKPLS APSWLENGDTYVVFTTPITPDQWGDPAIGRTAHMFQQRLDKEFEVRLTMVDGKAFP AAIHAHSDAARIDWRSDYDALTYSIPTVPQRVLTGARDLLRRLHLRYAALDFIVSPD GRWHFLEVNPNGQYGWIEEHTGQPISDAIADALTRKEN (SEQ ID NO: 115)

[0124] In some embodiments, the ATP-grasp enzyme comprises an amino acid sequencewith at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to ThfB (SEQ ID NO: 115 or 232). In some embodiments, the ATP-grasp enzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232).

[0125] In some embodiments, ThfB utilizes Thr, Ser, or Lys as a donor residue in theformation of a macrocyclic linkage. In some embodiments, ThfB utilizes Asp as an acceptor residue in the formation of macrocyclic linkages. Sequences, Vectors, and Host Cells

[0126] In yet another aspect, the disclosure provides a polynucleotide encoding a peptidesubstrate. In some embodiments, a sequence of the polynucleotide that encodes the loop of at least one core peptide unit of the peptide substrate comprises a restriction site.

[0127] As used herein, the term “polynucleotide” refers to a polymer comprising multiplenucleotide monomers (e.g., ribonucleotide monomers or deoxyribonucleotide monomers). “Polynucleotide” includes, for example, nucleic acid molecules such as DNA (e.g., genomic DNA and cDNA), RNA, and DNA-RNA hybrid molecules. Nucleic acid molecules can be naturally occurring, recombinant, or synthetic. In addition, nucleic acid molecules can be single-stranded, double-stranded or triple-stranded. In certain embodiments, nucleic acid molecules can be modified. In the case of a double-stranded polymer, “nucleic acid” can refer to either or both strands of the molecule.

[0128] The terms “nucleotide” and “nucleotide monomer” refer to naturally occurringribonucleotide or deoxyribonucleotide monomers, as well as non-naturally occurring derivatives and analogs thereof. Accordingly, nucleotides can include, for example, nucleotides comprising naturally occurring bases (e.g., adenosine, thymidine, guanosine, - 25 - 4144725.v15391.1039001 cytidine, uridine, inosine, deoxyadenosine, deoxythymidine, deoxyguanosine, or deoxycytidine) and nucleotides comprising modified bases known in the art.

[0129] As used herein, the term “identity” or “identical” refers to the extent to which twonucleotide sequences, or two amino acid sequences, have the same residues at the same positions when the sequences are aligned to achieve a maximal level of identity, expressed as a percentage. For sequence alignment and comparison, typically one sequence is designated as a reference sequence, to which test sequences are compared. The sequence identity between reference and test sequences is expressed as the percentage of positions across the entire length of the reference sequence where the reference and test sequences share the same nucleotide or amino acid upon alignment of the reference and test sequences to achieve a maximal level of identity. As an example, two sequences are considered to have 70% sequence identity when, upon alignment to achieve a maximal level of identity, the test sequence has the same nucleotide or amino acid residue at 70% of the same positions over the entire length of the reference sequence.

[0130] Alignment of sequences for comparison to achieve maximal levels of identity canbe readily performed by a person of ordinary skill in the art using an appropriate alignment method or algorithm. In some instances, the alignment can include introduced gaps to provide for the maximal level of identity. Examples include the local homology algorithm of Smith & Waterman, Adv. Appl. Math.2:482 (1981), the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol.48:443 (1970), the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), and visual inspection (see generally Ausubel et al., Current Protocols in Molecular Biology).

[0131] When using a sequence comparison algorithm, test and reference sequences areinput into a computer, subsequent coordinates are designated, if necessary, and sequence algorithm program parameters are designated. The sequence comparison algorithm then calculates the percent sequence identity for the test sequence(s) relative to the reference sequence, based on the designated program parameters. A commonly used tool for determining percent sequence identity is Protein Basic Local Alignment Search Tool (BLASTP) available through National Center for Biotechnology Information, National - 26 - 4144725.v15391.1039001 Library of Medicine, of the United States National Institutes of Health. (Altschul et al., 1990).

[0132] In various embodiments, two nucleotide sequences, or two amino acid sequences,can have at least, e.g., 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more, sequence identity. When ascertaining percent sequence identity to one or more sequences described herein, the sequences described herein are the reference sequences.

[0133] In yet another aspect, the disclosure provides a vector comprising a polynucleotideencoding a peptide substrate. In some embodiments, the vector further comprises an additional polynucleotide encoding an ATP-grasp enzyme. In some embodiments, the ATP- grasp enzyme is a graspetide synthase. In some embodiments, the ATP-grasp enzyme is encoded in the genome of Thermobifida fusca. In some embodiments, the ATP-grasp enzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232).

[0134] The terms “vector”, “vector construct” and “expression vector” mean the vehicleby which a DNA or RNA sequence (e.g. a foreign gene) can be introduced into a host cell, so as to transform the host and promote expression (e.g. transcription and translation) of the introduced sequence. Vectors typically comprise the DNA of a transmissible agent, into which foreign DNA encoding a protein is inserted by restriction enzyme technology. A common type of vector is a “plasmid”, which generally is a self-contained molecule of double-stranded DNA that can readily accept additional (foreign) DNA and which can readily introduced into a suitable host cell. A large number of vectors, including plasmid and fungal vectors, have been described for replication and / or expression in a variety of eukaryotic and prokaryotic hosts.

[0135] In yet another aspect, the disclosure provides a host cell comprising apolynucleotide encoding a peptide substrate or a vector. The vector comprises the polynucleotide encoding a peptide substrate, and may further comprise an additional polynucleotide encoding an ATP-grasp enzyme. In some embodiments, the host cell further comprises an additional vector encoding an ATP-grasp enzyme. In some embodiments, the ATP-grasp enzyme is a graspetide synthase. In some embodiments, the ATP-grasp enzyme is encoded in the genome of Thermobifida fusca. In some embodiments, the ATP-grasp enzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232). In some embodiments, the peptide substrate is expressed in the host cell. In - 27 - 4144725.v15391.1039001 some embodiments, the peptide substrate and the ATP-grasp enzyme are co-expressed in the host cell.

[0136] The terms “express” and “expression” mean allowing or causing the informationin a gene or DNA sequence to become manifest, for example producing a protein by activating the cellular functions involved in transcription and translation of a corresponding gene or DNA sequence. A DNA sequence is expressed in or by a cell to form an “expression product” such as a protein. The expression product itself, e.g. the resulting protein, may also be said to be “expressed” by the cell. A polynucleotide or polypeptide is expressed recombinantly, for example, when it is expressed or produced in a foreign host cell under the control of a foreign or native promoter, or in a native host cell under the control of a foreign promoter.

[0137] Gene delivery vectors generally include a transgene (e.g., nucleic acid encodingan enzyme) operably linked to a promoter and other nucleic acid elements required for expression of the transgene in the host cells into which the vector is introduced. Suitable promoters for gene expression and delivery constructs are known in the art.

[0138] For bacterial host cells, suitable promoters, include, but are not limited topromoters obtained from the E. coli lac operon, Streptomyces coelicolor agarase gene (dagA), Bacillus subtilis levansucrase gene (sacB), Bacillus licheniformis alpha-amylase gene (amyL), Bacillus stearothermophilus maltogenic amylase gene (amyM), Bacillus amyloliquefaciens alpha-amylase gene (amyQ), Bacillus licheniformis penicillinase gene (penP), Bacillus subtilis xy1A and xy1B genes, and prokaryotic beta-lactamase gene (See e.g., Villa-Kamaroff et al., Proc. Natl. Acad. Sci. USA 75: 3727-3731, 1978), as well as the tac promoter (See e.g., DeBoer et al., Proc. Natl. Acad. Sci. USA 80: 21-25, 1983). Examples of promoters for filamentous fungal host cells, include, but are not limited to promoters obtained from the genes for Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral alpha-amylase, Aspergillus niger acid stable alpha- amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (See e.g., WO 96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoters from the genes for Aspergillus niger neutral alpha-amylase and Aspergillus oryzae triose phosphate isomerase), and mutant, truncated, and hybrid promoters thereof. Examples of yeast cell promoters can be from the genes for Saccharomyces cerevisiae enolase (ENO-1), - 28 - 4144725.v15391.1039001 Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase. Other useful promoters for yeast host cells are known in the art (See e.g., Romanos et al., Yeast 8:423-488, 1992). The selection of a suitable promoter is within the skill in the art. The recombinant plasmids can also comprise inducible, or regulatable, promoters for expression of an enzyme in cells.

[0139] Although the genetic code is degenerate in that most amino acids are representedby multiple codons (called “synonyms” or “synonymous” codons), it is understood in the art that codon usage by particular organisms is nonrandom and biased towards particular codon triplets. Accordingly, in some embodiments, the vector includes a nucleotide sequence that has been optimized for expression in a particular type of host cell (e.g., through codon optimization). Codon optimization refers to a process in which a polynucleotide encoding a protein of interest is modified to replace particular codons in that polynucleotide with codons that encode the same amino acid(s), but are more commonly used / recognized in the host cell in which the nucleic acid is being expressed. In some aspects, the polynucleotides described herein are codon optimized for expression in a bacterial cell, e.g., E. coli.

[0140] Oligonucleotides used in the construction of polynucleotides, plasmids, andvectors disclosed herein are listed in Table 28. Amino acid sequences of peptides and proteins encoded by such polynucleotides, plasmids, and vectors are listed in Table 29.

[0141] Various gene delivery vehicles are known in the art and include both viral andnon-viral (e.g., naked DNA, plasmid) vectors. Viral vectors suitable for gene delivery are known to those skilled in the art. Such viral vectors include, e.g., vector derived from the herpes virus, baculovirus vector, lentiviral vector, retroviral vector, adenoviral vector and adeno-associated viral vector (AAV). Vectors derived from plant viruses can also be used, such as the viral backbones of the RNA viruses Tobacco mosaic virus (TMV), Potato virus X (PVX) and Cowpea mosaic virus (CPMV), and the DNA geminivirus Bean yellow dwarf virus. The viral vector can be replicating or non-replicating.

[0142] Non-viral vectors include naked DNA and plasmids, among others. Non-limitingexamples include pKK plasmids (Clonetech), pUC plasmids, pET plasmids (Novagen, Inc., Madison, Wis.), pRSET or pREP plasmids (Invitrogen, San Diego, Calif.), or pMAL plasmids (New England Biolabs, Beverly, Mass.), and such vectors may be introduced into many appropriate host cells, using methods disclosed or cited herein or otherwise known to those skilled in the relevant art. - 29 - 4144725.v15391.1039001

[0143] In certain embodiments, the vector comprises a transgene operably linked to apromoter. The transgene encodes a biologically active molecule, such as an enzyme described herein.

[0144] To facilitate the introduction of the gene delivery vector into host cells, the vectorcan be combined with different chemical means such as colloidal dispersion systems (macromolecular complex, nanocapsules, microspheres, beads) or lipid-based systems (oil-in- water emulsions, micelles, liposomes).

[0145] In some embodiments, the host cell is a bacterial cell. A wide variety of bacterialcells are suitable, such as cells of the genus Escherichia, including Escherichia coli; cells of the genus Bacillus, including Bacillus subtilis; cells of the genus Pseudomonas, including Pseudomonas aeruginosa; and cells of the genus Streptomyces, including Streptomyces griseus. In some embodiments, the hosts cells are cultured in a cell culture medium, such as a standard cell culture medium known in the art to be suitable for the particular host cell.

[0146] In yet another aspect, the disclosure provides a kit comprising any embodiment ofthe system for generating a macrocycle, the polynucleotide encoding the peptide substrate, the vector comprising the polynucleotide, or the host cell comprising the vector. Systems and Methods

[0147] In another aspect, the disclosure provides a system for generating a macrocycle,comprising a peptide substrate and an ATP-grasp enzyme.

[0148] In some embodiments, the macrocycle comprises one or more macrocycliclinkages. Each macrocyclic linkage is between a donor residue and an acceptor residue of a pair of macrocyclic linkage sites. In some embodiments, the one or more macrocyclic linkages are selected from ω-ester linkages, ω-amide linkages, ω-thioester linkages, and a combination of the foregoing. In some embodiments, the macrocyclic linkages are installed by the ATP-grasp enzyme.

[0149] In yet another aspect, the disclosure provides a method for generating amacrocycle, the method comprising contacting the peptide substrate with an ATP-grasp enzyme. An example mechanism for formation macrocyclic linkages is illustrated in FIG. 26A.

[0150] In yet another aspect, the disclosure provides a method for generating a head-to-tail cyclized macrocycle, the method comprising contacting the peptide substrate with an ATP-grasp enzyme, thereby producing a thioester-bearing macrocycle; and contacting the - 30 - 4144725.v15391.1039001 thioester-bearing macrocycle with trypsin. An example mechanism for the generation of a head-to-tail cyclized macrocycle is illustrated in FIG.36A. An example linear sequence corresponding to a head-to-tail cyclized macrocycle is SEQ ID NO: 81.

[0151] CGCMRVTYPDGQKPGQSDVEKD (SEQ ID NO: 81)

[0152] In some embodiments, the peptide substrate comprises at least two cysteine (Cys)donor residues on an N-terminal portion of the stem and at least one acceptor residue on a C- terminal portion of the stem. In some embodiments, the N-terminal portion of the stem comprises SEQ ID NO: 101. In some embodiments, the C-terminal portion of the stem comprises SEQ ID NO: 113.

[0153] In yet another aspect, the disclosure provides a method for generating a branched,cyclic, or multicyclic peptide, the method comprising contacting the peptide substrate with an ATP-grasp enzyme, thereby producing a thioester-bearing macrocycle; and performing native chemical ligation (NCL) on the thioester-bearing macrocycle. In some embodiments, the NCL is performed with free cysteine. In some embodiments, NCL is performed with 2-(4- sulfanylphenyl)acetic acid (MPAA). An example mechanism for performing NCL on a thioester-bearing macrocycle is illustrated in FIG.33A.

[0154] In some embodiments, the peptide substrate comprises at least two cysteine (Cys)donor residues on an N-terminal portion of the stem and at least one acceptor residue on a C- terminal portion of the stem. In some embodiments, the N-terminal portion of the stem comprises SEQ ID NO: 101. In some embodiments, the C-terminal portion of the stem comprises SEQ ID NO: 113. Example Embodiments

[0155] Disclosed is a recombinant method for the head-to-tail cyclization ofpolypeptides. Polypeptides as short as 4 amino acids and up to full globular proteins can be tagged at their N- and C- termini and co-expressed with an enzyme to effect their cyclization. The disclosed methodology provides a strategy for the preparation of libraries of cyclic peptides for drug discovery applications. At the protein level, it provides a method for the stabilization of an arbitrary protein which may result in increased thermostability of the protein and / or improved properties such as binding affinity or enzymatic activity.

[0156] Several aspects of the invention are described below, with reference to examplesfor illustrative purposes only. It should be understood that numerous specific details, relationships, and methods are set forth to provide a full understanding of the invention. One - 31 - 4144725.v15391.1039001 having ordinary skill in the relevant art, however, will readily recognize that the invention can be practiced without one or more of the specific details or practiced with other methods, protocols, reagents, cell lines and animals. The present invention is not limited by the illustrated ordering of acts or events, as some acts may occur in different orders and / or concurrently with other acts or events. Furthermore, not all illustrated acts, steps or events are required to implement a methodology in accordance with the present invention. Many of the techniques and procedures described, or referenced herein, are well understood and commonly employed using conventional methodology by those skilled in the art.

[0157] Unless otherwise defined, all terms of art, notations and other scientific terms orterminology used herein are intended to have the meanings commonly understood by those of skill in the art to which this invention pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art. It will be further understood that terms, such as those defined in commonly-used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and / or as otherwise defined herein.

[0158] The terminology used herein is for the purpose of describing particularembodiments only and is not intended to be limiting. As used herein, the indefinite articles “a”, “an” and “the” should be understood to include plural reference unless the context clearly indicates otherwise.

[0159] Disclosed herein is a method for the generation of peptide macrocycles. Moreparticularly, disclosed are compositions of matter in the form of specific macrocyclic peptides, as well as a methodology for the production of macrocyclic peptides. Also disclosed is the ability to covalently constrain a protein of interest in a cyclized form to improve its function and / or stability.

[0160] The disclosed recombinant method allows for the head-to-tail cyclization ofpolypeptides. Polypeptides as short as 4 amino acids and up to full globular proteins can be tagged at their N- and C- termini and co-expressed with an enzyme to effect their cyclization. The disclosed methodology provides a strategy for the preparation of libraries of cyclic peptides for drug discovery applications. At the protein level, it provides a method for the stabilization of an arbitrary protein which may result in increased thermostability of the protein and / or improved properties such as binding affinity or enzymatic activity. - 32 - 4144725.v15391.1039001

[0161] Cyclic peptides are an exciting and novel modality in drug discovery. Themethodology disclosed herein will enable the construction of libraries of cyclic peptides in a straightforward and cost-effective fashion to enable new screens in drug discovery.

[0162] An additional important feature disclosed herein is the stabilization of a protein ofinterest. Proteins with marginal stability can be rendered more stable via covalent cyclization between their N- and C- termini. Potential applications include the stabilization of proteins used as drugs as well as proteins used in industrial applications.

[0163] The disclosed approach provides a straightforward method of generating a cyclicpeptide or protein. A type of enzyme called ATP-grasp connects the N- and C-terminal portions of a protein / peptide. While other methods for peptide cyclization exist, such as solid- phase peptide synthesis, they are more difficult to scale up or utilize for the generation of libraries of compounds.

[0164] The disclosed compositions of matter include peptides and proteins that have beengenerated by recombinant expression in E. coli.

[0165] The disclosed approach also includes a process whereby a peptide or protein ofinterest is cloned into a recombinant DNA construct that adds tags to the N- and C-termini of the peptide / protein. This engineered peptide / protein is then co-expressed in E. coli with a specific ATP-grasp enzyme. This results in the formation of the cyclized peptide / protein in the cells. In an alternative process, the engineered peptide / protein of interest can be purified from E. coli first and then treated with the ATP-grasp enzyme in vitro to generate the cyclized peptide / protein.

[0166] Samples of cyclized peptides and proteins have been produced using the disclosedmethods, and one can also readily adapt other peptides or proteins of interest into the disclosed process.

[0167] The disclosed approach can be employed, inter alia, by pharmaceutical and / orbiotech companies who are engaged in macrocyclic peptide drug discovery, as well as pharma / biotech companies that produce therapeutic proteins. Other industries that utilize enzymes, such as the food industry or the personal care industry, may also find utility in the disclosed ability to stabilize enzymes. Laundry detergent, for example, utilizes proteolytic enzymes that can be stabilized for operation at higher temperatures using the disclosed technology.

[0168] Without limiting the scope of the claimed invention, example embodiments are setforth below. - 33 - 4144725.v15391.10390011. A cell, comprising:a first sequence configured to express a non-cyclized peptide or protein of interest having tags at its N- and C-termini; a second sequence configured to express an ATP-grasp enzyme configured to covalently connect at least one first residue at N-terminus with at least one second residue at the C-terminus.2. The cell of embodiment 1, wherein the cell is an E. coli cell.3. The cell of embodiment 1 or 2, wherein the ATP-grasp enzyme is ThfB.4. The cell of any one of embodiments 1-3, wherein the non-cyclized peptide or proteinof interest is fused to at least one additional non-cyclized peptide or protein of interest via at least one linker.5. The cell of any one of embodiment 1-4, wherein the first residue or the second residueincludes a Thr or Cys residue and the other of the first residue or the second residue includes an Asp residue.6. A method for generating a peptide macrocycle, comprising:providing a cell of any one of embodiments 1-5; coexpressing the non-cyclized peptide or protein of interest and the ATP-grasp enzyme; and stabilizing the peptide or protein by allowing the ATP-grasp enzyme to covalently connect the N- and C-termini, thereby forming a cyclized peptide or protein of interest.7. The method of embodiment 6, further comprising lysing the cell and collecting lysatecomprising the cyclized peptide or protein of interest.8. The method of embodiment 6 or 7, wherein the ATP-grasp enzyme is ThfB.9. A method for stabilizing a peptide or protein of interest, comprising:providing a non-cyclized peptide or protein of interest having tags at the N- and C-termini; and generating a cyclized peptide or protein by treating the non-cyclized peptide or protein of interest with an ATP-grasp enzyme, causing at least one residue at the N- terminus to covalently connect to a residue at the C-terminus.10. The method of embodiment 9, wherein providing the non-cyclized peptide or proteinof interest includes purifying lysate of a cell configured to express the non-cyclized peptide or protein of interest. - 34 - 4144725.v15391.103900111. The method of embodiment 9 or 10, wherein the ATP-grasp enzyme is ThfB.12. The method of any one of embodiment 9-11, wherein the tags include a first tagcomprising a Thr or Cys residue and a second tag comprising an Asp residue. EXEMPLIFICATION Introduction

[0169] Cyclic peptides form the basis of several existing drugs and are a promisingmodality for future drug discovery1-2. The cyclization of peptides often renders them more stable to proteolysis than their linear counterparts. Peptide cyclization can also serve to prepay some of the entropic penalty incurred upon binding a target, e.g., by engineering the conformational constraints required for the peptide to bind the target. Peptide cyclization can be achieved via synthetic methods3-4and is also a hallmark of many peptidic natural products5-7in which the cyclization is carried out by enzymes. The natural product superfamily of ribosomally synthesized and post-translationally modified peptides (RiPPs) includes many examples of cyclic peptides. One such RiPP family is the graspetides, in which peptide cyclization is achieved via ester and amide crosslinks between pairs of peptide side chains (FIG.1A)8-9. The formation of these crosslinks is catalyzed by an ATP-grasp enzyme that activates Glu or Asp residues for attack by nucleophilic side chains (Ser, Thr, or Lys). Several genome mining studies revealed thousands of putative graspetides encoded in bacterial genome10-13. Along with recent experimental studies across several classes of graspetides11, 13-23, their highly diverse sequence patterns implicate a wide variety of unique, complex structures of macrocyclic and multi-macrocyclic peptides. While widespread across different phyla of bacteria, ATP-grasp graspetide synthetases have shown a high degree of substrate fidelity, with little to no room for synthesizing peptides whose structures differ from their natural substrates17. Here, we demonstrate that the ATP-grasp enzyme for pre- fuscimiditide and fuscimiditide, graspetides encoded in the genome of Thermobifida fusca, exhibits unprecedented substrate promiscuity. This enzyme, designated ThfB, is able to synthesize a wide variety of macrocyclic and multi-macrocyclic peptide structures. Additionally, a recombinant strategy to leverage this enzyme to produce various branched and cyclic peptides is disclosed. - 35 - 4144725.v15391.1039001 Methods Plasmid Construction

[0170] All oligonucleotides and plasmids used in the present disclosure are described inTables 28 and 29, respectively.

[0171] For molecular cloning, Golden Gate assembly was used as a primary method toconstruct Thermobifida fusca fuscimiditide precursor (ThfA) variant plasmids, using the pBC108 vector generated in an amycolimiditide study (Choi et al.202221). Golden Gate assembly is a recombinant DNA assembly method that utilizes one or more type IIS restriction enzymes (e.g., BsaI, BsmBI, Esp3I, BbsI, SapI) and one or more DNA ligases (e.g., T4 DNA ligase, T7 DNA ligase) to generate constructs from DNA fragments. pBC108 was generated from pQE-80L (see, e.g., sequence and map published by QIAGEN), modified to contain the fast folding green fluorescent protein (ffGFP) constitutive expression cassette flanked by BsaI recognition sites under the control of PglpT promoter (Part: jtk2821; BBa_J72163 from iGEM Registry of Standard Biological Parts), and to mutate the arginine residue of the His6-tag to serine to suppress the background methylation of the N-terminus in E. coli. Inserts were prepared by polymerase chain reaction (PCR) amplification using the reagents (deoxynucleotide triphosphates (dNTP), Q5 High Fidelity DNA polymerase) purchased from New England Biolabs (NEB; Ipswich, MA, USA). As a template for PCR amplification, the plasmid encoding the thfA gene in multicloning site 1 (MCS1) of the pRSF-duet vector (see, e.g., Sigma-Aldrich Cat. No.71341) was used for single amino acid substitutions. For pBC267 encoding the ThfA A1C / T3C variant, the plasmid encoding the ThfA T3C variant (pBC255) generated in this study was used. Oligonucleotide primers for PCR amplification were designed with overhangs of the primers encoding the designed amino acid substitutions and purchased from Integrated DNA Technologies (IDT; Coralville, IA, USA). Amplified inserts were purified using ZymocleanTMGel DNA Recovery Kit (Zymo Research, Irvine, CA, USA) after gel electrophoresis. Short inserts (<61 bp) were prepared by oligonucleotide annealing using T4 polynucleotide kinase (NEB). Golden Gate assembly was carried out with BsaI-HF®v2 and T4 DNA ligase (NEB).

[0172] To clone the Thermobifida fusca ATP-grasp enzyme (graspetide synthetase)(ThfB) gene in the pRSF-duet vector (pBC262), for compatible co-expression with a ThfA variant plasmid, traditional restriction digestion and ligation was used. The insert was prepared by amplifying the ThfB gene in the pHE33 plasmid (Elashal et al.2022)20with the primers oBC352 and oBC035. The amplified insert was subjected to gel electrophoresis and - 36 - 4144725.v15391.1039001 purification and then digested with BsaI-HF®v2 restriction enzyme. The plasmid backbone compatible for ligation with the insert was prepared by digesting pRSF-duet with NcoI and HindIII restriction enzymes (NEB). The insert and the backbone were subjected to gel electrophoresis and purification, followed by ligation using T4 DNA ligase.

[0173] To propagate the assembled plasmids, E. coli XL1-BLUE® (Stratagene, AgilentTechnologies, Santa Clara, CA) cells were transformed using Mix & Go! Transformation Kit (Zymo Research) and grown in lysogeny broth (LB) medium (5 g L-1yeast extract (IBI Scientific, Dubuque, IA, USA), 10 g L-1tryptone (IBI Scientific), 10 g L-1sodium chloride, supplemented with 50 µg mL-1kanamycin or 100 µg mL-1ampicillin, as needed for selection). Plasmids were recovered and purified using QIAPREP®Spin Miniprep Kit (QIAGEN, Germantown, MD, USA). The DNA sequences of the inserts of all constructed plasmids were confirmed by Sanger Sequencing (GENEWIZ / Azenta Life Sciences, Burlington, MA, USA). ThfA and ThfB Coexpression in E. coli

[0174] E. coli BL21 (DE3) ΔslyD was transformed with a ThfA variant plasmid andpBC262 by electroporation. A transformed colony was grown initially in 6 mL LB medium supplemented with 50 µg mL-1kanamycin and 100 µg mL-1ampicillin overnight.500 mL of the same medium was inoculated with the overnight culture at optical density of 0.02 measured at a wavelength of 600 nm (OD600). The culture was incubated with shaking at 37 °C until reaching 0.5 OD600, spiked with 500 mL 1 M isopropyl β-D-thiogalactopyranoside (IPTG), and then further incubated with shaking at room temperature overnight (16-18 hours). Purification of ThfB-modified ThfA

[0175] A culture with E. coli BL21 (DE3) Δ slyD cells expressing ThfA and ThfB wascentrifuged at 4,000 x g for 15 minutes. The cell pellet was resuspended in 10 mL urea buffer (100 mM NaH2PO4, 10 mM Tris-base, 8 M urea) at pH 8.0 and then freeze-thawed once at - 80 °C and in water bath at room temperature. The cell lysate resulting from ThfA-ThfB coexpression was highly viscous compared to cell lysates prepared from expressing other graspimiditides (namely, amycolimiditide21and albusimiditide).13The cell lysate was centrifuged at maximum rotational speed (at least 16,000 x g) until clarified (usually 15 minutes). The clarified lysate containing the His-tagged ThfB-modified ThfA and ThfB was - 37 - 4144725.v15391.1039001 incubated with 1 mL nickel-nitrilotriacetic acid (Ni-NTA) resin (QIAGEN) at 4 °C for 1 hour. After the initial pass through a gravity column, the retained resin was washed with 10 mL urea buffer (pH 6.4), then with 10 mL urea buffer (pH 5.9). Protein elution was collected with 8 mL urea buffer (pH 4.5), fractionated by 1 mL. LC-MS Analysis of ThfB-modified ThfA as a Whole protein

[0176] For all liquid chromatography-mass spectrometry (LC-MS) or liquidchromatography-tandem mass spectrometry (LC-MS / MS) analyses, mass spectra were acquired using the positive ion mode for electrospray ionization (ESI+). The instrument was calibrated (LC-MS) or tuned (LC-MS / MS) prior to analysis using the “Mass Calibration / Check” method in the Agilent MassHunter Workstation Data Acquisition software.

[0177] To analyze the ThfB-modified ThfA as a whole protein, the second elutionfraction usually containing the highest concentration of eluted proteins was diluted 10-fold in ultrapure water and injected into XBRIDGE®Protein BEH C4 column (2.1 mm x 50 mm, 3.5 μm particle size; Waters Corporation, Milford, MA, USA) installed on Agilent 6530 Quadrupole Time of Flight (QTOF) mass spectrometer equipped with an Agilent 1260 Infinity II liquid chromatography (LC) unit. The following gradients using a binary mixture of solvents (Solvent A: ultrapure water with 1% formic acid; Solvent B: acetonitrile with 1% formic acid) were used for chromatography: 10% Solvent B, 0-2 minutes; 10-50% B, 1-15 minutes; 50-90% B, 15-20 minutes; 90% B, 20-30 minutes, flowing at 0.5 mL minutes-1. Deconvoluted mass spectra of the analyte (ThfB-modified ThfA variant) were acquired using the Agilent MassHunter Bioconfirm software. LC-MS and LC-MS / MS Analysis of a Core Peptide

[0178] The core peptide portion of a ThfB-modified ThfA variant was prepared bytrypsin digestion. For graspetide variants subjected to LC-MS analysis only, the second elution fraction containing the whole protein in 8 M urea buffer (pH 4.5) was diluted 10-fold in ultrapure water. For other variants to be further analyzed by a different experiment (e.g. A1C T3C variant for nuclear magnetic resonance (NMR) studies), all elution fractions were collected, buffer-exchanged in 100 mM phosphate buffer (pH 7.0) and concentrated using AMICON®Ultra-410 kDa cut-off centrifugal filters (Millipore, Burlington, MA, USA). The resulting protein sample prepared by either method was digested with 1:50 amount (by mass) - 38 - 4144725.v15391.1039001 of trypsin (Promega Corporation, Madison, WI, USA) in 50 mM ammonium phosphate (pH approximately 8 after preparation) at 37 °C for 15-30 minutes, which is long enough to digest the leader peptide. The reaction was quenched with 1% formic acid.

[0179] Either a crude trypsin digestion mixture or a high performance liquidchromatography (HPLC)-purified core peptide (see below) was injected into the ZORBAX®300SB-C18 (2.1 mm x 50 mm, 3.5 μm particle size; Agilent) column installed on the Agilent 1260 Infinity II LC-Agilent 6530 QTOF system. The following gradients using a binary mixture of solvents (Solvent A: ultrapure water with 1% formic acid; Solvent B: acetonitrile with 1% formic acid) were used for chromatography: 5% B, 0-1 minutes; 5-45% B, 1-20 minutes; 45-90% B, 20-25 minutes; 90% B, 25-30 minutes, flowing at 0.5 mL minutes-1. The resulting LC traces and mass spectra were extracted and analyzed using the Agilent MassHunter Qualitative Analysis software for manual inspection. The Agilent MassHunter Bioconfirm software was used to automate the identification of the tryptic fragments.

[0180] For LC-MS / MS analysis, collision-induced dissociation (CID) was performedwith 1.3 m / z isolation width and a defined collision energy (based on the following formula unless specified otherwise). V = 0.036 × (m / z) – 4.8 MS / MS spectra were extracted from the Agilent MassHunter Qualitative Analysis software and analyzed manually using mMass, a software tool for the annotation of cyclic peptide tandem mass spectra.36Fragment ions were identified and assigned based on the m / z error of the monoisotopic peak, ion intensity, and the isotopic distribution. HPLC Purification of a Core Peptide

[0181] The core peptide of a ThfB-modified ThfA variant was obtained by trypsindigestion (using 1:200 amount of trypsin) and purified by semi-preparative reverse-phase HPLC using the Agilent 1200 series instrument, equipped with the ZORBAX®300SB-C18 (9.4 mm x 250 mm, 5 μm) column and UV detector wavelength set at 215 nm. The following gradients using a binary mixture of solvents (Solvent A: ultrapure water with 1% trifluoroacetic acid (TFA); Solvent B: acetonitrile with 1% TFA) were used for chromatography: 10% B, 0-1 minutes; 10-50% B, 1-20 minutes; 50-90% B, 20-25 minutes; 90% B, 25-30 minutes; 90-10% B, 30-32 minutes, flowing at 4.0 mL minutes-1. The desired peak was collected, and the identity of the purified peptide was confirmed by LC-MS analysis. The purified peptide was frozen -80 °C and subsequently lyophilized (FreeZone - 39 - 4144725.v15391.1039001 Freeze Dry System; Labconco Corporation, Kansas City, Mo, USA). The resulting powder was resuspended in ultrapure water. Fluorescence Measurements

[0182] A protein construct containing red fluorescent protein mRuby2 was buffer-exchanged in a buffer containing 50 mM tris(hydroxymethyl)aminomethane hydrochloride (Tris-HCl), 200 mM NaCl, pH 8.0, and concentrated using AMICON®Ultra-410 kDa (or 30 kDa, as appropriate) cut-off centrifugal filters. The protein was further purified by size exclusion chromatography using an ÄKTApurifier chromatography system (GE HealthCare Technologies, Inc., Chicago, IL, USA), fitted with SUPERDEX®200 Increase 10 / 300 GL column (Cytiva, Marlborough, MA, USA) and run with a buffer containing 50 mM Tris-HCl, 200 mM sodium chloride at pH 8.0 at 1 mL minute-1flow rate. Fluorescence was measured in BioTek Synergy 4 plate reader (Agilent) with excitation wavelength at 559 nm and emission wavelength at 600 nm. Protein / Peptide Alkylation with Iodoacetamide

[0183] The experimental protocol was performed according to manufacturer instructions(Thermo Scientific, Waltham, MA, USA).50 mg of protein or peptide was mixed with 5 mL of 2% sodium dodecyl sulfate (SDS) and 45 mL of 200 mM ammonium bicarbonate. Ultrapure water was added to the resulting mixture up to 100 mL total volume.5 mL of 200 mM tris(2-carboxyethyl)phosphine (TCEP) was added afterwards and incubated at 55 °C for 1 hour. Then, 5 mL of 375 mM iodoacetamide solution prepared in the 200 mM ammonium bicarbonate solution was added, and the resulting mixture was incubated at room temperature for 30 minutes in absence of light. Peptide samples were immediately subjected to LC-MS analysis afterwards. Protein samples were subjected to trypsin digestion (as described above) prior to LC-MS analysis. Ester-selective Hydrazinolysis

[0184] 26 µL of an HPLC-purified peptide (at variable concentration, 0.1-3.0 mg mL-1)was mixed with the hydrazinolysis solution (10 µL of 35 wt% hydrazine solution in water (Sigma-Aldrich), 15 µL of 1X phosphate buffered saline (PBS), pH 5.0 (Cold Spring Harbor),371.0 µL 6 M HCl). The mixture was heated at 55 °C for 45 min and then subjected to LC-MS or LC-MS / MS analysis. - 40 - 4144725.v15391.1039001 Cysteine Native Chemical Ligation on the Pre-fuscimiditide T3C Variant

[0185] Pre-fuscimiditide (ThfA) T3C variant was prepared by coexpressing ThfA T3Cvariant with ThfB and HPLC-purifying the core peptide as described above. A buffer solution dedicated to this experiment (named “ligation buffer”) was prepared, containing 8 M urea, 20 mM TCEP, and 200 mM sodium phosphate, at final pH of approximately slightly under 7 (as indicated by a pH indicator strip).50 mM of the core peptide was mixed with 1 mM L- cysteine and, optionally, with 20 mM 4-mercaptophenylacetic acid (MPAA), in the ligation buffer. The reaction mixture was incubated at 37 °C for 16 hours, followed immediately by LC-MS analysis. For the reaction with MPAA, 2 mM TCEP was added prior to the LC-MS experiment. Overall, four LC-MS experiments were conducted on (1) the HPLC-purified pre- fuscimiditide T3C variant prior to native chemical ligation, and the resulting crude reaction mixture after native chemical ligation (2) without MPAA, (3) with MPAA, and (4) with MPAA followed by addition of 20 mM TCEP after the reaction. Head-to-tail Cyclized Pre-fuscimiditide A1C / T3C NMR

[0186] ThfA A1C / T3C variant was coexpressed with ThfB and purified from 15 L ofLB+Amp+Kan (with kanamycin and ampicillin) medium in total.11 mg of the core peptide was purified by trypsin digestion and HPLC, lyophilized, and then reconstituted in 600 mL 98:2 H2O:D2O mixture, resulting in the final peptide concentration of 7.7 mM. Two- dimensional (2D) NMR spectra, total correlation spectroscopy (TOCSY) and nuclear Overhauser effect spectroscopy (NOESY) of the peptide were acquired on a Bruker Avance III 800 MHz spectrometer (Bruker Corporation, Billerica, MA, USA) at 25 °C. The TOCSY experiment was performed with the mixing time of 80 ms, and two NOESY experiments were performed with the mixing times of 150 ms and 700 ms. The acquired 2D NMR spectra were processed and analyzed using MestReNova software (Mestrelab Research, Santiago de Compostela, Spain). Proton chemical shifts (listed in Table 27) were assigned manually using the TOCSY spectrum overlaid with the NOESY spectrum acquired with 700 ms mixing time. The NOESY spectrum acquired with 150 ms mixing time was used to integrate the assigned peaks, and the integrated volumes were used as through space distance restraints in CYANA 2.1 automated NMR protein structure calculation software40to construct the peptide structural model. Additional geometric restraints for the Thr7-Asp18 ω-ester linkage and Cys1-Asp22 isopeptide bond were added manually. Then, the top 20 structures acquired from - 41 - 4144725.v15391.1039001 CYANA 2.1. were further processed in Avogadro molecular editor41, energy-minimizing them using the MMFF94 force field42. The coordinates for the head-to-tail cyclized pre- fuscimiditide A1C / T3C variant have been deposited in Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB; ID No.8VYC) and Biological Magnetic Resonance Data Bank (BMRB Entry No.31145). In vitro Aspartimidylation of the Head-to-tail Cyclized Pre-fuscimiditide A1C / T3C Variant with Human PIMT (HsPIMT)

[0187] This procedure was adopted from a previous study on lihuanodin kinetics (Cao etal.2023)38and slightly modified in this study.20 µM HPLC-purified pre-fuscimiditide A1C T3C variant was mixed with 1.0 µM human protein l-isoaspartyl methyltransferase (HsPIMT), and 0.4 mM S-adenosyl methionine (SAM) in 100 mM phosphate buffer (pH 7.0). The reaction mixture was incubated at room temperature for 20 hours and then quenched by addition of 1% formic acid. Results Assessing graspetide synthetase substrate tolerance with amino acid substitutions at the core peptide

[0188] The pre-fuscimiditide peptide is a stem-loop structure with two ω-ester linkagesbetween Thr7-Asp18 and Thr3-Asp22, forming the bis-macrocyclic structure comprising the stem and loop macrocycles (FIG.1A). To assess the tolerance of pre-fuscimiditide to amino acid substitutions, ThfA variants with substitutions targeting the ester-forming amino acids (positions 3, 7, 18, and 22), stem (positions 4-6 and 19-21), and loop (positions 8-17) regions were generated. T3S, T7S, D18E, and D22E variants (SEQ ID NOs: 2-5, respectively) were tested for ThfB-mediated esterification to examine whether the enzyme can use alternative donor and acceptor residues to install macrocyclic linkages. The T3S and T7S variants exhibited complete two-fold dehydration (indicating two-fold esterification), whereas D18E and D22E variants only showed mostly one-fold dehydration (FIG.1B). The latter group of variants was further hydrazinolyzed and subjected to tandem mass spectrometry, revealing that an ester cross-link was missing at the substituted positions (FIGs.2, 3; Tables 7, 8). These results suggest that ThfB can tolerate Ser residues as donor residues for ω-ester linkages but not Glu residues as acceptor residues. - 42 - 4144725.v15391.1039001 Table 7: Ions detected from the MS / MS spectrum of the ester-hydrazinolyzed core peptide of mThfABD18E variant (see FIGs.2A-B).Table 8: Ions detected from the MS / MS spectrum of the ester-hydrazinolyzed core peptide of mThfABD22E variant (see FIG.3).

[0189] To test amino acid substitutions in the stem macrocycle, M4G / R5S / V6G (SEQ IDNO: 6) and V19G / E20S / K21G (SEQ ID NO: 7) variants were generated, each targeting the N- or C-terminal side of the pre-fuscimiditide core peptide. Upon coexpression with ThfB, both variants exhibited maximally two-fold dehydration with appreciable amounts of singly dehydrated and unreacted species present (FIG.1C). Arg5 and Glu20 are potentially salt- bridge-forming pairs, and the stem-substitution experiments support that these residues could be important for efficient substrate recognition of ThfB. Yet, both R5A and E20A variants exhibited complete two-fold dehydration after ThfB modification (FIG.4). After observing that the stem modification was somewhat tolerated for forming the bis-macrocyclic structure, we tried changing the size of the stem macrocycle. Each variant tested contains one or two of the esterification site residues (T3, T7, D18, and D22) swapped with an adjacent residue (see - 43 - 4144725.v15391.1039001 FIG.5). Out of the four variants tested, only the G2T / T3G variant was doubly esterified by ThfB, demonstrating that most examples of this type of modification to the stem are not well- tolerated.

[0190] Next, the loop region was substituted with varying lengths of GS-repeats, rangingfrom 0 (corresponding to loop deletion) to 72 amino acid (aa) residues (see SEQ ID NOs: 8- 17). As a point of reference, the loop size of native pre-fuscimiditide is 10 aa. Variants with extremely large or small substitutions (72-, 3-, 2-, and 0-residue substitutions) were mostly singly dehydrated by ThfB, while those with intermediate-sized substitutions (from 4- to 24- residue substitutions) were fully modified (two dehydrations) by the enzyme (FIG.1D). For variants with short loops, the residues corresponding to Thr7 and Asp18 in the wild-type core peptide may be too distance-constrained for esterification. Further MS / MS studies with the GSG-loop variant confirms that the singly dehydrated species only contains the outermost ω- ester linkage (FIGs.1E, 6A-B). On the other hand, harboring a large loop appears to interfere with esterification, as evidenced by the presence of singly esterified species with (GSSG)6- and (GSSG)18-loop variants. MS / MS analysis revealed the major product of the (GSSG)6- loop variant to contain a nonnative Ser-Asp ω-ester linkage (FIGs.1F, 7A-B). Nevertheless, the peptide was doubly esterified, suggesting that peptides with various sizes of stem and loop macrocycles can be designed by placing a Ser / Thr residue at the desired position. Table 9: Ions detected from the MS / MS spectrum of the singly esterified core peptide of the mThfABvariant with GSG loop (see FIG.6).Table 10: Ions detected from the MS / MS spectrum of the (GSSG)6loop mThfABvariant core peptide with net one-fold dehydration, peak 1 (see FIG.7B).- 44 - 4144725.v15391.1039001Table 11: Ions detected from the MS / MS spectrum of the (GSSG)6loop mThfABvariant core peptide with net one-fold dehydration, peak 2 (see FIG.7C).- 45 - 4144725.v15391.1039001 Table 12: Ions detected from the MS / MS spectrum of the (GSSG)6loop mThfABvariant core peptide with net two-fold dehydration (see FIG.7D).

[0191] Remarkably, the ability of ThfB to form the outermost ester crosslink withoutforming the innermost one beforehand is unique amongst graspetides. In general, cross- linking proceeds N-to-C-terminally following the acceptor (carboxylic acid) residues10, 13, 21,24, and mutational studies show that substitutions at the innermost ω-ester linkage abolishes all subsequent ester formation21, 24. A similar mutational study with ThfB was conducted, testing ThfA T7V / D18N (SEQ ID NO: 39) and T3V / D22N (SEQ ID NO: 40) variants. Both precursors were completely singly dehydrated, exhibiting esterification without following any strict order (FIG.8). Hence, ThfB may be the first graspetide synthetase to defy the common biosynthetic rule for graspetides, and this exception helps with repurposing the enzyme as general biocatalyst for designing multi-macrocyclic peptides.

[0192] After observing that deleting the entire loop of the core peptide still resulted inesterification, additional variants with further deletions targeting Thr7, Asp18, or both were tested (SEQ ID NOs: 41-43). Upon coexpression with ThfB, all variants remained unmodified (FIG.9), suggesting that the formation of the outermost ester cross-link requires at least eight residues in-between the sequence. - 46 - 4144725.v15391.1039001 Protein cyclization with an ω-ester linkage

[0193] The ability of ThfB to esterify a core peptide with a very large loop (e.g. 72 aa ofGlySer-repeats) suggests that the enzyme might be able to process a core peptide with a whole protein in the loop. As a proof-of-concept, variants with fluorescent proteins (abbreviated FP; mRuby2, ffGFP, or mTurquoise2) in the loop were tested. To minimize a possible steric hindrance of the FP insert during esterification, N- and C-terminal flexible (GS)x3 linkers flanking the FP sequence were added (FIG.10A). Mass spectrometry revealed a -22 Da shift from all FP-inserted ThfA constructs. While this mass shift is consistent with the formation of the chromophore in mRuby2, the chromophore in ffGFP and mTurquoise2 results in a -20 Da shift25-26. To determine whether a disulfide bond may be present in both ffGFP-inserted and mTurquoise2-inserted ThfA, the ThfB-modified versions of both proteins were incubated with dithiothreitol (DTT) and subjected them to mass spectrometry. Both resulted in a +2 Da shift before DTT treatment, resulting in a net -20 Da shift for the mass peaks corresponding to proteins without any ester (FIG.11). The ThfB-modified FP-inserted ThfA proteins exhibited additional mass peaks corresponding to one- or two-fold esterification (-18 or -36 Da shift) (FIG.10B). Up to 40-50% of ThfA variants with the mRuby2 or ffGFP loop were singly esterified, whereas significantly more of the ThfA variant with the mTurquoise2 loop were singly and doubly esterified (FIG.10B). No difference in fluorescence was detected between mRuby2-inserted ThfA, mRuby2-inserted mThfAB, and mRuby2 (FIG.12).

[0194] To map the ester linkages present in the ThfB-modified FP-inserted ThfAproteins, the tryptic fragments containing the esters were analyzed by mass spectrometry. After trypsin digestion, three tryptic fragments present in all three esterified FP-inserted ThfA proteins were identified (FIGs.10C, 10D). The identified tryptic fragment with the smallest mass, 1018.42 Da, matched the C-terminal FP-inserted ThfA sequence GSGSGSDVEKD (SEQ ID NO: 46) with one dehydration (1). The other tryptic fragments with higher masses, 1534.67 Da and 1552.68 Da (2 and 3, respectively), match the same C-terminal fragment (P2; SEQ ID NO: 46) connected to the N-terminal stem fragment, AGTMR (P1; SEQ ID NO: 27), via an ω-ester linkage; 2 harbors additional dehydration. All three ThfB-modified FP-inserted ThfA proteins predominantly yield 3 (FIGs.10C, 10D). The mass spectra suggest the highest proportion of 1 to be found in ffGFP-inserted mThfABand the highest proportion of 2 to be found in mTurquoise2-inserted mThfAB. Therefore, for mapping ω-ester linkages, MS / MS experiments for 1 and 2 were performed on the trypsinized ffGFP-inserted mThfABand - 47 - 4144725.v15391.1039001 mTurquoise2-inserted mThfAB, respectively. MS / MS for 3 was performed on the trypsinized mRuby2-inserted mThfAB.

[0195] In collision-induced dissociation (CID), 1 yielded y1 and y2 fragments, suggestingthat the last Asp residue is not esterified (FIG.13; Table 13). This implies that the first Asp residue is esterified instead. Additionally, y7and y9ions were detected with -18 Da shift, suggesting the last Ser residue in the GSGSGS (SEQ ID NO: 56) linker to be esterified. A b2ion was detected, consistent with the first Ser residue to be not esterified. While the b2and b3ions downshifted by 18 Da were also detected in lower intensity, these were probably the results of McLafferty rearrangement of the free Ser residue. The assignment of Ser6-Asp7 (i.e. the esterification of the third Ser residue and the first Asp residue) ω-ester linkage (FIG. 10E) is based on the highest peaks observed in the MS / MS spectrum. Table 13: Ions detected from the MS / MS spectrum of 1 (see FIG.13).

[0196] The MS / MS experiment on 2 suggested a different ω-ester linkage to be present inthe GSGSGSDVEKD (SEQ ID NO: 46) peptide chain. Instead of the Ser6-Asp7 linkage, Ser2-Asp7 linkage is implicated by the series of y ions detected in the MS / MS spectrum (FIGs.10F, 14; Table 14). Only the y(2)10:P1 ion, the y ion with 10 residues in the P2 chain (SGSGSDVEKD (SEQ ID NO: 57) from GSGSGSDVEKD (SEQ ID NO: 46)) connected to the P1 chain (AGTMR (SEQ ID NO: 27)), was found to be shifted by -18 Da, whereas no additional mass shift was observed in other y(2)n:P1 ions. Additionally, in comparison to the MS / MS of 3, y(2)6:P1, y(2)7:P1, and y(2)8:P1 ions are absent and the intensities of y(2)5:P1 and y(2)9:P1 ions are significantly lower (FIGs.14, 15; Table 15). These observations provide more evidence for the Ser2-Asp7 linkage in the P2 chain, besides the linkage between Thr3 from P1 and Asp11 from P2 suggested by the y(2)n:P1 ions detected in the MS / MS spectra of both 2 and 3 (FIGs.10F, 10G). - 48 - 4144725.v15391.1039001 Table 14: Ions detected from the MS / MS spectrum of 2 (see FIG.14).Table 15: Ions detected from the MS / MS spectrum of 3 (see FIG.15).- 49 - 4144725.v15391.1039001 Extending the fuscimiditide stem

[0197] As demonstrated previously, fuscimiditide offers a framework for designing amonocyclic or bicyclic peptide with varying sequence and size. To expand the structure of the peptide, a third ω-ester linkage was added, forming another macrocycle. In total, nine constructs were tested: 4-12 (FIGs.16A-B).4 (SEQ ID NO: 47) carries R5T / E20D substitutions, attempting to place the third ω-ester linkage at the center of the stem macrocycle. After coexpression with ThfB, 4 was only dehydrated twice (FIG.16A), unable to form three macrocycles. As this result suggested that an ester linkage cannot be formed inside the stem macrocycle, the new linkage was placed outside the stem macrocycle for the remaining constructs.5 (SEQ ID NO: 48) contains a G2T substitution and D23 insertion (denoted as G2T / +D23), designed to introduce the new crosslink immediately next to the T3- D22 ω-ester linkage. However, this construct was also only doubly dehydrated (FIG.16A), suggesting that ThfB cannot install ester linkages next to each other.6 (A1T / +S23 / +D24; SEQ ID NOs: 49 and 60) was designed to place the new linkage two residues away from the stem macrocycle. Adding the C-terminal interstitial residue Ser23 was based on the sequence of the fuscimiditide homolog from Thermobifida alba (SEQ ID NO: 59), which contains two extra C-terminal residues S23 and D24 and threonine in place of Ala1, compared to the sequence of fuscimiditide (SEQ ID NO: 58) (FIG.17).

[0198] Coexpressing 6 with ThfB yielded a major product with three-fold dehydrationcorresponding to the formation of three macrocycles (FIG.16A). To increase the size of the putative third macrocycle to be the same as the stem macrocycle, 7 (E- 2T / +V23 / +E24 / +K25 / +D26; SEQ ID NO: 50) was designed with the C-terminal interstitial residues being identical to that of the original stem macrocycle. However, only a very small amount of three-fold dehydrated ThfB-modified product was observed (FIG.16A). After suspecting that some residues in the N-terminal sequence (Arg-Ala-Gly) may interfere with the formation of the third ω-ester linkage, 8 (E-2T / R-1M / A1R / G2V / +V23 / +E24 / +K25 / +D26; SEQ ID NO: 51) was designed to harbor the same N-terminal interstitial residues Met-Arg- Val in the putative third macrocycle as in the stem macrocycle. ThfB predominantly modified 8 three-fold, consistent with the hypothesis that the interstitial residues affect the formation of the third ω-ester linkage.

[0199] After designing a construct that can be predominantly transformed into a tris-macrocyclic peptide with the macrocycles placed in series, the customizability of the amino acid sequence was examined. In designing the new constructs, the intermediate macrocycle - 50 - 4144725.v15391.1039001 (i.e. the original stem macrocycle) was targeted for amino acid substitution, and tested for whether ThfB could still install three ester linkages.9 (SEQ ID NO: 52) harbored Gly-Ser- Gly substituting the N-terminal interstitial residues of the intermediate macrocycle, in comparison to 8. Compared to 8, a smaller amount of 3-fold dehydrated species was observed in 9 after coexpression with ThfB (FIG.16B).10 (SEQ ID NO: 53) was designed with an additional Gly-Ser-Gly substituting the C-terminal interstitial residues of the intermediate macrocycle. ThfB could not form any appreciable amount of 3-fold dehydrated product of 10 (FIG.16B). Suspecting at least one C-terminal interstitial residue to be important for esterification, an tempt was made to restore the 3-fold esterification by reverting Ser20 to Glu20 (11; SEQ ID NO: 54). Partial restoration of the 3-fold esterification was observed (FIG.16B). A similar substitution S20Q was tested with 12 (SEQ ID NO: 55), but 3-fold esterification was barely observed, similar to 10 (FIG.16B). The results overall show that the intermediate macrocycle may not be amenable to complete substitutions targeting all six residues but can moderately tolerate partial substitutions.

[0200] All nine constructs 4-12 were designed to maintain or extend the stem-loophairpin structure of fuscimiditide. The connectivity of one of the constructs, 9, was verified by mass spectrometry. Digesting the ThfB-modified 9 whole protein with trypsin yielded one major species of singly esterified core peptide, two prevalent species of doubly esterified core peptide, and one major product of triply esterified core peptide (FIG.18). Each peptide was analyzed by tandem mass spectrometry to map the ω-ester linkage(s). The fragmentation patterns suggest the T7-D18 linkage to be present in the singly esterified species (FIG.19; Table 16). For the doubly esterified species, the major product additionally harbors the T3- D22 linkage (FIG.20; Table 17), whereas the minor product contains the S5-D22 nonnative linkage instead (FIG.21; Table 18). The fully modified, triply esterified peptide harbors the T-2-D26 linkage in addition to T3-D22 and T7-D18 linkages (FIG.22; Table 19). The results confirm that the three macrocycles are placed in series, predominantly forming the intended tris-macrocyclic stem-loop hairpin structure. Table 16: Ions detected from the MS / MS spectrum of the singly esterified core peptide of 9 (see FIG.19).- 51 - 4144725.v15391.1039001Table 17: Ions detected from the MS / MS spectrum of the major product of the doubly esterified core peptide of 9 (see FIG.20).Table 18: Ions detected from the MS / MS spectrum of the minor product of the doubly esterified core peptide of 9 (see FIG.21).- 52 - 4144725.v15391.1039001 Table 19: Ions detected from the MS / MS spectrum of the triply esterified core peptide of 9 (see FIG.22).Multivalent Core Peptides

[0201] While a subset of graspetides, such as plesiocin14, thuringinin16, andchryseoviridin22, are comprised of multiple macrocycles repeated in tandem, this type of multivalency is not observed in fuscimiditide and other graspetides containing an O- methyltransferase gene in the biosynthetic gene cluster (i.e. graspimiditides)13. The ability of the fuscimiditide ATP-grasp enzyme to construct core peptide units in tandem repeats was tested. First, divalent constructs were designed, in which the ThfA leader peptide was followed by the first core peptide unit, a linker sequence, and the second core peptide unit (FIG.23A). The GDPSAG linker sequence (SEQ ID NO: 63), derived from the linker sequences in plesiocin14, was tested. ThfB esterification of this construct yielded 3-fold and 4-fold dehydrated species, predominantly the former. Suspecting that the flexibility and the length of the linker sequence significantly affect the esterification efficiency, a panel of varying lengths of Ser-Gly-repeating linkers from 0 to 12 residues was then tested. The highest ratio of 4-fold to 3-fold esterification was observed from the divalent construct with the (SG)3linker. A significantly higher amount of 4-fold esterified species was observed in the construct with the (SG)3linker than that with the GDPSAG linker (SEQ ID NO: 63), suggesting that flexibility assists ThfB in making multivalent core peptides. The constructs with a longer or shorter linker than (SG)3exhibited a higher proportion of 3-fold esterification. Those with 1 aa (Gly) or no linker did not exhibit any appreciable amount of 4- fold esterification.

[0202] Following the success in synthesizing divalent, bis-macrocyclic peptide constructswith ThfB, construction of trivalent versions of the peptide was performed. For the first - 53 - 4144725.v15391.1039001 design, the wildtype ThfA was connected with two additional fuscimiditide core peptide units connected with Ser-Gly linkers (FIG.23B). After coexpression with ThfB, up to 5-fold dehydration was observed (FIG.23B), one dehydration less than desired for a trivalent bis- macrocyclic peptide. Likely, there were Thr- or Ser-to-Asp mispairings forming a different structure and preventing further esterification. Then, the design for a trivalent peptide was simplified by making each core repeat monocyclic (FIG.23C), substituting the outermost cross-linking Thr and Asp residues to Val and Asn, respectively. The internal trypsin cut sites were substituted with Ala residues. A Gly-Pro-Gly linker was placed in between the core repeats, expecting the proline effect27to assist with mapping the ester linkages by tandem mass spectrometry. ThfB modification resulted in a significant proportion of 3-fold dehydration, as desired for a trivalent, monocyclic peptide. Trypsin digestion yielded a three- fold dehydrated C-terminal peptide with cleavage at Arg5 (FIG.24), and MS / MS analysis of that peptide suggested all three Thr-Asp linkages at the expected locations (FIGs.25A-B). Hence, a simpler design with fewer possibilities for unintended linkages enabled multivalent peptides to be synthesized with ThfB more efficiently. Table 20: Ions detected from the MS / MS spectrum of the trivalent, monocyclic mThfABvariant core peptide (see FIGs.25A-B).- 54 - 4144725.v15391.1039001Alternative side chain-to-side chain linkages

[0203] To test whether the ATP-grasp enzyme ThfB catalyzes the formation of linkagesother than esters, variants of the ThfA precursor protein were generated with the following substitutions to the core peptide: T3C, T3K, T7C, and T7K. These variants were coexpressed with ThfB in E. coli and enabled determination of whether thioester or amide linkages could - 55 - 4144725.v15391.1039001 be installed in the graspetide pre-fuscimiditide (FIG.26A). All four ThfA variants were crosslinked by ThfB, although to different extents (FIG.26B). For the putative amide-linked variants, incomplete dehydration of the precursors was observed. Using tandem mass spectrometry and hydrazinolysis, the T3K variant was shown to generate some of the anticipated amide linkage (K3-D22) (FIGs.27, 28; Tables 21, 22), but most of the protein remains uncrosslinked (FIG.26B). The T7K variant formed a small amount of the anticipated linkage (K7-D18) (FIG.29; Table 23), but also formed an unexpected linkage (K13-D18) (FIG.30; Table 24), showing that ThfA T7K is not a good substrate for ThfB. In contrast to the Lys-substituted ThfA precursors, the Cys-substituted precursors were better substrates for the ThfB. About 50% of the T7C variant of ThfA was doubly dehydrated whereas essentially all the T3C variant was doubly dehydrated (FIG.26B). Table 21. Ions detected from the MS / MS spectrum of the doubly cross-linked core peptide from the ThfB-modified ThfA T3K variant after ester-selective hydrazinolysis. (see FIG.27)Table 22. Ions detected from the MS / MS spectrum of the singly cross-linked intermediate core peptide obtained from the ThfB-modified ThfA T3K variant. (see FIG.28)- 56 - 4144725.v15391.1039001Table 23. Ions detected from the MS / MS spectrum of the first core peptide obtained from the ThfB-modified ThfA T7K variant. (see FIG.29)Table 24. Ions detected from the MS / MS spectrum of the second core peptide obtained from the ThfB-modified ThfA T7K variant. (see FIG.30)

[0204] To confirm thioester formation, both proteins were treated with iodoacetamideand digested with trypsin. No alkylation was observed on the doubly dehydrated core peptides for both variants while an internal control (a fragment of the leader peptide with a single Cys) was successfully alkylated (FIGs.31, 32). These results show that alternative nucleophiles can be utilized by the ATP-grasp enzyme ThfB resulting in amide or thioester linkages. However, the extent of linkage formation depends both on nucleophilic sidechain and the position within the peptide. Without being bound by theory, the increased sidechain length of Lys may prevent it from being recognized effectively in the ATP-grasp active site. In contrast, the Cys sidechain is sterically similar to the native substrate Thr, so it is more effectively crosslinked than Lys. The disclosed findings are consistent with a recent report - 57 - 4144725.v15391.1039001 showing that a microviridin ATP-grasp enzyme could successfully mature chemically synthesized microviridin precursors harboring isosteric diaminopropionic acid or cysteine in place of serine.28Given the complete conversion of the T3C variant of ThfA, the T3C variant was used as the basis for further modification. ω-thioester enabled peptide ligation at the side chain

[0205] Thioesters allow for native chemical ligation (NCL) in which the electrophilicthioester is attacked by thiols followed by S-N acyl shifts.29-30NCL has been used to ligate multiple linear peptides generated by solid-phase peptide synthesis in protein total synthesis31-32and to ligate synthetic peptides and recombinant proteins in expressed protein ligation.33The thioester bearing graspetide disclosed herein was considered as a building block for branched and cyclic peptides accessible by NCL. As a first attempt to see if the thioester in the T3C variant of pre-fuscimiditide (13) could be elaborated using NCL, this peptide was treated with free cysteine (FIG.33A) and the ligation of cysteine to the graspetide was observed at modest conversion (FIG.33B, FIG.34), demonstrating the feasibility of this approach. The conversion of 13 to the cysteine adduct (14) could be substantially improved by employing MPAA to capture and activate the thioester (FIG.33A). Under the conditions described above but with an excess of MPAA, quantitative conversion to 14 was observed (FIG.33B, FIG.34). 14 was subjected to tandem mass spectrometry, which confirmed that Cys addition occurred at the C-terminus of the peptide (FIG.33C; Table 25). Without being bound by theory, it is expected that Cys was ligated to the Asp22 sidechain, forming an isopeptide bond (FIG.33A); however, the MS2 experiment does not differentiate between isopeptide-bonded Cys and conventional peptide-bonded Cys. Table 25. Ions detected from the MS / MS spectrum of the pre-fuscimiditide T3C variant ligated with L-cysteine (14). (see FIG.33C)- 58 - 4144725.v15391.1039001Intramolecular macrocyclic rearrangement enabled by ω-thioester

[0206] Encouraged by the success of the intermolecular NCL reaction between 1 andCys, a variant of pre-fuscimiditide that could react intramolecularly and generate a cyclic peptide was engineered. A variant of ThfA that harbors two Cys residues at the positions corresponding to aas 1 and 3 of pre-fuscimiditide, ThfA A1C T3C, was produced. After coexpression with ThfB (FIG.35), purification, and trypsin digest, the ThfA A1C T3C variant was expected to generate 15, a version of fuscimiditide with a thioester linkage and an N-terminal cysteine (FIG.36A). This species can rearrange via thioester exchange and N-S acyl shift to 16, a head-to-tail cyclized peptide with an amide bond between the N-terminus and the side chain of Asp22 (D22) (FIG.36A). To determine the ratio of 15 and 16, the ThfB-modified ThfA A1C T3C was subjected to tryptic digestion for 20 minutes. The peptides were purified via HPLC. The resulting peptides were treated with iodoacetamide since 15 had only one thiol but 16 had two thiols. Only a species with two alkylations was observed, suggesting full conversion to 16 (FIG.37). This data suggests that within the short 20 minute tryptic digest period, cleavage of the peptide, thioester exchange, and the N-S acyl shift were all complete. To confirm head-to-tail amide bond formation, hydrazine-treated 16 exhibited a mass shift of +32, corresponding to a single hydrazide, showing that 16 had only a single ester (FIG.38). Tandem mass spectrometry of the hydrazide of 16 showed that the ester in 16 was between the sidechains of Thr7 and Asp18, as expected (FIGs.39A-B; Table 26). The tandem mass spec provided clear evidence for connectivity between Cys1 and Asp22. Table 26. Ions detected from the MS / MS spectrum of the head-to-tail cyclized pre-- 59 - 4144725.v15391.1039001

[0207] 16 was also treated with human protein isoaspartyl methyltransferase (PIMT),34which methylates isoAsp residues, ultimately converting them into an aspartimide (FIG. 40A). Both methylated and aspartimidylated 16 were observed by mass spectrometry, providing strong evidence for isopeptide bond formation (FIG.40B).

[0208] The NMR structure of 16 was solved to provide direct evidence for isopeptidebond formation between the N-terminus of Cys1 and the Asp22 sidechain. An NMR sample (7.7 mM, 98:2 H2O:D2O) was generated and high quality TOCSY and NOESY spectra were acquired, allowing for full assignment of the peptide (FIGs. S34-43, Table 27). Table 27. Chemical shift assignments for the head-to-tail cyclized pre-fuscimiditide A1C / T3C variant (16).- 60 - 4144725.v15391.1039001- 61 - 4144725.v15391.1039001- 62 - 4144725.v15391.1039001

[0209] Clear evidence of isopeptide bond formation in the NOESY experiment wasobserved, with a crosspeak between the amide proton of Cys1 and the b protons of Asp22 (FIG.36B). The b proton of Thr7 was shifted downfield, a reliable marker for ester formation.9, 20-21, 35In comparing the proton chemical shifts of 16 to those of pre-fuscimiditide (PDB 7LI2), there is excellent agreement for residues in the loop region (residues 8-17) of both peptides (FIG.44). The chemical shift differences were larger in the residues that comprise the stem of fuscimiditide. Whereas pre-fuscimiditide is a bis-macrocyclic structure with 32 atoms in the loop and 34 atoms in the stem macrocycle, the rearrangement that generated 16 converted the stem into a longer, 39 atom macrocycle (FIG.45). Overall, the structure of 16 resembled the structure of pre-fuscimiditide; both structures exhibited b-strand character throughout the stem. Like pre-fuscimiditide, 16 did not exhibit extensive NOESY interactions through its macrocycles (FIGs. S35, 43), suggesting that these regions of the peptide are flexible (FIG.46).

[0210] This disclosure has demonstrated that graspetides can harbor thioesters in additionto the esters and amides that are natively found in the crosslinks of this family of RiPPs. The wild-type ATP-grasp enzyme was able to forge the thioester. Access to a graspetide with a sidechain-sidechain thioester linkage enabled the generation of both branched, cyclic, and multicyclic peptides via NCL. Of particular interest and possible utility is the disclosed system for generating cyclic peptides. Amide bonded cyclic peptides can be generated in two steps; recombinant protein expression followed by protease cleavage to reveal an N-terminal cysteine. Over the course of the protease cleavage, the peptide rearranges quantitatively. This two-step process provides an efficient method for accessing diverse cyclic peptides. - 63 - 4144725.v15391.1039001 Sequences- 64 - 4144725.v15391.1039001- 65 - 4144725.v15391.10390012, Reference 2 - 66 - 4144725.v15391.1039001 Table 29: Plasmids used for protein expression in this study.- 67 - 4144725.v15391.1039001- 68 - 4144725.v15391.1039001- 69 - 4144725.v15391.1039001- 70 - 4144725.v15391.1039001- 71 - 4144725.v15391.1039001- 72 - 4144725.v15391.1039001- 73 - 4144725.v15391.1039001Table 30. Vectors used in this study. Table 31. Ion-ladders from MS / MS analysis.Table 32. Fragment ions from MS / MS analysis.- 74 - 4144725.v15391.1039001Table 33. Tryptic fragments from MS / MS analysis.References

[0211] 1. Vinogradov, A. A.; Yin, Y.; Suga, H., Macrocyclic Peptides as DrugCandidates: Recent Progress and Remaining Challenges. J Am Chem Soc 2019, 141 (10), 4167-4181.

[0212] 2. Ji, X.; Nielsen, A. L.; Heinis, C., Cyclic Peptides for Drug Development.Angew Chem Int Ed Engl 2023, e202308251.

[0213] 3. Keyes, E. D.; Mifflin, M. C.; Austin, M. J.; Alvey, B. J.; Lovely, L. H.; Smith,A.; Rose, T. E.; Buck-Koehntop, B. A.; Motwani, J.; Roberts, A. G., Chemoselective, Oxidation-Induced Macrocyclization of Tyrosine-Containing Peptides. J Am Chem Soc 2023, 145 (18), 10071-10081.

[0214] 4. Malins, L. R.; deGruyter, J. N.; Robbins, K. J.; Scola, P. M.; Eastgate, M. D.;Ghadiri, M. R.; Baran, P. S., Peptide Macrocyclization Inspired by Non-Ribosomal Imine Natural Products. J Am Chem Soc 2017, 139 (14), 5233-5241.

[0215] 5. Schafer, M.; Schneider, T. R.; Sheldrick, G. M., Crystal structure ofvancomycin. Structure 1996, 4 (12), 1509-15.

[0216] 6. Imai, Y.; Meyer, K. J.; Iinishi, A.; Favre-Godal, Q.; Green, R.; Manuse, S.;Caboni, M.; Mori, M.; Niles, S.; Ghiglieri, M.; Honrao, C.; Ma, X.; Guo, J. J.; Makriyannis, - 75 - 4144725.v15391.1039001 A.; Linares-Otoya, L.; Bohringer, N.; Wuisan, Z. G.; Kaur, H.; Wu, R.; Mateus, A.; Typas, A.; Savitski, M. M.; Espinoza, J. L.; O'Rourke, A.; Nelson, K. E.; Hiller, S.; Noinaj, N.; Schaberle, T. F.; D'Onofrio, A.; Lewis, K., A new antibiotic selectively kills Gram-negative pathogens. Nature 2019, 576 (7787), 459-464.

[0217] 7. Dittmann, J.; Wenger, R. M.; Kleinkauf, H.; Lawen, A., Mechanism ofcyclosporin A biosynthesis. Evidence for synthesis via a single linear undecapeptide precursor. Journal of Biological Chemistry 1994, 269 (4), 2841-2846.

[0218] 8. Montalban-Lopez, M.; Scott, T. A.; Ramesh, S.; Rahman, I. R.; van Heel, A.J.; Viel, J. H.; Bandarian, V.; Dittmann, E.; Genilloud, O.; Goto, Y.; Grande Burgos, M. J.; Hill, C.; Kim, S.; Koehnke, J.; Latham, J. A.; Link, A. J.; Martinez, B.; Nair, S. K.; Nicolet, Y.; Rebuffat, S.; Sahl, H. G.; Sareen, D.; Schmidt, E. W.; Schmitt, L.; Severinov, K.; Sussmuth, R. D.; Truman, A. W.; Wang, H.; Weng, J. K.; van Wezel, G. P.; Zhang, Q.; Zhong, J.; Piel, J.; Mitchell, D. A.; Kuipers, O. P.; van der Donk, W. A., New developments in RiPP discovery, enzymology and engineering. Nat Prod Rep 2021, 38 (1), 130-239.

[0219] 9. Choi, B.; Link, A. J., Discovery, Function, and Engineering of Graspetides.Trends Chem 2023, 5 (8), 620-633.

[0220] 10. Lee, H.; Choi, M.; Park, J. U.; Roh, H.; Kim, S., Genome Mining RevealsHigh Topological Diversity of omega-Ester-Containing Peptides and Divergent Evolution of ATP-Grasp Macrocyclases. J Am Chem Soc 2020, 142 (6), 3013-3023.

[0221] 11. Ramesh, S.; Guo, X.; DiCaprio, A. J.; De Lio, A. M.; Harris, L. A.; Kille, B.L.; Pogorelov, T. V.; Mitchell, D. A., Bioinformatics-Guided Expansion and Discovery of Graspetides. ACS Chem Biol 2021, 16 (12), 2787-2797.

[0222] 12. Makarova, K. S.; Blackburne, B.; Wolf, Y. I.; Nikolskaya, A.; Karamycheva,S.; Espinoza, M.; Barry, C. E., 3rd; Bewley, C. A.; Koonin, E. V., Phylogenomic analysis of the diversity of graspetides and proteins involved in their biosynthesis. Biol Direct 2022, 17 (1), 7.

[0223] 13. Choi, B.; Acuña, A.; Koos, J. D.; Link, A. J., Large-scale Bioinformatic Studyof Graspimiditides and Structural Characterization of Albusimiditide. ACS Chem Biol 2023.

[0224] 14. Lee, H.; Park, Y.; Kim, S., Enzymatic Cross-Linking of Side Chains Generatesa Modified Peptide with Four Hairpin-like Bicyclic Repeats. Biochemistry 2017, 56 (37), 4927-4930. - 76 - 4144725.v15391.1039001

[0225] 15. Zhang, Y.; Li, K.; Yang, G.; McBride, J. L.; Bruner, S. D.; Ding, Y., Adistributive peptide cyclase processes multiple microviridin core peptides within a single polypeptide substrate. Nat Commun 2018, 9 (1), 1780.

[0226] 16. Roh, H.; Han, Y.; Lee, H.; Kim, S., A Topologically Distinct ModifiedPeptide with Multiple Bicyclic Core Motifs Expands the Diversity of Microviridin-Like Peptides. Chembiochem 2019, 20 (8), 1051-1059.

[0227] 17. Lee, C.; Lee, H.; Park, J. U.; Kim, S., Introduction of Bifunctionality into theMultidomain Architecture of the omega-Ester-Containing Peptide Plesiocin. Biochemistry 2020, 59 (3), 285-289.

[0228] 18. Unno, K.; Kaweewan, I.; Nakagawa, H.; Kodani, S., Heterologous expressionof a cryptic gene cluster from Grimontia marina affords a novel tricyclic peptide grimoviridin. Appl Microbiol Biotechnol 2020, 104 (12), 5293-5302.

[0229] 19. Unno, K.; Kodani, S., Heterologous expression of cryptic biosynthetic genecluster from Streptomyces prunicolor yields novel bicyclic peptide prunipeptin. Microbiol Res 2021, 244, 126669.

[0230] 20. Elashal, H. E.; Koos, J. D.; Cheung-Lee, W. L.; Choi, B.; Cao, L.; Richardson,M. A.; White, H. L.; Link, A. J., Biosynthesis and characterization of fuscimiditide, an aspartimidylated graspetide. Nat Chem 2022, 14 (11), 1325-1334.

[0231] 21. Choi, B.; Elashal, H. E.; Cao, L.; Link, A. J., Mechanistic Analysis of theBiosynthesis of the Aspartimidylated Graspetide Amycolimiditide. J Am Chem Soc 2022, 144 (47), 21628-21639.

[0232] 22. Zhao, G.; Kosek, D.; Liu, H. B.; Ohlemacher, S. I.; Blackburne, B.;Nikolskaya, A.; Makarova, K. S.; Sun, J.; Barry Iii, C. E.; Koonin, E. V.; Dyda, F.; Bewley, C. A., Structural Basis for a Dual Function ATP Grasp Ligase That Installs Single and Bicyclic omega-Ester Macrocycles in a New Multicore RiPP Natural Product. J Am Chem Soc 2021, 143 (21), 8056-8068.

[0233] 23. Wang, T.; Wang, X.; Zhao, H.; Huo, L.; Fu, C., Uncovering a Subtype ofMicroviridins via the Biosynthesis Study of FR901451. ACS Chem Biol 2022.

[0234] 24. Philmus, B.; Guerrette, J. P.; Hemscheidt, T. K., Substrate specificity andscope of MvdD, a GRASP-like ligase from the microviridin biosynthetic gene cluster. ACS Chem Biol 2009, 4 (6), 429-34. - 77 - 4144725.v15391.1039001

[0235] 25. Miyawaki, A.; Shcherbakova, D. M.; Verkhusha, V. V., Red fluorescentproteins: chromophore formation and cellular applications. Curr Opin Struct Biol 2012, 22 (5), 679-88.

[0236] 26. Cannon, J. R.; Kluwe, C.; Ellington, A.; Brodbelt, J. S., Characterization ofgreen fluorescent proteins by 193 nm ultraviolet photodissociation mass spectrometry. Proteomics 2014, 14 (10), 1165-73.

[0237] 27. Bleiholder, C.; Suhai, S.; Harrison, A. G.; Paizs, B., Towards understandingthe tandem mass spectra of protonated oligopeptides.2: The proline effect in collision- induced dissociation of protonated Ala-Ala-Xxx-Pro-Ala (Xxx = Ala, Ser, Leu, Val, Phe, and Trp). J Am Soc Mass Spectrom 2011, 22 (6), 1032-9.

[0238] 28. Patel, K. P.; Chen, W. T.; Delbecq, L.; Bruner, S. D., Alternative LinkageChemistries in the Chemoenzymatic Synthesis of Microviridin-Based Cyclic Peptides. Org Lett 2024.

[0239] 29. Dawson, P. E.; Muir, T. W.; Clark-Lewis, I.; Kent, S. B., Synthesis of proteinsby native chemical ligation. Science 1994, 266 (5186), 776-9.

[0240] 30. Thompson, R. E.; Muir, T. W., Chemoenzymatic Semisynthesis of Proteins.Chem Rev 2020, 120 (6), 3051-3126.

[0241] 31. Camarero, J. A.; Muir, T. W., Native chemical ligation of polypeptides. CurrProtoc Protein Sci 2001, Chapter 18, Unit184.

[0242] 32. Merrifield, R. B., Solid Phase Peptide Synthesis. I. The Synthesis of aTetrapeptide. Journal of the American Chemical Society 1963, 85 (14), 2149-2154.

[0243] 33. Severinov, K.; Muir, T. W., Expressed protein ligation, a novel method forstudying protein-protein interactions in transcription. J Biol Chem 1998, 273 (26), 16205-9.

[0244] 34. MacLaren, D. C.; Clarke, S., Expression and purification of a humanrecombinant methyltransferase that repairs damaged proteins. Protein Expr Purif 1995, 6 (1), 99-108.

[0245] 35. Ishitsuka, M. O.; Kusumi, T.; Kakisawa, H.; Kaya, K.; Watanabe, M. M.,Microviridin. A novel tricyclic depsipeptide from the toxic cyanobacterium Microcystis viridis. Journal of the American Chemical Society 1990, 112 (22), 8180-8182.

[0246] 36. Niedermeyer, T. H.; Strohalm, M., mMass as a software tool for theannotation of cyclic peptide tandem mass spectra. PLoS One 2012, 7 (9), e44913.

[0247] 37. Pbs(P). Cold Spring Harbor Protocols 2009, 2009 (1).- 78 - 4144725.v15391.1039001

[0248] 38. Cao, L.; Elashal, H. E.; Link, A. J., Kinetics of Aspartimide Formation andHydrolysis in Lasso Peptide Lihuanodin. Biochemistry 2023, 62 (3), 695-699.

[0249] 39. Ngoka, L. C.; Gross, M. L., A nomenclature system for labeling cyclic peptidefragments. J Am Soc Mass Spectrom 1999, 10 (4), 360-3.

[0250] 40. Güntert P, Buchner L. Combined automated NOE assignment and structurecalculation with CYANA. J Biomol NMR.2015 Aug;62(4):453-71. doi: 10.1007 / s10858- 015-9924-9. Epub 2015 Mar 24. PMID: 25801209.

[0251] 41. Hanwell MD, Curtis DE, Lonie DC, Vandermeersch T, Zurek E, HutchisonGR. Avogadro: an advanced semantic chemical editor, visualization, and analysis platform. J Cheminform.2012 Aug 13;4(1):17. doi: 10.1186 / 1758-2946-4-17. PMID: 22889332; PMCID: PMC3542060.

[0252] 42. Halgren TA. MMFF VI. MMFF94s option for energy minimization studies. JComput Chem.1999 May;20(7):720-729. doi: 10.1002 / (SICI)1096- 987X(199905)20:7<720::AID-JCC7>3.0.CO;2-X. PMID: 34376030.

[0253] The teachings of all patents, published applications and references cited herein areincorporated by reference in their entirety.

[0254] While example embodiments have been particularly shown and described, it willbe understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the embodiments encompassed by the appended claims. - 79 - 4144725.v1

Claims

5391.1039001 CLAIMS What is claimed is:

1. A peptide substrate for an ATP-grasp enzyme, the peptide substrate comprising atleast one core peptide unit, each core peptide unit comprising: a stem comprising at least one pair of macrocyclic linkage sites, each pair of macrocyclic linkage sites comprising a donor residue and an acceptor residue; and a loop comprising at least 4 amino acid residues.

2. The peptide substrate of claim 1, wherein the donor residue of at least one pair ofmacrocyclic linkage sites is selected from serine (Ser; S), valine (Val; V), cysteine (Cys; C), and Lysine (Lys; K).

3. The peptide substrate of claim 1 or 2, wherein at least one pair of macrocyclic linkagesites is a pair of esterification sites, and wherein the donor residue thereof is Ser, Val, or Thr.

4. The peptide substrate of any one of claims 1-3, wherein at least one pair ofmacrocyclic linkage sites is a pair of amidation sites, and wherein the donor residue thereof is Lys.

5. The peptide substrate of any one of claims 1-4, wherein at least one pair ofmacrocyclic linkage sites is a pair of thioesterification sites, and wherein the donor residue thereof is Cys.

6. The peptide substrate of any one of claims 1-5, wherein the acceptor residue of atleast one pair of macrocyclic linkage sites is selected from glutamic acid (Glu; E) and asparagine (Asn; N).

7. The peptide substrate of any one of claims 1-6, wherein the acceptor residue of atleast one pair of macrocyclic linkage sites is aspartic acid (Asp; D).

8. The peptide substrate of any one of claims 1-7, wherein the stem comprises an N-terminal portion connected to an N-terminal end of the loop and a C-terminal portion connected to a C-terminal end of the loop. - 80 - 4144725.v15391.10390019. The peptide substrate of any one of claims 1-8, wherein the donor residue of each pairof macrocyclic linkage sites is on the N-terminal portion of the stem, and wherein the acceptor residue of each pair of macrocyclic linkage sites is on the C-terminal portion of the stem.

10. The peptide substrate of any one of claims 1-9, wherein the stem comprises at leasttwo pairs of macrocyclic linkage sites.

11. The peptide substrate of any one of claims 1-10, wherein the stem comprises at leastthree pairs of macrocyclic linkage sites.

12. The peptide substrate of claim 10 or 11, wherein each donor residue is separated froman adjacent donor residue by 1-4 amino acid residues.

13. The peptide substrate of any one of claims 10-12, wherein each acceptor residue isseparated from an adjacent acceptor residue by 1-4 amino acid residues.

14. The peptide substrate of any one of claims 8-13, wherein the N-terminal portion of thestem of at least one core peptide unit comprises an amino acid sequence selected from SEQ ID NOs: 82-101.

15. The peptide substrate of any one of claims 8-14, wherein the C-terminal portion of thestem of at least one core peptide unit comprises an amino acid sequence selected from SEQ ID NOs: 102-114.

16. The peptide substrate of any one of claims 1-15, wherein the loop comprises at least 8amino acid residues.

17. The peptide substrate of any one of claims 1-16, wherein the loop comprises a peptideof interest or a protein of interest.

18. The peptide substrate of claim 17, wherein the loop further comprises at least onelinker connecting the peptide of interest or protein of interest to the stem.

19. The peptide substrate of claim 17 or 18, wherein the protein of interest is a fluorescentprotein. - 81 - 4144725.v15391.103900120. The peptide substrate of any one of claims 1-19, wherein the peptide substratecomprises at least two core peptide units, and wherein adjacent core peptide units are connected by a linker.

21. The peptide substrate of claim 20, wherein the peptide substrate comprises at leastthree core peptide units.

22. The peptide substrate of any one of claims 1-21, wherein the peptide substrate furthercomprises a leader peptide connected to an N-terminal end of a first core peptide unit.

23. The peptide substrate of any one of claims 1-22, wherein the peptide substrate furthercomprises a tail peptide connected to a C-terminal end of a last core peptide unit.

24. The peptide substrate of any one of claims 1-23, wherein the peptide substrate furthercomprises a detectable moiety.

25. The peptide substrate of claim 24, wherein the detectable moiety comprises a proteintag or a hydrazide.

26. The peptide substrate of any one of claims 1-25, wherein the peptide substrate isexpressed in a host cell.

27. The peptide substrate of claim 26, wherein the peptide substrate and the ATP-graspenzyme are co-expressed in a host cell.

28. The peptide substrate of claim 26 or 27, wherein the host cell is an E. coli cell.

29. The peptide substrate of any one of claims 1-28, wherein the ATP-grasp enzyme is agraspetide synthase.

30. The peptide substrate of any one of claims 1-29, wherein the ATP-grasp enzyme isencoded in the genome of Thermobifida fusca.

31. The peptide substrate of any one of claims 1-30, wherein the ATP-grasp enzymecomprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232).

32. A system for generating a macrocycle, comprising:- 82 - 4144725.v15391.1039001 the peptide substrate of any one of claims 1-31, and an ATP-grasp enzyme.

33. The system of claim 32, wherein the macrocycle comprises one or more macrocycliclinkages, and wherein each macrocyclic linkage is between a donor residue and an acceptor residue of a pair of macrocyclic linkage sites.

34. The system of claim 33, wherein the one or more macrocyclic linkages are selectedfrom ω-ester linkages, ω-amide linkages, ω-thioester linkages, and a combination of the foregoing.

35. The system of claim 33 or 34, wherein the macrocyclic linkages are installed by theATP-grasp enzyme.

36. A polynucleotide encoding the peptide substrate of any one of claims 1-31.

37. The polynucleotide of claim 36, wherein a sequence encoding the loop of at least onecore peptide unit of the peptide substrate comprises a restriction site.

38. A vector comprising the polynucleotide of claim 36 or 37.

39. The vector of claim 38, wherein the vector further comprises an additionalpolynucleotide encoding an ATP-grasp enzyme.

40. A host cell comprising the polynucleotide of claim 36 or 37 or the vector of claim 38or 39.

41. A host cell comprising the vector of claim 38 and further comprising an additionalvector encoding an ATP-grasp enzyme.

42. The vector of claim 39 or the host cell of claim 40 or 41, wherein the ATP-graspenzyme comprises an amino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232).

43. A kit comprising the system of any one of claims 32-35, the polynucleotide of claim36 or 37, the vector of any one of claims 38, 39, and 42 or the host cell of claim 41 or 42. - 83 - 4144725.v15391.103900144. A method for generating a macrocycle, the method comprising contacting the peptidesubstrate of any one of claims 1-31 with an ATP-grasp enzyme.

45. A method for generating a head-to-tail cyclized macrocycle, the method comprising:a) contacting the peptide substrate of any one of claims 1-31 with an ATP-graspenzyme, thereby producing a thioester-bearing macrocycle; and b) contacting the thioester-bearing macrocycle with trypsin.

46. A method for generating a branched, cyclic, or multicyclic peptide, the methodcomprising: a) contacting the peptide substrate of any one of claims 1-31 with an ATP-graspenzyme, thereby producing a thioester-bearing macrocycle; and b) performing native chemical ligation (NCL) on the thioester-bearingmacrocycle.

47. The method of claim 46, wherein the NCL is performed with free cysteine.

48. The method of claim 46 or 47, wherein the NCL is performed with 2-(4-sulfanylphenyl)acetic acid (MPAA).

49. The method of any one of claims 45-48, wherein the peptide substrate comprises atleast two cysteine (Cys) donor residues on an N-terminal portion of the stem and at least one acceptor residue on a C-terminal portion of the stem.

50. The method of any one of claims 45-49, wherein the N-terminal portion of the stemcomprises SEQ ID NO: 101.

51. The method of any one of claims 45-50, wherein the C-terminal portion of the stemcomprises SEQ ID NO: 113.

52. The method of any one of claims 45-51, wherein the ATP-grasp enzyme is agraspetide synthase.

53. The method of any one of claims 45-52, wherein the ATP-grasp enzyme is encoded inthe genome of Thermobifida fusca. - 84 - 4144725.v15391.103900154. The method of any one of claims 45-53, wherein the ATP-grasp enzyme comprises anamino acid sequence with at least 75% sequence identity to ThfB (SEQ ID NO: 115 or 232). - 85 - 4144725.v1