Methods of making cyclic peptides

The method of fusing linear peptides with G macrocyclase recognition sites and linkers, combined with RSI-TruD cyclodehydratase, addresses the inefficiencies of existing cyclic peptide production, achieving high-yield and diverse libraries suitable for therapeutic applications.

WO2025129035A9PCT designated stage expired Publication Date: 2025-08-07UNIV OF UTAH RES FOUND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/060088
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-12-13
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing methods for producing cyclic peptides face challenges such as slow reaction rates, substrate hydrolysis, and low yields, particularly in Escherichia coli expression libraries, due to the presence of N-terminal elements that need processing and cleavage, and unfavorable solubility and folding issues in vitro.

Method used

A method involving linear peptides fused with a G macrocyclase recognition site, a linker, and a G macrocyclase enzyme, followed by cyclization using RSI-TruD cyclodehydratase or under suitable conditions to generate cyclic peptides, optimizing the process for both in vivo and in vitro production.

Benefits of technology

This approach enhances the efficiency and yield of cyclic peptide production, enabling the generation of diverse libraries with high fidelity and suitability for screening and therapeutic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024060088_07082025_PF_FP_ABST
    Figure US2024060088_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein, are peptides comprising from N terminus to C terminus, a) a peptide of interest, wherein the peptide of interest having 6 to 20 amino acid residues and a C -terminal residue that facilitates cyclization; b) a G macrocyclase recognition site; c) a linker; and d) G macrocyclase; wherein the peptide is linear, wherein the peptide of interest is covalently bonded to the G macrocyclase recognition site, wherein the G macrocyclase recognition site is bond to the linker, and wherein the linker is bonded to the G macrocyclase, and methods of making cyclized peptides and generating libraries of cyclized peptides.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS OF MAKING CYCLIC PEPTIDES

[0002] CROSS REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Application No. 63 / 610,886, filed December 15, 2023. The content of this earlier filed application is hereby incorporated by reference herein in their entirety.

[0004] STATEMENT REGARDING FEDERALLY FUNDED RESEARCH

[0005] This invention was made with government support under grant number R35GM148283 awarded by the National Institutes of Health. The government has certain rights in the invention.

[0006] INCORPORATION OF THE SEQUENCE LISTING

[0007] The present application contains a sequence listing that is submitted concurrent with the filing of this application, containing the file name “21101_0473Pl_SL.xml” which is 147,456 bytes in size, created on December 6, 2024, and is herein incorporated by reference in its entirety.

[0008] SUMMARY

[0009] Disclosed herein are peptides comprising from N terminus to C terminus, a) a peptide of interest, wherein the peptide of interest having 6 to 20 amino acid residues and a C- terminal cysteine residue; b) a G macrocyclase recognition site; c) a linker; and d) G macrocyclase; wherein the peptide is linear, wherein the peptide of interest is covalently bonded to the G macrocyclase recognition site, wherein the G macrocyclase recognition site is bonded to the linker, and wherein the linker is bonded to the G macrocyclase.

[0010] Disclosed herein are methods of producing a cyclic peptide, the methods comprising: a) fusing a linear peptide comprising from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal cysteine residue, ii) a G macrocyclase recognition site, iii) a linker, iv) a G macrocyclase enzyme; wherein the peptide of interest is covalently bonded to the G macrocyclase recognition site, wherein the G macrocyclase recognition site is bond to the linker, and wherein the linker is bonded to the G macrocyclase; and b) contacting the linear peptide in a) with RSI-TruD cyclodehydratase, thereby generating the cyclic peptide.

[0011] Disclosed herein are methods of producing a cyclic peptide, the methods comprising: a) fusing a linear peptide comprising from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal cysteine residue, ii) a G macrocyclase recognition site, iii) a linker, iv) a G macrocyclase enzyme: wherein the peptide of interest is covalently bonded to the G macrocyclase recognition site, wherein the G macrocyclase recognition site is bond to the linker, and wherein the linker is bonded to the G macrocyclase; and b) contacting the linear peptide in a) with an enzyme or under suitable conditions to allow the linear peptide to cyclize, thereby generating the cyclic peptide.

[0012] Disclosed herein are methods of producing a cyclic peptide, the methods comprising: a) fusing a linear peptide comprising from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal residue that facilitates cyclization, ii) a G macrocyclase recognition site, iii) a linker, iv) a G macrocyclase enzyme; wherein the peptide of interest is covalently bonded to the G macrocyclase recognition site, wherein the G macrocyclase recognition site is bond to the linker, and wherein the linker is bonded to the G macrocyclase; and b) contacting the linear peptide in a) with an enzyme or under suitable conditions to allow the linear peptide to cyclize, thereby generating the cyclic peptide.

[0013] Other features and advantages of the present compositions and methods are illustrated in the description below, the drawings, and the claims.

[0014] BRIEF DESCRIPTION OF THE DRAWINGS

[0015] FIGS. 1 A-B show the process of covalently fusing PagG macrocyclase enzymes with prenylagramide core substrate. FIG. 1A shows protein constructs containing the artificial PagE precursor peptide fused with the PagG macrocyclase enzyme through flexible linkers with varying sizes. Treatment of these proteins with PatA protease cleaves the leader sequences, liberating the N-terminus of the core sequences, which then ultimately leads to the rapid release of cyclized products. FIG. IB depicts the LC-MS analysis of the reaction mixture of designed constructs showing cyclic prenylagramide (1) formation at tR = 3.15 mins, with the longest-size GSG linker providing the most efficient production. Scale is constant in the traces and represents total counts.

[0016] FIG. 2 shows a simplified construct dubbed as “auto-PagG” (SEQ ID NO: 6) that was designed to produce cyclic peptides in E. coll in vivo. A mixture of cyclic peptides containing the core and core with additional methionine was observed from organic extracts in this simplified protein construct, as shown by EIC peaks from LC-MS analysis. Scale is constant in the traces, showing relative abundance of each peak in the extract. FIGS. 3A-B show utilizing auto-PatG and RSI-TruD as control for peptide cyclization. FIG. 3A shows protein constructs containing patellin cores and RSIII (SYD) that were fused with the PatG macrocyclase enzyme through the optimized GSG flexible linker. Treatment of these constructs with RSI-TruD converts the C-term cysteine residue of the core into thiazoline ring, cascading into peptide cyclization. FIG. 3B depicts LC-MS analysis of the reaction mixture of designed protein constructs showing EIC peaks of cyclized products. Core peptide sequences used included the precursors to the natural products patellin 2 (TVPTLC; SEQ ID NO: 3) and patellin 3 (TLPVPTLC; SEQ ID NO: 4) , as well as a mutant of the patellin 2 sequence (TVPTVC; SEQ ID NO: 5).

[0017] FIGS. 4A-C show using auto-PatG and RSI-TruD to make an octapeptide library . FIG. 4A shows core sequences of substrate-fused Pat G were randomized from the Pl -7 position, allowing codons specific to certain amino acids. Coexpression with RSI-TruD enables the conversion of cysteine into thiazoline ring, which permits peptide cyclization. Mutant plasmids were transformed in BL21(DE3) E. coli cells and grown for 3 days at 37 °C. Cells were harv ested and treated with acetone, which is then used for LC-MS analysis. FIG. 4B shows the acceptance rate of selected amino acids in positions P1-P7 for peptide cyclization, generated by WebLogo (Crooks, G. E., et al. Genome Res. 2004, 14(6): 1188- 1190). FIG. 4C shows ‘H-NMR of cyclic TAYWTWIC (SEQ ID NO: 7) (7) in DMF- d7(500 MHz) at 25 °C.

[0018] FIGS. 5A-B show cloning strategies used in this study. FIG. 5A shows the Gibson reaction that was used to generate the auto-PagG in pet28b plasmid. Mutagenic primers were used to remove the leader peptide and PatA elements to generate a simplified “auto-PagG” plasmid. FIG. 5B show s that a similar Gibson strategy w as used to create the auto-PatG, to build the auto-PagG construct. Restriction cloning was used to transfer the RSI-TruD and auto-PatG gene in the pRSF-Duetl vector, followed by using mutagenic primers to remove unwanted codons in the N-terminus of the auto-PatG gene. This plasmid was used to create the octapeptide library7and to isolate the cyanobactin mutant.

[0019] FIG. 6 shows EIC peak comparison of formed monomer: dimer patellin 2 (4) while changing the auto-PatG and RSI-TruD molar ratio in the in vitro enzy me assay.

[0020] FIG. 7 depicts a heat map showing acceptance rate of selected amino acids overall and each position, P1-P7, for peptide cyclization. The heat map w as obtained using LC-MS analysis from bacterial cultures of each construct in the mutagenesis library .

[0021] FIG. 8 shows COSY. HMBC, and ROESY correlations confirm the structure of the isolated cyanobactin mutant (7) (TAYWTWIC; SEQ ID NO: 7). FIG. 9 shows LC-MS standard curve for cyclo[TAYWTWIC] (SEQ ID NO: 7) (7) quantification. The extracted ion chromatograms of cyclo[TAYWTWIC] (SEQ ID NO: 7) (7) in the organic extracts with increments of the purified standard were integrated and used to generate a standard curve. This was done in three independent replicates.

[0022] DETAILED DESCRIPTION

[0023] The present disclosure can be understood more readily by reference to the following detailed description of the invention, the figures and the examples included herein.

[0024] Before the present compositions and methods are disclosed and described, it is to be understood that they are not limited to specific synthetic methods unless otherwise specified, or to particular reagents unless otherwise specified, as such may, of course, vary. It is also to be understood that the terminology’ used herein is for the purpose of describing particular aspects only and is not intended to be limiting. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, example methods and materials are now’ described.

[0025] Moreover, it is to be understood that unless otherwise expressly stated, it is in no way intended that any method set forth herein be construed as requiring that its steps be performed in a specific order. Accordingly, where a method claim does not actually recite an order to be followed by its steps or it is not otherwise specifically stated in the claims or descriptions that the steps are to be limited to a specific order, it is in no way intended that an order be inferred, in any respect. This holds for any possible non-express basis for interpretation, including matters of logic with respect to arrangement of steps or operational flow, plain meaning derived from grammatical organization or punctuation, and the number or type of aspects described in the specification.

[0026] All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided herein can be different from the actual publication dates, which can require independent confirmation.

[0027] DEFINITIONS

[0028] As used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. The word “or” as used herein means any one member of a particular list and also includes any combination of members of that list.

[0029] Ranges can be expressed herein as from “about” or “approximately” one particular value, and / or to “about” or “approximately” another particular value. When such a range is expressed, a further aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” or “approximately,” it will be understood that the particular value forms a further aspect. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint and independently of the other endpoint. It is also understood that there are a number of values disclosed herein and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. It is also understood that each unit between two particular units is also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.

[0030] As used herein, the terms “optional” or “optionally” mean that the subsequently described event or circumstance may or may not occur and that the description includes instances where said event or circumstance occurs and instances where it does not.

[0031] Throughout the description and claims of this specification, the word “comprise” and variations of the word, such as “comprising” and “comprises,” means “including but not limited to,” and is not intended to exclude, for example, other additives, components, integers or steps. In particular, in methods stated as comprising one or more steps or operations it is specifically contemplated that each step comprises what is listed (unless that step includes a limiting term such as “consisting of’), meaning that each step is not intended to exclude, for example, other additives, components, integers or steps that are not listed in the step.

[0032] As used herein the terms “amino acid” and “amino acid identity” refers to one of the 20 naturally occurring amino acids or any non-natural analogues that may be in any of the antibodies, variants, or fragments disclosed. Thus “amino acid” as used herein means both naturally occurring and synthetic amino acids. For example, homophenylalanine, citrulline and norleucine are considered amino acids for the purposes of the invention. “Amino acid” also includes amino acid residues such as proline and hydroxyproline. The side chain may be in either the (R) or the (S) configuration. In some aspects, the amino acids are in the D- or L- configuration. If non-naturally occurring side chains are used, non-amino acid substituents may be used, for example to prevent or retard in vivo degradation. As used herein, the term “polypeptide"’ refers to a polymer composed of amino acid residues related naturally occurring structural variants, and synthetic non-naturally occurring analogs thereof linked via peptide bonds or modified peptide bonds (i.e., peptide isosteres), related naturally occurring structural variants, and synthetic non-naturally occurring analogs thereof, glycosylated polypeptides, and all “mimetic” and “peptidomimetic” polypeptide forms. Synthetic polypeptides can be synthesized, for example, using an automated polypeptide synthesizer. The term can refer to an oligopeptide, peptide, polypeptide, or protein sequence, or to a fragment, portion, or subunit of any of these. The term “protein” typically refers to large polypeptides. The term “peptide” ty pically refers to short polypeptides.

[0033] A “portion” of a polypeptide or protein means at least about three sequential amino acid residues of the polypeptide. It is understood that a portion of a polypeptide may include every amino acid residue of the polypeptide.

[0034] The term “fragment” can refer to a portion (e.g., at least 5, 10, 25, 50, 100, 125, 150, 200, 250. 300, 350, 400 or 500, etc. amino acids or nucleic acids) of a peptide that is substantially identical to a reference peptide and retains the biological activity of the reference peptide. In some aspects, the fragment or portion of a peptide retains at least 50%, 75%, 80%, 85%, 90%, 95% or 99% of the biological activity of the reference peptide described herein. A fragment of a referenced peptide can be a continuous or contiguous portion of the referenced polypeptide (e.g., a fragment of a reference peptide that is ten amino acids long can be any 2-9 contiguous residues within that reference peptide).

[0035] “Mutants,” “derivatives,” and “variants” of a polypeptide (or of the nucleic acid encoding the same) are polypeptides (or the nucleic acids) which may be modified or altered in one or more amino acids (or in one or more nucleotides) such that the peptide (or the nucleic acid) is not identical to the wild-type sequence, but has homology to the wild type polypeptide (or the nucleic acid).

[0036] The term “variant” can refer to a peptide or gene product that displays modifications in sequence and / or functional properties (i.e., altered characteristics) when compared to the wild-type peptide or gene product. In general, it is understood that one way to define any known variants and derivatives or those that might arise, of the disclosed genes and proteins herein, is through defining the variants and derivatives in terms of homology to specific known sequences. This identity of particular sequences disclosed herein is also discussed elsewhere herein. In general, variants of genes and proteins herein disclosed typically have at least, about 70. 71. 72. 73. 74. 75, 76, 77, 78, 79, 80, 81, 82, 83, 84. 85. 86. 87. 88. 89. 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99 percent homology to the stated sequence or the native sequence. Those of skill in the art readily understand how to determine the homology’ of two proteins or nucleic acids, such as genes. For example, the homology can be calculated after aligning the two sequences so that the homology is at its highest level. In an aspect, the term “variant” can mean a difference in some way from the reference sequence other than just a simple deletion of an N- and / or C-terminal amino acid residue or residues. In an aspect, a variant can include a substitution of an amino acid residue, the substitution can be considered conservative or non-conservative. Conservative substitutions are those within the following groups: Ser, Thr, and Cys; Leu, He, and Vai; Glu and Asp; Lys and Arg; Phe, Tyr, and Trp; and Gin, Asn, Glu, Asp, and His. Variants can include at least one substitution and / or at least one addition, there may also be at least one deletion. Variants can also include one or more non-naturally occurring residues. For example, they may include selenocysteine (e.g., seleno- L- cysteine) at any position, including in the place of cysteine. Many other “unnatural” amino acid substitutes are known in the art and are available from commercial sources. Examples of non-naturally occurring amino acids include D-amino acids, amino acid residues having an acetylaminomethyl group attached to a sulfur atom of a cysteine, a pegylated amino acid, and omega amino acids of the formula NH2(CH2)nCOOH wherein n is 2-6 neutral, nonpolar amino acids, such as sarcosine, t-butyl alanine, t-butyl glycine, N-methyl isoleucine, and norleucine. Phenylglycine may substitute for Trp, Tyr, or Phe; citrulline and methionine sulfoxide are neutral nonpolar, cysteic acid is acidic, and ornithine is basic. Proline may be substituted with hydroxyproline and retain the conformation conferring properties of proline.

[0037] As used herein, the term “substituted” is contemplated to include all permissible substituents of organic compounds. In a broad aspect, the permissible substituents include acyclic and cyclic, branched and unbranched, carbocyclic and heterocyclic, and aromatic and nonaromatic substituents of organic compounds. Illustrative substituents include, for example, those described below. The permissible substituents can be one or more, and the same or different for appropriate organic compounds. For purposes of this disclosure, the heteroatoms, such as nitrogen, can have hydrogen substituents and / or any permissible substituents of organic compounds described herein which satisfy the valences of the heteroatoms. This disclosure is not intended to be limited in any manner by the permissible substituents of organic compounds. Also, the terms “substitution” or “substituted with” include the implicit proviso that such substitution is in accordance with permitted valence of the substituted atom and the substituent, and that the substitution results in a stable compound, e.g., a compound that does not spontaneously undergo transformation such as by rearrangement, cyclization, elimination, etc. It is also contemplated that, in certain aspects, unless expressly indicated to the contrary, individual substituents can be further optionally substituted (i.e., further substituted or unsubstituted).

[0038] The term “heterologous” refers to an element that is not associated or linked to the subject feature in its natural environment. In other words, the association with a heterologous element is artificial and the element is associated or linked to the subject feature through human intervention.

[0039] Other objects, features and advantages of the present invention will become apparent from the following detailed description. It should be understood, however, that the detailed description and the specific examples, while indicating specific embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description.

[0040] Cyclic peptides have attracted attention as drug candidates because their properties bridge those of biologies and small molecules. As a result of their therapeutic advantages, cyclic peptides are growing rapidly as FDA-approved drugs and clinical trials candidates (Zhang, H. and Chen, S. RSC Chem. Biol. 2022, 3 (1), 18-31 ; and Costa, L., et al. Pharmaceuticals 2023, 16 (7), 996). Numerous cyclization strategies have been developed, including synthetic and biosynthetic methods, each of which has advantages in specific applications. For example, the biosynthetic approach enables genetic randomization to construct peptide libraries, which are compatible with selection- or phenotype-based screens directly within the producing organisms (Y oung, T. S., et al. Proc. Natl. Acad. Sci. 2011, 108 (27), 11052-11056; Horswill, A. R„ et al. Proc. Natl. Acad. Sci. 2004, 101 (44), 15591- 15596; Tavassoli, A. Curr. Opin. Chem. Biol. 2017. 38, 30-35; and Scott, C. P , et al. Proc. Natl. Acad. Sci. 1999, 96 (24). 13638-13643). For that reason, many genetic methods have been developed. Arguably, the most successful of these has been split-intein circular ligation of peptides and proteins (SICLOPPS) (Scott, C. P., et al. Proc. Natl. Acad. Sci. 1999, 96 (24), 13638-13643). Other tools have been developed, such as modified thioesterase domains (Kobayashi, M., et al. J. Am. Chem. Soc. 2023, 145 (6). 3270-3275; and Kohli, R. M.. et al. Nature 2002, 418 (6898), 658-661). Genetic code reprogramming supplements these methods by bringing in non-proteinogenic amino acids (Hipolito, C. J. and Suga, H. Curr. Opin. Chem. Biol. 2012, 16 (1-2), 196-203; and Tianero, Ma. D. B. et al. J. Am. Chem. Soc. 2012, 134 (1), 418-425). Each of these methods has strengths and limitations in the scope of substrate tolerance, the production of unwanted mixtures, and the relative yield produced in vivo.

[0041] The relaxed substrate promiscuity of ribosomal and post-translationally modified peptide (RiPP) pathways has enabled their application to library generation (Ruffner, D. E.; et al. ACS Synth. Biol. 2015, 4 (4), 482-492; Doma, M. S.; et al. Nat. Chem. Biol. 2006, 2 (12), 729-735; Pan, S. J. and Link, A. J. J. Am. Chem. Soc. 2011, 133 (13), 5016-5023; and Young. T. S., et al. Chem. Biol. 2012, 19 (12), 1600-1610). For example, cyanobactin-family macrocyclases, the G enzymes, show exceptionally broad substrate tolerance and high fidelity, efficiently generating a single cyclic peptide without observable linear side products. G enzymes are subtilisin-like proteases that, instead of hydrolysis, perform transpeptidation to form macrocycles. The native substrates for PatG and PagG macrocyclases are short peptides comprised of 6-10 amino acids in the hypervariable ‘'core peptide’’ region that encodes the final cyclized product, and 3-5 amino acids in the C-terminus that comprise the conserved “recognition sequence”. The core peptide is directed to the enzy me's active site through interactions with the recognition sequence. Once positioned, G enzymes cleave the peptide to form an enzyme-core peptide covalent intermediate. The core peptide’s free N- terminus attacks the ester intermediate, releasing a cyclized peptide product. In both PagG and PatG, a heterocyclic residue is important at the C-terminus of the core peptide, while other positions in the core are broadly substrate tolerant. In PagG, the preferred heterocycle is proline, while in PatG. an azoline is preferred; this azoline is usually produced enzymatically (Agarwal. V., et al. Chem. Biol. 2012, 19 (1 1), 1411-1422; Lee, J.; et al. J. Am. Chem. Soc. 2009, 131 (6), 2122-2124; Sarkar, S„ et al. ACS Catal. 2020, 10 (13), 7146-7153; and McIntosh, J. A., et al. J. Am. Chem. Soc. 2010, 132 (44), 15499-15501).

[0042] Because of their unusual fidelity and substrate tolerance, G enzymes have been widely applied to synthesizing cyclized peptide derivatives and large libraries (Donia, M. S., et al. Nat. Chem. Biol. 2006, 2 (12), 729-735; Pan, S. J. and Link, A. J. J. Am. Chem. Soc. 2011, 133 (13), 5016-5023; Young, T. S„ et al. Chem. Biol. 2012, 19 (12), 1600-1610; Agarwal, V., et al. Chem. Biol. 2012, 19 (11), 1411-1422; Lee, J., et al. J. Am. Chem. Soc. 2009, 131 (6), 2122-2124; Sarkar. S„ et al. ACS Catal. 2020. 10 (13), 7146-7153; McIntosh, J. A., et al. J. Am. Chem. Soc. 2010, 132 (44), 15499-15501; and Gu, W„ et al. The Biochemistry and Structural Biology' of Cyanobactin Pathways: Enabling Combinatorial Biosynthesis. In Methods in Enzymology; Elsevier, 2018; Vol. 604, pp 113-163). Experiment has demonstrated that G enzymes broadly accept many different substrates and are often capable of making millions of compounds, if just considering proteinogemc amino acids (Ruffner, D. E., et al. ACS Synth. Biol. 2015, 4 (4), 482-492; and Sarkar, S., et al. ACS Catal. 2020, 10 (13), 7146-7153). Adding to this diversity, G enzymes are compatible with D-amino acids (McIntosh, J. A., et al. J. Am. Chem. Soc. 2010, 132 (44), 15499-15501) and non-proteinogenic amino acids (Tianero, Ma. D. B., et al. J. Am. Chem. Soc. 2012, 134 (1), 418-425), making them amenable to generating compounds and libraries that bridge synthetic and genetic strategies.

[0043] However, several factors have limited their success. In Escherichia COII- )<\SQ expression libraries, instead of a simple core-recognition sequence substrate, the precursor peptides also contain N-terminal elements that must be processed and cleaved prior to the action of G enzymes, and the reactions are quite slow, creating several difficulties (Tianero, Ma. D., et al. Proc. Natl. Acad. Sci. 2016, 113 (7), 1772-1777). In vitro, the purified G enzymes and substrates are combined, but reactions are sometimes low-yield and do not always go to completion. This leaves a mixture that is not suitable for screening (McIntosh, J. A.; Robertson, C. R.; Agarwal, V.; Nair, S. K.; Bulaj, G. W.; Schmidt, E. W. Circular Logic: Nonribosomal Peptide-like Macrocyclization with a Ribosomal Peptide Catalyst. J. Am. Chem. Soc. 2010, 132 (44), 15499-15501). As described herein, it was tested whether the in vivo and in vitro limitations were largely caused by the substrates and not by the inherent efficiencies of the enzy mes. For example, in vivo different linear substrates may be readily hydrolyzed in E. coli in a sequence-dependent manner. Similarly, in vitro, short peptide substrates often have unfavorable solubility, oligomerization, or folding terms, which are also sequence-dependent.

[0044] Disclosed herein are methods for producing cyclized peptides in vivo and in vitro. In some aspects, the methods use cyanobacterial enzymes, such as patellamide biosynthesis enzymes. The methods and cyclized peptides can be useful for the production of peptidyl molecules, the biosynthesis and screening of candidate therapeutics, and nanotechnology applications.

[0045] COMPOSITIONS

[0046] Disclosed herein are linear peptides that can be used to generate one or more of the cyclized peptides disclosed. Disclosed herein are linear peptides that can be used to generate libraries comprising cyclized peptides.

[0047] Disclosed herein are peptides comprising from N terminus to C terminus: a) a peptide of interest; b) a G macrocyclase recognition site; c) a linker; and d) G macrocyclase. In some aspects, the peptide of interest has 6 to 20 amino acid residues and a C-terminal cysteine residue. In some aspects, the peptide can be linear. In some aspects, the peptide of interest can be covalently bonded to the G macrocyclase recognition site. In some aspects, the G macrocyclase recognition site can be bonded to the linker, and the linker can be bonded to the G macrocyclase.

[0048] In some aspects, the peptide can be MQAYLGIPLPFAGDDAEGSGGSGGSGGSGGSGGSGGSGGSGGSGGSGGSGGSGMPD LITIPGIPELWTQTKGDSRIKIAILDGAADLERACFKGAKITQFKPYWAEDIELNDEYY HYLKLATEFNQQQKAKKEDPDHDKEEAKKEREAFFKDFPEDIKRRIDLSSHATHISST ILGQHGSPVEGIAPNCTAINIPISFAGDDFISFVNLTHAINEALKAEVNIVHIAACHPTQ SGMAQEIFARAVKQCQDSNILIVAPGGNKDGECWCIPSILPDVLTVGAMRDDGQPFK FSNYGGEYQHKGVMANGENILGANPGTDEPVREKGTSCAAPIVTGISALLMSMQLQ RGEKPNAETVRQAILKSAIPCDQNEVEEPERCLLGKLNIPGAYNLLTGERLTTVKTSEI RQSEIT (SEQ ID NO: 6).

[0049] In some aspects, the peptide of interest can comprise a C-terminal cysteine residue. In some aspects, the peptide of interest can have a N-terminal methionine. In some aspects, the peptide of interest can comprise a C-terminal residue that facilitates cyclization. In some aspects, the C-terminal residue that facilitates cyclization can be a cysteine, a proline, an oxazoline, or a heterocycle. In some aspects, the oxazoline can be derived from a serine or a threonine. In some aspects, the heterocycle can be a triazole, an imidazole, or a pseudoproline. In some aspects, the peptide of interest can be any of the peptides listed in Table 6. In some aspects, the peptide of interest can have 6 to 20 amino acid residues. In some aspects, the peptide of interest can have 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid residues. In some aspects, the peptide of interest can be a cyanobactin sequence or a precursor thereof or a natural cyclotide sequence or a precursor thereof. In some aspects, the peptide of interest can be synthetic. In some aspects, the peptide of interest can be a heterologous sequence which is not normally associated with a G macrocyclase recognition site.

[0050] In some aspects, the peptide of interest can include modified amino acids, unmodified amino acids, heterocyclic amino acids, nonheterocyclic amino acids, naturally occurring amino acids, and / or non-naturally occurring amino acids. In some aspects, the peptide of interest can comprise heterocyclic amino acids using isolated cyanobacterial enzymes, and optionally the oxidation of the introduced heterocyclic amino acids. In some aspects, the peptide of interest can comprise 0, 1, 2, 3, 4, 5, 6, 7, 8 or more heterocyclic amino acids (Shin-ya, K. et al. J. Am. Chem. Soc. 2001, 123, 1262-1263). In some aspects, the residue directly N terminal to the cyclization signal in the peptide of interest can be a heterocyclic amino acid. For example, an amino acid selected from thiazoline (Tim), thiazole (Thz), oxazoline (Oxn), oxazole (Oxz), proline and pseudoproline (Pro).

[0051] Examples of peptides of interest include but are not limited to: QAYGIPLP (SEQ ID NO: 1), MQAYGIPLP (SEQ ID NO: 2), TVPTLC (SEQ ID NO: 3), TLPVPTLC (SEQ ID NO: 4), TVPTVC (SEQ ID NO: 5), TAYWTWIC (SEQ ID NO: 7), QAYLGIPLP (SEQ ID NO: 15), MQAYLGIPLP (SEQ ID NO: 16), WFTWALLC (SEQ ID NO: 49), ISYSILSC (SEQ ID NO: 50), TWASIIIC (SEQ ID NO: 53), TYVASVTC (SEQ ID NO: 54), SAYTLSLC (SEQ ID NO: 55), ATWTSISC (SEQ ID NO: 58), IISSFTFC (SEQ ID NO: 59), SAATVYLC (SEQ ID NO: 60), IWITVWTC (SEQ ID NO: 63), TTWAAWAC (SEQ ID NO: 64), ISFISSFC (SEQ ID NO: 65), SSVVWSVC (SEQ ID NO: 66), FWTLAVYC (SEQ ID NO: 67), ALTVVTTC (SEQ ID NO: 71), TAYWTWIC (SEQ ID NO: 73), WSVLWFLC (SEQ ID NO: 74), WTFYVLLC (SEQ ID NO: 76), TTYTVSVC (SEQ ID NO: 77), SLVASVFC (SEQ ID NO: 78), WWFVTFSC (SEQ ID NO: 83), VVLSSYIC (SEQ ID NO: 84), SIFLATFC (SEQ ID NO: 85), IASLSTIC (SEQ ID NO: 87), YSVFWAVC (SEQ ID NO: 89), SFIWFYAC (SEQ ID NO: 90), ISLLWVLC (SEQ ID NO: 91). IIWIILWC (SEQ ID NO: 94), YILWTTVC (SEQ ID NO: 96), WWTAFFTC (SEQ ID NO: 98), LFFVYTIC (SEQ ID NO: 99), IVWAWYWC (SEQ ID NO: 100), AVTALWAC (SEQ ID NO: 101), SLVLIYTC (SEQ ID NO: 102), SVIFVIYC (SEQ ID NO: 104), SLVWTVFC (SEQ ID NO: 112), SAYYTFAC (SEQ ID NO: 113), LLFWTVYC (SEQ ID NO: 114), LVIYFLIC (SEQ ID NO: 115), YTLYVWIC (SEQ ID NO: 116), TVFLITTC (SEQ ID NO: 117), TYFWYSAC (SEQ ID NO: 119), VIAS AYL C (SEQ ID NO: 120), WWTW1TIC (SEQ ID NO: 124), TSSFVSWC (SEQ ID NO: 125), FITLWVAC (SEQ ID NO: 131), VFSAVTIC (SEQ ID NO: 133), YTLYLILC (SEQ ID NO: 135), SIAAVSLC (SEQ ID NO: 136), FYIISTLC (SEQ ID NO: 137), ILALFAIC (SEQ ID NO: 139), TTTLIFVC (SEQ ID NO: 141). Other examples include but are not limited to TLAT1C (SEQ ID NO: 144). 1VPPFC (SEQ ID NO: 145), TTVTAC (SEQ ID NO: 146), VTPFVC (SEQ ID NO: 147), IPGSLC (SEQ ID NO: 148), TSIAPFC (SEQ ID NO: 149), TSIAPLC (SEQ ID NO: 150), TSIASFC (SEQ ID NO: 151), TVIAPFC (SEQ ID NO: 152), 1PISFPC (SEQ ID NO: 153), TLPVPTVC (SEQ ID NO: 154), TVPVPSFC (SEQ ID NO: 155), and VTACITFC (SEQ ID NO: 156). In some aspects, the disclosed peptides can further comprise thiazoline (Tim), thiazole (Thz), oxazoline (Oxn), oxazole (Oxz), proline and pseudoproline (Pro). In some aspects, one or more of the amino acids of the peptide of interest disclosed herein can have one or more amino acid substitutions. For example, one or more of the amino acids of the peptides disclosed in Table 6 can be substituted with thiazoline (Tim), thiazole (Thz), oxazoline (Oxn), oxazole (Oxz), proline or pseudoproline (Pro).

[0052] In some aspects, one or more residues in the peptide of interest can comprise a reactive functionality which may allow further chemical modification. Suitable residues that contain side chains with side chain linking groups include NH2, COOH, OH and SH. In some aspects, one or more residues in the peptide of interest can comprise NH2, COOH, OH or SH reactive groups attached to the peptide of interest.

[0053] In some aspects, the peptide of interest can further comprise a pro-sequence. In some aspects, the pro-sequence can be an N-tenninal leader sequence. In some aspects, the prosequence can be covalently bonded to the peptide of interest via a protease recognition site. A cyanobacterial protease is an enzyme from a cyanobacterium that can cleave a peptide at a protease recognition site. In some aspects, the cyanobacterial protease can be a PatA protease or a variant thereof. In some aspects, the sequence of the PatA protease can be Genbank AAY21150.0.

[0054] The peptide of interest can be covalently bonded to the G macrocyclase recognition site at the C-terminal cysteine residue. The G macrocyclase recognition site is located between the peptide of interest and the linker. The G macrocyclase recognition site can be covalently bonded to the linker at the C-terminal end of the G macrocyclase recognition site.

[0055] The G macrocyclase recognition site serves as the recognition site for the G macrocyclase. In some aspects, the G macrocyclase recognition site can be determined by the G macrocyclase being used. In some aspects, the G macrocyclase recognition site can comprise a small residue, a bulky residue, or an acidic residue. In some aspects, the G macrocyclase recognition site can be SYD, AYD, FAGDDAE (SEQ ID NO: 143), SYE, AYE, SFD, SVG, or AFD. In some aspects, the G macrocyclase recognition site comprises or consists of SYD. AYD, FAGDDAE (SEQ ID NO: 143), SYE, AYE, SFD, SVG. or AFD. In some aspects, the G macrocyclase recognition site can be the native recognition site of the G macrocyclase. In some aspects, the G macrocyclase recognition site can be a fragment of FAGDDAE (SEQ ID NO: 143).

[0056] In some aspects, the G macrocyclase can be a PatG macrocyclase and the G macrocyclase recognition site can be AYD or SYD. In some aspects, the G macrocyclase can be a PatG macrocyclase and the G macrocyclase recognition site can be SYE, AYE, SFD, SVG, or AFD. In some aspects, the G macrocyclase can be a PagG macrocyclase and the G macrocyclase recognition site can be FAGDDAE (SEQ ID NO: 143) or a fragment thereof. In some aspects, the G macrocyclase recognition site can be heterologous (e.g., not naturally associated with the peptide of interest). In some aspects, the G macrocyclase recognition site can be a synthetic or a modified G macrocyclase recognition site.

[0057] In some aspects, the G macrocyclase recognition site can be bonded to the linker and the linker can be bonded to the G macrocyclase. The linkers can be of any length, of a flexible sequence and not have any charges. In some aspects, the linker can be GSG or GS linkers that range between 3 and 40 amino acids. In some aspects, the linker can be GSGnn = 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12. In some aspects, the linker can be GSGs, GSGe, GSG$>, or GSG12. In some aspects, the linker can be GGGGSnn = 2, 3, 4, 5, 6, 7, or 8. In some aspects, the linker can be GSGGSGGSGGSGGSGGSGGSGGSGGSGGSGGSGGSG (SEQ ID NO: 30). In some aspects, the linker can be GSGGSGGSG (SEQ ID NO: 33). In some aspects, the linker can be GSGGSGGSGGSGGSGGSG (SEQ ID NO: 32). In some aspects, the linker can be GSGGSGGSGGSGGSGGSGGSGGSGGSG (SEQ ID NO: 31).

[0058] A G macrocyclase is a cyanobacterial enzyme that can catalyze the peptide of interest which contains a G macrocyclase recognition site. Examples of G macrocyclases include but are not limited to PatG macrocyclase, PagG macrocyclase, TruG macrocyclase, ThcG macrocyclase, OscG macrocyclase, TenG macrocyclase, LynG macrocyclase, McaG macrocyclase, AgeG macrocyclase, TriG macrocyclase, AcyG macrocyclase, BisG macrocyclase, and TrfG macrocyclase. In some aspects, the G macrocyclase can be PatG, PagG, a fragment thereof, or a variant thereof. In some aspects, the PatG macrocyclase sequence can be GenBank AAY21156.1. In some aspects, the PagG macrocyclase sequence can be GenBank AED99446. 1.

[0059] In some aspects, the amino acid sequence of the G macrocyclase or G macrocyclase recognition site and fragments thereof disclosed herein can include a peptide sequence that has some degree of identity or homology to any of sequences of the peptides disclosed herein. The degree of identity can vary and be determined by methods known to one of ordinary skill in the art. The terms “homology7'’ and "identity" each refer to sequence similarity between two polypeptide sequences. Homology and identity can each be determined by comparing a position in each sequence which can be aligned for purposes of comparison. When a position in the compared sequence is occupied by the same amino acid residue, then the polypeptides can be referred to as identical at that position; when the equivalent site is occupied by the same amino acid (e.g., identical) or a similar amino acid (e.g., similar in steric and / or electronic nature), then the molecules can be referred to as homologous at that position. A percentage of homology or identity between sequences is a function of the number of matching or homologous positions shared by the sequences. The peptides described herein can have at least or about 25%, 50%, 65%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity or homology to any of the G macrocyclases or G macrocyclase recognition sites disclosed herein.

[0060] Numerous variants of the G macrocyclase described herein are known and herein contemplated. Protein and peptide fragments, variants and derivatives are well understood to those of skill in the art and in can involve amino acid sequence modifications. For example, amino acid sequence modifications typically fall into one or more of three classes: substitutional, insertional or deletional variants. Insertions include amino and / or carboxyl terminal fusions as well as intrasequence insertions of single or multiple amino acid residues. Insertions ordinarily will be smaller insertions than those of amino or carboxyl terminal fusions, for example, on the order of one to four residues. Deletions are characterized by the removal of one or more amino acid residues from the peptide sequence. Typically, no more than about from 2 to 6 residues are deleted at any one site within the peptide. Amino acid substitutions are typically of single residues, but can occur at a number of different locations at once; insertions usually will be on the order of about from 1 to 10 amino acid residues; and deletions will range about from 1 to 30 residues. Deletions or insertions preferably are made in adjacent pairs, i.e., a deletion of 2 residues or insertion of 2 residues. Substitutions, deletions, insertions or any combination thereof may be combined to arrive at a final construct. Substitutional variants are those in which at least one residue has been removed and a different residue inserted in its place. Such substitutions generally are made in accordance with the following Tables 1 and 2 and are referred to as conservative substitutions.

[0061] Table 1: Amino Acid Abbreviations

[0062] Table 2: Amino Acid Substitutions

[0063] Substantial changes in function or immunological identity are made by selecting substitutions that are less conservative than those in Table 2, i.e., selecting residues that differ more significantly in their effect on maintaining (a) the structure of the polypeptide backbone in the area of the substitution, for example as a sheet or helical conformation, (b) the charge or hydrophobicity of the molecule at the target site or (c) the bulk of the side chain. The substitutions which in general are expected to produce the greatest changes in the protein properties will be those in which (a) a hydrophilic residue, e.g. seryl or threonyl, is substituted for (or by) a hydrophobic residue, e.g., leucyl, isoleucyl, phenylalanyl, valyl or alanyl; (b) a cysteine or proline is substituted for (or by) any other residue; (c) a residue having an electropositive side chain, e.g., lysyl, arginyl, or histidyl, is substituted for (or by) an electronegative residue, e g., glutamyl or aspartyl; or (d) a residue having a bulky side chain, e.g., phenylalanine, is substituted for (or by) one not having a side chain, e.g., glycine, in this case, (e) by increasing the number of sites for sulfation and / or glycosylation.

[0064] For example, the replacement of one amino acid residue with another that is biologically and / or chemically similar is known to those skilled in the art as a conservative substitution. For example, a conservative substitution would be replacing one hydrophobic residue for another or one polar residue for another. The substitutions include combinations such as, for example, Gly, Ala; Val, lie, Leu; Asp, Glu; Asn, Gin; Ser, Thr; Lys, Arg; and Phe, Tyr. Such conservatively substituted variations of each explicitly disclosed sequence are included within the mosaic polypeptides provided herein.

[0065] Substitutional or deletional mutagenesis can be employed to insert sites for N- glycosylation (Asn-X-Thr / Ser) or O-glycosylation (Ser or Thr). Deletions of cysteine or other labile residues also may be desirable. Deletions or substitutions of potential proteolysis sites (e.g., Arg), are accomplished for example by deleting one of the basic residues or substituting one by glutaminyl or histidyl residues.

[0066] Amino acid analogs and analogs and peptide analogs often have enhanced or desirable properties, such as. more economical production, greater chemical stability, enhanced pharmacological properties (half-life, absorption, potency, efficacy, etc.), altered specificity (e.g., a broad-spectrum of biological activities), reduced antigenicity, and others.

[0067] D-amino acids can be used to generate more stable peptides, because D amino acids are not recognized by peptidases and such. Systematic substitution of one or more amino acids of a consensus sequence with a D-amino acid of the same type (e.g., D-lysine in place of L-lysine) can be used to generate more stable peptides. Cysteine residues can be used to cyclize or attach two or more peptides together. This can be beneficial to constrain peptides into particular conformations. (Rizo and Gierasch Ann. Rev. Biochem. 61 :387 (1992), incorporated herein by reference).

[0068] The degree of identity can vary' and can be determined by methods well established in the art. “Homology” and “identity” each refer to sequence similarity between two polypeptide sequences, with identity being a stricter comparison. Homology and identity can each be determined by comparing a position in each sequence which may be aligned for purposes of comparison. When a position in the compared sequence is occupied by the same amino acid residue, then the polypeptides can be referred to as identical at that position; when the equivalent site is occupied by the- same amino acid (e.g., identical) or a similar amino acid (e.g., similar in steric and / or electronic nature), then the molecules can be referred to as homologous at that position. A percentage of homology or identity between sequences is a function of the number of matching or homologous positions shared by the sequences. A biologically active variant or a fragment of a peptide or polypeptide described herein can have at least or about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% identity or homology to a corresponding naturally occurring peptide or polypeptide.

[0069] In some aspects, the peptides disclosed herein can further comprise a detectable label. The detectable label can be any molecule, atom, ion or group which is detectable in vivo by a molecular imaging modality. Suitable detectable labels can include metals, radioactive isotopes and radio-opaque agents (e.g., gallium, technetium, indium, strontium, iodine, barium, bromine and phosphorus-containing compounds), radiolucent agents, contrast agents and fluorescent dyes. In some aspects, the detectable label can be FLAG™ tag, epitope or protein tags, such as myc tag, 6 His, and fluorescent fusion protein. In some aspects, the detectable label can be selected based on the molecular imaging modality being used. Examples of molecular imaging modalities include but are not limited to radiography, fluoroscopy, fluorescence imaging, high resolution ultrasound imaging, bioluminescence imaging, Magnetic Resonance Imaging (MRI), and nuclear imaging, for example scintigraphic techniques such as Positron Emission Tomography (PET) and Single Photon Emission Computerized Tomography (SPECT). In some aspects, the detectable label can be an imaging agent. Examples of imaging agents include, but are not limited to radionuclides, fluorescent dyes, chemiluminescent agents, colorimetric labels, and magnetic labels. In some aspects, the imaging agent can include a radiolabel that can be detected using gamma imaging wherein emitted gamma irradiation of the appropriate wavelength is detected. Methods of gamma imaging include, but are not limited to, SPECT and PET. For SPECT detection, the chosen radiolabel can lack a particular emission, but can produce a large number of photons in, for example, a 140-200 keV range. For PET detection, the radiolabel can be a positron-emitting moiety, such as 19F.

[0070] In some aspects, the detectable label can include an MRS / MRI radiolabel, including but not limited to gadolinium, 19F, 13C, that can be coupled (e.g., attached or complexed) with the composition using general organic chemistry techniques. The detectable label can also include radiolabels, such as 18F, 11 C, 75Br, or 76Br for PET by techniques well known in the art and are described by Fowler, J. and Wolf, A. in Positron Emission Tomography and Autoradiography (Phelps, M., Mazziota, J., and Schelbert, H. eds.) 391-450 (Raven Press, NY 1986) the content of which is hereby incorporated by reference. The imaging can also include 1231 for SPECT.

[0071] In some aspects, the detectable label can further include metal radiolabels. In some aspects, the radiolabel can be Technetium-99m (99mTc). Preparing radiolabeled derivatives of Tc99m is well known in the art. See, for example, Zhuang et al., “Neutral and stereospecific Tc-99m complexes: [99mTc]N-benzyl-3.4-di-(N-2-mercaptoethyl)-amino- pyrrolidines (P-BAT)” Nuclear Medicine & Biology 26(2):217-24, (1999); Oya et al., “Small and neutral Tc(v)O BAT, bisaminoethanethiol (N2S2) complexes for developing new brain imaging agents”, Nuclear Medicine & Biology 25(2): 135-40, (1998); and Hom et al., “Technetium-99m-labeled receptor-specific small-molecule radiopharmaceuticals: recent developments and encouraging results” Nuclear Medicine & Biology 24(6):485-98, (1997).

[0072] Suitable fluorescence detectable labels include but are not limited to fluorescein, phycoery thrin, Europium, TruRed, Allophycocyanin (APC), PerCP, Lissamine, Rhodamine, B X-Rhodamine, TRITC, BODIPY-FL, FluorX, Red 613, R-Phycoerythrin (PE), NBD, Lucifer yellow, Cascade Blue. Methoxy coumarin. Aminocoumarin, Texas Red, Hydroxy coumarin, Alexa Fluor™ dyes (Molecular Probes) such as Alexa FluorTM 350, Alexa Fluor™ 488, Alexa Fluor™ 546, Alexa Fluor™ 568, Alexa Fluor™ 633, Alexa Fluor™ 647, Alexa Fluor™ 660 and Alexa Fluor™ 700, sulfonate cyanine dyes (AP Biotech), such as Cy2, Cy3, Cy3.5, Cy5, Cy5.5 and Cy7, IRD41 IRD700 (Li-Cox, Inc.), NIR-1 (Dejindom. Japan). La Jolla Blue (Diatron), DyLight™ 405. 488, 549. 633, 649, 680 and 800 Reactive Dyes (Pierce / Thermo Fisher Scientific Inc) or LI-CORTM dyes, such as IRDyeTM (LI-CORTM Biosciences). Other suitable fluorescent detectable labels include but are not limited to lanthanide ions, such as terbium and europium. Lanthanide ions can be attached to the synaptotagmin polypeptide by means of chelates. Other suitable fluorescent detectable labels include but are not limited to quantum dots (e.g. QdotTM, Invitrogen).

[0073] Magnetic resonance image-based techniques create images based on the relative relaxation rates of water protons in unique chemical environments. Magnetic resonance imaging can include conventional magnetic resonance imaging (MRI), magnetization transfer imaging (MTI), magnetic resonance spectroscopy (MRS), diffusion-weighted imaging (DWI) and functional MR imaging (fMRI). Labels suitable for use as magnetic resonance imaging (MRI) labels can include but are not limited to paramagnetic or superparamagnetic ions, iron oxide particles, and water-soluble contrast agents. Superparamagnetic and paramagnetic ions can include transition, lanthanide and actinide elements such as iron, copper, manganese, chromium, erbium, europium, dysprosium, holmium and gadolinium. In some aspects, the paramagnetic detectable label can be gadolinium.

[0074] Any of the cyclic peptides disclosed herein or produced by the methods described herein can be attached to an antibody molecule, such as an antibody or antibody fragment or derivative, for example for use in antibody -directed drug therapies. Suitable techniques for the conjugation of cyclic peptides and antibodies are well known in the art.

[0075] Also disclosed herein are vectors comprising nucleic acid sequences capable of encoding one or more of the peptides disclosed herein. In some aspects, the vectors can comprise: a) a first nucleic acid capable of encoding a peptide comprising from N terminus to C terminus: a peptide of interest, wherein the peptide of interest having 6 to 20 amino acid residues and a C-terminal cysteine residue; a G macrocyclase recognition site; a linker; and G macrocyclase; and b) a second nucleic acid capable of encoding RSI-TruD cyclodehydratase. Also disclosed herein are vectors that comprise: a first nucleic acid capable of encoding a peptide comprising from N terminus to C terminus: a peptide of interest, wherein the peptide of interest having 6 to 20 amino acid residues and a C-terminal cysteine residue; a G macrocyclase recognition site; a linker; and G macrocyclase.

[0076] In some aspects, the vectors can comprise: a) a first nucleic acid capable of encoding a peptide comprising from N terminus to C terminus: a peptide of interest, wherein the peptide of interest having 6 to 20 amino acid residues and a C-terminal residue that facilitates cyclization, wherein the C-terminal residue that facilitates cyclization is a cysteine, a proline, an oxazoline. or a heterocycle; a G macrocyclase recognition site; a linker; and G macrocyclase; and b) a second nucleic acid capable of encoding RSI-TruD cyclodehydratase. Also disclosed herein are vectors that comprise: a first nucleic acid capable of encoding a peptide comprising from N terminus to C terminus: a peptide of interest, wherein the peptide of interest having 6 to 20 amino acid residues and a C-terminal residue that facilitates cyclization, wherein the C-terminal residue that facilitates cyclization is a cysteine, a proline, an oxazoline. or a heterocycle; a G macrocyclase recognition site; a linker; and G macrocyclase.

[0077] In some aspects, the first nucleic acid is operably linked to a first regulatory sequence (e.g., promoter) and the second nucleic acids is operably linked to a second regulatory sequence (e.g., promoter). In some aspects, the first nucleic acid and the second nucleic acid are operably linked to the same regulatory sequence. In some aspects, the first nucleic acid can be under the control of a separate regulatory sequence than the second nucleic acid. In some aspects, the first nucleic acid can be under the control of the same regulatory sequence as the second nucleic acid.

[0078] G-macrocyclases, oxidases, heterocyclases, proteases and other enzymes can be generated wholly or partly by recombinant techniques. For example, a nucleic acid encoding the enzyme can be expressed in a host cell and the expressed polypeptide isolated and / or purified from the cell culture. In some aspects, any of the enzymes disclosed herein can be expressed from nucleic acid which has been codon optimized for expression in E. coli.

[0079] Nucleic acid sequences and constructs as described herein can be comprised within an expression vector. Suitable vectors can be chosen or constructed, containing appropriate regulatory sequences, including promoter sequences, terminator fragments, polyadenylation sequences, enhancer sequences, marker genes and other sequences as appropriate. In some aspects, the vector contains appropriate regulatory sequences to drive the expression of the nucleic acid in a host cell. Suitable regulatory sequences to drive the expression of heterologous nucleic acid coding sequences in expression systems can include promoters including, but not limited to, constitutive promoters, for example viral promoters (e.g., CMV or SV40), and inducible promoters (e.g.. Tet-on controlled promoters). In some aspects, the vectors can also comprise sequences, such as origins of replication and selectable markers, which allow for its selection and replication and expression in bacterial hosts such as E coli and / or in eukaryotic cells.

[0080] METHODS

[0081] Disclosed herein, are methods of producing a cyclic peptide. In some aspects, the methods can comprise: a) fusing a linear peptide comprising from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C -terminal cysteine residue, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzyme; and b) contacting the linear peptide in a) with RSI-TruD cyclodehydratase, thereby generating the cyclic peptide. In some aspects, the peptide of interest can be covalently bonded to the G macrocyclase recognition site. In some aspects, the G macrocyclase recognition site can be bonded to the linker. In some aspects, the linker can be bonded to the G macrocyclase.

[0082] Disclosed herein, are methods of producing a cyclic peptide. In some aspects, the methods can comprise: a) fusing a linear peptide comprising from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C -terminal cysteine residue, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzyme; and b) optionally contacting the linear peptide in a) with RSI-TruD cyclodehydratase, thereby generating the cyclic peptide. In some aspects, the peptide of interest can be covalently bonded to the G macrocyclase recognition site. In some aspects, the G macrocyclase recognition site can be bonded to the linker. In some aspects, the linker can be bonded to the G macrocyclase.

[0083] Disclosed herein, are methods of producing a cyclic peptide. In some aspects, the methods can comprise: a) fusing a linear peptide comprising from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal residue that facilitates cyclization, wherein the C-terminal residue that facilitates cyclization is a cysteine, a proline, an oxazoline, or a heterocycle, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzy me; and b) optionally contacting the linear peptide in a) with RSI-TruD cyclodehydratase, thereby generating the cyclic peptide. In some aspects, the peptide of interest can be covalently bonded to the G macrocyclase recognition site. In some aspects, the G macrocyclase recognition site can be bonded to the linker. In some aspects, the linker can be bonded to the G macrocyclase.

[0084] Disclosed herein are methods of producing a cyclic peptide, the methods comprising: a) fusing a linear peptide comprising from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal cysteine residue, ii) a G macrocyclase recognition site, iii) a linker, iv) a G macrocyclase enzyme; wherein the peptide of interest is covalently bonded to the G macrocyclase recognition site, wherein the G macrocyclase recognition site is bond to the linker, and wherein the linker is bonded to the G macrocyclase; and b) contacting the linear peptide in a) with an enzy me or under suitable conditions to allow the linear peptide to cyclize, thereby generating the cyclic peptide. In some aspects, the peptide of interest can be covalently bonded to the G macrocyclase recognition site. In some aspects, the G macrocyclase recognition site can be bonded to the linker.

[0085] Disclosed herein are methods of producing a cyclic peptide, the methods comprising: a) fusing a linear peptide comprising from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal residue that facilitates cyclization, ii) a G macrocyclase recognition site, iii) a linker, iv) a G macrocyclase enzyme; wherein the peptide of interest is covalently bonded to the G macrocyclase recognition site, wherein the G macrocyclase recognition site is bond to the linker, and wherein the linker is bonded to the G macrocyclase; and b) optionally contacting the linear peptide in a) with an enzyme or under suitable conditions to allow the linear peptide to cyclize, thereby generating the cyclic peptide. In some aspects, the C-terminal residue that facilitates cyclization can be a cysteine, a proline, an oxazoline, or a heterocycle. In some aspects, the oxazoline can be derived from a serine or a threonine. In some aspects, the heterocycle can be a triazole, an imidazole, or a pseudoproline.

[0086] In some aspects, the cyclic peptide can be labeled with a detectable label.

[0087] In some aspects, the linker can be GSG3, GSGe, GSG9, or GSG12.

[0088] In some aspects, the G macrocyclase can be PatG, PagG, a fragment thereof, or a variant thereof.

[0089] Further disclosed herein are method of making a cyclized peptide. In some aspects, the methods can comprise: a) providing a vector comprising: a first nucleic acid capable of encoding a linear peptide; and a second nucleic acid capable of encoding RSI-TruD cyclodehydratase; b) transforming the vector in cells; c) culturing the transformed cells in b) under conditions to allow the vector to express a first construct and a second construct, wherein the first construct comprises the linear peptide, and the second construct comprises RSI-TruD cyclodehydratase, thereby generating one or more cyclized peptides; and d) isolating the one or more cyclized peptides. In some aspects, the linear peptide comprises from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal cysteine residue, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzyme. In some aspects, the linker can be GSG3, GSGe, GSG9, or GSG12. In some aspects, the G macrocyclase can be PatG, PagG, a fragment thereof, or a variant thereof.

[0090] Further disclosed herein are method of making a cyclized peptide. In some aspects, the methods can comprise: a) providing a vector comprising: a first nucleic acid capable of encoding a linear peptide; b) transforming the vector in cells; c) culturing the transformed cells in b) under conditions to allow the vector to express a first construct, wherein the first construct comprises the linear peptide, contacting the linear peptide with an enzyme or under suitable conditions to allow the linear peptide to cyclize, thereby generating the cyclic peptide, thereby generating one or more cyclized peptides; and d) isolating the one or more cyclized peptides. In some aspects, the linear peptide comprises fromN terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal cysteine residue, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzy me. In some aspects, the linear peptide comprises from N terminus to C terminus: i) a peptide of interest, wherein the C-terminal residue that facilitates cyclization is a cysteine, a proline, an oxazoline, or a heterocycle, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzyme. In some aspects, the linker can be GSG3, GSGe, GSG9, or GSG12. In some aspects, the G macrocyclase can be PatG, PagG, a fragment thereof, or a variant thereof.

[0091] Also disclosed herein are methods of generating a library comprising two or more cyclized peptides, the method comprising: a) providing two or more vectors, wherein each vector comprises: i) a first nucleic acid capable of encoding a linear peptide; and ii) a second nucleic acid capable of encoding RSI-TruD cyclodehydratase, wherein each of the linear peptides is a different peptide; b) transforming the two or vectors in cells; and c) culturing the transformed cells in b) under conditions to allow the two or more vectors to express to a first construct and a second construct, wherein the first construct comprises the different peptide, and the second construct comprises RSI-TruD cyclodehydratase, thereby generating two or more cyclized peptides; and thereby generating the library. In some aspects, the linear peptide comprises from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal cysteine residue, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzyme. In some aspects, the linear peptide comprises from N terminus to C terminus: i) a peptide of interest, wherein the C-terminal residue that facilitates cyclization is a cysteine, a proline, an oxazoline, or a heterocycle, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzyme. In some aspects, the linker can be GSG3, GSGe, GSG9, or GSG12. In some aspects, the G macrocyclase can be PatG, PagG, a fragment thereof, or a variant thereof. In some aspects, wherein each of the linear peptides is a different peptide, and the peptide of interest of the linear peptide can be different. In some aspects, the first nucleic acid capable of encoding the peptide capable of encoding a linear peptide and the second nucleic acid capable of encoding RSI-TruD cyclodehydratase can be in the same vector. In some aspects, the first nucleic acid capable of encoding the peptide capable of encoding a linear peptide and the second nucleic acid capable of encoding RSI-TruD cyclodehydratase can be in separate vectors.

[0092] Also disclosed herein are methods of generating a library comprising two or more cyclized peptides, the method comprising: a) providing two or more vectors, wherein each vector comprises: i) a first nucleic acid capable of encoding a linear peptide; wherein each of the linear peptides is a different peptide; b) transforming the two or vectors in cells; and c) culturing the transformed cells in b) under conditions to allow the two or more vectors to express to a first construct, and contacting the linear peptides with an enzyme or under suitable conditions to allow the linear peptide to cyclize, thereby generating two or more cyclized peptides; and thereby generating the library. In some aspects, the linear peptide comprises from N terminus to C terminus: i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal cysteine residue, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzyme. In some aspects, the linear peptide comprises from N terminus to C terminus: i) a peptide of interest, wherein the C-terminal residue that facilitates cyclization is a cysteine, a proline, an oxazoline, or a heterocycle, ii) a G macrocyclase recognition site, iii) a linker, and iv) a G macrocyclase enzyme. In some aspects, the linker can be GSG3, GSGe, GSG9, or GSG12. In some aspects, the G macrocyclase can be PatG, PagG, a fragment thereof, or a variant thereof. In some aspects, wherein each of the linear peptides is a different peptide, and the peptide of interest of the linear peptide can be different. In some aspects, the first nucleic acid capable of encoding the peptide capable of encoding a linear peptide and the second nucleic acid capable of encoding RSI-TruD cyclodehydratase can be in the same vector. In some aspects, the first nucleic acid capable of encoding the peptide capable of encoding a linear peptide and the second nucleic acid capable of encoding RSI-TruD cyclodehydratase can be in separate vectors.

[0093] EXAMPLES

[0094] Example 1: An autocatalytic peptide cyclase improves fidelity and yield of cyclized peptides in vivo and in vitro.

[0095] Peptide cyclization improves conformational ri gi di ty, providing favorable pharmacological properties such as proteolytic resistance, target specificity, and membrane permeability. Thus, many synthetic and biosynthetic peptide circularization strategies have been developed. PatG and related macrocyclases process diverse peptide sequences, generating millions of cyclic derivatives. However, the application of these cyclases is limited by low yields and the potential presence of unwanted intermediates. Disclosed herein are covalently fused G macrocyclase with substrates designed to efficiently and spontaneously release cyclic peptides. To increase the fidelity' of synthesis, an orthogonal control mechanism was developed enabling precision synthesis in Escherichia coli. As a result, a library comprising 4.8 million cyclic derivatives was constructed, producing an estimated 2.6 million distinct cyclic peptides with improved yield and fidelity.

[0096] It was tested whether covalent fusion of the substrate to the N-terminus of a G protein would significantly increase efficiency, yield, and fidelity. Further, developing orthogonal control of cyclization activity provides several advantages to in vivo and in vitro peptide cyclization. Disclosed herein are the design and implementation of the covalent fusion strategy', leading to an efficient synthesis of cyclic peptides in vivo and in vitro. A mutant library' with a high cyclization success rate was generated. Also, a large-scale culture was performed and one of the successful mutants was isolated, yielding enough material for full spectral analysis.

[0097] Development of autocatalytic G macrocyclase. PagG was used because its substrates do not require enzymatic modification prior to the cyclization step (Sarkar, S., et al. ACS Catal. 2020, 10 (13), 7146-7153). The precursor peptide, PagE (Donia, M. S. and Schmidt, E. W. Chem. Biol. 2011, 18 (4), 508-519). was fused through its C-terminus to different linkers, which in turn were fused with the PagG protease domain. The native core peptide (QAYGIPLP (SEQ ID NO: 1)) and C-terminal recognition sequence (“RSIII”) were used, but the N-terminal PagA protease recognition sequence was replaced with the PatA element (“RSIF’) since the PatA enzyme has been much more widely used (Sardar, D. and Schmidt, E. W. Curr. Opin. Chem. Biol. 2016, 31, 15—21). Cleavage of RSII would liberate a free N- terminus ready for cyclization (FIG. 1 A). Thus, in principle, after PatA action autocatalysis would afford the macrocycle. The cyclic peptide is the native PagG product and the precursor of the natural product, prenylagaramide from cyanobacteria (Murakami, M., et al. J. Nat. Prod. 1999, 62 (5). 752-755).

[0098] Constructs with variable GlySer linkers (GSGn n = 3 - 12 and GGGGSn n = 2 - 8) that range between 3-40 amino acids (Chen, X., et al. Adv. Drug Deliv. Rev. 2013, 65 (10), 1357— 1369; and Reddy Chichili, et al. Protein Sci. 2013, 22 (2), 153-167) were tested. Based upon available crystal structures (Agarwal, V., et al. Chem. Biol. 2012. 19 (11), 1411-1422; and Koehnke, J., et al. Nat. Struct. Mol. Biol. 2012, 19 (8), 767-772). the linker size ranges easily spanned the distance between the substrate and the active site of the enzymes. Eight enzy me constructs with variable linkers were expressed in A’ coll and purified. To confirm that the PagG enzyme constructs retained activity, an in-trans enzyme assay was performed using the PagG native substrate: a short peptide consisting of QAYGIPLP (SEQ ID NO: 1) fused to RSIII. The cyclic peptide was produced by each of the PagG constructs, indicating that the additional N-terminal sequences in the protein do not interfere to its catalytic activity. For example, an in-trans assay was performed to determine whether the additional PatA elements affect PagG activity. Truncated substrates containing the prenylagramide core (QAYLGIPLP; SEQ ID NO: 15) and the recognition sequence FAGDDAE (SEQ ID NO: 143) were tested, showing the engineered constructs still undergo cyclization.

[0099] The purified proteins were next treated with PatA protease, which cleaves the leader peptide and frees the N-terminus of the core peptide. In principle, this should lead to rapid cyclization and the release of cyclic products. In the event, it was found that cyclized peptides were detected in experiments using constructs with GSG linkers, but not those with GGGGS (SEQ ID NO: 13) linkers. Cyclization was most efficient with a 36 amino acid linker comprising 12 GSG repeats (SEQ ID NO: 30) (FIG. IB). Therefore, this linker length was used in further experiments.

[0100] To efficiently produce the cyclic products in vivo without any additional PatA or related protein, leader sequences were removed through mutagenesis, leaving a free N- terminus primed for immediate cyclization. This PagG protein construct, dubbed auto-PagG. was expressed in E. coli, and products were extracted from the resulting cell pellet using acetone. Abundant cyclic peptides were observed by UPLC-MS (FIG. 2). Interestingly, while the desired product cyclo[QAYGIPLP (SEQ ID NO: 1] was present in the mixture, by far the major product consisted of the cyclic peptide containing additional Met, cyclo[MQAYGIPLP (SEQ ID NO: 2)] . and the derivative with a spontaneously oxidized Met sulfoxide. The incorporation of Met in the final product suggests a competing reaction between the methionyl endopeptidase and PagG macrocyclization, indicating an unexpectedly rapid cyclization kinetics. This increased rate is analogous to what has been observed in some engineered protein kinases, in which phosphorylation is dramatically enhanced when the substrate is covalently linked to the enzyme, and the enhancement strongly depends on the linker length (Dyla, M. and Kjaergaard, M. Proc. Natl. Acad. Sci. 2020, 117 (35), 21413- 21419).

[0101] Applying RSI-TruD as an orthogonal control for autocyclization. A possible way to prevent unwanted Met incorporation would be to include PatA and the PatA RSII sequence in vivo. A disadvantage is that PatA releases the linear N-terminal sequence, which is potentially a complicating factor. To solve this problem, the PatG enzyme was developed. PatG native substrates contain Cys (instead of Pro) at the C-terminus; Cys is not a substrate for cyclization, but instead it must be heterocyclized to thiazoline by the action of D enzymes. Among these, TruD has been engineered to be constitutively active by fusing the enzyme N-terminus with a leader peptide, producing RSI-TruD (Donia, M. S. and Schmidt, E. W. Chem. Biol. 2011. 18 (4), 508-519; Sardar. D.. et al. J. Am. Chem. Soc. 2017, 139 (8). 2884-2887; and Koehnke, J., et al. Nat. Chem. Biol. 2015, 11 (8), 558-563), facilitating orthogonal control of enzyme action, in which a macrocycle would be observed in the presence of active RSI-TruD.

[0102] Therefore, in analogy to PagG. PatG variants were synthesized containing native PatG core peptides, the PatG RSIII, and the GSG12 sequence linked to the PatG protease domain (C -terminal His tag) (FIG. 3A). This construct was dubbed “auto-PatG”. Core peptide sequences used included the precursors to the natural products patellin 2 (TVPTLC; SEQ ID NO: 3) and patellin 3 (TLPVPTLC; SEQ ID NO: 4), as well as a mutant of the patellin 2 sequence (TVPTVC; SEQ ID NO: 5). The RSIII sequence was truncated to SYD. The three auto-PatG protein variants were expressed in E. coli and purified. The purified RSI-TruD enzyme was added, and the reaction mixtures were analyzed by UPLC-MS. No reaction products were observed unless RSI-TruD was included, demonstrating the need for thiazoline synthesis prior to cyclization (FIG. 3A). Efficient cyclic peptide formation was observed when RSI-TruD and ATP were included in the reaction mixture (FIG. 3B). Following these in vitro experiments, the same experiments were performed in E. coli, in which the auto-PatG and RSI-TruD constructs were simultaneously expressed from a single plasmid. The cell pellets were isolated, and their acetone extracts contained the expected cyclic peptides. For example, extracted ion chromatogram (EIC) and mass spectra of dimeric cyclic patellin 2 (tR = 8.00 min) and patellin 3 (tR = 6.98 min) were obtained from acetone extract of E. coli cultures containing the engineered pRSF-Duetl plasmid.

[0103] Interestingly, the patellin 2 core peptide led to two products, one of which was twice the expected mass of the TVPTLC (SEQ ID NO: 3) core. This dimer was observed in both in vitro and in vivo expression experiments. The presence of the dimeric patellin 2 derivative was highly sequence-dependent since it was not observed in a Leu-Val mutant TVPTVC (SEQ ID NO: 5), nor was it observed in the octapeptide patellin 3 product (FIG. 3B). In addition, dimers were not observed in larger libraries (see below). To investigate how these dimers are formed, the molar ratio of the auto-PatG and RSI-TruD vaned from 80: 1 to 1: 1. respectively (FIG. 6). Noticeably, as both enzy mes reach equal concentration, dimer product increased indicating that formation could be dependent upon the concentration of activated G protein available for cyclization. This implies that the product was synthesized by crossreaction between two auto-PatG constructs and not by more complex possibilities involving already-synthesized macrocycles. To test this latter idea, preformed macrocycles were added to enzyme reaction mixtures, but they had no effect on the dimer: monomer ratio. Moreover, since the dimer formation was restricted to the patellin 2 sequence, mechanistic assessments (e.g., by mixing enzy mes and observing unsymmetrical dimer formation) could not be carried out.

[0104] An efficient cyclic peptide library. The cyclic peptide library' was constructed in the duet vector encoding RSI-TruD and auto-PatG (FIG. 4). an octapeptide library was synthesized in which the N-terminal seven amino acids in the core peptide were randomized. The eighth amino acid, Cys was left intact to direct cyclization. The library was prepared using trimer phosphorami dites encoding Ala, Vai, He, Leu, Phe, Ser, Trp, Tyr, and Thr; these were selected because other amino acids exhibited greater positional dependence (Ruffner. D. E., et al. ACS Synth. Biol. 2015. 4 (4), 482-492) The resulting library of 79(4.8 x 106) theoretical derivatives was transformed into E. coli, and 100 individual colonies were picked randomly and sequenced. One of these anomalously contained Lys and was thus discounted in further experiments. The remaining 99 plasmids had the correct elements in place, achieving a relatively even incorporation of the codons: Ala (67). Vai (66), He (81), Leu (77), Phe (73), Ser (77), Trp (84), Tyr (64), and Thr (69).

[0105] The colonies were used in expression experiments at 25 mL scale. After 3 days of incubation, cultures were pelleted, extracted with acetone, and enriched by passage over a solid phase extraction resin. Analysis of the extract by UPLC-MS revealed that of the 99 mutants considered, 55 produced the desired cyclic peptide. Five of these were duplicated sequences, so 51 out of 94 peptide sequences were successfully turned over into products. This analysis was robust because every chromatogram served as a negative control for every’ other chromatogram analyzed. Some extracts contained two distinct, isobaric peaks, demonstrating spontaneous epimerization of the C-H protons adjacent to the thiazoline double bond

[0106] Every' amino acid appeared multiple times in each of the seven mutated positions. In comparison to the natural product sequence TLPVPTLC (patellin 3; SEQ ID NO: 4), highly different sequences such as YSVFWAVC (SEQ ID NO: 8). TWASIHC (SEQ ID NO: 9), SAYTLSLC (SEQ ID NO: 10), TTTLIFVC (SEQ ID NO: 11), and VVLSSYIC (SEQ ID NO: 12) were efficiently processed (Tables 6 and 7). Amino acids were not evenly accepted, with overall acceptance rates between 45-76% (FIG. 7). Thr is the preferred amino acid (76%). In contrast, aromatic amino acids are notably less accepted, with specific positions being less so. For example, Trp has an average acceptance rate of 45%, with the lowest at Pl at 27% (FIG. 7). Similarly, Phe has an average of 49% but the lowest acceptance rate at P7 at 19%. The apparent yield of the cyclic peptides varies depending on the core sequences (Table 6). Because of the methodology used, the area under the curve depends upon the ionizability of the individual peptide and, therefore is semiquantitative. Nonetheless, overall the analysis revealed that many cyclic peptides are noticeably abundant in the 25 mL cultures.

[0107] To confirm the abundance and identity of cyclic peptides, a plasmid encoding a sequence TAYWTWIC (SEQ ID NO: 7) was selected whose product would be readily isolatable due to aromatic residues that would make the product obvious even as a minor component in rich media The proteins were expressed at 12 L scale for 3 days. After extraction, the predicted peptide was clearly observed in the total ion chromatogram (not just in the mass-filtered chromatogram), indicating abundant production. After purification, cyclo| TAYWTWIC (SEQ ID NO: 7)] was obtained (4.9 mg), and its structure was confirmed using a full spectroscopic analysis (FIG. 8). This peptide was used as a standard to determine that the crude extract contained 17 mg / L before purification (FIG. 9).

[0108] The results described herein show a simple, cyanobactin-based method to synthesize cyclic peptides and cyclic peptide libraries in vivo and in vitro. The method uses a single plasmid and no special conditions for fermentation, in contrast to previously reported methods (Tianero, Ma. D., et al. Proc. Natl. Acad. Sci. 2016, 113 (7), 1772-1777). Finally, the process is compatible with other post-translational modifications, enabling combinatorial RiPP biosynthesis using multiple enzymes.

[0109] There are several other genetic methods to produce cyclic peptides and cyclic peptide libraries, dominated by intein-mediated cyclization and RiPP-based cyclization as found in broad-substrate pathw ays such as some of those to lanthipeptides, thiopeptides, cyclotides, and others (Knerr, P. J. and Van Der Donk, W. A. Annu. Rev. Biochem. 2012, 81 (1), 479- 505; Chan, D. C. K. and Burrows, L. L. J. Antibiot. (Toky o) 2021, 74 (3), 161-175; and De Veer, S. J., et al. Chem. Rev. 2019, 119 (24), 12375-12421). Previous studies took engineering approaches to improve the yield and scope of RiPP modifications. In these cases, RiPP enzy mes from cyanobactins, microviridins, and lanthipeptides were covalently attached to their leader peptides, allowing efficient leaderless processing (Koehnke, J., et al. Nat. Chem. Biol. 2015, 11 (8), 558-563; Reyna-Gonzalez. E., et al. Angew. Chem. 2016. 128 (32), 9544-9547; and Oman, T. J., et al. J. Am. Chem. Soc. 2012, 134 (16), 6952-6955). Instead of this approach, the G enzyme was redesigned to directly link to its substrate, which then independently released cyclic peptides. By doing so, the universal issues of substrate solubility, concentration, proteolytic stability, and so on are minimized. The methods disclosed herein can be applied to other ribosomal pathways since many operate under the same biosynthetic logic, allowing the RiPPs enzymes to be pliably engineered according to biotechnological needs.

[0110] Compared to other RiPP cyclases that usually make multiple products (both linear and cyclic) under the artificial conditions used to create libraries, PatG and TruG are thought to produce almost solely cyclized peptides (McIntosh, J. A., et al. J. Am. Chem. Soc. 2010, 132 (44), 15499-15501; and Nguyen, N. A., et al. ACS Chem. Biol. 2022, 17 (6). 1577-1585). This work demonstrates that G family enzymes are indeed exceptionally flexible in substrate tolerance, while simultaneously exhibiting high fidelity for RSIII requirement and for synthesizing one cyclized product. However, the enzy me itself certainly has preferences for specific core peptides in the natural setting. As evidence of this, perhaps the reasonable way a dimer of patellin 2 could form is if the patellin 2 core peptide itself has an especially stable interaction with the enzyme active site.

[0111] As disclosed herein, the engineered enzy mes described herein remove several practical limitations, enabling compatibility with numerous synthetic biology approaches to designed organisms or to in vitro synthetic methods. The results described herein demonstrate several of these applications by showing the advantages of the system in vivo and in vitro, such as the generation of high-quality cyclic peptide libraries that are compatible with cellbased functional screens. The results also demonstrated that the process is robust enough to produce enough material needed for applications such as pharmacological evaluation. Thus, the methods disclosed herein are a practical approach that is robust and predictably produces compounds and libraries at a practical scale.

[0112] Materials. DNA polymerases, restriction enzy mes, and other reagents for cloning were obtained from New England BioLabs Inc. Gene fragments for Gibson assemblies and mutagenesis library were purchased from Genewiz.

[0113] Cloning. To construct a plasmid that expresses the substrate-fused PagG, the PagG gene was recovered from an existing plasmid and it was linearized using PCR. Gene fragments containing leader peptide, recognition sequence for the PatA N-terminal proteases (RSII), prenylagaramide core, and recognition sequences for the PagG macrocyclase fused with tw o different linker types (GSG and GGGGS (SEQ ID NO: 13) in varying sizes were ordered from Genewiz. These two pieces were ligated together using Gibson assembly. The resulting plasmids were checked for successful ligation through Sanger sequencing.

[0114] To build a new plasmid containing the core sequences and the macrocyclase (auto- PagG), the initial plasmid was used as a template for PCR amplification using mutagenic primers that exclude the RSII and leader sequences. The linearized fragment was then recyclized for 1 hr at room temperature by adding T4 polynucleotide kinase and T4 DNA ligase. The ligation mixture was transformed into DH I () / > E. coll cells and grown in an LB + kanamycin plate overnight Colonies were picked and screened using Sanger sequencing (FIG. 5A).

[0115] Due to the results of using the auto-PagG plasmid, a new plasmid was designed containing the core sequences fused covalently with another G enzyme called PatG (auto- PatG). The plasmid containing the PatG gene was linearized using designed primers, followed by Gibson assembly of gene fragments containing core sequences, RSIII, and the [GSG]n=12 linker. Generated plasmids were confirmed through whole plasmid or Sanger sequencing.

[0116] To produce cyclic peptides containing thiazoline ring in E. coll in vivo, a plasmid was designed containing RSI-TruD and auto-PatG. The gene encoding the engineered leader peptide-TruD heterocyclase (RSI-TruD) was obtained from an existing plasmid. The RSI- TruD gene and an empty pRSFDuet-1 vector were digested with Ndel and Xhol, followed by gel purification and overnight ligation with T4 DNA ligase at 18 °C. The reaction mixture was transformed into DH I () / > E. coll competent cells and grown in an LB + kanamycin plate. Isolated colonies were picked in 2XYT liquid media and grown overnight. Plasmids were extracted and screened through digestion and Sanger sequencing.

[0117] Once the pRSFDuet-1 plasmid containing the RSI-TruD gene at 2ndmultiple cloning site (MCS2) was confirmed, the auto-PatG was inserted at MCS1. The generated pRSF plasmid and the auto-PatG gene were double-digested using Ncol and Notl restriction enzymes, followed by T4 DNA ligation overnight. Ligation using the Ncol restriction site leads to added undesired amino acids in the core sequence. Mutagenic primers were designed to remove unwanted codons, followed by plasmid linearization through PCR. The linearized plasmid was then ligated back by adding T4 polynucleotide kinase and T4 DNA ligase. The resulting mixture was transformed in DHIOP E. coli cells, which were then used to obtain colonies. Plasmids were extracted from picked colonies and then submitted for Plasmidsaurus for whole plasmid sequencing. To replace the sequence of the core of auto-PatG, different strategies, such as the use of mutagenic primers (TVPTLC (SEQ ID NO: 3) to TVPTVC (SEQ ID NO: 5)) or new Gibson fragment (TVPTLC (SEQ ID NO: 3) to TLPVPTLC (SEQ ID NO: 4)), were used (FIG. 5B).

[0118] Protein expression and purification. Expression plasmids were transformed in chemically competent BL21(DE3) E. coli cells. A single colony was inoculated in LB media (10 mL) containing antibiotic kanamycin (50 pg / mL) and grown overnight at 30 °C, 180 rpm. These starter cultures were added to Fembach flasks (2 L) containing 2XYT media (1 L) w ith kanamycin (50 pg / mL). The cultures were shaken at 35 °C and 200 rpm until it reached an ODeoo= 0.4-0.8, and 1 mM of P-D-l thiogalactopyranoside (IPTG) was added. Once IPTG was added, the temperature was lowered to 18 °C and incubated overnight. The cells were harvested through centrifuging at 3739 g for 30 min and stored at -80 °C. Cells were then thawed and suspended in lysis buffer (10 mM imidazole, 200 mM NaCl, 50 mM Tris, 5% glycerol, pH = 8.0) and stirred for 1 h with lysozyme (600 pg / mL) at 4 °C. After incubation, cells were sonicated using 30 s pulses, followed by the addition of deoxyribonuclease I (20 pg / mL) and 50 mM MgCh. Lysed cells were centrifuged at 28,928 g for 40 min at 4 °C. The supernatant was fdtered with 5 and 0.45 pm filters, and the resulting lysate was subjected to a pre-equilibrated Ni-NTA gravity column. Lysates were incubated with nickel resin at 4 °C for 2 h and washed with 100 mL wash buffer (1 M NaCl, 30 mM imidazole). Bound proteins were eluted with 20 mL elution buffer (1 M NaCl, 250 mM imidazole, pH 8.0) and analyzed using SDS-PAGE. Proteins were dialyzed and concentrated using a 30 kDa centrifugal filter with exchange buffer (50 mM Tris, 10% glycerol. pH = 5.5). Concentrated proteins were aliquoted in 100 pL volumes, flash-frozen, and stored at -80 °C until used for enzyme assays.

[0119] Enzyme assays. Assays using substrate-linked PagG enzymes were conducted at 25 °C in 10 mM MES buffer, pH = 7.0. An in-trans assay was performed to determine whether these engineered enzymes were active. The reaction mixture contained the substrate QAYLGIPLPFAGDDAE (SEQ ID NO: 14) (100 pM), DMSO (10%), MgCl2(10 mM), and substrate-fused PagG constructs (50 pM). Analyzing the effect of linker size on these PagG constructs, enzyme reactions with substrate-fused PagG constructs (100 pM), PatA protease (50 pM), MgCh(10 mM) in MES buffer (10 mM, pH = 7.0) were used. The reaction mixtures were incubated for 36 h and stopped by boiling for 1 min. Product formation was measured through liquid chromatography-mass spectrometry (LC-MS).

[0120] Auto-PatG (substrate-fused PatG) enzyme assays w ere conducted in Tris buffer (50 mM, pH = 7.5) at 37 °C containing auto-PatG (100 pM), RSI-TruD (5 pM), ATP (ImM) and MgCh (5 mM). The reaction mixtures were incubated for 24 h, followed by quenching the reaction by boiling for 1 min. Cyclic peptides were monitored through LC-MS. To interrogate the dimer formation, the molar ratio between auto-PatG and RSI-TruD was varied from 80: 1 to 1: 1 using the same reaction conditions: MgCh (5 mM), ATP (ImM) and Tris buffer (50 mM, pH = 7.5) at 37 °C.

[0121] Fermentation and extraction. Expression plasmids containing substrate-fused macrocyclases were transformed in chemically competent BL21 (DE3) E. coli cells. A single colony was transferred to LB (10 mL) supplemented with kanamycin (50 pg / mL). The cultures were grown overnight at 30 °C, with continuous shaking at 180 rpm. These starter cultures were introduced in 2.8 L Fembach flasks containing 2XYT media (1 L) with kanamycin (50 pg / mL). The flasks were agitated at 35 °C and 200 rpm until they reached an ODeoo of 0.4-0.8. Then, P-D-l thiogalactopyranoside (IPTG; 1 mM) was added. Cultures containing the auto-PagG plasmid were incubated for 6 days at 18 °C, while the pRSFDuet-1 plasmid containing the auto-PatG and RSI-TruD were grown at 30 °C for 5 days. After incubation, cells were harvested at 3739 g for 40 min and extracted with acetone. Cells were sonicated for 20 min, followed by centrifugation at 3739 g for 20 min. Extraction was repeated 3x, and the supernatant was dried in vacuo.

[0122] Mutagenesis library. Using the designed pRSFDuet-1 plasmid as a template, a mutagenesis library (Genewiz) was synthesized where the first 7 amino acids of the core sequence of patellin 3 (TLPVPTLC; SEQ ID NO: 4) were substituted with 9 proteinogenic amino acids in equal probabilities, namely Ala, Vai, He. Leu, Phe, Ser, Trp, Tyr, Thr. The resulting library was transformed into BL21 (DE3) E. coli cells. Single colonies were transferred in 24-well plates containing LB media (6 mL) with kanamycin (50 pg / mL). Plates were incubated overnight, and plasmids were harvested. Sequences of each colony were determined by designing a forward primer near but outside the N-terminus of the auto-PatG construct, which was then used for Sanger sequencing. Simultaneously, these colonies were also inoculated in 2XYT media (25 mL) with kanamycin (50 pg / mL) and IPTG (1 mM). These cultures were incubated for 3 days, 30 °C at 180 rpm. At the end of incubation, cells were centrifuged at 10 °C, 3739 g, and extracted with acetone. Organic extracts were dried and reconstituted with 100% MeOH. HP20 Diaion resin was added and lyophilized for 1 h. Resins were dried and washed with H2O to eliminate salts, followed by elution with 100% MeOH. 100% MeOH fractions were run in UPLC-MS to determine the presence of cyclic peptides.

[0123] LC-MS analysis. Aliquots (10 pL) of reaction mixtures from enzyme assay s / extracts were diluted in MS grade MeOH (90 pL). The diluted solutions (2 pL) were loaded in Acquity UPLC BEH C 18 (130 A, 1.7 pM. 2.1 mm x 50 mm) with a 0.4 mL / min flow rate. A gradient of 5% ACN (2 min), 5-100% ACN (8 min), 100% ACN (2.5 min) and 100-5% ACN (2.5 min) were used. The masses of the cyclic peptides at different adducts were searched within the mass range m / z 100-2000 Da. In the mutagenesis library, extract containing the EIC peak with the accurate mass of the cyclic product was considered as a hit.

[0124] Isolation of cyanobactin mutant from E. coli expressions. Cultures (12 L) of BL21 (DE3) E. coli containing a pRSF-plasmid with the core sequence TAYWTWIC (SEQ ID NO: 7) were grown in kanamycin (50 pg / mL) and IPTG (1 mM), followed by incubation for 3 days at 30 °C. At the end of the incubation period, cultures were centrifuged and extracted repeatedly with acetone. The resulting organic extracts were fractionated using reverse phase Cl 8 resin with elution of 30% MeOHiEEO and 100% MeOH. The 100% MeOH fraction was subjected to high performance liquid chromatography (HPLC) using Luna, 250 x 21.2 mM, 5 pM phenyl-hexyl preparative column on a gradient 50% ACN (5 min), 50% - 80% ACN (15 min), 80% (5 min) and 80%-50% (2.5 min) and 50% (2.5 min). The collected peak at tR= 14.5 min was further purified with a semi-preparative Luna, 250 x 10 mm, 5 pM phenylhexyl column using the solvent gradient: 47.5% ACN (5 min), 47.5%-52.5% ACN (10 min). 52.5% ACN (2.5 min) and 52.5%-47.5% ACN (2.5 min) to afford the cyclic TAYWTWIC (SEQ ID NO: 7) (7, tR= 8.3 min, 4.9 mgs).

[0125] Cyclic TAYWTWIC (SEQ ID NO: 7) (7): pale white solid; UV (MeOH): 221, 279 nm; 'H and13C NMR (Table 7); HRESIMS m / z 1007.4529 [M+H]+(calculated for C51H63N10O10S ,1007.4514, 5PPm = 6.54); m / z 1029.4253 [M+Na]+(calculated for C5iH62NioNaOioS ,1029.4263, 6ppm= -0.97).

[0126] Quantification of cyanobactin derivative from E. coli expressions using UPLC-MS. Cultures of BL21 (DE3) E. coli (2 L) containing a pRSF-plasmid with the core sequence TAYWTWIC (SEQ ID NO: 7) were grown out and extracted using similar conditions described. Organic extracts were diluted in methanol for a final concentration of 1.56 x 10'2mg / mL. Different known concentrations of the isolated compound were used as internal standard and spiked in the prepared methanolic extracts. The resulting solutions' aliquots (1 pL) were loaded in Acquity UPLC BEH C18 (130 A, 1.7 pM, 2. 1 mm x 50 mm) with a flow rate of 0.3 mL / min. A gradient of 5% ACN (2 min), 5-100% ACN (8 min), 100% ACN (2.5 min) and 100-5% ACN (2.5 min) were used. This was done in three independent replicates. Table 3. List of products from substrate-fused macrocyclases.

[0127] Table 4. Gene fragments used to generate substrate-fused macrocyclases.

[0128]

[0129] Table 5. List of proteins used in this study.

[0130] Table 6. Sequences of octapeptides tested in mutagenesis library'. Mass spectral analysis of octapeptide mutagenesis library. These represent extracted ion chromatograms (EIC) of the cyclic peptides. The amino acid composition of each core is determined through Sanger sequencing. The generated sequence was used to predict the mass of the cyclic product.

[0131] Table 7. NMR spectroscopic data for cyclo [TAYWTWIC] (SEQ ID NO: 7) (7) IN DMF-d7at 25 °C.

[0132] It will be apparent to those skilled in the art that various modifications and variations can be made in the present invention without departing from the scope or spirit of the invention. Other aspects of the invention will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the invention being indicated by the following claims.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A peptide comprising from N terminus to C terminus: a) a peptide of interest, wherein the peptide of interest having 6 to 20 amino acid residues and a C-terminal residue that facilitates cyclization, wherein the C- terminal residue that facilitates cyclization is a cysteine, a proline, an oxazoline, or a heterocycle; b) a G macrocyclase recognition site; c) a linker; and d) G macrocyclase; wherein the peptide is linear, wherein the peptide of interest is covalently bonded to the G macrocyclase recognition site, wherein the G macrocyclase recognition site is bonded to the linker, and wherein the linker is bonded to the G macrocyclase.

2. The peptide of claim 1, wherein the G macrocyclase recognition site consists of SYD, AYD, FAGDDAE (SEQ ID NO: 143), SYE, AYE, SFD, SVG, or AFD.

3. The peptide of claim 1, wherein the linker is GSG3, GSGe, GSG9, or GSG12.

4. The peptide of claim 1, wherein the G macrocyclase is PatG, PagG, TruG, ThcG, OscG, TenG, LynG, McaG, AgeG, TriG, AcyG, BisG, TrfG, a fragment thereof, or a variant thereof.

5. The peptide of claim 1, further comprising a detectable label.

6. A method of producing a cyclic peptide, the method comprising: a) fusing a linear peptide comprising from N terminus to C terminus:(i) a peptide of interest, wherein the peptide of interest has 6 to 20 amino acid residues and a C-terminal residue that facilitates cyclization, wherein the C-terminal residue that facilitates cyclization is a cysteine, a proline, an oxazoline, or a heterocycle,(ii) a G macrocyclase recognition site,(iii) a linker, and(iv) a G macrocyclase enzyme; wherein the peptide of interest is covalently bonded to the G macrocyclase recognition site, wherein the G macrocyclase recognition site is bonded to the linker, and wherein the linker is bonded to the G macrocyclase; and b) optionally contacting the linear peptide in a) with RSI-TruD cyclodehydratase, thereby generating the cyclic peptide.

7. The method of claim 6, wherein the cyclic peptide is labeled with a detectable label.

8. The method of claim 6, wherein the linker is GSG3, GSGe, GSG9, or GSG12.

9. The method of claim 6, wherein the G macrocyclase is PatG, PagG, a fragment thereof, or a variant thereof.

10. A vector comprising: a) a first nucleic acid capable of encoding the peptide of claim 1; and b) a second nucleic acid capable of encoding RSI-TruD cyclodehydratase.

11. A method of making a cyclized peptide, the method comprising: a) providing the vector of claim 10; b) transforming the vector in cells; c) culturing the transformed cells in b) under conditions to allow the vector to express to a first construct and a second construct, wherein the first constructcomprises the peptide of claim 1, and the second construct comprises RSI-TruD cyclodehydratase, thereby generating one or more cyclized peptides; and d) isolating the one or more cyclized peptides.

12. A method of generating a library comprising two or more cyclized peptides, the method comprising: a) providing two or more vectors, wherein each vector comprises:(i) a first nucleic acid capable of encoding the peptide of claim 1; and(ii) a second nucleic acid capable of encoding RSI-TruD cyclodehydratase, wherein each of the peptides of claim 1 is a different peptide; b) transforming the two or vectors in cells; and c) culturing the transformed cells in b) under conditions to allow the two or more vectors to express to a first construct and a second construct, wherein the first construct comprises the different peptide, and the second construct comprises RSI-TruD cyclodehydratase, thereby generating two or more cyclized peptides; and thereby generating the library.