Integrated and iterative methods for designing, identifying and optimizing lasso peptides
An integrated workflow optimizes lasso peptides through in silico modeling and experimental testing, addressing solubility and stability issues in peptide-based therapeutics, enhancing their clinical applicability.
Patent Information
- Application Number
- PCT/US2024/061377
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-20
- Publication Date
- 2025-08-07
AI Technical Summary
Existing peptide-based therapeutic compounds face challenges due to undesirable physicochemical and pharmacokinetic properties such as poor solubility, cell permeability, low bioavailability, and instability under physiological conditions, limiting their clinical use.
An integrated and iterative workflow involving structure-based design principles, in silico modeling, epitope grafting, virtual screening, genetic engineering, and experimental testing is used to optimize lasso peptides for improved binding selectivity and affinity, using methods like molecular dynamic simulations and epitope mapping to identify optimal conformational states and binding interactions.
The method enables the identification and optimization of lasso peptides with enhanced properties, facilitating rapid development of novel drugs by improving predictive accuracy and therapeutic potential.
Smart Images

Figure US2024061377_07082025_PF_FP_ABST
Abstract
Description
Attorney Docket No.: 200301-018002 / PCT INTEGRATED AND ITERATIVE METHODS FOR DESIGNING, IDENTIFYING AND OPTIMIZING LASSO PEPTIDES 1. CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 612,963, filed December 20, 2023, the entire contents of which is incorporated herein by reference. 2. FIELD
[0002] Provided herein are methods for discovering and optimizing lasso peptides using an integrated and iterative workflow involving structure-based design principles, in silico modeling, epitope grafting, virtual screening, genetic engineering, peptide evolution, and experimental testing. Also provided herein are lasso peptides that have been discovered and optimized for binding selectivity, binding affinity, biological activity, and other properties using the methods described herein. 3. BACKGROUND
[0003] Peptides serve as useful tools and leads for drug development since they often combine high affinity and specificity for their target receptor with low toxicity (Henninot, A., et al., J. Med Chem., 2018, 61, 1382-1414). However, the clinical use of peptides as efficacious drugs has been limited due to undesirable physicochemical and pharmacokinetic properties, including poor solubility and cell permeability, low bioavailability, and instability due to rapid proteolytic degradation under physiological conditions. Computational and structure-based methods have been developed for designing improved therapeutic peptides (Diller, D.J., et al., Future Med Chem, 2015, 7, 2173-2193; Ciemny, M., et al., Drug Discovery Today, 2018, 23, 1530-1537; Kruger, D.M., et al., J. Med. Chem., 2017, 60, 8982- 8988), and artificial intelligence and machine learning algorithms are showing promise for de novo discovery of peptide-based drugs (Guigere, S., et al., PLoS Comput Biol, 2015, 11(4): e1004074. doi:10.1371 / journal.pcbi.1004074; Tallorin, L., et al., Nature Comm., 2018, 9:5253, doi.org / 10.1038 / s41467-018-07717-6; Muller, A.T., et al., J. Chem. Inf. Model., 2018, 58, 2, 472-479). Unfortunately, accurately predicting peptide-protein interactions remains a challenge for linear and macrocyclic peptides due to the conformational flexibility of these systems. ACTIVE 705331286v1
[0004] There exists a need for new classes of peptide-based therapeutic compounds with improved properties and readily available methods for their identification and optimization. In particular, there exists the need for conformationally constrained peptide scaffolds that can enhance the predictive nature of in silico methods, and, which combined with evolution methods and high throughput screening technologies, can serve as a basis for rapid development of novel drugs for treating difficult diseases. The present disclosure provided herein meets these needs. 4. SUMMARY
[0005] Provided herein are methods for the design, identification, and optimization of lasso peptides. Also provided herein are methods for the iterative virtual and experimental screening of lasso peptides and lasso peptide libraries in order to identify candidates having desirable properties. Also provided herein are lasso peptides having desirable properties that have been optimized through integrated and iterative use of computational modeling, genetic evolution, and experimental testing methods.
[0006] Particularly, in one aspect of the present disclosure, provided herein is a computer-based method for identifying a lasso peptide having optimal binding with a target molecule. Such a method can include: (a) providing one or more three-dimensional (3D) model structures of lasso peptides; (b) performing conformational analysis on the one or more 3D model structures of lasso peptides using molecular dynamic simulation algorithms to obtain at least two conformational states of the one or more 3D model structures of lasso peptides; (c) docking the at least two conformational states of the one or more 3D model structures of lasso peptides onto a 3D model structure of the target molecule at a lasso-binding site of the target molecule to form at least two 3D model structures of the lasso peptide-target molecule complex; (d) performing conformational analysis on the at least two 3D model structures of the lasso peptide-target molecule complex using molecular dynamic simulation algorithms to obtain at least two conformational states of the 3D model structure of the lasso peptide-target molecule complex; and (e) selecting the 3D model structure of the lasso peptide-target molecule complex having a favored free energy, thereby identifying a lasso peptide binder candidate.
[0007] In some embodiments, the method provided herein further includes step (a-1): mapping a target-binding epitope onto different locations of the one or more 3D model structures of lasso peptides. In some embodiments, the method provided herein further 2 ACTIVE 705331286v1includes computationally grafting the target-binding epitope onto a mapped location of the selected lasso peptide, thereby generating a grafted lasso peptide binder candidate.
[0008] In some embodiments, for a method provided herein, the at least two conformational states of the one or more 3D model structures of lasso peptides are selected from 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 45 or 50 or more different conformational states.
[0009] In some embodiments, for a method provided herein, the favored free energy comprises a lower free energy compared to a different 3D model structure of the lasso peptide-target molecule complex.
[0010] In some embodiments, for a method provided herein, the favored free energy is the lowest free energy of the two or more 3D model structure of the lasso peptide-target molecule complex.
[0011] In some embodiments, for a method provided herein, step (a) further comprises retrieving one or more known 3D structures of lasso peptides from a protein structure database.
[0012] In some embodiments, for a method provided herein, step (a) further comprises computationally modeling the 3D structure of a lasso peptide based on atomic coordinates of the lasso peptide.
[0013] In some embodiments, for a method provided herein, the atomic coordinates of the lasso peptides are obtained from a protein structure database or scientific literature. In some embodiments, the protein structure database is selected from worldwide Protein Data Bank (wwPDB), Cambridge Structure Database, Molecular Model Database of National Center for Biotechnology Information (NCBI), and Biological Magnetic Resonance Data Bank (BMRB) database. In some embodiments, the atomic coordinates of the lasso peptide are obtained by subjecting the lasso peptide to nuclear magnetic resonance (NMR) analysis, X-ray crystallography, neutron diffraction, or 3-dimensional electron microscopy (3D-EM). In some embodiments, the X-ray crystallography is serial femtosecond crystallography. In some embodiments, the 3D-EM is cryogenic electron microscopy (cryo-EM).
[0014] In some embodiments, for a method provided herein, step (a) further comprises computationally modeling the 3D structure of the lasso peptide based on X-ray diffraction data and / or nuclear magnetic resonance (NMR) data of the lasso peptide and atomic coordinates of a reference lasso peptide, and wherein the lasso peptide has at least 50% amino acid sequence identity to the reference lasso peptide. In some embodiments, the X-ray diffraction data of the lasso peptide are obtained by subjecting a crystal of the lasso peptide to 3 ACTIVE 705331286v1X-ray crystallography analysis. In some embodiments, step (a) comprises: (i) obtaining atomic coordinates of the lasso peptide based on the X-ray diffraction data; and (ii) refining the atomic coordinates of the lasso peptide based on the atomic coordinates of the reference lasso peptide.
[0015] In some embodiments, for a method provided herein, the nuclear magnetic resonance (NMR) data of the lasso peptide comprises NMR chemical shift, J-coupling constant, and resonance intensity obtained by subjecting a solution of the lasso peptide to NMR analysis. In some embodiments, step (a) comprises: (i) creating an ensemble of structural models of the lasso peptide based on the NMR data; (ii) obtaining mean atomic coordinates of the lasso peptide based on the ensemble of structural models; and (iii) refining the atomic coordinates of the lasso peptide based on the atomic coordinates of the reference lasso peptide.
[0016] In some embodiments, for a method provided herein, step (a) comprises: computationally modeling the 3D structure of a lasso peptide based on the amino acid sequence of the lasso peptide and atomic coordinates of a reference lasso peptide; and wherein the lasso peptide has at least 50 percent (%) amino acid sequence identity to the reference lasso peptide. In some embodiments, computationally modeling the 3D structure of the lasso peptide is performed by homology modeling. In some embodiments, homology modeling is performed in combination with a protein structure prediction algorithm, wherein optionally the protein structure prediction algorithm is trRosetta or AlphaFold or AlphaFold 2.
[0017] In some embodiments, for a method provided herein, the one or more 3D model structures of lasso peptides are lasso backbone structures, and wherein step (a) further comprises computationally modeling the lasso backbone structures by removing side chains of each amino acid residue that is not an internal ring-forming residue from the 3D structure of lasso peptides.
[0018] In some embodiments, for a method provided herein, step (c) further comprises creating the 3D model structure of the target molecule before docking. In some embodiments, creating the model 3D structure of the target molecule comprises computationally modeling the 3D structure of the target molecule based on atomic coordinates of the target molecule.
[0019] In some embodiments, for a method provided herein, creating the 3D model structure of the target molecule comprises computationally modeling the 3D structure of the target molecule based on X-ray diffraction data and / or nuclear magnetic resonance data of the 4 ACTIVE 705331286v1target molecule and atomic coordinates of a reference polypeptide; wherein the reference polypeptide has at least 50% amino acid sequence identity to the reference polypeptide. In some embodiments, the X-ray diffraction data of the target molecule are obtained by subjecting a crystal of the target molecule to X-ray crystallography analysis. In some embodiments, creating the 3D model structure of the target molecule comprises: (i) obtaining atomic coordinates of the target molecule based on the X-ray diffraction data; (ii) refining the atomic coordinates of the target molecule based on the atomic coordinates of the reference polypeptide; and (iii) computationally modeling the 3D structure of the target molecule based on the refined atomic coordinates. In some embodiments, the nuclear magnetic resonance (NMR) data of the target molecule comprises NMR chemical shift, J-coupling constant, and resonance intensity obtained by subjecting a solution of the target molecule to NMR analysis. In some embodiments, creating the 3D model structure of the target molecule comprises: (i) creating an ensemble of structural models of the target molecule based on the NMR data; (ii) obtaining mean atomic coordinates of the target molecule based on the ensemble of structural models; (iii) refining the mean atomic coordinates of the target molecule based on the atomic coordinates of the reference polypeptide; and (iv) computationally modeling the 3D structure of the target molecule based on the refined mean atomic coordinates.
[0020] In some embodiments, for a method provided herein, creating the 3D model structure of the target molecule comprises computationally modeling the 3D structure of the target molecule based on the amino acid sequence of the target molecule and atomic coordinates of one or more reference polypeptide; and wherein the amino acid sequence of the target molecule is at least about 50% identical to the amino acid sequence of the reference polypeptide. In some embodiments, for a method provided herein, computationally modeling the 3D structure of the target molecule is performed using homology modeling. In some embodiments, for a method provided herein, homology modeling is performed in combination with a protein structure prediction algorithm, wherein optionally the protein structure prediction algorithm is trRosetta or AlphaFold or AlphaFold 2.
[0021] In some embodiments, for a method provided herein, the lasso-binding site is selected from a ligand binding site, a substrate binding site, a catalytic binding site, a co- factor binding site, a precursor binding site, an orthosteric binding site, an allosteric binding site, an open conformation binding site, an active conformation binding site, an inactive conformation binding site, and a closed conformation binding site of the target molecule.
[0022] In some embodiments, for a method provided herein, step (c) further comprises, for each 3D model structure of the lasso peptide: (c-1) positioning a docked portion of at least 5 ACTIVE 705331286v1one conformational state of the 3D model structure of the lasso peptide relative to the lasso- binding site of the 3D model structure of the target molecule in a first docked pose, thereby forming a docked system; (c-2) scoring the docked system based on structural complementarity between the docked portion and the lasso-binding site; (c-3) repositioning the docked portion relative to the lasso-binding site to a second docked pose, and repeating step (c-2); (c-4) repeating step (c-3) for one or more times; and (c-5) selecting the docked system having the highest score as the optimal docked system before proceeding to step (d). In some embodiments, the docked system comprises one or more complementary binding pairs; and wherein the complementary binding pair comprises a first binding moiety on the lasso peptide and a second binding moiety on the target molecule. In some embodiments, the first binding moiety is on a first amino acid residue of the lasso peptide, and the second binding moiety is on a second amino acid residue of the target molecule; and wherein the distance between any atom of the first amino acid residue and any atom of the second amino acid residue is less than about 5 Ångströms, and optionally less that about 4 Ångströms, and optionally less than about 3 Ångströms. In some embodiments, the complementary binding pair forms binding interaction that is selected from a hydrogen bonding interaction, an ionic or electrostatic bonding interaction, polar interaction, dipolar interaction, induced dipolar interaction, pi stacking interaction, hydrophobic interaction, and / or van der Waals interaction. In some embodiments, step (c-2) comprises calculating a total binding free energy of the docked system using one or more molecular mechanics force field functions; and assigning a score to the docked system based on the total binding free energy, and wherein the score relates to the total binding free energy. In some embodiments, the molecular mechanics force field functions are selected from Amber force fields AMBER99, Amber 10EHT, ff14SB, ff19SB, and the Merck molecular force field MMFF94x. In some embodiments, step (c-3) comprises identifying the second docked pose using an energy minimizing function before repositioning the docked portion into the second docked pose; wherein the docked system is predicted to have a lower binding free energy in the second docked pose than the first docked pose based on the energy minimizing function. In some embodiments, the energy minimizing function is selected from the steepest descent algorithm, conjugate gradients algorithm, L- BFGS (limited-memory Broyden-Fletcher-Goldfarb-Shanno) algorithm, and genetic algorithms. In some embodiments, the docked portion comprises at least one amino acid residue from the ring portion, the loop portion, and / or the tail portion of the lasso peptide.
[0023] In some embodiments, for a method provided herein, step (a-1) comprises: (a-1-1) in the optimal docked system, identifying one or more lasso-binding moieties in the lasso- 6 ACTIVE 705331286v1binding site of the target molecule; (a-1-2) selecting one or more amino acid residues comprising one or more target-binding moieties complementary to the one or more lasso- binding moieties; and (a-1-3) mapping an optimal set of positions in the amino acid sequence of the selected lasso peptide for grafting the one or more selected amino acid residues; wherein the grafting places the one or more target-binding moieties at suitable spatial locations and orientations for binding with the complementary lasso-binding moieties on the target molecule. In some embodiments, step (a-1-3) comprises: (a-1-3-1) computationally grafting the one or more selected amino acid residues into the amino acid sequence of the selected lasso peptide at a first set of positions; (a-1-3-2) calculating a total binding free energy of the optimal docked system using one or more molecular force field functions; (a-1- 3-3) modifying at least one position in the first set of positions thereby obtaining an adjusted set of positions, and computationally grafting the one or more selected amino acid residues into the amino acid sequence of the selected lasso peptide at the adjusted set of positions; (a- 1-3-4) repeating step (a-1-3-2); (a-1-3-5) repeating steps (a-1-3-3) and (a-1-3-4) sequentially for one or more times; and (a-1-3-6) selecting the adjusted set of positions associated with the lowest total binding free energy as the optimal map of positions. In some embodiments, the molecular mechanics force field functions are selected from Amber force fields AMBER99, Amber 10EHT, ff14SB, ff19SB, and the Merck molecular force field MMFF94x. In some embodiments, step (a-1-3-3) comprises identifying the adjusted set of positions using an energy minimizing function before modifying the first set of positions, wherein the docked system having the one or more selected amino acid residues grafted into the amino acid sequence of the selected lasso peptide at the adjusted set of positions is predicted to have a lower binding free energy than at the first set of positions based on the energy minimizing function. In some embodiments, the energy minimizing function is selected from the steepest descent algorithm, conjugate gradients algorithm, L-BFGS (limited-memory Broyden- Fletcher-Goldfarb-Shanno) algorithm, and genetic algorithms.
[0024] In some embodiments, for a method provided herein, the target-binding epitope corresponds to a fragment or fragments of a naturally-existing ligand of the target molecule, wherein the fragment or fragments are capable of binding with the lasso-binding site of the target molecule and forming a ligand-target interface. In some embodiments, step (a-1) comprises aligning the docked portion of the 3D model structure of the lasso peptide with a 3D model structure of the fragment or fragments in the ligand-target interface. In some embodiments, step (a-1) comprises adjusting the spatial position, conformation and / or 7 ACTIVE 705331286v1orientation of at least one target-binding moiety in the docked portion to mimic a corresponding binding moiety of the fragment or fragments in the ligand-target interface.
[0025] In some embodiments, for a method provided herein, the method further comprises: (f) mutating one or more amino acid residues of the lasso peptide binder candidate to produce a first set of lasso peptide binder variants; and (g) ranking the first set of lasso peptide binder variants based on a predicted binding affinity for binding with the target molecule. In some embodiments, step (f) further comprises adjusting conformation of the first set of lasso peptide binder variants to produce a second set of lasso peptide binder variants; and wherein step (g) comprises ranking the second set of lasso peptide binder variants based on the predicted binding affinity for binding with the target molecule. In some embodiments, in step (f), mutating the amino acid residue of the lasso peptide binder candidate comprises replacing the side chain of the amino acid residue of the lasso peptide binder candidate with the side chain of a second amino acid that is different from the mutated amino acid residue. In some embodiments, the second amino acid is a naturally-occurring or a non-natural amino acid. In some embodiments, in step (f), mutating the amino acid residue of the lasso peptide binder candidate comprises modifying one or more chemical moieties on the side chain of the mutated amino acid residue. In some embodiments, at least one modified chemical moiety is a target-binding moiety.
[0026] In some embodiments, for a method provided herein, in step (f), mutating the amino acid residue of the lasso peptide binder candidate comprises modifying one or more side chains of the lasso peptide binder candidate to complement one or more lasso-binding moieties on the target molecule; and wherein the modifying is selected from the group consisting of: (i) incorporating into the lasso peptide a neutral or basic side chain that is structurally opposite to an acidic lasso-binding moiety; (ii) incorporating into the lasso peptide a neutral or acid side chain that is structurally opposite to a basic lasso-binding moiety; (iii) incorporating into the lasso peptide a neutral or positively charged side chain that is structurally opposite to a negatively charged lasso-binding moiety; (iv) incorporating into the lasso peptide a neutral or negatively charged side chain that is structurally opposite to a positively charged lasso-binding moiety; (v) incorporating into the lasso peptide a sterically smaller side chain that is structurally opposite to a sterically larger lasso-binding moiety; (vi) incorporating into the lasso peptide a sterically larger side chain that is structurally opposite to a sterically smaller lasso-binding moiety; (vii) incorporating into the lasso peptide a hydrophobic side chain, preferably a similarly hydrophobic side chain, that is structurally opposite to a hydrophobic lasso-binding moiety; (viii) incorporating into the lasso peptide an 8 ACTIVE 705331286v1acidic, H-bond donor, positively charged, or aromatic side chain that is structurally opposite to an aromatic lasso-binding moiety; (ix) incorporating into the lasso peptide an electron-rich aromatic side chain that is structurally opposite to an electron-poor aromatic lasso-binding moiety; (x) incorporating into the lasso peptide an acid or electron-poor aromatic side chain that is structurally opposite to an electron-rich aromatic lasso-binding moiety; (xi) incorporating into the lasso peptide an inducible dipole or multipole side chain that is structurally opposite to an inducible dipole or multipole lasso-binding moiety; (xii) incorporating into the lasso peptide a permanent, inducible dipole or multipole side chain that is structurally opposite to an inducible or multipole lasso-binding moiety; (xiii) incorporating into the lasso peptide an H-bond acceptor side chain that is structurally opposite to an H-bond donor lasso-binding moiety, and (xiv) incorporating into the lasso peptide an H-bond donor side chain that is structurally opposite to an H-bond acceptor lasso-binding moiety.
[0027] In some embodiments, for a method provided herein, the mutated amino acid residue is in the: (i) ring portion of the lasso peptide; (ii) loop portion of the lasso peptide; (iii) tail portion of the lasso peptide; or (iv) any combination of (i) to (iii).
[0028] In some embodiments, for a method provided herein, the method further comprises: (h) synthesizing the lasso peptide binder candidate, grafted lasso peptide binder candidate or one or more lasso peptide binder variants having the highest rankings. In some embodiments, in step (h), synthesizing the lasso peptide binder candidate or the one or more lasso peptide binder variants is performed using a cell-free or cell-based synthesis method.
[0029] In some embodiments, for a method provided herein, the method further comprises creating a database of optimized lasso peptide structures or intermediates thereof generated in any one of steps (a) to (h). 5. BRIEF DESCRIPTION OF THE FIGURES
[0030] FIG.1 is a schematic illustration of the of a lasso peptide with the characteristic lasso (lariat) topology containing a loop, ring and tail. Amino acids are shown as balls.
[0031] FIG.2A is a schematic illustration of an iterative workflow including the steps generally referred to as “design,” “build,” “test,” and “evolve” according to the present disclosure, which workflow enables the creation of lasso peptide as therapeutic outputs from digital inputs through the integration of rational in silico peptide design and evolution methods.
[0032] FIG.2B is a schematic illustration of the workflow for computationally identifying lasso peptides having optimal binding with a target polypeptide. 9 ACTIVE 705331286v1
[0033] FIG.3A illustrates an example of grafting of a 5-amino acid linear epitope from the chemokine CCL5 into the ring portion of a lasso peptide structure.
[0034] FIG.3B illustrates examples of epitope grafting into the loop, ring and tail portions of a lasso peptide, respectively.
[0035] FIG.4 illustrates an example showing grafting a conformational or 3-dimensional (3D) epitope containing discontinuous segments of amino acid residues from a natural ligand (e.g., hormones, growth factors, chemokines, cytokines) into a lasso peptide structure.
[0036] FIG.5 is a schematic illustration of cell-free and cell-based methods for producing lasso peptides as part of an iterative workflow involving in silico modeling and experimental testing to validate and optimize biological activities and other properties of lasso peptides according to the present disclosure.
[0037] FIG.6 illustrates computationally modeled structures of a lasso peptide binding in a GPCR pocket from in silico modeling showing “tail down, loop up” (left) and “loop down, tail up” (right) binding configurations. This figure exemplifies an initial modeling result from a study aimed at determining the size and shape fit between 37 known lasso peptide structures and the receptor pocket.
[0038] FIG.7A shows a homology model of C-C chemokine receptor type 8 (CCR8) docked with its natural ligand C-C Motif Chemokine Ligand 1 (CCL1).
[0039] FIG.7B shows a homology model of C-C chemokine receptor type 4 (CCR4) docked with its natural ligand C-C Motif Chemokine Ligand 17 (CCL17).
[0040] FIG.8 shows computational models of all-Gly backbone structures for lasso peptides having SEQ ID NOS: 1 to 37, based on atomic coordinates obtained from the PDB database. SID stands for SEQ ID.
[0041] FIGS.9A and 9B illustrate two lasso peptides that are mapped onto C-C Motif Chemokine Ligand 17 (CCL17) in the binding pocket of C-C chemokine receptor 4 (CCR4) using in silico modeling. Particularly, FIG.9A shows a lasso peptide having SEQ ID NO:21 mapped onto the 30s loop of CCL17; and FIG.9B shows a lasso peptide having SEQ ID NO:2 mapped onto the N-terminus of CCL17.
[0042] FIG.10 illustrates plate-based calcium flux assay for testing functional inhibition of target proteins CCR8 and CCR4, showing selectivity for CCR8.
[0043] FIG.11A Illustrates the lasso peptide structure of SEQ ID NO: 76 showing the amino acids as balls and the critical W10 residue in the loop. 10 ACTIVE 705331286v1
[0044] FIG.11B shows a computational model of the complex between SEQ ID NO: 75 and ETBR, depicting the position of W10 in the energetically favored loop-down configuration.
[0045] FIG.12 illustrates the general two-step lasso peptide evolution process based on evolving the lasso peptide precursor peptide and subsequent cyclization using the lasso peptidase and lasso cyclase enzymes. Black balls represent mutated amino acid residues.
[0046] FIG.13 illustrates a five-step method involving computational modeling, conformational analysis, and model optimization for predicting lasso analogs with high integrin binding affinity.
[0047] FIG.14A shows ribbon diagrams of the apo form of the αvβ6 extracellular headpiece in a bent closed conformation where the computational model was constructed using atomic coordinates from PBD structure file 5FFG.
[0048] FIG.14B shows ribbon diagrams of the apo form of the αvβ6 extracellular headpiece in an open extended conformation where the computational model was constructed using atomic coordinates from PBD structure file 5NET
[0049] FIG.15A illustrates the lasso structure of SEQ ID NO: 109.
[0050] FIG.15B shows a model of the lasso peptide with SEQ ID NO: 109 docked into the ligand binding site of integrin αvβ6 and depicting the key binding interactions with the loop-grafted epitope RGDL.
[0051] FIG.16 illustrates iterative workflow for creating novel lasso peptides with high target affinity and optimized properties. 6. DETAILED DESCRIPTION
[0052] The features of this invention are set forth specifically in the appended claims. A better understanding of the features and benefits of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized. To facilitate a full understanding of the disclosure set forth herein, a number of terms are defined below. 6.1 General Techniques
[0053] Techniques and procedures described or referenced herein include those that are generally well understood and / or commonly employed using conventional methodology by those skilled in the art, such as, for example, the widely utilized methodologies described in 11 ACTIVE 705331286v1Sambrook et al., Molecular Cloning: A Laboratory Manual (4th ed.2012); Current Protocols in Molecular Biology (Ausubel et al. eds., 2003); Monoclonal Antibodies: Methods and Protocols (Albitar ed.2010). Antibody Engineering Vols 1 and 2 (Kontermann and Dübel eds., 2nd ed.2010). Molecular Biology of the Cell (6th Ed., 2014). Organic Chemistry, (Thomas Sorrell, 1999). March's Advanced Organic Chemistry (6thed.2007). Lasso Peptides, (Li, Y.; Zirah, S.; Rebuffet, S., Springer; New York, 2015). Natural Products in Medicinal Chemistry, Methods and Principles in Medicinal Chemistry (Hanessian, S., ed., Wiley-VCH; 1st edition, 2014). Basic Principles of Drug Discovery and Development (Blass, B. Academic Press; 1st edition, 2015). Computational Drug Discovery and Design, (Gore, M., Jagtap, U.B. (eds), Springer-Humana Press, 2018). Handbook of Biologically Active Peptides (Kastin, A. (ed), Academic Press, 2013). 6.2 Terminology
[0054] Unless described otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art. For purposes of interpreting this specification, the following description of terms will apply and whenever appropriate, terms used in the singular will also include the plural and vice versa. All patents, applications, published applications, and other publications are incorporated by reference in their entirety. In the event that any description of terms set forth conflicts with any document incorporated herein by reference, the description of term set forth below shall control.
[0055] The singular terms “a,” “an,” and “the” as used herein include the plural reference unless the context clearly indicates otherwise.
[0056] The term “about” or “approximately” means an acceptable error for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined. In certain embodiments, the term “about” or “approximately” means within 1, 2, 3, or 4 standard deviations. In certain embodiments, the term “about” or “approximately” means within 50%, 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, or 0.05% of a given value or range.
[0057] The term “substantially” means that something takes place, as a function or activity, to provide the expected outcome or result to a large degree and to a great extent, but still not to the fullest extent. For example, if a lasso peptide is substantially purified, the lasso peptide is isolated and purification steps afford the lasso peptide at purity level above 90% and as high as 99.99%. 12 ACTIVE 705331286v1
[0058] The term “substantially all” refers to at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or about 100%.
[0059] The terms “oligonucleotide” and “nucleic acid” refer to oligomers of deoxyribonucleotides (e.g., DNA) or ribonucleotides (e.g., RNA) and polymers thereof in either single- or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogues of natural nucleotides which have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless specifically limited otherwise, the term also refers to oligonucleotide analogs including PNA (peptidonucleic acid), and analogs of DNA used in antisense technology (phosphorothioates, phosphoroamidates, and the like). Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (including but not limited to, degenerate codon substitutions) and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer, M.A., et al., Nucleic Acid Res., 1991, 19, 5081-1585; Ohtsuka, E. et al., J. Biol. Chem., 1985, 260, 2605-2608; and Rossolini, G.M., et al., Mol. Cell. Probes, 1994, 8, 91-98). “Oligonucleotide,” as used herein, refers to short, generally single-stranded, synthetic polynucleotides that are generally, but not necessarily, fewer than about 200 nucleotides in length. The terms “oligonucleotide” and “polynucleotide” are not mutually exclusive. The description above for polynucleotides is equally and fully applicable to oligonucleotides. A cell that produces a lasso peptide of the present disclosure may include a bacterial and archaea host cell into which nucleic acids encoding the lasso peptide component have been introduced. Suitable host cells are disclosed below.
[0060] Unless specified otherwise, the left-hand end of any single-stranded polynucleotide sequence disclosed herein is the 5’ end; the left-hand direction of double- stranded polynucleotide sequences is referred to as the 5’ direction. The direction of 5’ to 3’ addition of nascent RNA transcripts is referred to as the transcription direction; sequence regions on the DNA strand having the same sequence as the RNA transcript that are 5’ to the 5’ end of the RNA transcript are referred to as “upstream sequences”; sequence regions on the DNA strand having the same sequence as the RNA transcript that are 3’ to the 3’ end of the RNA transcript are referred to as “downstream sequences.” 13 ACTIVE 705331286v1
[0061] The term “amino acid” refers to naturally occurring and non-naturally occurring alpha-amino acids, as well as alpha-amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring alpha-amino acids. Naturally encoded amino acids are the 22 common amino acids (alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid. glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, pyrrolysine and selenocysteine). Amino acid analogs or derivatives refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and a side chain R group, such as, homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (such as, norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes. The terms “non- natural amino acid” or “non-proteinogenic amino acid” or “unnatural amino acid” or “non- canonical” refer to alpha-amino acids that contain different side chains (different R groups) relative to those that appear in the twenty-two common or naturally occurring amino acids listed above. In addition, these terms also can refer to amino acids that are described as having D-stereochemistry or R-stereochemistry, rather than L-stereochemistry or S- stereochemistry of natural amino acids, despite the fact that some amino acids do occur in the D-stereochemical form in nature (e.g., D-alanine and D-serine).
[0062] The term “sterics” refers to the spatial volume occupied by an atom or a group of atoms in a molecule and as measured in cubic Angstroms. Experimentally, the sterics of an atom or groups of atoms are measured in various ways, such as how an atom or group of atoms influences the equilibrium between axial and equatorial conformations of monosubstituted cyclohexanes. The 20 common amino acids can be divided into five volume classes based on the sterics. Particularly, the “very large” class includes F, W, and Y having steric measurements in the range of about 189 to 228 cubic Angstroms. The “large” class includes I, L, M, K, and R having steric measurements in the range of about 162 to 174 cubic Angstroms. The “medium” class includes V, H, E, and Q having steric measurements in the range of about 138 to 154 cubic Angstroms, the “small” class includes C, P, T, D, and N having steric measurements in the range of about 108 to 117, and the “very small” class includes A, G, and S having steric measurements in the range of about 60 to 90 cubic 14 ACTIVE 705331286v1Angstroms. Hence, as used herein, a “sterically larger” molecule or chemical group refers to a molecule having a larger steric measurement relative to another molecule or chemical group, which is then referred to as a “sterically smaller” molecule or chemical group in the present disclosure.
[0063] The term “peptide” as used herein refers to a polymer chain containing between two and fifty (2-50) amino acid residues. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers in which one or more amino acid residues is a non- naturally occurring amino acid, e.g., an amino acid analog or non-natural amino acid.
[0064] The terms “polypeptide” and “protein” are used interchangeably herein to refer to a polymer of greater than about fifty (50) amino acid residues. That is, a description directed to a polypeptide applies equally to a description of a protein, and vice versa. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers in which one or more amino acid residues is a non-naturally occurring amino acid, e.g., an amino acid analog. As used herein, the terms encompass amino acid chains of any length, including full length proteins (i.e., antigens), wherein the amino acid residues are linked by covalent peptide bonds.
[0065] The terms “lasso peptide” and “lasso” are used interchangeably herein, and are used to refer to a class of peptide or polypeptide having the general lariat-like topology as exemplified in FIG.1. As shown in the figure, the lariat-like topology can be generally divided into a ring portion, a loop portion, and a tail portion. Particularly, a region on one end of the peptide forms the ring around the tail on the other end of the peptide, the tail is threaded through the ring, and a middle loop portion connects the ring and the tail, together forming the lariat-like topology. Particularly, the amino acid residues that are joined together to form the ring are herein referred to as the “ring-forming amino acid.” A ring-forming amino acid can located at the N- or C-terminus of the lasso peptide (“terminal ring-forming amino acid”), or in the middle (but not necessarily the center) of a lasso peptide (“internal ring-forming amino acid”). For example, the term “G1-D9 cyclized” as used herein when referring to a lasso peptide, means that the lasso peptide has a N-terminal ring-forming amino acid of a glycine residue (G1) and an internal ring-forming amino acid of an aspartate residue at position 9 (D9), where the amino group of G1 and the side chain carboxyl group of D9 form an isopeptide bond, thus forming the ring portion of the lasso peptide. The fragment of a lasso peptide between and including the two ring-forming amino acid residues is the ring portion; the fragment of a lasso peptide between the internal ring-forming amino acid and where the peptide threaded through the plane of the ring is the loop portion; and the 15 ACTIVE 705331286v1remaining fragment of a lasso peptide starting from where the peptide is threaded through the plane of the ring is the tail portion. In addition to the lariat-like topology, additional topological features of a lasso peptide may further include intra-peptide disulfide bonding, such as disulfide bond(s) between the tail and the ring, between the ring and the loop, and / or between different locations within the tail. As used herein, “lasso peptide” or “lasso” refers to both naturally-existing peptides and artificially designed or produced peptides that have the lariat-like topology as described herein. Similarly, “lasso peptide” or “lasso” also refers to non-naturally occurring analogs, derivatives, or variants of a naturally occurring lasso peptide, which analogs, derivatives or variants are also lasso peptides themselves.
[0066] The term “lasso precursor peptide” or “precursor peptide” as used herein refers to a precursor that is processed into or otherwise forms a lasso peptide. In some embodiments, a lasso precursor peptide comprises at least one a lasso core peptide portion. In some embodiments, a lasso precursor peptide comprises one or more amino acid residues or amino acid fragments that do not belong to a lasso core peptide, such as a leader sequence that facilitates recognition of the lasso precursor peptide by one or more lasso processing enzymes. In some embodiments, the lasso precursor peptide is enzymatically processed into a lasso peptide by removing the amino acid residues or fragments that do not belong to a lasso core peptide. In some embodiments, a lasso precursor peptide is the substrate of an enzyme that cleaves off the additional amino acid residues or fragments from a lasso precursor peptide to produce the lasso peptide. As used herein, the enzyme capable of catalyzing this reaction is referred to as the “lasso peptidase.”
[0067] The term “lasso core peptide” or “core peptide” refers to the peptide or the peptide segment of the precursor peptide that is processed into or otherwise forms a lasso peptide having the lariat-like topology. As used herein, the enzyme capable of catalyzing cyclization of the ring portion of a lasso core peptide is referred to as the “lasso cyclase.” As used herein, a core peptide may have the same amino acid sequence as a lasso peptide, but has not matured to have the lariat-like topology of a lasso peptide. In various embodiments, core peptides can have different lengths of amino acid sequences. In some embodiments, the core peptide is at least about 9 amino acid long. In some embodiments, the core peptide is at least about 10 amino acid long. In some embodiments, the core peptide is at least about 11 amino acid long. In some embodiments, the core peptide is at least about 12 amino acid long. In some embodiments, the core peptide is at least about 13 amino acid long. In some embodiments, the core peptide is at least about 14 amino acid long. In some embodiments, the core peptide is at least about 15 amino acid long. In some embodiments, the core peptide 16 ACTIVE 705331286v1is at least about 16 amino acid long. In some embodiments, the core peptide is at least about 17 amino acid long. In some embodiments, the core peptide is at least about 18 amino acid long. In some embodiments, the core peptide is at least about 19 amino acid long. In some embodiments, the core peptide is at least about 20 amino acid long. In some embodiments, the core peptide is at least about 25 amino acid long. In some embodiments, the core peptide is at least about 30 amino acid long. In some embodiments, the core peptide is at least about 35 amino acid long. In some embodiments, the core peptide is at least about 40 amino acid long. In some embodiments, the core peptide is at least about 45 amino acid long. In some embodiments, the core peptide is at least about 50 amino acid long. In some embodiments, the core peptide is at least about 55 amino acid long. In some embodiments, the core peptide is at least about 60 amino acid long. In some embodiments, the core peptide is at least about 65 amino acid long.
[0068] The term “lasso peptide backbone” as used herein, in the context of computational modeling, refers to a theoretical lasso peptide having glycine residues at each position of the peptide, except that the internal ring-forming amino acid residues is aspartate, glutamate or another non-naturally occurring amino acid residue having a side chain carboxyl group. In some embodiments, a computer modeling system defines a lasso peptide backbone by three parameters X, Y, Z, wherein X designates the number of glycine residues in the ring portion of the lasso peptide, Y designates the number of glycine residues in the loop portion of the lasso peptide, and Z designates the number of glycine residues in the tail portion of the lasso peptide. As such, in some embodiments, a computer modeling system is able to provide a plurality of lasso peptide backbone structures each having a different combination of X, Y, and Z values, resulting in different three-dimensional shapes. In some embodiments, a lasso peptide backbone structure can be created by first computationally modeling the 3D structure of a lasso peptide according to the present disclosure, and then computationally removing all side changes of amino acid residues in the lasso peptide except for the internal ring forming residue. In some embodiments, a computer modeling system uses a lasso peptide backbone as the starting structure for building a lasso peptide, or as a tool for simulating lasso-target interaction as described in various embodiments herein.
[0069] The terms “lasso peptide analog” or “lasso peptide variant” are used herein interchangeably and refer to a derivative of a natural lasso peptide that has been modified or changed relative to its original structure or atomic composition. In various embodiments, the lasso peptide analog can (i) have at least one amino acid substitution(s), insertion(s) or deletion(s) as compared to the sequence of a lasso peptide; (ii) have at least one different 17 ACTIVE 705331286v1modification(s) to the amino acids as compared to a lasso peptide, such modifications include but are not limited to acylation, biotinylation, O-methylation, N-methylation, amidation, glycosylation, pegylation, esterification, halogenation, amination, hydroxylation, dehydrogenation, prenylation, lipidoylation, heterocyclization, phosphorylation; (iii) have at least one unnatural amino acid(s) as compared to the sequence of a lasso peptide; (iv) have at least one different isotope(s) as compared to the lasso peptide molecule; or any combination of (i) to (iv). As used herein, the term of “lasso peptide analog” also includes a conjugate or fusion made of a lasso peptide or a lasso peptide analog and one or more additional molecule(s). In some embodiments, the additional molecule can be another peptide or protein, including but not limited a lasso peptide and a cell surface receptor or an antibody or an antibody fragment. In some embodiments, the additional molecule can be a non-peptidic molecule, such as a drug molecule, a fatty acid or lipid molecule, an isoprenoid molecule, or an oligonucleotide molecule. In some embodiments, the lasso peptide analogs retain the same general lasso topology as shown in FIG.1. In some embodiments, production of a lasso peptide analog may occur by introducing a modification into the gene of a lasso precursor or core peptide, followed by transcription and translation and cyclization using cell- free or cell-based methods, as described herein, leading to a lasso peptide containing that modification. In an alternative embodiment, production of a lasso peptide analog may occur by introducing a modification into a lasso precursor or core peptide, followed by cyclization of each using cell-free or cell-based methods, as described herein, leading to a lasso peptide containing that modification. In another embodiment, production of a lasso peptide analog may occur by introducing a modification into a pre-formed lasso peptide, leading to a lasso peptide containing that modification. In another embodiment, lasso peptide analogs are designed using structural information and in silico modeling algorithms, including docking and performing conformational analysis on 3D model structures using molecular dynamics simulation algorithms to obtain conformational states of 3D model structure of lasso peptides, and such “designed” lasso peptides may be produced by cell-free or cell-based methods. In another embodiment, lasso peptide analogs are designed de novo using computer algorithms, including artificial intelligence, machine learning, deep learning, neural nets, etc., and such “designed” lasso peptides may be produced by cell-free or cell-based methods.
[0070] The term “lasso peptide library” as used herein refers to a collection of at least two lasso peptides or lasso peptide analogs, or combinations thereof, which may be pooled together as a mixture or kept separated from one another. In some embodiments, the lasso peptide library is kept in vitro, such as in tubes or wells. In some embodiments, the lasso 18 ACTIVE 705331286v1peptide library may be created by biosynthesis of at least two lasso peptides or lasso peptide variants using a cell-free system. In some embodiments, the lasso peptide library may be created by biosynthesis of at least two lasso peptides or lasso peptide variants using a cell- based system. In some embodiments, the lasso peptides or lasso peptide variants of the library may be mixed with one or more component of the cell-free or cell-based systems. In other embodiments, the lasso peptides or lasso peptide variants may be purified from the cell- free or cell-based systems. In some embodiments, the lasso peptides or lasso peptide variants may be partially purified. In some embodiments, the lasso peptides or lasso peptide variants may be substantially purified. In some embodiments, the lasso peptides may be isolated. In some embodiments, the lasso peptide library may be created by isolating at least two lasso peptides from their natural environment. In some embodiments, the lasso peptides may be partially isolated. In some embodiments, the lasso peptides may be substantially isolated.
[0071] The term “isotopic variant” of a lasso peptide refers to a lasso peptide analog that contains an unnatural proportion of an isotope at one or more of the atoms that constitute such a peptide. In certain embodiments, an “isotopic variant” of a lasso peptide analog contains unnatural proportions of one or more isotopes, including, but not limited to, hydrogen (1H), deuterium (2H), tritium (3H), carbon-11 (11C), carbon-12 (12C) carbon-13 (13C), carbon-14 (14C), nitrogen-13 (13N), nitrogen-14 (14N), nitrogen-15 (15N), oxygen-14 (14O), oxygen-15 (15O), oxygen-16 (16O), oxygen-17 (17O), oxygen-18 (18O) fluorine-17 (17F), fluorine-18 (18F), phosphorus-31 (31P), phosphorus-32 (32P), phosphorus-33 (33P), sulfur-32 (32S), sulfur-33 (33S), sulfur-34 (34S), sulfur-35 (35S), sulfur-36 (36S), chlorine-35 (35Cl), chlorine-36 (36Cl), chlorine-37 (37Cl), bromine-79 (79Br), bromine-81 (81Br), iodine-123 (123I) iodine-125 (125I) iodine-127 (127I) iodine-129 (1291) and iodine-131 (131I). In certain embodiments, an “isotopic variant” of a lasso peptide is in a stable form, that is, non- radioactive. In certain embodiments, an “isotopic variant” of a lasso peptide contains unnatural proportions of one or more isotopes, including, but not limited to, hydrogen (1H), deuterium (2H), carbon-12 (12C), carbon-13 (13C), nitrogen-14 (14N), nitrogen-15 (15N), oxygen-16 (16O) oxygen-17 (17O), oxygen-18 (18O) fluorine-17 (17F), phosphorus-31 (31P), sulfur-32 (32S), sulfur-33 (33S), sulfur-34 (34S), sulfur-36 (36S), chlorine-35 (35Cl), chlorine-37 (37Cl), bromine-79 (79Br), bromine-81 (81Br), and iodine-127 (127I). In certain embodiments, an “isotopic variant” of a lasso peptide is in an unstable form, that is, radioactive. In certain embodiments, an “isotopic variant” of a compound contains unnatural proportions of one or more isotopes, including, but not limited to, tritium (3H), carbon-11 (11C), carbon-14 (14C), nitrogen-13 (13N), oxygen-14 (14O), oxygen-15 (15O), fluorine-18 (18F), phosphorus-32 (32P), 19 ACTIVE 705331286v1phosphorus-33 (33P), sulfur-35 (35S), chlorine-36 (36Cl), iodine-123 (123I) iodine-125 (125I), iodine-129 (129I) and iodine-131 (131I). It will be understood that, in a lasso peptide or lasso peptide analog as provided herein, any hydrogen can be2H, as example, or any carbon can be13C, as example, or any nitrogen can be15N, as example, and any oxygen can be18O, as example, where feasible according to the judgment of one of skill in the art. In certain embodiments, an “isotopic variant” of a lasso peptide contains an unnatural proportion of deuterium. Unless otherwise stated, structures of compounds (including peptides) depicted herein are also meant to include compounds that differ only in the presence of one or more isotopically enriched atoms. For example, compounds having the present structures including the replacement of hydrogen by deuterium or tritium, or the replacement of a carbon by a13C- or14C-enriched carbon are within the scope of this invention. Such compounds are useful, for example, as analytical tools, as probes in biological assays, or as therapeutic agents in accordance with the present invention.
[0072] The term “evolution” or “molecular evolution” or “evolve” as used herein refers to varying the sequence of a parent protein or peptide by introducing one or more mutations in the oligonucleotide sequence encoding the amino acid sequence of the protein or peptide. In some embodiments, a parent lasso peptide is evolved by introducing one or more mutations within the oligonucleotide sequence encoding a parent lasso peptide. In another embodiment, a parent lasso peptide is evolved by introducing one or more mutations within the amino acid sequence of the parent lasso peptide. In another embodiment, a parent lasso peptide is evolved by introducing one or more mutations within the amino acid sequence of the parent lasso peptide, including the introduction of natural or non-natural amino acids.
[0073] The term “biosynthetic gene cluster” as used herein refers to one or more nucleic acid molecule(s) independently or jointly comprising one or more coding sequences for a precursor and processing machinery capable of maturing the precursor into a biosynthetic end product. The coding sequences can comprise multiple open reading frames (ORFs) each independently coding for one component of the precursor and processing machinery. Alternatively, the coding sequences can comprise an ORF coding for two or more components of the precursor and processing machinery fused together, as further described herein. A biosynthetic gene cluster can be identified and isolated from the genome of an organism. Computer-based analytical tools can be used to mine genomic information and identify biosynthetic gene clusters encoding lasso peptides. For example, the genome-mining tool known as Rapid ORF Description and Evaluation Online (RODEO) has been used to identify more than a thousand of lasso biosynthetic gene clusters based on available genomic 20 ACTIVE 705331286v1information (Tietz et al. Nat Chem Biol.2017 May; 13(5): 470–478). Alternatively, a biosynthetic gene cluster can be assembled by artificially producing and combining the nucleic acid components of the gene cluster, using genetic manipulating methods and technology known in the art.
[0074] Some naturally existing lasso peptides are encoded by a lasso peptide biosynthetic gene cluster, which typically comprises three main genes: one encodes for a lasso precursor peptide (referred to as Gene A), and two encode for processing enzymes including a lasso peptidase (referred to as Gene B or B2) and a lasso cyclase (referred to as Gene C). The lasso precursor peptide comprises a lasso core peptide and additional peptidic fragments known as the “leader sequence” that facilitates recognition and processing by the processing enzymes. The leader sequence may determine substrate specificity of the processing enzymes. The processing enzymes encoded by the lasso peptide gene cluster convert the lasso precursor peptide into a matured lasso peptide having the lariat-like topology. Particularly, the lasso peptidase removes from the precursor peptide the additional portion that is not the lasso core peptide, and the lasso cyclase cyclizes a terminal portion of the core peptide around a terminal tail portion to form the lariat-like topology.
[0075] Some lasso gene clusters further encode for additional protein elements that facilitates the post-translational modification, including a facilitator protein known as the post-translationally modified peptide (RiPP) recognition element (RRE, also referred to as Gene E or B1). A lasso peptide biosynthetic gene clusters may encode two or more of lasso peptidase, lasso cyclase and RRE as different domains in the same protein. Some lasso gene clusters further encode for lasso peptide transporters, kinases, or proteins that play a role in immunity, such as isopeptidase. (Burkhart, B.J., et al., Nat. Chem. Biol., 2015, 11, 564–570; Knappe, T.A. et al., J. Am. Chem. Soc., 2008, 130, 11446-11454; Solbiati, J.O. et al. J. Bacteriol., 1999, 181, 2659-2662; Fage, C.D., et al., Angew. Chem. Int. Ed., 2016, 55, 12717 –12721; Zhu, S., et al., J. Biol. Chem.2016, 291, 13662–13678).
[0076] The term “lasso peptide biosynthesis component” as used herein refer to a protein comprising one or more of (i) a lasso peptidase, (ii) a lasso cyclase, and (iii) RRE.
[0077] The terms “cell-free biosynthesis” and “CFB” are used interchangeably herein and refer to an in vitro (outside the cell) biosynthetic process for the production of one or more peptides or proteins. In some embodiments, cell-free biosynthesis occurs in a “cell-free biosynthesis reaction mixture” or “CFB reaction mixture” which provides various components, such as RNA, proteins, enzymes, co-factors, natural products, small molecules, organic molecules, to carry out protein synthesis outside a living cell. In some embodiments, 21 ACTIVE 705331286v1the CFB reaction mixture can comprise one or more cell extracts or supplemented cell extracts, or commercially available cell-free reaction media (e.g. PURExpress®). Exemplary CFB methods and systems, including those involving the use of in vitro TX-TL, are described in Culler, S. et al., PCT Application WO2017 / 031399 A1, and is incorporated herein by reference. In some embodiments, cell-free biosynthesis involves the use of one or more isolated precursor peptide, core peptide, lasso peptidase, lasso cyclase, and lasso RRE. In some embodiments, cell-free biosynthesis involves the use of one or more isolated precursor peptide, core peptide, lasso peptidase, lasso cyclase, and lasso RRE, wherein an isolated precursor peptide or core peptide are produced either synthetically or biologically.
[0078] As used herein, the terms “in vitro transcription and translation” and “in vitro TX- TL” are used interchangeably and refer to a biosynthetic process outside an intact cell, where genes or oligonucleotides are transcribed into messenger ribonucleic acids (mRNAs), and mRNAs are translated into proteins or peptides. As used herein, the term “in vitro TX-TL machinery” refers to the components that act in concert to carry out the in vitro TX-TL. For the sole purpose of illustration, and by way of non-exhaustive and non-limiting examples, in some embodiments, an in vitro TX-TL machinery comprises enzyme(s) and co-factor(s) that carry out DNA transcription and / or mRNA translation. In some embodiments, an in vitro TX-TL machinery further comprises other small organic or inorganic molecules, such as amino acids, tRNAs or ATP, that facilitate the DNA transcription and / or mRNA translation. Various cellular components known to participate in in vivo transcription and translation can form part of the in vitro TX-TL machinery, see for example, Matsubayashi et al, “Purified cell-free systems as standard parts for synthetic biology.” Curr Opin Chem Biol.2014 Oct; 22:158-62; Li, et al. “Improved cell-free RNA and protein synthesis system.” PLoS One. 2014 Sep 2; 9 (9):e106232. In some embodiments, different components can be provided individually and combined to assemble the in vitro TX-TL machinery. Exemplary ways of providing the in vitro TX-TL machinery components include recombinant production, synthesis, and isolation from a cell. In some embodiments, the in vitro TX-TL machinery is provided in the form of one or more cell extract, or one or more supplemented cell extract that comprises the in vitro TX-TL machinery.
[0079] The term “condition suitable for lasso formation,” depending on the context, may refer to, for example, a condition suitable for the expression of one or more protein products in a bacterial host (e.g., a lasso precursor peptide, or a lasso processing enzyme). Exemplary suitable conditions included are not limited to a suitable culturing condition of the bacterial host that enable the protein synthesis and transportation in the host cell. Additionally, or 22 ACTIVE 705331286v1alternatively, depending on the context, the term “condition suitable for lasso formation” may refer to, for example, a condition suitable for post-translational modification of a lasso precursor peptide. Exemplary suitable conditions include but are not limited to a suitable temperature, pH, and / or incubation time for a lasso cyclase and / or lasso peptidase to process the lasso precursor into a matured lasso peptide.
[0080] The terms “microbial,” “microbial organism” or “microorganism” are intended to mean any organism that exists as a microscopic cell that is included within the domains of archaea, bacteria or eukarya. Therefore, the term is intended to encompass prokaryotic or eukaryotic cells or organisms having a microscopic size and includes bacteria, archaea and eubacteria of all species as well as eukaryotic microorganisms such as yeast and fungi. The term also includes cell cultures of any species that can be cultured for the production of a biochemical.
[0081] The term “naturally occurring” or “natural” or “native” when used in connection with naturally occurring biological materials such as nucleic acid molecules, oligonucleotides, amino acids, polypeptides, peptides, metabolites, small molecule natural products, host cells, and the like, refers to materials that are found in or isolated directly from Nature and are not changed or manipulated by humans. The term “natural” or “naturally occurring” refers to organisms, cells, genes, biosynthetic gene clusters, enzymes, proteins, oligonucleotides, and the like that are found in Nature and are unchanged relative to these components found in Nature. The term “wild-type” refers to organisms, cells, genes, biosynthetic gene clusters, enzymes, proteins, oligonucleotides, and the like that are found in Nature and are unchanged relative to these components found in Nature (in the wild).
[0082] The term “non-naturally occurring” or “non-natural” or “unnatural” or “non- native” or “non-canonical” as used herein refer to a material, substance, molecule, cell, enzyme, protein, peptide, or amino acid that is not known to exist or is not found in Nature or that has been structurally modified and / or synthesized by humans. The term “non-natural” or “unnatural” or “non-naturally occurring” when used in reference to a microbial organism or microorganism or cell extract or gene or biosynthetic gene cluster of the invention is intended to mean that the microbial organism or derived cell extract or gene or biosynthetic gene cluster has at least one genetic alteration not normally found in a naturally occurring strain or a naturally occurring gene or biosynthetic gene cluster of the referenced species, including wild-type strains of the referenced species. Genetic alterations include, for example, introduction of expressible oligonucleotides or nucleic acids encoding polypeptides, other nucleic acid additions, nucleic acid deletions and / or other functional disruption of the 23 ACTIVE 705331286v1microbial organism’s genetic material. Such modifications include, for example, nucleotide changes, additions, or deletions in the genomic coding regions and functional fragments thereof, used for heterologous, homologous or both heterologous and homologous expression of polypeptides. Additional modifications include, for example, nucleotide changes, additions, or deletions in the genomic non-coding and / or regulatory regions in which the modifications alter expression of a gene or operon. Exemplary polypeptides include enzymes, proteins, or peptides within a lasso peptide biosynthetic pathway. The terms “non- naturally occurring” or “non-natural” or “unnatural” or “non-native” or “non-canonical” are used to refer to amino acids that are introduced into a polypeptide sequence to modify the properties of the polypeptide.
[0083] The term “vector” refers to a substance that is used to carry or include a nucleic acid sequence, including for example, a nucleic acid sequence encoding a lasso precursor peptide, or lasso processing enzymes as described herein, in order to introduce a nucleic acid sequence into a host cell. Vectors applicable for use include, for example, expression vectors, plasmids, phage vectors, viral vectors, episomes, and artificial chromosomes, which can include selection sequences or markers operable for stable integration into a host cell’s chromosome. Additionally, the vectors can include one or more selectable marker genes and appropriate expression control sequences. Selectable marker genes that can be included, for example, provide resistance to antibiotics or toxins, complement auxotrophic deficiencies, or supply critical nutrients not in the culture media. Expression control sequences can include constitutive and inducible promoters, transcription enhancers, transcription terminators, and the like, which are well known in the art. When two or more nucleic acid molecules are to be co-expressed (e.g., both a lasso core peptide and a lasso cyclase), both nucleic acid molecules can be inserted, for example, into a single expression vector or in separate expression vectors. For single vector expression, the encoding nucleic acids can be operationally linked to one common expression control sequence or linked to different expression control sequences, such as one inducible promoter and one constitutive promoter. The introduction of nucleic acid molecules into a host cell can be confirmed using methods well known in the art. Such methods include, for example, nucleic acid analysis such as Northern blots or polymerase chain reaction (PCR) amplification of DNA, immunoblotting for expression of gene products, or other suitable analytical methods to test the expression of an introduced nucleic acid sequence or its corresponding gene product. It is understood by those skilled in the art that the nucleic acid molecules are expressed in a sufficient amount to produce a desired product (e.g., a lasso precursor peptide as described herein), and it is further understood that 24 ACTIVE 705331286v1expression levels can be optimized to obtain sufficient expression using methods well known in the art.
[0084] The term “encoding nucleic acid” or grammatical equivalents thereof as it is used in reference to nucleic acid molecule refers to a nucleic acid molecule in its native state or when manipulated by methods well known to those skilled in the art that can be transcribed to produce mRNA, which is then translated into a polypeptide and / or a fragment thereof. The antisense strand is the complement of such a nucleic acid molecule, and the encoding sequence can be deduced therefrom.
[0085] The term “exogenous” as it is used herein is intended to mean that the referenced molecule or the referenced activity is introduced into the host microbial organism. The molecule can be introduced, for example, by introduction of an encoding nucleic acid into the host genetic material such as by integration into a host chromosome or as non-chromosomal genetic material such as a plasmid. Therefore, the term as it is used in reference to expression of an encoding nucleic acid refers to introduction of the encoding nucleic acid in an expressible form into a microbial organism or into a cell extract for cell-free expression. When used in reference to a biosynthetic activity, the term refers to an activity that is introduced into the host reference organism or into a cell extract for cell-free activity. The source can be, for example, a homologous or heterologous encoding nucleic acid that expresses the referenced activity following introduction into the host microbial organism or into a cell extract for cell-free expression of activity. Therefore, the term “endogenous” refers to a referenced molecule or activity that is present in a microbial host. Similarly, the term when used in reference to expression of an encoding nucleic acid refers to expression of an encoding nucleic acid contained within the microbial organism or into a cell extract. The term “heterologous” refers to a molecule or activity derived from a source other than the referenced species whereas “homologous” refers to a molecule or activity derived from the host microbial organism or organism used to produce a cell-free extract. Accordingly, exogenous expression of an encoding nucleic acid of the invention can utilize either or both a heterologous or homologous encoding nucleic acid.
[0086] The term “isolated” when used in reference to a microbial organism or a biosynthetic gene, or a biosynthetic gene cluster, or a protein, or an enzyme, or a peptide, is intended to mean an organism, gene or biosynthetic gene cluster, protein, enzyme, or peptide that is substantially free of at least one component relative to the referenced microbial organism, gene, biosynthetic gene cluster, protein, enzyme, or peptide as it is found in nature or in its natural habitat. The term includes a microbial organism, gene, biosynthetic gene 25 ACTIVE 705331286v1cluster, protein, enzyme, or peptide that is removed from some or all components as it is found in its natural environment. Therefore, an isolated microbial organism, gene, biosynthetic gene cluster, protein, enzyme, or peptide is partly or completely separated from other substances as it is found in nature or as it is grown, stored, or subsisted in non-naturally occurring environments (e.g., laboratories). Specific examples of isolated microbial organisms, genes, biosynthetic gene clusters, proteins, enzymes, or peptides include partially pure microbes, genes, biosynthetic gene clusters, proteins, enzymes, or peptides, substantially pure microbes, genes biosynthetic gene clusters, proteins, enzymes, or peptides, and microbes cultured in a medium that is non-naturally occurring, or genes or biosynthetic gene clusters cloned in non-naturally occurring plasmids, or proteins, enzymes, or peptides purified from other components and substances present their natural environment, including other proteins, enzymes, or peptides.
[0087] An “isolated nucleic acid” is a nucleic acid, for example, an RNA, a DNA, or a mixed nucleic acid, which is substantially separated from other genome DNA sequences as well as proteins or complexes such as ribosomes and polymerases, which naturally accompany a native sequence. An “isolated” nucleic acid molecule is one which is separated from other nucleic acid molecules which are present in the natural source of the nucleic acid molecule. Moreover, an “isolated” nucleic acid molecule, such as a cDNA molecule, can be substantially free of other cellular material, or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. In a specific embodiment, one or more nucleic acid molecules encoding a lasso peptide precursor peptide as described herein are isolated or purified. The term embraces nucleic acid sequences that have been removed from their naturally occurring environment, and includes recombinant or cloned DNA isolates and chemically synthesized analogues or analogues biologically synthesized by heterologous systems. A substantially pure molecule may include isolated forms of the molecule.
[0088] The term “substantially anaerobic” when used in reference to a culture or growth condition is intended to mean that the amount of oxygen is less than about 10% of saturation for dissolved oxygen in liquid media. The term also is intended to include sealed chambers of liquid or solid medium maintained with an atmosphere of less than about 1% oxygen.
[0089] The terms “binds” or “binding” or “binding interaction” or “bonding interaction” refer to an interaction or set of interactions between molecules including, for example, to form a complex. The sum total of the binding interactions between molecules determines the binding free energy associated forming a complex between two or more molecules. 26 ACTIVE 705331286v1Interactions can be, for example, covalent binding interactions, or non-covalent interactions including hydrogen bonding interactions, ionic or electrostatic bonding interactions, polar interactions, dipolar interactions, induced dipolar interactions, pi stacking interactions, hydrophobic interactions, and / or van der Waals interactions. A complex can also include the binding of two or more molecules held together by covalent or non-covalent bonds, interactions, or forces. The strength of the total non-covalent interactions between a single target-binding site of a binding molecule (e.g., a ligand), and a single target site of a target molecule (e.g., a protein) is the affinity of the binding molecule or functional fragment for that target site. The ratio of dissociation rate (koff) to association rate (kon) of a binding molecule to a monovalent target site (koff / kon) is the dissociation constant KD, which is inversely related to affinity. The lower the KD value, the higher the affinity of the binding molecule. The value of KD varies for different complexes of lasso peptides or target proteins depends on both konand koff. The dissociation constant KDfor a binding molecule (e.g., a lasso peptide) provided herein can be determined using any method provided herein or any other method well known to those skilled in the art. The affinity at one binding site does not always reflect the true strength of the interaction between a binding molecule and the target molecule. When complex target molecule containing multiple, repeating target sites, such as a polyvalent target protein, come in contact with lasso peptides containing multiple target binding sites, the interaction of the lasso peptide with the target protein at one site will increase the probability of a reaction at a second site.
[0090] The term “binding affinity” generally refers to the strength of the sum total of noncovalent interactions between a single binding site of a molecule (e.g., a binding molecule such as a lasso peptide) and its binding partner (e.g., a target protein). Unless indicated otherwise, as used herein, “binding affinity” refers to intrinsic binding affinity which reflects a 1:1 interaction between members of a binding pair (e.g., lasso peptide and target protein). The affinity of a binding molecule X for its binding partner Y can generally be represented by the dissociation constant (KD). Affinity can be measured by common methods known in the art, including those described herein. Low-affinity lasso peptides generally bind target proteins slowly and tend to dissociate readily, whereas high-affinity lasso peptides generally bind target proteins faster and tend to remain bound longer. A variety of methods of measuring binding affinity are known in the art, any of which can be used for purposes of the present disclosure. Specific illustrative embodiments include the following. In one embodiment, the “KD” or “KDvalue” may be measured by assays known in the art, for example by a binding assay. The KD may be measured in a radioimmunoassay (RIA), for 27 ACTIVE 705331286v1example, performed with the lasso peptide of interest and its target protein. The KD or KD value may also be measured by using surface plasmon resonance assays by Biacore®, using, for example, a Biacore®TM-2000 or a Biacore®TM-3000, or by biolayer interferometry using, for example, the Octet®QK384 system. An “on-rate” or “rate of association” or “association rate” or “kon” may also be determined with the same surface plasmon resonance or biolayer interferometry techniques described above using, for example, a Biacore®T-200 or a Biacore®8K+, or the Octet®QK384 system.
[0091] The terms “target,” “target molecule,” “target polypeptide” and “target protein” are used interchangeably herein and refer to a peptide or polypeptide, for which it is desirable to produce a lasso peptide that specifically binds thereto.
[0092] The term “lasso-binding site” as used herein refers to the region on a target protein or peptide that contains at least one amino acid residue with which a lasso peptide interacts to form the binding with the target protein. According to the present disclosure, a lasso-binding site of a target protein can contain surface groupings of binding moieties (e.g., chemical active groups on amino acids or sugar side chains) responsible for binding with the lasso peptide. The binding moieties can locate on the same amino acid residue or different amino acid residues in the lasso-binding site. In some embodiments, the lasso-binding site of a target protein can have specific three-dimensional structural characteristics (e.g., a groove or pocket) where the binding moieties disperse. In some embodiments, different lasso peptides may bind to different target sites or compete for binding with the same target site of a target protein. In some embodiments, a lasso peptide specifically binds to a target molecule or a target site thereof. In various embodiments according to the present disclosure, a lasso- binding site can be a ligand binding site, a steric hindrance binding site, an orthosteric binding site, an allosteric control binding site, a conformational change binding site, a channel, or a catalytic site.
[0093] The term “orthosteric site” or “orthosteric binding site” is a term of art that refers to the primary binding site on a receptor that is recognized by the endogenous ligand or agonist for that receptor. For example, the orthosteric site in the M1 receptor is the site that acetylcholine binds. As used herein, the term “allosteric site” or “allosteric binding site” refers to the specific binding site other than the orthosteric site on a receptor molecule that, upon binding of an allosteric effector molecule, influences the function or activity of that receptor.
[0094] The term “epitope” and “binding epitope” are used interchangeably herein to indicate generally the set of one or more amino acid residues on a binding molecule (e.g., a 28 ACTIVE 705331286v1lasso peptide) that is responsible for the binding interaction between the binding molecule and a target molecule (e.g., a target protein). According to the present disclosure, an epitope can be linear or conformational, based on their structure and interaction with the binding target. A conformational epitope is formed by the three-dimensional conformation adopted by the interaction of discontinuous segments of amino acid residue(s). In contrast, a linear epitope is formed by the interaction of contiguous amino acid residues. In some embodiments, a linear epitope is not determined solely by the primary structure of the epitope amino acids directly involved in binding interactions. Residues that flank such amino acid residues, as well as more distant amino acid residues of the antigen can affect the ability of the primary structure epitope residues to adopt the epitope’s three-dimensional conformation required for optimal binding. According to the present disclosure, an epitope of a peptide or polypeptide can be “grafted” into another peptide or polypeptide, for example by inserting the sequence (e.g., in case of a linear epitope) or sequences (e.g., in case of a conformational epitope) of an epitope into the sequence of the peptide or polypeptide, such that the grafted peptide or polypeptide retains the binding capability and specificity of the epitope. Alternatively, an epitope of a peptide or polypeptide can be “grafted” into another peptide or polypeptide, for example by replacing a segment (e.g., in case of a linear epitope) or segments (e.g., in case of a conformation epitope) of the original sequence of the peptide or peptide with the sequence or sequences of the epitope, such that the grafted peptide or polypeptide retain the binding capability and specificity of the epitope. For example, arginine-glycine-aspartate (RGD) is a linear epitope composed of three amino acids that together bind to certain integrins on the surface of cells. The RGD epitope was grafted into the loop region of the lasso peptide microcin J25, and while the original lasso peptide microcin J25 did not bind to integrins, the RGD-grafted microcin J25 bound with high affinity to certain integrins (Hegemann, J.D., et al., J. Med. Chem.2014, 57, 5829−5834).
[0095] The term “target-binding epitope” as used herein refers to the amino acid residue or the group of amino acid residues of a lasso peptide that binds to a target molecule. A target-binding epitope as disclosed herein can be a linear epitope that contains a fragment of continuous amino acids (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 continuous amino acids) of the lasso peptide (FIG.3A and FIG.3B), or a conformational epitope that consists of amino acid residues in two or more non-continuous regions of the lasso peptide (FIG.4). According to the present disclosure, a target-binding epitope can contain surface groupings of binding moieties (e.g., chemical active groups on amino acids or sugar side chains) responsible for binding with the target protein. The binding moieties can locate on the same 29 ACTIVE 705331286v1amino acid residue or different amino acid residues forming the target-binding epitope. In some embodiments, the target-binding epitope of a lasso peptide can have specific three- dimensional structural characteristics (e.g., an edge or a hump) where the binding moieties are arranged and displayed. A lasso peptide can have more than one different target-binding epitopes that bind with different target sites of a target protein, or bind with different target proteins.
[0096] The term “structurally opposite to” refers to at least two amino acids residing on the target molecule and the lasso peptide, respectively, where the at least two amino acids are located and oriented in a 3-dimensional space in such a way that enables a binding interaction between the at least two amino acids. For example, a lysine residue on a target molecule may be located in space at the correct distance and orientation to engage in a binding interaction with an aspartate residue on the lasso peptide. The two amino acid residues in this example, lysine and aspartate, can referred to as being structurally opposite to each other.
[0097] A lasso peptide which binds a target molecule of interest is one that binds the target molecule with sufficient affinity such that the lasso peptide is useful, for example, as a diagnostic or therapeutic agent in targeting a cell or tissue expressing the target molecule, and does not significantly cross-react with other molecules. In such embodiments, the extent of binding of the lasso peptide to a “non-target” molecule will be less than about 10% of the binding of the lasso peptide to its particular target molecule, for example, as determined by fluorescence activated cell sorting (FACS) analysis or RIA.
[0098] With regard to the binding of a lasso peptide to a target molecule, the term “specific binding,” “specifically binds to,” or “is specific for” a particular polypeptide or a fragment on a particular polypeptide target means binding that is measurably different from a non-specific interaction. Specific binding can be measured, for example, by determining binding of a molecule compared to binding of a control molecule, which generally is a molecule of similar structure that does not have binding activity. For example, specific binding can be determined by competition with a control molecule that is similar to the target, for example, an excess of non-labeled target. In this case, specific binding is indicated if the binding of the labeled target to a probe is competitively inhibited by excess unlabeled target. The term “specific binding,” “specifically binds to,” or “is specific for” a particular polypeptide or a fragment on a particular polypeptide target as used herein refers to binding where a molecule binds to a particular polypeptide or fragment on a particular polypeptide without substantially binding to any other polypeptide or polypeptide fragment. In certain embodiments, a lasso peptide that binds to a target molecule has a dissociation constant (KD) 30 ACTIVE 705331286v1of less than or equal to 100 µM, 80 µM, 50 µM, 25 µM, 10 µM, 5 µM, 1 µM, 900 nM, 800 nM, 700 nM, 600 nM, 500 nM, 400 nM, 300 nM, 200 nM, 100 nM, 50 nM, 10 nM, 5 nM, 4 nM, 3 nM, 2 nM, 1 nM, 0.9 nM, 0.8 nM, 0.7 nM, 0.6 nM, 0.5 nM, 0.4 nM, 0.3 nM, 0.2 nM, 0.1 nM, 0.09 nM, 0.08 nM, 0.07 nM, 0.06 nM, 0.05 nM, 0.04 nM, 0.03 nM, 0.02 nM, or 0.01 nM.
[0099] In the context of the present disclosure, a target protein is said to specifically bind to a lasso peptide, for example, when the dissociation constant (KD) is ≤10-7M. In some embodiments, the lasso peptides specifically bind to a target protein with a KD of from about 10-7M to about 10-12M. In certain embodiments, the lasso peptides specifically bind to a target protein with high affinity when the KDis ≤10-8M or KDis ≤10-9M. In one embodiment, the lasso peptides may specifically bind to a purified human target protein with a KD of from 1 x 10-9M to 10 x 10-9M as measured by Biacore®. In another embodiment, the lasso peptides may specifically bind to a purified human target protein with a KDof from 0.1 x 10-9M to 1 x 10-9M as measured by KinExA™ (Sapidyne, Boise, ID). In yet another embodiment, the lasso peptides specifically bind to a target protein expressed on cells with a KDof from 0.1 x 10-9M to 10 x 10-9M. In certain embodiments, the lasso peptides specifically bind to a human target protein expressed on cells with a KDof from 0.1 x 10-9M to 1 x 10-9M. In some embodiments, the lasso peptides specifically bind to a human target protein expressed on cells with a KDof 1 x 10-9M to 10 x 10-9M. In certain embodiments, the lasso peptides specifically bind to a human target protein expressed on cells with a KDof about 0.1 x 10-9M , about 0.5 x 10-9M, about 1 x 10-9M, about 5 x 10-9M, about 10 x 10-9M, or any range or interval thereof. In still another embodiment, the lasso peptides specifically bind to a non-human target protein expressed on cells with a KDof 0.1 x 10-9M to 10 x 10-9M. In certain embodiments, the lasso peptides specifically bind to a non-human target protein expressed on cells with a KD of from 0.1 x 10-9M to 1 x 10-9M. In some embodiments, the lasso peptides specifically bind to a non-human target protein expressed on cells with a KD of 1 x 10-9M to 10 x 10-9M. In certain embodiments, the lasso peptides specifically bind to a non-human target protein expressed on cells with a KD of about 0.1 x 10-9M, about 0.5 x 10-9M, about 1 x 10-9M, about 5 x 10-9M, about 10 x 10-9M, or any range or interval thereof.
[0100] With regard to the binding of a lasso peptide to a target molecule, the term “preferential binding” or “preferentially binds to” a particular polypeptide or a fragment on a particular target molecule with respect to a reference molecule means binding of the target molecule is measurably higher than binding of the reference molecule, while the reference 31 ACTIVE 705331286v1molecule may or may not also bind to the lasso peptide. Preferential binding can be determined, for example, by determining the binding affinity. For example, a lasso peptide that preferentially binds to a target molecule over a reference molecule can bind to the target molecule with a KD less than the KD exhibited relative to the reference molecule. In some embodiments, the lasso peptide preferentially binds a target molecule with a KD less than half of the KDexhibited relative to the reference molecule. In some embodiments, the lasso peptide preferentially binds a target molecule with a KDat least 10 times less than the KDexhibited relative to the reference molecule. In some embodiments, the lasso peptide preferentially binds a target molecule with a KDwith KDthat is about 75%, about 50%, about 25%, about 10%, about 5%, about 2.5%, or about 1% of the KDexhibited relative to the reference molecule. In some embodiments, the ratio between the KD exhibited by the lasso peptide when binding to the reference molecule and the KD exhibited when binding to the target molecule is at least 2 fold, at least 3 fold, at least 4 fold, at least 5 fold, at least 10 fold, at least 20 fold, at least 100 fold, at least 500 fold, at least 103fold, at least 104fold, or at least 105fold.
[0101] A lasso peptide that specifically or preferentially binds to a target protein can be identified, for example, by immunoassays (e.g., ELISA, fluorescent immunosorbent assay, chemiluminescence immune assay, radioimmunoassay (RIA), enzyme multiplied immunoassay, solid phase radioimmunoassay (SPRIA), a surface plasmon resonance (SPR) assay (e.g., Biacore®), a biolayer interferometry assay, a fluorescence polarization assay, a fluorescence resonance energy transfer (FRET) assay, Dot-blot assay, fluorescence activated cell sorting (FACS) assay, or other techniques known to those of skill in the art. A lasso peptide binds specifically to a target protein when it binds to the target protein with higher affinity than to any cross-reactive target molecule as determined using experimental techniques, such as radioimmunoassays (RIA) and enzyme linked immunosorbent assays (ELISAs). Typically a specific or selective reaction will be at least twice background signal or noise and may be more than 10 times background.
[0102] The term “compete” when used in the context of lasso peptides (e.g., a lasso peptide and other binding proteins that bind to and compete for the same target molecule or target site on the target molecule) means competition as determined by an assay in which the lasso peptide (or binding fragment) thereof under study prevents or inhibits the specific binding of a reference molecule (e.g., a reference ligand of the target molecule) to a common target molecule. Numerous types of competitive binding assays can be used to determine if a test lasso peptide competes with a reference ligand for binding to a target molecule. 32 ACTIVE 705331286v1Examples of assays that can be employed include solid phase direct or indirect RIA, solid phase direct or indirect enzyme immunoassay (EIA), sandwich competition assay (see, e.g., Stahli et al., 1983, Methods in Enzymology 9:242-53), solid phase direct biotin-avidin EIA (see, e.g., Kirkland et al., 1986, J. Immunol.137:3614-19), solid phase direct labeled assay, solid phase direct labeled sandwich assay (see, e.g., Harlow and Lane, Antibodies, A Laboratory Manual (1988)), solid phase direct label RIA using I-125 label (see, e.g., Morel et al., 1988, Mol. Immunol.25:7-15), and direct labeled RIA (Moldenhauer et al., 1990, Scand. J. Immunol.32:77-82). Typically, such an assay involves the use of a purified target molecule bound to a solid surface, or cells bearing either of an unlabeled test target-binding lasso peptide or a labeled reference target-binding protein (e.g., reference target-binding ligand). Competitive inhibition may be measured by determining the amount of label bound to the solid surface in the presence of the test target-binding lasso peptide. Usually the test target-binding molecule is present in excess. Target-binding lasso peptides identified by competition assay (e.g., competing lasso peptides) include lasso peptides binding to the same target site as the reference and lasso peptides binding to an adjacent target site sufficiently proximal to the target site bound by the reference for steric hindrance to occur. Additional details regarding methods for determining competitive binding are described herein. Usually, when a competing lasso peptide is present in excess, it will inhibit specific binding of a reference to a common target molecule by at least 30%, for example 40%, 45%, 50%, 55%, 60%, 65%, 70%, or 75%. In some instance, binding is inhibited by at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more.
[0103] The term “blocking” lasso peptide or an “antagonist” lasso peptide is one which inhibits or reduces biological activity of the target molecule it binds. For example, blocking lasso peptide or antagonist lasso peptide may substantially or completely inhibit the biological activity of the target molecule.
[0104] The term “inhibition” or “inhibit,” when used herein, refers to partial (such as, 1%, 2%, 5%, 10%, 20%, 25%, 50%, 75%, 90%, 95%, 99%) or complete (i.e., 100%) inhibition. The term “IC50” refers an amount, concentration, or dosage of a compound that results in 50% inhibition of a maximal response in an assay that measures such response. The term “EC50” refers an amount, concentration, or dosage of a compound that results in for 50% of a maximal response in an assay that measures such response.
[0105] The term “attenuate,” “attenuation,” or “attenuated,” when used herein, refers to partial (such as, 1%, 2%, 5%, 10%, 20%, 25%, 50%, 75%, 90%, 95%, 99%) or complete (i.e., 100%) reduction in a property, activity, function, effect, or value. 33 ACTIVE 705331286v1
[0106] With regard to inhibition or attenuation or antagonism of a target molecule by a lasso peptide, the term “selective inhibition of,” “selectively inhibits,” “selective antagonism of,” “selectively antagonizes,” “selective attenuation of,” or “selectively attenuates” a target molecule or a signaling pathway mediated by a target molecule means inhibition of the target molecule activity is measurably stronger than inhibition of a reference molecule activity. Selective inhibition can be determined, for example, by determining the IC50value. For example, a lasso peptide that selectively inhibits or attenuates a target molecule can exhibit an IC50 value less than the IC50 exhibited relative to a reference molecule. In some embodiments, the lasso peptide selectively inhibits or attenuates a target molecule with an IC50less than half of the IC50exhibited relative to the reference molecule. In some embodiments, the lasso peptide selectively inhibits or attenuates a target molecule with an IC50 that is about 75%, about 50%, about 25%, about 10%, about 5%, about 2.5%, or about 1% of the IC50exhibited relative to the reference molecule. In some embodiments, the ratio between the IC50 exhibited by the lasso peptide with respect to the reference molecule and the IC50 exhibited with respect to the target molecule is at least 2 fold, at least 3 fold, at least 4 fold, at least 5 fold, at least 10 fold, at least 20 fold, at least 100 fold, at least 500 fold, at least 103fold, at least 104fold, or at least 105fold.
[0107] The phrase “substantially similar” or “substantially the same” denotes a sufficiently high degree of similarity between two numeric values (e.g., one associated with a lasso peptide of the present disclosure and the other associated with a reference ligand) such that one of skill in the art would consider the difference between the two values to be of little or no biological and / or statistical significance within the context of the biological characteristic measured by the values (e.g., KDvalues). For example, the difference between the two values may be less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 10%, or less than about 5%, as a function of the value for the reference ligand.
[0108] The phrase “substantially increased,” “substantially reduced,” or “substantially different,” as used herein, denotes a sufficiently high degree of difference between two numeric values (e.g., one associated with a lasso peptide of the present disclosure and the other associated with a reference ligand) such that one of skill in the art would consider the difference between the two values to be of statistical significance within the context of the biological characteristic measured by the values. For example, the difference between said two values can be greater than about 10%, greater than about 20%, greater than about 30%, 34 ACTIVE 705331286v1greater than about 40%, or greater than about 50%, as a function of the value for the reference ligand.
[0109] The term “assaying” is meant the creation of experimental conditions and the gathering of data regarding a particular result of the exposure to specific experimental conditions. For example, enzymes can be assayed based on their ability to act upon a detectable substrate. A lasso peptide can be assayed based on its ability to bind to a particular target molecule or molecules.
[0110] The term “modulating” or “modulate” as used herein refers to an effect of altering a biological activity (i.e. increasing or decreasing the activity), especially a biological activity associated with a particular biomolecule such as a, enzyme or cell surface receptor. For example, an inhibitor of a particular biomolecule modulates the activity of that biomolecule, e.g., an enzyme, by decreasing the activity of the biomolecule, such as an enzyme. Such activity is typically indicated in terms of an inhibitory concentration (IC50) of the compound for an inhibitor with respect to, for example, an enzyme or a cell surface receptor.
[0111] The term “contacting” and its grammatical variations, when used in reference to two or more components, refers to any process whereby the approach, proximity, mixture or commingling of the referenced components is promoted or achieved without necessarily requiring physical contact of such components, and includes mixing of solutions containing any one or more of the referenced components with each other. The referenced components may be contacted in any particular order or combination and the particular order of recitation of components is not limiting. For example, “contacting A with B and C” encompasses embodiments where A is first contacted with B then C, as well as embodiments where C is contacted with A then B, as well as embodiments where a mixture of A and C is contacted with B, and the like. Furthermore, such contacting does not necessarily require that the end result of the contacting process be a mixture including all of the referenced components, as long as at some point during the contacting process all of the referenced components are simultaneously present or simultaneously included in the same mixture or solution.
[0112] The term “partially” means that something takes place, as a function or activity, to provide the expected outcome or result in part and to a limited extent, not to the fullest extent. For example, if a lasso peptide is partially purified, the lasso peptide is isolated and purification steps afford the lasso peptide at purity level that is greater than about 20% and less than about 90%.
[0113] The terms “computational modeling,” “computational design,” “in silico modeling,” and “in silico design” may be used interchangeably and refer to methods of 35 ACTIVE 705331286v1applying computational algorithms or programs to predict, simulate, analyze, and assess the structure and properties of a molecule, especially with respect to the interaction between one molecule and another, such as a lasso peptide interacting with a target protein.
[0114] The term “docking” refers to the use of computer algorithms to fit a model ligand compound structure, such as a model three-dimensional structure of a lasso peptide, into the space of a binding site of a model target protein, and adjusting the models according to chemical and physical principles to simulate and predict the energy of their interactions. In some embodiments, amino acid residue interactions between the lasso peptide and the target protein are analyzed and energy is minimized to predict the lowest energy or highest-affinity binding. In some embodiments, the three-dimensional structure of the target protein is known or can be computationally created so that a plurality of model three-dimensional structures of the lasso peptide can be docked onto a known or predicted lasso-binding site of the target protein. In certain embodiments, interactions between amino acid residues of a plurality of lasso peptides and the target protein are analyzed and energy is minimized to predict and rank the lowest-energy or highest-affinity binding lasso peptide using scoring functions. In some embodiments, rigid docking methods are used to predict and rank the binding affinities of docked lasso peptides. In some embodiments, flexible docking methods are used to predict and rank the binding affinities of docked lasso peptides.
[0115] The term “structural complementarity” is used herein in the context of a three- dimensional structural feature of the accessible surface of a given peptide or protein, or an assembly of peptides and / or proteins. Structural complementarity forms the means to molecular recognition that allows the structure of one molecule to interact with another molecule similarly to matching pieces in a puzzle, to splines or tenons dovetailed into their corresponding grooves in a machine, or to a lock and its corresponding key. The phrase “structural complementarity” as used herein is meant to encompass chemical compatibility as well as spatial compatibility. Hence peptides or proteins in general, by virtue of their architectural and functional diversity, can interact with one another based on structural complementarity which combines (i) a suitable and matching placement or positioning of one or more chemically-corresponding associating groups which are selected as members of a binding pair that can form one or more bonds with their corresponding member of the binding pair, and (ii) the overall structural features of the compound which depend mostly on the core structure of the molecule to which the associating groups are attached.
[0116] The term “atomic coordinate” is a term of art and as used herein refers to a mathematical description of the location of an atom in 3-dimensional (3D) space using a 36 ACTIVE 705331286v1coordinate system. Exemplary types of atomic coordinates that can be used in connection with the present disclosure includes but are not limited to Cartesian coordinates (X, Y, Z), internal or “Z” matrix coordinates, redundant internal coordinates, symmetry-adapted coordinates, polar coordinates, Quaternion coordinates. In certain embodiments, atomic coordinates are used to define a model of atomic structure of a molecule (e.g., a peptide or a protein) that fits experimental data. For example, atomic structures archived in databases such as the wwPDB are models that have been constructed to be as consistent as possible with available experimental data. In some embodiments, atomic coordinates of a protein or peptide molecule can be obtained based on diffraction of X-rays by the atoms of a protein or peptide crystal. The diffraction data are typically used to calculate an electron density map, which is used to establish the positions of the individual atoms within the unit cell of the crystal. In some embodiments, atomic coordinates of a protein or peptide molecule can be derived from chemical shift, J coupling, and magnetic relaxation data based on the nuclear overhauser effect (NOE). In some embodiments, nuclear magnetic resonance (NMR) data regarding a given molecule (e.g., a peptide or polypeptide) provide information about distances between atoms and local conformations, which information allows an ensemble of atomic structure models to be constructed, each model with slightly different conformations and atomic coordinates, for the molecule in solution. Thus, in addition to atomic coordinates, NMR data affords information about the conformational flexibility of a molecule or portions of a molecule. Averaged atomic coordinates can be obtained from the ensemble of atomic structure models.
[0117] As used herein the general term “molecular modeling” is a term of art and refers to the use of a computer algorithm to generate a predicted structural model of a molecule, including a protein or peptide, or a set of interacting molecules, such as a lasso peptide and a target protein. A structural model generated as such is herein referred to as a “model structure,” such as a three-dimensional (3D) model structure. Molecular modeling can be performed with a collection of computer-based techniques for deriving, representing, simulating, predicting, and manipulating the structures and behaviors of molecules, as well as reactions and interactions between molecules, and those properties that are dependent on these three-dimensional atomic structures. Molecular modeling implements molecular mechanics calculations based on equilibrium bond lengths, bond angles, partial charge values, force constants, and van der Waals parameters, which collectively are referred to as force fields. Deviations from these equilibrium force field functions, along with non-bonded van der Waals and electrostatic interactions, will increase or decrease the energy of a system. 37 ACTIVE 705331286v1Energy minimization and optimization algorithms are used to find the lowest energy arrangements of interacting atoms or molecules that are defined by force fields.
[0118] The term “homology modeling” is a term of art and refers to a procedure that generates a previously unknown protein structure by “fitting” its sequence (target) into a known structure (template), given a certain level of sequence homology (at least 30%) between target and template. Homology modeling thus predicts the 3D structure of a query protein through the sequence and structure alignment with one or more template proteins of known 3D structure. Generally, the process of homology modeling involves four steps: template identification, sequence alignment, model building, and model refinement (Meier and Soding, PLoS Comput Biol.2015; 11(10): e1004343. doi: 10.1371 / journal.pcbi.1004343).
[0119] The term “protein structure prediction” refers to the use of algorithms for the prediction of tertiary protein structure based on primary sequence data, conformational modeling, and energy minimization. Numerous advanced algorithms have been developed for accurate protein structure prediction, including trRosetta (Yang, J., et al., Improved protein structure prediction using predicted interresidue orientations. Proc. Nat. Acad. Sci., USA.2020, 117(3), 1496-1503), AlphaFold (Senior, A.W., et al., Improved protein structure prediction using potentials from deep learning. Nature, 2020, 577, 706–710) and AlphaFold 2 (Skolnick et al., AlphaFold 2: Why It Works and Its Implications for Understanding the Relationships of Protein Sequence, Structure, and Function. J Chem Inf Model., 2021 Oct 25;61(10):4827-4831). Protein structure prediction algorithms may be used alone or in combination with homology modeling to create accurate 3D computational models of target proteins.
[0120] The term “molecular replacement” is a term of art and refers to a process implemented to overcome the phase problem associated with X-ray diffraction data regarding a molecule. In molecular replacement, a known structure model very similar in structure to the one crystallized, is computationally placed in the crystallographic unit cell. From this model, diffraction data phases can be calculated and used to start the process of interpreting the electron density map in order to solve the unknown structure and define atomic coordinates (Evans, P., McCoy, A. An introduction to molecular replacement. Acta Cryst. Bio. Cryst., 2008, D64, 1-10.).
[0121] The term “configuration” when used in the context of a three-dimensional (3D) structure of a molecule, refers to the fixed 3D relationship of atoms or groups of atoms of a molecule that is defined by the bonds between the atoms or groups of atoms. 38 ACTIVE 705331286v1
[0122] The term “conformation” when used in the context of a three-dimensional (3D) structure of a molecule, refers to the spatial arrangements among atoms or groups of atoms in a molecule that may adopt and convert between, such as by rotation about individual single bonds between the atoms or groups of atoms.
[0123] The term “orientation” when used in the context of spatial relationship between two or more molecules refers to the position of a molecule or a portion of the molecule in a three-dimensional (3D) space relative to another molecule or a portion of such other molecule.
[0124] The term “identity” refers to a relationship between the sequences of two or more polypeptide molecules or two or more nucleic acid molecules, as determined by aligning and comparing the sequences. “Percent (%) amino acid sequence identity” with respect to a reference polypeptide sequence is defined as the percentage of amino acid residues in a candidate sequence that are identical with the amino acid residues in the reference polypeptide sequence, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and not considering any conservative substitutions as part of the sequence identity. Alignment for purposes of determining percent amino acid sequence identity can be achieved in various ways that are within the skill in the art, for instance, using publicly available computer software such as BLAST, BLAST-2, ALIGN, or MEGALIGN (DNAStar, Inc.) software. Those skilled in the art can determine appropriate parameters for aligning sequences, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared. 6.3 Lasso Peptides
[0125] Bacterially-derived lasso peptides are emerging as a class of natural molecular scaffolds for drug design (Hegemann, J.D. et al., Acc. Chem. Res., 2015, 48, 1909−1919; Zhao, N., et al., Amino Acids, 2016, 48, 1347–1356; Maksimov, M.O., et al., Nat. Prod. Rep., 2012, 29, 996-1006). Lasso peptides are members of the larger class of natural ribosomally synthesized and post-translationally modified peptides (RiPPs). Lasso peptides are derived from a precursor peptide, comprising a leader sequence and core peptide sequence, which is cyclized through formation of an isopeptide bond between the N-terminal amino group of the linear core peptide and the side chain carboxyl groups of glutamate or aspartate residues located at positions 7, 8, 9, or 10 of the linear core peptide. The resulting macrolactam ring is formed around the C-terminal linear tail, which is threaded through the ring leading to the 39 ACTIVE 705331286v1characteristic lasso (also referred to as lariat) topology of general structure 1 as shown in FIG.1, which is held in place through sterically bulky side chains below and sometimes above the plane of the ring, and sometimes containing disulfide bonds between the tail and the ring or alternatively only in the tail.
[0126] Lasso peptide gene clusters typically consist of three main genes, one coding for the precursor peptide (referred to as Gene A), and two for the processing enzymes, a lasso peptidase (referred to as Gene B or B2) and a lasso cyclase (referred to as Gene C) that close the macrolactam ring around the tail to form the unique lariat structure. The precursor peptide consists of a leader sequence that binds to and directs the enzymes that carry out the cyclization reaction, and a core peptide sequence which contains the amino acids that together form the nascent lasso peptide upon cyclization. In addition, most lasso peptide gene clusters contain additional genes, such as those that encode for a small facilitator protein called a RIPP recognition element (RRE, also referred to as Gene E or B1), those that encode for lasso peptide transporters, those that encode for kinases, or those that encode proteins that are believed to play a role in immunity, such as an isopeptidase (Burkhart, B.J., et al., Nat. Chem. Biol., 2015, 11, 564–570; Knappe, T.A. et al., J. Am. Chem. Soc., 2008, 130, 11446- 11454; Solbiati, J.O. et al. J. Bacteriol., 1999, 181, 2659-2662; Fage, C.D., et al., Angew. Chem. Int. Ed., 2016, 55, 12717 –12721; Zhu, S., et al., J. Biol. Chem.2016, 291, 13662– 13678).
[0127] The ultimate lasso peptide directly derives from a core peptide that typically comprises a linear sequence ranging from about 11 to 50 amino acids in length. The macrolactam ring of a lasso peptide may contain 7, 8, 9, or 10 amino acids, while the loop and tail vary in length.
[0128] A genomic sequence mining algorithm called RODEO, has enabled identification of over 3000 entirely new lasso peptide gene clusters associated with a broad range of different bacterial species in the GenBank database, which is a vast increase over the 38 lasso peptides previously described in the literature (Tietz, J.I., et al., Nature Chem Bio, 2017, 13, 470-478; DiCaprio, A.J., et al., J. Am. Chem. Soc.2019, 141, 1, 290-297). Previous genome mining tools struggled to identify lasso peptide biosynthetic gene clusters due to the small size of the gene clusters and particularly the precursor peptide genes (Hegemann, J.D., et al., Biopolymers, 2013, 100, 527–542; Maksimov, M.O., et al., Proc. Nat. Acad. Sci., 2012, 109, 15223–15228). This study also demonstrated that lasso peptides are much more widespread in nature and more sequence diverse than previously expected. 40 ACTIVE 705331286v1
[0129] Lasso peptides represent a new type of molecular diversity which could serve as a vast source of novel products for use in the pharmaceutical, diagnostic, agricultural, and consumer industries. While the discovery of new activities and functions can occur by exploring the naturally occurring lasso peptides, a large percentage (>95%) of natural lasso peptides remain as prophetic entities predicted on the basis of genome sequence analyses. Lasso peptide development is constrained by the lack of effective methods to rapidly convert virtual lasso peptide biosynthetic gene cluster sequences into actual molecules that can be characterized and screened for biological activity and new methods for advancing lasso peptides are required.
[0130] To address this need, in one aspect of the present disclosure, provided herein are new methods for the identification, design, production, and optimization of lasso peptides. Particularly, the constrained nature of the three-dimensional lariat topology of a lasso peptide scaffold is suited for rational in silico design of non-natural sequences and structures that possess desirable properties. Thus, provided herein are an integrated set of technologies, methods, and tools that are implemented iteratively to develop new lasso peptides. Also provided herein is a design-build-test-evolve (DBTE) workflow that encompasses computational design, lasso peptide biosynthesis, testing of designed lasso peptide properties, and evolution of lasso peptides to optimize functions such as binding affinity and selectivity, and properties such as solubility, cell membrane permeability, metabolic stability, and pharmacokinetics. 6.4 Computational Methods for in silico Design of Lasso Peptides
[0131] In an effort to accelerate the drug discovery and optimization process, the pharmaceutical industry has been striving over the past forty years to develop and use in silico methods and algorithms to analyze structural information and assist in what is called computer-aided drug design (CADD).(Sliwoski, G., et al., Pharmacol Rev., 2014, 66:334– 395) CADD is capable of increasing the hit rate of new drug compounds because it uses a much more targeted search than traditional high throughput screening (HTS) and combinatorial chemistry. It not only aims to explain the molecular basis of therapeutic activity but also to predict possible derivatives that would improve activity. In a drug discovery campaign, CADD is usually used for three major purposes: (1) filter large compound libraries into smaller sets of predicted active compounds that can be tested experimentally; (2) guide the optimization of lead compounds, whether to increase its affinity 41 ACTIVE 705331286v1or optimize drug metabolism and pharmacokinetics (DMPK) properties including absorption, distribution, metabolism, excretion, and the potential for toxicity (ADMET); (3) design new compounds, either by “constructing” starting molecules one functional group at a time or by piecing together fragments into new chemotypes.
[0132] CADD can be classified into two general categories: structure-based and ligand- based. Structure-based CADD relies on the knowledge of the target protein structure to calculate interaction energies for all compounds tested, whereas ligand-based CADD exploits the knowledge of known active and inactive molecules through chemical similarity searches or construction of predictive, quantitative structure-activity relation (QSAR) models (Kalyaanamoorthy and Chen, Drug Discovery Today, 2011, 16(17-18), 831–839). Structure- based CADD is generally preferred where high-resolution structural data of the target proteins are available, e.g., for soluble proteins that can be structurally analyzed by X-ray crystallography or NMR spectrometry. Ligand-based CADD is generally preferred when no or little structural information is available, often for membrane protein targets. The central goal of structure-based CADD is to design compounds that bind tightly to the target, i.e., with large reduction in free energy, improved DMPK / ADMET properties, and are target specific, i.e., have reduced off-target effects (Pinzi, S., Rastelli, G. Int. J. Mol. Sci.2019, 20, 4331; doi:10.3390 / ijms20184331). A successful application of these methods will result in a compound that has been validated in vitro and in vivo and its binding location has been confirmed, ideally through a cocrystal structure.
[0133] One of the common uses in CADD is the screening of virtual compound libraries, also known as virtual high-throughput screening (vHTS). This allows experimentalists to focus resources on testing compounds likely to have any activity of interest. In this way, a researcher can identify an equal number of hits while screening significantly less compounds, because compounds predicted to be inactive with high confidence may be skipped. Avoiding a large population of inactive compounds saves money and time, because the size of the experimental HTS is significantly reduced without sacrificing a large degree of hits. (Ripphausen, P., et al., J. Med. Chem., 2010, .53(24), 8461–8467). The largest fraction of hits has been obtained thus far is for G-protein-coupled receptors (GPCRs) followed by kinases. vHTS comes in many forms, including chemical similarity searches by fingerprints or topology, selecting compounds by predicted biologic activity through QSAR models or pharmacophore mapping, and virtual docking of compounds into target of interest, known as structure-based docking. These methods allow the ranking of “hits” from the virtual compound library for acquisition. The ranking can reflect a property of interest such as 42 ACTIVE 705331286v1percent similarity to a query compound or predicted biologic activity, or in the case of docking, the lowest energy scoring poses for each ligand bound to the target of interest. Often initial hits are rescored and ranked using higher level computational techniques that are too time consuming to be applied to full-scale vHTS. It is important to note that vHTS does not aim to identify a drug compound that is ready for clinical testing, but rather to find leads with chemotypes that have not previously been associated with a target. This is not unlike a traditional HTS where a compound is generally considered a hit if its activity (e.g., Kd, IC50, or EC50) is close to 10 micromolar. Through iterative rounds of compound synthesis and in vitro testing, a potential drug is first developed into a “lead” with higher affinity, some understanding of its structure-activity-relation, and initial tests for DMPK / ADMET properties. Only after further iterative rounds of lead-to-drug optimization and in vivo testing does a compound reach a clinically appropriate potency and acceptable DMPK / ADMET properties.
[0134] The cost benefit of using computational tools in the lead optimization phase of drug development can be substantial. Development of new drugs can cost anywhere in the range of 500 million to 2 billion dollars, with synthesis and testing of lead analogs being a large contributor to that sum (Wouters, O.J., et al., JAMA.2020, 323(9), 844-853. doi:10.1001 / jama.2020.1166). Therefore, it is beneficial to apply computational tools in hit- to-lead optimization to cover a wider chemical space while reducing the number of compounds that must be synthesized and tested in vitro. The computational optimization of a hit compound can involve a structure-based analysis of docking poses and energy profiles for hit analogs, ligand-based screening for compounds with similar chemical structure or improved predicted biologic activity, or prediction of favorable DMPK / ADMET properties. The comparably low cost of CADD compared with chemical synthesis and biologic characterization of large libraries of compounds make these methods attractive to focus, reduce, and diversify the chemical space that is explored (Enyedy and Egan, J Comput Aided Mol Des., 2008; 22:161–168). De novo drug design is another tool in CADD methods, but rather than screening libraries of previously synthesized compounds, it involves the design of novel compounds. A structure generator is needed to sample the space of chemicals. Given the size of the total chemical search space of more than 1060molecules, and total available compound library size of over 95 million chemicals (Rifaioglu, A.S., et al., Briefings in Bioinformatics, 2019, 20(5), 1878-1912), heuristics are used to focus these algorithms on molecules that are predicted to be highly active, readily synthesizable, devoid of undesirable properties, often derived from a starting scaffold with demonstrated activity, etc. 43 ACTIVE 705331286v1Additionally, effective sampling strategies are used while dealing with large search spaces such as active-learning, evolutionary algorithms, metropolis search, or simulated annealing (Reker, D., et al., Chem. Sci., 2016, 7, 3919–3927). The construction algorithms are generally defined as either linking or growing techniques. Linking algorithms involve docking of small fragments or functional groups such as rings, acetyl groups, esters, etc., to particular binding sites followed by linking fragments from adjacent sites. Growing algorithms, on the other hand, begin from a single fragment placed in the binding site to which fragments are added, removed, and changed to improve activity. Similar to vHTS, the role of de novo drug design is not to design the single compound with nanomolar activity and acceptable DMPK / ADMET properties but rather to design a lead compound that can be subsequently improved.
[0135] Advances in the algorithms for in silico modeling have led to the creation of increasingly powerful tools for analyzing, designing, and optimizing potential new drugs. These new methods involve self-learning algorithms that are referred to by terms including artificial intelligence, machine learning, deep learning, adaptive learning, active learning, evolutionary learning, and deep reinforcement learning (Hessler, G., et al., Molecules, 2018, 23, 2520; doi:10.3390 / molecules23102520; Rifaioglu, A.S., et al., Briefings in Bioinformatics, 2019, 20(5), 1878-1912). These algorithms are enabling virtual screening (Carpenter, K.A., et al., Current Pharmaceutical Design, 2018, 24, 3347-3358) and de novo design (Schneider, G., Clark, D.E. Angew. Chemie Int. Ed. Engl., 2019, 58(32), 10792- 10803) of small molecule drugs based on extensive data that has been collected over decades of research. Importantly, as training datasets improve, learning algorithms are becoming increasingly proficient at predicting new molecular structures to explore. In certain embodiments of the present disclosure, in silico modeling methods can be applied to peptides and proteins, such as a lasso peptide and its target protein.
[0136] If a three-dimensional structure of a particular biological target is unavailable but one or more binding molecules are known, ligand-based design provides an alternative strategy. This scenario holds true for G-protein-coupled receptors (GPCRs), which are the most successful drug targets in terms of therapeutic benefit and sales. A ligand-based strategy, in contrast, can either consider the three-dimensional or the topological structure of one or more known ligands. When possible, target protein structures or homology models are preferred for defining binding sites and predicting ligand binding interactions.
[0137] Lasso peptides are scaffolds for amino acids and provide unique structural and geometric constraints in terms of arrangements of its amino acid residues in space for optimal interaction with a target protein. In some embodiments, lasso-binding sites are identified 44 ACTIVE 705331286v1through a search for target protein amino acids that are arranged in a manner that would allow engagement with complementary amino acids displayed on a 3D lasso peptide scaffold. Amino acids that are structurally complementary to a target protein set of amino acids in a predicted binding site are computationally arranged on a plurality of different lasso peptide scaffolds to predict scaffold-amino acid combinations that are scored to identify the lowest energy binding interactions (e.g., highest predicted binding high affinity).
[0138] Lasso peptides offer a uniquely constrained 3-dimensional structure that is amenable to designing new biologically active molecules through predictive in silico modeling. Provided herein are methods that utilize computational algorithms to design and optimize the properties of lasso peptides. In one aspect, lasso peptides having optimal binding are identified using computational algorithms that enable in silico docking of lasso peptide structures into a binding site of a target protein. In one aspect, lasso peptides are designed using computational algorithms that enable in silico docking of lasso peptide structures into a binding site of a target protein.
[0139] In one aspect, lasso peptides are identified using computational algorithms that enable in silico docking of lasso peptide structures into a lasso-binding site of protein structures (FIG.2B). In some embodiments, lasso peptides are identified and binding interactions are refined using computational algorithms that enable in silico docking and multiple rounds of conformational dynamic modeling of lasso peptide-protein target structures that are scored and ranked on the basis of predicted binding affinity of a lasso peptide bound to a protein structure. In some embodiments, lasso peptides are identified using computational algorithms that enable in silico docking of lasso peptide structures into a lasso-binding site of protein structures and the lasso peptide-protein interactions are further refined and ranked using artificial intelligence algorithms. In some embodiments, lasso peptides are identified and binding interactions are refined using computational and artificial intelligence algorithms that enable in silico docking and conformational dynamic modeling of lasso peptide structures that are scored and ranked on the basis of predicted binding affinity in a lasso-binding site of protein structures.
[0140] In some embodiments, lasso peptides are identified, and binding interactions are refined, using computational algorithms that enable in silico docking and conformational dynamic modeling of lasso peptide structures that are scored and ranked on the basis of predicted binding affinity in a lasso binding site of protein structures. In some embodiments, lasso peptides are identified using computational algorithms that enable in silico docking of lasso peptide structures into a lasso binding site of protein structures and the lasso peptide- 45 ACTIVE 705331286v1protein interactions are further refined and ranked using artificial intelligence algorithms. In some embodiments, lasso peptides are identified and binding interactions are refined using computational and artificial intelligence algorithms that enable in silico docking and conformational dynamic modeling of lasso peptide structures that are scored and ranked on the basis of predicted binding affinity in a lasso-binding site of protein structures.
[0141] In some embodiments, lasso peptides having optimal binding with a target molecule are identified using a combination of in silico docking algorithms and conformational analysis on 3D model structures of lasso peptides using molecular dynamics simulation algorithms. Such methods can include providing one or more three-dimensional (3D) model structures of lasso peptides, performing a first conformational analysis on these structures using molecular dynamics simulation algorithms, docking the conformational states of the 3D model structures on to a 3D model structure of a target molecule, performing a second conformational analysis on these structures using molecular dynamics simulation algorithms, and then selecting the 3D model structure of the lasso peptide-target molecule complex having a favored free energy. Such a method can identify lasso peptide binder candidates having optimal binding with a target molecule. These methods can be performed on multiple 3D model structures to provide at least two conformational states of the 3D model structures. Accordingly, in some embodiments, the at least two conformational states of the one ore more 3D model structures of lasso peptides can be selected from 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 24, 30, 35, 40, 45, or 50 or more different conformational states.
[0142] In some embodiments, the method for identifying a lasso peptide having optimal binding with a target molecule includes: (a) providing one or more three-dimensional (3D) model structures of lasso peptides; (b) performing conformational analysis on the one or more 3D model structures of lasso peptides using molecular dynamic simulation algorithms to obtain at least two conformational states of the one or more 3D model structures of lasso peptides; (c) docking the at least two conformational states of the one or more 3D model structures of lasso peptides onto a 3D model structure of the target molecule at a lasso- binding site of the target molecule to form at least two 3D model structures of the lasso peptide-target molecule complex; (d) performing conformational analysis on the at least two 3D model structures of the lasso peptide-target molecule complex using molecular dynamic simulation algorithms to obtain at least two conformational states of the 3D model structure of the lasso peptide-target molecule complex; and (e) selecting the 3D model structure of the lasso peptide-target molecule complex having a favored free energy, thereby identifying a lasso peptide binder candidate. 46 ACTIVE 705331286v1
[0143] In some embodiments, lasso peptides having optimal binding with a target molecule are identified, or lasso peptides are identified and binding interactions with target proteins are refined, using in silico docking algorithms (Pagadala, N.S., et al., Biophys Rev, 2017, 9, 91–102) including, but not limited to, CABS-dock (Kurcinski, M., et al., Protein Science, 2020, 29, 211–222), Macromodel (Mohamadi, F., et al., J. Comput. Chem, 1990, 11, 440-467), FlexPepDock (London, N. et al., Nucleic Acids Res.201039, W249–W253), DynaDock (Antes, I., Proteins, 2010, 78, 1084–1104), Autodock (Morris, G.M., et al., J Comput. Chem., 2009, 30, 2785-2791), MOE-Dock (Corbell, C.R., et al., J Comput Aided Mol Des., 2012, 26, 775-786), Surflex-dock (Jain, A.N., J Med Chem, 2003, 46, 499– 511), Glide (Friesner, R.A., et al., J. Med. Chem., 2004, 47, 1739-1749) AutoDock Vina (Trott, O.; Olson, A. J., J. Comput. Chem., 2010, 31 (2), 455−461), or ICM (Neves, M.A.C., et al., J Comput Aided Mol Des., 2012, 26, 675–686). In some embodiments, lasso peptides having optimal binding with a target molecule are identified, or lasso peptides are identified and binding interactions with target proteins are refined, using in silico docking algorithms together with conformational analysis on 3D model structures of lasso peptides using molecular dynamics simulation algorithms, including but not limited to the publicly available or commercial programs, GROMACS, AMBER, CHARMM, and Schroedinger’s MAESTRO.
[0144] In some embodiments, lasso peptides are identified and binding interactions with target proteins are modeled and analyzed by using artificial intelligence, deep learning, or machine learning algorithms for in silico virtual screening and / or de novo design to define the overall lasso peptide structural topologies encompassing the loop, ring, and tail size, as well as amino acid residues that suitably fit inside and engage the lasso-binding site of a target protein in order to optimize lasso peptide binding affinity predictions that may be refined further through iterative in silico docking and molecular dynamics simulations approaches. In some embodiments, one or more different artificial intelligence, deep learning, or machine learning algorithms are used, including but not limited to, support vector machines (Warmuth, M.K., et al., J. Chem. Inf. Comput. Sci., 2003, 43, 667-673), random forest (Deshmukh A.L., et al., Mol. Biosyst., 2017, 13, 1630–1639), k-nearest neighbors (Luo, M., et al., Mol. Inf., 2016, 35, 36 – 41), as well as neural networks with or without autoencoders, such as convolutional neural network (Jimenez, J., et al., J. Chem. Inf. Model., 2018, 58, 287−296), recurrent neural network (Sattarov, B., et al., J. Chem. Inf. Model.2019, 59, 1182−1196; Muller, A.T., et al., J. Chem. Inf. Model., 2018, 58, 2, 472-479), deep neural network (Ma, J., et al., J. Chem. Inf. Model., 2015, 55, 263−274), generative neural network 47 ACTIVE 705331286v1(Gupta, A., et al., Mol. Inf., 2018, 37, 1700111), and generative adversarial neural network (Prykhodko, O, et al., J Cheminform., 2019, 11:74; doi.org / 10.1186 / s13321-019-0397).
[0145] In some embodiments, lasso peptides are identified and binding interactions with target proteins are optimized using in silico docking algorithms together with conformational molecular dynamics algorithms and / or artificial intelligence, deep learning, or machine learning algorithms, and lasso peptides are simultaneously optimized for multiple properties, such as solubility, cell permeability, cellular activity, or stability using artificial intelligence algorithms that include, but are not limited to, fuzzy-logic design simulations (Warszawski, S., et al., J. Mol. Biol., 2014, 426, 4125–4138), iterative stochastic elimination (Stern, N., et al., Isr. J. Chem., 2014, 54, 1338 – 1357), and / or deep reinforcement learning (Popova, M., et al., Sci. Adv.2018;4: eaap7885; doi: 10.1126 / sciadv.aap7885).
[0146] In some embodiments, the structures of all known natural lasso peptide sequences are predicted using computational algorithms, including but not limited to PEP-FOLD3 (Lamiable, A., et al., Nucl. Acids Res., 2016, 44, doi: 10.1093 / nar / gkw329), PEPstrMOD (Singh, S., et al., Biology Direct, 2015, 10:73 doi: 10.1186 / s13062-015-0103-4), Rosetta and GenKIC (Ovchinnikov, S., et al., Proteins, 2018, 86, 113–121; Bhardwaj, G., et al., Nature, 2016, 538, 329–335), I-TASSER and QUARK (Zhang, C., et al., Proteins, 2018, 86, 136– 151).
[0147] In one aspect, provided herein are methods for identifying and grafting a target protein-binding epitope of a natural or synthetic polypeptide ligand into a lasso peptide topological structure (FIG.1). In some embodiments, lasso peptides containing epitope grafted segments are identified using computational algorithms that enable in silico docking of lasso peptide structures into a binding site of a target protein structures. In some embodiments, lasso peptides containing epitope grafted segments are identified and binding interactions are optimized using computational algorithms that enable in silico docking and conformational dynamic modeling of lasso peptide structures that are scored and ranked on the basis of predicted binding affinity in pocket of protein structures. In some embodiments, lasso peptides containing epitope grafted segments are identified using computational algorithms that enable in silico docking of lasso peptide structures into a binding site of a target protein structures and the lasso peptide-protein interactions are further refined and ranked using artificial intelligence algorithms. In some embodiments, lasso peptides containing epitope grafted segments are identified and binding interactions are optimized using computational and artificial intelligence algorithms that enable iterative in silico 48 ACTIVE 705331286v1docking and conformational dynamic modeling of lasso peptide structures that are scored and ranked on the basis of predicted binding affinity in a lasso binding site of protein structures.
[0148] In some embodiments, a binding epitope is computationally grafted into a lasso peptide structure. In some embodiments, epitope-grafted lasso peptides are used as a basis for in silico docking and virtual screening against protein targets of interest. In some embodiments, a linear or continuous binding epitope is computationally grafted into a lasso peptide structure. In some embodiments, a conformational epitope is computationally grafted into a lasso peptide structure.
[0149] In some embodiments, a plurality of target protein binding epitopes from the same or different polypeptide ligands are computationally grafted into a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of binding epitopes are computationally grafted into the loop of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of binding epitopes are computationally grafted into the ring of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of binding epitopes are computationally grafted into the tail of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of binding epitopes are computationally grafted into the loop, ring, and / or tail of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop and ring of a lasso peptide structure to create a library of lasso epitope graft variants In some embodiments, a plurality of conformational epitopes are computationally grafted into the ring and tail of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop and tail of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop, ring, and tail of a lasso peptide structures to create a library of lasso epitope graft variants.
[0150] In some embodiments, a plurality of target protein binding epitopes from different polypeptide ligands are experimentally grafted into a lasso peptide sequence to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are experimentally grafted into a lasso peptide structure to create a library of lasso epitope graft variants. 49 ACTIVE 705331286v1
[0151] In some embodiments, a plurality of target protein binding epitopes from the same or different polypeptide ligands are computationally grafted into a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are computationally grafted into a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop and ring of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the ring and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop, ring, and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants.
[0152] In one aspect, provided herein are computer-based methods for identifying a lasso peptide having optimal binding with a target molecule. In some embodiments, the method comprises the steps of (a) providing one or more three-dimensional (3D) model structures of lasso peptides; (b) docking the one or more 3D model structures of lasso peptides onto a 3D model structure of the target molecule at a lasso-binding site of the target molecule, thereby obtaining an optimal docked system; (c) selecting the 3D model structure of the lasso peptide forming the optimal docked system; (d) mapping a target-binding epitope onto the selected 3D model structure of the lasso peptide; and (e) computationally grafting the target-binding epitope onto the mapped location of the selected lasso peptide, thereby generating a lasso peptide binder candidate.
[0153] In one aspect, provided herein are computer-based methods for identifying a lasso peptide having optimal binding with a target molecule. In some embodiments, the method comprises the steps of (a) providing one or more three-dimensional (3D) model structures of lasso peptides; (a-1) mapping a target-binding epitope onto different locations of one or more 3D model structures of lasso peptides; (b) performing conformational analysis on the one or more 3D model structures of lasso peptides using molecular dynamic simulation algorithms to obtain at least two conformational states of the one or more 3D model structures of lasso peptides; (c) docking the at least two conformational states of the one or more 3D model 50 ACTIVE 705331286v1structures of lasso peptides onto a 3D model structure of the target molecule at a lasso- binding site of the target molecule to form at least two 3D model structures of the lasso peptide-target molecule complex; (d) performing conformational analysis on the at least two 3D model structures of the lasso peptide-target molecule complex using molecular dynamic simulation algorithms to obtain at least two conformational states of the 3D model structure of the lasso peptide-target molecule complex; and (e) selecting the 3D model structure of the lasso peptide-target molecule complex having a favored free energy, thereby identifying a lasso peptide binder candidate.
[0154] In one aspect, provided herein are computer-based methods for identifying a lasso peptide having optimal binding with a target molecule. In some embodiments, the method comprises the steps of (a) providing one or more three-dimensional (3D) model structures of lasso peptides; (a-1) mapping a target-binding epitope onto different locations of one or more 3D model structures of lasso peptides; (b) performing conformational analysis on the one or more 3D model structures of lasso peptides using molecular dynamic simulation algorithms to obtain at least two conformational states of the one or more 3D model structures of lasso peptides; (c) docking the at least two conformational states of the one or more 3D model structures of lasso peptides onto a 3D model structure of the target molecule at a lasso- binding site of the target molecule to form at least two 3D model structures of the lasso peptide-target molecule complex; (d) performing conformational analysis on the at least two 3D model structures of the lasso peptide-target molecule complex using molecular dynamic simulation algorithms to obtain at least two conformational states of the 3D model structure of the lasso peptide-target molecule complex; (e) selecting the 3D model structure of the lasso peptide-target molecule complex having a favored free energy; and (f) computationally grafting the target-binding epitope on a mapped location of the selected lasso peptide, thereby identifying a grafted lasso peptide binder candidate.
[0155] According to the present disclosure, providing 3D model structures of a protein or peptide can be performed by retrieving known structural information of the protein or peptide from a protein structure database, such as the worldwide Protein Data Bank (wwPDB) (website: www.wwpdb.org), the Cambridge Structure Database (CSD) (website: www.ccdc.cam.ac.uk / solutions / csd-core / components / csd), Molecular Model Database (MMDB) of National Center for Biotechnology Information (NCBI) (website: www.ncbi.nlm.nih.gov / Structure / MMDB / docs / mmdb_help.html), and the Biological Magnetic Resonance Data Bank (BMRB) (website: bmrb.io). In some embodiments, the protein structure database contains readily available modeled 3D structure of a peptide or 51 ACTIVE 705331286v1protein. In other embodiments, the protein structure database contains atomic coordinates of a peptide or protein that can be used to create a 3D model structure of the protein or peptide using a computer modeling algorithm.
[0156] Accordingly, in some embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of the present method comprises retrieving one or more known 3D structures of lasso peptide from a protein structure database (FIG.2B). In particular embodiments, the protein structure database is selected from the Protein Data Bank (wwPDB), the Cambridge Structure Database, the Molecular Model Database of National Center for Biotechnology Information (NCBI), and the Biological Magnetic Resonance Data Bank (BMRB) database.
[0157] In alternative embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of the present method comprises computationally modeling the 3D structure of a lasso peptide based on atomic coordinates of the lasso peptide. In particular embodiments, the atomic coordinates of the lasso peptide are obtained from a protein structure database or scientific literature. In yet particular embodiments, the protein structure database is selected from a protein structure database. In particular embodiments, the protein structure database is selected from the Protein Data Bank (wwPDB), the Cambridge Structure Database, the Molecular Model Database of National Center for Biotechnology Information (NCBI), and the Biological Magnetic Resonance Data Bank (BMRB) database.
[0158] In some embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of the present method comprises computationally modeling the 3D structure of a lasso peptide based on atomic coordinates of the lasso peptide. In particular embodiments, the atomic coordinates of the lasso peptide are obtained by subjecting the lasso peptide to nuclear magnetic resonance (NMR) analysis, X-ray crystallography, neutron diffraction, or 3-dimensional electron microscopy (3D-EM). In specific embodiments, the atomic coordinates of the lasso peptide are obtained by subjecting the lasso peptide to serial femtosecond crystallography. In specific embodiments, the atomic coordinates of the lasso peptide are obtained by subjecting the lasso peptide to cryogenic electron microscopy (cryo- EM).
[0159] In some embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of the present method comprises computationally modeling the 3D structure of a lasso peptide based on atomic coordinates of the lasso peptide. In some embodiments, the computational modeling is performed by molecular replacement modeling. In specific embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of 52 ACTIVE 705331286v1the present method comprises computationally modeling the 3D structure of the lasso peptide based on X-ray diffraction data and / or nuclear magnetic resonance (NMR) data of the lasso peptide and atomic coordinates of a reference lasso peptide.
[0160] In some embodiments, the X-ray diffraction data of the lasso peptide are obtained by subjecting a crystal of the lasso peptide to X-ray crystallography analysis. In particular embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of the present method comprises (i) obtaining atomic coordinates of the lasso peptide based on the X-ray diffraction data; and (ii) refining the atomic coordinates of the lasso peptide based on the atomic coordinates of the reference lasso peptide. In specific embodiments, the step of (ii) refining the atomic coordinates of the lasso peptide is performed using a molecular replacement method.
[0161] In some embodiments, the NMR data of the lasso peptide comprises NMR chemical shift, J-coupling constant, and resonance intensity obtained by subjecting a solution of the lasso peptide to NMR analysis. In particular embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of the present method comprises (i) creating an ensemble of structural models of the lasso peptide based on the NMR data; (ii) obtaining mean atomic coordinates of the lasso peptide based on the ensemble of structural models; and (iii) refining the atomic coordinates of the lasso peptide based on the atomic coordinates of the reference lasso peptide.
[0162] In some embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of the present method comprises computationally modeling the 3D structure of a lasso peptide based on the amino acid sequence of the lasso peptide and atomic coordinates of a reference lasso peptide, and wherein the lasso peptide has at least 50 percent (%) sequence identity to the reference lasso peptide.
[0163] In some embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of the present method comprises computationally modeling the 3D structure of a lasso peptide based on the amino acid sequence of the lasso peptide and atomic coordinates of a reference lasso peptide; and wherein the lasso peptide has at least 50 percent (%) amino acid sequence identity to the reference lasso peptide. In some embodiments, computationally modeling the 3D structure of the lasso peptide is performed by a homology modeling algorithm.
[0164] In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 50 percent (%) amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 53 ACTIVE 705331286v155% amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 60% amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 65% amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 70% amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 75% amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 80% amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 85% amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 90% amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 95% amino acid sequence identity to the reference lasso peptide. In specific embodiments involving a reference lasso peptide, the lasso peptide has at least about 97% amino acid sequence identity to the reference lasso peptide.
[0165] According to the present disclosure, in some embodiments, a 3D model structure of a lasso peptide can be a lasso peptide backbone structure. In some embodiments, a lasso peptide backbone structure can be constructed by computationally modeling the 3D structure of a lasso peptide and further computationally modeling the lasso backbone structure by removing side chains of each amino acid residue that is not an internal ring-forming residue from the 3D structure of lasso peptides.
[0166] Accordingly, in some embodiments, the step of (a) providing one or more 3D model structures of lasso peptides of the present method further comprises computationally modeling the lasso backbone structures by removing side chains of each amino acid residue that is not an internal ring-forming residue from the modeled 3D structure of lasso peptides before proceeding to step (c) of the present method.
[0167] In some embodiments, step (c) of the method comprises creating the 3D model structure of the target molecule before docking the one or more 3D model structure of lasso peptide onto the 3D model structure of the target molecule. In some embodiments, the target molecule has one or more lasso binding site on its surface, and step (c) of the method comprises docking the one or more 3D model structure of lasso peptide onto the lasso- binding site of the target molecule. 54 ACTIVE 705331286v1
[0168] In some embodiments, the target molecule is capable of binding to a ligand, and the lasso-binding site of the target molecule can be the ligand-binding site where the ligand binds to the target molecule.
[0169] In some embodiments, the target molecule is an enzyme capable of binding and catalyzing a reaction of the substrate. Particularly, the enzyme contains an active site which is the region where substrate molecules bind and undergo a chemical reaction. In some embodiments where the target molecule is an enzyme, the active site consists of amino acid residues that form temporary bonds with the substrate (substrate binding site) and residues that catalyze a reaction of that substrate (catalytic site). Accordingly, in some embodiment, the lasso-binding site of a target molecule can be a substrate binding site of the target molecule. In some embodiments, the lasso-binding site of a target molecule can be a catalytic site of the target molecule.
[0170] In some embodiments, the target molecule can be an enzyme or non-enzyme protein that uses one or more cofactor molecules to activate, inhibit, or otherwise perform its function. In some embodiments, the target molecule is capable of binding to the cofactor molecules, and the region in the target molecule where the cofactor binds is a cofactor binding site of the target molecule. Accordingly, in some embodiments, the lasso-peptide binding site of a target molecule can be the cofactor binding site of the target molecule.
[0171] In some embodiments, the target molecule is a receptor that has an orthosteric site where a ligand molecule binds. Accordingly, in some embodiments, the lasso-peptide binding site of a target molecule can be an orthosteric site of the target molecule. In some embodiment, the target molecule is a receptor that has at least one allosteric site where an allosteric effector molecule binds. Accordingly, in some embodiments, the lasso-peptide binding site of a target molecule can be an allosteric site of the target molecule.
[0172] In some embodiments, the target molecule is capable of switching between an active conformation (open conformation) and an inactive conformation (closed conformation) (FIG.14A and FIG.14B). Accordingly, in some embodiments, the lasso-peptide binding site of a target molecule can be a binding site that exists in the active or open conformation of the target peptide. In alternative embodiments, the lasso-peptide binding site of a target molecule can be a binding site that exists in the inactive or closed conformation of the target peptide.
[0173] In some embodiments, the step (c) of the present method comprises creating the 3D model structure of the target molecule before docking. Particularly, in some embodiments, creating the 3D model structure of the target molecule comprises 55 ACTIVE 705331286v1computationally modeling the 3D structure of the target molecule based on atomic coordinates of the target molecule. In some embodiments, the atomic coordinates of the target molecule are known. In some embodiments, the atomic coordinates of the target molecule can be obtained from a protein structure database or scientific literature. In some embodiments, the atomic coordinates of the target molecule can be obtained from protein structure database selected from the Protein Data Bank (wwPDB), the Cambridge Structure Database, the Molecular Model Database of National Center for Biotechnology Information (NCBI), and the Biological Magnetic Resonance Data Bank (BMRB) database.
[0174] In some embodiments, the atomic coordinates of the target molecule are obtained by subjecting the target molecule to nuclear magnetic resonance (NMR) analysis, X-ray crystallography, neutron diffraction analysis, or 3-dimensional electron microscopy (3D- EM). In particular embodiments, the atomic coordinates of the target molecule are obtained by subjecting the target molecule to serial femtosecond crystallography. In particular embodiments, the atomic coordinates of the target molecule are obtained by subjecting the target molecule to cryogenic electron microscopy (cryo-EM).
[0175] In some embodiments, step (c) of the present method comprises creating the 3D model structure of the target molecule before docking. Particularly, in some embodiments, creating the 3D model structure of the target molecule comprises computationally modeling the 3D structure of the target molecule based on X-ray diffraction data and / or nuclear magnetic resonance data of the target molecule and atomic coordinates of a reference polypeptide.
[0176] In some embodiments, the X-ray diffraction data of the target molecule are obtained by subjecting a crystal of the target molecule to X-ray crystallography analysis. In some embodiments, creating the 3D model structure of the target molecule comprises: (i) obtaining atomic coordinates of the target molecule based on the X-ray diffraction data; (ii) refining the atomic coordinates of the target molecule based on the atomic coordinates of the reference polypeptide; and (iii) computationally modeling the 3D structure of the target molecule based on the refined atomic coordinates. In particular embodiments, the refining step is performed using a molecular replacement algorithm. In particular embodiments, the atomic coordinates of the reference polypeptide are known. In particular embodiments, the atomic coordinates of the reference polypeptide are obtained from a protein structure database selected from the Protein Data Bank (wwPDB), the Cambridge Structure Database, the Molecular Model Database of National Center for Biotechnology Information (NCBI), and the Biological Magnetic Resonance Data Bank (BMRB) database. In some 56 ACTIVE 705331286v1embodiments, the atomic coordinates of the reference polypeptide are obtained from scientific literature.
[0177] In some embodiments, the nuclear magnetic resonance (NMR) data of the target molecule comprises NMR chemical shift, J-coupling constant, and resonance intensity obtained from subjecting a solution of the target molecule to NMR analysis. In some embodiments, creating the 3D model structure of the target molecule comprises (i) creating an ensemble of structural models of the target molecule based on the NMR data; (ii) obtaining mean atomic coordinates of the target molecule based on the ensemble of structural models; (iii) refining the mean atomic coordinates of the target molecule based on the atomic coordinates of the reference polypeptide; and (iv) computationally modeling the 3D structure of the target molecule based on the refined mean atomic coordinates. In particular embodiments, the refining step is performed using a molecular replacement algorithm. In particular embodiments, the atomic coordinates of the reference polypeptide are known. In particular embodiments, the atomic coordinates of the reference polypeptide are obtained from a protein structure database selected from the Protein Data Bank (wwPDB), the Cambridge Structure Database, the Molecular Model Database of National Center for Biotechnology Information (NCBI), and the Biological Magnetic Resonance Data Bank (BMRB) database. In some embodiments, the atomic coordinates of the reference polypeptide are obtained from scientific literature.
[0178] In some embodiments, step (c) of the present method comprises creating the 3D model structure of the target molecule before docking. In some embodiments, creating the 3D model structure of the target molecule comprises computationally modeling the 3D structure of the target molecule based on the amino acid sequence of the target molecule and atomic coordinates of one or more reference polypeptide. In some embodiments, computationally modeling the 3D structure of the target molecule is performed by homology modeling. In some embodiments, computationally modeling the 3D structure of the target molecule is performed by protein structure prediction. In some embodiments, computationally modeling the 3D structure of the target molecule is performed by a combination of protein structure prediction and homology modeling. In some embodiments, the atomic coordinates of the reference polypeptide are known. In some embodiments, the atomic coordinates of the reference polypeptide are obtained from a protein structure database or scientific literature. In some embodiments, the protein structural database is selected from worldwide Protein Data Bank (wwPDB), Cambridge Structure Database, 57 ACTIVE 705331286v1Molecular Model Database of National Center for Biotechnology Information (NCBI), and Biological Magnetic Resonance Data Bank (BMRB) database.
[0179] In specific embodiments involving a reference polypeptide, the target molecule has at least about 50 percent (%) amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 55% amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 60% amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 65% amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 70% amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 75% amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 80% amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 85% amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 90% amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 95% amino acid sequence identity to the reference peptide. In specific embodiments involving a reference polypeptide, the target molecule has at least about 97% amino acid sequence identity to the reference peptide.
[0180] In some embodiments, step (c) of the present method comprises docking the one or more 3D model structures of lasso peptides onto a 3D model structure of the target molecule at a lasso-binding site of the target molecule, thereby obtaining an optimal docked system. In some embodiments, the docking comprises, for each 3D model structure of the lasso peptide: (c-1) positioning a docked portion of at least one conformational state of the 3D model structure of the lasso peptide relative to the lasso-binding site of the 3D model structure of the target molecule in a first docked pose, thereby forming a docked system; (c-2) scoring the docked system based on structural complementarity between the docked portion and the lasso-binding site; (c-3) repositioning the docked portion relative to the lasso-binding site to a second docked pose, and repeating step (c-2); (c-4) repeating step (c-3) for one or more times; and (c-5) selecting the docked system having the highest score as the optimal docked system before proceeding to step (d). 58 ACTIVE 705331286v1
[0181] In some embodiments, the docked portion of the 3D model structure comprises at least one amino acid residue from the loop portion of the lasso peptide. In some embodiments, the docked portion of the 3D model structure comprises at least one amino acid residue from the ring portion of the lasso peptide. In some embodiments, the docked portion of the 3D model structure comprises at least one amino acid residue from the tail portion of the lasso peptide. In some embodiments, the docked portion of the 3D model structure comprises at least one amino acid residue from the loop portion of the lasso peptide and at least one amino acid residue from the ring portion of the lasso peptide. In some embodiments, the docked portion of the 3D model structure comprises at least one amino acid residue from the loop portion of the lasso peptide and at least one amino acid residue from the tail portion of the lasso peptide. In some embodiments, the docked portion of the 3D model structure comprises at least one amino acid residue from the ring portion of the lasso peptide and at least one amino acid residue from the tail portion of the lasso peptide. In some embodiments, the docked portion of the 3D model structure comprises at least one amino acid residue from the loop portion of the lasso peptide, at least one amino acid residue from the ring portion of the lasso peptide, and at least one amino acid residue from the tail portion of the lasso peptide.
[0182] In some embodiments, the docked system comprises one or more complementary binding pairs comprising a first binding moiety located on the lasso peptide, and a second binding moiety located on the target molecule. In some embodiments, in a docked system, the first and second binding moieties form binding interaction with one another. In some embodiments, the binding interaction between the first and second binding moieties is non- covalent interaction. In some embodiments, the binding interaction between the first and second binding moieties is selected from hydrogen bonding interactions, ionic or electrostatic bonding interactions, pi stacking interactions, polar interactions, dipolar interactions, induced dipolar interactions, hydrophobic interactions, and van der Waals interactions.
[0183] In some embodiments, in a docked system, the first binding moiety is located on a first amino acid residue of the lasso peptide, and the second binding moiety is on a second amino acid residue located on the target molecule. In some embodiments, the distance between any atom of the first amino acid residue and any atom of the second amino acid residue is less than about 5 Ångströms. In some embodiments, the distance between any atom of the first amino acid residue and any atom of the second amino acid residue is less than about 4 Ångströms. In some embodiments, the distance between any atom of the first 59 ACTIVE 705331286v1amino acid residue and any atom of the second amino acid residue is less than about 3 Ångströms.
[0184] In some embodiments, step (c-2) scoring the docked system based on structural complementarity between the docked portion and the lasso-binding site comprises calculating a total binding free energy of the docked system using one or more molecular mechanics force field functions, and assigning a score to the docked system based on the total binding free energy. In some embodiments, the assigned score to a docked system negatively relates to the total binding free energy of the docked system. In some embodiments, a relatively higher score is assigned to a docked system having a relatively lower total binding free energy. In some embodiments, a relatively lower score is assigned to a docked system having a relatively higher total binding free energy.
[0185] Various computational modeling platforms that are known in the art can be used for calculating the free energy differences between different docked molecules and / or poses and total binding free energy of a docked system using molecular mechanics force field functions and scoring the docked system. Exemplary computational modeling platforms include but are not limited to the MOE 2020.09 computational modeling platform (Molecular Operating Environment, Chemical Computing Group Inc. Canada). A variety of force fields and force field parameters can be used in connection with the present disclosure, including but not limited to Amber force fields AMBER99, Amber 10EHT, ff14SB, and ff19SB, with amino acids-specific side chain and protein backbone parameters (See: Ponder, J.W., Case, D.A. “Force fields for protein simulations.” Adv. Prot. Chem.2003, 66, 27-85; Tian, C., et al. “ff19SB: Amino-Acid-Specific Protein Backbone Parameters Trained against Quantum Mechanics Energy Surfaces in Solution.” J. Chem. Theory Comput., 2019, 16, 528-552; J.A. Maier, et al. “ff14SB: Improving the accuracy of protein side chain and backbone parameters from ff99SB.” J. Chem. Theory Comput., 2015, 11, 3696-3713) and the MMFF94x force field (See: Halgren, T.A. Merck molecular force field. I. Basis, form, scope, parameterization, and performance of MMFF94, J. Comput. Chem.; 1996; 490-519). Accordingly, in some embodiments, in step (c-2), scoring the docked system based on structural complementarity between the docked portion and the lasso-binding site comprises calculating a total binding free energy of the docked system using one or more molecular mechanics force field functions selected from Amber force fields AMBER99, Amber 10EHT, ff14SB, ff19SB, and the Merck force field MMFF94x, and assigning a score to the docked system based on the lower total binding free energy. In some embodiments, the score assigned to a docked system negatively relates to the total binding free energy. In some 60 ACTIVE 705331286v1embodiments, the favored free energy includes a lower free energy compared to a different 3D model structure of the lasso peptide-target molecule complex. In some embodiments, the favored free energy is the lowest free energy of the two or more 3D model structures of the lasso peptide-target molecule complex. In some embodiments, a relatively higher score is assigned to a docked system having a relatively lower total binding free energy. In some embodiments, a relatively lower score is assigned to a docked system having a relatively higher total binding free energy.
[0186] In some embodiments, step (c-3) repositioning the docked portion relative to the lasso-binding site to a second docked pose comprises identifying the second docked pose using an energy minimizing function before repositioning the docked portion, wherein the docked system is predicted to have a lower binding free energy in the second docked pose than in the first docked pose based on the energy minimizing function. Various energy minimizing function known in the art can be used in connection with the present disclosure, including but not limited to the steepest decent algorithm (Jaidhan, B.J., et al. Energy minimization and conformational analysis of molecules using steepest descent method, (IJCSIT) Int J Comp Sci Inf Tech, 2014, 5(3), 3525-3528), the conjugate gradient algorithm (Knyazev, A.V., Lashuk, I. “Steepest Descent and Conjugate Gradient Methods with Variable Preconditioning”. SIAM Journal on Matrix Analysis and Applications.2008, 29(4): 1267-1280), the truncated Newton algorithms (Nash, S. G. A survey of truncated-Newton methods. J. Comput. Appl. Math., 2000, 124(l-2): 45-59), the L-BFGS (limited-memory Broyden-Fletcher-Goldfarb-Shanno) algorithm (Byrd, R. H.; Lu, P.; Nocedal, J.; Zhu, C. A Limited Memory Algorithm for Bound Constrained Optimization. J. Sci. Comput. 1995, 16 (5): 1190–1208), and / or the genetic algorithm (Le Grand, S.M., et al. The Application of the Genetic Algorithm to the Minimization of Potential Energy Functions. J. Global Opt., 1993, 3(1), 49–66).
[0187] Accordingly, in some embodiments, the energy minimizing function is selected from the steepest descent algorithm, conjugate gradients algorithm, L-BFGS (limited- memory Broyden-Fletcher-Goldfarb-Shanno) algorithm, and genetic algorithms.
[0188] In some embodiments, step (a-1) mapping a target-binding epitope onto the selected 3D model structure of the lasso peptide comprises (a-1-1) in the optimal docked system, identifying one or more lasso-binding moieties in the lasso-binding site of the target molecule; (a-1-2) selecting one or more amino acid residues comprising one or more target- binding moieties complementary to the one or more lasso-binding moieties; and (a-1-3) mapping an optimal set of positions in the amino acid sequence of the selected lasso peptide 61 ACTIVE 705331286v1for grafting the one or more selected amino acid residues; wherein the grafting places the one or more target-binding moieties at suitable spatial locations and orientations for binding with the complementary lasso-binding moieties on the target molecule. In some embodiments, the binding interaction between the complementary lasso-binding moiety and the target-binding moiety is selected from covalent bonding interactions or non-covalent interactions including hydrogen bonding interactions, ionic or electrostatic bonding interactions, polar interactions, dipolar interactions, induced dipolar interactions, pi stacking interactions, H-bond-aromatic interactions, hydrophobic interactions, and van der Waals interactions.
[0189] In some embodiments, step (a-1-3) mapping an optimal set of positions in the amino acid sequence of the selected lasso peptide for grafting the one or more selected amino acid residues comprises: (a-1-3-1) computationally grafting the one or more selected amino acid residues into the amino acid sequence of the selected lasso peptide at a first set of positions; (a-1-3-2) calculating a total binding free energy of the optimal docked system using one or more molecular force field functions; (a-1-3-3) modifying at least one position in the first set of positions thereby obtaining an adjusted set of positions, and computationally grafting the one or more selected amino acid residues into the amino acid sequence of the selected lasso peptide at the adjusted set of positions; (a-1-3-4) repeating step (a-1-3-2); (a- 1-3-5) repeating steps (a-1-3-3) and (a-1-3-4) sequentially for one or more times; and (a-1-3- 6) selecting the adjusted set of positions associated with the lowest total binding free energy as the optimal map of positions.
[0190] In some embodiments, the molecular mechanics force field functions are selected from Amber force fields AMBER99, Amber 10EHT, ff14SB, ff19SB, and Merck force field MMFF94x.
[0191] In some embodiments, in step (a-1-3-5), wherein repeating steps (a-1-3-3) and (a- 1-3-4) sequentially for one or more times comprising repeating steps (a-1-3-3) and (a-1-3-4) for at least 1 time, at least twice, at least three times, at least four times, at least five times, at least six times, at least seven times, at least eight times, at least nine times, at least ten times, at least fifteen times, at least twenty times, or at least twenty-five times.
[0192] In some embodiments, step (a-1-3-3) comprises identifying the adjusted set of positions using an energy minimizing function before modifying the first set of positions to become the adjusted set of positions, wherein the docked system having the one or more selected amino acid residues grafted into the amino acid sequence of the selected lasso peptide at the adjusted set of positions is predicted to have a lower binding free energy than at the first set of positions based on the energy minimizing function. In some embodiments, the 62 ACTIVE 705331286v1energy minimizing function is selected from the steepest descent algorithm, conjugate gradients algorithm, L-BFGS (limited-memory Broyden-Fletcher-Goldfarb-Shanno) algorithm, and genetic algorithms.
[0193] In some embodiments, the target molecule has a naturally-existing ligand to which the target molecule binds. Accordingly, in some embodiments, the target-binding epitope to be mapped onto the 3D model structure of the lasso peptide corresponds to a continuous fragment of the naturally-existing ligand that binds to the target molecule. In some embodiments, the target-binding epitope is a linear epitope (FIG.3A). In some embodiments, the target-binding epitope has 100% amino acid sequence identity to the corresponding continuous fragment of the naturally–existing ligand of the target molecule. In some embodiments, the amino acid sequence of the target-binding epitope differ from the corresponding continuous fragment of the naturally-existing ligand of the target molecule by 1, 2, 3, 4, 5, or more than 5 amino acid residue.
[0194] Alternatively, in some embodiments, the target-binding epitope to be mapped onto the 3D model structure of the lasso peptide comprises multiple discontinuous fragments that respectively correspond to multiple discontinuous fragments of the naturally-existing ligand to which the target molecule binds (FIG.4). In some embodiments, the target-binding epitope is a conformational epitope. In some embodiments, the multiple discontinuous fragments of the target-binding epitope has 100% amino acid sequence identity to the corresponding multiple discontinuous fragments of the naturally-existing ligand. In some embodiments, the multiple discontinuous fragments of the target-binding epitope differ from the corresponding multiple discontinuous fragments of the naturally-existing ligand by 1, 2, 3, 4, 5, or more than 5 amino acid residue.
[0195] In some embodiments, the target molecule has a naturally-existing ligand, wherein the target-binding epitope corresponds to a fragment or fragments of a naturally-existing ligand of the target molecule that are capable of binding with the lasso-binding site of the target molecule and forming a ligand-target interface. In these embodiments, step (d) of the present method further comprises aligning the docked portion of the 3D model structure of the lasso peptide with a 3D model structure of the corresponding fragment or fragments of the naturally-existing ligand in the ligand-target interface. In particular embodiments, the aligning comprises adjusting the spatial position, conformation and / or orientation of at least one target-binding moiety in the docked portion of the lasso peptide to mimic the spatial position, conformation and / or orientation of the corresponding binding moiety of the fragment or fragments of the naturally-existing ligand in the ligand-target interface. 63 ACTIVE 705331286v1
[0196] In some embodiments, the present method further comprises (f) mutating one or more amino acid residues of the lasso peptide binder candidate to produce a first set of lasso peptide binder variants; and (g) ranking the first set of lasso peptide binder variants based on a predicted binding affinity for binding with the target molecule. In particular embodiments, step (f) further comprises adjusting conformations of the first set of lasso peptide binder variants to produce a second set of lasso peptide binder variants. In particular embodiments, step (g) further comprises ranking the second set of lasso peptide binder variants based on the predicted binding affinity for binding with the target molecule. In specific embodiments, in step (f), mutating the amino acid residue of the lasso peptide binder candidate comprises replacing the side chain of the amino acid residue of the lasso peptide binder candidate with the side chain of a second amino acid that is different from the mutated amino acid residue. In specific embodiments, the second amino acid is a naturally-occurring or a non-natural amino acid.
[0197] In some embodiments, in step (f), mutating the amino acid residue of the lasso peptide binder candidate comprises modifying one or more chemical moieties on the side chain of the mutated amino acid residue. In some embodiments, at least one modified chemical moiety is a target-binding moiety.
[0198] In some embodiments, in step (f), mutating the amino acid residue of the lasso peptide binder candidate comprises modifying one or more side chains of the lasso peptide binder candidate to complement one or more lasso-binding moieties on the target molecule; and wherein the modifying is selected from the group consisting of: (i) incorporating into the lasso peptide a neutral or basic side chain that is structurally opposite to an acidic lasso- binding moiety; (ii) incorporating into the lasso peptide a neutral or acid side chain that is structurally opposite to a basic lasso-binding moiety; (iii) incorporating into the lasso peptide a neutral or positively charged side chain that is structurally opposite to a negatively charged lasso-binding moiety; (iv) incorporating into the lasso peptide a neutral or negatively charged side chain that is structurally opposite to a positively charged lasso-binding moiety; (v) incorporating into the lasso peptide a sterically smaller side chain that is structurally opposite to a sterically larger lasso-binding moiety; (vi) incorporating into the lasso peptide a sterically larger side chain that is structurally opposite to a sterically smaller lasso-binding moiety; (vii) incorporating into the lasso peptide a hydrophobic side chain, preferably a similarly hydrophobic side chain, that is structurally opposite to a hydrophobic lasso-binding moiety; (viii) incorporating into the lasso peptide an acidic, positively charged, or aromatic side chain that is structurally opposite to an aromatic lasso-binding moiety; (ix) incorporating into 64 ACTIVE 705331286v1the lasso peptide an electron-rich aromatic side chain that is structurally opposite to an electron-poor aromatic lasso-binding moiety; (x) incorporating into the lasso peptide an acid or electron-poor aromatic side chain that is structurally opposite to an electron-rich aromatic lasso-binding moiety; (xi) incorporating into the lasso peptide an inducible dipole or multipole side chain that is structurally opposite to an inducible dipole or multipole lasso- binding moiety; (xii) incorporating into the lasso peptide a permanent, inducible dipole or multipole side chain that is structurally opposite to an inducible or multipole lasso-binding moiety; (xiii) incorporating into the lasso peptide an H-bond acceptor side chain that is structurally opposite to an H-bond donor lasso-binding moiety, and (xiv) incorporating into the lasso peptide an H-bond donor side chain that is structurally opposite to an H-bond acceptor lasso-binding moiety.
[0199] In some embodiments, the mutated amino acid residue is in the (i) ring portion of the lasso peptide; (ii) loop portion of the lasso peptide; (iii) tail portion of the lasso peptide; or (iv) any combination of (i) to (iii).
[0200] In some embodiments, the method further comprises step (h) synthesizing the lasso peptide binder candidate, grafted lasso peptide binder candidate or one or more lasso peptide binder variants having the highest rankings. In particular embodiments, in step (h), synthesizing the lasso peptide binder candidate or the one or more lasso peptide binder variants is performed using a cell-free or cell-based synthesis method.
[0201] In some embodiments, the method further comprises creating a database of optimized lasso peptide structures or intermediates thereof generated in any one of steps (a) to (h).
[0202] According to the present disclosure, de novo design of lasso peptides represents an alternative computational strategy for discovering new lasso peptides that can bind with high affinity to target proteins. De novo drug design is based on stochastic structure optimization, evolutionary fitness functions, and combinatorial design principles that potentially represent a vast chemical space. In order to manage the large numbers of possible structures, constraints are placed on the de novo design space. In general, receptor-based constrains or ligand-based constraints are implemented (Schneider, G., Fechner, U., Computer-based de novo design of drug-like molecules, Nat Rev Drug Discov, 2005, 4, 649- 663). All information that is related to the ligand–receptor interaction forms the primary target constraints for candidate compounds. Such constraints can be gathered both from the three-dimensional receptor structure and from the structures of known ligands of the particular target. Receptor-based design strategies starts with the determination of known 65 ACTIVE 705331286v1binding sites for known ligands and potential binding sites for new ligands. As complementarities in molecular shape and physical and chemical properties are important for specific binding, the binding site is then examined to derive shape constraints for a ligand, as well as specific non-covalent ligand–receptor interactions in the form of hypothetical interaction sites created through hydrogen bonding, van der Waals, electrostatic, and hydrophobic interactions. Receptor groups capable of hydrogen-bonding are of special interest owing to the strongly directional nature of the two interaction partners (i.e., hydrogen-bond acceptor and donor) and often form key interaction sites, especially in protein-protein or protein-peptide complexes. These constraints allow the assignment of ligand atom positions with a complementary hydrogen-bond type within a small region of space and a defined orientation. Key interaction sites have a major role in the effort to reduce the vast number of possible structures because they define strong and explicit requirements for successful receptor–ligand binding. For ligand generation, energetically favorable positions and orientations of functional groups in the binding site are determined. Functional groups (e.g., amino acids with side chains) are placed inside the binding pocket and fragments are then minimized simultaneously using a force field. Groups are discarded if the interaction energy between them and the protein is above a certain threshold. A de novo ligand design run yields a set of pre-docked fragments or sets of functional groups that can be further investigated through docking to choose the most promising ones. This placement of chemical groups provides a starting point for the assembly of complete ligands and scoring the ligand-receptor binding.
[0203] In some embodiments, the target binding epitope is designed de novo by creating a set of amino acids that complements those in a potential lasso peptide binding site of a target protein. In some embodiments, a linear or three-dimensional (3D) arrangement of a set of amino acids is identified as a potential lasso binding site in a target protein. In some embodiments, the linear or 3D arrangement consists of 2-10 amino acids. In some embodiments, a complementary set of amino acids are identified for interacting with the amino acids in the target binding site. In other embodiments, the complementary set of amino acids have properties that facilitate attractive or non-repulsive interactions with the amino acids in the target binding site. In other embodiments, the complementary set of amino acids is grafted onto a lasso peptide scaffold.
[0204] In one aspect, provided herein are methods to identify and computationally graft a target protein-binding epitope of a natural or synthetic polypeptide ligand into a lasso peptide 3-dimensional (3D) topological structure. In another aspect, provided herein are methods to 66 ACTIVE 705331286v1identify and experimentally graft a target protein-binding epitope of a natural or synthetic polypeptide ligand into a lasso peptide 3-dimensional (3D) topological structure. In some embodiments, lasso peptides containing epitope grafted segments are identified and modeled using computational algorithms that enable in silico docking of lasso peptide structures containing such epitope grafts into a lasso binding site of protein structures. In some embodiments, lasso peptides containing epitope grafted segments are modeled and binding interactions are optimized using computational algorithms that enable in silico docking and conformational molecular dynamic modeling of lasso peptide structures that are scored and ranked on the basis of predicted binding affinity in a lasso binding site of protein structures. In some embodiments, lasso peptides containing epitope grafted segments are identified using computational algorithms that enable in silico docking of lasso peptide structures into a lasso binding site of protein structures and the lasso peptide-protein interactions are further refined and ranked using artificial intelligence algorithms. In some embodiments, lasso peptides containing epitope grafted segments are identified and binding interactions are optimized using computational and artificial intelligence algorithms that enable in silico docking and conformational molecular dynamic modeling of lasso peptide structures that are scored and ranked on the basis of predicted binding affinity in a lasso binding site of protein structures.
[0205] In some embodiments, a binding epitope, which is identified by analyzing the binding interactions between a natural or synthetic polypeptide ligand and a target protein, is computationally grafted into a lasso peptide structure. In some embodiments a binding epitope is grafted into the loop of a lasso peptide. In some embodiments a binding epitope is grafted into the ring of a lasso peptide. In some embodiments a binding epitope is grafted into the tail of a lasso peptide. In some embodiments, a binding epitope is grafted into both the loop and ring of a lasso peptide. In some embodiments, a binding epitope is grafted into both the loop and tail of a lasso peptide. In some embodiments, a binding epitope is grafted into both the ring and tail of a lasso peptide. In some embodiments, a binding epitope is grafted into the loop, ring, and tail of a lasso peptide.
[0206] In some embodiments, a linear or continuous binding epitope is identified whereby contiguous sections (two or more amino acid residues) of a natural or synthetic polypeptide ligand bind to a target protein, and the linear epitope is grafted into a lasso peptide structure (FIG.3A and FIG.3B). In some embodiments, a conformational binding epitope is identified whereby two or more non-contiguous sections (one or more amino acid residues) of a natural or synthetic polypeptide ligand that binds to a target protein is grafted into a lasso peptide structure (FIG.4). In some embodiments, the grafting process involves a 67 ACTIVE 705331286v1replacement of specific amino acids in an existing natural or synthetic lasso peptide sequence with amino acids of a binding epitope. In some embodiments, the grafting process involves an insertion of amino acids into the sequence of an existing natural or synthetic lasso peptide sequence, thereby extending the length of the sequence. In some embodiments, the grafting process involves the de novo creation of lasso peptide sequences containing the binding epitopes of interest in the loop, ring, and / or tail of the lasso peptide.
[0207] In some embodiments, a plurality of target protein binding epitopes from the same or different polypeptide ligands are computationally grafted into a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of epitopes are computationally grafted into the loop of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of epitopes are computationally grafted into the ring of a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of epitopes are computationally grafted into the tail of a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of epitopes are computationally grafted into the loop and ring of a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of epitopes are computationally grafted into the loop and tail of a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of epitopes are computationally grafted into the ring and tail of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of epitopes are computationally grafted into the loop, ring, and tail of a lasso peptide structures to create a library of lasso epitope graft variants.
[0208] In some embodiments, a plurality of conformational epitopes are computationally grafted into a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop and ring of a lasso peptide structure to create a library of lasso epitope graft variants In some embodiments, a plurality of conformational epitopes are computationally grafted into the ring and tail of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop and tail of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of conformational epitopes are computationally grafted into the loop, ring, and tail of a lasso peptide structures to create a library of lasso epitope graft variants. 68 ACTIVE 705331286v1
[0209] In some embodiments, a plurality of target protein binding epitopes from different polypeptide ligands are experimentally grafted into a lasso peptide sequence to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are experimentally grafted into a lasso peptide structure to create a library of lasso epitope graft variants.
[0210] In some embodiments, a plurality of target protein binding epitopes from the same or different polypeptide ligands are computationally grafted into a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are computationally grafted into a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are computationally grafted into the loop and ring of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are computationally grafted into the ring and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are computationally grafted into the loop and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are computationally grafted into the loop, ring, and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are computationally grafted into the loop and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants.
[0211] In some embodiments, a plurality of target protein binding epitopes from the same or different polypeptide ligands are synthetically grafted into a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of binding epitopes are synthetically grafted into the loop of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into the ring of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into the tail of a lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of target protein binding epitopes from different polypeptide ligands are synthetically grafted into a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are 69 ACTIVE 705331286v1synthetically grafted into a lasso peptide structure to create a library of lasso epitope graft variants. In some embodiments, a plurality of target protein binding epitopes from the same or different polypeptide ligands are synthetically grafted into a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into the loop of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into the ring of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into the tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into the loop and ring of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into the loop and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into the ring and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants. In some embodiments, a plurality of 3D binding epitopes are synthetically grafted into the loop, ring, and tail of a plurality of lasso peptide structures to create a library of lasso epitope graft variants.
[0212] In another aspect, computationally identified lasso peptides are synthesized in a host organism that contains the enzymes of a lasso peptide biosynthetic pathway. In some embodiments, computationally identified lasso peptides are synthesized in a host organism that contains the enzymes of a lasso peptide biosynthetic pathway composed of a precursor peptide (A), and lasso peptidase (B or B2), a lasso cyclase (C), and a RiPP recognition sequence (E or B1). In some embodiments, computationally identified lasso peptides are synthesized by treating a precursor peptide (A), with at least one of a lasso peptidase (B or B2), a lasso cyclase (C), and a RiPP recognition sequence (E or B1). In some embodiments, computationally identified lasso peptides are synthesized by treating an isolated precursor peptide (A), with at least one of an isolated lasso peptidase (B or B2), an isolated lasso cyclase (C), and an isolated RiPP recognition sequence (E or B1). In some embodiments, computationally identified lasso peptides are synthesized by adding a precursor peptide (A), with host organism that contains at least one of a lasso peptidase (B or B2), a lasso cyclase 70 ACTIVE 705331286v1(C), and a RiPP recognition sequence (E or B1). In some embodiments, computationally identified lasso peptides are synthesized by contacting the genes for a precursor peptide (A), and at least one of an isolated lasso peptidase (B or B2), an isolated lasso cyclase (C), and an isolated RiPP recognition sequence (E or B1), with a cell extract to produce the identified lasso peptide in a cell-free biosynthesis process.
[0213] In another aspect, computationally modeled and identified lasso peptides are produced as described herein and screened for biological activity in a testing step that involves contacting such lasso peptides with biological targets, proteins, or receptors and measuring a property, such as cell penetration, binding affinity, or inhibition of function or activity. In some embodiments, the testing step allows the ranking of computationally identified and experimentally produced lasso peptides for desired biological activities and / or properties. In some embodiments, computationally designed and experimentally produced lasso peptides that are ranked highest are subjected to further optimization through iterative computational refinement based on the structure of the best ranked lasso peptides. In some embodiments, computationally identified and experimentally produced lasso peptides that are ranked highest are subjected to further optimization through directed or random peptide evolution methods. In some embodiments, computationally identified and experimentally produced lasso peptides that are ranked highest are subjected to further optimization through directed or random peptide evolution methods and lasso peptides that are ranked highest after evolution are further subjected to iterative computational optimization. 6.5 Methods of Producing Lasso Peptides and Lasso Peptide Analogs
[0214] In silico modeling is aimed at identifying molecules that are predicted to bind selectively and with high affinity to a target protein of interest. Once identified, the lasso peptide variants are synthesized, isolated, and partially or substantially purified to enable experimental testing for validation of biological activity and other properties. For lasso peptides, there are two main methods for producing the identified molecules, which involve: (i) cell-free biosynthesis technology, and / or (ii) cell-based production methods (FIG.5). 6.5.1 Cell-Free Biosynthesis of Lasso Peptides
[0215] In one aspect, provided herein are methods for producing lasso peptides, including identified lassos that are predicted using in silico modeling, comprising in vitro cell-free biosynthesis (CFB) methods. CFB methods employ the enzymes and the biosynthetic and 71 ACTIVE 705331286v1metabolic machinery present inside cells, but without using living cells. CFB methods allow rapid expression of natural biosynthetic genes and pathways and facilitate targeted or phenotypic activity screening of natural products, without the need for plasmid-based cloning or in vivo cellular propagation, thus enabling rapid process / product pipelines (e.g., creation of large quantity of lasso peptide in a short time). Features of the CFB methods for lasso peptide production include that oligonucleotides (linear or circular constructs of DNA or RNA) encoding a minimal set of lasso peptide biosynthesis pathway genes (e.g., Genes A-C in a lasso peptide biosynthetic gene cluster) may be added to a cell extract containing in vitro TX-TL machinery for transcribing and translating the genes into the functional enzymes and lasso precursor peptides for production of lasso peptides (See: Gagoski, D., et al., Biotechnol. Bioeng.2016;113: 292–300; Culler, S. et al., PCT Appl. No. WO2017 / 031399). Accordingly, the CFB methods can produce in a CFB reaction mixture at least one, two or more of the lasso peptide variants.
[0216] In some embodiments, the method for producing a lasso peptide comprises (a) providing a CFB system comprising a minimal set of lasso peptide biosynthesis components; and (b) incubating the CFB system under a suitable condition to produce the lasso peptide.
[0217] In some embodiments, the minimal set of lasso peptide biosynthesis components comprises one or more components functions to provide a lasso precursor peptide, and one or more components function to process the lasso precursor peptide into the lasso peptide. In some embodiments, the one or more components function to process the lasso precursor peptide into the lasso peptide consist of a lasso peptidase and a lasso cyclase. In some embodiments, the one or more components function to process the lasso precursor peptide into the lasso peptide consists of a lasso peptidase, a lasso cyclase and an RRE.
[0218] In some embodiments, the minimal set of lasso peptide biosynthesis components comprises one or more components functions to provide a lasso core peptide, and one or more components function to process the lasso core peptide into the lasso peptide. In some embodiments, the one or more components function to process the lasso core peptide into the lasso peptide comprises one or more selected from a lasso peptidase, a lasso cyclase and an RRE. In some embodiments, the one or more components function to process the lasso core into the lasso peptide consist of a lasso cyclase.
[0219] In various embodiments, the one or more components function to provide a peptide or protein (e.g., a lasso precursor peptide, a lasso core peptide, or lasso peptide biosynthetic enzymes and proteins) in a CFB system can be provided in the form of the peptide or protein are provided in the form of the peptide or protein per se. 72 ACTIVE 705331286v1
[0220] In some embodiments, at least some of the peptide or protein components in the CFB system can be natural peptides or polypeptides. In some embodiments, at least some of the peptide or protein components in the CFB system are derivatives of natural peptides or polypeptides. In some embodiments, at least some of the peptide or protein components in the CFB system are non-natural peptides. In some embodiments, the one or more peptide or protein components of the CFB system can be isolated from nature, such as isolated from microorganisms producing the lasso precursor peptides. In some embodiments, the one or more peptide or protein components of the CFB system can be synthetically or recombinantly produced, using methods known in the art. In some embodiments, the one or more peptide or protein components of the CFB system can be synthesized using the CFB system as described herein, followed by purifying the biosynthesized peptide or protein components from the CFB system.
[0221] In some embodiments, the CFB system comprises one or more fusion protein, or a polynucleotide encoding the fusion protein such that the CFB system is capable of producing the fusion protein through in vitro transcription and translation (TX-TL).
[0222] In some embodiments, the fusion protein comprised a lasso precursor peptide or a lasso core peptide fused to one or more lasso peptide biosynthesis components. In some embodiments, the one or more lasso peptide biosynthesis components are selected from (i) a lasso peptidase; (ii) a lasso cyclase; (iii) a RRE; or (iv) any combinations of (i) to (iii). In some embodiments, the one or more lasso peptide biosynthesis components are encoded by the same lasso peptide biosynthetic gene cluster. In other embodiments, the one or more lasso peptide biosynthesis components are encoded by different lasso peptide biosynthetic gene cluster.
[0223] In some embodiments, the fusion protein comprises an amino acid linker between the lasso peptidase or the lasso cyclase and the one or more additional peptide or polypeptide. In some embodiments, the fusion protein does not comprise an amino acid linker between the lasso peptidase or the lasso cyclase and the one or more additional peptide or polypeptide.
[0224] In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso peptidase. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso cyclase. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a RRE. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso peptidase and a lasso cyclase. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso peptidase and a RRE. In specific embodiments, the fusion protein comprises a lasso precursor 73 ACTIVE 705331286v1peptide fused to a lasso cyclase and a RRE. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso peptidase, a lasso cyclase and RRE. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso peptidase. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso cyclase. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a RRE. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso peptidase and a lasso cyclase. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso peptidase and a RRE. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso cyclase and a RRE. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso peptidase, a lasso cyclase and RRE.
[0225] In some embodiments, the fusion protein comprised a lasso precursor peptide or a lasso core peptide fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a peptide or polypeptide that facilitates production of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom through cell-free biosynthesis. Examples of peptide or polypeptide that can be fused with a lasso precursor peptide or a lasso core peptide according to the present disclosure include but are not limited to (i) a peptide or polypeptide that increases the level of transcription of the lasso precursor peptide or lasso core peptide in the CFB system; (ii) a peptide or polypeptide that increases the level of translation of the lasso precursor peptide or lasso core peptide in the CFB system; (iii) a peptide or polypeptide that facilitates the processing of the lasso precursor peptide or lasso core peptide into the lasso peptide; (iv) a peptide or polypeptide that improves stability of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; (v) a peptide or polypeptide that improves solubility of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; (vi) a peptide or polypeptide that enables or facilitates the detection of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; (vii) a peptide or polypeptide that enables or facilitates purification of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; (viii) a peptide or polypeptide that enables or facilitates immobilization of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; or (ix) any combination of (i) to (viii).
[0226] In some embodiments, the fusion protein comprised a lasso precursor peptide or a lasso core peptide fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a biologically 74 ACTIVE 705331286v1active peptide or polypeptide. Examples of biologically active peptide or polypeptide that can be fused with a lasso precursor peptide or lasso core peptide according to the present disclosure include but are not limited to (i) a peptide or polypeptide capable of binding to a target molecule (e.g., an antibody or an antigen); (ii) a peptide or polypeptide that enhance cell permeability of the fusion protein; (iii) a peptide or polypeptide capable of conjugating the fusion protein to at least one additional copy of the fusion protein; (iv) a peptide or polypeptide capable of linking the fusion protein to one or more peptidic or non-peptidic molecule; (v) a peptide or polypeptide capable of modulating activity of the lasso precursor peptide or lasso core peptide; (vi) a peptide or polypeptide capable of modulating activity of the lasso peptide derived from the lasso precursor peptide or the lasso core peptide; or (vii) any combinations of (i) to (vi).
[0227] In some embodiments, the fusion protein comprised a lasso peptidase or a lasso cyclase fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide is fused to the N-terminus of the lasso peptidase or the lasso cyclase. In some embodiments, the one or more additional peptide or polypeptide is fused at the C-terminus of the lasso peptidase or the lasso cyclase. In some embodiments, a polynucleotide encoding the fusion protein comprises a nucleic acid sequence encoding the lasso peptidase or the lasso cyclase, wherein the 5’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, a polynucleotide encoding the fusion protein comprises a nucleic acid sequence encoding the lasso peptidase or the lasso cyclase, wherein the 3’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, the fusion protein comprises an amino acid linker between the lasso peptidase or the lasso cyclase and the one or more additional peptide or polypeptide. In some embodiments, the fusion protein does not comprise an amino acid linker between the lasso peptidase or the lasso cyclase and the one or more additional peptide or polypeptide.
[0228] In some embodiments, the fusion protein comprised a lasso peptidase or a lasso cyclase fused to one or more additional peptide or polypeptide. In some embodiments, the more additional peptide or polypeptide comprises a peptide or polypeptide encoded by a lasso peptide biosynthetic gene cluster. Examples of peptide or polypeptide that can be fused with a lasso precursor peptide or a lasso core peptide according to the present disclosure include but are not limited to (i) a lasso precursor peptide; (ii) a lasso core peptide; (iii) a lasso peptidase; (iv) a lasso cyclase, (v) a RRE; or (vi) any combinations of (i) to (vi). In specific 75 ACTIVE 705331286v1embodiments, the fusion protein comprises at least one lasso cyclase and at least one lasso peptidase. In specific embodiments, the fusion protein comprises at least one lasso cyclase fused to a RRE. In specific embodiments, the fusion protein comprises at least one lasso peptidase fused to a RRE.
[0229] In some embodiments, the fusion protein comprised a lasso peptidase or a lasso cyclase fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a peptide or polypeptide that facilitates production of the lasso peptidase or lasso cyclase through cell-free biosynthesis. Examples of peptide or polypeptide that can be fused with the lasso peptidase or lasso cyclase according to the present disclosure include but are not limited to (i) a peptide or polypeptide that increases the level of transcription of the lasso peptidase or lasso cyclase in the CFB system; (ii) a peptide or polypeptide that increases the level of translation of the lasso peptidase or lasso cyclase in the CFB system; (iii) a peptide or polypeptide that improves stability of the lasso peptidase or lasso cyclase; (vi) a peptide or polypeptide that improves solubility of the lasso peptidase or lasso cyclase; (v) a peptide or polypeptide that enables or facilitates the detection of the lasso peptidase or lasso cyclase; (vi) a peptide or polypeptide that enables or facilitates purification of the lasso peptidase or lasso cyclase; (vii) a peptide or polypeptide that enables or facilitates immobilization of the lasso peptidase or lasso cyclase; or (viii) any combination of (i) to (vii).
[0230] In some embodiments, the fusion protein comprised a lasso peptidase or a lasso cyclase fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a biologically active peptide or polypeptide. Examples of biologically active peptide or polypeptide that can be fused with a lasso peptidase or a lasso cyclase according to the present disclosure include but are not limited to (i) a peptide or polypeptide capable of modulating the reaction catalyzing activity of the lasso peptidase or lasso cyclase; (ii) a peptide or polypeptide capable of modulating target specificity of the lasso peptidase or lasso cyclase; (iii) an enzyme having the same or different enzymatic activity as the lasso peptidase or lasso cyclase; or any combination of (i) to (iii).
[0231] In some embodiments, the fusion protein comprised a RIPP recognition element (RRE) fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide is fused to the N-terminus of the RRE. In some embodiments, the one or more additional peptide or polypeptide is fused at the C-terminus of the RRE. In some embodiments, a polynucleotide encoding the fusion protein comprises a 76 ACTIVE 705331286v1nucleic acid sequence encoding the RRE, wherein the 5’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, a polynucleotide encoding the fusion protein comprises a nucleic acid sequence encoding the RRE, wherein the 3’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, the fusion protein comprises an amino acid linker between the RRE and the one or more additional peptide or polypeptide. In some embodiments, the fusion protein does not comprise an amino acid linker between RRE and the one or more additional peptide or polypeptide.
[0232] In some embodiments, the fusion protein comprised a RIPP recognition element (RRE) fused to one or more additional peptide or polypeptide. In some embodiments, the more additional peptide or polypeptide comprises a peptide or polypeptide encoded by a lasso peptide biosynthetic gene cluster. Examples of peptide or polypeptide that can be fused with a lasso precursor peptide or a lasso core peptide according to the present disclosure include but are not limited to (i) a lasso precursor peptide; (ii) a lasso core peptide; (iii) a lasso peptidase; (iv) a lasso cyclase, (v) a RRE; or (vi) any combinations of (i) to (vi). In specific embodiments, the fusion protein comprises at least one lasso precursor peptide fused to a RRE. In specific embodiments, the fusion protein comprises at least one lasso core peptide fused to a RRE. In specific embodiments, the fusion protein comprises at least one lasso cyclase fused to a RRE. In specific embodiments, the fusion protein comprises at least one lasso peptidase fused to a RRE.
[0233] In some embodiments, the fusion protein comprised a RIPP recognition element (RRE) fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a peptide or polypeptide that facilitates production of the RRE through cell-free biosynthesis. Examples of peptide or polypeptide that can be fused with the RRE according to the present disclosure include but are not limited to (i) a peptide or polypeptide that increases the level of transcription of the RRE in the CFB system; (ii) a peptide or polypeptide that increases the level of translation of the RRE in the CFB system; (iii) a peptide or polypeptide that improves stability of the RRE; (vi) a peptide or polypeptide that improves solubility of the RRE; (v) a peptide or polypeptide that enables or facilitates the detection of the RRE; (vi) a peptide or polypeptide that enables or facilitates purification of the RRE; (vii) a peptide or polypeptide that enables or facilitates immobilization of the RRE; or (viii) any combination of (i) to (vii). 77 ACTIVE 705331286v1
[0234] In some embodiments, the fusion protein comprised a RIPP recognition element (RRE) fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a biologically active peptide or polypeptide. Examples of biologically active peptide or polypeptide that can be fused with a RRE according to the present disclosure include but are not limited to (i) a peptide or polypeptide capable of modulating the reaction catalyzing activity of the lasso peptidase or lasso cyclase; (ii) a peptide or polypeptide capable of modulating target specificity of the lasso peptidase or lasso cyclase; (iii) an enzyme having the same or different enzymatic activity as the lasso peptidase or lasso cyclase; or any combination of (i) to (iii).
[0235] In particular embodiments, the lasso precursor peptide genes are fused at the 5’- terminus of the DNA template strand of the gene to oligonucleotide sequences that encode peptides or proteins, such as sequences encoding maltose-binding protein (MBP) or small ubiquitin-like modifier protein (SUMO), which enhance the stability, solubility, and production of the desired TX-TL products (Marblestone, J.G., et al., Protein Sci, 2006, 15, 182–189). In particular embodiments, the lasso precursor peptides are fused at the C- terminus of the leader sequences to form conjugates with peptides or proteins, such as maltose-binding protein or small ubiquitin-like modifier protein, which enhance the stability, solubility, and production of the fused MBP-lasso or SUMO-lasso precursor peptide.
[0236] In particular embodiments, the lasso precursor peptide genes or lasso core peptide genes are fused at the 3’-terminus of the DNA template strand of the gene to oligonucleotide sequences that encode peptides or proteins, such as sequences encoding maltose-binding protein (MBP) or small ubiquitin-like modifier protein (SUMO), which enhance the stability, solubility, and production of the desired TX-TL products. In particular embodiments, the lasso precursor peptides, lasso core peptides, or lasso peptides are fused at the N-terminus to form conjugates with peptides or proteins, such as maltose-binding protein or small ubiquitin- like modifier protein, which enhance the stability, solubility, and production of the fused MBP-lasso or SUMO-lasso precursor peptide.
[0237] In particular embodiments, the lasso precursor peptide genes or lasso core peptide genes are fused at the 5’-terminus of the DNA template strand of the gene to oligonucleotide sequences that encode peptides or proteins, with or without a linker, such as sequences encoding peptide tags for affinity purification or immobilization, including his-tags, strep- tags, or FLAG-tags. In some embodiments, the lasso precursor peptides, lasso core peptides, or lasso peptides are fused at the C-terminus of the core peptides to form conjugates with 78 ACTIVE 705331286v1other peptides or proteins, with or without a linker, such as peptide tags for affinity purification or immobilization, including his-tags, strep-tags, or FLAG-tags.
[0238] In particular embodiments, lasso precursor peptides, lasso core peptides, or lasso peptides are fused to molecules that can enhance cell permeability or penetration into cells, for example through the use of arginine-rich cell-penetrating peptides such as TAT peptide, penetratin, and flock house virus (FHV) coat peptide (Brock, R., Bioconjug. Chem., 2014, 25, 863–868). In particular embodiments, a lasso precursor peptide gene or core peptide gene is fused at the 3’-terminus to oligonucleotide sequences that encode arginine-rich cell- penetrating peptides or proteins, including oligonucleotide sequences that encode penetratin, and flock house virus (FHV) coat peptide or similar peptides that contain guanidinium groups or a combination of lysine and guanidinium groups (Wender, P.A., et al., Adv. Drug Deliv. Rev., 2008, 60, 452–472). In particular embodiments, a lasso precursor peptide, lasso core peptide, or lasso peptide is fused at the C-terminus to peptides that promote cell penetration such as arginine-rich cell-penetrating peptides or proteins, including amino acid sequences that encode TAT peptide, penetratin, and flock house virus (FHV) coat peptide or similar peptides that contain guanidinium groups or a combination of lysine and guanidinium groups.
[0239] In particular embodiments, the lasso precursor peptide genes or lasso core peptide genes are fused at the 5’-terminus of the DNA template strand of the gene to oligonucleotide sequences that encode peptides or proteins, with or without a linker, such as sequences encoding natural or unnatural peptide epitopes that are known to bind with high affinity to antibodies, cell surface proteins, or cell surface receptors, including cytokine binding epitopes, integrin ligand binding epitopes, and the like. In particular embodiments, the lasso precursor peptides, lasso core peptides, or lasso peptides are fused at the C-terminus to peptides or proteins, with or without a linker, such as peptide epitopes that are known to bind with high affinity to antibodies, cell surface proteins, or cell surface receptors, including cytokine binding epitopes, chemokine binding epitopes, integrin ligand binding epitopes, and the like.
[0240] Additionally or alternatively, the one or more components function to provide a peptide or protein (e.g., a lasso precursor peptide, a lasso core peptide, or lasso peptide biosynthetic enzymes and proteins) in a CFB system can be provided in the form of a nucleic acid encoding the peptide or protein and in vitro TX-TL machinery capable of producing the peptide or protein via in vitro TX-TL of the coding sequences. In various embodiments, the coding nucleic acid can be DNA, RNA or cDNA. In various embodiments, one or more 79 ACTIVE 705331286v1coding nucleic acid sequences can be contained in the same nucleic acid molecule, such as a vector.
[0241] It is understood that when more than one coding nucleic acid sequences are included in a CFB system, such more than one encoding nucleic acid sequences can be introduced on separate nucleic acid molecules, on polycistronic nucleic acid molecules, or a combination thereof. For example, as disclosed herein, a microbial organism or a cell extract can be engineered to express two or more exogenous nucleic acids encoding lasso precursor peptide, lasso core peptide, lasso peptidase, lasso cyclase or RRE. In the case where two exogenous nucleic acids encoding a desired activity are introduced into a host microbial organism or into a cell extract, it is understood that the two exogenous nucleic acids can be introduced as a single nucleic acid, for example, on a single plasmid or as linear strands of DNA, or on separate plasmids, or can be integrated into the host chromosome at a single site or multiple sites, and still be considered as two exogenous nucleic acids. Similarly, it is understood that more than two exogenous nucleic acids can be introduced into a host organism or into a cell extract in any desired combination, for example, on a single plasmid, or on separate plasmids, or as linear strands of DNA, or can be integrated into the host chromosome at a single site or multiple sites.
[0242] In some embodiments, the in vitro TX-TL machinery is purified from a host cell. In some embodiments, the in vitro TX-TL machinery is provided in the form of a cell extract of a host cell. An exemplary procedure for obtaining a cell extract comprises the steps of (i) growing cells, (ii) breaking open or lysing the cells by mechanical, biological or chemical means, (iii) removing cell debris and insoluble materials e.g., by filtration or centrifugation, and (iv) optionally treating to remove residual RNA and DNA, but retaining the active enzymes and biosynthetic machinery for transcription and translation, and optionally the metabolic pathways for co-factor recycle, including but not limited to co-factors such as THF, S-adenosylmethionine, ATP, NADH, NAD and NADP and NADPH. In some embodiments, a cell extract may be further supplemented for improved performance in in vitro TX-TL.
[0243] In some embodiments, a cell extract can be further supplemented with some or all of the twenty proteinogenic naturally-occurring amino acids and corresponding transfer ribonucleic acids (tRNAs), and optionally, may be supplemented with additional components, including but not limited to: (1) glucose, xylose, fructose, sucrose, maltose, or starch, (2) adenosine triphosphate (ATP), and / or adenosine diphosphate (ADP), purine and guanidine nucleotides, adenosine triphosphate, guanosine triphosphate, cytosine triphosphate, and / or 80 ACTIVE 705331286v1uridine triphosphate, or combinations thereof, (3) cyclic-adenosine monophosphate (cAMP) and / or 3-phosphoglyceric acid (3-PGA), (4) nicotinamide adenine dinucleotides NADH and / or NAD, or nicotinamide adenine dinucleotide phosphates, NADPH, and / or NADP, or combinations thereof, (5) amino acid salts such as magnesium glutamate and / or potassium glutamate, (6) buffering agents such as HEPES, TRIS, spermidine, or phosphate salts, (7) inorganic salts, including but not limited to, potassium phosphate, sodium chloride, magnesium phosphate, and magnesium sulfate, (8) cofactors such as folinic acid and co- enzyme A (CoA), L(–)-5-formyl-5,6,7,8-tetrahydrofolic acid (THF), and / or biotin, (8) RNA polymerase, (9) 1,4-dithiothreitol (DTT), (10) magnesium acetate, and / or ammonium acetate, and / or (11) crowding agents such as PEG 8000, Ficoll 70, or Ficoll 400, or combinations thereof. In some embodiments, the cell extracts or supplemented cell extracts can be used as a reaction mixture to carry out in vitro TX-TL. In some embodiments, supplementations or adjustments can be made to the cell extract to provide a suitable condition for lasso formation.
[0244] In some embodiments, the in vitro TX-TL machinery is provided in the form of a cell extract or supplemented cell extract of a host cell. In some embodiments, the host cell is the cell of the same organism where the coding nucleic acid is derived from. For CFB of lasso peptides and related molecules thereof, the coding nucleic acid sequences can be identified using one or more computer-based genomic mining tools described herein or known in the art. For example, U.S. Provisional Application Nos.62 / 652,213 and 62 / 651,028 disclose thousands of sequences from lasso peptide biosynthetic gene clusters identified from various organisms, and provide GenBank accession numbers for various sequences for lasso precursor peptides, lasso peptidase, lasso cyclase and / or RRE. Host organisms where the lasso peptide biosynthetic gene clusters originate can be identified based on the GenBank accession numbers, including but not limited to Caulobacteraceae species (e.g., Caulobacter sp. K31, Caulobacter henricii), Streptomyces species (e.g. Streptomyces nodosus, Streptomyces caatingaensis), Burkholderiaceae species (e.g., Burkholderia thailandensis E264), Pseudomallei species, Bacillus species, Burkholderia species (e.g., Burkholderia thailandensis MSMB43, Burkholderia oklahomensis, Burkholderia pseudomallei), Sphingomonadaceae species (e.g., Sphingobium sp. YBL2, Sphingobium chlorophenolicum, Sphingobium yanoikuyae). In other embodiments, the host cell is a microbial organism known to be applicable to fermentation processes. Exemplary bacteria include species selected from Escherichia coli, Klebsiella oxytoca, Anaerobiospirillum succiniciproducens, Actinobacillus succinogenes, Mannheimia succiniciproducens, Rhizobium etli, Bacillus 81 ACTIVE 705331286v1subtilis, Corynebacterium glutamicum, Gluconobacter oxydans, Zymomonas mobilis, Lactococcus lactis, Lactobacillus plantarum, Streptomyces coelicolor, Streptomyces albus, Clostridium acetobutylicum, Vibrio natriegens, Pseudomonas fluorescens, and Pseudomonas putida. Exemplary yeasts or fungi include species selected from Saccharomyces cerevisiae, Schizosaccharomyces pombe, Kluyveromyces lactis, Kluyveromyces marxianus, Aspergillus terreus, Aspergillus niger and Pichia pastoris. E. coli is a particularly useful host organism since it is a well characterized microbial organism suitable for genetic engineering. Other particularly useful host organisms include Vibrio natriegens, and yeast such as Saccharomyces cerevisiae.
[0245] In some embodiments, the CFB system is configured to produce a lasso peptide. In specific embodiments, the CFB system comprises one or more components configured to provide (i) a lasso precursor peptide, (ii) a lasso peptidase, (iii) a lasso cyclase. In specific embodiments, the CFB system comprises one or more components configured to provide (i) a lasso core peptide, and (ii) a lasso cyclase. In some embodiments, the CFB system further comprises one or more components configured to provide (iv) an RRE. In some embodiments, all of (i) to (iv) above are provided in the CFB system as the corresponding peptide or protein. In alternative embodiments, at least one of (i) to (iv) above is provided in the CFB system as a nucleic acid encoding the corresponding protein, and the CFB system further comprises in vitro TX-TL machinery for producing the corresponding protein from the coding nucleic acid. In these embodiments, the CFB systems can be incubated under a condition suitable for lasso formation to produce the lasso peptide. The incubation condition can be designed and adjusted based on various factors known to skilled artisan in the art, including for example, condition suitable for maintain stability of components of the CFB system, conditions suitable for the lasso processing enzymes to exert enzymatic activities, and / or conditions suitable for the in vitro TX-TL of the coding sequences present in the CFB system.
[0246] Without being bound by the theory, it is contemplated that different lasso peptidases can process the same lasso precursor peptide into different lasso core peptide by recognizing and cleaving different leader peptide off the lasso precursor. Additionally, different lasso cyclase can process the same lasso core peptide into distinct lasso peptides by cyclizing the core peptide at different ring-forming amino acid residues. Additionally, different RREs can facilitate different processing by the lasso peptidase and / or lasso cyclase, and thus lead to formation of distinct lasso peptides from the same lasso precursor peptide. 82 ACTIVE 705331286v1
[0247] Accordingly, in some embodiments, to produce a natural lasso peptide, the CFB system comprises the lasso precursor peptide, lasso peptidase, and lasso cyclase produced from coding sequences of the same lasso peptide biosynthetic gene cluster (such as Genes A, B, and C of the same lasso peptide biosynthetic gene cluster). In some embodiments, to produce a natural lasso peptide, the CFB system comprises the lasso precursor peptide, lasso peptidase, lasso cyclase, and RRE produced from coding sequences of the same lasso peptide biosynthetic gene cluster.
[0248] In some embodiments, to produce a natural lasso peptide, the CFB system comprises the lasso core peptide, and lasso cyclase produced from coding sequences of the same lasso peptide biosynthetic gene cluster (such as Genes A and C of the same lasso peptide biosynthetic gene cluster). In some embodiments, to produce a natural lasso peptide, the CFB system comprises the lasso core peptide, lasso cyclase, and RRE produced from coding sequences of the same lasso peptide biosynthetic gene cluster.
[0249] In alternative embodiments, to produce a derivative of a natural lasso peptide, at least two of the lasso precursor peptide, lasso peptidase and lasso cyclase in the CFB system are produced from coding sequences of different lasso peptide biosynthetic gene clusters (such as Gene A from one, and Genes B and C from another, lasso peptide biosynthetic gene cluster). In alternative embodiments, to produce a derivative of a natural lasso peptide, at least two of the lasso precursor peptide, lasso peptidase, lasso cyclase and RRE in the CFB system are produced from coding sequences of different lasso peptide biosynthetic gene clusters.
[0250] In alternative embodiments, to produce a derivative of a natural lasso peptide, the lasso core peptide and lasso cyclase in the CFB system are produced from coding sequences of different lasso peptide biosynthetic gene clusters (such as Gene A from one, and Gene C from another, lasso peptide biosynthetic gene cluster). In alternative embodiments, to produce a derivative of a natural lasso peptide, at least two of the lasso core peptide, lasso cyclase and RRE in the CFB system are produced from coding sequences of different lasso peptide biosynthetic gene clusters.
[0251] In some embodiments, to produce a derivative of a natural lasso peptide, a lasso precursor peptide is modified at the core peptide sequence, while the leader sequence is maintained the same. The modified precursor peptide can then processed by corresponding lasso peptidase and / or lasso cyclase into a matured lasso peptide with modified amino acid sequence. 83 ACTIVE 705331286v1
[0252] In various embodiments, the contacting step (a) comprises adding a first nucleic acid sequence encoding the peptide into the cell-free biosynthesis reaction mixture, and where the cell-free biosynthesis reaction mixture comprises in vitro TX-TL machinery and is configured to express the peptide. In some embodiments, the contacting step (a) comprises adding a second nucleic acid sequence encoding the lasso peptide biosynthesis component to the cell-free biosynthesis reaction mixture, and where the cell-free biosynthesis reaction mixture comprises in vitro TX-TL machinery configured to express the lasso peptide biosynthesis component. In some embodiments, the lasso peptide biosynthesis component comprises a lasso peptidase. In some embodiments, the lasso peptide biosynthesis component comprises a lasso cyclase. In some embodiments, the lasso peptide biosynthesis component further comprises a post-translationally modified peptide (RiPP) recognition element (RRE).
[0253] More particularly, in some of those embodiments where the lasso peptide biosynthesis component comprises a lasso peptidase and a lasso cyclase, the contacting step (a) comprises adding the second nucleic acid sequence encoding the lasso cyclase and a third nucleic acid sequence encoding the lasso peptidase. In some of those embodiments where the lasso peptide biosynthesis component comprises a lasso cyclase and a post-translationally modified peptide (RiPP) recognition element (RRE), the contacting step (a) comprises adding the second nucleic acid sequence encoding the lasso cyclase and a fourth nucleic acid sequence encoding the RRE. In some of those embodiments where the lasso peptide biosynthesis component comprises a lasso peptidase, a lasso cyclase and a post- translationally modified peptide (RiPP) recognition element (RRE), and where the contacting step (a) comprises adding the second nucleic acid sequence encoding the lasso cyclase, a third nucleic acid sequence encoding the lasso peptidase and a fourth nucleic acid sequence encoding the RRE. In some embodiments, the cell-free biosynthesis reaction mixture comprises cell extract or supplemented cell extract.
[0254] Additional lasso peptide biosynthesis components and corresponding leader sequences are known in the art, such as those disclosed in PCT application publication numbers: WO2019 / 191571, which is incorporated herein by reference in its entirety.
[0255] In some embodiments, CFB reactions are conducted with a minimal set of lasso peptide biosynthesis components combined with genes that encode additional peptides, proteins or enzymes, including genes that encode RiPP recognition elements (RREs) or oligonucleotides that encode RREs that are fused to the 5’ or 3’ end of a lasso precursor peptide gene, a lasso core peptide gene, a lasso peptidase gene or a lasso cyclase gene. In other embodiments, CFB reactions are conducted with a minimal set of lasso peptide 84 ACTIVE 705331286v1biosynthesis components, including lasso precursor peptides, lasso peptidases, or lasso cyclase that are fused to RREs at the N-terminus or C-terminus. In other embodiments, CFB reactions are conducted with a minimal set of lasso peptide biosynthesis components combined and contacted with additional isolated proteins or enzymes, including RiPP recognition elements (RREs).
[0256] In some embodiments, CFB reactions are conducted with a minimal set of lasso peptide biosynthesis components combined and contacted with genes that encode additional proteins or enzymes, including genes that encode lasso peptide modifying enzymes such as N-methyltransferases, O-methyltransferases, biotin ligases, glycosyltransferases, esterases, acylases, acyltransferases, aminotransferases, amidases, hydroxylases, dehydrogenases, halogenases, kinases, RiPP heterocyclases, RiPP cyclodehydratases, peptidylarginine deiminase, and prenyltransferases.
[0257] In some embodiments, CFB reactions are conducted with a minimal set of lasso peptide biosynthesis components combined and contacted with additional isolated proteins or enzymes, including lasso peptide modifying enzymes such as N-methyltransferases, O- methyltransferases, biotin ligases, glycosyltransferases, esterases, acylases, acyltransferases, aminotransferases, amidases, hydroxylases, dehydrogenases, halogenases, kinases, RiPP heterocyclases, RiPP cyclodehydratases, peptidylarginine deiminase, and prenyltransferases.
[0258] CFB methods and systems provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are conducted in a CFB reaction mixture, comprising one or more cell extracts that are supplemented with all twenty proteinogenic naturally occurring amino acids and corresponding transfer ribonucleic acids (tRNAs). Cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components also may be supplemented with additional components, including but not limited to, glucose, xylose, fructose, sucrose, maltose, starch, adenosine triphosphate (ATP), and / or adenosine diphosphate (ADP), purine and guanidine nucleotides, adenosine triphosphate, guanosine triphosphate, cytosine triphosphate, and uridine triphosphate, cyclic-adenosine monophosphate (cAMP) and / or 3-phosphoglyceric acid (3-PGA), nicotinamide adenine dinucleotides NADH and / or NAD, or nicotinamide adenine dinucleotide phosphates, NADPH, and / or NADP, or combinations thereof, amino acid salts such as magnesium glutamate and / or potassium glutamate, buffering agents such as HEPES, TRIS, spermidine, or phosphate salts, inorganic salts, including but not limited to, potassium phosphate, sodium 85 ACTIVE 705331286v1chloride, magnesium phosphate, and magnesium sulfate, folinic acid and co-enzyme A (CoA), crowding agents such as PEG 8000, Ficoll 70, or Ficoll 400, L(–)-5-formyl-5,6,7,8- tetrahydrofolic acid, RNA polymerase, biotin, 1,4-dithiothreitol (DTT), magnesium acetate, ammonium acetate , or combinations thereof. For a general description of cell-free extract production and preparation, see: Krinsky, N., et al., PLoS ONE, 2016, 11(10): e0165137.
[0259] In alternative embodiments, the preparation CFB reaction mixtures and cell extracts employed for the CFB methods as provided herein, comprises characterization of the CFB reaction mixtures and cell extracts using proteomic approaches to assess and quantify the proteome available for the production of lasso peptides and related molecules thereof. In alternative embodiments,13C metabolic flux analysis (MFA) and / or metabolomics studies are conducted on CFB reaction mixtures and cell extracts to create a flux map and characterize the resulting metabolome of the CFB reaction mixture and cell extract or extracts.
[0260] In other embodiments, the CFB method is performed using: one or a combination of two or more cell extracts from various “chassis” organisms, such as E. coli, optionally mixed with one or a combination of two or more cell extracts derived from other species, e.g., a native lasso peptide-producing organism or relative. This can give the advantage of a robust transcription / translation machinery, combined with any unknown components of the native species that might be needed for proper protein folding or activity, or to supply precursors for the lasso peptide pathway. In alternative embodiments, if these factors are known they can be expressed in the chassis organism prior to making the cell extract or these factors can be isolated and purified and added directly to the CFB reaction mixture or cell extract.
[0261] In alternative embodiments, CFB methods and systems provided herein to produce lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, including the use of cell extracts for in vitro TX-TL systems express lasso peptide biosynthetic gene clusters without the regulatory constraints of the cell. In alternative embodiments, some or all of the lasso peptide pathway biosynthetic genes are refactored to remove native transcriptional and translational regulation. In alternative embodiments, some or all of the lasso peptide pathway biosynthetic genes are refactored and constructed into operons on plasmids.
[0262] In alternative embodiments, CFB methods, systems and processes, including in vitro TX-TL systems, provided herein to produce lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are cell-free platforms that can use whole cell, cytoplasmic or nuclear extract from a single organism such 86 ACTIVE 705331286v1as E.coli or Saccharomyces cerevisiae (S. cerevisiae) or from an organism of the Actinomyces genus, e.g., a Streptomyces. In alternative embodiments, CFB methods, systems and processes, including in vitro TX-TL systems, provided herein to produce lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are cell-free platforms that can use mixtures of whole cell, cytoplasmic, and / or nuclear extracts from the same or different organisms. In alternative embodiments, strain engineering approaches as well as modification of the growth conditions are used (on the organism from which at least one extract is derived) towards the creation of cell extracts as provided herein, to generate mixed cell extracts with varying proteomic and metabolic capabilities in the final CFB reaction mixture. In alternative embodiments, both approaches are used to tailor or design a final CFB reaction mixture for the purpose of synthesizing and characterizing lasso peptides, or for the creation of lasso peptide analogs through combinatorial biosynthesis approaches.
[0263] In alternative embodiments, cell extracts used in the CFB methods, provided herein to produce lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, comprise whole cell, cytoplasmic or nuclear extracts from a bacterial cell or eukaryotic cell, including insect, plant, fungal, yeast, or mammalian cells. In alternative embodiments, cell extracts used in the CFB methods, provided herein to produce lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, comprise whole cell, cytoplasmic or nuclear extracts from a bacterial cell or eukaryotic cell, including insect, plant, fungal, yeast, or mammalian cells, and are designed, produced and processed in a way to maximize efficacy and yield in the production of desired lasso peptides or related molecules thereof.
[0264] In an alternative embodiment, cell extracts used in the CFB methods, provided herein to produce lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, derive from at least two different bacterial cells, two different fungal cells; two different yeast cells, two different insect cells, two different plant cells or two different mammalian cells, or combinations of cell extracts from different species and genera thereof. In alternative embodiments, cell extracts used in the CFB methods, provided herein to produce lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, comprises an extract derived from: an Escherichia or a Escherichia coli (E. coli); a Streptomyces or an Actinobacteria; an Ascomycota, Basidiomycota, or a Saccharomycetales; a Penicillium or a Trichocomaceae; a 87 ACTIVE 705331286v1Spodoptera, a Spodoptera frugiperda, a Trichoplusia or a Trichoplusia ni; a Poaceae, a Triticum, or a wheat germ; a rabbit reticulocyte or a HeLa cell.
[0265] In alternative embodiments, cell extracts used in the CFB methods, provided herein to produce lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, comprises a cell extract from or comprises an extract derived from: any prokaryotic and eukaryotic organism including, but not limited to, bacteria, including Archaea, eubacteria, and eukaryotes, including yeast, plant, insect, animal, and mammal, including human cells. In alternative embodiments, at least one of the cell extracts used in the CFB methods provided herein comprises an extract from or comprises an extract derived from: Escherichia coli, Saccharomyces cerevisiae, Saccharomyces kluyveri, Candida boidinii, Clostridium kluyveri, Clostridium acetobutylicum, Clostridium beijerinckii, Clostridium saccharoperbutylacetonicum, Clostridium perfringens, Clostridium difficile, Clostridium botulinum, Clostridium tyrobutyricum, Clostridium tetanomorphum, Clostridium tetani, Clostridium propionicum, Clostridium aminobutyricum, Clostridium subterminale, Clostridium sticklandii, Ralstonia eutropha, Mycobacterium bovis, Mycobacterium tuberculosis, Porphyromonas gingivalis, Arabidopsis thaliana, Thermus thermophilus, Pseudomonas species, including Pseudomonas aeruginosa, Pseudomonas putida, Pseudomonas stutzeri, Pseudomonas fluorescens, Homo sapiens, Oryctolagus cuniculus, Rhodobacter spaeroides, Thermo-anaerobacter brockii, Metallosphaera sedula, Leuconostoc mesenteroides, Chloroflexus aurantiacus, Roseiflexus castenholzii, Erythrobacter, Simmondsia chinensis, Acinetobacter species, including Acinetobacter calcoaceticus and Acinetobacter baylyi, Porphyromonas gingivalis, Sulfolobus tokodaii, Sulfolobus solfataricus, Sulfolobus acidocaldarius, Bacillus subtilis, Bacillus cereus, Bacillus megaterium, Bacillus brevis, Bacillus pumilus, Rattus norvegicus, Klebsiella pneumonia, Klebsiella oxytoca, Euglena gracilis, Treponema denticola, Moorella thermoacetica, Thermotoga maritima, Halobacterium salinarum, Geobacillus stearothermophilus, Aeropyrum pernix, Sus scrofa, Caenorhabditis elegans, Corynebacterium glutamicum, Acidaminococcus fermentans, Lactococcus lactis, Lactobacillus plantarum, Streptococcus thermophilus, Enterobacter aerogenes, Candida, Aspergillus terreus, Pedicoccus pentosaceus, Zymomonas mobilus, Acetobacter pasteurians, Kluyveromyces lactis, Eubacterium barkeri, Bacteroides capillosus, Anaerotruncus colihominis, Natranaerobius thermophilusm, Campylobacter jejuni, Haemophilus influenzae, Serratia marcescens, Citrobacter amalonaticus, Myxococcus xanthus, Fusobacterium nuleatum, Penicillium chrysogenum, marine gamma proteobacterium, butyrate-producing bacterium, Nocardia 88 ACTIVE 705331286v1iowensis, Nocardia farcinica, Streptomyces griseus, Schizosaccharomyces pombe, Geobacillus thermoglucosidasius, Salmonella typhimurium, Vibrio cholera, Heliobacter pylori, Nicotiana tabacum, Oryza sativa, Haloferax mediterranei, Agrobacterium tumefaciens, Achromobacter denitrificans, Fusobacterium nucleatum, Streptomyces clavuligenus, Acinetobacter baumanii, Mus musculus, Lachancea kluyveri, Trichomonas vaginalis, Trypanosoma brucei, Pseudomonas stutzeri, Bradyrhizobium japonicum, Mesorhizobium loti, Bos taurus, Nicotiana glutinosa, Vibrio vulnificus, Vibrio natriegens, Selenomonas ruminantium, Vibrio parahaemolyticus, Archaeoglobus fulgidus, Haloarcula marismortui, Pyrobaculum aerophilum, Mycobacterium smegmatis MC2155, Mycobacterium avium subsp. paratuberculosis K-10, Mycobacterium marinum M, Tsukamurella paurometabola DSM 20162, Cyanobium PCC7001, Dictyostelium discoideum AX4.
[0266] In alternative embodiments, at least one cell, cytoplasmic or nuclear extract used in the CFB methods, provided herein to produce lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, comprises a cell extract from or comprises an extract derived from: Acinetobacter baumannii Naval-82, Acinetobacter sp. ADP1, Acinetobacter sp. strain M-1, Actinobacillus succinogenes 130Z, Allochromatium vinosum DSM 180, Amycolatopsis methanolica, Arabidopsis thaliana, Atopobium parvulum DSM 20469, Azotobacter vinelandii DJ, Bacillus alcalophilus ATCC 27647, Bacillus azotoformans LMG 9581, Bacillus coagulans 36D1, Bacillus megaterium, Bacillus methanolicus MGA3, Bacillus methanolicus PB1, Bacillus methanolicus PB-1, Bacillus selenitireducens MLS10 , Bacillus smithii, Bacillus subtilis , Burkholderia cenocepacia, Burkholderia cepacia, Burkholderia multivorans, Burkholderia pyrrocinia, Burkholderia stabilis, Burkholderia thailandensis E264, Burkholderiales bacterium Joshi_001, Butyrate-producing bacterium L2-50, Campylobacter jejuni, Candida albicans, Candida boidinii, Candida methylica, Carboxydothermus hydrogenoformans, Carboxydothermus hydrogenoformans Z-2901, Caulobacter sp. AP07, Chloroflexus aggregans DSM 9485, Chloroflexus aurantiacus J-10-fl, Citrobacter freundii, Citrobacter koseri ATCC BAA-895, Citrobacter youngae , Clostridium, Clostridium acetobutylicum, Clostridium acetobutylicum ATCC 824, Clostridium acidurici, Clostridium aminobutyricum, Clostridium asparagiforme DSM 15981, Clostridium beijerinckii , Clostridium beijerinckii NCIMB 8052, Clostridium bolteae ATCC BAA-613, Clostridium carboxidivorans P7, Clostridium cellulovorans 743B, Clostridium difficile, Clostridium hiranonis DSM 13275, Clostridium hylemonae DSM 15053, Clostridium kluyveri, Clostridium kluyveri DSM 555, 89 ACTIVE 705331286v1Clostridium ljungdahli, Clostridium ljungdahlii DSM 13528, Clostridium methylpentosum DSM 5476 , Clostridium pasteurianum, Clostridium pasteurianum DSM 525, Clostridium perfringens, Clostridium perfringens ATCC 13124, Clostridium perfringens str.13, Clostridium phytofermentans ISDg, Clostridium saccharobutylicum, Clostridium saccharoperbutylacetonicum, Clostridium saccharoperbutylacetonicum N1-4, Clostridium tetani, Corynebacterium glutamicum ATCC 14067, Corynebacterium glutamicum R, Corynebacterium sp. U-96, Corynebacterium variabile, Cupriavidus necator N-1, Cyanobium PCC7001, Desulfatibacillum alkenivorans AK-01, Desulfitobacterium hafniense, Desulfitobacterium metallireducens DSM 15288, Desulfotomaculum reducens MI-1, Desulfovibrio africanus str. Walvis Bay, Desulfovibrio fructosovorans JJ, Desulfovibrio vulgaris str. Hildenborough, Desulfovibrio vulgaris str. 'Miyazaki F', Dictyostelium discoideum AX4, Escherichia coli, Escherichia coli K-12 , Escherichia coli K-12 MG1655, Eubacterium hallii DSM 3353 , Flavobacterium frigoris, Fusobacterium nucleatum subsp. polymorphum ATCC 10953 , Geobacillus sp. Y4.1MC1, Geobacillus themodenitrificans NG80-2, Geobacter bemidjiensis Bem, Geobacter sulfurreducens, Geobacter sulfurreducens PCA, Geobacillus stearothermophilus DSM 2334, Haemophilus influenzae, Helicobacter pylori, Homo sapiens, Hydrogenobacter thermophilus, Hydrogenobacter thermophilus TK-6, Hyphomicrobium denitrificans ATCC 51888, Hyphomicrobium zavarzinii, Klebsiella pneumoniae, Klebsiella pneumoniae subsp. pneumoniae MGH 78578, Lactobacillus brevis ATCC 367, Leuconostoc mesenteroides, Lysinibacillus fusiformis, Lysinibacillus sphaericus, Mesorhizobium loti MAFF303099, Metallosphaera sedula, Methanosarcina acetivorans, Methanosarcina acetivorans C2A, Methanosarcina barkeri, Methanosarcina mazei Tuc01, Methylobacter marinus, Methylobacterium extorquens, Methylobacterium extorquens AM1, Methylococcus capsulatas, Methylomonas aminofaciens, Moorella thermoacetica, Mycobacter sp. strain JC1 DSM 3803, Mycobacterium avium subsp. paratuberculosis K-10, Mycobacterium bovis BCG, Mycobacterium gastri , Mycobacterium marinum M, Mycobacterium smegmatis, Mycobacterium smegmatis MC2155, Mycobacterium tuberculosis, Nitrosopumilus salaria BD31, Nitrososphaera gargensis Ga9.2, Nocardia farcinica IFM 10152, Nocardia iowensis (sp. NRRL 5646), Nostoc sp. PCC 7120, Ogataea angusta, Ogataea parapolymorpha DL-1 (Hansenula polymorpha DL-1), Paenibacillus peoriae KCTC 3763, Paracoccus denitrificans, Penicillium chrysogenum, Photobacterium profundum 3TCK, Phytofermentans ISDg, Pichia pastoris, Picrophilus torridus DSM9790, Porphyromonas gingivalis, Porphyromonas gingivalis W83, Pseudomonas aeruginosa PA01, Pseudomonas denitrificans, Pseudomonas knackmussii, Pseudomonas putida, Pseudomonas 90 ACTIVE 705331286v1sp, Pseudomonas syringae pv. syringae B728a, Pyrobaculum islandicum DSM 4184, Pyrococcus abyssi, Pyrococcus furiosus, Pyrococcus horikoshii OT3, Ralstonia eutropha, Ralstonia eutropha H16, Rhodobacter capsulatus, Rhodobacter sphaeroides, Rhodobacter sphaeroides ATCC 17025, Rhodopseudomonas palustris, Rhodopseudomonas palustris CGA009, Rhodopseudomonas palustris DX-1, Rhodospirillum rubrum, Rhodospirillum rubrum ATCC 11170, Ruminococcus obeum ATCC 29174, Saccharomyces cerevisiae, Saccharomyces cerevisiae S288c, Salmonella enterica, Salmonella enterica subsp. enterica serovar Typhimurium str. LT2, Salmonella enterica typhimurium , Salmonella typhimurium, Schizosaccharomyces pombe, Sebaldella termitidis ATCC 33386 , Shewanella oneidensis MR-1, Sinorhizobium meliloti 1021, Streptomyces coelicolor, Streptomyces griseus subsp. griseus NBRC 13350, Sulfolobus acidocalarius, Sulfolobus solfataricus P-2, Synechocystis str. PCC 6803, Syntrophobacter fumaroxidans, Thauera aromatica, Thermoanaerobacter sp. X514, Thermococcus kodakaraensis, Thermococcus litoralis, Thermoplasma acidophilum, Thermoproteus neutrophilus, Thermotoga maritima, Thiocapsa roseopersicina, Tolumonas auensis DSM 9187, Trichomonas vaginalis G3, Trypanosoma brucei, Tsukamurella paurometabola DSM 20162, Vibrio cholera, Vibrio harveyi ATCC BAA-1116, Vibrio natriegens, Xanthobacter autotrophicus Py2, Yersinia intermedia, or Zea mays.
[0267] In alternative embodiments, cell extracts used in the CFB methods and processes, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, e.g., including at least one of the cell, cytoplasmic or nuclear extracts, have added to them, or further comprise, supplemental ingredients, compositions or compounds, reagents, ions, trace metals, salts, or elements, buffers and / or solutions. In alternative embodiments, the CFB method and system of the present disclosure, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, use or fabricate environmental conditions to optimize the rate of formation or yield of a lasso peptide or related molecules thereof.
[0268] In alternative embodiments, CFB reaction mixtures and cell extracts used in the CFB methods and systems, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with a carbon source and other essential nutrients. The CFB production system, including cell extracts used in the CFB methods and processes, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, can include, for example, any carbohydrate 91 ACTIVE 705331286v1source. Such sources of sugars or carbohydrate substrates include glucose, xylose, maltose, arabinose, galactose, mannose, maltodextrin, fructose, sucrose and starch.
[0269] In alternative embodiments, CFB methods and systems provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are conducted in a CFB reaction mixture, comprising cell extracts that are supplemented with all twenty proteinogenic naturally occurring amino acids and corresponding transfer ribonucleic acids (tRNAs). In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with adenosine triphosphate (ATP), and / or adenosine diphosphate (ADP). In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with glucose, xylose, maltose, arabinose, galactose, mannose, maltodextrin, fructose, sucrose and / or starch. In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with purine and guanidine nucleotides, adenosine triphosphate, guanosine triphosphate, cytosine triphosphate, and uridine triphosphate. In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with cyclic-adenosine monophosphate (cAMP) and / or 3-phosphoglyceric acid (3-PGA). In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with nicotinamide adenine dinucleotides NADH and / or NAD, or nicotinamide adenine dinucleotide phosphates, NADPH, and / or NADP, or combinations thereof. In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with amino acid salts such as magnesium glutamate and / or potassium glutamate. In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with buffering agents such as HEPES, TRIS, spermidine, or 92 ACTIVE 705331286v1phosphate salts. In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with salts, including but not limited to, potassium phosphate, sodium chloride, magnesium phosphate, and magnesium sulfate. In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with folinic acid and co-enzyme A (CoA). In alternative embodiments, cell extracts used in the CFB reaction mixture, provided herein for the synthesis of lasso peptides and related molecules thereof from a minimal set of lasso peptide biosynthetic pathway components, are supplemented with crowding agents such as PEG 8000, Ficoll 70, or Ficoll 400, or combinations thereof. For a general description of cell-free extract production and preparation, see: Krinsky, N., et al., PLoS ONE, 2016, 11(10): e0165137.
[0270] In some embodiments, cell-free biosynthesis of lasso peptides or lasso peptide analogs is conducted with isolated peptide and enzyme components in standard buffered media, such as phosphate-buffered saline or tris-buffered saline, in each case containing salts, ATP, and co-factors required for lasso peptidase and lasso cyclase enzymatic activity. In some embodiments, cell-free biosynthesis of lasso peptides is conducted using genes that require transcription (TX) and translation (TL) to afford the lasso precursor peptide and / or lasso peptide biosynthetic enzymes in situ, and such in vitro biosynthesis processes are conducted in cell extracts derived from prokaryotic or eukaryotic cells (See: Gagoski, D., et al., Biotechnol. Bioeng.2016;113: 292–300; Culler, S. et al., PCT Appl. No. WO2017 / 031399). In alternative embodiments, a cell-free biosynthesis process for producing lasso peptide analogs is conducted by contacting an isolated, chemically- or biologically-synthesized precursor peptide with one or more isolated lasso peptidase, lasso cyclase, and lasso RRE. In other embodiments, a cell-free biosynthesis process for producing lasso peptide analogs is conducted by contacting an isolated, chemically- or biologically- synthesized core peptide with one or more isolated lasso peptidase, lasso cyclase, and lasso RRE. 6.5.2 Cell-based Biosynthesis of Lasso Peptides
[0271] In a related aspect, provided herein are also methods for producing lasso peptides and lasso peptide analogs, including lasso peptides identified using in silico modeling 93 ACTIVE 705331286v1methods, using a non-naturally occurring microbial organism. Certain cell-based production methods involve cultivating or fermenting a microbial organism that is a natural producer of a lasso peptide of interest. Alternative cell-based production methods involve cloning the genes encoding lasso peptide biosynthesis component into an appropriate vector or plasmid, introducing that vector or plasmid into a microorganism, and propagating or cultivating that organism with the necessary nutrients and under conditions for heterologous production of recombinant lasso peptides of interest (Zhang, Y., et al, Heterologous production of microbial ribosomally synthesized and post-translationally modified peptides, Front. Microbiol., 2018, doi: 10.3389 / fmicb.2018.01801).
[0272] Depending on the lasso peptide biosynthetic pathway constituents of a selected host microbial organism, the non-naturally occurring microbial organisms of the invention will include at least one exogenously expressed lasso peptide pathway-encoding nucleic acid and up to all encoding nucleic acids for one or more lasso peptide biosynthetic pathways. For example, lasso peptide biosynthesis can be established in a host deficient in a pathway enzyme or protein through exogenous expression of the corresponding encoding nucleic acid. In a host deficient in all enzymes or proteins of a lasso peptide pathway, exogenous expression of all enzyme or proteins in the pathway can be included, although it is understood that all enzymes or proteins of a pathway can be expressed even if the host contains at least one of the pathway enzymes or proteins. For example, exogenous expression of all enzymes or proteins in a pathway for production of a lasso peptide can be included, such as a lasso peptide precursor, a lasso peptide peptidase, a lasso peptide cyclase, and / or a lasso peptide RiPP recognition element (RRE).
[0273] Given the teachings and guidance provided herein, those skilled in the art will understand that the number of encoding nucleic acids to introduce in an expressible form will, at least, parallel the lasso peptide pathway deficiencies of the selected host microbial organism. Therefore, a non-naturally occurring microbial organism can have one, two, three, four, five, six, seven, eight, nine or ten, up to all nucleic acids encoding the enzymes or proteins constituting a lasso peptide biosynthetic pathway disclosed herein. In some embodiments, the non-naturally occurring microbial organisms also can include other genetic modifications that facilitate or optimize lasso peptide biosynthesis or that confer other useful functions onto the host microbial organism. One such other functionality can include, for example, augmentation of the synthesis of one or more of the lasso peptide pathway precursors, such as amino acids. 94 ACTIVE 705331286v1
[0274] Generally, a host microbial organism is selected such that it produces the lasso precursor peptide, either as a naturally produced molecule or as an engineered product that either provides de novo production of a desired biosynthesis precursor or increased production of a biosynthesis precursor naturally produced by the host microbial organism. For example, amino acids are produced naturally in a host organism such as E. coli. A host organism can be engineered to increase production of one or more amino acids in order to increase production of lasso precursor, as disclosed herein. Alternatively, a host organism can be engineered to produce a non-natural amino acid that is incorporated into the lasso precursor peptide (Piscotta, F.J., et al., Chem. Commun., 2015, 51, 409-412; Al-Toma, R.S., et al., ChemBioChem 2015, 16, 503 – 509). In addition, a microbial organism that has been engineered to produce a desirable lasso precursor peptide can be used as a host organism and further engineered to express enzymes or proteins that processes the lasso precursor peptide into matured lasso peptides containing non-natural amino acids.
[0275] In some embodiments, a non-naturally occurring microbial organism is generated from a host that contains the enzymatic capability to synthesize a lasso peptide or lasso peptide analog. In this specific embodiment it can be useful to increase the synthesis or accumulation of a lasso peptide pathway intermediate or product to, for example, drive lasso peptide pathway reactions toward lasso peptide production. Increased synthesis or accumulation can be accomplished by, for example, overexpression of nucleic acids encoding one or more of the above-described lasso peptide pathway enzymes or proteins. Over expression the enzyme or enzymes and / or protein or proteins of the lasso peptide pathway can occur, for example, through exogenous expression of the endogenous gene or genes, or through exogenous expression of the heterologous gene or genes. Therefore, naturally occurring organisms can be readily engineered to be non-naturally occurring microbial organisms for producing lasso peptides or lasso peptide analogs, through overexpression of one, two, three, four, five, six, seven, eight, nine, or ten, that is, up to all nucleic acids encoding a lasso peptide biosynthetic pathway enzymes or proteins. In addition, a non- naturally occurring organism can be generated by mutagenesis of an endogenous gene that results in an increase in activity of an enzyme in the lasso peptide biosynthetic pathway.
[0276] In particularly useful embodiments, exogenous expression of the encoding nucleic acids is employed. Exogenous expression confers the ability to custom tailor the expression and / or regulatory elements to the host and application to achieve a desired expression level that is controlled by the user. However, endogenous expression also can be utilized in other embodiments such as by removing a negative regulatory effector or induction of the gene’s 95 ACTIVE 705331286v1promoter when linked to an inducible promoter or other regulatory element. Thus, an endogenous gene having a naturally occurring inducible promoter can be up-regulated by providing the appropriate inducing agent (Daniel-Ivad, M. et al., ACS Chem. Biol.2017, 12, 628−634), or the regulatory region of an endogenous gene can be engineered to incorporate an inducible regulatory element, thereby allowing the regulation of increased expression of an endogenous gene at a desired time. Similarly, an inducible promoter can be included as a regulatory element for an exogenous gene introduced into a non-naturally occurring microbial organism.
[0277] It is understood that any of the one or more exogenous nucleic acids can be introduced into a microbial organism to produce a non-naturally occurring microbial organism of the invention. The nucleic acids can be introduced so as to confer, for example, a lasso peptide biosynthetic pathway onto the microbial organism. Alternatively, encoding nucleic acids can be introduced to produce an intermediate microbial organism having the biosynthetic capability to catalyze some of the required reactions to confer lasso peptide biosynthetic capability. For example, a non-naturally occurring microbial organism having a lasso peptide biosynthetic pathway can comprise at least one exogenous nucleic acid encoding desired enzymes or proteins, such as the linear lasso precursor peptide, or alternatively a combination of a lasso peptide peptidase and a lasso peptide cyclase. Thus, it is understood that any combination of one or more genes encoding one or more peptides, enzymes, or proteins of a biosynthetic pathway can be included in a non-naturally occurring microbial organism of the invention. Similarly, it is understood that any combination of two or more enzymes or proteins of a biosynthetic pathway can be included in a non-naturally occurring microbial organism of the invention, for example, lasso peptide peptidase and a lasso peptide cyclase, and so forth, as desired, so long as the combination of enzymes and / or proteins of the desired biosynthetic pathway results in production of the corresponding desired product.
[0278] In addition to the biosynthesis of lasso peptides as described herein, the non- naturally occurring microbial organisms also can be utilized in various combinations with each other and with other microbial organisms and methods well known in the art to achieve product biosynthesis by other routes. For example, one alternative to produce a lasso peptide other than use of the lasso peptide producers is through addition of another microbial organism capable of converting a lasso peptide pathway intermediate into a lasso peptide or lasso peptide analog. One such procedure includes, for example, the fermentation of a microbial organism that produces a linear lasso precursor peptide. The linear lasso precursor 96 ACTIVE 705331286v1peptide can then be used as a substrate for a second microbial organism that converts the linear lasso precursor peptide to a lasso peptide. The linear lasso precursor peptide can be added directly to another culture of the second organism or the original culture of the linear lasso precursor peptide producers can be depleted of these microbial organisms by, for example, cell separation, and then subsequent addition of the second organism to the fermentation broth can be utilized to produce the final product without intermediate purification steps. Alternatively, a lasso peptide also can be biosynthetically produced from microbial organisms through co-culture or co-fermentation using two organisms in the same vessel, where the first microbial organism produces a linear lasso precursor peptide and the second microbial organism converts the intermediate to a lasso peptide. Alternatively, a lasso peptide also can be biosynthetically produced by first chemically synthesizing the linear lasso precursor peptide, followed by addition of the chemically synthesized linear lasso precursor peptide to a fermentation broth using one or more organisms in the same vessel, where linear lasso precursor peptide is converted to a lasso peptide. Alternatively, a lasso peptide also can be biosynthetically produced from microbial organisms through cell-free biosynthesis of the linear lasso precursor peptide, followed by addition of the linear lasso precursor peptide fermentation broth using one or more organisms in the same vessel, where linear lasso precursor peptide is converted into a lasso peptide. Alternatively, a lasso peptide also can be biosynthetically produced by first chemically synthesizing the linear lasso precursor peptide, followed by addition of the chemically synthesized linear lasso precursor peptide to a broth containing the isolated biosynthetic enzymes, including but not limited to one or more of a lasso peptide peptidase, a lasso peptide cyclase, and lasso peptide RRE, wherein the linear lasso precursor peptide is converted to a lasso peptide. Alternatively, a lasso peptide also can be biosynthetically produced by first producing the linear lasso precursor peptide by cell- free biosynthesis methods, followed by addition of the linear lasso precursor peptide to a broth containing the isolated biosynthetic enzymes, including but not limited to one or more of a lasso peptide peptidase, a lasso peptide cyclase, and lasso peptide RRE, wherein the linear lasso precursor peptide is converted to a lasso peptide.
[0279] Given the teachings and guidance provided herein, those skilled in the art will understand that a wide variety of combinations and permutations exist for the non-naturally occurring microbial organisms, together with other microbial organisms, with the co-culture of other non-naturally occurring microbial organisms having sub-pathways and with combinations of other chemical and / or biochemical procedures well known in the art to produce a lasso peptide. 97 ACTIVE 705331286v1
[0280] In some embodiments, the fusion protein comprised a lasso precursor peptide or a lasso core peptide fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide is fused to the N-terminus of the lasso precursor peptide or lasso core peptide. In some embodiments, the one or more additional peptide or polypeptide is fused at the C-terminus of the lasso precursor peptide or lasso core peptide. In some embodiments, a polynucleotide encoding the fusion protein comprises a nucleic acid sequence encoding the lasso precursor peptide or the lasso core peptide, wherein the 5’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, a polynucleotide encoding the fusion protein comprises a nucleic acid sequence encoding the lasso precursor peptide or the lasso core peptide, wherein the 3’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, the fusion protein comprises an amino acid linker between the lasso precursor peptide or lasso core peptide and the one or more additional peptide or polypeptide. In some embodiments, the fusion protein does not comprise an amino acid linker between the lasso precursor peptide or lasso core peptide and the one or more additional peptide or polypeptide.
[0281] In some embodiments, the fusion protein comprised a lasso precursor peptide or a lasso core peptide fused to one or more lasso peptide biosynthesis components. In some embodiments, the one or more lasso peptide biosynthesis components are selected from (i) a lasso peptidase; (ii) a lasso cyclase; (iii) a RRE; or (iv) any combinations of (i) to (iii). In some embodiments, the one or more lasso peptide biosynthesis components are encoded by the same lasso peptide biosynthetic gene cluster. In other embodiments, the one or more lasso peptide biosynthesis components are encoded by different lasso peptide biosynthetic gene cluster.
[0282] In some embodiments, the fusion protein comprises an amino acid linker between the lasso peptidase or the lasso cyclase and the one or more additional peptide or polypeptide. In some embodiments, the fusion protein does not comprise an amino acid linker between the lasso peptidase or the lasso cyclase and the one or more additional peptide or polypeptide.
[0283] In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso peptidase. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso cyclase. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a RRE. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso peptidase and a lasso cyclase. In 98 ACTIVE 705331286v1specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso peptidase and a RRE. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso cyclase and a RRE. In specific embodiments, the fusion protein comprises a lasso precursor peptide fused to a lasso peptidase, a lasso cyclase and RRE. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso peptidase. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso cyclase. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a RRE. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso peptidase and a lasso cyclase. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso peptidase and a RRE. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso cyclase and a RRE. In specific embodiments, the fusion protein comprises a lasso core peptide fused to a lasso peptidase, a lasso cyclase and RRE.
[0284] In some embodiments, the fusion protein comprises a lasso precursor peptide or a lasso core peptide fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a peptide or polypeptide that facilitates production of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom through cell-free biosynthesis. Examples of peptide or polypeptide that can be fused with a lasso precursor peptide or a lasso core peptide according to the present disclosure include but are not limited to (i) a peptide or polypeptide that increases the level of transcription of the lasso precursor peptide or lasso core peptide in the microbial organism; (ii) a peptide or polypeptide that increases the level of translation of the lasso precursor peptide or lasso core peptide in the microbial organism; (iii) a peptide or polypeptide that facilitates the processing of the lasso precursor peptide or lasso core peptide into the lasso peptide; (iv) a peptide or polypeptide that improves stability of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; (v) a peptide or polypeptide that improves solubility of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; (vi) a peptide or polypeptide that enables or facilitates the detection of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; (vii) a peptide or polypeptide that enables or facilitates purification of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; (viii) a peptide or polypeptide that enables or facilitates immobilization of the lasso precursor peptide or lasso core peptide or the lasso peptide derived therefrom; or (ix) any combination of (i) to (viii). 99 ACTIVE 705331286v1
[0285] In some embodiments, the fusion protein comprises a lasso precursor peptide or a lasso core peptide fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a biologically active peptide or polypeptide. Examples of biologically active peptide or polypeptide that can be fused with a lasso precursor peptide or lasso core peptide according to the present disclosure include but are not limited to (i) a peptide or polypeptide capable of binding to a target molecule (e.g., an antibody or an antigen); (ii) a peptide or polypeptide that enhance cell permeability of the fusion protein; (iii) a peptide or polypeptide capable of conjugating the fusion protein to at least one additional copy of the fusion protein; (iv) a peptide or polypeptide capable of linking the fusion protein to one or more peptidic or non-peptidic molecule; (v) a peptide or polypeptide capable of modulating activity of the lasso precursor peptide or lasso core peptide; (vi) a peptide or polypeptide capable of modulating activity of the lasso peptide derived from the lasso precursor peptide or the lasso core peptide; or (vii) any combinations of (i) to (vi).
[0286] In some embodiments, the fusion protein comprises a lasso peptidase or a lasso cyclase fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide is fused to the N-terminus of the lasso peptidase or the lasso cyclase. In some embodiments, the one or more additional peptide or polypeptide is fused at the C-terminus of the lasso peptidase or the lasso cyclase. In some embodiments, a polynucleotide encoding the fusion protein comprises a nucleic acid sequence encoding the lasso peptidase or the lasso cyclase, wherein the 5’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, a polynucleotide encoding the fusion protein comprises a nucleic acid sequence encoding the lasso peptidase or the lasso cyclase, wherein the 3’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, the fusion protein comprises an amino acid linker between the lasso peptidase or the lasso cyclase and the one or more additional peptide or polypeptide. In some embodiments, the fusion protein does not comprise an amino acid linker between the lasso peptidase or the lasso cyclase and the one or more additional peptide or polypeptide.
[0287] In some embodiments, the fusion protein comprises a lasso peptidase or a lasso cyclase fused to one or more additional peptide or polypeptide. In some embodiments, the more additional peptide or polypeptide comprises a peptide or polypeptide encoded by a lasso peptide biosynthetic gene cluster. Examples of peptide or polypeptide that can be fused with 100 ACTIVE 705331286v1a lasso precursor peptide or a lasso core peptide according to the present disclosure include but are not limited to (i) a lasso precursor peptide; (ii) a lasso core peptide; (iii) a lasso peptidase; (iv) a lasso cyclase, (v) a RRE; or (vi) any combinations of (i) to (vi). In specific embodiments, the fusion protein comprises at least one lasso cyclase and at least one lasso peptidase. In specific embodiments, the fusion protein comprises at least one lasso cyclase fused to a RRE. In specific embodiments, the fusion protein comprises at least one lasso peptidase fused to a RRE.
[0288] In some embodiments, the fusion protein comprises a lasso peptidase or a lasso cyclase fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a peptide or polypeptide that facilitates production of the lasso peptidase or lasso cyclase through cell-free biosynthesis. Examples of peptide or polypeptide that can be fused with the lasso peptidase or lasso cyclase according to the present disclosure include but are not limited to (i) a peptide or polypeptide that increases the level of transcription of the lasso peptidase or lasso cyclase in the microbial organism; (ii) a peptide or polypeptide that increases the level of translation of the lasso peptidase or lasso cyclase in the microbial organism; (iii) a peptide or polypeptide that improves stability of the lasso peptidase or lasso cyclase; (vi) a peptide or polypeptide that improves solubility of the lasso peptidase or lasso cyclase; (v) a peptide or polypeptide that enables or facilitates the detection of the lasso peptidase or lasso cyclase; (vi) a peptide or polypeptide that enables or facilitates purification of the lasso peptidase or lasso cyclase; (vii) a peptide or polypeptide that enables or facilitates immobilization of the lasso peptidase or lasso cyclase; or (viii) any combination of (i) to (vii).
[0289] In some embodiments, the fusion protein comprises a lasso peptidase or a lasso cyclase fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a biologically active peptide or polypeptide. Examples of biologically active peptide or polypeptide that can be fused with a lasso peptidase or a lasso cyclase according to the present disclosure include but are not limited to (i) a peptide or polypeptide capable of modulating the reaction catalyzing activity of the lasso peptidase or lasso cyclase; (ii) a peptide or polypeptide capable of modulating target specificity of the lasso peptidase or lasso cyclase; (iii) an enzyme having the same or different enzymatic activity as the lasso peptidase or lasso cyclase; or any combination of (i) to (iii).
[0290] In some embodiments, the fusion protein comprises a RIPP recognition element (RRE) fused to one or more additional peptide or polypeptide. In some embodiments, the one 101 ACTIVE 705331286v1or more additional peptide or polypeptide is fused to the N-terminus of the RRE. In some embodiments, the one or more additional peptide or polypeptide is fused at the C-terminus of the RRE. In some embodiments, a polynucleotide encoding the fusion protein comprises a nucleic acid sequence encoding the RRE, wherein the 5’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, a polynucleotide encoding the fusion protein comprises a nucleic acid sequence encoding the RRE, wherein the 3’ end of the nucleic acid sequence is linked to a nucleic acid sequence encoding the one or more additional peptide or polypeptide. In some embodiments, the fusion protein comprises an amino acid linker between the RRE and the one or more additional peptide or polypeptide. In some embodiments, the fusion protein does not comprise an amino acid linker between RRE and the one or more additional peptide or polypeptide.
[0291] In some embodiments, the fusion protein comprises a RIPP recognition element (RRE) fused to one or more additional peptide or polypeptide. In some embodiments, the more additional peptide or polypeptide comprises a peptide or polypeptide encoded by a lasso peptide biosynthetic gene cluster. Examples of peptide or polypeptide that can be fused with a lasso precursor peptide or a lasso core peptide according to the present disclosure include but are not limited to (i) a lasso precursor peptide; (ii) a lasso core peptide; (iii) a lasso peptidase; (iv) a lasso cyclase, (v) a RRE; or (vi) any combinations of (i) to (vi). In specific embodiments, the fusion protein comprises at least one lasso precursor peptide fused to a RRE. In specific embodiments, the fusion protein comprises at least one lasso core peptide fused to a RRE. In specific embodiments, the fusion protein comprises at least one lasso cyclase fused to a RRE. In specific embodiments, the fusion protein comprises at least one lasso peptidase fused to a RRE.
[0292] In some embodiments, the fusion protein comprises a RIPP recognition element (RRE) fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a peptide or polypeptide that facilitates production of the RRE through cell-free biosynthesis. Examples of peptide or polypeptide that can be fused with the RRE according to the present disclosure include but are not limited to (i) a peptide or polypeptide that increases the level of transcription of the RRE in the microbial organism; (ii) a peptide or polypeptide that increases the level of translation of the RRE in the microbial organism; (iii) a peptide or polypeptide that improves stability of the RRE; (vi) a peptide or polypeptide that improves solubility of the RRE; (v) a peptide or polypeptide that enables or facilitates the detection of the RRE; (vi) a peptide or polypeptide 102 ACTIVE 705331286v1that enables or facilitates purification of the RRE; (vii) a peptide or polypeptide that enables or facilitates immobilization of the RRE; or (viii) any combination of (i) to (vii).
[0293] In some embodiments, the fusion protein comprised a RIPP recognition element (RRE) fused to one or more additional peptide or polypeptide. In some embodiments, the one or more additional peptide or polypeptide comprises a biologically active peptide or polypeptide. Examples of biologically active peptide or polypeptide that can be fused with a RRE according to the present disclosure include but are not limited to (i) a peptide or polypeptide capable of modulating the reaction catalyzing activity of the lasso peptidase or lasso cyclase; (ii) a peptide or polypeptide capable of modulating target specificity of the lasso peptidase or lasso cyclase; (iii) an enzyme having the same or different enzymatic activity as the lasso peptidase or lasso cyclase; or any combination of (i) to (iii).
[0294] In particular embodiments, the lasso precursor peptide genes or lasso core peptide genes are fused at the 5’-terminus of the DNA template strand of the gene to oligonucleotide sequences that encode peptides or proteins, with or without a linker, such as sequences encoding peptide tags for affinity purification or immobilization, including his-tags, a strep- tags, or FLAG-tags. In some embodiments, the lasso precursor peptides, lasso core peptides, or lasso peptides are fused at the C-terminus of the core peptides to form conjugates with other peptides or proteins, with or without a linker, such as peptide tags for affinity purification or immobilization, including his-tags, a strep-tags, or FLAG-tags.
[0295] In particular embodiments, lasso precursor peptides, lasso core peptides, or lasso peptides are fused to molecules that can enhance cell permeability or penetration into cells, for example through the use of arginine-rich cell-penetrating peptides such as TAT peptide, penetratin, and flock house virus (FHV) coat peptide (Brock, R., Bioconjug. Chem., 2014, 25, 863–868). In particular embodiments, a lasso precursor peptide gene or core peptide gene is fused at the 3’-terminus to oligonucleotide sequences that encode arginine-rich cell- penetrating peptides or proteins, including oligonucleotide sequences that encode penetratin, and flock house virus (FHV) coat peptide or similar peptides that contain guanidinium groups or a combination of lysine and guanidinium groups (Wender, P.A., ...
Claims
CLAIMS What is claimed is:
1. A computer-based method for identifying a lasso peptide having optimal binding with a target molecule, the method comprising: (a) providing one or more three-dimensional (3D) model structures of lasso peptides; (b) performing conformational analysis on the one or more 3D model structures of lasso peptides using molecular dynamic simulation algorithms to obtain at least two conformational states of the one or more 3D model structures of lasso peptides; (c) docking the at least two conformational states of the one or more 3D model structures of lasso peptides onto a 3D model structure of the target molecule at a lasso-binding site of the target molecule to form at least two 3D model structures of the lasso peptide-target molecule complex; (d) performing conformational analysis on the at least two 3D model structures of the lasso peptide-target molecule complex using molecular dynamic simulation algorithms to obtain at least two conformational states of the 3D model structure of the lasso peptide-target molecule complex; and (e) selecting the 3D model structure of the lasso peptide-target molecule complex having a favored free energy, thereby identifying a lasso peptide binder candidate.
2. The method of claim 1, further comprising step (a-1): mapping a target-binding epitope onto different locations of the one or more 3D model structures of lasso peptides.
3. The method of claim 2, further comprising computationally grafting the target- binding epitope onto a mapped location of the selected lasso peptide, thereby generating a grafted lasso peptide binder candidate.
4. The method of any one of claims 1-3, wherein the at least two conformational states of the one or more 3D model structures of lasso peptides are selected from 2, 3, 4, 5, 188 ACTIVE 705331286v16, 7, 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 45 or 50 or more different conformational states.
5. The method of any one of claims 1-4, wherein the favored free energy comprises a lower free energy compared to a different 3D model structure of the lasso peptide- target molecule complex.
6. The method of any one of claims 1-5, wherein the favored free energy is the lowest free energy of the two or more 3D model structure of the lasso peptide-target molecule complex.
7. The method of any one of claims 1 to 6, wherein step (a) comprises retrieving one or more known 3D structures of lasso peptides from a protein structure database.
8. The method of any one of claims 1 to 7, wherein step (a) comprises computationally modeling the 3D structure of a lasso peptide based on atomic coordinates of the lasso peptide.
9. The method of claim 8, wherein the atomic coordinates of the lasso peptides are obtained from a protein structure database or scientific literature.
10. The method of claim 7 or 9, wherein the protein structure database is selected from worldwide Protein Data Bank (wwPDB), Cambridge Structure Database, Molecular Model Database of National Center for Biotechnology Information (NCBI), and Biological Magnetic Resonance Data Bank (BMRB) database.
11. The method of claim 8, wherein the atomic coordinates of the lasso peptide are obtained by subjecting the lasso peptide to nuclear magnetic resonance (NMR) analysis, X-ray crystallography, neutron diffraction, or 3-dimensional electron microscopy (3D-EM).
12. The method of claim 11, wherein the X-ray crystallography is serial femtosecond crystallography. 189 ACTIVE 705331286v113. The method of claim 11, wherein the 3D-EM is cryogenic electron microscopy (cryo- EM).
14. The method of any one of claims 1 to 13, wherein step (a) comprises computationally modeling the 3D structure of the lasso peptide based on X-ray diffraction data and / or nuclear magnetic resonance (NMR) data of the lasso peptide and atomic coordinates of a reference lasso peptide, and wherein the lasso peptide has at least 50% amino acid sequence identity to the reference lasso peptide.
15. The method of claim 14, wherein the X-ray diffraction data of the lasso peptide are obtained by subjecting a crystal of the lasso peptide to X-ray crystallography analysis.
16. The method of claim 15, wherein step (a) comprises (i) obtaining atomic coordinates of the lasso peptide based on the X-ray diffraction data; and (ii) refining the atomic coordinates of the lasso peptide based on the atomic coordinates of the reference lasso peptide.
17. The method of claim 14, wherein the nuclear magnetic resonance (NMR) data of the lasso peptide comprises NMR chemical shift, J-coupling constant, and resonance intensity obtained by subjecting a solution of the lasso peptide to NMR analysis.
18. The method of claim 17, wherein step (a) comprises (i) creating an ensemble of structural models of the lasso peptide based on the NMR data; (ii) obtaining mean atomic coordinates of the lasso peptide based on the ensemble of structural models; and (iii) refining the atomic coordinates of the lasso peptide based on the atomic coordinates of the reference lasso peptide.
19. The method of any one of claims 14 to 18, wherein the atomic coordinates of the reference lasso peptide are obtained from a protein structure database or scientific literature. 190 ACTIVE 705331286v120. The method of claim 19, wherein the protein structure database is selected from worldwide Protein Data Bank (wwPDB), Cambridge Structure Database, Molecular Model Database of National Center for Biotechnology Information (NCBI), and Biological Magnetic Resonance Data Bank (BMRB) database.
21. The method of any one of claims 1 to 20, wherein step (a) comprises computationally modeling the 3D structure of a lasso peptide based on the amino acid sequence of the lasso peptide and atomic coordinates of a reference lasso peptide; and wherein the lasso peptide has at least 50 percent (%) amino acid sequence identity to the reference lasso peptide.
22. The method of claim 21, wherein the atomic coordinates of the reference lasso peptide are obtained from a protein structure database or scientific literature.
23. The method of claim 22, wherein the protein structure database is selected from worldwide Protein Data Bank (wwPDB), Cambridge Structure Database, Molecular Model Database of National Center for Biotechnology Information (NCBI), and Biological Magnetic Resonance Data Bank (BMRB) database.
24. The method of any one of claims 21 to 23, wherein computationally modeling the 3D structure of the lasso peptide is performed by homology modeling.
25. The method of claim 24, wherein homology modeling is performed in combination with a protein structure prediction algorithm, wherein optionally the protein structure prediction algorithm is trRosetta or AlphaFold or AlphaFold 2.
26. The method of any one of claims 1 to 25, wherein the one or more 3D model structures of lasso peptides are lasso backbone structures, and wherein step (a) further comprises computationally modeling the lasso backbone structures by removing side chains of each amino acid residue that is not an internal ring-forming residue from the 3D structure of lasso peptides.
27. The method of any one of claims 1 to 6, wherein step (c) further comprises creating the 3D model structure of the target molecule before docking. 191 ACTIVE 705331286v128. The method of claim 27, wherein creating the model 3D structure of the target molecule comprises computationally modeling the 3D structure of the target molecule based on atomic coordinates of the target molecule.
29. The method of claim 28, wherein the atomic coordinates of the target molecule are obtained from a protein structure database or scientific literature.
30. The method of claim 29, wherein the protein structure database is selected from worldwide Protein Data Bank (wwPDB), Cambridge Structure Database, Molecular Model Database of National Center for Biotechnology Information (NCBI), and Biological Magnetic Resonance Data Bank (BMRB) database.
31. The method of claim 29, wherein the atomic coordinates of the target molecule are obtained by subjecting the target molecule to nuclear magnetic resonance (NMR) analysis, X-ray crystallography, neutron diffraction analysis, or 3-dimensional electron microscopy (3D-EM).
32. The method of claim 31, wherein the X-ray crystallography is serial femtosecond crystallography.
33. The method of claim 31, wherein the 3D-EM is cryogenic electron microscopy (cryo- EM).
34. The method of claim 27, wherein creating the 3D model structure of the target molecule comprises computationally modeling the 3D structure of the target molecule based on X-ray diffraction data and / or nuclear magnetic resonance data of the target molecule and atomic coordinates of a reference polypeptide; wherein the reference polypeptide has at least 50% amino acid sequence identity to the reference polypeptide.
35. The method of claim 34, wherein the X-ray diffraction data of the target molecule are obtained by subjecting a crystal of the target molecule to X-ray crystallography analysis. 192 ACTIVE 705331286v136. The method of claim 34 or 35, wherein creating the 3D model structure of the target molecule comprises: (i) obtaining atomic coordinates of the target molecule based on the X-ray diffraction data; (ii) refining the atomic coordinates of the target molecule based on the atomic coordinates of the reference polypeptide; and (iii) computationally modeling the 3D structure of the target molecule based on the refined atomic coordinates.
37. The method of claim 34, wherein the nuclear magnetic resonance (NMR) data of the target molecule comprises NMR chemical shift, J-coupling constant, and resonance intensity obtained by subjecting a solution of the target molecule to NMR analysis.
38. The method of claim 34 or 37, wherein creating the 3D model structure of the target molecule comprises: (i) creating an ensemble of structural models of the target molecule based on the NMR data; (ii) obtaining mean atomic coordinates of the target molecule based on the ensemble of structural models; (iii) refining the mean atomic coordinates of the target molecule based on the atomic coordinates of the reference polypeptide; and (iv) computationally modeling the 3D structure of the target molecule based on the refined mean atomic coordinates.
39. The method of any one of claims 34 to 38, wherein the atomic coordinates of the reference polypeptide are obtained from a protein structure database or scientific literature.
40. The method of claim 39, wherein the protein structure database is selected from worldwide Protein Data Bank (wwPDB), Cambridge Structure Database, Molecular Model Database of National Center for Biotechnology Information (NCBI), and Biological Magnetic Resonance Data Bank (BMRB) database. 193 ACTIVE 705331286v141. The method of claim 27, wherein creating the 3D model structure of the target molecule comprises computationally modeling the 3D structure of the target molecule based on the amino acid sequence of the target molecule and atomic coordinates of one or more reference polypeptide; and wherein the amino acid sequence of the target molecule is at least about 50% identical to the amino acid sequence of the reference polypeptide.
42. The method of claim 41, wherein the atomic coordinates of the reference polypeptide are obtained from a protein structure database or scientific literature.
43. The method of claim 42, wherein the protein structure database is selected from worldwide Protein Data Bank (wwPDB), Cambridge Structure Database, Molecular Model Database of National Center for Biotechnology Information (NCBI), and Biological Magnetic Resonance Data Bank (BMRB) database.
44. The method of any one of claims 41 to 43, wherein computationally modeling the 3D structure of the target molecule is performed using homology modeling.
45. The method of claim 44, wherein homology modeling is performed in combination with a protein structure prediction algorithm, wherein optionally the protein structure prediction algorithm is trRosetta or AlphaFold or AlphaFold 2.
46. The method of any one of claims 1 to 45, wherein the lasso-binding site is selected from a ligand binding site, a substrate binding site, a catalytic binding site, a co-factor binding site, a precursor binding site, an orthosteric binding site, an allosteric binding site, an open conformation binding site, an active conformation binding site, an inactive conformation binding site, and a closed conformation binding site of the target molecule.
47. The method of any one of claims 1 to 46, wherein step (c) further comprises, for each 3D model structure of the lasso peptide: (c-1) positioning a docked portion of at least one conformational state of the 3D model structure of the lasso peptide relative to the lasso-binding site of the 3D model 194 ACTIVE 705331286v1structure of the target molecule in a first docked pose, thereby forming a docked system; (c-2) scoring the docked system based on structural complementarity between the docked portion and the lasso-binding site; (c-3) repositioning the docked portion relative to the lasso-binding site to a second docked pose, and repeating step (c-2); (c-4) repeating step (c-3) for one or more times; and (c-5) selecting the docked system having the highest score as the optimal docked system before proceeding to step (d).
48. The method of claim 47, wherein the docked system comprises one or more complementary binding pairs; and wherein the complementary binding pair comprises a first binding moiety on the lasso peptide and a second binding moiety on the target molecule.
49. The method of claim 48, wherein the first binding moiety is on a first amino acid residue of the lasso peptide, and the second binding moiety is on a second amino acid residue of the target molecule; and wherein the distance between any atom of the first amino acid residue and any atom of the second amino acid residue is less than about 5 Ångströms, and optionally less that about 4 Ångströms, and optionally less than about 3 Ångströms.
50. The method of claim 48 or 49, wherein the complementary binding pair forms binding interaction that is selected from a hydrogen bonding interaction, an ionic or electrostatic bonding interaction, polar interaction, dipolar interaction, induced dipolar interaction, pi stacking interaction, hydrophobic interaction, and / or van der Waals interaction.
51. The method of any one of claims 47 to 50, wherein step (c-2) comprises calculating a total binding free energy of the docked system using one or more molecular mechanics force field functions; and assigning a score to the docked system based on the total binding free energy, and wherein the score relates to the total binding free energy. 195 ACTIVE 705331286v152. The method of claim 51, wherein the molecular mechanics force field functions are selected from Amber force fields AMBER99, Amber 10EHT, ff14SB, ff19SB, and the Merck molecular force field MMFF94x.
53. The method of any one of claims 47 to 52, wherein step (c-3) comprises identifying the second docked pose using an energy minimizing function before repositioning the docked portion into the second docked pose; wherein the docked system is predicted to have a lower binding free energy in the second docked pose than the first docked pose based on the energy minimizing function.
54. The method of claim 52, wherein the energy minimizing function is selected from the steepest descent algorithm, conjugate gradients algorithm, L-BFGS (limited-memory Broyden-Fletcher-Goldfarb-Shanno) algorithm, and genetic algorithms.
55. The method of any one of claims 47 to 54, wherein the docked portion comprises at least one amino acid residue from the ring portion, the loop portion, and / or the tail portion of the lasso peptide.
56. The method of any one of claims 2 to 55, wherein step (a-1) comprises: (a-1-1) in the optimal docked system, identifying one or more lasso-binding moieties in the lasso-binding site of the target molecule; (a-1-2) selecting one or more amino acid residues comprising one or more target- binding moieties complementary to the one or more lasso-binding moieties; and (a-1-3) mapping an optimal set of positions in the amino acid sequence of the selected lasso peptide for grafting the one or more selected amino acid residues; wherein the grafting places the one or more target-binding moieties at suitable spatial locations and orientations for binding with the complementary lasso-binding moieties on the target molecule.
57. The method of claim 56, wherein step (a-1-3) comprises: (a-1-3-1) computationally grafting the one or more selected amino acid residues into the amino acid sequence of the selected lasso peptide at a first set of positions; (a-1-3-2) calculating a total binding free energy of the optimal docked system using one or more molecular force field functions; 196 ACTIVE 705331286v1(a-1-3-3) modifying at least one position in the first set of positions thereby obtaining an adjusted set of positions, and computationally grafting the one or more selected amino acid residues into the amino acid sequence of the selected lasso peptide at the adjusted set of positions; (a-1-3-4) repeating step (a-1-3-2); (a-1-3-5) repeating steps (a-1-3-3) and (a-1-3-4) sequentially for one or more times; and (a-1-3-6) selecting the adjusted set of positions associated with the lowest total binding free energy as the optimal map of positions.
58. The method of claim 57, wherein the molecular mechanics force field functions are selected from Amber force fields AMBER99, Amber 10EHT, ff14SB, ff19SB, and the Merck molecular force field MMFF94x.
59. The method of claim 57 or 58, wherein step (a-1-3-3) comprises identifying the adjusted set of positions using an energy minimizing function before modifying the first set of positions, wherein the docked system having the one or more selected amino acid residues grafted into the amino acid sequence of the selected lasso peptide at the adjusted set of positions is predicted to have a lower binding free energy than at the first set of positions based on the energy minimizing function.
60. The method of claim 59, wherein the energy minimizing function is selected from the steepest descent algorithm, conjugate gradients algorithm, L-BFGS (limited-memory Broyden-Fletcher-Goldfarb-Shanno) algorithm, and genetic algorithms.
61. The method of any one of claims 2 to 60, wherein the target-binding epitope corresponds to a fragment or fragments of a naturally-existing ligand of the target molecule, wherein the fragment or fragments are capable of binding with the lasso- binding site of the target molecule and forming a ligand-target interface.
62. The method of claim 61, wherein step (a-1) comprises aligning the docked portion of the 3D model structure of the lasso peptide with a 3D model structure of the fragment or fragments in the ligand-target interface. 197 ACTIVE 705331286v163. The method of claim 62, wherein step (a-1) comprises adjusting the spatial position, conformation and / or orientation of at least one target-binding moiety in the docked portion to mimic a corresponding binding moiety of the fragment or fragments in the ligand-target interface.
64. The method of any one of claims 3 to 58, further comprising: (f) mutating one or more amino acid residues of the lasso peptide binder candidate to produce a first set of lasso peptide binder variants; and (g) ranking the first set of lasso peptide binder variants based on a predicted binding affinity for binding with the target molecule.
65. The method of claim 64, wherein step (f) further comprises adjusting conformation of the first set of lasso peptide binder variants to produce a second set of lasso peptide binder variants; and wherein step (g) comprises ranking the second set of lasso peptide binder variants based on the predicted binding affinity for binding with the target molecule.
66. The method of claim 64 or 65, wherein in step (f), mutating the amino acid residue of the lasso peptide binder candidate comprises replacing the side chain of the amino acid residue of the lasso peptide binder candidate with the side chain of a second amino acid that is different from the mutated amino acid residue.
67. The method of claim 66, wherein the second amino acid is a naturally-occurring or a non-natural amino acid.
68. The method of claim 64 or 65, wherein in step (f), mutating the amino acid residue of the lasso peptide binder candidate comprises modifying one or more chemical moieties on the side chain of the mutated amino acid residue.
69. The method of claim 68, wherein at least one modified chemical moiety is a target- binding moiety.
70. The method of claim 64 or 65, wherein in step (f), mutating the amino acid residue of the lasso peptide binder candidate comprises modifying one or more side chains of the 198 ACTIVE 705331286v1lasso peptide binder candidate to complement one or more lasso-binding moieties on the target molecule; and wherein the modifying is selected from the group consisting of: (i) incorporating into the lasso peptide a neutral or basic side chain that is structurally opposite to an acidic lasso-binding moiety; (ii) incorporating into the lasso peptide a neutral or acid side chain that is structurally opposite to a basic lasso-binding moiety; (iii) incorporating into the lasso peptide a neutral or positively charged side chain that is structurally opposite to a negatively charged lasso-binding moiety; (iv) incorporating into the lasso peptide a neutral or negatively charged side chain that is structurally opposite to a positively charged lasso-binding moiety; (v) incorporating into the lasso peptide a sterically smaller side chain that is structurally opposite to a sterically larger lasso-binding moiety; (vi) incorporating into the lasso peptide a sterically larger side chain that is structurally opposite to a sterically smaller lasso-binding moiety; (vii) incorporating into the lasso peptide a hydrophobic side chain, preferably a similarly hydrophobic side chain, that is structurally opposite to a hydrophobic lasso- binding moiety; (viii) incorporating into the lasso peptide an acidic, H-bond donor, positively charged, or aromatic side chain that is structurally opposite to an aromatic lasso- binding moiety; (ix) incorporating into the lasso peptide an electron-rich aromatic side chain that is structurally opposite to an electron-poor aromatic lasso-binding moiety; (x) incorporating into the lasso peptide an acid or electron-poor aromatic side chain that is structurally opposite to an electron-rich aromatic lasso-binding moiety; (xi) incorporating into the lasso peptide an inducible dipole or multipole side chain that is structurally opposite to an inducible dipole or multipole lasso-binding moiety; (xii) incorporating into the lasso peptide a permanent, inducible dipole or multipole side chain that is structurally opposite to an inducible or multipole lasso-binding moiety; (xiii) incorporating into the lasso peptide an H-bond acceptor side chain that is structurally opposite to an H-bond donor lasso-binding moiety, and (xiv) incorporating into the lasso peptide an H-bond donor side chain that is structurally opposite to an H-bond acceptor lasso-binding moiety. 199 ACTIVE 705331286v171. The method of any one of claims 64 to 70, wherein the mutated amino acid residue is in the (i) ring portion of the lasso peptide; (ii) loop portion of the lasso peptide; (iii) tail portion of the lasso peptide; or (iv) any combination of (i) to (iii).
72. The method of any one of claims 1 to 71, further comprising: (h) synthesizing the lasso peptide binder candidate, grafted lasso peptide binder candidate or one or more lasso peptide binder variants having the highest rankings.
73. The method of claim 72, wherein in step (h), synthesizing the lasso peptide binder candidate or the one or more lasso peptide binder variants is performed using a cell-free or cell-based synthesis method.
74. The method of any one of claims 1 to 73, further comprising creating a database of optimized lasso peptide structures or intermediates thereof generated in any one of steps (a) to (h). 200 ACTIVE 705331286v1