Identifying amino acid sequences predicted to form a scaffold-bound peptide ligand that binds to a target molecule

A computer-implemented method using machine learning algorithms generates peptide sequences with high binding affinity and specificity to targets, addressing the inefficiencies of traditional peptide screening methods.

WO2026017809A1PCT designated stage Publication Date: 2026-01-22BICYCLETX LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/070522
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-19
Filing Date
2025-07-17
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing methods for identifying peptides that bind to scaffolds are time and resource intensive, particularly phage display-based screening.

Method used

A computer-implemented method for identifying amino acid sequences that form scaffold-bound peptide ligands by generating peptide sequences based on predefined sequence constraints and interaction characteristics, using machine learning algorithms to predict binding affinity and stability.

Benefits of technology

Facilitates the rapid identification of peptide sequences with high binding affinity and specificity to targets, reducing the time and resource requirements of traditional screening methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025070522_22012026_PF_FP_ABST
    Figure EP2025070522_22012026_PF_FP_ABST
Patent Text Reader

Abstract

According to an aspect of the disclosure, there is provided a computer implemented method for identifying amino acid sequences predicted to form a scaffold-bound peptide ligand that binds to a target, the method comprising: receiving as input a motif sequence, which is a first amino acid sequence describing at least part of a binding motif predetermined to interact with one or more predefined targets; generating a peptide sequence, which is a second amino acid sequence comprising the motif sequence, and describing a peptide, the second amino acid sequence satisfying one or more predefined sequence constraints such that the peptide binds to one of one or more predefined scaffolds in at least two locations, to form a scaffold-bound peptide ligand; determining at least one predicted interaction characteristic relating to an interaction between the peptide and the one or more predefined targets; determining a favourability of the at least one predicted interaction characteristic; and providing output data identifying one or more amino acid sequences predicted to form a scaffold-bound peptide ligand that binds to a target, based on the determined favourability.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] IDENTIFYING AMINO ACID SEQUENCES PREDICTED TO FORM A SCAFFOLD-BOUND PEPTIDE LIGAND THAT BINDS TO A TARGET MOLECULE

[0002] TECHNICAL FIELD

[0003] The present invention relates to computer implemented methods of identifying amino acid sequences likely to form scaffold-bound peptide ligands that bind to a target. Such methods may be used for identifying novel therapeutics in the form of scaffold-bound peptide ligands.

[0004] BACKGROUND ART

[0005] Cyclic peptides are able to bind with high affinity and target specificity to protein targets and hence are an attractive molecule class for the development of therapeutics. In fact, several cyclic peptides are already successfully used in the clinic, as for example the antibacterial peptide vancomycin, the immunosuppressant drug cyclosporine or the anticancer drug octreotide (Driggers et al. (2008), Nat Rev Drug Discov 7 (7), 608-24). Good binding properties result from a relatively large interaction surface formed between the peptide and the target as well as the reduced conformational flexibility of the cyclic structures. Typically, macrocycles bind to surfaces of several hundred square angstrom, as for example the cyclic peptide CXCR4 antagonist CVX15 (400 A2; Wu et al. (2007), Science 330, 1066-71), a cyclic peptide with the Arg-Gly-Asp motif binding to integrin aVb3 (355 A2) (Xiong et al. (2002), Science 296 (5565), 151-5) or the cyclic peptide inhibitor upain-1 binding to urokinase-type plasminogen activator (603 A2; Zhao et al. (2007), J Struct Biol 160 (1), 1-10).

[0006] Due to their cyclic configuration, peptide macrocycles are less flexible than linear peptides, leading to a smaller loss of entropy upon binding to targets and resulting in a higher binding affinity. The reduced flexibility also leads to locking target-specific conformations, increasing binding specificity compared to linear peptides. This effect has been exemplified by a potent and selective inhibitor of matrix metalloproteinase 8, MMP-8) which lost its selectivity over other MMPs when its ring was opened (Cherney et al. (1998), J Med Chem 41 (11), 1749-51). The favourable binding properties achieved through macrocyclization are even more pronounced in multicyclic peptides having more than one peptide ring as for example in vancomycin, nisin and actinomycin. Different research teams have previously tethered polypeptides with cysteine residues to a synthetic molecular structure (Kemp and McNamara (1985), J. Org. Chem; Timmerman et al. (2005), ChemBioChem). Meloen and co-workers had used tris(bromomethyl)benzene and related molecules for rapid and quantitative cyclization of multiple peptide loops onto synthetic scaffolds for structural mimicry of protein surfaces (Timmerman et al. (2005), ChemBioChem). Methods for the generation of candidate drug compounds wherein said compounds are generated by linking cysteine containing polypeptides to a molecular scaffold as for example tris(bromomethyl)benzene are disclosed in WO 2004 / 077062 and WO 2006 / 078161.

[0007] Phage display-based combinatorial approaches have been developed to generate and screen large libraries of bicyclic peptides to targets of interest (Heinis et al. (2009), Nat Chem Biol 5 (7), 502-7 and W02009 / 098450). Briefly, combinatorial libraries of linear peptides containing three cysteine residues and two regions of six random amino acids (Cys-(Xaa)6- Cys-(Xaa)6-Cys) were displayed on phage and cyclised by covalently linking the cysteine side chains to a small molecule (tris-(bromomethyl)benzene).

[0008] However, phage display-based screening is time and resource intensive. Therefore, there is a need for a computational approach to identifying peptides that bind to a scaffold and have favourable characteristics.

[0009] SUMMARY OF THE INVENTION

[0010] According to an aspect of the disclosure, there is provided a computer implemented method for identifying amino acid sequences predicted to form a scaffold-bound peptide ligand that binds to a target, the method comprising: receiving as input a motif sequence, which is a first amino acid sequence describing at least part of a binding motif predetermined to interact with one or more predefined targets; generating a peptide sequence, which is a second amino acid sequence comprising the motif sequence, and describing a peptide, the second amino acid sequence satisfying one or more predefined sequence constraints such that the peptide binds to one of one or more predefined scaffolds in at least two locations, to form a scaffold-bound peptide ligand; determining at least one predicted interaction characteristic relating to an interaction between the peptide and the one or more predefined targets; determining a favourability of the at least one predicted interaction characteristic; and providing output data identifying one or more amino acid sequences predicted to form a scaffold-bound peptide ligand that binds to a target, based on the determined favourability.

[0011] Optionally, the step of generating the peptide sequence comprises extending the motif sequence by adding at least one amino acid either side of the motif sequence. Optionally, the number of amino acids added either side of the motif sequence is determined randomly, subject to one or more of the sequence constraints.

[0012] Optionally, the motif sequence is first extended to form a peptide backbone.

[0013] Optionally, the method further comprises assigning one or more first residues to amino acid positions in the backbone, other than the residues forming the motif sequence, to satisfy one or more of the predefined sequence constraints, to form a partial peptide sequence.

[0014] Optionally, at least two the one or more first residues are configured to bind to one of one or more predefined scaffolds.

[0015] Optionally, at least one of the one or more first residues are assigned to random amino acids in the backbone, subject to the one or more predefined sequence constraints.

[0016] Optionally, the step of generating a peptide sequence comprises further assigning second residues to the remaining amino acid positions in the backbone to form a complete peptide sequence, such that the peptide sequence is configured to fold into a predefined the three- dimensional shape.

[0017] Optionally, the peptide backbone is assigned a predefined three-dimensional shape comprising a predefined three-dimensional shape of the motif sequence and which is suitable for the binding motif binding to the one or more predefined targets.

[0018] Optionally, the second residues are determined by an inverse folding algorithm. Optionally, the inverse folding algorithm is a machine learning algorithm configured to predict an amino acid sequence configured to fold into a predetermined three-dimensional structure.

[0019] Optionally, the step of generating an amino acid sequence, and subsequent steps, are repeated for a plurality of peptide sequences.

[0020] Optionally, the plurality of peptide sequences are generated based on different onedimensional backbones, different three-dimensional backbone shapes, different first residues assigned thereto to form different partial peptide sequences, and / or different second residues assigned thereto to form different complete peptide sequences.

[0021] Optionally, the output data comprises the peptide sequence.

[0022] Optionally, the one or more sequence constraints comprises one or more constraints that increase a likelihood that the peptide will bind to one of one or more predefined scaffolds.

[0023] Optionally, the one or more sequence constraints comprises one or more constraints that increase a likelihood that the peptide will form at least one loop structure when bound with one of the one or more predefined scaffolds.

[0024] Optionally, the one or more sequence constraints comprises a constraint on a total length of the peptide sequence. Optionally, the one or more sequence constraints comprises a constraint that the peptide sequence consists of from 6 to 30 amino acid residues.

[0025] Optionally, the one or more sequence constraints comprises a constraint that the peptide sequence comprises at least two cysteine residues.

[0026] Optionally, the one or more sequence constraints comprises a constraint that the peptide sequence comprises three cysteine residues.

[0027] Optionally, the one or more sequence constraints comprises a constraint on the relative positions of at least two cysteine residues. Optionally, the one or more sequence constraints comprises a constraint that at least one pair of successive cysteine residues are separated by no more than a first predefined number of amino acid residues.

[0028] Optionally, the one or more sequence constraints comprises a constraint that at least one pair of successive cysteine residues are separated by no fewer than a second predefined number of amino acid residues.

[0029] Optionally, the sequence comprises three cysteine residues, including two pairs of successive cysteine residues, and both pairs satisfy at least one of the constraints above regarding relative positions of cysteine residues.

[0030] Optionally, the one or more constraints comprises a constraint on the amino acid residue at one or more terminal positions in the sequence.

[0031] Optionally, the one or more constraints comprises a constraint that the amino acid residue at one or more terminal positions in the sequence is an alanine residue.

[0032] Optionally, the one or more interaction characteristics comprises a binding score based on the likelihood of the peptide binding to the target.

[0033] Optionally, the one or more interaction characteristics comprises a folding score based on the likelihood of the peptide being well-folded.

[0034] Optionally, the one or more interaction characteristics comprises a score for one or more residues of the target, based on the likelihood of each of these residue binding to the peptide.

[0035] Optionally, the one or more interaction characteristics comprises a score for one or more residues of the peptide, based on the likelihood of each of these residues binding to the target. Optionally, the one or more residues includes residues of the motif sequence. Optionally, the one or more interaction characteristics comprises a score for a predicted position of the binding motif relative to the target compared to a position at which the binding motif is predetermined to bind to the target.

[0036] Optionally, the one or more interaction characteristics is determined based on a predicted three-dimensional structure of the peptide and / or the target.

[0037] Optionally, the method further comprises a step of determining one or more structural characteristics of the peptide and a step of determining a favourability of the one or more structural characteristics.

[0038] Optionally, the one or more structural characteristics relate to a likelihood that the peptide will bind with one of one or more predefined molecular scaffolds.

[0039] Optionally, the one or more structural characteristics comprises a distance between two cysteine residues, in the predicted three-dimensional structure of the peptide.

[0040] Optionally, the one or more structural characteristics comprises a distance between respective sulphur atoms of two cysteine residues in the predicted three-dimensional structure of the peptide.

[0041] Optionally, the one or more structural characteristics comprises an angle formed between a first cysteine residue and two further cysteine residues, in the predicted three-dimensional structure of the peptide.

[0042] Optionally, the one or more structural characteristics comprises a dihedral angle formed between a first cysteine residue and a second cystine residue, in the predicted three- dimensional structure of the peptide.

[0043] Optionally, the predicted three-dimensional structure of the peptide is generated by a machine learning algorithm based on input data comprising the generated amino acid sequence. Optionally, the method further comprises a step of determining one or more functional characteristics of the peptide and a step of determining a favourability of the one or more functional characteristics.

[0044] Optionally, the one or more functional characteristics comprises one more of: membrane permeability of peptide and blood-brain barrier permeability of the peptide.

[0045] Optionally, the favourability is determined based on a comparison between the determined characteristics and predefined corresponding desired or optimal parameters.

[0046] Optionally, the favourability is determined by calculating a favourability score.

[0047] Optionally, the favourability score is calculated using an objective function comprising terms corresponding to each of the determined characteristics.

[0048] Optionally, the scaffold-bound peptide ligand is bicyclic, comprising two loop sequences, and is covalently bound to the scaffold at three locations.

[0049] Optionally, the one or more scaffolds comprise one or more of TATA, TATB, TCTZ, TBMB, TTZ, TBAB, TSTA, TCAZ, TCAN, and TCCU.

[0050] Optionally, the peptide forms a peptide ligand, and the peptide in the peptide ligand comprises at least three cysteine residues, separated by at least two loop sequences, and the scaffold forms covalent bonds with the cysteine residues of the peptide such that at least two peptide loops are formed on the scaffold.

[0051] Optionally, the target is a protein, a protein on a cell, a tumour antigen, a viral antigen, bacterial antigen.

[0052] According to a second aspect of the disclosure, there is provided computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of the preceding aspect. According to a third aspect of the disclosure, there is provided a data processing system comprising means for carrying out the method of the first aspect.

[0053] According to a fourth aspect of the disclosure, there is provided a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of the first aspect.

[0054] According to a fifth aspect of the disclosure, there is provided a method of identifying a therapeutic agent or diagnostic agent targeting a target, the therapeutic comprising a scaffold-bound peptide ligand, the method comprising using the computer implemented method of first aspect to identify a peptide forming the scaffold-bound peptide ligand.

[0055] According to a sixth aspect of the disclosure, there is provided a method of making a peptide or scaffold-bound peptide ligand, comprising using the computer implemented method of the first aspect to identify a peptide forming the scaffold-bound peptide ligand or the scaffold-bound peptide ligand.

[0056] According to a seventh aspect of the disclosure, there is provided a scaffold-bound peptide ligand made by the method of the sixth aspect.

[0057] BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Further features of the disclosure will be described below, by way of non-limiting examples and with reference to the accompanying drawings, in which:

[0059] Fig. 1 shows an example method or identifying amino acid sequences;

[0060] Fig. 2 shows an example method of generating a peptide sequences satisfying the sequence constraints;

[0061] Fig. 3 shows an example scaffold-bound peptide ligand;

[0062] Fig. 4 shows the functional activity of ligands in CHO-Ki cells expressing rat SSTR4;

[0063] Fig. 5 shows the percentage inhibition of binding by ligands in CHO-Ki cells expressing human SSTR subtypes. DETAILED DESCRIPTION

[0064] Peptide Ligands

[0065] Fig. 3 schematically shows an example scaffold-bound peptide ligand (also referred to herein as a peptide ligand) according the disclosure. The scaffold-bound peptide ligand comprises a peptide bound to a scaffold molecule.

[0066] A peptide ligand, as referred to herein, refers to a peptide, peptidic or peptidomimetic covalently bound to a molecular scaffold. Typically, such peptides, peptidics or peptidomimetics comprise a peptide having natural or non-natural amino acids, two or more reactive groups (i.e. cysteine residues) which are capable of forming covalent bonds to the scaffold, and a sequence subtended between said reactive groups which is referred to as the loop sequence, since it may form a loop when the peptide, peptidic or peptidomimetic is bound to the scaffold. The peptides, peptidics or peptidomimetics may comprise at least three cysteine residues, and form at least two loops on the scaffold, thus forming bicyclic peptides.

[0067] As shown in Fig. 3, the scaffold may be a trivalent molecule covalently bound, via bonds Li - L3, to cysteine reactive groups, and comprising loop sequences, Loop A and Loop B, subtended between said reactive groups.

[0068] Certain bicyclic peptides have a number of advantageous properties which enable them to be considered as suitable drug-like molecules for injection, inhalation, nasal, ocular, oral or topical administration. Such advantageous properties include:

[0069] - Species cross-reactivity. This is a typical requirement for preclinical pharmacodynamics and pharmacokinetic evaluation;

[0070] - Protease stability. Bicyclic peptide ligands should in most circumstances demonstrate stability to plasma proteases, epithelial ("membrane-anchored") proteases, gastric and intestinal proteases, lung surface proteases, intracellular proteases and the like. Protease stability should be maintained between different species such that a bicyclic peptide lead candidate can be developed in animal models as well as administered with confidence to humans;

[0071] - Desirable solubility profile. This is a function of the proportion of charged and hydrophilic versus hydrophobic residues and intra / inter-molecular H-bonding, which is important for formulation and absorption purposes;

[0072] - An optimal plasma half-life in the circulation. Depending upon the clinical indication and treatment regimen, it may be required to develop a bicyclic peptide with short or prolonged in vivo exposure times for the management of either chronic or acute disease states. The optimal exposure time will be governed by the requirement for sustained exposure (for maximal therapeutic efficiency) versus the requirement for short exposure times to minimise toxicological effects arising from sustained exposure to the agent.

[0073] Molecular Scaffold

[0074] A scaffold molecule may be any molecule configured to bind covalently to and orient a peptide. A scaffold molecule may be a non-aromatic molecular scaffold. The term “nonaromatic molecular scaffold” may refer to any molecular scaffold which does not contain an aromatic (i.e. unsaturated) carbocyclic or heterocyclic ring system.

[0075] Suitable examples of non-aromatic molecular scaffolds are described in Heinis et al (2014) Angewandte Chemie, International Edition 53(6) 1602-1606.

[0076] As noted in the foregoing documents, the molecular scaffold may be a small molecule, such as a small organic molecule. The molecular scaffold may comprise reactive groups that are capable of reacting with functional group(s) of the peptide to form covalent bonds.

[0077] The molecular scaffold may comprise chemical groups which form the linkage with a peptide, such as amines, thiols, alcohols, ketones, aldehydes, nitriles, carboxylic acids, esters, alkenes, alkynes, azides, anhydrides, succinimides, maleimides, alkyl halides and acyl halides. One example of a suitable scaffold is l,3,5-Triacryloylhexahydro-l,3,5-triazine (TATA) (Angewandte Chemie, International Edition (2014), 53(6), 1602-1606). Another example of a suitable scaffold is l,3,5-tris(bromomethyl)benzene (TBMB) (WO2016 / 067035 Al).

[0078] Other example scaffolds include TATB (l,T,l"-(l,3,5-triazinane-l,3,5-triyl)tris(2- bromoethan-l-one)), TCTZ (1,3,5-trichloromethyltriazine) (Kale, S.S., Villequey, C., Kong, XD. et al. Cyclization of peptides with two chemical bridges affords large scaffold diversities. Nature Chem 10, 715-723 (2018) https: / / doi.ors / lO 1038 s4155~'-018-0042-7), TTZ (2,4,6-Tris(Halomethyl)-l,3,5-Triazine), TBAB (2,4,6-Trioxo-l,3,5-triazinane-l,3,5- triyl)-trisethane-2,l-diyltriacrylate), TSTA (l,3,5-tri(ethenesulfonyl)-l,3,5-triazinane) (W02020 / 084305), TCAZ (l,4,7-tris(chloroacetyl)octahydro-l,4,7-triazonane), TCAN (N,N',N"-(nitrilotris(ethane-2,l-diyl))tris(2-chloroacetamide)), TCCU (1,1',1"-(1H,4H- 3a,6a-(methanoiminomethano)pyrrolo[3,4-c]pyrrole-2,5,8(3H,6H)-triyl)tris(2-chloroethan- 1-one)) (WO2018 / 197893).

[0079] Example Computer Implemented Method

[0080] The methods described herein may be executed by a computer, or other data processing system. The methods may be embodied by a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of the preceding claims. The methods may also be embodied by a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the methods. The methods may also be embodied by a data processing system comprising means for carrying out the method.

[0081] Fig. l is a flow chart illustrating a first example computer implemented method according to the disclosure, for identifying amino acid sequences likely to form scaffold-bound peptide ligands that bind to a target.

[0082] In a first step, SI a motif sequence, which is a first amino acid sequence describing a binding motif predetermined to interact with one or more predefined targets, is received. In a second step, S2, a peptide sequence, which is a second amino acid sequence describing a peptide is generated. The generated peptide sequence satisfies one or more predefined sequence constraints such that the peptide has a propensity to form a scaffold-bound peptide ligand when bound to one of one or more predefined scaffold molecules.

[0083] In a third step, S3, at least one predicted interaction characteristic of the peptide is determined. The interaction characteristic relates to an interaction between the peptide and one or more predefined targets.

[0084] In a fourth step, S4, a favourability of the at least one predicted interaction characteristic of the peptide is determined.

[0085] In a fifth step, S5, output data is provided by the method based on the favourability determined in step S4. The output data identifies one or more amino acid sequences likely to form a scaffold-bound peptide ligand with favourable characteristics.

[0086] As shown, in Fig. 1, the method may be configured to repeat the second to fourth, steps S2, S3 and S4 for a plurality of peptide sequences (e.g. a large number of diversified peptide sequences). In some examples, step S5 may be performed a plurality of times, e.g. with each iteration or only iterations for which the determined favourability is relatively high (e.g. above a predefined threshold). Accordingly, output data may be provided for a plurality of amino acid sequences. In some examples, step S4 may be performed once at the end of a predefined number of iterations.

[0087] Interaction Characteristics

[0088] Favourability of the at least one predicted interaction characteristic may be determined based on a comparison between the at least one predicted interaction characteristic and one or more predefined corresponding interaction requirements. The predefined corresponding interaction requirements may generally be desired (or optimal) parameters for the interaction characteristics.

[0089] In a preferred example, favourability may be determined by calculating a favourability score. The calculation may be made based on an objective function, such as a loss function. The objective function may comprise terms that mathematically compare the at least one predicted interaction characteristic and the one or more predefined corresponding interaction requirements. In a preferred example, a combined favourability for a plurality of predicted interaction characteristic is determined using a single objective function that calculates a single favourability score for each peptide, based on a plurality of predicted interaction characteristics.

[0090] In other examples, comparison with a predefined interaction requirement may be performed separately for each interaction characteristic and these may optionally be combined to provide a single favourability score.

[0091] In a preferred example, a predicted interaction characteristic may relate to the likelihood of the peptide binding to a target. For example, one predicted interaction characteristic may be a binding score, based on the likelihood of the peptide binding to a target. Favourability may be determined based on how the binding score compares to a desired (or optimal) binding score.

[0092] In a preferred example, a predicted interaction characteristic may relate to the stability of the peptide. For example, one predicted interaction characteristic may be a folding score based on the likelihood of the peptide being well-folded. Favourability may be determined based on how the folding score compares to a desired (or optimal) folding score.

[0093] Another example interaction characteristic may relate to the parts of the target at which the peptide is predicted to bind to the target. For example, a predicted interaction characteristic may comprise a score for specific residues of the target (or all residues), based on the likelihood of each of these residue binding to the peptide. Favourability may be determined based on the scores for these specific residues (or specific residues out of all the residues). The specific residues may correspond to residues located at a part of the target preferred for binding to the peptide, e.g. the part of the target predetermined to bind to the binding motif, with high favourability corresponding to a high likelihood of binding at these residues. It may be advantageous for the peptide to bind at a specific location on the target. Another example interaction characteristic may relate to the likelihood of one or more specific residues of the peptide binding to the target. For example, a predicted interaction characteristic may comprise a score for specific residues of the peptide (or all residues), based on the likelihood of each of these residues binding to the target. Favourability may be determined based on the scores for these specific residues (or specific residues out of all the residues). The specific residues may correspond to those of the predetermined binding motif. Alternatively, or additionally, the specific residues may correspond to residues likely not to bind to the target, with high favourability corresponding to a low likelihood of binding at these residues, for example but not limited to Lysine residues. Such residues, for example but not limited to Lysine residues, may provide cross-linking within the peptide, so it may be advantageous for these residues not to interact with the target in order to provide cross-linking.

[0094] At least some predicted interaction characteristics, including those described above, may be determined based on a predicted three-dimensional structure of the peptide and / or the target. The three-dimensional structure of the peptide may be determined by a machine learning algorithm, based on input data comprising the generated amino acid sequence. The three-dimensional structure of the target may be determined by a machine learning algorithm, based on input data comprising the amino acid sequence of a target protein. In a preferred example, the machine learning algorithm, in either case, may be the AlphaFold™ algorithm. In an example method, input data may be provided comprising data describing a predefined target. This data may comprise an amino acid sequence of the target and / or the three-dimensional structure of the target.

[0095] In a preferred example in which predicted interaction characteristics are determined using the AlphaFold™ algorithm, the predicted interaction characteristics may comprise a Predicted Aligned Error (PAE) score, such as an i_PAE score. A PAE score is an AlphaFold™ metric to assess the confidence in the relative positions of different parts of the three-dimensional models e.g., different domains. A specific example method uses the metric as a measure of confidence in the relative positioning of the target and peptide by extracting the interface PAE (i_PAE) confidence score, which relates to the relative positioning of the peptide to the target at their binding interface. For example, if a peptide has a central helix that interacts with the target while the flanking regions do not, then the i_PAE score would correspond to the confidence in the relative positioning of the central helix and the region of the target close to it, the non-interacting parts are not including in the final i_PAE score. It may be advantageous to maximize this score as it gives preference to an accurate model of the binding interface over the flanking non-binding regions. The i_PAE score relates to the likelihood of the peptide binding to the target.

[0096] In an example in which predicted interaction characteristics are determined using the AlphaFold™ algorithm, the predicted interaction characteristics may comprise a pLDDT score. The pLDDT may relate to the likelihood of the peptide forming a well-folded structure. The pLDDT score is an AlphaFold™ per-residue estimate confidence score (0- 100) corresponding to the IDDT-Ca metric. It is a superimposition-free measure to assess the local quality of the model based on the agreement between the model and the available PDB data. 1DDT, Mariani T et al, 2013, (https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC3799472 / ) measures how well the environment in a reference structure is reproduced in a protein model. AlphaFold’s pLDDT uses the alpha-carbon atoms 1DDT score for the pLDDT score.

[0097] In an example in which predicted interaction characteristics are determined using the AlphaFold™ algorithm, the predicted interaction characteristics may comprise a distogram. A distogram is a matrix of distances between all the CP atoms of the target and binder, and represents a contact map of the model. The distogram may be used to determine predicted interaction characteristics relating to where the peptide binds or does not bind to the target. For example, a relatively small distance between a target residue and a peptide residue may correspond with a relatively high likelihood of binding, and vice versa.

[0098] In a preferred example, another of the interaction characteristics may relate to the relative positions of the motif sequence of the peptide, when bound to the target, and the position of the motif sequence within a peptide or protein from which the motif sequence is isolated, when bound to the target. In an example, the characteristic may comprise a value calculated based on the root mean squared distance between the respective motif sequences relative to the target structure. For the peptide this may be estimated by modelling using AlphaFold™. For the predetermined structure, this may be determined based on a crystallographic structure. In a preferred example, the binding score may relate to a predicted binding affinity between the peptide and the target. In a preferred example, the binding affinity may be determined using the Prodigy algorithm (htt s : / / github . com / haddocking / rodi y ), which predicts binding affinity between protein-protein complexes, in this case the peptide and the target protein, based on the input structures. Optionally, the structure of the generated peptide may be relaxed by the OpenMM algorithm (https : / / github . com / openmm / openmm) before being provided to the Prodigy algorithm, to mitigate limitations of the structure predicted using AlphaFold™.

[0099] In an example method, input data may be provided comprising data describing the one or more interaction requirements.

[0100] Motif sequence

[0101] The motif sequence may be a sequence corresponding to predetermined binding motif, or a portion of a predefined binding motif, for a predetermined target. The binding motif may comprise a short amino acid sequence forming part of a longer peptide or protein. The binding motif may substantially comprise an amino acid sequence that interacts with the target, e.g. binds thereto. The motif sequence may be isolated from the longer peptide or protein by removing the amino acids either side of the motif sequence from the longer peptide or protein. In some examples, the binding motif may be discontinuous, i.e. the longer peptide or protein sequence may comprise portions which interact with the target separated by portions which do not interact with the target. In such cases, the motif sequence may correspond to a continuous portion of the binding motif. The motif sequence may be characterised by its sequence, and optionally its three-dimensional structure within the longer peptide or protein, e.g. including the backbone shape.

[0102] The motif sequence may be predetermined based on experimentation or computational methods for predicting protein / peptide target interactions. The motif sequence may be derived from a cyclic or non-cyclic peptide, which may or may not be bound to a molecular scaffold.

[0103] Sequence Constraints The one or more sequence constraints may comprise one or more constraints that increase the likelihood that the peptide will bind to one of one or more predefined scaffolds. For example, a sequence constraint may require that the sequence comprises at least one cysteine residue. A cysteine residue may bind with a scaffold molecule, for example.

[0104] Alternatively, or additionally, the sequence constraints may comprise a constraint on a total length of the sequence. For example, the total length may be constrained to be in a range of from 6 to 30 amino acid residues, e.g. 13 residues long.

[0105] In a preferred example, the one or more sequence constraints may comprise one or more constraints that increase a likelihood that the peptide will bind to one of the one or more predefined scaffolds at at least two locations. For example, a sequence constraint may require that the sequence comprises at least two cysteine residues. In some preferred examples, a sequence constraint may require that the sequence comprises three cysteine residues, for example to form scaffold-bound peptide ligands having a bicyclic structure.

[0106] In an example method, the one or more sequence constraints may comprise a constraint on the amino acid residues at one or more terminal positions in the sequence. For example, the one or more constraints may comprise a constraint that the amino acid residue at one or more terminal positions in the sequence is an alanine residue. In a preferred example, alanine residues may be fixed at both terminal positions.

[0107] The one or more sequence constraints may be based on the scaffold molecule or molecules to which the generated peptides are configured to bind. For example, the scaffold molecule may determine the specific residues required and / or their specific positions. For example, the scaffold molecule may determine the required number of cystine residues. In some cases, the scaffold molecule may determine the required distance between cystine residues.

[0108] Accordingly, in some examples, the one or more sequence constraints may comprise a constraint on the relative positions of at least two cysteine residues. For example, the one or more sequence constraints may comprise a constraint that at least one pair of successive cysteine residues are separated by no more than a first predefined number of amino acid residues and / or no fewer than a second predefined number of amino acid residues. When the sequence comprises three cysteine residues, including two pairs of successive cysteine residues, both pairs may satisfy at least one of the constraints above. Accordingly, in some examples, predetermined positions of one or more cysteine residues may satisfy one or more of the above constraints.

[0109] Sequence Generation

[0110] An example process for the step S2 of generating a peptide sequence is shown in Fig. 3. A first step of this process, Step S6 may comprise extending the motif sequence. In a preferred example, extending the motif sequence may comprise adding at least one amino acid residue either side of the motif sequence, subject to one or more relevant sequence constraints. For example, the sequence constrains relevant at this stage may relate to the total length of the peptide sequence.

[0111] In a preferred example, the motif sequence may be first extended to generate a peptide backbone comprising the motif sequence. The peptide backbone defines the length of the peptide sequence and the position of the motif sequence within the peptide sequence. The backbone may define the number of -N-C-C- units either side of the motif sequence. At this stage, the side-chains of amino acids on either side of the motif sequence are not important and may be modified in further steps. Accordingly, for example, these may be all be assigned uniform “dummy” residue or the positions may not be assigned any specific residue.

[0112] In a preferred example, a number of different peptide backbones may be generated having different lengths and / or different positions of the motif sequence, i.e. with different numbers of amino acids either side of the motif sequence. The number of amino acids added either side of the motif sequence may be determined randomly, subject to the one or more sequence constraints to generate the peptide backbones.

[0113] In Step S7, one or more first residues may be assigned to amino acid positions forming the backbone, other than the residues forming the motif sequence, to satisfy the one or more predefined sequence constraints, to generate a partial peptide sequence. At this stage, the relevant sequence constraints may relate to the presence, positions and / or relative positions of particular residues in the peptide sequence. For example, two or more cysteine residues may be assigned, which are configured to bind to one or more predefined scaffolds. Additionally, terminal alanine residues may be assigned.

[0114] First residues that do not have a requirement to be located at a specific position in the peptide sequence (e.g. at a terminal position) may be assigned randomly, subject to the sequence constraints. For example, cysteine residues may be assigned at random positions subject to any constraints on their relative positions. In some examples, a number of different partial peptide sequences may be generated with first residues assigned to different (e.g. randomly assigned) positions in the backbone, for each different backbone.

[0115] In Step S8, second residues may be assigned to the remaining amino acid positions of the backbone to generate a complete peptide sequence. The second residues may be assigned such that the peptide sequence is configured to fold into a predefined three-dimensional shape. Accordingly, a predefined three-dimensional shape may be assigned to each backbone or each partial sequence. More than one different peptide sequence may be configured to fold into the predefined three-dimensional shape (e.g. with varying accuracy or confidence). Accordingly, a plurality of different complete peptide sequence may be generated for each partial peptide sequence.

[0116] The predefined three-dimensional shape may comprise the predefined three-dimensional shape of the motif sequence. The predefined three-dimensional shape of the motif sequence may be the same shape of the motif sequence has within the larger peptide or protein from which it is isolated. The shape of the remaining portions of the peptide backbone may be determined such that the three-dimensional shape of the backbone is compatible with the binding motif binding to the one or more predefined targets. For example, the backbone may be required to fit within a binding pocket of the one or more predefined targets and / or the motif sequence may be required to be positioned relative to the one or more predefined targets where it is predetermined to bind to the one or more predefined targets.

[0117] In a specific example, suitable three-dimensional backbone shapes may be generated by computational methods. For example, machine learning algorithms configured to generate three-dimensional backbone shapes may be used, such as RFDiffusion, which is a diffusion-based backbone generation algorithm (https: / / github.com / RosettaCommons / RFdiffusion).

[0118] In a preferred example, prior to assigning second residues, a partial peptide sequence with a three-dimensional backbone structure assigned to it has been defined. In a preferred example, the second residues may be determined based on an inverse folding algorithm. The inverse folding algorithm may be a machine learning algorithm configured to predict an amino acid sequence configured to fold into a predetermined three-dimensional structure. The inverse folding algorithm may be configured to maintain the residues at positions defined by the partial peptide sequence. In a specific example, the inverse folding algorithm may be ProteinMPNN (htt s : / / github . com / dauparas / ProteinMPNN) .

[0119] As described above, the output of Step S8 and Step S2 may be at least one peptide sequence that satisfies the one or more sequence constraints and folds into a three- dimensional shape suitable to permit the binding motif to bind to the one or more predefined targets. In a preferred example, a large number of diversified peptide sequences (e.g. more than 10,000) may be generated in this way, e.g. with different lengths, motif sequence positions, and / or assigned first and second residues. This increases the likelihood of identifying one or more amino acid sequences predicted to form a scaffold-bound peptide ligand that binds to a target with a high confidence.

[0120] Structural characteristics

[0121] In some examples, in addition to determining one or more interaction characteristics, the method may further comprise determining one or more additional characteristics of the generated peptide. Further, favourability of any additional characteristics may be determined and the output data based on the favourability of the additional characteristics.

[0122] The favourability of additional characteristics may be performed in the same way as described above in relation to the interaction characteristics. For example, favourability may be determined based on a comparison between the additional characteristic and one or more predefined corresponding additional requirements, as described above; the predefined corresponding additional requirements may generally be desired (or optimal) parameters for the additional characteristics, as described above; and, favourability may be determined by calculating a favourability score, as described above.

[0123] The determination of additional characteristics may be performed as part of step S2 described above, for example. The determination of favourability of additional characteristics may be performed as part of step S3 described above, for example.

[0124] One example type of additional characteristics may be structural characteristics. The one or more structural characteristics may relate to a likelihood that the peptide will bind with one of one or more predefined molecular scaffolds. For example, one or more structural characteristics may relate to a likelihood that the peptide described by the sequence will bind to a scaffold at two locations, or alternatively any number of required locations greater than two.

[0125] One example structural characteristic may be a distance between two cysteine residues, in the predicted three-dimensional structure of the peptide. The distance may be between respective sulphur atoms of the cysteine residues. Favourability may be determined based on how the distance compares to a desired (or optimal) distance or distance range. For example, a desired (or optimal) distance range may be a range from 2 to 20 Angstrom. For particular scaffold molecules, such as TATA, a narrower range may be more desirable, such as a range of from 2.8 to 12.5 Angstrom. For particular scaffold molecules, such as TATB, a yet narrower range may be more desirable, such as a range of from 3.2 to 10 Angstrom. For particular scaffold molecules, such as TTZ, a yet narrower range may be more desirable, such as a range of from 4.5 to 8 Angstrom.

[0126] Another example structural characteristic may be an angle formed between a first cysteine residue and two further cysteine residues, in the predicted three-dimensional structure of the peptide. The angles may be between respective sulphur atoms of the cysteine residues. Favourability may be determined based on how the angle compares to a desired (or optimal) angle or range of angles. For example, a desired (or optimal) range of angles may be a range from 10 to 160 degrees. For particular scaffold molecules, such as TATA, a narrower range may be more desirable, such as a range of from 14 to 150 degrees. For particular scaffold molecules, such as TATB, a yet narrower range may be more desirable, such as a range of from 20 to 115 degrees. For particular scaffold molecules, such as TTZ, a yet narrower range may be more desirable, such as a range of from 25 to 85 degrees.

[0127] Another example structural characteristic may be a dihedral angle formed between pairs of cysteines. The dihedral angle may be defined as the angle between two planes defined by the quadruplet of positions Cal, SI, S2, Ca2; where Cal is the alpha carbon atom of the first cysteine residue, SI is the sulphur atom of the first cysteine residue, Ca2 is the alpha carbon atom of the second cysteine residue, S2 is the sulphur atom of the second cysteine residue. The dihedral angle may be calculated around the axis connecting positions SI and S2, as the angle between first and second planes defined by the positions Cal, SI, S2 and SI, S2, Ca2 respectively. Favourability may be determined based on how the angle compares to a desired (or optimal) angle or range of angles. For example, a desired (or optimal) range of angles may be a range of from -175 to 175 degrees. For particular scaffold molecules, such as TATB and / or TTZ, a narrower range may be more desirable, such as a range of from -170 to 170 degrees.

[0128] The predicted three-dimensional structure may be generated by a machine learning algorithm, based on input data comprising the generated amino acid sequence. The machine learning algorithm may be the AlphaFold™ algorithm, for example.

[0129] Functional characteristics

[0130] Another example type of additional characteristics may be functional characteristics. Examples of functional characteristics may include: a membrane permeability of the peptide and a blood-brain barrier permeability of the peptide. Corresponding functional requirements may include desired parameters for these characteristics. Favourability may be determined based on how each of these parameters compare to a desired (or optimal) parameter.

[0131] Functional characteristics may be determined based on the generated peptide sequence and / or a predicted three-dimensional structure of the peptide. This may be determined using known algorithms. For example, solubility, hydrophobicity, pKa, total charge, oxidation sensitivity, etc. (10.26434 / chemrxiv-2023-cwr53). Example

[0132] In a specific example, SSTR4 was selected as the target. A binding motif was isolated from SST-14 cyclic peptide (PDBID: 7xms) known to bind to SSTR4, with motif sequence FFWK. A diffusion-based backbone generation method, namely RFDiffusion, was used to generate 10,000 backbones by extending the FFWK motif by 4-7 residues on each side. At each iteration, the algorithm selected two numbers between 4 and 7 randomly and extended the N and C-termini by these numbers.

[0133] ProteinMPNN was used to generate sequences that fold into the generated backbones after randomly assigning cysteine residues to three positions, for each backbone with the condition that the distances between the cysteines should be greater than / equal to 2 residues. The positions assigned to cysteine and the motif are fixed while the rest of the sequence is generated using ProteinMPNN. To increase diversity, 8 sequences were generated for each backbone. This provided 10,000 x 8 = 80,000 sequences.

[0134] All of the 80,000 sequences generated were modelled using ColabFold - an open source version of AlphaFold2, to determine interface confidence (i_PAE). Additionally, the structures of the generated peptide sequences were relaxed using openMM. These relaxed structures were fed to the program Prodigy for an estimation of the binding affinity between the peptide sequence and the target. Finally, the generated peptide sequences were filtered using a set of selection criteria as follows:

[0135] 1. The interface confidence given by ColabFold is > 0.8

[0136] 2. The binding affinity calculated by Prodigy is < -8

[0137] 3. The RMSD between the motif position in the input structure (crystal structure), namely of SST-14 cyclic peptide, and the modelled structure is < 2

[0138] Alternative selection criteria may be applied in other examples, depending on requirements.

[0139] Fourteen bicyclic peptide ligands having sequences generated according to the above example method were tested, together with somatostatin- 14 as a control. Each of the peptides was bound to a TATB scaffold molecule and manufactured according to methods well known in the art.

[0140] The bicyclic peptide ligands were tested using CH0-K1 cells expressing human SSTR4 to assess both the agonist and antagonist response as shown in Table 1. Ligand biological efficacy was measured in CH0-K1 cells engineered to express human SSTR4 (SS4-R / SSR4) and using an adenyl yl cyclase activity, cyclic AMP (cAMP) fluorescence reporter assay as described by Engstom M. et al, 2005 (DOI: 10.1124 / jpet.104.075531, PubMed: 15333679).

[0141] Agonist biological activity is measured after a 10-minute incubation with increasing concentrations of ligand and measuring the induced activity of adenylyl cyclase through the concentration of fluorophore-labelled cAMP product, in comparison to a no ligand added negative control. The EC50 of a ligand is the concentration at which 50% agonist activity (half maximal response) is measured. The results are expressed as a percentage of a control agonist response (lOnM somatostatin- 14).

[0142] Antagonist activity was measured using the same cells and assay to assess the effect of the ligand on the stimulation response to 1 nM somatostatin- 14 over a range of ligand concentrations. Data was expressed as the percentage inhibition of the control agonist (somatostatin) response. IC50 values were calculated where possible from these values.

[0143] None of the tested molecules showed any measurable antagonism while agonist activity was observed for all with the EC50 values ranging from micromolar to sub-nanomolar. Accordingly, this verifies the method generates bicyclic peptides that bind to the selected target as intended.

[0144] Table 1

[0145] Ligands 01, 02, and 04 were further assayed in CHO-Ki cells expressing rat SSTR4 to assess functional agonist activity using an adenyl cyclase activity, cyclic AMP (cAMP) reporter assay, as shown in Table 2below. Fig. 4 shows the functional activity of ligands in CHO-Ki cells expressing rat SSTR4 respectively for Ligands 01, 02 and 04. This data demonstrates the method can generate ligands with cross species activity.

[0146] Table 2

[0147] Ligands 01-04 were further assayed in radioligand binding assays for affinity at the different hSSTRl-5 receptor subtypes and IC50 values were obtained, as shown in Table 3. Fig. 5 shows the percentage inhibition of binding by ligands in CHO-Ki cells expressing human SSTR subtypes respectively for Ligands 01, 02, and 04. This data demonstrates the method can generate ligands that show selectivity for hSSTR4 over other family subtypes.

[0148] Table 3

[0149] Example uses In a specific example use, methods described above may be used in a method of identifying a therapeutic agent or diagnostic agent targeting a target. The therapeutic agent or diagnostic agent may be a scaffold-bound peptide that binds to the target and the computer implemented methods described above may be used to identify a peptide forming the scaffold-bound peptide.

[0150] In another specific example use, methods described above may be used in a method of making a peptide or scaffold-bound peptide ligand. A computer implemented method may be used to identify a peptide forming a scaffold-bound peptide ligand, prior to making the peptide or the scaffold-bound peptide ligand.

[0151] It should be understood that variations of the above described examples are possible without departing from the spirit or scope of the invention.

Claims

1. CLAIMS1. A computer implemented method for identifying amino acid sequences predicted to form a scaffold-bound peptide ligand that binds to a target, the method comprising: receiving as input a motif sequence, which is a first amino acid sequence describing at least part of a binding motif predetermined to interact with one or more predefined targets; generating a peptide sequence, which is a second amino acid sequence comprising the motif sequence, and describing a peptide, the second amino acid sequence satisfying one or more predefined sequence constraints such that the peptide binds to one of one or more predefined scaffolds in at least two locations, to form a scaffold-bound peptide ligand; determining at least one predicted interaction characteristic relating to an interaction between the peptide and the one or more predefined targets; determining a favourability of the at least one predicted interaction characteristic; and providing output data identifying one or more amino acid sequences predicted to form a scaffold-bound peptide ligand that binds to a target, based on the determined favourability.

2. The method of claim 1, wherein the step of generating the peptide sequence comprises extending the motif sequence by adding at least one amino acid either side of the motif sequence.

3. The method of claim 2, wherein the number of amino acids added either side of the motif sequence is determined randomly, subject to one or more of the sequence constraints.

4. The method of claim 2 or 3, wherein the motif sequence is first extended to form a peptide backbone.

5. The method of claim 4, further comprising assigning one or more first residues to amino acid positions in the backbone, other than the residues forming the motif sequence,to satisfy one or more of the predefined sequence constraints, to form a partial peptide sequence.

6. The method of claim 5, wherein at least two the one or more first residues are configured to bind to one of one or more predefined scaffolds.

7. The method of claim 5 or 6 wherein at least one of the one or more first residues are assigned to random amino acids in the backbone, subject to the one or more predefined sequence constraints.

8. The method of claim 7, wherein the step of generating a peptide sequence comprises further assigning second residues to the remaining amino acid positions in the backbone to form a complete peptide sequence, such that the peptide sequence is configured to fold into a predefined the three-dimensional shape.

9. The method of claim 8, wherein the peptide backbone is assigned a predefined three-dimensional shape comprising a predefined three-dimensional shape of the motif sequence and which is suitable for the binding motif binding to the one or more predefined targets.

10. The method of claim 8 or 9, wherein the second residues are determined by an inverse folding algorithm.

11. The method of claim 10, wherein the inverse folding algorithm is a machine learning algorithm configured to predict an amino acid sequence configured to fold into a predetermined three-dimensional structure.

12. The method of any preceding claim, wherein the step of generating an amino acid sequence, and subsequent steps, are repeated for a plurality of peptide sequences.

13. The method of claim 12, wherein the plurality of peptide sequences are generated based on different one-dimensional backbones, different three-dimensional backbone shapes, different first residues assigned thereto to form different partial peptide sequences, and / or different second residues assigned thereto to form different complete peptide sequences.

14. The method of any preceding claim, wherein the output data comprises the peptide sequence.

15. The method of any preceding claim, wherein the one or more sequence constraints comprises one or more constraints that increase a likelihood that the peptide will bind to one of one or more predefined scaffolds.

16. The method of any preceding claim, wherein the one or more sequence constraints comprises one or more constraints that increase a likelihood that the peptide will form at least one loop structure when bound with one of the one or more predefined scaffolds.

17. The method of any preceding claim, wherein the one or more sequence constraints comprises a constraint on a total length of the peptide sequence.

18. The method of claim 17, wherein the one or more sequence constraints comprises a constraint that the peptide sequence consists of from 6 to 30 amino acid residues.

19. The method of any preceding claim, wherein the one or more sequence constraints comprises a constraint that the peptide sequence comprises at least two cysteine residues.

20. The method of any preceding claim, wherein the one or more sequence constraints comprises a constraint that the peptide sequence comprises three cysteine residues.

21. The method of any preceding claim, wherein the one or more sequence constraints comprises a constraint on the relative positions of at least two cysteine residues.

22. The method of any preceding claim, wherein the one or more sequence constraints comprises a constraint that at least one pair of successive cysteine residues are separated by no more than a first predefined number of amino acid residues.

23. The method of any preceding claim, wherein the one or more sequence constraints comprises a constraint that at least one pair of successive cysteine residues are separated by no fewer than a second predefined number of amino acid residues.

24. The method of claim 20, wherein the sequence comprises three cysteine residues, including two pairs of successive cysteine residues, and both pairs satisfy at least one of the constraints defined in claim 22 and claim 23.

25. The method of any preceding claim, wherein the one or more constraints comprises a constraint on the amino acid residue at one or more terminal positions in the sequence.

26. The method of any preceding claim, wherein the one or more constraints comprises a constraint that the amino acid residue at one or more terminal positions in the sequence is an alanine residue.

27. The method of any preceding claim, wherein the one or more interaction characteristics comprises a binding score based on the likelihood of the peptide binding to the target.

28. The method of any preceding claim, wherein the one or more interaction characteristics comprises a folding score based on the likelihood of the peptide being well- folded.

29. The method of any preceding claim, wherein the one or more interaction characteristics comprises a score for one or more residues of the target, based on the likelihood of each of these residue binding to the peptide.

30. The method of any preceding claim, wherein the one or more interaction characteristics comprises a score for one or more residues of the peptide, based on the likelihood of each of these residues binding to the target.

31. The method of claim 30, wherein the one or more residues includes residues of the motif sequence.

32. The method of any preceding claim, wherein the one or more interaction characteristics comprises a score for a predicted position of the binding motif relative to the target compared to a position at which the binding motif is predetermined to bind to the target.

33. The method of any preceding claim, wherein the one or more interaction characteristics is determined based on a predicted three-dimensional structure of the peptide and / or the target.

34. The method of any preceding claim, further comprising a step of determining one or more structural characteristics of the peptide and a step of determining a favourability of the one or more structural characteristics.

35. The method of claim 34, wherein the one or more structural characteristics relate to a likelihood that the peptide will bind with one of one or more predefined molecular scaffolds.

36. The method of claim 34 or 35, wherein the one or more structural characteristics comprises a distance between two cysteine residues, in the predicted three-dimensional structure of the peptide.

37. The method of any one of claims 34 to 36 wherein the one or more structural characteristics comprises a distance between respective sulphur atoms of two cysteine residues in the predicted three-dimensional structure of the peptide.

38. The method of any one of claims 34 to 37, wherein the one or more structural characteristics comprises an angle formed between a first cysteine residue and two further cysteine residues, in the predicted three-dimensional structure of the peptide.

39. The method of any one of claims 34 to 38, wherein the one or more structural characteristics comprises a dihedral angle formed between a first cysteine residue and a second cystine residue, in the predicted three-dimensional structure of the peptide.

40. The method of any one of claims 34 to 39, wherein the predicted three-dimensional structure of the peptide is generated by a machine learning algorithm based on input data comprising the generated amino acid sequence.

41. The method of any preceding claim, further comprising a step of determining one or more functional characteristics of the peptide and a step of determining a favourability of the one or more functional characteristics.

42. The method of claim 41, wherein the one or more functional characteristics comprises one more of: membrane permeability of peptide and blood-brain barrier permeability of the peptide.

43. The method of any preceding claim, where the favourability is determined based on a comparison between the determined characteristics and predefined corresponding desired or optimal parameters.

44. The method of claim 43, wherein the favourability is determined by calculating a favourability score.

45. The method of claim 44, wherein the favourability score is calculated using an objective function comprising terms corresponding to each of the determined characteristics.

46. The method of any preceding claim, wherein the scaffold-bound peptide ligand is bicyclic, comprising two loop sequences, and is covalently bound to the scaffold at three locations.

47. The method of any preceding claim, wherein the one or more scaffolds comprise one or more of TATA, TATB, TCTZ, TBMB, TTZ, TBAB, TSTA, TCAZ, TCAN, and TCCU.

48. The method of any preceding claim, wherein the peptide forms a peptide ligand, and the peptide in the peptide ligand comprises at least three cysteine residues, separated by at least two loop sequences, and the scaffold forms covalent bonds with the cysteine residues of the peptide such that at least two peptide loops are formed on the scaffold.

49. The method of any preceding claim, wherein the target is a protein, a protein on a cell, a tumour antigen, a viral antigen, bacterial antigen.

50. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of the preceding claims.

51. A data processing system comprising means for carrying out the method of any one of claims 1 to 49.

52. A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of any one of claims 1 to 49.

53. A method of identifying a therapeutic agent or diagnostic agent targeting a target, the therapeutic comprising a scaffold-bound peptide ligand, the method comprising using the computer implemented method of any one of claims 1 to 49 to identify a peptide forming the scaffold-bound peptide ligand.

54. A method of making a peptide or scaffold-bound peptide ligand, comprising using the computer implemented method of any one of claims 1 to 49 to identify a peptide forming the scaffold-bound peptide ligand or the scaffold-bound peptide ligand.

55. A scaffold-bound peptide ligand made by the method of claim 54.

Citation Information

Patent Citations

  • Method for selecting a candidate drug compound

    WO2004077062A2

  • Binding compounds, immunogenic compounds and peptidomimetics

    WO2006078161A1

  • Methods and compositions

    WO2009098450A2

  • Bicyclic peptide ligands specific for mt1-mmp

    WO2016067035A1

  • Bicyclic peptide ligands and uses thereof

    WO2018197893A1