A method for predicting interaction properties between proteins and peptides of interest

A ML-based method for predicting protein-peptide interactions addresses throughput and stability challenges, enabling efficient screening and reducing material consumption, thus improving therapeutic candidate development.

WO2026057688A1PCT designated stage Publication Date: 2026-03-19BIOCOPY AG
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Current technologies for predicting protein-peptide interactions, particularly in the context of peptide-HLA complexes, face limitations in throughput, material consumption, stability, and data evaluation, making it challenging to efficiently screen for off-target reactivities in therapeutic applications.

Method used

A computer-implemented method using machine learning (ML)-based models to predict protein-peptide interactions by training on sequence and structural data, combined with experimental binding data, enabling high-throughput prediction of binding properties and off-target risks.

Benefits of technology

Enables precise and efficient prediction of protein-peptide interactions, reducing material consumption and time, while providing comprehensive coverage of the human immunopeptidome, thereby enhancing the development of therapeutic candidates by identifying optimal and off-target interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025075872_19032026_PF_FP_ABST
    Figure EP2025075872_19032026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates in one aspect to a computer-implemented method for predicting the binding of at least one protein of interest, preferably at least one receptor, antibody, antibody fragment or equivalents thereof, or DARPin, to a peptide of interest, comprising a. providing sequence and / or structural data of at least one reference peptide, optionally providing sequence and / or structural data of at least one HLA molecule and / or of at least one HLA-peptide complex (pHLA), and / or the at least one reference peptide presented within a pHLA complex, b. providing binding data, preferably comprising the dissociation constant (KD), of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide complex (pHLA) with at least one protein of interest, c. training a machine learning (ML)-based model or artificial intelligence (AI) based on the data provided in a. and the binding data provided in b., d. providing sequence and / or structural data of at least one peptide of interest, e. employing the machine learning (ML)- based model or artificial intelligence (AI) trained in c. to predict the binding or interaction properties of the at least one peptide of interest, preferably when presented within a pHLA complex, to the at least one protein of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A method for predicting interaction properties between proteins and peptides of interest

[0002] DESCRIPTION

[0003] The invention lies in the field of biology, biotechnology and biochemistry, more particular in the field of molecular biology.

[0004] The invention relates in one aspect to a computer-implemented method for predicting the binding of at least one protein of interest, preferably at least one receptor, antibody, antibody fragment or equivalents thereof, or DARPin, to a peptide of interest, comprising a. providing sequence and / or structural data of at least one reference peptide, optionally providing sequence and / or structural data of at least one HLA molecule and / or of at least one HLA-peptide complex (pHLA), and / or the at least one reference peptide presented within a pHLA complex, b. providing binding data, preferably comprising the dissociation constant (KD), of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide complex (pHLA) with at least one protein of interest, c. training a machine learning (ML)-based model or artificial intelligence (Al) based on the data provided in a. and the binding data provided in b., d. providing sequence and / or structural data of at least one peptide of interest, e. employing the machine learning (ML)- based model or artificial intelligence (Al) trained in c. to predict the binding or interaction properties of the at least one peptide of interest, preferably when presented within a pHLA complex, to the at least one protein of interest.

[0005] BACKGROUND OF THE INVENTION

[0006] Recent developments in adoptive cell therapy (ACT) as well as in the field of bispecific T-cell engagers (BiTE) have shown that peptide-human leukocyte antigen (pHLA)-targeting therapies represent valid strategies to treat cancer, with a first bispecific T cell receptor (TCR)-based therapy recently approved in uveal melanoma.

[0007] Peptide-H LA class I complexes are trimeric complexes that consist of a polymorphic heavy chain (alpha-chain), the light chain beta-2 microglobulin ( / ?2m) and a peptide ligand, typically between 8 and 10 amino acids long and derived from cellular proteins by degradation. T cells can recognize specific peptide-HLA complexes with the TCR and initiate immune responses. Similarly, high- affinity TCR-based bispecific molecules or TCR-like antibodies bind to the respective pHLA target and trigger immune cell activity through a second binder, e.g. an anti-CD3 antibody that recruits T cells to the specific pHLA-expressing tumor cell. While such therapeutic approaches offer great potential, targeting pHLA molecules also bears the risk of off-target reactivities as experienced in some clinical trials.

[0008] Therefore, there is a clear need for comprehensive high-throughput off-target screenings for pH LA-targeting therapeutics, that can also be applied early in development to discriminate off- target profiles of different candidates. A great deal of energy has been invested in the development of soluble forms of the pH LA-reactive moiety of pHLA specific therapeutics, such as soluble T-cell receptors or TCR-like antibodies (TCRm), in order to enable interaction measurements at the molecular level.

[0009] Directly determining the kinetic parameters of TCR-pHLA interactions can be a valuable addition to complex and laborious cellular screening approaches that usually require less direct readouts like cytokine expression detection. Label-free measurement methods, such as surface plasmon resonance (SPR), biolayer interference (B LI) or reflectometric interferometry spectroscopy (RlfS) offer a convenient way to measure the kinetic parameters (kass, kdiss and KD). Grating-coupled interferometry (GCI) even demonstrated the possibility of determining the binding kinetics of ions to biomolecules. Consequently, these techniques have become a standard tool in the early phase of the development of novel therapeutics in recent years.

[0010] However, the immunopeptidome is estimated to contain at least 150.000 different pHLA complexes for the class I HLA molecules alone. This illustrates the strong need and potential for a new ultra-high throughput approach to characterize as many interactions in parallel as possible.

[0011] Currently, state of the art technologies can only realize such a high-throughput to a limited extend. Moritz et. al. 2019 were able to remove the obstacles of large scale high quality pHLA library generation for kinetic screening purposes by introducing disulfide-stabilized HLA molecules. The empty HLA complexes of said approach make it possible to bypass the timeconsuming step of refolding every individual pHLA complex with simple peptide loading reactions. Nevertheless, their measurements were limited to 16 different pHLA-analyte interactions in parallel using a BLI system. This throughput could at best be increased to 96 interactions in parallel on a different BLI machine with lower polling rates. Array-based SPR systems such as IBIS-MX96 have also been used successfully, but are limited to 96 parallel measurements, which is still far from a true ultra-high throughput approach. Even the newest Carterra instrument, the most advanced SPR systems for multiplexing, can only handle 192 measurements in parallel. Since microarrays were originally developed for high-throughput analysis, they are an excellent tool to determine interactions quickly and cost-efficiently.

[0012] Yet, standard microarrays also face some challenges. On the one hand, kinetic data is typically not easy to evaluate using microarrays. On the other hand, in the case of peptide-HLA complexes, the peptide exchange must be performed beforehand in a microtiter plate format, which often consumes large amounts of material. Additionally, with increased numbers of different pHLA molecules and spots, the spotting duration also increases, and molecule stability becomes a challenge particularly with comparatively unstable complexes like pHLAs. Furthermore, the storage of microarrays after their production and ensuring the functionality of the molecular complexes is an additional challenge that should not be underestimated.

[0013] SUMMARY OF THE INVENTION

[0014] In light of the prior art the technical problem underlying the present invention is to provide alternative and / or improved means for predicting the binding of protein of interest, preferably a receptor, DARPINs, an antibody or antibody fragment or equivalents thereof, to a peptide of interest.

[0015] This problem is solved by the features of the independent claims. Preferred embodiments of the present invention are provided by the dependent claims.

[0016] In one aspect the invention therefore relates to a computer-implemented method for predicting the binding of at least one protein of interest to at least one peptide of interest comprised within a HLA-peptide (pHLA) complex of interest, comprising a. providing sequence and / or structural data of at least one reference peptide, optionally providing sequence and / or structural data of at least one HLA molecule and / or of at least one HLA-peptide complex (pHLA), and / or the at least one reference peptide presented within a pHLA complex, b. providing binding data, preferably comprising the dissociation constant (KD), of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide complex (pHLA) with at least one protein of interest, c. training a machine learning (ML)-based model or artificial intelligence (Al) based on the data provided in a. and the binding data provided in b., d. providing sequence and / or structural data of at least one peptide of interest, e. employing the machine learning (ML)-based model or artificial intelligence (Al) trained in c. to predict the binding or interaction properties of the at least one peptide of interest, preferably when presented within a pHLA complex, to the at least one protein of interest.

[0017] In preferred embodiments the at least one protein of interest is at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest. In other words, in embodiments, the protein of interest, which specifically binds to a peptide of interest, is selected from a receptor, a designed ankyrin repeat protein (DARPin), an antibody, or antibody fragment or equivalent thereof. In embodiments the at least one protein of interest is at least one receptor, antibody, or antibody fragment or DARPin of interest. In embodiments the at least one protein of interest is at least one receptor, antibody or DARPin of interest. In embodiments, the at least one protein of interest is at least one receptor. In embodiments, the at least one protein of interest is at least one antibody of interest, or antibody fragment, or equivalent thereof. In embodiments, the at least one protein of interest is at least one immunoglobulin of interest. In embodiments the at least one protein of interest is at least one DARPin of interest. In some embodiments, the at least one protein of interest may be any protein or molecule of interest binding, or capable of binding to a pHLA complex.

[0018] In embodiments, the machine learning (ML)-based model or artificial intelligence (Al) comprises or is at least one neural network or deep learning neural network. In embodiments, the computer-implemented method for predicting the binding of a protein, preferably a receptor, an antibody or antibody fragment or equivalents thereof, or a DARPin, to a peptide of interest, comprises a. providing sequence and / or structural data of at least one reference peptide, optionally providing sequence and / or structural data of at least one HLA molecule and / or of at least one HLA-peptide complex (pHLA), and / or the at least one reference peptide presented within a pHLA complex, b. providing binding data, preferably at least comprising the dissociation constant (KD), of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide complex (pHLA) with at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest, c. training a machine learning (ML)-based model or artificial intelligence (Al) based on the data provided in a. and the binding data provided in b., d. providing sequence and / or structural data of at least one peptide of interest, e. employing the machine learning (ML)-based model or artificial intelligence (Al) trained in c. to predict the binding or interaction properties of the at least one peptide of interest, preferably when presented within a pHLA complex, to the at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest.

[0019] In embodiments, the binding data of the at least one reference peptide and / or the at least one reference peptide presented within an HLA-peptide complex (pHLA) provided in b. has been determined experimentally.

[0020] In embodiments, the experimental determination of the binding data provided in b. comprises one or more of the following steps: i. providing at least one reference peptide, ii. supplying HLA-molecules, preferably biotinylated HLA-molecules, to the at least one reference peptide and allowing the HLA-molecules to bind to the at least one reference peptide, thereby forming HLA-peptide (pHLA) complexes, iii. immobilizing the formed pHLA complexes on a solid support, e.g., on a streptavidin-coated solid support, iv. contacting the immobilized pHLA complexes with at least one protein of interest, preferably with at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest, v. determining binding data of the individual pHLA complexes with the at least one protein of interest, preferably with the at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest.

[0021] In some embodiments, the solid support is a microfluidic chip and the afore experimental determination of the binding data (in vitro screening) is performed on a microfluidic chip. In some of such embodiments the in vitro experimental determination of the binding data comprises the synthesis of the selected reference peptides and their incorporation into pH LA complexes. In embodiments, the pHLA complexes comprising the reference peptides are then immobilized on the solid support, namely a microfluidic chip. In embodiments, the protein of interest is then provided to the microfluidic chip, e.g., ‘flowed’ over the chip. In embodiments, kinetic interaction values for each pHLA complex are measured using a preferably label-free, image-based method, e.g., via proprietary reflectometric devices (SCORE). In embodiments, binding curves are generated from the measured data. In embodiments, a quality-control and / or extraction of interaction data, e.g., comprising dissociation constants (KD values), for each interaction between a protein of interest and a pHLA complex may be performed, in some specific embodiments using a proprietary in-house software package, such as the Anabel software package or equivalents.

[0022] In some embodiments, the experimental determination of the binding data comprises a feedback loop to further optimize the pHLA set for the in vitro interaction screening of the protein of interest.

[0023] In some of such embodiments the experimentally obtained interaction data, e.g., comprising KD values, may then be used to train ML-models, such as a pretrained protein language model, which was preferably originally trained on known peptide-H LA affinities and associated sequence data. In embodiments, such ‘updated’ ML-model, which may be trained (again based on updated or new data) can be (more) specific for a protein of interest, regarding one or more interaction properties, e.g., KD values, from the in vitro screening measurements as input and is preferably capable of predicting interaction properties, e.g., KD values, for unseen / unknown pHLA complexes that were not included in the in vitro experiment.

[0024] In embodiments, such process enables the construction of an affinity landscape in the representational space of the language model, connecting experimentally analysed / measured and unmeasured pHLA sequences with respect to their binding / interaction properties to a protein of interest. In embodiments, reference peptides that demonstrate high-affinity binding to the protein of interest may be grouped, e.g., based on their learned feature representations, with unmeasured peptide sequences exhibiting similar (biochemical, sequence or other applicable) characteristics. This significantly broadens the predictive coverage of the model across the human immunopeptidome.

[0025] In embodiments such model is capable of identifying peptides that are structurally unrelated to the original input sequence but are still predicted to bind the protein of interest.

[0026] In embodiments such newly identified peptides are subsequently printed on the solid support for another in vitro screening, e.g., a microfluidic chip, and experimentally tested (‘validated’) using the same kinetic in vitro assay. In embodiments, the resulting interaction data, e.g., comprising KD values, may then be used to assess the risk of cross-reactivity between the protein of interest and any peptide-H LA complex encoded within the human immunopeptidome.

[0027] In embodiments, the obtained (in vitro) binding data of the individual pHLA complexes with the at least one protein of interest, preferably obtained by high throughput binding experiments I screening, comprise: i) the association rate constant (kass), ii) the dissociation rate constant (kdiss) and / or iii) the dissociation constant (KD).

[0028] One advantage of using interaction / binding data from large in vitro interaction screenings according to the present invention over the use of general open source training data is that in vitro interaction screenings according to the present invention may in embodiments be performed in a customized manner, particularly with respect to specific proteins of interest, or specific groups of reference peptide, such that the computational model(s) may be subjected to a more targeted training, preferably also on both positive (binding) and negative interaction data (no binding / interaction) between a protein of interest and respective reference peptides, which are preferably comprised within a pHLA complex, leading to a more precise and reliable computational prediction of interactions between a protein of interest and a peptide of interest.

[0029] Complex protein molecules, such as immunoglobulins, antibodies, antibody fragments or DARPINs, have nearly infinite optimization possibilities, which makes them extremely hard to engineer for a certain purpose merely by single wet lab interaction screenings. The inventors herein propose that instead of performing an infinite amount of experiments using embodiments of the present computer-implemented method to predict the interaction and binding behavior of a molecule of interest, such as a protein of interest, allowing for in silico (computational) prediction and identification of the most promising molecule interactions, e.g., the interaction profile of a protein of interest as drug candidate, thereby facilitating massive cost and time savings over conventional wet lab-only approaches.

[0030] In general, it has been shown that a cornerstone of using Machine learning or Al in drug discovery is the availability of large, high-quality data sets for training purposes. An advantage of embodiments of the present invention is that the training data for the Al or ML model may in embodiments be obtained through a high throughput in vitro screening that enables the provision of large amounts of high-quality training data in high-throughput which may then be integrated with the computer-implemented methods according to the present invention thereby enhancing predictive modeling and streamline decision-making processes. The training data preferably comprises in embodiments antibody-pHLA binding data, complex pHLA interaction profiles an / or selected biochemical properties of one or more proteins of interest, such as immunoglobulins, antibodies, antibody fragments, or DARPins. In embodiments the present method therefore advantageously enables the provision and use of excessive, in vitro generated training data for the one or more Al or ML models, for generating a highly precise digital representation of a molecule or protein of interest. In embodiments, the integration and use of the in vitro generated training data into one or more Al and / or ML models according to the present invention enables also a continuous refining and enhancing of the employed Al or ML models through high quality and high throughput wet-lab data generation. This approach is considered to enable the training and provision / generation of unique Al or ML-models and thereby provide precise representations of the interaction properties (interactome) of a protein of interest that may not be generated by prior art methods.

[0031] In embodiments, the Al and / or ML-models used in the context of the present method enable specifically modeling the presentation of peptides of interest by HLA molecules (as pH LA complexes), based on the input data provided, e.g., structural, sequence and / or binding / interaction data. This approach significantly broadens the scope of predicted interactions / binding of a protein of interest, ensuring comprehensive coverage and precise annotation. In embodiments, the present method is particularly advantageous for predicting or estimating the scope of desired and undesired / off target interactions a protein of interest. This is particularly beneficial, e.g., in drug development and for selecting or modifying a protein of interest with respect to a desired interaction profile, e.g., when used as medication. Particularly pHLA interactions constitute one of the major challenges for therapeutic drugs in view of their side effects, toxicity and off-target binding, e.g., in the pH LA therapeutic space. In this context, the present method enables, e.g., due to the Al or ML-model trained in high throughput experimental interaction data, the accurate and reliable prediction of off-target binding for any peptide of interest (which e.g., may be an undesired off-target) to a protein of interest.

[0032] In embodiments, binding data provided in b. comprises interaction data of both positive (binding) and negative (does not bind) interaction partners (reference peptides) of a protein of interest (preferably a ‘balanced’ mix of both, e.g., comprising information of at least 10-65% negative / non- binding partners, e.g., about 40-60%). In some embodiments it may be advantageous, if a certain number of negative interaction partners (reference peptides) are similar or even highly similar (e.g., regarding their amino acid, and / or encoding nucleic acid sequence) to a known target binding partner (or target motif) of a protein of interest, thereby increasing the sensitivity of the model to differentiate between binding and non-binding partners. An advantage of such mixture of positive and negative binding data for the computational prediction of the binding of a peptide of interest to a protein of interest is that negative binding data is capable of preventing the used ML- model or Al from becoming overly optimistic (leading e.g., to false-positive predictions). This is, for example, how the present method overcomes a problem of most state of the art datasets, e.g., open source data sets of protein-protein interactions (e.g., derived from mass spectrometry data), which do not comprise data on negative (no binding) interactions thereby only allowing training and predictions based on positive binding I interaction data.

[0033] In embodiments, binding data provided in b. comprises interaction data on reference peptides (alone and / or presented in a pHLA complex), wherein at least some of the reference peptides comprise one or more mutation(s) with a distance of at least one to an optimal target peptide epitope of a protein of interest. Such data may assist the assessment of the importance of a certain amino acid of a reference peptide (sequence) for the binding to a protein of interest.

[0034] In embodiments, binding data provided in b. also comprises interaction data on reference peptides considered off-targets showing an unspecific or undesired binding or no binding to a protein of interest. In embodiments, the predicted binding or interaction properties in step e. comprise the binding strength, preferably comprising the dissociation constant (KD), of the at least one protein of interest, preferably of the at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest, to the at least one peptide of interest.

[0035] In embodiments, the predicted binding or interaction properties in step e. comprise: an indication of (e.g., numeric value indicative of) the probability or likelihood, the association rate constant (kass), the dissociation rate constant (kdiss) and / or the dissociation constant (KD) for the binding of the at least one protein of interest to the at least one peptide of interest.

[0036] In embodiments, the predicted binding or interaction properties in step e. comprise the binding strength, preferably comprising the dissociation constant (KD), and / or the probability or likelihood of binding of the at least one protein of interest to the at least one peptide of interest, preferably when presented within a pHLA complex.

[0037] In embodiments, predicting the binding or interaction properties in step e. comprises predicting the probability or likelihood of binding of the at least one protein of interest to the at least one peptide of interest, optionally represented by a numeric value or score indicative of the probability or likelihood of an interaction (e.g., any value between 0 (no binding) and 1 (binding)).

[0038] In embodiments, the at least one peptide of interest comprises at least one mutation, preferably with respect to at least one reference peptide.

[0039] In embodiments, the at least one peptide of interest comprises two or more mutations, preferably with respect to at least one reference peptide, such that information about synergistic effects of said mutations, e.g., on the binding or interaction properties, may be obtained. In embodiments synergistic effects may refer to any effect achieved by a certain combination of amino acids within an amino acid sequence of a peptide or protein, e.g., of amino acids in close proximity within an amino acid sequence or protein structure, which may be beneficial or disadvantageous for the binding or interaction to an interaction or binding partner, e.g., a protein or peptide of interest or reference peptide. In exemplary embodiments, a synergistic effect of an amino acid mutation may be an increased binding affinity to a target motif (e.g., of a peptide) or binding partner (e.g., protein of interest) due to an advantageously interaction (synergy, potentiation) of the physiochemical properties of certain amino acids within the peptide sequence leading to increased binding.

[0040] In embodiments, in step a. also data on physiochemical properties of the at least one reference peptide, and / or sequence and / or structural data of the at least one protein of interest and / or of the pHLA complex comprising the least one reference peptide are provided.

[0041] In embodiments, in step a. also data on physiochemical properties of the at least one reference peptide, and / or sequence and / or structural data of the at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest and / or of the pHLA complex comprising the least one reference peptide are provided. In embodiments, the data on physiochemical properties of the at least one reference peptide, and / or sequence and / or structural data of the at least one protein of interest (preferably at least one receptor, antibody or antibody fragment, or DARPin of interest) and / or of the pHLA complex comprising the least one reference peptide provided in e) has been provided or predicted by at least one machine learning (ML) model, artificial intelligence (Al) or at least one neural network (NN).

[0042] In embodiments, structural feedback is provided regarding the structural predicted target and molecule of interest, e.g., influencing the interaction of a predicted target and protein of interest; e.g., wherein the structural interaction is determined by modelling of in silica structural modelling, e.g., with open source tools, known to the skilled person, e.g., without limitation thereto AlphaFold, AlphaFold2, RoseTTAFold, ESMFold, PyMOL and / or AutoDock Vina / AutoDock-GPU.

[0043] In embodiments, the data provided in a. and / or d. has been determined experimentally and / or has been predicted by one or more machine learning (ML)-based model or artificial intelligence (Al).

[0044] In embodiments, the machine learning (ML)-based model or artificial intelligence (Al) comprises one or more machine learning model(s) and / or Al.

[0045] In embodiments, the machine learning (ML)-based model or artificial intelligence (Al) comprises one or more neural networks.

[0046] In embodiments, step e. further comprises: comparing the binding or interaction properties of the at least one protein of interest, preferably of the at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest, to

[0047] A) the at least one reference peptide and / or the at least one reference peptide presented within an HLA-peptide (pHLA) complex, with

[0048] B) the at least one peptide of interest, preferably when presented within a HLA- peptide (pHLA) complex, predicted in e.

[0049] In other words, in embodiments step e. further comprises a comparison of the (in vitro determined) binding or interaction properties of the at least one protein of interest to the at least one reference peptide (optionally presented in a pHLA complex) to the (in e. predicted) binding or interaction properties of the at least one protein of interest with the at least one peptide of interest (preferably when presented within a pHLA complex).

[0050] In embodiments, the input provided for the training in c. and / or prediction in e. are amino acid sequences. In embodiments, the input provided for the training in c. and / or prediction in e. are amino acid sequences which are converted to a numerical representation, preferably using an embedding model. In some of such embodiments, the numerical representation of amino acid sequences are then passed to a respective ML-model or Al for training and / or prediction, respectively.

[0051] In embodiments, various embedding models could be used interchangeably for embedding any data, for example, any substitution matrix like BLOSUM62, or any (pretrained) neural-networks, algorithms or data structures able to handle / process amino acid sequences, such as, without limitation ‘protein language models’ like ESM, ESM2 (Evolutionary Scale Modelling) or AntiBerty. In embodiments, a one hot encoding may be applied for embedding amino acid sequences, or any other dedicated embedding model. In embodiments, various methods to vectorize / encode amino acid sequences could be used in the context of the present method, and may be considered interchangeably.

[0052] In embodiments, the embedding model used for embedding input data is a substitution matrix, such as BLOSUM62. In embodiments, the embedding model used for embedding input data is an algorithm, neural network or data structure, such as a protein language model like ESM, ESM2 or similar. In embodiments, the embedding model used for embedding input data is an algorithm, neural network or data structure, such as a protein language model like AntiBerty.

[0053] In embodiments, the embedding model(s) may also be trained on experimental data and be incorporated as part of the ML- or Al model used for the final prediction, e.g., according to method step e.

[0054] In embodiments, the ML- or Al model may comprise or consist of at least one neural-network (NN) or gradient boosted model, and may have a variety of architectures such as a multi-layer perceptron (MLP), Transformer, or convolutional neural network (CNN). In embodiments, the ML- or Al model may comprise or consist of at least one neural-network (NN) or gradient boosted model, such as a multi-layer perceptron (MLP). In embodiments, the ML- or Al model may comprise or consist of at least one neural-network (NN) or gradient boosted model, and may have a variety of architectures such as a convolutional neural network (CNN).

[0055] The inventors have successfully used a number of different variants of computational models, such that the present method is considered to not depend on any specific computational implementation.

[0056] In specific embodiments, the present method is used to generate a molecule specific model for the prediction of a binding / interaction properties or other biochemical properties of a protein of interest. For example, in embodiments a training dataset comprising kinetic (interaction / binding) data of different pHLA combinations with a molecule of interest is used in a deep learning approach (e.g., step c.). In embodiments the input provided for the reference peptides (a.) and / or for a peptide of interest for the final prediction (e.) comprises at least respective amino acid sequences of respective peptides. In embodiments, such model or Al may be used for predicting the binding / interaction properties (e.g., binding I non-binding) of a peptide of interest

[0057] In specific embodiments, a ‘target specific model’ may be generated using at least one pretrained deep learning I ML-based model, such as protein language models like ESM (e.g., suitable for prediction of protein structures (ESMFold), MHC-peptide binding, CR-MHC binding, or protein generation (Genie2)) and AntiBerty (e.g., suitable for prediction of protein structure (IgFold), or thermostability) for embedding sequence and / or structural input data. In some embodiments, such pretrained models may be included / embedded into respective ML-model(s) / Al for the (final) prediction of the interaction / binding properties of a peptide of interest to a certain protein of interest (e.g., in step e.).

[0058] In specific embodiments, a ‘protein of interest-specific model’ may be generated employing a ML- model or Al that was trained, e.g., in a deep learning approach, according to the present method, for predicting the binding / interaction properties (e.g., binding I non-binding) for a given peptide sequence of a peptide of interest provided as input.

[0059] In embodiments, a K-Fold cross validation may be used as training method I approach.

[0060] However, any standard machine learning methodology may be used in the context of the present method, as, e.g., a key detail of the present method is amongst others the training data used to train the AL- or Al model(s) comprising a large number of specifically chosen peptides and / or comprising experimentally determined binding data according to the present method.

[0061] In embodiments, the machine learning (ML)-based model or artificial intelligence (Al) comprises an architecture of a CNN. In embodiments, the input data (comprising sequence and / or structural data of at least one reference peptide, and / or binding data of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide (pHLA) complex with at least one protein of interest) is embedded into a matrix format, such as (fixed-size) input tensors (e.g., with the dimension of 10x20), preferably using a substitution matrix or any other protein language embeddings, such as BLOSSUM62 or ESM2 or similar.

[0062] In embodiments, the architecture of a CNN comprises one or more (one-dimensional) convolutional layers, e.g., two one-dimensional convolutional layers. In embodiments, the one or more convolutional layers comprise ReLU (rectified linear units) activations.

[0063] In embodiments, the architecture of a CNN comprises one or more fully connected layers, e.g., one fully connected layer. In embodiments, the one or more convolutional layers, preferably comprising ReLU (rectified linear units) activations, are followed by one or more fully connected layer(s).

[0064] In a specific embodiment the architecture of a CNN comprises two convolutional layers with ReLU activations, which are followed by one fully connected layer.

[0065] In embodiments, the training in c. comprises an x-fold cross validation, preferably between a 5-15 fold, more preferably a 10-fld cross validation, optionally wherein each fold comprises a certain percentage or fraction of the input data and a remainder is left (held out) for validation.

[0066] In embodiments each fold of the training is trained between 1-10 times, preferably five times. In embodiments each fold of the training is trained with random seeds and / or data shuffling. In embodiments an ensemble of between one and ten models, preferably of five models, per fold is formed. In some embodiments between 10 and 100, preferably about 50 neural networks are generated / yielded per protein of interest.

[0067] In embodiments the training in c. uses a fixed number of epochs (‘rounds’ of training). In embodiments the training in c. uses binary cross entropy loss, e.g., to estimate the performance of the model, e.g., with regard to predicted vs. observed probability distributions.

[0068] In embodiments, the output is a single scalar. In embodiments, the output is a single scalar, passed through a sigmoid activation function, e.g., to yield a binding probability for binary classification.

[0069] In embodiments, step e. further comprises predicting the amino acid sequence of a (optimal) target binding motive of the at least one protein of interest, preferably of the at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest, preferably wherein the (optimal) target binding motive is predicted to have the highest binding affinity of all possible amino acid sequences to the at least one protein of interest, preferably to the at least one receptor, antibody, or antibody fragment or equivalent thereof, or DARPin of interest.

[0070] In embodiments, the predicted amino acid sequence of the (optimal) target binding motive is subsequently used to determine and / or compare potential off target-binding peptides to the at least one receptor, antibody or antibody fragment or equivalent thereof of interest. In embodiments, the predicted amino acid sequence of the target binding motive is subsequently used to determine and / or compare one or more peptides of interest potentially showing undesired or (disadvantageous) off target binding to the at least one protein of interest, preferably to the at least one receptor, immunoglobulin, antibody, or antibody fragment or equivalent thereof or DARPin of interest.

[0071] In embodiments, the predicted amino acid sequence of the (optimal) target binding motive is subsequently used to determine and / or compare potential off target-binding peptides from the entire or partial human peptidome to the at least one receptor, antibody or antibody fragment or equivalent thereof of interest.

[0072] In embodiments, sequence and / or structural data comprises information on the amino acid sequence, the three-dimensional structure, the two-dimensional structure, biochemical properties (like, e.g., electrostatic patches), the SMILES string, the InChi, the InChlKey, the chemical structure, any other biochemical property of a biomolecule, e.g. such as electrostatic patches, and / or other molecule features, optionally wherein any of the afore were predicted by one or more ML model or Al.

[0073] In embodiments, the pHLA complex is present in the human immunopeptidome or a synthetic pHLA complex. In embodiments, the HLA-peptide (pHLA) complex in which a reference peptide is presented and / or for which sequence and / or structural data is provided is either present in the human immunopeptidome or is a synthetic pHLA complex. In embodiments, HLA (proteins) and MHC (proteins) may be considered interchangeably. Without being bound by theory, HLA proteins may in embodiments be considered human MHC proteins. In embodiments, the present invention may be used for any MHC protein, peptide or complex.

[0074] In embodiments, the protein of interest is an antibody, preferably a TCR-like antibody (TCRm) or fragment thereof.

[0075] In embodiments, the receptor is a T cell receptor or fragment or equivalent thereof.

[0076] In embodiments, the receptor is a T cell receptor or fragment thereof for use in (adoptive) T cell therapy. In general, (adoptive) T cell therapy may comprise the use of patient-derived T cells, which are capable of recognizing tumor-specific antigens of a tumor of said patient. In some T cell therapies the T cells may be genetically engineered to express T-cell receptors (TCRs) with a particular high-affinity to a tumor antigen. In some embodiments the receptor is a TCR-like antibody or equivalent thereof, equivalent thereof, preferably capable of recognizing tumorspecific antigens of tumor cells.

[0077] In embodiments, the protein of interest may be a soluble pHLA binder (pHLA-binding molecule, molecule binding to pHLA complexes) or a membrane bound pHLA binder (pHLA-binding molecule, molecule binding to pHLA complexes). In embodiments, the protein of interest may be a soluble molecule binding to pHLA complexes (pHLA-binding molecule). In embodiments, the protein of interest may be a membrane-bound molecule binding to pHLA complexes (pHLA- binding molecule).

[0078] In embodiments, the protein of interest is a molecule capable of binding a pHLA complex. In some embodiments, the protein of interest is a genetically modified protein or antibody, or equivalent or antibody fragment.

[0079] In general, proteins of interest that are capable of binding a pHLA complex may be used in embodiments for mediating anti-tumor effects in a patient, such as mediating cell-mediated cytotoxicity (ADCC) of natural killer (NK) cells, complement-dependent cytotoxicity (CDC), antibody-dependent cellular phagocytosis (ADCP) by macrophages, or tumor cell apoptosis. Another application for proteins of interest may in embodiments be CAR-T cell therapy, or TCR-T cell therapy. Hence, the present method may be used in embodiments to develop, optimize or test therapeutic molecules of interest with regard to their binding / interaction properties to peptides or preferably pHLA complex-bound peptides of interest, e.g., in the context of developing new cancer therapeutics.

[0080] The proteins of interest that are capable of binding a pHLA complex may be used in embodiments for optimizing or developing research applications, e.g., when a protein of interest is conjugated to a fluorescent agent to detect the expression of target peptides or pHLA complexes on tumor cells.

[0081] In some embodiments, the equivalent of a receptor, an antibody or antibody fragment thereof is any biomolecule capable of recognizing and / or specifically binding antigens and / or capable of (specifically) binding to pHLA complexes. In embodiments, the term ‘any biomolecule capable of recognizing and / or specifically binding antigens and / or capable of (specifically) binding to pHLA complexes’ also comprises DARPins.

[0082] In specific embodiments, providing binding data of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide complex (pH LA) with at least one protein of interest comprises: a. selecting I defining a pHLA set to be experimentally analyzed regarding its binding to one or more protein(s) of interest.

[0083] In embodiments, a marker characterizing such a pHLA set of interest is the mutation distance to the target peptide sequence. Commonly, a distance of 1 means that the peptide (here preferably the peptide presented within a HLA-peptide) differs by 1 mutation (including insertions and deletions) from the target. Commonly, the maximum number of mutations is defined by the target length. In embodiments, it is preferred to include all possible peptides with a distance of 1 and an equal mix of all other mutations.

[0084] In embodiments, one or more of the following preferred 4 criteria for choosing the pHLA set to be experimentally analyzed may be applied:

[0085] 1 . The pHLA set may preferably be as diverse as possible, across all numbers of mutations and all amino acids. This may ensure that a pHLA set carries as little bias as possible.

[0086] 2. The minimal number of peptides in the pHLA set may preferably be at least at about 300 peptides.

[0087] 3. Preferably, the chosen peptides are of physiological relevance. In embodiments, peptides of physiological relevance are commonly found in the human peptidome.

[0088] 4. Preferably, a pHLA set may be composed of molecules binding and not-binding to an actual therapeutic molecule. However, in some embodiments, a pHLA set may be enlarged or redefined after the first round of wet lab experiments, as this may be difficult to determine in advance.

[0089] In embodiments, a pHLA set ma be composed of ~ 50% non-binder and ~ 50% binder to an actual therapeutic molecule.

[0090] In embodiments, a list, database or dataset of reference peptides may be generated or supplemented by reference peptides, e.g., of between 5-30, preferably between 10-15 amino acid length, generated from online data sources, such as UniProt, NCBI And / or Ensemble database. In embodiments such reference peptides generated from online data sources are generated by generating peptides using a sliding window approach, which has a respective length of amino acids, from the human proteome or a selection of the human proteome from such databases or online resources. Reference peptide obtained by such an approach may further be selected or filtered with an appropriate tool, e.g., NetMHCpan, respective their potential to be presented by HLA-A. Such reference peptides my also integrated into respective input matrix format, e.g., into an input substitution matrix, e.g., using BLOSSUM62. In embodiments, determining binding data of the at least one (reference) peptide and / or the at least one reference peptide presented within a HLA-peptide complex (pHLA) with at least one protein of interest, comprises selecting I defining a pHLA set to be experimentally analyzed regarding its binding to one or more protein(s) of interest.

[0091] In embodiments the at least one (reference) peptide is to be presented in the binding or interaction screening within a HLA-peptide (pHLA) complex.

[0092] In embodiments the (experimental) determination of the binding data in an binding or interaction screening comprises the following steps: i. providing at least one (reference) peptide, ii. supplying HLA-proteins, preferably biotinylated HLA-proteins, to the at least one (reference) peptide and allowing the HLA -proteins to bind to the at least one (reference) peptide, thereby forming pHLA complexes, iii. immobilizing the formed pHLA complexes on a solid support, preferably on a streptavidin-coated solid support, iv. contacting the immobilized pHLA complexes with at least one protein of interest, v. determining binding data of the individual pHLA complexes with the at least one protein of interest.

[0093] In embodiments the binding or interaction data (or properties) determined in the interaction screening comprise the binding strength, preferably comprising the dissociation constant (KD), of the at least one protein of interest to at least one (reference) peptide.

[0094] In embodiments the at least one (reference) peptide comprises at least one mutation with respect to a known (optimal) target peptide motif of the at least one protein of interest.

[0095] In embodiments, the experimental determination of the binding data of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide complex (pHLA) with at least one protein of interest comprises: b. performing high throughput binding experiments I analysis I screening of at least one protein of interest to the set of pH LAs selected previously.

[0096] A non-limiting technical example of such high throughput binding experiments I analysis / screening set up is shown in Figure 1.

[0097] In embodiments the high throughput binding experiments I analysis I screening comprise providing (e.g., spotting) at least one reference peptide, preferably different reference peptides, into individual cavities of a solid support, preferably a multi-well plate or microcavity chip, or equivalent thereof (e.g., step 1 Fig. 1). In embodiments the high throughput binding experiments I analysis I screening comprise, if the at least one reference peptide was provided in solution, drying the peptide solution within the individual cavities of the solid support (e.g., step 2 Fig. 1).

[0098] In embodiments the high throughput binding experiments I analysis I screening comprise an optional (long time) storage step of the solid support comprising the reference peptide(s) (e.g., step 3 Fig. 1 ; e.g., at RT or refrigerated, e.g., at ~ 4°C).

[0099] In embodiments the high throughput binding experiments I analysis I screening comprise providing I supplying (e.g., printing) biotinylated HLA-proteins into the individual cavities of the solid support, and preferably thereby to the at least one reference peptide, and allowing the HLA- molecules to bind to the at least one reference peptide, thereby forming pHLA complexes (e.g., step 4 Fig. 1).

[0100] In embodiments, the high throughput binding experiments I analysis I screening comprise immobilizing biotinylated HLA-molecules, preferably comprising the previously formed pHLA complexes, to a streptavidin (SA) coated solid support (e.g., step 5 Fig. 1).

[0101] In embodiments, the solid support may be used subsequently for a label-free analysis of different proteins of interest (e.g., step 6 Fig. 1) such that binding data of the at least one reference peptide and / or the at least one reference pHLA complex with at least one protein of interest, e.g., comprising the binding kinetics and / or dissociation constant (KD), can be determined, e.g., from the binding curves (e.g., step 7 Fig. 1).

[0102] In embodiments, the obtained I resulting (in vitro) binding data from high throughput binding experiments I analysis I screening comprise: i) the association rate constant (kass), ii) the dissociation rate constant (kdiss) and / or iii) the dissociation constant (KD).

[0103] In embodiments, the same or similar I equivalent values (e.g., IC50 or EC50) values obtained by ELISA, BLI or SPR may be used as well.

[0104] In embodiments, the training of the computational model (a machine learning (ML)-based model or artificial intelligence (Al), such as a (deep learning) neural network (NN) or a computational model with a NN architecture, may be performed based on the experimentally determined binding data.

[0105] In embodiments, different computational models, such as machine learning (ML)-based model(s), one or more artificial intelligence (Al) and / or one or more (deep learning) neural network(s) (NN) may be employed, however, as the field of machine learning and Al is constantly evolving, also other computational model architectures are also imagined and applicable.

[0106] In embodiments, the input (features) of the computational model (e.g., ML-model, Al and / or NN) may be one or more amino acid sequences, e.g., from at least one reference peptide, at least one HLA molecule and / or the pHLA complexes (pHLA set) selected for experimental analysis. In embodiments, the input (features) of the computational model (e.g., ML-model, Al and / or NN) for the training (c.) and / or for the final prediction in step e. comprises sequence and / or structural and / or binding data.

[0107] In embodiments, the input (features) of the computational model (e.g., ML-model, Al and / or NN) for the training (c.) comprises sequence and / or structural data of the at least one reference peptide and optionally sequence and / or structural data of at least one HLA molecule and / or the at least one reference peptide presented within a HLA-peptide (pHLA) complex (a.), and binding data of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide (pHLA) complex with at least one protein of interest (b.).

[0108] In embodiments, the input (features) of the computational model (e.g., ML-model, Al and / or NN) for the training (c.) and / or for the final prediction in step e. comprises one or more amino acid sequences, data on one or more protein structures, and / or binding data (preferably of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide (pHLA) complex with at least one protein of interest, preferably comprising the association rate constant (kass), the dissociation rate constant (kdiss) and / or the dissociation constant (KD).

[0109] In embodiments, for the training according to step c., the input comprises one or more amino acid sequences of and / or nucleic acid sequences encoding the at least one reference peptide and / or the at least one HLA molecule and / or the pHLA complexes comprising the at least one reference peptide.

[0110] In embodiments, the input (features) of the (trained) computational model (e.g., ML-model, Al and / or NN) for the final prediction in step e. comprises one or more amino acid sequences of the at least one peptide of interest, data on one or more protein structures of the at least one peptide of interest and / or nucleic acid sequences encoding the at least one peptide of interest.

[0111] In embodiments, labels may be applied to the input data, such as either 1 or 0, representing binding (1) and no binding (0), based on the fact that binding data, e.g., KD values, could be experimentally determined before, or not.

[0112] Hence, in embodiments the input data comprises a binary classification into: binding (1 ; e.g., Kd in vitro determined) or no binding (0; no Kd determined in vitro).

[0113] In embodiments, the output of the computational model (e.g., ML-model, Al and / or NN), e.g., in step e., may comprise a numerical value between 0 and 1 , representing a probability or likelihood (value) of expected binding (1), and no binding (expected) (0), and optionally any numeric value between 1 and 0 indicative of the probability or likelihood of binding, of the at least one protein of interest to the at least one peptide of interest. In some embodiments the classes (1 , binder and 0, non binder) are determined by a threshold of predicted binding probability / likelihood, e.g., in case of 2 classes 1 / 0 a threshold of 0.5.

[0114] In embodiments, the output of the computational model (e.g., ML-model, Al and / or NN), e.g., in step e., may comprise any other score suitable of indicating or representing a probability or likelihood of expected binding and no binding, and optionally any intermediate degree of probability or likelihood in between.

[0115] In embodiments, the predictivity of the model may be further boosted / improved by predicting more than 2 classes of output (e.g., indicating different degrees of binding, such as the scores: no-binding I weaker binding compared to target I similar binding as target I stronger binding compared to target).

[0116] Alternatively or in addition, further applicable input features are, e.g., structural embeddings of the pH LA complex or the protein of interest, physiochemical properties of the peptides and / or additional experimental data. Another option is the incorporating of large language models into the computational means like, e.g., the ESM 3.

[0117] In embodiments, after the training of the computational model in c., an additional in silico input data set may be generated / defined, for example, comprising 1000 or more peptide sequences. In embodiments, the in silico input data set may comprise, e.g., per mutation distance to a reference peptide and / or peptide of interest, 1000 or more respective peptide sequences that have been assembled / generated in silico.

[0118] In embodiments, the binding of some or all peptides comprised within the additional in silico input data set to the protein of interest may be predicted or calculated in silico. In embodiments, the binding data determined from such in silico predictions may be considered additionally during another training step of the computational model and / or the prediction performed in step e.

[0119] In embodiments, the prediction according to step e. further comprises the prediction of (the amino acid sequence of) a (accurate and / or optimal (high affinity)) target binding motive of the at least one protein of interest.

[0120] Identifying off-targets during the development of therapeutic molecules is crucial to ensure safety and efficacy by minimizing unintended interactions that could lead to adverse effects. In embodiments the present method may be used for predicting or estimating the toxicity and / or potential side effects and / or off-target binding of a protein of interest. In some embodiments the present method may be used for eliminating toxicity of a protein of interest, which may in embodiments constitute, e.g., a drug (candidate). Therefore, in embodiments extensive database searches (e.g., for undesired off-target peptides of interest associated with undesired site effects) may be combined with the present method to identify and reduce potential off-target binding of a protein of interest to a small selection of relevant (desired target) peptides, which may additionally subsequently be confirmed by wet lab experiments.

[0121] In embodiments the present method may further be used for predicting a target binding motif of a protein of interest. In embodiments the present method enables the prediction of binding motifs (target binding sequences) of a protein of interest, after training of the Al or ML-model on input data comprising binding / interaction data, preferably from in vitro interaction / binding screening experiments, of a protein of interest and sequence data of peptides used in said interaction screening, optionally in the context of pHLA complex presentation. From said learned interaction data between a protein of interest with different (preferably pH LA presented) peptides and their amino acid sequences the Al or ML-model is the able to predict a (desired, ideal or optimal) target binding motif of a protein of interest, e.g., with respect to Kass, Kdiss and / or KD values of the respective target motif. In embodiments, a particularly strong interaction / binding may be of interest, in other embodiments an interaction / binding with a specific kass, kdiss and / or KD value may be of interest.

[0122] In some embodiments, the present method may be used for predicting the structure of a protein of interest, such as an antibody, antibody fragment or DARPin of interest. In some of such embodiments an Al or ML-model of interest may be trained on different protein structures, and optionally their amino acid sequences, e.g., of VH and / or VL fragments of antibodies, or their fragments or equivalents, followed by the provision of one or more, optionally predicted, new structure(s) of protein(s) of interest. In embodiments, the present Al or ML-model may then be employed to predict in silica the binding / interaction of said new protein(s) of interest to different pHLA complexes or peptides of interest, e.g., using molecular operating environment (MOE) or equivalents. This approach may in embodiments comprise the in silica mutation of known proteins to generate in silica one mor more proteins of interest, whose binding / interaction to pH LA or peptide(s) of interest may then be predicted in silica according to the invention. Subsequently proteins of interest showing a desired interaction / binding profile may be synthesized and experimentally tested in vitro.

[0123] The present method may be particularly beneficial for determining a value for the specificity of a protein of interest, e.g., a therapeutic or diagnostic protein of interest, such as a therapeutic or diagnostic receptor, antibody, or antibody fragment or equivalent thereof, or DARPin. The present method may aid selecting promising candidate molecules for therapeutic or diagnostic application, e.g., a commonly initially a multitude of molecules is generated and has to be evaluated. In such scenarios, the present method enables informed decision making regarding which molecule to proceed with into an optimization phase or a potential preclinical and clinical study.

[0124] In embodiments the present invention further relates to a computer program (product) comprising instructions which, when executed by a computer, cause the computer to carry out the method according to the invention.

[0125] In one aspect the invention relates to a computer-readable medium or storage device comprising a computer program disclosed herein.

[0126] Each preferred or optional feature of the invention that is disclosed in the context of one aspect of the invention is herewith also disclosed in the context of the other aspects and embodiments of the invention described herein. All features disclosed in the context of the computer-implemented method according to the invention also relate to, and are herewith disclosed also in the context of the uses of said method and the computer program and storage devices disclosed herein, and vice versa. DETAILED DESCRIPTION

[0127] The present invention generally relates to computer-implemented methods for predicting the binding of at least one protein of interest to at least one peptide of interest, that is preferably comprised within a HLA-peptide (pHLA) complex of interest, and to computer-implemented methods for predicting the structure of a protein of interest or the binding motif of a protein of interest.

[0128] Herein, the input data for the computational model(s), e.g., ML-models or Al, used in the context of the present method preferably comprises sequence and / or structural data and / ore interaction / binding data.

[0129] In the context of the present application, ‘sequence data’ comprises any biological sequence data, e.g., amino acid sequence data and / or nucleic acid sequence data. Hence, in embodiments sequence data comprises data on the amnio acid sequence of a peptide or protein and / or nucleic acid sequence data of nucleic acids encoding a peptide or protein referred to herein. In embodiments sequence data comprises data on the amnio acid sequence of a peptide or protein of interest. In embodiments sequence data comprises nucleic acid sequence data of nucleic acids encoding a peptide or protein of interest.

[0130] In the context of the present invention, ‘structural data’ comprises any biological data on the structure, e.g., the 2- and / or 3 dimensional structure, of a protein or peptide.

[0131] In the context of the present invention, the sequence and / or structural data may comprise information on the amino acid sequence, the nucleic acid sequence encoding a respective peptide or protein, the three-dimensional (3D) structure, the two-dimensional (2D) structure, the SMILES string, the InChi, the InChlKey, the chemical structure, any other biochemical property of a biomolecule, and / or other molecule features.

[0132] In the context of the present invention, ‘interaction data’ or ‘binding data’ preferably comprises the experimentally determined interaction and / or binding data of at least one protein (e.g., a protein of interest) with another protein, peptide or preferably a peptide presented on a HLA protein / molecule forming a pHLA complex.

[0133] The ‘interaction data’ or ‘binding data’ or ‘data on the interaction I binding properties’ provided herein, that is preferably provided as an input to the computational model(s), preferably comprises experimentally determined data on the interaction and / or binding or interaction and / or binding properties of at least one protein to at least one peptide. The interaction or binding data comprises in preferred embodiments at least data indicating the binding strength, data regarding the positive or negative detection of binding / interaction, data on the duration of the binding, the association rate constant (kass), the dissociation rate constant (kdiss) and / or the dissociation constant (KD). In embodiments, the training of the Al or ML-model on positive and negative binding data is particularly advantageous to ensure that detected negative interactions (no binding) are true negatives and not assay artefacts. The ‘interaction data’ or ‘binding data’ or ‘data on the interaction I binding properties’ predicted herein, e.g., in method step e. according to the invention, preferably comprises data regarding the binding or interaction properties of the at least one peptide of interest, preferably when comprised within a pHLA complex, to the at least one protein of interest. In embodiments, the predicted data preferably comprises data indicating the binding strength, data regarding the positive or negative detection of binding / interaction, data on the duration of the binding, an estimated association rate constant (kass), an estimated dissociation rate constant (kdiss) and / or an estimated dissociation constant (KD). In embodiments, the ‘interaction data’ or ‘binding data’ or ‘data on the interaction I binding properties’ may further comprise one or more potential target motifs of a protein of interest, optionally also more than one proposed target motif, e.g., ranked according to expected binding properties to the respective protein of interest. In embodiments, binding or interaction properties comprise the binding strength, preferably comprising the dissociation constant (KD), and / or the probability or likelihood of binding of the at least one protein of interest to the at least one peptide of interest, preferably when presented within a pHLA complex. In embodiments the likelihood or probability predicted by the computational model(s), e.g., in step e., comprises a score or value indicative of the likelihood and / or strength of the predicted binding between a protein of interest to a peptide of interest, preferably when comprised or presented in a pHLA complex, e.g., a value between 1 (binding) and 0 (no binding), e.g., with a threshold of 0.5 between the two classes, or another threshold, e.g. if more than 2 classes are used.

[0134] Herein, the term ‘reference peptide’ preferably refers to one or more peptides that were comprised within the experimental determination of the interaction / binding data of the respective at least one protein of interest. In other words, a reference peptide is preferably a peptide for which interaction I binding data to a respective protein of interest is provided (as input to the computational model(s)), that has been determined before, preferably experimentally, and / or for which structural and / or sequence data is provided as input to the computational model(s). Hence, in embodiments for one or more reference peptides structural and / or sequence data and / or interaction I binding data to a respective protein of interest is provided as input to the computational model(s), preferably for training purposes of the computational model(s).

[0135] In general, the Major Histocompatibility complex (MHC) system is known in human as the human leukocyte antigen (HLA). The group of MHC / HLA genes in humans is located on chromosome 6 and encodes genes associated with immune system function in humans, such as cell-surface antigen-presenting proteins. The major HLA antigens are considered to be essential elements for immune function. Human leukocyte antigen molecules (HLA molecules) are proteins found on the surface of most cells in the human body that play a crucial role in the immune system. HLA molecules aid the immune system to distinguish self from non-self proteins, while specifically presenting protein fragments (peptides, antigens) to cells of the immune system, such as T cells. HLA molecules may be grouped into two major classes, namely HLA Class I and II. HLA antigens that correspond to MHC class I (HLA class I molecules) may be grouped into HLA-A, HLA-B or HLA-C 2. HLA class I molecules are found on all nucleated cells and present intracellular peptides (e.g., viral proteins) and are known to interact with CD8+ T cells (cytotoxic T cells). HLA Class II molecules (corresponding to MHC class II molecules) may be grouped into HLA-DP, HLA-DQ, HLA-DR. HLA Class II molecules are found mostly on immune cells (e.g., macrophages, dendritic cells, and B cells) and present extracellular peptides (e.g., such as bacterial proteins) and are known to interact with CD4+ T cells (helper T cells).

[0136] As used herein a HLA-peptide (pHLA) complex refers preferably to a molecular structure formed by a peptide binding to a HLA molecule. Commonly, such pHLA complexes are displayed on the surface of cells to be recognized by T cells of the immune system via their T-cell receptor (TCR).

[0137] In embodiments, the in vitro binding or interaction (high throughput) screening used for generating training data comprises a certain pHLA set, that is used for testing the binding or interaction of a protein of interest thereto. In embodiments, one marker characterizing a pHLA set used in said binding screen may be its mutation distance to a target peptide sequence, wherein a distance of 1 means that the peptide (sequence) differs by 1 mutation (e.g., 1 amino acid) from the target peptide (sequence).

[0138] Herein, a ‘target peptide sequence’ or ‘target sequence’ is preferably an amino acid sequence that is considered to be comprised within, or to be the (ideal / optimal) target binding sequence of a protein of interest, such as e.g., a target epitope of an antibody (fragment), a receptor (e.g., TCR) or DARPIN, or any other peptide sequence (of interest) that is considered a suitable target and should be used as basis for the analysis.

[0139] As used herein, an immunoassay or affinity-assay may refer in embodiments to a biochemical method employed to detect the presence or quantify the concentration of a macromolecule or polypeptide in a solution or to determine its binding strength and / or kinetics, or affinity to a protein of interest, e.g., utilizing said protein of interest, e.g., an antibody, antibody fragment, DARPin, receptor or immunoglobulin. Exemplary types of immunoassays include luminescence immunoassays (LIA), radioimmunoassays (RIA), chemiluminescence and fluorescence immunoassays, enzyme immunoassays (EIA), enzyme-linked immunosorbent assays (ELISA), luminescence-based bead arrays, magnetic bead-based arrays, protein microarray assays, rapid test formats, and rare cryptate assays. The term immunoassay encompasses a range of techniques, including, but not limited to, enzyme immunoassays (EIA) such as enzyme multiplied immunoassay technique (EMIT), enzyme-linked immunosorbent assays (ELISA), antigen capture ELISA, sandwich ELISA, IgM antibody capture ELISA (MAC ELISA), and microparticle enzyme immunoassays (MEIA); capillary electrophoresis immunoassays (CEIA); radioimmunoassays (RIA); immunoradiometric assays (IRMA); fluorescence polarization immunoassays (FPIA); lateral flow assays (LFA); turbidimetric assays; and chemiluminescence assays (CL) and label free binding kinetic techniques, such as SCORE, SPR and BLI.

[0140] Herein, the method may comprise or use computational means and / or computational models, such as one or more machine learning (ML)-based model(s), one or more artificial intelligence (Al) and / or one or more neural network (NN) or any combination thereof.

[0141] In general, machine learning comprises the field of supervised learning, namely the learning of a computational model or Al from labeled data, comprising linear regression, logistic regression, decision trees, random forests, k-nearest neighbors (k-NN), support vector machines (SVM), gradient boosting, as well as the field unsupervised learning, namely the learning of patterns from unlabeled data, comprising k-means clustering, hierarchical clustering, principal component analysis (PCA), DBSCAN and autoencoders. Another filed of machine learning is semisupervised learning that comprises a mix of labeled and unlabeled data, such as contrastive learning or contrastive learning objectives. Machine learning commonly further comprises the field of reinforcement learning (RL).

[0142] In embodiments herein the one or more machine learning (ML)-based model(s) or one or more artificial intelligence (Al) comprise or consist of one or more neural networks (NN), and optionally further (implemented) algorithms, Al or computational models. In embodiments the computational means referred to herein, e.g., as ‘Al’ or ‘at least one machine learning model’ may comprise numerous and diverse computational means that may work together and / or may be used in a computational pipeline when performing the present method.

[0143] In general, the term ‘neural networks’ comprises artificial neural network (ANN; in contrast to biological neural networks), such as convolutional neural networks (CNN), recurrent neural networks (RNN), transformer networks, generative adversarial networks (GAN), feedforward neural networks (FNN), spiking neural networks (SNN), multilayer perceptron (MLP) and autoencoders (encoder / decoder networks).

[0144] Commonly, generative adversarial networks (GAN) comprise two networks, comprising a generator and a discriminator. Commonly, autoencoders may be used for dimensionality reduction, denoising and anomaly detection. Commonly a FNN may be used for (simple) classification / regression. Recurrent neural networks (RNN) are commonly used to handle sequences, comprising LSTM (long short-term memory), and GRU (gated recurrent unit).

[0145] The present Al or ML-model is preferably able to perform tasks, such as regression modelling comprising predicting a value, such as Kd value, classification of data comprising predicting a class, such as binder or non-binder, and / or generation of data, comprising generating or predicting a new protein structure.

[0146] In embodiments of the present method different machine learning (ML) models may be used, and preferably trained on experimental binding / interaction and / or structure and / or sequence data of peptides (preferably comprised within pHLA complexes) and proteins of interest. Non-limiting and purely exemplary models comprise ESM and AntiBerty models, which are capable of predicting protein structures and interaction I binding based on learnings from training data.

[0147] In one non-limiting example, input data, comprising sequence data of reference peptides and / or proteins(s) of interest and / or in vitro binding / interaction data of the protein(s) of interest and the reference peptides is embedded (comprising the conversion of amino acid sequences to numerical representations) by ML-models, preferably pre-trained ML-models, such as ESM and / or AntiBerty. Subsequently the binding of a protein of interest may be predicted by the trained ML-model(s), e.g., based on sequence data of a peptide on interest provided to the ML-Model(s). In embodiments of the present invention various applicable embedding models could be used interchangeably, for example any substitution matrix like BLOSUM62, or any pretrained neural- networks that can handle sequences of amino acids, such as ESM2 (Evolutionary Scale Modelling). In embodiments, an embedding model may be employed and trained on experimental data and be incorporated as part of the ML- or Al model used for the final prediction in method step e. Similarly, in embodiments the ML- or Al model may comprise or consist of at least one neural-network or gradient boosted model, and may have a variety of architectures such as a multi-layer perceptron (MLP), Transformer, or convolutional neural network CNN. The inventors have successfully used a number of different variants of computational models, such that the present method is considered to not depend on any specific computational implementation.

[0148] In embodiments a protein of interest may be a receptor, an antibody, or an antibody fragment or any equivalent of an antibody or antibody fragment, or a designed ankyrin repeat protein (DARPin) or any comparable structure capable of specifically recognizing and binding antigens.

[0149] As utilized herein, the term "antibody" generally refers to a protein consisting of one or more polypeptide chains primarily encoded by immunoglobulin genes or their fragments. The term "antibody" is intended to encompass "antibody fragments" and any equivalents thereof as well. The term "antibody" is intended to encompass immunoglobulins, or may be used interchangeably with “immunoglobulin” herein.

[0150] In general, the immunoglobulin genes encompass the constant region genes for kappa, lambda, alpha, gamma, delta, epsilon, and mu, in addition to a diverse array of immunoglobulin variable region genes. Light chains are classified as either kappa or lambda, while heavy chains are categorized as gamma, mu, alpha, delta, or epsilon, which define the immunoglobulin isotypes IgG, IgM, IgA, IgD, and IgE, respectively. The fundamental structural unit of an antibody (immunoglobulin) is typically a tetramer or dimer. Each tetramer consists of two identical pairs of polypeptide chains, with each pair comprising one "light" (L) chain (~ 25 kDa) and one "heavy" (H) chain (~ 50-70 kDa). The N-terminal region of each chain defines a variable region, generally spanning approximately 100 to 110 amino acids, which is primarily responsible for antigen recognition. The terms "variable heavy chain" and "variable light chain" refer to the variable regions of the heavy and light chains, respectively. Optionally, the antibody or its immunologically active fragment may be chemically conjugated or expressed as a fusion protein with other proteins.

[0151] The term "specific binding", e.g., of a protein of interest, antibody, antibody fragment, receptor or DARPin, is intended to be interpreted as via a skilled person, who is clearly aware of different experimental procedures that may be used to test binding and binding specificity. While some degree of cross-reaction or background binding may occur in many protein-protein interactions, such occurrences do not undermine the "specificity" of the binding between an antibody and its epitope. Additionally, the term "directed against" is relevant in the context of "specificity" when describing the interaction between an antibody, antibody fragment, receptor or DARPin and its target epitope. Antibodies of the present invention include, but are not limited to, polyclonal, monoclonal, bispecific, human, or chimeric antibodies, single-domain antibodies (e.g., VHH fragments derived from nanobodies), single-variable fragments (ssFv), single-chain variable fragments (scFv), Fab fragments, F(ab')2 fragments, fragments produced via Fab expression libraries, anti-idiotypic antibodies, epitope-binding fragments, or any combination thereof, provided that they retain the original binding properties. Additionally, mini-antibodies and multivalent antibodies, such as diabodies, tetravalent antibodies, triabodies, and peptabodies, may also be employed in the methods described herein.

[0152] DARPins (Designed Ankyrin Repeat Proteins) are commonly considered a class of engineered proteins comprising ankyrin repeat domains, that constitute structural motifs consisting of tandem repeats, e.g., of ~ 33 amino acids. DARPins may be designed to bind with high specificity to a wide range of targets, including proteins, peptides, and other biomolecules, making them versatile tools in molecular biology, diagnostics, and therapeutics.

[0153] The immunoglobulin molecules referenced herein can belong to any class (i.e. , IgG, IgE, IgM, IgD, or IgA) or subclass of immunoglobulins. Accordingly, the term "antibody," as used in this context, also encompasses antibodies and antibody fragments that are either modified from whole antibodies or synthesized de novo using recombinant DNA technology.

[0154] A "variable region" of an antibody refers herein to the variable region of either the antibody light chain or the antibody heavy chain, individually or in combination. The variable regions of both the light and heavy chains comprise four framework regions (FRs) interspersed with three complementarity-determining regions (CDRs), also known as hypervariable regions. As used herein, a CDR may refer to CDRs defined by any suitable method. In embodiments, variations of variable region across different species, may be used and / or tested in the context of the present invention.

[0155] Herein a ‘solid support’ may be a microarray, glass plate, plastic plate or multi-well plate or any surface suitable for immobilizing a reference peptide, pHLA complex or protein of interest. In embodiments solid support of a certain kind may also be a streptavidin-coated solid support, or equivalents allowing for protein or peptide immobilization, optionally for covalent or non-covalent, permanent or transient immobilization.

[0156] Herein the term ‘physiochemical properties’ preferably comprises characteristics of amino acids influencing a proteins structure, such as hydrogen bonding, hydrophobicity, or charge. Inherently, the physiochemical properties of a peptide or protein are determined by the corresponding properties of its amino acids (sequence).

[0157] In general and in embodiments of the present invention, common physiochemical properties comprise, but are not limited to, molecular weight, LogP, LogD, pKa, aqueous solubility, hydrogen bond donors, hydrogen bond acceptors, topological polar surface area (TPSA), number of rotatable bonds, molecular volume, polarizability, refractivity, permeability, plasma protein binding, stability (chemical / metabolic), isoelectric point (pl), lipophilicity, charge at physiological pH. Herein, an optimal, high or even highest possible binding affinity (e.g., of all possible amino acid sequences, aa-motifs, peptides or proteins) refers to the binding affinity of a peptide to a protein of interest or vice versa, which is indicative of a positive binding of said peptide to a protein of interest or vice versa, namely a certain physicochemical affinity leading to their binding / interaction, which is preferably lower (the lower the KD value, the stronger the binding / interaction) than the binding affinity of the peptide or protein of interest to other proteins of interest or peptides, respectively.

[0158] In embodiments, the sequence and / or structural data may comprise information on the amino acid sequence, the nucleic acid sequence encoding a respective peptide or protein, the three- dimensional (3D) structure, the two-dimensional (2D) structure, the SMILES string, the InChi, the InChlKey, the chemical structure, any other biochemical property of a biomolecule, and / or other molecule features. In embodiments, the sequence and / or structural data may be I have been determined experimentally or may be I have been predicted by one or more computational model, such as a ML-model or Al.

[0159] The term ‘human immunopeptidome’ describes in general the entirety of all peptides that are being presented on HLA molecules on the surface of cells. These peptides originate from the digestion of cellular proteins. The immunopeptidome is a central part for the immune system. It provides it with the possibility to distinguish between foreign and self or healthy and sick.

[0160] FIGURES

[0161] The invention is further described by the disclosed figures. These are not intended to limit the scope of the invention but represent preferred embodiments of aspects of the invention provided for greater illustration of the invention described herein.

[0162] Figure 1 : In the depicted embodiments a binding analysis of the therapeutic molecules against the defined set of pHLAs is performed. Different peptides (from the set) are spotted into unique cavities of a microcavity chip (1). The peptide solution is dried out (2) and the arrays are ready for long time storage (3). To form the biomolecular complexes, biotinylated HLA is printed into the individual cavities and the complex formation takes place (4). The complexes are then transferred to a streptavidin (SA) coated substrate (5). The substrate can then be used for the label-free analysis of different analytes (6) and the kinetic values can be determined from the binding curves (7).

[0163] Figure 2: Exemplary, non-limiting embodiment of a work flow of the present method. (A) Depicted is a generalizing scheme of a possible architecture of the computational model and workflow of the present method. In the first depicted step (left), sequences of aminos acids (SEQ ID NO: 1) are converted to a numerical representation, preferably through an embedding model. The numerical representation is then passed to a ML or Al model that was trained on (in vitro) data to make a prediction. In embodiments the ML-model or Al was trained on datasets obtained by experimental interaction / binding assays to predict the binding / interaction properties of a protein of interest, e.g., the probability and / or strength of binding, e.g., antibody, antibody fragment, receptor or DARPin, binding to a peptide, preferably a peptide-H LA complex. Various applicable embedding models could be used interchangeably, for example any substitution matrix like BLOSUM62, or any pretrained neural-networks that can handle sequences of amino acids, such as ESM2 (Evolutionary Scale Modelling), see e.g., (B). In embodiments, the embedding model(s) may also be trained on experimental data and be incorporated as part of the ML- or Al model used for the final prediction, e.g., according to method step e. Similarly, the ML- or Al model may comprise or consist of at least one neural-network (NN) or gradient boosted model, and may have a variety of architectures such as a multi-layer perceptron (MLP), Transformer, or convolutional neural network (CNN). The inventors have successfully used a number of different variants of computational models, such that the present method is considered to not depend on any specific computational implementation. (B) Specific embodiment of (A) wherein the protein of interest is an antibody. For generation of a ‘target specific model’ pretrained deep learning I ML-based models, such as protein language models like ESM (e.g., suitable for prediction of protein structures (ESMFold), MHC-peptide binding, CR-MHC binding, or protein generation (Genie2)) and AntiBerty (e.g., suitable for prediction of protein structure (IgFold), or thermostability) were used und provided with sequences of reference peptides and at least parts of the sequence of a protein (here antibody) of interest as input (SEQ ID NO: 2-4). The pretrained models are included / embedded into the ML-model / Al for prediction of the interaction / binding properties of a peptide of interest to a certain protein of interest, e.g., also for sequence mutations of the protein (here antibody) of interest (top row of (B)). For generation of a ‘protein of interest-specific model’ a ML-model or Al that was trained, e.g., in a deep learning approach, according to the present method, may be used for predicting the binding / interaction properties (e.g., binding I non-binding) for a given peptide sequence provided as input (bottom row of (B)). (C) In embodiments, a K-Fold cross validation may be used as training method I approach. However, any standard machine learning methodology may be used in the context of the present method, as, e.g., a key detail of the present method is amongst others the training data used to train the AL- or Al model(s) comprising a large number of specifically chosen peptides and / or comprising experimentally determined binding data according to the present method.

[0164] Figure 3: Results of Example 1 for protein of interest “BCA0001”. For this experiment the experimental interaction data was split up into a set to be used either as input, or as a set of peptides of interest, for which only the sequence was provided and to be used for validation. (A) The graph shows the results for a binary classification of a set of peptides on interest, by the trained ML-models (ESM embeddings of peptides). The output is the predicted probability of a peptide of interest being a binder to the protein of interest, whereby the two classes (binder; 1 and non binder; 0) were determined by a threshold of a predicted binding probability of 0.5. (B) The graphs shows a ROC curve with a prediction accuracy of ROC AUC of 0.952 for the protein of interest, when compared to the experimentally determined binding data. (C) Theoretical principle of binding kinetics (kaand kd) of target peptide-pHLA complex (pHLA) and antibody (TCRm). (D) The graph shows the correlation (Pearson) of the predicted and experimentally measured Log KD values for the protein of interest (BCA0001) to peptides of interest. The Pearson correlation was R= 0.82. Figure 4: Predicted binding motifs for tested antibodies BCA0001 and BCA0002 (see Examples), wherein amino acids are depicted by one-letter code. The x-axis of the graphs respectively depict the amino acid position, y-axis depict the ‘information content’, wherein the size (height) of a depicted letter is representative of its likelihood of being present in the target motif at the respective sequence position. The best binding was predicted and determined in vitro for the motif ‘RMFPNAPYL’ (SEQ ID NO: 2).

[0165] Figure 5: Architecture and ML model training of one exemplary embodiments of the present invention. Peptide variants were BLOSUM62-encoded and processed through a through two dense layers, a ConvI D layer, and a sigmoid output ; training with 10-fold cross-validation and replicates produced ensemble models for binary binding classification.

[0166] Figure 6: Predicted antibody binding probabilities across sequence dissimilarity levels. Scatter plot of individual peptide-antibody pairs showing predicted binding probabilities (y-axis) as a function of sequence dissimilarity to the MAGE-A4 decapeptide, quantified by Hamming distance (x-axis; 1-10 substitutions). Each point represents a unique peptide differing by the indicated number of amino acids. Peptides with predicted binding probability >0.5 were classified as binders (above dashed line), whereas those <0.5 were classified as non-binders. The distribution demonstrates that predicted binders occur even among peptides with high sequence dissimilarity. A ‘SCORE’ in vitro screening validation of the predicted off-targets unrelated to target sequence was subsequently performed (results not shown).

[0167] Figure 7: Validation of Predicted Off-Target Peptide Binding in the T2 cell assay, a) T2 cells were loaded with 100 pM or 5 pM of either the wild-type decapeptide or off-target peptides predicted by the present method . To control for false negatives due to insufficient peptide presentation, peptide loading was verified by p2- microglobulin staining. Peptide PEP1990 could not be loaded even at 100 pM. b) Binding of anti- MAGE-A4 TCRms BCA0163 and BCA0165, and an isotype control antibody, to peptide-loaded T2 cells. Each triplet of bars depicts data for BCA0163 and BCA0165, and an isotype control from left to right, respectively, c) EC50dose-response titration of BCA0163 and BCA0165 on T2 cells loaded with PEP1984 and PEP0705.

[0168] Figure 8: Exemplary core workflow of one embodiments of the method according to the invention. A dataset of 533 peptides (343 ML-model based predicted off-targets and 190 X-scan variants (obtained by successively mutating one amino acid of the target peptide sequence)) was formatted as pHLA-A02:01 microarrays and profiled using SCORE. Binding data were analyzed with ANABEL (tool for determining biomolecular interaction I binding kinetics analysis), and datasets were stored in Genedata Biologies (a software platform suitable for managing biologic R&D workflows). In parallel, 316,855 predicted HLA-A02:01 binding / presented peptides were computationally generated from the human proteome (NetMHCpan rank <2 (= good to very strong binding); a presentation rank for the peptides, which is indicative, if peptides are actually presented in the binding groove of the HLA). The present ML-model(s) were trained on the experimental dataset and applied to this extended peptide space, with high-probability hits prioritized for confirmatory in vitro assays. EXAMPLES

[0169] The invention is further described by the disclosed examples. These are not intended to limit the scope of the invention but represent preferred embodiments of aspects of the invention provided for greater illustration of the invention described herein.

[0170] Example 1

[0171] The present example describes a specific embodiment of the present invention comprising using the data of an in vitro high-throughput interaction screening for the training of an Al or ML-model in the context of the present method.

[0172] Definition of a diverse pH LA set

[0173] First a pHLA set was selected. Choosing the correct pHLA set to be tested in the lab may represent the first, important step in a wet lab high throughput interaction screening. One important marker that characterizes a pHLA set may be its mutation distance to the target peptide sequence. A distance of 1 means that the peptide differs by 1 mutation from the target. The maximum number of mutations is defined by the target length. Typically, it is preferred to include all possible peptides with a distance of 1 and an equal mix of all other mutations.

[0174] In the present example, the following 4 criteria for choosing the pHLA set were applied:

[0175] 1 . The pHLA set should be as diverse as possible across all numbers of mutations and all amino acids. This will ensure that the set carries as little bias as possible. However, in other cases / applications it may be of greater importance or interest to also detect ‘likely to bind’ peptides (peptides that may be binders), instead of strictly excluding false positives.

[0176] 2. The minimal number of peptides in the pHLA set should be at least about 300 peptides, preferably between 300-5000 or more.

[0177] 3. Preferably, the chosen peptides are of physiological relevance. This means that they are commonly found in the human peptidome.

[0178] 4. Preferably, the pHLA set is composed of 50% non-binder and 50% binder to the actual therapeutic molecule. However, this can be hard to tell at this stage. Sometimes, a pHLA set needs to be enlarged or redefined after the first round of experiments.

[0179] Performance of high throughput binding experiments

[0180] Next a binding analysis of the protein of interest, in this example a therapeutic molecule, was performed against the defined set of pHLAs selected previously. A non-limiting example of such high throughput screening set up is shown in Figure 1 . Therein, in a first step different peptides (from the selected pHLA set) were spotted into unique cavities of a microcavity chip (step 1 Fig. 1 ; alternatively any multi-well plate may be used). The peptide solution was then dried out (step 2 Fig. 1) and the arrays were ready for long time storage (step 3 Fig. 1).

[0181] To form the biomolecular pH LA complexes, biotinylated HLAwas provided, here printed, into the individual cavities to enable complex formation between HLA and the peptides (step 4 Fig. 1 ). The pHLA complexes were then transferred to a streptavidin (SA) coated substrate / solid support (step 5 Fig. 1). The substrates (solid support) comprising the pHLA complexes were subsequently used for the label-free analysis of different analytes, namely proteins of interest, (step 6 Fig. 1) such that biding data, such as kinetic values, could be determined from the measured binding curves between a proteins of interest and respective pHLA complexes (step 7 Fig. 1).

[0182] The resulting binding / interaction data from such a screen included: i) kass: The association rate constant, ii) kdiss: The dissociation rate constant and / or iii) KD: The dissociation constant of each pHLA complex with a protein of interest.

[0183] Optionally, the same or similar I equivalent values (e.g., IC50 or EC50) may be obtained by ELISA, BLI or SPR, if of interest.

[0184] Training and test of the computational model

[0185] In the present example a neural net model architecture was used. However, the field of machine learning and Al is constantly evolving and other computartional models architectures are also imagined and applicable.

[0186] The input features to the model were the peptide sequences from of the selected pHLA set. The applied labels were either 1 or 0, representing binding and no binding, based on the fact that a KD value (see step 2 of Fig. 1), namely binding, could be obtained or not. The output comprised a value between 0 and 1 , for each respective combination of a peptide of interest and a protein of interest, representing the probability of expected binding or no binding.

[0187] In embodiments, the predictivity of the model may be boosted / improved by predicting more than 2 classes of output (no-binding I weaker binding compared to target I similar binding as target I stronger binding compared to target). Alternatively or in addition, further applicable input features are, e.g., structural embeddings of the pHLA complex or the antibody, physiochemical properties of the peptides and / or additional experimental data. Another option is the incorporating of large language models like, e.g., the ESM (e.g., ESM3).

[0188] In-silico experiments (performed predictions)

[0189] After the training and testing of the model, a large in silico peptide set (of peptides of interest) was defined, that consisted of a minimum of 1000 peptide sequences of interest per (mutation) distance to a specific target of the protein of interest. Thereafter, the binding of a protein of interest to all peptides of interest was predicted in silico. This dataset ultimately aided understanding the properties of the antibody which the model learned from the extensive in vitro training data.

[0190] In detail, for this experiment the used experimental interaction data was split up into i) a set to be used either as input for embedding and training the ML-model(s) (comprising binary interaction data with 1 =binding and 0= no binding), and ii) a data set of which only the sequence data of the peptides was provided to the model as peptides of interest, while the interaction data for said peptide sequences was used subsequently for evaluation I validation purposes only. The output of the in silico prediction by the trained ML-models (ESM embeddings of peptides) comprised the predicted probability of each peptide of interest regarding its binding to the protein of interest, whereby two classes (binder; 1 and non binder; 0) were determined by a threshold of a predicted binding probability of 0.5. A ROC curve analysis of the prediction accuracy of the method according to the used embodiment of the invention showed a ROC AUC of 0.952, when compared to the experimentally determined binding data. A correlation (Pearson) of the predicted and experimentally measured Log KD values for the protein of interest to peptides of interest was R= 0.82 (see Fig. 3).

[0191] When repeated for 3 further proteins of interest (BCA0094, BCA0097, BCA0099) an accuracy of between 0.892 and 0.932, a precision of 0.730 and 0.856, a ROC AUC of between 0.948 and 0.980 and an F1 of between 0.763 and 0.869 was achieved.

[0192] Moreover, in the context of the in silico predictions, also optimal (likely) binding motives of two antibodies (BCA0001 and BCA0002) were predicted and subsequently compared to in vitro experiments testing the binding affinity of the two antibodies to a large set of peptides comprising various mutations in the respective amino acid sequence over the target peptide ‘RMFPNAPYL’ (SEQ ID NO: 2). The binding motifs obtained by in silico predictions according to embodiment of the present invention revealed similar results to the in vitro binding high-throughput interaction screenings of the target peptide and its mutants.

[0193] Extraction and calculation of relevant data from predictions

[0194] From the predictions the inventors were also able to deduce an accurate binding motive of the tested protein of interest, here a therapeutic antibody, and were able to specify a value for the specificity of the protein of interest, e.g., by further evaluating potential off-target binding. This proofed to be particularly useful, since usually a multitude of candidates molecules is generated during the cause of a therapeutic development study. This information can be used to make informed decisions on which molecule, e.g., a therapeutic antibody, fragment thereof or DARPin, to proceed with into the optimization phase. It also provides information which of the candidate molecules might perform best in a potential preclinical study and clinical study, e.g., by having a high target specificity and less binding to potential off-target peptides.

[0195] The present example showed that the used ML models were able to predict binding for the analyzed specific antibody and the results surprisingly suggested that the models stay predictive even at large Hemming distances from the ideal target sequence. The inventors concluded that the present method may be successfully employed for determining detailed information, beyond experimentally obtained data, about a protein of interest.

[0196] Example 2

[0197] Methods

[0198] Identification of fully human TCR mimic anti-MAGE-A4 antibodies

[0199] Fully human TCR-mimic antibodies targeting MAGE-A4 were generated in collaboration with Biocytogen Pharmaceuticals Co., Ltd. (Beijing, China), using their RenMice-based TCR-mimic antibody platform, as previously described (M1). In this system, genetically engineered RenMice co-express a fully human antibody repertoire alongside HLA-A02:01 and were immunized with recombinant MAGE-A4 peptide-HLA-A02:01 complexes. B cells from immunized mice were subsequently screened using the Beacon on-chip platform to enable high-throughput discovery of TCR-mimic antibodies. Two highly specific anti-MAGE-A4 TCR-mimic antibodies, BCA0163 and BCA0165, were identified, demonstrating superior functional and biophysical properties.

[0200] Antibody and Fab Production

[0201] Recombinant full-length antibodies (BCA0163, BCA0165, and the isotype control Palivizumab, incorporating a human IgG 1 Fc LALA-PG mutation (M2)) and Fab fragments (BCA0097, BCA0099; C-terminal 3xMyc-1 xHis8 tag) were produced in collaboration with Biointron (Beijing, China). Heavy and light chain sequences were synthesized and cloned into pCDNA3.4 expression vectors, followed by co-transfection into CHO-K1 cells. Proteins were expressed in suspension culture for 4-6 days at 37 °C, 120 rpm, and 8% CO2. IgG antibodies were purified via Protein A affinity chromatography (MabSelect PrismA / MabSelect Sure; Cytiva Life Sciences, Marlborough, Massachusetts, USA), eluted with sodium acetate (pH 3.4), and dialyzed into PBS (pH 7.2-7.4). Fab fragments were purified using N NTA resin under standard conditions. Protein purity, monomeric state, and integrity were assessed by analytical size-exclusion chromatography (SEC).

[0202] Peptide-HLA array production and SCORE measurements

[0203] Peptide-H LA array production and SCORE measurements were performed essentially as described previously (M3). Peptides were purchased from Peptides & Elephants GmbH (Hennigsdorf, Germany), and biotinylated HLA-A*02:01 was obtained from Acrobiosystems (Basel, Switzerland).

[0204] BLI Binding Assay

[0205] BLI experiments were performed on an Octet R8 system (Sartorius). Biotinylated human HLA- A*02:01&p2M&MAGE-A4 (GVYDGREHTV; SEQ ID NO: 1) complex protein (Aero Biosystems, HLM-H82E5) was immobilized at 30 nM on Streptavidin (SA) Biosensors. Analytes and ligands were diluted in kinetics buffer (1 x DPBS, 0.1% casein, 0.05% Tween-20). Biosensors were hydrated in kinetics buffer for 10 min prior to use. Assays were conducted at 25 °C with shaking at 1000 rpm. Sensors were equilibrated in kinetics buffer for 60 s, loaded with pHLA for 200 s, followed by a 60s wash step, 120 s baseline, a 350 s association, and a 600 s dissociation phase. Three Fab concentrations (60, 20, and 6.7 nM) were tested. A pHLA-loaded sensor incubated in buffer served as a reference for subtraction. Data were analyzed using Octet Analysis Studio 13.0. Sensorgrams were aligned to the end of the baseline step, and inter-step correction was applied to match the start of the dissociation phase. Data were smoothed using Savitzky-Golay filtering and globally fitted with a 1 :1 Langmuir binding model.

[0206] T2 cell binding assay

[0207] Cultivation of T2 cells

[0208] T2 cells (DSMZ# ACC598) were cultivated at 37°C and 5% CO2 in RPMI-1640 medium (Thermo Fisher Scientific, Waltham. USA) supplemented with 10% heat-inactivated FCS.

[0209] Peptide Pulsing

[0210] For peptide pulsing, peptide solution and T2 cells were mixed in a 96-multiwell plates directly or pre-mixed in falcon tubes using AIM-V medium (Thermo Fisher Scientific). Peptide pulsing happened over night for 17-18 hours at 37°C and 5% CO2.

[0211] Preparation for Flow Cytometry and Live / Dead staining

[0212] The next day the multi-well plate was centrifuged at 700 x g for 1 min at 4 °C. The supernatant was discarded. Cells were washed once with PBS and LIVE / DEAD Fixable Violet stain (Thermo Fisher Scientific, 1 :1000 dilution in PBS) was added. The cells were resuspended and incubated on ice for 30 min with gentle shaking, protected from light. After staining, 170 pl of PBS was added per well, followed by centrifugation (1 min at 700 x g, 4 °C). The supernatant was discarded, and cells were washed twice with 200 pl of FACS buffer (PBS + 2% FCS).

[0213] Pulsing control antibody staining

[0214] For peptide pulsing control, an anti-B2M-FITC antibody (Sigma-Aldrich, Burlington, USA, #SAB4700012) or anti-B2M-AF647 antibody (Thermo Fisher Scientific, #MA5-18119) was diluted 1 :225 in FACS buffer and added to the wells. The cells were resuspended and incubated for 30 min on ice with gentle shaking.

[0215] TCRm antibody binding and detection

[0216] TCRm binding analysis was either performed in single dose measurement or full dose 8-point titration. After pulsing, cells were spun down, supernatant was removed and TCRm antibody dilutions were added. This was followed by gently resuspension and incubation on ice for 1 h and gentle shaking. After incubation, cells were washed once with FACS-buffer and secondary anti- Fc-detection antibody was added (Thermo Fisher, #A18818). This was followed by gently resuspension and incubation on ice for 1 h and gentle shaking. Subsequently, cells were washed twice with FACS-buffer. Measurement and Data analysis

[0217] Flow cytometry was performed using the Attune NxT flow cytometer (Thermo Fisher). Data analysis was performed with FlowJo (BD Life Sciences, Franklin Lakes, USA). Gating strategy for all experiments comprised pulse geometry gating and Live / Dead Gating. Compensation, if required, was done with FlowJo. Data visualization was done with GraphPad Prism (Dotmatics, Boston, USA).

[0218] Data Generation

[0219] Labeled kinetic dataset generation

[0220] To generate antibody-specific training data for the present method, the inventors employed a stepwise workflow beginning with in silico off-target prediction using an in-house computational pipeline. The present method identified candidate peptide sequences from the human proteome with up to five substitutions relative to the cognate MAGE-A4 epitope (GVYDGREHTV; SEQ ID NO: 1) filtered for predicted HLA-A*02:01 presentation and expression relevance. These candidate sequences were combined with the complete single-amino acid substitutional scan (X- scan) of the wild-type peptide, yielding a set of 533 unique peptides.

[0221] Custom peptide-HLA microarrays were generated from this panel as described above, and binding interactions with BCA0097 and BCA0099 were quantified using SCORE. Raw sensorgrams were processed with the ANABEL (R5) analysis framework, which performed automated curve fitting, rule-based quality control, and parameter extraction for kinetic association (kon), dissociation (koff), and affinity (KD). Complexes with reproducible KD values were designated as binders; all others were classified as non-binders. The resulting datasets (533 peptides per antibody) were curated and stored, establishing a labelled kinetic dataset that served as direct input for training the neural network models according to the present invention.

[0222] Unlabeled 10mer peptide dataset generation

[0223] To extend predictions beyond the experimentally profiled panel, the inventors generated a comprehensive in silico peptide space representing potential HLA-A*02:01 ligands across the human proteome. All canonical and isoform sequences from UniProt (release February 2025) were parsed with a sliding window of 10 amino acids (stride = 1), yielding a library of >50 million unique 10-mers.

[0224] This library was filtered for predicted HLA-A*02:01 presentation using NetMHCpan 4.1 (16), applying a rank threshold of <2 to select peptides with high likelihood of stable MHC binding. Redundant sequences and low-confidence predictions were removed, resulting in a final set of 316,855 unique peptides. Peptides were encoded using BLOSUM62 substitution matrices in a fixed-size tensor representation compatible with the framework of the present method. Unlike the labeled kinetic dataset, this in-silico collection was used without binding class annotations and served as the unlabeled prediction space for proteome-wide inference. Model architecture and training according to the invention

[0225] The neural network models were MultiLayer Perceptrons (MLPs) trained using python and the pytorch package. The MLP consisted of two layers with a ReLU activation in-between. A sigmoid activation was applied to the model output to produce probability scores for the binary classification task. Hyperparameters were optimized using five-fold cross validation with the binary cross-entropy loss function.

[0226] To minimize performance variance across training iterations, at each fold an ensemble of five MLPs were trained, with each model’s weights initialized to different values. The models were trained using the Adam optimizer and the One-Cycle learning rate scheduler. The full set of training parameters and final selected hyperparameters are provided in below.

[0227] The training procedure resulted in a nested ensemble of 25 MLPs, five models at each fold for five folds. During inference, the binding probability was calculated as arithmetic mean across all models (e.g., see Fig. 5).

[0228] Structure modeling with Chai-1 and MOE post-processing

[0229] Antibody-pH LA complexes for BCA0097 and BCA0099 with HLA-A*02:01 MAGE-A4 (230 239) were modeled using Chai-1 , a multimodal foundation model for biomolecular structure prediction that supports single-sequence inputs and optional constraint features (pocket / contact / docking) to encode prior knowledge about interfaces (M4, F2). The inventors ran Chai-1 in single-sequence mode, generated X predictions per complex with 10-20 recycles, and ranked outputs by pDockQ. To define a structural baseline and peptide pose, they superposed the modeled pHLA (MHC heavy chain + P2m + peptide) to the 8FJA cryo-EM structure (HLA-A*02:01 presenting MAGE-A4 230-239 bound by REGN6972) using an MHC + peptide alignment; antibody chains were not used for the alignment.

[0230] Because BCA0097 shows high CDR-L2 similarity to REGN6972 (8FJA), they prompted Chai-1 with contact restraints that mimic the 8FJA interface: E55 L2 -> R66 HLA and N53 L2 -> K69 HLA (Ca Ca maximum distance 5 A; soft penalty), and, consistent with the 8FJA peptide, the inventors encoded a putative CDR-H3 contact to the peptide arginine P6. Control runs without restraints were also executed to quantify the effect of conditioning on pose quality.

[0231] For each complex, MOE 2024.06 (Molecular Operating Environment, Chemical Computing Group ULC, Montreal, QC. Web. https: / / www.chemcomp.com / ) was used for preparation (QuickPrep: preserve sequence and neutralize; Protonate3D with ASN / GLN / HIS flips; remove waters farther than 4.5 A from either ligand or receptor; receptor tether strength 10, buffer 0.25; retain minimization restraints). Energy minimization employed AmberEHT with bonded, van der Waals, electrostatics, and restraints enabled; non-bonded cutoff on / off 8 / 10 A, reaction-field dielectric (interior 1 , exterior 80). The inventors report the minimization algorithm and RMS gradient used and treat interaction energies between antibody and pH LA (kcal / mol) as relative within this fixed setup.

[0232] Peptide backbone cp / i and sidechain x dihedrals were analyzed after MHC + peptide superposition to 8FJA to assess compatibility with the 8FJA-observed pose. Interface quality was quantified by DockQ, buried surface area, and shape complementarity; hydrogen-bond and saltbridge networks were enumerated, and per-residue interaction contributions were decomposed in MOE.

[0233] Results

[0234] Dataset construction

[0235] To generate a physiologically relevant dataset for fine-tuning Al models in TCRm-pHLA binding prediction, the inventors assembled a diverse panel of HLA-A*02:01 complexes centered on the MAGE-A4 decapeptide (230GVYDGREHTV239; comprising SEQ ID NO: 1), a well-characterized cancer-testis antigen (CTA) expressed across multiple tumor types (R1). The goal was to construct a biologically meaningful, sequence-diverse peptide set capable of yielding high- resolution kinetic interaction data for downstream machine learning (ML).

[0236] A dual strategy was employed. First, they generated a positional X-scan of the MAGE-A4 peptide, comprising all 190 single amino acid substitutions, to systematically explore TCRm sequence tolerance. Second, they applied a custom peptide-centric workflow, which scans the human proteome with a position-specific similarity scoring matrix to identify sequences related to MAGE- A4. For this study, peptides with up to four substitutions were included. Candidates were further filtered for predicted HLA-A*02:01 binding (NetMHCpan; R2) and tissue-specific expression (Human Protein Atlas; R3). This yielded 343 potential off-target peptides, which together with the 190 X-scan variants formed a dataset of 533 peptides. These peptides were then used to generate a custom peptide-HLA microarray, as previously described by Kramer et al., 2023 (R4, M3). Briefly, peptides were spotted into microcavity chips, complexed with biotinylated HLA- A*02:01 , and transferred to streptavidin-coated supports for profiling. High-throughput, label-free interaction profiling was then performed using SCORE (Single-Color Reflectometric Interference Detection) with two sequence-distinct TCRm antibodies, BCA0097 and BCA0099, both specific for the HLA-A*02:01-MAGE-A4 complex and exhibiting low-nanomolar affinities. The experimental ‘SCORE’-based workflow is summarized in Figure 1 . For each TCRm, kinetic parameters, association rate (kon), dissociation rate (koff), and equilibrium dissociation constant (KD), were measured across all 533 complexes. Complexes with measurable KD values were classified as binders; all others were designated non-binders. Data were processed and analysed using a customized version of ANABEL (F1 , R5), incorporating rule-based QC, curve fitting, and filtering. Final datasets (533 peptides per antibody) were stored in Genedata Biologies (GDB) (R6). In parallel, an unlabeled dataset was generated by systematically deriving all theoretical human proteome-encoded decapeptides (see below). This expanded peptide space was subsequently used for machine learning predictions according to an embodiment of the present method.

[0237] Neural Network Architecture and Training.

[0238] The tested embodiment of the present method comprises a supervised classifier trained independently for each TCRm antibody using SCORE-derived binding data. The model architecture is shown in Figure 5. Peptide sequences were encoded with BLOSUM62 and used as inputs to predict TCRm-pH LA binding probability. This framework enabled the identification of sequence-divergent off-target peptides, extending beyond simple motif recognition.

[0239] The neural network models were MultiLayer Perceptrons (MLPs) trained using python and the pytorch package. The MLP consisted of two layers with a ReLU activation in-between. A sigmoid activation was applied to the model output to produce probability scores for the binary classification task. Hyperparameters were optimized using five-fold cross validation with the binary cross-entropy loss function.

[0240] To minimize performance variance across training iterations, at each fold an ensemble of five MLPs were trained, with each model’s weights initialized to different values. The models were trained using the Adam optimizer and the One-Cycle learning rate scheduler. The full set of training parameters and final selected hyperparameters are provided in below.

[0241] The training procedure resulted in a nested ensemble of 25 MLPs, five models at each fold for five folds. During inference, the binding probability was calculated as arithmetic mean across all models. Proteome-Wide in Silico Screening.

[0242] To identify novel off-target pH LA complexes, the inventors then generated all possible 10-mer peptides from the human proteome using a sliding window (length = 10, stride = 1) across 34,184 UniProt protein sequences (including canonical entries and isoforms). The resulting peptide library was filtered using NetMHCpan 4.1 (R3) to retain only peptides predicted to be presented by HLA-A*02:01 , using a presentation rank threshold of <2. This yielded 316,855 unique peptides. These peptide sequences were encoded using BLOSUM62 substitution matrices and input into the trained classification models to predict TCRm binding probability.

[0243] Prediction and Experimental Validation.

[0244] The filtered peptidome, combined with the previously ML-based-predicted off-targets (with some overlap) and 190 X-scan variants, was used as input to the presently used models. For each TCRm, a second ensemble trained via 10-fold cross-validation was applied. Final predictions were obtained by averaging outputs from all models across folds. Peptides were then ranked by predicted binding probability, and top candidates were selected for experimental validation in vitro using SCORE microarrays and T2 cell-based flow cytometry assays.

[0245] TCRm-Specific Kinetic Interaction Profiling

[0246] High-throughput kinetic profiling of the two TCRms against 343 predicted off-target peptides revealed distinct kinetic interaction signatures, indicating unique off- target recognition profiles. Both antibodies exhibited favorable specificity, binding only a limited subset of the tested peptides.

[0247] Peptides bound with detectable affinity expressed as the ratio of the dissociation constant for each peptide relative to the WT MAGE-A4 peptide showed partial overlap in off-target binding between BCA0097 and BCA0099, alongside notable differences. Both antibodies bound PEP1087, derived from MAGE-A8, with affinities comparable to the WT peptide, consistent with the high sequence identity between MAGE-A8 and MAGE-A4 and their similar presentation by HLA-A*02:01 , as observed with other anti-MAGE-A4 antibodies. Similarly, both TCRms recognized other MAGE-family peptides, including MAGE-C2 (PEP1035) and MAGE-A11 (PEP0915), albeit with distinct affinity profiles.

[0248] Notably, BCA0097 uniquely bound several non-homologous peptides with low sequence similarity to MAGE-A4 such as PEP0726, PEP0921 , PEP0995, PEP1076, and PEP1148 while BCA0099 showed unique binding to PEP0999. These partially overlapping but divergent off-target profiles suggest that the two TCRms engage the MAGE-A4-pHLA complex via distinct paratope conformations and interaction modes, supported by differences in affinity, CDR sequences, and structural frameworks.

[0249] These findings were further supported by an X-scan peptide-HLA positional library microarray. A peptide panel was generated by systematically substituting each residue of the WT MAGE-A4 decapeptide with all 19 alternative amino acids (X-SCAN), and binding interactions with BCA0097 and BCA0099 were assessed. The two TCRms exhibited partially overlapping but distinct binding preferences for the MAGE-A4 peptide-HLA positional variants. Both antibodies required an arginine at position 6 for efficient binding. However, single substitutions at other positions had differential effects; for example, substitution at position 9 with positively charged residues enhanced binding of BCA0097 but had no effect on BCA0099.

[0250] Together, the kinetic and positional scanning data suggest that BCA0097 and BCA0099 possess distinct paratope characteristics, which may mediate differential recognition of additional, unseen peptide-HLA complexes.

[0251] Machine Learning-Based Prediction and Validation of TCRm-Specific Off-Target Peptides

[0252] The inventor’s objective was to identify cross-reactive off-target peptides with limited sequence similarity to the wild-type (WT) MAGE-A4 peptide. To this end, the inventors trained supervised machine learning models using the kinetic interaction profiles of the TCRm antibodies BCA0097 and BCA0099, as described above. Each TCRm-specific dataset was used independently to train a neural network employing BLOSUM62-based sequence encoding to predict the probability of TCRm-pHLA interactions.

[0253] The performance of the trained neural networks was evaluated using out-of-fold predictions from cross-validation training. Classification metrics were computed at a 50% probability threshold, with results summarized in the following table:

[0254] The models demonstrated high predictive performance, achieving F1 scores of 0.87 for BCA0097 and 0.76 for BCA099. However, the dataset consisting of peptides with few mutations (<5) likely resulted in overly optimistic performance estimates, particularly for peptides with a larger number of mutations.

[0255] The resulting models were applied to a predefined decapeptide sequence space comprising 316,855 peptides derived from the full human proteome, restricted to HLA-A*02:01 . The goal was to identify peptides predicted to exhibit binding behavior similar to that of the respective TCRm.

[0256] The employed classification model outputs a probability score reflecting the likelihood of peptide binding to a given TCRm based on the learned interaction landscape. A threshold probability of >50% was applied to extract potential off-target candidates. In total, 29 peptides were predicted as binders across both TCRms, with the majority being antibody-specific (Figure 6). Specifically, 18 peptides were predicted for BCA0097, 14 for BCA0099, and only one peptide, PEP1984, was predicted to bind both. When plotting predicted binding probability against sequence similarity (measured by the number of mismatches), several peptides with high predicted binding scores exhibited substantial sequence divergence from the WT peptide (Figure 6). Notably, the model identified peptides lacking the critical arginine at position 6, which prior X-scan analysis identified as essential for both BCA0097 and BCA0099 binding. These substitutions, lysine, asparagine, or glutamine, were not tolerated in single-position mutational scans, indicating that the model generalized beyond simple motif matching.

[0257] To experimentally validate the predicted off-targets, they performed SCORE-based microarray screening to assess kinetic interactions. For BCA0097, 8 of the 29 predicted peptides exhibited measurable binding, with dissociation constants in the high to mid-nanomolar range, consistent with specific interactions. In contrast, for BCA0099, only two peptides (PEP1984 and PEP1988) showed detectable binding among the 29 predicted candidates. Notably, these peptides were also validated for BCA0097, suggesting overlapping specificity. Sequence analysis of the validated binders highlighted the critical importance of arginine at position 6; peptides lacking this residue failed to bind, corroborating prior X-scan results.

[0258] Importantly, the present method identified peptides with up to 9 mismatches relative to the WT sequence that still showed specific binding to TCRms (PEP1988). This underscores the model's ability to capture complex, TCRm-specific binding determinants and to discover interactions with sequence-divergent, previously uncharacterized peptides highlighting its utility for off-target risk assessment and peptide discovery.

[0259] To assess peptide recognition in a cellular context, T2-cell binding assays were performed using predicted off-target peptides (Figure 7). T2 cells were pulsed with 100 pM of each peptide, a concentration sufficient for efficient loading of HLA-A*02 molecules in these TAP-deficient cells, as confirmed by increased p2-microglobulin surface expression (Figure 7a). Peptide-pulsed cells were incubated with BCA0163 (IgG-format of BCA0097) or BCA0165 (IgG-format of BCA0099), followed by staining with fluorophore-conjugated secondary antibodies. Binding was quantified by flow cytometry. All peptides previously validated in the peptide microarray also showed positive signals in the T2 assay (Figure 7b), consistent with the SCORE assay results. For PEP1984 (Figure 7c) and PEP1991 (not shown), complete EC50 dose-response curves were obtained, indicating strong and specific binding. Additional peptides recognized by BCA0163 exhibited signal increases at concentrations above 100 nM, consistent with lower-affinity interactions. Isotype control antibodies showed no detectable binding across all conditions, confirming the specificity of the observed TCRm-peptide interactions. In summary, T2-cell binding assays confirmed that machine learning-predicted peptides are recognized by TCRms in a physiologically relevant setting. These results support the conclusion that binding is not merely an artifact of recombinant pHLA complexes immobilized on microarrays.

[0260] Notably, the present method identified antibody-specific off-target peptides with minimal or no sequence homology to MAGE-A4 that are presented by HLA-A*02:01 in a cellular context, reinforcing its value for off-target prediction beyond conventional peptide-centric approaches. Structural Characterization of Machine Learning-Predicted Peptide-TCRm Interactions

[0261] To investigate off-target peptides predicted by the present method, the inventors examined structural binding preferences of TCRm antibodies BCA0097 and BCA0099. Because the present metho is able to map peptide sequences to antibody-specific binding probabilities, its predictions likely encode latent features of the paratope-epitope interaction landscape.

[0262] Baseline structural models of both antibodies with the WT MAGE-A4 peptide established their shared TCR-like binding mode: CDR-L3 and CDR-H3 dominated peptide contacts, while other CDR loops anchored the HLA helices. Both antibodies required Arg6 for binding yet adopted distinct paratope conformations. BCA0097 recognition involved CDR-L3 N93 to T9 and CDR-H1 N32 / K33 to D4, reinforced by ionic interactions of CDR-H2 E51 and CDR-H3 D99 with R6. By contrast, BCA0099 relied on a dominant CDR-H3 E99-R6 salt bridge, with supporting hydrogen bonds from CDR-H3 N102 and CDR-H2 S50 / S53.

[0263] Consistent with their distinct engagement modes, BCA0097 recognized six predicted peptides, whereas BCA0099 bound only PEP1984 and PEP1988, which were also validated for BCA0097. Despite substantial sequence divergence from MAGE-A4 (Hamming distances: 8 for PEP1984 and 9 for PEP1988) and strong differences between each other (Hamming distance PEP1984 vs. PEP1988: 8), both peptides adopted similar backbone conformations, exposing the conserved R6. This structural mimicry explains the observed shared cross-reactivity. For example, in PEP1984, BCA0097 retained strong ionic interactions with R6, complemented by hydrogen bonds to E4 and S9, whereas BCA0099 preserved its hallmark CDR-H3 E99-R6 salt bridge along with hydrogen bonds to F3, G5, and S7. In contrast, peptides lacking an exposed arginine for example PEP2004 (R6->K) failed to bind despite ML prediction. Structural modelling showed backbone displacement (>7 A at G5), burying Lys6 too deep in the groove for BCA0097 CDR-H2 E51 and BCA0099 CDR-H3 D99 to engage, leaving only a weak H1-D4 contact (not shown). Similarly, PEP1993 maintained R6 but diverged structurally at P3-P7; absence of Y3 prevented R6 stabilization, while steric clashes at W3 / E4 precluded binding (not shown).

[0264] Together, these analyses identify Arg6 as a central determinant of TCRm binding and highlight how structure-based modelling contextualizes ML predictions. Importantly, the present method was able to uncover off-target peptides with little sequence similarity yet structurally compatible features, demonstrating its utility for predicting paratope-epitope interactions beyond sequence homology. These insights may further guide optimization of TCRm specificity and inform safety assessments in therapeutic development.

[0265] Discussion

[0266] Comprehensive identification of potential off-target peptide-H LA complexes is a critical component of early-stage development for T cell receptor- based cancer therapeutics, particularly TCR-mimic antibodies, as unintended cross-reactivity can lead to dose-limiting toxicities, irreversible organ damage, and patient mortality. Although several preclinical screening technologies have been developed to characterize off-target reactivity, each is limited by inherent methodological constraints that affect sensitivity, specificity, or translational relevance. To enable TCRm-centric, in silico, proteome-wide off-target screening, the inventors developed the present method, which comprises in embodiments a machine learning framework trained on high-throughput kinetic data from X-scan mutagenesis and selected off-targets exhibiting up to four residue substitutions relative to the cognate epitope. The herein used embodiment of the present method learned an antibody-specific quantitative mapping from peptide sequence space to binding affinity, facilitating the prediction of pH LA interactions beyond sequence homology. This enabled identification of structurally and biophysically similar, yet sequence-dissimilar, off-targets across the human proteome. In the present study, the method was able to identify 29 candidate off-target peptides for two unrelated TCRm antibodies recognizing a MAG E-A4-de rived decapeptide, including variants with up to ten amino acid substitutions relative to the cognate epitope. Of these, 28 peptides successfully loaded onto T2 cells, supporting their compatibility with HLA-A*02:01 and underscoring the predictive accuracy of the algorithm. Nine peptides were further validated experimentally using in vitro ‘SCORE’-based kinetic assays and T2 cell-based binding assays. Among these, PEP1984 and PEP1988 were identified as shared binders for both antibodies, highlighting their potential as cross-reactive epitopes. Structural modelling revealed that, despite up to nine amino acid differences from the target peptide, both peptides adopted highly similar conformations when bound to HLA.

[0267] These results suggest that the present method is preferably able to encode latent structural information critical for TCRm-pHLA recognition, enabling the identification of structurally mimetic, sequence-dissimilar off-targets.

[0268] In the present example, the inventors applied an embodiment of the present method to two distinct T-cell receptor mimic (TCRm) antibodies targeting the cancer-testis antigen MAGE-A4. The model successfully predicted multiple off-targets with minimal sequence similarity to the intended epitope, many of which were experimentally validated via T2 cell binding assays. Interestingly, predicted off-target peptides were largely unique to each antibody, revealing crossreactivity patterns not evident from the target sequence alone. These findings evidence the present method to be a valuable tool for lead optimization of TCRms, enabling the identification of antibody-specific off-targets beyond the scope of traditional peptide-centric methods and supporting the preclinical de-risking of TCRm-based therapies.

[0269] In summary, the presently tested embodiment of the invention comprised a machine learning framework, which effectively identified target-unrelated and sequence-dissimilar off-target peptides in a TCRm-centric context. Importantly, predictions were experimentally validated as true binders using both SCORE-based interaction assays and T2-cell binding assays. The validated binders adopted conformations similar to the intended target when presented by HLA, suggesting that the present method is able to capture structurally relevant features despite relying primarily on sequence information. These findings highlight the present method as a valuable tool for early-stage assessment of off-target risk across the peptidome, extending beyond traditional sequence-similarity approaches. REFERENCES

[0270] Moritz, A. et al. High-throughput peptide-MHC complex generation and kinetic screenings of TCRs with peptide-receptive HLA-A* 02: 01 molecules. Sci. Immunol.4, eaav0860 (2019).

Claims

CLAIMS1 . A computer-implemented method for predicting the binding of at least one protein of interest to at least one peptide of interest comprised within a HLA-peptide (pH LA) complex of interest, comprising a. providing sequence and / or structural data of at least one reference peptide, optionally providing sequence and / or structural data of at least one HLA molecule and / or the at least one reference peptide presented within a HLA-peptide (pHLA) complex, b. providing binding data of the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide (pHLA) complex with at least one protein of interest, c. training a machine learning (ML)-based model or artificial intelligence (Al) based on the data provided in a. and the binding data provided in b., d. providing sequence and / or structural data of at least one peptide of interest, e. employing the machine learning (ML)-based model or artificial intelligence (Al) trained in c. to predict the binding or interaction properties of the at least one peptide of interest, preferably when comprised within a pHLA complex, to the at least one protein of interest.

2. The method according to claim 1 , wherein the binding data of the at least one reference peptide and / or the at least one reference peptide presented within an HLA-peptide (pHLA) complex provided in b. has been determined experimentally.

3. The method according to claim 1 or 2, wherein the at least one protein of interest is at least one receptor, antibody, or antibody fragment or equivalent thereof, or designed ankyrin repeat protein (DARPin) of interest.

4. The method according to claim 2, wherein the experimental determination of the binding data provided in b. comprises the following method steps: i. providing at least one reference peptide, ii. supplying HLA-molecules, preferably biotinylated HLA-molecules, to the at least one reference peptide and allowing the HLA-molecules to bind to the at least one reference peptide, thereby forming pHLA complexes, iii. immobilizing the formed pHLA complexes on a solid support, preferably on a streptavidin-coated solid support, iv. contacting the immobilized pHLA complexes with at least one protein of interest,v. determining binding data of the individual pHLA complexes with the at least one protein of interest.

5. The method according to any one of the preceding claims, wherein predicted binding or interaction properties in step e. comprise the binding strength, preferably comprising the dissociation constant (KD), and / or the probability or likelihood of binding of the at least one protein of interest to the at least one peptide of interest, preferably when presented within a pHLA complex.

6. The method according to any one of the preceding claims, wherein the at least one peptide of interest comprises at least one mutation with respect to the at least one reference peptide.

7. The method according to the preceding claim, wherein the at least one peptide comprises two or more mutations, such that information about synergistic effects may be obtained.

8. The method according to any one of the preceding claims, wherein in step a. also data on physiochemical properties of the at least one reference peptide, and / or sequence and / or structural data of the at least one protein of interest and / or of the pHLA complex comprising the least one reference peptide are provided.

9. The method according to the preceding claim, wherein the data provided in a. and / or d. has been determined experimentally and / or has been predicted by one or more machine learning (ML)-based model or artificial intelligence (Al).

10. The method according to any one of the preceding claims, wherein step e. further comprises: comparing the binding or interaction properties of the at least one protein interest toA) the at least one reference peptide and / or the at least one reference peptide presented within a HLA-peptide (pHLA) complex, withB) the least one peptide of interest, preferably when presented within a HLA-peptide(pHLA) complex, predicted in e.11 . The method according to any one of the preceding claims, wherein step e. further comprises predicting the amino acid sequence of a target binding motive of the at least one protein of interest, preferably wherein the target binding motive is predicted to have the highest binding affinity of all possible amino acid sequences to the at least one protein of interest.

12. The method according to claim 11 , wherein the predicted amino acid sequence of the target binding motive is subsequently used to determine and / or compare one or more peptides of interest potentially showing undesired or off target binding to the at least one protein of interest.

13. The method according to any one of the preceding claims, wherein sequence and / or structural data comprises information on the amino acid sequence, the nucleic acid sequence encoding the respective peptide or protein, the three-dimensional structure, the two-dimensional structure, the SMILES string, the InChi, the InChlKey, the chemical structure, any other biochemical property of a biomolecule, and / or other molecule features, optionally wherein any of the afore were predicted by one or more ML model or Al.

14. The method according to any one of the preceding claims, wherein the HLA-peptide (pHLA) complex in which a reference peptide is presented and / or for which sequence and / or structural data is provided is either present in the human immunopeptidome or is a synthetic pH LA complex.

15. The method according to any one of the preceding claims, wherein the protein of interest is an antibody or antibody fragment, preferably a TCR-like antibody (TCRm) or fragment thereof, or a receptor, preferably a T cell receptor or fragment thereof.

Citation Information

Cited By

  • Bidirectional reversible conversion method and system between peptide molecule SMILES and sequence expression

    CN121999858A