The generative design of peptide MHC (PMHC) binders
The generative protein design system using diffusion modeling and MPNN generates pMHC binders with high affinity and specificity, addressing the limitations of current TCR-like antibody methods by reducing off-target interactions and improving therapeutic efficacy.
Patent Information
- Application Number
- PCT/EP2025/071485
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-22
- Filing Date
- 2025-07-25
- Publication Date
- 2026-02-19
AI Technical Summary
Current methods for generating TCR-like antibodies or TCR-mimetics for targeting pMHC complexes in diseases such as cancers, immune disorders, and viral infections are labor-intensive, have low throughput, and suffer from inherent TCR cross-reactivity, low affinity, and on-target/off-tumor toxicities.
A generative protein design system using diffusion modeling and message passing neural networks (MPNN) to generate amino acid sequences for pMHC binders with high affinity and specificity, followed by in silico filtering and optimization to reduce off-target binding and cross-reactivity, and biosynthetic production of these binders.
The method produces pMHC binders with high specificity and affinity for target pMHC complexes, reducing off-target interactions and enhancing therapeutic efficacy.
Smart Images

Figure EP2025071485_19022026_PF_FP_ABST
Abstract
Description
[0001] THE GENERATIVE DESIGN OF PEPTIDE MHC (PMHC) BINDERS
[0002] FIELD
[0003] The present invention relates to the field of generative protein binder design, specifically peptide major histocompatibility complex (pMHC) binder design, using methods and computer implemented methods which relies on specifically trained neural networks using e.g., diffusion modelling and message passing neural networks (MPNN).
[0004] BACKGROUND
[0005] In the last decade the emergence of machine learning based protein structure prediction tools, like AlphaFold (AF), has revolutionized the field of structural biology, and computational protein structure prediction. Further advances in protein structure prediction have led to the emergence of further tools that enable not only protein structure prediction but also allow for de novo engineering of proteins with tailored functions, such as particular functionalities or binding characteristics.
[0006] With recent advances in the field of protein binder prediction, in particular with the emergence of MPNN, it has become possible to translate theoretical 3-dimensional envelopes into amino acid sequences that fit within these 3-dimensional structures, essentially providing a tool for generation of de novo protein sequences. These advances, together with the progress made within computational protein folding predictions, as evidenced by e.g., AlphaFold, open new and promising avenues for targeted approaches to de novo protein design.
[0007] Despite the tools for protein structure prediction and de novo protein design being now readily available, the diverse nature of proteins still requires specific algorithms to be trained for the particular purpose in hand in order to obtain the most reliable and useful results. Such training in itself remains insufficient, as the sheer possible variations in protein structures provide such a big data set, that in order for the algorithms to arrive at a meaningful conclusion within a meaningful timespan, the machine learning algorithms must further be guided by outside input.
[0008] Major histocompatibility complexes (MHC) are responsible for the presentation of short peptides on the surface of cells through direct interaction between the peptide and the MHC, thus forming a peptide-MHC complex (pMHC). MHC class I is expressed on all nucleated cells and presents 8-11-mer peptides derived from cytosolic proteins. MHC class II is expressed on antigen-presenting cells and presents 13-25 peptides derived from extracellular proteins, mainly derived from pathogens. Accordingly, pMHC complexes are a very diverse group of protein complexes. pMHCs are engaged by T cell receptors (TCRs) expressed on T cells. If the TCR recognizes its cognate pMHC, T cell activation may occur. MHC class I is bound by the pMHC-specific CD8+T cells and induces CD8-dependent cytotoxicity towards the pMHC-expressing cells, and MHC class II is bound by CD4+T cells, most commonly resulting in cytokine production by the CD4+T cells.
[0009] Due to the limited repertoire of TCRs, TCRs are inherently cross-reactive. This means that a single TCR can recognize and bind to multiple distinct pMHC complexes. This phenomenon arises from the inherent flexibility and degeneracy of the TCR's binding site, allowing it to interact with a diverse array of antigens.
[0010] The ability of TCRs to target disease related intracellular antigens has been explored in autoimmune diseases, in viral infections and in cancer immunotherapies (e.g. TCR- engineered T cell therapy), but their use has been hindered by inherent TCR cross-reactivity, low affinity for pMHC targets, low functional avidity and on-target / off-tumor toxicities. Unless substantially affinity optimized, TCRs will only function in a cell associated (i.e. multimerized) structure.
[0011] To overcome this problem, TCR-like antibodies or TCR-mimetics have been extensively explored as an alternative to TCR-based therapeutics. However, their generation mainly relies on labour-intensive and lengthy procedures like hybridoma technology. In addition, current pipelines for TCR-like antibodies requiring in vivo immunization have a low throughput exploration of cross-reactivity. These limitations call for alternative and innovative solutions to generate therapeutics able to bind pMHC with high affinity and specificity.
[0012] The present disclosure aims to overcome the common issues by a generative binder design approach, which allows for generation of high affinity and high specificity binders towards specific pMHC complexes implicated in diseases such as cancers, immune disorders, autoimmune targets, and viral infections.
[0013] SUMMARY
[0014] In a first aspect the present invention relates to a method for generating an amino acid sequence of a high affinity and high specificity pMHC binder candidate, wherein said method comprises a) providing a protein structure or structural model of a target pMHC and inputting the protein structure or structural model into a generative protein design system, i) generating a first set of putative pMHC binders targeting said target pMHC using generative protein design, ii) selecting from the first set of putative pMHC binders, all putative pMHC binders having both a pLDDT score of >85% and an interaction pAE (ipAE) if <12A, thereby generating a second set of putative pMHC binders, ill) optimizing the second putative set of pMHC binders in silico b) Selecting / Recovering only the optimised putative pMHC binders having a pLDDT of >90% and an ipAE of <7A, and c) outputting the amino acid sequences of the putative pMHC binders selected in step b), thereby providing amino acid sequences encoding pMHC binder candidates having high affinity and high specificity toward a target pMHC.
[0015] In a second aspect the present invention relates to a method for biosynthetic and / or synthetic production of a pMHC binder candidate comprising a) Obtaining a pMHC binder candidate amino acid sequence using the method as defined herein, b) Producing said pMHC binder candidate of said amino acid sequence using a recombinant expression system, protein synthesis, or a combination thereof, and optionally c) Purifying said pMHC binder candidate.
[0016] In a third aspect the present invention relates to a pMHC binder candidate obtainable from the methods disclosed herein.
[0017] In a fourth aspect the present invention relates to a nucleic acid expression vector comprising a nucleic acid sequence encoding a polypeptide construct encoding a) a pMHC binder candidate obtained according to the method disclosed herein, b) a polypeptide membrane anchor, c) a linker connecting said pMHC binder candidate and said membrane anchor, and d) a surface expression tag, and wherein said expression construct is under control of an expression element such as a promoter.
[0018] In a fifth aspect the present invention relates to a pMHC binder candidate defined herein or a nucleic acid expression vector as defined herein for use in the treatment of cancer.
[0019] BRIEF DESCRIPTION OF FIGURES
[0020] FIGURE 1
[0021] RFdiffusion-guided pMHC binder design campaign against SLLMWITQC / HLA-A*02:01 complex. Example of a design campaign involving an initial cycle RFdiffusion & dl- binder_design and a subsequent second cycle of dl-binder_design with a higher number of sequences in the ProteinMPNN step for previously filtered designs, a) Flow chart showing design campaign for designing pMHC binders for SLLMWITQC / HLA-A*02:01. b) ipAE and binder pLDDT of the 22000 initial RFdiffusion generated binder designs for the SLLMWITQC / HLA-A*02:01 complex (PDB: 2bnr) with all peptide residues (chain C) as ‘hotspot’ residues, c) Zoom in of Fig. 1a showing pMHC binders with an ipAE <12 A and binder pLDDT >88% filtered for unique initial RFdiffusion designs to maximize structural diversity (step not shown in graph), d) Scored designs after sequence diversification for the initially selected 109 designs combined (10900 designs), e) Zoom in of fig. 1 c within the cut-off with an ipAE <7 A and binder pLDDT >88% filtered for unique initial structural design.
[0022] FIGURE 2
[0023] Partial diffusion design sub-campaign for pMHC binder (3p9l_beta_429_dl_design_1 , SEQ ID NO: 377) of the campaign against murine pMHC SIINFEKL / H2-Kb. a) Steps of partial diffusion were highlighted in dark grey within the RFdiffusion schedule with refinement options listed as bullet points, b) Visualization of structural prediction of 3p9l_beta_429_dl_design_1 in complex with a cropped version of SIINFEKL / H2-kb comprised of positions A1-177 & F1-8 in the PDB ID 3p9l) included because of its uniquely high content of beta sheets as secondary structure elements, c) ipAE and total pLDDT of 300 partial diffused designs (sampling temperature: default, diffusion time steps: 20, structural designs: 100, sequences per structural design: 3). d) Zoom in on the structural visualization of the best scoring (with respect to ipAE) partially diffused binder. FIGURE S
[0024] Computational screening for potential cross-reactive binding of SLLMWITQC / HLA-A*02:01 binder against variants of the SLLMWITQC presented on HLA-A*02:01 . a) total count of unique matching (identity >85%) sequences in the human proteome to 3610 single & double mutated peptides (of SLLMWITQC) utilizing the web-based tool of NCBI pBLAST followed by filtering in NetMHCpan4.1 (accessible via DTU Health Tech Bioinformatics Services via https: / / services.healthtech.dtu.dk / services / NetMHCpan-4.1 / ). ‘SLLGNILRI’ obtained a ranking score within the highest 0.5-percentile and was considered as a peptide likely to be represented on the same HLA-A*02:01. b) Sequence logo displaying sequence variation of the 1128 single and double mutated peptides (out of 3610) which obtained better or equal predicted affinity towards HLA-A*02:01 as SLLMWITQC. c) Visualization of the pMHC binders against SLLMWITQC / HLA-A*02:01 passing the cross-reactivity screening with an ipAE of >25 A and <15 A for off-target and original target peptide(s) presented on the same HLA-A*02:01 modelled using Colabfold (Mirdita M. et al., ColabFold: making protein folding accessible to all, Nature Methods, 2022) and scored employing AF2 of the dl-binder design.
[0025] FIGURE 4
[0026] Multi-objective Bayesian optimization (MOBO) for selective binding of pMHC binders against SLLMWITQC / HLA-A*02:01. a) 25 (depicted as gray solid lines) of the previously selected 77 pMHC binders obtained low ipAE scores (<10 A) for the target SLLMWITQC / HLA-A*02:01 as well as single point-mutants (‘SLLAWITQC’ and ‘SLLMYITQC’) presented in HLA- A*02:01. b) The 25 binders with cross-reactivity underwent 8 rounds of MOBO. highlighted as dark lines are the best specificity matured binders that were able to retain a low ipAE with the target pMHC, while having high ipAE scores for the non-target pMHCs.
[0027] FIGURE 5
[0028] Computational screening for HLA-specific recognition of selected pMHC binder designed against RVTDESILSY / HLA-A*01 :01. Anchor residues were substituted (highlighted in bold) based on the stochastic frequency of amino acids at given positions for each of 12 common HLA subtypes (HLA-A*01:01, HLA-A*02:01, HLA-A*03:01, HLA-A*24:02, HLA-A*26:01, HLA-B*07:02, HLA-B*08:01 , HLA-B*27:05, HLA-B*39:01 , HLA-B*40:01 , HLA-B*15:01 (SEQ ID NO: 1-14 / 538-555)) and the pMHC binders were screened against these. None of the curated peptides were predicted by netMHCpan4.1 (with a cut-off el_rank of <2%) to bind on HLA-A*26:01 and HLA-B*39:01 within the applied cut-off, and thus, these HLA complexes were not included in the subsequent analysis. The 30 best scoring RVTDESILSY / HLA- A*01 :01 binders (SEQ ID NO 296-325) were re-aligned (using an inhouse pymol-based script) and modelled using AF3 for binding to each HLA-peptide pair. The plot shows the distribution of the ipAE of the pMHC binder-variant pMHC complexes (b-vMHC) sorted by vMHC as violin plots, whereby the thickness at a given y-axis value marks the relative abundance (using the Kernel Density Estimation) of respective ipAE scores obtained for all b-vMHC for each vMHC. White line: median of the distribution, thicker black bar: interquartile range, thinner black line: upper and lower adjacent values.
[0029] FIGURE 6
[0030] MHC-binder in vitro screening pipeline. Overview of pipeline for pMHC binder discovery using mammalian display on CD3-deficient Jurkat cells. CD3-deficient Jurkat cells transduced with libraries of membrane-anchored binder-encoding lentiviruses are stained with pMHC multimers with the target of interest in one fluorophore (fluorophore 1 ), and a library of pMHC multimers that are not supposed to be bound by the binders in another fluorophore (fluorophore 2). The binders that are positive for fluorophore 1 and negative for fluorophore 2 are sorted by FACS, and the lentivirus-integration is sequenced to identify the binders with high specificity.
[0031] FIGURE 7
[0032] Identification of SLLMWITQC / HLA-A*02:01 binders. Figure 7a depicts the representative flow plots of mammalian display-based expression of pMHC binders against SLLMWITQC / HLA-A*02:01 in CD3-deficient Jurkat cells. CD3-deficient Jurkat cells transduced with a library of 50 binder-encoding lentiviruses (top) or the empty backbone (bottom) were stained with SLLMWITQC / HLA-A*02:01 pMHC-multimers labelled with either PE and APC. Double positive cells were sorted by FACS and the lentivirus-integration was sequenced to identify the true binders. One minibinder sequence was enriched and identified as true binder in vitro (2bnr-27_7_3_9, SEQ ID NO: 341), as seen in the Log2FC plot in Figure 7b.
[0033] FIGURE S
[0034] Identification of RVTDESILSY / HLA-A*01 :01 binders. Figure 8a depicts sorting of pMHC binder-displaying CD3 KO Jurkat cells based on fluorescence signal resulting from the binding of the RVTDESILSY / HLA-A*01:01 pMHC-tetramers labelled with PE and APC. Double-positive cells were sorted by FACS and the lentivirus-integration was sequenced to identify the true binders. Figure 8b depicts that two pMHC binder sequences were significantly enriched and identified as true binder in vitro (AKAP9_binder_15 & AKAP9_binder_3_renumb_25_dldesign_0_af2pred, SEQ ID NO: 302 & SEQ ID NO: 665, respectively).
[0035] FIGURE 9
[0036] Figure 9 shows 2bnr-27_7_3_9 binds the WT peptide and some peptide variants. Fig 9A) 2bnr-27_7_3_9 expressing CD3 KO Jurkat cells stained with SLLMWITQC / HLA-A*02:01- APC tetramers and evaluated by flow cytometry. Fig 9B) 2bnr-27_7_3_9 expressing CD3 KO Jurkat cells double-stained with variant peptide / HLA-A2*02:01 PE tetramers (peptide above plots) and WT tetramers in APC and evaluated by flow cytometry. Fig 9C) 2bnr- 27_7_3_9 expressing CD3 KO Jurkat cells single-stained with variant peptide / HLA-A2*02:01 PE tetramers (peptide above plots) and evaluated by flow cytometry. In Fig 9B-C, cells were gated on CD3“ GFP+. Fig 9D) Barplot showing gMFI of variant peptide and WT peptide tetramers.
[0037] FIGURE 10
[0038] Figure 10 shows Cell-based killing assay for 2bnr-27_7_3_9-miBd CAR (SEQ ID NO: 76). T cells expressing the miBd-CAR induced rapid cell death of the A375 cell line (SLLMWITQC / HLA-A*02:01+) compared to the non-transduced control (UTD) at both 1 :1 and 0,5:1 effectontarget ratio.
[0039] FIGURE 11
[0040] Identification of SLLMWITQC / HLA-A*02:01 binders from diversification of pMHC binder 2bnr-27_7_3_9. A) Representative flow plots of mammalian display-based expression of pMHC binders against SLLMWITQC / HLA-A*02:01 in CD3 KO Jurkat cells. Cells were stained with PE- and APC- SLLMWITQC / HLA-A*02:01 pMHC-tetramers. Double positive and negative populations were sorted, and lentivirus-integration was sequenced to identify the true binders. B) Log2FC plot of sequencing data comparing double positive population vs negative. 40 binder sequences (triangles) were enriched (cut-off >1). C) Barplot depicting geometric Median Fluorescent Intensity (gMFI) of 32 pMHC binders with Log2FC>1 , individually cloned in CD3 KO Jurkat cells and stained with PE-Tetramer. The darker bar represents 2bnr-27_7_3_9, and UTD= untransduced control. DETAILED DESCRIPTION
[0041] Generative protein binder design
[0042] Generative protein design relates to the generation of proteins using a de novo protein generation algorithm, wherein protein structures are generated either purely de novo, or using a set of predefined parameters, such as protein length, amino acid sequence, volume, surface area or similar metrics. The novel approaches provided herein relates to the generation of protein binders, i.e., a proteinogenic molecules designed to bind to another protein through various protein: protein interactions.
[0043] The present disclosure describes how the amino acid sequence of polypeptides that binds to specific peptide presenting MHCs (pMHCs), herein referred to pMHC binders, may be generated using in silico generative protein design, based on actual or modelled pMHC structures and / or pMHC:TCR complexes. Example 1 provides several methods for generating the pMHC binders, using available methods such as e.g., RFdiffusion and ProteinMPNN.
[0044] Additionally, as unspecific binding for such pMHC binders is highly undesirable, as it may lead to side effects in cases where the pMHC binders are used in therapy. The present disclosure further describes how such pMHC binders that are more likely to cause unspecific binding can be filtered out. As a result, only pMHC binder that are more likely to have high specificity and affinity towards a target pMHC without unspecific binding are selected as candidate pMHCs for further testing and use.
[0045] Thus, in a first aspect, the present disclosure relates to a method for generating amino acid sequences encoding pMHC binder candidates having high affinity and high specificity toward a target pMHC.
[0046] In embodiments, the method for generating amino acid sequences encoding pMHC binder candidates having high affinity and high specificity toward a target pMHC comprises a) providing a protein structure or structural model of a target pMHC and inputting the protein structure or structural model into a generative protein design system, i) generating a first set of putative pMHC binders targeting said target pMHC using generative protein design, ii) selecting from the first set of putative pMHC binders, all putative pMHC binders having both a pLDDT score of >85% and an ipAE if >12A, thereby generating a second set of putative pMHC binders, ill) optimizing the second putative set of pMHC binders in silico b) Selecting / Recovering only the optimised putative pMHC binders having a pLDDT of >90% and an ipAE of <7A, and c) outputting the amino acid sequences of the putative pMHC binders selected in step b), thereby providing amino acid sequences encoding pMHC binder candidates having high affinity and high specificity toward a target pMHC.
[0047] The amino acid sequences encoding pMHC binder candidates may be further optimised in step iii) to ensure high target specificity and reduce off target binding and / or cross reactivity by subjecting the second set of putative pMHC binders to in silico testing of the risk for off target binding variant peptides and MHC complexes. Following such in silico testing only putative pMHC binders that show no indication of cross reactivity and / or off-target binding are selected for downstream testing.
[0048] Additionally, the binding specificity of putative pMHC candidates selected in step iii) may be further optimized by bayesian optimisation using the AF2 metric of ipAE as proxy for binding likelihood. In further embodiments, the method disclosed herein for generating amino acid sequences encoding pMHC binder candidates, is a computer implemented method.
[0049] Providing a protein structure or structural model of a pMHC
[0050] In the present context, providing a protein structure or structural model relates to the process of choosing a target pMHC and preparing the structural information about a given target pMHC that is necessary for a generative protein design system to proceed with generating amino acid sequences, which "encode" putative pMHC binders capable of binding the target pMHC.
[0051] In cases where the structure of a target pMHC have been solved and published at the protein databank (PDB, https: / / www.rcsb.org), these structures are cleaned up and used as a protein structure input for the generative protein design system. In cases where the available structures comprise one or more repetitions of a single pMHC, e.g., due to a repetitive crystal lattice, all the repetition besides one are removed so that only the single pMHC structure remains. Additionally, the MHC domains and subunits that are not involved in the formation of the peptide binding groove are removed to reduce the inference time of RFdiffusion during the generation step. For class I MHC crystal structures this means removal of the a3 domain, and P2 microglobulin subunit. For class II MHC crystal structures, this means removal of the a2 and P2 domains.
[0052] Alternatively, when no crystal structures are available for a given target pMHC, AlphaFold 2 (AF2), or a variant thereof, is used to generate a protein structure of the target pMHC, which is used as a structure model input for the generative protein design system. The structure of the pMHC generated by AF2 may either be generated based on all the extracellular domains and subunits of the target MHC and the target peptide bound to it or it may be generated based on only the peptide binding domains of target MHC and the target peptide. When the structure generated by AF2 is based on all the extracellular domains and subunits of the target MHC and target peptide, the structure is cleaned up prior to use, as described herein, by deletion of the subunits and domains that do not form the binding groove of the MHC.
[0053] The cleaned-up protein structures and or structural models of target pMHCs that are provided as discussed above are then used as input in step a).
[0054] Generating a first set of putative pMHC binders (step i)
[0055] The generation of the first set of putative pMHC binders in step i) is performed by generative protein design using diffusion modelling to generate a first set of putative pMHC binders. The first set of putative pMHC binders are generated by the diffusion modelling system based on the input provided in step a) and a set of additional criteria and / or limitations, which includes at least a definition of a desired protein length for the putative pMHC binders. Diffusion modelling systems that are of particular interest in the present context include RFdiffusion and may comprise further components such as e.g. message passing neural networks (MPNN). The generation of the first set of putative pMHC binders is performed in silico. In particular, the generative protein design is performed using sequential use of RFdiffusion and a message passing neural network (MPNN). Such sequential use may include application of RFdiffusion, followed by application of a message passing neural network (MPNN).
[0056] As used herein Diffusion modelling relates to RFdiffusion and its succeeding versions (including specifically RoseTTAFold All-Atom (Rohith Krishna et al., Generalized biomolecular modeling and design with RoseTTAFold All-Atom. Science Vol 384, 2024)), an open-source method for structure generation, published first in December 2022 by Baker et al., which is capable of addressing several aspects of protein design. Among these are motif scaffolding (generating a protein structure around a given (short) motif provided as amino acid sequence), symmetric multimer design, binder design to a provided target structure (in .pdb format) and a diversification of a given structure by ‘sampling’ similar protein structures after corruption and de-convolution of Gaussian noise (referred to as partial diffusion). In the application of pMHC binder design RFdiffusion is only used for the design of binding structures to the target pMHC as well as the generation of similar / slightly altered structures in partial diffusion.
[0057] On a more profound level the RoseTTAFold (RF) diffusion model uses the RF frame representation comprised of the Carbon alpha coordinates and N-Ca-C rigid orientation for each residue. The model was trained by Baker et al. utilising the public available structures Protein Data Bank (PDB), which were noised by perturbation of Calpha coordinates with threedimensional Gaussian noise (for translational dimensions) and simulated Brownian motion on the rational matrices of the backbone orientation, N-Ca-C. The noise was applied in a stepwise process with a random number of steps (up to 200), referred to as time steps, and the learning process of denoising scored for minimising a mean squared error (MSE) function between the ground truth and the denoised structure. To generate a new protein backbone, a random residue frame of specified length is initialised and RFdiffusion provides a denoised prediction. At the start the residue frame may allow manifold shapes and denoising trajectories which narrows down with each denoising prediction iteration. A new prediction iteration is initiated by the addition of Gaussian noise which is defined in direction and strength by the sampling temperature, the number of denoising time steps, and specific commands for scaling all defined by the end user (e.g., denoiser.noise_scale_ca=0, denoiser.noise_scale_frame=0).
[0058] As such the diffusion modelling is completed by the generation of a three dimensional "density map" representing the putative pMHC binder, which is populated by placeholder glycine residues. This "density map" is generally referred to herein as "backbone(s)" or "backbone structure(s)".
[0059] In interesting embodiments, the input provided in step a) is processed by RFdiffusion together with a set of additional criteria and / or limitations, to produce a structure backbone representing the structure conformation of a putative pMHC binder. This backbone is then introduced into a message passing neural network (MPNN), such as ProteinMPNN or ESM- IF, which is then used to generate a desired number of amino acid sequences that fit the backbone, thereby generating a first set of putative pMHC binders as defined in step i).
[0060] Accordingly, once the backbones have been generated by diffusion modelling, e.g. using hotspot-based RFdiffusion as described herein, about 1-10 sequences may be designed for each of the backbones using a message passing neural network (MPNN). The MPNN analyses the backbone structures and translate and / or populate the placeholder sequences with actual amino acid sequences that fits the boundaries of the backbone structures.
[0061] One option for performing this operation is to use ProteinMPNN (Bennett, N. R. et al., Improving de novo protein binder design with deep learning, Nat. Common. 14, 2625 (2023)) as the MPNN). In ProteinMPNN, the sampling temperature is typically set to 0.1 but often varied and is typically between 1 E-5 to 1 E-1 . Further, finetuned ProteinMPNNs, e.g., for immune evasion, or other similar sequence design approaches such as ESM-IF could be used (Hsu, Chloe et al., Learning inverse folding from millions of predicted structures, Proceedings of Machine Learning Research, Vol 162, 2022; Gasser, Hans-Christof et aL, Integrating MHC Class I visibility targets into the ProteinMPNN protein design process, BioRxiv, 2024).
[0062] Additional criteria / limitations may include a definition of hotspots to be taken into account during the generative protein design. In other embodiments, the additional criteria / limitations do not include a definition of hotspots to be taken into account during the generative protein design.
[0063] The RFdiffusion process generates a structure backbone representing the pMHC binder, that is based on the structural input and the additional criteria chosen, e.g., pMHC binder length and hotspots.
[0064] RFdiffusion (https: / / github.com / RosettaCommons / RFdiffusion) is a computational tool that enables design of binder proteins based on diffusion modelling and allows definition of the target specification by providing a pdb (in our case the pMHC structure), hotspot specification, number of binders, length of binders and noise scaling, as defined in Bennet et al. (2023), which scales the amount of gaussian noise added to the translations (noise_scale_ca) and rotations (noise_scale_frame). In the present context "hotspots" refers to one or more amino acid position in the bound peptide or the MHC part of a pMHC complex. Such hotspots may e.g. be residues which has been identified in structures known to be involved in binding of a pMHC to a binder, such as a TCR or a CAR, or which is suspected of being important for interactions between the pMHC and a given binder, such as a TCR or CAR.
[0065] In the bound peptide of the pMHC, hotspot amino acid positions are typically positions in the bound peptide that are not buried within the peptide binding groove of the pMHC and are exposed to the environment surrounding the pMHC. In some cases, hotspot amino acid positions may also partially interact with the MHC and be at least partially in contact with the environment surrounding the pMHC. In some cases, hotspot amino acids may also include all residues of the pMHC bound peptide.
[0066] In the MHC part of the pMHC, hotspots are typically amino acid positions on the MHC domains making up the peptide binding groove that are exposed to the environment surrounding the pMHC and located around and / or in the vicinity of the peptide binding groove. In some cases, hotspots in the MHC part of the pMHC may also include amino acid positions that are known to be involved in binding of TCRs, e.g., from existing TCR-pMHC crystal structures. In other cases, hotspots in the MHC part of the pMHC may also include amino acid positions that are suspected or predicted to be involved in binding of TCRs. Hotspots that are suspected or predicted to be involved in binding of TCRs, may be identified through computer modelling of TCR-pMHC interactions and / or computer modelling of TCR-pMHC crystal structures.
[0067] In some cases, it may be desirable to include only hotspot positions in the bound peptide in the RFdiffusion processes used for pMHC binder generation in step i) as described herein. It may be desirable to include all the amino acid positions of the bound peptide as hotspots in the RFdiffusion processes used for pMHC binder generation in step i) as described herein. It may also be desirable to include all the amino acid positions of the bound peptide as the only hotspots in the RFdiffusion processes used for pMHC binder generation in step i) as described herein. In other cases, it may be desirable to include a mixture of hotspots from the bound peptide and the MHC domains making up the peptide binding groove of the pMHC in the RFdiffusion processes used for pMHC binder generation in step i) as described herein. In yet other cases it may be desirable to include only hotspots from the domains making up the peptide binding groove in the RFdiffusion processes used for pMHC binder generation in step i) as described herein. For a given pMHC, hotspots in the bound peptide and MHC domains of the pMHC are identified by analysis of the pMHC crystal structure and / or structure model.
[0068] Selection of hotspots may involve identifying all residues of the peptide not directly involved in MHC binding, as these residues are most likely to be involved in TCR and pMHC binder recognition. Typically, to achieve the best result, several different hotspot combinations (with varying numbers and hotspot positions) for each target pMHC were applied in the RFdiffusion and evaluated accordingly.
[0069] In the context of pMHC binder generation, once one or more hotspots have been identified in the target pMHC, they are introduced into the RFdiffusion model together with the target pMHC protein structure or structure model, wherein the hotspots are interpreted by RFdiffusion as specifications of contact points between the target pMHC and pMHC binder that the RFdiffusion process is tasked to generate. In this way, the identified hotspots are used to guide the generative process of RFdiffusion.
[0070] Hotspots identified as described above may also be used to optimize target pMHC binder specificity. To optimize the specific pMHC specificity, similarly to the process described in the section ‘Cross-reactivity screening’, hotspots are selected based on the potential of forming unique interactions between the peptide’s characteristic amino acids and the designed binder. In this way, hotspots are filtered based on the reduction of ambiguity of binding.
[0071] Similar to the selection of hotspots the designed binder should interact with a unique set of amino acids of the peptide to maximize the chances of obtaining the desired cross-reactive and / or selective character of the designed binder. Here, we utilized an in-house derived script to automatically extract the putatively interacting amino acids of the pMHC and the binder and predict potential bonds based on atomic proximity (within 5A). Interacting amino acids were assigned based on their mutual distance. However, the modelling may be improved by calculating the difference in Gibbs free-energy of the particular bond(s).
[0072] Following backbone generation using RFdiffusion, the structure backbone is then imported into ProteinMPNN, which populates the backbone with potential amino acid sequences, which match the structure backbone, thereby generating a first set of putative pMHC binders as defined in step i). In particularly interesting embodiments, the input provided in step a) comprises a protein structure or a structural model of a target pMHC and is processed by RFdiffusion together with a set of additional criteria and / or limitations including at least a series of hotspots and a pMHC binder sequence length, to produce a structure backbone representing the general structure of a putative pMHC binder. This backbone is then introduced into a message passing neural network (MPNN), such as ProteinMPNN, which is then used to generate a desired number of amino acid sequences that fit the backbone, thereby generating a first set of putative pMHC binders as defined in step i).
[0073] De novo design of binders remains, in most cases, a two-step process of initial shape and subsequent sequence generation. As an alternative, the combination of AF2seq and ProteinMPNN can be used in a similar fashion to RFdiffusion and ProteinMPNN.
[0074] Alternatives to ProteinMPNN for sequence generation are:
[0075] • CarbonDesign: introduces Inverseformer, which learns representations from backbone structures and an amortized Markov random fields model for sequence decoding
[0076] • AntiFold: protein inverse folding model capable of generating diverse sequences folding into the same structure,
[0077] • SPDesign: protein sequence designer based on structural sequence profile using ultrafast shape recognition
[0078] • Evo: multimodal artificial intelligence model that can interpret and generate genomic sequences at a vast scale
[0079] • PocketGen: a deep generative model that produces residue sequence and atomic structure of the protein regions in which ligand interactions occur
[0080] • CoVES: Unsupervised approach termed CoVES (Combinatorial Variant Effects from Structure) based on the local structural contexts around a residue to predict mutation preferences
[0081] However, the generation of the first set of putative pMHC binders in step i) can be performed by alternative models. In this regard, the binder design modules can be regarded as alternatives to the shape generation of RFdiffusion, and to the sequence prediction of ProteinMPNN. The listing, definition and evaluation of the present alternatives are orientated towards the areas of (1) technology used (includes the underlying architecture, training data sets, and experimental validation), (2) potential for the generation of pMHC binders, and (3) differentiation from the combinatorial use of Rfdiffusion and proteinMPNN.
[0082] BindCraft
[0083] • Release date: Bindcraft was published as a preprint Oct. 21st, 2024. The preprint has been updated twice (latest: April 25th, 2025). The article has not been published in a peer-reviewed journal yet.
[0084] • Licensing and code availability: The code is publicly available and usable for commercial purposes.
[0085] • Claim: (1) Increasing beta-sheet content in secondary structure by penalizing alpha helicity; (2) User-friendly format with reduced human intervention.
[0086] • T echnology: A ColabFold implementation of AF2 is used for the generation of the de novo shape. The design process is initialised with a random sequence for the binder. Next, AF2 predicts the co-complex structure, and from the resulting confidence score, such as ipAE, pLDDT, and pAE of the binder, an error gradient is calculated that guides the iterative design process. Furthermore, the residue contacts within the binder and between the binder have a high proportional weight in the error gradient.
[0087] • Potential for design of pMHC binder: pMHC binder design can be conducted with BindCraft. The authors validated the binder design with BindCraft experimentally for 14 targeted molecules, including the PD-1 and PD-L1 immune checkpoint.
[0088] • Differentiation to RFDiffusion: (1 ) BindCraft repredicts the structure of the bindertarget complex at each iteration of the design process (as opposed to a fixed target structure in RFdiffusion). According to the authors, this leads to a higher flexibility of the sidechain (e.g., orientation, rotamers, angles). However, the authors also report 0.5-5.5 A RMSDCa difference of the target structure before and after the binder design. (2) BindCraft uses a combination of AF2 and MPNN for the de novo generation and optimisation of binder designs, allowing a binder backbone and sequence co-design.
[0089] Protpardelle
[0090] • Release date: A pre-print was released on May 25th 2023 and the peer-reviewed article is available since June 24th 2024, in Proceedings of the National Academy of Sciences doi: 10.1073 / pnas.2311500121 • Licensing and code availability: The code is publicly available and usable for commercial aspects.
[0091] • Claim: (1) Structure and sequence co-design algorithm (2) Generation of full-atomic structural model (opposing to RFdiffusion backbone design)
[0092] • Technology: Most current models, including RFdiffusion identify and fix the dimensionality of input and output prior to the diffusion process. For example the sequence length (defined by the amino-acod positions and the total number of atoms) and the shape (defined as acloud distribution of atomic coordinates or relative bond angles) can be regarded as dimensionality. The Protpardelle algorithm, the dimensionality varies between the diffusion steps. In general, all sidechain residues of the 20 canonical amino-acids are represented in the same model (the authors refer to it as ‘superposition’). In each sequence-generating step, one amino-acid is chosen. In every diffusion step, the backbone atoms are denoised and there atomic positions determined, while only the previously chosen sidechain is unmasked and denoised. All other sidechain residues for each amino-acid position remain masked are not updated in position.
[0093] • Potential for design of pMHC binder: Due to the focus on the integration of sidechain atoms the Protpardelle algorithm could have a significant potential for the use of pMHC binder design allowing to form and foster articular interactions towards the upward-facing sidechains of the peptide. However, the authors did not validate their de novo designs experimentally. Additionally, when evaluating the capability to ‘repack’ protein structures within a dataset that includes CASP14 and CASP15 structures, the Protpardelle mosel revealed higher clash scores than the competing algorithms as the authors report which renders.
[0094] • Differentiation to RFDiffusion: (1 ) Simultaneous design of sequences and structure of Protpadelle in comparison to RFdiffusion (2) No experimental validation of ProtPardelle while RFDiffusion has been tested extensively.
[0095] BoltzDesignl
[0096] • BoltzDesignl is referred to as an alternative design programme to RFdiffusion, ProteinMPNN, and Alphafold (see paragraph below on technology). Albeit, BoltzDesign also employs ProteinMPNN for sequence design and Alphafold3 for structural validation, which effectively renders BoltzDesignl an alternative to RFdiffusion in the context of pMHC binder design since only the initial shape generation step differentiates BoltzDesignl from the combination of RFdiffusion, ProteinMPNN, AF2 Jnitial guess as used in the underlying patent application. • Release date: BoltzDesignl was published as a preprint April 25th, 2025 on bioarxiv and has not been updated or published in a peer-reviewed journal yet. Moreover, the authors / inventors published the code on GithHub with a first contribution on April 2nd 2025. For a broader access even without computational resources to the user’s disposal, a version of BoltzDesignl has been integrated into a Google colab format.
[0097] • Licensing and code availability: The code is publicly available and usable for commercial purposes.
[0098] • Validity: While the code is freely available, it is currently (July 23rd 2025) referred to as in developmental stage and purposed for co-contribution and testing. The disclaimer extends to the aspect that no experimental validation has been published yet.
[0099] • Claim: (1 ) Optimising the probability distribution of atom distances allows for more efficient sampling of the sequence-structure landscape in comparison to diffusion and denoising step iteration of a random initial seed. (2) Similar to BindCraft, the antigen and the binder are co-predicted, allowing to account for an induced fit.
[0100] • Technology: Similar to RFdiffusion in combination with ProteinMPNN and AF2 initial guess as used in this patent application, BoltzDesignl is comprised of the same three subphases of (1 ) shape generation (using Boltzl or a newer version e.g., Boltz2), (2) sequence design (using ProteinMPNN or a similar MPNN), (3) structure validation (primarily, Alphafold3). In the third and last step, ipTM scores are calculated and recommend as filter step. In addition, AF3 calculates ipAE and pLDDT values by default and is capable of predicting ipSAE values.
[0101] Within the first subphase of shape generation, BoltzDesignl employs the probability distribution of distance pairs between atoms combined with a confidence model comprised of pLDDT, pAE, and pTM values. These scores are used to define a loss function and guide the sequence representation in the recycle process of the binder design. Once an initial binder shape and sequence is defined, LigandMPNN / ProteinMPNN or a similar MPNN optimise the sequence before structure validation using AF3.
[0102] • Potential for design of pMHC binder: pMHC binder design can be conducted with BoltzDesignl . The co-prediction of antigen and binder might hold an advantage in terms of a conformational change of especially the peptide in the MHC upon binder engagement and thus simulate an induced fit which is not feasible with RFdiffusion as it employs the provided structure representation of the antigen provided in the input. • Differentiation to RFdiffusion: (1 ) One of the main points of differentiation of BoltzDesign towards RFdiffusion (though not RFdiffusion3, which has not yet (July 23rd, 2025) been realised) is the capability to design binders towards RNA / DNA macromolecules, metal ions, and small organic compounds. However, concerning the design of pMHC binders, this focus area is irrelevant. (2) In BoltzDesign the atomic identity and spatial position are within a distogram of the probability distribution of atom pairs and, thus, not diffused as a three-dimensional structure, which, according to the authors, saves computational resources.
[0103] Selecting a second set of putative pMHC binders (step ii)
[0104] Following the generation of the first set of putative pMHC binders, this first set of putative pMHC binders is imported into AF2. In AF2, the amino acid sequences of the putative pMHC binders generated by ProteinMPNN are converted into folded protein structures and the binding of the folded protein structures to the target pMHC is evaluated through calculation of the predicted local distance difference test (pLDDT) for the folded protein structures and calculation of the predicted aligned error (pAE) for the complex between the folded protein structures and the pMHC.
[0105] Generally, pAE is an inherent measure of how confident AlphaFold (AF) (such as AlphaFold, AlphaFold 2 or AlphaFold 3 (AF3) or similar predictive algorithm such as ESM-fold) is in the relative position of two residues within the predicted structure. However, originally the pAE was defined as the expected positional error at residue X, measured in Angstroms (A), if the predicted and actual structures were superpositioned on residue Y. Therefore, pAE is effectively a measure of how confident AF2 is that the domains are well packed and that the relative placement of the domains in the predicted structure is correct. Accordingly, a low pAE suggests a high confidence structure, and a high pAE suggests a lower confidence structure. The pAE may therefore be calculated as a result of the complex between a binder and a pMHC (interaction pAE or ipAE), for the binder or pMHC separately, or for a selection of e.g., key residues in the structure (e.g. local pAE). For example, in some instances the suggested position or coordinates of in particular the residues directly involved in the interaction between the binder and the pMHC, from the binder, peptide and MHC may be selected and the pAE may be calculated on the basis of their respective positional errors, giving rise to a local pAE. In some instances, it is preferable that the evaluation of the binder considers both the local and interaction pAE. Accordingly, in embodiments, the calculated interaction pAE may be <10A, such as 10A or less, 9A or less, 8A or less, 7 A or less, 6A or less, or such as 5A or less, and the local pAE may be 5A or less, such as 5A or less, 4A or less, 3A or less, 2A or less, or 1 A or less, or such as <5A.
[0106] Generally, the pLDDT is a per-residue measure of local confidence. It is scaled from 0 to 100, with higher scores indicating higher confidence and usually a more accurate prediction. pLDDT measures confidence in the local structure, estimating how well the prediction would agree with an experimental structure. It is based on the local distance difference test Co (IDDT-Ca), which is a score that does not rely on superposition but assesses the correctness of the local distances.
[0107] On this basis, a pLDDT above 90 would be taken as the highest accuracy category, in which both the backbone and side chains are typically predicted with high accuracy. In contrast, a pLDDT above 70 usually corresponds to a correct backbone prediction with misplacement of some side chains. As such the pLDDT score can vary significantly throughout a protein sequence. This means that AF (such as AF, AF2 or AF3 or similar predictive algorithm) can be very confident in the structure of some parts of the protein, but less confident in other regions. Accordingly, as for the pAE, the pLDDT may be calculated for the entire complex structure, or for the binder alone, the pMHC alone or for selected regions in the respective structures. Usually, more unstructured regions give a lower pLDDT score, suggesting a lower confidence structure, while more structured or globular protein structures generally provides a higher pLDDT score. Generally, a pLDDT of 90% or above for a given structure is considered to suggest a high confidence structure. Accordingly, in a method as described herein, it may be favorable to use a higher pAE, such as a pAE in the range of 5-15A and lower pLDDT such as a pLDDT score in the range of 80-90%, in the initial selection step, to maintain a higher degree of structural and sequence diversity.
[0108] Putative binder selection may utilize alternative scoring metrics that predict protein-protein interactions. Examples include the predicted Template Modeling score (pTM) and the interface predicted Template Modeling score (ipTM), which are both derived from the Template Modeling (TM) score. The TM score measures the global structural accuracy of a predicted protein model (https: / / academic.oup.eom / bioinformatics / article / 26 / 7 / 889 / 213219). While pTM and ipTM are commonly used to evaluate interaction structures across the entire modeled complex, ipTM may produce unreliable results when full-length protein sequences are used. This unreliability arises from the inclusion of disordered regions or non-interacting domains, which can cause significant variations in ipTM scores despite consistent interface contacts. Consequently, ipTM may not accurately reflect interface quality in practical scenarios where interacting domains are not precisely known.
[0109] To address these limitations, the interaction prediction Score from Aligned Errors (ipSAE) has been developed (https: / / www.biorxiv.0rg / content / IO.l 101 / 2025.02.10.637595v1 ). ipSAE is calculated directly from predicted Alignment Error (pAE) values for interchain residue pairs, as provided in the output JSON files of AlphaFold2 or AlphaFold3. Unlike ipTM, ipSAE focuses specifically on high-confidence interface regions by: (1 ) including only residue pairs with low pAE values, thereby excluding poorly aligned regions; (2) modifying TM-score length normalization to account solely for high-confidence residues; and (3) using actual pAE values for residue-residue scoring, rather than inferred probabilities. This approach yields a more robust and domain-specific assessment of protein-protein interfaces.
[0110] Following the initial selection as described above, a second set of putative pMHC binders is then selected from the first set of putative pMHC binders.
[0111] The selection of the second set of putative pMHC binders may be based on pAE values and / or pLDDT values of hotspot positions, based on pAE values and / or pLDDT values of the putative pMHC binder, based on pAE values and / or pLDDT values of the putative pMHC binder-pMHC complex, and / or may be based on values for the residues positioned in the interface between the pMHC and putative pMHC binder.
[0112] The selection of the second set of putative pMHC binders according to pAE values and / or pLDDT values is based on pAE and / or pLDDT cutoff values. It is particularly desirable that the second set of putative pMHC binders in step ii) is selected to have a pLDDT value of above 80%, 81%, 82%, 83%, 84% or above 85%, or such as in the range of 80-90%, and / or a pAE value of below 10A, 11 A, 12A, 13A, 14A, or below 15A, or such as in the range of 1 - 15A, preferably in the range of 5-15A.
[0113] In some embodiments, the second set of putative pMHC binders in step ii) is selected on the basis of a pLDDT value of 75% or above, 76% or above, 77% or above, 78% or above, 79% or above, 80% or above, 81% or above, 82% or above, 83% or above, 84% or above, or 85% or above. In some embodiments, the second set of putative pMHC binders pMHC binders in step ii) is selected on the basis of a pLDDT value above 75%, above 80%, or above 85%. In some embodiments, the second set of putative pMHC binders in step ii) is selected to have a pLDDT value above 80%. In exemplary embodiments, the second set of putative pMHC binders in step ii) is selected to have a pLDDT value above 85%. In some embodiments, the second set of putative pMHC binders pMHC binders in step ii) is selected to have a pAE value below 10A, 11 A, 12A.13A, 14A, or below 15A. In some embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 10A or below 12A. In some embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 15A. In exemplary embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 12A.
[0114] In other embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 15A and a pLDDT value above 85%.
[0115] In other embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 12A and a pLDDT value above 85%.
[0116] In other embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 10A and a pLDDT value above 85%.
[0117] In other embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 10A and a pLDDT value above 88%.
[0118] In other embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 10A, 11 A, 12A.13A, 14A, or below 15A and a pLDDT value above 83%, 84%, 85%, 86%, 87% or above 88%.
[0119] In further embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 10A, 12A or below 15A and a pLDDT value above 80%, 81 %, 82%, 83%, 84%, or above 85%.
[0120] In further embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value in the range of 5-15A and a pLDDT value in the range of 80-99%.
[0121] In particularly interesting embodiments, the second set of putative pMHC binders in step ii) is selected to have a pAE value below 10A and a pLDDT value above 88%, preferably a pAE value below 7 A and a pLDDT value above 90% In embodiments, the second set of putative pMHC binders in step ii) is selected based on the pAE value of the hotspot positions, based on the pAE value of the putative pMHC binder, based on the pAE value of the putative pMHC binder-pMHC complex, and / or based on the pAE values for the residues positioned in the interface between the pMHC and putative pMHC binder.
[0122] In embodiments, the second set of putative pMHC binders in step ii) is selected based on the pLDDT value of the hotspot positions, based on the pLDDT value of the putative pMHC binder, based on the pLDDT value of the putative pMHC binder-pMHC complex, and / or based on the pLDDT values for the residues positioned in the interface between the pMHC and putative pMHC binder.
[0123] In embodiments, the second set of putative pMHC binders in step ii) is selected based on the pLDDT value and pAE value of the hotspot positions, based on the pLDDT value and pAE value of the putative pMHC binder, based on the pLDDT value and pAE value of the putative pMHC binder-pMHC complex, and / or based on the pLDDT value and pAE value for the residues positioned in the interface between the pMHC and putative pMHC binder.
[0124] The putative pMHC binders selected in step ii) and / or optimized in step iii), may be subjected to further testing and selection based on performance of the putative binders in a fine grained molecular dynamics model. In particular, such a model can be used to assess entropy and hydrogen bond formation and provide a score for a given putative pMHC. For entropy, low scores are better and for hydrogen bond formation, high scores are better, and ideally a desirable putative pMHC binder has both a low score for entropy and a high score for hydrogen bond formation. These scores can be used to rank the putative pMHC binders and be used to select which of the putative pMHC binders that should be used in downstream and eventually laboratory testing.
[0125] Optimization by partial diffusion (step iii)
[0126] The second set of putative pMHC binders that were generated in step ii) is further optimised in silico by subjecting said second set of putative pMHC binders to a partial diffusion process. As such the partial diffusion process restarts the generative design process from the protein structure solution represented by the second set of putative binders. This allows for the generative design process to re-evaluate, diversify, and optimise on the structure solution of the second set of putative pMHC binders by exploring the "possible folding space", i.e., exploring other alternative structures and sequences within the overall structure solution defined by the second set of putative pMHC binders, leading to a more complete and robust protein structure generation for the target pMHC binders.
[0127] The purpose of the partial diffusion process is to diversify the structure of the putative pMHC binders generated through the RFdiffusion process, when the generated putative pMHC binders as the point of origin for a second, but less extensive diffusion process, herein referred to as the partial diffusion. Essentially, this allows for a second more focused analysis of "sequence space" that explores the suitability of structures that are similar but not identical to the first predictive solution generated by the RFdiffusion.
[0128] The partial diffusion process is generally followed by a second round of amino acid sequence generation by a message passing neural network (MPNN), before employing a protein folding prediction software such as AF which then performs a further round of amino acid sequence to folded protein structure conversion.
[0129] The partial diffusion process may comprise at least 5, 10, 15, 20, 25, 30, 35, 40, 45 or a maximum of 50 noising time steps. Alternatively, the partial diffusion process may comprise at least 50 noising time steps.
[0130] As used herein the term “noising time step” relates to the iterations of corruption of structural coordinates and orientations of the residue frames and subsequent prediction of an updated version of a structure Xo. At any given point in the diffusion trajectory, may it be the initial diffusion campaign (starting from random noise Xtwith typically 200 noising time steps) or the partial diffusion re-interpretation originating from a previously obtained structure Xo_0id, noise as randomly sampled number in a given stochastic distribution (determined by the sampling temperature) around the previous value is added for all translational and rational matrices of all atoms in the residue frame. The newly noised input frame is the starting point for the re-interpretation of the new prediction Xo new in the denoising step. Generally, an increasing number of noising time steps in a partial diffusion allows to obtain a new predicted structure Xo new which increasingly unrelated (in structural properties such as secondary structural elements) from the input structure Xooid.
[0131] It has been found that about 20 noising time steps is particularly useful for the partial diffusion process as defined herein. In Example 1 (Structure Refinement) it is demonstrated how a partial diffusion comprising 20 noising steps can be used to optimize putative pMHC binders. In one or more embodiments, the partial diffusion process may be performed by RFdiffusion, the second round of amino acid sequence generation by a message passing neural network (MPNN) may be performed by ProteinMPNN, and the second of converting the amino acid sequences into folded protein structures may be performed by a version of AF or Colabfold (CF).
[0132] In one or more embodiments, the partial diffusion process is performed by RFdiffusion, the second round of amino acid sequence generation is performed by ProteinMPNN, and the second of converting the amino acid sequences into folded protein structures is performed by AF, or a version thereof.
[0133] In one or more embodiments, the partial diffusion process comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45 or a maximum of 50, wherein the partial diffusion process is performed by RFdiffusion, the second round of amino acid sequence generation is performed by ProteinMPNN, and the second of converting the amino acid sequences into folded protein structures is performed by AF, or a version thereof.
[0134] In one or more embodiments, the partial diffusion process comprises at least 20, 25, 30, 35, 40, 45 or at least 50, wherein the partial diffusion process is performed by RFdiffusion, the second round of amino acid sequence generation is performed by ProteinMPNN, and the second of converting the amino acid sequences into folded protein structures is performed by AF, or a version thereof.
[0135] In one or more embodiments, the partial diffusion process comprises 20 noising time steps, wherein the partial diffusion process is performed by RFdiffusion, the second round of amino acid sequence generation is performed by ProteinMPNN, and the second of converting the amino acid sequences into folded protein structures is performed by AF, or a version thereof.
[0136] Once the partial diffusion process is completed, the optimised set of putative pMHC binders may be selected using ipAE and pLDDT scores as calculated by the folding software. It is particularly desirable that the optimised set of putative pMHC binders are selected to have a pLDDT of above 91 %, 92%, 93%, 94%, or above 95% and an ipAE of below 5A, 6A, 7A, 8A or below 9A. As used herein the term “interaction pAE” (ipAE) relates to pAE but focuses on binder-target complexes. ipAE is calculated by taking the pAE between every binder-target residue pair and averaging these values. Lower ipAE values (typically below 10 A) at binder-target residues indicate a tighter and more accurate prediction of the binding sites, which can be interpreted as a proxy for higher binding affinity. This direct correlation with interaction quality makes ipAE a valuable metric for predicting the binding of proteins (Evans et al, Protein complex prediction with AlphaFold-Multimer, bioRxiv, 2022).
[0137] In some embodiments, the optimised set of putative pMHC binders in step iii) is selected to have a pLDDT value above 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, or above 95%, or a pLDDT in the range of 85-99%. In some embodiments, the optimised set of putative pMHC binders pMHC binders in step iii) is selected to have a pLDDT value above 85%, above 90%, or above 95%, or a pLDDT in the range of 85-99%. In some embodiments, the optimised set of putative pMHC binders in step iii) is selected to have a pLDDT value above 90%, preferably above 95%.
[0138] In some embodiments, the optimised set of putative pMHC binders pMHC binders in step iii) is selected to have a ipAE value below 2A, 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, or below 12A, or a ipAE in the range of 1-12A, preferably 1-7 A, more preferably 1-5A. In some embodiments, the optimised set of putative pMHC binders in step iii) is selected to have a ipAE value below 5A, 7 A or below 9A. In some embodiments, the optimised set of putative pMHC binders in step iii) is selected to have a ipAE value below 7A.
[0139] In other embodiments, the optimised set of putative pMHC binders in step iii) is selected to have a ipAE value below 10A, 11 A, 12A.13A, 14A, or below 15A and a pLDDT value above 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, or above 95%.
[0140] In further embodiments, the optimised set of putative pMHC binders in step iii) is selected to have a ipAE value below 5A, 7 A or below 9A and a pLDDT value above 90%, 91%, 92%, 93%, 94%, or above 95%.
[0141] In particularly interesting embodiments, the optimised set of putative pMHC binders in step iii) is selected to have a ipAE value below 7 A and a pLDDT value above 90%. Optimization by reduction of cross reactivity risk
[0142] The second set of putative pMHCs or the optimised set of putative pMHCs obtained from the partial diffusion process, can be further optimized by elimination of putative pMHCs that show risk of cross reactivity with variant pMHCs and / or unbound MHCs (MHCs that are not in complex with a peptide).
[0143] The risk of cross reactivity of putative pMHC binders with variant pMHCs can be assessed by first identifying a collection of variants to be tested and then determining a measure for how likely it is that a given variant pMHC will cross react with a putative pMHC binder, and then selecting only putative pMHC binders that are considered to have a sufficiently low likelihood of cross reacting with variant pMHCs. Similarly, binding of variant (or wildtype) peptides presented on other MHC allele variants can be assessed. Together, the crossreactivity testing can be performed against mutated peptides on same MHC alleles, same peptides on other MHC alleles, or mutated peptides on other MHC alleles.
[0144] One way of determining the likelihood that a putative pMHC will cross react with a variant pMHC comprises determining the ipAE of a variant pMHC with a putative pMHC binder and evaluating the ipAE value. The ipAE can also be determined for a large number of variants, such as a variant library, against one or more putative pMHC binders, such as against a library of putative pMHC binders. Once the ipAE is determined, all the putative pMHC binders that have a desirable ipAE against one or more variant pMHCs are selected. A desirable ipAE for a putative pMHC binder against a given variant pMHC may in this context be an ipAE of above 10A.
[0145] Alternatively, once the ipAE is determined, all the putative pMHC binders that have a desirable ipAE against one or more variant pMHCs and a desirable ipAE towards a target pMHC are selected. A desirable ipAE for a putative pMHC binder against a given variant pMHC may be an ipAE of above 25A and a desirable ipAE against a target pMHC may be an ipAE of below 15A.
[0146] Thus, a set of further optimized putative pMHC binders may be obtained through selection as described above and may eventually be used in downstream applications and / or in vitro testing.
[0147] Thus, in one or more embodiments, the optimization in step iii), further comprises -Identifying one or more potential variant peptides presented by said target pMHC (variant pMHC),
[0148] -Determining the ipAE of said variant pMHC: putative pMHC binder complexes,
[0149] -Selecting putative pMHC binders with ipAE >12A for the variant pMHC and an ipAE < 5-7A for the target pMHC, and
[0150] -Obtaining from said selection one or more further optimized putative pMHC binders.
[0151] Thus, in one or more embodiments, the optimization in step iii), further comprises -Identifying one or more potential variant peptides presented by said target pMHC (variant pMHC), -Determining the ipAE of said variant pMHC: putative pMHC binder complexes, -Selecting putative pMHC binders with ipAE > A for the variant pMHC and an ipAE < 7 A for the target pMHC, and -Obtaining from said selection one or more further optimized putative pMHC binders.
[0152] Such determination of the ipAE can be made using structural modelling program such as AF2 or CF. In one or more embodiments, the determination of the ipAE is made using AF2. In one or more embodiments, the determination of the ipAE is made using CF.
[0153] In similar fashion, the risk of cross reactivity of putative pMHC binders with unbound MHCs can be assessed by first identifying a collection of unbound MHCs to be tested and then determining a measure for how likely it is that a given unbound MHC will cross react with a putative pMHC binder, and then selecting only putative pMHC binders that are considered to have a sufficiently low likelihood of cross reacting with unbound pMHCs.
[0154] One way of determining the likelihood that a putative pMHC will cross react with an unbound MHC comprises determining the ipAE of a unbound MHC with a putative pMHC binder and evaluating the ipAE value. The ipAE can also be determined for several MHC alleles, such as an MHC allele library, against one or more putative pMHC binders, such as against a library of putative pMHC binders. Once the ipAE is determined, all the putative pMHC binders that have a desirable ipAE against one or more unbound MHCs are selected. A desirable ipAE for a putative pMHC binder against a given unbound MHC may in this context be a ipAE of above 10A.
[0155] Alternatively, once the ipAE is determined, all the putative pMHC binders that have a desirable ipAE against one or more unbound MHCs and a desirable ipAE towards a target pMHC are selected. A desirable ipAE for a putative pMHC binder against a given unbound MHC may be an ipAE of above 10A and a desirable ipAE against a target pMHC may be an ipAE of below 7A.
[0156] Thus, a set of further optimized putative pMHC binders may be obtained through selection as described above and may eventually be used in downstream applications and / or in vitro testing.
[0157] Thus, in one or more embodiments, the optimization in step iii), further comprises -Providing a structure of an unbound MHC, wherein the unbound MHC corresponds to the MHC part of the target pMHC, -Determining the ipAE of said unbound MHC:putative pMHC binder complexes, -Selecting putative pMHC binders with ipAE > 10-12A for the unbound MHC and an ipAE < 3-7A for the target pMHC, -Obtaining from said selection one or more further optimized putative pMHC binders.
[0158] In one or more embodiments, the optimization in step iii), further comprises -Providing a structure of an unbound MHC, wherein the unbound MHC corresponds to the MHC part of the target pMHC, -Determining the ipAE of said unbound MHC:putative pMHC binder complexes, -Selecting putative pMHC binders with ipAE > A for the unbound MHC and an ipAE < 7 A for the target pMHC, -Obtaining from said selection one or more further optimized putative pMHC binders.
[0159] Such determination of the ipAE can be made using software programs such as AF2 or CF. In one or more embodiments, the determination of the ipAE is made using AF2. In one or more embodiments, the determination of the ipAE is made using CF. Optimization by Bayesian optimization
[0160] The one or more putative pMHC binders and / or sets of putative pMHC binders identified and selected for further optimization under step ii) or ill) may further be optimized using Bayesian optimization. In particular, any putative pMHC binder that was optimized by either partial diffusion, elimination of putative pMHC binders with cross reactivity against variant pMHCs and / or elimination of putative pMHC binders with cross reactivity against unbound pMHCs was further optimized using Bayesian optimization. Bayesian optimization within the present context may comprise the use of Multi-Objective Bayesian Optimization (MOBO), MultiObjective Evolutionary Algorithms (MOEAs), Multi-Objective Particle Swarm Optimization (MOPSO), Multi-Objective Simulated Annealing (MOSA), Scalarization Techniques, Interactive Methods, Surrogate-assisted Optimization, and / or Deterministic Methods.
[0161] In one or more embodiments, the optimization of step iii) comprises the use of MultiObjective Bayesian Optimization (MOBO), Multi-Objective Evolutionary Algorithms (MOEAs), Multi-Objective Particle Swarm Optimization (MOPSO), Multi-Objective Simulated Annealing (MOSA), Scalarization Techniques, Interactive Methods, Surrogate-assisted Optimization, and / or Deterministic Methods.
[0162] In one or more embodiments, the optimization of step iii) comprises the use of Bayesian optimizations.
[0163] In one or more embodiments, the optimization step iii) comprises further optimizing any putative pMHC binder that was selected following optimization by either partial diffusion, elimination of putative pMHC binders with cross reactivity against variant pMHCs and / or elimination of putative pMHC binders with cross reactivity against unbound pMHCs, by using Multi-Objective Bayesian Optimization (MOBO), Multi-Objective Evolutionary Algorithms (MOEAs), Multi-Objective Particle Swarm Optimization (MOPSO), Multi-Objective Simulated Annealing (MOSA), Scalarization Techniques, Interactive Methods, Surrogate-assisted Optimization, and / or Deterministic Methods.
[0164] In one or more embodiments, the optimization step iii) comprises further optimizing any putative pMHC binder that was selected following optimization by either partial diffusion, elimination of putative pMHC binders with cross reactivity against variant pMHCs and / or elimination of putative pMHC binders with cross reactivity against unbound pMHCs, using Bayesian optimization.
[0165] Selecting from optimized set (step b)
[0166] Following optimization in step iii) by one or more of the optimization steps described herein, the resulting set of putative pMHC binders are subjected to a selection step, wherein only the putative pMHC binders having a pLDDT value above 90% and an ipAE value of below 7 A are selected. This selection step ensures, together with the optimization steps of iii), that only the putative pMHC binders with the highest specificity and affinity are selected.
[0167] Outputting pMHC binder candidates (step c)
[0168] Following the selection of step b, the sequences of the putative pMHC binders that were selected in b, are outputted as amino acid sequences, thereby generating a set of candidate pMHC binders.
[0169] Depending on the exact optimization steps each set of putative pMHC binders were subjected to, the exact properties may vary. In particular, a pMHC binder candidate obtained in step c) may have
[0170] -a pLDDT of >90%, and an ipAE of <7A against a target pMHC,
[0171] -a pLDDT of >90% against a target pMHC, an ipAE of <7A against a target pMHC, and an ipAE of >10A against variant pMHCs,
[0172] -a pLDDT of >90% against a target pMHC, an ipAE of <7A against a target pMHC, and an ipAE of >10A against unbound MHCs,
[0173] - a pLDDT of >90% against a target pMHC, an ipAE of <7A against a target pMHC, an ipAE of >10A against unbound MHCs, and / or an ipAE of >10A against variant pMHCs, or
[0174] -a pLDDT of >90% against a target pMHC, an ipAE of <7A against a target pMHC, an ipAE of >10A against unbound MHCs, and an ipAE of >10A against variant pMHCs.
[0175] Any candidate as defined above may be used downstream in in vitro assays to validate that the pMHC binder candidates are in fact pMHC binders having high affinity and high specificity toward a target pMHC. After having in vitro validated pMHC binders, it may be desirable to further optimise the in vitro validated pMHC binders through cross reactivity testing or partial diffusion sequence diversification based on the data obtained in the in vitro assays. The validated pMHC binders may also be subjected to further rounds of optimisation using the Bayesian optimization algorithm.
[0176] However, the data obtained for the in vitro validated pMHC binders may also be used to update the bayesian optimization algorithm. Updating the algorithm with data on pMHC binders that have been shown to bind a given pMHC may be used to further train the bayesian optimization algorithm narrow the directions the bayesian optimization process works in.
[0177] Similarly, providing the algorithm with in vitro validated data on pMHC candidates that did not bind to a given pMHC, may also be used to train and update the bayesian optimization algorithm and eventually help guide the bayesian optimization process.
[0178] Generally speaking, there is no limit to how often a given pMHC binder or pMHC binder candidate may be further optimized, even after succesful binding has been demonstrated. All information obtained regarding in vitro binding of a specific pMHC candidate may be used to train and update the algorithms used, and the updated algorithms may then be used to further explore alternative and / or optimized sequences based on the specific pMHC binder sequence. Any new data arising from those alternative and / or optimized sequences may then be used to further update the algorithms involved, and this information loop may be repeated for as many times as it is deemed necessary and / or beneficial.
[0179] The in vitro data used to update the algorithms could be based on affinity data, specific residues, specific sequences and / or specific sequence motifs, which according to the data may appear to play either a positive or a negative role in candidate pMHC binding.
[0180] In one or more embodiments pMHC binders confirmed in vitro can be further diversified through partial diffusion to generate a larger set of variants with improved affinity or decreased cross-reactivity.
[0181] In one or more embodiments in vitro binding data of identified pMHC binders and nonbinders can be used to improve the Bayesian Optimization algorithm with the use of in vitro affinity data in replacement of or in addition to the ipAE. AlphaFold
[0182] AlphaFold (AF) is an artificial intelligence (Al) tool based on a deep learning system that can be used for prediction of the three-dimensional structure of a protein from the amino acid sequence of the protein and aligned homologous sequences. In general, the AF consists of two stages, the neuronal network block (‘Evoformer’) and the structure module. Briefly, within the Evoformer neural network block, the analysis results of the multi-sequence alignment (MSA) and the residue pairs (all combinations of each residue with another) are represented in two separate arrays, a Nseq(number of aligned sequences of the processed MSA) x Nres (number of residues in the binder and pMHC complex) and a Nresx Nresarray. However, in the context of de novo designed sequences, which are generally more stable and regular in terms of the secondary structural elements (e. g..Table 7) than naturally occurring proteins, accurate prediction can be obtained from a few to a single sequence without the need for or only a reduced MSA. This significantly reduces the computational time and provides a property of AF which makes it well suited for high throughput scoring of tens of thousands of structural protein designs using a simple internalized error alignment matrix for scoring. The reduction of scoring matrices to only a few also entails the advantage of training further secondary algorithms (e.g., the ‘Bayesian Optimization’ algorithm for specificity enhancement).
[0183] In the second part of the AF architecture, the structural model uses a concrete three- dimensional backbone structure based on the pair representation and original sequence (first row of the MSA). Thereby, the backbone structure is represented as Nresindependent rotations and translations, representing the geometry of the N-CaiPha-C atoms, which vastly constrains the orientation of the side chain residues. Furthermore, after application of attention-based update mechanisms and including recycling (Bennett, N. R. et al.: Improving de novo protein binder design with deep learning, Nature Communnications, 2023), the relative positions of the backbone atoms and the angle of the side chain are calculated. In the entire process of de novo structure prediction of RFdiffusion, ProteinMPNN, and AF, this is the first time the orientations of the side chains of the binder are predicted. The final loss compares the predicted atom position xpto the ‘true’ position xtwhile aligned in a third position y with respect to its translation and orientation. When aligned in y all distances of atoms in the true and the predicted frame are calculated and the resulting Nres2distance matrix penalized with a loss function, which creates a high weight for atoms being correctly positioned in comparison to the local environment, and, thus, a heigh weight on the correct placement of side chains. The accurate placement and orientation of the peptide side chains in the predicted pMHC binder complex is especially crucial for the success rate for the translation of in silico results to in vitro testing.
[0184] In brief, the benefit of AF or similar structure prediction tools in the one-sided protein ligand design, and especially the pMHC are manifold. Firstly, AF has a high focus on correct alignment of side chains. Secondly, in de novo design it can perform structural prediction without sequence alignment e.g., MSA. Thirdly, the clear correlation of two matrices, ipAE and binder_pLDDT (also correlate with each other) with higher likelihood of in vitro success offer a simple scoring mechanism for high throughput pMHC binder design.
[0185] A system similar to AF2 are Colabfold (Mirdita, M. et al., ColabFold: making protein folding accessible to all, Nat Methods 19, 679-682 (2022)), ESMfold (Zeming Lin et al., Evolutionary-scale prediction of atomic-level protein structure with a language model, Science, 379,1123-1130, 2023), RoseTTAFold (Minkyung Baek et al., Accurate prediction of protein structures and interactions using a three-track neural network, Science, 373, p871- 876, 2021), OMEGAfold (Ruidong, W. et al., High-resolution de novo structure prediction from primary sequence, BioRxiv, 2022). Either one of AF2, Colabfold or a similar program may be used for calculation of pLDDT, pAE and ipAE values.
[0186] In one or more embodiments, either one of AF2 or Colabfold is used for calculation of pLDDT and ipAE values. In one or more embodiments, AF2 is used for calculation of pLDDT and ipAE values. In one or more embodiments, Colabfold is used for calculation of pLDDT and ipAE values.
[0187] Colabfold is a free protein folding platform for rapid folding of protein stuctures in silico. Colabfold combines AF2 and MMseqs2 to generate fast homology searches to enable faster protein folding in silico than AF2.
[0188] Identification of potential sterically clashes using the pymol script TopModel
[0189] TopModel is a pyMol based python script for the depiction of
[0190] 1 . Unnaturally occurring chirality of amino acids.
[0191] 2. Unnaturally occurring conformation of the amine bonds.
[0192] 3. Van der Waals radii clashes. In the design of pMHC binder, only the depiction of Van der Waals radii was considered. The python script is capable of using structure prediction of any pMHC & binder co-complex in the .pdb format as input and retrieves the identified clashes as text output. As default the T opModel detects Van der Waals clashes by calculating the distance between all pair of atoms that are within 5 A of each other. In the default version a clash is defined by: dAB < TA + TB - 0.5 A in which dAB is the distance of the atoms in the atom pair towards each other, and TA and TB are the radii of atom A and B. The penalizing term of 0.5 A was determined as useful by the authors of the code mostly based on antibody-antigen datasets. For the design minibinder the amount of clashes were too numerous and thus unsuited for deselecting the binder designs (and obtaining a reasonable selection size for in vitro testing). In consequence the term was manually increased for binder designs to be dAB < TA - 2 A.
[0193] TopModel is not part of any publication in a peer-reviewed scientific journal, however freely available under a default MIT license on the website GitHub under the following link: https: / / github.com / liedllab / TopModel.
[0194] Cross reactivity
[0195] To ensure successful binding to the desired target pMHC and minimize the risk of off-target binding, another screening step was introduced. MHC alleles are highly polymorphic and can bind a large repertoire of peptides (Radwan, J., Babik, W., Kaufman, J., Lenz, T. L. & Winternitz, J., Trends Genet. 36, 298-311 , 2020). Additionally, a particular MHC allele can have a broad peptide specificity. To minimize potential off-target interactions, it can be investigated whether the putative pMHC binders are highly specific in only binding to the target peptide (lower ipAE) and interact poorly with variant (c.f. below) peptides (higher ipAE score).
[0196] Variant peptides for use in cross reactivity testing may be obtained by creation of variant libraries comprising single- and double point mutations in the target peptide sequence (i.e. the peptide antigen of the target pMHC.
[0197] Point mutations may be applied to positions in the target peptide of the target pMHC, which are exposed to and / or expected to interact with the surface of a pMHC binder. Such positions should not include positions that are buried inside the pMHC molecule.
[0198] Since the MHC anchor positions are not usually exposed and not expected to interact with the pMHC binder, applying point mutations to the MHC anchor residues in the crossreactivity screen is unlikely to provide meaningful information about potential off target binding. Accordingly, in some cases, the MHC anchor positions are excluded from variant generation by point mutations, because these residues are not exposed and do not interact with the pMHC binder.
[0199] The variant libraries may then be blasted (Altschul et al. “Basic local alignment search tool”, Journal of Molecular Biology, volume 215(3), pages 403-410 (1990)) against the known human proteome, preferably with a high query coverage as a threshold as a high query coverage ensures that most, if not all, of the sequences in the proteome are being queried for potential matches. This blast allows for assessing which peptides are likely candidates for being presented by healthy human cells and possibly lead to off-target binding.
[0200] Since the pMHC binder might not only recognize consecutive amino acid positions of the target peptide, but several disconnected putatively more exposed amino acid, a ‘fingerprint’ for the binding mechanism is determined for each selected pMHC binder via AF2 scoring of binder and pMHC comprised of the same MHC and a single-point substituted peptides allowing to assess the binding contribution of each amino acid at each position of the target peptide. Once the scoring has been obtained the binding profile of each pMHC binder can be utilized to identify close matches in the human proteome for each individual pMHC binder. If a motif is shorter than eight amino acids and, therefore, unlikely to be presented in an MHC complex, the neighboring amino acids from the originating protein sequence on either side might be included, up to a total sequence length of specifically 8 or 9 amino acids.
[0201] The total number of amino acid sequences recovered from the human proteome blast to be tested for cross reactivity can be reduced by determining the affinity of the recovered amino acid sequences towards relevant MHC complexes. Those with a low affinity may be discarded, as they are not likely to be presented by the MHC complex as pMHCs, and therefore unlikely to result in off-target binding. Those with high affinity towards MHC are kept for further cross reactivity testing. This approach may be used to narrow down the total set of peptides to be assessed for potential cross reactivity. One possible approach for this narrowing of peptides is to use NetMHCpan4.1 to assess MHC affinity for each of the peptides recovered from the proteome blast and eliminate peptides with low MHC affinity. Other MHC alleles can also be considered within NetMHCpan4.1 and are taken into account based on target pMHC relevance (i.e. , if binding to other MHCs is likely and un / desirable). Similarly, this platform may be used to identify pMHC complexes that structurally resemble the target pMHC but have low sequence-similarity. The key metrics used in NetMHCpan4.0 for that analysis is the likelihood of binding relative to the training data and is determined in percentile. A score among the best 2%tile is considered a weak binding peptide (to the MHC) while a score below 0.5%tile is considered a strong binding peptide (to the MHC). In general, all peptides with a score within in 2%tile for the same MHC or another MHC within the following group (Considered MHC: HLA-A*01 :01 , HLA-A*02:01 , HLA-A*03:01 , HLA- A*24:02, HLA-A*26:01 , HLA-B*07:02, HLA-B*08:01 , HLA-B*27:05, HLA-B*39:01 , HLA- B*40:01 , HLA-B*58:01 , HLA-B*15:01 , which are included herein as SEQ ID NO: 538-555) have been included in further analysis.
[0202] To broaden the search all single-point peptide variants have also been directly applied to the netMHCpan scoring without prior blast against the human proteome. To increase the chances the peptide variants would be recognized and presented on different MHC complexes, several positions were modified according to the motif profiles of naturally presented peptides on a given MHC (e.g., for HLA-A*01 :01 the most predominant amino acid are D in position 3 and Y in position 9). Thus, a subset off single-point mutants with modifications for each MHC complex was provided for analysis with netMHCpan and only peptides scoring better than the original peptide were considered for further analysis.
[0203] Next, all pMHC complexes containing modified peptides of both the strategy involving human proteome blast and the strategy relying purely on netMHCpan were tested against the selected subset of pMHC binder by AF2 within dl-binder_design. If a pMHC obtained low ipAE values (below 20) for any the pMHC containing a modified peptide the PMHC binder was classified as being cross reactive and either not considered for further in vitro analysis or subjected to Bayesian Optimization for cross-reactivity reduction. A pMHC binder revealing only high ipAE values against all off-targets, but a low ipAE against a remodelled version (using AF2 and AF3) of the target pMHC, was considered ‘specific’ to the target and a valuable candidate for in vitro assessment.
[0204] Bayesian optimization
[0205] Bayesian optimization as discussed herein makes use of a Bayesian architecture that utilizes a machine learning-based approach to alter and evaluate proteins to enhance their properties systematically.
[0206] One approach to Bayesian optimization in the present disclosure, comprises optimizing a multi-objective Bayesian acquisition function, considering multiple goals simultaneously, and extending this by introducing diffusion-optimized sampling (NOS), an advanced technique for exploring potential protein designs more efficiently. NOS generates sequences that are more likely to occur naturally while also optimising the sequences against a specific objective (e.g., ipAE).
[0207] Additionally, the selection process is enhanced by using value saliency to identify the most impactful areas for modification and employ an ensemble-based approach for more accurate prediction of uncertain outcomes.
[0208] The optimization generally begins with the pMHC binder sequences post cross-reactivity scoring. However, the optimization may take place at any stage, e.g., following RFdiffusion or following partial diffusion. The selection process for modification within these sequences was guided by NOS, identifying the most impactful positions for potential alteration. This step ensured that subsequent modifications were focused on regions with the greatest potential for improving efficacy.
[0209] Putative pMHC binder sequences selected for optimization are encoded using an amino acid encoder, converting the biological sequences into a numerical format suitable for computational processing. This encoding captures essential biochemical properties within these sequences, simulating potential mutations and introducing variability to explore different sequence configurations. Central to the Bayesian architecture is a guided diffusion process, which employs several key elements:
[0210] 1. Timestep and positional embeddings: The architecture incorporates timestep embedding to track the progression of the diffusion process and sinusoidal positional embeddings to maintain the spatial context of each amino acid.
[0211] 2. Transformer Encoder: This component utilizes a transformer-based model to process and interpret the context of the encoded peptide sequences, facilitating a deeper understanding of sequence structure and potential interactions.
[0212] 3. Objective-driven guidance: The diffusion process is steered by specific objectives (high ipAE against variant peptides and low ipAE against original peptide), utilizing a regression head for quality assessment and a classifier head (MLM) for predicting amino acid probabilities. This guided approach ensures that the sequence modifications are aligned with the desired therapeutic characteristics.
[0213] 4. Iteration and Output Evaluation: The optimized sequences produced by the sampler are then evaluated against predefined objectives, such as decreased ipAE or increased pLDDT. The results of this evaluation inform subsequent iterations, allowing for continuous refinement and improvement of the peptide sequences. This iterative process is repeated until the optimized sequences have met the set objectives at a satisfactory level.
[0214] In the present context, an iterative loop involving BO and AF2 was used to refine binder sequences for target pMHC specificity. This is also illustrated in the herein-disclosed examples. Specifically, the data collected in the section above may be used as input information, i.e. ipAE and pLDDT against empty and non-target pMHCs. This allows for an optimization of the ipAE and pLDDT against the target pMHC, while worsening the scores against all non-target pMHCs and the empty target MHC. Thus, the binder sequences may be subjected to multi-objective Bayesian optimization (MOBO) with guided diffusion, as described herein. The method may thereby provide putatively optimized pMHC binder sequences. Such sequences may e.g., undergo a structural and interaction prediction using e.g., AF2, with an emphasis on their ipAE and pLDDT scores relative to each target type. Such data may be used as a guide for further subsequent MOBO optimization rounds, creating a feedback loop that iteratively refines the binder sequences. Accordingly, such iterative optimization cycle may be designed to incrementally improve the binder sequences in terms of the predicted specificity. Notably, ipAE and pLDDT are just some of many metrics that are relevant to optimize for. Specifically, LIS local interaction score, which uses a pae cutoff to determine which residues in the binder-target complex that are interacting, delta-G, entropy, number of hydrogen bonds formed, as well as in vitro data on binding, expression, specificity, etc.
[0215] MHC classes
[0216] The Human leukocyte antigens (HLA) are several large loci of polymorphic genes in vertebrates, which encode cell surface proteins called MHC molecules, that are essential for the adaptive immune system. These MHC genes and MHC molecules are divided into three classes (MHC class I, MHC class II and MHC class III) according to the function of MHC molecules encoded by the MHC genes. The role of the MHC molecules is to present peptide antigens on the surface of cells, allowing the adaptive immune system to differentiate between healthy and infected cells and trigger a variety of immune responses.
[0217] MHC class I molecules comprise a N-terminal extra-cellular heterodimer region composed of three domains, a1, a2, a3, and a P2 microglobulin subunit, a transmembrane helix and a short cytoplasmic tail. Two of the alpha domains, a1 and a2, form the peptide-binding groove between two long a-helices and the floor of the groove formed by eight p-strands. The peptide antigen is non-covalently bound to the MHC class I molecule in the peptide-binding groove. The immunoglobulin-like domain a3 is involved in the interaction with the CD8 coreceptor, present on the surface of CD8+T cells, and the P2 microglobulin subunit provides stability to the MHC molecule and participates in the recognition of pMHC class I complex by the CD8 co-receptor.
[0218] There are 3 primary MHC class I genes: HLA-A, HLA-B, and HLA-C (Table 1). These genes are highly polymorphic, that means that there are thousands of variants in humans with very different peptide-binding preferences. Accordingly, this contributes to a vast landscape of possible pMHC combinations across the population.
[0219] MHC class II molecules comprise an N-terminal extracellular heterodimer region composed of two chains: an alpha (a) chain and a beta (P) chain, each containing two domains, a1 and a2, and pi and P2, respectively. The peptide-binding groove is formed by the a1 and pi domains, with the groove consisting of two parallel a-helices that form the sides and a floor made up of eight p-strands. The peptide antigen is non-covalently bound within the peptide- binding groove. The a2 and P2 domains possess an immunoglobulin-like structure and are involved in stabilizing the overall conformation of the MHC class II molecule. Additionally, these domains contribute to interactions with the CD4 co-receptor, present on the surface of CD4+T cells. MHC class II molecules do not have a P2 microglobulin subunit, unlike MHC class I molecules. The transmembrane helices anchor the molecule in the cell membrane, and the short cytoplasmic tails of both chains participate in intracellular signalling and trafficking.
[0220] There are 3 primary MHC class II genes: HLA-DR, HLA-DQ and HLA-DP (Table 1). These encode proteins involved in peptide binding and presentation on the surface of cells. Non canonical MHC class II genes include HLA-DM, HLA-DOA and HLA-DOB. These encode intracellular proteins involved in peptide processing and loading. In mice, the MHC is known as the H-2 complex. The H-2 complex includes both MHC class I and MHC class II molecules. MHC class I molecules, such as H2-Kb and H2-Db, present endogenously derived peptides to CD8+T cells, while MHC class II molecules, such as l-A and l-E, present exogenously derived peptides to CD4+T cells (Table 1 ).
[0221] Table 1 - Primary human and mouse genes encoding for MHC class I and II proteins.
[0222] Unbound MHC
[0223] Generally, unbound MHCs are not found in nature as the MHC subunits do not assemble into MHC without coming into contact with a peptide. Using in silico methods, however it is possible to make a computational model of an unbound MHC.
[0224] An unbound MHC as used herein, refers to a computational MHC molecule that does not contain any peptide antigen. Thus, an unbound MHC molecule has a structure that is defined by the allele (e.g. HLA-A, HLA-B or HLA-C for MHC class I molecules) giving rise to the unbound MHC molecule. pMHC
[0225] As used herein the term "pMHC" refers to a peptide-MHC complex comprising an MHC molecule that binds and presents peptide antigens. In particular, the pMHC may be an MHC class I molecule that binds and presents a peptide antigen.
[0226] These peptide antigens are also referred to herein as bound peptide(s), when they are referred to as part of a pMHC.
[0227] Target pMHC
[0228] A target pMHC is a pMHC that binds and presents a target peptide antigen, where the target peptide antigen has been chosen as a target for generation of pMHC binders that are capable of binding the target pMHC (target peptide-MHC complex). pMHC binders
[0229] As used herein "pMHC binder" refers to a protein binding to MHCs that are presenting specific peptides (pMHCs). In particular, a pMHC binder may be a protein that binds to the peptide-MHC complex through binding to both the peptide part of the complex and binding to the MHC part of the complex. A pMHC binder may also be a protein that binds to the a1 and a2 peptide-binding groove and the specific peptide antigen of an MHC class I molecule.
[0230] Structure of pMHC binders
[0231] The pMHC binder may further comprise at least 80 % helical and / or beta-sheet structure and at most 20% random coil structure. Thus, in embodiments, the pMHC binder may comprise at least 80% helical and / or beta-sheet structure, such as at least 82%, 84%, 86%, 88%, 90%, 92% or 94% helical and / or beta-sheet structure. In embodiments, the pMHC binder may comprise at least 86% helical and / or beta-sheet structure. In embodiments, the pMHC binder may comprise at least 88% helical and / or beta-sheet structure. In embodiments, the pMHC binder may comprise at least 90% helical and / or beta-sheet structure. In embodiments, the pMHC binder may comprise at least 92% helical and / or beta-sheet structure. In embodiments, the pMHC binder may comprise at least 94% helical and / or betasheet structure.
[0232] In embodiments, the pMHC binder comprises a helix bundle comprising at least 3 helices and at least 80% helical and / or beta-sheet structure and at most 20% random coil structure.
[0233] In embodiments, the pMHC binder comprises a helix bundle comprising at least 3 helices and at least 80% helical and / or beta-sheet structure such as at least 82%, 84%, 86%, 88%, 90%, 92% or 94% helical and / or beta-sheet structure.
[0234] Preferably, N- and C-terminal of the pMHC binder are oriented so that at least one of the N- and C-terminal can be labelled with a protein and / or chemical tag.
[0235] In embodiments, the pMHC binder comprises a helix bundle comprising at least 3 helices, at least 80% helical and / or beta-sheet structure such as at least 82%, 84%, 86%, 88%, 90%, 92% or 94% helical and / or beta-sheet structure.
[0236] The pMHC binder may be less than 200 amino acids long, such as less than 175, 150, 125, or such as less than 100 amino acids long. In embodiments, the pMHC binder may be less than 200 amino acids long. In embodiments, the pMHC binder may be less than 175 amino acids long. In embodiments, the pMHC binder may be between 100 and 200 amino acids long. In embodiments, the pMHC binder may be between 100 and 200 amino acids long.
[0237] In embodiments, the pMHC binder may be between 125 and 175 amino acids long. In embodiments, the pMHC binder may be between 110 and 150 amino acids long.
[0238] In embodiments, the pMHC binder may be between 115 and 140 amino acids long. In embodiments, the pMHC binder may be between 125 and 150 amino acids long. In embodiments, the pMHC binder may be between 50 and 200 amino acids long.
[0239] According to the methods as disclosed herein pMHC binders were designed against pMHC SIINFEKL / H2-Kb, SLLMWITQC / HLA-A*02:01 , RVTDESILSY / HLA-A*01 :01 , NLFRRVWEL / HLA-B*08:01, and NLVPMVATV / HLA-A*01 :01. The amino acid and / or nucleic acid sequence of these pMHC binders are listed in SEQ ID NOs: 96-195, SEQ ID NOs: 19- 95, SEQ ID NOs: 296-325, SEQ ID NOs: 196-295, SEQ ID NOs: 435-503.
[0240] SEQ ID NOs: 96-195 are pMHC binders designed against SIINFEKL / H2-Kb, SEQ ID NOs: 19-95 are pMHC binders designed against SLLMWITQC / HLA-A*02:01 , SEQ ID NOs: 296- 325 are pMHC binders designed against RVTDESILSY / HLA-A*01 :01 , SEQ ID NOs: 196-295 are pMHC binders designed against SEQ ID NOs: 196-295.
[0241] In embodiments, the pMHC binder is selected as one or more from the list consisting of SEQ ID NOs: 96-195, SEQ ID NOs: 19-95, SEQ ID NOs: 296-325, SEQ ID NOs: 196-295, SEQ ID NOs: 435-503.
[0242] In one or more interesting embodiments, the pMHC binder is selected as one or more from the list consisting of SEQ ID NOs 19-95. In one or more interesting embodiments, the pMHC binder is selected as one or more from the list consisting of SEQ ID NOs: 96-195. In one or more interesting embodiments, the pMHC binder is selected as one or more from the list consisting of SEQ ID NOs: 196-295. In one or more interesting embodiments, the pMHC binder is selected as one or more from the list consisting of SEQ ID NOs: 296-325.
[0243] In one or more interesting embodiments, the pMHC binder is selected as one or more from the list consisting of SEQ ID NOs: 326-425. In one or more interesting embodiments, the pMHC binder is selected as one or more from the list consisting of SEQ ID NOs: 435-503.
[0244] In one or more interesting embodiments, the pMHC binder is selected as one or more from the list consisting of SEQ ID NOs: 670-761. In embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 96-195, SEQ ID NOs: 19-95, SEQ ID NOs: 296-325, SEQ ID NOs: 196-295, SEQ ID NOs: 435-503, SEQ ID NOs: 556-667 and SEQ ID NOs: 670-761.
[0245] In embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 96-195, SEQ ID NOs: 19-95, SEQ ID NOs: 296-325, SEQ ID NOs: 196-295, SEQ ID NOs: 435-503, and SEQ ID NOs: 556-667.
[0246] In embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 96-195, SEQ ID NOs: 19-95, SEQ ID NOs: 296-325, SEQ ID NOs: 196-295, SEQ ID NOs: 435-503, and SEQ ID NOs: 556-664.
[0247] In exemplary embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 96-195, SEQ ID NOs: 19-95, SEQ ID NOs: 296-325, SEQ ID NOs: 196-295, and SEQ ID NOs: 435-503.
[0248] In particularly interesting embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs 19-95. In particularly interesting embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 96-195. In particularly interesting embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 196-295. In particularly interesting embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 296-325. In particularly interesting embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 326-425. In particularly interesting embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 435-503. In particularly interesting embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 556-664. In particularly interesting embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list consisting of SEQ ID NOs: 670-761 . In one or more exemplary embodiments, the pMHC binder has at least 90% identity to one or more sequences selected from the list of SEQ ID NOs: 13-18, 22, 24-31 , 33, 35-38, 41- 42, 45-46, 48, 50, 54-55, 58-60, 62-63, 67-68, 71-72, 74-82, 84-87, and 95.
[0249] In one or more exemplary embodiments, the pMHC binder has at least 90% identity to at least one of SEQ ID NOs: 76, 302, 341 , 665, 666 and 667.
[0250] In one or more exemplary embodiments, the pMHC binder has at least 90% identity to at least one of SEQ ID NOs: 76, 302 and 665.
[0251] In one or more exemplary embodiments, the pMHC binder has at least 90% identity to at least one of SEQ ID NOs: 341 , 666 and 667.
[0252] In one or more exemplary embodiments, the pMHC binder has at least 90% identity to SEQ ID NO: 76.
[0253] In one or more exemplary embodiments, the pMHC binder has at least 90% identity to SEQ ID NO: 302.
[0254] In one or more exemplary embodiments, the pMHC binder has at least 90% identity to SEQ ID NO: 665.
[0255] In one or more exemplary embodiments, the pMHC binder has at least 90% identity to SEQ ID NO: 341.
[0256] In one or more exemplary embodiments, the pMHC binder has at least 90% identity to SEQ ID NO: 666.
[0257] In one or more exemplary embodiments, the pMHC binder has at least 90% identity to SEQ ID NO: 667.
[0258] Table 2 - Top 50 in vitro tested pMHC binders designed against SLLMWITQC / HLA-A*02:01
[0259] Table 3 - Overview of the sequences of the present disclosure
[0260] Sequence identity
[0261] The term "sequence identity" as used herein describes the relatedness between two amino acid sequences or between two nucleotide sequences, i.e. , a candidate sequence (e.g., a sequence of the invention) and a reference sequence (such as a prior art sequence) based on their pairwise alignment. For purposes disclosed herein, the sequence identity between two amino acid sequences is determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mo / . Biol. 48: 443-453) as implemented in the Needle program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16: 276-277,), preferably version 5.0.0 or later (available at https: / / www.ebi.ac.uk / Tools / psa / emboss needle / ). The parameters used are gap open penalty of 10, gap extension penalty of 0.5, and the EBLOSUM62 (EMBOSS version of 30 BLOSUM62) substitution matrix. The output of Needle labeled "longest identity" (obtained using the -nobrief option) is used as the percent identity and is calculated as follows: (Identical Residues x 100) / (Length of Alignment - Total Number of Gaps in Alignment).
[0262] For purposes disclosed herein, the sequence identity between two nucleotide sequences is determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1 970, supra) as implemented in the Needle program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16: 276- 277), 10 preferably version 5.0.0 or later. The parameters used are gap open penalty of 10, gap extension penalty of 0.5, and the DNAFULL (EMBOSS version of NCBI NUC4.4) substitution matrix. The output of Needle labelled "longest identity" (obtained using the - nobrief option) is used as the percent identity and is calculated as follows: (Identical Deoxyribonucleotides x 100) / (Length of Alignment — Total Number of Gaps in Alignment).
[0263] Putative pMHC binders
[0264] In the present context, the term putative pMHC binder(s) is used to refer to one or more pMHC binders and / or sets of binders that have been identified through a generative protein binder design process as putative binders of a target pMHC but is not yet sufficiently optimized to have been selected for downstream testing, e.g. in vitro or in vivo testing. pMHC binder candidates
[0265] In the present context, the term “pMHC binder candidate(s)” is used to refer to one more pMHC binders and / or sets of pMHC binders that have been identified through a generative protein binder design process and selected following an optimization step as defined in the method disclosed herein. In particular, pMHC binder candidates within the present context may be sequences that are outputted in step c of the method disclosed herein. Depending on the level of testing performed as part of step iii), a pMHC binder candidate may have a pLDDT of >90% against a target pMHC, an ipAE of <7A against a target pMHC, an ipAE of >10A against unbound MHCs, and / or an ipAE of >10A against variant pMHCs.
[0266] In one or more embodiments, a pMHC binder candidate has a pLDDT of >90%, and an ipAE of <7A against a target pMHC.
[0267] In one or more embodiments, a pMHC binder candidate has a pLDDT of >90% against a target pMHC, an ipAE of <7A against a target pMHC, and an ipAE of >10A against variant pMHCs.
[0268] In one or more embodiments, a pMHC binder candidate has a pLDDT of >90% against a target pMHC, an ipAE of <7A against a target pMHC, and an ipAE of >10A against unbound MHCs.
[0269] In one or more embodiments, a pMHC binder candidate has a pLDDT of >90% against a target pMHC, an ipAE of <7A against a target pMHC, an ipAE of >10A against unbound MHCs, and an ipAE of >10A against variant pMHCs. pMHC variants
[0270] A pMHC variant in the present context is a pMHC that has a similar, but not identical sequence compared to a target pMHC. Thus, a pMHC variant may be a peptide-MHC complex, wherein the bound peptide sequence of the variant pMHC has one or more amino acid substitutions compared to the bound peptide sequence of the target pMHC. A pMHC variant may also be a peptide-MHC complex, wherein the MHC allele of the variant pMHC is different from the MHC allele of the target pMHC (e.g., HLA-A vs HLA-B).
[0271] TCR
[0272] In the present disclosure, TCR is used as short for T cell receptor. T cell receptors are protein complexes that are found on the surface of T cells- TCRs are generated through V(D)J recombination in thymocytes during T cell development and are responsible for recognition of one or more peptide antigens bound to MHC molecules (pMHCs).
[0273] CAR
[0274] In the present disclosure, CAR is used as short for chimeric antigen receptor. Chimeric antigen receptors are in the present context understood as synthetic that are designed to bind specific antigens, such as certain proteins on the surface of cancer cells.
[0275] Soluble Blocker In the present disclosure, a soluble blocker refers to a biologically active molecule that inhibits the interaction between specific cellular receptors and their ligands. These blockers are soluble in bodily fluids and can circulate through the bloodstream to reach their target sites. Soluble blockers may be used to modulate immune responses or interfere with pathological processes by preventing the binding of natural ligands to their corresponding receptors (e.g. TCR-pMHC interactions)
[0276] Bispecific T cell engager (BITe)
[0277] In the present disclosure, BiTe refers to bispecific T cell engager. Bispecific T cell engagers are a class of engineered protein constructs (usually antibody constructs) designed to simultaneously bind to a T cell receptor and a tumor-associated antigen. By linking these two targets, BiTes facilitate the direct engagement and activation of T-cells to attack and eliminate cancer cells.
[0278] TCR-Like Molecule
[0279] In the present disclosure, TCR-like molecule refers to an engineered protein designed to recognize specific pMHC complexes on the surface of cells, similar to how natural TCRs identify antigens. Although it targets the same pMHC complexes as TCRs, the binding modality of a TCR-like molecule does not necessarily mimic that of a TCR.
[0280] TCR-Fusion Construct
[0281] In the present disclosure, TCR-fusion construct refers to a recombinant protein that combines a TCR or a TCR-like molecule with another functional domain, such as an effector protein, cytokine, or signalling molecule. These constructs are engineered to enhance the therapeutic potential of TCRs or TCR-like molecules by directing their activity, improving stability, or increasing their ability to modulate immune responses.
[0282] Drug-conjugates
[0283] In the present disclosure, drug-conjugates are therapeutic agents that consist of a cytotoxic drug linked to a targeting molecule, such as an antibody or peptide, that specifically binds to a target cell or receptor. The conjugate allows for the selective delivery of the cytotoxic drug to diseased cells, thereby minimizing systemic toxicity and improving therapeutic efficacy. Validation of pMHC binder candidates
[0284] The pMHC binder candidates identified in the method disclosed herein may be validated in vitro by various standard ligand binding assays. These ligand binding assays can be used to assess the affinity of a pMHC binder candidate to the target pMHC. A non-limiting example of such a ligand binding assay is an SPR assay.
[0285] Large numbers of pMHC binder candidates (e.g., candidate pMHC libraries) can be screened and validated rapidly using a combination of yeast display, or mammalian display and ligand binding assays. A particularly desirable method for screening and validation is a combination of a yeast display, or mammalian display technique combined with an SPR assay. One way of performing such a method is described in examples 2 and 3.
[0286] Affinity
[0287] A high affinity pMHC binder, as used herein, refers to a pMHC binder candidate that has a higher affinity towards the target pMHC and lower affinity against one of or both of a variant pMHC and unbound MHC.
[0288] The high affinity pMHC binder may have an in vitro affinity as measured by, e.g., surface plasmon resonance (SPR) against the target pMHC that is at least 2-fold higher, 1.75-fold higher, or 1.5-fold higher than against one of a variant pMHC and / or an unbound MHC. Additional methods for measuring the interactions affinity will be evident to the skilled person.
[0289] In one or more embodiments, the high affinity pMHC binder may have an in vitro affinity as measured by SPR against the target pMHC that is at least 2-fold higher, 1 .75-fold higher, or 1 .5-fold higher than against one of a variant pMHC and / or an unbound MHC. In one or more embodiments, the high affinity pMHC binder may have an in vitro affinity as measured by SPR against the target pMHC that is at least 2-fold higher, 1 .75-fold higher, or 1.5-fold higher than against one of a variant pMHC. Alternatively, the high affinity pMHC binder may have an in vitro affinity as measured by SPR of towards the target pMHC of less than 0.001 , 0.01 , 0.05, 0.1 , 0.7, 0.75. 0.8, 0.85, 0.9, 0.95, or less than 1 pM. Additionally, the high affinity pMHC binder may have an in vitro affinity as measured by SPR of towards the target pMHC of between 0.001-1 , 0.01-1 or 0.05-1. Preferably the affinity is less than 1 pM, 100nM, 10nM, 1 nM, 500pM or less than 10OpM.
[0290] In one or more embodiments, the high affinity pMHC binder may have an in vitro affinity as measured by SPR against the target pMHC that is at least 2-fold higher, 1 .75-fold higher, or 1 .5-fold higher than against unbound MHC. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR against the target pMHC that is at least 1 .5-fold higher than the pMHC binder affinity against to a nontarget or unbound pMHC. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR against the target pMHC that is at least 1.75-fold higher than the pMHC binder affinity against to a non-target or unbound pMHC. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR against the target pMHC that is at least 2-fold higher than the pMHC binder affinity against to a non-target or unbound pMHC. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 1 pM. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 0.5 pM. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 0.3 pM. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 0.1 pM. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 0.01 pM. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 0.007 pM. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 0.005 pM. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 0.01-1 pM. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 0.05-1 pM. In one or more embodiments, a high affinity pMHC binder is a pMHC binder that has an in vitro affinity as measured by SPR towards the target pMHC of less than 0.1-1 pM.
[0291] Nucleic acid construct
[0292] The present disclosure further relates to a nucleic acid expression vector or a nucleic acid sequence encoding a polypeptide construct comprising nucleic acids encoding
[0293] -a pMHC binder candidate as defined herein,
[0294] -a polypeptide membrane anchor, -a linker connecting said pMHC binder candidate and said membrane anchor, and -a surface expression tag.
[0295] The expression vector and / or nucleic acid further comprises an expression element such as a promoter. In one or more embodiments, the polypeptide construct is under the control of the expression element. In one or more embodiments, the polypeptide construct is under the control of one or more promoters.
[0296] The membrane anchor may further be linked to an intracellular signalling domain. Thus, in one or more embodiments, the polypeptide membrane anchor is linked to an intracellular signalling domain.
[0297] Production methods
[0298] The present disclosure further relates to methods for producing the pMHC binder candidate amino acid sequence(s) as disclosed herein. These pMHC binder candidate may be produced biosynthetically by expression of the outputted pMHC binder amino acid sequence or a nucleic acid construct as defined herein in a recombinant expression system. Alternatively, the pMHC binder candidate can be produced synthetically by protein synthesis. The pMHC binder candidate may also be produced by a combination of synthetic and biosynthetic production methods. After production of the pMHC binder candidates, it may be advantageous to employ one or more purification methods. Such methods are generally known in the prior art.
[0299] In one or more embodiments, the pMHC binder may be produced in host cell using a nucleic acid construct as defined herein. In one or more embodiments, the pMHC binder is produced using semi-synthesis. In one or more embodiments, the pMHC binder is synthetically produced.
[0300] Use of pMHC binders
[0301] The pMHC binders identified through the methods disclosed herein are expected to find use in the many applications making use of immunorecognition. These uses fall within fields such as immunodetection, diagnostics, and therapeutics such as cancer treatment and immunotherapies. It is envisioned that the pMHC binders identified can be used as soluble blockers of pMHC-TCR interactions, CARs, TCR-like molecules, TCR-like fusion constructs, BiTes, and drug-conjugates. The pMHC binders identified through the methods as disclosed herein may be advantageous compared to soluble blockers, CARs, TCR-like molecules, TCR-like fusion constructs, BiTes, and drug-conjugates developed through other methods, as the present binders can be proactively screened for cross-reactivity and therefore lead to development of safer drugs, therapies, and diagnostics with less side effects.
[0302] Medical use
[0303] In one or more embodiments, the present disclosure relates to a pMHC binder identified according to the methods defined herein for use as a medicament. In one or more embodiments, the present disclosure relates to a pMHC binder identified according to the methods defined herein for use in the treatment of cancer. In one or more embodiments, the present disclosure relates to a pMHC binder identified according to the methods defined herein for use in immunotherapy. In one or more embodiments, the present disclosure relates to a pMHC binder identified according to the methods defined herein for use in CAR T cell therapy. In one or more embodiments, the present disclosure relates to a pMHC binder identified according to the methods defined herein for use in the treatment of immune disorders. In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375, 435- 503, 665-667 and 670-761 for use as a medicament. In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375, 435-503 and 665-667 for use as a medicament. In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375, 435-503 for use as a medicament. In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375, 435- 503, 665-667 and 670-761 for use in the treatment of cancer. In one or more embodiments, the present disclosure relates to a pMHC binder the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375 and 665-667 for use in the treatment of cancer. In one or more embodiments, the present disclosure relates to a pMHC binder the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375 for use in the treatment of cancer. In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375, 435-503, 665-667 and 670-761 for use in immunotherapy. In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375 and 665-667 for use in immunotherapy. In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375 for use in immunotherapy. In one or more embodiments, the present disclosure relates to a pMHC binder identified according to the methods defined herein attached to an antibody, another pMHC binder, or other binding molecule, to generate a Bi-specific binder. In one or more embodiments, the present disclosure relates to a pMHC binder identified according to the methods defined herein attached to another pMHC binder or another class of binder to generate bi-specific binding molecules.
[0304] In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375 and 665-667 for use in CAR-T cell therapy.
[0305] In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375, 665-667 and 670- 761 for use in CAR-T cell therapy.
[0306] In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 19-95, 196-375 for use in CAR-T cell therapy.
[0307] In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 326-375 for use in CAR-T cell therapy.
[0308] In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 13-18, 22, 24-31 , 33, 35- 38, 41-42, 45-46, 48, 50, 54-55, 58-60, 62-63, 67-68, 71-72, 74-82, 84-87, and 95 for use in CAR-T cell therapy.
[0309] In one or more embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 670-761 for use in CAR-T cell therapy. In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 76, 302, 665, 341 , 666 and 667 for use in CAR-T cell therapy.
[0310] In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 76, 302 and 665 for use in CAR-T cell therapy.
[0311] In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 341 , 666 and 667 for use in CAR-T cell therapy.
[0312] In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 76 and 341 for use in CAR-T cell therapy.
[0313] In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 302 and 667 for use in CAR-T cell therapy.
[0314] In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to any one of SEQ ID NOs: 665 and 666 for use in CAR-T cell therapy.
[0315] In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to SEQ ID NO: 76 for use in CAR-T cell therapy.
[0316] In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to SEQ ID NO: 302 for use in CAR-T cell therapy.
[0317] In one or more exemplary embodiments, the present disclosure relates to a pMHC binder that has at least 90% sequence identity to SEQ ID NO: 665 for use in CAR-T cell therapy. Diagnostic use
[0318] Additionally, as briefly mentioned above, it is envisioned that the pMHC binders identified according to the method defined herein may be used a diagnostics tool. In particular it is envisioned that the pMHC binders may find use in immunoassays, immunoprecipitation assays and immunology-based isolation methods. Thus, the pMHC binders identified herein are envisioned for use in lateral flow assays, flow diagnostics, methods for isolating antigen presenting cells for future potentiation and differentiation, and genotyping assays.
[0319] In one or more embodiments, the pMHC binders identified according to the method defined herein may be of use in diagnostics if they are tagged with a detectable moiety. The detectable moiety can be selected from the list consisting of dyes, fluorescent tags, fluorophores and / or radionuclides.
[0320] It should be understood that any feature and / or aspect discussed above in connections with the compounds according to the invention apply by analogy to the methods described herein.
[0321] The following figures and examples are provided below to illustrate the present invention.
[0322] They are intended to be illustrative and are not to be construed as limiting in any way.
[0323] EXAMPLES
[0324] Example 1 - De novo generation of putative pMHC binders
[0325] Aim
[0326] The present examples aim to provide solutions to the problem of efficiently identifying and optimizing efficient binders towards pMHC complexes, thereby greatly improving the time spend on development of putative pMHC binders. Accordingly, the present example discloses the development of an in-silico platform for the generation of pMHC binders, in silico cross-reactivity screening to minimize risk of off-target effects and multi-objective Bayesian optimization to in silico iteratively refine binders.
[0327] Methods and materials
[0328] Class I pMHC Target preparation
[0329] For a given target class I pMHC the crystal structure was retrieved from the Protein Databank (PDB) website. When the crystal structure comprises one or more iterations of the structure of a single pMHC, all iterations were removed, so that only the crystal structure of a single pMHC remained. This single crystal structure was then further modified by removal of the Beta-2-microglobulin (P2 microglobulin) subunit and the alpha-3 (a3) domain of the MHC to reduce the inference time of RFdiffusion. In cases where no experimental structure was available, the pMHC structure was generated using AF2 (Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589 (2021) using the pMHC sequence as the input, as well as the .pdb of the most similar sequence and experimentally resolved pMHC complex as a template.
[0330] Generating binder backbones
[0331] RFdiffusion (Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089-1100 (2023); https: / / github.com / RosettaCommons / RFdiffusion) was used to generate the putative pMHC binders. In the generation the target region was specified with / O as a chain break to ensure the generated binder is a separate domain to the target. The binder length was defined as 100 to 150 amino acids. The script randomly picks a number within this range during each iteration, to yield a backbone of the selected length.
[0332] In the initial design several instances of the peptide residues of the pMHC targets were considered as hotspots that the designed binder should excerpt binding interaction to. For all target pMHCs, as defined in table 5 and Table , all peptide positions were defined as hotspots. For the target pMHCs RVTDESILSY / HLA-A*01:01 and NLFRRVWEL / HLA- B*08:01 the non-anchor residues, defined as not being involved in direct binding interactions towards the MHC, position 3-7 and 1-4 & 6-8 of the concerned peptide, respectively, were identified and employed as hotspots for the initial design of pMHC binders.
[0333] Generative sequence design
[0334] Once the backbones were generated using both hot and non-hotspot-based RFdiffusion, 1- 10 sequences were designed for each of the backbones using ProteinMPNN (Bennett, N. R. et al., Improving de novo protein binder design with deep learning, Nat. Common. 14, 2625, 2023). The sampling temperature is typically set to 0.1 but often varied and is typically between 10'4- 10’1. Subsequently, AF2 was employed to predict the complexes and calculate key scoring metrics, including the ipAE and pLDDT scores. A set of exemplary ProteinMPNN parameters used to generate the pMHC binders are provided in Table 4.
[0335] Table 4 - Example of parameters used for sequence design with ProteinMPNN.
[0336] Structure refinement
[0337] To further refine the binders, partial diffusion (Watson, J. L. et al., De novo design of protein structure and function with RFdiffusion, Nature 620, 1089-1100 (2023); Vazquez Torres, S. et al., De novo design of high-affinity binders of bioactive helical peptides, Nature 626, 435- 442 (2024)) (https: / / github.com / RosettaCommons / RFdiffusion) was employed to introduce diversity to backbone folds, in order to affinity mature promising pMHC binders by introducing minor changes. The input binder structures were noised up using a non-random starting point distribution with a user-specified time step instead of completing the full noising schedule, resulting in denoised structures that are structurally similar to the input. The putative pMHC binders were subjected to a variation of (10, 20, 30, and 40) noising time steps out of a total of 200 times steps in the noising schedule, and subsequently denoised. Approximately 5 partially diffused backbones were generated for each of the originally selected putative pMHC binders. The resulting putative pMHC binders then were used in a second sequence design step ProteinMPNN, followed structure prediction based on the sequence by AF2 with an initial guess (using the original binder-pMHC pdb complex); the initial guess increases the accuracy and speed of the prediction. The resulting binders were then filtered on from the AF2 metrics, and pMHC binders with a complex ipAE of 7 or below, and a complex pLDDT of 90 or above, were considered as suitable candidates for pMHC binders.
[0338] Following selection of the top candidates by the ipAE and pLDDT metrics, the selected binders were further filtered based on the interactions they were predicted to form with the peptide. To do so the top binder designs were both visually and automatically inspected to analyze the binder orientation and number and type of interaction points, N- and C-terminus accessibility and maximum number of contacts to the peptide.
[0339] Cross-reactivity screening
[0340] To minimize potential off-target interactions, it was investigated whether the binders were highly specific in only binding to the target peptide (ipAE) and interact poorly with modified (c.f. below) peptides (higher ipAE score).
[0341] Specifically, peptides that were similar (identity >85%) in sequence to the target peptide and have a strong binding affinity to the target pMHC were identified. To obtain the peptides to screen against, libraries of single- and double-point mutated peptides were generated. Point mutations are only applied to desired positions, which are exposed and have residues that interact with the binder surface. MHC anchor residue positions were avoided since they are not exposed and do not interact with the binder. To deselect peptides which are likely to be presented by healthy human cells the libraries were blasted using BLAST (https: / / blast.ncbi.nlm.nih.gov / Blast.cgi) against the known human proteome, using a high query coverage as a threshold. Peptide matches shorter than eight amino acids were embedded in the sequence of the originating protein sequence to obtain sequences of specifically eight or nine amino acids in length.
[0342] Once the peptides of the human proteome was filtered out, the predicted affinity of the mutated peptides to the target MHC was then determined (affinity prediction of peptide MHC affinities may e.g., be performed using NetMHCpan4.1 (Reynisson, B., Alvarez, B., Paul, S., Peters, B. & Nielsen, M., NetMHCpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data, Nucleic Acids Res. 48, W449-W454 (2020)), where peptides with weak affinities (e.g., weak binder as defined in NetMHCpan4.1 ) were considered unlikely to bind, and were therefore deselected, leaving only a few peptides to be analyzed further).
[0343] The remaining peptides were inserted into the pMHC structures, and the key metrics of the resulting complex was analyzed using AF2 to obtain the ipAE and pLDDT scores of the mutated pMHC structures.
[0344] To further enhance the in silico specificity screening, a secondary approach was introduced which allowed mapping of the interacting amino acid position of the peptide for each selected pMHC binder. Single-point mutants of the target peptide were generated and screened for their effect on the predicted ipAE score. Keeping the anchoring positions unchanged, the other positions were consecutively substituted with all canonical amino acids. The generated single-point mutant pMHC variants were modeled with AF2 and superimposed with the selected pMHC binders and subsequently as complex remodeled by AF2 within the RFdiffusion pipeline as described above. By aligning the predicted ipAE scores, the peptide positions and the favorable physicochemical properties of the amino acids that were predicted to be bound by each pMHC binder was accurately mapped, resulting in a primary risk assessment of cross-reactivity and, thereby, off-target binding.
[0345] Bayesian optimization
[0346] An iterative loop involving Bayesian optimization and AF2 was used to refine binder sequences for target pMHC specificity. Specifically, the data collected in the section above is used as input information, i.e. ipAE and pLDDT against empty and non-target pMHCs. The aim was to optimize ipAE and pLDDT against the target pMHC and drive towards worse scores against all non-target pMHCs and the empty target MHC. Thus, the binder sequences were subjected to multi-objective Bayesian optimization (MOBO) with guided diffusion. As output new putatively optimized pMHC binder sequences were produced. These sequences underwent structural and interaction prediction using AF2, focusing on their ipAE and pLDDT scores relative to each target type. These data then guided subsequent MOBO optimization rounds, creating a feedback loop that iteratively refines the binder sequences. This iterative optimization cycle was designed to incrementally improve the binder sequences in terms of their predicted specificity. Results
[0347] Generation of pMHC binders
[0348] About 40,000 binders targeting five different pMHC targets was generated as described above (see Table ).
[0349] For the melanoma neoantigen peptides RVTDESILSY / HLA-A*01 :01 and NLFRRVWEL / HLA-B*08:01 , the complex of the peptide with the MHC was modeled using AF2, while the structures of the SIINFEKL / H2-Kb, and SLLMWITQC / HLA-A*02:01 complexes were available from the pMHC:T-cell receptor co-crystal structures (see Table 55).
[0350] Table 5 - Selected pMHC targets
[0351] “protein structures may be retrieved using the provided PDB identifiers at: https: / / www.rcsb.org
[0352] Design process and selection
[0353] For each of the pMHCs, at least 2000, or 5500 in the case of the SLLMWITQC / HLA-A*02:01 complex, binder structures were generated using RFdiffusion as described under Material and Methods. For each backbone structures (as defined above), four amino acid sequences of putative pMHC binder were predicted using proteinMPNN (as described above) and scored by AF2 within the dl-binder_design module. The putative pMHC binders designs were reinterpreted into structures using AF2, and the key metrics i.e., the ipAE of the pMHC:binder complex and the folding confidence score (pLDDT) of each of the binders was extracted. All binders with an ipAE score below 12 A and a binder pLDDT score above 88% were selected for further optimization (Fig. 1 b and Fig. 1c). In the case of the target pMHC SLLMWITQC / HLA-A*02:01 , the subpopulation of selected designs surpassing the ipAE and binder pLDDT were further filtered for unique initial RFdiffusion designs to maximize the structural diversity, which resulted in a set of 109 binder designs (SEQ ID NOs: 556-664). Following the initial selection, the designed binders for each of the four peptides were subjected to a structure and / or sequence diversification.
[0354] The sequence diversification using ProteinMPNN and subsequent reinterpretation into structures by AF2 within the dl-binder_design module is a powerful method of sampling a high number of structurally very similar (in the protein backbone structure nearly identical) designs and, thus, identifying a variant with slightly better ipAE and binder pLDDT. A method especially applicable if the initial design and selection process could identify a set of binder designs (typically >100) that is larger than the final set of binders that would be tested in vitro. Since the initial design and selection process resulted in a high structural variety of >100 binder designs for the pMHC SLLMWITQC / HLA-A*02:01 , only sequence diversification was performed. At the default sampling temperature, 100 sequences were generated per initial binder design and reinterpreted into structures by AF2, whereafter the ipAE and pLDDT were extracted as described above. The best scoring variant for each initial binder design was identified and the subpopulation filtered based on an ipAE cut-off of 7 A and a pLDDT cut-off of 92% (Fig. 1d and Fig. 1e) which resulted in a subset of 77 pMHC binders.
[0355] In general, structural diversification is particular suited when the initial design and selection process culminated in a subset of designs smaller than the desired set of pMHC binders to be tested in vitro. In the particular case of the SIINFEKL / H2-Kb pMHC binder 3p9l_beta_429_dl_design_1 , structural diversity was introduced to obtain structurally similar but not identical variants of the subset of selected designs. While a score of >16 A (Fig. 2b) is not particularly desirable, the design was included because of its uniquely high content of beta sheets as secondary structure elements (refer to table 7 for comparison). In total 100 partially diffused structures were generated by RFdiffusion (at default sampling temperature and 20 diffusion time steps). Following the structural diversification, the variants were retranslated into protein sequence by ProteinMPNN with 3 sequences per variant and reinterpreted into structures by AF2, whereafter the ipAE and pLDDT were extracted as described above. The variant of 3p9l_beta_429_dl_design_1 (Fig. 2c and SEQ ID NO: 377) with the lowest ipAE of 9.58 A was selected and included in the final subset of pMHC binders tested in vitro. Table 6 - Overview of binder designs against four different pMHC targets.
[0356] Cross-reactivity screening
[0357] The selected 77 pMHC binders for the SLLMWITQC / HLA-A*02:01 complex were evaluated for their predicted binding capabilities against the MHC containing no or mutated variants of the target peptide as follows.
[0358] Firstly, all 3610 conceivable single and double-point mutated peptides were generated (while keeping the positions 1-3 and 9 unchanged, as described above). The first strategy employed entailed blasting of all 3610 peptides against the human proteome using NCBI pBlast (as described above under material as methods) which led to the identification of 16 hits with a length between 5 and 10 amino acids (at a 85% identity). Unsurprisingly, the predominant match was the target peptide itself (Fig. 2a). The MHC affinity of the 16 peptides was predicted (e.g., using NetMHCpan4.1) and only two peptides (SLLGNILRI and SLLGNILRII) scored equally well or better than the original peptide (SLLMWITQC). As can be seen the peptide “SLLGNILRI” is encompassed in the “SLLGNILRII” peptide, and accordingly only the latter was included in the following crossreactivity screening. As a second strategy of cross-reactivity screening all 3610 peptides were subjected to MHC affinity prediction (e.g., using NetMHCpan4.1 , MHCFurry 2.0, SMM, or similar tools (Kim, Y. et al., Derivation of an amino acid similarity matrix for peptide:MHC binding and its application as a Bayesian prior, BMC Bioinformatics 10, 2009)). Of the 3610 peptides, 1128 peptides received better ranking (NetMHCpan4.1) than the original peptide, indicating these were more likely to be presented on the HLA-A*02:01 complex. Overlaying the sequences (see figure 2b) it became apparent the residue in position 4, and to a lesser extend also the residues in positions 5, 7, and 8 were the least conserved residues with no or only a marginally dominating amino acid (Fig. 2b). To reduce the requirement for computational resources only six single- or double-point mutated peptides were chosen for structural modeling and subsequent cross-reactivity screening resulting in a total combination of 624 structures ((6 mutated peptides from strategy B, 1 peptide from strategy A, 1 remodeled original peptide) x 1 HLA-A*02:01 complex x 77 selected pMHC binders).
[0359] After remodeling and scoring the pMHC binders using the AF2 within the RFdiffusion pipeline, a set of 18 pMHC binders (SEQ-ID: 25, 30, 41 , 42, 45, 46, 48, 50, 58, 59, 63, 67, 76, 79, 81 , 85, 86, 95) passed an ipAE cut-off >25 or <15 for the mutated pMHC or the remodeled original pMHC, respectively (Fig. 2c).
[0360] It was observed that none of the resulting structures met the criteria of a ipAE < 10 A.
[0361] Lastly, based on the best ipAE score 26 (SEQ-ID: 22, 24, 27, 28, 29, 31 , 33, 35, 36, 37, 38, 54, 55, 60, 62, 68, 71 , 72, 74, 75, 77, 78, 80, 82, 84, 87) out of the 77 pMHC binders (not including the previously selected 18 supposedly non-cross-reactive pMHC binder), were selected for further in vitro validation (as described below).
[0362] Bayesian optimization of pMHC binders targeting the SLLMWITQC / HLA-A*02:01 complex Multi-objective Bayesian optimization (MOBO) was initially performed against three targets (SLLMWITQC / HLA-A*02:01, SLLAWITQC / HLA-A*02:01 , SLLMYITQC / HLA-A*02:01). As input, we provided ipAE (ipAE) scores for the top 20 binders against the SLLMWITQC / HLA- A*02:01 complex, as well as the scores against two-point mutants where position 4 and 5 were modified with amino acids that had similar properties (e.g., the substitution of tryptophan for tyrosine in position keeps the hydrophobic character while being slightly less bulky) specifically, the methionine in position 4 was mutated to an alanine for target SLLAWITQC / HLA-A*02:01 (SEQ ID NO 16), and in position 5 the tryptophan was substituted to a tyrosine for target SLLMYITQC / HLA-A*02:01 (SEQ ID NO 17). Generally, positions 4 and 5 were chosen as these displayed the highest degree of amino acid variance for the single-point mutations (as described above), that were still likely to be presented on the HLA-A*02:01 (Fig. 2b). The cross-reactivity screen demonstrated low ipAE scores against all three pMHC complexes, which is indicative of potential undesirable off- target binding. As a result, MOBO was employed to produce pMHC binder variants that would have a more favorable ipAE against SLLMWITQC / HLA-A*02:01 (preferably ipAE<10) and high ipAE against both mutants (preferably ipAE>10). As a result, six new pMHC binder variants were produced, which all achieved this showed an ipAE<10 towards
[0363] SLLMWITQC / HLA-A*02:01 and a ipAE>10 for the mutant peptides (Fig. 3). These six pMHC binders were all selected for in vitro characterization.
[0364] Structural properties pMHC binder
[0365] By visual inspection it was observed that the pMHC binders, independent of the specific target, was composed of multiple alpha helices with short linking regions (Fig. 4). Furthermore, the alpha helices generally aligned parallel to the flanking alpha-helices of the MHC and, thereby, encapsulate the peptide. The C- and N- termini were generally oriented parallel to the peptide and in proximity to each other while not part-taking in the direct interaction with the target peptide (See fig. 4).
[0366] To quantitatively assess the secondary structure configuration a Ramachandran plot was made for the selected pMHC binders by extracting the phi and psi angle between the protein backbone nitrogen atom and carbon alpha atom and carbon alpha atom and the carbon atom involved in the carboxyl group for each amino acid from all binder pdb structures, respectively (see Table 7). Utilizing the characteristic angle ranges for psi and phi, zones defining ‘beta sheets’, ‘right-handed alpha helices’ & ‘left-handed alpha helices’ was produced (see Table 7). For all pMHC binders, independent of the design campaign and the peptide they have been designed against, right-handed alpha helices were dominant in the secondary structure. Generally, about 91-97% of the amino acids in the pMHC binder were located in right-handed alpha helices (Table 7). Table 7 - Common properties ofpMHC binders surpassing the final filtering step in the design campaigns against fourpMHC complexes, i.e. SIINFEKL / H2-Kb (mouse model antigen from Hen Egg), SLLMWITQC / HLA-A*02:01 (Melanoma shared cancer antigen), RVTDESILSY / HLA-A*01 :01 (Melanoma neoantigen), and NLFRRVWEL / HLA-B*08:01 (Melanoma neoantigen).
[0367] Next, the orientation of the C-and N-termini of the pMHC binders was investigated by measuring the distance between the carbon alpha atoms of each terminus as well as the distance to the individual positions of the peptides. If the distance between the C- or N- terminus towards the peptide positions was increasing from the peptide’s C- to N-terminus, the orientation was classified as ‘dorsal’. If the distances decreased, the orientation would be classified as ‘frontal’. Lastly, if the distances led to indifferent conclusions, the orientation was classified as ‘lateral’, which entails every configuration of the C- and N-termini that is not strictly perpendicular to the peptide (Fig. 4).
[0368] Example 1.1 - Diversification of pMHC binder 2bnr-27_7_3_9 (SEQ ID NO: 341) after in vitro validation
[0369] To obtain a portfolio of refined variants of the in vitro verified pMHC binder 2bnr-27_7_3_9, SEQ ID NO: 341 against SLLMWITQC / HLA-A*02:01 , the AF2 prediction of the co-complex of the binder with SLLMWITQC / HLA-A*02:01 was employed as a blueprint for amino acid sequence as well as structure diversification. The diversification followed two aspects. One the one hand, solely the structure prediction of pMHC binder 2bnr-27_7_3_9 was only subjected to re-interpretation on the level of amino acid sequence using ProteinMPNN. One the other hand, in a second aspect the diversification also involved structural re-interpretation utilizing the previously described method of partial diffusion within RFdiffusion. For the sequence diversification, the initial binder-pMHC complex prediction was subjected to ProteinMPNN within the dl-design_binder pipeline to generate N=1000 sequences at each of the four arbitrary chosen sampling temperatures (Temp.: 0.01 , 0.03, 0.1 [default], 0.3). In the subsequent structure prediction step using AF2 initial guess, an ipAE and binder pLDDT score-based cut-offs (ipAE < 6 A and pLDDT binder > 90%)) were employed to deselect unsuitable designs. In a second branch of the pMHC binder 2bnr-27_7_3_9 refinement, the co-complex structure prediction (2bnr-27_7_3_9, SEQ ID NO: 341 against SLLMWITQC / HLA- A*02:01 ) was modified and re-interpreted by partial diffusion within RFdiffusion using T™ 10, 20 , 30 timesteps of denoising. The resulting structure predictions were further processed with ProteinMPNN (1 sequence per structure prediction) and finally re-folded and assessed by AF2 initial guess with slightly lower cut-off criteria (ipAE < 7 A and pLDDT binder > 85%). In summary, of the Nsumover all sampling temperatures = 4000 generated amino acid sequence-diversified variants of pMHC binder 2bnr-27_7_3_9, 257 binder designs passed the ipAE and pLDDT binder cuf-off. For the NSUm over all denoising stePs= 3000 structure- and sequence-diversified binder designs, 115 passed the cut-off criteria.
[0370] In a further aspect, since the selections made based on ipAE and pLDDT binder resulted in a number of binder designs beyond current capabilities of in vitro testing and to enhance chances of more binder designs revealing binding affinity to SLLMWITQC / HLA-A*02:01 , further filters were applied. Firstly, the binder predictions were filtered further to reduce the propensity of potential steric clashes (van der Waals radii overlap) in the structure predictions of the binder designs. For this purpose, a freely available pymol python script, referred to as TopModel, was locally installed and used to depict calculated clashes based on the input of the AF2 structure prediction and text output of the script. The identified clashes are artefacts of the AF2 (or also AF3, if used) structure prediction and might have benefitted an artificially low ipAE and pLDDT binder value for the particular binder design. To ensure the occurrence of binder designs with such artefacts is not enriched through the selection process with the aforementioned ipAE and pLDDT binder, TopModel was found to be a likely useful additional filter step. In short, for each binder design in a predicted co-complex with SLLMWITQC / HLA-A*02:01 , the atoms of the binder and pMHC in 5 A distance were determined and the theoretical van-der-Waals radii (inherent to the TopModel script) subtracted. If the sum of the radii plus a penalizing / buffer term of 2 A was larger than the distance of the atom pair, the atoms got depicted as a potential clash (dAB < TA + TB - 2 A, in which dAB is the distance of the atoms in the atom pair towards each other, and TA and TB are the radii of atom A and B).
[0371] Secondly, the designed variants were additionally filtered based on whether the residues potentially interacting with the peptide ‘SLLMWITQC’ of the pMHC SLLMWITQC / HLA-A*02:01 were conserved in comparison to initial in vitro confirmed pMHC binder 2bnr-27_7_3_9 (SEQ ID NO: 76). The identification of potential interaction was once again based solely on a 5 A cut-off between intermolecular atom pairs as previously described in the paragraph on TopModel. Accordingly, the residues in positions 11 (E, glutamic acid), 96 (E, glutamic acid), 107 (Y, tyrosine), and 119 (Q, glutamine) were considered.
[0372] After the additional filtering 92 pMHC binder designs against SLLMWITQC / HLA-A*02:01 were selected and subjected to in vitro assessment similar to the method described in example 2 (SEQ ID NOs: 670-761 AA, SEQ ID NOs: 762-853 DNA). Of these, 66 originated from the aspect of sequence diversification, while 26 are as well structurally diversified. 84% (77 / 92) of the pMHC binder designs have one or no substitutions at the four positions of 2bnr-27_7_3_9 that were considered as potentially interacting with the pMHC. For none of the pMHC binder designs in the final selection were any Van-der-Waals radii clashes reported. In pre-selection after ipAE and pLDDT binder cut-off and before TopModel filter application, 33 pMHC binder designs had reports of one or multiple clashes. Example 1.2 - Diversification of pMHC binder against RVTDESILSY / HLA-A*01 :01 before in vitro validation
[0373] Preceding the in vitro assessment of the 30 selected pMHC binder designs (Table 6, SEQ ID NO: 296-325), a diversification of the designs was introduced to increase chances of success, as defined by pMHC binder designs that have been tested in vitro and revealed binding affinity towards the pMHC RVTDESILSY / HLA-A*01 :01. An arbitrary number of 96 pMHC binder designs were chosen to be tested in vitro, as described in Example 2, leaving space for 66 further pMHC binder designs to be selected.
[0374] Similar to the diversification of the pMHC binder 2bnr-27_7_3_9, two aspects were executed. In the first aspect, all 30 selected pMHC binder designs against RVTDESILSY / HLA-A*01 :01 were re-interpreted in amino acid sequence level by utilization of ProteinMPNN at default sampling temperature t = 0.1 and N=30 sequences per initial binder design. In a second aspect, all 30 selected pMHC binder designs were re-interpreted in structure by partial diffusion with denoising steps T = 20 and, subsequently in sequence with N = 4 sequences per partially diffused structure binder design. Next, the newly generated variants of the 30 selected pMHC binder designs from both aspects of diversification were filtered based on an ipAE < 7 A and pLDDT > 80 %. 1117 pMHC binder designs passed the threshold of which 413 originated from the latter aspect, while 704 came from the first aspect. A further 127 of the 1117 pMHC binder designs revealed clash reports when analyzed by TopModel in similar manner as descripted previously. In total, 66 pMHC binder design were selected as diversified variants of the previously selected 30 pMHC binder designs against RVTDESILSY / HLA-A*01 :01 and screened in vitro.
[0375] Example 2 - In vitro screening of putative pMHC binders
[0376] Aim:
[0377] The aim of the present examples is to illustrate the setup for an in vitro platform for rapid screening of the pMHC binders using lentiviral membrane-tethering of the binders to mammalian cells (mammalian display) and screening for binding and specificity using FACS- based sorting after staining with single pMHC tetramers or pMHC multimer libraries.
[0378] Methods and Materials:
[0379] Generation of pMHC binder libraries
[0380] Peptide major histocompatibility complex (pMHC) binders may be screened in vitro using a T cell receptor-deficient Jurkat cell line assay. The in s / 7 / co-predicted pMHC binders described in Example 1 may be reverse translated and codon for mammalian expression. Sequences may be ordered as gene fragments (eBlocks, IDT) and cloned into a 2nd generation CAR construct by BsmBI-based Golden gate assembly at 5’ of a CD8 hinge / trans-membrane domain, a CD28 costimulatory domain, a CD3 activation domain, and a GFP marker, with the latter interspersed from the former components by a T2A element. Golden Gate assembly may be performed according to the manufacturer's protocol (NEB). Plasmids containing binders against the same pMHC may be pooled and transformed into NEB® Stable Competent E. coli bacteria (NEB, C3040H) for amplification. Pooled libraries may be purified by maxiprep and analyzed by next-generation sequencing to verify library completeness.
[0381] Cell culture
[0382] CD3 knockout Jurkat cell lines may be cultured at 37°C in a 5% CO2 atmosphere with 80% relative humidity. Culture media (R10) may consist of RPMI 1640 Medium supplemented with 1 % penicillin / streptomycin and 10% Heat Inactivated Fetal Bovine Serum. Cultures are to be maintained by adding or replacing fresh medium after centrifugation (1500 rpm, 5 min, 20°C) at a seeding density of 0.5x106cells / mL.
[0383] Virus production HEK-293T cells may be resuspended at 0.4x106cells / ml and 1 ml is plated on a T175 flask. Cells may then be incubated overnight for them to be 80-90% confluent the next day. Virus can be produced by using a 3rd generation lentiviral system: packaging plasmids [pRSV.REV (Addgene plasmid #12253), pMDLg / p.RRE (Addgene plasmid #12251 ), and pMD2.G (Addgene plasmid #12259)], alongside the pooled plasmid library, are to be transfected into HEK293T cells using Lipofectamine 3000 (Thermofisher Scientific) according to manufacturer’s protocols. 24 and 48 h post-transfection, supernatants are collected and clarified through polyethersulfone (PES) filters with 0.45 pm pore size. After that, supernatants may be mixed with LentiX Concentrator (Clonentech, #631232) at a ratio of 1 volume of reagent per 3 volumes of supernatant and incubated for 30 minutes at 4°C. The mixture may then be centrifuged at 1 .500G for 45 minutes. The supernatant is discarded, and the pellet is resuspended in a suitable volume of PBS and aliquoted. Vials are kept at -80°C until used for cell transduction. Titration of the virus may be performed before transduction to ensure a low multiplicity of infection (MOI) of one integration per cell. pMHC monomers production and staining pMHC monomers may be prepared by incubating peptide (200 pM) and biotinylated MHC (100 pg / ml) at room temperature (RT) for 30 min when using a stabilized Y84C variant of the MHC monomers. Alternatively, MHC monomers with a UV-cleavable peptide may be treated under UV exposure for 1 hour. pMHC are conjugated to labeled streptavidin using 18.04 pL / 100 pL pMHC to make the fluorochrome-labeled tetramers. PE and APC fluorochromes may be used to label tetramers with the intended pMHC target. If applicable, PE fluorochrome may be used to label tetramers with the intended pMHC target, and APC fluorochrome may be used to label a pool of pMHCs of unintended specificity.
[0384] The transduced CD3 KO Jurkat cells may be stained with pooled pMHC multimers for 15 min at RT, followed by 30 min on ice with a live dead marker or alternative surface markers (e.g. anti-CD3). The pMHC-binding Jurkat cells may be sorted on a FACSAria™ Fusion flow cytometer or a BD FACSDiscover™ S8 Cell Sorter (BD, USA). Cells with intended specificity and potential cross-reactive cells may be sorted in different tubes.
[0385] Genomic DNA from sorted cells may be purified using the QIAamp® DNA Micro Kit (Qiagen, #56304) and the integrated binder sequences may be PCR amplified and sequenced by next generation sequencing using Nanopore Minion or 250 bp paired-end Illumina sequencing.
[0386] Accordingly, thousands of binders may be screened against hundreds-thousands of pMHCs in each experiment (Andersen et al., “Parallel detection of antigen-specific T cell responses by combinatorial encoding of MHC multimers’’, Nature Protocols volume 7, pages 891-902 (2012), and Bentzen et al., “T cell receptor fingerprinting enables in-depth characterization of the interactions governing recognition of peptide-MHC complexes’’, Nature Biotechnology volume 36, pages 1191-1196 (2018)). To evaluate specificity, the Jurkat cells are stained with the intended target pMHC multimers in one fluorophore (PE) and a range of undesired pMHC targets using a pool of the undesired pMHC multimers in another fluorophore (APC). This allows for sorting the specific binders and sequencing thereof (Fig. 5). The described platform has been validated with membrane-tethering of in silico predicted binders designed to bind two different a-neurotoxins: the short-chain a-neurotoxin (ScNtx) and the a- Cobratoxin (aCBTX) using the designated plasmid (Baker et al. “De novo designed proteins neutralize lethal snake venom toxins” Res Sq [Preprint]. 2024 May 17:rs.3.rs-4402792 (2024)). Results
[0387] Identification of SLLMWITQC / HLA-A*02:01 binders
[0388] The mammalian display-based platform for pMHC-binder discovery was validated by assessing the surface display of in silico predicted pMHC binders for SLLMWITQC / HLA- A*02:01 as described above. A library of 50 predicted SLLMWITQC / HLA-A*02:01 binders (SEQ ID NO: 326-375) were cloned and transduced in CD3 KO Jurkat cells as a pool. SLLMWITQC / HLA-A*02:01 tetramers were separately labelled with PE and APC fluorochromes and pooled before cell staining. CD3 KO Jurkat cells were stained with a mouse anti-Human CD3 (BD, #562877). Thereafter, the cells were analyzed by flow cytometry, where binder expression was assessed by evaluating the GFP signal, while pMHC binding to the surface expressed pMHC binder was confirmed by the double PE and APC. Cells were therefore sorted for CD3', GFP+and double positive for PE and APC (Figure 7a). After genomic DNA extraction and PCR, integrated binder sequences were sequenced by next generation sequencing using Nanopore Minion. For data analysis, FASTQ read files obtained after sequencing were trimmed using an in-house Python script and matched against the library of 50 minibinders sequences. Read counts for each minibinder were normalized against total read counts in cells prior to sorting. For each minibinder, the Iog2-transformed fold change (log2FC) for enrichment was calculated between the sorted samples and the original library. From the 50 predicted SLLMWITQC / HLA-A*02:01 binders, the minibinder 2bnr-27_7_3_9 (SEQ ID NO: 341 , corresponding to the amino acid SEQ ID NO: 76) was identified as a true binder (Figure 7b).
[0389] Example 2.1 - In vitro screening of putative pMHC binders
[0390] Methods and Materials:
[0391] Generation of pMHC binder libraries
[0392] Peptide major histocompatibility complex (pMHC) binders may be screened in vitro using a T cell receptor-deficient Jurkat cell line assay. The in s / 7 / co-predicted pMHC binders described in Example 1 may be reverse translated and codon for mammalian expression. Sequences may be ordered as gene fragments (eBlocks, IDT) and cloned into a 2nd generation CAR construct by BsmBI-based Golden gate assembly at 5’ of a CD8 hinge / trans-membrane domain, a CD28 costimulatory domain, a CD3 activation domain, and a GFP marker, with the latter interspersed from the former components by a T2A element. Golden Gate assembly may be performed according to the manufacturer's protocol (NEB). Plasmids containing binders against the same pMHC may be pooled and transformed into NEB® Stable Competent E. coli bacteria (NEB, C3040H) for amplification. Pooled libraries may be purified by maxiprep and analyzed by next-generation sequencing to verify library completeness.
[0393] Cell culture
[0394] CD3 knockout Jurkat cell lines and the the A375 melanoma cell line may be cultured at 37°C in a 5% CO2 atmosphere with 80% relative humidity. Culture media (R10) may consist of RPMI 1640 Medium supplemented with 1 % penicillin / streptomycin and 10% Heat Inactivated Fetal Bovine Serum. Cultures are to be maintained by split every 2-4 days before confluence.
[0395] A375 may be modified to express mCherry and Luciferase by lentiviral transduction of mCherry and Luciferase encoding plasmids and sorted based on bright mCherry expression on an Aria Fusion (BD Bioscience) to generate mCherry+ A375 cells.
[0396] Virus production HEK-293T cells may be resuspended at 0.4x106cells / ml and 1 ml is plated on a T175 flask. Cells may then be incubated overnight for them to be 80-90% confluent the next day. Virus can be produced by using a 3rd generation lentiviral system: packaging plasmids [pRSV.REV (Addgene plasmid #12253), pMDLg / p.RRE (Addgene plasmid #12251 ), and pMD2.G (Addgene plasmid #12259)], alongside the pooled plasmid library, are to be transfected into HEK293T cells using Lipofectamine 3000 (Thermofisher Scientific) according to manufacturer’s protocols. 24 and 48 h post-transfection, supernatants are collected and clarified through polyethersulfone (PES) filters with 0.45 pm pore size. After that, supernatants may be mixed with LentiX Concentrator (Clonentech, #631232) at a ratio of 1 volume of reagent per 3 volumes of supernatant and incubated for 30 minutes at 4°C. The mixture may then be centrifuged at 1 .500G for 45 minutes. The supernatant is discarded, and the pellet is resuspended in a suitable volume of PBS and aliquoted. Vials are kept at -80°C until used for cell transduction. Titration of the virus may be performed before transduction to ensure a low multiplicity of infection (MOI) of one integration per cell.
[0397] Validation and cross reactivity of 2bnr-27_7_3_9 - Method
[0398] To confirm binding to SLLMWITQC / HLA-A*02:01 , The identified 2bnr-27_7_3_9 binder (SEQ ID NO: 341 , corresponding to the amino acid SEQ ID NO: 76) was cloned as a single construct, lentivirally transduced into CD3 KO Jurkat cells and stained with SLLMWITQC / HLA-A*02:01- PE tetramers as described above.
[0399] To evaluate specificity, a small cross-reactivity screen was performed by separately staining the transduced CD3 KO Jurkat cell lines with the target SLLMWITQC / HLA-A*02:01 tetramers in one fluorophore APC (WT Tetramer-APC) and a range of variant pMHC target tetramers in PE fluorochrome (Variant Tetramer-PE). In particular, the variant pMHC tetramers encompassed single- or double-point mutants of the WT peptide bound to HLA- A*02:01 , corresponding to SEQ ID NO 531-536 and SEQ ID NO 668-669, target sequences SLLAWITQC / HLA-A*02:01 to SEQ ID NO 16. 2bnr-27_7_3_9+CD3 KO Jurkat cells were single- and double-stained with WT Tetramer-APC and Variant Tetramer-PE as described above. Samples were sequentially stained with first the variant tetramer, then the WT tetramer in double-stains. Cross-reactivity was evaluated by comparing the geometric mean fluorescent intensity (gMFI) of PE signal in Variant Tetramer-PE single stains.
[0400] Incucyte killing assays - Method
[0401] The identified 2bnr-27_7_3_9 binder (SEQ ID NO: 341 , corresponding to the amino acid SEQ ID NO: 76) was cloned into the 2ndgeneration CAR construct used for mammalian display screening (2bnr-27_7_3_9-CD8a-CD28-CD3^ CAR). Primary T cells were isolated from PBMCs from healthy donors and activated overnight with 5 pg / ml anti-CD3 (platebound), 2 pg / ml anti-CD28 and 20 lU / ml of lnterleukin-2 (IL-2). The following day, activated T cells were transduced with the 2bnr-27_7_3_9-CD8a-CD28-CD3^ CAR construct. Primary human T cells were cultured in complete X-vivo-15 (Lonza) with 5% human serum and 1 % PenStrep at 37°C with 5% CO2. At day 10, transduced primary T cells were mixed 1 :1 and 0.5:1 T cells-to-target ratio with mCherry* SLLMWITQC / HLA-A*02:01+ target cells (A375 cancer cell line). The cytotoxicity of the 2bnr-27_7_3_9-CAR T cells was evaluated using an Incucyte Live Cell Analyzer (Sartorius), which took regular images every hour over 72 hours and identified the number of live target cells left in the culture over time. The number of live target cells left in the culture over time was normalised to the initial cell density of the individual cells.
[0402] Results
[0403] Identification of SLLMWITQC / HLA-A*02:01 binders from the diversification of pMHC binder
[0404] 2bnr-27 7 3 9 (SEQ ID NO: 341 ) A library of 92 predicted SLLMWITQC / HLA-A*02:01 binders (SEQ ID NO: 670-761 AA, 762- 853 DNA), selected from the diversification of 2bnr-27 7 3 9 (SEQ ID NO: 341), were cloned and transduced in CD3 KO Jurkat cells as a pool. SLLMWITQC / HLA-A*02:01 tetramers were separately labelled with PE and APC fluorochromes and pooled before cell staining. PE-APC double positive population and negative population were sorted separately (Figure 11 A), the genomic DNA was extracted, prepared for next generation sequencing using Nanopore Minion, sequenced, and analyzed using an in-house python script. For each minibinder, the log2FC for enrichment was calculated between the positive sorted samples and the negative sorted sample. From the 92 predicted SLLMWITQC / HLA-A*02:01 binders, 40 had a log2FC>1 and 32 were identified as true binders (35% success rate) with different binding avidities by SLLMWITQC / HLA-A*02:01 PE-tetramer staining. (Figure 11B, C). This success demonstrated the possibility of increasing the success rate of true binders by applying a second round of diversification to a previously identified true binder.
[0405] Identification of RVTDESILSY / HLA-A*01 :01 binders
[0406] Similar to the identification of pMHC binder against SLLMWITQC / HLA-A*02:01 , a library of 95 predicted RVTDESILSY / HLA-A*01 :01 binders were cloned transduced in CD3 KO Jurkat cells as a pool. RVTDESILSY / HLA-A*01 :01 tetramers were separately labelled with PE and APC fluorochromes and pooled before cell staining. As previously reported, the genomic DNA was extracted, prepared for next generation sequencing using Nanopore Minion, sequenced, and analyzed using an in-house python script. Two out of the 95 pMHC binder designs revealed a significant enrichment and are, therefore, considered as true binders against RVTDESILSY / HLA-A*01 :01 (Figure 8 A, B). One (AKAP9_binder_15, SEQ ID NO: 302) of the two pMHC binders originated from the initial selection of 30 pMHC binder designs, while the other was the partially diffused variant (AKAP9_binder_3_renumb_25_dldesign_0_af2pred, SEQ ID NO: 665) of another initial pMHC binder design.
[0407] This success demonstrated the possibility of applying the binder design platform to neoantigens on other HLAs than HLA-A*02:01 , and targets with no experimental structure.
[0408] Validation and cross reactivity of 2bnr-27_7_3_9 - results 2bnr-27_7_3_9-transduced CD3 KO Jurkat cells binding to SLLMWITQC / HLA-A*02:01 was confirmed by flow cytometry by pMHC tetramer staining (Figure 9A). To investigate the cross-reactivity profile of 2bnr-27_7_3_9, a small-scale cross-reactivity experiment was performed by staining the transduced CD3 KO Jurkat cells with selected pMHC tetramers. The ability of 2bnr-27_7_3_9 to bind 9 off-target peptides (single mutants or double mutants) in complex with HLA-A*02:01 was tested. 2bnr-27_7_3_9+CD3 KO Jurkat cells were cross- reactive with 3 out of 9 off-target peptides (Figure 9B-C), as exemplified by PE gMFI of Variant Tetramer-PE single stains (Figure 9D), which suggests the identified binder to be minimally cross-reactive.
[0409] Incucyte killing assays - results
[0410] The true binder 2bnr-27_7_3_9 (SEQ ID NO: 341 , corresponding to the amino acid SEQ ID NO: 76) was cloned as a single construct and introduced into primary human T cells by lentiviral transduction. T cells expressing the 2bnr-27_7_3_9-miBd CAR were evaluated for their capacity to kill target mCherry* NY-ESO-1+A375 melanoma cells. The miBd-CAR expressing T cells induced rapid cell death of the A375 cell line (Figure 10) compared to the non-transduced control (UTD) at both 1 :1 and 0.5:1 T cells-to-target ratio. This shows that the pMHC-binding miBds can be used in the context of a CAR to kill target cells.
[0411] Possible future results
[0412] The identified binder sequences may be cloned as single constructs and transduced by lentiviral transduction as described above to confirm binding for SLLMWITQC / HLA-A*02:01. To evaluate specificity, cross-reactivity screens may be performed by staining the Jurkat cells with the target SLLMWITQC / HLA-A*02:01 multimers in one fluorophore (PE) and a range of undesired pMHC targets using a pool of the undesired pMHC multimers in another fluorophore (APC). In particular, the undesired pMHC multimers may encompass single- or double-point mutated peptides bound to HLA-A*02:01 , or in combination with alternative HLAs.
[0413] To confirm minibinder functionality, a cell activation assay or a cell-based killing assay may be performed. For the cell activation assay, the identified true binders may be introduced in Jurkat cells with fluorescent reporter elements (NF-KB, NFAT and AP-1) that allow for evaluation of activation of the cells by flow cytometry. The cells will be mixed with target cells expressing SLLMWITQC / HLA-A*02:01 to ensure the presentation of the peptide on MHC. Hereafter, fluorescent signals from the reporters will be analyzed by flow cytometry. For the killing assay, the identified true binders may be introduced as minibinder-CARs in Jurkat cells or in primary CD8+T cells and mixed with fluorescently labelled SLLMWITQC / HLA- A*02:01+target cells. The cytotoxicity of the minibinder-CAR T cells may be evaluated using an Incucyte Live cell analyser set to take regular images every hour over 72h. This will evaluate the cytotoxic potential of the minibinder-CAR T cells by identifying the number of live target cells left in the culture over time.
[0414] Identified true binder sequences may be used to further select suitable pMHC binders. Additionally, the information obtained from the in vitro characterization may be used to run additional in silico campaigns for additional optimization of the pMHC binders (e.g. using Bayesian Optimization), in a recursive pMHC binder optimization method.
[0415] Example 3 - Affinity measurements of pMHC binders
[0416] Aim:
[0417] The aim of the present example is to illustrate how to validate the binding and affinity of the predicted pMHC binders towards the target pMHC. The following methods are exemplary and serves as illustrative methods for how the pMHC binder may be produced and affinity matured. Additional methods are evident to the skilled person, and as such the following examples is not to be considered limiting to the present disclosure.
[0418] Methods:
[0419] Expression and purification of best-binding pMHC binder
[0420] The pMHC binders may be produced and purified as described in the literature known to the skilled person (Ahmadi, S. et al. An in vitro methodology for discovering broadly-neutralizing monoclonal antibodies. Sci. Rep. 10, 10765 (2020)). Briefly, a pre-culture of about 50 ml comprising 2*TY medium, supplemented with 2% (w / v) glucose and 50 pg / mL kanamycin may be inoculated with a strain comprising an expression nucleic acid encoding one or more pMHC binder(s). The cultures may be incubated overnight at 37°C and 250 rpm. The following day, about 500-1000 mL of autoinduction medium may be inoculated with the overnight pre-cultures and further incubated overnight at 20-30°C and 100-200 rpm. After cultivation the cells may be harvested by centrifugation at 4300 * g for 10 min, and the supernatants may be discarded. The cell pellet is resuspended in e.g., 50 mL of TES-buffer (30 mM Tris-HCI pH 8.0, 1 mM EDTA, 20% sucrose (w / v)) containing 1.5 kU / mL of r- lysozyme (~ 70,000 U / mg). After about 20 min of incubation on ice, the cells may be centrifuged at 4,300 x g for 10 min, and the supernatant may be discarded. The cell pellet may be resuspended in 50 mL of 5 mM MgSO4 supplemented with the same amount of r- lysozyme as above and incubated on ice for 20 min. After centrifugation at 4300 x g for 10 min, the supernatant may be pooled with the supernatant from the previous step and kept on ice. The pooled supernatants are then centrifuged at 30,000 x g for 30 min. The produced pMHC binders may be purified using affinity chromatography or similar methods known to the skilled person.
[0421] Affinity measurement using biolayer interferometry / surface plasmon resonance
[0422] Precise binding affinities may be determined via biolayer interferometry (BLI) or surface plasmon resonance (SPR). BLI experiments may be performed on an Octet Red96 (ForteBio) instrument, with streptavidin coated tips (Sartorius Item no. 18-5019). Buffer may comprise 1X HBS-EP+ buffer (Cytiva BR100669) supplemented with 0.1 % w / v bovine serum albumin. Tips may be pre-incubated in the buffer for at least 10 minutes before use. Tips may be sequentially incubated in biotinylated pMHC target, buffer, designed pMHC binder, and buffer. SPR experiments may be conducted using a Biacore™ 8K instrument (Cytiva) and analyzed with the accompanying evaluation software. Biotinylated target pMHC may be immobilized on a streptavidin sensor chip (Cytiva). Alternatively, immobilization may involve the activation of carboxymethyl groups on a dextran-coated chip through reaction with N- hydroxysuccinimide. The ligands may be then covalently bonded to the chip surface via amide linkages, and excess activated carboxyls may be blocked with e.g., ethanolamine. For affinity measurements increasing concentrations of protein pMHC binders may be flown over the chip in 1X HBS-EP+ buffer (Cytiva BR100669).
[0423] Thermal stability of pMHC binders
[0424] The secondary structure content of the pMHC binders may be evaluated by circular dichroism (CD) in a Jasco J-1500 CD spectrometer coupled to a Peltier system (EXOS) for temperature control. The experiments may e.g., be performed on quartz cells with an optical path of 0.1 cm, covering a wavelength range from 200-260 nm. CD signal is generally reported as molar ellipticity [0]. The thermal unfolding experiments may be followed by a change in the ellipticity signal at 222 nm as a function of temperature. Proteins may be denatured by heating the proteins at 1°C / min from 20 to 95°C.
[0425] List of tables
[0426] Table 1 - Primary human and mouse genes encoding for MHC class I and II proteins.
[0427] Table 2 - Top 50 in vitro tested pMHC binders designed against SLLMWITQC / HLA-A*02:01 Table 3 - Overview of the sequences of the present disclosure
[0428] Table 4 - Example of parameters used for sequence design with ProteinMPNN. Table 5 - Selected pMHC targets
[0429] Table 6 - Overview of binder designs against four different pMHC targets.
[0430] Table 7 - Common properties of pMHC binders surpassing the final filtering step in the design campaigns against four pMHC complexes, i.e. SIINFEKL / H2-Kb (mouse model antigen from Hen Egg), SLLMWITQC / HLA-A*02:01 (Melanoma shared cancer antigen), RVTDESILSY / HLA-A*01:01 (Melanoma neoantigen), and NLFRRVWEL / HLA-B*08:01 (Melanoma neoantigen).
[0431] ITEMS
[0432] 1 . A method for generating an amino acid sequence of a high affinity and high specificity pMHC binder candidate, wherein said method comprises a) providing a protein structure or structural model of a target pMHC and inputting the protein structure or structural model into a generative protein design system, i) generating a first set of putative pMHC binders targeting said target pMHC using generative protein design, ii) selecting from the first set of putative pMHC binders, all putative pMHC binders having both a pLDDT score of >85% and an ipAE if >12A, thereby generating a second set of putative pMHC binders, ill) optimizing the second putative set of pMHC binders in silico b) Selecting / Recovering only the optimised putative pMHC binders having a pLDDT of >90% and an ipAE of <7A, and c) outputting the amino acid sequences of the putative pMHC binders selected in step b), thereby providing amino acid sequences encoding pMHC binder candidates having high affinity and high specificity toward a target pMHC.
[0433] 2. The method according to item 1 , wherein said optimization in ill) further comprises, a) Identifying one or more potential variant peptides presented by said target pMHC (variant pMHC), b) Determining the ipAE of said variant pMHC: putative pMHC binder complexes, c) Selecting putative pMHC binders with ipAE > 10A for the variant pMHC and an ipAE < 7 A for the target pMHC, d) Obtaining from said selection one or more further optimized putative pMHC binders.
[0434] 3. The method according to item 1 or 2, wherein said optimization in ill) further comprises, a) Providing a peptide unbound MHC structure of the target pMHC, b) Determining the ipAE of said peptide unbound MHC: putative pMHC binder complex, c) Selecting putative pMHC binders with ipAE > 10 A for the unbound pMHC and an ipAE < 10 A for the target pMHC, and d) Obtaining from said selection one or more further optimized putative pMHC binders. The method to any of the preceding items, wherein the putative pMHC binder generation in i) comprise the use of diffusion modelling, RFdiffusion and / or a message passing neural network (MPNN). The method according to any of the preceding items, wherein the putative pMHC binder generation in i) comprise the sequential use of diffusion modelling, RFdiffusion and a message passing neural network (MPNN). The method according to any one of the preceding items, wherein the putative pMHC binder generation in i) comprise the use of a novo shape generating model selected from BindCraft, Protpardelle and BoltzDesignl . The method according to any of the items 4-5, wherein the use of diffusion modelling, RFdiffusion and / or a message passing network comprises the use of ProteinMPNN, Carbondesign, AntiFold, SPDesign, Evo, PocketGen, and / or CoVES. The method according to any of the preceding items, wherein pLDDT and ipAE scores are calculated using a program selected from Alphafold, Alphafold 2 (AF2), Alphafold 3 (AF3) ColabFold, Boltz 1 , Boltz 2 and ESM-fold. The method according to any of the preceding items, wherein pLDDT and ipAE scores are calculated using Alphafold, ColabFold, and ESM-fold. The method according to any one of the preceding claims, wherein the ipAE score is an interaction prediction Score from Aligned Errors (ipSAE) ipSAE score. The method according to any of the preceding items, wherein said optimization in ill) comprises the use of Multi-Objective Bayesian Optimization (MOBO) and / or MultiObjective Evolutionary Algorithms (MOEAs), Multi-Objective Particle Swarm Optimization (MOPSO), Multi-Objective Simulated Annealing (MOSA), Scalarization Techniques, Interactive Methods, Surrogate-assisted Optimization, and Deterministic Methods. 12. The method according to any of the preceding items, wherein said optimization in iii) comprises the use of Bayesian optimization.
[0435] 13. The method according to any of the preceding items, wherein the amino acid sequence of said high affinity and high specificity protein pMHC binder candidate is <200 amino acids, such as less than 175, 150, 125, or such as less than 100 amino acids.
[0436] 14. The method according to any one of the preceding items, wherein the method is a computer implemented method.
[0437] 15. A pMHC binder candidate obtainable from the methods of any of the preceding items.
[0438] 16. The pMHC binder candidate according to item 15, wherein the in vitro affinity of the pMHC binder candidate towards the target pMHC when measured using SPR is <1uM.
[0439] 17. The pMHC binder candidate according to item 15 or 16, wherein the pMHC binder comprises at least 80% helical and / or beta-sheet structure and less than 20% random coil structure.
[0440] 18. The pMHC binder candidate according to any of items 15 to 17, wherein said pMHC binder candidate comprises a helix bundle comprising at least three helices.
[0441] 19. The pMHC binder candidate according to any of items 15 to 18, wherein the alpha carbon of the N- and C-terminal of said pMHC binder candidates are positioned within XA+Z- from each other.
[0442] 20. The pMHC binder candidate according to any of items 15 to 19, wherein the pMHC binder candidate has at least 2-fold target pMHC specificity when comparing the in vitro affinity towards the target pMHC compared to a variant pMHC or unbound MHC.
[0443] 21. The pMHC binder candidate according to any of items 15 to 20, wherein the pMHC binder candidate comprise or consists of less than 200 amino acids, such as less than 175, 150, 125, or such as less than 100 amino acids, or such as between 50 and 200 amino acids.
[0444] 22. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 19-325, 435-503, 556-667 and 670-761 .
[0445] 23. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 19-325, 435-503, and 556-667.
[0446] 24. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 19-325, 435-503, and 556-664.
[0447] 25. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 19-325, and 435-503.
[0448] 26. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 13-18, 22, 24-31 , 33, 35-38, 41-42, 45-46, 48, 50, 54-55, 58-60, 62-63, 67-68, 71-72, 74-82, 84-87, and 95.
[0449] 27. The pMCH binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 76, 302, 665, 341 , 666 and 667.
[0450] 28. pMCH binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 76, 302 and 665.
[0451] 29. pMCH binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 341 , 666 and 667.
[0452] 30. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 76 and 341.
[0453] 31. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to any one of SEQ ID NOs: 302 and 667.
[0454] 32. The pMHC binder candidate according to any one of items 15-21 that has at least
[0455] 90% identity to any one of SEQ ID NOs: 665 and 666. 33. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to SEQ ID NO: 76.
[0456] 34. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to SEQ ID NO: 302.
[0457] 35. The pMHC binder candidate according to any one of items 15-21 that has at least 90% identity to SEQ ID NO: 665.
[0458] 36. A nucleic acid expression vector comprising a nucleic acid sequence encoding a polypeptide construct comprising nucleic acids encoding a) a pMHC binder candidate according to any of items 15 to 35, b) a polypeptide membrane anchor, c) a linker connecting said pMHC binder candidate and said membrane anchor, and d) a surface expression tag, and wherein said expression construct in under control of an expression element such as a promoter.
[0459] 37. A method for biosynthetic and / or synthetic production of a pMHC binder candidate comprising a) Obtaining a pMHC binder candidate amino acid sequence using the method according to any of items 1-14, b) Producing said pMHC binder candidate of said amino acid sequence using a recombinant expression system, protein synthesis, or a combination thereof, and optionally c) Purifying said pMHC binder candidate.
[0460] 38. The method according to item 37, wherein the pMHC binder candidate is produced in a host cell.
[0461] 39. The method according to item 37, wherein the pMHC binder candidate is produced using semi-synthesis. 40. The method according to item 37, wherein the pMHC binder candidate is synthetically produced.
[0462] 41. A pMHC binder candidate produced according to any of items 37 to 40.
[0463] 42. A pMHC binder candidate according to any of items 15-35, a pMHC binder candidate according to item 41 , or a nucleic acid expression vector according to item 36, for use as a medicament.
[0464] 43. A pMHC binder candidate according to any of items 15-35, a pMHC binder candidate according to item 41 , or a nucleic acid expression vector according to item 36, for use in the treatment of cancer, viral infections, intracellular parasite infections, intracellular bacterial infections, and / or autoimmune diseases.
[0465] 44. A pMHC binder candidate according to any of items 15-35, a pMHC binder candidate according to item 41 , or a nucleic acid expression vector according to item 36, for use as a CAR, TCR-like molecule, BiTe, TCR-like-fusion construct, soluble blocker, and / or drug conjugate.
[0466] 45. A pMHC binder candidate according to any of items 15-35, a pMHC binder candidate according to item 41 , or a nucleic acid expression vector according to item 36, for use in cross reactivity testing.
[0467] 46. A pMHC binder candidate according to any of items 15-35, a pMHC binder candidate according to item 41 , or a nucleic acid expression vector according to item 36, for use in the treatment of cancer.
[0468] 47. Use of a pMHC binder candidate according to any of items 15-35, a pMHC binder candidate according to item 41 , or a nucleic acid expression vector according to item 36, in cancer cell targeting.
[0469] 48. Diagnostic use of a pMHC binder candidate according to any of items 15-35, a pMHC binder candidate according to item 41 , or a nucleic acid expression vector according to item 36. 49. Use of a pMHC binder candidate according to any of items 15-35, a pMHC binder candidate according to item 41 , or a nucleic acid expression vector according to item 36 as an affinity reagent.
[0470] 50. A method for generating an amino acid sequence of a high affinity and high specificity pMHC binder candidate, wherein said method comprises a) providing a protein structure or structural model of a target pMHC and inputting the protein structure or structural model into a generative protein design system, i) generating a first set of putative pMHC binders targeting said target pMHC using generative protein design, ii) selecting from the first set of putative pMHC binders, all putative pMHC binders having both a pLDDT score of >85% and an ipAE if >12A, thereby generating a second set of putative pMHC binders, ill) optimizing the second putative set of pMHC binders in silico b) Selecting / Recovering only the optimised putative pMHC binders having a pLDDT of >90% and an ipAE of <7A, and c) outputting the amino acid sequences of the putative pMHC binders selected in step b), thereby providing amino acid sequences "encoding" pMHC binder candidates having high affinity and high specificity toward a target pMHC.
[0471] 51. The method according to item 50, wherein said optimization in ill) further comprises, e) Identifying one or more potential variant peptides presented by said target pMHC (variant pMHC), f) Determining the ipAE of said variant pMHC:putative pMHC binder complexes, g) Selecting putative pMHC binders with ipAE > A for the variant pMHC and an ipAE < 7 A for the target pMHC, h) Obtaining from said selection one or more further optimized putative pMHC binders.
[0472] 52. The method according to item 50 or 51 , wherein said optimization in ill) further comprises, e) Providing a peptide unbound MHC structure of the target pMHC, f) Determining the ipAE of said peptide unbound MHC:putative pMHC binder complex, g) Selecting putative pMHC binders with ipAE > 10 A for the unbound pMHC and an ipAE < 10 A for the target pMHC, and h) Obtaining from said selection one or more further optimized putative pMHC binders.
[0473] 53. The method to any of items 50-52, wherein the putative pMHC binder generation in i) comprise the use of diffusion modelling, RFdiffusion and / or a message passing neural network (MPNN).
[0474] 54. The method according to any of items 50-53, wherein the putative pMHC binder generation in i) comprise the sequential use of diffusion modelling, RFdiffusion and a message passing neural network (MPNN).
[0475] 55. The method according to any of items 50-54, wherein pLDDT and ipAE scores are calculated using Alphafold or ColabFold.
[0476] 56. The method according to any of items 50-55, wherein said optimization in ill) comprises the use of Multi-Objective Bayesian Optimization (MOBO) and / or MultiObjective Evolutionary Algorithms (MOEAs), Multi-Objective Particle Swarm Optimization (MOPSO), Multi-Objective Simulated Annealing (MOSA), Scalarization Techniques, Interactive Methods, Surrogate-assisted Optimization, and Deterministic Methods.
[0477] 57. The method according to any of items 50-56, wherein said optimization in ill) comprises the use of Bayesian optimization.
[0478] 58. A method for biosynthetic and / or synthetic production of a pMHC binder candidate comprising d) Obtaining a pMHC binder candidate amino acid sequence using the method according to any of items 50-57, e) Producing said pMHC binder candidate of said amino acid sequence using a recombinant expression system, protein synthesis, or a combination thereof, and optionally f) Purifying said pMHC binder candidate.
[0479] 59. A pMHC binder candidate obtainable from the methods any of items 50-58. 60. The pMHC binder candidate according to item 59, wherein the in vitro affinity of the pMHC binder candidate towards the target pMHC when measured using SPR is <1uM.
[0480] 61. The pMHC binder candidate according to item 59 or 60, wherein the pMHC binder candidate comprises at least 80% helical and / or beta-sheet structure and less than 20% random coil structure and / or a helix bundle comprising at least three helices.
[0481] 62. The pMHC binder candidate obtained according to the method as defined in any one of items 50-58 or a pMHC binder candidate according to any one of items 59-61 , wherein the pMHC binder candidate has at least 90% identity to any one of SEQ ID NOs: 19-325, and 435-503.
[0482] 63. A nucleic acid expression vector comprising a nucleic acid sequence encoding a polypeptide construct comprising nucleic acids encoding a) a pMHC binder candidate according to any of items 60 to 62, b) a polypeptide membrane anchor, c) a linker connecting said pMHC binder candidate and said membrane anchor, and d) a surface expression tag, and wherein said expression construct in under control of an expression element such as a promoter.
[0483] 64. A pMHC binder candidate according to any of items 60-62, or a nucleic acid expression vector according to item 63, for use in the treatment of cancer, viral infections and / or autoimmune diseases.
Claims
CLAIMS1 . A method for generating an amino acid sequence of a high affinity and high specificity pMHC binder candidate, wherein said method comprises a) providing a protein structure or structural model of a target pMHC and inputting the protein structure or structural model into a generative protein design system, i) generating a first set of putative pMHC binders targeting said target pMHC using generative protein design, ii) selecting from the first set of putative pMHC binders, all putative pMHC binders having both a pLDDT score of >85% and an ipAE if >12A, thereby generating a second set of putative pMHC binders, ill) optimizing the second putative set of pMHC binders in silico b) Selecting / Recovering only the optimised putative pMHC binders having a pLDDT of >90% and an ipAE of <7A, and c) outputting the amino acid sequences of the putative pMHC binders selected in step b), thereby providing amino acid sequences encoding pMHC binder candidates having high affinity and high specificity toward a target pMHC.
2. The method according to claim 1 , wherein said optimization in ill) further comprises, i) Identifying one or more potential variant peptides presented by said target pMHC (variant pMHC), j) Determining the ipAE of said variant pMHC: putative pMHC binder complexes, k) Selecting putative pMHC binders with ipAE > A for the variant pMHC and an ipAE < 7 A for the target pMHC, l) Obtaining from said selection one or more further optimized putative pMHC binders.
3. The method according to claim 1 or 2, wherein said optimization in ill) further comprises, i) Providing a peptide unbound MHC structure of the target pMHC, j) Determining the ipAE of said peptide unbound MHC: putative pMHC binder complex, k) Selecting putative pMHC binders with ipAE > 10 A for the unbound pMHC and an ipAE < 10 A for the target pMHC, andI) Obtaining from said selection one or more further optimized putative pMHC binders.
4. The method to any of the preceding claims, wherein the putative pMHC binder generation in i) comprise the use of diffusion modelling, RFdiffusion and / or a message passing neural network (MPNN).
5. The method according to any of the preceding claims, wherein the putative pMHC binder generation in i) comprise the sequential use of diffusion modelling, RFdiffusion and a message passing neural network (MPNN).
6. The method according to any one of claims 1-3, wherein the putative pMHC binder generation in i) comprise the use of a novo shape generating model selected from BindCraft, Protpardelle and BoltzDesignl .
7. The method according to any of the preceding claims, wherein pLDDT and ipAE scores are calculated using a program selected from Alphafold, Alphafold 2 (AF2), Alphafold 3 (AF3) ColabFold, Boltz 1 , Boltz 2 and ESM-fold.
8. The method according to any of the preceding claims, wherein pLDDT and ipAE scores are calculated using Alphafold, ColabFold, and ESM-fold.
9. The method according to any one of the preceding claims, wherein the ipAE score is an interaction prediction Score from Aligned Errors (ipSAE) ipSAE score.
10. The method according to any of the preceding claims, wherein said optimization in ill) comprises the use of Multi-Objective Bayesian Optimization (MOBO) and / or MultiObjective Evolutionary Algorithms (MOEAs), Multi-Objective Particle Swarm Optimization (MOPSO), Multi-Objective Simulated Annealing (MOSA), Scalarization Techniques, Interactive Methods, Surrogate-assisted Optimization, and Deterministic Methods.
11. The method according to any of the preceding claims, wherein said optimization in ill) comprises the use of Bayesian optimization.
12. The method according to any of the preceding claims, wherein the amino acid sequence of said high affinity and high specificity protein pMHC binder candidate is <200 amino acids, such as less than 175, 150, 125, or such as less than 100 amino acids.
13. The method according to any one of the preceding claims, wherein the method is a computer implemented method.
14. The method according to any one of the preceding claims, wherein the method further comprises a) Obtaining pMHC binder candidate amino acid sequence(s) outputted in step c) of claim 1 b) Producing said pMHC binder candidate(s) of said amino acid sequence using a recombinant expression system, protein synthesis, or a combination thereof, and optionally c) Purifying said pMHC binder candidate.
15. The method according to claim 14, wherein the pMHC binder candidate is produced in a host cell.
16. The method according to claim 14, wherein the pMHC binder candidate is produced using semi-synthesis.
17. The method according to claim 14, wherein the pMHC binder candidate is synthetically produced.
18. A pMHC binder candidate obtained by the method of any one of the claims 1 -17.
19. The pMHC binder candidate according to claim 18, wherein the in vitro affinity of the pMHC binder candidate towards the target pMHC when measured using SPR is <1uM.
20. The pMHC binder candidate according to claim 18 or 19, wherein the pMHC binder comprises at least 80% helical and / or beta-sheet structure and less than 20% random coil structure.
21. The pMHC binder candidate according to any of claims 18 to 20, wherein said pMHC binder candidate comprises a helix bundle comprising at least three helices.
22. The pMHC binder candidate according to any of claims 18 to 21 , wherein the pMHC binder candidate has at least 2-fold target pMHC specificity when comparing the in vitro affinity towards the target pMHC compared to a variant pMHC or unbound MHC.
23. The pMHC binder candidate according to any of claims 18 to 22, wherein the pMHC binder candidate comprise or consists of less than 200 amino acids, such as less than 175, 150, 125, or such as less than 100 amino acids, or such as between 50 and 200 amino acids.
24. The pMHC binder candidate according to any one of claims 18-23 that has at least 90% identity to any one of SEQ ID NOs: 19-325, 435-503, 556-667 and 670-761 .
25. The pMCH binder candidate according to any one of claims 18-24 that has at least 90% identity to any one of SEQ ID NOs: 76, 302 and 665.
26. A pMHC binder candidate having at least 90% identity to any one of SEQ ID NOs: 19-325, 435-503, 556-667 and 670-761.
27. The pMCH binder candidate according to claim 26 that has at least 90% identity to any one of SEQ ID NOs: 76, 302 and 665.
28. A nucleic acid expression vector comprising a nucleic acid sequence encoding a polypeptide construct comprising nucleic acids encoding a) a pMHC binder candidate according to any of claims 18 to 27, b) a polypeptide membrane anchor, c) a linker connecting said pMHC binder candidate and said membrane anchor, and d) a surface expression tag, and wherein said expression construct in under control of an expression element such as a promoter.
29. The pMHC binder candidate according to any of claims 18-27, or the nucleic acid expression vector according to claim 28, for use as a medicament, for use in the treatment of cancer, viral infections, intracellular parasite infections, intracellular bacterial infections, and / or autoimmune diseases, for use as a CAR, TCR-like molecule, BiTe, TCR-I ike-fusion construct, soluble blocker, and / or drug conjugate, for use in cross reactivity testing, for use in cancer cell targeting, and / or for use as an affinity reagent.