MHC ligand identification and related systems and methods
A transformer encoder-decoder model and GAN-based approach improves MHC ligand prediction precision, addressing the limitations of existing tools and facilitating vaccine development by accurately identifying peptides that bind to MHC molecules.
Patent Information
- Application Number
- PCT/EP2025/057030
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-15
- Filing Date
- 2025-03-14
- Publication Date
- 2025-09-18
AI Technical Summary
Current prediction tools for MHC binding, such as NetMHC, MixMHCpred, and BERTMHC, do not provide optimum predictive values when evaluated against verified test data, necessitating improved methods for determining peptide presentation on MHC class I and class II molecules to facilitate vaccine development.
Utilization of a transformer encoder-decoder model and a conditional generative adversarial network (GAN) to identify peptides that bind and/or are presented by MHC molecules, involving an artificial neural network (ANN) architecture and positional embedding to enhance prediction precision.
The proposed method achieves hitherto unseen precision in predicting MHC ligand properties, enabling accurate identification of peptides that will be bound and/or presented by MHC molecules, thereby supporting effective vaccine development.
Smart Images

Figure EP2025057030_18092025_PF_FP_ABST
Abstract
Description
[0001] MHC LIGAND IDENTIFICATION AND RELATED SYSTEMS AND METHODS
[0002] FIELD OF THE INVENTION
[0003] The present invention relates to the field of immune therapy and immune modulation. In particular, the present invention relates to a novel method and system for MHC ligand identification as well as for a method of establishing the method and the system.
[0004] BACKGROUND OF THE INVENTION
[0005] T cells are some of the key players in the immune system. Cytotoxic (CD8+) T cells interact with the MHC class I molecule complexed with a peptide derived from a potential antigen. The role of the cytotoxic T cell is to recognize cells expressing potentially harmful antigens, e.g. virus-infected or cancerous cells. All nucleated cells express the MHC class I peptide complex as a snapshot of the proteome inside the cell. Helper (CD4+) T cells have other stimulatory purposes within the immune system and recognize peptides bound to the MHC class II molecule. A set of "professional" antigen presenting cells (APCs) express MHC class II, which functions as a snapshot of the various surveillance functions of the APCs. Examples of APCs include dendritic cells, macrophages, and B cells.
[0006] Predicting which peptides are subject to T-cell recognition - both CD4+and CD8+- is vital to any artificial intelligence (Al) based immunology pipeline. The first step in this interaction is the presentation of specific peptides on the cell surface. Both MHC class I and class II are encoded by extremely polymorphic genes, with 1,000s of known alleles. Specific alleles restrict the pool of peptides presented on the surface of the cell, and the specific combination of alleles (haplotype) in an individual determines the combined pool of peptides that can be presented complexed to MHC on cell surfaces and thereby enable recognition by the immune system.
[0007] Examples of currently available prediction tools for MHC binding are NetMHC (see reference [7]), MixMHCpred (see reference
[0011] ), MHCflurry (see reference [5]), and BERTMHC (see reference [9]). However, none of these tools provide for optimum predictive values when evaluated against verified test data.
[0008] Thus, to facilitate vaccine development, a state-of-the-art predictor of peptide presentation on MHC class I and class II is needed. OBJECT OF THE INVENTION
[0009] It is an object of embodiments of the invention to provide improved means and methods for determining the ability of peptides to act as ligands for MHC molecules. Further objects of embodiments of the invention are to provide improved methods for training ANN models that identify MHC ligands and to provide methods for identification of MHC ligands embedded in larger proteins.
[0010] SUMMARY OF THE INVENTION
[0011] It has been found by the present inventors that utilisation of a transformer encoder-decoder model, which has been properly trained provides for hitherto unseen precision in prediction of MHC ligand properties of peptides.
[0012] So, in a first aspect the present invention relates to a method for identifying peptides, which are bound and / or presented by MHC class I and / or MHC class II molecules in an individual, the method comprising a) providing a transformer encoder-decoder model, which evaluates whether amino acid sequences of peptides are ligands for selected MHC molecules, and which comprises an artificial neural network (ANN) architecture, which comprises an input layer, an output layer and one or more hidden layers, b) evaluating, in the transformer encoder-decoder model, data pairs, which each comprises 1) an MHC molecule amino acid sequence as well as 2) a peptide amino acid sequence, c) outputting from the transformer encoder-decoder model, for each inputted data pair, at least one quantitative assessment of the ability of each MHC molecule having the MHC molecule amino acid sequence of a data pair to bind and / or present the peptide having the peptide amino acid sequence of the same data pair, wherein the input layer receives and accepts data pairs each composed of an MHC molecule amino acid sequence and an amino acid sequence of a peptide, wherein the transformer encoder receives the MHC molecule amino acid sequence of a data pair with residue- and positional embedding as input and outputs a high-dimensional representation of the MHC amino acid sequence, wherein the transformer decoder receives the high-dimensional representation of the MHC molecule amino acid sequence and the peptide amino acid sequence of a data pair with residue embedding and positional encoding as input and outputs a high-dimensional representation of the relationship between the MHC molecule amino acid sequence and the peptide sequence, wherein the output from the transformer decoder is used to calculate the quantitative assessment of the ability of the MHC molecule having the MHC molecule amino acid sequence of a data pair to bind and / or present the peptide having the peptide amino acid sequence of the same data pair.
[0013] In a second aspect, the present invention relates to a method for training a predictor model capable of identifying peptides, which will be bound and / or presented by selected MHC class I and / or MHC class II molecules in an individual, the method comprising utilizing a conditional generative adversarial network (GAN) in the form of a generator model and a predictor model, wherein the predictor model evaluates whether inputted amino acid sequences of peptides are bound and / or presented by inputted MHC molecules, a) evaluating, in the predictor model, data pairs, which each comprises 1) an MHC molecule amino acid sequence as well as 2) a peptide amino acid sequence, wherein, in a fraction of said data pairs, the amino acid sequences in the data pairs are those of verified peptide-MHC complexes, and wherein, another fraction of said data pairs are random pairs of MHC molecule and peptide sequences, and wherein in the remaining data pairs, the amino acid sequences in the data pairs are from molecules that are generated by the generator model, b) outputting from the predictor model, for each inputted data pair, at least one quantitative and / or qualitative assessment of the ability of each MHC molecule having the MHC molecule amino acid sequence of a data pair to bind and / or present the peptide having the peptide amino acid sequence of the same data pair, c) outputting from the predictor model a quantitative and / or qualitative assessment of whether the peptide sequence of the inputted data pair is a generated sequence or a real, naturally occurring sequence, d) evaluating the output in steps b and c and adjusting weights of the neurons in the one or more hidden layers of the predictor model to minimize the error of the output or to optimize the performance of the model, e) evaluating the output in step c and adjusting weights of the neurons in the one or more hidden layers of the generator model to optimize the model's ability to generate a peptide sequence that the predictor predicts to be a real, naturally occurring peptide sequence, f) repeating steps a-e to train the predictor model to correctly identify peptides, which will be presented by selected MHC class I and / or MHC class II molecules in an individual, wherein the input layer in the predictor model receives and accepts data pairs composed of an amino acid sequence derived from an MHC molecule and an amino acid sequence of a peptide.
[0014] In a 3rdaspect, the present invention relates to a computer or computer system comprising a transformer encoder-decoder model as defined in the first aspect of the invention and any embodiments thereof disclosed herein.
[0015] Finally, in a 4thaspect, the present invention relates to a method of identifying protein- derived MHC interacting peptides, which will be bound and / or presented by at least one MHC molecule expressed by an individual, the method comprising a) determining the amino acid sequences of one or more MHC molecules expressed by said individual, b) deriving amino acid sequences of potential MHC interacting peptides from one or more proteins; c) evaluating pairwise the amino acid sequences from steps a and b by using said amino acid sequences as input in step b in the method of the first aspect of the invention and any embodiments thereof disclosed herein, and / or by querying a database comprising records consisting of interacting pairs of MHC molecules and peptides identified by the method of the first aspect of the invention and any embodiments thereof disclosed herein for the existence in the database of interacting pairs consisting of the MHC molecules and potential MHC interacting peptides from steps a and b respectively, and d) identifying those peptides that are ranked with the highest scores or which provide a positive query result as those that will be bound and / or presented by at least one MHC molecule in the individual LEGENDS TO THE FIGURE
[0016] Fig. 1 : Schematic depiction of an embodiment of the present invention.
[0017] GAN generator and predictor (critic) are set up as encoder-decoder transformer pairs. For the MHC sequence input, the sequence is encoded using positional embedding, and the peptide is encoded using positional encoding as the length will vary between peptides. The predictor splits into 2 terminal neural networks, one with 2 outputs for immunopeptidomics data and binding affinity data and the other output for the GAN output.
[0018] Fig. 2: The average precision scores "EvaxMHC4" (an embodiment of the invention), EvaxMHC3 (see below for details) and BERTMHC models per MHC allele dataset. Each point on the plots represents an MHC allele dataset. The y-axis shows the average precision of EvaxMHC4 on the given dataset, while the x-axis shows the average precision of EvaxMHC3 / BERTMHC. For all points above the dashed lines, EvaxMHC4 had superior performance and vice versa for points below the dashed lines.
[0019] Fig. 3: Summary of an MHC Class I dataset. See examples for details.
[0020] Fig. 4: Summary of an MHC Class II dataset. See examples for details.
[0021] Fig. 5: Graph showing rank calibration for 2 alleles each for MHC class I and II.
[0022] Fig. 6: Calibration results for 82 individual MHC class I allele clusters and 23 class II clusters.
[0023] Fig. 7: Graphs showing recall / precision, receiver operating characteristic, and detection error tradeoff curves for MHC class I prediction.
[0024] Fig. 8: Graphs showing recall / precision, receiver operating characteristic, and detection error tradeoff curves for MHC class II prediction.
[0025] Fig. 9 : Comparison of EvaxMHC4 to other MHC class I predictors. Values are calculated as the log ratio of EvaxMHC4 to other predictors.
[0026] Fig. 10 : Scatter plot of average precision, ROC AUC10 and recall for individual alleles predicted using the MHC class I predictors. Fig. 11 : Scatter plot of average precision, ROC AUC10 and recall for individual alleles predicted using the MHC class II predictors.
[0027] Fig. 12: Line graph showing development of implanted CT26 tumour volume upon immunization of BALB / c mice with pDNA constructs encoding predicted ERV sequences. Groups of mice were prophylactically immunized intramuscularly (i.m.) with 100 ug of ERV-encoding pDNA constructs delivered by 2 x EP + 3 x Kolliphor®. Each mouse was challenged with 2xl05in vitro grown CT26 cells on day 0. The lines represent group mean tumor growth curves (in mm3) ± standard error of the mean (SEM).
[0028] Fig. 13: Bar graph showing ex vivo immune reactivity in an IFNy ELISPOT assay upon restimulation with ERV peptide pools corresponding to the vaccine of each group.
[0029] Single-cell spleen suspensions from the two groups were pooled group-wise and subjected to re-stimulation with pools of ERV peptides. The cells were then subjected to IFNy ELISPOT for analysis of IFNy-secreting immune cells. Bars represent mean ± SD.
[0030] DETAILED DISCLOSURE OF THE INVENTION
[0031] Definitions
[0032] A "high dimensional representation" is in the present context a matrix or tensor containing a numerical encoding of one or more amino acid sequences. Each amino acid is represented by a vector, matrix, or tensor of floating-point numbers. The representation can either be learned by a machine learning model during the model training process or it can be calculated based on e.g. the physiochemical properties of the amino acids or based on evolutionary substitution e.g. BLOSUM, as described in reference
[0066] .
[0033] In the present context, an "MHC pseudosequence" is a subset of amino acid residues within the MHC molecule that have potential to interact with the amino acid residues of a peptide ligand being presented by the MHC molecule. The MHC pseudosequence for a given MHC allele can be determined by performing pairwise alignment of the MHC allele sequence to a reference allele and thereafter extracting a set of predefined positions, defined as the set of residues in contact with the peptide in a 3D crystal structure. Examples of how MHC pseudosequences can be defined are described in greater detail in references
[0067] and
[0068] .
[0034] "Positional embedding" denotes the mathematical process of preserving dependency information of each element (in the present case, an amino acid residue) in an input vector (in the present case an amino acid sequence) to an ANN. The preferred current technique for this purpose is frequency-based positional embedding, which integrates the position of the element, the maximum length of the vector representing the element in the input vector, and the indices of each embedding dimension of the element. An example of the use of positional embedding is described in reference
[0069] .
[0035] "Positional encoding" describes the mathematical process of assigning location or position of an entity in a sequence (traditionally a sequence of words in a sentence, but in the present context the amino acid sequence) so that each position is assigned a unique representation. Transformer models utilise a positional encoding scheme, where each position / index is mapped to a vector. Hence the output of the positional encoding layer is a matrix, where each row of the matrix represents an encoded object of the sequence summed with its positional information. An example of the use of positional encoding is described in reference [1].
[0036] An "MHC molecule" (major histocompatibility molecule) is a tissue antigen expressed by nucleated cells in vertebrates, which binds to peptide antigens and displays ("presents") the antigens to T-cells carrying T-cell receptors. MHC class I is expressed by all nucleated cells and primarily present proteolytically degraded protein fragments derived from proteins present in the cell. MHC class II is expressed by professional antigen presenting cells that typically take up extracellular protein, degrade it with lysosomal proteases, and present protein fragments on the surface. In humans, the MHC molecules are known as human leukocyte antigens (HLA), which in the present invention are the preferred MHC molecules to evaluate binding to.
[0037] A "T-cell epitope" is an MHC binding peptide, which is recognized as foreign (non-self) by a T- cell in a vertebrate due to specific binding between a T-cell receptor and the cell carrying the MHC-peptide complex on its surface. Hence, a peptide, which constitutes a T-cell epitope in one individual will not necessarily be a T-cell epitope in a different individual of the same species. First of all, two individuals having differing MHC molecules that bind different sets of peptides, do not necessarily present the same peptides complexed to MHC, and further, if a peptide is autologous in one of the individuals it may not be able to bind any T-cell receptor.
[0038] "EvaxMHC4" denotes an embodiment of the present invention, i.e. the ANN-based prediction model described in detail herein. However, the term is used interchangeably with the broader concepts falling under the various aspects of the present invention. Specific embodiments of the invention
[0039] First aspect of the invention and embodiments thereof.
[0040] The first aspect of the invention relates to a method for identifying peptides, which are bound and / or presented by MHC class I and / or MHC class II molecules in an individual, the method comprising a) providing a transformer encoder-decoder model, which evaluates whether amino acid sequences of peptides are ligands for selected MHC molecules, and which comprises an artificial neural network (ANN) architecture, which comprises an input layer, an output layer and one or more hidden layers, b) evaluating, in the transformer encoder-decoder model, data pairs, which each comprises 1) an MHC molecule amino acid sequence as well as 2) a peptide amino acid sequence, c) outputting from the transformer encoder-decoder model, for each inputted data pair, at least one quantitative assessment of the ability of each MHC molecule having the MHC molecule amino acid sequence of a data pair to bind and / or present the peptide having the peptide amino acid sequence of the same data pair, wherein the input layer receives and accepts data pairs each composed of an MHC molecule amino acid sequence and an amino acid sequence of a peptide, wherein the transformer encoder receives the MHC molecule amino acid sequence of a data pair with residue embedding and positional embedding as input and outputs a highdimensional representation of the MHC amino acid sequence, wherein the transformer decoder receives the high-dimensional representation of the MHC molecule amino acid sequence and the peptide amino acid sequence of a data pair with residue embedding and positional encoding as input and outputs a high-dimensional representation of the relationship between the MHC molecule amino acid sequence and the peptide sequence, wherein the output from the transformer decoder is used to calculate the quantitative assessment of the ability of the MHC molecule having the MHC molecule amino acid sequence of a data pair to bind and / or present the peptide having the peptide amino acid sequence of the same data pair. It is convenient that the quantitative assessment output from the transformer encoderdecoder model is normalized and transformed into at least one probability indication, e.g. a real number between 0 and 1. However, any probability indication can be used (such as a percentage) as long as it translates into a probability ranging between zero probability and complete certainty.
[0041] In preferred embodiments, the at least one quantitative assessment comprises or consists of a calculation of affinity between the peptide having the peptide amino acid sequence of the data pair and the MHC molecule having the MHC molecule amino acid sequence of the data pair. In the connection, the "affinity" is a measure of the strength of binding between the two constituents of the data pair and will typically be expressed as a binding constant (or its inverse, the dissociation constant).
[0042] The method may also include that at least one quantitative assessment comprises or consists of a calculation of probability of presentation of the peptide on any cells that express the MHC molecule of the same data pair. This calculation may be combined with the calculation of the affinity, meaning that preferred embodiments comprise that the quantitative assessment comprises or consists of both a calculation of affinity and a calculation of the presentation probability.
[0043] The data pairs may be input to the input layer of the transformer encoder-decoder model individually or in the form of a matrix or vector, which comprises multiple data pairs. This choice mainly dictates how the subsequent actions by the model are carried out (in sequence or as an integrated operation etc).
[0044] In the embodiments described above, it is possible to allow one or more peptide amino acid sequences be paired with one or more MHC molecule amino acid sequences; put differently, the method allows that the same peptide amino acid sequence is paired with several different MHC molecules and vice versa.
[0045] As will be explained in the example below, the method of the first aspect of the invention and its various embodiments allows for evaluation of binding of peptides to both MHC class I and MHC class II molecules; hence, the MHC molecule amino acid sequences are MHC class I amino acid sequences, MHC class II amino acid sequences, or both.
[0046] In order to ensure that evaluation of presentation by / binding to both MHC Class I and II is optimized, the transformer encoder-decoder model has preferably been trained on peptide:MHC datasets comprising peptides paired with MHC Class I and peptides paired with MHC Class II. Further, one embodiment of the invention utilises that the transformer encoder-decoder model has been (at least partially) trained on peptide-MHC datasets comprising data pairs generated by another prediction model capable of predicting peptides that will be bound and / or presented by MHC molecules.
[0047] A useful feature of the present approach is that the quantitative assessment for each peptide-MHC molecule data pair can be used to rank the list of data pairs. In turn, this provides useful information, for instance when it is desired to develop a vaccine that will induce immunity against polypeptides that comprise peptides evaluated according to the method of the first aspect of the invention. After the evaluation and ranking according to a probability score, a vaccine can be prepared which relies on induction of immunity against those peptides that are ranked with the highest probabilities of being binders to / presented by MHC molecules in a given patient.
[0048] As will be clear from the following, the 1stand 2ndaspects of the invention are closely related. Hence, the transformer encoder-decoder model detailed above is in preferred embodiments an artificial neural network (ANN) capable of identifying peptides, which is provided according to the method the 2ndaspect of the invention or any of the embodiments thereof discussed below.
[0049] 2ndaspect of the invention and embodiments thereof.
[0050] This aspect relates to a method for training a predictor model capable of identifying peptides, which will be bound and / or presented by selected MHC class I and / or MHC class II molecules in an individual, the method comprising utilizing a conditional generative adversarial network (GAN) in the form of a generator model and a predictor model, wherein the predictor model evaluates whether inputted amino acid sequences of peptides are bound and / or presented by inputted MHC molecules, the method comprising a) evaluating, in the predictor model, data pairs, which each comprises 1) an MHC molecule amino acid sequence as well as 2) a peptide amino acid sequence, wherein, in a fraction of said data pairs, the amino acid sequences in the data pairs are those of verified peptide-MHC complexes, and wherein another fraction of said data pairs are random pairs of MHC molecule and peptide sequences, and wherein, in the remaining data pairs, the amino acid sequences in the data pairs are from molecules that are generated by the generator model, b) outputting from the predictor model, for each inputted data pair, at least one quantitative and / or qualitative assessment of the ability of each MHC molecule having the MHC molecule amino acid sequence of a data pair to bind and / or present the peptide having the peptide amino acid sequence of the same data pair, c) outputting from the predictor model a quantitative and / or qualitative assessment of whether the peptide sequence of the inputted data pair is a generated sequence or a real, naturally occurring sequence, d) evaluating the output in steps b and c and adjusting weights of the neurons in the one or more hidden layers of the predictor model to minimize the error of the output or to optimize the performance of the model, e) evaluating the output in step c and adjusting weights of the neurons in the one or more hidden layers of the generator model to optimize the model's ability to generate a peptide sequence that the predictor predicts to be a real, naturally occurring peptide sequence, f) repeating steps a-e to train the predictor model to correctly identify peptides, which will be bound and / or presented by selected MHC class I and / or MHC class II molecules in an individual, wherein the input layer in the predictor model receives and accepts data pairs composed of an amino acid sequence derived from an MHC molecule and an amino acid sequence of a peptide.
[0051] The data pairs used in the model training typically comprises peptide:MHC datasets comprising peptides paired with MHC Class I and peptides paired with MHC Class II, since this allows the predictor model to reliably output predictions in respect of binding to both MHC classes. At any rate, data pairs used in at least part of the model training comprises data pairs generated by another prediction model capable of predicting peptides that will be bound and / or presented by MHC molecules.
[0052] Typically, the random pairs of MHC molecule and peptide sequences are pairs each consisting of a randomly sampled naturally occurring peptide sequence and a randomly sampled MHC molecule sequence. With respect to the output in step b, this is typically selected from i) a quantitative or qualitative assessment of whether the peptide having the peptide amino acid of a data pair is a generated peptide sequence or a naturally occurring peptide sequence, and / or ii) a quantitative assessment of affinity between the peptide having the peptide amino acid sequence of the data pair and the MHC molecule having the MHC molecule amino acid sequence of the data pair, and / or iii) a quantitative assessment of (the likelihood of) presentation of the peptide having the peptide amino acid sequence of the data pair by the MHC molecule having the MHC molecule amino acid sequence of the data pair.
[0053] The predictor model typically comprises an artificial neural network (ANN) architecture comprising an input layer, an output layer and one or more hidden layers, hence the use of the terms "neurons" and "input layer", "output layer", and "hidden layers".
[0054] The predictor model in preferred embodiments comprises a transformer encoder-decoder model as described above in relation to the method of the 1staspect of the invention and the various embodiments thereof.
[0055] Also the generator model may (preferably) comprise an artificial neural network architecture comprising an input layer, an output layer and one or more hidden layers.
[0056] In preferred embodiments the generator model comprises a transformer encoder-decoder architecture that pairwise encodes MHC molecule amino acid sequence data and randomly sampled digits, which are decoded to output a generated peptide amino acid sequence defined by the encoded MHC molecule amino acid sequence.
[0057] In some embodiments of particular importance, the generator- and predictor models encode MHC molecule amino acid sequence input using positional embedding and encode either peptide sequence input or randomly sampled digits using positional encoding and feed the encoded information to be decoded to their respective decoders.
[0058] In one convenient set of embodiments where all of assessments i)-iii) are carried out, the predictor model comprises 2 terminal neural networks, one which outputs the quantitative assessments described above in points II and iii, and one which outputs the quantitative or qualitative assessment described above in point i.
[0059] In cases during training of the models where the input data for the predictor model comprises more than one MHC molecule amino acid sequence per peptide ( / .e. several pairs are input, where the same peptide is paired with different MHC molecules), then the model weights are only adjusted based on the output composed of the amino acid sequences from the most likely peptide-MHC molecule complex for each peptide. Phrased differently, it is ensured that less probable peptide-MHC molecule complexes do not influence the adjustment of weights during the training of the model.
[0060] In preferred embodiments of the 2ndaspect of the invention, the outputs from the predictor model are normalized and transformed into probabilities, preferably a number between 0 and 1 but into any format that can be transformed into a probability between zero probability and full certainty. Further, a rank calibration is typically carried out by generating raw predictions for a plurality of random peptides, sorting the raw predictions, selecting a number of reference points from the raw prediction to rank the list, and using these reference points to interpolate a rank score for any given raw prediction. Typically, the rank calibration is carried out for the plurality of random peptides in respect of each of a plurality of MHC molecules; see also the example for more details in a specific embodiment of the 2ndaspect of the present invention.
[0061] When the outputs from the predictor model are normalized and transformed into probabilities, a probability calibration can additionally be performed by generating raw predictions on an external calibration dataset, comprising data pairs composed of an amino acid sequence derived from an MHC molecule and an amino acid sequence of a peptide, where this external calibration dataset comprises known positive and negative data points not used in the training, sorting the predictions and calculating local predictive performance metric scores in rolling windows along the sorted list of raw predictions, then fitting a mathematical function to the local performance metric scores such that any raw prediction can be transformed to a performance score. An example of a predictive performance metric score is the precision score, but as set forth in the Example, other predictive performance scores can be applied (such as threshold, Fl, Recall, ROC AUC, and ROC AUC10; see example for details). Also, a joint probability calibration can be carried out using raw predictions from the predictor model, combined with expression values of the peptide or expression values of the source protein of the peptide; here the rationale is to be able to exclude peptides that might be excellent binders of MHC, but which are expressed at very low levels in the target tissue.
[0062] Such a joint probability calibration can according to embodiments of the 2ndaspect be carried out by generating raw predictions on an external calibration dataset with known expression values, comprising both positive and negative data points not used in the training, then a) creating bins of raw predictions and expression values, b) sorting each peptide-MHC data point in the external calibration dataset into a raw prediction bin as well as an expression value bin, c) creating a matrix of pairs of raw prediction and expression value bins, d) for each pair of bins in the matrix, calculating a predictive performance metric score for all peptides that fall into both the raw prediction bin and the expression value bin, and e) fitting a mathematical function to the predictive performance metric scores for each pair of bins, such that any pair of raw prediction and expression value can be transformed to a predictive performance score.
[0063] Also in this case, a useful example of a predictive performance metric score is the precision score, but as set forth in the previous probability calibration, other predictive performance scores can be applied (such as threshold, Fl, Recall, ROC AUC, and ROC AUC10; see example for details).
[0064] Also, in the embodiments where a probability calibration is performed, it can preferably be carried out for the plurality of random peptides in respect of each of a plurality of MHC molecules.
[0065] 3rdaspect of the invention and embodiments thereof
[0066] This aspect relates to a computer or computer system comprising a transformer encoderdecoder model as defined above in respect of the 1staspect of the invention.
[0067] As such the computer and computer system therefore is one which implements the method of the first aspect of the present invention.
[0068] In preferred embodiments, the executable code is an ANN as defined above in embodiments of the first aspect of the invention. Such an ANN is typically programmed to - at least during training - adjust the weight functions that define each node (or neuron) in the network. The number of hidden layers in such an ANN can vary.
[0069] The interface(s) for inputting amino acid sequences of peptides is / are typically selected from any device for inputting data into a computer memory or storage medium: in principle, a simple keyboard connected to the computer can serve this purpose, but typically amino acid sequence data will be read from an external data carrier or data source by a connected disk drive or other data carrier (a memory stick, memory card, network associated storage) or via a network or internet connection and a suitable protocol for file transfer (FTP, FTPS, SFTP, CSP, HTTP or HTTPS, AS2, 3-, and -4, or PeSIT).
[0070] Likewise, the computer will comprise storage in the form of any convenient data carrier or storage medium (a hard drive, a solid-state hard drive, a memory stick) but also directly in the memory (RAM) of the computer or computer system. The storage format can be any convenient format such as in the form of records in a relation database (both row-oriented and column-oriented), an object database, but also as entries in text files (e.g. as comma separated values or a suitable XML format), or as a simple file system or other similar root- and-tree structure.
[0071] The executable code(s) in the computer or computer system is capable of accessing the linked input devices and storage media as well as the computer working memory in order to perform the necessary operations specified in the first aspect of the invention.
[0072] Output from the computer can also take any convenient form, but will typically be output to a storage medium or external data carrier.
[0073] 4thaspect of the invention and embodiments thereof
[0074] This aspect relates to a method of identifying protein-derived MHC interacting peptides, which will be bound and / or presented by at least one MHC molecule expressed by an individual, the method comprising a) determining the amino acid sequences of one or more MHC molecules expressed by said individual, b) deriving amino acid sequences of potential MHC interacting peptides from one or more proteins; c) evaluating pairwise the amino acid sequences from steps a and b by using said amino acid sequences as input in step b in the method of the 1staspect of the invention and any embodiment thereof, and / or by querying a database comprising records consisting of interacting pairs of MHC molecules and peptides identified by the method according the method of the 1staspect of the invention and any embodiment thereof for the existence in the database of interacting pairs consisting of the MHC molecules and potential MHC interacting peptides from steps a and b respectively, d) identifying those peptides that are 1) ranked with the highest scores or 2) which provide a positive query result as those that will be bound and / or presented by at least one MHC molecule in the individual.
[0075] The protein(s) in step b could be derive from a malignant neoplasm in the individual or from an infectious agent, such as a virus, thus facilitating identification of vaccine candidate peptides derived from the individual's malignant neoplasm or the infectious agent. A particular interesting embodiment in this context arises when the amino acid sequences of potential MHC interacting peptides identified in step b are potential neoantigens or neoepitopes, since neoantigens / neoepitopes generally constitute safe vaccine agents that are unlikely to cause adverse events due to induced autoimmunity. Along the same lines of reasoning, in some embodiments of the 4thaspect, the amino acid sequences of potential MHC interacting peptides identified in step b are derived from antigens expressed specifically in the malignant neoplasm, such as antigens derived from endogenous retroviruses or endogenous retroviral elements.
[0076] In one set of embodiments of the 4thaspect, the presence of the one or more proteins are determined from the transcriptome of malignant cells in the individual. Also, step a and / or step b can comprise sequencing of material in a sample from the individual.
[0077] EXAMPLE 1
[0078] Establishment and test of prediction model for MHC Class I and II binding and presentation
[0079] Curating MHC ligand- and peptide-MHC binding affinity datasets for model training :
[0080] In vitro peptide-MHC binding affinity data (IC5o) was curated from IEDB www.iedb.org')
[0012] and O'Donnell et al. (2020) [5]. For peptide-MHC pairs with multiple measurements, the median value was used to represent the binding affinity. IC5o measurements were transformed using 'l-log50000(IC5o nM)' as described in
[0013] .
[0081] Immunopeptidomics datasets were generated and curated from several public sources. Generated MHC class I datasets comprise immunopeptidomics of murine cell lines (CT26, B16F10 and 4T1) and transformed human C1R cell lines. Peptide elutions on the murine cell lines were done using MHC locus-specific antibodies. The C1R transformants express HLA- A02:01, HLA-B07:02, HLA-B15 :25, HLA-B41:02, HLA-B44:02, HLA-B44:05, HLA-B51:06, and HLA-C16:02, respectively, and were eluted using a pan-HLAI antibody (H / 6 / 32). Generated MHC class II datasets comprise immunopeptidomics of the human cell lines 9013-SCHU, 9022, 9031, 9087, C1R, JESI, JESTHOM, and LCL-721.221 and were eluted using pan-HLAII and HLA-locus specific (DR, DP, DQ) antibodies. Details on the MHC Class II datasets can be found in Garde et al. (2019)
[0013] . The immunopeptidomics datasets were generated according to the protocols described by Purcell et al. (2019)
[0014] and Garde et al. (2019)
[0013] .
[0082] The sources of immunopeptidomics include IEDB, the HLA ligand atlas [5] and the supporting material from several scientific publications [6, 8, and 16-56].
[0083] The dataset was subset to peptide-MHC pairs with peptides of lengths 8 through 12 amino acids for MHC class I and 13 through 19 amino acids for MHC class II. The negative complement to the immunopeptidomics dataset was represented by randomly sampled decoy peptides. The decoy peptides were sampled from the human proteome in a ratio of 50 : 1 relative to the immunopeptidomics data using a uniform peptide length distribution spanning from 8-12-mers for MHC class I and 13-19-mers for MHC class II as described in
[0013] .
[0084] Dataset partitioning
[0085] The data were split into an MHC class I dataset and an MHC class II dataset. Each of these datasets was partitioned by first computing all possible 9-mer peptide binding cores as described in
[0057] and
[0013] . Briefly, for MHC class I, 9-mer cores were generated by wildcard amino acid insertions for peptides of length 8, while deletions representing bulging peptides were imposed on peptides longer than 9 amino acids. For MHC class II, 9-mer cores were generated by 9-mer sliding windows along the peptide. Clustering was then performed based on shared 9-mer peptide binding cores, followed by aggregation into 6 partitions. This partitioning scheme ensures that there are no 9-mer peptide binding core shared between partitions.
[0086] Fig. 3 and Fig. 4 show summary statistics on the dataset in respect of MHC class I and - II, respectively. Partition 0 to 4 in both datasets are used as training and validation sets while partition 5 is reserved as a test set. Panel A) in Fig. 3 shows the count of mono- and poly-allelic ligands (immunopeptidomics derived) in each partition. Poly-allelic denotes ligands derived from cell lines or samples which contain multiple MHC molecules on their surface, precipitated with a pan-specific antibody. IC5o data is from affinity assays. Panel B) shows sampled negative "ligand" data, where data is sampled with identical MHC allele frequency as positive data. Panel C) shows the count (Y-axis) of each allele (X-Axis). For poly-allelic data, each allele is counted. Finally, panel D) shows the peptide length distribution for each partition. MHC Class I ligand length is naturally biased towards 9-mers. Peptides shorter than 8 or longer than 12 are excised from the dataset.
[0087] Panel A) in Fig. 4 shows the count of mono- and polyallelic ligands (immunopeptidomics derived) in each partition. Poly-allelic denotes ligands derived from cell lines or samples, which contain multiple MHC molecules on their surface, precipitated with a pan-specific antibody. IC5o data is from affinity assays. Panel B) shows negative "ligand" data that are sampled with identical MHC allele frequency as positive data. Panel C) shows the count (Y- axis) of each allele (X-Axis). For poly-allelic data, each allele is counted. Finally, panel D) shows peptide length distribution for each partition. MHC class II ligand length is naturally biased towards 15-mers with a broader distribution than MHC class I. Peptides shorter than 13 or longer than 19 are excised from the dataset.
[0088] Developing a GAN + Transformer Framework for Peptide-MHC Prediction
[0089] In order to set up a framework for peptide-MHC prediction, the transformer architecture
[0050] was used. An encoder-decoder setup was used, where the encoder consumes the sequence of the MHC, and the decoder receives the coded MHC and the peptide.
[0090] To improve performance, a conditional GAN
[0058] was set up. The generator takes an MHC sequence and random noise and transforms it into a generated "fake" peptide. The predictor (critic) predicts whether the input peptide is real or fake using the Wasserstein criterion
[0059] , as well as whether the peptide is bound and / or presented by the provided MHC molecule.
[0091] The model is coded in PyTorch (version 1.10) using the native transformer implementation. The full architecture is shown in Fig. 1. Peptides and MHC sequences are encoded using a BLOSUM encoding scheme. MHC sequences correspond to MHC pseudo-sequences as defined above.
[0092] A training loop was implemented where the predictor is trained every iteration and the generator is trained every second iteration. The generator loss is defined as the negative mean of the GAN output. The generator and the predictor use separate RMSprob optimizers.
[0093] The predictor loss is the sum of the mean of the GAN output and the predictor output multiplied by a variable weight and a gradient penalty
[0065] .
[0094] Training MHC Class I Using Multi-Instance Learning
[0095] While most data points are mono-allelic and consist of a defined peptide-MHC pair, some data points are poly-allelic and consist of a peptide sequence which may have been presented by one or more MHC alleles. Predictions are created for each potential allele binding to a peptide, and a maxpool is used to choose the most likely peptide-MHC pair. Gradients are thus propagated only on the best prediction for a given peptide. When generating fake peptides using poly-allelic data, a random allele is chosen from the possible set of alleles.
[0096] The first 30 epochs of training is reserved to training only the GAN, after which the weights on the peptide-MHC ligand prediction are increased slowly until it dominates the loss function over 100 epochs. To prime the critic for poly-allelic data, a mono-allelic burn-in period of 15 epochs is used.
[0097] Training MHC Class II Models Using Knowledge Transfer
[0098] The model used for MHC class II predictions is identical to the model used for MHC class I. The training is however slightly different, as the MHC class II model fitting suffered from overfitting. To alleviate the overfitting and boost the performance, a knowledge-transfer strategy was used. Predictions from the previous version of the tool (EvaxMHC3) was used as a target for the first 190 epochs, after which mono-allelic data was used as training data for 32 epochs and finally poly-allelic data was used for 8 epochs. The key aspects of the EvaxMHC3 model and training are described in reference
[0013] .
[0099] Knowledge transfer is implemented by extracting 50x 10® peptides from the UniRef30 database
[0060] , a clustering of the Uniprot database. The UniRef30 database is sampled by successively sampling by random a cluster, a sequence within the cluster and finally 5 peptides of length 12-18 from the protein sequence. Sampling from a clustered database ensures a fair representation of the "global" protein space. Randomly selected MHC alleles are assigned to each peptide and the synthetic dataset is created by setting the raw EvaxMHC3 prediction output as the eluted ligand and affinity outputs. 105random MHC class II peptides are selected as the first 105of the 50x10® sampled peptides. 10® random MHC class I peptides are selected as the first 10® of the 50x 10® sampled peptides including cutting the length to 8-12 amino acids. Rank calibration
[0100] A common method to normalize between alleles is the concept of rank calibration. Rank normalization in EvaxMHC4 is accomplished by generating predictions for 1,000,000 random peptides drawn from the UniRef30
[0060] database with a flat distribution of peptide lengths (8-11 for MHC class I and 12-18 for class II). Predictions are sorted and 25 log-spaced points are picked out as representative of the calibration curve. Calibrations are performed individually for each MHC allele. See Fig. 5, which shows Examples of 2 alleles each for MHC class I and II. Raw prediction output is fit on a sorted list of a million random peptides. Points marked on the curves are extracted and used to calculate the rank based on the prediction value through linear interpolation.
[0101] Probability Calibration
[0102] The method of the invention exploits data-driven calibrated output. More precisely, EvaxMHC4 introduces an allele-specific precision calibration. Alleles for which data exists are calibrated by scanning a window across predictions from the test set and calculating the precision for each window. More precisely, raw predictions are created for each allele in the test partition of the data set. For each allele, predictions are sorted and a rolling window of size 50 is used to calculate the average prediction value and the local precision for each window. The function ligand_prob(raw_prediction) = a ■ e-b(raw-prediction+c)is fitted to the window values of each allele and the variables a, b, c are saved and used for calculating the estimated probability given a raw prediction. To calculate the generic calibration curve, the same calculation is done for all alleles in a single dataset.
[0103] Alleles are clustered by their intersection-over-union (loll) similarity (also known as the Jaccard index) to each other using the top 20,000 predictions of 1,000,000 random predictions (using the same peptides as in the rank calibration). Alleles with an loll higher than 0.5 are clustered together and a common probability calibration curve is calculated. Alternatively, alleles missing a calibration curve are filled using the closest allele measured by loll similarity, provided similarity is higher than 0.5. Alleles which are not clustered with any other allele with calibration data available are given a global calibration curve calculated over all samples. Fig. 6 shows an example of results of calibration data for 82 individual MHC Class I allele clusters and 23 MHC Class II clusters. The generic calibration curve is shown in lighter shading.
[0104] Probability Calibration with Expression Data
[0105] Abundance of proteins being processed to be expressed on MHC class I directly affect the probability of a ligand being detectable, both in an immunopeptidomics setting and as a ligand presented to the immune system. Therefore, an additional calibration can be performed when expression data is available. Based on paired immunopeptidomics data from a dataset, a calibration matrix can be constructed where a ligand probability and an expression value can be mapped to a shared probability.
[0106] First, expression levels are mapped to a probability through a rolling window of sorted expression values of known eluted ligands, random non-ligand peptides and associated expression values:
[0107] To estimate the probability of a correct ligand prediction given only the expression value, positive peptides are extracted from the dataset, as well as random peptides in 1 :50 ratio. Each peptide is associated with the gene expression value taken from the same sample. Peptides are sorted by expression value and a rolling window of 5,000 is used to calculate the mean expression value and the precision for each window. The function ligand_prob(expression) = a ■ tanh(b ■ expression) is fitted to the window values and the a, b is saved to convert expression values to a ligand probability.
[0108] Finally, ligand probability and expression probability are integrated into a matrix mapping to a joint probability:
[0109] The probability curves from the ligand predictions and expression values are binned into 50 and 19 values, respectively. Each peptide is binned into a matrix of pairs of bins, and for each pair a probability value is calculated from the peptides that fall into each bin. To ensure the calibration is monotonically increasing in both expression values and ligand predictions, each element in the matrix is set to the maximum value of all indexes lower or equal than the current index. To calculate a joint probability from expression probability and ligand probability, a 2D linear interpolation is used. Performance Evaluation
[0110] To benchmark the performance of EvaxMHC4, a hold-out validation set was created when setting up the dataset (see remarks above in relation to Figs. 3 and 4). Hold-out sets were created by partitioning, in order to prevent ligands to have overlapping 9-mers between partitions. Fig. 7 shows EvaxMHC4 compared to other MHC class I ligand predictors— MHCflurry [5], MixMHCpred
[0061] and NetMHCpan
[0062] as well as EvaxMHC3 described above. EvaxMHC4 outperforms all other predictors on average precision (Fig. 7A), ROC AUC (Fig. 7B) as well as other metrics shown in the following table:
[0111] MHC class II prediction performance is shown in Fig. 8, where EvaxMHC4 performance is compared to EvaxMHC3, MixMHC2pred
[0063] , NetMHCIIpan
[0062] , and BERTMHC
[0064] . Compared to MHC class I, MHC class II performance is in general lower across all predictors, nevertheless with EvaxMHC4 as the best performing model. The table below shows performance metrics for MHC class II predictors:
[0112] Fig. 9 shows performance of EvaxMHC4 calculated per allele relative to other predictors which shows that EvaxMHC4 consistently outperforms other models in average precision and ROC AUC. Last panel of Fig. 9 shows that recall is generally higher for EvaxMHC4, with some predictors having a slightly higher performance for a few alleles.
[0113] Fig. 10 shows that the highest performance increases are for alleles which generally performed worse on other MHC ligand predictors. This may suggest the architecture in EvaxMHC4 is better at capturing general patterns in the MHC sequence across alleles. Fig. 11 shows a similar picture for MHC class II alleles, with most alleles present in the upper left corner indicating EvaxMHC4 has higher performance for the majority of the alleles evaluated.
[0114] Comparison of EvaxMHC4 with other state-of-the-art models trained on identical datasets
[0115] The above-described datasets were used to train the EvaxMHC4 models ( / .e. the models according to an embodiment of the present invention), comprising transformer encoderdecoder models using a GAN framework as described above. For comparison, additional models were trained on exactly the same datasets.
[0116] The performance of each trained model was evaluated on the hold-out validation dataset. The hold-out dataset was split into smaller datasets defined by MHC alleles. Predictions were generated for each data point and performance metrics were calculated for each MHC allele dataset. As summarized in Fig. 2, the average precision of the EvaxMHC4 model was superior to the prior art MHC prediction models. This improvement was seen for both MHC class I and -II allele datasets and was statistically significant.
[0117] EXAMPLE 2
[0118] Immunogenicity of epitopes identified according to the invention compared to ...
[0119] Animals
[0120] 6 to 8 weeks old female BALB / c mice were acquired from Janvier Labs (France). The mice were acclimated for one week before initiation of experiments. The experiment was conducted under license 2017-15-0201-01209 from the Danish Animal Experimentation Inspectorate in accordance with the Danish Animal Experimentation Act (BEK nr. 12 of 7 / 01 / 2016), which is compliant with the European Directive (2010 / 63 / EU). Vaccine formulation
[0121] The pTVG4 vector is generated from the standard plasmid pUMVC3 (cat #4010) acquired from Aldevron (North Dakota, USA) (containing a CMV-driven expression vector and kanamycin resistance gene for selection), by cloning two copies of a 36-bp CpG-rich immunostimulatory sequence (ISS) containing the 5'-GTCGTT-3' motif downstream of the multi-cloning site (cf. US 2004 / 0142890).
[0122] Plasmid inserts contained a start codon, the epitope insert relevant for each construct (see tables below) and a stop codon. Each of the 13 epitopes encoded by each DNA construct were of 27mer length. The DNA was prepared according to formulation calculations, was aliquoted and stored at -20 °C until use.
[0123] The epitopes identified with EvaxMHC3 and -4, respectively, and encoded by the 2 plasmids, were the following :
[0124] Sequences identified with EvaxMHC3 SEQ ID NO: Position in plasmid
[0125] NDEEEAATTSEVSPPSPMVSRLRGRRD 1 1
[0126] PRSQGPQRYGNRFVRTQEAVREATQED 2 2
[0127] EPPPWVKPFVSPKLSPSPTAPILPSGP 3 3
[0128] FHDPEMSKFTNSPSLQAHLQALQAVQR 4 4
[0129] GILEGDPHISSPRTLTLAANQALQKVE 5 5
[0130] GKPRLLKTDNGPAYTSQKFQQFCRQMD 6 6
[0131] AASDFLLMPQMSIQPVPVEPIPSLPPG 7 7
[0132] MGSDNGPAFVSQVSQGLARQLGTNWKL 8 8
[0133] DPLNMAAITQPPTPQVPLTITPEIPSR 9 9
[0134] RILRYSDEMGTRLSPFPAARAKRATVA 10 10
[0135] SSHDFCVMIQLLPRVYYHPASSLEESY 11 11
[0136] TLFNWGPDQQKAYQEIKQALLTAPALG 12 12
[0137] DKTQECWLCLVSGPPYYEGVAVLGTYS 13 13
[0138] Sequences identified with EvaxMHC4 SEQ ID NO: Position in plasmid
[0139] SGIQGPYLNMIKAIYSKPVANIKVNGE 14 1 VGAKRKYGSQNKYTGLSKGLEPEEKFR 15 2 RQMDVTHLTGLPYNPQRQGIVERAHRT 16 3 ADQTNYHWGAYAQISSTAIRAWKALSR 17 4 GILEGDPHISSPRTLTLAANQALQKVE 5 5
[0140] AALSS PVE AARN FH N N FH VTAETLRSR 18 6 H WG AYAQVS STAI RAW KALS RAG ETTG 19 7 SILEGDPHISSPRALTLAANQALQKVE 20 8 Sequences identified with EvaxMHC4 SEQ ID NO: Position in plasmid
[0141] TYHSPSYVYHQFERRAKYKREPVSLTL 21 9
[0142] LKIVRKEVWEQLKETYVAGDTQVPHQF 22 10
[0143] EKSGTQGPYLSIVKAIYSKQVSNIKLN 23 11
[0144] LILSNENAPSGGYSAKAKNIMAKMGYK 24 12
[0145] RISVVQALVLTQQYHQLKSIGPEKVKS 25 13
[0146] It is noted that both identification methods resulted in inclusion of SEQ ID NO: 5 in the corresponding vaccine plasmid.
[0147] For the two first immunizations, electroporation was used for the in vivo delivery of DNA. For the last three immunizations, the different DNA solutions were mixed with in-house prepared 9% w / v pololoxamer 188 (Kolliphor®) in 3xPBS. Mixing was performed in a ratio of 2 parts of DNA (100 pl) to 1 part of Kolliphor® (50 pl). In the final solution, there was a concentration of 3% w / v P188 in l xPBS and a dosing volume of 150 pl per mouse. Mice were immunized in left and right tibialis anterior muscles (i.m.) with 2x75 pl. Kolliphor was supplied by BASF (Germany). Until use, Kolliphor® was stored at 4°C and shed from light.
[0148] Immunizations and challenge
[0149] A total of 5 i.m. immunizations were made in 39 mice: on days -14, -7, 3, 9 and 16 relative to a primary challenge with CT26 tumour cells (2 x 105cells per mouse). DNA vector was, as mentioned above, delivered in immunizations 1 and 2 using electroporation, whereas the last three immunizations used DNA formulated with Kolliphor® instead. 13 mice received a vector encoding EvaxMHC4 predicted ERV peptides from CT26, 13 mice received a vector encoding EvaxMHC3 predicted ERV sequences from CT26, and 13 mice received a mock plasmid. Spleens were isolated after the study and processed into single-cell suspensions which were then used for ex vivo analyses. Tumour volume was measured after tumour inoculation.
[0150] IFNy enzyme-linked immune absorbent spot (ELISPOT) assay
[0151] To investigate whether the different treatment groups harbour immune recognition of ERVs encoded by the different DNA constructs, 5xl05viable splenocytes of single-cell suspensions from individual mice were pooled group-wise and plated on the wells of an ELISPOT plate coated with anti-IFNy capture antibody. Cells were then re-stimulated with the ERV peptide pools encoded by the respective DNA vaccine. The final peptide concentration for each peptide stimuli in the pool was 5 pg / ml. Each sample was also incubated with RIO medium only (unstimulated control). Cells were subsequently incubated with the peptide stimulants or RIO media at 37°C overnight. During the incubation, IFNy released by re-activation of peptide-specific T cells is captured on the bottom of the plate. The cytokine was visualized the following day through addition of anti-IFNy detection antibody, followed by addition of streptavidin-HRP enzyme and AEC chromogen substrate. ELISPOT plates were air dried and read on ELISPOT analyzer to count Spot Forming Units (SFUs) in each well. All counts were normalized to SFUs per 105splenocytes.
[0152] Results
[0153] Fig. 12 shows the development of CT26 tumour growth in 3 groups of mice after challenge with CT26 cells. As is clear, the group immunized with the plasmid encoding EvaxMHC4 predicted epitopes shows complete tumour growth prevention,, as opposed to the 2 other groups.
[0154] Likewise, Fig. 13 shows that splenocytes from mice immunized with the vector encoding the EvaxMHC4 predicted ERV epitopes exhibits a >8 times higher reactivity when re-stimulated with the peptides encoded by the vector.
[0155] Both results underline the superior performance of the method of the invention in predicting true ligand of MHC molecules.
[0156] LIST OF REFERENCES
[0157] [1] Ashish Vaswani et al., 2017, doi: 10.48550 / ARXIV.1706.03762.
[0158] [2] Mehdi Mirza and Simon Osindero, 2014, doi: 10.48550 / ARXIV.1411.1784.
[0159] [3] Martin Arjovsky et al., 2017, doi: 10.48550 / ARXIV.1701.07875.
[0160] [4] Milot Mirdita et al., Nucleic Acids Research, 45(D1) :D17O-D176, November 2016, doi: 10.1093 / nar / gkwl081.
[0161] [5] Timothy J. O'Donnell et al., Cell Systems, ll(l) :42-48.e7, July 2020, doi: 10.1016 / j.cels.2020.06.010.
[0162] [6] Michal Bassani-Sternberg et al., PLOS Computational Biology, 13(8) :el005725, August 2017, doi : 10.1371 / journal.pcbi.1005725. [7] Birkir Reynisson et al., Nucleic Acids Research, 48(W1):W449-W454, May 2020, doi: 10.1093 / nar / gkaa379.
[0163] [8] Julien Racle et al., Nature Biotechnology, 37(11): 1283-1286, October 2019, doi: 10.1038 / S41587-019-0289-6.
[0164] [9] Jun Cheng et al., Bioinformatics, June 2021., doi: 10.1093 / bioinformatics / btab422.
[0165]
[0010] Ishaan Gulrajani et al., 2017, doi: 10.48550 / ARXIV.1704.00028.
[0166]
[0011] David Gfeller et al., Cell Syst. 2023 Jan 18; 14(1): 72-83. e5.
[0167]
[0012] Randi Vita et al., Nucleic Acids Research, 47(D1) :D339-D343, October 2018. doi: 10.1093 / nar / gkyl006.
[0168]
[0013] Christian Garde et al., Immunogenetics, 71(7) :445-454, June 2019, doi: 10.1007 / s00251-019-01122-z.
[0169]
[0014] Anthony W. Purcell et al., Nature Protocols, 14(6) : 1687-1707, May 2019, doi: 10.1038 / s41596-019-0133-y.
[0170]
[0015] Ana Marcu et al., Journal for ImmunoTherapy of Cancer, 9(4) :e002071, April 2021, doi: 10.1136 / jitc-2020-002071.
[0171]
[0016] M. Bassani-Sternberg et al. Mol Cell Proteomics, 14(3) :658-673, Mar 2015.
[0172]
[0017] A. Sofron et al., Eur J Immunol, 46(2) :319-328, Feb 2016.
[0173]
[0018] D. Ritz et al., Proteomics, 16(10) : 1570-1580, May 2016.
[0174]
[0019] A. Gloger et al., Cancer Immunol Immunother, 65(11) : 1377-1393, Nov 2016.
[0175]
[0020] Q. Wang et al., J Proteome Res, 16(1) : 122-136, Jan 2017.
[0176]
[0021] K. P. Karunakaran et al., J Proteome Res, 16(l) :298-306, Jan 2017.
[0177]
[0022] D. Ritz et al., Proteomics, Jan 2017.
[0178]
[0023] M. Bassani-Sternberg et al., Nat Commun, 7: 13404, Nov 2016.
[0179]
[0024] A. Alplzar et al., Mol Cell Proteomics, 16(2) : 181-193, Feb 2017.
[0180]
[0025] T. Fugmann et al. J Immunol, 198(3) : 1357-1364, Feb 2017.
[0181]
[0026] P. Pymm et al, Nat Struct Mol Biol, 24(4) :387-394, Apr 2017.
[0027] J. G. Abelin et al., Immunity, 46(2) :315-326, Feb 2017.
[0182]
[0028] M. S. Khodadoust et al, Nature, 543(7647) :723-727, Mar 2017.
[0183]
[0029] G. Kaur et al., Nat Common, 8: 15924, Jun 2017.
[0184]
[0030] D. Ritz et al., Proteomics, Oct 2017; 17(19) : 10.1002 / pmic.201700177. doi: 10.1002 / pmic.201700177.
[0185]
[0031] J. I. Mobbs et al., J Biol Chem, 292(42) : 17203-17215, Oct 2017.
[0186]
[0032] E. M. Scholz et al., Front Immunol, 8:984, 2017.
[0187]
[0033] M. Di Marco et al., .J Immunol, 199(8) :2639-2651, Oct 2017.
[0188]
[0034] C. Chong at al., Mo! Cell Proteomics, 17(3) : 533-548, Mar 2018.
[0189]
[0035] D. Ritz et al., Proteomics, 18(12) :el700246, Jun 2018.
[0190]
[0036] Y. T. Ting et al., J Biol Chem, 293(9) :3236-3251, Mar 2018.
[0191]
[0037] S. Yair-Sabag et al., Proteomics, 18(9) :el700249, May 2018.
[0192]
[0038] S. H. Ramarathinam et al., Proteomics, 18(12):el700253, Jun 2018.
[0193]
[0039] J. Lanoix et al., Proteomics, 18(12) :el700251, Jun 2018.
[0194]
[0040] M. C. Neidert et al., Acta Neuropathol, 135(6) :923-938, Jun 2018.
[0195]
[0041] C. I. DeVette et al., Immunol Res, 6(6) :636-644, Jun 2018.
[0196]
[0042] A. Sanz-Bravo et al., Mol Cell Proteomics, 17(7): 1308-1323, Jul 2018.
[0197]
[0043] A. Nelde et al., Oncoimmunology, 7(4) :el316438, 2018.
[0198]
[0044] L. Komov et al., Proteomics, 18(12) :el700248, Jun 2018.
[0199]
[0045] A. Sanz-Bravo et al., Mol Cell Proteomics, 17(8) : 1564-1577, Aug 2018.
[0200]
[0046] B. Shraibman et al., Mol Cell Proteomics, 17(11) :2132-2145, Nov 2018.
[0201]
[0047] H. Schuster et a / ., Sci Data, 5: 180157, Aug 2018.
[0202]
[0048] R. Chen et a!., Ana! Chem 90(19) : 11409-11416, Oct 2018.
[0049] P. Faridi et al., Sci Immunol, Vol 3, Issue 28, Oct 2018, DOI: 10.1126 / sciimmunol.aar3947.
[0203]
[0050] P. T. Illing et al., Nat Commun, 9(1):4693, Nov 2018.
[0204]
[0051] D. Gfeller et al., J Immunol, 201(12) :3705-3716, Dec 2018.
[0205]
[0052] S. M. Jensen et al., Front Immunol, 9:2697, 2018.
[0206]
[0053] M. van Lummel et al., Diabetes, 68(4) :787-795, Apr 2019.
[0207]
[0054] N. P. Croft et a / ., Proc Natl Acad Sci U S A, 116(8):3112-3117, Feb 2019.
[0208]
[0055] J. G. Abelin et al., Immunity, 51(4) :766-779, Oct 2019.
[0209]
[0056] S. Sarkizova et al., Nat Biotechnol, 38(2) : 199-209, Feb 2020.
[0210]
[0057] Massimo Andreatta and Morten Nielsen. Gapped sequence alignment using artificial neural networks: application to the MHC class i system. Bioinformatics, 32(4) :511-517, October 2015.
[0211]
[0058] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. 2017, doi: 10.48550 / ARXIV.1706.03762.
[0212]
[0059] Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets. 2014, doi: 10.48550 / ARXIV.1411.1784.
[0213]
[0060] Milot Mirdita et. al. , Lars von den Driesch et al., Nucleic Acids Research, 45(D1) :D17O- D176, November 2016, doi: 10.1093 / nar / gkwl081.
[0214]
[0061] Michal Bassani-Sternberg et al., PLOS Computational Biology, 13(8) :el005725, August 2017, doi : 10.1371 / journal.pcbi.1005725.
[0215]
[0062] Birkir Reynisson et al., Nucleic Acids Research, 48(W1) :W449-W454, May 2020, doi: 10.1093 / nar / gkaa379.
[0216]
[0063] Julien Racle et al., Nature Biotechnology, 37(11) : 1283-1286, October 2019, doi: 10.1038 / S41587-019-0289-6.
[0217]
[0064] Jun Cheng et al., Bioinformatics, June 2021, doi: 10.1093 / bioinformatics / btab422.
[0218]
[0065] Ishaan Gulrajani et a!., 2017, doi: 10.48550 / ARXIV.1704.00028.
[0066] Hesham ElAbd et al., BMC Bioinformatics 21(1) :235, June 2020, doi: 10.1186 / S12859- 020-03546-x.
[0219]
[0067] Morten Nielsen et al., PLoS One 2(8):e796, August 2007, doi: 10.1371 / journal. pone.0000796.
[0068] Edita Karosiene et al., Immunogenetics 65(10) :711-724, October 2013, doi:
[0220] 10.1007 / s00251-013-0720-y.
[0221]
[0069] Jonas Gehring et al., 2017, doi: 10.48550 / arXiv.1705.03122
Claims
CLAIMS1. A method for identifying peptides, which are bound and / or presented by MHC class I and / or MHC class II molecules in an individual, the method comprising a) providing a transformer encoder-decoder model, which evaluates whether amino acid sequences of peptides are ligands for selected MHC molecules, and which comprises an artificial neural network (ANN) architecture, which comprises an input layer, an output layer and one or more hidden layers, b) evaluating, in the transformer encoder-decoder model, data pairs, which each comprises 1) an MHC molecule amino acid sequence as well as 2) a peptide amino acid sequence, c) outputting from the transformer encoder-decoder model, for each inputted data pair, at least one quantitative assessment of the ability of each MHC molecule having the MHC molecule amino acid sequence of a data pair to bind and / or present the peptide having the peptide amino acid sequence of the same data pair, wherein the input layer receives and accepts data pairs each composed of an MHC molecule amino acid sequence and an amino acid sequence of a peptide, wherein the transformer encoder receives the MHC molecule amino acid sequence of a data pair with residue- and positional embedding as input and outputs a high-dimensional representation of the MHC amino acid sequence, wherein the transformer decoder receives the high-dimensional representation of the MHC molecule amino acid sequence and the peptide amino acid sequence of a data pair with residue embedding and positional encoding as input and outputs a high-dimensional representation of the relationship between the MHC molecule amino acid sequence and the peptide sequence, and wherein the output from the transformer decoder is used to calculate the quantitative assessment of the ability of the MHC molecule having the MHC molecule amino acid sequence of a data pair to bind and / or present the peptide having the peptide amino acid sequence of the same data pair.
2. The method according claim 1, wherein the quantitative assessment output from the transformer encoder-decoder model is normalized and transformed into at least one probability indication.
3. The method according to claim 2, wherein the at least one probability indication is a real number between 0 and 1.
4. The method according to any one of claims 1-3, wherein the at least one quantitative assessment comprises a calculation of affinity between the peptide having the peptide amino acid sequence of the data pair and the MHC molecule having the MHC molecule amino acid sequence of the data pair.
5. The method according to any one of claims 1-3, wherein the at least one quantitative assessment comprises a calculation of probability of presentation of the peptide on any cells that express the MHC molecule of the same data pair.
6. The method according any one of claims 1-3, wherein the at least one quantitative assessment comprises the calculations of claims 4 and 5.
7. The method according to any one of the preceding claims, wherein the data pairs are fed to the input layer of the transformer encoder-decoder model individually or in the form of a matrix or vector, which comprises multiple data pairs.
8. The method according to any one of the preceding claims, wherein one or more peptide amino acid sequences are paired with one or more MHC molecule amino acid sequences.
9. The method according to any one of the preceding claims, wherein the MHC molecule amino acid sequences are MHC class I amino acid sequences, MHC class II amino acid sequences, or both.
10. The method according to any one of the preceding claims, wherein the transformer encoder-decoder model has been trained on peptide:MHC datasets comprising peptides paired with MHC Class I and peptides paired with MHC Class II.
11. The method according to any one of the preceding claims, where the transformer encoder-decoder model has been trained on peptide-MHC datasets comprising data pairs generated by another prediction model capable of predicting peptides that will be bound and / or presented by MHC molecules.
12. The method according to any one of the preceding claims, where the quantitative assessment for each peptide-MHC molecule data pair is used to rank the list of data pairs.
13. A method for training a predictor model capable of identifying peptides, which will be presented by selected MHC class I and / or MHC class II molecules in an individual, the method comprising utilizing a conditional generative adversarial network (GAN) in the form of a generator model and a predictor model, wherein the predictor model evaluates whether amino acid sequences of peptides are ligands for inputted MHC molecules, a) evaluating, in the predictor model, data pairs, which each comprises 1) an MHC molecule amino acid sequence as well as 2) a peptide amino acid sequence, wherein, in a fraction of said data pairs, the amino acid sequences in the data pairs are those of verified peptide-MHC complexes, and wherein, another fraction of said data pairs are random pairs of MHC molecule and peptide sequences, and wherein in the remaining data pairs, the amino acid sequences in the data pairs are from molecules that are generated by the generator model, b) outputting from the predictor model at least one quantitative and / or qualitative assessment of the ability of each peptide amino acid sequence of a data pair to be presented by the MHC molecule from which the MHC derived amino acid sequence is included in the same data pair, c) outputting from the predictor model a quantitative and / or qualitative assessment of whether the peptide sequence of the inputted data pair is a generated sequence or a real, naturally occurring sequence, d) evaluating the output in steps b and c and adjusting weights of the neurons in the one or more hidden layers of the predictor model to minimize the error of the output or to optimize the performance of the model, e) evaluating the output in step c and adjusting weights of the neurons in the one or more hidden layers of the generator model to optimize the model's ability to generate a peptide sequence that the predictor predicts to be a real, naturally occurring peptide sequence, and f) repeating steps a-e to train the predictor model to correctly identify peptides, which will be presented by selected MHC class I and / or MHC class II molecules in an individual, wherein the input layer in the predictor model receives and accepts data pairs composed of an amino acid sequence derived from an MHC molecule and an amino acid sequence of a peptide.
14. The method according to claim 13, wherein the data pairs used in the model training comprises peptide:MHC datasets comprising peptides paired with MHC Class I and peptides paired with MHC Class II.
15. The method according to claim 13 or 14, wherein the data pairs used in the model training comprises data pairs generated by another prediction model capable of predicting peptides that will be bound and / or presented by MHC molecules.
16. The method according to any one of claims 13-15 wherein the random pairs of MHC molecule and peptide sequences are pairs each consisting of a randomly sampled naturally occurring peptide sequence and a randomly sampled MHC molecule sequence.
17. The method according to any one of claims 13-16, wherein the output in step b is i. a quantitative or qualitative assessment of whether the peptide having the peptide amino acid of a data pair is a generated peptide sequence or a naturally occurring peptide sequence, and / orII. a quantitative assessment of affinity between the peptide having the peptide amino acid sequence of the data pair and the MHC molecule having the MHC molecule amino acid sequence of the data pair, and / or ill. a quantitative assessment of presentation of the peptide having the peptide amino acid sequence of the data pair by the MHC molecule having the MHC molecule amino acid sequence of the data pair.
18. The method according to any one of claims 13-17, wherein the predictor model comprises an artificial neural network architecture comprising an input layer, an output layer and one or more hidden layers.
19. The method according to any one of claims 13-18, wherein the predictor model comprises a transformer encoder-decoder model according to any of claims 1-12.
20. The method according to any one of claims 13-19, wherein the generator model comprises an artificial neural network architecture comprising an input layer, an output layer and one or more hidden layers.
21. The method according to any one of claims 13-20, wherein the generator model comprises a transformer encoder-decoder architecture that pairwise encodes MHC molecule amino acid sequence data and randomly sampled digits, which are decoded to output a generated peptide amino acid sequence defined by the encoded MHC molecule amino acid sequence.
22. The method according to any one of claims 13-21 wherein the generator- and predictor models encode MHC molecule amino acid sequence input using positional embedding and encode either peptide sequence input or randomly sampled digits using positional encoding and feed the encoded information to be decoded to their respective decoders.
23. The method according to any one of claims 13-22, wherein the predictor model comprises 2 terminal neural networks, one which outputs the quantitative assessments described in points II and ill, and one which outputs the quantitative or qualitative assessment described in point i.
24. The method according to any one of claims 13-23, wherein, if the input data for the predictor model comprises more than one MHC molecule amino acid sequence per peptide, then the model weights are only adjusted based on the output composed of the amino acid sequences from the most likely peptide-MHC molecule complex for each peptide.
25. The method according to any one of claims 13-24, wherein outputs from the predictor model are normalized and transformed into probabilities, preferably a number between 0 and 1.
26. The method according to claim 25, wherein a rank calibration is carried out by generating raw predictions for a plurality of random peptides, sorting the predictions, selecting a number of reference points from the raw prediction to rank the list, and using these reference points to interpolate a rank score for any given raw prediction.
27. The method according to claim 26, wherein the rank calibration is carried out for the plurality of random peptides in respect of each of a plurality of MHC molecules.
28. The method according to any one of claims 25-27 wherein a probability calibration is performed by generating raw predictions on an external calibration dataset, comprising known positive and negative data points not used in the training, sorting the predictions and calculating local predictive performance metric scores in rolling windows along the sorted listof raw predictions, then fitting a mathematical function to the local performance metric scores such that any raw prediction can be transformed to a performance score.
29. The method according to claim 28 wherein the predictive performance metric score is the precision score.
30. The method according to claim 28 or 29, wherein a joint probability calibration is performed using raw predictions from the predictor model, combined with expression values of the peptide.
31. The method according to claim 30, wherein the joint probability calibration is carried out by generating raw predictions on an external calibration dataset with known expression values, comprising both positive and negative data points not used in the training, then a) creating bins of raw predictions and expression values, b) sorting each peptide-MHC data point in the external calibration dataset into a raw prediction bin as well as an expression value bin, c) creating a matrix of pairs of raw prediction and expression value bins, d) for each pair of bins in the matrix, calculating a predictive performance metric score for all peptides that fall into both the raw prediction bin and the expression value bin, and e) fitting a mathematical function to the predictive performance metric scores for each pair of bins, such that any pair of raw prediction and expression value can be transformed to a predictive performance score.
32. The method according to claim 31 wherein the predictive performance metric score is the precision.
33. The method according to any of claims 28-32, wherein the probability calibration is carried out for the plurality of random peptides in respect of each of a plurality of MHC molecules.
34. The method according to any one of claims 1-12, wherein transformer encoderdecoder model is an artificial neural network (ANN) capable of identifying peptides, which is provided according to the method of any one of claims 13-33.
35. A computer or computer system comprising a transformer encoder-decoder model as defined in any one of claims 1-12 and 34.
36. A method of identifying protein-derived MHC interacting peptides, which will be bound and / or presented by at least one MHC molecule expressed by an individual, the method comprising a) determining the amino acid sequences of one or more MHC molecules expressed by said individual, b) deriving amino acid sequences of potential MHC interacting peptides from one or more proteins; c) evaluating pairwise the amino acid sequences from steps a and b by using said amino acid sequences as input in step b in the method according of any one of claims 1-12 or 34, and / or by querying a database comprising records consisting of interacting pairs of MHC molecules and peptides identified by the method according of any one of claims 1-12 or 34 for the existence in the database of interacting pairs consisting of the MHC molecules and potential MHC interacting peptides from steps a and b respectively, and d) identifying those peptides that are ranked with the highest scores or which provide a positive query result as those that will be bound and / or presented by at least one MHC molecule in the individual.
37. The method according to claim 36, wherein the protein in derived from a malignant neoplasm.
38. The method according to claim 36 or 37, wherein the amino acid sequences of potential MHC interacting peptides identified in step b are potential neoantigens or neoepitopes.
39. The method according to claim 36 or 37, wherein the amino acid sequences of potential MHC interacting peptides identified in step b are derived from antigens expressedspecifically in the malignant neoplasm, such as antigens derived from endogenous retroviruses or endogenous retroviral elements.
40. The method according to any one of claims 36-39, wherein the presence of the one or more proteins are determined from the transcriptome of malignant cells in the individual.
41. The method according to any one of claims 36-40, wherein step a and / or step b comprises sequencing of material in a sample from the individual.
Citation Information
Patent Citations
Methods and compositions for treating prostate cancer using DNA vaccines
US20040142890A1
Generation of protein sequences using machine learning techniques
US11587645B2
Attention-based neural network to predict peptide binding, presentation, and immunogenicity
US20220122690A1
Method, System and Computer Program Product for Determining Presentation Likelihoods of Neoantigens
US20230298692A1